All articles
Engineering·July 5, 2026·9 min read

Ultranaturalism: Inside Neyra's Six-Layer Prompt Engine

Why our frames don't look 'AI' — the physics, light and optics stack that writes cinematography instead of describing it.

By Neyra Engineering
Ultranaturalism: Inside Neyra's Six-Layer Prompt Engine

There is a look that gives AI video away instantly. The skin is too smooth. The light comes from nowhere. Fabric moves like liquid metal. Everything is beautiful, and nothing is true.

The industry calls it the plastic problem. We call it a prompting problem — and it is the single reason we spent eighteen months building what is now the core of the Neyra engine: the Six-Layer Prompt Architecture.


Adjectives don't light a scene

Open any prompt guide for today's video models and you will find the same vocabulary: cinematic, photorealistic, 8K, masterpiece, dramatic lighting. These are adjectives. A gaffer cannot rig dramatic. A DP cannot expose for masterpiece.

Real cinematography is not described — it is specified. A frame is the product of measurable decisions: the height of the sun, the diffusion on the key, the focal length, the shutter angle, the density of the atmosphere between lens and subject. When those decisions are missing, the model invents them — differently in every shot. That's why raw model output drifts, flickers and feels synthetic.

Neyra never sends an adjective where a measurement belongs.


The six layers

Every shot that leaves the Neyra pipeline is compiled — not written — through six stacked layers. Each layer is authored by a specialist agent, validated against the project's canon, and only then fused into the final generation instruction.

Layer 1 — World physics. Gravity, mass, momentum, wind load, fluid behavior. A coat has weight. Rain has terminal velocity. A car that brakes transfers weight to the front axle. This layer is why motion in Neyra footage reads as filmed, not simulated.

Layer 2 — Light transport. Sources, color temperature, falloff, bounce, occlusion. We specify the key, the fill, the practicals and the ambient dome the way a lighting plot does — in kelvins and stops, not vibes. Shadows agree with each other because they come from the same declared sources.

Layer 3 — Optics. The virtual camera is a real camera. Focal length, aperture, focus distance, sensor size, anamorphic squeeze, lens character — including the imperfections: breathing, chromatic fringing at the edges, the specific way a 35mm spherical rolls off focus. Perfection is what looks fake; calibrated imperfection is what looks photographed.

Layer 4 — Material truth. Subsurface scattering on skin, micro-roughness on fabric, wear on metal, moisture on stone. Ultranaturalism lives in the half-millimeter: pores, flyaway hairs, dust in a light shaft. This layer forces the model to render surfaces, not textures.

Layer 5 — Composition mathematics. Golden-ratio framing, headroom, lead room, horizon discipline, blocking geometry that matches the previous and the next shot. Composition is computed against the storyboard so that cuts land where an editor expects them to.

Layer 6 — Emotional direction. The layer that makes the other five matter. Performance beats, micro-expression timing, the tempo of a look. The emotional curve of the scene — inherited from the script agent — modulates everything above it: a tense scene tightens the lens, cools the fill, shortens the cut.


Compiled, checked, regenerated

Because the six layers are structured data — not a paragraph of prose — the engine can do something no manual prompter can: verify the output against the input. After generation, a QC pass measures the frame: Did the key stay at 5600K? Did the horizon hold? Did the character's wardrobe survive the shot? Anything that fails is regenerated with the deviation pinned — automatically, before a human ever sees it.

This is the quiet advantage of running a brain instead of a node graph. A node passes text forward and hopes. A brain writes the specification, checks the result and argues with the model until the physics agree.


What it looks like on screen

Side by side, the difference is not subtle. Ultranatural footage holds up where AI video traditionally collapses: slow push-ins on faces, hands interacting with objects, backlit atmospherics, water, glass, crowds. The image has what colorists call bite — real micro-contrast from real (declared) light, not sharpening.

And because every layer is versioned in the project's canon, shot 47 obeys the same sun as shot 3 — even if they were generated weeks apart.

Adjectives ask a model to imagine a film. Specifications instruct it to shoot one. That is the whole philosophy, in one sentence.

The Six-Layer Prompt Engine ships inside every Neyra production run today — from 15-second brand spots to long-form narrative. You never see it. You only see the frame.

ultranaturalismprompt-enginephysicslightingcinematographyengine
Join Neyra

Bring AI footage into a real cinematic pipeline.

28 specialist agents, EXR proxies, ASC-CDL grades and 4K masters — built for studios, agencies and serious creators.

Inquiries