Back to Journal

Real-Time Shadows and Shadow Maps

A real-time shadow is a question asked twice: what can the light see, and what can the camera see. Shadow mapping answers the first with a depth image, and every artifact it produces follows from that image's limits.

A shadow is an absence, and absences are surprisingly hard to compute. To know whether a point on the ground lies in shadow, the renderer must know whether anything stands between that point and the light, which is a question about the whole scene, not about the point itself. A fragment shader, working on one pixel at a time and seeing only its own interpolated data, has no natural way to look across the world and ask what might be blocking the sun.

The technique that most games use to answer that question is called shadow mapping, described by Lance Williams in 1978 and still, in refined forms, the foundation of real-time shadows in Unity and nearly every other engine. Its idea is disarmingly direct: render the scene once from the light's point of view, remember how far the light can see in every direction, and later compare each pixel the camera sees against that memory.

The idea is simple, and almost every problem with real-time shadows arises from the way it is simple. A shadow map is a finite image with finite precision, stretched over a scene that may be very large, and each of its limits shows itself as a recognisable artifact. Following those artifacts one by one is the clearest route to understanding the settings that Unity exposes, and why changing one of them so often worsens another.

Seeing the World From the Light

In the first pass, the renderer places a virtual camera at the light, pointing the way the light shines, and draws every shadow casting object into a depth texture. No colour is written; each texel stores only the distance from the light to the nearest surface along that direction. For a spot light this camera uses a perspective projection shaped like the cone of the light; for a directional light such as the sun, whose rays are parallel, it uses an orthographic projection.

In the second pass, the ordinary one, each pixel the main camera draws is transformed into the light's coordinate system. That transformation yields two things: the position in the shadow map that corresponds to this point, and the distance from the light to this point. The shader samples the shadow map at that position and compares. If the stored depth is smaller than the point's own distance, something nearer to the light lies in the way, and the point is in shadow.

The cost of this scheme lies mostly in the first pass. Every object that casts shadows must be drawn again, from the light's viewpoint, with its own draw calls, and a scene with several shadowed lights draws its casters several extra times. This is why the Cast Shadows setting on a Mesh Renderer is worth considering object by object: small debris, distant props and objects that never block meaningful light can often be excluded without anyone noticing, and the savings add up.

Acne, Bias and Floating Shadows

The first artifact appears on surfaces that should be fully lit. The shadow map records depth only at the centre of each texel, but a single texel covers a patch of the surface that slopes away from the light. Points within that patch lie slightly in front of or slightly behind the stored depth, and those behind it conclude, wrongly, that they are shadowed by themselves. The result is a pattern of dark stripes or speckles across lit surfaces known as shadow acne.

The standard cure is bias. Before comparing, the shader pushes the point slightly toward the light, or treats the stored depth as slightly farther away, so that small discrepancies no longer register as occlusion. Unity's Light component exposes a depth Bias and a Normal Bias; the second pushes the sampled position along the surface normal, which helps particularly on surfaces seen at steep angles to the light, where acne is worst because each texel spans the greatest change in depth.

Too much bias produces the opposite problem. If the offset is larger than the thickness of an object or the gap between an object and the ground, the shadow begins only some distance from where the object actually touches the surface, and a character appears to hover slightly above its own shadow. Graphics programmers call this peter panning, after the boy whose shadow came loose. Tuning bias is a matter of finding the narrow band between acne and floating shadows for the scene at hand.

Resolution and Distance

A shadow map has a fixed size, commonly 1024, 2048 or 4096 texels on a side, and that size must be spread across whatever area the light covers. The arithmetic is unforgiving. If a 2048 texel map of the sun must cover a stretch of world 100 meters across, each meter receives about 20 texels, and a shadow edge can be no sharper than about five centimetres. Shrink the covered area to 25 meters and each meter receives about 82 texels.

This is why Unity, and every other engine, imposes a Shadow Distance, set in the Quality settings or in the URP asset, beyond which the main light casts no realtime shadows at all. Lowering the distance concentrates the available texels on the area near the camera and sharpens what the player sees closely; raising it lets shadows reach farther across a wide vista but spreads the resolution thinner, so that nearby shadows grow blocky and their edges begin to shimmer as the camera moves.

Filtering softens the blockiness without adding resolution. Instead of a single comparison per pixel, the shader takes several samples of the shadow map around the target position and averages their results, a technique called percentage closer filtering. The edge becomes a short gradient rather than a staircase. Unity's soft shadow options rely on this approach, which costs a few extra texture samples per pixel but disguises the grid of texels beneath. It cannot invent detail the map never recorded, however, so a very coarse map filtered heavily simply becomes a vague, smeared shadow rather than a sharp one.

Cascades for the Sun

The sun poses a particular difficulty because its shadows must cover everything the camera can see, from the cobbles at the player's feet to the hills on the horizon. Perspective makes nearby objects large on screen and distant ones small, so a uniform shadow map wastes texels in the distance and starves the foreground. Cascaded shadow maps address this by splitting the view frustum into several slices along its depth and giving each slice a shadow map of its own.

The nearest cascade covers a short distance and therefore packs its texels densely, producing crisp shadows where the player looks most closely. Each following cascade covers a longer stretch at lower effective resolution, which suits objects that are themselves smaller on screen. Unity allows up to four cascades for the main directional light, and the split distances can be adjusted so that the boundaries fall where they are least noticeable in a particular game's camera.

Cascades bring their own trade-offs. Each one is an additional rendering of the shadow casters, so four cascades can cost close to four times the first pass of a single map. At the boundary between two cascades, the change in resolution can be visible as a line where shadows suddenly soften, and some pipelines blend across the boundary to hide it. A camera that sits low over a busy scene benefits greatly from cascades, while a high, top down camera may need fewer.

Point Lights and Cube Maps

A spot light looks in one direction and can be served by one perspective shadow map. A point light shines in every direction at once, and no single projection can capture a full sphere of view. The usual solution is a cube map: the renderer places six cameras at the light, each with a ninety degree field of view, pointing up, down and along the four horizontal directions, and renders the shadow casters into each face in turn.

Six renders per light make point light shadows the most expensive kind in common use. A scene with a dozen shadowed point lights must redraw its casters as many as seventy two times, quite apart from the main camera's work, and every face consumes memory at whatever resolution is chosen. Engines therefore cap the number of shadowed additional lights, pack their maps into a shared atlas, and give lower resolutions to lights that are far away or small on screen.

The practical rule for designers is to treat a shadowed point light as a luxury. Many lamps can light a scene without casting shadows, their absence barely noticed, while a few chosen ones, the lantern the player carries or a fire at the centre of a square, carry shadows that matter to the mood. Torchlight against a palisade at night, of the sort a settlement in Crown & Ashes might rely on, is exactly where a designer must decide which flames deserve the cost.

Living With Imperfection

Every real-time shadow is a compromise, assembled from a depth image that is too coarse, a bias that is a little wrong in one direction or the other, and a distance beyond which the shadows simply stop. The skill lies in arranging those compromises so that they fall where the eye is least likely to notice. A shadow that is slightly soft near the horizon offends no one; a character floating above the ground beneath their own feet breaks the spell at once.

Profiling helps make those choices honestly. Unity's Frame Debugger shows each shadow pass as a separate step, listing the draw calls spent on casters, and a GPU capture reveals how much time the shadow maps consume. It is not unusual to find that a single shadowed point light, placed casually, costs more than all the other lighting in a scene, or that excluding a few hundred small props from casting saves more than any change of resolution.

It is easy to forget, among biases, cascades and atlases, how strange the underlying trick is. To draw darkness, the renderer first pretends to be the light, looking out into the world and memorising the nearest surface in each direction, and then, from the camera, asks of each pixel whether the light would have seen it. A shadow in a game is the record of that second glance falling short of the first.