






1 / 8
If you keep only one of a vision model's learned concepts switched on and look at a face through it alone, what do you see? And if the face turns even slightly, does that seeing survive?
I light a face with one feature the model grew on its own, an eye. Where it responds the surface brightens, and everywhere else sinks into dark. The lighting makes no new pixels. It only redivides the light that is already there. The faces themselves are made by a 3D generator, and no real person is among them.
I hold one face and turn it through a full circle. The feature that burned brightest head-on falls below half its strength with the smallest turn of the head. Both eyes are still plainly in frame. The eye does not actually leave view until the face is nearly in profile, yet the feature dies long before that. And even head-on it does not stay on the eye. I chose it for selectivity, not purity, so it was never clean. It fires on eyebrows, on the rims of glasses, on the edge of a cheek. There was no clean eye to begin with.
I do not test the whole model. I turn one carved lens up until it breaks, and watch where it fails. So this one, at least, is not a concept that knows an eye is there. It is a narrow seeing that answers only to a frontal appearance. Partial, tilted to one side, bound to a single point of view.
The work enlarges the same reduction. From one eye on one face, it changes parts and cuts a person into pieces, cuts a crowd down to one shared trait, and finally scatters a single face into hundreds of fragments from the dictionary. Breaking a face into a sum of parts has a long lineage. Where the eigenface once stood, a learned dictionary now takes its place. What remains is darkness.
The image is, throughout, a halftone of dots. That the machine reads a face as separate fragments is itself made into the form. Under the dots a darkness always stirs, holding what the reduction discarded submerged rather than erased. The sound was not composed. The harder a feature fires, the denser its grain, and when the image sinks to black the sound sinks with it. These are images made to be measured, not to be seen. We hear the measurement.
The grid of faces comes from August Sander, who believed a person could be filed under a type. Latens moves that question one layer in. Not what data it learned from, but what the carved lens finally holds. The Latin latens means lying hidden. It is also a latent lens. The work turns that one lens up until it breaks, so that the act of naming can be seen.
The piece runs 95 seconds and is best met as a single sitting in a dark room with sound. Watch one thing across the seven movements. Not the face, but the light on it, and the angle at which that light gives out.
Latens has no photographic input. Every face is synthesized by PanoHead, a 3D-aware GAN for full 360-degree head synthesis, which is rendered in four ways. A per-identity yaw orbit, three fixed yaws of the same identity, an identity-morph sequence, and a 600-image frontal corpus.
Each rendered frame is passed through DINOv2 (ViT-B/14) at 518 px, giving a 37x37 grid of patch tokens. On the patch tokens of the 600-image corpus I train a Top-K sparse autoencoder with 8,192 dictionary features and 32 active features per patch, reaching about 5% fractional variance unexplained. Each dictionary column is a direction in token space, and its per-patch response is what the work calls a lens. Region-level names are assigned by seeding with masks over eyes, brows, nose, mouth, skin and hair, then keeping the single most selective feature per region. The features are picked for selectivity, not for purity, and their impurity is part of what the work shows.
The render path contains no generative model. A feature's 37x37 activation map is smoothed, upsampled to frame size, and used as a light-only modulation of the frame, controlled by three parameters (gain, gamma, floor) that cycle slowly between hard isolation and bleed. Bright means the feature fired there. Nothing is painted in.
The measurement in the piece is made the same way it is shown. Over a single-identity 360-degree orbit I take the mean of a feature's ten strongest patch responses per frame as its response curve, and independently run a classical Haar cascade eye detector on the same frames. The eye feature falls below half its frontal peak within roughly 12 degrees of head rotation, while the eye stays detectable out to roughly 106 degrees. The claim is therefore about frontal appearance rather than occlusion, and it is bounded. One identity, one dictionary, and a death angle that depends on the chosen threshold.
Seven movements compose the arc. One face under the eye lens, a fixed-yaw triptych, a wall of many identities reduced to one shared trait, a taxonomy in which the lens changes region, an inverted pass lit only where the feature fails, the entire dictionary rendered as a 24x13 field of fragments over one head, and a final collapse into dark. Movements are cut at 30 fps, joined with fade-through-black crossfades, and graded with a lifted cool black point.
A final pass rewrites every frame as a halftone. Cell size is set per movement, from 3 px where small faces must stay legible to 10 px on single portraits, dot radius follows local luminance, and the ground is a slowly evolving deep indigo field rather than flat black. The 4K master doubles the cell and upscales.
Nothing in the soundtrack is composed by hand. Per-frame mean brightness and spatial standard deviation are read back off the finished video and drive detuned-saw pads over a fixed chord progression, a shimmer pad, and a granular cloud whose density follows on-screen activity, with a simple tapped reverb. Audio is synthesized in numpy and muxed with ffmpeg, so the image sinking to black is also what makes the sound recede.
Output is a single-channel 16:9 video of 94.7 seconds, mastered at 3840x2160 and 1920x1080, H.264 with stereo AAC. Implementation is PyTorch, Hugging Face transformers, OpenCV and ffmpeg, on one A6000.
Joonhyung Bae
Korea Advanced Institute of Science and Technology (KAIST)
Joonhyung Bae (b. 1995) is a multimodal AI researcher and art-technology practitioner whose work examines the relationship between humans and society within technological environments. Across games, VR, performance, and installation, he designs the conditions from which narrative emerges, inviting audiences into that relationship through playful, embodied experience rather than explanation. He holds a BFA in Design from Korea University and an MS and PhD from KAIST, where he is a postdoctoral researcher in the Music and Audio Computing Lab. He has published at NeurIPS, SIGGRAPH Asia, CHI, and ISMIR, and exhibited at the Asia Culture Center and Unfold X.

