










1 / 12
Buddhist sculptures communicate through a culturally specific visual language of mudras, bodily attributes, drapery, and ornament. Many museum visitors find these features compelling, yet their meanings can remain distant when conveyed through static wall labels alone. Commissioned for the Temple Room gallery at the San Diego Museum of Art, this interactive interpretive station approaches interaction as a form of interpretation.
The station offers two complementary modes. In Learn (Attributes of the Buddha), visitors explore a real-time 3D model of a seventeenth-century Amida Buddha, digitized on site using mobile photogrammetry. Touch-sensitive hotspots direct attention to specific iconographic details and explain their meanings. In Play (Mudra of Meaning), visitors perform mudras in front of a camera. A hybrid gesture-recognition system, which extends MediaPipe with a custom classifier for overlapping two-hand gestures, responds with generative animations developed from sculptures in the museum’s collection.
The project brings close looking and embodied learning together. Guided interaction encourages visitors to attend to details they might otherwise overlook, while performing the gestures allows iconographic knowledge to be understood through the body. As one visitor described the experience, it is a form of “learning by doing.” The project also demonstrates how culturally specific, artifact-centered interaction can be developed within the practical conditions of a museum, including immovable objects, variable lighting, and visitors of different ages, bodies, and levels of familiarity with the subject.
The work draws from interactive museum interpretation, digital heritage, and embodied and participatory learning. It enters into dialogue with projects such as the Smithsonian’s Cosmic Buddha in 3D, as well as the San Diego Museum of Art’s Learn-and-Play interaction model developed through its Art and Empathy strategy. It also advances a specific approach to generative AI. Here, generative media does not serve as visual decoration. Instead, it functions through what we call co-imaging, a method for producing interpretive responses grounded in particular artifacts and collections.
Located at the entrance to the gallery, the bilingual English and Spanish station is designed as a walk-up experience for visitors of all ages. No prior knowledge of Buddhist iconography is required. Visitors can begin by exploring the sculpture’s attributes, then try forming the gestures with their own hands. The mudra experience concludes with an optional photo keepsake, allowing visitors to carry a trace of the encounter beyond the gallery.
This commission is made possible with special thanks to Polly Liew (金蓉蓉) for their sponsorship, and to JoJo Ni (張瓊芳) for initially proposing the project.
The installation operates as a single real-time TouchDesigner system running on gallery hardware. It receives input through two channels: touch events from a vertical touchscreen and an RGB video stream from a webcam mounted above the display and calibrated for the gallery’s lighting conditions. Both modes share one screen. In the touch-based mode, visitors navigate a real-time 3D model of the digitized sculpture, open close-up views through interactive hotspots, and read bilingual interpretive text. In the gesture mode, recognized mudras trigger interpretive animations and captions. Visitors may also choose to receive a photograph of their interaction by email.
Three main processes support the system. The first is an offline asset-production pipeline. Because the seventeenth-century sculpture could not be moved, it was digitized in situ using smartphone photogrammetry. More than 1,700 photographs were captured with portable LED panels and an extension stick to accommodate restricted viewpoints. Denser photographic coverage was given to iconographically important areas, including the head, hands, and drapery. The reconstructed model was then manually cleaned and decimated to maintain a stable frame rate without losing the detail needed for close viewing. We also evaluated sparse-view synthesis and single-image-to-3D generation. Sparse-view synthesis produced artifacts in areas with limited photographic coverage, while single-image methods struggled to maintain consistency across viewpoints. Neither approach offered sufficient reliability for this installation.
The second process is a hybrid gesture-recognition pipeline. Twenty-one landmarks from each detected hand are streamed into TouchDesigner frame by frame, where two classifiers interpret them. A custom one-hand classifier recognizes three mudra classes and an explicit “none” class. Rather than using random negative examples, the “none” dataset was assembled from poses that caused false recognition during early testing. It therefore acts as an exclusion class that reduces accidental triggering.
A separate supervised classifier recognizes the overlapping two-hand welcoming mudra, which cannot be reliably interpreted by standard single-hand recognition. The landmark sets from both hands are combined into a single 42-keypoint feature vector and classified as one coordinated pose. The model runs through real-time Python inference within TouchDesigner. Runtime logic connects the one-hand and two-hand classifiers. One-hand recognition first identifies the mudra set being attempted, while the distance between the wrists determines when the system should enter two-hand recognition mode. For paired one-hand gestures, both poses must appear together and maintain a stable spatial relationship.
Several additional measures help the system remain reliable in a public gallery. Hands are ordered consistently according to wrist position, and landmark coordinates are normalized using wrist anchoring and palm scale. Strict confidence thresholds and sustained-state triggering prevent brief detection errors or single-frame fluctuations from activating a response.
The third process produces the interpretive animations. Each recognized mudra triggers a short, collection-specific sequence created through a multimodal generative workflow. Storyboards and shot breakdowns were developed with the assistance of a language model. Keyframes were then produced by image-editing photographs of sculptures from the museum’s collection. Video segments were generated from first and last frames, with the final frame of each segment used to initiate the next. This chained structure helped maintain visual continuity and limit drift. The sequences were then refined through conventional postproduction.
Mingyong Cheng is a new media artist, researcher, creative technologist, and assistant professor in the School of Arts, Entertainment, and Creative Technologies at the Georgia Institute of Technology. She holds a PhD in Visual Arts from the University of California, San Diego, and an MFA from Duke University. Her practice and research explore human–AI co-creation through immersive installations, interactive systems, and multimedia performance, with a focus on generative AI, cultural memory, ecological imagination, and embodied experience. Cheng serves as the project’s Artistic Director and Creative Technologist, leading its artistic vision, research, system design, and technical development.
Lucas Justinien Pérez holds a Bachelor of Fine Arts from Pratt Institute and a Master of Fine Arts from Parsons School of Design, and brings a multidisciplinary background spanning fine art, museum practice, digital media, and archival research. For this project, Perez served as Digital Producer & Project Manager, overseeing the development and production of its digital content and interactive media while helping shape the project's overall visitor experience.
Cameron Surh is the Digital Project Manager in the Digital Department at Lincoln Center for the Performing Arts and an artist based in New York City. Surh serves as the Creative Producer for this project, supporting its creative development, technical production, and installation from concept through final presentation.
Xuexi Dang is a Ph.D. candidate in Art History, Theory, and Criticism in the Department of Visual Arts at the University of California, San Diego. She holds an M.A. in East Asian Languages and Civilizations from the University of Pennsylvania. Her research examines modern Chinese visual culture, focusing on the development of satire, transnational exchange, and visual modernism across print media and exhibition contexts from the 1930s to the 1950s. She is also interested in the resilience of traditional Chinese aesthetics within contemporary digital and AI-generated art — work that has been presented at venues including the College Art Association, SIGGRAPH Asia Art Gallery, the NeurIPS Creative AI Track, and ISEA. Dang serves as the external Educational Content Editor, supporting content editing for the installation.
JoJo Ni
San Diego Museum of Art
Jojo Ni is an independent curator and cultural producer who explores traditional Chinese culture and art through a modern lens. As the project's initiator, she authored the initial concept draft for the digital interactive panel and successfully secured the foundational funding for its acquisition in the SDMA Temple Room. Her work focuses on inviting contemporary interpretations of classical traditions, utilizing innovative technology and cross-disciplinary collaboration to foster connections between Eastern and Western communities. She holds a Postgraduate Diploma in Asian Art History from the University of London and a Master of Business Administration from San Francisco State University

