Digital opera generation and engagement system supporting multi-modal interaction

By integrating perception, generation, and interaction design, and combining facial recognition and hand gesture recognition, an immersive experience is achieved in the digital opera system where users play roles and lead the plot. This solves the problem of weak interactivity in traditional opera and is applicable to various digital fields.

CN120931776BActive Publication Date: 2026-04-07WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies struggle to provide a complete experience for users to play opera roles and drive the plot in digital opera systems. The lack of deep integration between user input and plot logic results in weak interactivity and insufficient participation.

Method used

It adopts an integrated design of perception, generation, and interaction, combining facial recognition, generative image modeling, spatial sensing interaction, and plot logic-driven modules. It achieves hand motion recognition through a specially designed flashlight and Leap Motion sensing module, and supports multimodal interaction by combining digital face generation and plot control.

Benefits of technology

It enables users to have an immersive experience in digital opera, enhances the interactive dimension and sense of participation, supports multiple endings, breaks through the physical limitations of traditional opera stages, and is suitable for digital venues such as museums, exhibition halls and cultural tourism spaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931776B_ABST
    Figure CN120931776B_ABST
Patent Text Reader

Abstract

This invention discloses a digital opera generation and participation system supporting multimodal interaction, comprising an interactive device and a multimodal interaction subsystem. The interactive device includes: a light-shielding shell and a conical light outlet, used to define a conical area along the projection direction as a identifiable interactive area; and a motion-sensing control module, used to collect hand skeleton and motion trajectory information within the identifiable interactive area. The multimodal interaction subsystem includes: a facial recognition module, a digital opera character generation module, an interactive storyline control module, and a rendering engine execution module. This invention has the following advantages: 1. It accurately captures user hand movements, allowing users to select storyline paths, enhancing the interactivity and immersion of digital opera performances, and strengthening the gamified narrative capabilities of opera; 2. The interaction method is simple and intuitive, reducing the system threshold and difficulty of use, and has stronger scalability and universality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of human-computer interaction systems and digital arts performance technology, and in particular to a digital opera generation and participation system that supports multimodal interaction. Background Technology

[0002] With the continuous development of artificial intelligence and immersive interactive technologies, the digital transformation of cultural content is ushering in new opportunities. Especially in the field of intangible cultural heritage dissemination, how to utilize advanced technologies to achieve modern expression and innovative inheritance of traditional arts has become one of the important research directions. Traditional stage arts, represented by opera, often struggle to attract sustained attention and deep understanding from young audiences due to their highly formulaic nature and weak interactivity, becoming one of the pain points in the contemporary integration of culture and tourism.

[0003] Currently, various digital display methods for traditional Chinese opera have emerged in the market, mainly falling into two categories: "video playback" and "motion-sensing interactive" methods.

[0004] Video playback products mostly use multimedia projection, screen display or AR head-mounted devices to digitally present opera content, which has strong display and visualization effects, but its interactive form is simple, the audience lacks a sense of participation, and it is difficult to achieve a personalized immersive experience.

[0005] Motion-sensing interactive products use motion capture devices such as Kinect to allow users to imitate opera movements in front of the screen, enhancing their sense of participation. However, they usually only focus on imitating opera movements and lack a sense of character immersion and plot-driven mechanisms, making it difficult to build a complete closed loop of theatrical experience.

[0006] In addition, although some products have introduced AI face-swapping or virtual avatar technology, they are mostly static filters or preset templates, making it difficult to generate realistic and stylistically consistent opera characters based on the user's facial features, and also difficult to achieve dynamic performances that are synchronized with the plot.

[0007] In summary, due to the lack of deep integration between user input and plot logic, there is currently no systematic solution that can simultaneously achieve a complete experience flow of "users playing opera roles" and "users leading the plot direction". Summary of the Invention

[0008] The purpose of this invention is to provide a digital opera generation and participation system that supports multimodal interaction. Based on the design concept of "perception-generation-interaction" and integrated hardware and software, it combines facial recognition, generative image modeling, spatial sensing interaction and plot logic-driven modules. The aim is to provide the audience with an immersive digital opera experience that can be played and performed without changing the core art form and cultural heritage of traditional opera.

[0009] According to one aspect of the present invention, a digital opera generation and participation system supporting multimodal interaction is provided, comprising an interaction device and a multimodal interaction subsystem, wherein,

[0010] The interactive device includes: a light-shielding shell and a conical light outlet, used to define a conical area along the projection direction as a recognizable interactive area; and a motion control module, used to collect information on the hand skeleton and motion trajectory within the recognizable interactive area.

[0011] The multimodal interaction subsystem includes:

[0012] The facial recognition module is used to acquire and preprocess facial images in real time, and extract key facial feature points from the preprocessed facial images.

[0013] The digital opera character generation module is used to generate a fused digital face image based on the extracted key facial feature points and the input conditions of the digital face generation model, and output it to the rendering execution engine.

[0014] The interactive plot control module is used to control plot nodes, schedule NPC character performances, and drive user branch selections based on the constructed opera drama interaction engine; at the same time, it performs three-dimensional coordinate recognition and event control mapping based on the hand skeleton and motion trajectory information collected by the interactive device, and realizes closed-loop interaction through hardware triggering mechanism.

[0015] The rendering execution engine module is used to trigger the plot based on spatial interaction location and user interaction behavior, launch audio and video resources of preset plot nodes, and make plot branch selection and ending interpretation based on interaction prompts.

[0016] As a further technical solution, the interactive device also includes:

[0017] The aperture adjustment knob is used to rotate and adjust the divergence angle of the beam formed by the conical light exit.

[0018] The Z-axis displacement sensor is used to detect the forward and backward sliding displacement of the aperture adjustment knob and transmit the detected data to the central control unit to obtain the user's interaction intention in the Z-axis direction.

[0019] The central control unit integrates UWB and IMU modules for spatial positioning and inertial measurement, respectively; it is also connected to the aperture adjustment knob, Z-axis displacement sensor and motion control module to achieve interactive integrated control.

[0020] As a further technical solution, the interactive device also includes:

[0021] The physical trigger switch is used to switch the selection state, and transmits the switching trigger signal to the multimodal interaction subsystem and binds it to the input signal in the opera repertoire interaction engine.

[0022] As a further technical solution, the facial recognition module includes:

[0023] The camera unit is used to capture real-time images of the user's face;

[0024] The image preprocessing unit is used to perform face localization, feature normalization, and pose correction on the acquired user facial images;

[0025] The feature extraction unit is used to output the user's expression code and facial geometric parameters based on the preprocessing results and using a multi-layer neural network structure.

[0026] As a further technical solution, the digital opera character generation module includes:

[0027] A face matching database is used to build and store a structured traditional opera face tag system, and to design a digital face generation model that takes the user's original facial image, facial feature vector, face style tag and facial key point heat map as input conditions to obtain the preset digital face.

[0028] The digital face generation unit is used to map features from the face recognition module to preset digital faces to generate a fused digital face image.

[0029] The output unit is used to load the generated fused digital face image into the rendering execution engine.

[0030] As a further technical solution, the digital opera character generation module also includes:

[0031] The digital face image binding unit is used to establish a mapping relationship between the generated fused digital face image and the user's original facial image after loading the generated fused digital face image into the rendering execution engine, and bind it to the 3D mesh model.

[0032] The expression parameter-driven mechanism building unit is used to extract multiple facial geometric features through expression recognition and convert them into standardized expression parameters, which are then bound to the animation controller or deformation node in the rendering execution engine.

[0033] As a further technical solution, the interactive storyline control module includes:

[0034] The opera repertoire interaction engine is used to build the main scene container node, which contains a stage background layer, plot node structure, NPC digital character resources and user personalized image, and sets plot trigger conditions and multi-branch plot structure.

[0035] The spatial recognition unit is used to construct a three-dimensional interactive coordinate model based on the hand skeleton and motion trajectory information collected by the interactive device, and to realize event control mapping.

[0036] As a further technical solution, the interactive storyline control module also includes:

[0037] The hardware trigger control unit is used to determine the specific trigger event type after receiving the user's pressing action signal, combined with the current three-dimensional coordinate information, and to provide real-time feedback on the operation result on the screen in a visual / auditory manner.

[0038] As a further technical solution, the interactive storyline control module also includes:

[0039] The path recording unit is used to automatically form a path tree structure from the user's participation process, and play the complete ending segment and display the generated subtitles at the final node.

[0040] As a further technical solution, the rendering execution engine module includes:

[0041] The plot node triggering module is used to activate the NPC singing playback logic bound to the plot node based on the detected spatial interaction location and user interaction behavior.

[0042] The singing segment resource scheduling module is used to synchronously load multimodal performance resources corresponding to the current singing segment performance content when NPC singing segments are played.

[0043] The branch selection and jump unit is used to automatically provide multiple optional branch prompts after each plot node is completed, and to jump to the plot node corresponding to the selected branch after selecting the target branch prompt area through the interactive device.

[0044] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0045] (1) By embedding an interactive device (special flashlight) with a motion control module (Leap Motion infrared sensor), the user's hand movements can be accurately recognized. The interaction range is limited by the cone-shaped beam area, which improves the pointing accuracy and the naturalness of the interactive control.

[0046] (2) By combining a modified flashlight with the TouchDesigner visualization engine (a visualization node-based multimedia interactive creation tool), real-time interactive operation can be achieved in the stage screen, enabling users to have an immersive experience of "acting with light" in digital opera.

[0047] (3) By extracting the XYZ coordinate information of the hand in real time, users can trigger multi-level interactive operations such as browsing, selection, and advancement at different spatial distances, which greatly expands the interactive dimensions of the traditional opera viewing method;

[0048] (4) The system supports branching plot structure design. Multiple ending paths can be set for each node. Users can influence the plot through the "choice of light" to achieve a participatory and controllable opera performance experience.

[0049] (5) Cross-module linkage is achieved through OSC signals. The user's selection result can simultaneously trigger multiple visual and auditory modules such as stage background, NPC singing segments, and lyrics subtitles, making the virtual theater performance richer and more expressive.

[0050] (6) The system empowers users to make multiple ending choices from the perspective of the characters, promoting the transformation of opera art from "watching" to "participation", and opening up a new interactive path for opera education, cultural display and digital intangible cultural heritage dissemination;

[0051] (7) The interactive digital opera form proposed in this invention breaks through the physical limitations of traditional opera stage and has good display adaptability and commercial value in digital fields such as museums, exhibition halls, and cultural tourism spaces. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 This is an overall schematic diagram of the system provided in an embodiment of the present invention;

[0054] Figure 2 A schematic diagram of a special flashlight for a system provided in an embodiment of the present invention;

[0055] In the image: 101, Special flashlight device; 102, UWB sensor module assembly; 103, Light display assembly; 104, Front-end camera capture assembly; 105, Dual-screen display assembly;

[0056] 1. Aperture adjustment knob; 2. Directional conical light outlet (lens assembly); 3. Housing; 4. Main control circuit board (integrated UWB and IMU modules); 5. Leap Motion sensing module; 6. Battery compartment; 7. Z-axis displacement sensor; 8. USB charging and debugging interface. Detailed Implementation

[0057] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. In addition, the technical features of the various embodiments or individual embodiments provided by the present invention can be arbitrarily combined to form new technical solutions. Such combinations are not bound by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0059] The digital opera generation and participation system for supporting multimodal interaction provided in this embodiment of the invention includes an interactive device and a multimodal interaction subsystem.

[0060] 1. The interactive device adopts a specially designed flashlight device 101.

[0061] like Figure 2 As shown, the specially designed flashlight device 101 includes an aperture adjustment knob 1, a directional conical light outlet 2, a housing 3, a main control circuit board 4, an embedded Leap Motion sensing module 5, a battery compartment 6, a Z-axis displacement sensor 7, a USB charging and debugging interface 8, and a signal transmission module. The Leap Motion sensing module 5 is located in the middle section of the specially designed flashlight device 101 and is connected to the signal transmission module through a data interface.

[0062] An aperture adjustment knob 1 is located at the front of the flashlight. Rotating it adjusts the divergence angle of the beam formed by the directional conical light outlet 2, allowing the user to dynamically control the aperture's contraction and diffusion effects during interaction. The directional conical light outlet 2 is a lens structure used to focus the light source and guide the beam direction, providing precise visual feedback for spatially directional interaction. A Z-axis displacement sensor 7 is connected to its rear side to detect the forward and backward sliding displacement of the aperture adjustment knob 1, acquiring the user's interaction intention in the Z-axis direction and transmitting the sensor data to the main control circuit board 4 for processing. The outer casing 3 serves as the supporting structure of the device, housing the main control circuit board 4, battery compartment 6, and USB charging and debugging interface 8.

[0063] The main control circuit board 4 serves as the central control unit, integrating UWB (Ultra Wide Band) and IMU (Inertial Measurement Unit) modules for spatial positioning and inertial measurement, respectively. The main control circuit board 4 is also electrically connected to components such as the aperture adjustment knob 1, Z-axis displacement sensor 7, and Leap Motion sensing module 5, and is integrated and controlled via a program. The battery compartment 6 houses a rechargeable battery, providing power to the main control circuit board 4 and other functional modules. A USB charging and debugging interface 8 is located at the rear of the device, used for charging and connecting to external computing devices, facilitating program debugging and data exchange.

[0064] like Figure 1 As shown, the UWB sensing module component 102 mainly consists of an external antenna, enabling communication between it and the UWB module integrated in the main control circuit board 4 to obtain the flashlight's precise position in three-dimensional space in real time. Simultaneously, this component works in conjunction with the IMU module integrated in the main control circuit board 4 to improve the accuracy of interaction trajectory and spatial positioning. The light display component 103 is located at the front end of the dual-screen display component 105, and the front-end camera capture component 104 is installed below the dual-screen display component 105 to capture dynamic images of the user's face and gestures in real time, supporting subsequent image recognition and posture analysis. The dual-screen display component 105 includes two synchronized displays, left and right, which can be used to display real-time interactive content, user feedback interfaces, and system status information, enhancing overall interactive visibility and system feedback efficiency. The dual-screen display component 105 receives user gesture information through the Leap Motion sensing module 5, enabling complex human-computer interaction operations.

[0065] II. The multimodal interaction subsystem includes a facial recognition module, a digital opera character generation module, an interactive plot control module, and a rendering execution engine module.

[0066] (a) The facial recognition module consists of two parts: facial image acquisition and preprocessing, and key facial feature point extraction.

[0067] 1. Face image acquisition and preprocessing

[0068] (1) Face image acquisition

[0069] This invention uses a common RGB camera (such as a USB webcam, a laptop built-in webcam, or a mobile device front-facing camera) to capture facial images in real time. The captured image resolution must be no less than 640×480 pixels, and the frame rate no less than 24 frames per second to ensure that facial feature changes (such as expressions and postures) can be recorded in real time. The front-end camera capture component 104 should be positioned directly in front of the user, with an angle of less than 15° to the horizontal viewing angle of the face, and a distance between 30cm and 80cm to avoid facial obstruction, strong backlighting, or excessive facial rotation.

[0070] (2) Image stability judgment and dynamic frame selection

[0071] To improve the accuracy of subsequent facial recognition, the system performs stability analysis on the acquired image stream. The specific method is as follows:

[0072] ① Set image quality detection thresholds, including indicators such as brightness range, sharpness score, and inter-frame change rate.

[0073] ② If the blur of the current frame image is lower than the preset threshold, or the facial area jitter exceeds a certain pixel distance (Δx>5px, Δy>5px), then the frame is marked as an "invalid frame" and discarded.

[0074] ③ The system uses a sliding window averaging algorithm to denoise the acquired frame sequence, taking 5 frames as a window to average the center point and size changes of the facial area, in order to smooth out the effects of slight shaking or light changes caused by the user during operation.

[0075] (3) Image cropping and face region localization

[0076] First, the Haar Cascade face detection algorithm in OpenCV is used to locate faces in the image. In the specific implementation, a pre-trained classifier model file is loaded; the object detection method of the classifier is called to perform multi-scale scanning detection on the input image; the input of this method is the grayscale image to be detected.

[0077] Configurable parameters include, but are not limited to: the scaling factor of the scanning window, the minimum number of neighborhood rectangles required to form an effective detection target, detection process control flags (e.g., skipping smooth regions, returning only the largest target, early termination conditions, etc.), and minimum and maximum target size limits;

[0078] This method returns a matrix containing the location and size information of the detected target.

[0079] The extracted regions are cropped and scaled at a uniform ratio to reduce the face region to a fixed size, ensuring consistency of input to the subsequent neural network model. If multiple faces are detected, the face region with the highest confidence score is selected as the primary input source by default.

[0080] (4) Image preprocessing and enhancement

[0081] To improve the model's adaptability to different skin tones, lighting conditions, and facial features of different ethnicities, a series of preprocessing operations are performed on the cropped face images:

[0082] ① Image normalization: Normalize the pixels of an image, mapping pixel values ​​to the range of [0,1] or [-1,1].

[0083] ② Color space conversion: Convert the RGB image to YCrCb or Lab space to facilitate the subsequent separation of skin color and texture features.

[0084] (5) Multi-frame fusion and caching mechanism

[0085] The system by default acquires N consecutive frames (e.g., 15 frames) of images, extracts and fuses the average features of the facial regions, eliminating instantaneous errors and jitter, and improving the accuracy of feature point extraction. Each successfully acquired image, along with its corresponding timestamp and facial detection status, is stored in a temporary cache pool for subsequent modules to retrieve.

[0086] 2. Key Facial Feature Point Extraction

[0087] (1) MediaPipe Face Mesh

[0088] MediaPipe is an open-source multimedia machine learning model application framework. The MediaPipe FaceLandmarker task allows for the detection of images and videos, which can be used to recognize human facial expressions, apply facial filters and effects, and create virtual avatars. This task outputs a 3D face marker. The MediaPipe face landmark detection model contains 468 3D landmarks.

[0089] After receiving a user's facial image frame, the system invokes the MediaPipe Face Mesh module to automatically detect the face region and construct a 3D mesh model of the face. Through this process, the system can obtain a set of high-precision facial key points. , used to describe the structural features of the current user's face in space.

[0090] (2) Coordinate projection and attitude calibration

[0091] This invention further performs two-dimensional mapping and pose normalization on the aforementioned three-dimensional feature point set to adapt it to subsequent image fusion and style transfer algorithms. Specifically, it includes:

[0092] By using the affine transformation image processing algorithm, three-dimensional coordinate points are mapped to a two-dimensional coordinate system in the standard image space, ensuring that the point positions can be accurately corresponded to the pixels of the static image.

[0093] The system combines head pose estimation (such as pose calculation methods based on the PnP problem) to adjust and calibrate the mapped face pose. The system controls the facial orientation within a standard range, such as a frontal view (yaw = 0°) or a semi-lateral view (|yaw| ≤ 30°), to ensure the frontal visual consistency of the generated character.

[0094] If the posture deviation is large, the system will automatically prompt the user to perform face alignment, or perform inverse transformation processing on the facial image through a posture normalization algorithm to restore the pseudo-positive viewpoint.

[0095] After pose calibration, the system obtains a set of aligned two-dimensional facial feature points. This serves as the foundational anchor point for subsequent image fusion steps.

[0096] (3) Face Embedding

[0097] To achieve personalized driving and style mapping control of facial images, this invention further extracts facial feature vectors based on standardized facial images. These facial feature vectors are implemented using the FaceNet model. By inputting a standardized image or its keypoint set, it outputs a set of high-dimensional feature embedding vectors to characterize the user's facial identity features and local differences. Its core components include a triplet input mechanism, a deep architecture, L2 norm normalization, embedding vector generation, and a triplet loss function optimization strategy. The working mechanism and parameter settings of each module are explained in detail below:

[0098] To achieve precise facial discrimination and inter-class difference learning, the model input is a triplet structure, meaning each training set consists of three images: an anchor image (A), a positive image (P) (belonging to the same subject as A), and a negative image (N) (belonging to a different subject than A). During training, the batch size is controlled by the parameter batch_size. For example, when batch_size=5, it means that 5 sets of triplets are input each time, for a total of 15 images, to improve training efficiency and the generalization ability of feature learning.

[0099] This facial feature vector will serve as a conditional embedding in the subsequent digital opera character generation model. It will guide the image generation model to integrate opera character facial features while preserving the user's facial features, thereby achieving a unity of "formal resemblance" and "spiritual resemblance" in the digital character.

[0100] (ii) The digital opera character generation module consists of three parts: digital face generation model input construction, integrated digital face image generation, and dynamic binding and real-time driving.

[0101] 1. Input Construction for Digital Face Generation Model

[0102] First, a traditional opera makeup style tag library is established. This sub-step aims to construct a structured traditional opera facial makeup style tag system to support conditional control during style transfer. Taking Huagu Opera as an example, it includes:

[0103] ① Collect materials

[0104] First, collect facial makeup and costume images corresponding to classic Huagu Opera plays (such as "Standing by the Flower Wall", "Liu Hai Cuts Wood", "Wang Zhaojun" etc.).

[0105] ②Classification labeling

[0106] The roles are categorized and labeled based on the play, role type (male, female, painted-face, clown), gender, personality traits (gentle, strong, cunning, etc.), and historical context. For example, the character Chunxiang in "Standing by the Flower Wall" is labeled as follows:

[0107] └── Standing on the Flower Wall

[0108] └── Chunxiang (Dan)

[0109] ├── Gender: Female

[0110] ├── Personality: Gentle and resilient

[0111] └── Historical Context: Folk Legends of the Qing Dynasty

[0112] ③ Establish tag space

[0113] Style embeddings are extracted, and a label space is built using a combination of image clustering (such as K-Means) and manual semantic annotation. Each style label vector contains attribute fields such as color palette, line pattern, and position weight of local feature regions (such as the area around the eyes, forehead, and lips).

[0114] (2) Image generation model input condition construction (Conditional Input Packaging)

[0115] To achieve end-to-end face style generation, the method of this invention designs the model input as a multimodal fusion structure, which includes the following four main inputs:

[0116] ① User's original facial image (Image Input)

[0117] As the main input channel in the generative model, it provides users with basic facial texture and shape information.

[0118] ② Facial Embedding

[0119] Using a pre-trained face recognition network, we extract and characterize the "identity semantics" of a user's face in the face space.

[0120] ③Style Condition

[0121] The input is the style label vector generated in the previous steps; it is used as the input to the conditional control module in the image generation model and is integrated into the conditional mapping network of StyleGAN to guide the generated results to present a specific color distribution and geometric layout style.

[0122] ④ Facial landmark heatmap

[0123] Based on the MediaPipe facial landmark detection model, the coordinates of facial landmarks are extracted and converted into a two-dimensional heat map. This heat map is used to constrain the alignment of key elements of the face (such as eyebrows, eyes, and lip lines) with the user's facial structure, thereby improving the fit.

[0124] 2. Fusion Digital Face Image Generation

[0125] (1) Constructing a fusion-based image generation architecture

[0126] A generative network is built based on the Stable Diffusion model, which embeds opera mask style information into the image while preserving the user's original facial structure (especially the eyes, nose, and mouth).

[0127] The DDIM Inversion technique is used to encode user images into the latent space. The style embedding vector is used as a conditional input and fused with the latent space features to guide the generation of digital faces with opera-style makeup.

[0128] (2) Introduce a condition control module (such as ControlNet)

[0129] To ensure that the generated image preserves the user's facial structure, a ControlNet conditional control structure is adopted. As a downstream module of Stable Diffusion, ControlNet can guide the generated result towards "fidelity + style fusion" by injecting feature maps.

[0130] The input conditions include: a facial landmark heatmap extracted from the original image and an inflated contour map used to enhance facial structures.

[0131] 3. Dynamic binding and real-time driving

[0132] (1) Importing and binding digital face images

[0133] The digital opera facial makeup images obtained through the aforementioned diffusion-based image generation model are loaded into the real-time rendering engine TouchDesigner to construct a facial rigging model to support subsequent real-time expression-driven rendering.

[0134] Specifically, the facial image is first imported into the rendering engine as a texture resource. By establishing a UV mapping relationship with the user's original facial image, it is ensured that the digital facial image is strictly aligned geometrically with the user's facial features (eyes, nose, mouth, etc.). Subsequently, by binding it to the 3D mesh model Skinned Mesh Renderer, the image becomes a dynamically rendered object that can be controlled in real time, providing the image foundation and structural anchor points for realizing dynamic expression mapping.

[0135] (2) Construction of expression parameter-driven mechanism

[0136] To achieve synchronized expression driving based on user facial movements, the system integrates a real-time expression recognition module, preferably using the MediaPipe Face Mesh framework, to extract multiple facial geometric features from the video stream, including but not limited to eye opening and closing angles, eyebrow displacement amplitude, and degree of mouth corner stretching.

[0137] After the features are converted into standardized expression parameters, they are bound to the animation controller or deformation node in the rendering engine to drive the dynamic changes of corresponding areas (such as periorbital texture and lip lines) in the digital face. Furthermore, to achieve delicate and natural expression performance, the system sets different response weights and driving mapping strategies for different local areas (such as the forehead, eyes, and mouth) in the face image, enabling flexible projection of facial expressions into the stylized image. In addition, the expression parameters can be configured with a trigger mechanism; when a specific expression pattern (such as a smile or surprise) is detected, the visual effects module is automatically activated to achieve immersive interactive feedback such as local texture dynamics, color changes, or environmental feedback.

[0138] (III) The interactive plot control module shown includes two parts: the construction of the opera repertoire interactive engine and the flashlight interactive device and spatial recognition.

[0139] 1. Construction of an interactive engine for traditional Chinese opera performances

[0140] First, a multi-ending interactive engine for traditional Chinese opera performances is built based on TouchDesigner. This engine serves as the rendering control module within the interactive opera experience system, managing plot nodes, scheduling NPC performances, and driving user branching choices. By establishing a visual scene container with conditional control and branching capabilities, modular interpretation and structured interaction of traditional opera content are achieved.

[0141] (1) Construction of the main scene container

[0142] In the TouchDesigner platform, construct a main scene container node (Container COMP). This node serves as the central control module for system rendering and logic scheduling, and internally contains the following sub-structures:

[0143] ① Stage background layer: Used to showcase the overall stage design and environment of the opera performance, including static or dynamic background videos, lighting effects layers, stage element textures, etc.

[0144] ② Plot node structure: The structure information of each plot segment is maintained by using tabular data (DAT Table). Each node includes fields such as node number, corresponding character ID, plot description, and jump target.

[0145] ③ NPC Digital Character Resources: Configure a set of digital characters (Non-PlayerCharacter) for each plot node, including their visual appearance (video material or 3D model), vocal audio resources (.wav / .mp3 format), and motion control parameters (such as entry, performance, exit).

[0146] ④ User Personalized Image: Used to load digital character images generated by users through the opera face mask generation system. It supports avatar overlay or full-body image import, and realizes expression changes and co-performance with the character based on the subsequent facial driving module.

[0147] (2) Setting of plot trigger conditions

[0148] Set at least one "trigger condition" for each plot node, which includes, but is not limited to:

[0149] ① Spatial positioning trigger: Based on Leap Motion or infrared vision sensors, the position coordinates (X, Y, Z) of the user's hand or interactive device are obtained, and it is determined whether the user has entered a preset hot zone (Trigger Zone).

[0150] ②Time synchronization control: Utilize the frame synchronization nodes (Time COMP, Timer CHOP) in TouchDesigner to achieve precise synchronization and status judgment of the plot's audio and video, ensuring that the NPC's singing performance and the visual performance proceed in a timely and coordinated manner.

[0151] ③ Combinatorial logic judgment: When the user's action meets both spatial and temporal conditions, the system automatically triggers the character performance and plot progression in the current node to avoid accidental triggering and conflict behavior.

[0152] (3) Multi-branch plot structure setting

[0153] For each plot node, the system sets two or more "selection branch paths," with each branch corresponding to a jump link to a subsequent plot node, forming a "plot tree" structure. The branch paths include:

[0154] ① Visual prompting area: Several interactive options are preset in the screen or projection image, and prompts are given to the user through visual highlighting, icon marking, role guidance, etc.

[0155] ② User pointing confirmation mechanism: Combined with a special flashlight device, it detects whether the light cone area overlaps with the branch area, and sets a time confirmation mechanism (such as a 1-second pause is considered a selection).

[0156] ③ Plot jumps and ending settings: Based on the path selected by the user, the system automatically jumps to the plot node corresponding to the selected branch until the preset ending node is finally reached.

[0157] 2. Flashlight Interaction Device and Spatial Recognition

[0158] (1) Design of spatial recognition and sensing structure for modified flashlight

[0159] This invention presents a spatial recognition and interactive selection device suitable for interactive opera experience systems, specifically a specially designed flashlight interaction device based on Leap Motion infrared sensing technology. This device embeds a Leap Motion module inside the flashlight body, and, in conjunction with a light-shielding shell structure and a conical light outlet design, limits the recognizable interactive area to the area of ​​the conical beam projected in the forward direction, thereby eliminating irrelevant interference information other than that from the hand.

[0160] The Leap Motion module is installed in the middle of the flashlight body. By modifying the rear shell to block the field of view except for the cone angle directly in front, it only collects information on the hand skeleton and movement trajectory within the cone angle area. This structure enables precise perception of hand movements in the area pointed by the flashlight, ensuring the accuracy and stability of interactive operation.

[0161] (2) Three-dimensional coordinate recognition and event control mapping

[0162] Based on the above structure, this step further extracts the position information of key skeletal points of the hand in real time using the Leap Motion SDK, constructs a three-dimensional interactive coordinate model, and implements the following mapping strategy:

[0163] The X and Y coordinates of the hand correspond to the two-dimensional projection coordinates on the screen and are used to identify the interactive area node pointed to by the user.

[0164] The Z-axis coordinate (i.e., the depth distance between the hand and the sensor) is used to control the size of the interactive aperture and introduces the concept of event trigger levels. Specifically, this includes:

[0165] ① Long-distance area (greater than preset threshold): Triggers "Browse" state, only lighting up the relevant area or displaying the prompt message;

[0166] ② Close range (less than the threshold): Triggers the "Select" state, executing a branch plot jump or character response.

[0167] By driving interactive parameters in real time through three-dimensional skeleton points, continuous perception and dynamic control of gestures within the light cone can be achieved, effectively supporting multi-level and multi-state interactive logic.

[0168] (3) Hardware triggering mechanism and closed-loop interaction implementation

[0169] To further enhance the initiative and selectivity of operation, the flashlight interaction device is equipped with a physical trigger switch, which the user can explicitly switch the current "selection state" by pressing with their thumb or finger. This trigger signal can be connected to the main control computing device through a digital signal port and bound to the oscillator (OSC) signal input module in graphics engines such as TouchDesigner, forming a complete input feedback closed loop.

[0170] After receiving a user's press action, the system will combine the current three-dimensional coordinate information to determine the specific trigger event type (such as plot node jump, character action start, sound effect feedback, etc.) and provide real-time feedback on the operation result on the screen in a visual / auditory manner, forming an immersive, multi-branching, and user-participatory opera interactive experience process.

[0171] (iv) The rendering execution engine module includes two parts: plot triggering and NPC singing playback, and branch selection and ending interpretation.

[0172] 1. Plot triggers and NPC singing segments

[0173] (1) Triggering plot nodes

[0174] The method of this invention is based on spatial interaction location and user operation behavior recognition to trigger the plot in an interactive digital opera system, thereby enabling the activation of audio and video resources for preset plot nodes.

[0175] In this method, the user holds an interactive flashlight device with positioning recognition function and turns it on in a preset plot node area. When the system detects that the X / Y coordinates of the area pointed by the user enter the set recognition range of the corresponding plot node, and the Z-axis depth judgment is a valid trigger state (such as pointing at a close distance and pressing the button), the NPC singing playback logic bound to the plot node is immediately started.

[0176] This logic includes retrieving the vocal audio resources and character animation frames of the digital character (NPC) corresponding to the node, ensuring a natural transition in the plot performance and a complete dramatic rhythm.

[0177] (2) Singing segment resource allocation

[0178] While the NPC's singing segment is playing, the system will simultaneously load the multimodal performance resources corresponding to that performance segment, specifically including:

[0179] The stage visual background scene required for the plot node is automatically switched according to the development of the plot; the corresponding lyrics subtitle text is presented word by word in rhythm according to the audio timeline; the body and facial performance movement trajectory of the NPC digital character; the performance action interaction between the user character (i.e. the character generated by mapping the user's digital face image) and the NPC, such as the linkage mechanism of synchronized turning, bowing, and staring.

[0180] The above resources are uniformly scheduled and played through the TouchDesigner rendering engine to ensure that the performance process has complete stage design, consistent rhythm and natural interaction.

[0181] (3) Playback dynamic monitoring and branch response mechanism

[0182] To enhance the variability of the storyline and user engagement, the system continuously monitors user behavior during NPC singing segments to determine if the following situations exist:

[0183] The user moves to other story nodes;

[0184] The user triggers the selection state again (e.g., by pressing the interaction button again).

[0185] Users perform actions such as terminating or fast-forwarding (e.g., long press, continuous waving gestures).

[0186] The system evaluates the current plot playback status in real time based on the above monitoring results. If a change in user behavior is detected, the current song playback can be interrupted, and the user can jump to the new plot branch selected by the user or enter the next stage node in advance, thus realizing a plot advancement mechanism with a high degree of freedom.

[0187] 2. Branching Choices and Ending Deduction

[0188] (1) Visual cues and presentation mechanisms of branch options

[0189] This method selects plot branches based on interactive prompts. After each plot node is completed, 2 to 3 optional branch prompts automatically appear on the main screen or projection device. These prompts are presented in one or more of the following visual forms, including but not limited to: simulated light spots (apertures) mapped to different branch directions; graphic or icon-style buttons, combined with sound effects to enhance guidance; and short singing / recitation prompts, representing the emotions or event types of the subsequent plot direction.

[0190] This hint mechanism is implemented by TouchDesigner or an equivalent rendering engine, and binds branch hint resources to plot nodes to ensure that the hint content is consistent with the plot logic.

[0191] (2) User orientation selection and branch path jump

[0192] While the branch prompt interface is displayed, the user points to the target branch prompt area using a handheld interactive flashlight device. When Leap Motion or a similar sensing system recognizes that the user's X / Y spatial coordinates are continuously and stably pointing to a certain option area, and the set minimum pointing time threshold is met (e.g., 1 second to 3 seconds), and the Z-axis depth is within the "confirm operation" distance range, the system determines that the user has completed the current branch selection.

[0193] The system then immediately jumps to the subsequent story node corresponding to the selected branch, including: loading the corresponding NPC's singing audio and video resources; switching the stage background and props; and executing character animations and interaction settings.

[0194] This path-jumping mechanism is responsive in real time and supports highly immersive story branch transitions in a short period of time.

[0195] (3) User path recording and ending rendering output

[0196] Throughout the entire story interaction process, the system continuously records the user's choices at each story node and dynamically constructs the choice sequence using a branch tree structure.

[0197] The path record structure includes: node number and timestamp; coordinate position and judgment result of each user selection; trigger branch identifier and target plot number.

[0198] After the user reaches the final node, the system automatically generates the final ending scenario based on the path information and outputs the corresponding: ending plot animation segments, ending song subtitle text, and optional branch path visualization diagram, which can be used to review the user's choice process.

[0199] As a preferred embodiment, this invention uses the interactive creation of the Jingzhou Flower Drum Opera "Standing on the Flower Wall" as an example to illustrate the specific working method of the invention.

[0200] During the interactive performance of this play, users can engage in multiple rounds of interaction using a flashlight-based interactive device. First, the user illuminates the initial startup position set by the system, and the other end of the screen slowly lights up, revealing Yang Yuchun, one of the main characters. The user, playing the role of Chunxiang, then interacts for the first time with Yang Yuchun outside the flower wall. A scene of Yang Yuchun striking a wooden fish appears on the screen, as he sings of longing. The user then uses the flashlight-based interactive device to shine it on a specific area of ​​the screen to trigger a dialogue between Yang Yuchun and "Chunxiang."

[0201] The user illuminates the location again according to the guide to advance the plot, the screen changes scenery, and the image of Wang Meirong, one of the main characters, lights up. The system plays a classic aria from the dialogue between the two, and the user's "Chunxiang" character can interact with Wang Meirong and try to persuade her to meet at the flower wall. The user needs to move back and forth in the flower wall area using the flashlight interaction device to trigger the meeting between Wang Meirong and Yang Yuchun.

[0202] When the drama reaches a crucial point, multiple dialogue options will be displayed on the screen. The user selects one using the flashlight beam, triggering different plot branches. The system accurately records the user's choice path using the flashlight's location and arranges subsequent interactive content based on the path. Each choice made by the user affects the subsequent plot development, increasing the diversity of interaction and the user's immersion. In "Standing on the Flower Wall," the user (Chunxiang) will play a crucial intermediary role in the segment where Miss Wang Meirong and Yang Yuchun express their feelings. After the flower wall interaction, options will pop up on the screen, allowing the user to choose between "encouraging Miss" or "dissuading Miss": if the user chooses to encourage, Wang Meirong and Yang Yuchun can continue to communicate and confide in each other, and the user will watch their duet and emotional outpouring, experiencing their emotional entanglement; if the user chooses to dissuade, the plot shifts to Chunxiang trying to remind Wang Meirong of the difference in their social status, and the interaction stops. Users will gradually unfold different story endings by choosing to assist Wang Meirong in visiting Yang Yuchun in prison, prevent her from visiting, or help the two escape and elope. If users choose to assist in visiting, the system will guide them to use a flashlight interactive device to simulate bribing the prison guards and successfully enter the prison to meet Yang Yuchun. If users choose to prevent visiting, it will affect the emotional development of Wang Meirong and Yang Yuchun, causing them to drift apart.

[0203] Users can trigger the appearance of multiple characters sequentially on the screen. When a user points their flashlight beam at an area, the corresponding character will instantly appear on the screen, playing their lines, actions, or songs to create a realistic dialogue scene. When the plot progresses to key points, the system displays multiple plot options on the screen. Users make branching choices by shining their light on different options, and the system uses these choices to invoke a preset non-linear plot path, determining the subsequent story direction and the characters' fates, thus forming a dynamic interactive process. During the interaction, user-generated Huagu Opera characters will appear on the screen as "supporting roles" to participate in the interactive performance, lowering the user experience threshold for those with little knowledge of Huagu Opera while giving them an immersive sense of participation.

[0204] In summary, the present invention discloses a digital opera generation and participation system that supports multimodal interaction, which has the following advantages: 1. It achieves accurate capture of user hand movements, allows users to select plot paths, enhances the interactivity and immersion of digital opera performances, and strengthens the gamified narrative ability of opera; 2. The interaction method is simple and intuitive, reducing the system threshold and difficulty of use, and has stronger scalability and universality.

[0205] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A digital opera generation and participation system supporting multimodal interaction, characterized in that: Includes interactive devices and multimodal interaction subsystems, among which, The interactive device includes: a light-shielding shell and a conical light outlet, used to define a conical area along the projection direction as a recognizable interactive area; a motion control module, used to collect hand skeleton and motion trajectory information within the recognizable interactive area; the interactive device also includes: an aperture adjustment knob, used to rotate and adjust the divergence angle of the light beam formed by the conical light outlet; a Z-axis displacement sensor, used to detect the forward and backward sliding displacement of the aperture adjustment knob and transmit the detected data to the central control unit to obtain the user's interactive intention in the Z-axis direction; the central control unit integrates UWB and IMU modules, used for spatial positioning and inertial measurement respectively; it is also connected to the aperture adjustment knob, Z-axis displacement sensor and motion control module respectively to achieve integrated interactive control; The multimodal interaction subsystem includes: The facial recognition module is used to acquire and preprocess facial images in real time, and extract key facial feature points from the preprocessed facial images. The digital opera character generation module is used to generate a fused digital face image based on the extracted key facial feature points and the input conditions of the digital face generation model, and output it to the rendering execution engine. The interactive plot control module is used to control plot nodes, schedule NPC character performances, and drive user branch selections based on the constructed opera drama interaction engine; at the same time, it performs three-dimensional coordinate recognition and event control mapping based on the hand skeleton and motion trajectory information collected by the interactive device, and realizes closed-loop interaction through hardware triggering mechanism. The rendering execution engine module is used to trigger the plot based on spatial interaction location and user interaction behavior, launch audio and video resources of preset plot nodes, and make plot branch selection and ending interpretation based on interaction prompts.

2. The digital opera generation and participation system supporting multimodal interaction according to claim 1, characterized in that, The interactive device also includes: The physical trigger switch is used to switch the selection state, and transmits the switching trigger signal to the multimodal interaction subsystem and binds it to the input signal in the opera repertoire interaction engine.

3. The digital opera generation and participation system supporting multimodal interaction according to claim 1, characterized in that, The facial recognition module includes: The camera unit is used to capture real-time images of the user's face; The image preprocessing unit is used to perform face localization, feature normalization, and pose correction on the acquired user facial images; The feature extraction unit is used to output the user's expression code and facial geometric parameters based on the preprocessing results and using a multi-layer neural network structure.

4. The digital opera generation and participation system supporting multimodal interaction according to claim 1, characterized in that, The digital opera character generation module includes: A face matching database is used to build and store a structured traditional opera face tag system, and to design a digital face generation model that takes the user's original facial image, facial feature vector, face style tag and facial key point heat map as input conditions to obtain the preset digital face. The digital face generation unit is used to map features from the face recognition module to preset digital faces to generate a fused digital face image. The output unit is used to load the generated fused digital face image into the rendering execution engine.

5. The digital opera generation and participation system supporting multimodal interaction according to claim 4, characterized in that, The digital opera character generation module also includes: The digital face image binding unit is used to establish a mapping relationship between the generated fused digital face image and the user's original facial image after loading the generated fused digital face image into the rendering execution engine, and bind it to the 3D mesh model. The expression parameter-driven mechanism building unit is used to extract multiple facial geometric features through expression recognition and convert them into standardized expression parameters, which are then bound to the animation controller or deformation node in the rendering execution engine.

6. The digital opera generation and participation system supporting multimodal interaction according to claim 1, characterized in that, The interactive storyline control module includes: The opera repertoire interaction engine is used to build the main scene container node, which contains a stage background layer, plot node structure, NPC digital character resources and user personalized image, and sets plot trigger conditions and multi-branch plot structure. The spatial recognition unit is used to construct a three-dimensional interactive coordinate model based on the hand skeleton and motion trajectory information collected by the interactive device, and to realize event control mapping.

7. The digital opera generation and participation system supporting multimodal interaction according to claim 6, characterized in that, The interactive story control module also includes: The hardware trigger control unit is used to determine the specific trigger event type after receiving the user's pressing action signal, combined with the current three-dimensional coordinate information, and to provide real-time feedback on the operation result on the screen in a visual / auditory manner.

8. The digital opera generation and participation system supporting multimodal interaction according to claim 6, characterized in that, The interactive story control module also includes: The path recording unit is used to automatically form a path tree structure from the user's participation process, and play the complete ending segment and display the generated subtitles at the final node.

9. The digital opera generation and participation system supporting multimodal interaction according to claim 1, characterized in that, The rendering execution engine module includes: The plot node triggering module is used to activate the NPC singing playback logic bound to the plot node based on the detected spatial interaction location and user interaction behavior. The singing segment resource scheduling module is used to synchronously load multimodal performance resources corresponding to the current singing segment performance content when NPC singing segments are played. The branch selection and jump unit is used to automatically provide multiple optional branch prompts after each plot node is completed, and to jump to the plot node corresponding to the selected branch after selecting the target branch prompt area through the interactive device.

Citation Information

Patent Citations

  • Branch plot generation method and device based on motion capture

    CN116243790A

  • Human body posture multi-view visual identification AI training data set automatic generation and identification method based on simulation environment

    CN119888024A