Virtual Reality Reproduction Method of Opera Performance Based on Motion Capture Technology
By constructing a method for multi-source sensor data acquisition and graph convolutional neural network driving, the application of virtual reality in opera performances was solved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU FOOD & PHARMA SCI COLLEGE
- Filing Date
- 2026-02-25
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies struggle to accurately reproduce the dynamic details and rhythms of traditional opera performances. Furthermore, virtual characters and cultural elements are treated in a fragmented manner, lacking cross-modal temporal alignment mechanisms, resulting in a lack of realism and immersion in virtual representations.
By constructing a multi-source heterogeneous sensor data acquisition system and integrating inertial, optical, eye-tracking, and facial expression data, a mapping model between the semantics of opera movements and the movement of virtual characters is established. A graph convolutional neural network is used to drive a high-fidelity digital human model, achieving high-precision reproduction of opera performances and simultaneously rendering stage lighting and sound effects.
It achieves high-fidelity reproduction of opera performances, enhances the realism and immersion of virtual inheritance, provides a quantifiable and interactive technical platform, and overcomes the problems of motion distortion, stiff facial expressions, and costume distortion.
Smart Images

Figure CN122086246A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer and virtual reality technology, specifically relating to a method for virtual reality reproduction of opera performances based on motion capture technology. Background Technology
[0002] With the deepening application of virtual reality technology in cultural heritage and digital art, the digital reproduction of traditional opera performances has become an important direction for the protection of intangible cultural heritage. Opera art, with its highly stylized movements, gestures, footwork, and unique costumes and props, constitutes a complete performance vocabulary system, its movements containing rich rhythms and cultural symbolism. Currently, virtual reality content production largely relies on general motion capture systems or pre-set animation sequences. While these can achieve three-dimensional reconstruction of basic human postures, they struggle to accurately reproduce the millimeter-level dynamic details and inherent rhythmic characteristics of opera movements, resulting in virtual character performances lacking realism and artistic appeal.
[0003] The virtual reproduction method of traditional Chinese opera based on motion capture technology aims to drive virtual characters to reproduce actors' performances through sensor data. Existing technologies generally employ a single type of capture method, which is susceptible to occlusion, electromagnetic interference, or drift errors in complex stage environments, making it impossible to stably capture details such as the high-frequency vibrations of flowing sleeves and the micro-amplitude resonances of feather movements. Furthermore, traditional Chinese opera performance is a highly integrated whole of singing, recitation, acting, and acrobatics, with movement rhythms tightly coupled with vocal temporal rhythms, facial makeup meanings, and lighting and sound effects. However, current virtual systems often process movement, audio, and visual elements separately, lacking cross-modal temporal alignment mechanisms, resulting in a fragmented virtual performance that is "similar in form but different in spirit."
[0004] Existing technologies for virtual reproduction of traditional Chinese opera suffer from the following problems: insufficient motion capture precision, making it difficult to simultaneously capture large-scale stage movements and detailed local movements; lack of modeling methods for the dynamic characteristics of traditional Chinese opera, making it impossible to extract and reproduce the rhythmic tension and aesthetic rhythm of movements from raw data; lack of semantic-level linkage between virtual character behavior and cultural elements such as singing style and lighting, weakening the cultural depth of the immersive experience; and limited audience interaction design to fixed perspectives or simple switching, failing to dynamically adjust the observation focus according to the user's gestures and lacking an adaptive balance between overall scene staging and close-up details.
[0005] The aforementioned problems severely restrict the high-fidelity inheritance and interactive dissemination of traditional opera in virtual space, and there is an urgent need for a virtual reality reproduction solution that deeply integrates motion capture, cultural semantics and temporal synchronization mechanisms. Summary of the Invention
[0006] This invention provides a virtual reality reproduction method for opera performances based on motion capture technology, aiming to solve the technical problem in existing technologies that fail to specifically integrate motion capture technology to achieve accurate virtual reality reproduction of opera performances, resulting in insufficient realism in the virtual inheritance of opera. This method constructs a motion capture data acquisition system oriented towards the artistic characteristics of opera performances, integrates multi-source heterogeneous sensor data, establishes a mapping model between the semantics of opera movements and the movement of three-dimensional virtual characters, and synchronously drives a high-fidelity digital human model in a virtual reality environment, achieving high-precision and high-fidelity reproduction of elements such as body posture, footwork, gestures, eye contact, and stage choreography in opera performances.
[0007] This invention provides a method for virtual reality reproduction of traditional Chinese opera performances based on motion capture technology, comprising: The system uses an array of inertial motion capture sensors placed on parts of the performer's body to collect real-time data on the movement of all joints during the performance; it also uses an optical motion capture system placed on the stage area to collect the performer's global position coordinates, movement trajectory, and orientation information in the stage space; it uses an eye-tracking device to collect data on the performer's gaze direction, fixation point, and blink frequency during the performance; and it uses a facial expression capture device to collect data on the performer's facial muscle movements. A 4D temporal action dataset containing spatial pose, limb movement, facial expression, and eye movement behavior is constructed; based on a stylized action library of traditional Chinese opera performance, the 4D temporal action dataset is semantically annotated. The labeled motion data is input into the pre-trained opera motion-digital human driving mapping model. The opera motion-digital human driving mapping model adopts a graph convolutional neural network architecture, with the human skeleton topology as the graph node and the joint rotation angle, displacement vector and facial expression weight coefficient as the graph edge features, and outputs a sequence of driving parameters that are adapted to the target virtual character skeleton binding system. A high-fidelity digital human model of traditional Chinese opera is loaded into a virtual reality rendering engine. The high-fidelity digital human model of traditional Chinese opera has the characteristics of traditional opera roles, such as clothing textures, makeup details, headdress structure and dynamic fabric physical properties. Based on the driving parameter sequence, the digital human model is driven in real time in the virtual reality environment to perform opera movements that are completely consistent with the original performance, and the stage lights, background scenery and sound effects are rendered simultaneously to generate an immersive opera performance virtual reality scene; the immersive opera performance virtual reality scene is presented to the user through a virtual reality head-mounted display device, and the user can switch viewing angles, zoom distance and replay performance segments in the virtual stage space.
[0008] Preferably, an array of inertial motion capture sensors arranged on parts of the performer's body is used to collect real-time data on the movement of all joints of the opera actor during the performance, including: The inertial motion capture sensor array uses a nine-axis microelectromechanical system sensor, including a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer; Each sensor is encapsulated in a flexible fabric strap that conforms to the curves of the human body. It transmits data to the central processing unit via a wireless communication module, using the low-latency Bluetooth 5.0 standard as the communication protocol. The collected data includes the triaxial linear acceleration, triaxial angular velocity, and triaxial magnetic field strength of each joint in the local coordinate system, which are used to calculate the three-dimensional rotational attitude of the joint through complementary filtering or Kalman filtering algorithms.
[0009] Preferably, an optical motion capture system deployed in the stage area is used to simultaneously collect the performer's global position coordinates, movement trajectory, and orientation information within the stage space, including: The optical motion capture system consists of multiple infrared high-speed cameras, which are arranged around and above the stage to form a three-dimensional visual field covering the entire performance area. Each camera is equipped with an active infrared LED marker recognition module to track reflective spheres attached to performers’ costumes or sensor housings; The three-dimensional coordinates of the marker points in the global space are calculated using the principle of triangulation. Then, multiple marker points are combined into rigid body units including the head, torso, and limbs using a rigid body fitting algorithm to derive the overall pose.
[0010] Preferably, data on the performer's gaze direction, fixation point, and blink frequency during the performance are collected using an eye-tracking device, including: The eye-tracking device uses a combination of a near-infrared light source and a high frame rate image sensor, and is mounted on a lightweight headband. Calculate the line-of-sight vector using the pupil-corneal reflex principle; A blink event is determined by the instantaneous decrease in the visible area of the pupil. When the area is below a threshold and the duration is between 100 and 400 milliseconds, it is recorded as a complete blink.
[0011] Preferably, facial muscle movement data of the performer is collected using a facial expression capture device, including: The facial expression capture device uses a combination of a structured light projector and a depth camera to project an coded grating pattern onto the performer's face and reconstructs a three-dimensional deformable mesh of the face through a phase shift algorithm. The reconstructed 3D mesh is output per frame, with each vertex containing X, Y, Z coordinates and normal vector information; The system pre-collects a reference grid of the performer's neutral facial expressions. In each subsequent frame, vertex displacement difference calculation is performed between the grid and the reference grid to obtain the facial expression variation vector. After dimensionality reduction by principal component analysis, the first 50 principal components of the facial expression variation vector are extracted as facial expression feature coefficients.
[0012] Preferably, a 4-dimensional temporal action dataset is constructed, including spatial pose, limb movement, facial expression, and eye movement, comprising: Timestamp alignment is achieved through a hardware synchronization trigger signal, with all sensors receiving the same pulse signal upon startup; Coordinate system one transforms the local coordinate systems of each sensor to a global right-handed coordinate system with the stage center as the origin through a calibration matrix; All data were resampled to 200 Hz along a unified time axis to form structured data frames. Each frame contains quaternions of full-body joint rotation, global position and orientation of the torso, principal component coefficients of facial expressions, eye movement direction vectors, and blinking status.
[0013] Preferably, based on a stylized action library of traditional Chinese opera performances, the 4-dimensional temporal action dataset is annotated with action semantics, including: The stylized movement library for opera performance contains no fewer than 500 standard opera movement templates. Each template defines the movement name, the role it belongs to, the applicable repertoire, the starting posture, the ending posture, the keyframe sequence, the rhythm and beat, and the emotional tag. The annotation process adopts a semi-automatic workflow: the input action sequence is matched with the templates in the library using a dynamic time warping algorithm, the top three candidate templates with the highest matching degree are selected, the final labels are confirmed by opera experts and the start and end time boundaries of the actions are manually adjusted; The annotation results are appended to the original data frame as structured metadata, with zero or more action semantic tags associated with each time point, supporting the overlay annotation of compound actions.
[0014] Preferably, the labeled motion data is input into a pre-trained opera motion-digital human driven mapping model, including: The graph convolutional neural network contains five graph convolutional layers, each followed by batch normalization and modified linear unit activation functions; The model input consists of 60 frames of 4D temporal motion data captured by a sliding window; The graph nodes include multiple key joints of the human body, and the graph edges are constructed based on the anatomical connections of the human body to form an adjacency matrix; The output layer is a fully connected layer that embeds and maps the final graph into the driving parameter space, including bone rotation quaternions, face blendshape weights, eye rotation angles, and cloth control parameters.
[0015] Preferably, loading a high-fidelity opera digital human model into a virtual reality rendering engine includes: The high-fidelity opera digital human model uses physically based rendering materials, including diffuse, normal, roughness, metallicity and transparency maps. The model's skeletal system contains 120 drive joints, and the number of facial blendshapes is no less than 150. The headdress and armor components were simulated using rigid body dynamics, while the water sleeves and robes were simulated using finite element fabric simulation.
[0016] Preferably, the immersive opera performance virtual reality scene is presented to the user through a virtual reality head-mounted display device, and the user can switch viewing angles, zoom distances, and replay performance segments in the virtual stage space, including: Users wear virtual reality head-mounted displays that support six degrees of freedom tracking, with a refresh rate of 90 Hz; The system tracks the user's head position and orientation in real time and dynamically updates the rendering perspective; The interactive functions are implemented through a handheld controller, allowing users to trigger menus to select the playback start point, adjust the viewing distance, switch to a fixed camera position, or enable slow motion playback.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention is the first to construct a multimodal motion capture and virtual reproduction technology system oriented towards the characteristics of opera performance art. By integrating four types of high-precision sensor data, namely inertial, optical, eye-tracking and facial expression, it can completely capture all the artistic elements of "hand, eye, body, method and step" in opera performance. 2. By using graph convolutional neural networks to establish a nonlinear mapping relationship between stylized movements in traditional Chinese opera and driving parameters of digital humans, the technical problem that general motion capture systems cannot accurately reproduce the unique body movements and rhythms of traditional Chinese opera is solved. 3. The constructed high-fidelity digital human model strictly follows the costumes, makeup, and dynamic physical characteristics of traditional opera roles. Combined with precisely synchronized stage lighting and sound effects, it achieves a highly realistic reproduction of opera performance in a virtual reality environment. 4. This method not only enhances the realism and immersion of digital inheritance of traditional opera, but also provides a new type of technology platform that is quantifiable, traceable, and interactive for opera teaching, research and dissemination, overcoming the problems of motion distortion, stiff facial expressions, costume distortion and lack of stage atmosphere in existing virtual reproduction technology. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the overall technical solution architecture of the present invention; Figure 2 This is a schematic diagram of the core principle framework of the opera action-digital human driven mapping model in this invention; Figure 3 This is a logical flowchart of the multi-source heterogeneous motion capture data fusion processing in this invention; Figure 4This is a logical flowchart of the semantic annotation and driving parameter generation of stylized actions in traditional Chinese opera in this invention. Figure 5 This is a logical flow diagram of the synchronous rendering of the high-fidelity opera digital human model and the virtual stage environment in this invention. Figure 6 This is a schematic diagram of the multi-level interaction relationship and data flow of the four types of sensing systems in this invention: inertial, optical, eye-tracking, and facial expression. Detailed Implementation
[0019] refer to Figures 1 to 6 This invention provides a virtual reality reproduction method for opera performances based on motion capture technology. It aims to solve the technical problem in existing technologies where motion capture technology is not specifically integrated to achieve accurate virtual reality reproduction of opera performances, resulting in insufficient realism in virtual opera transmission. This method constructs a motion capture data acquisition system oriented towards the characteristics of opera performance art, integrates multi-source heterogeneous sensor data, establishes a mapping model between the semantics of opera movements and the movement of three-dimensional virtual characters, and synchronously drives a high-fidelity digital human model in a virtual reality environment. This achieves high-precision and high-fidelity reproduction of elements such as body posture, footwork, gestures, eye contact, and stage choreography in opera performances.
[0020] The method includes the following steps: S1, through an array of inertial motion capture sensors placed on parts of the performer's body, collects real-time data on the movement of all joints of the opera actor during the performance. S2, through an optical motion capture system deployed in the stage area, synchronously collects the global position coordinates, movement trajectory and orientation information of the performers in the stage space; S3 uses an eye-tracking device to collect data on the performer's gaze direction, fixation point, and blink frequency during the performance. S4, collects facial muscle movement data of the performer through a facial expression capture device; S5 performs timestamp alignment, coordinate system unification, and data fusion processing on multi-source heterogeneous motion capture data to form a 4-dimensional temporal motion dataset containing spatial pose, limb movement, facial expression, and eye movement behavior. S6. Based on the stylized action library of traditional Chinese opera performance, perform action semantic annotation on the 4-dimensional temporal action dataset; S7, input the labeled motion data into the pre-trained opera motion-digital human driving mapping model, and output the driving parameter sequence adapted to the target virtual character skeleton binding system; S8 loads a high-fidelity opera digital human model into the virtual reality rendering engine; S9, based on the driving parameter sequence, drive the digital human model in real time in the virtual reality environment to perform opera movements that are completely consistent with the original performance, and simultaneously render stage lights, background scenery and sound effects to generate an immersive opera performance virtual reality scene. S10 presents the immersive opera performance virtual reality scene to the user through a virtual reality head-mounted display device, and supports the user to freely switch viewing angles, zoom distances, and replay specific performance segments in the virtual stage space.
[0021] In step S1, an array of inertial motion capture sensors, deployed on various parts of the performer's body, collects real-time data on the joint movements of the entire body during the performance. These parts include the head, shoulders, elbows, wrists, hips, knees, ankles, and finger joints. The inertial motion capture sensor array employs a nine-axis microelectromechanical system (MEMS) sensor, comprising a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer, with a sampling frequency of at least 200 Hz and an angular resolution of at least 0.1 degrees. Each sensor is encapsulated within a flexible fabric strap, conforming to the body's curvature to ensure it does not slip or detach during large-scale opera movements such as rolling and leaping. The sensors transmit data to the central processing unit via a wireless communication module, using the low-latency Bluetooth 5.0 standard, with a maximum transmission distance of 10 meters and a data packet loss rate of less than 1 / 10000. The collected data includes the three-axis linear acceleration, three-axis angular velocity, and three-axis magnetic field strength of each joint in the local coordinate system, which are used to subsequently calculate the three-dimensional rotational attitude of the joints using complementary filtering or Kalman filtering algorithms. All sensors must undergo static calibration before startup to eliminate zero bias error and scale factor deviation. The calibration process lasts for 30 seconds, during which the actors remain upright and still.
[0022] In step S2, an optical motion capture system deployed in the stage area synchronously acquires the performer's global position coordinates, movement trajectory, and orientation information within the stage space. This optical motion capture system consists of no fewer than 12 high-speed infrared cameras, positioned on the stage walls and ceiling trusses to form a stereoscopic visual field covering the entire performance area. The camera frame rate is no less than 120 frames per second, and the spatial positioning accuracy is no less than 0.5 millimeters. Each camera is equipped with an active infrared LED marker recognition module to track reflective spheres attached to the performer's clothing or sensor housing. The system calculates the three-dimensional coordinates of the markers in global space using triangulation principles and combines multiple markers into rigid body units such as the head, torso, and limbs using a rigid body fitting algorithm, thereby deriving the overall pose. The optical system and inertial system are triggered by the same hardware synchronization signal to ensure strict alignment of the two systems on the time axis. The stage area is pre-calibrated spatially. Images are acquired at multiple locations using calibration rods of known size. The intrinsic and extrinsic parameters of the cameras are calculated, and a right-handed global coordinate system is established with the center of the stage as the origin, the X-axis pointing in front of the stage, the Y-axis pointing to the left side of the stage, and the Z-axis pointing vertically upward.
[0023] In step S3, an eye-tracking device collects data on the performer's gaze direction, fixation point, and blink frequency during the performance. The eye-tracking device combines a near-infrared light source with a high frame rate image sensor, mounted on a lightweight headband, positioned approximately 3 cm below the performer's brow bone and close to the eyeball. The device calculates the gaze vector based on the pupil-corneal reflection principle, with a sampling frequency of at least 60 Hz and a gaze tracking accuracy of at least 0.5 degrees. The near-infrared light source emits invisible light with a wavelength of 850 nanometers to avoid interfering with the performer's visual perception. The image sensor captures grayscale images of the eye area at 120 frames per second, extracting the coordinates of the pupil center and corneal reflection point using edge detection and ellipse fitting algorithms, and then calculating the unit direction vector of the gaze in three-dimensional space. Blink events are determined by the instantaneous decrease in the visible pupil area; when the area falls below a threshold and the duration is between 100 and 400 milliseconds, it is recorded as a complete blink. Eye-tracking data is output as an independent time series, including timestamps, gaze direction vectors, gaze point world coordinates (obtained by fusion with the stage coordinate system), and blink status flags.
[0024] In step S4, facial muscle movement data of the performer is collected using a facial expression capture device, including movement parameters of the brow ridge, orbicularis oculi muscle, orbicularis oris muscle, zygomaticus major muscle, and mandible. The facial expression capture device uses a combination of a structured light projector and a depth camera. The projector projects an coded grating pattern onto the performer's face, and the depth camera simultaneously captures the deformed grating image. A three-dimensional deformable mesh of the face is reconstructed using a phase shift algorithm, with at least 5000 vertices and an update frequency of at least 30 Hz. The projector and camera are fixed on an adjustable bracket 3 meters in front of the stage to ensure the performer's face is always within the effective working distance. The reconstructed three-dimensional mesh is output frame by frame, with each vertex containing X, Y, Z coordinates and normal vector information. The system pre-collects a reference mesh for the performer's neutral expression. For each subsequent frame, vertex displacement difference calculations are performed between the mesh and the reference mesh to obtain the facial expression deformation vector. After dimensionality reduction using principal component analysis, the first 50 principal components of this facial expression deformation vector are extracted as expression feature coefficients, which are used to drive the digital human face blendshape weights.
[0025] In step S5, the aforementioned multi-source heterogeneous motion capture data undergoes timestamp alignment, coordinate system one, and data fusion processing to form a 4D temporal motion dataset containing spatial pose, limb movement, facial expression, and eye movement. Timestamp alignment is achieved through a hardware synchronization trigger signal; all sensors receive the same pulse signal upon startup, ensuring consistent data acquisition start times and a time synchronization error of less than 1 millisecond. Coordinate system one transforms the local coordinate systems of each sensor to a global right-handed coordinate system with the stage center as the origin using a calibration matrix. Specifically, the attitude data from the inertial sensors is first corrected for heading angle using a magnetometer, and then, combined with the rigid body pose of the torso provided by the optical system, the rotation matrix of each limb segment in the global coordinate system is solved using least-squares optimization. The gaze direction vector in the eye movement data is transformed to the global coordinate system using the rigid body transformation matrix of the headgear. The reference coordinate system of the facial expression mesh is aligned to the global system using the rigid body pose of the head. Finally, all data were resampled to 200 Hz along a unified time axis to form structured data frames. Each frame contains: full-body joint rotation quaternions (26 joints in total), global trunk position and orientation, facial expression principal component coefficients (50 dimensions), eye movement gaze direction vector, and blink status. This 4-dimensional temporal action dataset is stored as a continuous temporal tensor with dimensions of time step × feature dimension, used for subsequent semantic annotation and model driving.
[0026] In step S6, based on the stylized movement library of Chinese opera performances, semantic annotation of actions is performed on the 4D temporal action dataset. The action semantics include typical Chinese opera postures such as Qiba, Zoubian, Tangma, Shuixiu, Yunshou, Shanbang, and Liangxiang. The stylized movement library of Chinese opera performances contains no less than 500 standard Chinese opera action templates. Each template defines the action name, the acting role, the applicable repertoire, the starting posture, the ending posture, the key frame sequence, the rhythm beat, and the emotion label. The annotation process adopts a semi-automatic process: First, the input action sequence is matched with the templates in the library through the dynamic time warping algorithm, and the top three candidate templates with the highest matching degree are selected; Subsequently, the Chinese opera experts confirm the final label in the graphical interface and manually adjust the start and end time boundaries of the action. The annotation results are attached to the original data frame as structured metadata. Each time point can be associated with zero or more action semantic labels, supporting the superimposed annotation of composite actions such as "Yunshou + Liangxiang". The rhythm beat information is extracted through peak detection of the action speed curve, and the emotion label is jointly determined according to the action amplitude, speed, and facial expressions.
[0027] In step S7, the annotated action data is input into a pre-trained mapping model of Chinese opera actions - digital human driving. This mapping model of Chinese opera actions - digital human driving adopts a graph convolutional neural network architecture. Taking the human body bone topology structure as the graph nodes and the joint rotation angle, displacement vector, and expression weight coefficient as the graph edge features, it outputs a sequence of driving parameters adapted to the target virtual character's bone binding system. The graph convolutional neural network contains five graph convolutional layers, each followed by batch normalization and a rectified linear unit activation function. The input dimension is the total number of joint degrees of freedom multiplied by the time step, and the output dimension is the number of driving channels of the target digital human bone system. The model input is 60-frame 4D temporal action data intercepted by a sliding window, corresponding to a 300-millisecond action segment. The graph nodes include 26 main human body joint points, and the graph edges construct an adjacency matrix based on the human anatomy connection relationship. Each graph convolutional operation updates the current node representation by aggregating the features of neighbor nodes, gradually extracting hierarchical features from local joints to the overall posture. The output layer is a fully connected layer that maps the final graph embedding to the driving parameter space, including bone rotation quaternions, facial blendshape weights, eyeball rotation angles, and cloth control parameters. During the training process, the mean squared error loss function and the action smoothness regularization term are jointly optimized. The loss function is defined as: ; is the th real driving parameter, , , , are the , , , Each model predicts a value. The total number of samples, For time step, and The weighting coefficients are set to 1.0 and 0.1 respectively. The regularization term constrains the second-order difference of the predicted sequence, suppressing high-frequency jitter and ensuring smooth operation.
[0028] In step S8, a high-fidelity opera digital human model is loaded into the virtual reality rendering engine. This digital human model possesses clothing textures, makeup details, headdress structures, and dynamic fabric physical properties that conform to the characteristics of traditional opera roles. The high-fidelity opera digital human model uses physically based rendering materials, with clothing texture resolution of 8192 pixels × 8192 pixels, including diffuse, normal, roughness, metallicity, and transparency maps. The model's skeletal system includes 120 driven joints, and the number of facial blendshapes is no less than 150, covering all typical opera expressions. The headdress and armor components are simulated using rigid body dynamics, with collision detection to prevent clipping; the water sleeves and robes are simulated using finite element fabric simulation, with a drag coefficient set to 0.35, an elastic modulus set to 1.2 MPa, and a Poisson's ratio set to 0.4. The digital human models are stored according to their respective roles, including five major categories: Sheng (male lead), Dan (female lead), Jing (painted face), Mo (old man), and Chou (clown). Each category contains multiple role variations, such as Lao Sheng (old male lead), Xiao Sheng (young male lead), Qing Yi (young female lead), and Hua Dan (young female lead), ensuring the historical accuracy of costumes and makeup.
[0029] In step S9, based on the driving parameter sequence, the digital human model is driven in real-time within the virtual reality environment to perform opera movements completely identical to the original performance. Simultaneously, stage lighting, background scenery, and sound effects are rendered to generate an immersive virtual reality scene of the opera performance. The virtual reality rendering engine uses a forward rendering pipeline, supports real-time ray tracing shadows and ambient occlusion, maintains a stable frame rate of over 90 frames per second, and has a latency of less than 20 milliseconds. The stage lighting system replicates the top lighting, front lighting, side lighting, and follow spot configuration of a traditional opera stage, with a color temperature range of 2700 Kelvin to 6500 Kelvin and an adjustable illuminance range of 100 lux to 2000 lux. The background scenery uses high dynamic range panoramic images or 3D modeled scenes, automatically switching according to the repertoire. The sound system synchronously plays the original performance recording or multi-channel surround sound mix, with the sound source position bound to the digital human's mouth position to achieve audio-visual synchronization. The driving parameter sequence is input to the rendering engine on a per-frame basis. The model geometry is updated through skeletal skinning and blendshape interpolation, and the cloth and rigid body system respond to physical constraints in real time.
[0030] In step S10, the immersive opera performance virtual reality scene is presented to the user through a virtual reality head-mounted display device, allowing the user to freely switch viewing angles, zoom distances, and replay specific performance segments within the virtual stage space. The user wears a virtual reality head-mounted display device supporting six degrees of freedom tracking, with a refresh rate of 90 Hz and a field of view of no less than 110 degrees. The system tracks the user's head position and orientation in real time, dynamically updating the rendering perspective. Interactive functions are implemented through a handheld controller, allowing the user to trigger menus to select the playback starting point, adjust the viewing distance (ranging from 1 meter to 20 meters), switch fixed camera positions (such as the audience seats, stage wing, or aerial view), or enable slow-motion playback (adjustable speed ranging from 0.1x to 2.0x). All interactive operations do not affect the background data stream processing, ensuring the integrity and consistency of the performance reproduction.
[0031] The system comprises an inertial motion capture data acquisition unit, an optical motion capture data acquisition unit, an eye-tracking behavior data acquisition unit, a facial expression data acquisition unit, a multi-source data fusion processing unit, a traditional Chinese opera action semantic annotation unit, a digital human driving parameter generation unit, a virtual reality scene rendering unit, and an immersive presentation unit. The inertial motion capture data acquisition unit consists of multiple nine-axis microelectromechanical system (MEMS) sensors, positioned on various parts of the performer's body, responsible for collecting full-body joint motion data. The optical motion capture data acquisition unit consists of multiple infrared high-speed cameras, covering the stage area, used to acquire global pose information. The eye-tracking behavior data acquisition unit integrates a near-infrared light source and image sensors, mounted on a head-mounted device, to collect gaze and blink data. The facial expression data acquisition unit uses structured light projection and depth imaging technology to reconstruct three-dimensional facial deformation. The multi-source data fusion processing unit performs time synchronization, coordinate transformation, and data resampling to generate a unified 4D time-series dataset. The traditional Chinese opera action semantic annotation unit semantically labels action sequences based on a stylized action library. The digital human driving parameter generation unit runs a graphical convolutional neural network model and outputs driving instructions. The virtual reality scene rendering unit loads high-fidelity digital human models and performs real-time rendering. The immersive presentation unit delivers the final experience to the user through a virtual reality headset and supports interactive control. All units are interconnected via a high-speed local area network, and data transmission uses a time-sensitive network protocol to ensure controllable end-to-end latency.
[0032] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0033] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for virtual reality reproduction of traditional Chinese opera performances based on motion capture technology, characterized in that: include: By deploying an array of inertial motion capture sensors on parts of the performer's body, real-time data on the movement of all joints of the opera actor during the performance is collected. An optical motion capture system deployed in the stage area is used to simultaneously collect the performer's global position coordinates, movement trajectory, and orientation information in the stage space; an eye-tracking device is used to collect data on the performer's gaze direction, fixation point, and blink frequency during the performance. Facial muscle movement data of performers is collected using a facial expression capture device; A 4D temporal action dataset containing spatial pose, limb movement, facial expression, and eye movement behavior is constructed; based on a stylized action library of traditional Chinese opera performance, the 4D temporal action dataset is semantically annotated. The labeled motion data is input into the pre-trained opera motion-digital human driving mapping model. The opera motion-digital human driving mapping model adopts a graph convolutional neural network architecture, with the human skeleton topology as the graph node and the joint rotation angle, displacement vector and facial expression weight coefficient as the graph edge features, and outputs a sequence of driving parameters that are adapted to the target virtual character skeleton binding system. A high-fidelity digital human model of traditional Chinese opera is loaded into a virtual reality rendering engine. The high-fidelity digital human model of traditional Chinese opera has the characteristics of traditional opera roles, such as clothing textures, makeup details, headdress structure and dynamic fabric physical properties. Based on the driving parameter sequence, the digital human model is driven in real time in the virtual reality environment to perform opera movements that are completely consistent with the original performance, and the stage lights, background scenery and sound effects are rendered simultaneously to generate an immersive opera performance virtual reality scene; the immersive opera performance virtual reality scene is presented to the user through a virtual reality head-mounted display device, and the user can switch viewing angles, zoom distance and replay performance segments in the virtual stage space.
2. The virtual reality reproduction method for opera performance based on motion capture technology according to claim 1, characterized in that, By deploying an array of inertial motion capture sensors on various parts of the performer's body, real-time data on the movement of all joints during the performance is collected, including: The inertial motion capture sensor array uses a nine-axis microelectromechanical system sensor, including a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer; Each sensor is encapsulated in a flexible fabric strap that conforms to the curves of the human body. It transmits data to the central processing unit via a wireless communication module, using the low-latency Bluetooth 5.0 standard as the communication protocol. The collected data includes the triaxial linear acceleration, triaxial angular velocity, and triaxial magnetic field strength of each joint in the local coordinate system, which are used to calculate the three-dimensional rotational attitude of the joint through complementary filtering or Kalman filtering algorithms.
3. The virtual reality reproduction method for opera performance based on motion capture technology according to claim 2, characterized in that, An optical motion capture system deployed in the stage area synchronously collects the performer's global position coordinates, movement trajectory, and orientation information within the stage space, including: The optical motion capture system consists of multiple infrared high-speed cameras, which are arranged around and above the stage to form a three-dimensional visual field covering the entire performance area. Each camera is equipped with an active infrared LED marker recognition module to track reflective spheres attached to performers’ costumes or sensor housings; The three-dimensional coordinates of the marker points in the global space are calculated using the principle of triangulation. Then, multiple marker points are combined into rigid body units including the head, torso, and limbs using a rigid body fitting algorithm to derive the overall pose.
4. The virtual reality reproduction method for opera performance based on motion capture technology according to claim 3, characterized in that, Data on the performer's gaze direction, fixation point, and blink frequency during the performance are collected using an eye-tracking device, including: The eye-tracking device uses a combination of a near-infrared light source and a high frame rate image sensor, and is mounted on a lightweight headband. Calculate the line-of-sight vector using the pupil-corneal reflex principle; A blink event is determined by the instantaneous decrease in the visible area of the pupil. When the area is below a threshold and the duration is between 100 and 400 milliseconds, it is recorded as a complete blink.
5. The virtual reality reproduction method for opera performance based on motion capture technology according to claim 4, characterized in that, Facial muscle movement data of performers is collected using facial expression capture devices, including: The facial expression capture device uses a combination of a structured light projector and a depth camera to project an coded grating pattern onto the performer's face and reconstructs a three-dimensional deformable mesh of the face through a phase shift algorithm. The reconstructed 3D mesh is output per frame, with each vertex containing X, Y, Z coordinates and normal vector information; The system pre-collects a reference grid of the performer's neutral facial expressions. In each subsequent frame, vertex displacement difference calculation is performed between the grid and the reference grid to obtain the facial expression variation vector. After dimensionality reduction by principal component analysis, the first 50 principal components of the facial expression variation vector are extracted as facial expression feature coefficients.
6. The virtual reality reproduction method for opera performance based on motion capture technology according to claim 5, characterized in that, A 4-dimensional temporal motion dataset containing spatial pose, limb movement, facial expression, and eye movement behavior was constructed, including: Timestamp alignment is achieved through a hardware synchronization trigger signal, with all sensors receiving the same pulse signal upon startup; Coordinate system one transforms the local coordinate systems of each sensor to a global right-handed coordinate system with the stage center as the origin through a calibration matrix; All data were resampled to 200 Hz along a unified time axis to form structured data frames. Each frame contains quaternions of full-body joint rotation, global position and orientation of the torso, principal component coefficients of facial expressions, eye movement direction vectors, and blinking status.
7. The virtual reality reproduction method for opera performance based on motion capture technology according to claim 6, characterized in that, Based on a stylized action library from traditional Chinese opera performances, semantic annotation of the 4-dimensional temporal action dataset is performed, including: The stylized movement library for opera performance contains no fewer than 500 standard opera movement templates. Each template defines the movement name, the role it belongs to, the applicable repertoire, the starting posture, the ending posture, the keyframe sequence, the rhythm and beat, and the emotional tag. The annotation process adopts a semi-automatic workflow: the input action sequence is matched with the templates in the library using a dynamic time warping algorithm, the top three candidate templates with the highest matching degree are selected, the final labels are confirmed by opera experts and the start and end time boundaries of the actions are manually adjusted; The annotation results are appended to the original data frame as structured metadata, with zero or more action semantic tags associated with each time point, supporting the overlay annotation of compound actions.
8. The virtual reality reproduction method for opera performance based on motion capture technology according to claim 7, characterized in that, The labeled motion data is input into a pre-trained opera motion-digital human driven mapping model, including: The graph convolutional neural network contains five graph convolutional layers, each followed by batch normalization and modified linear unit activation functions; The model input consists of 60 frames of 4D temporal motion data captured by a sliding window; The graph nodes include multiple key joints of the human body, and the graph edges are constructed based on the anatomical connections of the human body to form an adjacency matrix; The output layer is a fully connected layer that embeds and maps the final graph into the driving parameter space, including bone rotation quaternions, face blendshape weights, eye rotation angles, and cloth control parameters.
9. The virtual reality reproduction method for opera performance based on motion capture technology according to claim 8, characterized in that, Loading high-fidelity opera digital human models into a virtual reality rendering engine includes: The high-fidelity opera digital human model uses physically based rendering materials, including diffuse, normal, roughness, metallicity and transparency maps. The model's skeletal system contains 120 drive joints, and the number of facial blendshapes is no less than 150. The headdress and armor components were simulated using rigid body dynamics, while the water sleeves and robes were simulated using finite element fabric simulation.
10. The virtual reality reproduction method for opera performance based on motion capture technology according to claim 9, characterized in that, The immersive opera performance virtual reality scene is presented to the user through a virtual reality head-mounted display device, and the user can switch viewing angles, zoom distances, and replay performance clips in the virtual stage space, including: Users wear virtual reality head-mounted displays that support six degrees of freedom tracking, with a refresh rate of 90 Hz; The system tracks the user's head position and orientation in real time and dynamically updates the rendering perspective; The interactive functions are implemented through a handheld controller, allowing users to trigger menus to select the playback start point, adjust the viewing distance, switch to a fixed camera position, or enable slow motion playback.