Immersive deduction system and method based on 3D projection fusion
By using a perception module and NeRF geometric correction technology, combined with edge computing rendering, the problem of projection effect adaptation deviation caused by the height difference of actors was solved, achieving a highly synchronized and low-latency immersive performance effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-07
AI Technical Summary
Existing 3D projection fusion technology fails to effectively consider the height differences of actors in immersive performance scenarios such as martial arts and dance, resulting in large discrepancies between the projection effects and the body trajectories, severe delays in motion prediction, and negatively impacting the viewing experience.
The system uses a perception module to collect actors' fast-moving data and height parameters in real time. A mapping relationship between height and movement amplitude and trajectory span is established through a motion association sub-model. Combined with NeRF geometric correction and edge computing rendering, the system achieves precise synchronization of projection effects.
It achieves precise synchronization between projection effects and actor movements, improves the synchronicity and integrity of the immersive experience, reduces rendering latency, adapts to the movement characteristics of different actors, and ensures long-term stable operation of the system.
Smart Images

Figure CN121810992A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of 3D projection technology, and in particular relates to an immersive performance system and method based on 3D projection fusion. Background Technology
[0002] In immersive performance scenarios such as martial arts and dance, the synergy between actors' rapid movements and projection effects is key to enhancing the viewing experience. However, existing 3D projection fusion technology has many shortcomings: Traditional systems often use fixed-parameter rendering, failing to consider the impact of actors' height differences on the range and trajectory of their movements. This results in significant discrepancies between the projection effects and the limb trajectories when actors of different heights perform the same action, leading to a fragmented sense of immersion. Furthermore, existing motion capture and rendering processes suffer from noticeable delays. In fast-paced action scenes, projection effects often lag behind the actors' movements. Moreover, motion prediction relies solely on single temporal data without incorporating height-related features, resulting in prediction errors often exceeding 1cm. This fails to meet the requirements for precise linkage and severely impacts the viewing experience, necessitating urgent improvements. Therefore, we propose an immersive performance system and method based on 3D projection fusion. Summary of the Invention
[0003] In view of the problem that the projection effects are out of sync with the actors' movements in the above-mentioned or existing technologies, this invention is proposed.
[0004] In view of this, the present invention provides an immersive performance system based on 3D projection fusion, including a perception module for collecting actors' rapid movement data and height parameters. The rapid movement data includes limb trajectories, joint angles, and movement speed data in martial arts and dance scenes. A motion prediction module is communicatively connected to the perception module and includes a motion association sub-model and a motion timing prediction sub-module. The motion association sub-model establishes a mapping relationship between the actor's height and the amplitude and trajectory span of the movements. The motion timing prediction sub-module predicts the motion trajectory of the next 1-2 frames based on preceding motion data and height parameters. A rendering output module is also included. The rendering output module is communicatively connected to the motion prediction module, including a NeRF geometric correction unit, an edge computing rendering unit, and a projector array. The NeRF geometric correction unit generates a 3D dense mesh model of the stage carrier and dynamically adjusts the projection parameters. The edge computing rendering unit adopts a layered rendering strategy of keyframe pre-rendering + non-keyframe lightweight compression. The projector array synchronously projects the rendered projection effects content. An optimization module is communicatively connected to the perception module and the motion prediction module, respectively, and is used to incrementally train and optimize the parameters of the motion association sub-model and the motion timing prediction sub-module based on the deviation data between the actual motion and the predicted trajectory.
[0005] As a preferred embodiment of the immersive performance system based on 3D projection fusion of the present invention, the perception module includes an optical motion capture camera and an inertial motion capture sensor. The optical motion capture camera and the inertial motion capture sensor work together to achieve the complete acquisition of rapid motion data. The perception module also includes an actor height input unit, which is used to input actor height data and associate it with the corresponding actor's motion acquisition data stream.
[0006] As a preferred embodiment of the immersive performance system based on 3D projection fusion of the present invention, the mapping relationship of the action association sub-model satisfies the following: under the same action, the limb swing amplitude and trajectory span of the tall actor are greater than those of the short actor. The mapping relationship is generated by training a labeled dataset containing actors of different heights and different types of fast actions. The action types in the dataset cover the core movements of martial arts and dance, such as arm swinging, kicking, rotating and jumping.
[0007] As a preferred embodiment of the immersive performance system based on 3D projection fusion of the present invention, the motion timing prediction submodule inputs the actor's motion sequence data of the previous 60 frames and height parameters, and outputs the motion trajectory prediction result for the next 16.7-33.4ms, with the deviation between the predicted trajectory and the actual motion ≤0.5cm.
[0008] As a preferred embodiment of the immersive performance system based on 3D projection fusion of the present invention, the NeRF geometric correction unit includes a lidar and an RGB-D camera. The generated 3D dense mesh model can dynamically adapt to the subtle deformation of the stage carrier and adjust the projection matrix parameters in real time to ensure that the projection effects are accurately matched with the stage carrier.
[0009] As a preferred embodiment of the immersive performance system based on 3D projection fusion of the present invention, the edge computing rendering unit includes at least 4 edge computing nodes.
[0010] As a preferred embodiment of the immersive performance system based on 3D projection fusion of the present invention, the number of projectors in the projector array is N≥6, the overlap rate of the projection range of adjacent projectors is 15%-20%, and all projectors achieve synchronous projection of images through a synchronous control module to ensure seamless presentation of projection effects throughout the stage.
[0011] As a preferred embodiment of the immersive performance system based on 3D projection fusion of the present invention, the optimization module optimizes every 3 minutes. It obtains deviation data by calculating the Euclidean distance between the predicted trajectory and the actual action trajectory. When the mean deviation data is >0.5cm, it initiates incremental training of the sub-model parameters. The training iterations are 100-200 times until the mean deviation data is ≤0.5cm.
[0012] As a preferred embodiment of the immersive performance method based on 3D projection fusion of the present invention, wherein: S1: Input the height data of the participating actors and train the motion association sub-model and motion timing prediction sub-module; S2: Rapid motion and height data collection. The sensor module collects the rapid motion data of the actors in action and dance scenes in real time, and synchronously associates the actors' height parameters to form a fused data stream. S3: Height-adaptive motion trajectory prediction. The fused data stream is input into the motion prediction module, the motion features are corrected through the motion association sub-model, and the motion temporal prediction sub-module outputs the motion trajectory prediction results for the next 1-2 frames. S4: Low-latency rendering and projection output, generating projection effect adaptation instructions based on predicted trajectory, NeRF geometric correction unit dynamically adjusts projection parameters, edge computing rendering unit performs layered rendering and pre-rendering processing, and projector array synchronously projects projection effects. S5: Compare the actual collected motion trajectory with the predicted trajectory, calculate the deviation data, and incrementally train and optimize the sub-model parameters based on the deviation data to continuously improve the prediction accuracy.
[0013] As a preferred embodiment of the immersive performance method based on 3D projection fusion of the present invention, in step S4, the edge computing rendering unit pre-renders the key frames corresponding to the predicted trajectory and stores them in a high-speed cache, and renders the non-key frames using the H.265 lightweight compression algorithm. The rendering data is transmitted through the UDP protocol and error retransmission mechanism to ensure zero-time-difference linkage between projection effects and actor movements.
[0014] The beneficial effects of this invention are: The sensing module can accurately and comprehensively collect the fast movement data and height parameters of actors in martial arts and dance scenes, providing data for subsequent movement prediction and special effects adaptation, and avoiding the problem of projection effects being out of sync with the actors' movements due to missed movement collection or insufficient data accuracy. Furthermore, the motion prediction module can establish a mapping relationship between height and the range and span of motion, and predict the motion trajectory of the next 1-2 frames, so that the projection effects can be adapted to the actors' movements in advance, thereby reducing the delay in the presentation of effects from the source and improving the synchronization of the immersive experience. It can also combine NeRF geometric correction technology and layered rendering strategy through the rendering output module, which not only ensures the precise fit between the projection effects and the stage carrier, but also improves the rendering efficiency through lightweight processing. Combined with the projector array, it can achieve seamless projection across the entire area, enhancing the integrity and realism of the effects presentation. The optimization module continuously optimizes the sub-model parameters through deviation data, enabling the system to adapt to the movement characteristics of different actors, continuously improve the accuracy of movement prediction, and ensure the reliability of the system's long-term stable operation. Attached Figure Description
[0015] Figure 1 This is a flowchart of the immersive performance system and method based on 3D projection fusion proposed in this invention; Detailed Implementation
[0016] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0017] In the description of this application, it should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. For ease of description, the dimensions of the various parts shown in the drawings are not drawn to actual scale. Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered part of the specification. In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values. It should be noted that similar reference numerals and letters in the following drawings denote similar items; therefore, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.
[0018] It should be noted that the terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and are not limited in number; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0019] It should be noted that in the description of this application, the directional terms such as "front, back, up, down, left, right", "horizontal, vertical, horizontal" and "top, bottom" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description. Unless otherwise stated, these directional terms do not indicate or imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the scope of protection of this application. The directional terms "inner" and "outer" refer to the inner and outer contours relative to the outline of each component itself.
[0020] It should be noted that, in this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0021] Example 1, referring to Figure 1An immersive performance system based on 3D projection fusion includes a perception module for collecting actors' rapid movement data and height parameters. This rapid movement data includes limb trajectories, joint angles, and movement speed data in martial arts and dance scenes. A motion prediction module communicates with the perception module and includes a motion association sub-model and a motion timing prediction sub-module. The motion association sub-model establishes a mapping relationship between the actor's height and the amplitude and span of their movements. The motion timing prediction sub-module predicts the motion trajectory of the next 1-2 frames based on preceding motion data and height parameters. Finally, a rendering output module is included. The system communicates with the motion prediction module and includes a NeRF geometric correction unit, an edge computing rendering unit, and a projector array. The NeRF geometric correction unit generates a dense 3D mesh model of the stage carrier and dynamically adjusts the projection parameters. The edge computing rendering unit adopts a layered rendering strategy of keyframe pre-rendering + non-keyframe lightweight compression. The projector array synchronously projects the rendered projection effects. The optimization module communicates with the perception module and the motion prediction module respectively. It is used to incrementally train and optimize the parameters of the motion association sub-model and the motion timing prediction sub-module based on the deviation data between the actual motion and the predicted trajectory.
[0022] This application uses a perception module to capture actors' movements in real time via a high frame rate acquisition device, focusing on collecting rapid motion data such as limb trajectories, joint angles, and movement speeds in martial arts and dance scenes. Simultaneously, it collects or records actors' height parameters. After preliminary integration of multi-dimensional data, it transmits the data to subsequent modules, building a high-quality data input foundation. After receiving the output data from the perception module, the motion prediction module uses a motion association sub-model to call a preset height-motion feature mapping relationship to correct the input motion data. Then, the motion timing prediction sub-module uses a timing prediction algorithm to predict the motion trajectory of the next 1-2 frames based on the corrected preceding motion sequence, achieving motion prediction. Simultaneously, the rendering output module... After receiving the predicted trajectory data, the NeRF geometric correction unit first generates a 3D dense mesh model of the stage carrier using 3D reconstruction technology. Based on this, it dynamically adjusts the projection parameters to match the stage shape. The edge computing rendering unit uses a layered strategy that combines keyframe pre-rendering with non-keyframe lightweight compression to process the special effects rendering task. Finally, the projector array synchronously projects the rendered special effects content onto the stage. At the same time, the optimization module collects the actual motion data of the perception module and the predicted trajectory data of the motion prediction module in real time, calculates the deviation between the two, and feeds the deviation data as an optimization signal to the motion association sub-model and the motion timing prediction sub-module. Through incremental training, the model parameters are adjusted to achieve dynamic optimization of system performance.
[0023] In the example of this application, the perception module includes an optical motion capture camera and an inertial motion capture sensor. The optical motion capture camera and the inertial motion capture sensor work together to achieve the complete acquisition of fast motion data. The perception module also includes an actor height input unit, which is used to input actor height data and associate it with the corresponding actor's motion acquisition data stream.
[0024] As a preferred example of the present invention, the optical motion capture camera achieves spatial positioning of the actor's movements through multi-view image acquisition, and the inertial motion capture sensor collects dynamic data of acceleration and angular velocity of the movements through sensors worn on key parts of the actor's limbs. The data of the two are fused and complemented in real time to make up for the acquisition defects of the single capture method in occluded and fast-moving scenarios, ensuring that no fast-moving data is missed. At the same time, the actor's height input unit obtains the actor's height data through manual input or automatic measurement, and assigns a unique identifier to the height data, which is then bound to the corresponding actor's motion acquisition data stream. This enables subsequent modules to accurately associate the height and motion data of the same actor, providing data for height-adapted motion prediction.
[0025] It is worth noting that the optical motion capture cameras are deployed in a distributed manner using eight cameras with multiple viewing angles, and the three-dimensional coordinates of the limbs are calculated using triangulation. Highly reflective markers were affixed to key skeletal joints of the actor's limbs (21 in total, such as the head, shoulders, elbows, wrists, and hips). The intrinsic and extrinsic parameters of the eight cameras were calibrated in advance: the intrinsic parameters included focal length f, pixel center coordinates (u0, v0), and distortion coefficients, and the extrinsic parameters included the rotation matrix Ri and translation vector ti (i=1,2,...,8) of each camera relative to the world coordinate system. The calibration accuracy was ≤0.01mm. Two-dimensional image coordinate extraction: Each camera synchronously acquires images containing markers. The two-dimensional pixel coordinates (ui, vi) of the markers are extracted through threshold segmentation and contour detection algorithms. Marker points that are missed by a single camera due to occlusion are temporarily marked as invalid coordinates and subsequently filled in by inertial data. For multiple sets of 3D coordinates of the same marker point calculated by 8 cameras, a weighted average method is used for optimization (the weight is inversely proportional to the distance between the camera and the marker point). Finally, the accurate 3D coordinates of the joint point are output. The calculation frequency is consistent with the camera sampling rate, and the calculation delay of the 3D coordinates of a single marker point is ≤1ms.
[0026] The inertial motion capture sensor uses a 17-axis distributed inertial measurement unit: A total of 17 inertial sensors were deployed, corresponding to 21 key joints on the human body. Specific locations are as follows: One head sample: located at the center of the forehead, to collect angular velocity and acceleration data of the head posture; Two trunk points were collected: one at the upper end of the sternum and one at the lower end of the lumbar vertebrae. Data on trunk twisting and pitching movements were collected. Six upper limbs: one each on the left and right acromion, the outer side of the left and right elbows, and the back of the left and right wrists, covering the flexion, extension, and rotation movements of the shoulder, elbow, and wrist joints; Eight exercises for the lower limbs: one each for the outer sides of the left and right hips, the middle of the left and right thighs, the outer sides of the left and right knees, and the outer sides of the left and right ankles, covering the flexion, extension, abduction, and adduction movements of the hip, knee, and ankle joints; The sensor is fixed with an elastic strap to ensure no relative displacement with the limb, and the tightness of the strap is set so as not to affect the movement of the limb. Each sensor incorporates a three-axis accelerometer, a three-axis gyroscope, a three-axis magnetometer, and a temperature compensation module. The raw sensor data is first filtered by Kalman filtering to eliminate motion noise, and then the acceleration and angular velocity data are converted into Euler angles of the joints by the attitude calculation algorithm quaternion method, outputting the joint angle change data.
[0027] In this system, the actor height input unit uses a combination of manual input and unique identifier binding. The specific operation process is as follows: The system backend has a preset input interface, which includes fields for actor information: name, actor ID, and height. The actors' net height was measured in advance using a standard height measuring instrument to ensure the accuracy of the measurement data; Operators log into the system, enter their accounts, and access the actor height management module; Enter the actor's unique actor ID and fill in the actor's name; Enter your measured net height in the height input field and click the confirm input button; Once the data is entered, the system will bind the height data to the actor's ID and simultaneously link it to the subsequent motion capture data stream.
[0028] In the example of this application, the mapping relationship of the action association sub-model satisfies the following: under the same action, the limb swing amplitude and trajectory span of the tall actor are greater than those of the short actor. The mapping relationship is generated by training a labeled dataset containing actors of different heights and different types of fast actions. The action types in the dataset cover core martial arts and dance movements such as arm swinging, kicking, rotating and jumping.
[0029] As a preferred example of the present invention, the core of the action association sub-model is a mapping relationship model trained on a labeled dataset. This dataset contains labeled data of actors of different heights performing various core martial arts and dance movements such as arm swinging, kicking, rotating, and jumping. During model training, the differences in the amplitude of limb swing and trajectory span characteristics when actors of different heights perform the same movement are statistically analyzed to construct a quantitative mapping relationship between height and action features and store it in the model. When receiving action data and height parameters, the model matches the corresponding mapping rules according to the input height parameters and performs quantitative correction on the amplitude of movement and trajectory span characteristics, so that the action data is more in line with the actual action characteristics of the actor of that height and improves the accuracy of subsequent predictions.
[0030] In the example of this application, the motion timing prediction submodule takes into account the actor's motion sequence data of the first 60 frames and height parameters, and outputs the motion trajectory prediction result for the next 16.7-33.4ms. The deviation between the predicted trajectory and the actual motion is ≤0.5cm.
[0031] As a preferred example of this invention, the action timing prediction submodule uses 60 frames as an action sequence window, extracts the actor's action data and corresponding height parameters within this window as input features, and models the temporal correlation of the action sequence through a deep learning temporal model. Based on the action timing evolution rules learned during training, combined with the action feature constraints corresponding to the height parameters, the model calculates the action trajectory coordinates for the next 16.7-33.4ms (i.e., 1-2 frames). By setting strict model training objectives, the deviation between the output predicted trajectory and the subsequently actually collected action trajectory is ensured to be controlled within 0.5cm, reserving sufficient processing time for the rendering output module and ensuring the synchronization of special effects and actions.
[0032] In the example of this application, the NeRF geometry correction unit includes a lidar and an RGB-D camera. The generated 3D dense mesh model can dynamically adapt to the subtle deformations of the stage carrier and adjust the projection matrix parameters in real time to ensure that the projection effects are accurately aligned with the stage carrier.
[0033] As a preferred example of the present invention, the NeRF geometric correction unit integrates a lidar and an RGB-D camera. The lidar is responsible for collecting the three-dimensional spatial distance information of the stage carrier, while the RGB-D camera simultaneously collects the color image and depth information of the stage. After the data from both are fused, three-dimensional reconstruction is performed using the NeRF algorithm to generate a high-precision 3D dense mesh model of the stage carrier. During system operation, the lidar and RGB-D camera continuously collect stage data and update the 3D dense mesh model in real time, dynamically capturing subtle deformations of the stage carrier. At the same time, based on the updated mesh model, the projection parameters of each projector are adjusted through a projection matrix calculation algorithm to ensure that the projected special effects image can accurately fit the real-time shape of the stage carrier and avoid projection offset.
[0034] It is worth noting that this system employs a Kalman filter fusion algorithm to achieve complementary advantages between optically captured spatial coordinate data and inertial captured kinematic data. The specific fusion process is as follows: Data time synchronization: Using the sampling frequency of the optical camera as a reference, the high-frequency data of the inertial sensor is downsampled, and the time synchronization error between the two sets of data is ensured to be ≤1ms by aligning the timestamps.
[0035] Kalman filter fusion model construction: State equation: Define the state vector of the joint as Xk=[X,Y,Z,X˙,Y˙,Z˙]T (position + velocity), construct the state transition matrix Fk based on the kinematic data (acceleration, angular velocity) of the inertial sensor, and predict the state of the joint at the next moment; Observation equation: Using the three-dimensional coordinates calculated by the optical camera as the observation value Zk, an observation matrix Hk is constructed to map the predicted state to the observation space.
[0036] Scenario-based fusion strategy: Normal observation scenario: The optical camera is unobstructed, both sets of data are valid, and Kalman filtering predicts the state and observation values through weighted fusion, outputting the optimal key point coordinates and motion parameters; Occlusion scenario: When the optical camera misses a marker point and the observation value is invalid, the system automatically switches to inertial data extrapolation mode. Based on the previously fused state data and the real-time kinematic data of the inertial sensor, the system extrapolates and completes the trajectories of the joints during the occlusion period.
[0037] Output of merged data: The fused data, after being processed by denoising, outlier removal, and coordinate normalization, is transmitted to the action prediction module in JSON format.
[0038] In the example of this application, the edge computing rendering unit includes at least 4 edge computing nodes.
[0039] As a preferred example of this invention, the edge computing rendering unit adopts a distributed architecture, consisting of a rendering cluster composed of at least four edge computing nodes. Based on the complexity of the rendering task, the system breaks down the overall rendering task into multiple sub-tasks, which are then distributed to different edge computing nodes for parallel processing using a load balancing algorithm. This significantly improves rendering efficiency compared to single-node rendering, meeting the requirements for low-latency rendering. Simultaneously, the system incorporates a built-in node fault detection and redundancy backup mechanism. When a single edge computing node fails, its rendering sub-tasks are quickly switched to other working nodes, preventing rendering interruptions and ensuring the continuity of projection effects output.
[0040] In the example of this application, the number of projectors in the projector array is N≥6, the overlap rate of the projection range of adjacent projectors is 15%-20%, and all projectors achieve synchronous projection of the image through a synchronization control module to ensure seamless presentation of projection effects across the entire stage.
[0041] As a preferred example of the present invention, the projector array is configured with no fewer than 6 projectors according to the stage size and projection coverage requirements. The installation position and projection angle of each projector are determined through pre-calibration to ensure that there is a 15%-20% overlap in the projection range of adjacent projectors. The synchronization control module uses a hardware synchronization or network synchronization protocol to send a unified synchronization trigger signal to all projectors, controlling each projector to start projecting at the same time. For the overlapping area, color correction and brightness blending algorithms are used for transition processing to eliminate splicing marks, ultimately achieving seamless projection coverage of the entire stage area and ensuring that audiences in different areas receive a consistent immersive experience.
[0042] It is worth noting that the fused motion data, in conjunction with a projector array, achieves an immersive 3D projection effect. The specific process is as follows: I. Motion prediction and special effects command generation: The motion prediction module receives the fused data, corrects the motion amplitude and trajectory span to match the actor's height through the motion association sub-model, and then predicts the motion trajectory of the next 1-2 frames through the improved LSTM network. 1. Action-related sub-model: Core objective: To adjust the range and trajectory of movements based on the actor's height; Key basis: Based on the annotation data of the core martial arts and dance movements (arm swings, kicks, etc.) of 180 actors of different heights, a quantitative mapping formula was established; 2. Action Timing Prediction Submodule: Core objective: To predict motion trajectories in the next 1-2 frames based on a corrected 60-frame motion sequence; Core model: Improved LSTM network, input includes 60 frames of joint coordinates + height parameters, output 1-2 frames of joint 3D coordinates; Based on the predicted trajectory, the system generates corresponding 3D special effects instructions (such as weapon lighting and shadows in martial arts scenes and body motion blur in dance scenes). The instructions include the spatial position, shape, and color parameters of the special effects.
[0043] II. NeRF Geometric Correction: The NeRF geometry correction unit generates a dense 3D mesh model of the stage using LiDAR and an RGB-D camera. Based on the predicted actor movement trajectory, the relative position of the special effects in the stage space is calculated, and the projection matrix parameters of each projector, including projection angle, scaling ratio, and offset, are dynamically adjusted to ensure that the special effects are precisely matched with the actor's movements and the stage shape.
[0044] III. Edge computing and layered rendering: The edge computing rendering unit distinguishes between key frames and non-key frames based on special effects instructions: when the movement amplitude of adjacent frames is ≥5° or the movement speed is ≥1m / s, it is determined to be a key frame. Key frames are pre-rendered and stored in SSD high-speed cache; non-key frames are rendered using H.265 lightweight compression algorithm. The rendering cluster processes rendering tasks in parallel using a load balancing algorithm.
[0045] IV. Synchronous projection using a projector array: The synchronization control module uses the IEEE1588PTP protocol to send a unified synchronization trigger signal to ≥6 projectors, ensuring that all projectors project images at the same time. The overlap rate of the projection range of adjacent projectors in the projector array is 15%-20%. The Delta-E calibration method is used for color correction and the Gaussian attenuation function is used for brightness fusion in the overlapping area to eliminate splicing marks. During the projection process, the system receives deviation data from the optimization module in real time to predict the Euclidean distance between the predicted trajectory and the actual trajectory, and fine-tunes the projection parameters to ensure that the 3D effects and the actors' movements are linked in real time, ultimately achieving spatial matching between the actors' movements and the projection effects, presenting an immersive 3D projection effect.
[0046] In the example of this application, the optimization module optimizes every 3 minutes. It obtains the deviation data by calculating the Euclidean distance between the predicted trajectory and the actual action trajectory. When the mean deviation data is >0.5cm, it starts the incremental training of the sub-model parameters. The training iterations are 100-200 times until the mean deviation data is ≤0.5cm.
[0047] As a preferred example of the present invention, the optimization module sets a fixed optimization cycle of 3 minutes. Within each cycle, the distance between the predicted trajectory and the corresponding key points in the actual action trajectory is first calculated using the Euclidean distance formula. The average distance of all key points is then used as the deviation data. The average deviation data is compared with a preset threshold (0.5cm). If the average is greater than the threshold, the incremental training process is initiated: the deviation data is used as the loss function input, and the parameters of the action association sub-model and the action timing prediction sub-module are fine-tuned using the gradient descent algorithm. After iterative training for 100-200 times, the average deviation data is recalculated. If the average is ≤0.5cm, training is stopped, and one optimization is completed. If the target is not met, iteration continues until the accuracy requirement is met, thereby achieving accurate optimization of model parameters and efficient utilization of resources.
[0048] Example 2, refer to Figure 1 This is the second embodiment of the present invention. Unlike the previous embodiment, this embodiment provides an immersive performance method based on 3D projection fusion, including the immersive performance system based on 3D projection fusion described in the previous embodiments. S1: Input the height data of the participating actors and train the motion association sub-model and motion timing prediction sub-module; S2: Rapid motion and height data collection. The sensor module collects the rapid motion data of the actors in action and dance scenes in real time, and synchronously associates the actors' height parameters to form a fused data stream. S3: Height-adaptive motion trajectory prediction. The fused data stream is input into the motion prediction module, the motion features are corrected through the motion association sub-model, and the motion temporal prediction sub-module outputs the motion trajectory prediction results for the next 1-2 frames. S4: Low-latency rendering and projection output, generating projection effect adaptation instructions based on predicted trajectory, NeRF geometric correction unit dynamically adjusts projection parameters, edge computing rendering unit performs layered rendering and pre-rendering processing, and projector array synchronously projects projection effects. S5: Compare the actual collected motion trajectory with the predicted trajectory, calculate the deviation data, and incrementally train and optimize the sub-model parameters based on the deviation data to continuously improve the prediction accuracy.
[0049] Specifically, in step S4, the edge computing rendering unit pre-renders the key frames corresponding to the predicted trajectory and stores them in a high-speed cache. Non-key frames are rendered using the H.265 lightweight compression algorithm. The rendering data is transmitted through the UDP protocol and error retransmission mechanism to ensure that the projection effects and the actors' movements are linked in real time.
[0050] The embodiments of this application have been described above with reference to the accompanying drawings. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. This application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An immersive performance system based on 3D projection fusion, characterized in that, include: The sensing module is used to collect actors' fast movement data and height parameters. The fast movement data includes limb trajectory, joint angle and movement speed data in martial arts and dance scenes. The motion prediction module is communicatively connected to the perception module and includes a motion association sub-model and a motion timing prediction sub-module. The motion association sub-model establishes a mapping relationship between the actor's height and the amplitude and trajectory span of the motion. The motion timing prediction sub-module predicts the motion trajectory of the next 1-2 frames based on the preceding motion data and height parameters. The rendering output module is communicatively connected to the motion prediction module. It includes a NeRF geometric correction unit, an edge computing rendering unit, and a projector array. The NeRF geometric correction unit generates a 3D dense mesh model of the stage carrier and dynamically adjusts the projection parameters. The edge computing rendering unit adopts a layered rendering strategy of keyframe pre-rendering + non-keyframe lightweight compression. The projector array synchronously projects the rendered projection effects content. An optimization module is communicatively connected to the perception module and the action prediction module, respectively, and is used to incrementally train and optimize the parameters of the action association sub-model and the action temporal prediction sub-module based on the deviation data between the actual action and the predicted trajectory.
2. The immersive performance system based on 3D projection fusion according to claim 1, characterized in that, The sensing module includes an optical motion capture camera and an inertial motion capture sensor. The optical motion capture camera and the inertial motion capture sensor work together to achieve the complete acquisition of fast motion data. The sensing module also includes an actor height input unit, which is used to input actor height data and associate it with the corresponding actor's motion acquisition data stream.
3. The immersive performance system based on 3D projection fusion according to claim 2, characterized in that, The mapping relationship of the action association sub-model satisfies the following: under the same action, the limb swing amplitude and trajectory span of tall actors are greater than those of short actors. The mapping relationship is generated by training a labeled dataset containing actors of different heights and different types of fast actions. The action types in the dataset cover core martial arts and dance movements such as arm swinging, kicking, spinning and jumping.
4. The immersive performance system based on 3D projection fusion according to claim 3, characterized in that, The motion timing prediction submodule takes into account the actor's motion sequence data for the first 60 frames and height parameters, and outputs the motion trajectory prediction result for the next 16.7-33.4ms. The deviation between the predicted trajectory and the actual motion is ≤0.5cm.
5. The immersive performance system based on 3D projection fusion according to claim 4, characterized in that, The NeRF geometric correction unit includes a lidar and an RGB-D camera. The generated 3D dense mesh model can dynamically adapt to the subtle deformations of the stage carrier and adjust the projection matrix parameters in real time to ensure that the projection effects are accurately matched with the stage carrier.
6. The immersive performance system based on 3D projection fusion according to claim 5, characterized in that, The edge computing rendering unit includes at least four edge computing nodes.
7. The immersive performance system based on 3D projection fusion according to claim 6, characterized in that, The projector array contains N≥6 projectors, with an overlap rate of 15%-20% between the projection ranges of adjacent projectors. All projectors are synchronized through a synchronization control module to ensure seamless presentation of projection effects across the entire stage.
8. The immersive performance system based on 3D projection fusion according to claim 7, characterized in that, The optimization module optimizes every 3 minutes. It calculates the Euclidean distance between the predicted trajectory and the actual action trajectory to obtain deviation data. When the mean deviation data is >0.5cm, it initiates incremental training of the sub-model parameters. The training iterations are 100-200 times until the mean deviation data is ≤0.5cm.
9. An immersive performance method based on 3D projection fusion, characterized in that, An immersive performance system based on 3D projection fusion, as described in any one of claims 1-8, comprises the following steps: S1: Input the height data of the participating actors and train the motion association sub-model and motion timing prediction sub-module; S2: Rapid motion and height data collection. The sensor module collects the rapid motion data of the actors in action and dance scenes in real time, and synchronously associates the actors' height parameters to form a fused data stream. S3: Height-adaptive motion trajectory prediction. The fused data stream is input into the motion prediction module, the motion features are corrected through the motion association sub-model, and the motion temporal prediction sub-module outputs the motion trajectory prediction results for the next 1-2 frames. S4: Low-latency rendering and projection output, generating projection effect adaptation instructions based on predicted trajectory, NeRF geometric correction unit dynamically adjusts projection parameters, edge computing rendering unit performs layered rendering and pre-rendering processing, and projector array synchronously projects projection effects. S5: Compare the actual collected motion trajectory with the predicted trajectory, calculate the deviation data, and incrementally train and optimize the sub-model parameters based on the deviation data to continuously improve the prediction accuracy.
10. The immersive performance method based on 3D projection fusion according to claim 9, characterized in that, In step S4, the edge computing rendering unit pre-renders the key frames corresponding to the predicted trajectory and stores them in a high-speed cache. Non-key frames are rendered using the H.265 lightweight compression algorithm. The rendering data is transmitted through the UDP protocol and error retransmission mechanism to ensure that the projection effects and the actors' movements are linked in real time.