Image data processing method and system for dance editing

Through multi-view optical equipment and light field phase difference processing, combined with frequency domain suppression of Doppler shift characteristics, a high-precision three-dimensional model of dance movements is generated, which solves the problems of noise accumulation and difficulty in distinguishing dynamic noise in existing technologies, and realizes high-precision and real-time collaborative choreography of dance directors.

CN120766348AInactive Publication Date: 2025-10-10YANGTZE NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510870501.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing technologies, optical systems are susceptible to occlusion by marker points or interference from ambient light, and multi-perspective reconstruction algorithms lack accuracy in capturing fast-moving or low-texture areas, resulting in noise accumulation in dance movements and the inability to distinguish between effective frequency bands and dynamic noise, affecting the accuracy and real-time performance of choreography.

Method used

The dynamic image data of dancers is collected synchronously through multi-view optical equipment, and interference processing is performed using light field phase difference adjustment to enhance bone edge details and suppress noise. Doppler frequency shift characteristics are combined for frequency domain suppression. The de-noised spatiotemporal correlation data and bone positioning information are integrated to generate a high-precision three-dimensional model of dance movements and display the collaborative choreography effect of multiple dancers' movements in real time.

Benefits of technology

It achieves high-precision dance motion capture, eliminates data loss caused by perspective occlusion and lighting interference, improves the precision and real-time performance of dance choreography, and supports the accuracy and visualization of collaborative choreography for multiple dancers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766348A_ABST
    Figure CN120766348A_ABST
Patent Text Reader

Abstract

The invention provides an image data processing method and system for dance editing. The method comprises the steps of performing light field interference processing on multi-view-angle dynamic image data by adjusting light field phase differences of different view angles, generating enhanced motion capture data, performing frequency domain suppression processing on dynamic range noise based on Doppler frequency shift characteristics generated by movement of a dancer, and generating multi-view-angle space-time associated data after noise reduction; fusing the noise-reduced multi-view space-time associated data with the skeleton positioning information to generate a dance movement three-dimensional model; and outputting the dance movement three-dimensional model to a dance editing cooperation terminal, so that the dance editing cooperation terminal displays the cooperation editing effect of the multi-dancer movement in real time according to the joint movement track of the dance movement three-dimensional model. According to the method, the three-dimensional action model is constructed by fusing the noise-reduced multi-view space-time associated data and the skeleton positioning information, millimeter-level precision capture of the joint movement track of the dancer can be realized, and the accuracy of the collaborative arrangement effect is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image data processing, and in particular to an image data processing method and system for dance choreography. Background Art

[0002] Dance choreography involves designing and arranging dance movements to guide dancers in their artistic performances. With the advancement of dance art, choreography has placed higher demands on the precision of movement design, the real-time coordination of multiple dancers, and the authenticity of movement reproduction. Traditional methods that rely on manual observation and 2D video analysis are no longer sufficient for complex choreography. Intelligent motion capture and 3D modeling technologies are urgently needed to achieve high-precision movement recording, real-time collaborative choreography, and virtual rehearsal.

[0003] To meet these requirements, existing technologies primarily utilize optical motion capture systems or multi-view camera arrays to capture dance movements. For example, optical systems based on infrared markers generate skeletal trajectory data by tracking reflective markers, while computer vision methods reconstruct 3D motion sequences from multi-view images. Furthermore, frequency-domain filtering algorithms are used to suppress dynamic noise, and inverse kinematics algorithms are used to generate ergonomic 3D skeletal models. These technologies enable, to a certain extent, the digital recording and basic analysis of dance movements.

[0004] However, existing technologies still have the following defects: the optical system is easily blocked by markers or interfered by ambient light, resulting in data loss or noise accumulation; the multi-perspective reconstruction algorithm has insufficient accuracy in capturing fast-moving or low-texture areas (such as limb edges); the frequency domain denoising method with a fixed threshold has difficulty distinguishing between the effective frequency bands of dance movements (such as periodic swings) and dynamic noise. In addition, traditional three-dimensional modeling may generate non-physiological motion trajectories, affecting the authenticity and collaborative efficiency of the choreography. Therefore, there is an urgent need for a solution that can integrate light field interferometry enhancement, adaptive frequency domain denoising and multi-perspective spatiotemporal optimization to improve the accuracy and real-time performance of dance choreography. Summary of the Invention

[0005] The present application provides an image data processing method and system for dance choreography, which is used to solve the problems in the prior art such as severe noise accumulation, inability to effectively distinguish the effective frequency band of dance movements from dynamic noise, and low reliability of model reconstruction, which lead to low accuracy of collaborative choreography effects.

[0006] In a first aspect, the present application provides an image data processing method for dance choreography, comprising:

[0007] Acquiring multi-perspective dynamic image data of a dancer captured by optical motion capture devices distributed at different spatial positions, wherein the multi-perspective dynamic image data includes skeleton positioning information;

[0008] By adjusting the phase difference of light fields at different viewing angles, light field interference processing is performed on the multi-view dynamic image data to generate enhanced motion capture data;

[0009] Based on the Doppler frequency shift characteristics generated by the dancer's movements in the enhanced motion capture data, dynamic range noise is suppressed in the frequency domain to generate noise-reduced multi-view spatiotemporal correlation data;

[0010] fusing the noise-reduced multi-view spatiotemporal correlation data with the skeleton positioning information to generate a three-dimensional dance movement model;

[0011] The dance movement three-dimensional model is output to the choreography collaboration terminal, so that the choreography collaboration terminal displays the collaborative choreography effect of multiple dancers' movements in real time according to the joint movement trajectory of the dance movement three-dimensional model.

[0012] Optionally, the fusing the noise-reduced multi-view spatiotemporal correlation data with the skeleton positioning information to generate a three-dimensional dance movement model includes:

[0013] Extracting joint topological relationships from the multi-view dynamic image data, and constructing a hierarchical kinematic constraint model of skeletal nodes based on the skeletal positioning information and the joint topological relationships;

[0014] Mapping the spatiotemporal coordinates of each view in the denoised multi-view spatiotemporal correlation data to a corresponding skeleton node, and extracting the local spatial features of each skeleton node under the corresponding view;

[0015] Establishing a dynamic motion equation of the skeletal node according to the geometric constraint relationship between the local spatial features and the skeletal nodes in the hierarchical kinematic constraint model;

[0016] Dynamically adjust the weight coefficient of each perspective data in the dynamic motion equation according to the frequency domain feature weights of different perspectives to generate a global optimized three-dimensional trajectory of the skeleton node;

[0017] Calculating the rigid body motion parameters of the skeletal chain based on the global optimized three-dimensional trajectory and the joint rotational degrees of freedom defined in the hierarchical kinematic constraint model;

[0018] A three-dimensional motion sequence of a skeleton chain is generated according to the rigid body motion parameters, and the three-dimensional motion sequence constitutes a three-dimensional dance movement model.

[0019] Optionally, establishing the dynamic motion equation of the skeletal nodes according to the geometric constraint relationship between the local spatial features and the skeletal nodes in the hierarchical kinematic constraint model includes:

[0020] Determining the relative translation of the parent node and the child node in three-dimensional space based on the parent-child hierarchical relationship between the skeletal nodes in the hierarchical kinematic constraint model;

[0021] Convert the displacement vectors of the skeleton nodes in the local spatial features at each viewing angle into the observation residuals of the relative translation amount to construct a multi-view observation constraint item for each skeleton node;

[0022] Constructing rigid geometric constraint terms between bone nodes based on fixed geometric lengths between bone nodes in the hierarchical kinematic constraint model;

[0023] The multi-view observation constraint term and the rigid geometric constraint term are combined into a dynamic motion equation of the skeleton node.

[0024] Optionally, the step of performing light field interference processing on the multi-view dynamic image data by adjusting the light field phase difference of different view angles to generate enhanced motion capture data includes:

[0025] Based on the spatial distribution position of the optical motion capture device, calculating the light field propagation path difference between adjacent perspectives to generate an initial light field phase difference;

[0026] Identifying the movement speed of the skeleton nodes according to the temporal motion trajectory of the skeleton nodes in the multi-view dynamic image data, and generating a dynamic light field phase difference compensation amount according to the initial light field phase difference and a mapping relationship between the movement speed and the phase difference compensation amount;

[0027] Based on the dynamic light field phase difference compensation amount, adjusting the initial light field phase difference frame by frame to achieve a target light field phase difference;

[0028] Based on the target light field phase difference, the multi-view dynamic image data is superimposed in the spatial domain to obtain a superposition result, and the intensity distribution of the interference area is calculated according to the superposition result to extract the interference feature of the bone edge;

[0029] The interference feature is fused with the skeleton positioning information to generate enhanced motion capture data.

[0030] Optionally, the identifying the movement speed of the skeleton nodes according to the temporal motion trajectory of the skeleton nodes in the multi-view dynamic image data, and generating the dynamic light field phase difference compensation amount according to the initial light field phase difference and the mapping relationship between the movement speed and the light field phase difference compensation amount includes:

[0031] Based on the temporal motion trajectory of the skeleton node, the displacement change of the skeleton node between adjacent frames is calculated to generate the instantaneous motion speed parameter of the skeleton node;

[0032] According to the instantaneous motion speed parameter and the motion direction of the limb where the skeletal node is located, a nonlinear mapping relationship between the motion speed and the light field phase difference compensation amount is established to generate a phase difference adjustment function;

[0033] Based on the phase difference adjustment function, the instantaneous motion speed parameter is converted into a time-series continuous dynamic phase difference compensation amount sequence;

[0034] According to the motion acceleration of the skeleton nodes between adjacent frames, the dynamic phase difference compensation amount sequence is subjected to temporal smoothing processing to generate the dynamic light field phase difference compensation amount.

[0035] Optionally, the performing frequency domain suppression processing on dynamic range noise based on Doppler frequency shift characteristics generated by dancer's movement in the enhanced motion capture data to generate noise-reduced multi-view spatiotemporal correlation data includes:

[0036] Performing frequency domain decomposition on the temporal motion trajectory of the skeletal nodes in the enhanced motion capture data to generate Doppler frequency shift features;

[0037] generating a dynamic noise spectrum model based on the periodic distribution law of the Doppler frequency shift characteristics in the frequency domain;

[0038] constructing a frequency domain suppression mask according to a boundary threshold in the dynamic noise spectrum model;

[0039] The frequency domain suppression mask is applied to the frequency domain components of the enhanced motion capture data to reconstruct the time domain signal and generate multi-view spatiotemporal correlation data after noise reduction.

[0040] Optionally, generating a dynamic noise spectrum model based on the periodic distribution law of the Doppler frequency shift feature in the frequency domain includes:

[0041] Based on the periodic distribution law of the Doppler frequency shift characteristics in the frequency domain, determining the effective signal frequency band composed of the main frequency and harmonic components of the dance movement;

[0042] According to the discrete degree of frequency domain energy distribution, the noise energy candidate area is extracted from the non-effective signal frequency band;

[0043] Dynamically adjusting the boundary threshold between the effective signal frequency band and the noise energy candidate area according to the temporal variation of the movement speed of the skeleton node;

[0044] The energy distribution of the noise energy candidate area is associated with the boundary threshold to generate a dynamic noise spectrum model.

[0045] In a second aspect, the present application provides an image data processing system for dance choreography, comprising:

[0046] an acquisition module, configured to acquire multi-perspective dynamic image data of a dancer collected by optical motion capture devices distributed at different spatial positions, wherein the multi-perspective dynamic image data includes skeleton positioning information;

[0047] an adjustment module, configured to perform light field interference processing on the multi-view dynamic image data by adjusting the phase difference of the light fields at different viewpoints to generate enhanced motion capture data;

[0048] A generation module, configured to perform frequency domain suppression processing on dynamic range noise based on Doppler frequency shift characteristics generated by dancer movements in the enhanced motion capture data, and generate noise-reduced multi-view spatiotemporal correlation data;

[0049] A fusion module, configured to fuse the noise-reduced multi-view spatiotemporal correlation data with the skeleton positioning information to generate a three-dimensional dance movement model;

[0050] The output module is used to output the dance movement three-dimensional model to the choreography collaboration terminal, so that the choreography collaboration terminal can display the collaborative choreography effect of multiple dancers' movements in real time according to the joint movement trajectory of the dance movement three-dimensional model.

[0051] In a third aspect, the present application provides a computing device comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement an image data processing method for dance choreography as described in any one of the first aspects.

[0052] In a fourth aspect, the present application provides a computer storage medium storing a computer program, which, when executed by a computer, implements an image data processing method for dance choreography as described in any one of the first aspects.

[0053] In the present application, a method for processing image data for dance choreography is provided, the method comprising: obtaining multi-perspective dynamic image data of a dancer collected by optical motion capture devices distributed at different spatial positions, wherein the multi-perspective dynamic image data includes skeletal positioning information; performing light field interference processing on the multi-perspective dynamic image data by adjusting the light field phase difference of different perspectives to generate enhanced motion capture data; performing frequency domain suppression processing on dynamic range noise based on the Doppler frequency shift characteristics generated by the dancer's movement in the enhanced motion capture data to generate de-noised multi-perspective spatiotemporal correlation data; fusing the de-noised multi-perspective spatiotemporal correlation data with the skeletal positioning information to generate a three-dimensional model of dance movements; and outputting the three-dimensional model of dance movements to a choreography collaboration terminal so that the choreography collaboration terminal can display the collaborative choreography effect of multiple dancers' movements in real time according to the joint motion trajectory of the three-dimensional model of dance movements.

[0054] This application uses optical motion capture devices distributed in different spatial locations to synchronously capture multi-perspective dynamic image data of dancers, providing skeletal positioning information covering all angles, eliminating the blind spots of single-perspective observation, and providing a high-precision, high-integrity raw data foundation for subsequent processing. Multi-perspective interference superposition based on light field phase difference adjustment enhances the high-frequency detail resolution of skeletal edges and micro-movements, suppresses local data loss caused by perspective occlusion or ambient light interference, and generates enhanced data containing submillimeter motion features. Doppler shift features are used to identify and filter dynamic range noise, retain the frequency domain energy distribution of effective motion signals, improve the spatiotemporal consistency of multi-perspective data, and provide low-noise input for three-dimensional reconstruction. Combining the de-noised spatiotemporal correlation data with skeletal positioning information, through inverse kinematic constraints and multi-perspective registration algorithms, a three-dimensional dance movement model that conforms to biomechanical laws and can reflect the skeletal motion chain is constructed, accurately restoring the joint motion trajectory. The joint motion trajectory of the three-dimensional model is transmitted in real time to the choreography terminal. Then, through spatiotemporal trajectory overlap analysis methods and collision detection algorithms, the coordinated phase differences and conflict areas of multiple dancers' movements can be dynamically displayed, supporting real-time adjustment of formation and rhythm.

[0055] Furthermore, joint topological relationships are extracted from multi-view dynamic image data, and a hierarchical kinematic constraint model of skeletal nodes is constructed based on skeletal positioning information. The spatiotemporal coordinates of each view in the denoised spatiotemporal correlation data are mapped to skeletal nodes to extract local spatial features. Multi-view observation constraint terms are constructed based on the relative translation of parent-child nodes and combined with rigid geometric constraints to form a dynamic motion equation. The data weight coefficients in the equation are dynamically adjusted according to the frequency domain feature weights of different viewpoints to optimize the generation of a global three-dimensional trajectory. The rigid body motion parameters of the skeletal chain are calculated based on the rotational degrees of freedom of the joints, ultimately generating a dance movement model that conforms to biomechanical laws. Through the joint optimization of multi-view observation residual constraints and rigid bone length constraints, the problem of local data loss and multi-view motion trajectory conflicts in occluded scenes is resolved, ensuring that the three-dimensional motion sequence meets both data fitting accuracy and biophysical rationality. Combined with the adaptive adjustment of frequency domain weights, the interference of low-quality viewpoint data on trajectory reconstruction can be suppressed, improving the model reconstruction stability and motion smoothness of complex dance movements, providing a highly accurate and physically verifiable movement data foundation for multi-dancer collaborative choreography.

[0056] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings described below are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of the present application.

[0058] Figure 1 A flow chart of an image data processing method for dance choreography provided by an embodiment of the present application;

[0059] Figure 2 A structural schematic diagram of an image data processing system for dance choreography provided by an embodiment of the present application;

[0060] Figure 3 A structural schematic diagram of a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0061] In order to make the person skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.

[0062] In some of the processes described in the specification and claims of the present application and the above drawings, a plurality of operations appear in a specific order, but it should be clearly understood that these operations can be executed or performed in parallel or in a different order from that in which they appear in the text. The serial numbers of the operations, such as 11, 12, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and the operations can be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in the text are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence. Also, "first" and "second" are not of different types.

[0063] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of the present application.

[0064] In order to solve the problems of low accuracy of collaborative choreography caused by serious noise accumulation, inability to effectively distinguish the effective frequency band of dance movements from dynamic noise, and low reliability of model reconstruction in the prior art, an embodiment of the present application provides an image data processing method for dance choreography, which adopts the following ideas: synchronously collect multi-perspective dynamic image data of dancers through multi-perspective optical equipment, and use the light field phase difference to enhance the bone edge details and suppress occlusion noise; identify the frequency domain difference between motion signals and dynamic noise based on Doppler frequency shift characteristics, and realize noise reduction of multi-perspective spatiotemporal data through adaptive filtering; combine bone positioning information with the multi-perspective spatiotemporal correlation data after noise reduction, and reconstruct a high-precision three-dimensional skeleton model through the fusion of inverse kinematic constraints and multi-perspective registration algorithms; finally, map the model joint motion trajectory to the choreography terminal in real time, realize dynamic visualization of collaborative choreography of multiple dancers' movements, and form a full-process closed loop from data acquisition, interference noise reduction, three-dimensional modeling to real-time collaboration, which can take into account motion capture accuracy, noise robustness and choreography efficiency at the same time.

[0065] Figure 1 This is a flowchart of an image data processing method for dance choreography provided in an embodiment of the present application, such as Figure 1 As shown, the method includes:

[0066] S11. Acquire multi-perspective dynamic image data of a dancer collected by optical motion capture devices distributed at different spatial positions, where the multi-perspective dynamic image data includes skeleton positioning information.

[0067] Optical motion capture equipment, such as cameras, uses infrared light sources and high-speed camera arrays to capture reflected signals from markers, thereby tracking the spatial position of moving objects. Multi-view dynamic image data refers to a sequence of images captured continuously from multiple angles, including time, the spatial coordinates of each dancer, and skeletal positioning information. Skeletal positioning information refers to biomechanical parameters such as joint position and rotation angle obtained by inverting the coordinates of the markers.

[0068] In this embodiment, optical motion capture devices located at different spatial locations first collect multi-perspective dynamic image data of a dancer. The optical motion capture devices capture reflected light signals from markers or skeletal nodes attached to key body parts of the dancer, generating dynamic image data containing skeletal positioning information. This data is synchronized and spatially aligned across the multi-perspective images, providing the foundational input for subsequent processing.

[0069] S12. Performing light field interference processing on the multi-view dynamic image data by adjusting the light field phase difference at different viewpoints to generate enhanced motion capture data.

[0070] Light field phase difference refers to the phase offset between light waves at different viewpoints, and is used to control the interference enhancement effect. Light field interferometry processing can refer to phase enhancement processing of multi-view images, which can improve image resolution. Enhanced motion capture data refers to high-resolution motion data processed by light field interferometry, including millimeter-level micro-movement characteristics of dancers.

[0071] In the embodiments of this application, light field interferometry processing is performed on the acquired multi-view dynamic image data by adjusting the phase difference of the light fields at different viewpoints. Specifically, by calculating the coherent superposition of the light fields at different viewpoints, the motion detail and spatial resolution in the motion capture data are enhanced, noise caused by view occlusion or uneven lighting is eliminated, and ultimately enhanced motion capture data containing high-precision spatiotemporal information is generated.

[0072] S13. Based on the Doppler frequency shift characteristics generated by the dancer's movements in the enhanced motion capture data, the dynamic range noise is suppressed in the frequency domain to generate denoised multi-view spatiotemporal correlation data.

[0073] The Doppler shift feature refers to the physical quantity of the change in the frequency of reflected light caused by the dancer's movement. Multi-view spatiotemporal correlation data refers to a motion trajectory dataset that retains temporal synchronization and spatial consistency after noise reduction.

[0074] In this embodiment, the Doppler shift characteristics caused by the dancer's movements are determined from enhanced motion capture data. A frequency-domain filtering algorithm is used to suppress dynamic range noise, retaining the spatiotemporal correlation information of the effective frequency band. Ultimately, de-noised multi-view spatiotemporal correlation data is generated, ensuring motion consistency across different perspectives.

[0075] S14. Fuse the denoised multi-view spatiotemporal correlation data with the skeleton positioning information to generate a three-dimensional dance movement model.

[0076] Among them, the three-dimensional model of dance movements refers to a digital representation of human movements that integrates skeletal constraints and kinematic algorithms.

[0077] In an embodiment of the present application, the generated denoised multi-view spatiotemporal correlation data is fused with the skeleton positioning information. After the fusion is completed, the inverse kinematics algorithm and point cloud registration technology are used, combined with the skeleton constraint model, to reconstruct the dancer's three-dimensional skeleton motion trajectory, and finally generate a highly accurate three-dimensional model of dance movements, accurately reflecting joint rotation, displacement and posture changes.

[0078] S15. Outputting the dance movement three-dimensional model to the choreography collaboration terminal, so that the choreography collaboration terminal displays the collaborative choreography effect of multiple dancers' movements in real time according to the joint motion trajectory of the dance movement three-dimensional model.

[0079] The choreography collaboration terminal refers to an interactive platform that supports real-time multi-user editing and rendering of dance choreography. The joint motion trajectory refers to the displacement and rotation path of the joint in three-dimensional space that changes over time.

[0080] In this embodiment of the application, the generated 3D dance movement model is output to a choreography collaboration terminal via a communication interface. Based on the joint motion trajectory data of the model, the terminal drives the movement of the virtual character through a real-time rendering engine and uses a collaborative optimization algorithm to dynamically display the collaborative choreography effect in a multi-dancer scene.

[0081] The following is a specific example: In a dance studio, 12 optical motion capture devices are deployed in a circular arrangement to capture multi-perspective images of dancers' jumping movements in real time; a phase modulator is used to adjust the phase difference of the light fields of adjacent cameras to generate enhanced data and detect the 10Hz Doppler frequency shift caused by arm swinging; wavelet threshold filtering is used to eliminate ambient light noise, and the denoised spatiotemporal data is fused with the coordinates of the skeleton nodes, and a three-dimensional model is reconstructed through inverse kinematics; the model is transmitted to the choreography terminal via the 5G network, rendering the coordinated kicking movements of the three virtual dancers in real time, and the trajectory overlap analysis indicates movement synchronization errors.

[0082] By executing S11 to S15, the embodiment of the present application enhances the resolution of multi-perspective dynamic image data through light field interference processing, suppresses dynamic noise in combination with Doppler frequency shift characteristics, and fuses skeletal positioning information with spatiotemporal correlation data to construct a high-precision three-dimensional model of dance movements, ultimately achieving millimeter-level motion capture and real-time visualization of collaborative choreography of multiple dancers, thereby improving dance creation efficiency and movement restoration.

[0083] In a possible embodiment, S14, fusing the noise-reduced multi-view spatiotemporal correlation data with the skeleton positioning information to generate a three-dimensional dance movement model, includes:

[0084] Step 141: Extract joint topological relationships from multi-view dynamic image data, and construct a hierarchical kinematic constraint model of the skeleton nodes based on the skeleton positioning information and the joint topological relationships.

[0085] The joint topology refers to the connections and hierarchical structure between joints, which can be defined using a joint point connection matrix. A hierarchical kinematic constraint model is a mathematical model that defines the parent-child relationships and range of motion of skeletal nodes. Parent-child node relationships can be stored in a tree structure.

[0086] Specifically, the hierarchical structure between bone nodes is defined based on human anatomy, the graph theory algorithm is used to describe the joint topology, and the physical constraint parameters are set to form a hierarchical kinematic constraint model to provide biomechanical constraints for subsequent motion equations.

[0087] Step 142: Map the spatiotemporal coordinates of each view in the denoised multi-view spatiotemporal correlation data to the corresponding skeleton node, and extract the local spatial features of each skeleton node at the corresponding view.

[0088] Spatiotemporal coordinates refer to data structures containing timestamps and three-dimensional spatial positions (x, y, z). Skeletal nodes are abstract representations of key points in the human skeletal system within a digital model. Local spatial features are used to reflect the motion properties of skeletal nodes within a specific viewpoint coordinate system, including their displacement vectors within that viewpoint coordinate system.

[0089] In an embodiment of the present application, the local spatial features of each skeletal node under the corresponding perspective can be extracted through principal component analysis. It should be noted that this embodiment does not specifically limit the expression form of the local spatial features.

[0090] Step 143: Establish the dynamic motion equations of the skeletal nodes based on the local spatial features and the geometric constraint relationships between the skeletal nodes in the hierarchical kinematic constraint model.

[0091] Optionally, the process of establishing the dynamic motion equation is as follows: ~ As shown in step a4, it will not be repeated here.

[0092] Alternatively, the dynamic motion equation is composed of the equations of the physical dynamics layer and the observation constraints of the multi-view data fusion layer. The equations of the physical dynamics layer are differential equations used to describe the motion laws of the skeletal nodes. For example, the equations of the physical dynamics layer can specifically adopt the Newton-Euler equation, which has the formula: τ = I·α + C·ω 2 , where τ is the net joint torque applied to the skeletal node during motion, I is the inertial property of the skeletal node mass distribution, α is the angular acceleration of the skeletal node around the rotation axis, C is the Coriolis force term coefficient, and ω is the instantaneous angular velocity of the skeletal node around the rotation axis. The multi-view data fusion layer converts the multi-view observation data into input constraints for the physical dynamics layer, and the weight coefficient takes effect at this layer. It should be noted that this embodiment does not specifically limit the specific expression of this equation.

[0093] Step 144: Dynamically adjust the weight coefficient of each perspective data in the dynamic motion equation according to the frequency domain feature weights of different perspectives to generate a global optimized three-dimensional trajectory of the skeleton node.

[0094] The frequency domain feature weight refers to the confidence coefficient assigned to data from different perspectives based on the energy distribution of the frequency band. The globally optimized 3D trajectory refers to the optimal motion path of the skeleton nodes after integrating the weights of multiple perspectives.

[0095] In the embodiment of the present application, there can be a corresponding mapping relationship between the frequency domain feature weight and the weight coefficient. Based on the mapping relationship, it can be known that the frequency domain feature weight of the high-confidence view angle is greater than the frequency domain feature weight of the low-confidence view angle, and correspondingly, the weight coefficient of the high-confidence view angle is also greater than the weight coefficient of the low-confidence view angle. Solving the global optimization three-dimensional trajectory of the skeleton node can improve the proportion of the high-confidence view angle data on the final trajectory.

[0096] For example, when the dancer's motion state changes dynamically, camera No. 5 is in a motion tangential view angle, and the data collected by camera No. 5 presents prominent harmonic continuity characteristics. This view angle shows high energy concentration and continuous harmonic distribution in the frequency domain characteristics in the main frequency band of the dance action, which conforms to the physical law of high-speed rotation. The increase in the weight increases the contribution of the observation displacement data of camera No. 5 in the objective function.

[0097] Step 145, based on the global optimization three-dimensional trajectory and the joint rotation freedom defined in the hierarchical kinematics constraint model, the rigid body motion parameters of the skeleton chain are calculated.

[0098] The joint rotation freedom refers to the angle range of the joint rotation around a specific axis. The rigid body motion parameter refers to the rigid motion description parameter including displacement, rotation and velocity.

[0099] In the embodiment of the present application, an interpolation algorithm can be used to calculate displacement, rotation and velocity.

[0100] Step 146, generating a three-dimensional motion sequence of the skeleton chain according to the rigid body motion parameters, and the three-dimensional motion sequence constitutes a dance action three-dimensional model.

[0101] The three-dimensional motion sequence of the skeleton chain refers to a set of skeleton node pose data arranged in time sequence.

[0102] In the embodiment of the present application, according to the rigid body motion parameters, the three-dimensional motion sequence of the skeleton chain can be calculated frame by frame through the forward kinematics chain rule. Specifically, the local transformation matrix of each joint can be multiplied in hierarchical order to generate the skeleton pose data at consecutive time stamps, and finally to constitute a dance action three-dimensional model. The model data format can be a Bio Vision Hierarchy (BVH) file, which contains, for example, a 600-frame motion sequence of 30 skeleton nodes.

[0103] The following is a specific example: In the dance rehearsal room, 12 optical motion capture devices are deployed in a circular arrangement to collect multi-perspective dynamic image data of the synchronized rotational movements of three dancers in real time; a hierarchical kinematic constraint model is constructed through joint topology analysis; the denoised spatiotemporal correlation data is mapped to the skeleton nodes, and the local spatial velocity characteristics of the elbow joint from the perspective of camera 3 are extracted; a dynamic equation is established based on the motion constraints, and the three-dimensional trajectory of the shoulder joint is optimized according to the weight coefficients of cameras 2 and 5 in the high-frequency motion bands of the limbs; a skeletal chain motion sequence is generated through rigid body motion parameter calculation, driving the virtual dancer model to complete the formation transformation; and finally, the formation phase difference is displayed in real time through the choreography collaboration terminal, and the conflicting areas of the movements are marked for the choreographer to adjust.

[0104] By executing steps 141 to 146, the embodiment of the present application realizes high-precision three-dimensional motion sequence generation of dance movements through joint topology constraints and multi-view frequency domain weight optimization, combined with rigid body kinematic chain calculation, improves movement smoothness and physical compliance, effectively solves the problem of multi-view data conflict, and supports real-time collaborative arrangement and physical rationality verification of complex choreography movements.

[0105] In a possible embodiment, step 143, establishing the dynamic motion equations of the skeletal nodes based on the local spatial features and the geometric constraint relationships between the skeletal nodes in the hierarchical kinematic constraint model, includes:

[0106] Step a1: Based on the parent-child hierarchical relationship between skeletal nodes in the hierarchical kinematic constraint model, determine the relative translation amount between the parent node and the child node in the three-dimensional space.

[0107] The parent-child hierarchical relationship between skeletal nodes refers to the subordinate connection relationship between joints. The relative translation in three-dimensional space refers to the displacement vector of the child node relative to the parent node.

[0108] In an embodiment of the present application, based on the parent-child hierarchical relationship between skeletal nodes in a hierarchical kinematic constraint model, a graph traversal algorithm can be used to traverse the skeletal topology tree and determine the relative translation of the parent and child nodes in three-dimensional space. The specific calculation method is: the world coordinates of the child node = the world coordinates of the parent node × the local transformation matrix + the relative translation vector, where the relative translation vector is determined by the anatomical length constraint of the skeleton and provides a theoretical displacement reference value for subsequent observation constraints.

[0109] Step a2: Convert the displacement vectors of the bone nodes in the local spatial features at each viewing angle into observation residuals of relative translation to construct multi-view observation constraints for each bone node.

[0110] The displacement vector is the change in position of a skeletal node at a specific viewpoint. The observation residual is the square of the difference between the actual observed displacement and the theoretically predicted displacement. The multi-view observation constraint is the optimization objective term that integrates the multi-view residuals to reduce data conflicts.

[0111] In an embodiment of the present application, the displacement vector of the skeletal node in the local spatial feature at each viewing angle is compared with the calculated theoretical value of the relative translation to calculate the observation residual. This embodiment does not specifically limit the calculation formula of the observation residual. A multi-view observation constraint term is constructed by the least squares method or linear weighting method. The expression of the constraint term can be: observation constraint = Σ(observation residual at each viewing angle × weight coefficient), and the weight coefficient is obtained by dynamically allocating the quality of the corresponding viewing angle data.

[0112] Step a3: Based on the fixed geometric lengths between the bone nodes in the hierarchical kinematic constraint model, construct rigid geometric constraint terms between the bone nodes.

[0113] The fixed geometric length between bone nodes refers to the physical length of the bones as defined by anatomical constraints. The rigid geometric constraint is an optimization penalty term that forces the bone length to remain constant, preventing unphysiological deformation.

[0114] In this embodiment of the present application, a rigid geometric constraint is constructed by calculating the difference between the actual distance between adjacent skeletal nodes and a fixed length. For example, the rigid geometric constraint = (actual distance - fixed length)². This constraint enforces the skeletal chain to maintain physiological plausibility during motion. It should be noted that this embodiment does not specifically limit the expression of the rigid geometric constraint.

[0115] However, in practical applications, geometric constraints refer to mathematical restrictions imposed when computing skeletal poses or joint angles to ensure that the generated motion conforms to human anatomical limitations. These constraints typically do not have a single, fixed, universal formula; their specific form depends on: the constraint type, the joint type, the optimization framework, and the modeling accuracy requirements.

[0116] The constraint type refers to the specific objective of the constraint, such as limiting joint angle range, preventing bone penetration, maintaining distance, and preventing self-intersection. Different joint types, such as ball-and-socket joints and hinge joints, have different degrees of freedom and therefore require different constraint representations. The optimization framework determines how the constraints are integrated into the optimization objective function for pose solving, acting as hard constraints, soft constraints, or penalties. Furthermore, depending on the modeling accuracy requirements, different constraints can be used for simplified and complex biomechanical models.

[0117] Step a4: Combine the multi-view observation constraint terms and the rigid geometric constraint terms into the dynamic motion equation of the skeleton node.

[0118] In an embodiment of the present application, multiple constraints exist simultaneously to ensure that the generated motion conforms to the anatomical limitations of the human body, and the multi-view observation constraint terms and the rigid geometric constraint terms are combined into the dynamic motion equation of the skeletal node.

[0119] Here's a specific example:

[0120] In the dance studio, the dancer's arm swinging movement data is collected through the process's optical motion capture equipment; based on the parent-child hierarchical relationship of the shoulder, elbow, and wrist, the theoretical translation of the elbow joint relative to the shoulder joint is calculated; the elbow joint displacement observed by camera No. 3 is compared with the theoretical value, generating a residual of 0.0001; the forearm bone length is constrained to be fixed at 0.3m; finally, the observation and geometric constraints are combined to optimize the generation of dynamic motion equations, driving the virtual character to accurately restore the arm swing angle error less than 1° action sequence.

[0121] By executing steps a1 to a4, the embodiment of the present application solves the problems of perspective data conflicts and non-physical motion in motion capture through the joint optimization of multi-perspective observation residual constraints and rigid bone length constraints, improves the biomechanical rationality and motion smoothness of the three-dimensional model of dance movements, and supports high-precision virtual character driving.

[0122] To enhance dance movement data, the present application embodiment may use devices such as spatial light modulators and liquid crystal arrays for phase modulation. In one possible embodiment, S12, by adjusting the phase difference of the light fields at different viewing angles, performs light field interference processing on the multi-view dynamic image data to generate enhanced motion capture data, including:

[0123] Step 121: Based on the spatial distribution position of the optical motion capture device, calculate the light field propagation path difference between adjacent perspectives to generate an initial light field phase difference.

[0124] The spatial distribution position refers to the installation coordinates and orientation parameters of the optical device in three-dimensional space. The light field propagation path refers to the geometric path of light from the light source, through the target, and to the camera. The initial light field phase difference refers to the static phase offset determined by the spatial layout of the device.

[0125] For example, the path difference ΔL = the distance between the optical axes of adjacent cameras × sin(angle) + the distance difference between the target skeletal node and the camera. The initial light field phase difference Δφ0 = 2π × ΔL / λ, where λ is the wavelength of the light source. This generates an initial light field phase difference that reflects the spatial layout characteristics and provides a reference value for dynamic compensation. This embodiment does not specifically limit the calculation formula for the light field propagation path difference between adjacent perspectives.

[0126] Step 122: Identify the movement speed of the skeleton nodes according to the temporal movement trajectory of the skeleton nodes in the multi-view dynamic image data, and generate the dynamic light field phase difference compensation amount according to the mapping relationship between the initial light field phase difference and the movement speed and the phase difference compensation amount.

[0127] The mapping relationship can be obtained by looking up a table. As for the specific expression of the mapping relationship between the motion speed and the phase difference compensation amount, the embodiments of the present application do not specifically limit it. The temporal motion trajectory refers to the continuous position sequence of the skeletal nodes changing over time. The phase difference compensation amount refers to the dynamic phase correction value that needs to be added to offset the motion error, which can correct the phase offset caused by the dancer's movement.

[0128] In this embodiment, the temporal motion trajectory of the skeleton nodes is extracted from the multi-view dynamic image data, and the motion velocity v is calculated as displacement / time interval. Based on the preset phase difference compensation mapping relationship, the dynamic light field phase difference compensation value is generated to offset the phase mismatch caused by motion.

[0129] Step 123: Based on the dynamic light field phase difference compensation amount, the initial light field phase difference is adjusted frame by frame to achieve the target light field phase difference, wherein the target light field phase difference refers to the phase difference value used for actual interference calculation after optimization.

[0130] In the embodiment of the present application, the initial light field phase difference is adjusted frame by frame based on the generated dynamic light field phase difference compensation. A linear interpolation algorithm is used to transition the current frame phase difference to the target light field phase difference, ensuring continuous phase changes between adjacent frames and avoiding interference fringe jumps.

[0131] Step 124 : Based on the target light field phase difference, the multi-view dynamic image data is superimposed in the spatial domain to obtain a superposition result, and the intensity distribution of the interference area is calculated according to the superposition result to extract the interference feature of the bone edge.

[0132] The spatial domain refers to the coordinate system of the two-dimensional image plane, and the superposition result refers to the composite light intensity distribution after multi-view light field interference. The intensity distribution of the interference region refers to the brightness variation pattern of the interference fringes in the image, and the interference feature of the bone edge refers to the bone contour information characterized by the sharpness variation of the interference fringes.

[0133] In an embodiment of the present application, multi-view dynamic image data is superimposed in the spatial domain according to the target light field phase difference, and the intensity distribution of the interference area is extracted by the peak detection algorithm to generate the interference characteristics of the bone edge.

[0134] Step 125: Fuse the interference features and the skeleton positioning information to generate enhanced motion capture data.

[0135] In this embodiment of the present application, the extracted bone edge interference features are fused with the bone positioning information. A feature weighting algorithm is used: enhanced data = bone positioning coordinates × 0.6 + interference feature coordinates × 0.4. This generates enhanced motion capture data with a resolution of 0.1mm, eliminating joint drift caused by motion blur.

[0136] Here's a specific example:

[0137] During the dance motion capture process, 12 optical motion capture devices are arranged in a ring to calculate the initial phase difference Δφ0=π / 3; for the dancer's rapidly rotating arm movements, a compensation amount Δφ'=0.2π is generated; the phase difference is adjusted to Δφ_target=0.95π frame by frame; the density of interference fringes at the edge of the elbow joint is detected to be improved through the spatial superposition algorithm; and finally, enhanced data is generated by fusion to accurately restore the wrist micro-movement of 2 rotations per second. The embodiment of the present application does not specifically limit the specific expression form of the spatial superposition algorithm.

[0138] By executing steps 121 to 125, the embodiment of the present application improves the ability to capture submillimeter details of bone edges through dynamic light field phase compensation and spatial interference superposition, and combines the timing optimization of motion trajectories to effectively suppress motion blur under high-speed movements and achieve high-accuracy dance movement data enhancement.

[0139] In a possible embodiment, step 122, identifying the motion speed of the skeleton nodes according to the temporal motion trajectory of the skeleton nodes in the multi-view dynamic image data, and generating the dynamic light field phase difference compensation amount according to the initial light field phase difference and the mapping relationship between the motion speed and the light field phase difference compensation amount, includes:

[0140] Step b1: Based on the temporal motion trajectory of the skeleton node, calculate the displacement change of the skeleton node between adjacent frames and generate the instantaneous motion speed parameter of the skeleton node.

[0141] The displacement change refers to the change in the 3D distance between two adjacent frames. The instantaneous motion speed parameter refers to the motion rate of a skeleton node within a single frame.

[0142] Step b2: According to the instantaneous motion speed parameter and the motion direction of the limb where the skeleton node is located, a nonlinear mapping relationship between the motion speed and the light field phase difference compensation amount is established to generate a phase difference adjustment function.

[0143] The nonlinear mapping relationship refers to a non-proportional functional relationship between velocity and phase compensation, and can be implemented using a lookup table, a fitted polynomial, a machine learning model, or the like. The phase difference adjustment function is a mathematical expression that converts velocity into phase compensation. This embodiment does not specifically limit the nonlinear mapping relationship or the expression of the phase difference adjustment function.

[0144] Step b3: Based on the phase difference adjustment function, the instantaneous motion speed parameter is converted into a temporally continuous dynamic phase difference compensation amount sequence.

[0145] The dynamic phase difference compensation amount sequence refers to a set of phase compensation amounts arranged in chronological order.

[0146] Step b4: performing temporal smoothing processing on the dynamic phase difference compensation amount sequence according to the motion acceleration of the skeleton nodes between adjacent frames to generate the dynamic light field phase difference compensation amount.

[0147] Among them, time series smoothing processing refers to eliminating high-frequency jitter components in the compensation amount sequence through a filtering algorithm.

[0148] The following is a specific example: In the continuous rotation motion capture of a dancer, the displacement change of the wrist joint between adjacent frames is calculated as 0.15m, generating an instantaneous velocity v = 1.5m / s; the rotational tangential motion is determined based on the direction cosines and the velocity is fitted; the velocity sequence is converted into a compensation sequence [0.3π, 0.35π, ...]; and a sudden acceleration increase of 2m / s is detected in the fifth frame. 2 , the compensation amount is smoothed from 0.5π to 0.48π through Kalman filtering, ultimately improving the stability of the light field interference fringes.

[0149] By executing steps b1 to b4, the embodiment of the present application realizes adaptive adjustment of the light field phase compensation amount under high-speed motion through nonlinear mapping of velocity and phase and acceleration-driven timing smoothing, effectively suppresses motion blur and phase jumps, and improves the stability of dynamic light field interference and the accuracy of bone edge reconstruction.

[0150] In one possible embodiment, S13, based on the Doppler frequency shift characteristics generated by the dancer's movements in the enhanced motion capture data, frequency domain suppression processing is performed on the dynamic range noise to generate noise-reduced multi-view spatiotemporal correlation data, including:

[0151] Step 131: Perform frequency domain decomposition on the temporal motion trajectory of the skeleton nodes in the enhanced motion capture data to generate Doppler frequency shift features.

[0152] Frequency domain decomposition is the process of converting a time domain signal into frequency components, revealing the frequency composition of the signal. For example, the time domain motion trajectory of a skeletal node is converted into a frequency domain signal via a short-time Fourier transform, and the frequency offset is extracted as the Doppler shift feature.

[0153] Step 132: Generate a dynamic noise spectrum model based on the periodic distribution law of the Doppler frequency shift characteristics in the frequency domain.

[0154] The periodic distribution law refers to the repetitive energy concentration phenomenon of the signal in the frequency domain. The dynamic noise spectrum model is a probabilistic model that describes the noise frequency range and energy distribution.

[0155] Step 133: Construct a frequency domain suppression mask according to the boundary threshold in the dynamic noise spectrum model.

[0156] The boundary threshold refers to the frequency range boundary that separates noise from valid signals. The frequency domain suppression mask is a binary matrix used to filter out signal components in a specific frequency band. For example, a binary frequency domain suppression mask can be constructed using the following mask construction rules: for frequency points in the noise energy candidate region, the mask value is set to 0; for frequency points in the valid signal band, the mask value is set to 1.

[0157] Step 134 : Apply the frequency domain suppression mask to the frequency domain components of the enhanced motion capture data to reconstruct the time domain signal and generate multi-view spatiotemporal correlation data after noise reduction.

[0158] The frequency domain component of enhanced motion capture data refers to the mathematical representation of high-precision skeletal motion data processed through light field interferometry in the frequency domain. It consists of the amplitude, phase, and energy distribution of each frequency component. Reconstructing the time domain signal is the process of inversely converting the filtered frequency domain data into time series data.

[0159] In an embodiment of the present application, the specific process of the frequency domain suppression mask acting on the frequency domain components of the enhanced motion capture data is as follows: first, the time domain signal of the enhanced motion capture data is divided into several time windows by short-time Fourier transform, and the signal in each window is converted into a frequency domain component; then, the binary frequency domain suppression mask is multiplied by the frequency domain component frequency point by frequency point to suppress the energy of the noise frequency band; then, the filtered frequency domain component is inversely short-time Fourier transformed, and the time domain signal of each window is reconstructed into a continuous time series signal by overlapping and adding; finally, the multi-view spatiotemporal correlation data after noise reduction is generated, the details of the skeleton trajectory of the effective motion frequency band are retained, and the jitter caused by periodic noise is eliminated. In other words, the time domain motion trajectory of the skeleton node in the enhanced motion capture data is input into the short-time Fourier transform algorithm, the signal segment is intercepted according to a fixed time window, the Fourier transform is applied to each window, the time domain signal is decomposed into frequency domain components, and a frequency domain amplitude-phase matrix containing the fundamental frequency, harmonics and noise components is generated.

[0160] The following is a specific example: in a dance studio, the enhanced motion capture data of the dancer's rapid arm movement shows that the wrist joint trajectory has a 50Hz periodic jitter; through short-time Fourier transform decomposition, it is found that the amplitude of 55Hz frequency point is abnormal; Gaussian mixture model modeling determines that 45-55Hz is the noise band; the binary mask of this frequency band is constructed; after applying the mask, the jitter amplitude of the reconstructed time domain signal is reduced, generating stable space-time data that can be used for three-dimensional modeling.

[0161] By performing steps 131-134, the embodiment of the application accurately identifies and suppresses the periodic interference frequency band through Doppler shift feature analysis and dynamic noise spectrum modeling, improves the time domain smoothness and multi-view data consistency of the skeletal motion trajectory, and provides high signal-to-noise ratio input for three-dimensional model reconstruction.

[0162] In one possible embodiment, step 132, based on the periodic distribution law of Doppler shift features in the frequency domain, generates a dynamic noise spectrum model, comprising:

[0163] Step c1, based on the periodic distribution law of Doppler shift features in the frequency domain, determines the effective signal frequency band composed of the main frequency and harmonic components of the dance action.

[0164] Wherein, the main frequency of the dance action refers to the fundamental frequency of the periodic motion of the dance action. The harmonic component refers to the integer multiple frequency component of the main frequency. The effective signal frequency band refers to the frequency range containing the main frequency and the harmonic, representing the effective motion information, indicating the frequency energy concentration interval corresponding to the periodic motion of the skeletal node.

[0165] Step c2, according to the dispersion degree of frequency energy distribution, extract noise energy candidate area from non-effective signal frequency band.

[0166] Wherein, the dispersion degree refers to the dispersion degree of frequency energy distribution, which is quantified by variance or standard deviation. The non-effective signal frequency band refers to the frequency interval without harmonic law and energy distribution. The higher the dispersion degree, the larger the variance or standard deviation, the greater the noise probability. The noise candidate area represents the frequency energy distribution interval that is not associated with the harmonic of the effective signal frequency band.

[0167] Step c3, according to the time sequence change of the motion speed of the skeletal node, dynamically adjust the boundary threshold value between the effective signal frequency band and the noise energy candidate area.

[0168] Wherein, the boundary threshold value refers to the energy critical value for distinguishing between effective signal and noise, and the threshold value expands the coverage range of the effective signal frequency band as the motion speed increases.

[0169] Specifically, an instantaneous velocity sequence is extracted from the time-domain motion trajectory of the skeletal node, and acceleration, which is the temporal variation of the motion velocity, is calculated. Based on the acceleration and boundary threshold adjustment formula, the boundary threshold is calculated. For example, the product of the normalized acceleration and the base threshold is used as the boundary threshold. It should be noted that the embodiments of this application do not specifically limit the expression of the boundary threshold adjustment formula.

[0170] Step c4: Associating the energy distribution of the noise energy candidate area with the boundary threshold to generate a dynamic noise spectrum model.

[0171] The energy distribution of the noise energy candidate area refers to the statistical characteristics such as the mean and variance of the candidate noise frequency band.

[0172] Specifically, a Gaussian mixture model is established based on the energy distribution of the noise energy candidate area and used as the dynamic noise spectrum model. The model is used to characterize the spectral characteristics of the noise. The model parameters include the weight, mean and standard deviation of each Gaussian component.

[0173] The following is a specific example: During the dancer's continuous accelerated rotation, short-time Fourier transform decomposition analysis shows that the main frequency increases from 3Hz to 6Hz; the 60-80Hz frequency band is detected with a variance of 200 and discontinuous energy, and is marked as a noise candidate area; the boundary threshold is dynamically adjusted from 0.6 to 0.9; finally, the energy mean of the 60Hz frequency band of 85 is lower than the energy threshold corresponding to the boundary threshold of 0.9×100=90, and is not included in the noise model, while the energy mean of the 70Hz frequency band of 95>90 is filtered out, thus retaining the true motion frequency band.

[0174] By executing steps c1 to c4, the embodiment of the present application adaptively distinguishes motion signals from noise frequency bands through dynamic threshold adjustment and noise energy distribution correlation modeling, solves the problem of over-suppression or under-suppression of the fixed threshold algorithm in speed change scenarios, and improves the robustness and accuracy of frequency domain noise reduction.

[0175] Figure 2 This is a structural diagram of an image data processing system for dance choreography provided in an embodiment of the present application, such as Figure 2 As shown, the system includes:

[0176] The acquisition module 21 is used to acquire multi-view dynamic image data of a dancer collected by optical motion capture devices distributed in different spatial positions, wherein the multi-view dynamic image data includes skeleton positioning information.

[0177] The adjustment module 22 is configured to perform light field interference processing on the multi-view dynamic image data by adjusting the light field phase difference of different view angles to generate enhanced motion capture data.

[0178] The generation module 23 is used to perform frequency domain suppression processing on the dynamic range noise based on the Doppler frequency shift characteristics generated by the dancer's movement in the enhanced motion capture data, and generate noise-reduced multi-view spatiotemporal correlation data.

[0179] The fusion module 24 is used to fuse the denoised multi-view spatiotemporal correlation data with the skeleton positioning information to generate a three-dimensional dance movement model.

[0180] The output module 25 is used to output the dance movement three-dimensional model to the choreography collaboration terminal, so that the choreography collaboration terminal can display the collaborative choreography effect of multiple dancers' movements in real time according to the joint movement trajectory of the dance movement three-dimensional model.

[0181] Figure 2 The image data processing system for dance choreography can be executed Figure 1 The implementation principles and technical effects of the image data processing method for dance choreography described in the illustrated embodiment are not further elaborated. The specific manner in which the various modules and units perform operations in the image data processing system for dance choreography in the aforementioned embodiment have been described in detail in the relevant embodiments of the method and will not be further elaborated here.

[0182] In one possible design, Figure 2 The image data processing system for dance choreography of the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32 .

[0183] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .

[0184] The processing component 32 is used to: obtain multi-perspective dynamic image data of a dancer collected by optical motion capture devices distributed in different spatial positions, wherein the multi-perspective dynamic image data includes skeletal positioning information; perform light field interference processing on the multi-perspective dynamic image data by adjusting the light field phase difference of different perspectives to generate enhanced motion capture data; perform frequency domain suppression processing on the dynamic range noise based on the Doppler frequency shift characteristics generated by the dancer's movement in the enhanced motion capture data to generate noise-reduced multi-perspective spatiotemporal correlation data; fuse the noise-reduced multi-perspective spatiotemporal correlation data with the skeletal positioning information to generate a three-dimensional model of dance movements; and output the three-dimensional model of dance movements to a choreography collaboration terminal so that the choreography collaboration terminal can display the collaborative choreography effect of multiple dancers' movements in real time according to the joint motion trajectory of the three-dimensional model of dance movements.

[0185] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.

[0186] The storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as random access memory (RAM), static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0187] Of course, a computing device may also include other components, such as input / output interfaces, display components, communication components, etc.

[0188] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.

[0189] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.

[0190] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.

[0191] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The illustrated embodiment provides an image data processing method for dance choreography.

[0192] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0193] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0194] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for processing image data for dance choreography, characterized in that: include: Acquiring multi-perspective dynamic image data of a dancer captured by optical motion capture devices distributed at different spatial positions, wherein the multi-perspective dynamic image data includes skeleton positioning information; By adjusting the phase difference of light fields at different viewing angles, light field interference processing is performed on the multi-view dynamic image data to generate enhanced motion capture data; Based on the Doppler frequency shift characteristics generated by the dancer's movements in the enhanced motion capture data, dynamic range noise is suppressed in the frequency domain to generate noise-reduced multi-view spatiotemporal correlation data; fusing the noise-reduced multi-view spatiotemporal correlation data with the skeleton positioning information to generate a three-dimensional dance movement model; The dance movement three-dimensional model is output to the choreography collaboration terminal, so that the choreography collaboration terminal displays the collaborative choreography effect of multiple dancers' movements in real time according to the joint movement trajectory of the dance movement three-dimensional model.

2. The method according to claim 1, characterized in that The step of fusing the noise-reduced multi-view spatiotemporal correlation data with the skeleton positioning information to generate a three-dimensional dance movement model includes: Extracting joint topological relationships from the multi-view dynamic image data, and constructing a hierarchical kinematic constraint model of skeletal nodes based on the skeletal positioning information and the joint topological relationships; Mapping the spatiotemporal coordinates of each view in the denoised multi-view spatiotemporal correlation data to a corresponding skeleton node, and extracting the local spatial features of each skeleton node under the corresponding view; Establishing a dynamic motion equation of the skeletal node according to the geometric constraint relationship between the local spatial features and the skeletal nodes in the hierarchical kinematic constraint model; Dynamically adjust the weight coefficient of each perspective data in the dynamic motion equation according to the frequency domain feature weights of different perspectives to generate a global optimized three-dimensional trajectory of the skeleton node; Calculating the rigid body motion parameters of the skeletal chain based on the global optimized three-dimensional trajectory and the joint rotational degrees of freedom defined in the hierarchical kinematic constraint model; A three-dimensional motion sequence of a skeleton chain is generated according to the rigid body motion parameters, and the three-dimensional motion sequence constitutes a three-dimensional dance movement model.

3. The method according to claim 2, characterized in that The step of establishing the dynamic motion equation of the skeletal nodes according to the geometric constraint relationship between the local spatial features and the skeletal nodes in the hierarchical kinematic constraint model includes: Determining the relative translation of the parent node and the child node in three-dimensional space based on the parent-child hierarchical relationship between the skeletal nodes in the hierarchical kinematic constraint model; Convert the displacement vectors of the skeleton nodes in the local spatial features at each viewing angle into the observation residuals of the relative translation amount to construct a multi-view observation constraint item for each skeleton node; Constructing rigid geometric constraint terms between bone nodes based on fixed geometric lengths between bone nodes in the hierarchical kinematic constraint model; The multi-view observation constraint term and the rigid geometric constraint term are combined into a dynamic motion equation of the skeleton node.

4. The method according to claim 1, wherein The method of performing light field interference processing on the multi-view dynamic image data by adjusting the light field phase difference of different view angles to generate enhanced motion capture data includes: Based on the spatial distribution position of the optical motion capture device, calculating the light field propagation path difference between adjacent perspectives to generate an initial light field phase difference; Identifying the movement speed of the skeleton nodes according to the temporal movement trajectory of the skeleton nodes in the multi-view dynamic image data, and generating a dynamic light field phase difference compensation amount according to the initial light field phase difference and a mapping relationship between the movement speed and the phase difference compensation amount; Based on the dynamic light field phase difference compensation amount, adjusting the initial light field phase difference frame by frame to achieve a target light field phase difference; Based on the target light field phase difference, the multi-view dynamic image data is superimposed in the spatial domain to obtain a superposition result, and the intensity distribution of the interference area is calculated according to the superposition result to extract the interference characteristics of the bone edge; The interference feature is fused with the skeleton positioning information to generate enhanced motion capture data.

5. The method according to claim 4, characterized in that The method includes: identifying the movement speed of the skeleton nodes according to the time-series movement trajectory of the skeleton nodes in the multi-view dynamic image data; and generating the dynamic light field phase difference compensation amount according to the initial light field phase difference and the mapping relationship between the movement speed and the light field phase difference compensation amount. Based on the temporal motion trajectory of the skeleton node, the displacement change of the skeleton node between adjacent frames is calculated to generate the instantaneous motion speed parameter of the skeleton node; According to the instantaneous motion speed parameter and the motion direction of the limb where the skeletal node is located, a nonlinear mapping relationship between the motion speed and the light field phase difference compensation amount is established to generate a phase difference adjustment function; Based on the phase difference adjustment function, the instantaneous motion speed parameter is converted into a time-series continuous dynamic phase difference compensation amount sequence; According to the motion acceleration of the skeleton nodes between adjacent frames, the dynamic phase difference compensation amount sequence is subjected to temporal smoothing processing to generate the dynamic light field phase difference compensation amount.

6. The method according to claim 1, characterized in that The method of performing frequency domain suppression processing on dynamic range noise based on Doppler frequency shift characteristics generated by dancer's movement in the enhanced motion capture data to generate noise-reduced multi-view spatiotemporal correlation data includes: Performing frequency domain decomposition on the temporal motion trajectory of the skeletal nodes in the enhanced motion capture data to generate Doppler frequency shift features; Based on the periodic distribution law of the Doppler frequency shift characteristics in the frequency domain, a dynamic noise spectrum model is generated; constructing a frequency domain suppression mask according to a boundary threshold in the dynamic noise spectrum model; The frequency domain suppression mask is applied to the frequency domain components of the enhanced motion capture data to reconstruct the time domain signal and generate multi-view spatiotemporal correlation data after noise reduction.

7. The method according to claim 6, characterized in that The generating of a dynamic noise spectrum model based on the periodic distribution law of the Doppler frequency shift characteristics in the frequency domain includes: Based on the periodic distribution law of the Doppler frequency shift characteristics in the frequency domain, determining the effective signal frequency band composed of the main frequency and harmonic components of the dance movement; According to the discrete degree of frequency domain energy distribution, the noise energy candidate area is extracted from the non-effective signal frequency band; Dynamically adjusting the boundary threshold between the effective signal frequency band and the noise energy candidate area according to the temporal variation of the movement speed of the skeleton node; The energy distribution of the noise energy candidate area is associated with the boundary threshold to generate a dynamic noise spectrum model.

8. An image data processing system for dance choreography, characterized in that: include: an acquisition module, configured to acquire multi-perspective dynamic image data of a dancer collected by optical motion capture devices distributed at different spatial positions, wherein the multi-perspective dynamic image data includes skeleton positioning information; an adjustment module, configured to perform light field interference processing on the multi-view dynamic image data by adjusting the phase difference of the light fields at different viewpoints to generate enhanced motion capture data; A generation module, configured to perform frequency domain suppression processing on dynamic range noise based on Doppler frequency shift characteristics generated by dancer movements in the enhanced motion capture data, thereby generating noise-reduced multi-view spatiotemporal correlation data; A fusion module, configured to fuse the noise-reduced multi-view spatiotemporal correlation data with the skeleton positioning information to generate a three-dimensional dance movement model; The output module is used to output the dance movement three-dimensional model to the choreography collaboration terminal, so that the choreography collaboration terminal can display the collaborative choreography effect of multiple dancers' movements in real time according to the joint movement trajectory of the dance movement three-dimensional model.

9. A computing device, characterized in that The method comprises a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement an image data processing method for dance choreography as described in any one of claims 1 to 7.

10. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a computer, the image data processing method for dance choreography according to any one of claims 1 to 7 is implemented.