A method and system for detecting multi-target motion synchronization in a group dance performance
By using multi-view video processing and a neural network for skeletal coordinate extraction, a three-dimensional skeletal coordinate sequence and a phase synchronization index are calculated, solving the problem of high cost in existing technologies for detecting the synchronization of group dance movements and achieving high-precision synchronization assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MINZU UNIVERSITY OF CHINA
- Filing Date
- 2026-04-16
- Publication Date
- 2026-05-29
AI Technical Summary
Existing group dance performance synchronization detection technology relies on high-precision optical motion capture systems, which are costly and complex to deploy, making them difficult to adapt to conventional stages and rehearsal scenarios, and unable to achieve accurate quantification and standardized evaluation.
By acquiring and spatially calibrating multi-view video, a skeletal coordinate extraction neural network is used to calculate a three-dimensional skeletal coordinate sequence and a phase synchronization index. Combined with the spatial coincidence index, multi-target action synchronization detection results are generated.
It achieves high integrity and high accuracy detection of the synchronization of group dance movements in conventional stage and rehearsal scenarios, generates objective and quantifiable synchronization evaluation indicators, and eliminates the influence of individual physiological and spatial deviations.
Smart Images

Figure CN122116483A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of multi-target motion synchronization detection, and in particular relates to a method and system for multi-target motion synchronization detection in group dance performances. Background Technology
[0002] As one of the core forms of stage performance art, the synchronization of movements in group dance is a core standard for judging performance quality, directly determining the presentation effect and artistic appeal of the stage performance. For a long time, the judgment of the synchronization of group dance performances has relied entirely on the subjective scoring of professional judges. This model has inherent flaws that are difficult to avoid: the judgment results are easily affected by non-technical factors such as the judge's personal professional experience, subjective preferences, visual fatigue, and the on-site environment, making it impossible to accurately quantify and locate the deviation in movement synchronization, and also making it difficult to form standardized judgment results that are reproducible and comparable across different audiences. At the same time, in daily rehearsal scenarios, the limits of human visual perception cannot capture subtle temporal and spatial deviations, making it difficult to achieve refined improvement in the quality of group dance performances. This industry pain point constitutes the core demand background and underlying driving force for the development of group dance movement synchronization detection technology.
[0003] With the iterative development of motion capture, computer vision, and signal processing technologies, the technology for detecting the synchronization of group dance movements has gradually achieved technological leaps from contact motion capture to markerless visual detection, from two-dimensional to three-dimensional, and from single-target to multi-target. However, existing motion synchronization detection technologies often rely on high-precision optical motion capture systems, which require attaching special optical markers to the dancers' bodies and using multiple high-speed cameras to capture the spatial position changes of the markers to achieve digital acquisition and synchronization analysis of movements. However, due to the high cost of the equipment and the complex deployment process, it is difficult to adapt to conventional stage performances and daily rehearsals, making it difficult to achieve widespread application. Summary of the Invention
[0004] Therefore, it is necessary to provide a labelless method and system for detecting the synchronization of multi-target movements in group dance performances, which can accurately locate movement deviations, in order to address the above-mentioned technical problems.
[0005] Firstly, this application provides a method for detecting the synchronization of multi-target movements in a group dance performance, including:
[0006] Acquire multi-view video of a group dance performance, and perform spatial calibration and temporal synchronization on the multi-view video of the group dance performance to obtain standard multi-view video and camera calibration parameters;
[0007] The camera calibration parameters and multi-view images at each time point in the standard multi-view video are input into the skeletal coordinate extraction neural network to obtain the calibrated 3D skeletal coordinates of each dancer at each time point. Based on the calibrated 3D skeletal coordinates of each dancer at each time point, the 3D skeletal coordinate sequence and 3D skeletal motion sequence of each dancer are constructed.
[0008] The phase sequence of dance movements is extracted based on the three-dimensional skeletal motion sequence, and the phase synchronization index is calculated based on the instantaneous phase difference between each pair of dancers' dance movement phase sequences at each time stamp.
[0009] The three-dimensional skeletal coordinate sequences of each dancer are mapped to the standard spatial coordinate system, the three-dimensional skeletal coordinate sequences of each dancer are registered, and the spatially aligned three-dimensional skeletal coordinate sequences of each dancer are obtained. Based on the spatially aligned three-dimensional skeletal coordinate sequences of each dancer, the spatial overlap index is calculated.
[0010] By integrating the phase synchronization index and the spatial coincidence index, multi-target action synchronization detection results are generated.
[0011] In one embodiment, a three-dimensional skeletal coordinate sequence and a three-dimensional skeletal motion sequence for each dancer are constructed based on the calibrated three-dimensional skeletal coordinates of each dancer at each timestamp, including:
[0012] The three-dimensional skeletal coordinate sequence of each dancer is constructed according to the calibrated three-dimensional skeletal coordinates of each dancer at each time stamp in chronological order;
[0013] Temporal difference operation is performed on the calibrated 3D skeletal coordinates of each time point in the 3D skeletal coordinate sequence to calculate the 3D skeletal point motion vectors of each skeletal key point of each dancer at each time point;
[0014] By splicing the 3D skeletal motion vectors of the key points of the same skeletal structure of the same dancer at the same time stamp, the 3D skeletal motion matrix of each dancer at each time stamp is obtained.
[0015] The 3D skeletal motion sequence of each dancer is constructed by chronologically constructing the 3D skeletal motion matrix of each dancer at each time stamp.
[0016] In one embodiment, a dance movement phase sequence is extracted based on a three-dimensional skeletal motion sequence, including:
[0017] The three-dimensional skeleton motion matrix in the three-dimensional skeleton motion sequence is weighted and fused to obtain the action time sequence;
[0018] The action time sequence is orthogonally transformed by Hilbert transform to generate a conjugate action time sequence that is completely orthogonal to the action time sequence; the action time sequence is used as the real part and the conjugate action time sequence is used as the imaginary part to construct a complex analytic action time sequence.
[0019] Solve the instantaneous phase of each timestamp in the complex analytic motion temporal sequence operator, and construct the dance motion phase sequence of each dancer according to the instantaneous phase of each timestamp in temporal order.
[0020] In one embodiment, the expression for the phase synchronization index is:
[0021]
[0022] In the formula, For phase synchronization index, For the first The timestamp of the first The dancer and the first The instantaneous phase difference of movements between individual dancers Solving for the modulus operator, The total number of timestamps. The total number of dancers.
[0023] In one embodiment, camera calibration parameters and multi-view images at each time stamp in a standard multi-view video are input into a skeletal coordinate extraction neural network to obtain the calibrated three-dimensional skeletal coordinates of each dancer at each time stamp, including:
[0024] The multi-view images at each time stamp in the standard multi-view video are input into the two-dimensional skeletal coordinate extraction module in the skeletal coordinate extraction neural network to extract the two-dimensional skeletal coordinate set of each dancer in each multi-view image;
[0025] Based on the two-dimensional skeletal coordinate sets of each dancer in each multi-view image, and combined with camera calibration parameters, the three-dimensional skeletal coordinate sets corresponding to the two-dimensional skeletal coordinate sets of each dancer in each multi-view image are reconstructed.
[0026] Based on the spatial distance between the three-dimensional skeletal coordinate sets in each multi-view image, the three-dimensional skeletal coordinate sets of each dancer at each time stamp in the standard multi-view video are matched.
[0027] The camera calibration parameters and the set of 3D skeletal coordinates of each dancer at each time stamp are input into the 3D skeletal coordinate fusion and calibration module in the skeletal coordinate extraction neural network to obtain the calibrated 3D skeletal coordinates of each dancer at each time stamp.
[0028] Secondly, this application also provides a system for detecting the synchronization of multi-target movements in a group dance performance, including:
[0029] The performance video acquisition module is used to acquire multi-view videos of group dance performances, and to perform spatial calibration and temporal synchronization on the multi-view videos of group dance performances to obtain standard multi-view videos and camera calibration parameters;
[0030] The skeleton coordinate extraction module is used to input the camera calibration parameters and multi-view images of each time stamp in the standard multi-view video into the skeleton coordinate extraction neural network to obtain the calibrated three-dimensional skeleton coordinates of each dancer at each time stamp, and to construct the three-dimensional skeleton coordinate sequence and three-dimensional skeleton motion sequence of each dancer based on the calibrated three-dimensional skeleton coordinates of each dancer at each time stamp.
[0031] The phase synchronization index calculation module is used to extract the phase sequence of dance movements based on the three-dimensional skeletal motion sequence, and to calculate the phase synchronization index based on the instantaneous phase difference between each pair of dancers' dance movement phase sequences at each time stamp.
[0032] The spatial overlap index calculation module is used to map the three-dimensional skeletal coordinate sequence of each dancer to the standard spatial coordinate system, register the three-dimensional skeletal coordinate sequence of each dancer, obtain the spatially aligned three-dimensional skeletal coordinate sequence of each dancer, and calculate the spatial overlap index based on the spatially aligned three-dimensional skeletal coordinate sequence of each dancer.
[0033] The detection result generation module integrates the phase synchronization index and the spatial coincidence index to generate multi-target action synchronization detection results.
[0034] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method as described in any of the first aspects of this application.
[0035] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any of the first aspects of this application.
[0036] The aforementioned method and system for detecting multi-target motion synchronization in group dance performances acquires a complete video sequence of the group dance performance by deploying multi-viewpoint image acquisition devices. It performs unified spatial calibration and temporal synchronization processing on the multi-viewpoint videos, establishing a stable and unified global spatiotemporal reference and eliminating spatial coordinate deviations and temporal sequence errors between multi-viewpoint devices. By inputting the spatiotemporally synchronized camera calibration parameters and multi-viewpoint images into a dedicated skeletal coordinate extraction neural network, it can completely reconstruct the three-dimensional skeletal spatial position changes and motion state changes of each dancer throughout the entire performance cycle, accurately representing the continuous dynamic process of the dancers' dance movements and generating highly complete and accurate digital representation data of the dancers' movements. Based on the dancers' three-dimensional skeletal motion sequence, it extracts the motion features of key points of each skeleton and fuses them to generate the main motion temporal signal. After signal processing and conversion into a complex analytic signal, it calculates the continuous dance movement phase sequence. Then, based on the instantaneous phase difference between each pair of dancers at each timestamp, it performs group phase aggregation and statistics to calculate the phase synchronization index, which can separate the amplitude of movements and individual dynamics. This method accurately depicts the temporal progression and rhythmic changes of each dancer's movements throughout the entire performance cycle, quantifies the degree of temporal deviation in the movements of each dancer within the group, and generates objective and quantifiable evaluation indicators for the temporal synchronization of group movements. By mapping the three-dimensional skeletal coordinate sequences of each dancer to a unified standard spatial coordinate system, and performing rigid body transformation and scale normalization operations to complete the global registration of each dancer's skeletal coordinates, a spatially aligned three-dimensional skeletal coordinate sequence is obtained, and the spatial overlap index is calculated. Then, the phase synchronization index and the spatial overlap index are integrated to generate multi-target movement synchronization detection results. This method can eliminate non-movement-related spatial deviations such as individual physiological differences, stance offsets, and different limb orientations among dancers, and accurately quantifies the degree of overlap of the movements of each dancer within the group in the spatial posture dimension. By combining the phase synchronization index in the time dimension to achieve a two-dimensional comprehensive evaluation, it can comprehensively and objectively reflect the overall synchronization status of multi-target movements in group dance performances, generating complete detection results containing information in both time and space dimensions, and achieving full-dimensional and comprehensive detection of the synchronization of group dance movements. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 A flowchart illustrating a method for detecting the synchronization of multi-target movements in a group dance performance, provided as an embodiment of this application;
[0039] Figure 2This is a schematic diagram of the structure of a multi-target motion synchronization detection system in a group dance performance, provided as an embodiment of this application. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0041] In one exemplary embodiment of this application, such as Figure 1 As shown, a method for detecting the synchronization of multi-target movements in a group dance performance is provided. This embodiment illustrates the application of this method to a synchronization detection terminal. It is understood that this method can also be applied to a synchronization detection server, and further to a synchronization detection system including a synchronization detection terminal and a synchronization detection server, and is implemented through the interaction between the synchronization detection terminal and the synchronization detection server. In this embodiment, the method includes the following steps:
[0042] Step S101: Acquire multi-view video of group dance performance, and perform spatial calibration and temporal synchronization on the multi-view video of group dance performance to obtain standard multi-view video and camera calibration parameters.
[0043] Specifically, after authorization and verification, the synchronization detection terminal can establish communication connections with multiple image acquisition devices deployed in the stage area of the group dance performance. The synchronization detection terminal can acquire video data of the entire group dance performance in real time through each image acquisition device. Each image acquisition device can acquire video frames at a unified acquisition frequency, and the video data acquired by each device can be single-viewpoint video. The synchronization detection terminal can then construct a multi-viewpoint video of the group dance performance based on these single-viewpoint videos.
[0044] Optionally, the synchronization detection terminal can perform spatial calibration on the acquired multi-viewpoint video of the group dance performance, construct a unified global world coordinate system for the stage, and, combined with the physical installation location information of the image acquisition devices, calculate the camera calibration parameters of each image acquisition device through a camera calibration algorithm. The camera calibration parameters may include, but are not limited to, the intrinsic parameter matrix, extrinsic parameter matrix, and distortion correction coefficients of the image acquisition devices.
[0045] Furthermore, the synchronization detection terminal can extract the timestamp information of all image frames from each viewpoint video in a multi-view video of a group dance performance. Based on the main time reference, it calculates the difference between the timestamps of each image acquisition device and the main time reference, and then uses this difference to calibrate and compensate the timestamps of the image frames in each individual viewpoint video. The synchronization detection terminal can then set the multi-view video of a group dance performance that has completed spatial calibration and temporal synchronization as a standard multi-view video.
[0046] Step S102: Input the camera calibration parameters and the multi-view images of each time stamp in the standard multi-view video into the skeletal coordinate extraction neural network to obtain the calibrated three-dimensional skeletal coordinates of each dancer at each time stamp, and construct the three-dimensional skeletal coordinate sequence and three-dimensional skeletal motion sequence of each dancer based on the calibrated three-dimensional skeletal coordinates of each dancer at each time stamp.
[0047] Specifically, the synchronization detection terminal can extract multi-view images from standard multi-view videos according to their timestamp order. The terminal can then input each individual viewpoint image from each timestamp's multi-view images into the 2D skeletal coordinate extraction module of the skeletal coordinate extraction neural network to identify the dancer's skeletal keypoints in each individual viewpoint image, thereby extracting the 2D skeletal coordinate set for each dancer. The 2D skeletal coordinate set for a particular dancer at a specific timestamp can include the 2D pixel coordinates of all skeletal keypoints of that dancer at that timestamp in the multi-view images.
[0048] Furthermore, the synchronization detection terminal can, based on the two-dimensional skeletal coordinate sets of each dancer in each multi-view image and combined with camera calibration parameters, complete the preliminary reconstruction of the three-dimensional skeletal coordinates through a multi-view stereo vision reconstruction algorithm. The synchronization detection terminal can combine the two-dimensional skeletal coordinate sets of the same dancer in each multi-view image with the corresponding camera calibration parameters, and using the principle of triangulation, convert the two-dimensional pixel coordinates into three-dimensional spatial coordinates in the stage global world coordinate system, thus reconstructing the three-dimensional skeletal coordinate sets corresponding to the two-dimensional skeletal coordinate sets of each dancer in each multi-view image. The three-dimensional skeletal coordinate set of a particular dancer can include the three-dimensional spatial coordinates of all the dancer's skeletal keypoints in the stage global world coordinate system.
[0049] Optionally, the synchronization detection terminal can perform dancer identity matching processing on the 3D skeletal coordinate sets in various multi-view images at the same time stamp, based on the spatial distance between the 3D skeletal coordinate sets of each individual view image in the multi-view image. The synchronization detection terminal can calculate the overall spatial distance between any two 3D skeletal coordinate sets in different multi-view images at the same time stamp. The synchronization detection terminal determines that two 3D skeletal coordinate sets with an overall spatial distance less than a distance threshold belong to the same dancer, thereby matching the 3D skeletal coordinate sets of each dancer at each time stamp in the standard multi-view video.
[0050] Optionally, the overall spatial distance between any two 3D skeleton coordinate sets in different multi-view images at the same time stamp can be the average of the 3D spatial distances of each skeleton key point in the two 3D skeleton coordinate sets.
[0051] For example, the synchronization detection terminal can input the camera calibration parameters and the three-dimensional skeletal coordinate sets of each dancer at each time stamp into the three-dimensional skeletal coordinate fusion calibration module in the skeletal coordinate extraction neural network, perform distortion correction and error compensation on the coordinate data in each three-dimensional skeletal coordinate set, and fuse multiple three-dimensional skeletal coordinate sets of the same dancer to obtain the calibrated three-dimensional skeletal coordinates of each dancer at each time stamp.
[0052] Schematic illustration: The synchronization detection terminal can construct a sequence of three-dimensional skeletal coordinates for each dancer based on their calibrated three-dimensional skeletal coordinates at each time stamp. Each dancer in a group dance performance corresponds to a unique sequence of three-dimensional skeletal coordinates. This sequence of three-dimensional skeletal coordinates can be used to characterize the spatial position changes of key points on a dancer's limbs throughout the entire performance cycle.
[0053] Optionally, the synchronization detection terminal can perform temporal difference operations on the calibrated three-dimensional skeletal coordinates of each time point in the three-dimensional skeletal coordinate sequence to calculate the three-dimensional skeletal point motion vectors of each skeletal key point of each dancer at each time point.
[0054] For example, the three-dimensional skeletal point motion vector may include a three-dimensional skeletal point velocity vector and / or a three-dimensional skeletal point acceleration vector. The synchronization detection terminal can calculate the coordinate difference of each skeletal keypoint of each dancer at two adjacent time stamps, and then divide the coordinate difference by the time interval between the two adjacent time stamps to obtain the three-dimensional skeletal point velocity vector of the skeletal keypoint at the corresponding time stamp. The synchronization detection terminal can also calculate the velocity vector difference of each skeletal keypoint of each dancer at two adjacent time stamps, and then divide the velocity vector difference by the time interval between the two adjacent time stamps to obtain the three-dimensional skeletal point acceleration vector of the skeletal keypoint at the corresponding time stamp. The three-dimensional skeletal point motion vector can characterize the motion direction and amplitude of the skeletal keypoint, reflecting the instantaneous motion state of the skeletal keypoint.
[0055] Optionally, the synchronization detection terminal can stitch together the three-dimensional skeletal motion vectors of each skeletal key point of the same dancer at the same time stamp to obtain the three-dimensional skeletal motion matrix of each dancer at each time stamp.
[0056] Step S103: Extract the dance movement phase sequence based on the three-dimensional skeletal motion sequence, and calculate the phase synchronization index based on the instantaneous phase difference between each pair of dancers' dance movement phase sequences at each time stamp.
[0057] Specifically, the synchronization detection terminal can perform feature fusion on the 3D skeletal motion sequences of each dancer to obtain a motion timing sequence. The terminal calculates the magnitude of the motion vector of each 3D skeletal point in the 3D skeletal motion matrix, and then performs a weighted sum of the magnitudes of all key skeletal points according to preset weighting coefficients to obtain the motion feature fusion value at each timestamp. The terminal then arranges these motion feature fusion values in chronological order to obtain the motion timing sequence. This motion timing sequence can characterize the core rhythmic variation patterns of the dancers' movements.
[0058] Schematic, when the motion vector of a three-dimensional skeleton point can include a three-dimensional skeleton point velocity vector or a three-dimensional skeleton point acceleration vector, the action timing sequence can be a one-dimensional timing sequence; when the motion vector of a three-dimensional skeleton point can include a three-dimensional skeleton point velocity vector and a three-dimensional skeleton point acceleration vector, the action timing sequence can be a two-dimensional timing sequence.
[0059] Optionally, the synchronization detection terminal can perform an orthogonal transformation on the action time sequence using a Hilbert transform to generate a conjugate action time sequence that is completely orthogonal to the original action time sequence. The synchronization detection terminal can input the action time sequence as a real value into the Hilbert transform operator to generate the conjugate action time sequence. The conjugate action time sequence remains completely orthogonal to the original action time sequence, with a fixed phase offset only.
[0060] Furthermore, the synchronization detection terminal can construct a complex analytic action time sequence by using the action time sequence as the real part and the conjugate action time sequence as the imaginary part. The magnitude of the complex analytic action time sequence corresponds to the instantaneous amplitude of the original action time sequence, and the argument of the complex analytic action time sequence corresponds to the instantaneous phase of the original action time sequence, thus allowing focus on the temporal progression and rhythmic changes of the action.
[0061] Optionally, the synchronization detection terminal can solve for the instantaneous phase of each timestamp in the complex analytic motion time sequence operator to obtain the instantaneous phase of each dancer at each timestamp. The synchronization detection terminal can extract the argument angle of the complex values of each timestamp in the complex analytic motion time sequence using the argument angle solver to obtain the instantaneous phase of the dancer at that timestamp. The instantaneous phase can be used to characterize the temporal progression state of the dancer's dance movement at the corresponding timestamp within its complete motion cycle, and the rate of change of the instantaneous phase can reflect the rhythmic speed of the dancer's movements.
[0062] Optionally, the synchronization detection terminal can construct a dance movement phase sequence for each dancer based on the instantaneous phase of each dancer at each time stamp in a temporal sequence. The synchronization detection terminal can calculate the phase synchronization index based on the instantaneous phase difference between each pair of dancers' dance movement phase sequences at each time stamp.
[0063] For example, the synchronization detection terminal can perform a complex exponential transformation on the instantaneous phase difference of the movements between all pairs of dancers at each timestamp, and calculate... The instantaneous phase difference of each movement is converted into a unit vector on the complex plane, and all the converted complex exponential vectors under each timestamp are summed. The summation result is then divided by the product of the total number of dancers and the total number of dancers minus one. This completes the normalization of the group pairing dimension.
[0064] Furthermore, the synchronization detection terminal can calculate the modulus of the summation of the normalized complex exponential vectors at each timestamp using a modulus calculation operator. This modulus result characterizes the phase consistency of the group dance performance at the corresponding timestamp. A modulus result closer to 1 indicates a more consistent instantaneous phase of movement among all dancers within the group at that timestamp, signifying stronger temporal synchronization of the group's movements. Conversely, a modulus result closer to 0 indicates a more dispersed instantaneous phase of movement among dancers within the group at that timestamp, signifying weaker temporal synchronization of the group's movements. The synchronization detection terminal can sum the modulus calculation results for all timestamps in the standard multi-view video and then divide the sum by the total number of timestamps. After normalizing the time dimension, the phase synchronization index of the group dance performance is obtained. Phase synchronization index The value ranges from 0 to 1. The closer the value is to 1, the stronger the overall synchronization of the group movements in the group dance performance, and the higher the consistency of the rhythm and timing of each dancer's dance movements. The closer the value is to 0, the weaker the overall synchronization of the group movements in the group dance performance, and the greater the deviation in the timing of each dancer's dance movements.
[0065] For example, the expression for the phase synchronization index can be:
[0066]
[0067] In the formula, For phase synchronization index, For the first The timestamp of the first The dancer and the first The instantaneous phase difference of movements between individual dancers Solving for the modulus operator, The total number of timestamps. The total number of dancers.
[0068] Step S104: Map the three-dimensional skeletal coordinate sequence of each dancer to the standard spatial coordinate system, register the three-dimensional skeletal coordinate sequence of each dancer, obtain the spatially aligned three-dimensional skeletal coordinate sequence of each dancer, and calculate the spatial overlap index based on the spatially aligned three-dimensional skeletal coordinate sequence of each dancer.
[0069] Specifically, the synchronization detection terminal can construct a standard spatial coordinate system and accurately map the calibrated three-dimensional skeletal coordinates in the three-dimensional skeletal coordinate sequence of each dancer from the global world coordinate system of the stage to the standard spatial coordinate system.
[0070] Optionally, the synchronization detection terminal uses the center point of the torso of the standard human body model as the origin of the standard spatial coordinate system, and the vertical direction of the torso of the standard human body model as the longitudinal axis, the left and right direction as the transverse axis, and the front and back direction as the longitudinal axis.
[0071] Optionally, the synchronization detection terminal can extract the coordinates of each dancer's torso reference center point and limb proportion features, calculate a scale normalization coefficient based on the limb proportions of a standard human body model, and map the dancer's limb scale to the uniform scale of the standard human body model using this scale normalization coefficient. The synchronization detection terminal can also calculate the 3D rotation transformation matrix for each dancer, perform rotation transformations on the calibrated 3D skeletal coordinates of each dancer, and precisely align the dancer's torso orientation to the reference orientation of the standard spatial coordinate system. Finally, the synchronization detection terminal can use translation transformations to map the dancer's torso reference center point coordinates to the origin of the standard spatial coordinate system.
[0072] Optionally, the synchronization detection terminal can arrange the skeletal coordinate data of each dancer after spatial registration according to the order of timestamps to obtain a spatially aligned 3D skeletal coordinate sequence for each dancer. The synchronization detection terminal can then calculate the spatial overlap index based on the spatially aligned 3D skeletal coordinate sequence of each dancer.
[0073] For example, the synchronization detection terminal can extract the 3D coordinates of skeletal keypoints from the spatial alignment 3D skeletal coordinate sequence of each dancer at each time stamp in a standard multi-view video, and calculate the Euclidean distance between the 3D coordinates of the skeletal keypoints of each pair of dancers at that time stamp, thus obtaining the spatial position deviation of each skeletal keypoint. The synchronization detection terminal can perform a weighted summation of the spatial position deviations of all skeletal keypoints between each pair of dancers at each time stamp, and then divide the summation result by the product of the total number of dancers and the total number of dancers minus one, thus normalizing the group pairing dimension. The synchronization detection terminal can then divide the normalized result of the group pairing dimension by the total number of skeletal keypoints involved in the group dance performance, thus normalizing the skeletal keypoint dimension, and obtaining the average group spatial deviation at that time stamp. The synchronization detection terminal can perform a weighted summation of the average group spatial deviation at each time stamp in the standard multi-view video, and then divide the weighted summation result by the total number of time stamps, thus obtaining the overall average spatial deviation of the entire group dance performance. The synchronization detection terminal can convert the overall spatial average deviation value into a spatial overlap index by subtracting 1 from the spatial average deviation value. The spatial overlap index can range from 0 to 1. The closer the spatial overlap index is to 1, the stronger the spatial consistency of the group movements in the group dance performance, and the higher the uniformity of the dance postures and amplitudes of each dancer. The closer the spatial overlap index is to 0, the weaker the spatial consistency of the group movements in the group dance performance, and the greater the deviation of the dance postures and amplitudes of each dancer.
[0074] Step S105: Integrate the phase synchronization index and the spatial coincidence index to generate multi-target action synchronization detection results.
[0075] Specifically, the synchronization detection terminal can either perform a weighted summation of the phase synchronization index and the spatial coincidence index to obtain the multi-target action synchronization detection result, or it can splice the phase synchronization index and the spatial coincidence index and perform a weighted summation to obtain the multi-target action synchronization detection result; neither is limited here.
[0076] In the aforementioned method for detecting the synchronization of multi-target movements in a group dance performance, multi-viewpoint image acquisition devices are deployed to obtain the entire video sequence of the group dance performance. Unified spatial calibration and temporal synchronization processing are performed on the multi-viewpoint videos, establishing a stable and unified global spatiotemporal reference and eliminating spatial coordinate deviations and temporal errors between multi-viewpoint devices. By inputting the spatiotemporally synchronized camera calibration parameters and multi-viewpoint images into a dedicated skeletal coordinate extraction neural network, the changes in the three-dimensional skeletal spatial position and motion state of each dancer throughout the entire performance cycle can be fully reconstructed, accurately representing the continuous dynamic process of the dancers' movements and generating highly complete and accurate digital representation data of the dancers' movements. Based on the dancers' three-dimensional skeletal motion sequence, the motion features of each skeletal key point are extracted and fused to generate the main movement temporal signal. After signal processing and conversion into a complex analytical signal, the continuous dance movement phase sequence is calculated. Then, based on the instantaneous phase difference between each pair of dancers at each timestamp, group phase aggregation and statistics are performed to calculate the phase synchronization index, which can isolate the amplitude of movements and individual physiological factors. Non-temporal interference factors such as structural differences are eliminated, accurately depicting the temporal process and rhythmic changes of each dancer's movements throughout the entire performance cycle. The degree of temporal deviation in the movements of each dancer within the group is quantified, generating objective and quantifiable evaluation indicators for the temporal synchronization of group movements. By mapping the three-dimensional skeletal coordinate sequences of each dancer to a unified standard spatial coordinate system, and completing global registration of each dancer's skeletal coordinates through rigid body transformation and scale normalization operations, a spatially aligned three-dimensional skeletal coordinate sequence is obtained, and the spatial overlap index is calculated. The phase synchronization index and spatial overlap index are then integrated to generate multi-target movement synchronization detection results. This eliminates non-movement-related spatial deviations such as individual physiological differences, stance shifts, and different limb orientations among dancers, accurately quantifying the degree of overlap of each dancer's movements in the spatial posture dimension within the group. By combining the phase synchronization index in the temporal dimension to achieve a comprehensive two-dimensional evaluation, the overall synchronization status of multi-target movements in group dance performances can be comprehensively and objectively reflected, generating complete detection results containing both temporal and spatial information, achieving full-dimensional and comprehensive detection of the synchronization of group dance movements.
[0077] In an optional embodiment of this application, constructing a three-dimensional skeletal coordinate sequence and a three-dimensional skeletal motion sequence for each dancer based on the calibrated three-dimensional skeletal coordinates of each dancer at each timestamp may include:
[0078] Specifically, the synchronization detection terminal can construct a sequence of three-dimensional skeletal coordinates for each dancer based on the calibrated three-dimensional skeletal coordinates of each dancer at each timestamp in a chronological order.
[0079] Specifically, the synchronization detection terminal can perform temporal difference operations on the calibrated three-dimensional skeletal coordinates of each time stamp in the three-dimensional skeletal coordinate sequence to calculate the three-dimensional skeletal point motion vectors of each skeletal key point of each dancer at each time stamp.
[0080] Specifically, the synchronization detection terminal can stitch together the three-dimensional skeletal motion vectors of the key points of the same skeletal structure of the same dancer at the same time stamp to obtain the three-dimensional skeletal motion matrix of each dancer at each time stamp.
[0081] Specifically, the synchronization detection terminal can construct the three-dimensional skeletal motion sequence of each dancer according to the three-dimensional skeletal motion matrix of each dancer at each time stamp in a temporal sequence.
[0082] In an optional embodiment of this application, extracting the dance movement phase sequence based on the three-dimensional skeletal motion sequence may include:
[0083] Specifically, the synchronization detection terminal can weightedly fuse the three-dimensional skeletal motion matrix in the three-dimensional skeletal motion sequence to obtain the action timing sequence.
[0084] Specifically, the synchronization detection terminal can perform an orthogonal transformation on the action time sequence using the Hilbert transform to generate a conjugate action time sequence that is completely orthogonal to the original action time sequence. Using the original action time sequence as the real part and the conjugate action time sequence as the imaginary part, a complex analytic action time sequence is constructed.
[0085] Specifically, the synchronization detection terminal can solve the instantaneous phase of each timestamp in the complex analytical action time sequence operator, and construct the dance action phase sequence of each dancer according to the instantaneous phase of each timestamp in time sequence.
[0086] In an optional embodiment of this application, the expression for the phase synchronization index can be:
[0087]
[0088] In the formula, For phase synchronization index, For the first The timestamp of the first The dancer and the first The instantaneous phase difference of movements between individual dancers Solving for the modulus operator, The total number of timestamps. The total number of dancers.
[0089] In an optional embodiment of this application, inputting camera calibration parameters and multi-view images at each time stamp in a standard multi-view video into a skeletal coordinate extraction neural network to obtain the calibrated three-dimensional skeletal coordinates of each dancer at each time stamp may include:
[0090] Specifically, the synchronization detection terminal can input the multi-view images of each time stamp in the standard multi-view video into the two-dimensional skeletal coordinate extraction module in the skeletal coordinate extraction neural network to extract the two-dimensional skeletal coordinate set of each dancer in each multi-view image.
[0091] Specifically, the synchronization detection terminal can reconstruct the three-dimensional skeletal coordinate set corresponding to the two-dimensional skeletal coordinate set of each dancer in each multi-view image based on the two-dimensional skeletal coordinate set of each dancer in each multi-view image, combined with camera calibration parameters.
[0092] Specifically, the synchronization detection terminal can match the three-dimensional skeletal coordinate sets of each dancer at each time stamp in the standard multi-view video based on the spatial distance between the three-dimensional skeletal coordinate sets in each multi-view image.
[0093] Specifically, the synchronization detection terminal can input the camera calibration parameters and the set of three-dimensional skeletal coordinates of each dancer at each time stamp into the three-dimensional skeletal coordinate fusion calibration module in the skeletal coordinate extraction neural network to obtain the calibrated three-dimensional skeletal coordinates of each dancer at each time stamp.
[0094] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0095] Based on the same inventive concept, this application also provides a system for detecting the synchronization of multiple targets' movements in a group dance performance, used to implement the aforementioned method for detecting the synchronization of multiple targets' movements in a group dance performance. The solution provided by this system is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the system for detecting the synchronization of multiple targets' movements in a group dance performance provided below can be found in the limitations of the method for detecting the synchronization of multiple targets' movements in a group dance performance described above, and will not be repeated here.
[0096] In one exemplary embodiment, such as Figure 2 As shown, a multi-target motion synchronization detection system 200 is provided for group dance performances, including:
[0097] The performance video acquisition module 201 can be used to acquire multi-view videos of group dance performances, and perform spatial calibration and temporal synchronization on the multi-view videos of group dance performances to obtain standard multi-view videos and camera calibration parameters.
[0098] The skeleton coordinate extraction module 202 can be used to input the camera calibration parameters and the multi-view images of each time stamp in the standard multi-view video into the skeleton coordinate extraction neural network to obtain the calibrated three-dimensional skeleton coordinates of each dancer at each time stamp, and construct the three-dimensional skeleton coordinate sequence and three-dimensional skeleton motion sequence of each dancer based on the calibrated three-dimensional skeleton coordinates of each dancer at each time stamp.
[0099] The phase synchronization index calculation module 203 can be used to extract the phase sequence of dance movements based on the three-dimensional skeletal motion sequence, and calculate the phase synchronization index based on the instantaneous phase difference between each pair of dancers' dance movement phase sequences at each time stamp.
[0100] The spatial overlap index calculation module 204 can be used to map the three-dimensional skeletal coordinate sequence of each dancer to a standard spatial coordinate system, register the three-dimensional skeletal coordinate sequence of each dancer, obtain the spatially aligned three-dimensional skeletal coordinate sequence of each dancer, and calculate the spatial overlap index based on the spatially aligned three-dimensional skeletal coordinate sequence of each dancer.
[0101] The detection result generation module 205 can be used to integrate the phase synchronization index and the spatial coincidence index to generate multi-target action synchronization detection results.
[0102] In an optional embodiment of this application, the phase synchronization index calculation module 203 can also be used for:
[0103] The three-dimensional skeletal coordinate sequence of each dancer is constructed by chronologically constructing the calibrated three-dimensional skeletal coordinates of each dancer based on each timestamp.
[0104] The time-series difference operation is performed on the calibrated 3D skeletal coordinates of each time stamp in the 3D skeletal coordinate sequence to calculate the 3D skeletal point motion vectors of each skeletal key point of each dancer at each time stamp.
[0105] By splicing the 3D skeletal motion vectors of key points of the same dancer at the same time stamp, the 3D skeletal motion matrix of each dancer at each time stamp is obtained.
[0106] The 3D skeletal motion sequence of each dancer is constructed by chronologically constructing the 3D skeletal motion matrix of each dancer at each time stamp.
[0107] In an optional embodiment of this application, the phase synchronization index calculation module 203 can also be used for:
[0108] The three-dimensional skeletal motion matrix in the three-dimensional skeletal motion sequence is weighted and fused to obtain the action time sequence.
[0109] The action time sequence is orthogonally transformed using the Hilbert transform to generate a conjugate action time sequence that is completely orthogonal to the original action time sequence. Using the original action time sequence as the real part and the conjugate action time sequence as the imaginary part, a complex analytic action time sequence is constructed.
[0110] Solve the instantaneous phase of each timestamp in the complex analytic motion temporal sequence operator, and construct the dance motion phase sequence of each dancer according to the instantaneous phase of each timestamp in temporal order.
[0111] In an optional embodiment of this application, the skeletal coordinate extraction module 202 can also be used for:
[0112] The multi-view images at each time stamp in the standard multi-view video are input into the two-dimensional skeletal coordinate extraction module in the skeletal coordinate extraction neural network to extract the two-dimensional skeletal coordinate set of each dancer in each multi-view image.
[0113] Based on the two-dimensional skeletal coordinate sets of each dancer in each multi-view image, and combined with camera calibration parameters, the three-dimensional skeletal coordinate sets corresponding to the two-dimensional skeletal coordinate sets of each dancer in each multi-view image are reconstructed.
[0114] Based on the spatial distance between the 3D skeletal coordinate sets in each multi-view image, the 3D skeletal coordinate sets of each dancer at each time stamp in the standard multi-view video are matched.
[0115] The camera calibration parameters and the set of 3D skeletal coordinates of each dancer at each time stamp are input into the 3D skeletal coordinate fusion and calibration module in the skeletal coordinate extraction neural network to obtain the calibrated 3D skeletal coordinates of each dancer at each time stamp.
[0116] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the method for detecting the synchronization of multiple target movements in a group dance performance as described above.
[0117] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0118] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0119] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A method for detecting the synchronization of multi-target movements in a group dance performance, characterized in that, The method includes: Acquire multi-view video of a group dance performance, and perform spatial calibration and temporal synchronization on the multi-view video of the group dance performance to obtain standard multi-view video and camera calibration parameters; The camera calibration parameters and the multi-view images of each time stamp in the standard multi-view video are input into the skeletal coordinate extraction neural network to obtain the calibrated three-dimensional skeletal coordinates of each dancer at each time stamp. Based on the calibrated three-dimensional skeletal coordinates of each dancer at each time stamp, a three-dimensional skeletal coordinate sequence and a three-dimensional skeletal motion sequence of each dancer are constructed. Based on the three-dimensional skeletal motion sequence, a dance movement phase sequence is extracted, and based on the instantaneous phase difference between each pair of dancers in the dance movement phase sequence at each timestamp, a phase synchronization index is calculated. The three-dimensional skeletal coordinate sequences of each dancer are mapped to a standard spatial coordinate system, and the three-dimensional skeletal coordinate sequences of each dancer are registered to obtain the spatially aligned three-dimensional skeletal coordinate sequences of each dancer. Based on the spatially aligned three-dimensional skeletal coordinate sequences of each dancer, the spatial overlap index is calculated. The phase synchronization index and the spatial coincidence index are integrated to generate multi-target action synchronization detection results.
2. The method according to claim 1, characterized in that, The process of constructing a three-dimensional skeletal coordinate sequence and a three-dimensional skeletal motion sequence for each dancer based on the calibrated three-dimensional skeletal coordinates of each dancer according to each timestamp includes: The three-dimensional skeletal coordinate sequence of each dancer is constructed according to the calibrated three-dimensional skeletal coordinates of each dancer based on the timestamps in chronological order; A temporal difference operation is performed on the calibrated three-dimensional skeletal coordinates of each time stamp in the three-dimensional skeletal coordinate sequence to calculate the three-dimensional skeletal point motion vectors of each skeletal key point of each dancer at each time stamp. By concatenating the three-dimensional skeletal motion vectors of the skeletal key points of the same dancer at the same timestamp, a three-dimensional skeletal motion matrix of each dancer at each timestamp is obtained. The three-dimensional skeletal motion sequence of each dancer is constructed according to the three-dimensional skeletal motion matrix of each dancer based on the timestamp.
3. The method according to claim 2, characterized in that, The extraction of the dance movement phase sequence based on the three-dimensional skeletal motion sequence includes: The three-dimensional skeleton motion matrix in the three-dimensional skeleton motion sequence is weighted and fused to obtain the action timing sequence; The action time sequence is orthogonally transformed by Hilbert transform to generate a conjugate action time sequence that is completely orthogonal to the action time sequence; the action time sequence is used as the real part and the conjugate action time sequence is used as the imaginary part to construct a complex analytic action time sequence. Solve for the instantaneous phase of each timestamp in the complex analytical action temporal sequence operator, and construct the dance action phase sequence of each dancer according to the instantaneous phase of each timestamp in temporal order.
4. The method according to claim 3, characterized in that, The expression for the phase synchronization index is: In the formula, The phase synchronization index is... For the first The timestamp of the first The dancer and the first The instantaneous phase difference of movements between individual dancers Solving for the modulus operator, The total number of timestamps. The total number of dancers mentioned.
5. The method according to any one of claims 1 to 4, characterized in that, The step of inputting the camera calibration parameters and the multi-view images at each time stamp in the standard multi-view video into the skeletal coordinate extraction neural network to obtain the calibrated three-dimensional skeletal coordinates of each dancer at each time stamp includes: The multi-view images of each time stamp in the standard multi-view video are input into the two-dimensional skeletal coordinate extraction module in the skeletal coordinate extraction neural network to extract the two-dimensional skeletal coordinate set of each dancer in each multi-view image. Based on the two-dimensional skeletal coordinate sets of each dancer in each of the multi-view images, and in conjunction with the camera calibration parameters, a three-dimensional skeletal coordinate set corresponding to the two-dimensional skeletal coordinate sets of each dancer in each of the multi-view images is reconstructed. Based on the spatial distance between the three-dimensional skeletal coordinate sets in each of the multi-view images, the three-dimensional skeletal coordinate sets of each dancer in each of the timestamps in the standard multi-view video are matched to obtain the three-dimensional skeletal coordinate sets of each dancer. The camera calibration parameters and the set of three-dimensional skeletal coordinates of each dancer at each time stamp are input into the three-dimensional skeletal coordinate fusion calibration module in the skeletal coordinate extraction neural network to obtain the calibrated three-dimensional skeletal coordinates of each dancer at each time stamp.
6. A system for detecting the synchronization of multi-target movements in a group dance performance, characterized in that, The system includes: The performance video acquisition module is used to acquire multi-view videos of group dance performances, and to perform spatial calibration and temporal synchronization on the multi-view videos of group dance performances to obtain standard multi-view videos and camera calibration parameters; The skeleton coordinate extraction module is used to input the camera calibration parameters and the multi-view images of each time stamp in the standard multi-view video into the skeleton coordinate extraction neural network to obtain the calibrated three-dimensional skeleton coordinates of each dancer at each time stamp, and to construct the three-dimensional skeleton coordinate sequence and three-dimensional skeleton motion sequence of each dancer based on the calibrated three-dimensional skeleton coordinates of each dancer at each time stamp. The phase synchronization index calculation module is used to extract the dance movement phase sequence based on the three-dimensional skeletal motion sequence, and calculate the phase synchronization index based on the instantaneous phase difference of the dance movement phase sequence between each pair of the dancers at each timestamp. The spatial overlap index calculation module is used to map the three-dimensional skeletal coordinate sequences of each dancer to a standard spatial coordinate system, register the three-dimensional skeletal coordinate sequences of each dancer, obtain the spatially aligned three-dimensional skeletal coordinate sequences of each dancer, and calculate the spatial overlap index based on the spatially aligned three-dimensional skeletal coordinate sequences of each dancer. The detection result generation module is used to integrate the phase synchronization index and the spatial coincidence index to generate multi-target action synchronization detection results.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.