Dancing movement space-time integration evaluation method and system based on monocular camera shooting
Through monocular photography and audio acquisition technology, dance action images are processed in frames and key points are identified, and timing integration is combined with audio data, which solves the problems of large labor consumption and subjective influence in existing dance evaluations, and achieves efficient and objective dance evaluation.
Patent Information
- Application Number
- CN202510346584.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-08-15
AI Technical Summary
The existing dance evaluation methods rely on manual scoring, which leads to large workload, high physical consumption and susceptible to subjective factors, and cannot be scored using unified standards, which reduces the efficiency and credibility of the evaluation.
Continuous shooting and audio acquisition are performed through monocular imaging, dance motion images are processed in frames, key points of the human body are identified, audio data is combined for timing integration, the differences with standard movements are compared, and unified standards are used for scoring.
It achieves efficient, objective and credible scoring of dance evaluation, reduces manpower consumption, and improves the uniformity and accuracy of evaluation.
Smart Images

Figure CN120496162A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent visual recognition, and in particular to a method and system for temporal and spatial integration evaluation of dance movements based on monocular photography. Background Art
[0002] Dance, as a sport that involves the entire body, requires not only the correctness of body movements but also the presentation of corresponding dynamic beauty. Competitive events such as dance competitions usually require the scoring of the entire dance process of the contestants. In order to ensure the objectivity and fairness of the scoring, multiple judges will score the entire dance process at the same time to generate the final dance score result. The above-mentioned dance evaluation method relies heavily on professionals with dance-related knowledge to score, and the judges are required to give detailed and independent scores to each dance move during the scoring process, which requires the judges to be in a highly concentrated state throughout the dance process, increasing the human workload and physical exertion of dance scoring. In addition, the above-mentioned dance evaluation method is inevitably affected by subjective factors, and it is impossible to use a unified standard to score all contestants, which reduces the efficiency and credibility of dance evaluation. Summary of the Invention
[0003] The object of the present invention is to provide a dance movement spatiotemporal integration evaluation method and system based on monocular camera, which performs continuous monocular shooting and audio collection on the dance process of a target object to obtain corresponding dance movement dynamic images and background audio data, and performs frame processing on the dance movement dynamic images to facilitate subsequent refined and independent evaluation of the target object's dance movements; performs human body key point recognition and spatial integration processing on all dance movement image frames to obtain human body key point posture data, and based on the background audio data, performs temporal integration processing on the human body key point posture data corresponding to all dance movement image frames to obtain human body key point full-process posture change information, and effectively refines the dance movement characteristics of the target object; furthermore, compares the human body key point full-process posture change information with the standard dance movement posture change information to obtain dance movement difference information, thereby evaluating the real-time dance of the target object, obtaining a dance movement scoring result, adopting a unified standard for dance scoring, and improving the efficiency, objectivity and credibility of dance evaluation.
[0004] The present invention is achieved through the following technical solutions:
[0005] The method for evaluating the spatiotemporal integration of dance movements based on monocular videography includes:
[0006] Continuously monocularly photographing and collecting audio of the target subject's dance process to obtain corresponding dance movement dynamic images and background audio data; performing frame processing on the dance movement dynamic images to obtain a plurality of dance movement image frames, and preprocessing all the dance movement image frames;
[0007] Performing human body key point recognition and spatial integration processing on all dance action image frames to obtain human body key point posture data of the target object; performing temporal integration processing on the human body key point posture data corresponding to all dance action image frames based on the background audio data to obtain the full-process human body key point posture change information of the target object during the dance process;
[0008] The whole-process posture change information of the key points of the human body is compared with the posture change information of the standard dance movement to obtain the dance movement difference information of the target object; based on the dance movement difference information, the actual dance movement of the target object is evaluated to obtain the corresponding dance movement scoring result.
[0009] Optionally, the target subject's dance process is continuously photographed and audio is collected monocularly to obtain corresponding dance movement dynamic images and background audio data; the dance movement dynamic images are frame-processed to obtain a plurality of dance movement image frames, and all dance movement image frames are pre-processed, including:
[0010] The target subject's dance process is synchronously and continuously photographed and audio is collected by a monocular camera to obtain corresponding dance movement images and background audio data; the background audio data is compared with the background music played during the dance process, thereby converting the playback time axis of the background music into the collection time axis of the background audio data;
[0011] The dance action dynamic image is subjected to frame processing to obtain a plurality of dance action image frames, and all the dance action image frames are subjected to pixel edge sharpening preprocessing and image noise removal preprocessing.
[0012] Optionally, after continuously monocularly photographing and collecting audio of the target subject's dance process to obtain corresponding dance movement dynamic images and background audio data, the method further includes:
[0013] Step S1, using the following formula (1), based on the dance movement dynamic images and background audio data obtained by continuous monocular shooting and audio collection, dynamic features of the dance movement are extracted.
[0014]
[0015] In the above formula (1), D represents the dynamic characteristic value of the dance movement; T represents the total number of image frames contained in the dance movement dynamic image; I(t) represents the image matrix data of the t-th image frame; I(t-1) represents the image matrix data of the t-1-th image frame; W(t) represents the weight function of the image matrix data of the t-th image frame; β represents the audio influence factor; S(t) represents the instantaneous value of the audio signal strength; γ represents the adjustment factor; R(t) represents the current environmental noise intensity; R th represents the noise threshold;
[0016] Step S2, using the following formula (2), according to the dynamic feature value of the dance movement, obtain the audio-visual fusion feature value of the dance movement,
[0017]
[0018] In the above formula (2), F represents the audio-visual fusion feature value of the dance action; A(t) represents the instantaneous volume intensity value of the corresponding audio data; α(t) represents the time-dependent weight factor; δ(t) represents the time attenuation factor;
[0019] Step S3, using the following formula (3), based on the dynamic feature value of the dance movement and the audio-visual fusion feature value, a dynamic control feedback adjustment signal is generated.
[0020]
[0021] In the above formula (3), C represents the dynamic control feedback adjustment signal; k represents the control gain; N represents the complexity of the dance movement; M represents the preset maximum movement complexity; F0 represents the preset audio-visual benchmark feature value of the dance movement; and λ represents the adjustment factor for controlling the impact of environmental noise on the feedback signal.
[0022] Optionally, all dance action image frames are subjected to human body key point recognition and spatial integration processing to obtain human body key point posture data of the target object; based on the background audio data, the human body key point posture data corresponding to all dance action image frames are subjected to temporal integration processing to obtain the whole-process human body key point posture change information of the target object during the dance process, including:
[0023] Performing human body contour recognition processing on each dance action image frame to obtain head contour information and limb contour information corresponding to the target object in each dance action image frame; determining the contour shape and depth feature information of the head and limbs of the target object based on the head contour feature information and the limb contour information; then selecting a plurality of head key points and a plurality of limb key points from the head and limbs respectively based on the contour shape and depth feature information; and integrating the position information of all the head key points and all the limb key points during the dance process into the same three-dimensional spatial coordinate system to obtain the human body key point posture data of the target object during the dance process;
[0024] Based on the generated clock signal relationship between the background audio data and the dance movement dynamic image, the acquisition time axis of the background audio data is converted to generate a dance movement generation time axis corresponding to the dance movement dynamic image; based on the dance movement generation time axis, the human body key point posture data corresponding to all dance movement image frames are time-series integrated to obtain the full-process posture change information of the target object's human body key points during the dance process.
[0025] Optionally, the whole-process posture change information of the key points of the human body is compared with the posture change information of the standard dance movement to obtain dance movement difference information of the target object; based on the dance movement difference information, the actual dance movement of the target object is evaluated to obtain a corresponding dance movement scoring result, including:
[0026] The whole-process posture change information of the key points of the human body is divided into the whole-process posture change information of the key points of the head and the whole-process posture change information of the key points of the limbs, and the whole-process posture change information of the key points of the head and the whole-process posture change information of the key points of the limbs are respectively compared with the posture change information of the standard dance movements to obtain the head movement posture deviation information and the limb movement posture deviation information of the target object during the dance process; wherein the head movement posture deviation information includes the posture angle deviation value and the displacement deviation value of the head movement compared with the head sub-movement in the standard dance movement; the limb movement posture deviation information includes the posture angle deviation value and the displacement deviation value of the limb movement compared with the limb sub-movement in the standard dance movement;
[0027] Based on the head movement posture deviation information and the limb movement posture deviation information, the actual dance movement of the target object is evaluated for the coordination degree of the head and limb movements and the cumulative degree of movement posture deviations to obtain a corresponding dance movement scoring result.
[0028] The dance movement spatiotemporal integration evaluation system based on monocular camera includes:
[0029] The dance process acquisition module is used to continuously capture the target object's dance process with a single camera and collect audio data to obtain the corresponding dance movement dynamic images and background audio data;
[0030] a dance image processing module, configured to perform frame processing on the dance action dynamic image to obtain a plurality of dance action image frames, and perform pre-processing on all dance action image frames;
[0031] A key point posture data generation module is used to perform human key point recognition and spatial integration processing on all dance action image frames to obtain human key point posture data of the target object;
[0032] A key point posture data time series integration module is used to perform time series integration processing on the human body key point posture data corresponding to all dance action image frames based on the background audio data, so as to obtain the whole-process posture change information of the human body key points of the target object during the dance process;
[0033] a dance movement difference determination module, configured to compare the whole-process posture change information of the key points of the human body with the posture change information of the standard dance movement to obtain dance movement difference information of the target object;
[0034] The dance movement scoring module is used to evaluate the actual dance movement of the target object based on the dance movement difference information to obtain a corresponding dance movement scoring result.
[0035] Optionally, the dance process acquisition module is used to continuously monocularly capture and capture audio of the target object's dance process to obtain corresponding dance movement dynamic images and background audio data, including:
[0036] The target subject's dance process is synchronously and continuously photographed and audio is collected by a monocular camera to obtain corresponding dance movement images and background audio data; the background audio data is compared with the background music played during the dance process, thereby converting the playback time axis of the background music into the collection time axis of the background audio data;
[0037] The dance image processing module performs frame processing on the dance action dynamic image to obtain a plurality of dance action image frames, and pre-processes all the dance action image frames, including:
[0038] The dance action dynamic image is subjected to frame processing to obtain a plurality of dance action image frames, and all the dance action image frames are subjected to pixel edge sharpening preprocessing and image noise removal preprocessing.
[0039] Optionally, the key point posture data generation module is used to perform human key point recognition and spatial integration processing on all dance action image frames to obtain human key point posture data of the target object, including:
[0040] Performing human body contour recognition processing on each dance action image frame to obtain head contour information and limb contour information corresponding to the target object in each dance action image frame; determining the contour shape and depth feature information of the head and limbs of the target object based on the head contour feature information and the limb contour information; then selecting a plurality of head key points and a plurality of limb key points from the head and limbs respectively based on the contour shape and depth feature information; and integrating the position information of all the head key points and all the limb key points during the dance process into the same three-dimensional spatial coordinate system to obtain the human body key point posture data of the target object during the dance process;
[0041] The key point posture data time series integration module is used to perform time series integration processing on the human body key point posture data corresponding to all dance action image frames based on the background audio data to obtain the full-process posture change information of the target object's human body key points during the dance process, including:
[0042] Based on the generated clock signal relationship between the background audio data and the dance movement dynamic image, the acquisition time axis of the background audio data is converted to generate a dance movement generation time axis corresponding to the dance movement dynamic image; based on the dance movement generation time axis, the human body key point posture data corresponding to all dance movement image frames are time-series integrated to obtain the full-process posture change information of the target object's human body key points during the dance process.
[0043] Optionally, the dance movement difference determination module is configured to compare the whole-process posture change information of the human body key points with the standard dance movement posture change information to obtain the dance movement difference information of the target object, including:
[0044] The whole-process posture change information of the key points of the human body is divided into the whole-process posture change information of the key points of the head and the whole-process posture change information of the key points of the limbs, and the whole-process posture change information of the key points of the head and the whole-process posture change information of the key points of the limbs are respectively compared with the posture change information of the standard dance movements to obtain the head movement posture deviation information and the limb movement posture deviation information of the target object during the dance process; wherein the head movement posture deviation information includes the posture angle deviation value and the displacement deviation value of the head movement compared with the head sub-movement in the standard dance movement; the limb movement posture deviation information includes the posture angle deviation value and the displacement deviation value of the limb movement compared with the limb sub-movement in the standard dance movement;
[0045] The dance movement scoring module is used to evaluate the actual dance movement of the target object based on the dance movement difference information to obtain a corresponding dance movement scoring result, including:
[0046] Based on the head movement posture deviation information and the limb movement posture deviation information, the actual dance movement of the target object is evaluated for the coordination degree of the head and limb movements and the cumulative degree of movement posture deviations to obtain a corresponding dance movement scoring result.
[0047] Compared with the prior art, the present invention has the following beneficial effects:
[0048] The dance movement spatiotemporal integration evaluation method and system based on monocular camera provided in the present application performs continuous monocular shooting and audio collection on the dance process of the target object to obtain corresponding dance movement dynamic images and background audio data, and performs frame processing on the dance movement dynamic images to facilitate subsequent refined and independent evaluation of the target object's dance movements; human body key points are identified and spatially integrated on all dance movement image frames to obtain human body key point posture data, and based on the background audio data, time-series integration processing is performed on the human body key point posture data corresponding to all dance movement image frames to obtain human body key point posture change information throughout the process, and effectively refine the dance movement characteristics of the target object; the human body key point posture change information throughout the process is also compared with the standard dance movement posture change information to obtain dance movement difference information, thereby evaluating the real-time dance of the target object and obtaining a dance movement scoring result, and using a unified standard for dance scoring to improve the efficiency, objectivity and credibility of dance evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. Among them:
[0050] Figure 1 This is a flow chart of the method for temporal and spatial integration evaluation of dance movements based on monocular photography provided by the present invention.
[0051] Figure 2 This is a schematic diagram of the structure of the monocular camera-based dance movement spatiotemporal integration evaluation system provided by the present invention. DETAILED DESCRIPTION
[0052] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are described in detail below in conjunction with the accompanying drawings. It will be understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only some, rather than all, structures related to the present application are shown in the accompanying drawings. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0053] As used herein, the terms "comprise," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0054] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0055] See also Figure 1 As shown, an embodiment of the present application provides a method for evaluating dance movement spatiotemporal integration based on monocular photography. The method for evaluating dance movement spatiotemporal integration based on monocular photography includes:
[0056] Continuously monocularly photograph and collect audio of the target subject's dance process to obtain corresponding dance movement dynamic images and background audio data; frame-process the dance movement dynamic images to obtain several dance movement image frames, and pre-process all dance movement image frames;
[0057] All dance action image frames are subjected to human key point recognition and spatial integration processing to obtain the human key point posture data of the target object; based on the background audio data, the human key point posture data corresponding to all dance action image frames are subjected to temporal integration processing to obtain the human key point posture change information of the target object during the dance process;
[0058] The whole-process posture change information of the key points of the human body is compared with the posture change information of the standard dance movement to obtain the dance movement difference information of the target object; based on the dance movement difference information, the actual dance movement of the target object is evaluated to obtain the corresponding dance movement score result.
[0059] The beneficial effects of the above embodiments are as follows: the dance movement spatiotemporal integration evaluation method based on monocular camera performs continuous monocular shooting and audio collection on the dance process of the target object to obtain corresponding dance movement dynamic images and background audio data, and performs frame processing on the dance movement dynamic images to facilitate subsequent refined and independent evaluation of the target object's dance movements; human body key point recognition and spatial integration processing are performed on all dance movement image frames to obtain human body key point posture data, and based on the background audio data, time series integration processing is performed on the human body key point posture data corresponding to all dance movement image frames to obtain human body key point posture change information throughout the process, and effectively refine the dance movement characteristics of the target object; the human body key point posture change information throughout the process is also compared with the standard dance movement posture change information to obtain dance movement difference information, thereby evaluating the real-time dance of the target object and obtaining a dance movement scoring result, using a unified standard for dance scoring to improve the efficiency, objectivity and credibility of dance evaluation.
[0060] In another embodiment, a target subject's dance process is continuously photographed and audio is collected by a monocular camera to obtain a corresponding dance motion dynamic image and background audio data; the dance motion dynamic image is frame-processed to obtain a plurality of dance motion image frames, and all dance motion image frames are pre-processed, including:
[0061] The target subject's dance process is synchronously and continuously photographed and audio is collected by a monocular camera to obtain corresponding dance movement images and background audio data; the background audio data is compared with the background music played during the dance process, thereby converting the playback timeline of the background music into the collection timeline of the background audio data;
[0062] The dance action dynamic image is frame-processed to obtain a number of dance action image frames, and all the dance action image frames are pre-processed with pixel edge sharpening and picture noise removal.
[0063] The beneficial effect of the above embodiment is that the target object will play the accompaniment music synchronously during the dance process. Ideally, the dance movements made by the target object should match the accompaniment music, so the playback time attribute of the accompaniment music should be consistent with the execution time attribute of the dance movement. In the actual dance evaluation, the entire dance process of the target object is synchronized with continuous monocular shooting and audio collection, that is, a monocular camera and a microphone array are used to perform linked synchronous operations to achieve continuous monocular shooting and audio collection, and obtain corresponding dance movement images and background audio data (i.e., audio data of the accompaniment music). The background audio data is compared with the background music (i.e., accompaniment music) played corresponding to the dance process, so that the playback time axis of the background music is converted into the collection time axis of the background audio data. In this way, the collection timing state of the background audio data can be calibrated, and the time axis comparison information is provided for the subsequent timing integration processing of the key point posture data of the human body. The dance movement dynamic image is also frame-processed to obtain several dance movement image frames, and all dance movement image frames are pre-processed with pixel edge sharpening and image noise removal. This can enable individual and detailed recognition of each dance movement made by the target object during the entire dance process and improve the recognition accuracy of subsequent dance movement features.
[0064] In another embodiment, after continuously monocularly photographing and collecting audio of the target subject's dance process to obtain corresponding dance movement dynamic images and background audio data, the method further includes:
[0065] Step S1, using the following formula (1), based on the dance movement dynamic images and background audio data obtained by continuous monocular shooting and audio collection, dynamic features of the dance movement are extracted.
[0066]
[0067] In the above formula (1), D represents the dynamic feature value of the dance movement; T represents the total number of image frames contained in the dance movement dynamic image; I(t) represents the image matrix data of the t-th image frame; I(t-1) represents the image matrix data of the t-1-th image frame; W(t) represents the weight function of the image matrix data of the t-th image frame, which represents the importance of the image in the dynamic feature extraction; β represents the audio influence factor, which represents the contribution of audio to the dynamic feature; S(t) represents the instantaneous value of the audio signal strength; γ represents the adjustment factor; R(t) represents the current environmental noise intensity; R th represents the noise threshold, which corresponds to the ambient noise standard for optimal dance performance;
[0068] Step S2, using the following formula (2), according to the dynamic feature value of the dance movement, obtain the audio-visual fusion feature value of the dance movement,
[0069]
[0070] In the above formula (2), F represents the audio-visual fusion feature value of the dance action; A(t) represents the instantaneous volume intensity value of the corresponding audio data; α(t) represents the time-dependent weight factor, whose magnitude changes with time; δ(t) represents the time attenuation factor, whose magnitude gradually decreases with time;
[0071] Step S3, using the following formula (3), based on the dynamic feature value of the dance movement and the audio-visual fusion feature value, a dynamic control feedback adjustment signal is generated.
[0072]
[0073] In the above formula (3), C represents the dynamic control feedback adjustment signal, which is used to adjust the dance movement performance; k represents the control gain, which characterizes the strength of the dynamic control feedback adjustment signal; N represents the complexity of the dance movement; M represents the preset maximum movement complexity, which is used as the normalization standard; F0 represents the preset audio-visual benchmark feature value of the dance movement; λ represents the adjustment factor for controlling the influence of environmental noise on the feedback signal.
[0074] Dynamic control feedback adjustment signal C is used to adjust the dancer's performance in real time. For example, if F is lower than the target F0, the dancer is prompted to speed up the rhythm or enhance the expression through sound or visual signals, taking into account the environmental noise and movement complexity to achieve more accurate feedback.
[0075] The beneficial effects of the above embodiment are as follows: using the above formula (1), based on the dance movement dynamic image and background audio data obtained by continuous monocular shooting and audio collection, the dynamic feature extraction of the dance movement is performed, thereby calculating the dynamic feature value of each frame of the dance movement and considering the weighted average of multiple factors to ensure the accuracy of the calculation; then using the above formula (2), based on the dynamic feature value of the dance movement, the audio-visual fusion feature value of the dance movement is obtained, thereby combining the dynamic feature of the dance movement and the audio feature to calculate the fusion feature value, and introducing a time attenuation factor. By introducing multiple variables and time dynamic factors, the accuracy and flexibility of the dance movement analysis and optimization are enhanced; finally, using the above formula (3), based on the dynamic feature value of the dance movement and the audio-visual fusion feature value, a dynamic control feedback adjustment signal is generated, and then the dance performance is adjusted in real time through the feedback mechanism of the audio-visual feature. Combined with environmental factors, the dynamic adjustment can more comprehensively reflect the various factors in the dance performance and improve the performance quality.
[0076] In another embodiment, all dance action image frames are subjected to human body key point recognition and spatial integration processing to obtain human body key point posture data of the target object; based on the background audio data, the human body key point posture data corresponding to all dance action image frames are subjected to temporal integration processing to obtain the full-process posture change information of the human body key points of the target object during the dance process, including:
[0077] Performing human body contour recognition processing on each dance action image frame to obtain head contour information and limb contour information corresponding to the target object in each dance action image frame; determining the contour shape and depth feature information of the target object's head and limbs based on the head contour feature information and the limb contour information; then selecting a plurality of head key points and a plurality of limb key points from the head and limbs, respectively, based on the contour shape and depth feature information; and integrating the position information of all the head key points and all the limb key points during the dance process into the same three-dimensional spatial coordinate system to obtain the human body key point posture data of the target object during the dance process;
[0078] Based on the generated clock signal relationship between the background audio data and the dance movement dynamic image, the acquisition time axis of the background audio data is converted to generate a dance movement generation time axis corresponding to the dance movement dynamic image; based on the dance movement generation time axis, the human body key point posture data corresponding to all dance movement image frames are time-series integrated to obtain the full-process posture change information of the target object's human body key points during the dance process.
[0079] The beneficial effect of the above embodiment is that every time the target object makes a dance move during the entire dance process, the head and limbs of the target object will present corresponding posture angles and displacement amplitudes. In order to accurately and comprehensively evaluate the dance moves of the target object, it is necessary to conduct a linkage evaluation of the movement forms of the head and limbs of the target object during the entire dance process. To this end, human body contour recognition processing is performed on each dance move image frame to obtain the head contour information and limb contour information corresponding to the target object in each dance move image frame, thereby determining the contour shape and depth feature information of the head and limbs of the target object. If the action recognition is performed directly on the entire head and limbs of the target object, it is necessary to use the external contours of the entire head and limbs as a reference for recognition, which not only requires a large amount of computational workload, but also increases the difficulty of evaluating the subsequent action posture angles and displacement amplitudes. Based on the contour shape and depth feature information, a number of head key points and a number of limb key points are selected from the head and limbs respectively. With these head key points and limb key points as a reference, the contours of the head and limbs of the target object are simplified and represented, effectively reducing the subsequent computational workload. The positional information of all key points of the head and all key points of the limbs during the dance process is integrated into the same three-dimensional spatial coordinate system to obtain the human body key point posture data of the target object during the dance process. This allows the position angles and displacement amplitudes of the target object's head and limbs in the same dance movement to be unified into a unified space, achieving a comprehensive representation of the dance movement form corresponding to each dance movement image. Furthermore, based on the generated clock signal relationship between the background audio data and the dance movement dynamic image, the acquisition time axis of the background audio data is converted to generate the dance movement generation time axis corresponding to the dance movement dynamic image. This allows the human body key point posture data corresponding to all dance movement image frames to be time-series integrated and processed to obtain the full-process posture change information of the target object's human body key points during the dance process, thereby representing the continuous change state of the target object's head and limb movement posture throughout the entire dance process.
[0080] In another embodiment, the whole-process posture change information of the key points of the human body is compared with the posture change information of the standard dance movement to obtain dance movement difference information of the target object; based on the dance movement difference information, the actual dance movement of the target object is evaluated to obtain a corresponding dance movement scoring result, including:
[0081] The whole-process posture change information of the key points of the human body is divided into the whole-process posture change information of the key points of the head and the whole-process posture change information of the key points of the limbs, and the whole-process posture change information of the key points of the head and the whole-process posture change information of the key points of the limbs are respectively compared with the posture change information of the standard dance movements to obtain the head movement posture deviation information and the limb movement posture deviation information of the target object during the dance process; wherein the head movement posture deviation information includes the posture angle deviation value and the displacement deviation value of the head movement compared with the head sub-movement in the standard dance movement; the limb movement posture deviation information includes the posture angle deviation value and the displacement deviation value of the limb movement compared with the limb sub-movement in the standard dance movement;
[0082] Based on the head movement posture deviation information and the limb movement posture deviation information, the actual dance movement of the target object is evaluated for the degree of coordination between the head and limb movements and the cumulative degree of movement posture deviations to obtain a corresponding dance movement scoring result.
[0083] The beneficial effects of the above embodiment are as follows: the entire posture change information of the key points of the human body is divided into the entire posture change information of the key points of the head and the entire posture change information of the key points of the limbs, and the entire posture change information of the key points of the head and the entire posture change information of the key points of the limbs are respectively compared with the posture change information of the standard dance movements to obtain the head movement posture deviation information and the limb movement posture deviation information of the target subject during the dance process. In this way, the movement differences between the actual dance movements performed by the target subject during the entire dance process and the standard dance movements in the head and limbs can be quantitatively characterized. In addition, based on the head movement posture deviation information and the limb movement posture deviation information, the actual dance movements of the target subject are evaluated for the degree of coordination between the head and limb movements and the cumulative degree of movement posture deviation, and the corresponding dance movement scoring results are obtained. The evaluation of the degree of coordination between the head and limb movements and the cumulative degree of movement posture deviation can be achieved using models such as convolutional neural network models, which are conventional technical means in this field and will not be described in detail here.
[0084] See also Figure 2 As shown, an embodiment of the present application provides a dance movement spatiotemporal integration evaluation system based on monocular camera. The dance movement spatiotemporal integration evaluation system based on monocular camera includes:
[0085] The dance process acquisition module is used to continuously capture the target object's dance process with a single camera and collect audio data to obtain the corresponding dance movement dynamic images and background audio data;
[0086] The dance image processing module is used to perform frame processing on the dance action dynamic image to obtain a number of dance action image frames, and pre-process all the dance action image frames;
[0087] The key point posture data generation module is used to perform human key point recognition and spatial integration processing on all dance action image frames to obtain the human key point posture data of the target object;
[0088] A key point posture data time series integration module is used to perform time series integration processing on the human body key point posture data corresponding to all dance action image frames based on the background audio data, and obtain the full-process posture change information of the target object's human body key points during the dance process;
[0089] A dance movement difference determination module is used to compare the whole-process posture change information of the key points of the human body with the posture change information of the standard dance movement to obtain the dance movement difference information of the target object;
[0090] The dance movement scoring module is used to evaluate the actual dance movement of the target object based on the dance movement difference information to obtain a corresponding dance movement scoring result.
[0091] The beneficial effects of the above embodiments are as follows: the dance movement spatiotemporal integration evaluation system based on monocular camera performs continuous monocular shooting and audio collection on the dance process of the target object to obtain corresponding dance movement dynamic images and background audio data, and performs frame processing on the dance movement dynamic images to facilitate subsequent refined and independent evaluation of the target object's dance movements; human body key point recognition and spatial integration processing are performed on all dance movement image frames to obtain human body key point posture data, and based on the background audio data, time series integration processing is performed on the human body key point posture data corresponding to all dance movement image frames to obtain human body key point posture change information throughout the process, and effectively refine the dance movement characteristics of the target object; the human body key point posture change information throughout the process is also compared with the standard dance movement posture change information to obtain dance movement difference information, thereby evaluating the real-time dance of the target object and obtaining a dance movement scoring result, using a unified standard for dance scoring to improve the efficiency, objectivity and credibility of dance evaluation.
[0092] In another embodiment, the dance process acquisition module is used to continuously monocularly capture and audio capture the target object's dance process to obtain corresponding dance movement dynamic images and background audio data, including:
[0093] The target subject's dance process is synchronously and continuously photographed and audio is collected by a monocular camera to obtain corresponding dance movement images and background audio data; the background audio data is compared with the background music played during the dance process, thereby converting the playback timeline of the background music into the collection timeline of the background audio data;
[0094] The dance image processing module performs frame processing on the dance action dynamic image to obtain a number of dance action image frames, and pre-processes all the dance action image frames, including:
[0095] The dance action dynamic image is frame-processed to obtain a number of dance action image frames, and all the dance action image frames are pre-processed with pixel edge sharpening and picture noise removal.
[0096] The beneficial effect of the above embodiment is that the target object will play the accompaniment music synchronously during the dance process. Ideally, the dance movements made by the target object should match the accompaniment music, so the playback time attribute of the accompaniment music should be consistent with the execution time attribute of the dance movement. In the actual dance evaluation, the entire dance process of the target object is synchronized with continuous monocular shooting and audio collection, that is, a monocular camera and a microphone array are used to perform linked synchronous operations to achieve continuous monocular shooting and audio collection, and obtain corresponding dance movement images and background audio data (i.e., audio data of the accompaniment music). The background audio data is compared with the background music (i.e., accompaniment music) played corresponding to the dance process, so that the playback time axis of the background music is converted into the collection time axis of the background audio data. In this way, the collection timing state of the background audio data can be calibrated, and the time axis comparison information is provided for the subsequent timing integration processing of the key point posture data of the human body. The dance movement dynamic image is also frame-processed to obtain several dance movement image frames, and all dance movement image frames are pre-processed with pixel edge sharpening and image noise removal. This can enable individual and detailed recognition of each dance movement made by the target object during the entire dance process and improve the recognition accuracy of subsequent dance movement features.
[0097] In another embodiment, the key point posture data generation module is used to perform human key point recognition and spatial integration processing on all dance action image frames to obtain the human key point posture data of the target object, including:
[0098] Performing human body contour recognition processing on each dance action image frame to obtain head contour information and limb contour information corresponding to the target object in each dance action image frame; determining the contour shape and depth feature information of the target object's head and limbs based on the head contour feature information and the limb contour information; then selecting a plurality of head key points and a plurality of limb key points from the head and limbs, respectively, based on the contour shape and depth feature information; and integrating the position information of all the head key points and all the limb key points during the dance process into the same three-dimensional spatial coordinate system to obtain the human body key point posture data of the target object during the dance process;
[0099] The key point posture data time series integration module is used to perform time series integration processing on the human body key point posture data corresponding to all dance action image frames based on the background audio data, and obtain the full-process posture change information of the target object's human body key points during the dance process, including:
[0100] Based on the generated clock signal relationship between the background audio data and the dance movement dynamic image, the acquisition time axis of the background audio data is converted to generate a dance movement generation time axis corresponding to the dance movement dynamic image; based on the dance movement generation time axis, the human body key point posture data corresponding to all dance movement image frames are time-series integrated to obtain the full-process posture change information of the target object's human body key points during the dance process.
[0101] The beneficial effect of the above embodiment is that every time the target object makes a dance move during the entire dance process, the head and limbs of the target object will present corresponding posture angles and displacement amplitudes. In order to accurately and comprehensively evaluate the dance moves of the target object, it is necessary to conduct a linkage evaluation of the movement forms of the head and limbs of the target object during the entire dance process. To this end, human body contour recognition processing is performed on each dance move image frame to obtain the head contour information and limb contour information corresponding to the target object in each dance move image frame, thereby determining the contour shape and depth feature information of the head and limbs of the target object. If the action recognition is performed directly on the entire head and limbs of the target object, it is necessary to use the external contours of the entire head and limbs as a reference for recognition, which not only requires a large amount of computational workload, but also increases the difficulty of evaluating the subsequent action posture angles and displacement amplitudes. Based on the contour shape and depth feature information, a number of head key points and a number of limb key points are selected from the head and limbs respectively. With these head key points and limb key points as a reference, the contours of the head and limbs of the target object are simplified and represented, effectively reducing the subsequent computational workload. The positional information of all key points of the head and all key points of the limbs during the dance process is integrated into the same three-dimensional spatial coordinate system to obtain the human body key point posture data of the target object during the dance process. This allows the position angles and displacement amplitudes of the target object's head and limbs in the same dance movement to be unified into a unified space, achieving a comprehensive representation of the dance movement form corresponding to each dance movement image. Furthermore, based on the generated clock signal relationship between the background audio data and the dance movement dynamic image, the acquisition time axis of the background audio data is converted to generate the dance movement generation time axis corresponding to the dance movement dynamic image. This allows the human body key point posture data corresponding to all dance movement image frames to be time-series integrated and processed to obtain the full-process posture change information of the target object's human body key points during the dance process, thereby representing the continuous change state of the target object's head and limb movement posture throughout the entire dance process.
[0102] In another embodiment, the dance movement difference determination module is used to compare the whole-process posture change information of the key points of the human body with the standard dance movement posture change information to obtain the dance movement difference information of the target object, including:
[0103] The whole-process posture change information of the key points of the human body is divided into the whole-process posture change information of the key points of the head and the whole-process posture change information of the key points of the limbs, and the whole-process posture change information of the key points of the head and the whole-process posture change information of the key points of the limbs are respectively compared with the posture change information of the standard dance movements to obtain the head movement posture deviation information and the limb movement posture deviation information of the target object during the dance process; wherein the head movement posture deviation information includes the posture angle deviation value and the displacement deviation value of the head movement compared with the head sub-movement in the standard dance movement; the limb movement posture deviation information includes the posture angle deviation value and the displacement deviation value of the limb movement compared with the limb sub-movement in the standard dance movement;
[0104] The dance movement scoring module is used to evaluate the actual dance movement of the target object based on the dance movement difference information to obtain a corresponding dance movement scoring result, including:
[0105] Based on the head movement posture deviation information and the limb movement posture deviation information, the actual dance movement of the target object is evaluated for the degree of coordination between the head and limb movements and the cumulative degree of movement posture deviations to obtain a corresponding dance movement scoring result.
[0106] The beneficial effects of the above embodiment are as follows: the entire posture change information of the key points of the human body is divided into the entire posture change information of the key points of the head and the entire posture change information of the key points of the limbs, and the entire posture change information of the key points of the head and the entire posture change information of the key points of the limbs are respectively compared with the posture change information of the standard dance movements to obtain the head movement posture deviation information and the limb movement posture deviation information of the target subject during the dance process. In this way, the movement differences between the actual dance movements performed by the target subject during the entire dance process and the standard dance movements in the head and limbs can be quantitatively characterized. In addition, based on the head movement posture deviation information and the limb movement posture deviation information, the actual dance movements of the target subject are evaluated for the degree of coordination between the head and limb movements and the cumulative degree of movement posture deviation, and the corresponding dance movement scoring results are obtained. The evaluation of the degree of coordination between the head and limb movements and the cumulative degree of movement posture deviation can be achieved using models such as convolutional neural network models, which are conventional technical means in this field and will not be described in detail here.
[0107] In general, the dance movement spatiotemporal integration evaluation method and system based on monocular camera continuously monocularly shoots and collects audio of the target object's dance process to obtain corresponding dance movement dynamic images and background audio data, and frames the dance movement dynamic images for subsequent refined and independent evaluation of the target object's dance movements; human body key points are identified and spatially integrated for all dance movement image frames to obtain human body key point posture data, and based on the background audio data, the human body key point posture data corresponding to all dance movement image frames are temporally integrated to obtain human body key point posture change information throughout the process, and the dance movement characteristics of the target object are effectively refined; the human body key point posture change information throughout the process is also compared with the standard dance movement posture change information to obtain dance movement difference information, thereby evaluating the real-time dance of the target object and obtaining the dance movement scoring result, and using a unified standard for dance scoring to improve the efficiency, objectivity and credibility of dance evaluation.
[0108] The above is only a specific embodiment of the present invention, and any other improvements made based on the concept of the present invention are considered to be within the scope of protection of the present invention.
Claims
1. A method for temporal and spatial integration evaluation of dance movements based on monocular camera, characterized by: include: Continuously capture the target subject's dance process with a single camera and collect audio to obtain the corresponding dance movement dynamic images and background audio data; Performing frame processing on the dance action dynamic image to obtain a plurality of dance action image frames, and preprocessing all the dance action image frames; Performing human body key point recognition and spatial integration processing on all dance action image frames to obtain human body key point posture data of the target object; Based on the background audio data, time-series integration processing is performed on the human body key point posture data corresponding to all dance action image frames to obtain the full-process posture change information of the human body key points of the target object during the dance process; The whole-process posture change information of the key points of the human body is compared with the posture change information of the standard dance movement to obtain the dance movement difference information of the target object; based on the dance movement difference information, the actual dance movement of the target object is evaluated to obtain the corresponding dance movement scoring result.
2. The method for temporal and spatial integration of dance movements based on monocular photography as claimed in claim 1, wherein: Continuously capture the target subject's dance process with a single camera and collect audio to obtain the corresponding dance movement dynamic images and background audio data; The dance action dynamic image is subjected to frame processing to obtain a plurality of dance action image frames, and all the dance action image frames are pre-processed, including: The target subject's dance process is synchronously and continuously photographed and audio is collected by a monocular camera to obtain corresponding dance movement images and background audio data; the background audio data is compared with the background music played during the dance process, thereby converting the playback time axis of the background music into the collection time axis of the background audio data; The dance action dynamic image is subjected to frame processing to obtain a plurality of dance action image frames, and all the dance action image frames are subjected to pixel edge sharpening preprocessing and image noise removal preprocessing.
3. The method for temporal and spatial integration of dance movements based on monocular camera according to claim 1, wherein: After continuously shooting the target subject's dance and collecting audio data through a monocular camera to obtain the corresponding dance motion dynamic images and background audio data, the following steps are also included: Step S1, using the following formula (1), based on the dance movement dynamic images and background audio data obtained by continuous monocular shooting and audio collection, dynamic features of the dance movement are extracted. In the above formula (1), D represents the dynamic characteristic value of the dance movement; T represents the total number of image frames contained in the dance movement dynamic image; I(t) represents the image matrix data of the t-th image frame; I(t-1) represents the image matrix data of the t-1-th image frame; W(t) represents the weight function of the image matrix data of the t-th image frame; β represents the audio influence factor; S(t) represents the instantaneous value of the audio signal strength; γ represents the adjustment factor; R(t) represents the current environmental noise intensity; R th represents the noise threshold; Step S2, using the following formula (2), according to the dynamic feature value of the dance movement, obtain the audio-visual fusion feature value of the dance movement, In the above formula (2), F represents the audio-visual fusion feature value of the dance action; A(t) represents the instantaneous volume intensity value of the corresponding audio data; α(t) represents the time-dependent weight factor; δ(t) represents the time attenuation factor; Step S3, using the following formula (3), based on the dynamic feature value of the dance movement and the audio-visual fusion feature value, a dynamic control feedback adjustment signal is generated. In the above formula (3), C represents the dynamic control feedback adjustment signal; k represents the control gain; N represents the complexity of the dance movement; M represents the preset maximum movement complexity; F0 represents the preset audio-visual benchmark feature value of the dance movement; and λ represents the adjustment factor for controlling the impact of environmental noise on the feedback signal.
4. The method for temporal and spatial integration of dance movements based on monocular photography as claimed in claim 1, wherein: Performing human body key point recognition and spatial integration processing on all dance action image frames to obtain human body key point posture data of the target object; performing temporal integration processing on the human body key point posture data corresponding to all dance action image frames based on the background audio data to obtain the full-process human body key point posture change information of the target object during the dance process, including: Performing human body contour recognition processing on each dance action image frame to obtain the head contour information and limb contour information corresponding to the target object in each dance action image frame; determining the contour shape and depth feature information of the head and limbs of the target object based on the head contour feature information and the limb contour information; and selecting a plurality of head key points and a plurality of limb key points from the head and limbs respectively based on the contour shape and depth feature information; and Integrating the position information of all key points of the head and all key points of the limbs during the dance into the same three-dimensional space coordinate system to obtain the key point posture data of the target object during the dance; Based on the generated clock signal relationship between the background audio data and the dance movement dynamic image, the acquisition time axis of the background audio data is converted to generate a dance movement generation time axis corresponding to the dance movement dynamic image; based on the dance movement generation time axis, the human body key point posture data corresponding to all dance movement image frames are time-series integrated to obtain the full-process posture change information of the target object's human body key points during the dance process.
5. The method for temporal and spatial integration of dance movements based on monocular photography as claimed in claim 1, wherein: Comparing the whole-process posture change information of the key points of the human body with the posture change information of the standard dance movement to obtain dance movement difference information of the target object; based on the dance movement difference information, evaluating the actual dance movement of the target object to obtain a corresponding dance movement scoring result, including: The whole-process posture change information of the key points of the human body is divided into the whole-process posture change information of the key points of the head and the whole-process posture change information of the key points of the limbs, and the whole-process posture change information of the key points of the head and the whole-process posture change information of the key points of the limbs are respectively compared with the posture change information of the standard dance movements to obtain the head movement posture deviation information and the limb movement posture deviation information of the target object during the dance process; wherein the head movement posture deviation information includes the posture angle deviation value and the displacement deviation value of the head movement compared with the head sub-movement in the standard dance movement; the limb movement posture deviation information includes the posture angle deviation value and the displacement deviation value of the limb movement compared with the limb sub-movement in the standard dance movement; Based on the head movement posture deviation information and the limb movement posture deviation information, the actual dance movement of the target object is evaluated for the coordination degree of the head and limb movements and the cumulative degree of movement posture deviations to obtain a corresponding dance movement scoring result.
6. A dance movement spatiotemporal integration evaluation system based on monocular camera, characterized by: include: The dance process acquisition module is used to continuously capture the target object's dance process with a single camera and collect audio data to obtain the corresponding dance movement dynamic images and background audio data; a dance image processing module, configured to perform frame processing on the dance action dynamic image to obtain a plurality of dance action image frames, and perform pre-processing on all dance action image frames; A key point posture data generation module is used to perform human key point recognition and spatial integration processing on all dance action image frames to obtain human key point posture data of the target object; A key point posture data time series integration module is used to perform time series integration processing on the human body key point posture data corresponding to all dance action image frames based on the background audio data, so as to obtain the whole-process posture change information of the human body key points of the target object during the dance process; a dance movement difference determination module, configured to compare the whole-process posture change information of the key points of the human body with the posture change information of the standard dance movement to obtain dance movement difference information of the target object; The dance movement scoring module is used to evaluate the actual dance movement of the target object based on the dance movement difference information to obtain a corresponding dance movement scoring result.
7. The dance movement spatiotemporal integration evaluation system based on monocular camera according to claim 6, characterized in that: The dance process acquisition module is used to continuously monocularly capture and audio capture the target object's dance process to obtain corresponding dance movement dynamic images and background audio data, including: The target subject's dance process is synchronously and continuously photographed and audio is collected by a monocular camera to obtain corresponding dance movement images and background audio data; the background audio data is compared with the background music played during the dance process, thereby converting the playback time axis of the background music into the collection time axis of the background audio data; The dance image processing module performs frame processing on the dance action dynamic image to obtain a plurality of dance action image frames, and pre-processes all the dance action image frames, including: The dance action dynamic image is subjected to frame processing to obtain a plurality of dance action image frames, and all the dance action image frames are subjected to pixel edge sharpening preprocessing and image noise removal preprocessing.
8. The dance movement spatiotemporal integration evaluation system based on monocular camera according to claim 6, characterized in that: The key point posture data generation module is used to perform human key point recognition and spatial integration processing on all dance action image frames to obtain human key point posture data of the target object, including: Performing human body contour recognition processing on each dance action image frame to obtain the head contour information and limb contour information corresponding to the target object in each dance action image frame; determining the contour shape and depth feature information of the head and limbs of the target object based on the head contour feature information and the limb contour information; and selecting a plurality of head key points and a plurality of limb key points from the head and limbs respectively based on the contour shape and depth feature information; and Integrating the position information of all key points of the head and all key points of the limbs during the dance into the same three-dimensional space coordinate system to obtain the key point posture data of the target object during the dance; The key point posture data time series integration module is used to perform time series integration processing on the human body key point posture data corresponding to all dance action image frames based on the background audio data to obtain the full-process posture change information of the target object's human body key points during the dance process, including: Based on the generated clock signal relationship between the background audio data and the dance movement dynamic image, the acquisition time axis of the background audio data is converted to generate a dance movement generation time axis corresponding to the dance movement dynamic image; based on the dance movement generation time axis, the human body key point posture data corresponding to all dance movement image frames are time-series integrated to obtain the full-process posture change information of the target object's human body key points during the dance process.
9. The dance movement spatiotemporal integration evaluation system based on monocular camera according to claim 6, characterized in that: The dance movement difference determination module is used to compare the whole-process posture change information of the key points of the human body with the posture change information of the standard dance movement to obtain the dance movement difference information of the target object, including: The whole-process posture change information of the key points of the human body is divided into the whole-process posture change information of the key points of the head and the whole-process posture change information of the key points of the limbs, and the whole-process posture change information of the key points of the head and the whole-process posture change information of the key points of the limbs are respectively compared with the posture change information of the standard dance movements to obtain the head movement posture deviation information and the limb movement posture deviation information of the target object during the dance process; wherein the head movement posture deviation information includes the posture angle deviation value and the displacement deviation value of the head movement compared with the head sub-movement in the standard dance movement; the limb movement posture deviation information includes the posture angle deviation value and the displacement deviation value of the limb movement compared with the limb sub-movement in the standard dance movement; The dance movement scoring module is used to evaluate the actual dance movement of the target object based on the dance movement difference information to obtain a corresponding dance movement scoring result, including: Based on the head movement posture deviation information and the limb movement posture deviation information, the actual dance movement of the target object is evaluated for the coordination degree of the head and limb movements and the cumulative degree of movement posture deviations to obtain a corresponding dance movement scoring result.
Citation Information
Cited By
Human body action evaluation method and system based on camera recognition
CN121482873A