Human action evaluation method and system based on camera recognition
By generating motion axes through camera recognition and comparing videos, the problem of identifying differences in rhythm and posture in group dances has been solved. This has enabled intelligent and efficient evaluation of group dance movements, providing personalized correction criteria and improving the efficiency and accuracy of group dance training and evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies struggle to effectively determine whether each dancer has rhythm or posture issues in group dance scenarios, especially when multiple dancers overlap or movements switch rapidly. The extraction of temporal features is chaotic, and there is a lack of accurate modeling of the relationship between movement timing and musical rhythm.
By using a camera-based recognition method, the movement axes of the lead dancer and the group dancers are generated, and video comparison is performed to identify time periods of difference. The difference display image slot filling design is used to realize the quantitative recognition and display of beat differences and posture differences.
It achieves intelligent and efficient evaluation of group dance movements throughout the entire process, enabling rapid identification of common movement problems in the group, providing personalized correction criteria, and significantly improving the efficiency and accuracy of group dance training and evaluation.
Smart Images

Figure CN121482873B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a human action evaluation method and system based on camera recognition. BACKGROUND
[0002] In group dance performance and training scenarios, the standardization and consistency of dancer's action posture directly determine the artistic presentation effect of the performance and the efficiency of the training quality improvement. Whether it is the stage rehearsal of a professional dance troupe, the teaching evaluation of an art college, or the fair evaluation of a dance competition, it is necessary to accurately and objectively evaluate the action posture of each dancer in the group. Such evaluation not only judges whether the individual dancer's action conforms to the standard norm, but also considers the action synchronization and posture coordination in the group to ensure the artistic characteristics of uniformity and rhythm unity pursued by group dance performance, so the detailed evaluation of each dancer's individual action posture has become one of the core needs in the field of group dance.
[0003] Technical research on dance action evaluation has made certain progress, and existing technologies mainly fall into two categories: traditional manual evaluation and computer vision-based technical solutions. The traditional method relies on professional judges or teachers to observe by eye and subjectively score the dancer's action based on their own dance experience. This method is greatly influenced by the evaluator's subjective cognition, fatigue level, and other factors, making it difficult to achieve standardization and consistency of evaluation results. The computer vision-based solution captures dance video data through a camera, extracts key skeletal node information of the dancer using human pose estimation algorithms (such as OpenPose, AlphaPose, and other models based on deep learning), and then compares the angle and distance between skeletal nodes with standard action templates to achieve quantitative evaluation of action posture. Some technologies also introduce action timing analysis to attempt to judge the continuity of the action through feature sequence matching.
[0004] Although existing technologies provide a quantitative approach to dance action evaluation, there are still obvious limitations in group dance scenarios, making it difficult to effectively determine whether each dancer has a rhythm problem or a posture problem. In terms of rhythm judgment, existing technologies mainly focus on spatial dimension comparison of action features, lack precise modeling of the correlation between action timing and music rhythm, and cannot accurately identify the problem of dancers' action being ahead of or lagging behind the rhythm, especially when multiple dancers overlap or action switches quickly, which can lead to confusion in extracting timing features. SUMMARY
[0005] Based on the above problems, the present application is proposed to provide a human action evaluation method and system based on camera recognition to overcome the above problems or at least partially solve the above problems.
[0006] According to one aspect of the present application, a human action evaluation method based on camera recognition is provided, comprising the following steps:
[0007] Tracking and shooting different dancers located in each collection and shooting group based on each collection unit, and generating a main action axis and each slave action axis based on the obtained leading dance video and each group dance video corresponding to the leading dancer and each group dancer;
[0008] Video comparison is performed between the leading dance video and each group dance video to obtain each difference period corresponding to each group dance video, and each vertical difference axis corresponding to each group dancer generated based on each difference period is added to each first filling slot in the difference display image for filling each vertical difference axis;
[0009] When the management end interacts with the difference axis part corresponding to any difference period located in any vertical difference axis based on the difference display image, the vertical difference axis and the difference axis part are determined as the target difference axis and the target axis part, and the target difference segment corresponding to the target axis part is determined based on the group dance video corresponding to the target difference axis;
[0010] Each group dance difference segment corresponding to each difference axis part having an axis overlap relationship with the target axis part in each vertical difference axis is determined based on the remaining group dance videos, and the target difference segment and each group dance difference segment are added to each second filling slot in the difference display image for filling each difference segment.
[0011] Optionally, in the method according to the present application, tracking and shooting different dancers located in each collection and shooting group based on each collection unit, and generating a main action axis and each slave action axis based on the obtained leading dance video and each group dance video corresponding to the leading dancer and each group dancer, comprises:
[0012] When the light collection value output by the light sensor arranged on the stage platform at any time is less than a first preset light value, the time is determined as a waiting time;
[0013] When the light collection value output by the light sensor at any time after the waiting time is greater than a second preset light value, the time is determined as a collection time, and the overhead unit arranged above the stage is controlled to collect images of the stage to obtain a stage overhead view;
[0014] Image recognition is performed on the stage overhead view to obtain each collection and shooting group corresponding to different collection units, including different dancers, wherein the dancers include the leading dancer and the group dancers;
[0015] Tracking and shooting different dancers located in each collection and shooting group based on each collection unit, and generating a main action axis and each slave action axis based on the obtained leading dance video and each group dance video corresponding to the leading dancer and each group dancer.
[0016] Optionally, in the method according to the present application, the stage view is subjected to image recognition to obtain each collection shooting group including different dancers corresponding to different collection units, wherein the dancers include a lead dancer and group dancers, comprising:
[0017] The stage view is divided based on a collection angle corresponding to each collection unit to obtain a collection region corresponding to each collection unit;
[0018] The stage view is subjected to image recognition to determine each dancer region corresponding to each dancer located on the stage view, wherein the dancers include a lead dancer and group dancers;
[0019] Each dancer region having a complete overlap relationship with the same collection region is divided into the same region group, and each dancer pixel point constituting each dancer region is determined;
[0020] Based on the stage view, each region connection line connecting each dancer pixel point corresponding to each dancer region corresponding to the same region group to a collection center point of the collection unit corresponding to the region group is determined, and each line segment pixel point constituting each region connection line is subjected to pixel connection to obtain each connection region;
[0021] Each dancer region corresponding to the connection region and not having an overlap relationship with any other dancer region is determined as an effective region, and the dancer corresponding to the effective region is divided into a collection shooting group corresponding to the collection unit of the collection region corresponding to the region group including the connection region.
[0022] Optionally, in the method according to the present application, the method further comprises:
[0023] Based on the stage view, each dancer connection line connecting a dancer center point corresponding to each dancer region to a collection center point of the collection unit corresponding thereto is generated, and each connection length corresponding to each dancer connection line is determined;
[0024] If each collection shooting group corresponding to multiple collection units includes the same dancer, each dancer connection line corresponding to the smallest connection length in each dancer connection line corresponding to the dancer is determined as a first connection line;
[0025] The dancer is removed from each collection shooting group based on each collection unit corresponding to each dancer connection line except the first connection line.
[0026] Optionally, in the method according to the present application, each dancer located in each collection shooting group is subjected to tracking shooting based on each collection unit, and a master action axis and each slave action axis are generated based on a lead dance video and each group dance video corresponding to the lead dancer and each group dancer, comprising:
[0027] Each acquisition unit tracks and films different dancers located in each acquisition and shooting group to obtain each acquisition video;
[0028] Image recognition is performed on the video image frames that are at the beginning of each video capture to obtain the human body regions located in each video image frame;
[0029] Based on the stage view, each human body region that has the same regional position as each dancer region corresponding to the acquisition and shooting group of the acquisition unit is identified as a target region in each human body region.
[0030] Determine the number of each target in each target region included in each video image frame, and copy each captured video based on the number of each target to obtain each captured video corresponding to each dancer in each target region.
[0031] Based on each captured video, determine the corresponding lead dancer video and group dance video for each lead dancer and group dancer, and generate the main motion axis and the secondary motion axis based on each lead dancer video and group dance video.
[0032] Optionally, in the method according to the present invention, determining the lead dancer video and each group dance video corresponding to the lead dancer and each group dancer based on each acquired video, and generating the main motion axis and each secondary motion axis based on each lead dancer video and each group dance video, includes:
[0033] Based on each video image frame that makes up each captured video, determine the contour of each target region corresponding to it, and generate the bounding rectangle of each contour corresponding to each region contour.
[0034] Based on the preset magnification factor, the bounding rectangles of each contour are magnified, and based on the updated bounding rectangles of each contour, each video image frame is cropped. Face recognition is performed on the updated captured videos, and the recognition results are compared with the preset lead dancer information.
[0035] The videos that show consistent information after comparison are identified as lead dance videos corresponding to the lead dancers, and the videos that show discrepancies after comparison are identified as group dance videos corresponding to the group dancers.
[0036] The main motion axis and the secondary motion axis are generated based on the videos of each lead dancer and each group dance.
[0037] Optionally, in the method according to the present invention, the lead dancer video is compared with each group dance video to obtain the time periods in which the group dancers and the lead dancer differ in each group dance video, including:
[0038] Acquire the individual lead dancer image frames that make up the lead dancer video, and determine the corresponding lead dancer limb movements for each lead dancer image frame;
[0039] Starting with the lead dancer image frame at the beginning, determine frame by frame whether the lead dancer's body movements are the same for all adjacent lead dancer image frames, and divide each lead dancer image frame that is arranged continuously and whose corresponding lead dancer body movements are the same into a lead dancer segment.
[0040] Based on the lead dancer's video, determine the standard movement time periods for each lead dancer segment corresponding to different lead dancer's body movements;
[0041] Based on each group dance video, determine each group dance segment corresponding to each standard movement time period, and determine the group dance limb movements corresponding to the first and last group dance image frames that make up each group dance segment.
[0042] If no group dance limb movement is present in the first or last group dance image frame corresponding to any group dance segment, it is determined that the group dance segment has a beat difference with slow or fast beat attributes, and the beat difference time period corresponding to the group dance segment with beat difference is determined based on the group dance video corresponding to the group dance segment.
[0043] Optionally, in the method according to the invention, the method further includes:
[0044] Acquire the group dance image frames that make up each group dance video, and determine the group dance limb movements corresponding to each group dance image frame;
[0045] Starting with the first group dance image frame, determine frame by frame whether the group dance limb movements corresponding to all adjacent group dance image frames are the same, and divide the group dance image frames that are arranged continuously and whose corresponding group dance limb movements are the same into a group dance segment.
[0046] Based on the videos of each group dance, the time periods of each group dance movement in each group dance segment corresponding to different group dance body movements are determined;
[0047] Each group dance segment is identified, and a preset standard segment corresponding to the group dance segment is retrieved based on the identification results;
[0048] Each group dance segment is compared with a preset standard segment corresponding to the group dance segment, and the time periods of each group dance movement with posture differences are determined as posture difference time periods based on the comparison results of each group dance video.
[0049] Optionally, in the method according to the present invention, each group dance segment is compared with a preset standard segment corresponding to the group dance segment, and the time periods of each group dance movement where the comparison result shows a difference in posture are determined as the time periods of each posture difference based on each group dance video, including:
[0050] Determine the sequence number of each group dance frame corresponding to each group dance image frame that makes up each group dance segment, and determine the sequence number of each posture image frame corresponding to each preset standard segment that makes up each group dance segment.
[0051] Image recognition is performed on each group dance image frame and each standard posture image frame to obtain the group dance limb regions and each standard posture regions corresponding to the group dance limb movements and preset standard segments.
[0052] Each group dance image frame that makes up each group dance segment is added to the retrieved image comparison layer, and based on each image comparison layer, the standard pose image that has the same frame number as the group dance image frame located in the image comparison layer is superimposed on the image comparison layer from each standard pose image frame corresponding to the group dance segment.
[0053] The percentage of overlap between the limb regions of each group dance and the regions of each standard posture is determined based on the image comparison layers.
[0054] When the area overlap ratio of any image comparison layer is within a preset ratio range or less than a preset ratio threshold, the group dance segment corresponding to the group dance image frame of the image comparison layer is determined to have a posture difference with flawed or incorrect attributes, and the group dance action time period corresponding to the group dance segment with posture difference is determined as the posture difference time period based on the group dance video corresponding to the group dance segment.
[0055] According to another aspect of the present invention, a human motion evaluation system based on camera recognition is provided, comprising:
[0056] The generation module is configured to track and film different dancers in each acquisition and shooting group based on each acquisition unit, and generate main motion axes and secondary motion axes based on the obtained lead dancer videos and group dance videos corresponding to the lead dancer and each group dancer.
[0057] The comparison module is configured to compare the lead dancer video with each group dance video to obtain the difference time periods corresponding to each group dance video, and add the vertical difference axis corresponding to each group dancer generated based on each difference time period to each first filling slot located in the difference display image for filling each vertical difference axis.
[0058] The interaction module is configured to, when the management terminal interacts with the difference axis portion corresponding to any difference time period on any vertical difference axis based on the difference display image, determine the vertical difference axis and the difference axis portion as the target difference axis and the target axis portion, and determine the target difference segment corresponding to the target axis portion based on the group dance video corresponding to the target difference axis.
[0059] The module is configured to determine, based on the remaining group dance videos, the group dance difference segments corresponding to the difference axis portions that have an axial overlap relationship with the target axis portion in each vertical difference axis, and add the target difference segment and each group dance difference segment to the second fill slots located in the difference display image for filling each difference segment.
[0060] According to the present invention, the entire process of group dance movement evaluation is made intelligent and efficient, possessing significant practical value and technical advantages. On the one hand, based on the tracking and shooting of the acquisition unit and the generation of each movement axis, the dance videos of the lead dancer and group dancers are transformed into structured temporal movement data, realizing the quantitative identification of rhythm and posture differences. On the other hand, the slot-filling design of the vertical difference axis and the difference display image makes the difference periods of each group dancer clear at a glance, so that the management end does not need to review lengthy videos frame by frame, but can retrieve the target difference segment corresponding to each group dancer through simple interaction, greatly reducing the time cost of problem investigation. At the same time, it can batch retrieve and centrally display the difference segments of each group dance corresponding to the difference axis parts with axis overlap, helping the management end to quickly locate common movement problems of the group. The present invention can provide personalized correction basis for each dancer, significantly improving the efficiency and accuracy of group dance training and evaluation. Attached Figure Description
[0061] Figure 1 A flowchart of a human motion evaluation method based on camera recognition according to an embodiment of the present invention is shown;
[0062] Figure 2 A schematic diagram of the acquisition area according to an embodiment of the present invention is shown;
[0063] Figure 3 A structural block diagram of a human motion evaluation system based on camera recognition according to another embodiment of the present invention is shown. Detailed Implementation
[0064] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0065] To address the problems existing in the aforementioned background art, the inventors proposed the solution of this invention. One embodiment of this invention provides a human motion evaluation method based on camera recognition, which can be executed in a computing device.
[0066] Figure 1A flowchart of a camera-based human motion evaluation method according to an embodiment of the present invention is shown, the method being adapted to be executed in a computing device.
[0067] like Figure 1 As shown, the human motion evaluation method based on camera recognition proposed in this embodiment begins with step S102, which includes the following:
[0068] Each acquisition unit tracks and films different dancers located in each acquisition and shooting group, and generates main motion axes and secondary motion axes based on the obtained lead dancer videos and group dance videos corresponding to the lead dancer and each group dancer.
[0069] For example, in this embodiment, the acquisition unit can be understood as a video acquisition device located at the front of the stage for capturing videos of the dancers on the stage.
[0070] Since there may be many dancers on stage at the same time, it is not possible to assign a specific capture unit to each dancer. Therefore, the server will group each dancer into capture and shooting groups corresponding to each capture unit, and each capture and shooting group contains different dancers.
[0071] In a typical dance troupe, there is a lead dancer and multiple group dancers. Therefore, the server can control each acquisition unit to track and film different dancers in each acquisition and shooting group, and then obtain the lead dancer video corresponding to the lead dancer and the group dance videos corresponding to each group dancer.
[0072] To facilitate quick and intuitive viewing of each video later, the server will generate the main motion axis and each secondary motion axis based on the lead dancer video and each group dance video.
[0073] Furthermore, the aforementioned "tracking and filming different dancers located in each acquisition and shooting group based on each acquisition unit, and generating main motion axes and secondary motion axes based on the obtained lead dancer videos and group dance videos corresponding to the lead dancer and each group dancer" also includes the following steps:
[0074] If the light sensor installed on the stage surface outputs a light collection value that is less than the first preset light value at any given moment, that moment is determined as a waiting moment.
[0075] When the light sensor outputs a light acquisition value greater than the second preset light value at any time after the waiting time, the moment is determined as the acquisition moment, and the upward viewing unit set above the stage is controlled to acquire an image of the stage to obtain a view of the stage.
[0076] Image recognition is performed on the stage view to obtain various acquisition and shooting groups corresponding to different acquisition units, including different dancers, wherein the dancers include lead dancers and group dancers;
[0077] Each acquisition unit tracks and films different dancers located in each acquisition and shooting group, and generates main motion axes and secondary motion axes based on the obtained lead dancer videos and group dance videos corresponding to the lead dancer and each group dancer.
[0078] For example, in this embodiment, when the light sensor set on the stage surface outputs a light collection value that is less than the first preset light value at any time, it indicates that the stage lights on the current stage are dim, that is, the dancer may be in the standing preparation stage, so the server will determine this moment as the waiting moment.
[0079] Then, when the light sensor outputs a light collection value greater than the second preset light value at any time after the waiting time, it indicates that the stage lights on the current stage are bright, that is, the dancer is about to perform a dance. At this time, the server will determine the time as the collection time and control the upward viewing unit set above the stage to collect images of the stage and obtain a view of the stage.
[0080] Next, the server will perform image recognition on the view on the stage to obtain various capture and shooting groups corresponding to different capture units, including different dancers, among which the dancers include lead dancers and group dancers.
[0081] Finally, the server will control each acquisition unit to track and film different dancers in each acquisition and shooting group, thereby obtaining the lead dancer video corresponding to the lead dancer and the group dance videos corresponding to each group dancer. In order to facilitate quick and intuitive viewing of each video later, the server will generate the main motion axis and each secondary motion axis based on the lead dancer video and each group dance video.
[0082] Furthermore, the aforementioned "performing image recognition on the stage view to obtain various acquisition and shooting groups corresponding to different acquisition units, including different dancers, wherein the dancers include lead dancers and ensemble dancers" also includes the following steps:
[0083] The stage view is divided based on the acquisition angle of each acquisition unit to obtain the acquisition area of each acquisition unit.
[0084] Image recognition is performed on the view on the stage to determine the dancer area corresponding to each dancer in the view on the stage, wherein the dancers include lead dancers and group dancers;
[0085] Each dancer region that completely overlaps with the same acquisition area is divided into the same region group, and each dancer pixel that makes up each dancer region is determined.
[0086] Based on the stage view, connect the dancer pixels corresponding to each dancer area in the same area group to the acquisition center point of the acquisition unit corresponding to the area group, and connect the pixel points of each line segment that makes up each area connection line to obtain each connection area.
[0087] Dancer areas whose corresponding connecting regions do not overlap with any other dancer areas are identified as valid areas, and dancers corresponding to valid areas are assigned to acquisition and shooting groups corresponding to acquisition units of acquisition areas that include the connecting regions.
[0088] For example, in this embodiment, since each acquisition unit is installed at a different location, each acquisition unit located at a different location has a different acquisition angle. Therefore, the server will first divide the stage view based on the acquisition angle of each acquisition unit, so as to obtain the acquisition area of each acquisition unit in the stage view.
[0089] Assuming three acquisition units are deployed at the front of the stage, acquisition unit 1 has an acquisition angle of 45° to the left side of the stage, acquisition unit 2 has an acquisition angle of the front of the stage, and acquisition unit 3 has an acquisition angle of 45° to the right side of the stage, the server will divide the view on the stage into three acquisition areas: left, center, and right, corresponding to the three acquisition units, based on the acquisition angles and coverage of these three acquisition units. Then, the server will use a human contour recognition algorithm to scan the view on the stage and determine the dancer area for each dancer on the stage.
[0090] To ensure that subsequent acquisition units can completely capture each dancer within the acquisition and shooting group corresponding to that acquisition unit, the server will group dancers whose areas completely overlap with the same acquisition area into the same area group, such as... Figure 2 The shaded area filled with polka dots shown is the acquisition area corresponding to the acquisition unit, and each dancer area completely overlaps with the acquisition area. Therefore, the server will classify both dancer areas into the area group corresponding to the acquisition area.
[0091] Next, the server will determine the individual dancer pixels that make up each dancer region. Then, it will connect the individual dancer pixels in each region group to the acquisition center point of the acquisition unit corresponding to that region group. Finally, it will connect the individual pixel segments of each line that make up each region connection line to obtain each connected region.
[0092] When the connecting area corresponding to any dancer area does not overlap with any other dancer area, it means that the dancer in that dancer area is not obstructed by other dancers and can be directly captured by the acquisition unit. Therefore, the server will determine the dancer area whose connecting area does not overlap with any other dancer area as a valid area, and assign the dancers corresponding to the valid area to the acquisition and shooting group corresponding to the acquisition unit of the acquisition area that includes the area group of the connecting area, to ensure that the dancers in each acquisition and shooting group are within the optimal shooting range of the corresponding acquisition unit.
[0093] Furthermore, the above method also includes the following steps:
[0094] Based on the stage view, generate connection lines for each dancer, starting from the center point of each dancer's area and moving towards the center point of the corresponding acquisition unit, and determine the connection length corresponding to each connection line.
[0095] If each acquisition and shooting group corresponding to multiple acquisition units includes the same dancer, the dancer connection line with the smallest connection length among the dancer connection lines corresponding to the dancer is determined as the first connection line;
[0096] The dancers are removed from each acquisition unit's acquisition and shooting group based on each dancer's connection line except for the first connection line.
[0097] For example, in this embodiment, it can be understood that the same dancer may be located in multiple acquisition units at the same time. In order to reduce the amount of data processing on the server and avoid the problem of repeated tracking and resource waste in subsequent video acquisition, the server will first generate dancer connection lines in the stage view, starting from the center point of the dancer in each dancer area and connecting to the acquisition center point of the corresponding acquisition unit, and determine the connection length corresponding to each dancer connection line.
[0098] The shorter the connection length, the closer the dancer is to the acquisition unit. When multiple acquisition units correspond to the same dancer in each acquisition and shooting group, the server will identify the dancer connection line with the smallest connection length among the dancer connection lines corresponding to that dancer as the first connection line. The dancer will then be removed from each acquisition and shooting group of each acquisition unit corresponding to each dancer connection line other than the first connection line. This allows the dancer to be filmed based on the acquisition unit with the smallest corresponding connection length. This not only ensures the filming effect of the dancer and improves the quality of video data, but also avoids the interference of duplicate video data on subsequent analysis, thus improving the efficiency of data processing.
[0099] Furthermore, the aforementioned "tracking and filming different dancers in each acquisition and shooting group based on each acquisition unit, and generating main motion axes and secondary motion axes based on the obtained lead dancer videos and group dance videos corresponding to the lead dancer and each group dancer" also includes the following steps:
[0100] Each acquisition unit tracks and films different dancers located in each acquisition and shooting group to obtain each acquisition video;
[0101] Image recognition is performed on the video image frames that are at the beginning of each video capture to obtain the human body regions located in each video image frame;
[0102] Based on the stage view, each human body region that has the same regional position as each dancer region corresponding to the acquisition and shooting group of the acquisition unit is identified as a target region in each human body region.
[0103] Determine the number of each target in each target region included in each video image frame, and copy each captured video based on the number of each target to obtain each captured video corresponding to each dancer in each target region.
[0104] Based on each captured video, determine the corresponding lead dancer video and group dance video for each lead dancer and group dancer, and generate the main motion axis and the secondary motion axis based on each lead dancer video and group dance video.
[0105] For example, in this embodiment, the server controls each acquisition unit to track and film different dancers located in each acquisition and filming group, thereby obtaining each acquisition video. The tracking and filming method can be to enable the function of the portrait tracking frame when filming, and continuously lock the different dancers located in each acquisition and filming group through the portrait tracking frame to achieve tracking and filming.
[0106] Since the video captured by the corresponding acquisition unit may include dancers other than those in the corresponding acquisition and shooting group, the server will perform image recognition on the video image frames that make up each video to obtain the human body regions in each video image frame. Then, in the stage view, the human body regions that have the same region position as the dancer regions corresponding to the acquisition and shooting group are determined as the target regions. That is, the dancers in the target regions are located in the acquisition and shooting group corresponding to the acquisition unit.
[0107] Next, the server will determine the number of each target in each target region included in each video image frame, and copy each captured video based on the number of each target to obtain each captured video corresponding to each dancer in each target region.
[0108] For example, if the target quantity is 5, the server will copy the captured video into 5 copies, that is, get 5 identical captured videos, and bind each captured video to each dancer;
[0109] Finally, the server will determine the lead dancer videos and group dance videos corresponding to each lead dancer and group dancer based on each captured video, and generate the main motion axis and the secondary motion axis based on each lead dancer video and group dance video.
[0110] Furthermore, the aforementioned "determining the lead dancer videos and group dance videos corresponding to each lead dancer and group dancer based on each captured video, and generating the main motion axis and secondary motion axes based on each lead dancer video and group dance video" also includes the following steps:
[0111] Based on each video image frame that makes up each captured video, determine the contour of each target region corresponding to it, and generate the bounding rectangle of each contour corresponding to each region contour.
[0112] Based on the preset magnification factor, the bounding rectangles of each contour are magnified, and based on the updated bounding rectangles of each contour, each video image frame is cropped. Face recognition is performed on the updated captured videos, and the recognition results are compared with the preset lead dancer information.
[0113] The videos that show consistent information after comparison are identified as lead dance videos corresponding to the lead dancers, and the videos that show discrepancies after comparison are identified as group dance videos corresponding to the group dancers.
[0114] The main motion axis and the secondary motion axis are generated based on the videos of each lead dancer and each group dance.
[0115] For example, in this embodiment, the server first determines the contour of each target area corresponding to each video image frame that makes up each captured video, and then generates the bounding rectangle of each contour corresponding to each area contour to ensure that the bounding rectangle can completely wrap the dancer's whole body.
[0116] The server has a preset magnification factor, for example, a preset magnification factor of 1.1x. The server will magnify the bounding rectangle of each outline proportionally according to the preset magnification factor to avoid missing the dancer's body movements during subsequent cropping. After the server magnifies the bounding rectangle of each outline based on the preset magnification factor, it will crop each video image frame based on the updated bounding rectangle of each outline to obtain the updated captured video.
[0117] Next, the server will perform facial recognition on each of the updated captured videos, extract the facial feature points of the dancers in the videos, and compare them with the facial features corresponding to the lead dancers in the preset lead dancer information database. For example, when the facial feature matching degree is greater than 95%, the server will determine the comparison result as consistent, indicating that the captured video is the lead dance video corresponding to the lead dancer. Conversely, the server will determine the comparison result as inconsistent, that is, the captured video is not the lead dance video corresponding to the lead dancer.
[0118] Finally, the server will generate main motion axes and secondary motion axes based on the obtained lead dancer videos and group dance videos, which will facilitate subsequent quantitative comparison of rhythm and posture differences and ensure the objectivity of the evaluation results.
[0119] Step S104 includes the following:
[0120] The lead dancer video is compared with each group dance video to obtain the difference time periods for each group dance video. The vertical difference axis corresponding to each group dancer, generated based on each difference time period, is added to the first fill slots in the difference display image to fill each vertical difference axis.
[0121] For example, in this embodiment, the server compares the lead dancer video with each group dance video frame by frame to obtain the time difference and posture difference between each group dance video and the lead dancer video. The time difference can be understood as the presence of a rushed or slow shot, and the posture difference can be understood as the movement being too large or too small.
[0122] Subsequently, the server will generate a vertical difference axis for each group of dancers based on the time intervals of these difference periods, and mark the vertical difference axis with different preset colors according to different difference situations. For example, the axis part corresponding to the beat difference will be set to red, and the axis part corresponding to the posture difference will be set to blue.
[0123] Next, the server creates a difference display image, which includes first fill slots for filling each vertical difference axis and second fill slots for filling each difference segment.
[0124] Finally, the server adds each vertical difference axis to the first fill slot in the difference display image, making it easy for the management to quickly locate the dancer's movement problems based on each vertical difference axis. This makes the problems of each dancer clear at a glance and greatly reduces the time cost for the management to troubleshoot problems.
[0125] Furthermore, the aforementioned "comparing the lead dancer's video with each group dance video to identify the time periods where differences exist between the group dancers and the lead dancer in each group dance video" also includes the following steps:
[0126] Acquire the individual lead dancer image frames that make up the lead dancer video, and determine the corresponding lead dancer limb movements for each lead dancer image frame;
[0127] Starting with the lead dancer image frame at the beginning, determine frame by frame whether the lead dancer's body movements are the same for all adjacent lead dancer image frames, and divide each lead dancer image frame that is arranged continuously and whose corresponding lead dancer body movements are the same into a lead dancer segment.
[0128] Based on the lead dancer's video, determine the standard movement time periods for each lead dancer segment corresponding to different lead dancer's body movements;
[0129] Based on each group dance video, determine each group dance segment corresponding to each standard movement time period, and determine the group dance limb movements corresponding to the first and last group dance image frames that make up each group dance segment.
[0130] If no group dance limb movement is present in the first or last group dance image frame corresponding to any group dance segment, it is determined that the group dance segment has a beat difference with slow or fast beat attributes, and the beat difference time period corresponding to the group dance segment with beat difference is determined based on the group dance video corresponding to the group dance segment.
[0131] For example, in this embodiment, the server will first obtain each lead dancer image frame that makes up the lead dancer video. The corresponding lead dancer limb movements can be determined by the scene corresponding to each image frame. The server will then take the lead dancer image frame at the beginning as the starting point and determine whether the lead dancer limb movements corresponding to all adjacent lead dancer image frames are the same according to the time sequence. The lead dancer image frames that are arranged continuously and correspond to the same lead dancer limb movements will be divided into a lead dancer segment.
[0132] For example, the lead dancer video has 5 lead dancer image frames, and the corresponding lead dancer body movements are head dance movements, head dance movements, leg dance movements, leg dance movements, and head dance movements. Then the first and second frames can be divided into a lead dancer segment, the third and fourth frames can be divided into a lead dancer segment, and the fifth frame is a lead dancer segment. Thus, lead dancer segments corresponding to different lead dancer body movements can be obtained.
[0133] Next, the server will determine the standard action time period for each dance segment corresponding to different lead dancer body movements based on the lead dancer video. For example, if the time interval corresponding to a lead dance segment with a corresponding leg dance movement is determined to be 00:00:20-00:00:22 based on the lead dance video, then 00:00:20-00:00:22 is the standard action time period corresponding to that lead dance segment.
[0134] Next, the server will identify the group dance segments corresponding to the standard action time periods in each group dance video. For example, if 00:00:20-00:00:22 is the standard action time period corresponding to a lead dancer segment, the server will extract the group dance segments in each group dance video that are consistent with the time of that standard action time period.
[0135] Subsequently, the server will determine the group dance limb movements corresponding to the first and last group dance image frames that make up each group dance segment. If there is no group dance limb movement in the first group dance image frame corresponding to any group dance segment, it means that the group dancers corresponding to that group dance segment are in a slow beat during the standard movement period. At this time, the server will determine that there is a beat difference with slow beat attributes in that group dance segment.
[0136] If the last group dance image frame of any group dance segment does not contain any group dance limb movements, it means that the group dancers corresponding to that group dance segment have rushed the beat during the standard movement period. In this case, the server will determine that there is a beat difference with the rushing attribute in the group dance segment.
[0137] Finally, the server will determine the time period of the beat difference corresponding to the group dance segment with the beat difference based on the group dance video corresponding to the group dance segment;
[0138] This invention can accurately capture and determine whether there are rhythm differences in group dance segments, such as slow or fast beats, thereby improving the accuracy of rhythm evaluation.
[0139] Furthermore, the above method also includes the following steps:
[0140] Acquire the group dance image frames that make up each group dance video, and determine the group dance limb movements corresponding to each group dance image frame;
[0141] Starting with the first group dance image frame, determine frame by frame whether the group dance limb movements corresponding to all adjacent group dance image frames are the same, and divide the group dance image frames that are arranged continuously and whose corresponding group dance limb movements are the same into a group dance segment.
[0142] Based on the videos of each group dance, the time periods of each group dance movement in each group dance segment corresponding to different group dance body movements are determined;
[0143] Each group dance segment is identified, and a preset standard segment corresponding to the group dance segment is retrieved based on the identification results;
[0144] Each group dance segment is compared with a preset standard segment corresponding to the group dance segment, and the time periods of each group dance movement with posture differences are determined as posture difference time periods based on the comparison results of each group dance video.
[0145] For example, in this embodiment, in addition to the difference in rhythm, the dancers may also have differences in posture. In order to accurately capture the differences in posture of the dancers, the server will first obtain the group dance image frames that make up each group dance video, and then determine the group dance limb movements corresponding to each group dance image frame.
[0146] Next, the server will also start from the first group dance image frame and determine whether the group dance limb movements corresponding to all adjacent group dance image frames are the same, and divide the group dance image frames that are arranged continuously and whose corresponding group dance limb movements are the same into a group dance segment.
[0147] Next, the server will determine the time period of each group dance segment corresponding to different group dance body movements based on each group dance video. For example, based on the group dance video of a certain group dancer, the time period of the group dance segment corresponding to the head group dance body movements is determined to be 00:01:00-00:01:04.
[0148] Subsequently, the server will identify each group dance segment and retrieve the preset standard segment corresponding to the group dance segment from the preset action standard segment library. Then, each group dance segment will be compared with the preset standard segment corresponding to the group dance segment.
[0149] If the comparison result shows that there is a difference in posture, the server will determine the time period of each group dance movement that shows a difference in posture based on each group dance video.
[0150] This embodiment can more accurately determine whether there are postural differences among each group dancer by retrieving preset standard segments, thereby improving the standardization of postural evaluation.
[0151] Furthermore, the aforementioned "comparing each group dance segment with a preset standard segment corresponding to the group dance segment, and determining the time periods of each group dance movement with posture differences based on the comparison results" also includes the following steps:
[0152] Determine the sequence number of each group dance frame corresponding to each group dance image frame that makes up each group dance segment, and determine the sequence number of each posture image frame corresponding to each preset standard segment that makes up each group dance segment.
[0153] Image recognition is performed on each group dance image frame and each standard posture image frame to obtain the group dance limb regions and each standard posture regions corresponding to the group dance limb movements and preset standard segments.
[0154] Each group dance image frame that makes up each group dance segment is added to the retrieved image comparison layer, and based on each image comparison layer, the standard pose image that has the same frame number as the group dance image frame located in the image comparison layer is superimposed on the image comparison layer from each standard pose image frame corresponding to the group dance segment.
[0155] The percentage of overlap between the limb regions of each group dance and the regions of each standard posture is determined based on the image comparison layers.
[0156] When the area overlap ratio of any image comparison layer is within a preset ratio range or less than a preset ratio threshold, the group dance segment corresponding to the group dance image frame of the image comparison layer is determined to have a posture difference with flawed or incorrect attributes, and the group dance action time period corresponding to the group dance segment with posture difference is determined as the posture difference time period based on the group dance video corresponding to the group dance segment.
[0157] For example, in this embodiment, the server will determine the sequence number of each group dance frame corresponding to each group dance image frame that makes up each group dance segment, retrieve each preset standard segment corresponding to each group dance segment, and further obtain the sequence number of each posture frame corresponding to each standard posture image frame that makes up each preset standard segment.
[0158] Next, the server will perform image recognition on each group dance image frame and each standard posture image frame, extracting key limb regions such as "head", "arms", "legs" and "torso", thereby obtaining each group dance limb region and each standard posture region corresponding to the group dance limb movements and preset standard segments.
[0159] Next, the server will add each group dance image frame that makes up each group dance segment to the retrieved image comparison layer. That is, each group dance image frame corresponds to an image comparison layer. The server will overlay the standard pose images with the same frame number as the group dance image frames located in each image comparison layer from the standard pose image frames corresponding to the group dance segment to the image comparison layer.
[0160] Subsequently, the server will determine the percentage overlap between the group dance limb area and the standard posture area based on each image comparison layer, and retrieve the preset percentage range and preset percentage threshold. The preset percentage range can be 80%-90%, and the preset percentage threshold can be 80%.
[0161] When the overlap ratio of any region in the image comparison layer is within the preset ratio range, it indicates that the movement amplitude of the group dance image frame corresponding to the overlap ratio of that region is slightly lacking. At this time, the server will determine the group dance segment corresponding to the group dance image frame located in the image comparison layer as having a posture difference with flawed attributes.
[0162] When the overlap ratio of any region in the image comparison layer is less than the preset threshold, it indicates that the group dance image frame corresponding to the overlap ratio of that region has a large difference from the standard pose image frame. At this time, the server will determine the group dance segment corresponding to the group dance image frame located in the image comparison layer as having a pose difference with incorrect attributes.
[0163] Finally, the server will determine the period of group dance movements corresponding to the group dance segments with posture differences as the posture difference period based on the group dance video corresponding to the group dance segments.
[0164] Step S106 includes the following:
[0165] When the management terminal interacts with the difference axis portion corresponding to any difference time period on any vertical difference axis based on the difference display image, it determines the vertical difference axis and the difference axis portion as the target difference axis and the target axis portion, and determines the target difference segment corresponding to the target axis portion based on the group dance video corresponding to the target difference axis.
[0166] For example, in this embodiment, when the management terminal views the difference display image, after interacting with the difference axis portion corresponding to any difference time period located on any vertical difference axis, the server will determine the vertical difference axis as the target difference axis and the interacted difference axis portion as the target axis portion. For example, when the management terminal interacts with the difference axis portion corresponding to the difference time period 00:00:30-00:00:32, the server will determine the difference axis portion as the target axis portion.
[0167] Subsequently, the server will retrieve the group dance video of the dancers corresponding to the target difference axis, and extract the video segment of the difference time period "00:00:30-00:00:32" corresponding to the target axis. This video segment is the target difference segment corresponding to the target axis, which makes it easy for the management terminal to intuitively view the target difference segments of the dancers corresponding to different difference time periods.
[0168] Step S108 includes the following:
[0169] Based on the remaining group dance videos, identify the group dance difference segments corresponding to the difference axis portions that have an axial overlap relationship with the target axis portions in each vertical difference axis, and add the target difference segments and each group dance difference segment to the second fill slots in the difference display image for filling each difference segment.
[0170] For example, in this embodiment, the server will determine the difference segments of each group dance based on the remaining group dance videos, which are located on each vertical difference axis and have an axial overlap with the target axis. For example, after determining that the difference time period corresponding to the target difference segment is 00:00:30-00:00:32, the server will traverse the vertical difference axis corresponding to each of the remaining group dancers and search for the difference axis portion whose time interval overlaps with the difference time period.
[0171] If there exists a vertical difference axis for one group dancer with a difference axis portion of "00:00:29-00:00:31", and another group dancer with a vertical difference axis portion of "00:00:30-00:00:33", then both of these difference axis portions have a time axis overlap relationship with the target axis portion. To avoid redundant retrieval of irrelevant segments, it can be set to only determine the axis overlap relationship when the intersection of the time intervals of the time axis overlap is greater than 50%.
[0172] Subsequently, the server will extract video clips from the corresponding time intervals of the group dance videos of the two group dancers as the difference segments of the group dance;
[0173] Finally, the server will add the target difference segment and the obtained group dance difference segments to the preset second fill slot in the difference display image, so that the management terminal can view the movement differences of multiple dancers at the same time.
[0174] This embodiment enables batch retrieval of different segments from multiple dancers within the same time period. This allows the management end to simultaneously compare the movement deviations of multiple dancers, analyze whether there are common problems, and improve the corresponding evaluation efficiency.
[0175] According to the present invention, the entire process of group dance movement evaluation is made intelligent and efficient, possessing significant practical value and technical advantages. On the one hand, based on the tracking and shooting of the acquisition unit and the generation of each movement axis, the dance videos of the lead dancer and group dancers are transformed into structured temporal movement data, realizing the quantitative identification of rhythm and posture differences. On the other hand, the slot-filling design of the vertical difference axis and the difference display image makes the difference periods of each group dancer clear at a glance, so that the management end does not need to review lengthy videos frame by frame, but can retrieve the target difference segment corresponding to each group dancer through simple interaction, greatly reducing the time cost of problem investigation. At the same time, it can batch retrieve and centrally display the difference segments of each group dance corresponding to the difference axis parts with axis overlap, helping the management end to quickly locate common movement problems of the group. The present invention can provide personalized correction basis for each dancer, significantly improving the efficiency and accuracy of group dance training and evaluation.
[0176] Another embodiment of the present invention provides a human motion evaluation system based on camera recognition. Figure 3 Its corresponding system block diagram includes:
[0177] The generation module is configured to track and film different dancers in each acquisition and shooting group based on each acquisition unit, and generate main motion axes and secondary motion axes based on the obtained lead dancer videos and group dance videos corresponding to the lead dancer and each group dancer.
[0178] The comparison module is configured to compare the lead dancer video with each group dance video to obtain the difference time periods corresponding to each group dance video, and add the vertical difference axis corresponding to each group dancer generated based on each difference time period to each first filling slot located in the difference display image for filling each vertical difference axis.
[0179] The interaction module is configured to, when the management terminal interacts with the difference axis portion corresponding to any difference time period on any vertical difference axis based on the difference display image, determine the vertical difference axis and the difference axis portion as the target difference axis and the target axis portion, and determine the target difference segment corresponding to the target axis portion based on the group dance video corresponding to the target difference axis.
[0180] The module is configured to determine, based on the remaining group dance videos, the group dance difference segments corresponding to the difference axis portions that have an axial overlap relationship with the target axis portion in each vertical difference axis, and add the target difference segment and each group dance difference segment to the second fill slots located in the difference display image for filling each difference segment.
[0181] In the specification provided herein, the algorithms and displays are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used with the examples of this invention. The required structure for constructing such systems is apparent from the above description. Furthermore, this invention is not directed to any particular programming language. It should be understood that the contents of the invention described herein can be implemented using various programming languages, and the above description of specific languages is for the purpose of disclosing preferred embodiments of the invention.
[0182] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0183] Similarly, it should be understood that, in order to streamline this disclosure and aid in understanding one or more of the various aspects of the invention, in the above description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof.
[0184] Those skilled in the art will understand that modules, units, or components of the devices disclosed in the examples herein can be arranged in the devices described in this embodiment, or alternatively, can be located in one or more devices different from the devices in this example. The modules in the foregoing examples can be combined into a single module or, in addition, can be divided into multiple sub-modules.
[0185] Those skilled in the art will understand that the modules in the device of the embodiment can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiment can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components.
[0186] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of the invention and form different embodiments.
[0187] Furthermore, some of the embodiments described herein are methods or combinations of method elements that can be implemented by a processor of a computer system or by other means of performing the functions. Therefore, a processor having the necessary instructions for implementing the methods or method elements forms means for implementing the methods or method elements. Furthermore, the elements described herein in the apparatus embodiments are examples of means for implementing the functions performed by elements for the purposes of carrying out the invention.
[0188] As used herein, unless otherwise specified, the use of ordinal numbers such as “first,” “second,” “third,” etc., to describe ordinary objects merely indicates different instances of similar objects and is not intended to imply that the objects being described must have a given order in time, space, ordering, or any other manner.
[0189] Although the invention has been described with respect to a limited number of embodiments, those skilled in the art will understand from the foregoing description that other embodiments are conceivable within the scope of the invention described herein. Furthermore, it should be noted that the language used in this specification has been chosen primarily for readability and edibility purposes, and not for the purpose of explaining or limiting the subject matter of the invention.
Claims
1. A method for evaluating human motion based on camera recognition, characterized in that, include: Based on each acquisition unit, different dancers located in each acquisition and shooting group are tracked and filmed, and the main motion axis and each secondary motion axis are generated based on the obtained lead dance video and group dance video corresponding to the lead dancer and each group dancer. The lead dancer's video was compared with each group dance video to identify the time periods of difference for each group dance video, including: Acquire the individual lead dancer image frames that make up the lead dancer video, and determine the corresponding lead dancer limb movements for each lead dancer image frame; Starting with the lead dancer image frame at the beginning, determine frame by frame whether the lead dancer's body movements are the same for all adjacent lead dancer image frames, and divide each lead dancer image frame that is arranged continuously and whose corresponding lead dancer body movements are the same into a lead dancer segment. Based on the lead dancer's video, determine the standard movement time periods for each lead dancer segment corresponding to different lead dancer's body movements; Based on each group dance video, determine each group dance segment corresponding to each standard movement time period, and determine the group dance limb movements corresponding to the first and last group dance image frames that make up each group dance segment. If no group dance limb movement is present in the first or last group dance image frame corresponding to any group dance segment, it is determined that the group dance segment has a beat difference with slow or fast beat attributes, and the beat difference time period corresponding to the group dance segment with beat difference is determined based on the group dance video corresponding to the group dance segment. Acquire the group dance image frames that make up each group dance video, and determine the group dance limb movements corresponding to each group dance image frame; Starting with the first group dance image frame, determine frame by frame whether the group dance limb movements corresponding to all adjacent group dance image frames are the same, and divide the group dance image frames that are arranged continuously and whose corresponding group dance limb movements are the same into a group dance segment. Based on the videos of each group dance, the time periods of each group dance movement in each group dance segment corresponding to different group dance body movements are determined; Each group dance segment is identified, and a preset standard segment corresponding to the group dance segment is retrieved based on the identification results; Each group dance segment is compared with the preset standard segment corresponding to the group dance segment, and the time periods of each group dance movement with posture differences are determined as the posture difference time periods based on the comparison results of each group dance video. And the vertical difference axis corresponding to each group of dancers generated based on each difference time period is added to each of the first fill slots in the difference display image for filling each vertical difference axis; When the management terminal interacts with the difference axis portion corresponding to any difference time period on any vertical difference axis based on the difference display image, it determines the vertical difference axis and the difference axis portion as the target difference axis and the target axis portion, and determines the target difference segment corresponding to the target axis portion based on the group dance video corresponding to the target difference axis. Based on the remaining group dance videos, identify the group dance difference segments corresponding to the difference axis portions that have an axial overlap relationship with the target axis portions in each vertical difference axis, and add the target difference segments and each group dance difference segment to the second fill slots in the difference display image for filling each difference segment.
2. The method according to claim 1, characterized in that, Each acquisition unit tracks and films different dancers located in each acquisition and shooting group, and generates main motion axes and secondary motion axes based on the obtained lead dancer videos and group dance videos corresponding to the lead dancer and each group dancer, including: If the light sensor installed on the stage surface outputs a light collection value that is less than the first preset light value at any given moment, that moment is determined as a waiting moment. When the light sensor outputs a light acquisition value greater than the second preset light value at any time after the waiting time, the moment is determined as the acquisition moment, and the upward viewing unit set above the stage is controlled to acquire an image of the stage to obtain a view of the stage. Image recognition is performed on the stage view to obtain various acquisition and shooting groups corresponding to different acquisition units, including different dancers, wherein the dancers include lead dancers and group dancers; Each acquisition unit tracks and films different dancers located in each acquisition and shooting group, and generates main motion axes and secondary motion axes based on the obtained lead dancer videos and group dance videos corresponding to the lead dancer and each group dancer.
3. The method according to claim 2, characterized in that, Image recognition is performed on the stage view to obtain various acquisition and shooting groups corresponding to different acquisition units, each including different dancers. The dancers include lead dancers and ensemble dancers. The stage view is divided based on the acquisition angle of each acquisition unit to obtain the acquisition area of each acquisition unit. Image recognition is performed on the view on the stage to determine the dancer area corresponding to each dancer in the view on the stage, wherein the dancers include lead dancers and group dancers; Each dancer region that completely overlaps with the same acquisition area is divided into the same region group, and each dancer pixel that makes up each dancer region is determined. Based on the stage view, connect the dancer pixels corresponding to each dancer area in the same area group to the acquisition center point of the acquisition unit corresponding to the area group, and connect the pixel points of each line segment that makes up each area connection line to obtain each connection area. Dancer areas whose corresponding connecting regions do not overlap with any other dancer areas are identified as valid areas, and dancers corresponding to valid areas are assigned to acquisition and shooting groups corresponding to acquisition units of acquisition areas that include the connecting regions.
4. The method according to claim 3, characterized in that, The method further includes: Based on the stage view, generate connection lines for each dancer, starting from the center point of each dancer's area and moving towards the center point of the corresponding acquisition unit, and determine the connection length corresponding to each connection line. If each acquisition and shooting group corresponding to multiple acquisition units includes the same dancer, the dancer connection line with the smallest connection length among the dancer connection lines corresponding to the dancer is determined as the first connection line; The dancers are removed from each acquisition unit's acquisition and shooting group based on each dancer's connection line except for the first connection line.
5. The method according to claim 3, characterized in that, Each acquisition unit tracks and films different dancers located in each acquisition and shooting group, and generates main motion axes and secondary motion axes based on the obtained lead dancer videos and group dance videos corresponding to the lead dancer and each group dancer, including: Each acquisition unit tracks and films different dancers located in each acquisition and shooting group to obtain each acquisition video; Image recognition is performed on the video image frames that are at the beginning of each video capture to obtain the human body regions located in each video image frame; Based on the stage view, each human body region that has the same regional position as each dancer region corresponding to the acquisition and shooting group of the acquisition unit is identified as a target region in each human body region. Determine the number of each target in each target region included in each video image frame, and copy each captured video based on the number of each target to obtain each captured video corresponding to each dancer in each target region. Based on each captured video, determine the corresponding lead dancer video and group dance video for each lead dancer and group dancer, and generate the main motion axis and the secondary motion axis based on each lead dancer video and group dance video.
6. The method according to claim 5, characterized in that, Based on the captured videos, corresponding lead dancer videos and group dance videos are determined for each lead dancer and each group dancer. Then, based on each lead dancer video and group dance video, primary motion axes and secondary motion axes are generated, including: Based on each video image frame that makes up each captured video, determine the contour of each target region corresponding to it, and generate the bounding rectangle of each contour corresponding to each region contour. Based on the preset magnification factor, the bounding rectangles of each contour are magnified, and based on the updated bounding rectangles of each contour, each video image frame is cropped. Face recognition is performed on the updated captured videos, and the recognition results are compared with the preset lead dancer information. The videos that show consistent information after comparison are identified as lead dance videos corresponding to the lead dancers, and the videos that show discrepancies after comparison are identified as group dance videos corresponding to the group dancers. The main motion axis and the secondary motion axis are generated based on the videos of each lead dancer and each group dance.
7. The method according to claim 1, characterized in that, Each group dance segment is compared with a preset standard segment corresponding to the group dance segment, and based on each group dance video, the time periods of each group dance movement that show posture differences are determined as posture difference time periods, including: Determine the sequence number of each group dance frame corresponding to each group dance image frame that makes up each group dance segment, and determine the sequence number of each posture image frame corresponding to each preset standard segment that makes up each group dance segment. Image recognition is performed on each group dance image frame and each standard posture image frame to obtain the group dance limb regions and each standard posture regions corresponding to the group dance limb movements and preset standard segments. Each group dance image frame that makes up each group dance segment is added to the retrieved image comparison layer, and based on each image comparison layer, the standard pose image that has the same frame number as the group dance image frame located in the image comparison layer is superimposed on the image comparison layer from each standard pose image frame corresponding to the group dance segment. The percentage of overlap between the limb regions of each group dance and the regions of each standard posture is determined based on the image comparison layers. When the area overlap ratio of any image comparison layer is within a preset ratio range or less than a preset ratio threshold, the group dance segment corresponding to the group dance image frame of the image comparison layer is determined to have a posture difference with flawed or incorrect attributes, and the group dance action time period corresponding to the group dance segment with posture difference is determined as the posture difference time period based on the group dance video corresponding to the group dance segment.
8. A human motion evaluation system based on camera recognition, characterized in that, include: The generation module is configured to track and film different dancers in each acquisition and shooting group based on each acquisition unit, and generate main motion axes and secondary motion axes based on the obtained lead dancer videos and group dance videos corresponding to the lead dancer and each group dancer. The comparison module is configured to compare the lead dancer's video with each group dance video to obtain the time periods of difference for each group dance video, including: Acquire the individual lead dancer image frames that make up the lead dancer video, and determine the corresponding lead dancer limb movements for each lead dancer image frame; Starting with the lead dancer image frame at the beginning, determine frame by frame whether the lead dancer's body movements are the same for all adjacent lead dancer image frames, and divide each lead dancer image frame that is arranged continuously and whose corresponding lead dancer body movements are the same into a lead dancer segment. Based on the lead dancer's video, determine the standard movement time periods for each lead dancer segment corresponding to different lead dancer's body movements; Based on each group dance video, determine each group dance segment corresponding to each standard movement time period, and determine the group dance limb movements corresponding to the first and last group dance image frames that make up each group dance segment. If no group dance limb movement is present in the first or last group dance image frame corresponding to any group dance segment, it is determined that the group dance segment has a beat difference with slow or fast beat attributes, and the beat difference time period corresponding to the group dance segment with beat difference is determined based on the group dance video corresponding to the group dance segment. Acquire the group dance image frames that make up each group dance video, and determine the group dance limb movements corresponding to each group dance image frame; Starting with the first group dance image frame, determine frame by frame whether the group dance limb movements corresponding to all adjacent group dance image frames are the same, and divide the group dance image frames that are arranged continuously and whose corresponding group dance limb movements are the same into a group dance segment. Based on the videos of each group dance, the time periods of each group dance movement in each group dance segment corresponding to different group dance body movements are determined; Each group dance segment is identified, and a preset standard segment corresponding to the group dance segment is retrieved based on the identification results; Each group dance segment is compared with the preset standard segment corresponding to the group dance segment, and the time periods of each group dance movement with posture differences are determined as the posture difference time periods based on the comparison results of each group dance video. And the vertical difference axis corresponding to each group of dancers generated based on each difference time period is added to each of the first fill slots in the difference display image for filling each vertical difference axis; The interaction module is configured to, when the management terminal interacts with the difference axis portion corresponding to any difference time period on any vertical difference axis based on the difference display image, determine the vertical difference axis and the difference axis portion as the target difference axis and the target axis portion, and determine the target difference segment corresponding to the target axis portion based on the group dance video corresponding to the target difference axis. The module is configured to determine, based on the remaining group dance videos, the group dance difference segments corresponding to the difference axis portions that have an axial overlap relationship with the target axis portion in each vertical difference axis, and add the target difference segment and each group dance difference segment to the second fill slots located in the difference display image for filling each difference segment.
Citation Information
Patent Citations
Dance posture action evaluation method and device, electronic equipment and storage medium
CN116978110A
Dance posture recognition and interaction method and system based on binocular data acquisition equipment
CN119007302A