Body measurement data evaluation method and system based on image acquisition
Through image acquisition and model training, a physical measurement data evaluation system was established, which solved the subjective problem of movement evaluation, achieved accurate and objective evaluation of students' movements, and provided an efficient feedback mechanism.
Patent Information
- Application Number
- CN202510612806.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-05-13
AI Technical Summary
In the existing technology, the evaluation of students' movements is subjective, which makes it difficult to improve the accuracy of the evaluation and leads to inaccurate feedback.
Through image acquisition technology, the body measurement data evaluation model is obtained and trained. The camera is used to collect training action video samples of different people. Combined with the body shape data, a key point detection and action node recognition model is established, and the standard and non-standard action loss functions are determined to achieve objective evaluation of the actions of the people being tested.
It improves the accuracy of movement evaluation, reduces subjectivity, provides students with objective and accurate feedback, and improves the objectivity and consistency of evaluation.
Smart Images

Figure CN120612731A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a body measurement data evaluation method and system based on image acquisition. Background Art
[0002] In related technologies, when students are training and testing certain movements (for example, gymnastics), coaches or referees usually evaluate the students' movements. However, the evaluation by coaches or referees usually contains certain human factors, which makes the evaluation subjective and difficult to improve the accuracy of the evaluation, thus making it difficult to provide accurate feedback to students. Summary of the Invention
[0003] The present invention provides a method and system for evaluating body measurement data based on image acquisition, which can solve the technical problems that the evaluation of personnel movements in related technologies is somewhat subjective, making it difficult to improve the accuracy of the evaluation and provide accurate feedback to trainees.
[0004] According to a first aspect of the present invention, there is provided a method for evaluating body measurement data based on image acquisition, comprising:
[0005] collecting a first training action video sample through a camera, wherein the first training action video sample is a video shot by a first person who has mastered the standard action at a specified position, and the content of the first training action video sample is the first person performing the specified action;
[0006] Capturing a second training action video sample using a camera, wherein the second training action video sample is a video shot of a second person who is not proficient in the standard action and is at a designated position, wherein the second training action video is of the second person performing the designated action, and the second person and the first person have different body shape data;
[0007] Training a physical measurement data evaluation model based on the first training action video sample, the second training action video sample, the body shape data of the first person, and the body shape data of the second person to obtain a trained physical measurement data evaluation model;
[0008] Acquire a video of the person to be tested's action to be tested through a camera, wherein the video of the action to be tested is a video shot with the person to be tested at a specified position, and the content of the video of the action to be tested is the person to be tested performing the specified action;
[0009] Obtaining the body shape data of the person to be tested;
[0010] The trained physical measurement data evaluation model is used to process the body shape data and the action video of the person to be tested to obtain the physical measurement evaluation information of the person to be tested.
[0011] According to the present invention, the first training video sample includes a standard training video of the first person performing a standard action, and a non-standard training video of the first person performing a non-standard action;
[0012] Training a physical measurement data evaluation model based on the first training action video sample, the second training action video sample, the body shape data of the first person, and the body shape data of the second person to obtain a trained physical measurement data evaluation model includes:
[0013] Determine the standard motion loss function of the physical measurement data evaluation model using the standard training video and the body shape data of the first person;
[0014] Determine a non-standard motion loss function of a physical measurement data evaluation model using the non-standard training video and the body shape data of the first person;
[0015] Determining a test loss function of a physical measurement data evaluation model based on a second training action video sample and the body shape data of a second person;
[0016] The physical measurement data evaluation model is trained according to the standard action loss function, the non-standard action loss function and the test loss function to obtain the trained physical measurement data evaluation model.
[0017] According to the present invention, a standard motion loss function of a physical measurement data evaluation model is determined using a standard training video and the physical shape data of a first person, including:
[0018] Processing the standard training video using a key point detection sub-model of the body measurement data evaluation model to obtain first position information of key points in a plurality of first training video frames in the standard training video;
[0019] Inputting first position information of key points corresponding to a plurality of first training video frames into the action node recognition sub-model to determine a first action node video frame;
[0020] Determining first node position information of a key point in a first action node video frame according to the first position information;
[0021] Inputting the first position information, the sequence number of the first action node video frame, the first key position information, and the body shape data of the first person into an evaluation sub-model of a physical measurement data evaluation model to obtain a first training physical measurement score;
[0022] According to the first training physical test score, the standard action loss function is determined.
[0023] According to the present invention, determining a standard action loss function based on the first training physical test score includes:
[0024] According to the formula
[0025] LOSSS =1-S1
[0026] Determine the standard action loss function LOSS S , where S1 is the first training physical test score.
[0027] According to the present invention, determining a non-standard motion loss function of a physical measurement data evaluation model using a non-standard training video and the physical shape data of a first person includes:
[0028] Processing the non-standard training video using a key point detection sub-model of the physical measurement data evaluation model to obtain second position information of key points in a plurality of second training video frames in the non-standard training video;
[0029] Inputting second position information of key points in the plurality of second training video frames into the action node recognition sub-model to determine a second action node video frame;
[0030] Determining second node position information of a key point in a second action node video frame according to the second position information;
[0031] Obtaining a first reference score for the non-standard training video according to the sequence number of the first action node video frame, the sequence number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information;
[0032] Get the first annotation score of the non-standard training video;
[0033] Obtaining a first control score for the non-standard training video according to the first reference score and the first labeled score;
[0034] Inputting the body shape data of the first person, the sequence number of the second action node video, the second node position information, and the second position information into the evaluation sub-model of the physical measurement data evaluation model to obtain a second training physical measurement score;
[0035] A non-standard action loss function of the physical measurement data evaluation model is determined based on the second training physical measurement score, the first control score, the serial number of the first action node video frame, the serial number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information.
[0036] According to the present invention, obtaining a first reference score of a non-standard training video according to the sequence number of the first action node video frame, the sequence number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information includes:
[0037] Obtaining a first action node sequence number vector according to the sequence number of the first action node video frame;
[0038] Obtaining a second action node sequence number vector according to the sequence number of the second action node video frame;
[0039] Obtain the similarity between the first action node sequence number vector and the first node sequence number vector of the second action node sequence number;
[0040] Set basic key points among multiple key points;
[0041] Determining a first node motion vector between the basic key point and other key points other than the basic key point according to the first node position information;
[0042] Determining a second node motion vector between the basic key point and other key points other than the basic key point according to the second node position information;
[0043] Determine an average of similarities between the first node action vector and the second node action vector corresponding to multiple other key points at multiple action nodes as the first node action similarity;
[0044] Uniformly sampling the first training video frames between the first action node video frames to obtain first sampled video frames, and uniformly sampling the second training video frames between the second action node video frames to obtain second sampled video frames;
[0045] determining, according to first position information of the first sampled video frame, a first motion vector between the basic key point and other key points other than the basic key point;
[0046] determining, according to second position information of the second sampled video frame, a second motion vector between the basic key point and other key points other than the basic key point;
[0047] Determine an average of similarities between the first motion vector and the second motion vector corresponding to multiple other key points at multiple sampling points as a first non-node motion similarity;
[0048] A first reference score is determined according to the first node sequence number similarity, the first node action similarity, and the first non-node action similarity.
[0049] According to the present invention, a non-standard motion loss function of a physical test data evaluation model is determined based on the second training physical test score, the first control score, the sequence number of the first action node video frame, the sequence number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information, including:
[0050] According to the formula
[0051]
[0052] Determine the non-standard motion loss function LOSS of the physical test data evaluation model N , where S 1,c is the first control score, S2 is the second training physical test score, sim SN,1 is the similarity of the first node number, sim S,1 is the first node action similarity, sim N,1 is the first non-node action similarity.
[0053] According to the present invention, determining the test loss function of the physical measurement data evaluation model using the second training action video sample and the body shape data of the second person includes:
[0054] Processing the second training action video sample using a key point detection sub-model of the physical measurement data evaluation model to obtain third position information of key points in a plurality of third training video frames in the second training action video sample;
[0055] Inputting third position information of key points in the plurality of third training video frames into the action node recognition sub-model to determine a third action node video frame;
[0056] Determining third node position information of a key point in a third action node video frame according to the third position information;
[0057] Obtaining a second reference score for the second training action video sample according to the sequence number of the first action node video frame, the sequence number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person, and the body shape data of the second person;
[0058] Obtain a second annotation score for the second training action video sample;
[0059] Obtaining a second control score for the second training action video sample based on the second reference score and the second labeled score;
[0060] Inputting the second person's body shape data, the sequence number of the third action node video, the third node position information, and the third position information into the evaluation sub-model of the physical measurement data evaluation model to obtain a third training physical measurement score;
[0061] Determine the test loss function of the physical measurement data evaluation model based on the third training physical measurement score, the second control score, the serial number of the first action node video frame, the serial number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person and the body shape data of the second person.
[0062] According to the present invention, obtaining a second reference score for a second training action video sample based on the sequence number of the first action node video frame, the sequence number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person, and the body shape data of the second person includes:
[0063] Obtaining a first action node sequence number vector according to the sequence number of the first action node video frame;
[0064] Obtaining a third action node sequence number vector according to the sequence number of the third action node video frame;
[0065] Obtaining the similarity of the second node sequence number of the first action node sequence number vector and the third action node sequence number vector;
[0066] Set basic key points among multiple key points;
[0067] Determining a first node motion vector between the basic key point and other key points other than the basic key point according to the first node position information;
[0068] Determining a third-node motion vector between the basic key point and other key points other than the basic key point according to the third-node position information;
[0069] Determining a first body shape vector based on the first body shape data, and determining a second body shape vector based on the second body shape data;
[0070] determining a body shape similarity based on the first body shape vector and the second body shape vector;
[0071] Determine the second node action similarity by taking the average value of the similarities between the first node action vector and the third node action vector corresponding to the plurality of other key points at the plurality of action nodes and the ratio of the similarities to the body shape similarity;
[0072] Uniformly sampling the first training video frames between the first action node video frames to obtain first sampled video frames, and uniformly sampling the third training video frames between the third action node video frames to obtain third sampled video frames;
[0073] determining, according to first position information of the first sampled video frame, a first motion vector between the basic key point and other key points other than the basic key point;
[0074] determining, according to third position information of the third sampled video frame, a third motion vector between the basic key point and other key points other than the basic key point;
[0075] Determine the second non-node action similarity by taking the ratio of the average value of the similarities between the first motion vector and the third motion vector corresponding to the plurality of other key points at the plurality of sampling points to the body shape similarity;
[0076] A second reference score is determined according to the second node sequence number similarity, the second node action similarity, and the second non-node action similarity.
[0077] According to a second aspect of the present invention, there is provided a body measurement data evaluation system based on image acquisition, comprising:
[0078] a first acquisition module configured to acquire a first training action video sample through a camera, wherein the first training action video sample is a video shot of a first person who has mastered a standard action at a specified position, and the content of the first training action video sample is the first person performing the specified action;
[0079] a second acquisition module configured to acquire a second training action video sample using a camera, wherein the second training action video sample is not a video of a second person who masters the standard action and is at a specified position, wherein the second training action video is a video of the second person performing the specified action, and the second person and the first person have different body shape data;
[0080] a training module, configured to train a physical measurement data evaluation model based on the first training action video sample, the second training action video sample, the body shape data of the first person, and the body shape data of the second person, to obtain a trained physical measurement data evaluation model;
[0081] A shooting module is used to obtain a video of the person to be tested performing a test action through a camera, wherein the video of the test action is a video shot when the person to be tested is at a specified position, and the content of the video of the test action is the person to be tested performing a specified action;
[0082] Body shape data module, used to obtain body shape data of the person to be tested;
[0083] The evaluation module is used to process the body shape data and the action video of the person to be tested through the trained physical measurement data evaluation model to obtain the physical measurement evaluation information of the person to be tested.
[0084] By adopting the above technical solution, the present invention can achieve the following technical effects:
[0085] According to the present invention, the physical test data evaluation model can be trained by different video samples of various types of people, so that the physical test data evaluation model can identify the movements of people of various body types, and then determine whether the movements of the person to be tested are standard, providing an objective and accurate basis for judging and evaluating the movements of the person to be tested, reducing subjectivity, improving evaluation accuracy, and providing accurate feedback for the person to be tested. When determining the standard action loss function, when the first training video sample is a standard training video of the first person performing a standard action, the first position information, the sequence number of the first action node video frame, the first key position information and the body shape data of the first person are obtained, and the standard action loss function is determined by the error between the theoretical value of the first training physical test score and the output value of the evaluation sub-model to improve the evaluation accuracy of the evaluation sub-model for the standard action. When determining the non-standard action loss function, the minimum value among the first node sequence number similarity, the first node action similarity and the first non-node action similarity can be used to determine the most important difference between the non-standard action and the standard action, and the error between the first control score and the second training physical test score can be specifically amplified, thereby specifically improving the training intensity of the evaluation sub-model for the evaluation of the most important difference and improving the accuracy of the evaluation sub-model. When determining the second control score, the body shape difference between the first person and the second person can be considered when determining the two-node action similarity and the second non-node action similarity. Therefore, by calculating the body shape similarity of the first person and the second person, the body shape difference can be eliminated. Only the similarity between the second person's action and the standard action at the action node and the similarity between the second person's action and the standard action during the action can be described, thereby improving the accuracy of action evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 A schematic flow chart of a method for evaluating body measurement data based on image acquisition according to an embodiment of the present invention is exemplarily shown;
[0087] Figure 2 A flowchart of training a physical measurement data evaluation model according to an embodiment of the present invention is exemplarily shown;
[0088] Figure 3 A block diagram of a body measurement data evaluation system based on image acquisition according to an embodiment of the present invention is exemplarily shown. DETAILED DESCRIPTION
[0089] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0090] Figure 1 A flow chart of a method for evaluating body measurement data based on image acquisition according to an embodiment of the present invention is exemplarily shown. The method includes:
[0091] Step S1, capturing a first training action video sample through a camera, wherein the first training action video sample is a video shot of a first person who has mastered a standard action at a specified position, and the content of the first training action video sample is the first person performing the specified action;
[0092] Step S2, capturing a second training action video sample using a camera, wherein the second training action video sample is not a video of a second person who masters the standard action and is at a designated position, the second training action video is about the second person performing the designated action, and the second person and the first person have different body shape data;
[0093] Step S3, training a physical measurement data evaluation model based on the first training action video sample, the second training action video sample, the body shape data of the first person, and the body shape data of the second person to obtain a trained physical measurement data evaluation model;
[0094] Step S4, obtaining a video of the person to be tested performing the action to be tested by a camera, wherein the video of the action to be tested is a video shot with the person to be tested at a specified position, and the content of the video of the action to be tested is the person to be tested performing the specified action;
[0095] Step S5, obtaining the body shape data of the person to be measured;
[0096] Step S6: Processing the body shape data of the person to be measured and the video of the action to be measured by the trained body measurement data evaluation model to obtain body measurement evaluation information of the person to be measured.
[0097] According to the physical measurement data evaluation method based on image acquisition in an embodiment of the present invention, the physical measurement data evaluation model can be trained through different video samples of various types of people, so that the physical measurement data evaluation model can recognize the movements of people of various body types, and then judge whether the movements of the people to be tested are standard, which provides an objective and accurate basis for judging and evaluating the movements of the people to be tested, reduces subjectivity, improves evaluation accuracy, and can provide accurate feedback to the people to be tested.
[0098] According to one embodiment of the present invention, in step S1, the designated position can be a point or an area, and the camera's field of view includes the designated position. In the example, the designated position is an area, and the distance and angle between the camera and the centroid of the area are fixed values. For example, the initial position of the person when performing the designated action is located at the centroid of the area and faces the lens of the camera. The designated action is an action that the person to be tested needs to practice or perform, such as gymnastics, etc. The standard action is a designated action when there is no action error. The first person can be a coach, etc. The first person can be proficient in the standard action. The first person can be at the designated position and use the camera to shoot a video of the first person performing the standard action. The first person can also shoot a video with action errors as a sample. For example, some common wrong actions can be shot to provide students with attention to avoidance. The video sample of the first person taken by the camera is the first training action video sample. The first training action video sample may include two types of first training video samples, including standard training videos of the first person performing standard actions and non-standard training videos of the first person performing non-standard actions. Non-standard actions are actions with action errors, such as the above-mentioned erroneous actions, etc. The two types of first training video samples are used to expand the number and diversity of samples, which helps to improve the performance of the physical measurement data evaluation model during training.
[0099] According to one embodiment of the present invention, in step S2, the second person is a person who does not master the standard movements, for example, a trainee, and there may be some errors in the movements of the second person. In addition, the body shape data of the second person is different from the body shape data of the first person. The video of the second person performing the specified movement taken by the camera is used as a sample (i.e., the second training movement video sample). This can improve the recognition accuracy of the physical measurement data evaluation model for movements performed by people of various body shapes during training, help reduce the possibility of a decrease in the accuracy of the physical measurement data evaluation model due to movement and posture errors caused by body shape differences, and help improve the robustness of the physical measurement data evaluation model.
[0100] According to one embodiment of the present invention, in step S3, the physical measurement data evaluation model is a combination of multiple sub-models, including a key point detection sub-model, an action node recognition sub-model, and an evaluation sub-model. The key point detection sub-model and the action node recognition sub-model can be pre-trained, and the evaluation sub-model is trained using the trained key point detection sub-model and the trained action node recognition sub-model, as well as the first training video sample and the second training action video sample, so that the physical measurement data evaluation model has the ability to accurately evaluate the accuracy of the person's movements based on the action video and body shape data.
[0101] Figure 2 The flowchart of training a physical measurement data evaluation model according to an embodiment of the present invention is exemplarily shown.
[0102] According to one embodiment of the present invention, in step S3, the physical measurement data evaluation model is trained according to the first training action video sample, the second training action video sample, the body shape data of the first person and the body shape data of the second person to obtain the trained physical measurement data evaluation model, including: step S31, determining the standard action loss function of the physical measurement data evaluation model through the standard training video and the body shape data of the first person; step S32, determining the non-standard action loss function of the physical measurement data evaluation model through the non-standard training video and the body shape data of the first person; step S33, determining the test loss function of the physical measurement data evaluation model through the second training action video sample and the body shape data of the second person; step S34, training the physical measurement data evaluation model according to the standard action loss function, the non-standard action loss function and the test loss function to obtain the trained physical measurement data evaluation model.
[0103] According to one embodiment of the present invention, in step S31, the standard training video is a video sample taken when a first person who has mastered the standard action performs the action without error, which can be used to enable the physical measurement data evaluation model to obtain the data of the standard action, so as to judge whether the action of the person to be tested is standard based on the data. The standard training video and the body shape data of the first person are used to determine the standard action loss function of the physical measurement data evaluation model, including: processing the standard training video through the key point detection submodel of the physical measurement data evaluation model to obtain the first position information of the key points in multiple first training video frames in the standard training video; inputting the first position information of the key points corresponding to the multiple first training video frames into the action node recognition submodel to determine the first action node video frame; determining the first node position information of the key points in the first action node video frame based on the first position information; inputting the first position information, the sequence number of the first action node video frame, the first key position information and the body shape data of the first person into the evaluation submodel of the physical measurement data evaluation model to obtain a first training physical measurement score; and determining the standard action loss function based on the first training physical measurement score.
[0104] According to one embodiment of the present invention, a standard training video can be parsed to obtain multiple first training video frames of the standard training video, and the first position information of the key points of the first person in each first training video frame can be detected by a key point detection sub-model. The key point detection sub-model is a convolutional neural network model, which can be used to detect the position coordinates of multiple key points of the body of the person in the image. The multiple key points may include head key points, neck key points, shoulder key points, elbow key points, hand key points, waist key points, knee key points, ankle key points, etc. The posture of the first person in each first training video frame can be determined based on the first position information of each key point, and the posture change of the first person can be determined based on the timestamp of the first training video frame, thereby analyzing the movement of the first person.
[0105] According to one embodiment of the present invention, the first position information of the key points in the first training video frame can be arranged in a sequence according to the order of the first training video frame, and the sequence of the first position information is input into the action node recognition sub-model to determine the first action node video frame. The action node recognition sub-model is a BP neural network model, which can identify the action nodes in the action process based on the position of the key points. For example, when performing an action, before reaching the action node, the movement direction of each key point does not change significantly, but after reaching the action node, the movement direction of at least one key point changes significantly. For example, in the action of raising an arm to a certain height and then lowering it, before the action node (i.e., before the arm is raised to a certain height), the curves of the hand key points and the elbow key points are relatively smooth, that is, the hand key points and the elbow key points rise along a smooth trajectory without a large angle, and after reaching the action node, the movement trajectories of the hand key points and the elbow key points descend. Therefore, the trajectories of the hand key points and the elbow key points before reaching the action node and after reaching the action node have a large angle (e.g., greater than or equal to 90°). The action node recognition sub-model can monitor the motion trajectory of each key point, and when the motion trajectory of at least one key point appears at a larger angle, it records the moment and determines the first training video frame corresponding to the moment, which is the first action node video frame. In the example, the action node recognition sub-model can calculate the motion vector between the first position information of the key point in adjacent first training video frames. If there is at least one key point, the angle between the motion vector between the first position information in the A-1th first training video frame and the first position information in the Ath first training video frame, and the angle between the motion vector between the first position information in the Ath first training video frame and the first position information in the A+1th first training video frame is greater than or equal to 90°, then the Ath first training video frame is the first action node video frame, where A is a positive integer and A is less than the total number of first training video frames.
[0106] According to one embodiment of the present invention, when performing a specified action, the action node is a node that is about to change its motion state. The posture at the action node has a significant impact on the standardization and aesthetics of the overall action. For example, when performing gymnastics or other movements, the posture at the action node can affect the integrity and stretchability of the movement. Therefore, after determining the first action node video frame, the first node position information of the key point in the first action node video frame can be determined, and then the posture at the action node can be determined. When performing a standard action, this posture is the standard posture.
[0107] According to one embodiment of the present invention, the first position information, the serial number of the first action node video frame, the first key position information and the physical data of the first person (for example, height, arm span, leg length and other data) can be input into the evaluation sub-model of the physical measurement data evaluation model to obtain the first training physical measurement score. The physical measurement data evaluation model is a BP neural network model, which can process the above-mentioned multiple data and output the first training physical measurement score of the standard training video. In theory, the first training physical measurement score is 1, that is, the standard training video is a video shot when the first person performs a standard action, so the first training physical measurement score is theoretically full marks. However, there may be errors in the evaluation sub-model, resulting in the first training physical measurement score output by the evaluation sub-model not being equal to 1.
[0108] According to one embodiment of the present invention, determining the standard action loss function according to the first training body measurement score includes: determining the standard action loss function LOSS according to formula (1) S ,
[0109] LOSS S =1-S1(1)
[0110] Among them, S1 is the physical test score of the first training.
[0111] According to one embodiment of the present invention, as described above, the first training body measurement score is theoretically 1, and the first training body side score output by the evaluation sub-model may have an error. Therefore, the error 1-S1 is used as the standard action loss function, and the loss function is back-propagated during training to reduce the error, so that the evaluation sub-model can output a first training body side score that is closer to the theoretical value 1 when processing the first position information corresponding to the standard action, the serial number of the first action node video frame, the first key position information and the body shape data of the first person, so as to improve the accuracy of the evaluation sub-model.
[0112] According to one embodiment of the present invention, the first person may include multiple people. When the multiple first persons perform standard actions respectively, the first training action video samples can be shot respectively, and the first position information, the serial number of the first action node video frame, the first key position information and the body shape data of the first person can be obtained respectively, and the above training can be iteratively performed to improve the accuracy of the evaluation sub-model.
[0113] In this way, when the first training video sample is a standard training video of the first person performing a standard action, the first position information, the serial number of the first action node video frame, the first key position information and the body shape data of the first person can be obtained, and the standard action loss function can be determined by the error between the theoretical value of the first training physical test score and the output value of the evaluation sub-model to improve the evaluation accuracy of the evaluation sub-model for the standard action.
[0114] According to one embodiment of the present invention, in step S32, the non-standard training video is a video sample shot when the first person performs a non-standard action, for example, a video sample shot when the first person demonstrates a common mistake.
[0115] According to one embodiment of the present invention, for non-standard actions, that is, actions with errors during execution, corresponding points can be deducted for these errors during the training process, and referees or coaches can give expert scores to non-standard actions. Video samples of non-standard actions can also be analyzed to automatically obtain scores, and the expert scores and automatically obtained scores are combined to obtain a first labeled score for training the evaluation sub-model, which is then compared with the score output by the evaluation sub-model to determine the non-standard action loss function.
[0116] According to one embodiment of the present invention, in step S22, a non-standard action loss function of a physical measurement data evaluation model is determined by using a non-standard training video and the body shape data of a first person, including: processing the non-standard training video by a key point detection sub-model of the physical measurement data evaluation model to obtain second position information of key points in a plurality of second training video frames in the non-standard training video; inputting the second position information of key points in the plurality of second training video frames into an action node recognition sub-model to determine a second action node video frame; determining the second node position information of the key points in the second action node video frame according to the second position information; determining the second node position information of the key points in the second action node video frame according to the serial number of the first action node video frame, the serial number of the second action node video, the first node position information, the second node position information, the second node position information, the second action node video frame ... The first position information and the second position information are used to obtain a first reference score of the non-standard training video; a first annotation score of the non-standard training video is obtained; according to the first reference score and the first annotation score, a first comparison score of the non-standard training video is obtained; the body shape data of the first person, the serial number of the second action node video, the second node position information and the second position information are input into the evaluation sub-model of the physical measurement data evaluation model to obtain the second training physical measurement score; according to the second training physical measurement score, the first comparison score, the serial number of the first action node video frame, the serial number of the second action node video, the first node position information, the second node position information, the first position information and the second position information, a non-standard action loss function of the physical measurement data evaluation model is determined.
[0117] According to one embodiment of the present invention, similar to the method for obtaining the first position information, multiple second training video frames of the non-standard training video can be processed using a key point detection sub-model to obtain second position information of key points in each second training video frame. Similar to the first action node video frame, the second position information of key points in multiple second training video frames can be processed using an action node identification sub-model to obtain second action node video frames, and the second node position information of key points in the second action node video frames can be obtained. The specific processing process will not be repeated here.
[0118] According to one embodiment of the present invention, the first reference score is a score obtained by analyzing the difference in action and posture between the second training video frame and the first training video frame. According to the serial number of the first action node video frame, the serial number of the second action node video, the first node position information, the second node position information, the first position information and the second position information, the first reference score of the non-standard training video is obtained, including: according to the serial number of the first action node video frame, obtaining the first action node serial number vector; according to the serial number of the second action node video frame, obtaining the second action node serial number vector; obtaining the first node serial number similarity between the first action node serial number vector and the second action node serial number vector; setting a basic key point among multiple key points; according to the first node position information, determining the first node action vector between the basic key point and other key points except the basic key point; according to the second node position information, determining the second node action vector between the basic key point and other key points except the basic key point; and assigning the first node action vector corresponding to multiple other key points at multiple action nodes to the second node position information. The first non-node action similarity is determined by taking the average value of the similarity between the first training video frame and the second node action vector; uniformly sampling the first training video frame between the first action node video frame to obtain the first sampling video frame, and uniformly sampling the second training video frame between the second action node video frame to obtain the second sampling video frame; determining the first action vector between the basic key point and other key points other than the basic key point according to the first position information of the first sampling video frame; determining the second action vector between the basic key point and other key points other than the basic key point according to the second position information of the second sampling video frame; determining the average value of the similarity between the first action vector and the second action vector corresponding to multiple other key points at multiple sampling points as the first non-node action similarity; determining the first reference score according to the first node sequence number similarity, the first node action similarity and the first non-node action similarity.
[0119] According to one embodiment of the present invention, as described above, the posture at the action node has a greater impact on the standardization and aesthetics of the overall action. Therefore, the first action node video frame and the second action node video frame are both relatively important data. The sequence number of the first action node video frame and the sequence number of the second action node video frame are also relatively important data. The time interval between adjacent first action node video frames is a fixed value. Therefore, the gap between the sequence numbers of adjacent first action node video frames can reflect the duration of the execution process of the action between adjacent action nodes, and can also reflect the execution speed of the action between adjacent action nodes. Similarly, the gap between the sequence numbers of adjacent second action node video frames can reflect the duration of the execution process of the action between adjacent action nodes, and can also reflect the execution speed of the action between adjacent action nodes. Therefore, the sequence number of the first action node video frame and the sequence number of the second action node video frame can be used to determine the consistency of the execution speed of each action, which can be used to determine whether the execution speed of non-standard actions and standard actions is similar.
[0120] According to one embodiment of the present invention, the sequence numbers of the first action node video frames can be combined to obtain a first action node sequence number vector, that is, each component in the first action node sequence number vector is the sequence number of the first action node video frame. The sequence numbers of the second action node video frames can also be combined to obtain a second action node sequence number vector, and each component in the second action node sequence number vector is the sequence number of the second action node video frame. If the dimensions of the first action node sequence number vector and the second action node sequence number vector are different (for example, the difference between standard action and non-standard action is large), then zero padding is performed to make the dimensions of the two vectors after zero padding the same. For example, if the dimension of the first action node sequence number vector is greater than the dimension of the second action node sequence number vector, then zero padding is performed at the end of the second action node sequence number vector so that the dimension of the second action node sequence number vector after zero padding is equal to the dimension of the first action node sequence number vector. For another example, if the dimension of the first action node sequence number vector is less than the dimension of the second action node sequence number vector, then zero padding is performed at the end of the first action node sequence number vector so that the dimension of the first action node sequence number vector after zero padding is equal to the dimension of the second action node sequence number vector. After performing the above dimensional alignment process, the first node number similarity between the first action node number vector and the second action node number vector can be determined. For example, the cosine similarity between the first action node number vector and the second action node number vector can be calculated as the first node number similarity. The first node number similarity can be used to determine whether the execution speed of non-standard actions and standard actions is similar.
[0121] According to one embodiment of the present invention, the basic keypoint can be a neck keypoint or a waist keypoint. The present invention does not limit the setting of the basic keypoint. At the action node, a first-node action vector is determined based on the first-node position information. For example, the basic keypoint points to the first-node action vector of other keypoints. Similarly, a second-node action vector can be determined.
[0122] According to one embodiment of the present invention, the first node position information can be arranged in sequence, and the second node position information can be arranged in sequence, and then the first node position information and the second node position information with the same order can be matched. If the number of first node position information and second node position information is different (for example, the standard action and the non-standard action are quite different), a zero-padding process similar to the above can be performed. For example, if the number of first node position information is more than the number of second node position information, the second node position information can be zero-padding. For example, if the number of first node position information is 20 and the number of second node position information is 18, the second node position information can be zero-padding, and the 19th and 20th second node position information can all be set to (0, 0), that is, the coordinates of all key points are set to (0, 0), and then the second node action vector is solved.
[0123] According to one embodiment of the present invention, the similarity between the first node action vector corresponding to the first node position information and the second node action vector corresponding to the second node position information with the same ranking can be solved. For example, the cosine similarity between the first node action vector between the head key point and the neck key point (basic key point) in the first first node position information and the first node action vector between the head key point and the neck key point in the first second node position information is determined; the cosine similarity between the first node action vector between the shoulder key point and the neck key point in the first first node position information and the first node action vector between the shoulder key point and the neck key point in the first second node position information is determined... and these cosine similarities are averaged to determine the action similarity between the first first node position information and the first second node position information. In a similar manner to the above, the action similarity between the second first node position information and the second second node position information, the action similarity between the third first node position information and the third second node position information can be obtained... and then the action similarities can be averaged to obtain the first node action similarity, which can be used to describe the similarity between standard actions and non-standard actions when describing action nodes.
[0124] According to one embodiment of the present invention, there may be multiple first training video frames between adjacent first action node video frames, and these first training video frames may be uniformly sampled. For example, there are 8 first training video frames between adjacent first action node video frames, and these 8 first training video frames may be uniformly sampled, for example, the 3rd and 6th first training video frames are taken. Since the speed of non-standard actions may be inconsistent with the speed of standard actions, the number of second training video frames between second action node video frames with the same sequence number may be different from the number of first training video frames. For example, the number of second training video frames between adjacent second action node video frames is 5, then when uniform sampling is performed, the 2nd and 4th second training video frames may be selected as second sampled video frames. The above sampling quantity and sampling interval are only examples. When sampling, it is only necessary to make the sampling quantity of the first training video frames between adjacent first action node video frames equal to the sampling quantity of the second training video frames between second action node video frames with the same sequence number. The first sampled video frames between adjacent first action node video frames and the second sampled video frames between second action node video frames with the same sequence number are in a one-to-one correspondence.
[0125] According to one embodiment of the present invention, in the first sampling video frame, based on the first position information, the first motion vector between the basic key point and other key points other than the basic key point is determined. The determination method is similar to the first node motion vector and will not be repeated here. Similarly, the second motion vector in the second sampling video frame can be determined. The cosine similarity between each first motion vector in the first sampling video frame and the corresponding second motion vector in the corresponding second sampling video frame can be determined, and the cosine similarity between each first motion vector and the second motion vector can be averaged to obtain the motion similarity of the person in the first sampling video frame and the second sampling video frame. Furthermore, the motion similarity of the person in each first sampling video frame and the corresponding second sampling video frame can be averaged to obtain the first non-node motion similarity, which can be used to describe the similarity between standard motion and non-standard motion when the action node is not reached (or when the action is in progress).
[0126] According to one embodiment of the present invention, the first node sequence number similarity, the first node action similarity and the first non-node action similarity can be weighted averaged to determine the first reference score. The first reference score can be combined with the consistency of action speed between action nodes, the action similarity at the action node and the action similarity during the action to comprehensively evaluate the similarity between non-standard actions and standard actions, thereby improving the accuracy and comprehensiveness of the evaluation.
[0127] According to one embodiment of the present invention, the first reference score obtained above based on the analysis of the first and second training video frames can also be manually scored to obtain a first annotated score. For example, this can be performed by professionals such as referees or coaches to obtain the first annotated score. Furthermore, a weighted average of the first annotated score and the first reference score can be taken to obtain a first comparison score for the non-standard training video. This allows for combining expert experience with objective analysis of the video frames to obtain a comprehensive first comparison score, improving its accuracy and objectivity for use in training the evaluation sub-model.
[0128] According to one embodiment of the present invention, the evaluation sub-model may evaluate a video of a first person performing a non-standard action. For example, the evaluation sub-model may input the first person's body shape data, the sequence number of the second action node video, the second node position information, and the second position information into the evaluation sub-model to obtain a second training physical fitness score. The second training physical fitness score may contain errors, and the first control score may be used as the accurate score to determine the error of the second training physical fitness score.
[0129] According to one embodiment of the present invention, a non-standard motion loss function of a physical test data evaluation model is determined based on the second training physical test score, the first control score, the sequence number of the first action node video frame, the sequence number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information, including: determining the non-standard motion loss function LOSS of the physical test data evaluation model according to formula (2) N ,
[0130]
[0131] Among them, S 1,c is the first control score, S2 is the second training physical test score, sim SN,1 is the similarity of the first node number, sim S,1 is the first node action similarity, sim N,1 is the first non-node action similarity.
[0132] According to one embodiment of the present invention, in formula (2), |S 1,c -S2| is the error between the first control score and the second training physical test score. During the training process, this error can be reduced, thereby improving the accuracy of the evaluation sub-model of the physical test data evaluation model. min(sim SN,1 ,sim S,1 ,sim N,1) is the minimum value among the first node number similarity, the first node action similarity and the first non-node action similarity. Since the first node number similarity, the first node action similarity and the first non-node action similarity are all values less than or equal to 1, the minimum value of the three can be used as the denominator to amplify the error between the first control score and the second training physical test score. Moreover, the most important difference between the non-standard action and the standard action can be determined by the minimum value among the first node number similarity, the first node action similarity and the first non-node action similarity. For example, if the minimum value of the three is the first node number similarity, the most important difference is the difference in action speed. If the minimum value of the three is the first node action similarity, the most important difference is the difference in action similarity at the action node. If the minimum value of the three is the first non-node action similarity, the most important difference is the difference in action similarity during the action process. Based on the most important difference, the error between the first control score and the second training physical test score can be targetedly amplified, thereby targetedly improving the evaluation accuracy of the evaluation sub-model for the most important difference, thereby improving the overall evaluation accuracy. During the training process, the above non-standard action loss function can be back-propagated to adjust the accuracy of the evaluation sub-model, and can also be targeted to improve the evaluation accuracy of the evaluation sub-model for the most important differences.
[0133] In this way, the most important difference between non-standard actions and standard actions can be determined by the minimum value of the first node sequence number similarity, the first node action similarity and the first non-node action similarity, and the error between the first control score and the second training physical test score can be amplified in a targeted manner, thereby specifically improving the training intensity of the evaluation sub-model for the evaluation of the most important differences and improving the accuracy of the evaluation sub-model.
[0134] According to one embodiment of the present invention, in step S33, in order to adapt the physical measurement data evaluation model to the evaluation of individuals of various body types and improve its applicability and robustness, a test loss function may be obtained using a second training action video sample of a second individual and the body shape data of the second individual to train the physical measurement data evaluation model. The second individual may include multiple individuals, each having different body shape data than the first individual, and each of the second individuals having different levels of mastery of the designated action.
[0135] According to one embodiment of the present invention, in step S33, the test loss function of the physical measurement data evaluation model is determined by the second training action video sample and the body shape data of the second person, including: processing the second training action video sample by the key point detection submodel of the physical measurement data evaluation model to obtain the third position information of the key points in multiple third training video frames in the second training action video sample; inputting the third position information of the key points in multiple third training video frames into the action node recognition submodel to determine the third action node video frame; determining the third node position information of the key points in the third action node video frame according to the third position information; determining the third node position information of the key points in the third action node video frame according to the serial number of the first action node video frame, the serial number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the first person The second person's body shape data and the second person's body shape data are input into the evaluation sub-model of the physical measurement data evaluation model to obtain the third training physical measurement score; the test loss function of the physical measurement data evaluation model is determined according to the third training physical measurement score, the second control score, the serial number of the first action node video frame, the serial number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person and the body shape data of the second person.
[0136] According to one embodiment of the present invention, the method for obtaining the third position information, the third action node video frame and the third node position information is similar to the method for obtaining the first position information, the first action node video frame and the first node position information mentioned above, and will not be repeated here.
[0137] According to one embodiment of the present invention, a second reference score of a second training action video sample is obtained based on the serial number of the first action node video frame, the serial number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person and the body shape data of the second person, including: obtaining a first action node serial number vector according to the serial number of the first action node video frame; obtaining a third action node serial number vector according to the serial number of the third action node video frame; obtaining a second node serial number similarity between the first action node serial number vector and the third action node serial number vector; setting a basic key point among multiple key points; determining a first node action vector between the basic key point and other key points except the basic key point according to the first node position information; determining a third node action vector between the basic key point and other key points except the basic key point according to the third node position information; determining a first body shape vector according to the first body shape data, and determining a second body shape vector according to the second body shape data; determining a first body shape vector according to the first body shape vector and the second body shape vector shape vector, determine the body shape similarity; determine the second node action similarity by the average value of the similarities between the first node action vector and the third node action vector corresponding to multiple other key points at multiple action nodes and the ratio of the body shape similarity; uniformly sample the first training video frame between the first action node video frames to obtain the first sampling video frame, and uniformly sample the third training video frame between the third action node video frames to obtain the third sampling video frame; determine the first action vector between the basic key point and the other key points other than the basic key point according to the first position information of the first sampling video frame; determine the third action vector between the basic key point and the other key points other than the basic key point according to the third position information of the third sampling video frame; determine the second non-node action similarity by the average value of the similarities between the first action vector and the third action vector corresponding to the multiple other key points at the multiple sampling points and the ratio of the body shape similarity; determine the second reference score according to the second node sequence number similarity, the second node action similarity and the second non-node action similarity.
[0138] According to one embodiment of the present invention, the method for obtaining the second node number similarity is similar to the method for obtaining the first node number similarity, and the meaning is also similar, indicating the similarity between the second person's movement speed and the standard movement speed.
[0139] According to one embodiment of the present invention, the method for obtaining the first node motion vector is the same as above, and the method for obtaining the third node motion vector is similar to the method for obtaining the second node motion vector above, and will not be repeated here. However, the difference between the first node motion vector and the third node motion vector, in addition to the difference in motion, also includes the difference in body shape between the first person and the second person. Therefore, when judging whether the second person's motion is standard, the interference of the body shape difference can be eliminated, and only whether the second person's motion is different from the standard motion can be judged. Therefore, the first body shape vector can be determined based on the first body shape data. For example, the first body shape data such as height, arm span, and leg length can be combined to obtain the first body shape vector, and each component of the first body shape vector is a type of first body shape data. Similarly, the second body shape vector can be determined based on the second body shape data. Furthermore, the cosine similarity of the first body shape vector and the second body shape vector can be determined as the body shape similarity, and when determining the second node action similarity, the action similarity of the first node position information and the third node position information with the same serial number can be determined in a similar manner to the above, and the average value of each action similarity can be obtained, and the ratio of the average value of the action similarity to the body shape similarity can be obtained, thereby excluding the body shape difference and obtaining the second node action similarity, which can be used only to describe the similarity between the second person's action and the standard action at the action node.
[0140] According to one embodiment of the present invention, the first sampling video frame and the first motion vector are obtained in the same manner as above, and the third sampling video and the third motion vector are obtained in a similar manner to the second sampling video and the second motion vector, which will not be repeated here. Similarly, body shape differences can be excluded, and only the similarity between the second person's action and the standard action during the action is obtained. The action similarity of the person in each first sampling video frame and the corresponding third sampling video frame can be obtained in a manner similar to the above, and averaged. Further, the ratio of the average value to the body shape similarity is determined as the second non-node action similarity, thereby excluding body shape differences and only used to describe the similarity between the second person's action and the standard action during the action.
[0141] According to one embodiment of the present invention, the second node sequence number similarity, the second node action similarity and the second non-node action similarity can be weighted averaged to obtain a second reference score, and then weighted averaged with the second labeled score obtained based on the expert experience of the coach or referee to obtain a second control score as an accurate score.
[0142] According to one embodiment of the present invention, the evaluation sub-model can process the body shape data of the second person, the serial number of the third action node video, the third node position information and the third position information, and output a third training physical fitness score, which may contain errors.
[0143] According to one embodiment of the present invention, a test loss function can be obtained in a manner similar to formula (2), that is, the error between the third training body test score and the second control score is solved, and the minimum value among the second node sequence number similarity, the second node action similarity, and the second non-node action similarity is solved, and then the ratio of the error to the minimum value is used as the test loss function to specifically improve the evaluation accuracy of the evaluation sub-model for the most important differences.
[0144] In this way, the body shape difference between the first person and the second person can be considered when determining the similarity of two-node actions and the similarity of the second non-node action. Therefore, by calculating the body shape similarity of the first person and the second person, the body shape difference can be eliminated. Only the similarity between the second person's action and the standard action at the action node and the similarity between the second person's action and the standard action during the action can be described, thereby improving the accuracy of action evaluation.
[0145] According to one embodiment of the present invention, in step S34, the physical measurement data evaluation model can be trained using the standard action loss function, the non-standard action loss function and the test loss function, that is, one of the three loss functions can be selected based on the data input into the physical measurement data evaluation model, and the parameters of the evaluation sub-model can be adjusted so that all three loss functions are reduced, and after multiple trainings, the trained physical measurement data evaluation model is obtained.
[0146] According to one embodiment of the present invention, in step S4, a video of the person to be tested performing the action to be tested can be captured by a camera, and in step S5, the body shape data of the person to be tested can be obtained. Then, in step S6, the body shape data and the video of the person to be tested can be processed by the trained body measurement data evaluation model to obtain a body measurement score for the person to be tested, and body measurement evaluation information can be determined based on the body measurement score. For example, the body measurement score can be compared with a scoring threshold. If the body measurement score is higher than or equal to the scoring threshold, the body measurement evaluation information is qualified; otherwise, it is unqualified. The present invention does not limit the specific content of the body measurement evaluation information.
[0147] According to the physical measurement data evaluation method based on image acquisition of an embodiment of the present invention, the physical measurement data evaluation model can be trained through different video samples of various types of people, so that the physical measurement data evaluation model can recognize the movements of people of various body shapes, and then judge whether the movements of the person to be tested are standard, providing an objective and accurate basis for judging and evaluating the movements of the person to be tested, reducing subjectivity, improving evaluation accuracy, and being able to provide accurate feedback for the person to be tested. When determining the standard action loss function, when the first training video sample is a standard training video of the first person performing a standard action, the first position information, the serial number of the first action node video frame, the first key position information and the body shape data of the first person can be obtained, and the standard action loss function can be determined by the error between the theoretical value of the first training physical measurement score and the output value of the evaluation sub-model to improve the evaluation accuracy of the evaluation sub-model for the standard action. When determining the non-standard action loss function, the most important difference between the non-standard action and the standard action can be determined by taking the minimum value of the first node sequence number similarity, the first node action similarity, and the first non-node action similarity, and the error between the first control score and the second training physical test score can be specifically amplified, thereby specifically improving the training intensity of the evaluation sub-model for the evaluation of the most important difference and improving the accuracy of the evaluation sub-model. When determining the second control score, the body shape difference between the first person and the second person can be considered when determining the two-node action similarity and the second non-node action similarity, thereby excluding the body shape difference by calculating the body shape similarity of the first person and the second person, and only describing the similarity between the second person's action and the standard action at the action node and the similarity between the second person's action and the standard action during the action, thereby improving the accuracy of the action evaluation.
[0148] Figure 3 A block diagram of a body measurement data evaluation system based on image acquisition according to an embodiment of the present invention is exemplarily shown. The system includes:
[0149] a first acquisition module configured to acquire a first training action video sample through a camera, wherein the first training action video sample is a video shot of a first person who has mastered a standard action at a specified position, and the content of the first training action video sample is the first person performing the specified action;
[0150] a second acquisition module configured to acquire a second training action video sample using a camera, wherein the second training action video sample is not a video of a second person who masters the standard action and is at a specified position, wherein the second training action video is a video of the second person performing the specified action, and the second person and the first person have different body shape data;
[0151] a training module, configured to train a physical measurement data evaluation model based on the first training action video sample, the second training action video sample, the body shape data of the first person, and the body shape data of the second person, to obtain a trained physical measurement data evaluation model;
[0152] A shooting module is used to obtain a video of the person to be tested performing a test action through a camera, wherein the video of the test action is a video shot when the person to be tested is at a specified position, and the content of the video of the test action is the person to be tested performing a specified action;
[0153] Body shape data module, used to obtain body shape data of the person to be tested;
[0154] The evaluation module is used to process the body shape data and the action video of the person to be tested through the trained physical measurement data evaluation model to obtain the physical measurement evaluation information of the person to be tested.
[0155] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0156] Those skilled in the art will appreciate that the embodiments of the present invention described above and shown in the accompanying drawings are intended to be illustrative only and are not intended to limit the present invention. The objectives of the present invention have been fully and effectively achieved. The functional and structural principles of the present invention have been demonstrated and illustrated in the embodiments. Any variations or modifications may be made to the embodiments of the present invention without departing from the principles described.
[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for evaluating body measurement data based on image acquisition, characterized in that: include: collecting a first training action video sample through a camera, wherein the first training action video sample is a video shot by a first person who has mastered the standard action at a specified position, and the content of the first training action video sample is the first person performing the specified action; Capturing a second training action video sample using a camera, wherein the second training action video sample is a video shot of a second person who is not proficient in the standard action and is at a designated position, wherein the second training action video is of the second person performing the designated action, and the second person and the first person have different body shape data; Training a physical measurement data evaluation model based on the first training action video sample, the second training action video sample, the body shape data of the first person, and the body shape data of the second person to obtain a trained physical measurement data evaluation model; Acquire a video of the person to be tested's action to be tested through a camera, wherein the video of the action to be tested is a video shot with the person to be tested at a specified position, and the content of the video of the action to be tested is the person to be tested performing the specified action; Obtaining the body shape data of the person to be tested; The trained physical measurement data evaluation model is used to process the body shape data and the action video of the person to be tested to obtain the physical measurement evaluation information of the person to be tested.
2. The body measurement data evaluation method based on image acquisition according to claim 1, characterized in that: The first training video sample includes a standard training video of the first person performing a standard action, and a non-standard training video of the first person performing a non-standard action; Training a physical measurement data evaluation model based on the first training action video sample, the second training action video sample, the body shape data of the first person, and the body shape data of the second person to obtain a trained physical measurement data evaluation model includes: Determine the standard motion loss function of the physical measurement data evaluation model using the standard training video and the body shape data of the first person; Determine a non-standard motion loss function of a physical measurement data evaluation model using the non-standard training video and the body shape data of the first person; Determining a test loss function of a physical measurement data evaluation model based on a second training action video sample and the body shape data of a second person; The physical measurement data evaluation model is trained according to the standard action loss function, the non-standard action loss function and the test loss function to obtain the trained physical measurement data evaluation model.
3. The method for evaluating body measurement data based on image acquisition according to claim 2, characterized in that: Using standard training videos and the body shape data of the first person, the standard motion loss function of the physical measurement data evaluation model is determined, including: Processing the standard training video using a key point detection sub-model of the body measurement data evaluation model to obtain first position information of key points in a plurality of first training video frames in the standard training video; Inputting first position information of key points corresponding to a plurality of first training video frames into the action node recognition sub-model to determine a first action node video frame; Determining first node position information of a key point in a first action node video frame according to the first position information; Inputting the first position information, the sequence number of the first action node video frame, the first key position information, and the body shape data of the first person into an evaluation sub-model of a physical measurement data evaluation model to obtain a first training physical measurement score; According to the first training physical test score, the standard action loss function is determined.
4. The body measurement data evaluation method based on image acquisition according to claim 3, characterized in that: According to the first training physical test score, determine the standard action loss function, including: According to the formula LOSS S =1-S1 Determine the standard action loss function LOSS S , where S1 is the first training physical test score.
5. The method for evaluating body measurement data based on image acquisition according to claim 3, characterized in that: Using non-standard training videos and the body shape data of the first person, a non-standard motion loss function of the physical measurement data evaluation model is determined, including: Processing the non-standard training video using a key point detection sub-model of the physical measurement data evaluation model to obtain second position information of key points in a plurality of second training video frames in the non-standard training video; Inputting second position information of key points in the plurality of second training video frames into the action node recognition sub-model to determine a second action node video frame; Determining second node position information of a key point in a second action node video frame according to the second position information; Obtaining a first reference score for the non-standard training video according to the sequence number of the first action node video frame, the sequence number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information; Get the first annotation score of the non-standard training video; Obtaining a first control score for the non-standard training video according to the first reference score and the first labeled score; Inputting the body shape data of the first person, the sequence number of the second action node video, the second node position information, and the second position information into the evaluation sub-model of the physical measurement data evaluation model to obtain a second training physical measurement score; A non-standard action loss function of the physical measurement data evaluation model is determined based on the second training physical measurement score, the first control score, the serial number of the first action node video frame, the serial number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information.
6. The body measurement data evaluation method based on image acquisition according to claim 5, characterized in that: Obtaining a first reference score of the non-standard training video according to the sequence number of the first action node video frame, the sequence number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information, including: Obtaining a first action node sequence number vector according to the sequence number of the first action node video frame; Obtaining a second action node sequence number vector according to the sequence number of the second action node video frame; Obtain the similarity between the first action node sequence number vector and the first node sequence number vector of the second action node sequence number; Set basic key points among multiple key points; Determining a first node motion vector between the basic key point and other key points other than the basic key point according to the first node position information; Determining a second node motion vector between the basic key point and other key points other than the basic key point according to the second node position information; Determine an average of similarities between the first node action vector and the second node action vector corresponding to multiple other key points at multiple action nodes as the first node action similarity; Uniformly sampling the first training video frames between the first action node video frames to obtain first sampled video frames, and uniformly sampling the second training video frames between the second action node video frames to obtain second sampled video frames; determining, according to first position information of the first sampled video frame, a first motion vector between the basic key point and other key points other than the basic key point; determining, according to second position information of the second sampled video frame, a second motion vector between the basic key point and other key points other than the basic key point; Determine an average of similarities between the first motion vector and the second motion vector corresponding to multiple other key points at multiple sampling points as a first non-node motion similarity; A first reference score is determined according to the first node sequence number similarity, the first node action similarity, and the first non-node action similarity.
7. The body measurement data evaluation method based on image acquisition according to claim 6, characterized in that: Determining a non-standard motion loss function of a physical test data evaluation model according to the second training physical test score, the first control score, the sequence number of the first action node video frame, the sequence number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information includes: According to the formula Determine the non-standard motion loss function LOSS of the physical test data evaluation model N , where S 1,c is the first control score, S2 is the second training physical test score, sim SN,1 is the similarity of the first node number, sim S,1 is the first node action similarity, sim N,1 is the first non-node action similarity.
8. The method for evaluating body measurement data based on image acquisition according to claim 3, characterized in that: Determine the test loss function of the physical measurement data evaluation model using the second training action video sample and the body shape data of the second person, including: Processing the second training action video sample using a key point detection sub-model of the physical measurement data evaluation model to obtain third position information of key points in a plurality of third training video frames in the second training action video sample; Inputting third position information of key points in the plurality of third training video frames into the action node recognition sub-model to determine a third action node video frame; Determining third node position information of a key point in a third action node video frame according to the third position information; Obtaining a second reference score for the second training action video sample according to the sequence number of the first action node video frame, the sequence number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person, and the body shape data of the second person; Obtain a second annotation score for the second training action video sample; Obtaining a second control score for the second training action video sample based on the second reference score and the second labeled score; Inputting the second person's body shape data, the sequence number of the third action node video, the third node position information, and the third position information into the evaluation sub-model of the physical measurement data evaluation model to obtain a third training physical measurement score; Determine the test loss function of the physical measurement data evaluation model based on the third training physical measurement score, the second control score, the serial number of the first action node video frame, the serial number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person and the body shape data of the second person.
9. The body measurement data evaluation method based on image acquisition according to claim 8, characterized in that: Obtaining a second reference score for a second training action video sample according to the sequence number of the first action node video frame, the sequence number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person, and the body shape data of the second person, including: Obtaining a first action node sequence number vector according to the sequence number of the first action node video frame; Obtaining a third action node sequence number vector according to the sequence number of the third action node video frame; Obtaining the similarity of the second node sequence number of the first action node sequence number vector and the third action node sequence number vector; Set basic key points among multiple key points; Determining a first node motion vector between the basic key point and other key points other than the basic key point according to the first node position information; Determining a third-node motion vector between the basic key point and other key points other than the basic key point according to the third-node position information; Determining a first body shape vector based on the first body shape data, and determining a second body shape vector based on the second body shape data; determining a body shape similarity based on the first body shape vector and the second body shape vector; Determine the second node action similarity by taking the average value of the similarities between the first node action vector and the third node action vector corresponding to the plurality of other key points at the plurality of action nodes and the ratio of the similarities to the body shape similarity; Uniformly sampling the first training video frames between the first action node video frames to obtain first sampled video frames, and uniformly sampling the third training video frames between the third action node video frames to obtain third sampled video frames; determining, according to first position information of the first sampled video frame, a first motion vector between the basic key point and other key points other than the basic key point; determining, according to third position information of the third sampled video frame, a third motion vector between the basic key point and other key points other than the basic key point; Determine the second non-node action similarity by taking the ratio of the average value of the similarities between the first motion vector and the third motion vector corresponding to the plurality of other key points at the plurality of sampling points to the body shape similarity; A second reference score is determined according to the second node sequence number similarity, the second node action similarity, and the second non-node action similarity.
10. A body measurement data evaluation system based on image acquisition, characterized in that: include: a first acquisition module configured to acquire a first training action video sample through a camera, wherein the first training action video sample is a video shot of a first person who has mastered a standard action at a specified position, and the content of the first training action video sample is the first person performing the specified action; a second acquisition module configured to acquire a second training action video sample using a camera, wherein the second training action video sample is not a video of a second person who masters the standard action and is at a specified position, wherein the second training action video is a video of the second person performing the specified action, and the second person and the first person have different body shape data; a training module, configured to train a physical measurement data evaluation model based on the first training action video sample, the second training action video sample, the body shape data of the first person, and the body shape data of the second person, to obtain a trained physical measurement data evaluation model; A shooting module is used to obtain a video of the person to be tested performing a test action through a camera, wherein the video of the test action is a video shot when the person to be tested is at a specified position, and the content of the video of the test action is the person to be tested performing a specified action; Body shape data module, used to obtain body shape data of the person to be tested; The evaluation module is used to process the body shape data and the action video of the person to be tested through the trained physical measurement data evaluation model to obtain the physical measurement evaluation information of the person to be tested.
Citation Information
Patent Citations
User action evaluation method and device, and readable medium
CN109635644A
Professional dance evaluation method for realizing human body posture detection based on deep transfer learning
CN112560665A
Physical fitness evaluation method based on parameterized human body model
CN116824637A
Action analysis system and action analysis method
JP2024057773A
System and method for effectively providing analysis results on the user's exercise behavior in visual aspect to support decision-making related to exercise and medical services
KR102689063B1