Method and system for evaluating anthropometric data based on image acquisition
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]本发明提供一种基于图像采集的体测数据评价方法和系统,能够解决相关技术对于人员动作的评价具有一定的主观性,难以提升评价的准确性,也难以对学员提供准确的反馈得技术问题
[0085]According to the present invention, a physical fitness assessment model can be trained using different video samples from various types of personnel. This enables the model to recognize the movements of individuals with different body types, thereby determining whether the movements of the tested personnel are standard. This provides an objective and accurate basis for evaluating the movements of the tested personnel, reduces subjectivity, improves evaluation accuracy, and provides accurate feedback to the tested personnel. When determining the standard movement loss function, the first training video sample can be a standard training video of a first person performing a standard movement. First position information, the sequence number of the first action node video frame, first key position information, and the body type data of the first person can be obtained. The standard movement loss function is determined by the error between the theoretical value of the first training physical fitness assessment score and the output value of the evaluation sub-model, thereby improving the evaluation accuracy of the evaluation sub-model for standard movements. When determining the non-standard movement loss function, the minimum value among the first node sequence number similarity, the first node action similarity, and the first non-node action similarity can be used to determine the most significant difference between the non-standard and standard movements. The error between the first control score and the second training physical fitness assessment score can be amplified in a targeted manner, thereby specifically enhancing the training strength of the evaluation sub-model for evaluating the most significant differences and improving the accuracy of the evaluation sub-model. When determining the second control score, the body shape difference between the first and second personnel can be considered when determining the similarity of the two-node action and the similarity of the second non-node action. By calculating the body shape similarity between the first and second personnel, the body shape difference can be excluded. Only the similarity between the second person's action and the standard action at the action node and the similarity between the second person's action and the standard action during the action can be described, thereby improving the accuracy of the action evaluation.
Smart Images

Figure CN120612731B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method and system for evaluating physical fitness data based on image acquisition. Background Technology
[0002] In related technologies, when trainees are training and testing a certain movement (e.g., gymnastics), the trainees' movements are usually evaluated by coaches or referees. However, the evaluations by coaches or referees are usually influenced by human factors, making the evaluations subjective and difficult to improve their accuracy, thus making it difficult to provide accurate feedback to the trainees. Summary of the Invention
[0003] This invention provides a method and system for evaluating physical fitness data based on image acquisition, which can solve the technical problems that related technologies have a certain degree of subjectivity in evaluating human movements, making it difficult to improve the accuracy of evaluation and provide accurate feedback to trainees.
[0004] According to a first aspect of the present invention, a method for evaluating physical fitness data based on image acquisition is provided, comprising:
[0005] The first training action video sample is captured by a camera. The first training action video sample is a video taken of the first person who has mastered the standard action at a designated position. The content of the first training action video sample is the first person performing the designated action.
[0006] The second training action video sample is captured by a camera. The second training action video sample is a video taken at a designated position by a second person who does not know the standard action. The content of the second training action video is the second person performing the designated action, and the body data of the second person is different from that of the first person.
[0007] Based on the first training action video sample, the second training action video sample, the body shape data of the first person and the body shape data of the second person, the physical fitness test data evaluation model is trained to obtain the trained physical fitness test data evaluation model.
[0008] The test involves acquiring a video of the subject's actions using a camera. The video is taken from a designated location and shows the subject performing a specified action.
[0009] Obtain the body shape data of the person being tested;
[0010] The trained physical fitness test data evaluation model processes the body shape data and video of the test subjects to obtain their physical fitness test evaluation information.
[0011] According to the present invention, the first training video sample includes standard training videos of the first person performing standard actions and non-standard training videos of the first person performing non-standard actions;
[0012] Based on the first training action video samples, the second training action video samples, and the body shape data of the first and second personnel, the physical fitness test data evaluation model is trained to obtain the trained physical fitness test data evaluation model, including:
[0013] By using standard training videos and body shape data of the first person, the standard motion loss function of the physical test data evaluation model is determined;
[0014] By using non-standard training videos and body shape data of the first person, the non-standard motion loss function of the physical test data evaluation model was determined;
[0015] The test loss function of the physical fitness data evaluation model is determined by using the second training action video samples and the body shape data of the second person.
[0016] The physical fitness test data evaluation model is trained using the standard motion loss function, the non-standard motion loss function, and the test loss function to obtain the trained physical fitness test data evaluation model.
[0017] According to the present invention, a standard motion loss function for a physical fitness data evaluation model is determined using standard training videos and body shape data of a first person, including:
[0018] The key point detection sub-model of the physical test data evaluation model is used to process the standard training video to obtain the first position information of key points in multiple first training video frames in the standard training video.
[0019] Input the first position information of key points corresponding to multiple first training video frames into the action node recognition sub-model to determine the first action node video frame.
[0020] Based on the first position information, determine the first node position information of the key point in the video frame of the first action node;
[0021] The first position information, the sequence number of the first action node video frame, the first key position information, and the body shape data of the first person are input into the evaluation sub-model of the physical test data evaluation model to obtain the first training physical test score.
[0022] Based on the first training physical test score, the standard movement loss function is determined.
[0023] According to the present invention, determining the standard movement loss function based on a first training fitness test score includes:
[0024] According to the formula
[0025] LOSSS =1-S1
[0026] Determine the standard action loss function (LOSS) S S1 is the score of the first training physical test.
[0027] According to the present invention, a non-standard motion loss function for a physical fitness data evaluation model is determined using non-standard training videos and body shape data of a first person, including:
[0028] The key point detection sub-model of the physical test data evaluation model is used to process non-standard training videos to obtain the second position information of key points in multiple second training video frames in the non-standard training videos.
[0029] Input the second position information of key points in multiple second training video frames into the action node recognition sub-model to determine the second action node video frame;
[0030] Based on the second position information, determine the second node position information of the key points in the video frame of the second action node;
[0031] Based on the sequence number of the first action node video frame, the sequence number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information, a first reference score for the non-standard training video is obtained.
[0032] Obtain the first labeled score of non-standard training videos;
[0033] Based on the first reference score and the first annotation score, a first comparative score for the non-standard training video is obtained.
[0034] The body shape data of the first person, the sequence number of the second action node video, the location information of the second node, and the location information of the second position are input into the evaluation sub-model of the physical test data evaluation model to obtain the second training physical test score.
[0035] Based on the second training physical test score, the first control score, the sequence number of the first action node video frame, the sequence number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information, the non-standard action loss function of the physical test data evaluation model is determined.
[0036] According to the present invention, a first reference score for a non-standard training video is obtained based on the sequence number of the first action node video frame, the sequence number of the second action node video frame, the first node position information, the second node position information, the first position information, and the second position information, including:
[0037] Based on the sequence number of the video frame of the first action node, obtain the sequence number vector of the first action node;
[0038] Based on the sequence number of the video frame of the second action node, obtain the sequence number vector of the second action node;
[0039] Obtain the similarity of the first node index between the first action node index vector and the second action node index vector;
[0040] Set basic key points among multiple key points;
[0041] Based on the first node position information, determine the first node action vector between the basic key point and other key points other than the basic key point;
[0042] Based on the second node position information, determine the second node action vector between the basic key point and other key points other than the basic key point;
[0043] The average similarity between the first node action vector and the second node action vector corresponding to multiple other key points at multiple action nodes is determined as the first node action similarity.
[0044] Uniformly sample the first training video frames between the first action node video frames to obtain the first sampled video frame, and uniformly sample the second training video frames between the second action node video frames to obtain the second sampled video frame.
[0045] Based on the first position information of the first sampled video frame, determine the first motion vector between the basic key point and other key points besides the basic key point;
[0046] Based on the second position information of the second sampled video frame, determine the second motion vector between the basic key point and other key points other than the basic key point;
[0047] The average similarity between the first action vector and the second action vector corresponding to multiple other key points at multiple sampling points is determined as the first non-node action similarity.
[0048] The first reference score is determined based on the similarity of the first node sequence number, the similarity of the first node action, and the similarity of the first non-node action.
[0049] According to the present invention, a non-standard motion loss function for the physical fitness test data evaluation model is determined based on a second training physical fitness test score, a first control score, the sequence number of the first action node video frame, the sequence number of the second action node video frame, the first node position information, the second node position information, the first position information, and the second position information, including:
[0050] According to the formula
[0051]
[0052] Determine the non-standard motion loss function (LOSS) for the physical fitness test data evaluation model. N , of which S 1,c S1 is the first control score, and S2 is the second training fitness test score. SN,1 For the similarity of the first node's sequence number, sim S,1 For the action similarity of the first node, sim N,1 The similarity is for the first non-node action.
[0053] According to the present invention, the test loss function of the physical fitness data evaluation model is determined by using second training motion video samples and body shape data of a second person, including:
[0054] The key point detection sub-model of the physical test data evaluation model is used to process the second training action video samples to obtain the third position information of key points in multiple third training video frames in the second training action video samples.
[0055] Input the third position information of key points in multiple third training video frames into the action node recognition sub-model to determine the third action node video frame;
[0056] Based on the third position information, determine the third node position information of the key points in the video frame of the third action node;
[0057] Based on the sequence number of the first action node video frame, the sequence number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person and the body shape data of the second person, a second reference score for the second training action video sample is obtained.
[0058] Obtain the second labeled score of the second training action video sample;
[0059] Based on the second reference score and the second annotation score, a second control score is obtained for the second training action video sample;
[0060] Input the body shape data of the second person, the sequence number of the third action node video, the third node position information, and the third position information into the evaluation sub-model of the physical test data evaluation model to obtain the third training physical test score.
[0061] Based on the third training physical test score, the second control score, the sequence number of the first action node video frame, the sequence number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person and the body shape data of the second person, the test loss function of the physical test data evaluation model is determined.
[0062] According to the present invention, a second reference score for the second training action video sample is obtained based on the sequence number of the first action node video frame, the sequence number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person, and the body shape data of the second person, including:
[0063] Based on the sequence number of the video frame of the first action node, obtain the sequence number vector of the first action node;
[0064] Based on the sequence number of the video frame of the third action node, obtain the sequence number vector of the third action node;
[0065] Obtain the similarity of the second node indices between the first action node index vector and the third action node index vector;
[0066] Set basic key points among multiple key points;
[0067] Based on the first node position information, determine the first node action vector between the basic key point and other key points other than the basic key point;
[0068] Based on the third node position information, determine the third node action vector between the basic key point and other key points other than the basic key point;
[0069] Based on the first body shape data, determine the first body shape vector, and based on the second body shape data, determine the second body shape vector;
[0070] Determine the body shape similarity based on the first body shape vector and the second body shape vector;
[0071] The ratio of the average similarity between the first node action vector and the third node action vector corresponding to multiple other key points at multiple action nodes to the body shape similarity is determined as the second node action similarity.
[0072] Uniformly sample the first training video frames between the first action node video frames to obtain the first sampled video frame, and uniformly sample the third training video frames between the third action node video frames to obtain the third sampled video frame.
[0073] Based on the first position information of the first sampled video frame, determine the first motion vector between the basic key point and other key points besides the basic key point;
[0074] Based on the third position information of the third sampled video frame, determine the third motion vector between the basic key point and other key points besides the basic key point;
[0075] The ratio of the average similarity of the first action vector and the third action vector corresponding to multiple other key points at multiple sampling points to the body shape similarity is determined as the second non-node action similarity.
[0076] The second reference score is determined based on the similarity of the second node sequence number, the similarity of the second node action, and the similarity of the second non-node action.
[0077] According to a second aspect of the present invention, a physical fitness data evaluation system based on image acquisition is provided, comprising:
[0078] The first acquisition module is used to acquire first training action video samples through a camera. The first training action video sample is a video taken by the first person who has mastered the standard action at a designated position. The content of the first training action video sample is the first person performing the designated action.
[0079] The second acquisition module is used to acquire second training action video samples through a camera. The second training action video samples are videos taken by a second person who has mastered the standard action at a designated position. The content of the second training action video is the second person performing the designated action, and the body data of the second person is different from that of the first person.
[0080] The training module is used to train the physical fitness test data evaluation model based on the first training action video sample, the second training action video sample, the body shape data of the first person and the body shape data of the second person, so as to obtain the trained physical fitness test data evaluation model.
[0081] The shooting module is used to acquire a video of the subject's actions through a camera. The video of the subject's actions is a video taken when the subject is at a designated location, and the content of the video of the subject's actions is the subject performing a designated action.
[0082] The body shape data module is used to acquire the body shape data of the person being tested;
[0083] The evaluation module is used to process the body shape data and video of the test subjects' movements using the trained physical test data evaluation model to obtain the physical test evaluation information of the test subjects.
[0084] By adopting the above technical solution, the present invention can achieve the following technical effects:
[0085] According to the present invention, a physical fitness assessment model can be trained using different video samples from various types of personnel. This enables the model to recognize the movements of individuals with different body types, thereby determining whether the movements of the tested personnel are standard. This provides an objective and accurate basis for evaluating the movements of the tested personnel, reduces subjectivity, improves evaluation accuracy, and provides accurate feedback to the tested personnel. When determining the standard movement loss function, the first training video sample can be a standard training video of a first person performing a standard movement. First position information, the sequence number of the first action node video frame, first key position information, and the body type data of the first person can be obtained. The standard movement loss function is determined by the error between the theoretical value of the first training physical fitness assessment score and the output value of the evaluation sub-model, thereby improving the evaluation accuracy of the evaluation sub-model for standard movements. When determining the non-standard movement loss function, the minimum value among the first node sequence number similarity, the first node action similarity, and the first non-node action similarity can be used to determine the most significant difference between the non-standard and standard movements. The error between the first control score and the second training physical fitness assessment score can be amplified in a targeted manner, thereby specifically enhancing the training strength of the evaluation sub-model for evaluating the most significant differences and improving the accuracy of the evaluation sub-model. When determining the second control score, the body shape difference between the first and second personnel can be considered when determining the similarity of the two-node action and the similarity of the second non-node action. By calculating the body shape similarity between the first and second personnel, the body shape difference can be excluded. Only the similarity between the second person's action and the standard action at the action node and the similarity between the second person's action and the standard action during the action can be described, thereby improving the accuracy of the action evaluation. Attached Figure Description
[0086] Figure 1 An exemplary flowchart of a body measurement data evaluation method based on image acquisition according to an embodiment of the present invention is shown.
[0087] Figure 2 A flowchart of a training body measurement data evaluation model according to an embodiment of the present invention is shown as an example;
[0088] Figure 3 A block diagram of an image acquisition-based physical assessment data evaluation system according to an embodiment of the present invention is shown as an example. Detailed Implementation
[0089] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0090] Figure 1 An exemplary flowchart of a physical fitness data evaluation method based on image acquisition according to an embodiment of the present invention is shown, the method comprising:
[0091] Step S1: Collect a first training action video sample using a camera. The first training action video sample is a video taken by the first person who has mastered the standard action at a designated position. The content of the first training action video sample is the first person performing the designated action.
[0092] Step S2: Collect a second training action video sample using a camera. The second training action video sample is a video taken at a designated location by a second person who has mastered the standard action. The content of the second training action video is the second person performing the designated action, and the body shape data of the second person is different from that of the first person.
[0093] Step S3: Based on the first training action video sample, the second training action video sample, the body shape data of the first person and the body shape data of the second person, train the physical test data evaluation model to obtain the trained physical test data evaluation model.
[0094] Step S4: Acquire a video of the subject's actions using a camera. The video of the subject's actions is a video taken at a designated location, and the content of the video is the subject performing a designated action.
[0095] Step S5: Obtain the body shape data of the person to be tested;
[0096] Step S6: The trained physical test data evaluation model processes the body shape data and the video of the test action of the test subject to obtain the physical test evaluation information of the test subject.
[0097] According to an embodiment of the present invention, the image-based physical fitness data evaluation method can train the physical fitness data evaluation model through different video samples of various types of people, enabling the physical fitness data evaluation model to recognize the movements of people of various body types, and then determine whether the movements of the person being tested are standard. This provides an objective and accurate basis for judging and evaluating the movements of the person being tested, reduces subjectivity, improves evaluation accuracy, and can provide accurate feedback to the person being tested.
[0098] According to one embodiment of the present invention, in step S1, the designated location can be a point or a region, and the camera's field of view includes the designated location. In the example, the designated location is a region, and the distance and angle between the camera and the centroid of the region are fixed values. For example, the initial position of the person performing the designated action is located at the centroid of the region and facing the camera lens. The designated action is an action that the person to be tested needs to practice or perform, such as a gymnastic movement. The standard action is the designated action without any errors. The first person can be a coach or similar personnel who is proficient in the standard action. The first person can be in the designated location, and the camera can record a video of the first person performing the standard action. The first person can also record videos with errors as samples, for example, some common incorrect actions can be recorded to provide trainees with points to avoid. The video sample of the first person captured by the camera is the first training action video sample. The first training action video sample can include two types: standard training video of the first person performing standard actions and non-standard training video of the first person performing non-standard actions. Non-standard actions are actions with errors, such as the aforementioned incorrect actions. By using two types of first training video samples, the number and diversity of samples are increased, which helps to improve the performance of the physical test data evaluation model during training.
[0099] According to one embodiment of the present invention, in step S2, the second person is a person who does not master the standard movements, such as a trainee. The second person's movements may contain some errors, and the body shape data of the second person is different from that of the first person. Using the video of the second person performing the specified movements captured by the camera as a sample (i.e., the second training movement video sample) can improve the accuracy of the body measurement data evaluation model in recognizing the movements of people with various body shapes during training. This helps to reduce the possibility of the body measurement data evaluation model's accuracy decreasing due to movement and posture errors caused by body shape differences, and also helps to improve the robustness of the body measurement data evaluation model.
[0100] According to an embodiment of the present invention, in step S3, the physical fitness test data evaluation model is a combination model of multiple sub-models, including a key point detection sub-model, a motion node recognition sub-model, and an evaluation sub-model. The key point detection sub-model and the motion node recognition sub-model can be pre-trained, and the evaluation sub-model is trained using the trained key point detection sub-model and the trained motion node recognition sub-model, as well as the aforementioned first training video samples and second training motion video samples, thereby enabling the physical fitness test data evaluation model to accurately evaluate the motion accuracy of the test subject based on the motion video and body shape data.
[0101] Figure 2 A flowchart of a training performance data evaluation model according to an embodiment of the present invention is shown as an example.
[0102] According to an embodiment of the present invention, in step S3, the physical fitness test data evaluation model is trained based on the first training action video sample, the second training action video sample, the body shape data of the first person, and the body shape data of the second person to obtain the trained physical fitness test data evaluation model, including: step S31, determining the standard action loss function of the physical fitness test data evaluation model using the standard training video and the body shape data of the first person; step S32, determining the non-standard action loss function of the physical fitness test data evaluation model using the non-standard training video and the body shape data of the first person; step S33, determining the test loss function of the physical fitness test data evaluation model using the second training action video sample and the body shape data of the second person; step S34, training the physical fitness test data evaluation model based on the standard action loss function, the non-standard action loss function, and the test loss function to obtain the trained physical fitness test data evaluation model.
[0103] According to an embodiment of the present invention, in step S31, the standard training video is a video sample taken when a first person who has mastered the standard action performs the action without error. This video sample can be used to enable the physical fitness data evaluation model to obtain data on the standard action, thereby determining whether the action of the person being tested is standard. The standard action loss function of the physical fitness data evaluation model is determined using the standard training video and the body shape data of the first person. This includes: processing the standard training video using the key point detection sub-model of the physical fitness data evaluation model to obtain the first position information of key points in multiple first training video frames in the standard training video; inputting the first position information of the key points corresponding to the multiple first training video frames into the action node recognition sub-model to determine the first action node video frame; determining the first node position information of the key points in the first action node video frame based on the first position information; inputting the first position information, the sequence number of the first action node video frame, the first key position information, and the body shape data of the first person into the evaluation sub-model of the physical fitness data evaluation model to obtain the first training physical fitness score; and determining the standard action loss function based on the first training physical fitness score.
[0104] According to one embodiment of the present invention, a standard training video can be parsed to obtain multiple first training video frames. A keypoint detection sub-model is then used to detect the first position information of key points of a first person in each first training video frame. The keypoint detection sub-model is a convolutional neural network model, which can be used to detect the position coordinates of multiple key points on a person's body in an image. These multiple key points may include head key points, neck key points, shoulder key points, elbow key points, hand key points, waist key points, knee key points, ankle key points, etc. The posture of the first person in each first training video frame can be determined based on the first position information of each key point, and the posture changes of the first person can be determined based on the timestamps of the first training video frames, thereby analyzing the actions of the first person.
[0105] According to one embodiment of the present invention, the first position information of key points in the first training video frames can be arranged into a sequence according to the order of the first training video frames, and the sequence of the first position information can be input into the action node recognition sub-model to determine the first action node video frame. The action node recognition sub-model is a BP neural network model, which can identify action nodes in the action process based on the position of key points. For example, when performing an action, the movement direction of each key point will not change significantly before reaching the action node, but after reaching the action node, the movement direction of at least one key point will change significantly. For example, in the action of raising an arm to a certain height and then lowering it, before the action node (i.e., before the arm is raised to a certain height), the curves of the hand key point and the elbow key point are relatively smooth, that is, the hand key point and the elbow key point rise along a smooth trajectory without a large angle. However, after reaching the action node, the movement trajectory of the hand key point and the elbow key point descends. Therefore, there is a large angle (e.g., greater than or equal to 90°) between the trajectory of the hand key point and the trajectory of the elbow key point before reaching the action node and the trajectory after reaching the action node. The action node recognition sub-model monitors the motion trajectory of each keypoint and records the moment when the motion trajectory of at least one keypoint exhibits a large angle, thus determining the first training video frame corresponding to that moment as the first action node video frame. In the example, the action node recognition sub-model calculates the motion vector between the first position information of the keypoint in adjacent first training video frames. If at least one keypoint has a motion vector between its first position information in the (A-1)th first training video frame and the first position information in the Ath first training video frame, and the motion vector between its first position information in the Ath first training video frame and the first position information in the (A+1)th first training video frame, the Ath first training video frame is the first action node video frame, where A is a positive integer and A is less than the total number of first training video frames.
[0106] According to one embodiment of the present invention, when performing a specified action, the aforementioned action node is a node about to change its motion state. The posture at the action node has a significant impact on the standardization and aesthetics of the overall action. For example, when performing gymnastics, the posture at the action node can affect the completeness and fluidity of the action. Therefore, after determining the first action node video frame, the first node position information of the key points in the first action node video frame can be determined, thereby determining the posture at the action node. When performing a standard action, this posture is the standard posture.
[0107] According to one embodiment of the present invention, first position information, the sequence number of the first action node video frame, first key position information, and the body shape data of the first person (e.g., height, arm span, leg length, etc.) can be input into the evaluation sub-model of the physical fitness test data evaluation model to obtain a first training physical fitness test score. The physical fitness test data evaluation model is a BP neural network model, which can process the above-mentioned various data and output the first training physical fitness test score of the standard training video. Theoretically, the first training physical fitness test score is 1, that is, the standard training video is a video taken when the first person performs the standard action, therefore, the first training physical fitness test score is theoretically full marks. However, the evaluation sub-model may have errors, causing the first training physical fitness test score output by the evaluation sub-model to be not equal to 1.
[0108] According to one embodiment of the present invention, determining the standard movement loss function based on a first training fitness test score includes: determining the standard movement loss function LOSS according to formula (1). S ,
[0109] LOSS S =1-S1(1)
[0110] S1 is the score of the first training physical test.
[0111] According to an embodiment of the present invention, as described above, the first training body measurement score is theoretically 1, but the first training body measurement score output by the evaluation sub-model may have errors. Therefore, the error 1-S1 is used as the standard action loss function, and the loss function is backpropagated during training to reduce the error. This enables the evaluation sub-model to output a first training body measurement score that is closer to the theoretical value of 1 when processing the first position information, the sequence number of the first action node video frame, the first key position information, and the body shape data of the first person corresponding to the standard action, thereby improving the accuracy of the evaluation sub-model.
[0112] According to one embodiment of the present invention, the first person may include multiple people. When the multiple first people perform standard actions, they can respectively capture video samples of the first training actions, and can respectively obtain the first position information, the sequence number of the first action node video frame, the first key position information and the body shape data of the first person, and iteratively execute the above training to improve the accuracy of the evaluation sub-model.
[0113] In this way, when the first training video sample is a standard training video of the first person performing a standard action, the first position information, the sequence number of the first action node video frame, the first key position information, and the body shape data of the first person can be obtained. The standard action loss function is determined by the error between the theoretical value of the first training body measurement score and the output value of the evaluation sub-model, so as to improve the evaluation accuracy of the evaluation sub-model for the standard action.
[0114] According to one embodiment of the present invention, in step S32, the non-standard training video is a video sample taken when the first person performs a non-standard action, for example, a video sample taken when the first person demonstrates a common mistake.
[0115] According to one embodiment of the present invention, for non-standard actions, i.e., actions with errors during execution, points can be deducted accordingly during training. Referees or coaches can provide expert scoring for non-standard actions, and video samples of non-standard actions can be analyzed to automatically obtain scores. The expert scores and automatically obtained scores are combined to obtain a first labeled score for training the evaluation sub-model, which is then compared with the score output by the evaluation sub-model to determine the non-standard action loss function.
[0116] According to an embodiment of the present invention, in step S22, determining the non-standard motion loss function of the physical fitness data evaluation model using non-standard training videos and body shape data of a first person includes: processing the non-standard training videos using a key point detection sub-model of the physical fitness data evaluation model to obtain second position information of key points in multiple second training video frames in the non-standard training videos; inputting the second position information of key points in multiple second training video frames into an action node recognition sub-model to determine second action node video frames; determining the second node position information of key points in the second action node video frames based on the second position information; and determining the second node position information of key points in the second action node video frames based on the sequence number of the first action node video frame, the sequence number of the second action node video frame, the first node position information, and the second node position information. Information, first position information, and second position information are used to obtain a first reference score for the non-standard training video; a first labeled score for the non-standard training video is obtained; based on the first reference score and the first labeled score, a first control score for the non-standard training video is obtained; the body shape data of the first person, the sequence number of the second action node video, the second node position information, and the second position information are input into the evaluation sub-model of the physical fitness test data evaluation model to obtain a second training physical fitness test score; based on the second training physical fitness test score, the first control score, the sequence number of the first action node video frame, the sequence number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information, the non-standard action loss function of the physical fitness test data evaluation model is determined.
[0117] According to one embodiment of the present invention, similar to the method for obtaining the first location information, multiple second training video frames of a non-standard training video can be processed by a keypoint detection sub-model to obtain the second location information of keypoints in each second training video frame. Similar to the first action node video frame, the second location information of keypoints in multiple second training video frames can be processed by an action node recognition sub-model to obtain a second action node video frame, and the second node location information of keypoints in the second action node video frame can be obtained. The specific processing steps are not detailed here.
[0118] According to one embodiment of the present invention, the first reference score is a score obtained by analyzing the differences in actions and poses between the second training video frame and the first training video frame. The first reference score for a non-standard training video is obtained based on the sequence number of the first action node video frame, the sequence number of the second action node video frame, the first node position information, the second node position information, the first position information, and the second position information. This includes: obtaining a first action node sequence number vector based on the sequence number of the first action node video frame; obtaining a second action node sequence number vector based on the sequence number of the second action node video frame; obtaining the first node sequence number similarity between the first action node sequence number vector and the second action node sequence number vector; setting basic keypoints among multiple keypoints; determining the first node action vector between the basic keypoints and other keypoints besides the basic keypoints based on the first node position information; determining the second node action vector between the basic keypoints and other keypoints besides the basic keypoints based on the second node position information; and mapping the first node action vectors corresponding to the multiple other keypoints at the multiple action nodes to... The average similarity between the quantity and the action vector of the second node is determined as the first node action similarity; uniform sampling is performed on the first training video frames between the first action node video frames to obtain the first sampled video frames, and uniform sampling is performed on the second training video frames between the second action node video frames to obtain the second sampled video frames; based on the first position information of the first sampled video frames, the first action vector between the basic key point and other key points other than the basic key point is determined; based on the second position information of the second sampled video frames, the second action vector between the basic key point and other key points other than the basic key point is determined; the average similarity between the first action vector and the second action vector corresponding to multiple other key points at multiple sampling points is determined as the first non-node action similarity; based on the first node sequence similarity, the first node action similarity, and the first non-node action similarity, the first reference score is determined.
[0119] According to an embodiment of the present invention, as described above, the posture at an action node has a significant impact on the standardization and aesthetics of the overall action. Therefore, both the first and second action node video frames are relatively important data, as are their sequence numbers. The time interval between adjacent first action node video frames is a fixed value; therefore, the difference between the sequence numbers of adjacent first action node video frames can reflect both the duration and speed of the action execution process between adjacent action nodes. Similarly, the difference between the sequence numbers of adjacent second action node video frames can reflect both the duration and speed of the action execution process between adjacent action nodes. Therefore, the sequence numbers of the first and second action node video frames can be used to determine the consistency of the execution speed of each action, and can be used to determine whether the execution speeds of non-standard and standard actions are similar.
[0120] According to one embodiment of the present invention, the sequence numbers of the first action node video frames can be combined to obtain a first action node sequence number vector, that is, each component in the first action node sequence number vector is the sequence number of the first action node video frame. The sequence numbers of the second action node video frames can also be combined to obtain a second action node sequence number vector, where each component is the sequence number of the second action node video frame. If the dimensions of the first action node sequence number vector and the second action node sequence number vector are different (e.g., the difference between standard and non-standard actions is significant), zero-padding is performed to make the dimensions of the two vectors the same. For example, if the dimension of the first action node sequence number vector is greater than the dimension of the second action node sequence number vector, zeros are padded at the end of the second action node sequence number vector so that the dimension of the zero-padding second action node sequence number vector is equal to the dimension of the first action node sequence number vector. Similarly, if the dimension of the first action node sequence number vector is less than the dimension of the second action node sequence number vector, zeros are padded at the end of the first action node sequence number vector so that the dimension of the zero-padding first action node sequence number vector is equal to the dimension of the second action node sequence number vector. After performing the above dimensional alignment processing, the similarity of the first node indices of the first action node indices and the second action node indices can be determined. For example, the cosine similarity of the first and second action node indices can be calculated as the first node indices similarity. The first node indices similarity can be used to determine whether the execution speeds of non-standard actions and standard actions are similar.
[0121] According to one embodiment of the present invention, the basic keypoints may be neck keypoints or waist keypoints; the present invention does not limit the setting of the basic keypoints. At the action node, a first node action vector is determined based on the first node position information; for example, the first node action vector is a basic keypoint pointing to other keypoints. Similarly, a second node action vector can be determined.
[0122] According to one embodiment of the present invention, the first node position information can be arranged sequentially, and the second node position information can be arranged sequentially. Then, the first node position information and the second node position information with the same order can be matched. If the number of the first node position information and the second node position information are different (for example, the difference between standard actions and non-standard actions is large), zero-padding processing similar to the above can be performed. For example, if the number of the first node position information is greater than the number of the second node position information, the second node position information can be zero-padding processing. For example, if the number of the first node position information is 20 and the number of the second node position information is 18, the second node position information can be zero-padding processing. The 19th and 20th second node position information can all be set to (0, 0), that is, the coordinates of all key points are set to (0, 0), and then the second node action vector can be solved.
[0123] According to one embodiment of the present invention, the similarity between the first node action vector corresponding to the first node position information and the second node action vector corresponding to the second node position information with the same order can be solved. For example, the cosine similarity between the first node action vector between the head key point and the neck key point (basic key point) in the first first node position information and the first node action vector between the head key point and the neck key point in the first second node position information can be determined. The cosine similarity between the first node action vector between the shoulder key point and the neck key point in the first first node position information and the first node action vector between the shoulder key point and the neck key point in the first second node position information can be determined. And these cosine similarities are averaged to determine the action similarity between the first first node position information and the first second node position information. By using a similar method as above, the action similarity between the second first node position information and the second second node position information can be obtained, the action similarity between the third first node position information and the third second node position information, and so on. Then, the action similarity of each action can be averaged to obtain the action similarity of the first node, which can be used to describe the similarity between standard and non-standard actions when describing action nodes.
[0124] According to one embodiment of the present invention, multiple first training video frames may exist between adjacent first action node video frames. These first training video frames can be uniformly sampled. For example, if there are 8 first training video frames between adjacent first action node video frames, these 8 first training video frames can be uniformly sampled, for example, the 3rd and 6th first training video frames can be selected. Since the speed of non-standard actions may not be consistent with the speed of standard actions, the number of second training video frames between second action node video frames with the same sequence number may be different from the number of first training video frames. For example, if the number of second training video frames between adjacent second action node video frames is 5, then during uniform sampling, the 2nd and 4th second training video frames can be selected as the second sampled video frames. The above sampling quantity and sampling interval are only examples. During sampling, it is only necessary to ensure that the sampling quantity of the first training video frames between adjacent first action node video frames is equal to the sampling quantity of the second training video frames between second action node video frames with the same sequence number. The first sampled video frames between adjacent first action node video frames and the second sampled video frames between second action node video frames with the same sequence number have a one-to-one correspondence.
[0125] According to one embodiment of the present invention, in the first sampled video frame, based on the first position information, a first motion vector is determined between the basic key point and other key points besides the basic key point. The determination method is similar to that of the first node motion vector, and will not be repeated here. Similarly, a second motion vector in the second sampled video frame can be determined. The cosine similarity between each first motion vector in the first sampled video frame and the corresponding second motion vector in the corresponding second sampled video frame can be determined, and the cosine similarity between each first motion vector and the second motion vector can be averaged to obtain the motion similarity of the person in the first sampled video frame and the second sampled video frame. Further, the motion similarity of the person in each first sampled video frame and the corresponding second sampled video frame can be averaged to obtain a first non-node motion similarity, which can be used to describe the similarity between standard and non-standard actions when an action node is not reached (or during the action process).
[0126] According to one embodiment of the present invention, a first reference score can be determined by weighted averaging of the first node sequence similarity, the first node action similarity, and the first non-node action similarity. The first reference score can comprehensively evaluate the similarity between non-standard actions and standard actions by combining the consistency of action speed between action nodes, the action similarity at the action node, and the action similarity during the action process, thereby improving the accuracy and comprehensiveness of the evaluation.
[0127] According to one embodiment of the present invention, the first reference score obtained above based on the analysis of the first and second training video frames can also be obtained by manual scoring, such as by referees, coaches, or other professionals. Furthermore, a weighted average can be taken between the first reference score and the first reference score to obtain a first control score for the non-standard training video. This allows for the combination of expert experience and objective analysis of video frames to obtain a comprehensive first control score, improving the accuracy and objectivity of the first control score for use in training the evaluation sub-model.
[0128] According to one embodiment of the present invention, the evaluation sub-model can evaluate a video of a first person performing non-standard movements. For example, the evaluation sub-model can be input with the first person's body shape data, the sequence number of the second action node video, the second node position information, and the second position information to obtain a second training physical fitness test score. The second training physical fitness test score may contain errors; therefore, a first control score can be used as the accurate score to determine the error of the second training physical fitness test score.
[0129] According to one embodiment of the present invention, the non-standard motion loss function of the physical test data evaluation model is determined based on the second training physical test score, the first control score, the sequence number of the first action node video frame, the sequence number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information, including: determining the non-standard motion loss function LOSS of the physical test data evaluation model according to formula (2). N ,
[0130]
[0131] Among them, S 1,c S1 is the first control score, and S2 is the second training fitness test score. SN,1 For the similarity of the first node's sequence number, sim S,1 For the action similarity of the first node, sim N,1 The similarity is for the first non-node action.
[0132] According to an embodiment of the present invention, in formula (2), |S 1,c -S2| represents the error between the first control score and the second training physical test score. During training, this error can be reduced, thereby improving the accuracy of the evaluation sub-model of the physical test data evaluation model. min(sim SN,1 ,sim S,1 ,sim N,1The minimum value among the similarity of the first node number, the similarity of the first node action, and the similarity of the first non-node action is used. Since the similarity of the first node number, the similarity of the first node action, and the similarity of the first non-node action are all less than or equal to 1, the minimum value of the three can be used as the denominator to amplify the error between the first control score and the second training physical test score. Furthermore, the minimum value among the first node number similarity, the first node action similarity, and the first non-node action similarity can determine the most significant difference between non-standard and standard actions. For example, if the minimum value of the three is the similarity of the first node number, the most significant difference is the difference in action speed; if the minimum value of the three is the similarity of the first node action, the most significant difference is the difference in action similarity at the action node; and if the minimum value of the three is the similarity of the three is the similarity of action during the action, the most significant difference is the difference in action similarity during the action. Based on the most significant difference, the error between the first control score and the second training physical test score can be amplified in a targeted manner, thereby improving the evaluation accuracy of the evaluation sub-model for the most significant difference and thus improving the overall evaluation accuracy. During training, the above non-standard action loss function can be backpropagated to adjust the accuracy of the evaluation sub-model, and the evaluation accuracy of the evaluation sub-model for the most important differences can be improved in a targeted manner.
[0133] In this way, the minimum value among the first node sequence similarity, the first node action similarity, and the first non-node action similarity can be used to determine the most significant difference between non-standard and standard actions. This allows for targeted amplification of the error between the first control score and the second training physical test score, thereby enhancing the training strength of the evaluation sub-model for evaluating the most significant difference and improving the accuracy of the evaluation sub-model.
[0134] According to one embodiment of the present invention, in step S33, in order to make the physical fitness data evaluation model applicable to the evaluation of people with various body types and improve its applicability and robustness, a test loss function can be obtained using the second training action video samples of the second person and the body type data of the second person to train the physical fitness data evaluation model. The second person may include multiple people, each with different body type data, which is different from the body type data of the first person, and the degree of mastery of the specified action by the second person is different for each person.
[0135] According to an embodiment of the present invention, in step S33, determining the test loss function of the physical fitness data evaluation model using the second training action video samples and the body shape data of the second person includes: processing the second training action video samples using the key point detection sub-model of the physical fitness data evaluation model to obtain the third position information of key points in multiple third training video frames in the second training action video samples; inputting the third position information of key points in multiple third training video frames into the action node recognition sub-model to determine the third action node video frame; determining the third node position information of key points in the third action node video frame based on the third position information; and determining the third node position information of key points in the third action node video frame based on the sequence number of the first action node video frame, the sequence number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, and the first person's... The body shape data of the first person and the body shape data of the second person are used to obtain a second reference score for the second training action video sample; a second labeled score for the second training action video sample is obtained; based on the second reference score and the second labeled score, a second control score for the second training action video sample is obtained; the body shape data of the second person, the sequence number of the third action node video, the third node position information, and the third position information are input into the evaluation sub-model of the physical fitness test data evaluation model to obtain a third training physical fitness test score; based on the third training physical fitness test score, the second control score, the sequence number of the first action node video frame, the sequence number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person and the body shape data of the second person, the test loss function of the physical fitness test data evaluation model is determined.
[0136] According to one embodiment of the present invention, the method of obtaining the third location information, the third action node video frame and the third node location information is similar to the method of obtaining the first location information, the first action node video frame and the first node location information described above, and will not be repeated here.
[0137] According to one embodiment of the present invention, a second reference score is obtained for a second training action video sample based on the sequence number of the first action node video frame, the sequence number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person, and the body shape data of the second person. This includes: obtaining a first action node sequence number vector based on the sequence number of the first action node video frame; obtaining a third action node sequence number vector based on the sequence number of the third action node video frame; obtaining a second node sequence number similarity between the first action node sequence number vector and the third action node sequence number vector; setting basic keypoints among multiple keypoints; determining a first node action vector between the basic keypoint and other keypoints besides the basic keypoint based on the first node position information; determining a third node action vector between the basic keypoint and other keypoints besides the basic keypoint based on the third node position information; determining a first body shape vector based on the first body shape data, and determining a second body shape vector based on the second body shape data; determining a second body shape vector based on the first body shape vector and the second body shape data; and determining a second body shape vector based on the first body shape vector and the second body shape data. The following steps are taken: First, body shape similarity is determined using a shape vector. Second, the average similarity of the first and third node action vectors corresponding to multiple other key points at multiple action nodes is compared to the body shape similarity, and this ratio is used to determine the second node action similarity. Third, uniform sampling is performed on the first training video frames between the first action node video frames to obtain the first sampled video frame, and uniform sampling is performed on the third training video frames between the third action node video frames to obtain the third sampled video frame. Based on the first position information of the first sampled video frame, the first action vector between the basic key point and other key points is determined. Based on the third position information of the third sampled video frame, the third action vector between the basic key point and other key points is determined. Fourth, the average similarity of the first and third action vectors corresponding to multiple other key points at multiple sampling points is compared to the body shape similarity, and this ratio is used to determine the second non-node action similarity. Finally, a second reference score is determined based on the second node sequence similarity, the second node action similarity, and the second non-node action similarity.
[0138] According to one embodiment of the present invention, the method for obtaining the similarity of the second node number is similar to the method for obtaining the similarity of the first node number, and the meaning is also similar, indicating the similarity between the action speed of the second person and the standard action speed.
[0139] According to one embodiment of the present invention, the method for obtaining the first node action vector is the same as described above, and the method for obtaining the third node action vector is similar to that for obtaining the second node action vector, and will not be repeated here. However, the difference between the first node action vector and the third node action vector includes not only the difference in action but also the difference in body shape between the first person and the second person. Therefore, when judging whether the action of the second person is standard, the interference of body shape difference can be excluded, and only the difference between the action of the second person and the standard action can be judged. Therefore, the first body shape vector can be determined based on the first body shape data. For example, the first body shape data such as height, arm span, and leg length can be combined to obtain the first body shape vector, and each component of the first body shape vector is a first body shape data. Similarly, the second body shape vector can be determined based on the second body shape data. Furthermore, the cosine similarity between the first body shape vector and the second body shape vector can be determined as the body shape similarity. When determining the action similarity of the second node, the action similarity of the first node position information and the third node position information with the same sequence number can be determined in a similar manner as above, and the average value of each action similarity can be obtained. The ratio of the average action similarity to the body shape similarity can be obtained, thereby eliminating body shape differences and obtaining the action similarity of the second node. This can be used only to describe the similarity between the action of the second person and the standard action when describing the action node.
[0140] According to one embodiment of the present invention, the acquisition method of the first sampled video frame and the first motion vector is the same as described above, and the acquisition method of the third sampled video and the third motion vector is similar to the acquisition method of the second sampled video and the second motion vector described above, and will not be repeated here. Similarly, body shape differences can be excluded, and only the similarity between the second person's actions and the standard actions during the action is acquired. The similarity of the actions of the person in each first sampled video frame and the corresponding third sampled video frame can be obtained in a similar manner as described above, and then averaged. Further, the ratio of the average value to the body shape similarity is determined as the second non-node action similarity, thereby excluding body shape differences and only using it to describe the similarity between the second person's actions and the standard actions during the action.
[0141] According to one embodiment of the present invention, a second reference score can be obtained by weighted averaging the second node sequence similarity, the second node action similarity, and the second non-node action similarity. This second reference score can then be weighted averaging with a second annotation score obtained based on the expert experience of a coach or referee to obtain a second control score, which serves as the accurate score.
[0142] According to one embodiment of the present invention, the evaluation sub-model can process the body shape data of the second person, the sequence number of the third action node video, the third node position information, and the third position information to output a third training physical test score, which may contain errors.
[0143] According to one embodiment of the present invention, the test loss function can be obtained in a manner similar to formula (2), that is, the error between the third training test score and the second control score is solved, and the minimum value among the second node number similarity, the second node action similarity and the second non-node action similarity is solved, and then the ratio of the error to the minimum value is used as the test loss function, so as to improve the evaluation accuracy of the evaluation sub-model for the most important differences in a targeted manner.
[0144] In this way, the body shape difference between the first and second personnel can be considered when determining the similarity between two-node actions and the second non-node actions. By calculating the body shape similarity between the first and second personnel, the body shape difference can be excluded. Only the similarity between the second person's action and the standard action at the action node and the similarity between the second person's action and the standard action during the action can be described, thus improving the accuracy of action evaluation.
[0145] According to an embodiment of the present invention, in step S34, the physical test data evaluation model can be trained using a standard action loss function, a non-standard action loss function, and a test loss function. That is, based on the data input to the physical test data evaluation model, one of the three loss functions can be selected to adjust the parameters of the evaluation sub-model so that all three loss functions are reduced, and after multiple training sessions, the trained physical test data evaluation model is obtained.
[0146] According to one embodiment of the present invention, in step S4, a video of the subject's actions to be tested can be captured by a camera, and in step S5, the subject's body shape data can be obtained. Then, in step S6, the body shape data and the video of the subject's actions are processed by a trained physical fitness assessment model to obtain the subject's physical fitness score, and physical fitness assessment information can be determined based on the score. For example, the physical fitness score can be compared with a scoring threshold; if the score is higher than or equal to the threshold, the physical fitness assessment information is considered qualified; otherwise, it is considered unqualified. The present invention does not limit the specific content included in the physical fitness assessment information.
[0147] According to an embodiment of the present invention, the image-based physical fitness data evaluation method can train a physical fitness data evaluation model using different video samples of various types of people. This enables the model to recognize the movements of people with various body types, thereby determining whether the movements of the person being evaluated are standard. This provides an objective and accurate basis for evaluating the movements of the person being evaluated, reduces subjectivity, improves evaluation accuracy, and provides accurate feedback to the person being evaluated. When determining the standard movement loss function, when the first training video sample is a standard training video of the first person performing a standard movement, the first position information, the sequence number of the first action node video frame, the first key position information, and the body type data of the first person can be obtained. The standard movement loss function is determined by the error between the theoretical value of the first training physical fitness score and the output value of the evaluation sub-model, thereby improving the evaluation accuracy of the evaluation sub-model for standard movements. When determining the loss function for non-standard movements, the minimum value among the first node sequence similarity, first node movement similarity, and first non-node movement similarity can be used to identify the most significant difference between non-standard and standard movements. This allows for targeted amplification of the errors in the first control score and the second training physical test score, thereby enhancing the training strength of the evaluation sub-model for assessing the most significant differences and improving its accuracy. When determining the second control score, the body shape difference between the first and second personnel can be considered when determining the similarity between the two node movements and the second non-node movements. By calculating the body shape similarity between the first and second personnel, body shape differences can be excluded. Only the similarity between the second person's movement and the standard movement at the movement node and the similarity between the second person's movement and the standard movement during the movement process can be described, improving the accuracy of the movement evaluation.
[0148] Figure 3 An exemplary block diagram of a physical fitness data evaluation system based on image acquisition according to an embodiment of the present invention is shown, the system comprising:
[0149] The first acquisition module is used to acquire first training action video samples through a camera. The first training action video sample is a video taken by the first person who has mastered the standard action at a designated position. The content of the first training action video sample is the first person performing the designated action.
[0150] The second acquisition module is used to acquire second training action video samples through a camera. The second training action video samples are videos taken by a second person who has mastered the standard action at a designated position. The content of the second training action video is the second person performing the designated action, and the body data of the second person is different from that of the first person.
[0151] The training module is used to train the physical fitness test data evaluation model based on the first training action video sample, the second training action video sample, the body shape data of the first person and the body shape data of the second person, so as to obtain the trained physical fitness test data evaluation model.
[0152] The shooting module is used to acquire a video of the subject's actions through a camera. The video of the subject's actions is a video taken when the subject is at a designated location, and the content of the video of the subject's actions is the subject performing a designated action.
[0153] The body shape data module is used to acquire the body shape data of the person being tested;
[0154] The evaluation module is used to process the body shape data and video of the test subjects' movements using the trained physical test data evaluation model to obtain the physical test evaluation information of the test subjects.
[0155] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0156] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are merely examples and do not limit the present invention. The objectives of the present invention have been fully and effectively achieved. The functions and structural principles of the present invention have been demonstrated and explained in the embodiments, and any variations or modifications may be made to the implementation of the present invention without departing from the stated principles.
[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for evaluating physical fitness data based on image acquisition, characterized in that, include: The first training action video sample is captured by a camera. The first training action video sample is a video taken of the first person who has mastered the standard action at a designated position. The content of the first training action video sample is the first person performing the designated action. The second training action video sample is captured by a camera. The second training action video sample is a video taken at a designated position by a second person who does not know the standard action. The content of the second training action video is the second person performing the designated action, and the body data of the second person is different from that of the first person. Based on the first training action video sample, the second training action video sample, the body shape data of the first person and the body shape data of the second person, the physical fitness test data evaluation model is trained to obtain the trained physical fitness test data evaluation model. The test involves acquiring a video of the subject's actions using a camera. The video is taken from a designated location and shows the subject performing a specified action. Obtain the body shape data of the person being tested; The trained physical fitness test data evaluation model processes the body shape data and video of the test subjects to obtain the physical fitness test evaluation information of the test subjects. The first training video sample includes standard training videos of the first person performing standard movements, and non-standard training videos of the first person performing non-standard movements; Based on the first training action video samples, the second training action video samples, and the body shape data of the first and second personnel, the physical fitness test data evaluation model is trained to obtain the trained physical fitness test data evaluation model, including: By using standard training videos and body shape data of the first person, the standard motion loss function of the physical test data evaluation model is determined; By using non-standard training videos and body shape data of the first person, the non-standard motion loss function of the physical test data evaluation model was determined; The test loss function of the physical fitness data evaluation model is determined by using the second training action video samples and the body shape data of the second person. The physical fitness test data evaluation model is trained using the standard motion loss function, the non-standard motion loss function, and the test loss function to obtain the trained physical fitness test data evaluation model.
2. The method for evaluating physical fitness data based on image acquisition according to claim 1, characterized in that, Using standard training videos and body shape data from a first-person participant, the standard motion loss function for evaluating the physical fitness test data model was determined, including: The key point detection sub-model of the physical test data evaluation model is used to process the standard training video to obtain the first position information of key points in multiple first training video frames in the standard training video. Input the first position information of key points corresponding to multiple first training video frames into the action node recognition sub-model to determine the first action node video frame. Based on the first position information, determine the first node position information of the key point in the video frame of the first action node; The first position information, the sequence number of the first action node video frame, the first key position information, and the body shape data of the first person are input into the evaluation sub-model of the physical test data evaluation model to obtain the first training physical test score. Based on the first training physical test score, the standard movement loss function is determined.
3. The method for evaluating physical fitness data based on image acquisition according to claim 2, characterized in that, Based on the first training fitness test score, the standard movement loss function is determined, including: According to the formula ; Determine the standard motion loss function ,in, The score for the first training physical test.
4. The method for evaluating physical fitness data based on image acquisition according to claim 2, characterized in that, By using non-standard training videos and body shape data of the first person, the non-standard movement loss function of the physical fitness test data evaluation model was determined, including: The key point detection sub-model of the physical test data evaluation model is used to process non-standard training videos to obtain the second position information of key points in multiple second training video frames in the non-standard training videos. Input the second position information of key points in multiple second training video frames into the action node recognition sub-model to determine the second action node video frame; Based on the second position information, determine the second node position information of the key points in the video frame of the second action node; Based on the sequence number of the first action node video frame, the sequence number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information, a first reference score for the non-standard training video is obtained. Obtain the first labeled score of non-standard training videos; Based on the first reference score and the first annotation score, a first comparative score for the non-standard training video is obtained. The body shape data of the first person, the sequence number of the second action node video, the location information of the second node, and the location information of the second position are input into the evaluation sub-model of the physical test data evaluation model to obtain the second training physical test score. Based on the second training physical test score, the first control score, the sequence number of the first action node video frame, the sequence number of the second action node video, the first node position information, the second node position information, the first position information, and the second position information, the non-standard action loss function of the physical test data evaluation model is determined.
5. The method for evaluating physical fitness data based on image acquisition according to claim 4, characterized in that, Based on the sequence number of the first action node video frame, the sequence number of the second action node video frame, the position information of the first node, the position information of the second node, the first position information, and the second position information, a first reference score for the non-standard training video is obtained, including: Based on the sequence number of the video frame of the first action node, obtain the sequence number vector of the first action node; Based on the sequence number of the video frame of the second action node, obtain the sequence number vector of the second action node; Obtain the similarity of the first node index between the first action node index vector and the second action node index vector; Set basic key points among multiple key points; Based on the first node position information, determine the first node action vector between the basic key point and other key points other than the basic key point; Based on the second node position information, determine the second node action vector between the basic key point and other key points other than the basic key point; The average similarity between the first node action vector and the second node action vector corresponding to multiple other key points at multiple action nodes is determined as the first node action similarity. Uniformly sample the first training video frames between the first action node video frames to obtain the first sampled video frame, and uniformly sample the second training video frames between the second action node video frames to obtain the second sampled video frame. Based on the first position information of the first sampled video frame, determine the first motion vector between the basic key point and other key points besides the basic key point; Based on the second position information of the second sampled video frame, determine the second motion vector between the basic key point and other key points other than the basic key point; The average similarity between the first action vector and the second action vector corresponding to multiple other key points at multiple sampling points is determined as the first non-node action similarity. The first reference score is determined based on the similarity of the first node sequence number, the similarity of the first node action, and the similarity of the first non-node action.
6. The method for evaluating physical fitness data based on image acquisition according to claim 5, characterized in that, Based on the second training physical test score, the first control score, the sequence number of the first action node video frame, the sequence number of the second action node video frame, the first node position information, the second node position information, the first position information, and the second position information, the non-standard action loss function of the physical test data evaluation model is determined, including: According to the formula ; Determine the non-standard movement loss function of the physical test data evaluation model. ,in, As the first control score, For the second training physical test score, The similarity is based on the first node's sequence number. For the similarity of actions at the first node, The similarity is for the first non-node action.
7. The method for evaluating physical fitness data based on image acquisition according to claim 2, characterized in that, Using second training video samples and body shape data of a second person, the test loss function for the physical fitness data evaluation model is determined, including: The key point detection sub-model of the physical test data evaluation model is used to process the second training action video samples to obtain the third position information of key points in multiple third training video frames in the second training action video samples. Input the third position information of key points in multiple third training video frames into the action node recognition sub-model to determine the third action node video frame; Based on the third position information, determine the third node position information of the key points in the video frame of the third action node; Based on the sequence number of the first action node video frame, the sequence number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person and the body shape data of the second person, a second reference score for the second training action video sample is obtained. Obtain the second labeled score of the second training action video sample; Based on the second reference score and the second annotation score, a second control score is obtained for the second training action video sample; Input the body shape data of the second person, the sequence number of the third action node video, the third node position information, and the third position information into the evaluation sub-model of the physical test data evaluation model to obtain the third training physical test score. Based on the third training physical test score, the second control score, the sequence number of the first action node video frame, the sequence number of the third action node video, the first node position information, the third node position information, the first position information, the third position information, the body shape data of the first person and the body shape data of the second person, the test loss function of the physical test data evaluation model is determined.
8. The method for evaluating physical fitness data based on image acquisition according to claim 7, characterized in that, Based on the sequence number of the first action node video frame, the sequence number of the third action node video, the position information of the first node, the position information of the third node, the first position information, the third position information, the body shape data of the first person, and the body shape data of the second person, a second reference score is obtained for the second training action video sample, including: Based on the sequence number of the video frame of the first action node, obtain the sequence number vector of the first action node; Based on the sequence number of the video frame of the third action node, obtain the sequence number vector of the third action node; Obtain the similarity of the second node indices between the first action node index vector and the third action node index vector; Set basic key points among multiple key points; Based on the first node position information, determine the first node action vector between the basic key point and other key points other than the basic key point; Based on the third node position information, determine the third node action vector between the basic key point and other key points other than the basic key point; Based on the first body shape data, determine the first body shape vector, and based on the second body shape data, determine the second body shape vector; Determine the body shape similarity based on the first body shape vector and the second body shape vector; The ratio of the average similarity between the first node action vector and the third node action vector corresponding to multiple other key points at multiple action nodes to the body shape similarity is determined as the second node action similarity. Uniformly sample the first training video frames between the first action node video frames to obtain the first sampled video frame, and uniformly sample the third training video frames between the third action node video frames to obtain the third sampled video frame. Based on the first position information of the first sampled video frame, determine the first motion vector between the basic key point and other key points besides the basic key point; Based on the third position information of the third sampled video frame, determine the third motion vector between the basic key point and other key points besides the basic key point; The ratio of the average similarity of the first action vector and the third action vector corresponding to multiple other key points at multiple sampling points to the body shape similarity is determined as the second non-node action similarity. The second reference score is determined based on the similarity of the second node sequence number, the similarity of the second node action, and the similarity of the second non-node action.
9. A physical fitness data evaluation system based on image acquisition, characterized in that, The system is used to perform the method as described in any one of claims 1-8, comprising: The first acquisition module is used to acquire first training action video samples through a camera. The first training action video sample is a video taken by the first person who has mastered the standard action at a designated position. The content of the first training action video sample is the first person performing the designated action. The second acquisition module is used to acquire second training action video samples through a camera. The second training action video samples are videos taken by a second person who has mastered the standard action at a designated position. The content of the second training action video is the second person performing the designated action, and the body data of the second person is different from that of the first person. The training module is used to train the physical fitness test data evaluation model based on the first training action video sample, the second training action video sample, the body shape data of the first person and the body shape data of the second person, so as to obtain the trained physical fitness test data evaluation model. The shooting module is used to acquire a video of the subject's actions through a camera. The video of the subject's actions is a video taken when the subject is at a designated location, and the content of the video of the subject's actions is the subject performing a designated action. The body shape data module is used to acquire the body shape data of the person being tested; The evaluation module is used to process the body shape data and video of the test subjects' movements using the trained physical test data evaluation model to obtain the physical test evaluation information of the test subjects.
Citation Information
Patent Citations
User action evaluation method and device, and readable medium
CN109635644A
Professional dance evaluation method for realizing human body posture detection based on deep transfer learning
CN112560665A