A training data monitoring method based on deep learning

Through a deep learning-based method, the training image features are extracted using convolutional neural network and the closed area of the action is corrected, which solves the problem of insufficient subjectivity and accuracy of traditional training monitoring methods, and achieves high-precision and real-time feedback of action recognition, improving training effect and safety.

CN120148124BActive Publication Date: 2025-07-22CHENGDU AERONAUTIC POLYTECHNIC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510632084.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-22
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

Traditional training monitoring methods have problems such as strong subjectivity, insufficient accuracy and poor real-time performance. It is difficult to accurately evaluate subtle changes in the movement and promptly feedback deviations, which affects the training effect and safety.

Method used

Using a deep learning-based method, training image features are extracted through convolutional neural networks, target intervals are constructed and closed areas are corrected, and action compliance is determined using similarity comparison and geometric constraints, and objective quantitative indicators are provided.

Benefits of technology

It improves the accuracy and real-time nature of action recognition, reduces subjective errors, can instantly determine the action failure and provide instant reminders, improving training efficiency and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148124B_ABST
    Figure CN120148124B_ABST
Patent Text Reader

Abstract

The present invention discloses a training data monitoring method based on deep learning, which relates to the technical field of action recognition and includes the following steps: S1, collecting training images of a user; S2, constructing a target interval according to the training feature map corresponding to the training image, and extracting the training action closed area of the training image; S3, using the corner points of the training action closed area to correct the training action closed area to obtain a smooth training action closed area; S4, comparing the similarity between the smooth training action closed area and the standard training action, and when the similarity is lower than a set threshold, determining that the training action of the user is unqualified. This training data monitoring method based on deep learning provides an objective quantitative index by comparing with the standard training action, reduces subjective errors, immediately determines unqualified when the similarity is lower than the threshold, and supports instant reminder.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of action recognition, and particularly relates to a training data monitoring method based on deep learning. Background Art

[0002] With the rapid development of artificial intelligence technology, deep learning has shown great application potential in the fields of action recognition, pose estimation, and training monitoring. Traditional training monitoring methods mainly rely on manual observation or simple sensor data, and have problems such as strong subjectivity, low accuracy, and poor real-time performance. For example, in fitness training, rehabilitation training, or sports skill training, it is difficult for coaches or doctors to capture the subtle deviations of actions in real time, resulting in inaccurate training effect evaluation or increased risk of training injuries. Therefore, it is of great practical significance to develop an automated training data monitoring method based on deep learning.

[0003] The disadvantages of traditional training monitoring methods are as follows: 1. Strong subjectivity, manual observation depends on experience, and different evaluators may have different evaluations of the same action; 2. Insufficient accuracy, traditional methods are difficult to capture the subtle changes of actions, resulting in rough evaluation results; 3. Poor real-time performance, manual monitoring cannot provide real-time feedback on action deviations, affecting training efficiency. Deep learning has shown many advantages in the field of action monitoring. For example, for high-precision feature extraction, deep learning models (such as convolutional neural network CNN) can automatically learn the complex features in images and accurately identify action details. Therefore, the present invention proposes an action monitoring technology based on deep learning. Summary of the Invention

[0004] In order to solve the above problems, the present invention proposes a training data monitoring method based on deep learning.

[0005] The technical solution of the present invention is: a training data monitoring method based on deep learning includes the following steps:

[0006] S1. Collect the training images of the user;

[0007] S2. According to the training feature map corresponding to the training image, construct a target interval, and extract the training action closed area of the training image;

[0008] S3. Use the corner points of the training action closed area to correct the training action closed area to obtain a smooth training action closed area;

[0009] S4. Compare the similarity between the smooth training action closed area and the standard training action. When the similarity is lower than the set threshold, it is determined that the training action of the user is unqualified.

[0010] In S4, the structural similarity index (SSIM) or cosine similarity can be used to take into account geometric and feature consistency.

[0011] Furthermore, S2 includes the following sub-steps:

[0012] S21. Use a convolution kernel to slide through the user's training image to obtain a training feature map;

[0013] S22. Based on the training feature map, construct a target interval for the training image;

[0014] S23. Based on the target interval, construct a target function for the training image;

[0015] S24. Extract the corner points of the training image and construct constraint conditions;

[0016] S25. Use the target function and constraint conditions of the training image to generate a target constraint model;

[0017] S26. Take the pixel points in the training image whose pixel values are greater than the target constraint model as the training action closed area.

[0018] The beneficial effects of the above further solution are as follows: In the present invention, the convolution kernel (or filter) is the core component of the convolutional neural network (CNN). Through the sliding window mechanism, local features of the input image are extracted to generate a feature map, where each element is the eigenvalue calculated by the convolution kernel in a certain local area. The target interval is adaptively generated based on the training feature map to enhance the robustness of the target constraint model to individual differences; by combining the target function with geometric constraint conditions, the accuracy and stability of the closed area extraction are improved.

[0019] Furthermore, in S22, take the mean value of all elements in the training feature map as the target weight, take the product of the target weight and the maximum pixel value in the training image as the right endpoint of the target interval, and take the product of the target weight and the minimum pixel value in the training image as the left endpoint of the target interval.

[0020] The beneficial effects of the above further solution are as follows: In the present invention, using the mean value of the feature map as the weight and combining with the extreme values of the image to generate an interval can adapt to the gray-scale distribution of different training images and avoid the sensitivity of the fixed threshold to pixel values.

[0021] Furthermore, in S23, the target function has the following expression:

[0022] ;

[0023] ;

[0024] In the formula, represents the number of elements in the target interval that are greater than the pixel value of the th pixel point in the training image, and represents the The deviation degree between a pixel point and the target interval denotes the right endpoint of the target interval denotes the left endpoint of the target interval denotes the Euclidean distance between the pixel point of the th pixel point in the training image and the pixel point of the maximum curvature point in the training image denotes the pixel point of the maximum curvature point in the training image denotes the pixel value of the

[0025] The beneficial effect of the above further solution is that in the present invention, the maximum curvature point is introduced as a distance reference to strengthen the association between the closed area and the key structure. As a geometric anchor point, it ensures the alignment of the closed area with the key parts of the action (such as joint points). Pixel points belonging to the target interval in the training feature map can be extracted as elements.

[0026] Furthermore, in S24, the constraint condition has the expression of ; in the formula, denotes the number of elements in the target interval that are greater than the pixel value of the th pixel point in the training image, denotes the deviation degree between the th pixel point and the target interval.

[0027] The beneficial effect of the above further solution is that in the present invention, the constraint condition can constrain the pixel deviation, prevent overfitting or under-segmentation, and try to ensure the geometric continuity of the closed area boundary and the action structure (such as limb contour).

[0028] Furthermore, in S25, the result of multiplying the constraint condition by the Lagrange multiplier is added to the partial derivative of the objective function as the target constraint model of the training image.

[0029] The Lagrange multiplier is used to balance the objective function and the constraint condition.

[0030] Furthermore, S3 includes the following sub-steps:

[0031] S31. Extract all corner points included in the closed area of the training action;

[0032] S32. Correct the corner points with a curvature greater than the set curvature threshold to obtain a smooth closed area of the training action.

[0033] Furthermore, in S32, the coordinate position of the corrected corner point is ; in the formula, denotes the abscissa of the original coordinate of the corner point represents the ordinate of the original coordinates of the corner point, represents the gradient magnitude of the first four-neighborhood pixel point near the corner point, represents the gradient magnitude of the second four-neighborhood pixel point near the corner point, represents the gradient magnitude of the third four-neighborhood pixel point near the corner point, represents the gradient magnitude of the fourth four-neighborhood pixel point near the corner point.

[0034] The beneficial effects of the above further solution are: In the present invention, the corner point is a key geometric feature of the action closed area. The four-neighborhood gradient average is used to correct the corner point position, enhancing the geometric continuity of the boundary. The gradient magnitude reflects the pixel change intensity, which can be used to guide the corner point to move towards the real boundary. The average value operation smooths local mutations and avoids excessive offset of the corner point.

[0035] The beneficial effects of the present invention are:

[0036] (1) The deep learning-based training data monitoring method adaptively generates a target interval based on the training feature map, enhancing the adaptability of the target constraint model to individual differences. By combining the objective function and constraint conditions, it can accurately locate the action closed area as much as possible;

[0037] (2) The deep learning-based training data monitoring method eliminates sharp noise and enhances boundary continuity through curvature screening and gradient weighted correction. It uses the four-neighborhood gradient average to enhance the robustness of the corner point position, adapts to dynamic deformation, and corrects key structures (such as joint points). The high-gradient area has a greater weight, guiding the corner point to move towards the real boundary; ensuring the accuracy of the training action closed area;

[0038] (3) The deep learning-based training data monitoring method compares with the standard training action, provides objective quantitative indicators, reduces subjective errors, immediately determines non-compliance when the similarity is lower than the threshold, and supports instant reminders. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flowchart of the deep learning-based training data monitoring method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The following further describes the embodiments of the present invention with reference to the accompanying drawings.

[0041] As Figure 1 shown, the present invention provides a deep learning-based training data monitoring method, including the following steps:

[0042] S1. Collect the training images of the user;

[0043] S2. Construct a target interval based on the training feature map corresponding to the training image, and extract the training action closed region of the training image;

[0044] S3. Use the corner points of the training action closed region to correct the training action closed region to obtain a smoothed training action closed region;

[0045] S4. Compare the similarity between the smoothed training action closed region and the standard training action. When the similarity is lower than the set threshold, it is determined that the user's training action is unqualified.

[0046] In S4, the structural similarity index (SSIM) or cosine similarity can be used to take into account both geometric and feature consistency.

[0047] In the embodiment of the present invention, S2 includes the following sub-steps:

[0048] S21. Use a convolution kernel to slide through the user's training image to obtain a training feature map;

[0049] S22. Construct a target interval for the training image according to the training feature map;

[0050] S23. Based on the target interval, construct a target function for the training image;

[0051] S24. Extract the corner points of the training image and construct constraint conditions;

[0052] S25. Use the target function and constraint conditions of the training image to generate a target constraint model;

[0053] S26. Take the pixel points in the training image whose pixel values are greater than the target constraint model as the training action closed region.

[0054] In the present invention, the convolution kernel (or filter) is the core component of the convolutional neural network (CNN). By means of the sliding window mechanism, local features of the input image are extracted to generate a feature map, where each element is the feature value calculated by the convolution kernel in a certain local area. The target interval is adaptively generated based on the training feature map to enhance the robustness of the target constraint model to individual differences; combining the target function with geometric constraint conditions can improve the accuracy and stability of the closed region extraction.

[0055] In the embodiment of the present invention, in S22, the mean value of all elements in the training feature map is used as the target weight, the product of the target weight and the maximum pixel value in the training image is used as the right endpoint of the target interval, and the product of the target weight and the minimum pixel value in the training image is used as the left endpoint of the target interval.

[0056] In the present invention, the mean value of the feature map is used as the weight, combined with the image extreme value generation interval, to adapt to the gray-scale distribution of different training images and avoid the sensitivity of the fixed threshold to pixel values.

[0057] In the embodiment of the present invention, in S23, the objective function has the following expression:

[0058] ;

[0059] ;

[0060] In the formula, represents the number of elements in the target interval that are greater than the pixel value of the th pixel point in the training image, represents the deviation degree between the th pixel point and the target interval, represents the right endpoint of the target interval, represents the left endpoint of the target interval, represents the Euclidean distance between the th pixel point in the training image and the pixel point of the maximum curvature point in the training image, represents the pixel point of the maximum curvature point in the training image, represents the pixel value of the th pixel point in the training image.

[0061] In the present invention, the maximum curvature point is introduced as a distance reference to strengthen the correlation between the closed region and the key structure. As a geometric anchor point, it ensures the alignment of the closed region with the key parts of the action (such as joint points).

[0062] In the embodiment of the present invention, in S24, the constraint condition has the following expression ; In the formula, represents the number of elements in the target interval that are greater than the pixel value of the th pixel point in the training image, represents the deviation degree between the th pixel point and the target interval.

[0063] In the present invention, the constraint condition can constrain the pixel deviation, prevent overfitting or under-segmentation, and try to ensure the geometric continuity of the closed region boundary and the action structure (such as limb contour).

[0064] In the embodiment of the present invention, in S25, the result of multiplying the constraint condition by the Lagrange multiplier is added to the partial derivative of the objective function to serve as the target constraint model of the training image.

[0065] Lagrange multipliers are used to balance the objective function and the constraints.

[0066] In the embodiment of the present invention, S3 includes the following sub-steps:

[0067] S31. Extract all corner points included in the closed area of the training action;

[0068] S32. Correct the corner points with curvature greater than the set curvature threshold to obtain a smooth closed area of the training action.

[0069] In the embodiment of the present invention, in S32, the coordinate position of the corrected corner point is ; where represents the abscissa of the original coordinate of the corner point, represents the ordinate of the original coordinate of the corner point, represents the gradient magnitude of the first four-neighborhood pixel point near the corner point, represents the gradient magnitude of the second four-neighborhood pixel point near the corner point, represents the gradient magnitude of the third four-neighborhood pixel point near the corner point, represents the gradient magnitude of the fourth four-neighborhood pixel point near the corner point.

[0070] In the present invention, the corner point is a key geometric feature of the closed area of the action. The position of the corner point is corrected by the average of the four-neighborhood gradients to enhance the geometric continuity of the boundary. The gradient magnitude reflects the intensity of pixel change and can be used to guide the corner point to move towards the real boundary. The averaging operation smooths local mutations and avoids excessive deviation of the corner point.

[0071] Those of ordinary skill in the art will realize that the embodiments described herein are to assist the reader in understanding the principles of the present invention and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations without departing from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A training data monitoring method based on deep learning, characterized in that, It includes the following steps: S1. Collect the training images of the user; S2. Based on the training feature maps corresponding to the training images, construct a target interval and extract the training action closed region of the training images; S3. Use the corner points of the training action closed region to correct the training action closed region to obtain a smooth training action closed region; S4. Compare the similarity between the smooth training action closed region and the standard training action. When the similarity is lower than the set threshold, it is determined that the user's training action is unqualified; The S2 includes the following sub-steps: S21. Use a convolutional kernel to slide through the training images of the user to obtain training feature maps; S22. Based on the training feature maps, construct a target interval for the training images; S23. Based on the target interval, construct an objective function for the training images; S24. Extract the corner points of the training images and construct constraint conditions; S25. Use the objective function and constraint conditions of the training images to generate an objective constraint model; S26. Take the pixel points in the training images with pixel values greater than the objective constraint model as the training action closed region; In the S22, take the mean value of all elements in the training feature maps as the target weight, take the product of the target weight and the maximum pixel value in the training images as the right endpoint of the target interval, and take the product of the target weight and the minimum pixel value in the training images as the left endpoint of the target interval; In S23, the objective function has the following expression: ; ; In the formula, represents the number of elements in the target interval that are greater than the pixel value of the -th pixel point in the training image, represents the deviation degree between the -th pixel point and the target interval, represents the right endpoint of the target interval, represents the left endpoint of the target interval, represents the Euclidean distance between the -th pixel point in the training image and the pixel point of the maximum curvature point in the training image, represents the pixel point of the maximum curvature point in the training image, represents the pixel value of the -th pixel point in the training image; In the S24, the constraint condition has an expression of ; in the formula, represents the number of elements in the target interval that are greater than the pixel value of the -th pixel point in the training image, and represents the deviation degree between the -th pixel point and the target interval; In the S25, add the result of multiplying the constraint conditions by the Lagrange multiplier to the partial derivative of the objective function as the objective constraint model of the training images.

2. The training data monitoring method based on deep learning according to claim 1, characterized in that, The S3 includes the following sub-steps: S31. Extract all the corner points included in the training action closed region; S32. Correct the corner points with curvature greater than the set curvature threshold to obtain a smooth training action closed region.

3. The training data monitoring method based on deep learning according to claim 2, characterized in that In S32, the coordinate position of the corrected corner point is ; where represents the abscissa of the original coordinate of the corner point, represents the ordinate of the original coordinate of the corner point, represents the gradient magnitude of the first four-neighborhood pixel point near the corner point, represents the gradient magnitude of the second four-neighborhood pixel point near the corner point, represents the gradient magnitude of the third four-neighborhood pixel point near the corner point, represents the gradient magnitude of the fourth four-neighborhood pixel point near the corner point.

Citation Information

Patent Citations

  • Low-dose CT multi-target image reconstruction method of unsupervised learning based on ADMM

    CN118298051A

  • Medical ultrasonic image recognition system and method based on deep learning

    CN118334417A