3D motion dangerous action monitoring method based on image reconstruction

The method addresses inefficiencies in manual motion evaluation by using video preprocessing, dual encoders, and dynamic time warping for real-time detection of dangerous motions, enhancing automation and reducing resource consumption.

CN120318906APending Publication Date: 2025-07-15FUZHOU INSTITUE OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510433954.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art lacks automated intelligent motion hazard monitoring methods in users' daily exercise training. Relying on manual judgment leads to strong subjectivity, human resources consuming and difficult to meet the needs of diverse and complex sports forms, and lightweight models are difficult to ensure practicality and efficiency in sports hazard analysis.

Method used

Using a 3D motion hazardous motion monitoring method based on image reconstruction, through video preprocessing, feature extraction, image reconstruction, 3D pose prediction and motion injury risk assessment, a multi-layer perceptron network and F-DTW algorithm are used for pose estimation and hazardous motion recognition, and an unsupervised learning model is constructed to reduce dependence on the labeled data set.

Benefits of technology

It realizes efficient and accurate motion posture estimation and dangerous action recognition in complex motion scenarios, and can instantly identify and warn of dangerous actions on the mobile terminal, reduce the complexity of model training, and improves the intelligence and practicality of motion monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318906A_ABST
    Figure CN120318906A_ABST
Patent Text Reader

Abstract

The invention claims to protect a 3D motion dangerous action monitoring method based on image reconstruction. A system architecture comprises a video preprocessing module, a feature extraction module, an image reconstruction module, a 3D attitude prediction module and a motion injury risk evaluation module. In order to solve the problems that a traditional manual judgment technology is limited, a 3D motion posture data set with labels is lacked and a lightweight model is required, a 3D posture estimation method based on image reconstruction and unsupervised learning is adopted, a motion video is analyzed to obtain motion sequence data, and the motion sequence data is analyzed to obtain the motion sequence data. And whether the exercise posture of the user is standard or not and whether the user has an injury risk or not are judged by using an exercise injury risk evaluation method. The invention provides a lightweight and high-precision motion dangerous action monitoring method for a user under the conditions of multiple scenes, multiple targets, multiple views and the like, and has very high application requirements and popularization values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the field of computer vision technology, and specifically provides a 3D motion dangerous action monitoring method based on image reconstruction. Background Art

[0002] In the daily exercise training of users, the evaluation of the accuracy, placeability, and completion quality of actions almost entirely depends on human eye observation and empirical judgment. The manual evaluation method is not comprehensive enough, subject to a large subjective influence, and consumes a lot of human resources. Due to problems such as the high-dimensional characteristics of the human body and the diversity and complexity of sports actions, there is a lack of an automated intelligent sports efficient evaluation scheme for the testing and evaluation of sports actions. The efficient and intelligent motion dangerous action monitoring method is restricted by the following aspects: (1) Limitations of the 3D motion pose annotated dataset. Currently, most of the mainstream annotated 2D and 3D datasets are used for the model training of pose estimation, and when evaluating the standardization of relevant motion poses, they mostly rely on temporary manually annotated datasets. The traditional manually annotated datasets cannot meet the task requirements of various sports forms. The repetitive and simple work greatly wastes human resources, and at the same time, the annotation quality of the dataset is subject to a large subjective influence, affecting the model analysis effect.

[0003] (2) High demand for lightweight models. For the motion dangerous action analysis model, the lightweight model is an important means to improve the practicality of the model. A complex video action analysis model will increase the computing power requirements of the deployment device and reduce the efficiency of model analysis. Once a dangerous action occurs, if it cannot be stopped immediately, it may cause serious injuries to the exerciser in a short time.

[0004] To provide a lightweight and high-precision motion dangerous action monitoring method for users in multi-scene, multi-target, and multi-view situations, allowing users to record actions, identify postures, and warn of dangers during daily exercise, which has important practical significance for a healthy life and also has strong application requirements and promotion value. Summary of the Invention

[0005] The technical solution of the present invention aims at the problems restricted by the above traditional technologies, and provides a 3D motion dangerous action monitoring method based on image reconstruction, allowing users to record actions, identify postures, and warn of dangers during daily exercise.

[0006] The technical solution adopted by the present invention to solve the above technical problems is as follows: A 3D motion dangerous action monitoring method based on image reconstruction, characterized by including the following steps: S1. Video preprocessing: Preprocess the input motion video to extract the human motion action frame sequence and background pictures.

[0007] S2. Feature extraction: By using two independent encoders to decompose the feature of the human motion action frame sequence, the time-varying feature and time-invariant feature of the motion posture are obtained.

[0008] S3. Image reconstruction: The time-varying feature and time-invariant feature of the extracted motion posture and the background picture are used for image reconstruction through an image generator to optimize the extraction of the time-varying feature of the motion posture.

[0009] S4. 3D pose prediction: The time-varying feature of the motion posture is input into a multi-layer perceptron network for 3D pose estimation to generate a joint point coordinate sequence, that is, the actual action sequence, so as to realize the capture of 3D key points of the moving human body. The obtained joint point coordinate sequence is subjected to geometric projection mapping to obtain 2D pose estimation.

[0010] S5. Motion injury risk assessment: Construct a multi-dimensional template sequence of the actual action sequence and the action sequence in the standard template library, compare the two multi-dimensional template sequences, calculate the improved dynamic time warping distance, that is, the F-DTW value, and judge their similarity for injury risk assessment.

[0011] Furthermore, in step S1, the preprocessing specifically includes: S1-1: The RGB frame is converted into the YUV color space by using an adaptive grayscale algorithm and the luminance channel is extracted; S1-2: The human foreground is segmented by using the GrabCut algorithm, and the human motion action frame sequence and the background picture .

[0012] Furthermore, in step S1-2, when multiple targets are detected to overlap, the human body bounding box tracking based on Kalman filtering is used to realize ID continuous tracking through the similarity matching of joint point heat maps.

[0013] Furthermore, in step S2, the feature decomposition includes performing feature decomposition on the input cropped image and decomposing the feature into two parts, the time-varying feature and time-invariant feature of the motion posture, by respectively using two independent encoders.

[0014] Furthermore, to make the extracted pose features contain more effective information, two independent encoders are respectively constructed to extract the time-varying feature and time-invariant feature of the motion posture in the video, that is, the image is mapped into a set of time-varying features to obtain the components that change with time in the image , and then mapped into a set of time-invariant components to extract the features in the video that are less affected by time changes . Among them and Is the encoding function for two independent encoders.

[0015] Further, in step S3, image reconstruction is performed through a neural network Combines the obtained time-varying features and time-invariant features of the motion posture with the background image through an image generator to reconstruct the image at the current moment , .

[0016] Even further, in step S3, by calculating the reconstructed image The pixel-level difference from the input image is used as an unsupervised loss. By minimizing the image reconstruction loss, the network's latent representation is encouraged to focus on the time-varying features of the foreground object's motion posture, realizing the extraction optimization of the time-varying features of the motion posture in the feature extraction module. Among them, the training loss function of the image reconstruction network is:

[0017] Further, in step S4, the time-varying features of the motion posture Are input into a multi-layer perceptron network For 3D pose estimation. Apply spatial softmax on the 3D voxel grid to locate 17 human body key points, and optimize the joint angles through kinematic chain constraints to obtain a sequence of 3D joint point coordinates.

[0018] Even further, in step S4, the sequence of 3D joint point coordinates obtained from 3D pose estimation is geometrically projected to obtain a 2D pose prediction structure.

[0019] Even further, in step S4, the projected 2D pose data and the existing annotated 2D standard template library action sequence data are input together into a two-dimensional pose discriminator for training until convergence, so that the geometric projection effect truly reflects the pose features.

[0020] Further, in step S5, the construction of the multi-dimensional template sequence includes the 2D pose data, motion rate, and direction of the actual action sequence and the standard template library action sequence.

[0021] Among them, considering that in the actual action sequence and the standard template library action sequence, the same action may have different motion rates due to individual differences, the time axis is stretched or compressed so that actions with different rates can be effectively compared and matched.

[0022] Even further, in step S5, the two multi-dimensional template sequence comparison methods include copying points on the time series and then assigning corresponding values, comparing the multi-dimensional template sequences of the actual action sequence and the standard template library action sequence, and are used to calculate the F-DTW value of the two sequences. The steps include: S5-1: Distance matrix construction. Calculate the distances between all pairs of points in two time series (i.e., the standard template and the action to be recognized), and form a distance matrix.

[0023] S5-2: Find the optimal path. Using dynamic programming techniques, start from the upper left corner of the matrix and end at the lower right corner to find a path that minimizes the sum of the distances of the point pairs on the path. This path represents the best alignment between the two sequences.

[0024] S5-3: Calculate the total distance. The sum of the distances of the point pairs on the path is the F-DTW distance between the two sequences, and this distance reflects their similarity. The F-DTW distance calculation formula is: , where A and B are two action sequences to be compared (A is the standard template, B is the actual action); i and j are the indices of the time points in the sequences (i ∈ A, j ∈ B); γ is the direction weight coefficient (direction consistency constraint).

[0025] Furthermore, in step S5, the method for judging similarity includes comparing the calculated F-DTW value with a preset threshold. If the F-DTW value is less than or equal to this threshold, it is considered that the actual action is similar to the standard action, that is, the action execution is standard and there is no risk of injury; if it is greater than the threshold, it is considered that there is a risk of danger and a warning prompt is given.

[0026] Compared with the prior art, the beneficial effects of the present invention are: The present invention proposes a 3D pose estimation scheme based on unsupervised learning. Compared with existing pose estimation methods, by constructing unsupervised learning, the dependence of model training on the 3D motion pose estimation annotation data set is reduced, and by constructing the loss function of image reconstruction to optimize the extraction of pose time-varying features, the effectiveness of feature extraction for estimation is further improved, and the robustness of the algorithm in complex motion scenarios is enhanced. At the same time, the three-dimensional motion data obtained according to the present invention will also be processed by a pose analysis algorithm to realize the motion correction of non-compliant postures of athletes during training.

[0027] (2) The present invention proposes a pose estimation and anomaly detection algorithm based on computer vision and motion action sequences, which can significantly improve the operation efficiency of motion pose estimation on the basis of ensuring accuracy. The algorithm model based on sequence classification can meet the real-time requirements in actual motion monitoring, identify dangerous actions in less than one second, and can be deployed on mobile devices for wide applications.

[0028] The present invention will be explained in detail below in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings

[0029] Figure 1System architecture diagram of a 3D motion dangerous action monitoring method based on image reconstruction proposed by the present invention Figure 2 Flowchart of the working process of a 3D motion dangerous action monitoring method based on image reconstruction proposed by the present invention Figure 3 Flowchart of image preprocessing proposed by the present invention Figure 4 Flowchart of feature extraction based on image reconstruction proposed by the present invention Figure 5 Flowchart of 3D pose prediction proposed by the present invention Figure 6 Flowchart of motion injury risk assessment proposed by the present invention Specific implementation manner

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] As Figure 1 shown, a 3D motion dangerous action monitoring method based on image reconstruction, its system architecture includes five modules: video preprocessing module, feature extraction module, image reconstruction module, 3D pose prediction module, and motion injury risk assessment module.

[0032] The described 3D motion dangerous action monitoring method based on image reconstruction, as Figure 2 shown, includes the following steps: S1. Video preprocessing: Preprocess the input motion video to extract the human motion action frame sequence and background pictures.

[0033] S2. Feature extraction: Decompose the features of the human motion action frame sequence by using two independent encoders to obtain the time-varying features and time-invariant features of the motion posture.

[0034] S3. Image reconstruction: Reconstruct the extracted time-varying features and time-invariant features and the background pictures through an image generator to optimize the extraction of the time-varying features of the motion posture.

[0035] S4. 3D pose prediction: Input the time-varying features of the motion posture into the multi-layer perceptron network for 3D pose estimation, generate the joint point coordinate sequence, that is, the actual action sequence, and realize the capture of 3D key points of the moving human body. The obtained joint point coordinate sequence is subjected to geometric projection mapping to obtain 2D pose estimation.

[0036] S5, Motion Injury Risk Assessment: Construct a multi-dimensional template sequence of the actual action sequence and the action sequence in the standard template library, compare the two multi-dimensional template sequences, calculate the F-DTW value, and judge their similarity for injury risk assessment.

[0037] The video preprocessing module described above, as Figure 3 shown, includes preprocessing the input motion video, converting the RGB frame to the YUV color space and extracting the luminance channel using an adaptive grayscale algorithm, performing human foreground segmentation through the GrabCut algorithm, and retaining the human motion action frame sequence and the background picture .

[0038] When multi-object overlap is detected, use human bounding box tracking based on Kalman filtering to achieve continuous ID tracking through joint point heat map similarity matching.

[0039] The feature extraction module described above, as Figure 4 shown, to make the extracted pose features contain more effective information, the input human motion action frame sequence is decomposed into features by using two independent encoders. The feature decomposition is carried out by using an encoder to decompose the features into two parts, the time-varying features and the time-invariant features of the motion pose. The image encoding process is . Among them, represents using a neural network to map the image to a set of time-varying features to obtain the components that change with time in the image. represents using a neural network to map the image to time-invariant components and extract the features in the video that are less affected by time changes.

[0040] The image reconstruction module described above, as Figure 4 shown, combines the obtained time-varying features and time-invariant features of the motion pose with the background picture through a picture generator by a neural network to reconstruct the image at the current moment . .

[0041] Calculate the pixel-level difference between the reconstructed image and the input image as the unsupervised loss , and encourage the network latent representation to focus on the time-varying features of the foreground object's motion pose by minimizing the image reconstruction loss, so as to realize the extraction optimization of the time-varying features in the feature extraction module.

[0042] The 3D pose prediction module described above, asFigure 5 As shown, the time-varying features of the motion postures extracted in the feature extraction module are input into the multi-layer perceptron network for 3D pose estimation. The spatial softmax is applied on the 3D voxel grid to locate 17 human body key points, and the joint angles are optimized through kinematic chain constraints to obtain the 3D joint point coordinate sequence.

[0043] The 3D joint point coordinate sequence is geometrically projected to map and obtain the 2D pose estimation. The projected 2D pose data and the existing annotated 2D standard template library action sequence data are input into the two-dimensional pose discriminator for training until convergence, so that the geometric projection effect can truly reflect the pose features.

[0044] The motion injury risk assessment module described above, as Figure 6 shown, includes constructing a multi-dimensional template sequence, that is, the 2D pose data, motion rate and direction of the actual action sequence and the standard template library action sequence.

[0045] Considering that in the actual action sequence and the standard template library action sequence, the same action may have different motion rates due to individual differences, the time axis is stretched or compressed so that actions with different rates can be effectively compared and matched.

[0046] The two multi-dimensional template sequences are compared. After copying the points on the time series and assigning corresponding values, the multi-dimensional template sequences of the actual action sequence and the standard template library action sequence are compared, and the F-DTW value of the two sequences is calculated. The calculation steps include: S1 Distance matrix construction: Calculate the distances between all pairs of points between the two time series to form a distance matrix.

[0047] S2: Finding the optimal path: Using dynamic programming techniques, starting from the upper left corner of the matrix to the lower right corner, find a path that minimizes the sum of the distances of the pairs of points on the path. This path represents the best alignment between the two sequences.

[0048] S3: Calculating the total distance: The sum of the distances of the pairs of points on the path is the F-DTW distance between the two sequences, and this distance reflects their similarity. The F-DTW distance calculation formula is: , where A and B are the two action sequences to be compared (A is the standard template, B is the actual action); i and j are the indices of the time points in the sequences (i ∈ A, j ∈ B); γ is the direction weight coefficient (direction consistency constraint).

[0049] According to the calculated F-DTW value, compare it with a preset threshold to determine the similarity between the actual action and the standard action. If the F-DTW value is less than or equal to this threshold, it is considered that the actual action is similar to the standard action, that is, the action execution is standardized and there is no risk of injury; if it is greater than the threshold, it is considered that there is a risk of danger and a warning prompt is given.

[0050] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A 3D motion dangerous action monitoring method based on image reconstruction, whose system architecture includes five modules: video preprocessing module, feature extraction module, image reconstruction module, 3D posture prediction module, and sports injury risk assessment module. The characteristics are: The video preprocessing module is used to preprocess the input motion video and extract the human motion action frame sequence and background picture. The feature extraction module is used to perform feature decomposition on the human motion action frame sequence extracted by preprocessing by using two independent encoders to obtain time-varying features and time-invariant features of the motion posture. The image reconstruction module is used to reconstruct the time-varying features and time-invariant features of the motion posture obtained by the feature extraction module and the background image obtained by the video preprocessing module through an image generator, and optimize the extraction of the time-varying features of the motion posture by minimizing the image reconstruction loss. The 3D posture prediction module is used to input the time-varying features of the motion posture into the regression network for 3D posture estimation, generate a joint point coordinate sequence, i.e., an actual action sequence, and realize 3D key point capture of the moving human body. The obtained joint point coordinate sequence is subjected to geometric projection mapping to obtain 2D posture estimation. The sports injury risk assessment module is used to construct a multi-dimensional template sequence of the actual action sequence and the standard template library action sequence, and compare the two multi-dimensional template sequences to calculate the improved dynamic time warping distance, i.e., the F-DTW value, and compare it with the preset threshold. If the F-DTW value is less than or equal to the threshold, it is considered that the actual action is similar to the standard action, that is, the action is executed in a standard manner and there is no risk of injury; if it is greater than the threshold, it is considered that there is a dangerous risk and an early warning prompt is issued.

2. The 3D motion dangerous action monitoring method based on image reconstruction according to claim 1, wherein: The preprocessing uses an adaptive grayscale algorithm to convert the video RGB frame into the YUV color space and extract the luminance channel; the human foreground is segmented by the GrabCut algorithm, and the human motion action frame sequence I is retained. crop and the background picture B. When there are multiple targets overlapping, the human body bounding box tracking based on Kalman filtering is used, and ID continuous tracking is achieved through the similarity matching of joint point heat maps.

3. A 3D motion dangerous action monitoring method based on image reconstruction according to claim 1, characterized in that: The feature decomposition module extracts motion posture features, including constructing two independent encoders, which respectively map the image I crop to a set of time-varying features of the motion posture to obtain the components that change with time in the image, that is, I tv = E1(I crop ), and then map it to a set of time-invariant components to extract the features in the video that are less affected by time changes, that is, I ti = E2(I crop ).

4. A 3D motion dangerous action monitoring method based on image reconstruction according to claim 1, characterized in that: The image reconstruction is performed through a neural network D The obtained time-varying features and time-invariant features of the motion posture are combined with the background image through an image generator to reconstruct the image at the current moment Calculate the reconstructed image The pixel-level difference between the reconstructed image and the input image is used as an unsupervised loss By minimizing the image reconstruction loss, the network latent representation is encouraged to focus on the time-varying features of the motion posture of the foreground object, and the extraction optimization of the time-varying features I of the motion posture in the feature extraction module is realized tv of the extraction optimization 5. A 3D motion dangerous action monitoring method based on image reconstruction according to claim 1, characterized in that: 3D pose estimation is to input the time-varying features I of the motion pose tv into a multi-layer perceptron network for 3D pose estimation, apply spatial softmax on the 3D voxel grid to locate 17 human body key points, optimize the joint angles through kinematic chain constraints, and obtain the 3D joint point coordinate sequence.

6. The 3D motion dangerous action monitoring method based on image reconstruction according to claim 5, wherein: The 3D joint point coordinate sequence is used for geometric projection and mapping to obtain 2D posture estimation. At this time, the projected 2D posture data and the existing annotated 2D standard template library action sequence data are input into the two-dimensional posture discriminator training until convergence.

7. The 3D motion dangerous action monitoring method based on image reconstruction according to claim 1, wherein: The construction of the multi-dimensional template sequence includes 2D posture data, movement speed and direction of the actual action sequence and the standard template library action sequence.

8. The 3D motion dangerous action monitoring method based on image reconstruction according to claim 7, wherein: Considering the actual action sequence and the standard template library action sequence, the same action may result in different movement rates due to individual differences. By stretching or compressing the time axis, actions of different rates can be effectively compared and matched.

9. A 3D motion dangerous action monitoring method based on image reconstruction according to claim 1, characterized in that: The two multi-dimensional template sequence comparison method includes copying points on the time series and then assigning corresponding values, comparing the actual action sequence with the multi-dimensional template sequence of the standard template library action sequence, and calculating the F-DTW values of the two sequences.

10. A 3D motion dangerous action monitoring method based on image reconstruction according to claim 9, characterized in that: Calculating the F-DTW value of two sequences includes: S1: distance matrix construction, calculate the distances of all point pairs between two time series to form a distance matrix. S2: Find the best path, using dynamic programming techniques, starting from the upper left corner of the matrix and ending at the lower right corner, to find a path that minimizes the sum of the distances between the points on the path. This path represents the best alignment between the two sequences. S3: Calculate the total distance. The sum of the distances between point pairs on the path is the F-DTW distance between the two sequences, and this distance reflects their similarity. The formula for calculating the F-DTW distance is as follows: F-DTW_Distance = ∑(|A i - B j | + γ·|cosθ A - cosθ B |), where A and B are two action sequences to be compared (A is the standard template and B is the actual action); i and j are the indices of time points in the sequences (i ∈ A, j ∈ B); γ is the direction weight coefficient (direction consistency constraint).