Automatic body posture recognition system based on video analysis
By combining parallel detection and information fusion, human skeleton tracking, and temporal convolution and spatiotemporal graph convolution, the stability and consistency problems of traditional body pose recognition methods in complex environments are solved, and high-precision and robust body pose recognition is achieved.
Patent Information
- Application Number
- CN202511475910.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2025-11-14
AI Technical Summary
Existing body pose recognition methods lack stability and temporal consistency in complex environments, motion blur, and camera perspective changes, making it difficult to meet the requirements of real-time performance, high robustness, and interpretability. Furthermore, traditional methods are susceptible to environmental interference, have simple matching strategies, and involve frequent identity switching, resulting in broken pose trajectories and poor action recognition performance.
By employing parallel detection and information fusion, the system utilizes the precise location of high-resolution features and the rich semantics of low-resolution features, combined with the topological structure and key point locations of the human skeleton for tracking. It also combines temporal convolution and spatiotemporal graph convolution to jointly model spatial and temporal features, generating stable and continuous human skeleton motion trajectories and performing cross-view fusion.
It significantly improves the accuracy of human key point detection and the robustness of the network to environmental changes, generates more stable and continuous human skeleton motion trajectories, ensures the overall ability to judge global actions, and improves the stability, accuracy and adaptability of body posture recognition.
Smart Images

Figure CN120954103A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of body posture recognition technology, specifically to an automatic body posture recognition system based on video analysis. Background Technology
[0002] Existing body pose recognition methods mostly rely on single-frame keypoint detection. In complex environments, with motion blur, and changes in camera perspective, their detection stability and temporal consistency are insufficient, making it difficult to meet the requirements of real-time performance, high robustness, and interpretability during motion. Traditional keypoint detection methods typically downsample the image multiple times, drastically reducing the image's spatial resolution, leading to irreversible loss of feature information and limited keypoint localization accuracy. Traditional target tracking methods rely on appearance features, are easily affected by environmental interference, and have simple matching strategies with frequent identity switching, resulting in broken pose trajectories. Traditional pose recognition methods are mostly based on local fragment feature classification, easily losing global action context, leading to poor action recognition performance during motion. Summary of the Invention
[0003] To address the aforementioned issues and overcome the shortcomings of existing technologies, this invention provides an automatic body pose recognition system based on video analysis. Traditional keypoint detection methods typically involve multiple downsampling of images, drastically reducing spatial resolution and leading to irreversible loss of feature information and limited keypoint localization accuracy. This solution employs parallel detection and information fusion, leveraging the precise location of high-resolution features and the rich semantics of low-resolution features to significantly improve the detection accuracy of human keypoints. It also enhances the network's generalization ability to scale changes, occlusion, and abnormal poses. Furthermore, it provides higher-precision keypoint coordinates and confidence levels through probability heatmaps, facilitating further operations by subsequent modules. Finally, it addresses the limitations of traditional target tracking methods, which rely on appearance features, are susceptible to environmental interference, and have simple matching strategies. Frequent identity switching can lead to broken posture trajectories. This solution tracks targets based on the topological structure and key point positions of the human skeleton, exhibiting strong robustness to changes in lighting and clothing. By measuring the skeleton similarity across different frames, it effectively improves the accuracy and reliability of matching, generating more stable and continuous human skeleton motion trajectories. Addressing the issue that traditional posture recognition methods often rely on local fragment feature classification, easily losing global action context and resulting in poor action recognition during motion, this solution combines temporal convolution and spatiotemporal graph convolution to jointly model spatial and temporal features, obtaining a complete action representation and outputting the action category and its confidence level. This ensures the overall ability to discriminate global actions, effectively improving the stability, accuracy, and adaptability of body posture recognition.
[0004] The present invention provides an automatic body posture recognition system based on video analysis, including a video acquisition module, a preprocessing module, a key point detection module, a skeleton tracking module, an action recognition module, and a cross-view fusion module;
[0005] The video acquisition module acquires multiple video streams containing the athlete's body posture from different perspectives, performs standardization processing on the multiple video streams, and outputs the original frame stream.
[0006] The preprocessing module optimizes the original frame stream, extracts the athlete's human body region, and outputs a human body image frame stream.
[0007] The key point detection module performs feature extraction and human key point detection on the human image frame stream based on a convolutional neural network, and outputs a two-dimensional key point set containing the coordinates of human key points.
[0008] The skeleton tracking module detects the human skeleton in each frame of the human image frame stream based on a two-dimensional key point set, calculates the skeleton similarity, assigns a unique ID to each human skeleton, and outputs the motion trajectory of the human skeleton.
[0009] The action recognition module smooths the motion trajectory of the human skeleton, identifies the action category of the entire human image frame stream, and outputs the action category label and its confidence level.
[0010] The cross-view fusion module associates human skeletons with the same ID from different viewpoints, performs weighted fusion of the action category labels of the same athlete, and outputs the final body posture category.
[0011] Furthermore, the key point detection module includes a feature extraction unit, a feature fusion unit, a heatmap prediction unit, and a coordinate regression unit;
[0012] The feature extraction unit constructs a deep convolutional layer based on the ResNet structure, and extracts multi-level features of the human image frame stream through layer-by-layer convolution and pooling operations.
[0013] The feature fusion unit adopts the HRNet structure to perform parallel detection and information fusion of multi-level features. It fuses contextual semantics with high-resolution features of small-scale joints in human image frame streams, preserves spatial information for low-resolution features of large-scale motion, obtains fused multi-scale features, and outputs a fused feature map.
[0014] The heatmap prediction unit sets the number of human key points to be detected to N, uses a 1×1 convolutional layer to output a feature map with N channels from the fused feature map, each channel corresponding to a human key point, and uses the sigmoid function to normalize the feature map of each channel into a probability heatmap, where the value of each pixel on the probability heatmap represents the probability that the corresponding human key point is located at that position.
[0015] The coordinate regression unit converts the probability heatmap into human keypoint coordinates and confidence scores: it performs peak detection on the probability heatmap, obtains the peak coordinates as human keypoint coordinates, and uses the peak as the confidence score of the human keypoint, outputting a set containing N human keypoint coordinates, which is a two-dimensional keypoint set describing human posture.
[0016] Furthermore, the skeleton tracking module includes a topology construction unit, a trajectory prediction unit, a skeleton matching unit, and a target tracking unit;
[0017] The topology building unit acquires the MPII human anatomy topology standard, defines the skeleton connection set according to the MPII human anatomy topology standard, and constructs a structured human skeleton by spatial matching the two-dimensional keypoint set, in the following form: ;
[0018] In the formula, Represents a frame. Represents the human skeleton. and These represent the beginning and end points of the bone, respectively. and Represents the coordinates of the skeleton. Represents the join function. Represents a skeleton connection set;
[0019] The trajectory prediction unit detects the human skeleton appearing in each frame of the human image frame stream, models the motion trajectory of the human skeleton through Kalman filtering, and predicts the position of the human skeleton in the current frame.
[0020] The skeleton matching unit matches the detection result of the current frame with the human skeleton of the previous frame and calculates the skeleton similarity using weighted Euclidean distance.
[0021] The target tracking unit assigns an ID to the human skeleton detected in each frame based on the calculated skeleton similarity. When a new target appears, a new ID is created. When the human skeleton leaves the field of view, it is marked as disappeared, and the motion trajectory of the human skeleton is output.
[0022] Furthermore, the action recognition module includes a temporal smoothing unit, a temporal convolution unit, a spatiotemporal graph convolution unit, and an action output unit;
[0023] The temporal smoothing unit uses an exponentially weighted moving average method to process the motion trajectory of the human skeleton, and performs a weighted average of the state of each ID in the current frame and the previous frame to output the temporally smoothed trajectory.
[0024] The temporal convolutional unit constructs a temporal convolutional network, using a one-dimensional convolutional kernel that slides along the time dimension to capture the dynamic pattern of the action after temporally smoothing the trajectory for each ID. The formula used is as follows: ;
[0025] In the formula, This represents the length of the one-dimensional convolution kernel. Represents the kernel index. Indicates network layer index, Indicates the first Layered networks in the first layer Frame output, Indicates the first Layered networks in convolutional kernels The weight at the offset, Indicates the first Layered networks in The characteristic input of the frame, Indicates the bias term. Represents a non-linear activation function;
[0026] The spatiotemporal graph convolutional unit models the human skeleton as a graph structure, constructs a spatiotemporal graph convolutional network, and performs convolution operations simultaneously between human key points within the same frame and between adjacent frames. This learns the coordinated motion patterns between the limb joints of the human skeleton in space and time, outputting a high-level feature sequence. The formula used is as follows: ; ;
[0027] In the formula, For graph structure, Represents a set of key points on the human body. Represents a set of connected bones. Represents the neighbor information aggregation matrix. and They represent the first Layer and first The node features output by the layer. Indicates the first The trainable weight matrix of the layer;
[0028] The action output unit constructs a global pooling layer and a fully connected classification layer. The global pooling layer compresses the high-level feature sequence into a single feature vector, representing the action content of the entire human image frame stream. The fully connected classification layer outputs action category labels, and the action category labels are converted into confidence scores through the softmax function.
[0029] Furthermore, the cross-view fusion module performs geometric transformation correction on the motion trajectories of the human skeleton from different viewpoints, projects the two-dimensional human key points onto a unified coordinate system using camera calibration parameters, outputs a set of human key points in the unified coordinate system, and weights the action category labels from different viewpoints according to confidence level to output the final body posture category. The formula used is as follows: ; ;
[0030] In the formula, This represents the scaling factor in homogeneous coordinates. , The coordinates of key points on the human body are in two dimensions. , , To unify the three-dimensional coordinates of the coordinate system, Represents the rotation matrix. Represents the translation vector. This represents the extrinsic parameter matrix of the camera. This represents the intrinsic parameter matrix of the camera. Indicates camera calibration parameters, Indicates the view index. Indicates the total number of viewpoints. Indicates the confidence weight. Indicates the first Action category tags from different perspectives This indicates the final body posture category.
[0031] The beneficial effects achieved by the present invention using the above solution are as follows:
[0032] (1) Traditional key point detection methods usually downsample the image multiple times, which drastically reduces the spatial resolution of the image, leading to irreversible loss of feature information and limited accuracy of key point localization. This solution adopts parallel detection and information fusion, which utilizes the precise location of high-resolution features and the rich semantics of low-resolution features to significantly improve the detection accuracy of human key points, enhance the network's generalization ability to scale changes, occlusion and abnormal poses, and provide higher accuracy key point coordinates and confidence through probability heatmaps, which facilitates further operations by subsequent modules.
[0033] (2) Traditional target tracking methods rely on appearance features, are easily affected by environmental interference, and have simple matching strategies, frequent identity switching, and pose trajectory breakage problems. This solution tracks the target based on the topology and key point positions of the human skeleton, and has strong robustness to changes in lighting and clothing. By measuring the skeleton similarity of different frames, the accuracy and reliability of matching are effectively improved, and more stable and continuous human skeleton motion trajectories are generated.
[0034] (3) In view of the problem that traditional pose recognition methods are mostly based on local fragment feature classification, which easily loses the global action context and leads to poor action recognition effect during the motion process, this scheme combines temporal convolution and spatiotemporal graph convolution to jointly model the features of spatial and temporal domains, obtain complete action representation, output action category and its confidence, ensure the overall discrimination ability of global actions, and effectively improve the stability, accuracy and adaptability of body pose recognition. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of an automatic body posture recognition system based on video analysis proposed in this invention.
[0036] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation
[0037] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0038] Example 1, see Figure 1 The present invention provides an automatic body posture recognition system based on video analysis, including a video acquisition module, a preprocessing module, a key point detection module, a skeleton tracking module, an action recognition module, and a cross-view fusion module;
[0039] The video acquisition module acquires multiple video streams containing the athlete's body posture from different perspectives, performs standardization processing on the multiple video streams, and outputs the original frame stream.
[0040] The preprocessing module optimizes the original frame stream, extracts the athlete's human body region, and outputs a human body image frame stream.
[0041] The key point detection module performs feature extraction and human key point detection on the human image frame stream based on a convolutional neural network, and outputs a two-dimensional key point set containing the coordinates of human key points.
[0042] The skeleton tracking module detects the human skeleton in each frame of the human image frame stream based on a two-dimensional key point set, calculates the skeleton similarity, assigns a unique ID to each human skeleton, and outputs the motion trajectory of the human skeleton.
[0043] The action recognition module smooths the motion trajectory of the human skeleton, identifies the action category of the entire human image frame stream, and outputs the action category label and its confidence level.
[0044] The cross-view fusion module associates human skeletons with the same ID from different viewpoints, performs weighted fusion of the action category labels of the same athlete, and outputs the final body posture category.
[0045] Example 2, see Figure 1 This embodiment is based on the above embodiment. The video acquisition module is equipped with front, side and top view cameras to acquire multiple video streams, capture the athlete's body posture from different perspectives, decode the multiple video streams and align them with timestamps, and convert them into RGB format to output the original frame stream.
[0046] Example 3, see Figure 1 This embodiment is based on the above embodiment, and the preprocessing module includes a noise reduction unit, a brightness adjustment unit, a motion detection unit, and a resolution adjustment unit;
[0047] The noise reduction unit removes image noise from the original frame stream using Gaussian filtering.
[0048] The brightness adjustment unit adaptively adjusts the brightness and contrast of the original frame stream to maintain visual consistency under different lighting conditions.
[0049] The motion detection unit uses the frame difference method on the original frame stream to separate the athlete's body area from the complex background and use it as the ROI (region of interest).
[0050] The resolution adjustment unit trims the original frame stream, uniformly trimming the ROI region to 256×256, and finally outputs the trimmed human image frame stream.
[0051] Example 4, see Figure 1 This embodiment is based on the above embodiment, and the key point detection module includes a feature extraction unit, a feature fusion unit, a heatmap prediction unit, and a coordinate regression unit;
[0052] The feature extraction unit constructs deep convolutional layers based on the ResNet structure, and extracts multi-level features from the human image frame stream through layer-by-layer convolution and pooling operations. The formula used is as follows: ;
[0053] In the formula, This represents the input human body image frame stream. Represents network parameters, This indicates that convolution and pooling operations are performed on the human body image frame stream. For the multi-level features of the output;
[0054] The feature fusion unit adopts the HRNet structure to perform parallel detection and information fusion of multi-level features. It fuses contextual semantics with high-resolution features of small-scale joints in the human image frame stream, and preserves spatial information for low-resolution features of large-scale motion, thereby obtaining fused multi-scale features and outputting a fused feature map. The formula used is as follows: ;
[0055] In the formula, Indicates the resolution branch, The total number of branches representing the resolution of multi-level features. Represents resolution The following features This represents the multi-scale features after fusion. The fusion weights represent the features at different resolutions;
[0056] The heatmap prediction unit sets the number of human key points to be detected to N, uses a 1×1 convolutional layer to output a feature map with N channels from the fused feature map, each channel corresponding to a human key point, and uses the sigmoid function to normalize the feature map of each channel into a probability heatmap, where the value of each pixel on the probability heatmap represents the probability that the corresponding human key point is located at that position.
[0057] The coordinate regression unit converts the probability heatmap into human keypoint coordinates and confidence scores: peak detection is performed on the probability heatmap, and the obtained peak coordinates are the human keypoint coordinates. The peak value is used as the confidence score of the human keypoint, and the output is a set containing N human keypoint coordinates, which is the two-dimensional keypoint set describing human posture, as shown below: ;
[0058] In the formula, A two-dimensional set of key points representing human posture. , Indicates coordinate position, Indicates the confidence level of key points.
[0059] By performing the aforementioned operations, this solution addresses the problem that traditional keypoint detection methods typically downsample images multiple times, drastically reducing the spatial resolution of the image and leading to irreversible loss of feature information and limited keypoint localization accuracy. Instead, it employs parallel detection and information fusion, leveraging the precise location of high-resolution features and the rich semantics of low-resolution features to significantly improve the detection accuracy of human keypoints. This enhances the network's generalization ability to scale changes, occlusion, and abnormal poses, and provides higher-precision keypoint coordinates and confidence levels through probabilistic heatmaps, facilitating further operations by subsequent modules.
[0060] Example 5, see Figure 1 This embodiment is based on the above embodiment, and the skeleton tracking module includes a topology construction unit, a trajectory prediction unit, a skeleton matching unit, and a target tracking unit;
[0061] The topology building unit acquires the MPII human anatomy topology standard, defines the skeleton connection set according to the MPII human anatomy topology standard, and constructs a structured human skeleton by spatial matching the two-dimensional keypoint set, in the following form: ;
[0062] In the formula, Represents a frame. Represents the human skeleton. Represents the join function. and These represent the beginning and end points of the bone, respectively. and Represents the coordinates of the skeleton. Represents a skeleton connection set;
[0063] The trajectory prediction unit detects the human skeleton appearing in each frame of the human image frame stream, models the motion trajectory of the human skeleton using Kalman filtering, and predicts the position of the human skeleton in the current frame. The formula used is as follows: ;
[0064] In the formula, Represents a frame. This indicates the predicted location of the tracked target. This indicates the state of the target in the previous frame. This represents the state transition matrix that describes the motion law of the tracked target. Represents the control matrix. Indicates the input control quantity. Indicates noise;
[0065] The skeleton matching unit matches the detection result of the current frame with the human skeleton of the previous frame, and calculates the skeleton similarity using weighted Euclidean distance. The formula used is as follows: ;
[0066] In the formula, This indicates the detection result for the current frame. This represents the human skeleton in the previous frame. Indicates skeletal similarity. Indicates the key point index. Indicates the number of key points on the human body. Weights for different key points on the human body and Represents the coordinates of the detection result in the current frame. and Represents the coordinates of the human skeleton in the previous frame;
[0067] The target tracking unit assigns an ID to the human skeleton detected in each frame based on the calculated skeleton similarity. When a new target appears, a new ID is created. When the human skeleton leaves the field of view, it is marked as disappeared, and the motion trajectory of the human skeleton is output.
[0068] By performing the aforementioned operations, this solution addresses the problems of traditional target tracking methods that rely on appearance features, are susceptible to environmental interference, have simple matching strategies, and suffer from frequent identity switching and pose trajectory breakage. Instead, it tracks targets based on the topological structure and key point positions of the human skeleton, exhibiting strong robustness to changes in lighting and clothing. By measuring the skeleton similarity of different frames, it effectively improves the accuracy and reliability of matching, generating a more stable and continuous human skeleton motion trajectory.
[0069] Example 6, see Figure 1 This embodiment is based on the above embodiment, and the action recognition module includes a temporal smoothing unit, a temporal convolution unit, a spatiotemporal graph convolution unit, and an action output unit;
[0070] The temporal smoothing unit uses an exponentially weighted moving average method to process the motion trajectory of the human skeleton, and performs a weighted average of the state of each ID in the current frame and the previous frame to output the temporally smoothed trajectory.
[0071] The temporal convolutional unit constructs a temporal convolutional network, using a one-dimensional convolutional kernel that slides along the time dimension to capture the dynamic pattern of the action after temporally smoothing the trajectory for each ID. The formula used is as follows: ;
[0072] In the formula, This represents the length of the one-dimensional convolution kernel. Represents the kernel index. Indicates network layer index, Indicates the first Layered networks in the first layer Frame output, Indicates the first Layered networks in convolutional kernels The weight at the offset, Indicates the first Layered networks in The characteristic input of the frame, Indicates the bias term. Represents a non-linear activation function;
[0073] The spatiotemporal graph convolutional unit models the human skeleton as a graph structure, constructs a spatiotemporal graph convolutional network, and performs convolution operations simultaneously between human key points within the same frame and between adjacent frames. This learns the coordinated motion patterns between the limb joints of the human skeleton in space and time, outputting a high-level feature sequence. The formula used is as follows: ; ;
[0074] In the formula, For graph structure, Represents a set of key points on the human body. Represents a set of connected bones. Represents the neighbor information aggregation matrix. and They represent the first Layer and first The node features output by the layer. Indicates the first The trainable weight matrix of the layer;
[0075] The action output unit constructs a global pooling layer and a fully connected classification layer. The global pooling layer compresses the high-level feature sequence into a single feature vector, representing the action content of the entire human image frame stream. The fully connected classification layer outputs action category labels, and the action category labels are converted into confidence scores through the softmax function.
[0076] By performing the aforementioned operations, this solution addresses the problem that traditional pose recognition methods often rely on local segment feature classification, which easily loses the global action context and leads to poor action recognition performance during motion. Instead, it combines temporal convolution and spatiotemporal graph convolution to jointly model features in both spatial and temporal domains, obtaining a complete action representation and outputting the action category and its confidence level. This ensures the overall ability to discriminate global actions and effectively improves the stability, accuracy, and adaptability of body pose recognition.
[0077] Example 7, see Figure 1 This embodiment is based on the above embodiment. The cross-view fusion module performs geometric transformation correction on the motion trajectory of the human skeleton from different viewpoints, projects the two-dimensional human key points to a unified coordinate system through camera calibration parameters, outputs a set of human key points in the unified coordinate system, and weights the action category labels from different viewpoints according to confidence level to output the final body posture category. The formula used is as follows: ; ;
[0078] In the formula, This represents the scaling factor in homogeneous coordinates. , The coordinates of key points on the human body are in two dimensions. , , To unify the three-dimensional coordinates of the coordinate system, Represents the rotation matrix. Represents the translation vector. This represents the extrinsic parameter matrix of the camera. This represents the intrinsic parameter matrix of the camera. Indicates camera calibration parameters, Indicates the view index. Indicates the total number of viewpoints. Indicates the confidence weight. Indicates the first Action category tags from different perspectives This indicates the final body posture category.
[0079] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0080] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
[0081] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.
Claims
1. An automatic body posture recognition system based on video analysis, characterized in that: It includes a video acquisition module, a preprocessing module, a key point detection module, a skeleton tracking module, an action recognition module, and a cross-view fusion module; The video acquisition module acquires multiple video streams containing the athlete's body posture from different perspectives, performs standardization processing on the multiple video streams, and outputs the original frame stream. The preprocessing module optimizes the original frame stream, extracts the athlete's human body region, and outputs a human body image frame stream. The key point detection module performs feature extraction and human key point detection on the human image frame stream based on a convolutional neural network, and outputs a two-dimensional key point set containing the coordinates of human key points. The skeleton tracking module detects the human skeleton in each frame of the human image frame stream based on a two-dimensional key point set, calculates the skeleton similarity, assigns a unique ID to each human skeleton, and outputs the motion trajectory of the human skeleton. The action recognition module smooths the motion trajectory of the human skeleton, identifies the action category of the entire human image frame stream, and outputs the action category label and its confidence level. The cross-view fusion module associates human skeletons with the same ID from different viewpoints, performs weighted fusion of the action category labels of the same athlete, and outputs the final body posture category.
2. The automatic body posture recognition system based on video analysis according to claim 1, characterized in that: The key point detection module includes a feature extraction unit, a feature fusion unit, a heatmap prediction unit, and a coordinate regression unit; The feature extraction unit constructs a deep convolutional layer and extracts multi-level features of the human image frame stream through layer-by-layer convolution and pooling operations. The feature fusion unit performs parallel detection and information fusion of multi-level features, fuses contextual semantics with high-resolution features of small-scale joints in human image frame streams, preserves spatial information for low-resolution features of large-scale motion, obtains fused multi-scale features, and outputs a fused feature map. The heatmap prediction unit sets the number of human key points to be detected, uses a 1×1 convolutional layer to output a multi-channel feature map from the fused feature map, with each channel corresponding to a human key point, and uses the sigmoid function to normalize the feature map of each channel into a probability heatmap, where the value of each pixel on the probability heatmap represents the probability that the corresponding human key point is located at that position. The coordinate regression unit converts the probability heatmap into human keypoint coordinates and confidence scores: it performs peak detection on the probability heatmap, obtains the peak coordinates as human keypoint coordinates, and uses the peak as the confidence score of the human keypoint, outputting a set containing human keypoint coordinates, which is a two-dimensional keypoint set describing human posture.
3. The automatic body posture recognition system based on video analysis according to claim 1, characterized in that: The skeleton tracking module includes a topology construction unit, a trajectory prediction unit, a skeleton matching unit, and a target tracking unit; The topology building unit defines a skeleton connection set, and constructs a structured human skeleton by spatial matching the set of two-dimensional key points; The trajectory prediction unit detects the human skeleton appearing in each frame of the human image frame stream, models the motion trajectory of the human skeleton through Kalman filtering, and predicts the position of the human skeleton in the current frame. The skeleton matching unit matches the detection result of the current frame with the human skeleton of the previous frame and calculates the skeleton similarity. The target tracking unit assigns an ID to the human skeleton detected in each frame based on the calculated skeleton similarity. When a new target appears, a new ID is created. When the human skeleton leaves the field of view, it is marked as disappeared, and the motion trajectory of the human skeleton is output.
4. The automatic body posture recognition system based on video analysis according to claim 1, characterized in that: The action recognition module includes a temporal smoothing unit, a temporal convolution unit, a spatiotemporal graph convolution unit, and an action output unit; The temporal smoothing unit uses an exponentially weighted moving average method to process the motion trajectory of the human skeleton, and performs a weighted average of the state of each ID in the current frame and the previous frame to output the temporally smoothed trajectory. The temporal convolutional unit constructs a temporal convolutional network, which uses a one-dimensional convolutional kernel to slide in the time dimension to capture the dynamic pattern of the action after the temporal smoothing of the trajectory of each ID. The spatiotemporal graph convolutional unit models the human skeleton as a graph structure, constructs a spatiotemporal graph convolutional network, and performs convolution operations simultaneously between human key points in the same frame and between adjacent frames to learn the coordinated motion laws between the limb joints of the human skeleton in space and time, and outputs high-level feature sequences. The action output unit constructs a global pooling layer and a fully connected classification layer. The global pooling layer compresses the high-level feature sequence into a single feature vector, representing the action content of the entire human image frame stream. The fully connected classification layer outputs action category labels, and the action category labels are converted into confidence scores through the softmax function.
5. The automatic body posture recognition system based on video analysis according to claim 1, characterized in that: The cross-view fusion module performs geometric transformation correction on the motion trajectory of the human skeleton from different viewpoints, projects the two-dimensional human key points to a unified coordinate system through camera calibration parameters, outputs a set of human key points in the unified coordinate system, and weights the action category labels from different viewpoints according to confidence level to output the final body posture category.
Citation Information
Patent Citations
Skeleton human body behavior recognition method based on comprehensive space-time representation
CN118038539A
Deep learning-based basketball player posture detection method and system
CN120340121A
Apparatus and method for tracking object by using skeleton analysis
WO2022080844A1
Cited By
AI-based motion posture recognition system
CN121564446A
3D human body posture estimation and multi-key-point time sequence analysis method
CN122024331A