Video stream-based attitude feature recognition method

By preprocessing and extracting features from the video stream and combining the temporal continuity of local adjacent frames and key frames, an enhanced feature representation is generated. This solves the problem of insufficient accuracy in posture estimation and action recognition in existing technologies and achieves high-precision posture estimation and action recognition in complex dynamic environments.

CN120635989AActive Publication Date: 2025-09-12北京汇畅数宇科技发展有限公司

Patent Information

Application Number
CN202510810234.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-12
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

In the existing technology, the posture estimation method of video stream faces problems such as rapid motion and occlusion in dynamic environments, which leads to the degradation of visual information. In addition, the human action recognition model outputs a multi-modal distribution when predicting a single input motion sequence, resulting in classification ambiguity and insufficient accuracy.

Method used

The video stream is preprocessed to extract keyframes and adjacent frames. An object detection algorithm is used to extract human body regions, and a feature extraction module is constructed to obtain global and local frames. Semantic association information is constructed by leveraging the temporal continuity between local adjacent frames and the current keyframe. A conditional feature aggregation algorithm is then used to fuse visual context and semantic association information to generate an enhanced feature representation. A deep learning algorithm is then used to selectively enhance human features, obtaining detailed pose features. These features are then integrated into keyframes based on temporal features to generate pose sequence data.

Benefits of technology

It improves the accuracy and robustness of posture estimation in complex dynamic environments, enhances the model's ability to understand complex motion patterns, and ensures the temporal consistency and natural fluency of posture estimation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635989A_ABST
    Figure CN120635989A_ABST
Patent Text Reader

Abstract

The invention discloses a posture feature recognition method based on a video stream, and the method comprises the steps: carrying out the preprocessing of a continuous video stream, obtaining video frame training data, extracting a key frame and an adjacent frame in each frame of image, constructing a feature extraction module for a human body region, and obtaining a global frame, performing local extraction on the human body area by using adjacent frames on the left side and the right side to obtain local frames, and constructing semantic association information for the global frame through time sequence continuity between the local adjacent frames and the current key frame; acquiring enhanced feature representation by adopting a conditional feature aggregation algorithm; obtaining attitude sequence data through the attitude detail features; the method comprises the following steps: establishing three-dimensional coordinates, adaptively extracting posture change data by adopting a human body motion decoupling model, predicting human body posture characteristics through a smooth optimization strategy, and introducing a cross attention mechanism to realize deep fusion of spatio-temporal characteristics, so that the understanding ability of the model to a complex action mode is enhanced; and the attitude expression capability of the model in a sheltered or fuzzy region is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video gesture feature recognition, and in particular to a gesture feature recognition method based on video stream. Background Art

[0002] Human pose estimation uses computer vision and deep learning technologies to accurately detect and identify the main key points of the human body, such as the position of the head, neck, hips, etc. from images or videos, and then infer the human posture information. Human pose estimation is a basic technology in understanding human motion, and is also the foundation for action recognition, motion prediction, and behavior understanding.

[0003] Human posture estimation technology has a wide range of applications, covering motion analysis, medical rehabilitation, behavior recognition, etc. In sports, dance, fitness and other fields, human posture estimation algorithms can help analyze and evaluate athletes' movement skills and posture correctness by detecting joint positions, posture angles, etc. in real time, providing athletes with more scientific and personalized training and guidance, thereby promoting sports training and improving competitive level; in the field of intelligent security, through posture estimation technology, the security system can analyze human posture in the security area, detect whether anyone enters the restricted area or performs abnormal activities, and once the security system detects abnormal behavior or area intrusion, it can immediately trigger a security alarm and send early warning information to relevant personnel or emergency agencies, and implement corresponding safety measures to ensure regional safety and prevent potential safety hazards.

[0004] Compared to posture estimation, human action recognition focuses more on the sequence of human motion over a period of time. Based on posture estimation, action recognition detects and tracks the positions of human joints in a video. It extracts and identifies specific action patterns from dynamic posture sequences and determines the type of action the sequence belongs to. While action recognition has direct real-world applications, current posture feature recognition methods based on video streams still have the following problems: (1) In human pose estimation, rapid motion and pose occlusion in dynamic video environments cause the visual information of video frames to degrade, making pose estimation on blurred frames, fast motion frames, and occluded frames extremely difficult. However, existing human pose estimation research does not make sufficient use of global semantic information and local visual environment information, and cannot effectively solve the problem of visual degradation of video frames. (2) In human motion recognition, when the recognition model predicts the action category of a single input motion sequence, the output is often multimodal, resulting in high probability values ​​in multiple action categories rather than concentrated in one category. This situation causes classification ambiguity and makes the model's accurate classification ability insufficient. Summary of the Invention

[0005] The purpose of the present invention is to provide a posture feature recognition method based on video stream to solve the technical problems in the prior art of insufficient utilization of the global semantic information and local visual environment information of the video stream, inability to effectively solve the problem of visual degradation of video frames, and insufficient accurate classification ability of the human motion recognition model.

[0006] In order to solve the above technical problems, the present invention specifically provides the following technical solutions: The present invention provides a method for identifying gesture features based on video stream, comprising the following steps: Preprocessing a continuous video stream to obtain video frame training data, extracting key frames and adjacent frames in each frame of the video frame training data, and using a target detection algorithm to extract human body regions in the video stream; Constructing a feature extraction module for the human body region to obtain a global frame, performing local extraction on the human body region using adjacent frames on both sides to obtain a local frame, and constructing semantic association information for the global frame based on the temporal continuity between the local adjacent frames and the current key frame; Using a conditional feature aggregation algorithm to fuse the visual context information of the local frame with the semantic association information to obtain an enhanced feature representation; Selectively enhancing human body features using a deep learning algorithm on the enhanced feature representation to obtain posture detail features, integrating the posture detail features into the key frames based on time series features to obtain posture sequence data; Three-dimensional coordinates are established for the posture sequence data according to the joint nodes of the human body, a human motion decoupling model is used to adaptively extract posture change data, and human posture characteristics are predicted through a smoothing optimization strategy.

[0007] As a preferred solution of the present invention, a continuous video stream is preprocessed to obtain video frame training data, key frames and adjacent frames in each frame image of the video frame training data are extracted, and a target detection algorithm is used to extract human body regions in the video stream, including: The continuous video stream is sampled at a set frame rate, divided into a video frame sequence arranged in time order, and each frame of the video frame sequence is encoded and normalized to obtain video frame training data; A key frame is selected from the video frame sequence using a temporal averaging strategy, and several adjacent frames are selected within t time windows before and after the key frame to construct a frame set centered on the key frame and containing local temporal context information; Performing human body detection on each frame of the video frame sequence using the Faster R-CNN detection model to identify human targets in the video frame sequence images and obtain corresponding human body bounding boxes; Identifying a human body region image within the human body boundary frame, performing human body image cropping and normalization processing on the human body region image, and obtaining a human body region; The human body region images of the key frame and its adjacent frames are combined into a multi-frame input sample, and the annotation information of the key points corresponding to the time series is used to construct a video frame training data set.

[0008] As a preferred solution of the present invention, a feature extraction module is constructed for the human body region to obtain a global frame, and local extraction of the human body region is performed using adjacent frames on both sides to obtain a local frame, including: A human pose estimation algorithm is used to extract a global frame of person i in a human region from the video frame training dataset, a high-level semantic feature map is extracted using a convolutional neural network, a global feature representation with rich semantic information is obtained, and the semantic feature map is used as a global frame feature; Obtain key frames in the global frame according to the motion characteristics of the character , select the adjacent frames within t time windows before and after the key frame, and extract the corresponding human body area images respectively; The local visual features of each frame are extracted through a convolutional network with shared weights, and the features of adjacent frames are combined to form a local frame feature set to model the local motion changes and visual context information between the frames before and after the key frame.

[0009] As a preferred solution of the present invention, semantic association information is constructed for the global frame through temporal continuity between local adjacent frames and the current key frame, including: Selecting local adjacent frames from the local frame feature set in a time series to perform temporal alignment with the current key frame, and using a variable convolution model to align the features of multiple frames within a time window of t to obtain temporal continuity between the local adjacent frames and the current key frame; An attention mechanism is used to fuse the local motion changes and visual context information between the local adjacent frames and the key frame to obtain the semantic similarity and motion continuity of the time sequence before and after the key frame; The local frame features are weightedly aggregated according to the semantic similarity of the key frames to generate enhanced context-aware features and obtain enhanced semantic association information.

[0010] As a preferred solution of the present invention, a conditional feature aggregation algorithm is used to fuse the visual context information of the local frame with the semantic association information, including: Associating the visual context information of the local frame with the semantic association information by using the temporal distance between the local adjacent frame and the key frame as a correlation metric; The global features and key frame features are input into the shared embedding feature layer of the variability convolutional model, the original feature space is converted into an embedding representation, and the embedding information representation of each frame is obtained. 、 ; The embedded information is represented as 、 Obtain the semantic correlation between global features and key frame features through matrix dot product , and its correlation calculation expression is:

[0011] in, Function represents the feature embedding operation, represents the matrix product operation, represents the global features of person i extracted from the global frame, Represents the key frame features of person i.

[0012] As a preferred solution of the present invention, according to the semantic relevance The feature aggregation model is used to perform conditional feature aggregation on key frame features, local features, and global features to obtain enhanced feature representation, including: Input the key frame features, local features and global features into a convolutional neural network for encoding processing to obtain a feature vector of unified measurement, input the key frame features into a feature encoder for encoding, and generate weight matrices for different regions of the human body; Using the semantic correlation based on time series The local features and global features are weighted respectively, and are connected with the key frame features according to the time series to obtain the local feature information adjacent in time; The weight matrix is ​​processed by a Sigmoid function to obtain an enhancement point at each pixel position in the key frame feature, and the local features, global features and key frame features of each frame image are weightedly fused according to a time series to obtain fused feature information; The fused feature information is aggregated and transformed with the original key frame features to obtain enhanced feature vectors of different regions of the human body and generate enhanced feature representations.

[0013] As a preferred solution of the present invention, the enhanced feature representation is selectively enhanced using a deep learning algorithm to obtain posture detail features, including: Extracting visual context information represented by the enhanced features, generating a visual feature sequence using a convolutional neural network, introducing learnable category labels, and calculating spatial similarity within each video frame by cascading the visual feature sequence; Classify and label the corresponding visual feature sequence according to the spatial similarity, perform matrix multiplication with the visual feature sequence to obtain a human body mask, and perform element-by-element dot multiplication of the human body mask with the corresponding visual feature sequence to obtain coarse-grained features; The key frames are marked by the enhanced feature representation, and are spliced ​​with the coarse-grained features along the time series dimension of the key point marks to form a multi-frame feature sequence; Inputting the multi-frame feature sequence into a deep learning model for feature learning, separating visual features and key point labels frame by frame, and obtaining multi-frame features and multi-frame key point labels; After transposing the multi-frame features, matrix multiplication is performed with the corresponding multi-frame key point labels to generate a human key point confidence map matrix; The softmax function is used to normalize the weights of the elements in the human key point confidence map matrix to generate a key point mask, and the enhanced posture detail features are obtained by performing element-by-element multiplication of the key point mask with the multi-frame features.

[0014] As a preferred solution of the present invention, the posture detail features are integrated into the key frames according to the time sequence features to obtain posture sequence data, including: A self-attention mechanism is used to model the detailed features of the posture of multiple consecutive frames, capture the movement pattern and evolution trend of human joints in the time dimension, and obtain dynamic feature representations with temporal relationships; A P3D-Resnet spatiotemporal feature extraction network is constructed to map the pose detail features extracted from non-keyframes to the keyframe spatial coordinates. The mapped pose detail features are weightedly fused with the original pose features of the keyframe through an attention mechanism to enhance the pose expression capability of occluded or blurred areas in the current frame, introduce dynamic detail cues from adjacent frames, and generate enhanced keyframe pose features. The key frame posture features are arranged in chronological order to construct continuous human posture sequence data.

[0015] As a preferred solution of the present invention, three-dimensional coordinates are established for the posture sequence data according to each joint node of the human body, and the posture change data is adaptively extracted using a human motion decoupling model, including: Extracting key nodes of the human body according to the human body posture sequence data, assigning initial three-dimensional coordinates to each key node of the human body, classifying and identifying the human body posture sequence data using a cross entropy loss function, and optimizing the three-dimensional coordinate position of each joint point by minimizing the cross entropy loss; Constructing a recursive neural network to train the human body posture sequence data and extract different types of motion patterns of human body postures; By introducing the attention mechanism to train the recursive neural network, representative motion features are extracted, dynamic features are captured, posture change data are extracted, and a posture sequence dataset is generated.

[0016] As a preferred solution of the present invention, the posture change data between the previous and next frames in the posture sequence data set are processed by a smoothing optimization strategy to predict the human body posture features, including: Inputting the posture sequence data set into a convolutional neural network as a posture sequence to extract a pattern of posture changes over time; Acquiring spatial feature information S-Transformer and temporal feature information T-Transformer in the convolutional neural network; The posture features containing coordinate and joint information and the fine-grained contour features are convolved and input into the S-Transformer to obtain spatial features at different time scales. The temporal posture features are input into T-Transformer, and then the cross attention mechanism is introduced to fuse the spatiotemporal features; The corresponding posture features are fused with the appearance contour features to identify the posture features of multi-feature spatiotemporal fusion.

[0017] Compared with the prior art, the present invention has the following beneficial effects: The present invention extracts high-level semantic feature maps as global frame features through convolutional neural networks, and uses neighboring frames to capture local motion changes, providing rich visual context information. It performs weighted aggregation of local frame features based on temporal distance and semantic similarity to generate enhanced context-aware features, thereby enhancing the model's ability to understand complex action patterns.

[0018] The self-attention mechanism and neural network are used to capture the movement patterns of human joints in the temporal dimension, ensuring the temporal smoothness and consistency of the posture sequence. The S-Transformer and T-Transformer are used to process the information in the spatial and temporal dimensions respectively, and the cross-attention mechanism is introduced to achieve deep fusion of spatiotemporal features, which improves the model's ability to understand dynamic changes. The cross-entropy loss function is used to optimize the three-dimensional coordinate position of each joint point to ensure that the posture estimation results are both in line with actual physical constraints and as accurate as possible. It can capture different types of movement patterns and enhance the model's ability to understand and predict human motion. The attention mechanism is used to perform weighted fusion of the mapped posture detail features and the original features of the key frames, significantly improving the model's posture expression ability in occluded or blurred areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other implementation drawings based on the provided drawings without inventive effort.

[0020] Figure 1 This is a flow chart of a method for posture feature recognition based on video streams provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0022] like Figure 1 As shown, the present invention provides a posture feature recognition method based on video stream, comprising the following steps: Preprocessing a continuous video stream to obtain video frame training data, extracting key frames and adjacent frames in each frame of the video frame training data, and using a target detection algorithm to extract human body regions in the video stream; Constructing a feature extraction module for the human body region to obtain a global frame, performing local extraction on the human body region using adjacent frames on both sides to obtain a local frame, and constructing semantic association information for the global frame based on the temporal continuity between the local adjacent frames and the current key frame; In this embodiment, by combining the temporal continuity between local adjacent frames and key frames and the semantic association information in the global frame, an enhanced feature representation containing rich visual context and global semantic information can be created, which helps to improve the robustness of the model in complex scenes, such as dealing with occlusion, background clutter and other problems.

[0023] Using a conditional feature aggregation algorithm to fuse the visual context information of the local frame with the semantic association information to obtain an enhanced feature representation; Selectively enhancing human body features using a deep learning algorithm on the enhanced feature representation to obtain posture detail features, integrating the posture detail features into the key frames based on time series features to obtain posture sequence data; In this embodiment, a deep learning algorithm is used to selectively enhance the enhanced feature representation, which can effectively highlight those detailed features that are critical for recognition, even if these features may be ignored or difficult to detect in the original image. This method can significantly improve the accuracy of pose estimation.

[0024] Three-dimensional coordinates are established for the posture sequence data according to the joint nodes of the human body, a human motion decoupling model is used to adaptively extract posture change data, and human posture characteristics are predicted through a smoothing optimization strategy.

[0025] In this embodiment, the human motion decoupling model is used to adaptively extract posture change data, so that the system can better understand and predict complex motion patterns, which is particularly suitable for challenging scenarios such as multi-person interaction and rapid motion changes.

[0026] In this embodiment, by applying a smoothing optimization strategy, a coherent prediction of human posture features can be achieved between previous and subsequent frames, thereby ensuring the temporal consistency and natural fluency of the posture sequence and reducing the posture jump phenomenon caused by single-frame errors.

[0027] In this embodiment, the posture feature recognition method covers the entire process from video preprocessing to final posture prediction, including multiple subtasks such as target detection, feature extraction, posture estimation, three-dimensional coordinate construction, and motion prediction, providing a one-stop solution that is suitable for application in a variety of practical scenarios, such as security monitoring, virtual reality, sports analysis, etc. It not only improves the accuracy and reliability of human posture estimation, but also enhances the adaptability and practicality of the system, providing strong support for solving the problem of understanding human motion in complex dynamic environments.

[0028] Preprocessing a continuous video stream to obtain video frame training data, extracting key frames and adjacent frames from each frame of the video frame training data, and using a target detection algorithm to extract human body regions in the video stream, including: The continuous video stream is sampled at a set frame rate, divided into a video frame sequence arranged in time order, and each frame of the video frame sequence is encoded and normalized to obtain video frame training data; In this embodiment, the continuous video stream is sampled by setting the frame rate, and encoding and standardization processing are performed on each frame, thereby ensuring the consistency and comparability of the data and reducing unnecessary computational burden. This approach helps to improve the efficiency of subsequent processing steps.

[0029] A key frame is selected from the video frame sequence using a temporal averaging strategy, and several adjacent frames are selected within t time windows before and after the key frame to construct a frame set centered on the key frame and containing local temporal context information; In this embodiment, the method of selecting key frames by adopting the temporal averaging strategy and selecting adjacent frames in the time window before and after them can effectively capture the subtle changes and dynamic characteristics in human body movements. It not only provides rich local temporal context information, but also enhances the model's ability to understand complex motion patterns.

[0030] Performing human body detection on each frame of the video frame sequence using the Faster R-CNN detection model to identify human targets in the video frame sequence images and obtain corresponding human body bounding boxes; In this embodiment, the Faster R-CNN detection model can efficiently and accurately locate the human body area in the video frame, which can significantly improve the detection accuracy, especially when dealing with problems such as occlusion and scale changes.

[0031] Identifying a human body region image within the human body boundary frame, performing human body image cropping and normalization processing on the human body region image, and obtaining a human body region; The human body region images of the key frame and its adjacent frames are combined into a multi-frame input sample, and the annotation information of the key points corresponding to the time series is used to construct a video frame training data set.

[0032] In this embodiment, the human body area images of the key frame and its adjacent frames are combined into multi-frame input samples, and the corresponding key point annotation information is combined to construct a training data set. This approach is conducive to training a more robust and more generalizable model, allowing the model to learn the transition details between different postures, thereby providing more accurate posture estimation results in practical applications.

[0033] Constructing a feature extraction module for the human body region to obtain a global frame, and performing local extraction on the human body region using adjacent frames on both sides to obtain a local frame, including: A human pose estimation algorithm is used to extract a global frame of person i in a human region from the video frame training dataset, a high-level semantic feature map is extracted using a convolutional neural network, a global feature representation with rich semantic information is obtained, and the semantic feature map is used as a global frame feature; In this embodiment, high-level semantic feature maps are extracted as global frame features through convolutional neural networks, which can capture the overall structural information and contextual semantics of the human body in the key frame, provide a stable and semantically rich foundation for subsequent posture estimation, and help the model better understand the overall picture of the human body posture in the current frame.

[0034] Obtain key frames in the global frame according to the motion characteristics of the character , select the adjacent frames within t time windows before and after the key frame, and extract the corresponding human body area images respectively; In this embodiment, the information of the adjacent frames on the left and right sides is used to help restore the posture information lost due to occlusion or image quality degradation. At the same time, this joint modeling method of the previous and next frames also improves the temporal consistency of the posture estimation results, avoiding abrupt or unnatural posture jumps.

[0035] The local visual features of each frame are extracted through a convolutional network with shared weights, and the features of adjacent frames are combined to form a local frame feature set to model the local motion changes and visual context information between the frames before and after the key frame.

[0036] In this embodiment, adjacent frames in the time window before and after the key frame are selected, and the local visual features of each frame are extracted through a convolutional network with shared weights to construct a local frame feature set. This can effectively model the local motion changes and visual context information around the key frame. This method enhances the model's ability to perceive dynamic details, especially when dealing with challenges such as occlusion and blur.

[0037] Constructing semantic association information for the global frame through temporal continuity between local neighboring frames and the current key frame, including: Selecting local adjacent frames from the local frame feature set in a time series to perform temporal alignment with the current key frame, and using a variable convolution model to align the features of multiple frames within a time window of t to obtain temporal continuity between the local adjacent frames and the current key frame; In this embodiment, the deformable convolution model is used to align the features of local adjacent frames and key frames within the time window, which can effectively capture the motion changes and spatial displacement relationships between frames. This alignment method improves the robustness of the model in complex dynamic scenes, and is particularly suitable for posture estimation in fast motion or occlusion situations.

[0038] An attention mechanism is used to fuse the local motion changes and visual context information between the local adjacent frames and the key frame to obtain the semantic similarity and motion continuity of the time sequence before and after the key frame; In this embodiment, the attention mechanism is used to fuse the visual context information and local motion changes between local frames and key frames, so that the model can adaptively select the most useful information between different time points. This approach not only enhances the understanding of action details, but also improves the performance of the model in the face of background interference, posture ambiguity and other situations.

[0039] The local frame features are weightedly aggregated according to the semantic similarity of the key frames to generate enhanced context-aware features and obtain enhanced semantic association information.

[0040] In this embodiment, local frame features are weightedly aggregated based on the semantic similarity of key frames, which means that the model can decide which information is more important based on the "degree of correlation" between the current frame and other frames. This mechanism is similar to the attention mechanism of humans when understanding continuous actions, which helps to improve the accuracy and stability of posture estimation.

[0041] A conditional feature aggregation algorithm is used to fuse the visual context information of the local frame with the semantic association information, including: Associating the visual context information of the local frame with the semantic association information by using the temporal distance between the local adjacent frame and the key frame as a correlation metric; In this embodiment, by using the temporal distance between local neighboring frames and key frames as a correlation metric, the dynamic characteristics of human motion over time can be more accurately captured. This method helps improve the model's ability to understand complex motion patterns, which is particularly important when dealing with fast movements or non-rigid deformations.

[0042] The global features and key frame features are input into the shared embedding feature layer of the variability convolutional model, the original feature space is converted into an embedding representation, and the embedding information representation of each frame is obtained. 、 ; The embedded information is represented as 、 Obtain the semantic correlation between global features and key frame features through matrix dot product , and its correlation calculation expression is:

[0043] in, Function represents the feature embedding operation, represents the matrix product operation, represents the global features of person i extracted from the global frame, Represents the key frame features of person i.

[0044] In this embodiment, the shared embedding feature layer of the deformable convolutional model is used to convert the original feature space into a more abstract and high-level embedding representation. This not only effectively integrates the rich information in the global features and keyframe features, but also helps the model learn more discriminative feature expressions, thereby improving the accuracy of pose estimation.

[0045] According to the semantic relevance The feature aggregation model is used to perform conditional feature aggregation on key frame features, local features, and global features to obtain enhanced feature representation, including: Input the key frame features, local features and global features into a convolutional neural network for encoding processing to obtain a feature vector of unified measurement, input the key frame features into a feature encoder for encoding, and generate weight matrices for different regions of the human body; In this embodiment, the three types of features of key frames, local neighboring frames and global frames are uniformly encoded into feature vectors in a unified metric space, breaking the limitations of traditional single-frame modeling and achieving a leap from "static perception" to "dynamic understanding". This multi-source information fusion strategy significantly improves the comprehensiveness and robustness of feature representation.

[0046] Using the semantic correlation based on time series The local features and global features are weighted respectively, and are connected with the key frame features according to the time series to obtain the local feature information adjacent in time; In this embodiment, local features and global features are weighted separately using semantic correlation based on time series, so that the model can automatically identify which adjacent frames’ information is more important for the pose estimation of the current key frame. This approach is similar to the ability of humans to focus on key context frames when observing actions, thereby improving the model’s discriminative ability and adaptability.

[0047] The weight matrix is ​​processed by a Sigmoid function to obtain an enhancement point at each pixel position in the key frame feature, and the local features, global features and key frame features of each frame image are weightedly fused according to a time series to obtain fused feature information; In this embodiment, weight matrices for different regions of the human body are generated based on key frame features, and the resulting data are processed using a Sigmoid function to serve as enhancement points. This achieves adaptive enhancement of key joints or occluded areas in human posture. Compared with the overall enhancement method, this fine-grained spatial attention mechanism can better highlight local areas that are of high value for posture estimation.

[0048] In this embodiment, weighted fusion of local, global and key frame features is performed on each frame image based on the time series, which not only enhances the expressive ability of the current frame, but also ensures the smoothness and consistency of the entire posture sequence in the time dimension, avoiding inter-frame jumping and prediction jitter problems.

[0049] The fused feature information is aggregated and transformed with the original key frame features to obtain enhanced feature vectors of different regions of the human body and generate enhanced feature representations.

[0050] In this embodiment, the fused feature information is aggregated with the original key frame features through transformations such as splicing, addition, and channel attention, and the feature representation is further optimized through learnable parameters, so that the model has stronger nonlinear modeling capabilities and is suitable for diversified posture estimation tasks in complex scenarios.

[0051] The enhanced feature representation uses a deep learning algorithm to selectively enhance human features to obtain posture detail features, including: Extracting visual context information represented by the enhanced features, generating a visual feature sequence using a convolutional neural network, introducing learnable category labels, and calculating spatial similarity within each video frame by cascading the visual feature sequence; In this embodiment, the visual context information in the enhanced feature representation is encoded through a convolutional neural network to generate a visual feature sequence, which enables the model to more accurately understand the spatial structure and motion state of the human body in each frame of the image, providing a high-quality input basis for subsequent key point positioning and feature enhancement.

[0052] Classify and label the corresponding visual feature sequence according to the spatial similarity, perform matrix multiplication with the visual feature sequence to obtain a human body mask, and perform element-by-element dot multiplication of the human body mask with the corresponding visual feature sequence to obtain coarse-grained features; In this embodiment, the introduction of learnable category labels and the combination of cascaded frame-by-frame spatial similarity calculation help the model automatically identify the semantic relevance and motion continuity between different frames. This approach enhances the model's ability to understand cross-frame information, especially when dealing with complex dynamic scenes.

[0053] In this embodiment, visual features are classified and labeled according to spatial similarity, and a human body mask is generated through matrix multiplication, thereby achieving adaptive attention to the human body area in the image. This spatial attention mechanism can effectively suppress background interference and focus on areas related to posture estimation, thereby improving the accuracy of feature extraction.

[0054] The key frames are marked by the enhanced feature representation, and are spliced ​​with the coarse-grained features along the time series dimension of the key point marks to form a multi-frame feature sequence; Inputting the multi-frame feature sequence into a deep learning model for feature learning, separating visual features and key point labels frame by frame, and obtaining multi-frame features and multi-frame key point labels; In this embodiment, the enhanced features of the key frames are spliced ​​with the coarse-grained features along the time dimension to form a multi-frame feature sequence. This not only retains the detailed information of the current frame, but also integrates the context clues of the previous and next frames, enhances the model's ability to understand the motion evolution process, and improves the temporal consistency of posture estimation.

[0055] In this embodiment, a multi-frame feature sequence is input into a deep learning model for feature learning, and visual features and key point markers are separated frame by frame, so that the model can independently and coherently learn subtle changes in posture in each frame, greatly improving the ability to capture posture details.

[0056] After transposing the multi-frame features, matrix multiplication is performed with the corresponding multi-frame key point labels to generate a human key point confidence map matrix; The softmax function is used to normalize the weights of the elements in the human key point confidence map matrix to generate a key point mask, and the enhanced posture detail features are obtained by performing element-by-element multiplication of the key point mask with the multi-frame features.

[0057] In this embodiment, the softmax function is used to normalize the weights of the confidence map matrix to generate a key point mask, which is then applied to multi-frame features through element-by-element dot multiplication to further enhance the feature information related to the key points while suppressing irrelevant noise, thereby achieving selective enhancement of posture details.

[0058] Integrating the posture detail features into the key frame according to the time sequence features to obtain posture sequence data, including: A self-attention mechanism is used to model the detailed features of the posture of multiple consecutive frames, capture the movement pattern and evolution trend of human joints in the time dimension, and obtain dynamic feature representations with temporal relationships; In this embodiment, the posture detail features of multiple consecutive frames are modeled through the self-attention mechanism, which can effectively capture the movement patterns and evolution trends of human joints in the time dimension. This global perspective time modeling enhances the model's overall understanding of human movements, and is particularly suitable for the recognition and prediction of complex and continuous movements.

[0059] A P3D-Resnet spatiotemporal feature extraction network is constructed to map the pose detail features extracted from non-keyframes to the keyframe spatial coordinates. The mapped pose detail features are weightedly fused with the original pose features of the keyframe through an attention mechanism to enhance the pose expression capability of occluded or blurred areas in the current frame, introduce dynamic detail cues from adjacent frames, and generate enhanced keyframe pose features. In this embodiment, a P3D-ResNet spatiotemporal feature extraction network is constructed to map the detailed posture features in non-key frames to the spatial coordinate system of key frames. This solves the problem of spatial misalignment between different frames caused by perspective, displacement or deformation. This step ensures the effective fusion of information from adjacent frames and improves the accuracy of posture estimation.

[0060] The key frame posture features are arranged in chronological order to construct continuous human posture sequence data.

[0061] In this embodiment, the mapped posture detail features and the original posture features of the key frames are weightedly fused through the attention mechanism, so that the model can automatically identify and supplement the information of occluded or blurred areas. This "outside-in" enhancement method greatly improves the robustness and generalization ability of the model in complex scenarios. It pays attention to both the spatial structure within a single frame and the temporal evolution between frames, and realizes multi-scale modeling from local details to global motion patterns. It is suitable for various tasks such as static posture recognition, dynamic motion classification, and long-term motion prediction.

[0062] In this embodiment, dynamic detail clues from the previous and next adjacent frames are introduced into the key frame, which not only makes up for the limitations of single-frame information, but also provides rich local temporal context support for pose estimation. This approach helps to recover key pose information lost due to rapid motion or image quality degradation.

[0063] Establishing three-dimensional coordinates for the posture sequence data according to each joint node of the human body, and adaptively extracting posture change data using a human motion decoupling model, including: Extracting key nodes of the human body according to the human body posture sequence data, assigning initial three-dimensional coordinates to each key node of the human body, classifying and identifying the human body posture sequence data using a cross entropy loss function, and optimizing the three-dimensional coordinate position of each joint point by minimizing the cross entropy loss; In this embodiment, the posture sequence data is classified and identified through the cross entropy loss function, and the loss is minimized to optimize the three-dimensional coordinate position of each joint point, which can ensure that the generated posture conforms to the actual physical constraints and is as accurate as possible, and can effectively improve the accuracy of posture estimation, especially in complex scenes such as occlusion, perspective changes, etc., which helps to obtain a more natural and accurate representation of the human body posture.

[0064] Constructing a recursive neural network to train the human body posture sequence data and extract different types of motion patterns of human body postures; In this embodiment, a recursive neural network is constructed to train human posture sequence data, so that the model can capture different types of motion patterns of human posture that change over time. This method is particularly suitable for modeling long-term and short-term dependencies and is very effective for understanding continuous action sequences.

[0065] By introducing the attention mechanism to train the recursive neural network, representative motion features are extracted, dynamic features are captured, posture change data are extracted, and a posture sequence dataset is generated.

[0066] In this embodiment, the introduction of the attention mechanism during the training process allows the model to automatically focus on the most representative motion features, thereby more effectively capturing dynamic features. This is particularly important for improving the performance of the model in complex movements, and can help the model ignore noise and focus on truly important information.

[0067] The posture change data between the previous and next frames in the posture sequence data set are processed by a smoothing optimization strategy to predict the human body posture features, including: Inputting the posture sequence data set into a convolutional neural network as a posture sequence to extract a pattern of posture changes over time; Acquiring spatial feature information S-Transformer and temporal feature information T-Transformer in the convolutional neural network; In this embodiment, S-Transformer focuses on processing information in the spatial dimension and can capture the relative positional relationship between human joints at different time points, maintaining high accuracy even in complex or rapidly changing movements; T-Transformer is specifically used to process information in the temporal dimension and can capture dynamic change trends in posture sequences, which helps the model understand the development process of the movement and is particularly important for predicting future posture changes.

[0068] The posture features containing coordinate and joint information and the fine-grained contour features are convolved and input into the S-Transformer to obtain spatial features at different time scales. In this embodiment, a cross-attention mechanism is used to establish a connection between the spatial and temporal dimensions to achieve deeper feature fusion. This approach allows the model to dynamically adjust its focus as needed, thereby improving its ability to understand complex actions.

[0069] The temporal posture features are input into T-Transformer, and then the cross attention mechanism is introduced to fuse the spatiotemporal features; The corresponding posture features are fused with the appearance contour features to identify the posture features of multi-feature spatiotemporal fusion.

[0070] In this embodiment, the powerful feature extraction capability of the convolutional neural network (CNN), the advantages of the Transformer in processing long sequence data, and the flexible feature fusion strategy provided by the cross-attention mechanism are comprehensively utilized to ensure that the model can maintain high accuracy and good robustness in the face of various challenges such as occlusion, blur, and fast motion.

[0071] The present invention extracts high-level semantic feature maps as global frame features through convolutional neural networks, and uses neighboring frames to capture local motion changes, providing rich visual context information. It performs weighted aggregation of local frame features based on temporal distance and semantic similarity to generate enhanced context-aware features, thereby enhancing the model's ability to understand complex action patterns.

[0072] The self-attention mechanism and neural network are used to capture the movement patterns of human joints in the temporal dimension, ensuring the temporal smoothness and consistency of the posture sequence. The S-Transformer and T-Transformer are used to process the information in the spatial and temporal dimensions respectively, and the cross-attention mechanism is introduced to achieve deep fusion of spatiotemporal features, which improves the model's ability to understand dynamic changes. The cross-entropy loss function is used to optimize the three-dimensional coordinate position of each joint point to ensure that the posture estimation results are both in line with actual physical constraints and as accurate as possible. It can capture different types of movement patterns and enhance the model's ability to understand and predict human motion. The attention mechanism is used to perform weighted fusion of the mapped posture detail features and the original features of the key frames, significantly improving the model's posture expression ability in occluded or blurred areas.

[0073] The above embodiments are merely exemplary embodiments of the present application and are not intended to limit the scope of the present application. The scope of protection of the present application is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions to the present application within the essence and scope of protection of the present application, and such modifications or equivalent substitutions shall also be deemed to fall within the scope of protection of the present application.

Claims

1. A method for posture feature recognition based on video stream, characterized in that: The following steps are involved: Preprocessing a continuous video stream to obtain video frame training data, extracting key frames and adjacent frames in each frame of the video frame training data, and using a target detection algorithm to extract human body regions in the video stream; Constructing a feature extraction module for the human body region to obtain a global frame, performing local extraction on the human body region using adjacent frames on both sides to obtain a local frame, and constructing semantic association information for the global frame based on the temporal continuity between the local adjacent frames and the current key frame; Using a conditional feature aggregation algorithm to fuse the visual context information of the local frame with the semantic association information to obtain an enhanced feature representation; Selectively enhancing human body features using a deep learning algorithm on the enhanced feature representation to obtain posture detail features, integrating the posture detail features into the key frames based on time series features to obtain posture sequence data; Three-dimensional coordinates are established for the posture sequence data according to the joint nodes of the human body, a human motion decoupling model is used to adaptively extract posture change data, and human posture characteristics are predicted through a smoothing optimization strategy.

2. A method for identifying posture features based on video stream according to claim 1, characterized in that: Preprocessing a continuous video stream to obtain video frame training data, extracting key frames and adjacent frames from each frame of the video frame training data, and using a target detection algorithm to extract human body regions in the video stream, including: The continuous video stream is sampled at a set frame rate, divided into a video frame sequence arranged in time order, and each frame of the video frame sequence is encoded and normalized to obtain video frame training data; A key frame is selected from the video frame sequence using a temporal averaging strategy, and several adjacent frames are selected within t time windows before and after the key frame to construct a frame set centered on the key frame and containing local temporal context information; Performing human body detection on each frame of the video frame sequence using the Faster R-CNN detection model to identify human targets in the video frame sequence images and obtain corresponding human body bounding boxes; Identifying a human body region image within the human body boundary frame, performing human body image cropping and normalization processing on the human body region image, and obtaining a human body region; The human body region images of the key frame and its adjacent frames are combined into a multi-frame input sample, and the annotation information of the key points corresponding to the time series is used to construct a video frame training data set.

3. A method for identifying posture features based on video stream according to claim 2, characterized in that: Constructing a feature extraction module for the human body region to obtain a global frame, and performing local extraction on the human body region using adjacent frames on both sides to obtain a local frame, including: A human pose estimation algorithm is used to extract a global frame of person i in a human region from the video frame training dataset, a high-level semantic feature map is extracted using a convolutional neural network, a global feature representation with rich semantic information is obtained, and the semantic feature map is used as a global frame feature; Obtain key frames in the global frame according to the motion characteristics of the character , select the adjacent frames within t time windows before and after the key frame, and extract the corresponding human body area images respectively; The local visual features of each frame are extracted through a convolutional network with shared weights, and the features of adjacent frames are combined to form a local frame feature set to model the local motion changes and visual context information between the frames before and after the key frame.

4. A method for posture feature recognition based on video stream according to claim 3, characterized in that: Constructing semantic association information for the global frame through temporal continuity between local neighboring frames and the current key frame, including: Selecting local adjacent frames from the local frame feature set in a time series to perform temporal alignment with the current key frame, and using a variable convolution model to align the multi-frame features within a time period window of t to obtain temporal continuity between the local adjacent frames and the current key frame; An attention mechanism is used to fuse the local motion changes and visual context information between the local adjacent frames and the key frame to obtain the semantic similarity and motion continuity of the time sequence before and after the key frame; The local frame features are weightedly aggregated according to the semantic similarity of the key frames to generate enhanced context-aware features and obtain enhanced semantic association information.

5. The method for posture feature recognition based on video stream according to claim 3, characterized in that: A conditional feature aggregation algorithm is used to fuse the visual context information of the local frame with the semantic association information, including: Associating the visual context information of the local frame with the semantic association information using the temporal distance between the local adjacent frame and the key frame as a correlation metric; The global features and key frame features are input into the shared embedding feature layer of the variability convolutional model, the original feature space is converted into an embedding representation, and the embedding information representation of each frame is obtained. 、 ; The embedded information is represented as 、 Obtain the semantic correlation between global features and key frame features through matrix dot product , and its correlation calculation expression is: in, Function represents the feature embedding operation, represents the matrix product operation, represents the global features of person i extracted from the global frame, Represents the key frame features of person i.

6. A method for posture feature recognition based on video stream according to claim 5, characterized in that: According to the semantic relevance The feature aggregation model is used to perform conditional feature aggregation on key frame features, local features, and global features to obtain enhanced feature representation, including: Input the key frame features, local features and global features into a convolutional neural network for encoding processing to obtain a feature vector of unified measurement, input the key frame features into a feature encoder for encoding, and generate weight matrices for different regions of the human body; Using the semantic correlation based on time series The local features and global features are weighted respectively, and are connected with the key frame features according to the time series to obtain the local feature information adjacent in time; The weight matrix is ​​processed by a Sigmoid function to obtain an enhancement point at each pixel position in the key frame feature, and the local features, global features and key frame features of each frame image are weightedly fused according to the time series to obtain fused feature information; The fused feature information is aggregated and transformed with the original key frame features to obtain enhanced feature vectors of different regions of the human body and generate enhanced feature representations.

7. A method for posture feature recognition based on video stream according to claim 6, characterized in that: The enhanced feature representation uses a deep learning algorithm to selectively enhance human features to obtain posture detail features, including: Extracting visual context information represented by the enhanced features, generating a visual feature sequence using a convolutional neural network, introducing learnable category labels, and calculating the spatial similarity within each video frame by cascading the visual feature sequence; Classify and label the corresponding visual feature sequence according to the spatial similarity, perform matrix multiplication with the visual feature sequence to obtain a human body mask, and perform element-by-element dot multiplication of the human body mask with the corresponding visual feature sequence to obtain coarse-grained features; The key frames are marked by the enhanced feature representation, and are spliced ​​with the coarse-grained features along the time series dimension of the key point marks to form a multi-frame feature sequence; Inputting the multi-frame feature sequence into a deep learning model for feature learning, separating visual features and key point labels frame by frame, and obtaining multi-frame features and multi-frame key point labels; After transposing the multi-frame features, matrix multiplication is performed with the corresponding multi-frame key point labels to generate a human key point confidence map matrix; The softmax function is used to normalize the weights of the elements in the human key point confidence map matrix to generate a key point mask, and the enhanced posture detail features are obtained by performing element-by-element multiplication of the key point mask with the multi-frame features.

8. The method for posture feature recognition based on video stream according to claim 7, characterized in that: Integrating the posture detail features into the key frame according to the time sequence features to obtain posture sequence data, including: A self-attention mechanism is used to model the detailed features of the posture of multiple consecutive frames, capture the movement pattern and evolution trend of human joints in the time dimension, and obtain dynamic feature representations with temporal relationships; A P3D-Resnet spatiotemporal feature extraction network is constructed to map the pose detail features extracted from non-keyframes to the keyframe spatial coordinates. The mapped pose detail features are weightedly fused with the original pose features of the keyframe through an attention mechanism to enhance the pose expression capability of occluded or blurred areas in the current frame, introduce dynamic detail cues from adjacent frames, and generate enhanced keyframe pose features. The key frame posture features are arranged in chronological order to construct continuous human posture sequence data.

9. A method for posture feature recognition based on video stream according to claim 8, characterized in that: Establishing three-dimensional coordinates for the posture sequence data according to each joint node of the human body, and adaptively extracting posture change data using a human motion decoupling model, including: Extracting key nodes of the human body according to the human body posture sequence data, assigning initial three-dimensional coordinates to each key node of the human body, classifying and identifying the human body posture sequence data using a cross entropy loss function, and optimizing the three-dimensional coordinate position of each joint point by minimizing the cross entropy loss; Constructing a recursive neural network to train the human body posture sequence data and extract different types of motion patterns of human body postures; By introducing the attention mechanism to train the recursive neural network, representative motion features are extracted, dynamic features are captured, posture change data are extracted, and a posture sequence dataset is generated.

10. The method for posture feature recognition based on video stream according to claim 9, characterized in that: The posture change data between the previous and next frames in the posture sequence data set are processed by a smoothing optimization strategy to predict the human body posture features, including: Inputting the posture sequence data set into a convolutional neural network as a posture sequence to extract a pattern of posture changes over time; Acquiring spatial feature information S-Transformer and temporal feature information T-Transformer in the convolutional neural network; The posture features containing coordinate and joint information and the fine-grained contour features are convolved and input into the S-Transformer to obtain spatial features at different time scales. The temporal posture features are input into T-Transformer, and then the cross attention mechanism is introduced to fuse the spatiotemporal features; The corresponding posture features are fused with the appearance contour features to identify the posture features of multi-feature spatiotemporal fusion.

Citation Information

Patent Citations

  • A video behavior recognition method based on spatio-temporal fusion features and attention mechanism

    CN109101896A

  • Gesture recognition method based on deep neural network and attention mechanism

    CN113378641A

  • Double-flow global-local action recognition method, system and equipment based on video input and storage medium

    CN116311495A

  • Human body posture estimation method and device in motion scene, equipment and storage medium

    CN116386089A

  • Image recognition method and apparatus, computer-readable storage medium, and electronic device

    US20220172518A1

Cited By

  • Eye using and sitting posture monitoring and adjusting method and system based on multi-sensor fusion

    CN121059111A

  • Physical training posture correction method based on machine vision

    CN121281141A

  • Intelligent video customer service access model establishment method based on 5G

    CN121526625A

  • Gesture recognition method and system based on video stream analysis

    CN121564802A

  • Animal three-dimensional attitude estimation method based on feature screening

    CN121616645A