Method and system for two-dimensional pose estimation in an occlusion scenario based on spatio-temporal information

By constructing a two-dimensional pose estimation model of spatiotemporal information fusion, the problem of key points loss in occlusion scenarios is solved, and the accuracy and smoothness of pose estimation are achieved, and high-frequency noise interference and calculation complexity are reduced.

CN119625846BActive Publication Date: 2025-06-03HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510169107.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-03
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

Traditional two-dimensional pose estimation methods are difficult to deal with key point feature loss in occlusion scenarios, resulting in poor accuracy and continuity of estimation results, and poor performance when facing complex dynamic scenarios, making it difficult to reduce high-frequency noise interference and calculation complexity.

Method used

A two-dimensional pose estimation model including input module, visibility prediction and masking module, key point inference module, space-time dependency module and regression layer is constructed. Through the fusion of spatiotemporal information and the completion of key points, the smoothness and accuracy of pose estimation are improved, and high-frequency noise interference is reduced through discrete cosine filtering module.

Benefits of technology

Accurately infer the key points of occlusion in occlusion scenarios, improve the smoothness and accuracy of posture estimation, reduce high-frequency noise interference, ensure model accuracy, and reduce calculation complexity while meeting real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625846B_ABST
    Figure CN119625846B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for two-dimensional pose estimation in occluded scenarios based on spatio-temporal information, which relates to the field of computer vision. The method includes: S1, obtaining a video including multi-person body information; S2, constructing a two-dimensional pose estimation model, where the model includes an input module, a visibility prediction and masking module, a key point inference module, a spatio-temporal dependence module, and a regression layer; S3, inputting the obtained video into the two-dimensional pose estimation model for training to obtain a trained two-dimensional pose estimation model; S4, using the obtained trained two-dimensional pose estimation model to perform two-dimensional pose estimation in occluded scenarios based on spatio-temporal information. The two-dimensional pose estimation model constructed by the present invention can accurately infer occluded key points, improve the smoothness and accuracy of pose estimation, reduce the interference of high-frequency noise, ensure the accuracy of the model while reducing the computational complexity to meet the real-time requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular, to a method and system for two-dimensional pose estimation in occluded scenes based on spatio-temporal information. Background Art

[0002] Traditional two-dimensional pose estimation methods usually have difficulty in dealing with the problem of lost key-point features in occluded scenes, resulting in poor accuracy and continuity of the estimation results, which limits their effectiveness in practical applications. Especially in the fields of real-time monitoring, motion analysis, etc., traditional methods often perform unsatisfactorily when facing challenges such as invisible key points and high-frequency noise interference in complex dynamic scenes. Therefore, how to improve the robustness and accuracy of the model in occluded environments has become a major problem faced by pose estimation technology. With the continuous development of deep learning and computer vision technologies, spatio-temporal information modeling has gradually become an important research direction for solving occlusion problems. The introduction of deep learning provides a new solution idea for the dynamic reasoning of key-point features and the fusion of spatio-temporal information. By combining time-series information with spatial structure relationships, deep learning methods can significantly enhance the model's adaptability to occluded scenes, thereby improving the accuracy of pose estimation. However, existing technologies still face challenges in constructing an effective spatio-temporal feature fusion framework. Especially in occluded scenes, how to accurately reason about occluded key points and improve the smoothness and accuracy of pose estimation remains an urgent problem to be solved. In addition, existing methods have limited ability to handle high-frequency noise interference, resulting in poor smoothness of pose sequences, further affecting the application value of the estimation results. At the same time, how to reduce the computational complexity while ensuring the accuracy of the model to meet real-time requirements is also an important problem faced by current technologies. Summary of the Invention

[0003] In view of the above problems, the present invention proposes a method and system for two-dimensional pose estimation in occluded scenes based on spatio-temporal information. By constructing a two-dimensional pose estimation model including an input module, a visibility prediction and masking module, a key-point reasoning module, and a spatio-temporal dependency module, the present invention solves the problems of lost key-point features and low pose estimation accuracy in existing video two-dimensional pose estimation technologies in occluded scenes, can accurately reason about occluded key points, improve the smoothness and accuracy of pose estimation, reduce high-frequency noise interference, and reduce the computational complexity while ensuring the accuracy of the model to meet real-time requirements.

[0004] On the one hand, a method for two-dimensional pose estimation in occluded scenes based on spatio-temporal information is as follows:

[0005] S1, obtain a video including multi-person body information;

[0006] S2. Build a two-dimensional pose estimation model; the model includes an input module, a visibility prediction and masking module, a key point inference module, a spatio-temporal dependence module, and a regression layer;

[0007] The input module identifies and crops individuals in the input video to obtain individual video clips for each individual, processes them one by one, and uses a convolutional backbone network to extract the key point features of each frame of the current individual video clip; the key points are human joints.

[0008] The visibility prediction and masking module performs visibility prediction on the input key point features to obtain mask features.

[0009] The key point inference module inputs the mask features, completes the occluded key points, and outputs enhanced key point features.

[0010] The spatio-temporal dependence module processes the input enhanced key point features frame by frame, extracts the time-dependent update features of each key point in the current frame; extracts the spatial-dependent update features between each key point in the current frame, and respectively fuses the time-dependent update features and spatial-dependent update features of each key point, and outputs enhanced spatio-temporal features frame by frame.

[0011] The regression layer processes the enhanced spatio-temporal features frame by frame for two-dimensional pose estimation to obtain the two-dimensional pose estimation sequence of the current individual; processes each individual to obtain the two-dimensional pose estimation sequences of all individuals; the two-dimensional pose estimation model is optimized using a residual likelihood estimation loss function.

[0012] S3. Input the obtained video into the two-dimensional pose estimation model to obtain a trained two-dimensional pose estimation model.

[0013] S4. Use the obtained trained two-dimensional pose estimation model for two-dimensional pose estimation in an occlusion scenario.

[0014] Preferably, the two-dimensional pose estimation model further includes a discrete cosine filtering module; the discrete cosine filtering module takes the two-dimensional pose estimation sequence as input, converts the two-dimensional pose estimation sequence into discrete cosine coefficients, uses a low-pass filter to retain a preset number of low-frequency discrete cosine coefficients, and then obtains a two-dimensional pose sequence with high-frequency noise removed through an inverse discrete cosine transform, and outputs the two-dimensional pose sequence with high-frequency noise removed as the output of the two-dimensional pose estimation model.

[0015] Preferably, the visibility prediction and masking module performs visibility prediction on the input key point features to obtain mask features, and the specific steps are as follows:

[0016] First, use a lightweight multi-layer perceptron module to perform visibility prediction on the key point features to obtain the Score ;

[0017] According to the score of each key point of the individual, a visibility mask is obtained based on a preset classification threshold, and the formula is as follows:

[0018] ;

[0019] Wherein, represents the visibility mask; represents function; represents a lightweight multi-layer perceptron module; represents the key point feature; represents the preset classification threshold;

[0020] Then, the visibility mask and the key point feature are multiplied element by element to obtain a masked feature, and the formula is as follows:

[0021] ;

[0022] Wherein, represents the masked feature; represents element-by-element multiplication.

[0023] Preferably, the key point inference module adopts network; the network is composed of 4 identical blocks stacked in sequence; each block includes a multi-head self-attention module and a multi-layer perceptron module; the processing steps of the network are specifically as follows:

[0024] First, the masked feature is input into the multi-head self-attention module; the self-attention heads of the multi-head self-attention module are represented as:

[0025] ;

[0026] Wherein, represents the self-attention head; represents function; , and respectively represent linear projection layers; represents the dimension of each masked feature; represents transpose of;

[0027] The multi-head self-attention module combines self-attention heads, represented as:​

[0028] ;

[0029] Among them, represents multi-head self-attention; represents the m-th self-attention head; represents a linear projection layer;

[0030] Secondly, perform residual connection and layer normalization, and the formula is as follows:

[0031] ;

[0032] Among them, represents the output result of residual connection and layer normalization; represents the layer normalization operation;

[0033] Then, input the multi-layer perceptron module, which is expressed as:

[0034] ;

[0035] Among them, represents the result of performing a feed-forward neural network operation on the input vector ; and respectively represent linear transformation matrices; and respectively represent bias terms; represents the activation function;

[0036] Finally, perform residual connection and layer normalization to obtain enhanced key-point features, and the formula is as follows:

[0037] ;

[0038] Among them, represents the enhanced key-point features.

[0039] Preferably, the enhanced spatio-temporal features are expressed as , specifically as follows:

[0040] , ;

[0041] Among them, represents the spatio-temporal aggregation feature of joint j of individual i in the current frame , , represents the time span, represents the set of all joints; represents the concatenation operation; Denote the time-dependent updated feature of joint j of individual i at the current frame ; Denote the space-dependent updated feature of joint j of individual i at the current frame ; n represents the number of joints.

[0042] Preferably, the first local perception attention module is used to obtain the time-dependent updated feature; the first local perception attention module includes a second Transfomer network; the second Transfomer network includes 4 identical blocks stacked in sequence, and each block includes a multi-head self-attention module and a multi-layer perception module; it is expressed as follows:

[0043] ;

[0044] Among them, denotes the enhanced key-point feature of the j-th joint of individual ; denotes local perception attention.

[0045] Preferably, the second local perception attention module is used to obtain the space-dependent updated feature; the second local perception attention module includes a third network; the third network includes 4 identical blocks stacked in sequence, and each block includes a multi-head self-attention module and a multi-layer perception module; specifically as follows:

[0046] Based on the semantic structure of human postures, the enhanced key-point features are divided into groups, and local perception attention is used for each group respectively, which is expressed as follows:

[0047] ; ;

[0048] Among them, denotes the set of group joint indicators; denotes the space-dependent updated feature of the k-th group; denotes local perception attention; denotes the enhanced key-point feature of joint j of individual i at the current frame ; denotes the of the k-th group.

[0049] Preferably, the regression layer includes fully connected layers equal in number to the number of joints. Each fully connected layer independently processes the features of one joint without sharing network parameters with other fully connected layers. The regression layer outputs the coordinates of each joint in the current frame, completing the two-dimensional pose estimation.

[0050] On the other hand, a system for two-dimensional pose estimation in an occlusion scenario based on spatio-temporal information includes the following:

[0051] A data acquisition module for acquiring a video including multi-person body information;

[0052] A model construction module for constructing a two-dimensional pose estimation model; the model includes an input module, a visibility prediction and masking module, a key point inference module, a spatio-temporal dependence module, and a regression layer;

[0053] A model training module for inputting the acquired video into the two-dimensional pose estimation model for training to obtain a trained two-dimensional pose estimation model;

[0054] A two-dimensional pose estimation module for performing two-dimensional pose estimation in an occlusion scenario using the obtained trained two-dimensional pose estimation model;

[0055] The input module identifies and crops individuals in the input video to obtain individual video segments of each individual, and sequentially extracts the key point features of each frame of each individual video segment using a convolutional backbone network, and outputs them one by one to the visibility prediction and masking module;

[0056] The visibility prediction and masking module performs visibility prediction on the input key point features to obtain mask features, and outputs them to the key point inference module;

[0057] The key point inference module complements the occluded key points based on the input mask features to obtain enhanced key point features, and outputs them to the spatio-temporal dependence module;

[0058] The spatio-temporal dependence module processes the input enhanced key point features frame by frame, extracts the time-dependent update features of each key point in the current frame and the space-dependent update features between each key point, and respectively fuses the time-dependent update features and space-dependent update features of each key point to obtain enhanced spatio-temporal features, and outputs them frame by frame;

[0059] The regression layer performs two-dimensional pose estimation based on the enhanced spatio-temporal features processed frame by frame to obtain a two-dimensional pose estimation sequence of the current individual; processes each individual to obtain two-dimensional pose estimation sequences of all individuals.

[0060] Compared with the prior art, the present invention has the following beneficial effects:

[0061] By constructing a two-dimensional pose estimation model including an input module, a visibility prediction and masking module, a key point inference module, a spatio-temporal dependence module, a regression layer, and a discrete cosine filtering module, the present invention can accurately infer occluded key points, improve the smoothness and accuracy of pose estimation, reduce the interference of high-frequency noise, ensure the accuracy of the model while reducing the computational complexity to meet the real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] The present invention will be further described in detail below with reference to the accompanying drawings;

[0063] Figure 1 is a flowchart of a method for two-dimensional pose estimation in an occlusion scenario based on spatio-temporal information according to an embodiment of the present invention;

[0064] Figure 2 is a schematic flow diagram of a method for two-dimensional pose estimation in an occlusion scenario based on spatio-temporal information according to an embodiment of the present invention;

[0065] Figure 3 is a structural block diagram of a system for two-dimensional pose estimation in an occlusion scenario based on spatio-temporal information according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] The present invention will be further described below through specific embodiments.

[0067] As Figure 1 shown, a method for two-dimensional pose estimation in an occlusion scenario based on spatio-temporal information:

[0068] S1. Obtain a video including multi-person body information;

[0069] The video is represented as , where is a video frame including multiple persons at a given time , and is a predefined time span.

[0070] S2. Construct a two-dimensional pose estimation model; the model includes an input module, a visibility prediction and masking module, a key point inference module, a spatio-temporal dependence module, and a regression layer. See the two-dimensional pose estimation model in Figure 2 shown.

[0071] Input module: The input module identifies and crops individuals in the input video to obtain individual video segments of each individual, and uses a convolutional backbone network to sequentially extract key point features of each frame of each individual video segment, and outputs them to the visibility prediction and masking module one by one, specifically as follows:

[0072] The input module receives the video . First, use a human body detector in Identify individuals and expand the bounding boxes of each detected individual by 25% to extract the same individual in the sequence, forming a cropped video segment , where is an individual in the video. The Extract the key point features of the individual through a convolutional backbone network , where , represents the key point feature of the th joint of the individual pose at the video frame of time , then represents all the key point features of the th joint of the individual pose at the video frame of time .

[0073] Specifically, the convolutional backbone network is the core part of the convolutional neural network (CNN) in deep learning, responsible for extracting features from the input data. It usually includes multiple convolutional layers, activation functions, and pooling layers, and extracts features from low-level to high-level through layer-by-layer processing of the image.

[0074] Visibility prediction and masking module, the visibility prediction and masking module performs visibility prediction on the input key point features, obtains the mask features, and outputs them to the key point reasoning module, specifically as follows:[[]]

[0075] First, use a lightweight multi-layer perceptron module to capture the visibility state of a single key point, that is, the lightweight multi-layer perceptron module predicts the scores of the visibility of all key points given the input . By using the function. The formula is as follows:[[]] ;

[0076] ;

[0077] where is an unknown. Convert the score to a probability, where a value close to 0 indicates an invisible key point, and a value close to 1 indicates a visible key point. According to the obtained scores of each key point, apply a classification threshold of 0.5 to obtain the visibility mask, which is represented as . The visibility mask is a -dimensional vector. The visibility mask function is as follows:[[]]

[0078] ;

[0079] Then, apply the visibility mask to the key-point features to mask the corresponding features of the invisible key points. That is, multiply the visibility mask element-wise with the key-point features to obtain the masked features, which are called . The formula is as follows:

[0080] ;

[0081] where, means multiplying each pair of corresponding elements in and . The obtained masked features mask the influence of the features of the occluded key points, improving the accuracy of pose estimation.

[0082] Specifically, the lightweight multi-layer perceptron module adopts , which is designed specifically for predicting the visibility state (visible or occluded) of human key points. By performing binary classification on the input key-point features, it outputs the visibility probability of each key point, which is used to generate the visibility mask, thereby masking the features of the occluded key points and avoiding interference with the model inference. By explicitly introducing visibility information, it can effectively reduce the negative impact of occlusion on pose estimation and significantly improve the prediction accuracy and robustness in complex scenarios.

[0083] The key-point inference module, based on the input masked features, completes the occluded key points, obtains enhanced key-point features, and outputs them to the spatio-temporal dependence module, as follows:

[0084] The key-point inference module takes the masked features as input and uses network to infer the occluded key points. The finally output features contain the features of all key points, where the occluded key points are inferred and the visible key points are optimized. Taking the masked features as input and outputting the enhanced key-point features through visibility-aware attention. Among them, network uses four original identical blocks stacked in sequence as network. Each block contains a multi-head self-attention module and a multi-layer perceptron module. The self-attention head is expressed as:

[0085] ;

[0086] ;

[0087] where, , , is a linear projection layer, is the dimension of each masked feature .

[0088] The multi-head self-attention module combines self-attention heads, denoted as:

[0089] ;

[0090] where, is the dimension of each self-attention head, is a linear projection layer.

[0091] Perform residual connection and layer normalization:

[0092] ;

[0093] where, is the output result of the residual connection and layer normalization, is the layer normalization operation.

[0094] The multi-layer perceptron module is denoted as:

[0095] ;

[0096] where, represents the result of performing a feed-forward neural network operation on the input vector . , is a linear transformation matrix; , is a bias term. represents the activation function.

[0097] Perform residual connection and layer normalization to obtain enhanced key-point features , denoted as:

[0098] ;

[0099] For the current individual over the time span , the enhanced key-point features can also be denoted as:

[0100] ;

[0101] where, represents the th pose of individual ​Enhanced key point features of a joint, then denotes all key point features of the individuals at the video frame in time for a joint.

[0102] The spatio-temporal dependence module processes the input enhanced key point features frame by frame, extracts the time-dependent update features of each key point in the current frame and the spatial dependence update features between each key point, and respectively fuses the time-dependent update features and the spatial dependence update features of each key point to obtain enhanced spatio-temporal features, which are output frame by frame, specifically as follows:

[0103] The spatio-temporal dependence module separately models the time dynamic dependence relationship and the spatial structure dependence relationship according to the feature embedding of the joint over the time span , that is module and module, so as to capture the time and spatial dependence relationships between joints, and then by fusing the captured spatial and time information, the aggregated spatio-temporal features of each joint in the current frame are derived, expressed as:

[0104] ;

[0105] where denotes the concatenation operation, which is respectively applied to each pair of corresponding updated feature tags related to each joint.

[0106] module. The module recognizes the time dependence of each joint, thus generating another updated feature for each joint in the current frame. In the module, the time dynamic dependence relationship of each joint in the current frame is captured. The local perception attention module is selectively applied over the time span . Among them, the configuration of follows the

[0107] network in the key point inference module: four identical

[0108] blocks are stacked in sequence, and each block contains a multi-head self-attention module and a multi-layer perception module. The input features pass through these modules in sequence, and each module generates an updated version as the input for the subsequent layer. That is: , denotes the individual Enhanced key-point features of the j-th joint representing the current video frame on the th joint's updated features, encoding the temporal dependency information of this joint embedded in the sequence .

[0109] module The module also utilizes the local perception attention module to learn the spatial dependency relationships between adjacent joints and generate updated features for each joint in the current frame accordingly. In the module, the spatial structural dependencies between joints within the current frame are captured. According to the semantic structure of the human pose, the joints are divided into groups, and for each group, is used respectively, expressed as:

[0110] , ;

[0111] wherein, represents the set of group joint indices, representing the updated feature label of the th joint in the current video frame , containing the spatial structural dependency of this joint in the pose

[0112] Through and the module , the spatio-temporal context of each joint in the current frame is captured, obtaining the corresponding updated features and . Therefore, the spatio-temporal aggregation features of each joint under the current frame can be explicitly calculated as:

[0113] , ;

[0114] The spatio-temporal aggregation features output combine temporal dynamics and spatial structural information, enabling subsequent models to more comprehensively understand the input data

[0115] Regression layer. The regression layer performs two-dimensional pose estimation based on the enhanced spatio-temporal information features processed frame by frame, obtaining the two-dimensional pose estimation sequence of the current individual; processing each individual separately to obtain the two-dimensional pose estimation sequences of all individuals, specifically as follows:

[0116] The regression layer is a fully connected network, with the input spatio-temporal aggregation features , the output is the current frame in the coordinates of the th joint, that is:

[0117] ;

[0118] Among them, is a per-joint fully connected layer, that is, for each joint a separate fully connected layer is used for processing . The fully connected layers of each joint work independently, only processing the features related to that joint and not sharing network parameters with the features of other joints. Process the spatio-temporal aggregation features output by the spatio-temporal dependence module frame by frame to obtain the pose estimation coordinates of all video frames in the video segment in the time span , that is, the two-dimensional pose sequence .

[0119] The two-dimensional pose estimation model also includes a discrete cosine filtering module. The input of the discrete cosine filtering module is the two-dimensional pose sequence , among which, is the length of the two-dimensional pose sequence. Convert the two-dimensional pose sequence into discrete cosine coefficients, denoted by . Then, use a low-pass filter for each joint trajectory, only retaining the first low-frequency discrete cosine coefficients in the sequence, discarding the coefficients representing high-frequency noise, ensuring that key time information is retained and excluding high-frequency components that may cause data distortion or interference. Perform an inverse discrete cosine transform on to obtain the filtered two-dimensional pose sequence. The discrete cosine filtering module removes unnecessary high-frequency noise and maximally retains the time information and dynamic characteristics of the original sequence, making the processed data smoother and more accurate.

[0120] S3, input the acquired video into the two-dimensional pose estimation model to obtain the trained two-dimensional pose estimation model;

[0121] S4, use the obtained trained two-dimensional pose estimation model to perform two-dimensional pose estimation for occlusion scenarios based on spatio-temporal information.

[0122] As Figure 3 shown, the present invention also discloses a system for two-dimensional pose estimation of occlusion scenarios based on spatio-temporal information, including:

[0123] A data acquisition module 301 for acquiring a video including multi-person body information.

[0124] The model construction module 302 is used to construct a two-dimensional pose estimation model; the model includes an input module, a visibility prediction and masking module, a key point inference module, a spatio-temporal dependency module, and a regression layer.

[0125] The model training module 303 is used to input the obtained video into the two-dimensional pose estimation model for training to obtain a trained two-dimensional pose estimation model.

[0126] The two-dimensional pose estimation module 304 is used to perform two-dimensional pose estimation for an occlusion scenario using the obtained trained two-dimensional pose estimation model.

[0127] The input module identifies and crops individuals in the input video to obtain individual video segments of each individual, and sequentially extracts key point features of each frame of each individual video segment using a convolutional backbone network, and outputs them one by one to the visibility prediction and masking module.

[0128] The visibility prediction and masking module performs visibility prediction on the input key point features to obtain mask features, and outputs them to the key point inference module.

[0129] The key point inference module complements the occluded key points based on the input mask features to obtain enhanced key point features, and outputs them to the spatio-temporal dependency module.

[0130] The spatio-temporal dependency module processes the input enhanced key point features frame by frame, extracts the time-dependent update features of each key point in the current frame and the spatial-dependent update features between each key point, and respectively fuses the time-dependent update features and spatial-dependent update features of each key point to obtain enhanced spatio-temporal features, and outputs them frame by frame.

[0131] The regression layer performs two-dimensional pose estimation based on the spatio-temporal features enhanced by frame-by-frame processing to obtain a two-dimensional pose estimation sequence of the current individual; processes each individual to obtain two-dimensional pose estimation sequences of all individuals.

[0132] The specific implementation of the system for two-dimensional pose estimation in an occlusion scenario based on spatio-temporal information is the same as that of the method for two-dimensional pose estimation in an occlusion scenario based on spatio-temporal information, and will not be repeated in this embodiment.

[0133] The above is only the specific implementation manner of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantive modification made to the present invention using this concept shall fall within the scope of infringement of the protection scope of the present invention.

Claims

1. A method for estimating two-dimensional pose of an occluded scene based on spatiotemporal information, characterized in that: The steps include: S1, obtaining a video including multiple human body information; S2, constructing a two-dimensional posture estimation model; the model includes an input module, a visibility prediction and shielding module, a key point reasoning module, a spatiotemporal dependency module and a regression layer; The input module identifies and crops individuals in the input video to obtain individual video clips of each individual, uses a convolutional backbone network to extract key point features of each frame of each individual video clip in turn, and outputs them one by one to the visibility prediction and shielding module; The visibility prediction and masking module performs visibility prediction on the input key point features to obtain mask features and outputs them to the key point reasoning module; The key point reasoning module completes the occluded key points based on the input mask features, obtains enhanced key point features, and outputs them to the spatiotemporal dependency module; The spatiotemporal dependency module processes the input enhanced key point features frame by frame, extracts the time-dependent update features of each key point in the current frame and the space-dependent update features between each key point, and respectively fuses the time-dependent update features and space-dependent update features of each key point to obtain enhanced spatiotemporal features, and outputs them frame by frame; The regression layer performs two-dimensional posture estimation based on the enhanced spatiotemporal features processed frame by frame to obtain a two-dimensional posture estimation sequence of the current individual; and processes each individual one by one to obtain a two-dimensional posture estimation sequence of all individuals; S3, inputting the acquired video into a two-dimensional posture estimation model for training to obtain a trained two-dimensional posture estimation model; S4, using the trained two-dimensional pose estimation model to perform two-dimensional pose estimation of the occluded scene; The key point reasoning module adopts a Transformer network; the Transformer network is composed of 4 identical Transformer blocks stacked in sequence; each Transformer block includes a multi-head self-attention module and a multi-layer perception module; the processing steps of the Transformer network are as follows: First, the mask features are input into the multi-head self-attention module; the self-attention head of the multi-head self-attention module is expressed as: in, represents the self-attention head; softmax(·) represents the softmax function; W Q , W K and W V They represent the linear projection layer respectively; d represents the dimension of each mask feature; express The transpose of The multi-head self-attention module combines h self-attention heads, expressed as: in, It indicates long self-attention; represents the mth self-attention head; W P Represents a linear projection layer; Secondly, perform residual connection and layer normalization, the formula is as follows: in, Represents the output result of residual connection and layer normalization; LayerNorm represents the layer normalization operation; Then, the multi-layer perception module is input, which is expressed as: in, Represents the input vector The result of the feedforward neural network operation; W1 and W2 represent linear transformation matrices; b1 and b2 represent bias terms; max(0,·) represents the ReLU activation function; Finally, residual connection and layer normalization are performed to obtain enhanced key point features. The formula is as follows: in, Represents enhanced key point features.

2. The method for estimating two-dimensional posture of an occluded scene based on spatiotemporal information according to claim 1, characterized in that: The two-dimensional posture estimation model also includes a discrete cosine filtering module; the discrete cosine filtering module takes a two-dimensional posture estimation sequence as input, converts the two-dimensional posture estimation sequence into discrete cosine coefficients, uses a low-pass filter to retain a preset number of low-frequency discrete cosine coefficients, and then obtains a two-dimensional posture sequence with high-frequency noise removed through a discrete inverse cosine transform, and outputs the two-dimensional posture sequence with high-frequency noise removed as the output of the two-dimensional posture estimation model.

3. The method for estimating two-dimensional posture of an occluded scene based on spatiotemporal information according to claim 1, characterized in that: The visibility prediction and masking module performs visibility prediction on the input key point features to obtain mask features. The specific steps are as follows: First, a lightweight multi-layer perceptron module is used to train key point features. Perform visibility prediction and obtain the logit score of each key point of the individual According to the logit score of each key point of the individual, the visibility mask is obtained based on the preset classification threshold. The formula is as follows: in, represents visibility mask; sigmoid(·) represents sigmoid function; V(·) represents a lightweight multilayer perceptron module; Indicates key point features; 0.5 indicates the preset classification threshold; Then, the visibility mask is multiplied element-by-element with the key point feature to obtain the mask feature, as shown below: in, represents mask features; ⊙ represents element-by-element multiplication.

4. The method for estimating two-dimensional posture of an occluded scene based on spatiotemporal information according to claim 1, characterized in that: The enhanced spatiotemporal features are expressed as The details are as follows: Among them, f i j (t) represents the spatiotemporal aggregate features of joint j of individual i in the current frame t, t∈[tT,t+T], T represents the time span, represents the set of all joints; Indicates a serial operation; Represents the time-dependent update feature of joint j of individual i at the current frame t; represents the spatially dependent update feature of joint j of individual i in the current frame t; n represents the number of joints.

5. The method for estimating two-dimensional posture of an occluded scene based on spatiotemporal information according to claim 4, characterized in that: The first local perception attention module is used to acquire time-dependent update features; the first local perception attention module includes a second Transformer network; the second Transformer network includes 4 identical Transformer blocks stacked in sequence, each Transformer block includes a multi-head self-attention module and a multi-layer perception module; it is expressed as follows: in, represents the enhanced keypoint features of the j-th joint of individual i; ATT(·) represents local-aware attention.

6. The method for estimating two-dimensional posture of an occluded scene based on spatiotemporal information according to claim 4, characterized in that: A second local perception attention module is used to acquire spatial dependency update features; the second local perception attention module includes a third Transformer network; the third Transformer network includes 4 identical Transformer blocks stacked in sequence, each Transformer block includes a multi-head self-attention module and a multi-layer perception module; specifically as follows: Based on the semantic structure of human posture, the enhanced key point features are divided into K groups, and local-aware attention is used for each group, which is expressed as follows: Among them, G(k) represents a set of k sets of joint indicators; represents the spatially dependent updated features of the kth group; ATT(·) represents local perceptual attention; represents the enhanced key point features of joint j of individual i in the current frame t, represents the kth group 7. The method for estimating two-dimensional posture of an occluded scene based on spatiotemporal information according to claim 1, characterized in that: The regression layer includes fully connected layers equal to the number of joints. Each fully connected layer independently processes the features of a joint and does not share network parameters with other fully connected layers. The regression layer outputs the coordinates of each joint in the current frame to complete two-dimensional posture estimation.

8. A system for estimating two-dimensional pose of an occluded scene based on spatiotemporal information, comprising: A data acquisition module, used for acquiring a video including a plurality of human body information; A model building module, used to build a two-dimensional posture estimation model; the model includes an input module, a visibility prediction and shielding module, a key point reasoning module, a spatiotemporal dependency module and a regression layer; A model training module is used to input the acquired video into a two-dimensional posture estimation model for training to obtain a trained two-dimensional posture estimation model; A two-dimensional pose estimation module is used to perform two-dimensional pose estimation of occluded scenes using a trained two-dimensional pose estimation model; The input module identifies and crops individuals in the input video to obtain individual video clips of each individual, uses a convolutional backbone network to extract key point features of each frame of each individual video clip in turn, and outputs them one by one to the visibility prediction and shielding module; The visibility prediction and masking module performs visibility prediction on the input key point features to obtain mask features and outputs them to the key point reasoning module; The key point reasoning module completes the occluded key points based on the input mask features, obtains enhanced key point features, and outputs them to the spatiotemporal dependency module; The spatiotemporal dependency module processes the input enhanced key point features frame by frame, extracts the time-dependent update features of each key point in the current frame and the space-dependent update features between each key point, and respectively fuses the time-dependent update features and space-dependent update features of each key point to obtain enhanced spatiotemporal features, and outputs them frame by frame; The regression layer performs two-dimensional posture estimation based on the enhanced spatiotemporal features processed frame by frame to obtain a two-dimensional posture estimation sequence of the current individual; and processes each individual one by one to obtain a two-dimensional posture estimation sequence of all individuals; The key point reasoning module adopts a Transformer network; the Transformer network is composed of 4 identical Transformer blocks stacked in sequence; each Transformer block includes a multi-head self-attention module and a multi-layer perception module; the processing steps of the Transformer network are as follows: First, the mask features are input into the multi-head self-attention module; the self-attention head of the multi-head self-attention module is expressed as: in, represents the self-attention head; softmax(·) represents the softmax function; W Q , W K and W V They represent the linear projection layer respectively; d represents the dimension of each mask feature; express The transpose of The multi-head self-attention module combines h self-attention heads, expressed as: in, It indicates long self-attention; represents the mth self-attention head; W P Represents a linear projection layer; Secondly, perform residual connection and layer normalization, the formula is as follows: in, Represents the output result of residual connection and layer normalization; LayerNorm represents the layer normalization operation; Then, the multi-layer perception module is input, which is expressed as: in, Represents the input vector The result of the feedforward neural network operation; W1 and W2 represent linear transformation matrices; b1 and b2 represent bias terms; max(0,·) represents the ReLU activation function; Finally, residual connection and layer normalization are performed to obtain enhanced key point features. The formula is as follows: in, Represents enhanced key point features.

Citation Information

Patent Citations

  • Video three-dimensional human body posture estimation method and system based on multistage supervision graph convolution

    CN114694261A

  • Three-dimensional human body posture estimation method based on multi-scale space-time encoder network

    CN118674778A