Multi-modal human body posture estimation model training method based on space-time Transform

By constructing a multimodal human pose estimation model, a text feature extraction network was introduced for global feature comparison learning and optimization, which solved the problem of insufficient feature capture in space-time Transformer in human pose estimation, and achieved more efficient pose capture and more accurate three-dimensional pose estimation.

CN120387028APending Publication Date: 2025-07-29CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510463984.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing space-time Transformer architecture cannot effectively capture more features in human pose estimation, resulting in inaccurate pose capture.

Method used

A multimodal human pose estimation model was constructed, a text feature extraction network was introduced to perform global text features and global pose features comparison learning, and a comparison loss was used to optimize the network parameters of the shallow space-time Transformer cascade network and text feature extraction network. The global pose features of the shallow space-time Transformer cascade network were compared and learned, and the Token pruning and recovery module optimization model was combined.

Benefits of technology

The accuracy and efficiency of human body's three-dimensional posture capture is improved. Through global feature alignment and local feature optimization, the model can identify key action features earlier, improving the multimodal alignment effect and overall model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387028A_ABST
    Figure CN120387028A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of posture capture, and provides a space-time Transform-based multi-modal human body posture estimation model training method, which comprises the following steps that: a multi-modal human body posture estimation network comprises a shallow layer space-time Transform cascade network and a deep layer space-time Transform cascade network; obtaining a sample pair set; iterative training is carried out on the text feature extraction network and the multi-modal human body posture estimation network based on the sample pair set, and comparative learning is carried out on global posture features obtained by the shallow space-time Transform cascade network and global text features obtained by the text feature extraction network; optimizing network parameters of the shallow space-time Transform cascade network and the text feature extraction network based on comparison loss, and optimizing network parameters of the visual projection layer, the deep space-time Transform cascade network and the attitude output layer based on joint position errors; the invention further discloses a space-time Transform-based multi-modal human body posture estimation method, a computer program product and electronic equipment. The posture estimation accuracy is improved by the space-time Transform-based multi-modal human body posture estimation method and the space-time Transform-based multi-modal human body posture estimation device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of gesture capture, and particularly to a training method for a multi-modal human pose estimation model based on spatio-temporal Transformer. Background Art

[0002] Transformer has gradually become the preferred architecture in the field of deep learning due to its high efficiency, scalability, and powerful modeling capabilities, and has been successfully applied to 3D human pose estimation. Through the self-attention mechanism, Transformer can effectively capture long-range dependencies and complex spatio-temporal features, especially suitable for processing tasks with complex spatial structures and time dependencies. Alternately using spatial Transformer and temporal Transformer can effectively capture the action sequence of the human body in three-dimensional space and handle dynamic changes in the time dimension. Compared with traditional geometric prior-based human pose estimation models, Transformer can automatically learn complex motion patterns and semantic information from data, significantly improving the generalization ability and flexibility of the model.

[0003] There are still certain limitations in the spatio-temporal Transformer architecture during the spatio-temporal modeling of human poses. Although alternately using spatial Transformer and temporal Transformer enables feature sharing at different stages, the feature interaction between the two is relatively limited, and more features cannot be captured, resulting in inaccurate human pose capture. Summary of the Invention

[0004] This application aims to at least solve the technical problems existing in the prior art and provides a training method for a multi-modal human pose estimation model based on spatio-temporal Transformer.

[0005] In a first aspect, the present application provides a method for training a multi-modal human pose estimation model based on spatio-temporal Transformer, including: constructing a text feature extraction network and a multi-modal human pose estimation network, where the multi-modal human pose estimation network includes a 2D skeleton joint extraction network, a visual projection layer, a shallow spatio-temporal Transformer cascade network, a deep spatio-temporal Transformer cascade network, and a pose output layer connected in sequence; obtaining a set of sample pairs, each sample pair including consecutive video frames and a pose description text; iteratively training the text feature extraction network and the multi-modal human pose estimation network in batches based on the set of sample pairs. When the iterative training stop condition is reached, the trained multi-modal human pose estimation network is used as the multi-modal human pose estimation model; in each batch of iterative training, the consecutive video frames of the sample pair are input into the multi-modal human pose estimation network, and the pose description text of the sample pair is input into the text feature extraction network; performing contrastive learning on the global pose features obtained by the shallow spatio-temporal Transformer cascade network and the global text features obtained by the text feature extraction network to obtain a contrastive loss; calculating the joint position error according to the three-dimensional estimated human pose output by the pose output layer; optimizing the network parameters of the shallow spatio-temporal Transformer cascade network and the text feature extraction network based on the contrastive loss; and optimizing the network parameters of the visual projection layer, the deep spatio-temporal Transformer cascade network, and the pose output layer based on the joint position error.

[0006] In a second aspect, the present application provides a multi-modal human pose estimation method based on spatio-temporal Transformer, including: obtaining consecutive video frames to be estimated, and inputting the consecutive video frames to be estimated into the multi-modal human pose estimation model, where the multi-modal human pose estimation model outputs a three-dimensional estimated human pose; and the multi-modal human pose estimation model is trained according to the method for training a multi-modal human pose estimation model based on spatio-temporal Transformer described in the first aspect of the present application.

[0007] In a third aspect, the present application provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, the steps of the method described in the first aspect or the second aspect of the present application are implemented.

[0008] In a fourth aspect, the present application provides an electronic device, where the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the steps of the method described in the first aspect or the second aspect of the present application.

[0009] Advantageous technical effects of the present application:

[0010] In the training of the multi-modal human pose estimation network, a text feature extraction network is introduced to perform contrastive learning between global text features and global pose features, and the contrastive loss is calculated. The network parameters of the shallow spatio-temporal Transformer cascade network and the text feature extraction network are optimized using the contrastive loss to achieve the alignment of global pose features and global text features, with semantic consistency, which helps the multi-modal human pose estimation model to understand and fuse features of different modalities and improve the accuracy of human three-dimensional pose capture;

[0011] In the contrastive learning, only the global pose features of the shallow spatio-temporal Transformer cascade network are used for contrastive learning, rather than the pose features output by the deep spatio-temporal Transformer cascade network, because the shallow spatio-temporal Transformer cascade network can more effectively capture short-term dynamic features and local details related to actions, which is crucial for understanding and aligning the relationship between text and pose. Although the deep spatio-temporal Transformer cascade network can extract global dependencies and abstract features, the shallow spatio-temporal Transformer cascade network can identify key action features earlier and more accurately, thus improving the effect of multi-modal alignment and the accuracy of the overall model; and using global pose features and global text features for contrastive learning, rather than local features, enables the model to automatically focus on important joint information and key time steps and ignore unimportant information, thereby improving the accuracy of pose estimation. Brief Description of the Drawings

[0012] Figure 1 is a schematic flowchart of the training method of the multi-modal human pose estimation model in a preferred embodiment of the present invention;

[0013] Figure 2 is a schematic flowchart of the iterative training for each batch in a preferred embodiment of the present invention;

[0014] Figure 3 is a network architecture diagram of the training method of the multi-modal human pose estimation model in a preferred embodiment of the present invention;

[0015] Figure 4 is a schematic diagram of the network structure of the spatio-temporal Transformer module in a preferred embodiment of the present invention;

[0016] Figure 5 is a schematic diagram of the processing process of the Token pruning module in a preferred embodiment of the present invention;

[0017] Figure 6It is a schematic diagram of the processing process of the Token recovery module in a preferred embodiment of the present invention;

[0018] Figure 7 It is a schematic diagram of the network structure of the feature fusion module in a preferred embodiment of the present invention;

[0019] Figure 8 It is a schematic diagram of the structure of the multi-modal human pose estimation network in a preferred embodiment of the present invention;

[0020] Figure 9 It is a schematic diagram of the network structure of the text feature extraction network in a preferred embodiment of the present invention;

[0021] Figure 10 It is a schematic diagram of the network for training the multi-modal human pose estimation model in a preferred embodiment of the present invention;

[0022] Figure 11 It is a schematic diagram of the structure of the electronic device in a preferred embodiment of the present invention. Specific embodiments

[0023] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.

[0024] The execution subject of the chest radiograph lesion detection method based on enhanced feature extraction and fusion or the multi-modal human pose estimation method based on spatio-temporal Transformer provided by the present invention includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiments of the present application. In other words, the chest radiograph lesion detection method based on enhanced feature extraction and fusion or the multi-modal human pose estimation method based on spatio-temporal Transformer can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0025] The present invention provides a method for training a multi-modal human pose estimation model based on spatio-temporal Transformer. In a preferred embodiment, please refer to Figure 1, the training method includes:

[0026] Step S1, construct a text feature extraction network and a multi-modal human pose estimation network. The multi-modal human pose estimation network includes a 2D skeleton joint extraction network, a visual projection layer, a shallow spatio-temporal Transformer cascade network, a deep spatio-temporal Transformer cascade network, and a pose output layer connected in sequence; obtain a set of sample pairs, and each sample pair includes consecutive video frames and a pose description text.

[0027] In this embodiment, Figure 3 a training network structure framework diagram is given. The text feature extraction network is used to process the pose description text of the sample pair, and existing CLIP text encoders, Word2Vec encoders, or BERT encoders can be used. The 2D skeleton joint extraction network is preferably but not limited to an HRNet (High-Resolution Net) network or a Stacked Hourglass Network network. The 2D skeleton joint extraction network can be pre-trained. The 2D skeleton joint extraction network is used to extract a sequence of 2D skeleton joint points from consecutive video frames. One set of 2D skeleton joint points is extracted from each frame of the image, and each set contains 17 joint points in total. Each joint point is represented by its two-dimensional coordinates, which are used to describe the human pose in each frame of the image. The sequence of 2D skeleton joint points of the consecutive video frames extracted is denoted as K in ∈R F*N*2 , where F represents the number of time frames, that is, the number of consecutive video frames. Preferably, F = 243; N = 17 represents the number of joints, and 2 is the number of channels, representing the two-dimensional coordinate components (x, y) of each joint point. The visual projection layer is preferably a linear projection layer, and a fully connected network can be used. The two-dimensional coordinates of each joint point in each frame of the image are projected to a high-dimensional embedding feature space through the visual projection layer to obtain the joint point embedding feature K p ∈

[0028] R F*N*dim , where dim is the embedding dimension and is a positive integer. Processing through the visual projection layer can enhance the feature expression ability.

[0029] In this embodiment, the shallow spatio-temporal Transformer cascade network includes two or more cascaded first spatio-temporal Transformer modules. The first spatio-temporal Transformer module includes a first spatial Transformer unit and a first temporal Transformer unit connected in sequence. The deep spatio-temporal Transformer cascade network includes two or more cascaded second spatio-temporal Transformer modules. The second spatio-temporal Transformer module includes a second spatial Transformer unit and a second temporal Transformer unit connected in sequence. The first spatio-temporal Transformer module and the second spatio-temporal Transformer module adopt Figure 4 the spatio-temporal Transformer module network structure shown. The first spatial Transformer unit and the second spatial Transformer unit adopt Figure 4 the Spatial Transformer Model network in Figure 4 . The first temporal Transformer unit and the second temporal Transformer unit adopt

[0030] the Temporal Transformer Model network in Figure 8 and Figure 10 . The shallow spatio-temporal Transformer cascade network includes n1 first spatio-temporal Transformer modules, where n1 is a positive integer and n1 = depth / 4, and depth is the total depth of the multi-modal human pose estimation network.

[0031] In this embodiment, preferably, see Figure 8 and Figure 10 . A position embedding unit is provided before the first spatial Transformer unit of the first first spatio-temporal Transformer module. The position embedding unit is used to embed a learnable spatial position embedding matrix in the joint embedding features output by the visual projection layer. A corresponding Token is set for each joint in each frame in the spatial position embedding matrix.

[0032] In this embodiment, to enhance the ability of the first spatio-temporal Transformer module in the shallow spatio-temporal Transformer cascade network to perceive the spatio-temporal structure when processing the joint embedding feature sequence, before the joint embedding feature sequence enters the first spatial Transformer unit and the first temporal Transformer unit, spatial position information (i.e., spatial position embedding matrix) and temporal position information (i.e., temporal position embedding matrix) are respectively embedded. The spatial position information enables the model to distinguish different position relationships. A learnable spatial position embedding matrix: P s ∈R 1*N*dim , and each joint (Token) in each frame has an independent embedding vector representing its topological relationship in the skeleton. The spatial position embedding matrix is added to the projected layer joint embedding feature K p output by the visual projection layer to obtain the feature K s : K s = K p + P s . Through the broadcast mechanism, P s is extended to the shape (F * N * dim), and each frame shares the same spatial position information embedding. The temporal position information embedding is used to represent the temporal position of the current frame in the sequence, helping the model understand the order and dynamic characteristics of the sequence. Similarly, a learnable temporal position embedding matrix: P t ∈R F*1*dim is created, and each frame has an independent temporal position encoding representing its position in the time series. This temporal position embedding is added to the input feature of the first temporal Transformer unit to obtain the final embedding feature: K st = K p + P s + P t .

[0033] In this embodiment, the spatio-temporal Transformer module is an extension based on the Transformer architecture for deep modeling of the spatio-temporal features of the input data. This module alternately uses spatial Transformer units and temporal Transformer units to respectively model the spatial dependency relationships between joints and the dynamic changes in the temporal dimension. By using multiple layers of spatio-temporal Transformer modules to capture the deeper relationships of spatial and temporal features layer by layer, this alternating processing method refines the spatio-temporal features layer by layer and enhances the understanding of human actions layer by layer. It also significantly improves the stability of cross-frame motion modeling, enabling the model to accurately capture complex pose changes.

[0034] In this embodiment, see Figure 4, the spatial Transformer unit captures the dependencies between joints within each frame of the video through the self-attention mechanism and generates high-dimensional feature representations for each joint. By calculating the correlations between joint points in the input feature map, the spatial Transformer unit can adaptively weight the spatial positions and enhance the expressive power of key regions. These relationship weights are represented by attention scores, where higher scores indicate a greater influence of the joint on other joints. During the calculation, the high-dimensional features obtained by processing each joint point through the visual projection layer are regarded as a Token, and the high-dimensional features of the 2D coordinates of each joint generate three vectors: query vector (Query, Q), key vector (Key, K), and value vector (Value, V). The attention score calculation formula is as follows:

[0035]

[0036] where Q, K, and V are the query, key, and value matrices respectively, and dim is the scaling factor of the feature dimension to prevent the dot product calculation value from being too large and causing unstable gradients. Summing up these weighted outputs for all joints generates high-dimensional feature representations for each joint, which can reflect the spatial relationships between joints.

[0037] In this embodiment, please see Figure 4 , the temporal Transformer unit focuses on capturing the dynamic changes of human joint points in the time series. The joint features of each time frame output by the previous temporal Transformer unit are regarded as a Token, and the temporal Transformer unit calculates the dependencies between time steps through the temporal self-attention mechanism to capture the temporal consistency of human motion patterns. Similar to the spatial Transformer unit, the temporal Transformer unit calculates the correlations between different time steps through the attention mechanism and adaptively adjusts the weights to capture the dynamic changes and cross-frame temporal dependencies in the time series.

[0038] In this embodiment, to ensure the continuity of the information flow and alleviate the problem of gradient vanishing or gradient explosion in the deep network, please see Figure 4 , the present invention introduces residual connections (Residual Connection) in both the temporal Transformer unit and the temporal Transformer unit, adding the attention output and the input features through the residual connection. In addition, layer normalization (Layer Normalization, LN) is applied after calculating the attention to normalize the feature dimensions of each sample to reduce the internal covariate shift, stabilize the gradient flow, thereby improving the training stability, accelerating convergence, and enhancing the generalization ability of the model. The formula is as follows:

[0039] K out1 = LayerNorm(K in1 + Attention(Q, K, V) (1.2)

[0040] where K in1 is the input feature of the temporal Transformer unit or the temporal Transformer unit, such as the output of the previous spatio-temporal Transformer module. Attention(Q, K, V) is the output after being processed by the attention mechanism.

[0041] In this embodiment, the 2D skeleton joint extraction network extracts a sequence of 2D skeleton joint points from consecutive video frames. The visual projection layer projects the sequence of 2D skeleton joint points into a high-dimensional embedded feature space to obtain joint point embedded features. In the output features of the last first temporal Transformer unit of the shallow spatio-temporal Transformer cascade network, each time frame corresponds to a local pose feature, and the local pose feature is expressed as where T represents the number of time frames. Here, T = 243. These local features represent the two-dimensional positions of each joint and their dynamic information in the time series. According to experimental verification, the number of frames 243 can cover a typical action cycle. The local pose features are globally weighted and summed to generate the global pose feature h< POSE> , and the model can automatically focus on important joint information and key time steps and ignore unimportant information. The local pose features of consecutive frames continue to be input into the deep spatio-temporal Transformer cascade network for processing to obtain the final pose features.

[0042] In this embodiment, the pose output layer includes a cascaded normalization layer and a fully connected layer. The input features of the pose output layer are normalized by the normalization layer to stabilize the training and improve the generalization ability of the model. The normalized features are mapped to a three-dimensional output space through the fully connected layer to obtain the final three-dimensional coordinate output K out ∈ R F*N*3 , where 3 represents the 3 components (x, y, z) of the three-dimensional coordinates of each key point, that is, the pose information of the human body in the 3D space.

[0043] In this embodiment, in step S1, a set of sample pairs is obtained. Each sample pair includes consecutive video frames and pose description texts, and also includes a true similarity label indicating whether the consecutive video frames and pose description texts of the sample pair match, and a sequence of true three-dimensional coordinates of human body joint points corresponding to the consecutive video frames. If the consecutive video frames and pose description texts of the sample pair match, such as Figure 10Among them, if the consecutive video frames of a person standing match the pose description text "Standing in a neutral pose", the sample pair is a positive sample pair, and the true similarity label of the positive sample pair is 1. If the consecutive video frames of the sample pair do not match the pose description text, the sample pair is a negative sample pair, and the true similarity label of the negative sample pair is 0.

[0044] Step S2: Iteratively train the text feature extraction network and the multi-modal human pose estimation network in batches based on the sample pair set. When the iterative training stop condition is reached, use the trained multi-modal human pose estimation network as the multi-modal human pose estimation model.

[0045] In this embodiment, the sample pair set is divided into a training sample pair set, a test sample pair set, and a validation sample pair set. Use the training sample pair set to iteratively train the text feature extraction network and the multi-modal human pose estimation network in batches, and then use the test sample pair set and the validation sample pair set to test and validate the trained multi-modal human pose estimation network respectively. After passing the test and validation, use the trained multi-modal human pose estimation network as the multi-modal human pose estimation model. Otherwise, continue to iteratively train the text feature extraction network and the multi-modal human pose estimation network in batches using the training sample pair set. The iterative training stop condition is preferably but not limited to that the number of iterative training reaches the preset maximum training times, or the contrast loss and the joint position error converge.

[0046] Please see Figure 2 , in each batch of iterative training, execute:

[0047] Step S21: Input the consecutive video frames of the sample pair into the multi-modal human pose estimation network, and input the pose description text of the sample pair into the text feature extraction network.

[0048] Step S22: Perform contrast learning on the global pose feature obtained by the shallow spatio-temporal Transformer cascade network and the global text feature obtained by the text feature extraction network to obtain a contrast loss; calculate the joint position error according to the three-dimensional estimated human pose output by the pose output layer.

[0049] Step S23: Optimize the network parameters of the shallow spatio-temporal Transformer cascade network and the text feature extraction network based on the contrast loss; optimize the network parameters of the visual projection layer, the deep spatio-temporal Transformer cascade network, and the pose output layer based on the joint position error.

[0050] In this embodiment, the three-dimensional estimated human pose output by the pose output layer is in the specific form of a sequence of estimated three-dimensional coordinates of joint points, and each frame of image has N three-dimensional coordinates of joint points. In step S22, the joint position error includes the mean per-joint position error (MPJPE) and / or the reconstructed joint position error (PA-MPJPE). The mean per-joint position error (MPJPE) refers to the average Euclidean distance between the estimated three-dimensional coordinates and the true three-dimensional coordinates of each joint, and can be expressed as:

[0051]

[0052] and p i respectively represent the true three-dimensional coordinates and the estimated three-dimensional coordinates of joint point i.

[0053] The reconstructed joint position error (PA-MPJPE) is an improved MPJPE. By performing Procrustes alignment on the predicted three-dimensional coordinates of joint points and the true three-dimensional coordinates of joint points, the global translation, rotation, and scaling factors in the pose are eliminated, and the joint positions can be compared more accurately. The calculation formula is:

[0054]

[0055] The unit of the MPJPE index is millimeters (mm). The smaller the value of this index, the better the performance of the model and the more accurate the prediction result. Num represents the number of joint points in the sequence of estimated three-dimensional coordinates of joint points. and p i ' respectively represent the true three-dimensional coordinates and the estimated three-dimensional coordinates of joint point i after Procrustes alignment.

[0056] Multi-modal Alignment Prediction is a mechanism that aligns data of different modalities (such as global text features and global pose features) in a shared feature space through contrastive learning. Its main goal is to ensure semantic consistency between different pose text descriptions and pose features, thereby helping the model understand and fuse data of different modalities and improving the performance of the task. Therefore, it is very important to achieve the alignment of global text features and global pose features.

[0057] In a preferred embodiment, the shallow spatio-temporal Transformer cascade network and the text feature extraction network are synchronously optimized by means of contrastive learning to enhance the alignment effect. In step S22, contrastive learning is performed on the global pose features obtained by the shallow spatio-temporal Transformer cascade network and the global text features obtained by the text feature extraction network to obtain a contrastive loss, including:

[0058] Step S2201: Map the global pose feature and the global text feature to a shared feature space. Specifically, the global pose feature and the global text feature are respectively mapped to the shared feature space through a linear layer.

[0059] Step S2202: Calculate the similarity between the mapped global pose feature and the mapped global text feature of the sample pair to obtain the pose-text similarity. Existing Euclidean distance, Manhattan distance, or cosine similarity can be used to characterize the similarity between the global pose feature and the global text feature. Preferably, for sample pair j, its pose-text similarity The calculation formula is:

[0060]

[0061] where p represents the pose and w represents the text; represents the global pose feature obtained by processing the consecutive video frames of sample pair j through the shallow spatio-temporal Transformer cascade network; represents the global text feature of the pose description text (i.e., the positive sample text) in the positive sample pair containing the consecutive video frames of sample pair j in the sample pair set. τ is the temperature parameter. represents the global text feature of the pose description text (i.e., the positive sample text) of the i-th negative sample pair among the K + 1 negative sample pairs containing the consecutive video frames of sample pair j in the sample pair set. K is a positive integer and i ∈ [0, 1, … K]. In this way, the contrast intensity of the positive and negative sample pairs is controlled by adjusting the scale of the similarity.

[0062] Step S2203: Calculate the focal loss weight of the sample pair according to the pose-text similarity of the sample pair. The focal loss weight of the sample pair is negatively correlated with the pose-text similarity. For example, the focal loss weight is the pose-text similarity multiplied by a negative constant. Preferably, the focal loss weight of sample pair j is where, represents the pose-text similarity of sample pair j, γ represents the focal loss adjustment parameter, which is a hyperparameter and adopts a power function form to make its value stable and continuous; j is the sample pair index and is a positive integer.

[0063] Step S2204: Calculate the ratio of the true similarity label of the sample pair to the pose-text similarity, and perform a logarithmic function processing on this ratio to obtain the divergence loss factor. Exemplarily, the divergence loss factor of sample pair j is: where y i represents the true similarity label of sample pair j, which is 0 or 1.

[0064] Step S2205: Multiply the focal loss weight, the true similarity label, and the divergence loss factor of the sample pair to obtain the contrast loss component of the sample pair. Exemplarily, the contrast loss component of sample pair j is:

[0065] Step S2206, accumulate the contrast loss components of the sample pairs in the current batch to obtain the contrast loss. Suppose the current batch has M sample pairs, then the contrast loss is:

[0066]

[0067] In this embodiment, the focal loss weight helps the model to focus more on those negative sample pairs with low similarity and difficult to distinguish. In iterative training, the optimization objective is to minimize the contrast loss, so that the model is trained by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs at the same time. The focal loss weight enables the model to pay more attention to those sample pairs that are difficult to distinguish, promoting the effectiveness of contrast learning; while the divergence loss factor is used to optimize the distribution alignment of the model to ensure the precise alignment of text and pose features in the shared feature space. Finally, through this optimization process, the model can achieve efficient alignment of text and pose features.

[0068] In a preferred embodiment, please refer to Figure 8 and Figure 10 , a Token Pruning Model is arranged between the shallow spatio-temporal Transformer cascade network and the deep spatio-temporal Transformer cascade network, and a Token Recovery Model is arranged between the deep spatio-temporal Transformer cascade network and the pose output layer. The Token Pruning Model measures the importance of time frames, screens video frames with higher semantic representativeness, removes redundant frames, reduces the number of video frames from 243 to 81, reduces the computational overhead while retaining key information, and improves the inference speed. In addition, to compensate for the spatio-temporal information loss caused by frame pruning, this application further designs a Token Recovery Model based on the multi-head attention mechanism to ensure that the model can still completely model the spatio-temporal features of human poses while maintaining efficient computation.

[0069] In this embodiment, the architecture of the Token Pruning Model is as Figure 5 shown. The output feature representation of the n1-th first time Transformer unit (where n1 = depth / 4 and depth is the model depth) of the shallow spatio-temporal Transformer cascade network is: Among them, N is the number of joints, F is the number of time frames, and dim is the dimension of the embedded feature. To reduce computational redundancy and retain key pose information, the Token pruning module adopts a Token selection strategy based on KNN density peak clustering to screen and prune the time frames.

[0070] First, perform average spatial pooling on the joint features of each frame to aggregate all joint features of each frame into an overall representation, reduce the data dimension, and retain the global pose information to obtain the pooled features:

[0071] Next, to screen representative time frames, use the KNN-based effective density peak clustering algorithm to cluster the pooled time frame features. By calculating the local density ρ i and the density peak distance δ i , further calculate the clustering center score of each feature to evaluate its possibility as a clustering center. The local density ρ i measures the density of the feature in the feature space, and the calculation formula is as follows:

[0072]

[0073] where kNN(x i ) represents the k nearest neighbors of the time frame x i , represents the Euclidean distance between the time frames x i and x j . The density peak distance δ i reflects the minimum Euclidean distance between the current time frame and the nearest time frame with a higher density ρ i , and the calculation formula is as follows:

[0074]

[0075] Finally, the importance score of each time frame is defined as:

[0076] score i = ρ i * δ i (1.5)

[0077] The higher the clustering center score of the time frame, the better the key features can cover the overall feature space. Therefore, according to the sorting of the clustering center scores, select the top F' (F' = 81) time frames as the clustering centers, discard the unselected features, and reduce redundant calculations. Finally, the pruned features are (where F' = 81, indicating the time frames retained after pruning).

[0078] In this embodiment, the Token Recovery Module architecture is as follows Figure 6 shown. This module aims to reconstruct the time frame information lost due to Token pruning, ensuring that the model can retain a complete spatio-temporal feature representation and improving the integrity and consistency of temporal information.

[0079] Initialize a zero-valued query vector for each time frame, with its shape being Q ∈ R N*F*dim . The query vector serves as the query input (Query) of the recovery module and is used to fill in the time frame information lost during the pruning process in the multi-head attention calculation. Its size is consistent with the features before pruning to ensure that the recovery process can operate in the same dimension as the original feature space. Take the output features of the last second temporal Transformer unit in the deep spatio-temporal Transformer cascade network as the input of the Token Recovery Module, where F' is the number of time frames retained after pruning. The output features are used as the key (Key) and value (Value), which are used to guide the reconstruction of the lost time frames in the multi-head attention mechanism. To effectively recover the spatio-temporal information lost due to Token pruning, through the multi-head attention module (MCA), according to the correlation between the query (Q) and the key-value pair (K, V), the missing time frame features are reconstructed to complete the spatio-temporal representation of the original video sequence. Its calculation formula is:

[0080]

[0081] The output obtained by the MCA calculation can be regarded as the repaired time frame representation, which can effectively complete the spatio-temporal information missing caused by Token pruning. To further alleviate the information loss and ensure that the recovered time frame features still maintain the spatial distribution consistency with the original features, a residual connection is used to add the recovered time frame representation to the original query vector to enhance the gradient flow and maintain the information integrity. Finally, the recovered time frame sequence is obtained:

[0082]

[0083] where x' is the query Token initialized to zero; is the output of the last second temporal Transformer unit, which is used as the key and value.

[0084] It should be noted that the Token pruning module is inserted after the n1-th first-time Transformer unit of the model, where n1 = depth / 4 and depth is the total depth of the model. That is, the Token pruning module is inserted into the middle part of the model to ensure that after the shallow spatio-temporal Transformer cascade network in the first half fully models the initial global spatio-temporal features, Token reduction is carried out, thereby reducing redundant calculations and improving computational efficiency while retaining key information. In addition, the pruned Tokens can still further model the temporal dependencies in the remaining second spatio-temporal Transformer module to enhance the model's understanding ability of long time series. The Token restoration module is inserted after the last second-time Transformer unit of the model, that is, before the pose output layer, to restore the extracted deep temporal features. This enables the Token restoration module to fully utilize the pruned time frame information and combine the global context after the entire Transformer calculation is completed, fill in the lost spatio-temporal information, reduce unnecessary waste of computational resources, and ensure the integrity of the finally output time frame features.

[0085] The Token pruning module and the Token restoration module proposed in this embodiment can adaptively select key time frames, reduce redundant calculations, and effectively restore the pruned spatio-temporal features. While maintaining the high-precision 3D pose estimation ability, it greatly improves the model inference efficiency and provides a better solution for application scenarios with limited computational resources (such as mobile devices, real-time pose estimation).

[0086] In a preferred embodiment, as Figure 8 and Figure 10 , in the deep spatio-temporal Transformer cascade network, a feature fusion module is further provided after at least one second spatio-temporal Transformer module.

[0087] Please refer to Figure 7 , the feature fusion module includes:

[0088] A spatial feature normalization layer that normalizes the output features of the second spatial Transformer unit of the second spatio-temporal Transformer module adjacent to the feature fusion module to obtain spatially normalized features.

[0089] A temporal feature normalization layer that normalizes the output features of the second temporal Transformer unit of the second spatio-temporal Transformer module adjacent to the feature fusion module to obtain temporally normalized features.

[0090] The spatial features and temporal features are respectively standardized through a spatial feature normalization layer and a temporal feature normalization layer, so that their means are zeroed and variances are normalized, ensuring that the features are within the same distribution range and improving the stability of the fusion process.

[0091] The cross-attention unit is used to perform cross-attention processing on the spatially normalized features and temporally normalized features to obtain the fused feature fused output , and information interaction is carried out through the cross-attention mechanism:

[0092]

[0093] Among them, q s , k s , v s , q t , v, k t respectively represent the query, key, and value of the spatial features and temporal features. out s→t represents the weighted representation of the spatial features on the temporal features, and out t→s represents the weighted representation of the temporal features on the spatial features. The two are added to obtain the final fused feature:

[0094] fused output = out s→t + out t→s (1.10)

[0095] The residual unit performs a residual connection between the output features of the second spatial Transformer unit of the previous second spatio-temporal Transformer module adjacent to the feature fusion module and the fused feature to obtain the enhanced feature.

[0096] The multi-layer perceptron unit is used to process the enhanced feature. To enhance the representation ability of the feature while retaining the original spatial and temporal information, the residual unit uses residual connection for feature enhancement. Finally, the multi-layer perceptron unit (Multilayer Perceptron, MLP) further extracts high-level features from the enhanced feature to improve the learning ability and spatio-temporal modeling ability of the model, thereby optimizing the prediction effect of the 3D pose estimation task.

[0097] This embodiment proposes a feature fusion module based on the cross-attention mechanism, which effectively combines spatial information and temporal information, not only improving the modeling ability of the 3D pose estimation model for complex spatio-temporal relationships, but also achieving a good balance between computational efficiency and prediction accuracy, providing a feasible solution for efficient and accurate 3D pose estimation.

[0098] In a preferred embodiment, see Figure 9, the text feature extraction network includes a text projection layer, a word position encoding embedding unit, and multiple Transformer modules connected in sequence.

[0099] In this embodiment, the text projection layer performs Tokenization on the pose description text, splitting the pose description text into multiple sub-words, words, or characters for further processing. In one example, for the pose description text "Standing in a neutral pose", the following words are obtained after Tokenization: ["standing", "in", "a", "neutral", "pose"]. Tokenization is an important step in natural language processing (NLP), which assigns an identifier to each word or sub-word, enabling the computer to understand and process the text. The text projection layer also converts each Token into a high-dimensional word vector.

[0100] In this embodiment, the word position encoding embedding unit is used to embed the position encoding of the word in the high-dimensional word vector, providing different encodings for the index of each word, enabling the model to understand the order information of the words, and thus performing temporal modeling on the text.

[0101] In this embodiment, after being processed by the word position encoding embedding unit, the text vector is input into the multiple Transformer modules. In this module, three consecutive Transformer modules are used to process the input text vector. By calculating the influence weights of each Token on other Tokens through the self-attention mechanism, the model can adaptively adjust the importance of each Token in the context, helping the model effectively understand the relationships between different Tokens in the pose description text, especially long-distance dependencies. The multiple Transformer structure helps to perform layer-by-layer abstraction and information enhancement on the input text. Each layer can capture semantic and context information at different levels, making the representation of the text more and more semantically rich and deep.

[0102] After being processed by the multiple Transformer modules, the model will finally generate the global representation h <cls>< / cls> (global text feature) of the pose description text. This representation aggregates the context features of all Tokens in the text and can capture the semantic information of the entire sentence or text. Each Token (such as a word) in the text sequence has a corresponding embedding vector, denoted as These vectors capture the context relationships and semantic dependencies between the words in the text, and N1 represents the number of words in the text.

[0103] In this embodiment, the text feature extraction network is based on the advanced Transformer architecture, focusing on extracting high-quality semantic representations from action description texts, and can effectively capture the context relationships and semantic dependencies between words, thereby obtaining richer and more accurate text feature representations.

[0104] In the multi-modal human pose estimation network of this application, the global pose features extracted by the shallow spatio-temporal Transformer cascade network are utilized, and combined with the global text features generated by the text feature extraction network. The model maps these two types of features into a shared feature space for alignment. In this shared space, the model calculates the similarity between the global pose features and the global text features, thereby ensuring their consistency at the semantic level. Through this alignment mechanism, the model can effectively capture the internal correlation between the text description and the pose data, and further improve the performance of multi-modal tasks such as 3D pose estimation and action recognition.

[0105] The multi-modal human pose estimation network provided by this application alternately uses spatial Transformers and temporal Transformers to fully integrate spatio-temporal features. The spatial Transformer focuses on modeling the spatial correlations between joints and extracting the structural information of human poses in a single frame; the temporal Transformer is responsible for capturing the temporal dynamic changes in the action sequence, modeling the cross-frame motion trajectories, and enhancing the ability to understand action patterns. To further optimize the model efficiency and enhance the feature expression ability, a Token pruning module is inserted after the n1-th spatio-temporal Transformer module (where n1 = depth / 4) to dynamically screen key spatio-temporal features, reduce redundant data, and improve the computational efficiency. A Token recovery module is inserted after the last spatio-temporal Transformer module to reconstruct the pruned Tokens and restore the complete spatio-temporal feature representation, ensuring the integrity and accuracy of pose estimation. To further strengthen the synergistic effect of spatial and temporal information, a feature fusion module is added after the spatio-temporal Transformer module after Token pruning to alleviate the disconnection problem between spatial and temporal features and further improve the model's ability to model complex spatio-temporal relationships.

[0106] The present invention also discloses a multi-modal human pose estimation method based on spatio-temporal Transformers, including: obtaining continuous video frames to be estimated, inputting the continuous video frames to be estimated into a multi-modal human pose estimation model, and the multi-modal human pose estimation model outputs the three-dimensional estimated human pose, where the multi-modal human pose estimation model is trained according to the above-mentioned training method of the multi-modal human pose estimation model based on spatio-temporal Transformers.

[0107] The present invention also discloses a computer program product, including a computer program which, when executed by a processor, implements the steps of the above-mentioned chest radiograph lesion detection method based on enhanced feature extraction and fusion or the multi-modal human pose estimation method based on spatio-temporal Transformer provided by the present invention. The computer program product should be understood as a software product that mainly implements its solution through a computer program, such as a program product integrated in the cloud or a software library.

[0108] The present invention also discloses an electronic device. In one embodiment, the electronic device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein,

[0109] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the chest radiograph lesion detection method based on enhanced feature extraction and fusion or the multi-modal human pose estimation method based on spatio-temporal Transformer provided by the present invention.

[0110] As Figure 11 shown, it is a schematic structural diagram of an electronic device for the chest radiograph lesion detection method based on enhanced feature extraction and fusion provided by an embodiment of the present invention. The electronic device may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a program for the chest radiograph lesion detection method based on enhanced feature extraction and fusion or a program for the multi-modal human pose estimation method based on spatio-temporal Transformer.

[0111] Among them, the processor 10 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as executing the chest radiograph lesion detection method based on enhanced feature extraction and fusion or the multi-modal human pose estimation method based on spatio-temporal Transformer, etc.), and calling the data stored in the memory 11, to execute various functions of the electronic device and process data.

[0112] The memory 11 includes at least one type of readable storage medium, which includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device, such as the mobile hard disk of the electronic device. In other embodiments, the memory 11 can also be an external storage device of the electronic device, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device. Further, the memory 11 can also include both the internal storage unit and the external storage device of the electronic device. The memory 11 can be used not only to store application software installed on the electronic device and various types of data, such as the code of the chest radiograph lesion detection method based on enhanced feature extraction and fusion or the program of the multi-modal human pose estimation method based on spatio-temporal Transformer, etc., but also to temporarily store the data that has been output or will be output.

[0113] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to realize the connection and communication between the memory 11 and at least one processor 10, etc.

[0114] The communication interface 13 is used for the communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between this electronic device and other electronic devices. The user interface can be a display, an input unit (such as a keyboard), and optionally, the user interface can also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display can also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device and to display the visual user interface.

[0115] Figure 11 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 11The structures shown do not constitute a limitation on the electronic device, and may include fewer or more components than those shown, or combine certain components, or have different component arrangements.

[0116] It should be understood that the embodiments are for illustrative purposes only and are not limited by this structure in the scope of the patent application.

[0117] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A training method for a multi-modal human pose estimation model based on spatio-temporal Transformer, characterized in that Including: Construct a text feature extraction network and a multi-modal human pose estimation network. The multi-modal human pose estimation network includes a 2D skeleton joint extraction network, a visual projection layer, a shallow spatio-temporal Transformer cascade network, a deep spatio-temporal Transformer cascade network, and a pose output layer connected in sequence; Obtain a set of sample pairs, where each sample pair includes consecutive video frames and pose description texts; Iteratively train the text feature extraction network and the multi-modal human pose estimation network in batches based on the set of sample pairs. When the iterative training stop condition is reached, use the trained multi-modal human pose estimation network as the multi-modal human pose estimation model; In each batch of iterative training, input the consecutive video frames of the sample pair into the multi-modal human pose estimation network, and input the pose description text of the sample pair into the text feature extraction network; perform contrastive learning on the global pose features obtained by the shallow spatio-temporal Transformer cascade network and the global text features obtained by the text feature extraction network to obtain a contrastive loss; calculate the joint position error based on the three-dimensional estimated human pose output by the pose output layer; optimize the network parameters of the shallow spatio-temporal Transformer cascade network and the text feature extraction network based on the contrastive loss; optimize the network parameters of the visual projection layer, the deep spatio-temporal Transformer cascade network, and the pose output layer based on the joint position error.

2. The training method of the multi-modal human pose estimation model based on spatio-temporal Transformer according to claim 1, wherein, The performing contrastive learning on the global pose features obtained by the shallow spatio-temporal Transformer cascade network and the global text features obtained by the text feature extraction network to obtain a contrastive loss includes: Map the global pose features and the global text features to a shared feature space; Calculate the similarity between the mapped global pose features and the mapped global text features of the sample pair to obtain a pose-text similarity; Calculate the focal loss weight of the sample pair according to the pose-text similarity of the sample pair. The focal loss weight of the sample pair is negatively correlated with the pose-text similarity; Calculate the ratio of the true similarity label of the sample pair to the pose-text similarity, and perform a logarithmic function process on the ratio to obtain a divergence loss factor; Multiply the focal loss weight, the true similarity label, and the divergence loss factor of the sample pair to obtain the contrastive loss component of the sample pair; Accumulate the contrastive loss components of the sample pairs in the current batch to obtain a contrastive loss.

3. The training method of the multi-modal human pose estimation model based on spatio-temporal Transformer according to claim 2, wherein The focal loss weight of sample pair j is where represents the pose text similarity of sample pair j, γ represents the focal loss adjustment parameter; j is the sample pair index, which is a positive integer.

4. The training method of the multi-modal human pose estimation model based on spatio-temporal Transformer according to any one of claims 1-3, characterized in that, The shallow spatio-temporal Transformer cascade network includes two or more cascaded first spatio-temporal Transformer modules. The first spatio-temporal Transformer module includes a first spatial Transformer unit and a first temporal Transformer unit connected in sequence; A position embedding unit is arranged before the first spatial Transformer unit of the first first spatio-temporal Transformer module. The position embedding unit is used to embed a learnable spatial position embedding matrix in the joint embedding features output by the visual projection layer. A corresponding Token is set for each joint of each frame in the spatial position embedding matrix; Before the first - time Transformer unit of the Transformer module at the first first - blank, a time - embedding unit is provided. The time - embedding unit is used to embed a learnable time - position embedding matrix in the output features of the first - space Transformer unit of the Transformer module at the first first - blank. In the time - position embedding matrix, a time - position encoding is set for each frame.

5. The training method of the multi-modal human pose estimation model based on spatio-temporal Transformer according to claim 4, wherein, A Token pruning module is provided between the shallow spatio - temporal Transformer cascade network and the deep spatio - temporal Transformer cascade network, and a Token recovery module is provided between the deep spatio - temporal Transformer cascade network and the pose output layer.

6. The training method of the multi-modal human pose estimation model based on spatio-temporal Transformer according to claim 5, wherein, The deep spatio - temporal Transformer cascade network includes two or more cascaded second spatio - temporal Transformer modules. The second spatio - temporal Transformer module includes a second - space Transformer unit and a second - time Transformer unit connected in sequence. A feature fusion module is further provided after at least one second spatio - temporal Transformer module; The feature fusion module includes: A spatial feature normalization layer that normalizes the output features of the second - space Transformer unit of the previous second spatio - temporal Transformer module adjacent to the feature fusion module to obtain spatially - normalized features; A temporal feature normalization layer that normalizes the output features of the second - time Transformer unit of the previous second spatio - temporal Transformer module adjacent to the feature fusion module to obtain temporally - normalized features; A cross - attention unit that performs cross - attention processing on the spatially - normalized features and the temporally - normalized features to obtain fused features; A residual unit that performs a residual connection between the output features of the second - space Transformer unit of the previous second spatio - temporal Transformer module adjacent to the feature fusion module and the fused features to obtain enhanced features; A multi - layer perceptron unit that processes the enhanced features.

7. The training method of the multi-modal human pose estimation model based on spatio-temporal Transformer according to claim 1 or 2 or 3 or 5 or 6, characterized in that, The text feature extraction network includes a text projection layer, a word - position encoding embedding unit, and a multi - layer Transformer module connected in sequence.

8. A multi-modal human pose estimation method based on spatio-temporal Transformer, characterized in that, It includes: Obtain consecutive video frames to be estimated, and input the consecutive video frames to be estimated into the multi - modal human pose estimation model. The multi - modal human pose estimation model outputs a three - dimensional estimated human pose; The multi - modal human pose estimation model is trained according to the training method of the multi - modal human pose estimation model based on spatio - temporal Transformer as described in any one of claims 1 - 7.

9. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 - 8.

10. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 8.