Gesture feature-guided fine-grained character interaction relationship detection method
Through the pose feature guidance method, image features are extracted and fused, redundant interaction pairs are filtered, and the interaction relationship between people and objects is identified using the pose structure decoder, solving the accuracy problem of interaction relationship detection in complex scenes, and achieving efficient identification and filtering of fine-grained interactions.
Patent Information
- Application Number
- CN202510009531.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to realize accurate character interaction relationship detection in complex scenarios, mainly due to the ambiguity of information and the complexity of the scene, which makes it difficult to deal with subtle differences between similar relationships and non-interaction interference.
A fine-grained character interaction relationship detection method guided by posture feature is adopted. By extracting image features, fusing human posture and visual features, filtering redundant interaction pairs, and inputting them to the attitude structure decoder module for processing to identify the interaction relationship between people and objects in the image.
The model's ability to identify fine-grained interactions in complex scenarios is improved, and human poses and feature information is fully utilized, interaction details are captured, the model's ability to filter error interactions is improved, and the model's generalization ability is improved.
Smart Images

Figure CN119942074A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a posture feature-guided fine-grained human interaction relationship detection method. Background Art
[0002] The automation of human-object interaction (HOI) detection is an important goal in the field of computer vision. As one of the basic tasks of scene understanding, human-object interaction relationship detection can automatically identify the interaction relationship between people and objects in the image, avoiding the consumption of huge manpower and material resources. When specific behaviors or important events appear in complex scenes, human-object interaction information is automatically used as a key clue, linked to system status information, and sent to the control center. Therefore, HOI detection is particularly important in the fields of complex scene understanding, intelligent robot interaction, video surveillance, and autonomous driving technology.
[0003] At present, human interaction relationship detection has made significant progress in the field of computer vision. However, it is still challenging to achieve accurate HOI detection in complex real-world scenes. This is mainly due to the ambiguity of information and the complexity of the scene, which makes HOI detection face the challenges of subtle differences between similar relationships and interference from non-interactive pairs, making it difficult to learn interactive relationships.
[0004] Traditional methods for detecting human interaction relationships usually use multi-branch streams to predict interactions separately, relying mainly on the visual features and position information of objects. These methods lack understanding of context and have difficulty distinguishing fine-grained interactions. In addition, although posture information can provide more fine-grained interaction details, traditional models cannot effectively capture and utilize this information, resulting in low recognition accuracy between similar interaction relationships. Summary of the invention
[0005] The embodiment of the present invention provides a posture feature-guided fine-grained human interaction relationship detection method, which is used to solve the problems existing in the prior art.
[0006] In order to achieve the above object, the present invention adopts the following technical scheme.
[0007] A posture feature-guided fine-grained human interaction relationship detection method, comprising:
[0008] S1 extracts features of the image to be identified, performs target classification operations on the extracted features, and obtains the bounding boxes and corresponding categories of people and objects in the image to be identified;
[0009] S2 pairs the people and objects in the extracted features to obtain human-object interaction pairs; and performs posture feature and visual feature fusion operations on the obtained human-object interaction pairs;
[0010] S3 performs a filtering operation on redundant human-object interaction pairs on the execution result of step S2 to obtain the top K human-object interaction pairs with higher interaction possibilities;
[0011] S4 inputs the first K human-object interaction pairs with higher interaction possibilities into the posture structure decoder module for processing to obtain the category detection results of human-object interaction;
[0012] The obtained category detection results of human-object interaction are used to identify the interaction relationship between people and objects in the image.
[0013] Preferably, extracting the features of the image to be identified is implemented using a pre-trained backbone network, the backbone network comprising a 7×7 convolution layer, a 3×3 pooling layer, and a plurality of residual groups, the plurality of residual groups being arranged in sequence along the direction of the data flow;
[0014] The first residual group includes 3 residual blocks, each of which includes a convolution kernel of 1×1×64, a convolution kernel of 3×3×64, and a convolution kernel of 1×1×256; the second residual group includes 4 residual blocks, each of which includes a convolution kernel of 1×1×128, a convolution kernel of 3×3×128, and a convolution kernel of 1×1×512; the third residual group includes 23 residual blocks, each of which includes a convolution kernel of 1×1×256, a convolution kernel of 3×3×256, and a convolution kernel of 1×1×1024; the fourth residual group includes 3 residual blocks, each of which includes a convolution kernel of 1×1×512, a convolution kernel of 3×3×512, and a convolution kernel of 1×1×2048.
[0015] Preferably, step S1 is performed using a DETR target detector, and the execution process includes:
[0016] Use CNN to extract features from the image to be identified to obtain high-dimensional features;
[0017] The obtained high-dimensional features are flattened into a sequence form and input into the Transformer encoder. Through the multi-head self-attention mechanism, the high-dimensional features are globally modeled to obtain the long-distance dependency relationship between the various parts in the image to be identified;
[0018] A set of learnable target queries is received through the Transformer encoder, and the target queries interact with the features output by the Transformer encoder to decode and obtain the category and location of the target;
[0019] Pass-through
[0020] p=(c x ,c y ,W x ,Hy ,c)
[0021] Calculate the bounding boxes and corresponding categories of people and objects in the image to be identified; where c x ,c y Represents the coordinates of the center point of the bounding box, W x ,H y represents the width and height of the bounding box, and c represents the category of the instance;
[0022] Through the Hungarian matching algorithm, the bounding boxes of people and objects in the image to be identified and the corresponding categories are associated with the real targets.
[0023] Preferably, step S2 comprises:
[0024] Calculate the human posture attention map distribution through the graph attention network;
[0025] Pass-through
[0026] Q all =RELU(FC2(RELU(FC1([f h ||f o ||f p ||f s ||H (2) ]))))
[0027] The human visual feature f h , object visual features f o , spatial features f s , the position data of the pose node f p , posture feature H (2) The attention score Q is obtained by fusing the linear layer and the convolution operation. all ; In the formula, FC1 and FC2 refer to two fully connected layers, where FC1 receives the concatenated multiple features as input and performs linear transformation, and FC2 performs a second linear transformation on the output of FC1. Both fully connected layers cooperate with the RELU activation function.
[0028] Preferably, step S2 further includes:
[0029] Constructing graph G based on human posture p =(V,E); where V is the pose node set and E is the edge set; each pose node v in the pose node set i ∈V represents a body joint and has a corresponding feature vector h i ;
[0030] Through the feature matrix X∈R N×F Characterize the features of all posture nodes; where N is the number of posture nodes and F is the dimension of each node feature;
[0031] Pass-through
[0032] h′ i =W (k) h i
[0033] Perform linear transformation on the feature vector of the posture node; where W (k) is a learnable full-weight matrix;
[0034] Pass-through
[0035] α ij =softmax j (LeakyReLU(a (k)· [h′ i ∣∣h′ j ]))
[0036] Calculate the attention coefficient α of node i relative to its neighbor node j ij ; In the formula, || represents the connection operation of feature vectors;
[0037] By weighted average calculation
[0038]
[0039] Calculate the new feature vector of each node; where α ij is the attention coefficient, σ is the activation function;
[0040] Use multiple attention heads and pass
[0041]
[0042] Calculate the nodes of all postures; where (∥) means connecting the outputs of multiple attention heads. is the attention coefficient of the kth head, N(i) represents the set of neighbor nodes of node i;
[0043] Pass-through
[0044] H (2) =AGGREGATE(H (1) ,edge index)
[0045] Calculate the posture feature H (2) ; Where AGGREGATE represents the attention aggregation function of the k-th head, and edge_index is the edge index of the pose node.
[0046] Preferably, step S3 comprises:
[0047] Pass-through
[0048] s ij =σ(W2(ReLU(W1Q all +b1))+b2)
[0049] Calculate the score of the interaction relationship between each pair of characters; where W1 and W2 are two learnable weight matrices used to perform linear transformation on the input features; b1 and b2 are corresponding bias terms used to adjust the flexibility of feature mapping;
[0050] Based on the score of each pair of characters’ interaction relationship, the top K character-character interaction pairs with the highest score are screened.
[0051] Preferably, in step S4, the posture structure decoder module includes 6 layers of stacked decoders, each layer of the decoder includes a self-attention calculation submodule, a feedforward neural network and a posture-guided cross-attention calculation submodule sequentially arranged along the data flow direction;
[0052] The self-attention calculation submodule has a multi-head self-attention mechanism, through the formula
[0053]
[0054] MHSA(Q,K,V)=Concat(head1,…,head h )W O
[0055] MHSA(Q,K,V)=Concat(head1,…,head h )W O
[0056] To achieve; In the formula, W i K , is the weight matrix of different attention heads, W O is the output projection matrix, K T is the transpose of the key-value matrix K, which is multiplied by the query matrix Q to calculate the attention weights; d k is the dimension of the key vector, used to scale the result of the dot product attention;
[0057] The feedforward neural network is used to perform nonlinear transformation and mapping on the output of the self-attention calculation submodule, including two layers of fully connected networks, with an intermediate layer between the two layers of fully connected networks, and the intermediate layer adopts a nonlinear activation function;
[0058] The posture-guided cross-attention calculation submodule includes a convolutional layer and a cross-attention mechanism calculation layer, which is used to associate the target query with the output features of the encoder and fuse the posture information; the processing process of the posture-guided cross-attention calculation submodule includes:
[0059] Determine the spatial positions of background, human body, object and pose nodes in the entire image to be processed, and assign layout labels l to each corresponding component i,j ∈{0,1,2,3,4,5}; where 0 represents the background, 1 represents the human body, 2 represents the object, 3 represents the joint area, 4 is the interaction area, and 5 represents the human posture area;
[0060] Using layout tags i,j ∈{0,1,2,3,4,5}, construct the posture space mask matrix and generate the posture space graph;
[0061] The posture space graph is combined with the features extracted by the backbone network f=(f1,...f n ) and process the features through the perceptron to extract the posture spatial structure knowledge a p =(FC2(ReLU(FC1(f j ‖E lay (l ij )))); where,
[0062] E lay represents the spatial mask matrix embedding that incorporates pose information, a p is the posture space structure knowledge, W q and W k are the weight matrices for query and key, respectively, used to map input features to query space and key space, q i and k j denote the i-th query vector and the j-th key vector, respectively, and d key is the dimension of the key vector, used to normalize the attention scores;
[0063] Through the cross-attention mechanism, the query vector is refined to gradually combine the pose and spatial information;
[0064] The objective function of minimizing the loss is
[0065] FL(p t )=-α t (1-p t ) γ log(p t )
[0066]
[0067] Guide model learning; where p t is the probability that the model predicts the positive class, α tis a factor that balances the weights of positive and negative samples, γ is a factor that adjusts the weights of difficult-to-classify samples, L1 loss is the filtering loss of redundant person pairs, which is used to reduce irrelevant or invalid interaction pairs, L2 loss is the interaction classification loss of the posture space guided decoder, which is used to accurately classify the interaction type; FL is the cross entropy loss function, which is used to calculate the loss between the interaction category predicted by the model and the true label; N represents the number of samples; represents the prediction score of the redundant character pair screening network for the interactivity of the i-th character pair; z i Represents the true value of the i-th sample; is to normalize the predictive scores so that the sum of the prediction scores of all samples is 1; N is the number of samples; C is the number of categories, indicating different interaction types; y ic The true label value of sample i in category c is a binary value, indicating whether the sample belongs to this category; The predicted value of sample i in category c represents the interactive prediction probability of the model; It is to normalize the real labels.
[0068] It can be seen from the technical solutions provided by the above-mentioned embodiments of the present invention that the present invention provides a fine-grained human interaction relationship detection method guided by posture features, including: S1 extracts the features of the image to be identified, performs target classification operations on the extracted features, and obtains the bounding boxes and corresponding categories of people and objects in the image to be identified; S2 pairs the people and objects in the extracted features to obtain human-object interaction pairs; performs posture feature and visual feature fusion operations on the obtained human-object interaction pairs; S3 performs redundant human-object interaction pair filtering operations on the execution results of step S2 to obtain the top K human-object interaction pairs with higher interaction possibilities; S4 inputs the top K human-object interaction pairs with higher interaction possibilities into the posture structure decoder module for processing to obtain the category detection results of human-object interactions. The method provided by the present invention has the following advantages:
[0069] 1. Improved the model’s ability to recognize fine-grained interactions in complex scenes.
[0070] 2. Make full use of human posture and feature information to enhance the capture of interaction details.
[0071] 3. A posture feature guidance mechanism was designed, and posture information was used in multiple modules of the model. The posture is not only used in the decoder setting, but also in the interaction pair screening mechanism, which improves the model's ability to filter erroneous interactions; a simple and effective architecture design is adopted to balance performance and efficiency.
[0072] 4. By increasing the weight of human body information in the decoder, the generalization ability of the model is improved.
[0073] Additional aspects and advantages of the present invention will be given in part in the following description, which will become obvious from the following description, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0075] Figure 1 A processing flow chart of a posture feature-guided fine-grained person interaction relationship detection method provided by the present invention;
[0076] Figure 2 A schematic diagram of the overall architecture of a posture feature-guided fine-grained person interaction relationship detection method provided by the present invention;
[0077] Figure 3 A schematic diagram of a spatial mask of a posture feature-guided fine-grained person interaction relationship detection method provided by the present invention;
[0078] Figure 4 A schematic diagram of an image to be detected in a posture feature-guided fine-grained person interaction relationship detection method provided by the present invention;
[0079] Figure 5 A schematic diagram of a self-attention calculation module of a posture feature-guided fine-grained person interaction relationship detection method provided by the present invention;
[0080] Figure 6 A framework diagram of a posture-guided cross-attention calculation submodule of a posture feature-guided fine-grained person interaction relationship detection method provided by the present invention. DETAILED DESCRIPTION
[0081] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention.
[0082] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The term "and / or" used herein includes any unit and all combinations of one or more associated listed items.
[0083] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless defined as herein.
[0084] To facilitate understanding of the embodiments of the present invention, several specific embodiments will be further explained below with reference to the accompanying drawings, and each embodiment does not constitute a limitation on the embodiments of the present invention.
[0085] The present invention provides a posture feature-guided fine-grained person interaction relationship detection method, which is used to solve the following technical problems existing in the prior art:
[0086] In recent years, Transformer-based models have made significant progress in detecting human interaction relationships. Using Transformer to classify human interaction relationships makes HOI detection more efficient and the detection performance improves rapidly. This is mainly due to the Transformer's powerful feature representation ability and self-attention mechanism, which can capture global context information. However, such methods also face a series of challenges.
[0087] First, most human interaction relationship detection models need to pair people and objects in the image to form all possible human pairs. This method will generate a large number of non-interactive human pairs in practical applications, because not all human pairs have interactive relationships. The existence of these non-interactive human pairs not only increases the computational complexity of the model, but also introduces a large amount of noise data, resulting in reduced prediction efficiency and poor results. The model needs to learn from a large number of negative samples, which places higher requirements on the model's discriminative ability.
[0088] Second, existing Transformer models usually use simple query mechanisms and original decoder designs, which have limited ability to reason about fine-grained interaction relationships. Simple queries may not fully capture the complex and subtle interaction details between humans and objects, such as similar actions but different interaction intentions. The original decoder design lacks deep mining of detailed features, which limits the model's ability to discern fine-grained interactions and makes it difficult to distinguish similar but essentially different interaction behaviors.
[0089] In addition, complex environments also bring great challenges to task interaction relationship detection. Complex scenes contain many factors, such as background clutter, occlusion, the presence of multiple people and multiple objects, etc. These factors will interfere with the model's extraction of human and object features, making the identification of interaction relationships more difficult. Especially in scenes with multiple people and multiple objects, the relationship between people and objects becomes more complex, and the diversity and uncertainty of interaction relationships increase. It is difficult for traditional models to accurately capture these complex interaction patterns.
[0090] In addition, there are some existing interactive methods that detect interactions by capturing contextual semantics and spatial multi-scale features, such as Zhang Yuqi et al.'s "Research on Human Interaction Detection Algorithm Based on Complex Background". However, although the DETR process, HOI detection and fine-grainedness are disclosed in this solution, and human posture information is also used, the posture information is only used for interaction intention branches and cannot filter interaction errors.
[0091] Therefore, how to reduce the dependence on object categories while making full use of human feature information and improving the model's ability to distinguish fine-grained interactive relationships has become an urgent problem to be solved in the current HOI detection field. This requires the introduction of more effective feature fusion and attention mechanisms in model design, fully tapping into human posture and behavior characteristics, and enhancing the model's ability to understand and recognize interactive relationships in complex environments in real life.
[0092] See also Figure 1 The present invention provides a posture feature-guided fine-grained character interaction relationship detection method, comprising the following steps:
[0093] S1 extracts features of the image to be identified, performs target classification operations on the extracted features, and obtains the bounding boxes and corresponding categories of people and objects in the image to be identified;
[0094] S2 pairs the people and objects in the extracted features to obtain human-object interaction pairs; and performs posture feature and visual feature fusion operations on the obtained human-object interaction pairs;
[0095] S3 performs a filtering operation on redundant human-object interaction pairs on the execution result of step S2 to obtain the top K human-object interaction pairs with higher interaction possibilities;
[0096] S4 inputs the first K human-object interaction pairs with higher interaction possibilities into the posture structure decoder module for processing to obtain the category detection results of human-object interactions.
[0097] The obtained human-object interaction category detection results are used to identify the interaction relationship between people and objects in the image, such as the automatic recognition of human actions in video images by a monitoring system.
[0098] In the preferred embodiment provided by the present invention, the specific execution process of the above steps is as follows:
[0099] 1. Use existing target detection to detect people and objects from the image to be detected;
[0100] The image to be detected refers to an image that needs to be detected. In order to fully explain the human interaction detection method provided in this embodiment, an image is used for illustration.
[0101] As a preferred embodiment, Figure 4 The image to be detected is shown.
[0102] Specifically, the image to be detected I∈R H×W×C As the input of the target detector DETR, R represents the matrix dimension, H, W and C represent the height, width and number of channels of the image to be detected respectively. After DETR extraction, the output result p = (c x ,c y ,W x ,H y ,c), where c x ,c y Represents the center point coordinates of the bounding box, W x ,H y represents the width and height of the bounding box, and c represents the category of the instance.
[0103] As a preferred embodiment, DETR is an end-to-end target detection model based on the Transformer architecture. Its overall architecture consists of a superposition of 6 layers of Transformer encoders and decoders. First, the input image is extracted through CNN to obtain a high-dimensional feature representation; then, these features are flattened into a sequence form and input into the Transformer encoder. The encoder uses a multi-head self-attention mechanism to globally model the features and capture the long-distance dependencies between various parts of the image. Next, the Transformer decoder receives a set of learnable object queries and decodes the category and position of the target by interacting with the features output by the encoder. Finally, the model uses the Hungarian matching algorithm to solve the set prediction problem in target detection and realizes the end-to-end target detection process. Specifically, the results extracted by DETR at each stage are associated with the real target through the Hungarian matching algorithm. This process is mainly responsible for the best matching of the predicted results with the real annotations in the target detection task to ensure the efficiency and accuracy of the detection. Through this algorithm process, each target in the detection process (such as the person in the image) is uniquely matched and output, and such processing provides a clear target object for subsequent interactive relationship prediction.
[0104] Specifically, the backbone network (Backbone) of the DETR target detector is a pre-trained ResNet-101.
[0105] As a preferred embodiment, ResNet-101 is a feature extraction backbone network consisting of a 7×7 convolution layer, a 3×3 pooling layer, and four residual groups.
[0106] Among them, the first residual group contains 3 residual blocks, each residual block consists of 3 layers of convolution, namely, 1×1×64 convolution kernel, 3×3×64 convolution kernel and 1×1×256 convolution kernel; the second residual group has 4 residual blocks, each residual block also consists of 3 layers of convolution, namely, 1×1×128 convolution kernel, 3×3×128 convolution kernel and 1×1×512 convolution kernel. Unlike ResNet-50, the third residual group in ResNet-101 has been increased to 23 residual blocks, each of which consists of a 1×1×256 convolution kernel, a 3×3×256 convolution kernel, and a 1×1×1024 convolution kernel; the fourth residual group still has 3 residual blocks, each of which consists of a 1×1×512 convolution kernel, a 3×3×512 convolution kernel, and a 1×1×2048 convolution kernel. The specific composition and parameter settings can be seen in Table 1 below.
[0107] Table 1 ResNet-101 parameter settings
[0108]
[0109] Second, pair the people and objects extracted from the image to generate candidate character pairs. Extract and fuse character postures and visual features.
[0110] Specifically, the human visual features, object visual features, spatial features, etc. are obtained through the above ResNet-101, and then the human posture attention map distribution is calculated through the graph attention network; then the human visual features f h , object visual features f o , spatial features f s , the position data of the pose node f p , posture feature H (2) Etc. are fused through linear layers and convolution operations.
[0111] Q all =RELU(FC2(RELU(FC1([f h ||f o ||f p ||f s ||H (2) ]))))
[0112] In the formula, FC1 and FC2 refer to two fully connected layers, which play the role of feature fusion and mapping in this attention calculation formula. FC1 receives the concatenated multiple features as input and performs linear transformation, while FC2 performs a second linear transformation on the output of FC1. Both fully connected layers cooperate with the RELU activation function to map the input features to the new feature space through the learnable weight matrix, and finally obtain the attention score Q all .Q all It is a fused feature, which is used to characterize the attention score in the embodiment provided by the invention.
[0113] The above-mentioned Graph Attention Network (GAT) is used in the present invention to process human posture data. Each body joint is regarded as a node in the graph, and the connection relationship between nodes is defined according to the human body structure. The Graph Attention Network module generates an output H (2) , contains the attention distribution information of each body part. Then the output is combined with the extracted visual feature f h 、f o , the position data of the pose node f p and the spatial property f s Integration, get Q all The characteristics of person-object pairs (such as Figure 2 For Q allSubsequent operations involve filtering redundant character pairs and calculating the interaction decoder, thereby achieving accurate positioning and identification of character interaction relationships.
[0114] Specifically, we construct a graph G based on human body postures. p =(V,E), where V is a set of pose nodes and E is a set of edges. i ∈V represents a body joint, such as a wrist, elbow, or knee, with a specific feature vector h i , these features can be expressed as the feature matrix X∈R N×F , where N is the number of pose nodes and F is the dimension of each node feature.
[0115] In GAT, the attention mechanism allows each pose node to update its features based on the features of its neighbor nodes and the relationship weights between nodes. First, the features of each pose node undergo a linear transformation by combining them with a learnable weight matrix W. (k) Multiply them to get the transformed eigenvector h' i :
[0116] h′ i =W (k) h i
[0117] Next, calculate the attention coefficient α of node i relative to its neighbor node j ij This involves a shared attention mechanism, implemented by concatenating the feature vectors of nodes i and j and applying the LeakyReLU activation function:
[0118] α ij =softmax j (LeakyReLU(a (k)· [h′ i ∣∣h′ j ]))
[0119] where || represents the concatenation operation of the feature vector. The softmax function ensures that the sum of the attention coefficients of all neighbors of node i is 1.
[0120] Then, the new feature vector of each node is calculated by weighted average of its neighbors’ features, with the weight determined by the attention coefficient α ij is determined and processed by an activation function σ (such as ReLU):
[0121]
[0122] To enhance the expressiveness of the model, in this embodiment, multiple such attention heads are used in parallel, each with its own weight matrix W (k) and attention mechanism a(k) The final pose node representation is obtained by concatenating the outputs of all attention heads.
[0123] In the implementation, two GATConv layers are used. The first GATConv uses 8 attention heads, and its process can be expressed as:
[0124]
[0125] Where (∥) means connecting the outputs of multiple attention heads. is the attention coefficient of the kth head, and N(i) represents the set of neighbor nodes of node i.
[0126] The input of the second GATConv is H (1) , the formula is:
[0127] H (2) =AGGREGATE(H (1) ,edge index)
[0128] Here, AGGREGATE represents the attention aggregation function of the k-th head, and edge_index is the edge index of the pose node.
[0129] Finally, the GAT module generates the output H (2) , contains the attention distribution information of each body part. Then the output is combined with the extracted visual feature f h 、f o , the position data of the pose node f p and the spatial property f s Integration, get Q all The characteristics of the person-object pair. all Subsequent operations involve filtering redundant character pairs and calculating the interaction decoder, thereby achieving accurate positioning and identification of character interaction relationships.
[0130] 3. Use a two-layer MLP network to predict whether there is interaction between character pairs. The top K character pairs with the highest scores are used as the initialization of the query vector for the pose space guide decoder.
[0131] In this embodiment, a network architecture based on a two-layer multi-layer perceptron is used to predict whether there is an interaction relationship between pairs of characters. The input feature vector represents the posture, position information, and other scene context information of each character. Through the two-layer MLP network, the model can learn high-dimensional feature representations to score the interaction relationship of each pair of characters. Specifically, for each pair of characters (i, j), by inputting feature Q all , the model calculates the interaction score s ij, which represents the probability of interaction between character i and character j. The score is expressed by the following formula:
[0132] s ij =σ(W2(ReLU(W1Q all +b1))+b2)
[0133] Where W1 and W2 are two learnable weight matrices used to perform linear transformation on the input features; b1 and b2 are corresponding bias terms used to adjust the flexibility of feature mapping; Q all is the fusion feature obtained by the fully connected layer. The whole formula uses ReLU as the intermediate activation function, and the outermost layer uses the σ (sigma) function for normalization, and finally obtains the score s of the interaction relationship between each pair of characters. ij .
[0134] After obtaining the interaction scores of all person pairs, the top K person pairs with the highest scores are selected and their feature vectors are used as the initialization of the query vectors for guiding the decoder in the pose space. The decoder then uses these initialized query vectors to perform more refined reasoning and decoding of the poses in the entire scene.
[0135] Fourth, the pose space embedding is used in the pose space guided decoder to calculate the cross attention mechanism. With the help of the pose-guided cross attention calculation mechanism, multiple potential interaction relationships between people and objects are further encoded, thereby realizing the detection of human interaction.
[0136] The pose space guided decoder is composed of 6 layers of decoders, each of which includes a self-attention calculation module, a feedforward neural network, and a pose-guided cross-attention calculation module. The query initialization of the decoder comes from the fused feature Q all , including the character’s posture information, scene context information, and interaction features of character pairs. The entire decoding process updates the query representation layer by layer to capture the complex interaction information between characters, and is ultimately used for accurate posture decoding and reasoning.
[0137] 1. Self-attention calculation module
[0138] The first step of each decoder layer is to process the query through the self-attention calculation module. This module uses the Multi-Head Self-Attention mechanism (MHSA) to achieve mutual communication between query vectors and capture the global dependencies within the query. Specifically, for each layer, the query Q, key K, and value V are the same query vector, and the output Q′ is calculated through the multi-head attention mechanism:
[0139]
[0140] In the multi-head self-attention mechanism, multiple attention heads are calculated in parallel and can learn feature representations from different subspaces:
[0141] MHSA(Q,K,V)=Concat(head1,…,head h )W O
[0142]
[0143] In the above three formulas, W i K , is the weight matrix of different attention heads, W O is the output projection matrix, K T is the transpose of the key-value matrix K, which is multiplied by the query matrix Q to calculate the attention weights; d k is the dimension of the key vector, which is used to scale the result of the dot product attention so that the distribution of the attention scores is smoother. This scaling operation is to ensure that the input to the softmax function does not have too much variance when the dimension is large.
[0144] Figure 5 The schematic diagram of the framework structure of the self-attention calculation module is shown. As shown in the figure, the query Q, key K and value V are respectively processed through three parallel linear transformation layers. The three-way calculation results are uniformly processed through the dot product attention calculation layer, and finally processed through concatenation (layer) and linear transformation (layer) in turn.
[0145] 2. Feedforward Neural Network (FFN)
[0146] The query Q′ updated by the self-attention module is passed to the feedforward neural network module for nonlinear transformation and mapping. The feedforward neural network usually consists of two layers of fully connected networks, and the middle layer uses a nonlinear activation function (such as ReLU). The role of this module is to further enhance the expressiveness of the model so that it can learn more complex feature representations. Figure 2 Shown is the placement of the feed-forward neural network in the pose structure decoder.
[0147] 3. Posture-guided cross-attention calculation module
[0148] like Figure 6 As shown in Figure 1, the posture-guided cross-attention calculation module consists of a convolutional layer and a cross-attention mechanism calculation. First, the spatial positions of the background, human body, object, and posture nodes are determined in the entire image, and a layout label l is assigned to each corresponding component. i,j∈{0,1,2,3,4,5}, where 0 represents background, 1 represents human body, 2 represents object, 3 represents joint area, 4 represents interaction area, and 5 represents human posture area. In the cross-attention calculation process, these layout labels are used to construct the posture space mask matrix to generate the posture space graph. The posture space graph is compared with the feature f=(f1,...f n ) and process the features through the perceptron (MLP) to extract the posture spatial structure knowledge a p Subsequently, the query vector is refined under the action of the cross-attention mechanism, gradually combining posture and spatial information to improve the model's accurate recognition of the interaction relationship between people and objects.
[0149] a p =(FC2(ReLU(FC1(f j ‖E lay (l ij ))))
[0150]
[0151] Among them, E lay represents the spatial mask matrix embedding that incorporates the posture information. The posture space graph is combined with the feature f extracted by the backbone network. j Combined with the perceptron (MLP), the features are processed to extract the posture spatial structure knowledge. p .W q and W k are the weight matrices for query and key, respectively, used to map input features to query space and key space; q i and k j denote the i-th query vector and the j-th key vector respectively; d key is the dimension of the key vector, used to normalize the attention scores.
[0152] 5. Use relevant loss functions to guide model learning.
[0153] The objective function of the constructed posture feature-guided fine-grained human interaction relationship detection method is shown in the following formula, with the ultimate goal of minimizing the sum of the Focal Loss of the objective functions L1 and L2:
[0154] L=L1+L2
[0155]
[0156] Specifically, L1 loss is the filtering loss of redundant person pairs, which is mainly used to reduce irrelevant or invalid interaction pairs and ensure that the model only focuses on valid person-object interactions. Where N represents the number of samples; represents the prediction score of the redundant character pair screening network for the interactivity of the i-th character pair; z i Represents the true value of the i-th sample; The predictive scores are normalized so that the sum of the prediction scores of all samples is 1. This is similar to a weighting mechanism, ensuring that the score of each sample is compared relative to all samples.
[0157] L2 loss is the interaction classification loss of the posture space guided decoder, which is used to accurately classify the interaction type and improve the accuracy of interaction recognition through posture space information. Where N represents the number of samples; C is the number of categories, indicating different interaction types; y ic The true label value of sample i in category c is a binary value, indicating whether the sample belongs to this category; The predicted value of sample i in category c represents the interactive prediction probability of the model; It is to normalize the true labels, which is similar to the weighted summation of the labels of each sample in different interaction categories.
[0158] Furthermore, Focal Loss is a loss function suitable for unbalanced data sets, which can effectively deal with the problem of imbalanced ratio of positive and negative samples. Standard cross entropy loss can easily cause the model to be biased towards the negative class when dealing with mostly negative samples and a small number of positive samples, while Focal Loss gives higher weights to samples that are difficult to classify, thereby strengthening the model's learning of these samples. The expression of Focal Loss is:
[0159] FL(p t )=-α t (1-p t ) γ log(p t )
[0160] Among them, p t is the probability that the model predicts the positive class, α t is a factor that balances the weights of positive and negative samples, and γ is a factor that adjusts the weights of difficult-to-classify samples. By setting an appropriate γ, Focal Loss, or FL, is a cross-entropy loss function, which is a variant of the standard cross-entropy loss. In this formula, FL is used to calculate the loss between the interaction category predicted by the model and the true label, which can reduce the impact on easy-to-classify samples and focus on samples that are difficult to distinguish, which helps improve the performance of the model in fine-grained interaction detection. The effectiveness of this method will be verified on the benchmark datasets HICO-Det and V-COCO to ensure its reliability in practical applications.
[0161] In the task of detecting human interaction relationships, two commonly used and important datasets are V-COCO and HICO-DET. The effectiveness of the present invention is verified based on the benchmark data HICO-Det and V-COCO.
[0162] V-COCO is a subset of MS-COCO, containing 10,346 images and 16,199 labeled person instances. Each labeled person has 26 binary action labels. The mean average precision (mAP) is used as the evaluation metric. The images in the dataset are divided into two scenes according to whether there is occlusion, and different scoring criteria are used in these two scenes: Scenes and Scene. Scene If the action prediction is correct and the intersection over union (IoU) of the person exceeds 0.5, the prediction result is considered correct, including occluded or missing characters. No occluded objects are involved.
[0163] HICO-DET is another extended dataset designed for HOI detection, containing a total of 47,776 images, of which 38,118 are used for training and 9,658 are used for testing. Similar to the V-COCO dataset, HICO-DET also covers 80 object categories and 117 action categories, with a total of 600 HOI triplets. HICO-DET calculates the mAP of HOI in two different evaluation settings: the default setting and the known object setting. The known object setting focuses on evaluating detections that only contain object categories in the image, while the default setting evaluates all images, regardless of whether they contain specific object categories. In addition, based on the number of images, the HOI categories are divided into three categories: Full, Rare, and Non-Rare, and these categories are evaluated accordingly.
[0164]
[0165] Table 2 Comparison results of the present invention with other algorithms on the V-COCO dataset.
[0166]
[0167]
[0168] Table 3 Comparison results of the present invention with other algorithms on the HICO-DET dataset.
[0169] In summary, the present invention provides a posture feature-guided fine-grained human interaction relationship detection method, including: S1 extracts the features of the image to be identified, performs target classification operations on the extracted features, and obtains the bounding boxes and corresponding categories of people and objects in the image to be identified; S2 pairs the people and objects in the extracted features to obtain human-object interaction pairs; performs posture feature and visual feature fusion operations on the obtained human-object interaction pairs; S3 performs redundant human-object interaction pair filtering operations on the execution results of step S2 to obtain the top K human-object interaction pairs with higher interaction possibilities; S4 inputs the top K human-object interaction pairs with higher interaction possibilities into the posture structure decoder module for processing to obtain the human-object interaction category detection results. The method provided by the present invention has the following advantages:
[0170] 1. Improved the model’s ability to recognize fine-grained interactions in complex scenes.
[0171] 2. Make full use of human posture and feature information to enhance the capture of interaction details.
[0172] 3. A posture feature guidance mechanism was designed, and posture information was used in multiple modules of the model. The posture is not only used in the decoder setting, but also in the interaction pair screening mechanism, which improves the model's ability to filter erroneous interactions; a simple and effective architecture design is adopted to balance performance and efficiency.
[0173] 4. By increasing the weight of human body information in the decoder, the generalization ability of the model is improved.
[0174] Those skilled in the art can understand that the accompanying drawings are only schematic diagrams of an embodiment, and the modules or processes in the accompanying drawings are not necessarily required to implement the present invention.
[0175] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention or certain parts of the embodiments.
[0176] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0177] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A posture feature-guided fine-grained person interaction relationship detection method, characterized in that: include: S1 extracts features of the image to be identified, performs target classification operations on the extracted features, and obtains the bounding boxes and corresponding categories of people and objects in the image to be identified; S2 pairs the people and objects in the extracted features to obtain human-object interaction pairs; and performs posture feature and visual feature fusion operations on the obtained human-object interaction pairs; S3 performs a filtering operation on redundant human-object interaction pairs on the execution result of step S2 to obtain the top K human-object interaction pairs with higher interaction possibilities; S4 inputs the first K human-object interaction pairs with higher interaction possibilities into a posture structure decoder module for processing to obtain a human-object interaction category detection result; The obtained category detection results of human-object interaction are used to identify the interaction relationship between people and objects in the image.
2. The method according to claim 1, characterized in that The feature extraction of the image to be identified is implemented using a pre-trained backbone network, the backbone network includes a 7×7 convolution layer, a 3×3 pooling layer and a plurality of residual groups, and the plurality of residual groups are arranged in sequence along the data flow direction; The first residual group includes 3 residual blocks, each residual block includes a convolution kernel of 1×1×64, a convolution kernel of 3×3×64, and a convolution kernel of 1×1×256; the second residual group includes 4 residual blocks, each residual block includes a convolution kernel of 1×1×128, a convolution kernel of 3×3×128, and a convolution kernel of 1×1×512; the third residual group includes 23 residual blocks, each residual block includes a convolution kernel of 1×1×256, a convolution kernel of 3×3×256, and a convolution kernel of 1×1×1024; the fourth residual group includes 3 residual blocks, each residual block includes a convolution kernel of 1×1×512, a convolution kernel of 3×3×512, and a convolution kernel of 1×1×2048.
3. The method according to claim 2, characterized in that Step S1 is performed using the DETR target detector, and the execution process includes: Use CNN to extract features from the image to be identified to obtain high-dimensional features; The obtained high-dimensional features are flattened into a sequence form and input into the Transformer encoder. Through the multi-head self-attention mechanism, the high-dimensional features are globally modeled to obtain the long-distance dependency relationship between the various parts in the image to be identified; A set of learnable target queries is received through the Transformer encoder, and the target query interacts with the features output by the Transformer encoder to decode and obtain the category and position of the target; Pass-through p=(c x ,c y ,W x ,H y ,c) Calculate the bounding boxes and corresponding categories of people and objects in the image to be identified; where c x ,c y Represents the coordinates of the center point of the bounding box, W x ,H y represents the width and height of the bounding box, and c represents the category of the instance; Through the Hungarian matching algorithm, the bounding boxes of people and objects in the image to be identified and the corresponding categories are associated with the real targets.
4. The method according to claim 3, characterized in that Step S2 includes: Calculate the human posture attention map distribution through the graph attention network; Pass-through Q all =RELU(FC2(RELU(FC1([f h ||f o ||f p ||f s ||H (2) ])))) The human visual feature f h , object visual features f o , spatial features f s , the position data of the pose node f p , posture feature H (2) The attention score Q is obtained by fusing the linear layer and the convolution operation. all ; In the formula, FC1 and FC2 refer to two fully connected layers, where FC1 receives the concatenated multiple features as input and performs linear transformation, and FC2 performs a second linear transformation on the output of FC1. Both fully connected layers cooperate with the RELU activation function.
5. The method according to claim 4, characterized in that Step S2 also includes: Constructing graph G based on human posture p =(V,E); where V is the pose node set and E is the edge set; each pose node v in the pose node set i ∈V represents a body joint and has a corresponding feature vector h i ; Through the feature matrix X∈R N×F Characterize the features of all posture nodes; where N is the number of posture nodes and F is the dimension of each node feature; Pass-through h′ i =W (k) h i Perform linear transformation on the feature vector of the posture node; where W (k) is a learnable full-weight matrix; Pass-through α ij =softmax j (LeakyReLU(a ( k )· [h′ i ∣∣h′ j ])) Calculate the attention coefficient α of node i relative to its neighbor node j ij ; In the formula, || represents the connection operation of feature vectors; By weighted average calculation Calculate the new feature vector of each node; where α ij is the attention coefficient, σ is the activation function; Use multiple attention heads and pass Calculate the nodes of all postures; where (∥) means connecting the outputs of multiple attention heads. is the attention coefficient of the kth head, N(i) represents the set of neighbor nodes of node i; Pass-through H (2) =AGGREGATE(H (1) ,edge index) Calculate the posture feature H (2) ; Where AGGREGATE represents the attention aggregation function of the k-th head, and edge_index is the edge index of the pose node.
6. The method according to claim 4, characterized in that Step S3 includes: Pass-through s ij =σ(W2(ReLU(W1Q all +b1))+b2) Calculate the score of the interaction relationship between each pair of characters; where W1 and W2 are two learnable weight matrices used to perform linear transformation on the input features; b1 and b2 are corresponding bias terms used to adjust the flexibility of feature mapping; Based on the score of each pair of characters’ interaction relationship, the top K character-character interaction pairs with the highest score are screened.
7. The method according to claim 4, characterized in that In step S4, the posture structure decoder module includes 6 layers of stacked decoders, and each layer of the decoder includes a self-attention calculation submodule, a feedforward neural network, and a posture-guided cross-attention calculation submodule arranged in sequence along the data flow direction; The self-attention calculation submodule has a multi-head self-attention mechanism, through the formula MHSA(Q,K,V)=Concat(head1,…,head h )W O MHSA(Q,K,V)=Concat(head1,…,head h )W O To achieve; In the formula, is the weight matrix of different attention heads, W O is the output projection matrix, K T is the transpose of the key-value matrix K, which is multiplied by the query matrix Q to calculate the attention weight; d k is the dimension of the key vector, used to scale the result of the dot product attention; The feedforward neural network is used to perform nonlinear transformation and mapping on the output of the self-attention calculation submodule, including two layers of fully connected networks, with an intermediate layer between the two layers of fully connected networks, and the intermediate layer adopts a nonlinear activation function; The posture-guided cross-attention calculation submodule includes a convolutional layer and a cross-attention mechanism calculation layer, which is used to associate the target query with the output features of the encoder and fuse the posture information; The processing process of the posture-guided cross-attention calculation submodule includes: Determine the spatial positions of background, human body, object and pose nodes in the entire image to be processed, and assign layout labels l to each corresponding component i,j ∈{0,1,2,3,4,5}; where 0 represents the background, 1 represents the human body, 2 represents the object, 3 represents the joint area, 4 is the interaction area, and 5 represents the human posture area; Using the layout tag l i,j ∈{0,1,2,3,4,5}, construct the posture space mask matrix and generate the posture space graph; The posture space graph is compared with the feature f=(f1,...f n ) and process the features through the perceptron to extract the posture spatial structure knowledge a p =(FC2(ReLU(FC1(f j ‖E lay (l ij )))); where, E lay represents the spatial mask matrix embedding that incorporates pose information, a p is the posture space structure knowledge, W q and W k are the weight matrices for query and key, respectively, used to map input features to query space and key space, q i and k j denote the i-th query vector and the j-th key vector, respectively, and d key is the dimension of the key vector, used to normalize the attention scores; Through the cross-attention mechanism, the query vector is refined to gradually combine the pose and spatial information; The objective function of minimizing the loss is FL(p t )=-a t (1-p t ) γ log(p t ) Guide model learning; where p t is the probability that the model predicts the positive class, α t is a factor that balances the weights of positive and negative samples, γ is a factor that adjusts the weights of difficult-to-classify samples, L1 loss is the filtering loss of redundant person pairs, which is used to reduce irrelevant or invalid interaction pairs, L2 loss is the interaction classification loss of the posture space guided decoder, which is used to accurately classify the interaction type; FL is the cross entropy loss function, which is used to calculate the loss between the interaction category predicted by the model and the true label; N represents the number of samples; represents the prediction score of the redundant character pair screening network for the interactivity of the i-th character pair; z i Represents the true value of the i-th sample; is to normalize the predictive scores so that the sum of the prediction scores of all samples is 1; N is the number of samples; C is the number of categories, indicating different interaction types; y ic The true label value of sample i in category c is a binary value, indicating whether the sample belongs to this category; The predicted value of sample i in category c represents the interactive prediction probability of the model; It is to normalize the real labels.