Character interaction detection method based on multi-scale semantic attention mechanism
By introducing PVT multi-scale fusion module and relative position coding in HOI detection, combined with the interactive semantic attention mechanism, the attention mechanism of the Transformer decoder is improved, and the bottleneck problem of multi-scale feature utilization and long-distance interaction detection is solved, and high-accurate character interaction detection is achieved.
Patent Information
- Application Number
- CN202510392330.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
The existing HOI detection methods have bottlenecks in the use of multi-scale feature and long-distance interaction detection. The local receptive field of CNN limits the global semantic capture ability. The Transformer self-attention mechanism is prone to feature confusion when dealing with extreme scale differences.
A multi-scale fusion module based on PVT is introduced to enhance feature extraction, combining relative position coding and interactive semantic attention mechanism, improve the attention mechanism of Transformer decoder, decode interactive semantic features through self-attention and perform cross-attention fusion.
Multi-scale and high-accurate character interaction detection is realized, which improves the model's semantic understanding ability and resource utilization efficiency, and improves the detection effect of complex scenarios.
Smart Images

Figure CN120259952A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision. Background Art
[0002] Human-Object Interaction Detection (HOI), as a core task in the field of computer vision, has a development process closely related to the evolution of object detection technology. Early object detection focused on single-object recognition and localization, making it difficult to analyze the semantic associations between multiple objects in an image. With the increasing demand for scene understanding, the HOI task emerged. Its core is to construct a <human, verb, object> triple model to reveal the action relationship between humans and objects in an image. This task not only expands the dimension of object detection but also endows the computer with the ability to understand the deep semantics of images, becoming a key technical support in fields such as intelligent monitoring, human-computer interaction, and autonomous driving.
[0003] The core challenge of HOI lies in its cross-modal and multi-dimensional complex characteristics: on the one hand, it is necessary to accurately locate the spatial positions of humans and objects; on the other hand, it is necessary to determine the interaction action category through semantic reasoning. For example, "a person rides a bicycle" and "a person pushes a bicycle" are highly similar in visual features but have significant differences in action semantics. Traditional methods based on handcrafted features (such as Histogram of Oriented Gradients, Local Binary Pattern) rely on manually designed features and are difficult to handle the interaction diversity in complex scenes. After the rise of deep learning, methods based on convolutional neural networks (CNNs) (such as HO-RCNN, ICAN) extract local features through multi-stream networks or attention mechanisms, significantly improving the detection accuracy. For example, HO-RCNN realizes the preliminary prediction of interaction categories through the fusion of human, object, and interaction three-stream features and the modeling of spatial position relationships; ICAN introduces an instance-centered attention mechanism, dynamically focusing on the interaction area and enhancing the fine-grained feature discrimination ability. However, the local receptive field characteristic of CNNs limits its ability to capture global semantics and is vulnerable to background noise interference. With the breakthrough of Transformer technology, current HOI detection methods achieve end-to-end detection through the Transformer encoder-decoder architecture, avoiding the limitations of traditional anchor box designs.
[0004] Although significant progress has been made in Transformer methods, there are still bottlenecks in the utilization of multi-scale features and the detection of long-distance interactions. The current mainstream one-stage and two-stage Transformer methods both face the following challenges: one-stage methods usually adopt single-scale feature extraction, making it difficult to balance the detailed features of small objects and the global semantics of large objects; two-stage methods generate candidate pairs through region proposals, but lack multi-scale context guidance, resulting in low feature fusion efficiency. In addition, although the self-attention mechanism of Transformer can capture global dependencies, when dealing with extreme scale differences, due to the lack of local feature enhancement, feature confusion is likely to occur. To improve the impact of the above problems on the HOI detection task, the present invention proposes a person interaction detection method based on an interactive semantic attention mechanism to solve the above problems. Summary of the Invention
[0005] To solve the above problems, we propose a person interaction detection method based on a multi-scale semantic attention mechanism. By introducing a multi-scale fusion module based on PVT at the backbone feature extraction network of the Transformer model, the detection ability of person interaction is enhanced. At the same time, the attention mechanism is improved in the Transformer decoder part. The interactive semantic features are decoded through self-attention, and the multi-scale granular features are fused through the cross-attention mechanism to decode the person interaction prediction results, realizing a multi-scale and high-accuracy person interaction detection algorithm. In summary, the algorithm steps of the present invention are as follows:
[0006] Step 1: Input an image, and extract multi-level feature maps from the image through an improved backbone feature extraction network.
[0007] Step 2: Fuse the multi-level feature maps through a multi-scale fusion module and superimpose position encoding.
[0008] Step 3: Pass the image through the DETR model to obtain the detection results of humans and objects, perform semantic feature modeling on the results, and extract interactive semantic features.
[0009] Step 4: Fuse the interactive semantic features into the self-attention of the Transformer decoder for decoding.
[0010] Step 5: Fuse the multi-scale features into the Transformer decoder to complete cross-attention calculation, and finally obtain the ternary prediction results of person, object, and action through a perceptron prediction. Description of the Drawings
[0011] Appendix Figure 1 : The overall framework diagram of the network adopted by the present invention
[0012] Appendix Figure 2 : The PVT multi-level feature extraction diagram of the network adopted by the present invention
[0013] Appendix Figure 3 : Flow chart of multi-scale fusion adopted by the present invention
[0014] Appendix Figure 4 : Schematic diagram of interactive semantic features adopted by the present invention
[0015] Appendix Figure 5 : Transformer attention structure diagram adopted by the present invention Detailed implementation manners
[0016] In the present invention, the multi-scale semantic fusion module adopts the Pyramid Vision Transformer (PVT) and the cascaded feature fusion network to achieve the efficient fusion of multi-scale context information. At the same time, this module models the humans and objects detected by DETR, and then extracts the interactive semantic features therefrom. Finally, the present invention integrates these interactive semantic features and multi-scale features into the self-attention and cross-attention mechanisms of the decoder, and obtains the predicted results of HOI through decoding. The specific process is as follows:
[0017] Step 1: Construct multi-level network features
[0018] The present invention improves the backbone feature extraction network into a PVT network capable of extracting multi-scale features. The PVT network includes four stages, and each stage generates feature maps of different scales. We will use these feature maps to construct multi-level features. Each stage extracts multi-scale feature maps through the patch embedding layer and the Transformer encoder. Assuming that the input image size is H×W×C, first, it is divided into patches of size 4×4, and each patch is flattened and linearly projected to obtain an embedding vector. After being processed by the Transformer Encoder of layer L1, the output feature map Then, F1 is input into the network of the second stage to continue this step.
[0019] Generally speaking, for an image The present invention generates four feature maps of different dimensions after the image is inferred through the 4 stages of PVT for subsequent fusion.
[0020] Step 2: Multi-scale feature fusion module
[0021] The 4-scale feature maps output by PVT The sizes are respectively to The multi-scale feature fusion module mentioned in the present invention is shown in Appendix Figure 2As shown, the module starts from the deep feature map which has the lowest resolution and the strongest semantic information. Channel reduction is performed on it using a 1×1 convolution. The specific formula is as follows, where W is the 1×1 convolution kernel and c3 is the number of channels of the previous-level feature map. The operation results in the feature map
[0022]
[0023] After obtaining the feature map of the convolution operation, bilinear interpolation is used for upsampling to match its size with the previous-level feature map F3, resulting in In the subsequent channel fusion part, the upsampled F4″ and the upper-level feature map F3 are concatenated in the channel dimension to achieve feature fusion at different levels.
[0024] X3 = concat([F3, UpSample(F4″)]) (2)
[0025] Through continuous operation and fusion, a feature map that integrates the semantic features of images from 4 different stages is finally obtained Finally, these feature maps are flattened into one-dimensional vectors in the H×W dimension to obtain a set of feature vectors
[0026] Absolute position encoding assigns an independent encoding vector to each position in the sequence, usually generated based on a fixed function, such as the sine and cosine functions. In a sequence of length n and dimension d, the encoding vector PE of position pos pos is calculated as follows:
[0027]
[0028] where i ∈ {0, 1, …, [d / 2]}. Although this method is straightforward, it lacks flexibility. Sequences of different lengths need to be recalculated or the encoding adjusted, which limits the model's ability to process variable-length sequences and makes it difficult to capture the dynamic changes in the relative position relationships between elements.
[0029] The present invention selects to use relative position encoding to be fused with the image vector sequence to improve the network. The traditional attention calculation score A ij is as shown in the following formula:
[0030]
[0031] After introducing relative position encoding, the relative position matrix R is defined n×n , and its element R ij represents the relative relationship between positions i and j. The new attention score is as shown in Equation (6):
[0032]
[0033] The relative position encoding has advantages such as strong adaptability and excellent dynamic capture ability. It can focus on the relative relationships between elements and is more flexible when dealing with sequences of different lengths. At the same time, it can accurately capture the dynamic information of the relative position changes of sequence elements, helping the model reflect the changes in position relationships. The relative position encoding also improves the accuracy of the model's semantic understanding and context awareness ability. In addition, the relative position encoding has computational efficiency. In the processing of long sequences, the absolute position encoding increases the computational amount and memory occupancy as the encoding vector increases with the growth of the sequence. The relative position encoding only considers the relative distance between elements, and the computational amount grows smoothly. Especially when dealing with ultra-long sequences (such as the processing of high-resolution image sequences), it can greatly reduce computational redundancy, optimize resource utilization, and improve the model training and inference efficiency, having certain advantages in resource-constrained scenarios.
[0034] Step 3: Semantic Feature Modeling Module for DETR Detection Results
[0035] In the present invention, for each interaction pair detected by DETR, its features are represented from three aspects: the appearance features, spatial features, and semantic features of the person and the object. The appearance features are directly represented as the concatenation of the instance feature vectors of the person and the object extracted from DETR. After being detected by DETR, the category of the object in the candidate pair can be obtained. The present invention encodes the category label to obtain a semantic feature representation of the candidate pair.
[0036] F = [distance x , distance y , distance, E] (7)
[0037]
[0038] By defining the normalized center coordinates of the person and object bounding boxes as and The present invention proposes a representation form of joint spatial features, and the specific process is as follows:
[0039] As shown in Equation (7), the present invention defines a spatial distance feature between a person and an object, where is calculated through the center point of the human body coordinates and the center point of the object . distance x and distance y are the offsets of the person and the object in the X and Y directions, representing the relative position differences. The distance in the equation is the straight-line distance between the person and the object. These features can clarify the position changes of the person relative to the object in the horizontal and vertical directions and the distance between the two.
[0040] During interaction detection, the relative direction between a person and an object belongs to the implicit spatial semantic information in the image. Since the direction of the person relative to the object is different, the likelihood of interaction also varies. Therefore, the present invention uses a method based on quadrant coordinates for direction encoding. Specifically, according to the signs of distance x and distance y , the plane is divided into four quadrants to complete the direction encoding. The specific rules are shown in Equation (8). This encoding form has a low computational complexity and can efficiently provide direction information for the model.
[0041] After completing the DETR detection, both the person and the object will be assigned corresponding detection categories. Using these category labels, the semantic features of the person-object interaction pair can be constructed. To convert the text into a vector form, the present invention uses an Embedding layer to perform word vector embedding operations on the text labels. The specific process is as follows:
[0042]
[0043] As shown in Equations (9) and (10), after the image is detected by DETR, the probability distribution vector of each label is calculated using softmax. In this vector, p i represents the probability that the target belongs to the i-th class. Subsequently, according to Equation (11), the category index k with the highest probability can be calculated. Based on the index k, a One-Hot encoding vector of length C is constructed. The initial elements of this vector are 0, and the element at the index k is set to 1. The detailed situation is shown in Equation (12).
[0044]
[0045] After obtaining the One-Hot encoding, the next step is to use the word embedding layer to convert these encodings into semantic vector representations. Essentially, the word embedding layer is a weight matrix that can be learned through training. Assume that C represents the size of the set of object category labels, and d represents the dimension of the output embedding vector.
[0046] The specific calculation process of the word embedding layer can be referred to Equation (13). Through this process, the label name of the object is finally successfully encoded into an image semantic feature vector, so that in subsequent analysis and processing, these semantic information can be more effectively utilized to improve the execution effect of related tasks.
[0047] Step Four: Transformer Decoder Self-Attention Module
[0048] In existing person-object interaction detection tasks, the Transformer uses the self-attention mechanism to convert the query vector q, key vector k, and value vector v into an output sequence to complete the set prediction of task interaction. Its basic principle is shown in the following equation:
[0049]
[0050] s.t. Q = XW Q , K = XW K , V = XW V (15)
[0051] where W Q 、W K 、W V are weight matrices, X is the input feature vector, and the Q, K, and V matrices are obtained through projection mapping by the trainable weight matrices. Q i and K j are the i-th and j-th row vectors in the Q and K matrices respectively, d k represents the dimension of the key vector K, and the attention score e ij is obtained after the above calculation.
[0052]
[0053] Then, the attention scores are normalized by the softmax function to obtain the attention weights α ij .
[0054]
[0055] Finally, V is weighted and summed according to the attention weights, and the final output Y = [Y1, Y2, L, Y n is obtained.
[0056] When an ordinary Transformer decoder performs set prediction, it uses a randomly initialized interactive query vector, and this method cannot effectively mine the prior knowledge contained in the human interactions in the image. In view of this, the present invention proposes an interactive semantic self-attention module, which can carry out set prediction work for human interaction detection based on the prior semantic feature knowledge of the image. Based on the deep semantic features extracted by the above modules, the present invention constructs a new self-attention mechanism based on interactive semantics, and its specific calculation formula is as follows:
[0057]
[0058] Specifically, in step three, the interactive semantic feature vector F of the human body and objects detected by DETR is obtained, and the vector of the i-th human pair is F i . We use F i as the query vector, key vector, and value vector of self-attention, and through the projection operation of the Transformer. Thus, the interactive self-attention score of the image human pair is calculated.
[0059] Step Five: Transformer Decoder Cross-Attention Module
[0060] The improved cross-attention mechanism of the present invention adopts a unique information fusion strategy. It sets the output of the self-attention module as the query vector. In this scenario, the self-attention module deeply mines the semantic features of the interaction candidate pairs, and its output contains rich interaction semantic information. At the same time, the key-value vector is derived from the result of the fusion of the image feature map and the position encoding. The specific formula is as follows:
[0061]
[0062] where q i is the result of the self-attention output and serves as the query vector for cross-attention. The pixel sequence obtained by the multi-scale feature fusion network is x i , and in the multi-scale feature extraction module in Step Two, position encoding pos i is added to each pixel vector, and finally the cross-attention score is calculated. Finally, the attention result is input into the fully connected layer for prediction to obtain the final interaction prediction result.
[0063] Although the above describes the illustrative specific embodiments of the present invention for the convenience of those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. Any equivalent substitution or equivalent replacement, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
Claims
1. The present invention proposes a method for human interaction detection based on a multi-scale semantic attention mechanism. The steps of the method are as follows: Step 1: Input an image and extract multi-level feature maps from the image through an improved backbone feature extraction network. Step 2: Fuse the multi-level feature maps through a multi-scale fusion module and superimpose position encoding. Step 3: Pass the image through the DETR model to obtain the detection results of humans and objects, perform semantic feature modeling on the results, and extract interactive semantic features. Step 4: Fuse the interactive semantic features into the self-attention of the Transformer decoder for decoding. Step 5: Fuse the multi-scale features into the Transformer decoder to complete cross-attention calculation, and finally obtain the ternary prediction results of humans, objects, and actions through a perceptron prediction.
2. The method for detecting human interaction based on a multi-scale semantic attention mechanism according to claim 1, characterized in that The improved backbone feature extraction network in Step 1 is specifically the Pyramid Vision Transformer (PVT) network. The PVT network includes four stages, and each stage extracts multi-scale feature maps through a patch embedding layer and a Transformer encoder. The specific implementation process is as follows: 2.1 Input image preprocessing The input image is segmented into patches of size 4×4. After each patch is flattened, it is linearly projected to obtain an embedding vector E∈RN×C, where N is the number of patches and C is the embedding dimension. 2.2 Multi-stage progressive processing Take the feature map F1 output in the first stage as the input and sequentially input it into the networks of the second, third, and fourth stages. Each subsequent stage gradually reduces the patch size and increases the number of channels to gradually reduce the spatial resolution of the feature map and gradually enhance the semantic information. For example, from the first stage to the second stage, the output feature map changes from to And so on to the fourth stage. 2.3 Feature map output After four stages of processing, four feature maps F1, F2, F3, F4 with different dimensions are finally generated. Their spatial sizes are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image respectively, and the number of channels increases sequentially (C1 = 64, C2 = 128, C3 = 256, C4 = 512), which are used for the subsequent multi-scale feature fusion module.
3. The method for detecting human interaction based on a multi-scale semantic attention mechanism according to claim 1, wherein The present invention fuses multi-level feature maps and introduces relative position encoding. The specific steps of the multi-scale fusion module in Step 2 are as follows: 3.1 Deep feature map dimensionality reduction The module starts from the deep feature map which has the lowest resolution and the strongest semantic information. Channel reduction is performed on it using 1×1 convolution. The specific formula is as follows, where W is the 1×1 convolution kernel and c3 is the number of channels of the previous-level feature map, and the resulting feature map is obtained through the operation 3.2 Bilinear interpolation upsampling After obtaining the feature map of the convolution operation, bilinear interpolation is used for upsampling to match its size with the previous-level feature map F3, obtaining X3 = concat([F3, UpSample(F4″)]) 3.3 Multi-stage iterative splicing fusion In the subsequent channel fusion part, the upsampled F4″ is concatenated with the upper-layer feature map F3 in the channel dimension to achieve feature fusion at different levels. Repeat the above steps of dimensionality reduction, upsampling, and concatenation to fuse the feature maps F2 and F1 in turn, and finally generate the feature map F that integrates the semantic features of the four stages. final The fused feature map F final is flattened into a one-dimensional vector in the H×W dimension to obtain the feature vector group v. 3.4 Relative position encoding superposition Replace the traditional absolute position encoding with relative position encoding, define the relative position matrix R, and its element R i,j represents the relative relationship between positions i and j. The improved attention score calculation is as follows: This encoding method dynamically captures the relative position relationship between elements, improving the model's context awareness ability and long sequence processing efficiency.
4. According to the method for human interaction detection based on a multi-scale semantic attention mechanism described in claim 1, it is characterized in that the semantic feature modeling in Step 3 includes the following contents: Define the normalized center coordinates of the bounding boxes of people and objects as and The present invention proposes a representation form of combined spatial features: F = [distance x , distance y , distance, E] Among them Through the center point of the human body coordinates And the center point of the object It is calculated and obtained. distance x And distance y Are the offsets of the human and the object in the X and Y directions, representing the relative position differences. The distance in the formula is the straight-line distance between the human and the object. These features can clarify the position changes of the human relative to the object in the horizontal and vertical directions and the distance between the two. In addition, the present invention performs direction encoding based on quadrant coordinates. Specifically, according to the positive and negative of, the plane is divided into four quadrants to complete direction encoding. This encoding form has a low computational complexity and can efficiently provide direction information for the model. After completing the DETR detection, both humans and objects will be assigned corresponding detection categories. Using these category labels, the semantic features of human interaction pairs can be constructed. In order to convert the text into a vector form, the present invention uses an Embedding layer to perform word vector embedding operations on the text labels. The specific process is as follows: After the image is detected by DETR, the probability distribution vector of each label is calculated using softmax. In this vector, p i represents the probability that the target belongs to the i-th class. Subsequently, the class index k with the highest probability can be calculated. Based on the index k, a One-Hot encoding vector of length C is constructed. The initial elements of this vector are 0, and the element at index k is set to 1. For details, see the equation. After obtaining the One-Hot encoding, the next step is to use the word embedding layer to convert these encodings into semantic vector representations. Essentially, the word embedding layer is a weight matrix that can be learned through training. Suppose C represents the size of the set of object class labels, and d represents the dimension of the output embedding vector. Through this process, the label name of the object is finally successfully encoded into the image semantic feature vector.
5. The method for detecting human interaction based on a multi-scale semantic attention mechanism according to claim 1, characterized in that, The self-attention module of the Transformer decoder in step 4 adopts an interactive semantic self-attention mechanism, and the specific implementation steps are as follows: In step three, the interactive semantic feature vector F of the human body and objects detected by DETR is obtained, where the vector of the i-th pair of characters is F i . We use F i as the query vector, key vector, and value vector of self-attention, and through the projection operation of the Transformer, the interactive self-attention score of the image character pair is calculated. Specifically, it is shown as follows:
6. The method for detecting human interaction based on a multi-scale semantic attention mechanism according to claim 1, wherein, The specific implementation of the cross-attention module of the Transformer decoder in step 5 is as follows: The improved cross-attention mechanism of the present invention adopts a unique information fusion strategy. It sets the output of the self-attention module as the query vector. In this scenario, the self-attention module deeply mines the semantic features of the interaction candidate pairs, and its output contains rich interaction semantic information. At the same time, the key-value vector comes from the result of the fusion of the image feature map and the position encoding. Among them, q i is the result of the self-attention output and serves as the query vector of the cross-attention. The pixel sequence obtained by the multi-scale feature fusion network is x i , and the position encoding pos i is added to each pixel vector in the multi-scale feature extraction module in step two, and finally the cross-attention score is calculated. Finally, the attention result is input into the fully connected layer for prediction to obtain the final interaction prediction result.
Citation Information
Cited By
Spatial perception multi-view anomaly detection method based on meta-view representation
CN121074512A
Dialogue abstract and key intention extraction method, system and equipment based on large model and multi-dimensional acoustic features
CN122157662A