Character interaction detection method based on multi-modal information fusion
By introducing methods of text label semantic enhancement and multimodal feature fusion, the semantic understanding and long-tail distribution of existing character interaction detection methods in complex scenarios is solved, and the detection accuracy and detection effect of rare interaction categories are improved.
Patent Information
- Application Number
- CN202510392332.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
The existing character interaction detection methods have weak semantic understanding ability in complex scenarios, image features are susceptible to occlusion and lighting interference, classifier semantic clues are insufficiently utilized, and long-tail distribution problems are prominent, resulting in unsatisfactory detection results.
A text tag semantic enhancement classifier was introduced, multimodal features were fused into the character interaction detection classification header, and a multimodal classifier was implemented using CLIP, and the results of the two classifiers were finally fused, and feature extraction and prediction were performed through the Transformer encoder-decoder.
It improves the accuracy of character interaction detection, enhances the semantic understanding of complex scenes, and improves the detection performance of rare interaction categories.
Smart Images

Figure CN120259953A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision. Background Art
[0002] Human-Object Interaction Detection (HOI), as a core task in the field of computer vision, aims to reveal the action relationship between humans and objects in images by constructing a <human, verb, object> triple model. This technology has wide application value in fields such as intelligent monitoring, human-computer interaction, and autonomous driving. However, the detection performance of existing methods still faces significant challenges in complex scenarios. With the development of deep learning technology, Transformer-based methods (such as HOTR) have improved detection performance through a global self-attention mechanism.
[0003] Although Transformer methods have made significant progress, there are still the following bottlenecks. For example, the semantic understanding ability is weak. Existing methods mainly rely on image features, which are vulnerable to factors such as occlusion and illumination and are difficult to accurately capture the semantic essence of interaction behaviors. For example, "a person riding a bicycle" and "a person pushing a bicycle" are highly similar in visual features, but the action semantics are significantly different, and existing methods are prone to confusing such scenarios. In addition, existing methods do not make full use of semantic clues of classifiers: HOI labels are composed of combinations of verbs and objects, and there is semantic correlation between labels. However, existing classifiers are mostly based on randomly initialized weights and fail to effectively mine this semantic association, resulting in limited classification performance. The long-tail distribution problem is prominent: the number of samples of rare interaction categories in the dataset is small, and the generalization ability of traditional methods is insufficient, and the detection effect is not ideal. Summary of the Invention
[0004] In view of the above problems, the present invention proposes a method for human-object interaction detection based on multi-modal information fusion. By introducing text label semantics to enhance the classifier, multi-modal features are fused into the human-object interaction detection classification head. At the same time, a multi-modal classifier is implemented based on CLIP. Finally, the results of the two classifiers are fused to obtain a higher-precision human-object interaction detection result. The specific steps of the method are as follows:
[0005] Step 1: Input an image and extract image features through a backbone network.
[0006] Step 2: Perform set prediction through a Transformer encoder-decoder.
[0007] Step 3: Extract text embedding features through a multi-modal label semantic embedding module.
[0008] Step 4: Use the text embedding vector as the initial weight of the linear classification layer for interaction prediction.
[0009] Step 5: Use the multi-modal instance semantic classifier for prediction.
[0010] Step 6: Use the multi-source fusion module to fuse the interactive prediction results. Description of the Drawings
[0011] Appendix Figure 1 : Overall framework diagram of the network adopted by the present invention
[0012] Appendix Figure 2 : Multi-modal label semantic embedding module diagram adopted by the present invention
[0013] Appendix Figure 3 : Structure diagram of the semantic instance classifier adopted by the present invention
[0014] Appendix Figure 4 : Schematic diagram of the multi-source fusion classification method adopted by the present invention Detailed Implementation Manner
[0015] The specific process of a method for detecting human interaction based on multi-modal information fusion proposed by the present invention is as follows:
[0016] Step 1: Extract image features through the backbone network
[0017] In the process of extracting the backbone feature network of the present invention. The input image x ∈ R 3×H×W , where H and W respectively represent the height and width of the image, and 3 is the number of input channels. The convolution kernel is a four-dimensional tensor with a shape of K × K × C in × C out , where K is the convolution kernel size and C out is the number of output channels. The intermediate features are generated through the following convolution operation formula:
[0018]
[0019] After the convolution operation, an activation function is usually applied to introduce non-linear features. A common activation function is ReLU (Rectified Linear Unit), and its formula is as follows:
[0020] Z i,j,c = max(0, Y i,j,c )
[0021] Here, Z i,j,c is the value of the output feature map at position (i, j) and channel c after being processed by the activation function.
[0022] Pooling operations can reduce the spatial size of the feature map and reduce the amount of calculation. Max pooling is a commonly used pooling method, and its formula is as follows:
[0023]
[0024] As shown in the above formula, P i,j,c is the value of the pooled feature map at position (i, j) and channel c, and S is the size of the pooling window
[0025] The CNN contains multiple convolutional layers, activation layers, and pooling layers. By stacking these layers layer by layer, different levels of features can be extracted, gradually transforming the original image data into high-level semantic features. After several layers of processing, the finally obtained feature map will be flattened into a one-dimensional vector as the input of the subsequent Transformer module.
[0026] Step 2: Perform set prediction through the Transformer encoder-decoder
[0027] In this step, the Transformer encoder-decoder architecture is used to extract and process the features of the input image to achieve effective detection of human-object interaction (HOI). The specific process is as follows:
[0028] Input the feature map output by the backbone network into a projection convolutional layer with a convolutional kernel size of 1×1 to reduce the dimension from D b to D c , obtaining
[0029] The Transformer encoder processes the dimension-reduced feature map z c based on the self-attention mechanism to generate a feature map with richer context information. At the same time, a fixed position encoding p is additionally input to the encoder to supplement the position information that the self-attention mechanism itself cannot inherently integrate. Finally, the encoded feature map is obtained, and the calculation formula is z e = f enc (z c , p), where f enc (.,.) represents a set of stacked Transformer encoder layers.
[0030] The Transformer decoder uses the attention mechanism to convert a set of learnable query vectors by referring to the encoded feature map z e into a set of embeddings containing the global context information of the image for HOI detection That is, D = f dec (z e , p, Q), where f dec(.,.,.) is a set of stacked Transformer decoder layers. Here, N q is the number of query vectors and is set large enough to ensure that it is always greater than the number of actual human-object pairs in the image. When designing the query vectors, they are designed to capture at most one human-object pair and the interaction between them.
[0031] Through the collaborative work of the above Transformer encoder-decoder and the interaction detection head, the collective prediction of human-object interactions in the image is achieved, and the prediction of the interaction results will be carried out in the subsequent steps.
[0032] Step 3: Extract text embedding features through the multi-modal label semantic embedding module
[0033] In the human-object interaction detection task, the HOI triple is composed of a verb and an object combination, making them show extremely strong semantic correlation with each other. However, traditional image backbone networks face great difficulties in learning such complex semantic structures. The CLIP model, through joint training on large-scale image-text pairs, has a very powerful semantic representation ability in its language model, and can efficiently encode the structural information of HOI classes. By using language embeddings to initialize the linear classification layer of the detection head part, we can incorporate these semantic clues into the classifier. This can guide the model to learn during the training process, and at the same time provide more appropriate weight initial values for the few-shot classes with relatively few samples in the training data.
[0034] Since the original HOI object labels and interaction actions exist in the form of independent words, to enhance the text semantic expression, we add specific guiding information on the basis of each HOI triple description to associate the two, so as to generate a more task-oriented text description and accurately convey the semantics of HOI instances. The specific operation is to transform the HOI triple, converting the original label to [guiding word][person][action][item]. This transformation method helps the language model better understand and process the HOI class semantics, and meets the input requirements of CLIP pre-trained on natural language text.
[0035] After completing the corresponding transformation, the processed text is sent into the pre-trained CLIP text encoder. The language model of CLIP will generate corresponding language embeddings for these transformed prompts. The CLIP text encoder will perform operations such as word segmentation and feature extraction on the input text in sequence, and finally output a feature vector that can represent the text semantics. In the word segmentation stage, the CLIP text encoder adopts the byte pair encoding algorithm, which continuously merges high-frequency byte pairs according to the occurrence frequency of characters or sub-words in the text to construct a vocabulary, and then cuts the text into appropriate sub-word units based on this vocabulary.
[0036] The token sequence is encoded using a multi - layer Transformer architecture, and finally the hidden state representation of each token is generated. Here, the final hidden state corresponding to the [EOS] token vector in the CLIP model is selected as the feature representation of the entire text, denoted as h text . Specifically, assuming that the output after encoding the input text by CLIP is H, the formula is as follows, where h i represents the hidden state vector of the i - th token, and n is the number of tokens. The embedding vector h text =h [EOS] , that is, the vector corresponding to the [EOS] token is taken out from H and used as the final text embedding.
[0037] H = {h [SOS] , h1, h2, …, h n , h [EOS]}
[0038] Step Four: The text embedding vector is used as the initial weight of the linear classification layer for interaction prediction.
[0039] To make different text embeddings have the same dimension and comparability in subsequent calculations and comparisons, the obtained text embedding vector needs to be normalized to keep it consistent in scale. The normalized embedding vector will be used as the initial weight of the linear classification layer. Assuming that the weight matrix of the linear classification layer is, where the i - th row vector corresponds to the weight of the i - th HOI class, the specific formula is as follows::
[0040]
[0041] Here, is the normalized text embedding vector of the i - th HOI class after the above - mentioned processing. During the training process, for the i - th class, the output calculation formula of the linear classifier is:
[0042] S i =γx T w i +b i
[0043] In this module, to guide the classifier to better learn the semantic relationships between HOI classes during training, the method of integrating the semantic information of language embeddings into the initial weights of the classifier is adopted. b i exists as a bias term, γ is a scalar hyperparameter that can control the output range, and x is the image feature vector from the backbone network.
[0044] Step Five: Multimodal Instance Semantic Classifier
[0045] In the process of constructing the classifier, a crucial step is to transform the input image and its corresponding HOI text description into feature vectors. For the input HOI text description, in step three, we use the text encoder of CLIP to process it. This encoder can mine the semantic information in the text and finally obtain text feature vectors with semantic guidance ability, which provide semantic-level guidance for subsequent classification. For the input image, we directly input it into the image encoder. The CLIP image encoder trained with a large amount of data has good performance and can extract rich visual features from the image. These features not only contain basic visual information such as the appearance and contour of objects, but also contain the context relationship and semantic clues between objects in the image, laying a foundation for accurately judging the interaction information in the image.
[0046] After obtaining the text feature vector and the image instance feature vector, the module needs to calculate the classification score. Here, cosine similarity is used to measure the similarity between the image feature and the text feature. Specifically, let the feature vector obtained by processing the image through the CLIP image encoder be I feature , and the feature vector obtained by processing the text through the CLIP text encoder be T feature . To make the features more suitable for the classification task, we perform a feature projection operation on these two feature vectors. Assuming the linear projection function is Linear, the projected vectors are as follows:
[0047] T′ = Linear(T feature )
[0048] I′ = Linear(I feature )
[0049] Linear projection is implemented through a fully connected layer, and the weights and biases of the fully connected layer are learnable parameters. In the training stage of the model, these parameters will be continuously updated according to the feedback of the loss function to improve the performance of the model. To avoid the influence of the feature vector length on the similarity calculation and ensure the fairness and stability of different feature vectors when calculating the similarity, we perform L2 normalization on the projected image feature vector and text feature vector. The formula for L2 normalization is:
[0050]
[0051] where PI′P2 and PT′P2 represent the L2 norms of I′ and T′ respectively. After normalization, the normalized image feature vector and the normalized text feature vector and
[0052] When calculating the cross-modal classification score, we use cosine similarity to evaluate the similarity between the normalized image feature vector and the normalized text feature vector. The original formula for cosine similarity is:
[0053]
[0054] S cosine The score is used to measure the semantic matching between the HOI instance in the image and the given text description, and its value ranges from [-1, 1]. When it is closer to 1, it indicates a higher semantic fit between the interaction instance in the image and the corresponding HOI text description, that is, the closer their semantics are.
[0055] To convert the adjusted similarity score into a probability value, the Sigmoid function is used for processing. The formula for the Sigmoid function is:
[0056]
[0057] S final ranges from [-1, 1], which represents the probability that the HOI instance in the image belongs to the corresponding text description category. The higher the value, the greater the possibility that the image shows the corresponding human interaction behavior.
[0058] Step 6: Use the multi-source fusion module to fuse the interaction prediction results
[0059] To accurately calculate the interaction probability of the HOI category, the present invention fuses the scores obtained in Step 4 and Step 5 in Step 6. In the previous steps, the present invention has done two aspects of work: on the one hand, integrating the HOI label semantics into the classification prediction head of the HOI detector to obtain the probability score of the interaction prediction; on the other hand, using the CLIP image-text encoder to further obtain the interaction probability score from a multi-modal perspective. The score of the interaction classification prediction head obtained in Step 4 is denoted as S label . The calculation process of this score is to first process the HOI text description by the CLIP text encoder and then calculate it in combination with the linear classification layer of the network. The score obtained by the classifier based on the multi-modal instance semantics in Step 5 is denoted as S multi-modal . This score utilizes the multi-modal characteristics of CLIP. By calculating the cosine similarity between the image features and the text features and then through a series of processes, it reflects the model's evaluation of the interaction probability from a multi-modal perspective.
[0060] To effectively fuse these two scores, this module adopts a weighted fusion strategy. Introducing a weight coefficient, the fused interaction probability score is calculated according to the following formula:
[0061] S fusion = αSlabel +(1 - α)S multi-modal
[0062] By calculating the interaction probability scores through the above method, the interaction behavior with the highest interaction score can finally be found, and then the final triple can be predicted.
[0063] Although the illustrative specific embodiments of the present invention have been described above for the understanding of those skilled in the art of the present technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. Any equivalent replacement or equivalent substitution, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
Claims
1. The present invention proposes a method for human interaction detection based on multi-modal information fusion, and the steps of the method are as follows: Step 1: Input an image and extract image features through a backbone network. Step 2: Perform set prediction through a Transformer encoder-decoder. Step 3: Extract text embedding features through a multi-modal label semantic embedding module. Step 4: Use the text embedding vector as the initial weight of the linear classification layer for interaction prediction. Step 5: Use a multi-modal instance semantic classifier for prediction. Step 6: Use a multi-source fusion module to fuse the interaction prediction results.
2. The method for detecting human interaction based on multi-modal information fusion according to claim 1, wherein The backbone network in Step 1 adopts a CNN network structure including multiple convolutional layers, activation layers, and pooling layers, and different levels of features are extracted by stacking these layers layer by layer: During the extraction process of the backbone feature network of the present invention, the input image x ∈ R 3×H×W . The convolutional kernel is a four-dimensional tensor with a shape of K × K × C in × C out , where K is the convolutional kernel size and C out is the number of output channels. The intermediate features are generated through the following convolutional operation formula: After the convolutional operation, an activation function is usually applied to introduce non-linear features, and its formula is as follows: Z i,j,c = max(0, Y i,j,c ) The Z here i,j,c is the value of the output feature map at position (i, j) and channel c after being processed by the activation function. The pooling operation can reduce the spatial size of the feature map and reduce the amount of calculation. Its formula is as follows:
3. The method for detecting human interaction based on multi-modal information fusion according to claim 1, wherein In Step 3, guiding information is added to the HOI triple and converted into the form of a natural language description prompt, specifically as follows: Since the original HOI object label and interaction action exist in the form of independent words, to enhance the text semantic expression, we add specific guiding information on the basis of each HOI triple description to associate the two. The specific operation is to transform the HOI triple and convert it into the form of a natural language description prompt, that is, [guiding word][person][action][item]. This transformation method helps the language model better understand and process HOI class semantics and meets the input requirements of CLIP pre-trained on natural language text. The token sequence is encoded using a multi-layer Transformer architecture to finally generate the hidden state representation of each token. Here, the final hidden state corresponding to the [EOS] token vector in the CLIP model is selected as the feature representation of the entire text, denoted as h text . Specifically, assuming that the output after CLIP encoding of the input text is H, the formula is as follows, where h i represents the hidden state vector of the i-th token, and n is the number of tokens. The embedding vector h text = h [EOS] , that is, the vector corresponding to the [EOS] token is taken from H and used as the final text embedding. H = {h [SOS] , h1, h2, L, h n , h [EOS]}.
4. The method for detecting human interaction based on multi-modal information fusion according to claim 1, wherein, The present invention integrates the label embedding in Step 3 into the classifier: For the obtained text embedding vector h text Perform normalization processing to make it consistent in scale. Let the weight matrix of the linear classification layer be W, where the i-th row vector w i Corresponds to the weight of the i-th HOI class, and the specific formula is as follows: The h here texti is the normalized text embedding vector of the i-th HOI class after the above processing. During the training process, for the i-th class, the output calculation formula of the linear classifier is: S i = γx T w i + b i To guide the classifier to better learn the semantic relationships between HOI classes during training, a method of integrating the semantic information of language embeddings into the initial weights of the classifier is adopted. Here, b i exists as a bias term, γ is a scalar hyperparameter that can control the output range, and x is the image feature vector from the backbone network.
5. The method for detecting human interaction based on multi-modal information fusion according to claim 1, characterized in that In Step 5, through a multi-modal instance semantic classifier, multi-modal instance semantic interaction classification is realized, specifically as follows: After obtaining the text feature vector and the image instance feature vector, the module needs to calculate the classification score. Here, cosine similarity is used to measure the similarity between the image features and the text features. Specifically, let the feature vector obtained after the image is processed by the CLIP image encoder be I feature , and the feature vector obtained after the text is processed by the CLIP text encoder be T feature . To make the features more suitable for the classification task, we perform a feature projection operation on these two feature vectors. Assuming the linear projection function is Linear, the projected vectors are as follows: When calculating the interaction classification score, we use the cosine similarity to evaluate the similarity between the normalized image feature vector and the normalized text feature vector. The original calculation formula of the cosine similarity is: To convert the adjusted similarity score into a probability value, the Sigmoid function is used for processing. The calculation formula of the Sigmoid function is: S final The value range of S is between [-1, 1], which represents the probability that the HOI instance in the image belongs to the corresponding text description category. The higher the value, the greater the possibility that the image shows the corresponding human interaction behavior.
6. The method for detecting human interaction based on multimodal information fusion according to claim 1, characterized in that The weight coefficient in Step 6 is optimized and adjusted according to the loss function during the training process to achieve the effective fusion of the two scores, specifically as follows: Let the interaction classifier prediction score be S label , and the score is calculated by processing the HOI text description through the CLIP text encoder and then combining it with the linear classification layer of the network. Let the classifier score based on multi-modal instance semantics be S multi-modal , which utilizes the multi-modal characteristics of CLIP, is obtained by calculating the cosine similarity between the image feature and the text feature and through a series of processes, and reflects the model's evaluation of the interaction probability from a multi-modal perspective. To achieve the effective fusion of the two scores, this module adopts a weighted fusion strategy. The weight coefficient α(0 ≤ α ≤ 1) is introduced, and the calculation formula of the fused interaction probability score is as follows: S fusion = αS label +(1 - α)S multi-modal By calculating the interaction probability score in the above manner, the interaction behavior with the highest interaction score is finally obtained, and the final triple is predicted.
Citation Information
Cited By
Video manuscript abnormal data detection method and system based on artificial intelligence
CN120599383A
CLIP adaptive optimization method based on diffusion feedback driving and related equipment
CN121708602A