Person interaction relation detection method based on detection Transform

By designing a parallel inference network, using two independent predictors to handle instance-level positioning and interaction-level semantic understanding, the problem that existing detection Transformers cannot effectively understand complex character interactions is solved, and better detection performance is achieved, especially when dealing with sparse set interactions.

CN120107991APending Publication Date: 2025-06-06YANGTZE DELTA REGION INST OF UNIV OF ELECTRONICS SCI & TECH OF CHINE (HUZHOU)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311645369.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing detection Transformers cannot effectively understand complex character interactions in character interaction relationship detection, and the attention and field of view of instance detection and interaction relationship understanding are inconsistent, resulting in poor detection results.

Method used

A parallel reasoning network (PR-Net) was designed, which contains two independent predictors: instance-level predictors and interactive relationship-level predictors, respectively responsible for positioning and semantic understanding. Through parallel predictors and trident-type non-maximum suppression technology, instance information and interaction relationship information are decoupled to improve the model's understanding of complex interaction relationships.

Benefits of technology

On the HICO-DET and V-COCO datasets, PR-Net significantly improves the performance of character interaction detection, especially when detecting character interactions in common and rare sets, and can more accurately identify complex interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107991A_ABST
    Figure CN120107991A_ABST
Patent Text Reader

Abstract

The invention provides a new person interaction relation detector parallel reasoning network (PR-Net) based on detection Transform, which is composed of an instance level predictor and an interaction relation level predictor so as to solve the problem of inconsistent attention areas of attention views between instance level prediction and interaction relation level prediction. In addition, the parallel reasoning network provided by the invention also realizes better balance between detection of two sub-task instances of character interaction relationship detection and interaction relationship understanding. Besides, in combination with interaction relationship consistency loss and a trigeminal non-maximum suppression module, the parallel reasoning network provided by the invention can greatly exceed the most advanced character interaction relationship detection method in the past, which verifies the effectiveness of the parallel reasoning network on the character interaction relationship detection task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of character interaction relationship detection in deep learning, and designs a different character interaction relationship detection model based on the existing Transformer character interaction relationship detector. Background Art

[0002] Human interaction detection is one of the important tasks in visual understanding. It requires detecting people and objects in the image scene and identifying the categories of interaction relationships between people and objects. The specific process of human interaction detection is to first input an image containing human interaction relationships, use the object detector to identify and locate all people and objects in the image, and then use the interaction analyzer to identify all <people, interaction relationships, objects> triplets in the image. Compared with the object detection task, the human interaction detection task has higher requirements for image cognition and understanding. It is necessary to explore "what exactly happened in the image" rather than just "what exists in the image" to achieve a higher level of analysis and understanding of the image. The difficulty of the human interaction detection task also lies in the fact that there are multiple reasonable human interaction relationships between people and objects of the same category. To accurately determine the specific interaction relationship between people and objects of the same category, causal reasoning needs to be performed based on various information in the scene. For example, whether a person is standing still on a skateboard or sliding on a skateboard, the semantics of these two interaction relationships are different, and it is necessary to infer whether the person is on a slope in the scene and the person's body movement information. Moreover, the task of detecting person interaction relationships also involves multiple interaction relationships between a person and the same object, different types of interactions between the same person and multiple objects, and different types of interactions between multiple people and the same object. The existence of these special cases makes person interaction relationship detection a complex multi-label understanding task. Solving this task also requires a more comprehensive and in-depth understanding of visual semantics.

[0003] In the early years of the field of human interaction detection, most human interaction detection models used methods based on convolutional neural networks (CNNs) to detect people, objects and their interactions in images. However, these methods based on convolutional neural networks have three insurmountable defects. First, due to the local connectivity of CNN itself, the CNN-based methods cannot directly use the global context features at the image level; second, the CNN-based methods need to integrate global context features through manually designed feature aggregation methods; third, the CNN-based methods will inevitably mix multiple different human interaction features when predicting human interaction. Global context features are very important for human interaction detection tasks. For example, if the image context scene is known to be the seaside, the human interaction in the image is likely to be a person holding a skateboard or surfing on the sea. Therefore, it is a better choice to use the Transformer method that can effectively aggregate global context features to detect human interaction. Moreover, the Transformer method usually combines set prediction to independently process multiple different human interaction features in the image, thereby avoiding the mixing and interference of multiple different human interaction features.

[0004] The methods for detecting human interaction relationships that have emerged in recent years can be divided into the following four categories: methods based on multi-stream neural networks, methods based on graph neural networks, methods based on key points, and methods based on Transformer. Methods based on multi-stream neural networks can use information from multiple sources to obtain richer human interaction relationship features, which may include human feature information, object feature information, relative spatial information of human characters, posture information, behavioral semantic information, and human gaze angle information, so as to achieve the goal of improving the prediction performance of human interaction relationships. Methods based on graph neural networks regard multiple instances or multiple human interaction relationships in an image as a node in a graph, and use the characteristics of efficient information interaction of graph neural networks to allow each instance or human interaction relationship in the image to obtain rich contextual information, so as to better predict human interaction relationships. Methods based on key points will regard each interaction relationship as a key point in the image, directly predict the key points of the interaction relationship, thereby identifying and locating the interaction relationship, and then use the matching method of interaction relationship and person to match the detection results of the person with the detection results of the interaction relationship one by one. The Transformer-based method can make full use of the global context information in the image to obtain image features with richer semantics, and use the Transformer decoder to perform adaptive set predictions on the interaction relationships between people, reducing the reliance on manual design and complex post-processing, thereby efficiently performing person interaction relationship detection.

[0005] The existing detection Transformer cannot understand complex character interaction relationships well. The previous detection Transformer for the character interaction relationship detection task simultaneously detected the bounding boxes of people and objects in the interaction relationship and identified the interaction relationship between the two. This resulted in the two character interaction relationship detection subtasks, overall instance detection and interaction relationship understanding, being bundled together. The attention field of these two tasks is not consistent, resulting in the final character interaction relationship detection effect is not good enough, and it is impossible to effectively understand some complex and confusing interaction relationships.

[0006] The present invention solves the problem of inconsistent attention fields of multiple tasks based on a parallel reasoning network, thereby enabling the model to have a deeper understanding of complex character interactions and better overcome the above-mentioned shortcomings. Summary of the invention

[0007] In order to overcome the shortcomings of the above-mentioned prior art, the present invention proposes a new detection Transformer-based method, named Parallel Reasoning Network (PR-Net). PR-Net also contains two independent predictors for instance-level positioning and interactive relationship-level semantic understanding. The former focuses on instance-level positioning by perceiving the end area of ​​the instance. The latter diffuses the field of view to the interactive relationship area to better understand the interactive relationship-level semantics. Detailed experiments and analysis on the HICO-DET dataset prove that the PR-Net proposed in this chapter can effectively alleviate these problems. PR-Net can also achieve quite superior performance on HICO-DET and V-COCO.

[0008] Figure 1 The overall framework of PR-Net is shown in . Structurally, PR-Net mainly includes an image global feature extraction and information interaction module and two parallel predictor modules (instance-level predictor and interaction-relationship-level predictor). The two predictors are used to decouple instance information (such as human bounding box, object bounding box, object category) and interaction relationship information (interaction relationship bounding box, interaction relationship category). After that, the loss functions for instance level and interaction relationship level are used to learn the location of the instance and the interaction relationship between each person-object pair. Finally, trident-type non-maximum suppression is introduced to effectively filter out repeated person interaction relationship predictions. Figure 1The overall framework in the paper consists of four parts: image feature extraction module, instance-level predictor and interaction-level predictor, training and post-processing techniques. First, the structure of CNN and Transformer is used to extract serialized visual features. Then, the instance-level predictor is used to predict the categories and bounding boxes of people and objects, and the interaction-level predictor is used to predict the categories and joint boxes of interaction relationships. During the training process, a multi-task training joint method composed of multiple loss functions is used, and the trident non-maximum suppression method is used during the evaluation process.

[0009] The technical solution adopted by the present invention is:

[0010] Step 1: The global feature extraction and information interaction module of the entire image includes a standard convolutional neural network backbone (CNNBackbone) c and a Transformer encoder f e The conventional convolutional neural network backbone takes the input image x∈R 3×H×W Transformed into a global context feature map z∈R c×H′×W′ , where the image is downsampled to a shape with channel dimension c and spatial size (H′, W′). Then, the global context feature map is serialized into a continuous tag sequence, where the spatial structure of the feature map is collapsed into a tag sequence of dimension H′×W′. Afterwards, the tag sequence is linearly mapped to T={t i |t i ∈R c′}, where N q =H′×W′. Finally, these mapped tag sequences are fed into the Transformer encoder;

[0011] Step 2: The instance-level predictor consists of a three-layer standard Transformer decoder and three small Feed Forward Networks (FFNs). ip The visual memory E is initialized with a series of randomly learnable instance-level query vectors Decoding is performed, where each query vector is embedded with a sinusoidal position The instance-level query vector is used to train and learn more accurate instance locations, focusing more on local information about instance locations. Three small feedforward neural networks (FFNs) include the human bounding box FFNφ hb , object bounding box FFNφ ob , object category FFNφ oc , which are for the human body bounding box b h , object bounding box b o and object category co feature transformation.

[0012] Step 3: In order to obtain the final character interaction relationship detection results, this chapter needs to use the instance-level predictor to output the human bounding box, object bounding box and object category, and use the interaction relationship level predictor to output the interaction relationship category. Based on the above prediction, the trident-type non-maximum suppression post-processing module proposed in the present invention performs repeated prediction filtering. Specifically, if the trident intersection-over-union ratio between the i-th and i-th character interaction relationship predictions is higher than the overlap threshold h, the predictions with lower character interaction relationship prediction scores will be filtered out.

[0013] Compared with the prior art, the present invention has the following beneficial effects:

[0014] (1) The detection results of the proposed parallel reasoning network under two different evaluation criteria of HICO-DET's full set and non-sparse set are the best among all methods, which shows that the parallel reasoning network of the present invention is more competitive than previous methods in detecting the vast majority of common character interaction relationships.

[0015] (2) The parallel reasoning network also performs very well in detecting rare sets of character interaction relationships (character interaction relationship categories with less than 10 training instances). This performance can be attributed to the fact that the parallel reasoning network can transfer the knowledge understanding of interaction relationships in non-scarce sets to rare sets of samples, so that complex interaction relationships in rare sets can also be detected. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 : The overall architecture diagram of the present invention.

[0017] Figure 2 : The attention fields corresponding to the two level predictors in PR-Net.

[0018] Figure 3 Figure 2: Visualization of some interaction relationship instances (Top1 results) detected by the parallel inference network on the HICO-DET test set.

[0019] Figure 4 For: Performance comparison of PR-Net with previous character interaction methods on the HICO-DET test set.

[0020] Figure 5 For: Performance comparison of PR-Net with previous character interaction methods on the V-COCO test set.

[0021] Figure 6 Figure 2: Ablation experiment results of PR-Net using ResNet-101 as the backbone network on the HICO-DET test set. DETAILED DESCRIPTION

[0022] The present invention is further described below in conjunction with the accompanying drawings.

[0023] Step 1: Character interaction detection model based on detection transformer

[0024] This chapter proposes a parallel reasoning network (PR-Net) to solve the problem of inconsistent multi-task attention fields, with the goal of better understanding some complex and confusing interactive relationships without increasing the amount of computation. Specifically, PR-Net contains two parallel predictors, which focus on instance-level positioning and relationship-level understanding respectively. In addition, the instance-level query vector of the instance-level predictor and the interactive relationship-level query vector of the relationship-level predictor in this chapter are one-to-one corresponding, so there is no need for any instance-relation matching procedure between them, which greatly reduces the computational burden. In addition, this chapter designs a consistency constraint loss between the person-object box and the relationship box during the training process, so that the joint area of ​​the person-object pair positioning box is as consistent as possible with the overall attention area of ​​the relationship understanding, avoiding the relationship understanding from deviating from the person-object pair itself. In the reasoning process, this chapter designs a trident non-maximum suppression post-processing module (Trident-NMS) to comprehensively consider the overlap of the person, object and relationship regions, so as to achieve better post-processing effects.

[0025] exist Figure 1 The overall framework of PR-Net is shown in . Structurally, PR-Net mainly includes an image global feature extraction and information interaction module and two parallel predictor modules (instance-level predictor and interaction-relationship-level predictor). The two predictors are used to decouple instance information (such as human bounding box, object bounding box, object category) and interaction relationship information (interaction relationship bounding box, interaction relationship category). After that, the loss functions for instance level and interaction relationship level are used to learn the location of the instance and the interaction relationship between each person-object pair. Finally, trident-type non-maximum suppression is introduced to effectively filter out repeated person interaction relationship predictions. Figure 1 The overall framework in the paper consists of four parts: image feature extraction module, instance-level predictor and interaction-level predictor, training and post-processing techniques. This chapter first uses the structure of CNN and Transformer to extract serialized visual features. Then, the instance-level predictor is used to predict the categories and bounding boxes of people and objects, and the interaction-level predictor is used to predict the categories and joint boxes of interaction relationships. During the training process, a multi-task training joint method composed of multiple loss functions is used. During the evaluation process, this chapter uses the trident non-maximum suppression method.

[0026] The feed-forward neural networks (FFNs) for predicting human bounding boxes, object bounding boxes, and interaction bounding boxes are all set to three fully connected layers equipped with ReLU nonlinear mapping functions, while the feed-forward neural networks (FFNs) for predicting object categories and interaction categories are set to one fully connected layer equipped with Softmax and Sigmoid nonlinear mapping functions respectively. During the training process, this chapter uses the DETR parameters trained on the MS-COCO dataset to initialize the network weights of the parallel inference network. In this chapter, the loss function weight coefficients and matching cost weight coefficients of bounding box regression (including human, object, and interaction), bounding box generalized intersection-over-union (including human, object, and interaction), object category, interaction category, and interaction consistency are set to 2.5, 1, 11, 1, and 0.5, respectively. The above weight coefficient settings are basically consistent with QPIC. In this chapter, the entire parallel inference network is optimized by the AdamW optimizer, and the weight decay is set to 10-4. This chapter trains the model for 150 epochs, with the learning rate of the backbone feature network set to 10-5 and the learning rate of the rest of the network set to 10-4. The learning rate is reduced by 10 times at the 100th and 130th epochs respectively. The batch size of each iteration during the entire training process is set to 16. All model training and evaluation experiments are conducted on 8 NVIDIA TelsaA100 GPUs, PyTorch1.8.1 and CUDA11.2.

[0027] Step 2: Image global feature extraction method and information interaction module

[0028] The entire image global feature extraction and information interaction module includes a standard convolutional neural network backbone (CNNBackbone) c and a Transformer encoder f e The conventional convolutional neural network backbone takes the input image x∈R 3×H×W Transformed into a global context feature map z∈R c×H′×W′ , where the image is downsampled to a shape with channel dimension c and spatial size (H′, W′). Then, the global context feature map is serialized into a continuous tag sequence, where the spatial structure of the feature map is collapsed into a tag sequence of dimension H′×W′. Afterwards, the tag sequence is linearly mapped to T={t i |t i ∈R c′}, where N q=H′×W′. Finally, these mapped token sequences are fed into the Transformer encoder. For the Transformer encoder, each encoder layer follows the standard Transformer architecture, including a multi-head self-attention module (MSA) and a feedforward neural network (FFN). Additional position embedding q e ∈R c′×H′×W′ It will also be added to the serialized continuous tags to supplement the position information. Based on the self-attention interaction layer, the encoder can map the global context feature map output by the previous CNN into a feature map with richer context information. Finally, these encoded image feature sets {d i |d i ∈R c′} will be represented as visual memory E = f e (T,q e ). This visual memory E includes complete and rich contextual information in the image.

[0029] ResNet-50 and ResNet-101 are used as the backbone feature extractors of the parallel reasoning network. The Transformer encoder in the parallel reasoning network is 6 layers, and the number of heads of the Multi-Head SelfAttention Module in each layer is set to 8. The number of Transformer layers in the instance-level predictor and the interaction-level predictor is set to 3. The hidden layer dimension of the visual memory in the parallel reasoning network is set to 256. For the HICO-DET and V-COCO datasets, the number of instance-level query vectors and interaction-level query vectors is set to 100.

[0030] Step 3: Parallel reasoning module for instance localization and relation identification

[0031] The interaction relationship understanding problem is decoupled from the human interaction relationship detection (HOIDetection) problem, and an interaction relationship level predictor is used to infer the interaction relationship from the semantics of a larger scale. This chapter proposes to use an interaction relationship bounding box to guide the interaction relationship level predictor to perceive the semantic relationship between people and objects. The interaction relationship level predictor contains a three-layer standard Transformer decoder f rp and two small feed-forward neural networks FFNs. Interactive relation-level query vector Initialized randomly, the interaction level position embedding is set to a sinusoidal mode, and the two are added together and sent to the Transformer decoder f rp In , it is used to decode the visual memory E at the interactive relationship level to obtain the interactive relationship level features in the image. Then the Transformer decoder f rp The output interaction relationship features are respectively sent to the interaction relationship bounding box FFNφ rb and the interaction relationship category FFNφ rc , which can decouple the interaction relationship bounding box br and the interaction relationship category cr respectively. Specifically, it can be expressed as follows:

[0032] br=Φ rb (f rp (Q r ,P r ,E))

[0033] cr=Φ rc (f rc (Q r ,P r ,E))

[0034] Interaction-level query vector Q r It will also pay more attention to the entire area where people and objects interact, rather than being limited to a certain person or object in the interaction. Therefore, compared with previous interaction relationship prediction methods, the interaction relationship predictor in this chapter can understand the relationship semantics more comprehensively and carefully, and thus can more accurately identify complex interaction relationships.

[0035] In addition, in order to convert the interaction relationship category c output by the interaction relationship level predictor r Volume bounding box output by instance-level predictor Object Bounding Box and object category c o To match, this chapter adopts the interactive relationship level query vector Q r and the instance-level query vector Q i Predictions are made one-to-one in a sequential bundling manner. Specifically, for the i-th instance output of the instance-level predictor and the ith interaction output c of the interaction level predictor ir , the corresponding character interaction relationship label will be the same, c ir Describes the human body frame and object frame Such a simple design allows predictors at different levels to focus on different sizes of attention fields, while discarding the previous complex instance and interaction relationship matching steps such as HO Pointer in HOTR, thereby achieving better person interaction relationship detection performance with less computational effort.

[0036] from Figure 2 and Figure 3 It can be seen from the figure that the parallel reasoning network (PR-Net) can accurately detect the human bounding box, object bounding box, interaction bounding box and their corresponding interaction relationships at the same time. Figure 2 and Figure 3 The second prediction example in the first row shows that the parallel inference network can accurately identify that the person in the image is riding a horse, which can be seen from Figure 2 and Figure 3 The second row of prediction examples shows that the parallel reasoning network can accurately detect those character interactions with small interaction areas that are difficult to distinguish. Figure 4 The third row of prediction examples shows that the parallel reasoning network can also detect the interaction between people in grayscale images very well. Therefore, from the results of the prediction quality analysis, the parallel reasoning network proposed in this chapter can well detect those complex and difficult interaction between people.

[0037] Figure 6 The figure and the graph show the contribution of each part to the final person interaction relationship detection performance. Among these three parts, the parallel predictor is the core method of this chapter. With the help of the parallel predictor alone, a mAP performance gain of 1.72 can be obtained on the HICO-DET test set. This result shows that the parallel reasoning structure proposed in this chapter can significantly improve the instance localization ability and interaction relationship understanding ability of the HOI detection model.

[0038] The above description is only a specific implementation mode of the present invention. Any feature disclosed in this specification, unless otherwise stated, can be replaced by other equivalent or alternative features with similar purposes; all disclosed features, or all methods or steps in the process, except for mutually exclusive features and steps, can be combined in any way.

Claims

1. A method for detecting human interaction relationships based on detection transformer, The following steps are involved: Step 1: The global feature extraction and information interaction module of the entire image includes a standard convolutional neural network backbone (CNNBackbone) c and a Transformer encoder f e , the conventional convolutional neural network backbone takes the input image x∈R 3×H×W Transformed into a global context feature map z∈R c×H′×W′ , where the image is downsampled to a shape with channel dimension c and spatial size (H′, W′), then the global context feature map is serialized into a continuous tag sequence, where the spatial structure of the feature map is folded into a tag sequence of dimension H′×W′, and then the tag sequence is linearly mapped to T={t i |t i ∈R c′ }, where N q =H′×W′, finally, these mapped tag sequences will be sent to the Transformer encoder; Step 2: The instance-level predictor consists of a three-layer standard Transformer decoder and three small feed-forward neural networks (FFNs). ip The visual memory E is initialized with a series of randomly learnable instance-level query vectors Q i = {q i |q i ∈R c′ } i=1 N q Decoding is performed, where each query vector is embedded with a sinusoidal position embedding p i ∈R c′×Nq , the instance-level query vector is used to train and learn more accurate instance locations, focusing more on local information about instance locations. Three small feedforward neural networks (FFN) include the human bounding box FFNφ hb , object bounding box FFNφ ob , object category FFNφ oc , which are for the human body bounding box b h , object bounding box b o and object category c o feature transformation. Step 3: In order to obtain the final character interaction relationship detection results, this chapter needs to use the instance-level predictor to output the human bounding box, object bounding box and object category, and use the interaction relationship level predictor to output the interaction relationship category. Based on the above prediction, the trident-type non-maximum suppression post-processing module proposed in the present invention performs repeated prediction filtering. Specifically, if the trident intersection-over-union ratio between the i-th and i-th character interaction relationship predictions is higher than the overlap threshold h, the predictions with lower character interaction relationship prediction scores will be filtered out.

2. The method according to claim 1, It is characterized in that The Transformer attention mechanism in step 2 focuses on instance-level localization by perceiving the end regions of the instance.

3. The method according to claim 1, It is characterized in that In step 3, the field of view is expanded to the interactive relationship area to better understand the interactive relationship level semantics.