A hierarchical human interaction detection method based on human interaction intention information

By introducing multi-grained human interaction intention information into character interaction detection, the HII-Net algorithm significantly improves detection accuracy and robustness, solves the shortcomings of existing methods in terms of accuracy and robustness, and realizes high-performance detection in complex scenarios.

CN116311518BActive Publication Date: 2025-06-10BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310266335.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-20
Publication Date
2025-06-10
Estimated Expiration
2043-03-20

AI Technical Summary

Technical Problem

The existing character interaction detection methods have shortcomings in terms of accuracy and robustness, especially when dealing with interactive behaviors at different angles and focal lengths and overlapping occlusion scenes of human body, the detection accuracy is not high and the robustness is poor.

Method used

A hierarchical character interaction detection algorithm HII-Net is proposed. By introducing macroscopic human body gaze information, microscopic human joint information and meso-human body part information, we construct multi-grained complementary human interaction intention information to improve detection accuracy and robustness.

Benefits of technology

On the HICO-DET and V-COCO datasets, HII-Net significantly improves the accuracy of character interaction detection. Compared with the latest PMFNet algorithm, HII-Net's mAP on the HICO-DET dataset has increased by 7.56%, 9.62% in rare categories, and 52.54mAP on the V-COCO dataset, proving its superior performance in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116311518B_ABST
    Figure CN116311518B_ABST
Patent Text Reader

Abstract

The present invention discloses a hierarchical human interaction detection method based on human interaction intention information, which is divided into 1) object detection: detecting all object instances in the input image. 2) Human interaction detection: performing human interaction detection on all <human-object> pair instances in the image. The human gaze information is abstracted through the design of visual features to model the context area concerned by the interaction participants; a human body pose graph construction oriented to human interaction intention is proposed to optimize the differential information of body movement for interaction detection; the distance-feature between the person and the object is used as a guide to optimize the visual distance feature, so as to improve the performance of the human interaction detection algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and human interaction detection, and studies a new human interaction detection method. Background Art

[0002] Object detection and recognition are the basis and prerequisite for image analysis and image understanding. Its purpose is to locate the position of objects in an image or video sequence and identify the corresponding object categories. It is one of the important basic problems in computer vision. However, in order to better understand the visual world, the computer should not only accurately detect individual object instances in the scene, but also further understand the interaction methods between people and objects in the scene. Human-object interaction detection (HOI Detection) is to further understand the behavior of humans in the scene at a higher level. It requires the model to accurately locate the positions of people and objects in the scene and correctly predict various interaction behaviors existing between them. By studying human-object interaction detection to better understand the interaction mode between humans and the world, enabling the machine to have a mechanism to observe and understand the surrounding environment like humans and make quick decisions, can promote the development of robot technologies such as intelligent security and intelligent service robots. Thus, the human-object interaction detection technology not only has important theoretical research significance and practical value, but also has broad development potential.

[0003] Early human interaction detection methods mainly rely on visual features to capture the contextual relationship between people and objects, or use the spatial position relationship of structured bodies and objects to train human interaction models. For example, Chao et al. proposed the network HO-RCNN, which first used object spatial relationships and instance features of people for human interaction detection. Liao et al. proposed to regard human interaction detection as a key point detection and group matching problem, and used pure visual features to predict interaction categories. However, the detection methods that only rely on rough visual information and spatial relationships need to improve their accuracy. To further improve the detection accuracy, the instance-centered attention network iCAN added an attention mechanism in the human flow and object flow to automatically extract important contextual features. Shen et al. proposed a HOI decomposition model and constructed a high-precision human interaction detection network composed of a set of visual feature extraction layers, verb separation, and object detection networks. However, human interaction detection is to detect the action behaviors of people on objects. The above methods only extract rough visual features and spatial features of human and object regions, lacking the extraction and use of human body information, resulting in low accuracy of these algorithms. Therefore, some researchers proposed to use human body pose information in the human interaction detection task. Li et al. used a pose estimation network and a human pose branch to construct an interaction detection network and distinguished the interactivity of the network, enabling the model to learn interactivity knowledge. Liang et al. standardized the center of the human body bounding box, constructed absolute spatial pose features, and proposed a modular network PMN based on human body pose. Kim et al. found that there are natural correlations or oppositions between interaction actions and proposed a new pairwise HOI recognition framework for body part attention. However, due to the large differences in the poses of the same human interaction behavior captured from different directions and angles with different focal lengths, and as the number of human instances in the scene increases, the probability of human body overlap and occlusion appears. Due to the lack of complementarity of granularity between layers of the detection network, simply integrating coarse-grained human body pose information into the training of the model from a macroscopic perspective does not greatly improve the model accuracy and the model robustness is not high, and the overall performance still faces the problem of low accuracy.

[0004] To solve the above problems and obtain higher accuracy in human interaction detection, the present invention proposes a hierarchical human interaction detection algorithm enhanced by human interactive intention information, namely HII-Net (Hierarchical HOI Detection Framework Augmented by Human Interactive Intention), to achieve human interaction detection based on multi-granularity complementary human interactive intention information. Compared with the human interaction detection algorithm PMFNet, HII-Net introduces macroscopic human gaze information, microscopic human joint information, and mesoscopic human body part information between the two. They jointly provide visual interaction clues beyond layout and appearance for HOI detection, making the model prediction more accurate. Compared with the algorithm proposed by Xu et.al, HII-Net innovatively constructs a hierarchical HOI detection network with granularities from low to high being the spatial layout layer, the interactive intention layer, and the spatial layout layer. It uses the Zoom In method to extract highly complementary interactive information layer by layer to optimize features and improve the prediction accuracy.

[0005] Experiments were conducted on the existing human interaction detection datasets HICO-DET and the subset V-COCO (Verbs in COCO) of the Microsoft COCO dataset to verify the effectiveness of the proposed method. Specifically, HII-Net achieved an accuracy of 23.24 mAP on the HICO-DET dataset and an accuracy of 19.45 mAP on rare categories. Compared with the latest human interaction detection algorithm PMFNet, the relative gains were 7.56% and 9.62% respectively; it achieved an accuracy of 52.54 mAP on the V-COCO dataset. Summary of the Invention

[0006] To further improve the performance of the human interaction detection algorithm based on the traditional multi-stream architecture, the present invention proposes a new hierarchical human interaction detection algorithm based on human interactive intention information, namely HII-Net. HII-Net proposes to abstract human gaze information through the design of visual features to model the context area of interest of interaction participants; and proposes to construct a human pose graph for human interactive intention to optimize the differential information of body movement for interaction detection; at the same time, it proposes to use the distance-feature between the person and the object as a guide to optimize the visual distance feature and improve the performance of the human interaction detection algorithm.

[0007] To this end, the key technical problems to be solved include: the representation of human gaze information based on human interaction intentions, so that different visual features can better represent the context areas concerned by interaction participants; the construction of a human pose graph for human key points and its operation mode, to enhance the model's learning ability of human information; the construction of an interaction distance-feature network based on human body part information, to abstract the distance position information between people through the interaction relationship between human body parts and objects.

[0008] The proposed person interaction detection algorithm based on the traditional multi-stream architecture can be decomposed into two stages. 1) Object detection: Detect all object instances in the input image. 2) Person interaction detection: Perform person interaction detection on all <person-object> pair instances in the image. The HII-Net network structure is designed as Figure 1 shown.

[0009] 1) Object detection: This part is mainly responsible for performing object detection on object instances in the RGB image to obtain the object bounding box, object category, and object detection score, and providing them to the person interaction detection stage for person interaction detection of <person-object> pair instances. In the experiment, we use Faster R-CNN pre-trained on the COCO dataset as the object detector.

[0010] 2) Person interaction detection: The scenarios of person interaction are complex. When people participate in different interactions, they will have different action postures. For the instances of "person eating an apple" and "person making a phone call", their actions are very similar, except that the object in the former is closer to the person's mouth and the object in the latter is closer to the person's ear. Therefore, simply using visual and spatial features cannot obtain high-performance predictions. The essence of person interaction detection is to detect human interaction intentions (i.e., actions). We design to use multi-granularity human information to construct an interaction intention layer to supplement the details of the original spatial semantic information. At the same time, considering the diversity of the sizes of person interaction detection data instances, if a hierarchical interaction detection network with complementary granularities is constructed, on the overall three-layer framework, it can achieve the complementarity of interaction information at three different granularities: macroscopic, mesoscopic, and microscopic. This can not only improve the overall performance but also take into account the person interaction detection performance in complex scenarios. Based on these motivations, we propose the person interaction detection algorithm HII-Net based on RGB images.

[0011] The innovations of HII-Net mainly include the following three points. First, considering that the gaze area of the human body contains key regional information for HOI detection, that is, the human gaze information provides important clues for human interaction intentions, we model the regional information of the gaze of the interaction participants and use the gaze clues of the human body to guide the context area of human attention in complex scenes. Second, considering that for most interaction behaviors, using the pose clues of people can have a positive impact on the results, we propose the construction of a human pose graph for human interaction detection and fuse it with the spatial layout branch to jointly form a spatial & pose branch flow to optimize the differential information of body movement on interaction detection. Third, since different interaction behaviors (such as eating, riding, opening, etc.) are only associated with certain parts of the human body, we add more fine-grained human body information to the network to construct a distance-feature map of human body parts, so that the visual features focus on the regional features more relevant to human interaction behaviors and filter out the regional features irrelevant to human interaction behaviors, further improving the performance of HII-Net in human interaction detection.

[0012] The input of the human interaction detection algorithm HII-Net is the RGB image x i , the detection box information of people and the detection box information of objects The output is the interaction behavior scores of all <person-object> pair instances in the image It is described by the formula as follows:

[0013]

[0014] Among them, is the set of m people in the image , is the set of n objects in the image , and the function corresponds to the HII-Net algorithm model in this paper. Therefore, represents the interaction behavior scores of mn <person-object> pair instances generated by the interaction of m people and n objects.

[0015] The human interaction detection algorithm HII-Net adopts a hierarchical network structure design, which consists of a spatial layout layer, an interaction intention layer, and an objective appearance layer. For clarity, we introduce the network structure and the overall algorithm process layer by layer according to the hierarchical network architecture, and then introduce the construction and operation modes of interaction intention clues such as human gaze features and human pose features involved in each layer in the form of subsections.

[0016] i. Hierarchical network structure

[0017] The hierarchical network structure consists of a spatial layout layer, an interaction intention layer, and an objective appearance layer. To obtain the input features of each branch, we use the Residual Network ResNet50 to extract the required visual features. The original input image first undergoes object detection by the Faster R-CNN object detection network. After obtaining the information of the detection bounding boxes of people and objects in the original input image, the image marked with the positions of people and objects is input into the ResNet50 network to extract the global feature map of the image. We adopt a two-stage HOI detection process. Given the input image x i , this framework first obtains the result S of the interaction judgment stage through the spatial layout branch, the human gaze branch, and the appearance flow j . The person-object pairs with judgment results higher than the threshold will enter the interaction classification stage. The candidate human-object pairs entering the second stage continue to pass through the spatial layout flow, the appearance flow, the pose flow, and the body part flow to obtain the final HOI classification result S c . Among them, the human joint positions and gaze information are obtained through transfer learning from other social activity datasets. Next, we will introduce each layer of the network HII-Net one by one. The overall flowchart of HII-Net is as shown in Figure 2 .

[0018] Spatial layout layer: The purpose of the spatial layout layer is to obtain interactive spatial layout information at the macroscopic interaction level. Since position information between instances is required in the HOI detection task, the spatial position relationship can be used to locate the positions of instances in the scene. Since we focus on the spatial positions of people and objects, the input of this layer ignores pixel values and only uses the position information of the bounding boxes. For the spatial flow branch of the spatial layout layer, the input is the spatial feature map M encoded by the position information of all <person, object> pairs output by the object detection sp . The encoding rule of the spatial feature map M sp is as follows: The pixel values within the bounding box of the instance are set to 1, and other values outside the instance bounding box are set to 0 in two channels. Therefore, for a given pair of human-object bounding boxes, its spatial interaction mapping is defined as a binary image with two channels: the first channel corresponds to the binary pattern of the human, and the second channel corresponds to the object. This representation enables the neural network to learn two-dimensional filters to respond to two-dimensional human spatial interaction patterns. In addition, the binary map needs to be scaled to a fixed size m×m.. M sp undergoes feature extraction through a shallow convolutional neural network, and two max pooling convolutional layers and two fully connected layers are used to extract the features f of the spatial layout flow sp , which participate in the final interaction category classification and are described by the formula as follows:

[0019] f sp = W sp1 f cnn (Msp ) (2)

[0020] S sp = Sigmoid(W sp2 f sp ) (3)

[0021] where, denotes the fully connected layer parameter matrix, and uses the Sigmoid non-linear activation function to perform human interaction classification on the <person-object> pair spatial features. S sp represents the probability scores of the spatial flow features on each interaction category.

[0022] Objective appearance layer: The objective appearance layer contains a human flow branch and a logistics branch, providing pixel-level appearance prediction information at the micro-interaction level for the network. We use a residual module containing global average pooling to extract the visual features f h and f o of people and objects from the global appearance features. Then we scale the extracted F h and F o to a fixed size p×p. And calculate the probability scores S h and S h of the human flow features and logistics features on the interaction category after feature enhancement through two fully connected layers. Described by the formula as follows:

[0023] S h = Sigmoid(W h f h ) (4)

[0024] S o = sigmoid(W o f o ) (5)

[0025] where, formulas (4) and (5) respectively represent the operations of two fully connected layers, and W h and W o denote the fully connected layer parameter matrices. For convenience, hereinafter we use to represent the above-mentioned fully connected operations.

[0026] Interaction intention layer: Although the spatial layout layer and the objective appearance layer provide interaction prediction information from the macroscopic and microscopic perspectives respectively, there is still a lack of interaction intention information at the mesoscopic level. The interaction intention layer can mine information from the mesoscopic perspective, reduce some unpredictable situations that may occur from the macroscopic and microscopic perspectives, and thus can mine deeper interaction intention information. Inspired by the above reasons, we construct an interaction intention layer driven by human interaction intention, providing a new computational perspective to utilize three visually observable forms of human intention. It consists of the following three branches:

[0027] (1) Human gaze tributary: Human gaze information can be considered as a macro-level interaction intention. The area where humans gaze usually contains key area information for HOI detection, which clearly conveys the interaction intention. When a person wants to pick up an object, he usually looks in the direction of the object while reaching out to pick it up. Perceiving potential targets through the eyes can facilitate the inference of interaction.

[0028] We use a pre-trained two-stream model to obtain the human gaze region. The pre-trained gaze prediction model takes as input the input image I and the location of the human eye center calculated by the human pose estimation network, and outputs a probability density map G of a fixed region. The word order network consists of a saliency path and a gaze path. The gaze path only has access to a close-up image of the person's head and its location, and obtains a heat map M (x h ,x p ). The saliency path takes the full image as input and obtains another heatmap H(x i ), which can learn the importance of objects. Combining these two results with the element-wise product, we get the following equation:

[0029]

[0030] Among them, x i is the input image, x h is a cropped close-up image of human features, x p is the quantized spatial position of the human head.

[0031] For each human instance in the image, select k candidate object regions b = (b 1 ,...,b k ). For each candidate region b∈b, we calculate its attention weight g b . Where g b It is calculated by adding the values ​​of the density map G in b and then normalizing the area of ​​b to:

[0032]

[0033] Then, we select the k candidate regions with the largest g b The region R is taken as the region where people are looking. For the selected region R, we define its corresponding feature vector f g ={f a , f l , f c}. Among them, f a is the appearance vector of object b, f l Is included x , l y , lw , l h Four-dimensional vector, where l x , l y Specifies the coordinate distance of the bounding box, l w , l h Specifies the height / width in the logarithmic space, f c is the target classifier score vector.

[0034] (2) Human keypoint branch: As a micro-level interaction intention, the joint attributes of humans have strong interaction expression capabilities. In fact, for most interactions, the ability to accurately distinguish them also requires specific pose information. Due to the weak binding force of the binary space pattern and the appearance features of humans and object instances, we use the positions of human keypoints to obtain pose position information and enhance the position constraint to reduce the differences in body movement on interaction detection.

[0035] First, we use human pose estimation to estimate 17 human keypoints of the human body. Then, we connect the 17 keypoints with lines of different gray values (0.15 - 0.95) and set other regions to 0 to construct a pose map. Since the line segments with different gray values represent different body regions, this modeling method can implicitly encode pose features. We use two max-pooling convolutional layers and two FC layers to connect the pose map with the layout maps of humans and objects to extract the features f of the spatial pose stream sp , the specific process is as Figure 2 shown.

[0036] To make the four influencing factors affect the interaction judgment process, we connect f g , f sp , f h and f o to obtain the joint overall vector f hol . Then we set the prediction score S j in the interaction judgment stage to:

[0037]

[0038] In the two-dimensional probability vector output in this stage, the first dimension is the probability of the existence of interaction, and the second dimension is the probability of the non-existence of interaction. We use the threshold δ to define whether the interaction exists. If the probability value is higher than δ, there is an interaction, otherwise there is no interaction. As mentioned above, for a given pair of human objects (b h , b o ), we first need to judge whether the interaction exists. Only the human-object pairs judged to have an interaction can enter the following classification stage.

[0039] (3) Body part tributaries: In actual scenarios, some interaction behaviors are only related to relatively subtle local parts and joints of the human body. For example, the actions of "grasping" and "cutting" are only related to our hands. For example, for the interaction <eating, hamburger>, the distance between the hand and the hamburger and the microscopic local features of the hand are crucial. As an intention clue at the mesoscopic level, the position information between the human body part and the object obtained through combination can reflect the interaction intention clue at the mesoscopic level.

[0040] First, we respectively construct a 2-channel distance map between each body part and the object as Figure 3 shown. Specifically, we define the position coordinates of the center point in the global features of the object as <h x ,h y >, and define humans as <o x ,o y >, where x and y are the coordinates in the x and y directions respectively. The length and width of the global appearance features are set to H and W. Then, we define two position vectors a and b. Vector a is defined from <o x ,o y > to <h x ,h y >; vector b is defined as each pixel <h x ,h y > in the human map.

[0041] Next, we construct a 2-channel distance map with dimensions of H×W. First, we use the cosine distance to reflect the relative difference in direction between the body and the object, which can reflect the position relationship between them. Therefore, the pixel value of the human body box in the first channel is the cosine distance between vector a and vector b. However, the cosine distance cannot distinguish the distance between vectors in the same direction, so we introduce the Euclidean distance to capture the absolute distance difference between the two vectors. Therefore, the pixel value of the human body box in the second channel is the Euclidean distance between vector a and vector b. Finally, we set the pixel values outside the human body boxes in the two channels to 0. In this way, we can use the 2-channel distance map to model the position distance relationship between the human body part and the object.

[0042] After obtaining the distance map, we concatenate the constructed distance map with the global appearance feature map in the channel dimension to obtain our distance-feature map as Figure 3 shown.

[0043] We use existing pose estimation methods to obtain the positions of key points (human joints), and then determine the corresponding human body parts based on the key points. Specifically, we construct 17 rectangles with each human key point as the center, and the size of each rectangle is 1 / 10 of the area of the original input image. The obtained rectangular regions are used as the regions where the body parts are located. Therefore, after obtaining the bounding boxes representing each part of the human body, based on the bounding boxes of each body part, we use ROI pooling to extract the corresponding regions from the distance-feature map and scale them to q×q. To combine the influence of each body part on the interaction detection, we concatenate the distance-feature maps of all the human body parts obtained and use the FC layer to convert them into the feature vector f of the body part branches part , and the specific process is as Figure 3 shown

[0044] v. Model Optimization and Interaction Score Fusion

[0045] Loss function: Since the overall structure of our method includes two stages, it should be considered separately when calculating the loss. Considering that the classification loss is usually calculated using the cross-entropy loss function CE():

[0046]

[0047] where, Y ij is the true action label, S ij is the predicted action score, is the average value of M batches of samples. To consider the influence of each stream and make the loss function converge better, we add the loss of each stream and the overall loss of the interaction detection stage, which enables more effective updating of the parameters of each stream

[0048] Therefore, the loss of the interaction judgment stage can be calculated according to the following formula

[0049]

[0050] where, Y J and S J are the true label and the final predicted score of the interaction judgment module. To consider the influence of each stream and make the loss function converge better, we add the loss of each stream and the overall loss of the interaction detection stage, which enables more effective updating of the parameters of each stream. The overall loss function of the interaction classification module is as follows

[0051]

[0052] where, α and β are the branch loss coefficient and the total loss coefficient of all streams, and Y C is the label of the interaction classification stage. When jointly training the network, the total loss function L of the two stages is used to update the parameters

[0053]

[0054] Interaction score fusion: Since the feature vector f of the body part stream part not only contains the fine features of the human body but also reflects the positional relationship between the person and the object, we regard it as an independent stream. We use a late fusion strategy to fuse the four streams. First, we use the FC layer to map the feature vector of each stream to the interaction prediction score S h 、S o 、S p and S part , and then we fuse the prediction scores of each stream. Considering that the spatial information and the human body information are complementary, we first fuse S h 、S o and S part , and then we multiply it by S p . Therefore, the final interaction prediction score vector S c of the final detection result is set to:

[0055]

[0056] where the operations and represent element-wise sum and multiplication respectively. The combination of the hierarchical idea and the two-stage strategy can better utilize the features at each pixel level and further improve the effect of interactive detection.

[0057] 3) Experimental details: In our experiment, to detect objects and extract features, we use Faster R-CNN and VGG16 as the feature backbones, which are performed on the pre-trained MS-COCO dataset. The object detection framework predicts the bounding boxes (b h ,b o ) and the confidence levels (S h ,S o ) of people and objects, and keeps the person bounding box with S h >0.6, so the object bounding box with >0.6 has the appearance features of the role. When obtaining the appearance features of people, we scale the extracted visual features to a fixed size of 7×7 (p = 7); when obtaining the features of human body parts, we scale the extracted features to 5×5 (q = 5). The binary spatial pattern map and the pose map of the person and the object are scaled to 64×64.

[0058] The human and object streams of the interaction judgment module and the interaction classification module consist of a residual block with global average pooling and four FC layers with an output dimension of n = 1024. The spatial layout branch stream consists of two convolutional layers with max pooling and two FC layers with n = 1024. In the interaction judgment module, the feature vectors of each stream are connected and mapped through two FC layers to predict the scores. The algorithm is implemented using TensorFlow and deployed on a machine with a single Nvidia 3090 GPU. Stochastic Gradient Descent (SGD) is used to train the network. We set the initial learning rate to 1×10-4 and the weight decay to 1×10-4. When testing the model, the probability threshold δ of the interaction judgment module is set to 0.3, and the loss coefficients α and β of the interaction classification module are set to 1.3 and 0.7 respectively. Description of the Drawings

[0059] Figure 1 Overall flowchart of HII-Net.

[0060] Figure 2 Network structure design diagram of HII-Net.

[0061] Figure 3 Process diagram for obtaining the distance-feature map. Detailed Implementation Manner

[0062] The following is a detailed description in conjunction with the drawings and embodiments.

[0063] To verify the actual effect of HII-Net, we used the publicly available human interaction detection datasets HICO-DET and V-COCO to evaluate the performance of human interaction detection. We followed the evaluation method of predecessors and used the average precision AP to evaluate the precision of each type of human interaction behavior, and then averaged the APs of all classes to obtain the final mean average precision mAP.

[0064] For a person-object pair instance in an image, if the intersection over union (IoU) of the detection box of the person and the detection box of the object with their respective ground truth bounding boxes is greater than 0.5, and the predicted human interaction class label of the current person-object pair is correct, then the current person-object pair is a positive sample.

[0065] To illustrate the positive effects of the present invention, we compared the proposed HII-Net with the latest human interaction detection methods: iCAN, iHOI, Inteactiveness, and PMFNet, etc. It can be seen from Table 1 and Table 2 that our method achieved higher precision.

[0066] Table 1 Performance of different methods on the HICO-DET test set

[0067]

[0068] Table 2 Performance of Different Methods on the V-COCO Test Set

[0069] Paper mAP(Sc.1) mAP(Sc.2) InteractNet 40.0 47.98 GPNN 44.0 - iCAN 45.3 52.4 iHOI 48.3 Xuet.al 45.9 - Interactiveness 47.8 54.2 PMFNet 52.0 - HII-Net(Ours) 52.54 59.71

[0070] Meanwhile, to verify the effectiveness of each part of the model, we conducted a comparative experiment on the V-COCO dataset. The results of the comparative experiment are shown in Table 3. Among them, we define the baseline model HII-Net[B] of HII-Net as a model composed of a simple human stream, object stream, and spatial layout stream. At this time, the performance of human interaction detection on the V-COCO dataset is 49.82 mAP. For the sake of convenience of expression, we represent the Baseline, Gaze Stream, Joint-pose Stream, and Body partStream of HII-Net as B, G, H, and P respectively.

[0071] Table 3 Performance of Comparative Experiments on the V-COCO Dataset

[0072] Model mAP(Sc.1) HII-Net[B] 49.76 HII-Net[BG] 50.83 HII-Net[BGH] 51.37 HII-Net[BGHP](Ours) 52.54

[0073] HII-Net[BG]: To verify the influence of human gaze information on human interaction intention, we modeled the area where the interaction participants gaze and used human gaze information to guide the context area that the human body focuses on in complex scenarios. Compared with the HII-Net[B] model, the performance of the HII-Net[BG] model increased from 49.76 mAP to 50.83 mAP, with a gain of 1.07 mAP.

[0074] HII-Net[BGH]: To verify the influence of human joint information on the performance of human interaction detection, we proposed the construction of a human body pose graph for human interaction detection and fused it with the spatial layout branch to jointly form a spatial & pose branch flow to optimize the differential information of body movement on interaction detection. Compared with the HII-Net[BG] model, the performance of the HII-Net[BGH] model increased from 50.83 mAP to 51.37 mAP, with a gain of 0.54 mAP.

[0075] HII-Net[BGHP]: To make the visual features focus on more discriminative location features of different human interaction behaviors and ignore irrelevant location features, we propose to use the distance-feature between humans and objects as the feature optimization to guide the visual branch, so that the visual features focus on the regional features more relevant to human interaction behaviors and filter out the regional features irrelevant to human interaction behaviors. Compared with the HII-Net[BGH] model, the performance of the HII-Net[BGHP] model increases from 51.37 mAP to 52.54 mAP, with a gain of 1.17 mAP.

[0076] In summary, the human interaction detection algorithm HII-Net proposed in the present invention incorporates the interaction intention information of real-life scenarios into the visual features, and proposes to abstract the human gaze information through the design of visual features to model the context area concerned by interaction participants; and proposes the construction of a human body pose graph for human interaction detection to optimize the differential information of body movement for interaction detection; at the same time, proposes to use the distance-feature between humans and objects as the optimization to guide the visual features, jointly completing the further improvement of the performance of human interaction detection. HII-Net has achieved the best current results in the detection performance on the HICO-DET dataset and its rare categories.

[0077] Table 4 Comparison of results of different HOI algorithms on the V-COCO dataset

[0078] HOI Class #pos iCAN InteractNet HII-Net(Ours) hold-obj 3608 29.06 37.33 42.52 sit-instr 1916 26.04 31.62 43.26 ride-instr 556 61.90 66.28 72.38 look-obj 3347 26.49 32.25 35.63 hit-instr 349 74.11 74.40 76.87 hit-obj 349 46.13 52.59 53.45 eat-obj 521 37.73 39.14 42.68 eat-instr 521 8.26 9.40 17.28 jump-instr 635 51.45 53.83 52.64 lay-instr 387 22.40 29.57 34.41 talk_on_phone 285 52.81 53.59 53.89 carry-obj 472 32.02 40.82 42.74 throw-obj 244 40.62 43.27 44.67 catch-obj 246 47.61 48.38 48.69 cut-instr 269 37.18 41.63 43.32 cut-obj 269 34.76 40.14 38.68 work_on_comp 410 56.29 65.51 66.43 ski-instr 424 41.69 49.95 47.24 surf-instr 486 77.15 79.70 78.75 HIIteboard-instr 417 79.35 83.39 87.95 drink-instr 82 32.19 34.36 42.61 kick-obj 180 66.89 66.26 64.86 read-obj 111 30.74 29.94 39.82 snowboard-instr 277 74.35 71.59 72.64 Average mAP 682 45.30 48.96 52.54

Claims

1. A hierarchical human interaction detection method based on human interaction intention information, characterized in that, the method includes: 1) Object detection: Detect all object instances in the input image; 2) Human interaction detection: Perform human interaction detection on all <person-object> pair instances in the image; 1) Object detection is responsible for detecting object instances in the RGB image to obtain the object bounding box, object category, and object detection score, and providing them to the human interaction detection stage for human interaction detection of <person-object> pair instances; 2) Human interaction detection: Use multi-granularity human information to construct an interaction intention layer to supplement the details of the original spatial semantic information; Considering the diversity of the sizes of human interaction detection data instances, construct a hierarchical interaction detection network with complementary granularity, and on a three-layer framework, realize the complementarity of three different granularities of interaction information: macroscopic, mesoscopic, and microscopic; The input of the human-object interaction detection method is the RGB image x i , the detection box information of people , the detection box information of objects The output is the interaction behavior scores of all <human-object> pair instances in the image It is described by the formula as follows: Among them, is the set of m individuals in the image , and is the set of n objects in the image ; represents the interaction behavior scores of the mn <person-object> pair instances generated by the interaction between m individuals and n objects;​ The hierarchical network structure consists of a spatial layout layer, an interaction intention layer, and an objective appearance layer; in order to obtain the input features of each branch, the residual network ResNet50 is used to extract the required visual features; the original input image first undergoes object detection by the Faster R-CNN object detection network. After obtaining the information of the detection boxes of people and objects in the original input image, the image marked with the positions of people and objects is input into the ResNet50 network to extract the global feature map of the image; a two-stage HOI detection process is adopted, given the input image x i , and the result S of the interaction judgment stage is first obtained through the spatial layout branch, the human gaze branch, and the appearance branch J . The person-object pairs with judgment results higher than the threshold will enter the interaction classification stage; the candidate human-object pairs continue to pass through the spatial layout branch, the appearance branch, the pose branch, and the body part branch to obtain the final HOI classification result S c .

2. The hierarchical human interaction detection method based on human interaction intention information according to claim 1, characterized in that, The purpose of the spatial layout layer is to obtain interactive spatial layout information at the macro interaction level; locate the positions of instances in the scene by means of spatial position relationships; for the spatial flow branch of the spatial layout layer, the input is the spatial feature map M encoded by the position information of all <person, object> pairs output by object detection. sp ; The spatial feature map M sp is encoded as follows: the pixel values within the bounding box of the instance are set to 1, and other values outside the bounding box of the instance are set to 0 in two channels; for a given pair of human object bounding boxes, its spatial interaction map is defined as a binary image with two channels: the first channel corresponds to the binary pattern of the human, and the second channel corresponds to the object; enabling the neural network to learn two-dimensional filters to respond to two-dimensional human spatial interaction patterns; using two max-pooling convolutional layers and two fully connected layers to extract the feature f of the spatial layout flow sp , which participates in the final interaction category classification, described as follows: Among them, and represents the fully connected layer parameter matrix, f cnn represents the convolution operation; and the Sigmoid non-linear activation function is used to perform human-object interaction classification on the <human-object> pair spatial features, f sp represents the spatial feature vector, S sp represents the probability scores of the spatial flow features on each interaction category.

3. The hierarchical human interaction detection method based on human interaction intention information according to claim 1, characterized in that, The objective appearance layer contains a human flow branch and a logistics flow branch, providing pixel-level appearance prediction information at the micro interaction level; a residual module containing global average pooling is used to extract the visual features f of people and objects from the global appearance features h and f o , and the extracted F h and F o are scaled to a fixed size of p×p; and after feature enhancement through two fully connected layers, the probability scores S h and S o of the human flow features and the logistics flow features in the interaction category are calculated, which is described by the formula as follows: S h = Sigmoid(W h f h ) (4) S o = Sigmoid(W o f o ) (5) Among them, formulas (4) and (5) respectively represent the operations of two fully connected layers, W h and W o represent the parameter matrices of the fully connected layers, and f h and f o represent the visual features of humans and objects respectively.

4. The hierarchical human interaction detection method based on human interaction intention information according to claim 1, characterized in that, The interaction intention layer mines mesoscopic perspective information, constructs an interaction intention layer driven by human interaction intention, and provides a computational perspective to utilize three forms of human intention visually available, which consists of the following three branches: (1) Human gaze branch: The human gaze area is obtained using a pre-trained two-stream model. The pre-trained gaze prediction model takes the input image I and the position of the human eye center calculated by the human pose estimation network as input, and outputs a gaze probability density map G of a fixed area. The word order network consists of a saliency path and a gaze path. The gaze path can only access close-up images of the person’s head and its position, and obtains a heat map M (x h ,x p ); the saliency path transforms the complete image x i As input, and obtain another heat map H(x i ); combining these two results with the element-wise product yields the following equation: where x i is the input image, x h is the close-up image of the cropped human feature, x p is the quantized spatial position of the human head, and G is the output fixation probability density map; For each human instance in the image, select k candidate object regions b = (b 1 ,..., b k ); for each candidate region b ∈ b, calculate its fixation weight g b ; where g b is obtained by summing the values of the density map G in b and then normalizing by the area of b: Among them, area b represents the candidate region b, G x,y is the fixation probability density map obtained for this region; Then, select the region R with the maximum g from the k candidate regions as the region of human fixation; for the selected region R, we define its corresponding feature vector f b ={f g , f a , f l , f c}; where f a is the appearance vector of object b, f l is a four-dimensional vector containing l x , l y , l w , l h , where l x , l y specify the horizontal and vertical coordinate distances of the object bounding box, l w , l h are the height and width in the specified logarithmic space, and f c is the target classifier feature vector; (2) Human pose branch: Utilize human pose estimation to estimate 17 human key points of the human body, connect the 17 key points with lines of different gray values, and set other areas to 0 to construct a pose map; use two max-pooling convolutional layers and two fully connected layers to connect the pose map with the layout map of the person and the object to extract the features f of the spatial pose flow sp ; The human gaze feature f g , the spatial pose feature f sp , the human appearance feature f h and the object appearance feature f o are connected to obtain the joint vector f hol , and then the prediction score S j in the interaction judgment stage is set to: Among them, represents a fully connected operation; In the two-dimensional probability vector output at this stage, the first dimension is the probability of the existence of an interaction, and the second dimension is the probability of the non-existence of an interaction; a threshold δ is used to define whether an interaction exists; if the probability value is higher than δ, there is an interaction, otherwise there is no interaction; for a given pair of human objects (b h , b o ), it is first necessary to determine whether an interaction exists; only the person-object pairs judged to have an interaction can enter the following classification stage; (3) Body key part branches: Construct 2-channel distance maps between each part of the body and the object respectively, and define the position coordinates of the center point in the global features of the object as <h x ,h y >, define humans as <o x ,o y >, where x and y are the coordinates in the x and y directions respectively; set the length and width of the global appearance features as h and W; then, define two position vectors a and b, vector a starts from <o x ,o y > and is defined as <h x ,h y >; vector b is defined as each pixel <h x ,h y > in the human map; Construct a 2-channel distance map with dimensions of H×W; Use cosine distance to reflect the relative difference in direction between the body and the object. The pixel value of the human body box in the first channel is the cosine distance between vector a and vector b; Cosine distance cannot distinguish the distance between vectors in the same direction, so Euclidean distance is introduced to capture the absolute distance difference between two vectors; The pixel value of the human body box in the second channel is the Euclidean distance between vector a and vector b; Set the pixel values outside the human body box of the two channels to 0; Use the 2-channel distance map to model the positional distance relationship between human body parts and objects; After obtaining the distance map, concatenate the constructed distance map with the global appearance feature map in the channel dimension to obtain a distance-feature map; Use existing pose estimation methods to obtain the positions of key points, and then determine the corresponding human body parts based on the key points; for each human key point, construct 17 rectangles with a size of 1 / 10 of the area of the original input image centered around the key point, and the obtained rectangular regions are used as the regions where the body parts are located; after obtaining the bounding boxes representing the regions of each part of the human body, based on the bounding boxes of each body part, use the region of interest pooling operation to extract the corresponding regions from the distance-feature map and scale them to q×q; in order to combine the influence of each body part on the interaction detection, connect the distance-feature maps of all the human body parts obtained, and use a fully connected layer to convert them into a feature vector f of the key parts of the human body part .

5. The hierarchical human interaction detection method based on human interaction intention information according to claim 4, characterized in that, It should be considered separately when calculating the loss. Considering that the classification loss usually uses the cross-entropy loss function Calculate: Among them, Y ij is the real action label, S ij is the predicted action score, is the average value of M batches of samples, i represents the interaction candidate, and j represents the interaction candidate object; Therefore, the loss in the interaction judgment stage is calculated according to the following formula: Among them, Y J and S J are the final predicted scores of the true label and interaction judgment module, is the cross-entropy loss; considering the influence of each stream and making the loss function converge better, the loss of each stream and the overall loss in the interaction detection stage are added, and the overall loss function of the interaction classification module is expressed as follows: where α and β are the branch loss coefficient and the total loss coefficient of all branches, Y C is the label of the interaction classification stage, S h is the probability score of the human body appearance branch, S o is the probability score of the object appearance branch, S p is the probability score of the human body pose and space joint branch, S part is the probability score of the human body key part branch, S c is the total probability score of the interaction classification stage; when performing joint training, the total loss function of the two stages is used to update the parameters:

6. The hierarchical human interaction detection method based on human interaction intention information according to claim 5, characterized in that, Interaction score fusion: Since the feature vector f of the body part stream part not only contains the fine features of the human body but also reflects the positional relationship between the human and the object, it is regarded as an independent stream; a late fusion strategy is used to fuse the four streams; a fully connected layer is used to map the feature vector of each stream to the interaction prediction probability score S h 、S o 、S p and S part , and the prediction scores of each stream are fused; first, the probability scores S of the human appearance branch h , the probability scores S of the object appearance branch o and the probability scores S of the human key part branch part are fused, and then we multiply it by the probability scores S of the human pose and spatial joint branch p ; therefore, the final interaction prediction score vector S of the final detection result c is set to: Among them, the operations and respectively represent element-wise sum multiplication; the combination of the hierarchical idea and the two-stage strategy can better utilize the features at each pixel level and improve the effect of interactive detection.

Citation Information

Patent Citations

  • Method and device for detecting human-object interaction relationship in video

    CN112464875A

  • Single-view human-object interaction identification method of three-stage network framework

    CN113887468A