Person Interaction Detection Method, Device, Equipment and Medium Based on Cascade Parallel
By adopting a cascading parallel detection method in character interaction detection, the feature learning paths of instance detection and interactive classification tasks are separated and optimized, and interactive feature expression is enhanced through multiple relationship construction modules, the problems of inefficiency, multi-task learning conflict and model optimization in the existing technology are solved, and efficient and accurate character interaction detection is achieved.
Patent Information
- Application Number
- CN202510336186.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The prior art has problems such as inefficiency, multi-task learning conflicts and model optimization difficulties in character interaction detection tasks.
The character interaction detection method based on cascading parallelism is adopted, and the feature learning paths of the instance detection and interaction classification tasks are separated through iterative output modules, query vector modules, human body detection modules, object detection modules, initial interaction modules, relationship construction modules, enhancement modules and final interaction modules, and the multiple relationship characteristics between people, objects and interactions are mined through multiple relationship construction modules.
Effectively separate and optimize the feature learning paths of instance detection and interactive classification tasks, significantly reduce inference time, improve detection accuracy, enhance the semantic expression ability of interactive behavior, and improve classification accuracy in complex interactive scenarios.
Smart Images

Figure CN119851318B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and more particularly, to a method, apparatus, device, and medium for human interaction detection based on cascaded parallelism. Background Art
[0002] In the field of artificial intelligence, human interaction detection technology has extensive application requirements in multiple fields such as video analysis, intelligent monitoring, human-computer interaction, and behavior recognition. The key technology in these application scenarios lies in whether human interaction behavior can be accurately and efficiently recognized. Accurate human interaction detection enables the computer vision system to better understand the scene semantics, thereby improving the intelligence and practicality of the system.
[0003] To solve the problem of detecting the interaction relationship between the human body and objects, the existing technologies are mainly divided into two-stage detection methods and one-stage detection methods. The two-stage detection method adopts a strategy of first detecting and then classifying. A large number of negative samples are screened out through instance detection, and then the remaining positive samples are classified for interaction behavior. Although this method can effectively utilize additional context information to improve the detection results, its serial structure leads to low efficiency, a large amount of computing resources are wasted in the generated proposals, and the inference time is also long. The one-stage detection method regards the human interaction detection task as a whole, and performs instance detection and interaction classification in real time through a parallel detector, and uses a special post-processing method to merge the prediction results. However, this method faces challenges in multi-task learning. Due to the significant differences in visual features required for human-object detection and interaction classification tasks, the performance of the model is easily restricted. In addition, traditional Transformer-based object detection models regard object queries as a set of learnable embeddings, but these embeddings usually lack clear physical meanings and cannot intuitively explain the positions that the detector should focus on, resulting in increased difficulty in the model optimization process.
[0004] In summary, the existing technologies have problems such as low efficiency, multi-task learning conflicts, and difficulty in model optimization in the human interaction detection task, which limit the performance of human interaction detection technology in practical applications. Summary of the Invention
[0005] The present invention provides a method, apparatus, device, and medium for human interaction detection based on cascaded parallelism to improve at least one of the above technical problems.
[0006] In a first aspect, the present invention provides a method for human interaction detection based on cascaded parallelism, which includes:
[0007] S01. Obtain an image to be recognized.
[0008] S02. Extract image features from the image to be recognized and generate position information of the image features.
[0009] S03. Iteration Perform the following steps multiple times, and predict the person interaction detection result based on the human feature vector, object feature vector, and interaction feature vector output in the last iteration.
[0010] S04. Generate a human query vector and an object query vector based on the image features and the prior bounding boxes of the person pair to obtain a query vector group. Among them, the prior bounding boxes of the person pair are obtained by random initialization in the first iteration, and are updated based on the human decoded features and object decoded features of the previous iteration in subsequent iterations.
[0011] S05. Decode the human decoded features through a human detector based on the image features, the position information, and the human query vector.
[0012] S06. Update the object query vector based on the human decoded features, and then decode the object decoded features through an object detector based on the updated object query vector, the image features, and the position information.
[0013] S07. Input the image features, the human decoded features, and the object decoded features into a classification decoder to obtain initial interaction decoded features.
[0014] S08. Input the human decoded features, the object decoded features, and the initial interaction decoded features into a multi-relationship construction module to obtain unary relationship features, pairwise relationship features, and ternary relationship features.
[0015] S09. Embed the unary relationship features, the pairwise relationship features, and the image features into the ternary relationship features through an attention module for fusion.
[0016] S10. Enhance the initial interaction decoded features through the fused ternary relationship features to obtain an interaction feature vector.
[0017] In an optional implementation, the prior bounding boxes of the person pair include a human prior bounding box and an object prior bounding box. The query generator, the human detector, and the object detector are all stacked, and the number of stacked layers is layers.
[0018] In an optional implementation, step S04 specifically includes steps S041 to S046.
[0019] S041. Obtain human region features and object region features from the image features based on the human prior bounding box and the object prior bounding box, and merge the two region features according to the feature dimension to obtain person pair region features.
[0020] S042. Predict the human key points and object key points respectively through a convolutional layer, a multi-layer perceptron, and an activation function according to the regional characteristics of the person pair.
[0021] S043. Perform bilinear interpolation sampling operations on the person pair regional features according to the human key points and the object key points, and extract the human local visual features and object local visual features respectively.
[0022] S044. Aggregate the human local visual features and the object local visual features through a fully connected layer to obtain the human visual features and object visual features.
[0023] S045. Obtain the position embedding corresponding to each query according to the sine position encoding of the person pair prior bounding box.
[0024] S046. Obtain the initial query vector of the query generator of the current layer, and fuse the position embedding to obtain the final query vector of the current layer. Among them, the initial query vector of the query generator of the first layer uses the human visual features and the object visual features. The initial query vectors of the query generators of the remaining layers need to fuse the feature vector output by the detector of the previous layer with the human visual features and the object visual features.
[0025] In an optional implementation manner, the calculation model of the final query vector is:
[0026] .
[0027] .
[0028] In the formula, is the serial number of the layer, is the final query vector of the human body of the th layer, is the final query vector of the object of the th layer, is the initial query vector of the human body detector of the th layer, is the initial query vector of the object detector of the th layer, is the human body position embedding, is the object position embedding, is the th layer of the feature vector output by the human body detector, is the th layer of the feature vector output by the object detector.
[0029] In an alternative embodiment, both the human body detector and the object detector are stacked, and the number of stacked layers is layers. Each layer of the human body detector and each layer of the object detector is a Transformer decoder.
[0030] In an alternative embodiment, the model decoded by the human body detector is:
[0031] .
[0032] In the formula, is the feature vector output by the -th layer of the human body detector, is the -th layer of the human body detector, is the image feature, is the position information, is the final query vector of the human body for the -th layer.
[0033] In an alternative embodiment, the update model for updating the object query vector according to the decoded human body feature is:
[0034] .
[0035] In the formula, is the updated object query vector for the -th layer, is the final query vector of the object for the -th layer, is the feature vector output by the -th layer of the human body detector.
[0036] In an alternative embodiment, the model decoded by the object detector is:
[0037] .
[0038] In the formula, is the feature vector output by the -th layer of the object detector, is the -th layer of the object detector, is the image feature, is the position information, is the updated object query vector for the -th layer.
[0039] In an alternative embodiment, the number of stacked layers of the interaction classifier is layers, and each layer of the interaction classifier includes a classification decoder.
[0040] In an alternative embodiment, the classification decoder model is:
[0041] .
[0042] .
[0043] Wherein, is the number of stacked layers, is the serial number of the layer, is the classification decoding output by the classification decoder of the th layer, is the classification decoder of the th layer, is the image feature, is the position information, is the person pair decoding feature obtained by fusing and adding the human body decoding feature and the object decoding feature, is the feature vector output by the th layer decoder of the human body detector, is the feature vector output by the th layer decoder of the object detector, is the th layer of interactive decoding feature.
[0044] In an alternative embodiment, step S08 specifically includes steps S081 to S083.
[0045] S081. Respectively use the human body decoding feature, the object decoding feature, and the initial interactive decoding feature as unary relationship features.
[0046] S082. Combine the human body decoding feature, the object decoding feature, and the initial interactive decoding feature pairwise and merge them in the channel dimension, and then perform dimensionality reduction through a linear layer to obtain the pairwise relationship feature.
[0047] S083. Combine the human body decoding feature, the object decoding feature, and the initial interactive decoding feature as a ternary relationship feature.
[0048] The pairwise relationship feature model is:
[0049] .
[0050] .
[0051] .
[0052] The ternary relationship feature model is:
[0053] .
[0054] Wherein, is the paired relationship feature of the human body decoding feature and the object decoding feature of the th layer, is the paired relationship feature of the human body decoding feature and the initial interaction decoding feature of the th layer, is the paired relationship feature of the object decoding feature and the initial interaction decoding feature of the th layer, is a linear layer, is the feature vector output by the th layer decoder of the human body detector, is the feature vector output by the th layer decoder of the object detector, is the initial interaction decoding feature output by the th layer classification decoder, is the ternary relationship feature of the th layer, represents a merging operation.
[0055] In an optional implementation manner, step S09 specifically includes steps S091 to S095.
[0056] S091. Obtain the context content of the unary relationship feature through the self-attention mechanism. Wherein, . Wherein, is the context content, represents the self-attention mechanism.
[0057] S092. Use the cross-attention mechanism to embed the context content into the ternary relationship feature to obtain the first embedded feature. Wherein, . Wherein, is the first embedded feature, represents the cross-attention mechanism, is the ternary relationship feature of the th layer.
[0058] S093. Obtain the paired relationship of the paired relationship feature through the self-attention mechanism. Wherein, . Wherein, is the paired relationship.
[0059] S094. Use the cross-attention mechanism to embed the paired relationship into the first embedded feature to obtain the second embedded feature. Wherein, . Wherein, is the second embedded feature.
[0060] S095. Embed the image features into the second embedded features using the cross-attention mechanism to obtain enhanced ternary relationship features. Among them, . In the formula, is the enhanced ternary relationship feature, is the image feature.
[0061] In an alternative embodiment, step S10 is specifically: Convert the enhanced ternary relationship feature and the initial interactive decoding feature through a linear layer. The conversion model is:
[0062] .
[0063] .
[0064] In the formula, is the interactive feature vector of the th layer, is the initial interactive decoding feature output by the classification decoder of the th layer, is the conversion coefficient, represents element-wise multiplication, is the enhanced ternary relationship feature, is the activation function.
[0065] In an alternative embodiment, step S02 specifically includes steps S021 to S023.
[0066] S021. Extract a convolutional feature map from the image to be recognized through a convolutional neural network module.
[0067] S022. Generate fixed-position encoding of the same size as the convolutional feature map as the position information of the feature map.
[0068] S023. Input the position information and the convolutional feature Figure 1 into the Transformer encoder together to obtain the encoded image features.
[0069] In an alternative embodiment, step S03 is specifically: The feature vector output by the last layer of the decoder is predicted through a head forward network to obtain the final human interaction detection result.
[0070] .
[0071] .
[0072] .
[0073] 。
[0074] In the formula, is the human body bounding box, is the forward network for prediction, is the human body feature vector output by the last layer of the human body detector, is the human body position information obtained by the last iteration update, is the object bounding box, is the forward network for prediction, is the object feature vector output by the last layer of the object detector, is the object position information obtained by the last iteration update, is the object category, is the forward network for prediction, is the interaction category, is the forward network for prediction, is the interaction decoding feature output by the last layer of the interaction classifier, is the activation function, is the activation function.
[0075] In a second aspect, the present invention provides a person interaction detection device based on cascaded parallelism, which includes an image acquisition module, a feature extraction module, an iterative output module, a query vector module, a human body detection module, an object detection module, an initial interaction module, a relationship construction module, an enhancement module, and a final interaction module.
[0076] The image acquisition module is used to acquire the image to be recognized.
[0077] The feature extraction module is used to extract image features according to the image to be recognized and generate the position information of the image features.
[0078] The iterative output module is used to iterate less than
[0079] times the following steps, and predict the person interaction detection result according to the human body feature vector, object feature vector, and interaction feature vector output by the last iteration.
[0080] The human body detection module is used to decode the human body detection features through a human body detector according to the image features, the position information, and the human body query vector.
[0081] The object detection module is used to update the object query vector according to the human body detection features, and then decode the object detection features through an object detector according to the updated object query vector, the image features, and the position information.
[0082] The initial interaction module is used to input the image features, the human body detection features, and the object detection features into a classification decoder to obtain initial interaction decoding features.
[0083] The relationship construction module is used to input the human body detection features, the object detection features, and the initial interaction decoding features into a multiple relationship construction module to obtain unary relationship features, pairwise relationship features, and ternary relationship features.
[0084] The enhancement module is used to fuse the unary relationship features, the pairwise relationship features, and the image features into the ternary relationship features through an attention module.
[0085] The final interaction module is used to enhance the initial interaction decoding features through the fused ternary relationship features to obtain an interaction feature vector.
[0086] In a third aspect, the present invention provides a person interaction detection device based on cascaded parallelism, which includes a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement a person interaction detection method based on cascaded parallelism as described in any paragraph of the first aspect.
[0087] In a fourth aspect, the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute a person interaction detection method based on cascaded parallelism as described in any paragraph of the first aspect.
[0088] By adopting the above technical solutions, the present invention can achieve the following technical effects:
[0089] The inventive solution effectively separates the feature learning paths of instance detection and interaction classification tasks through a cascaded parallel decoding architecture, enabling each subtask to be independently optimized and reducing interference between tasks. This architecture combines the feature separation advantages of two-stage methods with the parallel processing efficiency of one-stage methods, significantly reducing the inference time while improving the detection accuracy, thus achieving a balance between model performance and computational resource consumption.
[0090] In addition, the query optimization generation mechanism provides the decoder with query vectors having clear physical meanings, guiding the model to focus on the target area, reducing the interference of irrelevant areas, and further accelerating the model convergence process.
[0091] By introducing an interactive content context learning module, this solution fully explores the multiple relationship features among people, objects, and interactions, enhancing the semantic expression ability of interactive behaviors. This module fuses global encoding features and local fine-grained context information, strengthens the associative reasoning between triples, and effectively improves the classification accuracy in complex interactive scenarios. Experimental results show that the performance indicators of this method on mainstream data sets are superior to existing technical solutions, verifying its effectiveness and robustness in practical applications. Brief Description of the Drawings
[0092] Figure 1 is a schematic flow diagram of a person interaction detection method based on cascade parallelism.
[0093] Figure 2 is a schematic diagram of the detection results of a person interaction detection method based on cascade parallelism.
[0094] Figure 3 is a network structure diagram of a person interaction detection method based on cascade parallelism.
[0095] Figure 4 is a network structure diagram of a feature extractor.
[0096] Figure 5 is a connection relationship diagram between a query generator, a human body detector, an object detector, and an interaction classifier.
[0097] Figure 6 is a network structure diagram of a query generator.
[0098] Figure 7 is a network structure diagram of an interaction classifier.
[0099] Figure 8 is a network structure diagram of a multiple relationship construction module. Detailed Embodiments
[0100] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0101] Embodiment 1. Please refer to Figures 1 to 8 , the embodiment of the present invention provides a person interaction detection method based on cascade parallelism, as shown in Figure 3As shown in the figure, the network structure for implementing the person interaction detection method based on cascaded parallel mainly includes five modules: a feature extractor, a query generator, a human body detector, an object detector, and an interaction classifier. Among them, the query generator, the human body detector, the object detector, and the interaction classifier are all stacked, and the number of stacked layers is layers. is the layer number. Each layer of the interaction classifier contains a classification decoder. Therefore, the classification decoders are also stacked with layers.
[0102] Each layer of the human body detector and each layer of the object detector are a Transformer decoder. Therefore, the human body detector is also called the human body decoder, and the object detector is also called the object decoder.
[0103] The data processing flow of the person interaction detection method based on cascaded parallel is outlined as follows.
[0104] First, for a given image , first obtain the image features through the feature extractor . At the same time, randomly initialize a set of human body prior bounding boxes and object prior bounding boxes , which are collectively called person pair prior bounding boxes .
[0105] Then, the query generator, the human body detector, the object detector, and the interaction classifier are iterated to obtain the human body decoded features, the object decoded features, and the interaction decoded features. During the iteration process, the image features , the human body prior bounding boxes and the object prior bounding boxes are jointly input into the query generator. The query generator generates the human body query vectors and the object query vectors for each layer of the human body detector and the object detector, as well as the human body position embedding and the object position embedding . Then, the human body detector and the object detector generate the corresponding human body feature vectors and the object feature vectors according to the input image features , the query vectors and their position embeddings. Then, through the human body feature vectors and the object feature vectors , update the prior bounding boxes and obtain the decoded interaction feature vectors in the interaction classifier. Among them, the interaction feature vectors output by the last layer are defined as the interaction decoded features.
[0106] Finally, the interactive decoding features are processed through a classification head to obtain a prediction result. Among them, the interactive decoding features output by the last-layer interactive classifier are used to predict the interaction category, and the human feature vector output by the last-layer human detector is used to predict the human position, and the feature vector output by the last-layer object detector is used to predict the position and category of the object.
[0107] The data processing flow of the person interaction detection method based on cascaded parallel is introduced in detail below. It can be executed by a person interaction detection device based on cascaded parallel (hereinafter referred to as: person interaction detection device). In particular, it is executed by one or more processors in the person interaction detection device to implement steps S01 to S10.
[0108] S01. Obtain an image to be recognized.
[0109] Specifically, the image to be recognized is used for human-object interaction detection. The concept of human-object interaction (HOI, Human-Object Interaction) detection is proposed to solve the problem of detecting the interaction relationship between a human body and an object. HOI detection aims to conduct a more in-depth analysis of the human-object behavior in the image and learn how humans interact with the environment. These interaction relationships reflect human behavior and scene semantics, and are of great significance for improving the intelligence and practicality of computer vision systems.
[0110] It can be understood that the person interaction detection device can be an electronic device with computing performance such as a portable notebook computer, a desktop computer, a server, a smart phone, or a tablet computer.
[0111] S02. Extract image features from the image to be recognized and generate position information of the image features. In an optional implementation manner, step S02 specifically includes steps S021 to S023.
[0112] S021. Extract a convolutional feature map from the image to be recognized through a convolutional neural network module.
[0113] S022. Generate a fixed position encoding of the same size as the convolutional feature map as the position information of the feature map.
[0114] S023. Input the position information and the convolutional feature Figure 1 into a Transformer encoder together to obtain the encoded image features.
[0115] Specifically, as Figure 4As shown, the feature extractor part consists of a convolutional neural network module and a Transformer encoder, which is used to extract image features. For an input image , it will first pass through the convolutional neural network module to obtain a convolutional feature map . Then, for the convolutional feature map , fixed-position encodings of the same size are generated , as the spatial position information of the convolutional feature map , and they are input into the Transformer encoder together. Finally, the encoded image features are obtained.
[0116] S03. Iterate steps S04 to S10 times, and predict the human-object interaction detection result based on the human feature vector, object feature vector, and interaction feature vector output in the last iteration. Among them, .
[0117] There are some similarities between the existing state-of-the-art methods for two-stage and one-stage HOI detection (CDN, Mining the Benefits of two-stage and One-stage HOI Detection) and the embodiments of the present invention.
[0118] Based on CDN, the embodiments of the present invention optimize the design of the query vector. Each query in the query vector group is regarded as a candidate human-object interaction detection result, thereby introducing prior human-object pair bounding boxes and generating query vectors from them. By designing query vectors with clear physical meanings, the decoder can focus on specific regions and accelerate the convergence speed of the model.
[0119] Based on CDN, the embodiments of the present invention further optimize the sub-task division of HOI detection, splitting a single instance decoder into a human detection decoder and an object detection decoder. The instance detection task is completed by the parallel operation of the two decoders, and the interaction classification decoder is connected in series. This architecture effectively separates the feature learning paths of object detection and interaction classification.
[0120] In addition, the embodiments of the present invention innovatively propose an interaction classification decoder, which mines the relationship context information between multiple relationships of the HOI triple, enhances the expression form of interaction features through multi-layer relationship reasoning, and improves the detection accuracy.
[0121] The human detection decoder, object detection decoder, and interaction classification decoder will be described in detail below through steps S04 to S10.
[0122] S04. Extract a human query vector and an object query vector from the prior bounding boxes according to the image features and the person, and obtain a query vector group.
[0123] Among them, the prior bounding box of the person is obtained by random initialization in the first iteration. In the subsequent iteration process, the prior bounding box of the person is updated according to the human decoded features and the object decoded features of the previous iteration. The prior bounding box of the person includes a human prior bounding box and an object prior bounding box.
[0124] As Figure 5 and Figure 6 shown, this embodiment uses a query generator to generate two groups of query vectors input to each layer of the human detector and the object detector. The query optimization generator inputs and . Among them, represents the th human query vector, represents the th object query vector, is the total number of query vectors.
[0125] As Figure 6 shown, obtain the corresponding region of interest (ROIs) features according to the prior box, then predict the human key points and object key points from the region features, then sample the visual features using the key points, and finally fuse the visual features of the points to generate the human query vector and the object query vector . In addition, since the generated query vectors come from the prior bounding boxes, the length of the query vector group depends on the number of prior bounding boxes.
[0126] In an alternative embodiment, step S04 specifically includes steps S041 to S046.
[0127] S041. According to the human prior bounding box and the object prior bounding box, obtain the human region feature and the object region feature from the image features, and merge the two region features according to the feature dimension to obtain the person-object pair region feature.
[0128] Obtaining the region of interest (ROIs) features, according to the human prior bounding box and the object prior bounding box obtain the human region feature from the image features ) and the object region feature , and merge the two region features according to the feature dimension to obtain the person-object pair region feature .
[0129] The calculation model is:
[0130] 。
[0131] 。
[0132] 。
[0133] In the formula, is the human body region feature, is the th human body region feature corresponding to the query vector, is the object region feature, is the th object region feature corresponding to the query vector, is the total number of query vectors, is the human-object pair region feature, is the image feature, is the human body prior bounding box, is the object prior bounding box, represents the sampling operation for obtaining the region feature from the encoded image feature, is the merging operation.
[0134] S042. According to the human-object pair region feature, predict the human body key points and object key points through a convolutional layer, a multi-layer perceptron, and an activation function respectively.
[0135] Prediction of key points. For each region feature and , predict the human body key points and object key points through a convolutional layer (Conv) and a multi-layer perceptron (MLP) respectively. Among them, is the th key point of , is the th key point of , is the total number of key points.
[0136] Considering that during the feature sampling process, the coordinates of the key points need to be normalized to [-1, 1], so when outputting the key point prediction value, an activation function is used for non-linear transformation to ensure the output coordinate range.
[0137] The prediction model of the key points is:
[0138] 。
[0139] 。
[0140] Among them, represents a convolution operation, represents a multi-layer perceptron (MLP), represents an activation function. In this embodiment, the MLP is composed of three fully connected layers and is used to map the input features to two-dimensional point information.
[0141] S043. According to the human key points and the object key points, perform a bilinear interpolation sampling operation on the region features of the person-object pair, and respectively extract the human local visual features and the object local visual features.
[0142] For each instance, multiple key points will be predicted. By fusing the features sampled from these key points, the visual features of this instance are constructed. The key point feature sampling process includes using the human key points and the object key points to perform a bilinear interpolation sampling operation on the person-object pair region features so as to respectively extract the human local visual features and the object local visual features. Through this method, the local information of each key point can be effectively obtained from the region features.
[0143] The model for extracting local visual features is as follows:
[0144] ;
[0145] ;
[0146] In the formula, is the human local visual feature, is the object local visual feature, represents the sampling operation.
[0147] S044. Aggregate the human local visual features and the object local visual features through a fully connected layer to obtain the human visual features and the object visual features.
[0148] After extracting the local visual features, the local visual features in the same instance are aggregated through a fully connected layer to obtain the human visual features and the object visual features.
[0149] 。
[0150] 。
[0151] In the formula, represents the fully connected layer.
[0152] S045. Obtain the position embedding corresponding to each query according to the sine position encoding of the prior bounding box for the person.
[0153] Specifically, use the sine position encoding of the prior box as the position embedding corresponding to each query.
[0154] The calculation model of the position embedding is:
[0155] .
[0156] .
[0157] In the formula, is the human body position embedding, is the object position embedding, represents the sine embedding function, is the human body prior bounding box, is the object prior bounding box. Among them, the sine embedding function is used to convert point coordinates into feature vectors, which is the same as that used in object detection.
[0158] S046. Obtain the initial query vector of the query generator of the current layer, and fuse the position embedding to obtain the final query vector of the current layer. Among them, the initial query vector of the query generator of the first layer uses the human visual feature and the object visual feature. The initial query vectors of the query generators of the remaining layers need to fuse the feature vector output by the detector of the previous layer with the human visual feature and the object visual feature.
[0159] Since the human detector and the object detector are composed of stacked decoders, the corresponding query generators are also stacked.
[0160] Except that the query generator of the first layer directly uses the visual features of the human body and the object as the query vectors of the human detector and the object detector, the queries of each remaining layer need to fuse the decoded feature vector of the previous layer to generate the final query of the current layer , so as to retain the effective information of the previous layer. Among them, represents the th human body final query vector. represents the th object final query vector.
[0161] Preferably, the calculation model of the final query vector is:
[0162] .
[0163] .
[0164] In the formula, is the serial number of the layer, is the final query vector of the human body for the layer, is the final query vector of the object for the layer, is the initial query vector of the human body detector for the layer, is the initial query vector of the object detector for the is the human body position embedding, is the object position embedding, is the feature vector output by the human body detector for the layer, is the feature vector output by the object detector for the
[0165] The initial query vector of the current layer is
[0166]
[0167]
[0168] In the formula, and represent the initial query vectors of the human body and object detectors for the layer.
[0169] Specifically, both the human body detector and the object detector are composed of stacked Transformer decoders. The query vector group, image features, and corresponding position information generated by the query optimizer are respectively input into the corresponding decoders, and the decoders output the decoded feature vectors. The bounding box information is updated through the human body decoded feature vector and the object decoded feature vector, and is simultaneously input into the interactive context content learning module for learning of interactive features.
[0170] S05. According to the image features, the position information, and the human body query vector, the human body decoded feature is obtained through decoding by the human body detector.
[0171] Preferably, the model for decoding by the human body detector is:
[0172] .
[0173] In the formula, is the feature vector output by the human body detector for the layer, is the human body detector for the is the image feature, is the position information, is the final query vector of the human body for the layer.
[0174] S06. Update the object query vector according to the human body decoding feature, and then decode the object decoding feature through an object detector according to the updated object query vector, the image feature, and the position information.
[0175] The update model for updating the object query vector according to the human body decoding feature is:
[0176] .
[0177] The model for decoding through the object detector is:
[0178] .
[0179] In the formula, is the updated object query vector for the layer, is the final query vector of the object for the layer, is the feature vector output by the human body detector for the layer, is the feature vector output by the object detector for the layer, is the object detector for the layer, is the image feature, is the position information, is the updated object query vector for the layer.
[0180] Specifically, after each layer of the human body detector and the object detector outputs a feature vector, it will be used to update the bounding box of the prior human-object pair:
[0181] .
[0182] .
[0183] In the formula, is the prior bounding box of the human body for the layer, is the feature vector output by the human body detector for the layer, is the prior bounding box of the human body for the layer, is the prior bounding box of the object for the layer, is the feature vector output by the object detector for the layer, For the prior bounding box of the object at the th layer, For the activation function, For the forward network for prediction and For the forward network for prediction .
[0184] Considering the key role of the relationship context among people, objects and interactions in the HOI instance detection, this embodiment designs an interaction content context learning module (i.e., interaction classification module). This module mines the multiple relationship contexts in the triple to enhance the expression ability of the interaction features. The multiple relationships include unary, pairwise and triple relationship contexts, where the unary relationships are person, object, interaction features, the pairwise relationships are person-object, person-interaction, object-interaction features, and the triple relationships cover person-object-interaction features.
[0185] The unary and pairwise relationships focus on the task understanding at the fine-grained level and provide the fine-grained context information of the HOI instance. While the triple relationship provides the context information of the overall person, thus enabling the overall-level understanding of the HOI instance. To make full use of these overall and fine-grained context information, the unary and pairwise relationships are embedded into the context of the triple relationship to effectively generate enhanced interaction features. Its specific structure is as Figure 7 shown.
[0186] First, the decoded features output by the human body detector and object detector at the jth layer and the image encoded features are input into the classification decoder module to obtain the initial interaction decoded features (i.e., step S07). Then, the human body decoded features, object decoded features and initial interaction decoded features are input into the multiple relationship construction module to generate unary relationship features, pairwise relationship features and triple relationship features (i.e., step S08 and step S09), as Figure 7 shown.
[0187] Next, the context information of the unary relationship is extracted through the self-attention mechanism, and the unary relationship context is embedded into the triple relationship features by using the cross-attention mechanism. Similarly, the context content of the pairwise relationship is extracted by using the self-attention mechanism and embedded into the triple relationship context through the cross-attention. In addition, the global image encoded features are also embedded into the triple relationship context through the cross-attention calculation. Finally, by using the enhanced triple relationship context information, the expression ability of the interaction features is strengthened.
[0188] S07. Input the image features, the human body decoded features and the object decoded features into the classification decoder to obtain the initial interaction decoded features.
[0189] Preferably, the classification decoder model is:
[0190] 。
[0191] 。
[0192] Wherein, is the number of stacked layers, is the serial number of the layer, is the classification decoding output by the classification decoder of the th layer, is the classification decoder of the th layer, is the image feature, is the position information, is the person pair decoding feature obtained by fusing and adding the human body decoding feature and the object decoding feature, is the feature vector output by the th layer decoder of the human body detector, is the feature vector output by the th layer decoder of the object detector, is the interactive decoding feature of the th layer.
[0193] S08. Input the human body decoding feature, the object decoding feature and the initial interactive decoding feature into the multi-relationship construction module to obtain the unary relationship feature, the pairwise relationship feature and the ternary relationship feature.
[0194] Preferably, merging is performed in the channel dimension, and the dimension is reduced from a high dimension to a low dimension through an MLP (multiple linear layers). Step S08 specifically includes steps S081 to S083.
[0195] S081. Respectively use the human body decoding feature, the object decoding feature and the initial interactive decoding feature as the unary relationship features. As Figure 8 shown.
[0196] S082. Combine the human body decoding feature, the object decoding feature and the initial interactive decoding feature pairwise, merge them in the channel dimension, and then perform dimensionality reduction through a linear layer to obtain the pairwise relationship feature.
[0197] Preferably, the pairwise relationship feature model is:
[0198] 。
[0199] 。
[0200] 。
[0201] In the formula, is the paired relationship feature of the human body decoding feature and the object decoding feature of the th layer, is the paired relationship feature of the human body decoding feature and the initial interaction decoding feature of the th layer, is the paired relationship feature of the object decoding feature and the initial interaction decoding feature of the th layer, is a linear layer, is the feature vector output by the th layer decoder of the human body detector, is the feature vector output by the th layer decoder of the object detector, is the initial interaction decoding feature output by the th layer classification decoder.
[0202] S083. Combine the human body decoding feature, the object decoding feature, and the initial interaction decoding feature as a ternary relationship feature.
[0203] Preferably, the ternary relationship feature model is:
[0204] .
[0205] In the formula, is the ternary relationship feature of the th layer, is a linear layer, represents a merging operation, is the feature vector output by the th layer decoder of the human body detector, is the feature vector output by the th layer decoder of the object detector, is the th layer initial interaction decoding feature output by the classification decoder.
[0206] S09. Embed the unary relationship feature, the paired relationship feature, and the image feature into the ternary relationship feature through an attention module for fusion.
[0207] In an alternative embodiment, as Figure 7 shown, step S09 specifically includes steps S091 to S095.
[0208] S091. Obtain the context content of the unary relationship feature through a self-attention mechanism.
[0209] .
[0210] In the formula, is the context content, indicating the self-attention mechanism.
[0211] S092. Use the cross-attention mechanism to embed the context content into the ternary relationship feature to obtain the first embedded feature.
[0212] .
[0213] In the formula, is the first embedded feature, indicating the cross-attention mechanism, is the layer's ternary relationship feature.
[0214] S093. Obtain the pairwise relationship of the pairwise relationship feature through the self-attention mechanism.
[0215] .
[0216] In the formula, is the pairwise relationship.
[0217] S094. Use the cross-attention mechanism to embed the pairwise relationship into the first embedded feature to obtain the second embedded feature.
[0218] .
[0219] In the formula, is the second embedded feature.
[0220] S095. Use the cross-attention mechanism to embed the image feature into the second embedded feature to obtain the enhanced ternary relationship feature.
[0221] .
[0222] In the formula, is the enhanced ternary relationship feature, is the image feature.
[0223] S10. Enhance the initial interactive decoding feature through the fused ternary relationship feature to obtain the interactive feature vector.
[0224] Preferably, step S10 is specifically: through The linear layer converts the enhanced ternary relationship feature and the initial interactive decoding feature, and selects the necessary context information for them to generate rich interactive features.
[0225] The conversion model is:
[0226] .
[0227] 。
[0228] In the formula, is the interaction feature vector of the th layer, is the initial interaction decoding feature output by the th layer classification decoder, is the conversion coefficient, represents element-wise multiplication, is the enhanced ternary relationship feature, is the activation function.
[0229] Based on the above embodiments, in an optional embodiment of the present invention, as Figure 3 shown, after iterating times from step S04 to step S10, step S03 is specifically: according to the feature vector output by the last layer decoder, perform prediction through the head forward network to obtain the final human interaction detection result.
[0230] The prediction model based on the forward network is:
[0231] 。
[0232] 。
[0233] 。
[0234] 。
[0235] In the formula, is the activation function, is the activation function,
[0236] is the human body bounding box, is the forward network for predicting , is the human body feature vector output by the last layer human body detector, is the human body position information obtained by the last iteration update, is the object bounding box, is the forward network for predicting , is the object feature vector output by the last layer object detector, is the object position information obtained by the last iteration update, is the object category, is the forward network for predicting The forward network of is the interaction category, is for prediction the forward network of is the interaction decoding feature output by the last-layer interaction classifier.
[0237] Specifically, the , , output by the last layer will pass through the head forward network for prediction to obtain the human body bounding box , object bounding box , object category and interaction category of the person interaction prediction result.
[0238] The person interaction detection method based on cascade parallel of the present invention effectively solves some problems existing in the person interaction detection method based on the encoder-decoder structure, and realizes efficient person interaction behavior detection. It has the following advantages.
[0239] 1. Design a query optimization generator to generate query vectors related to the target region, guide the decoder to focus on specific regions of the image, reduce the interference of irrelevant regions to the decoder, thereby improving the training efficiency of the model and accelerating the convergence process. At the same time, for the query vectors of each subtask, a forward-guided query mechanism is adopted to enhance the correlation between the three subtask branches of HOI, thereby promoting the information flow and mutual cooperation between tasks.
[0240] 2. Design a cascade parallel decoding Transformer architecture to effectively separate the learning of instance features and interaction features. This architecture can not only enable each subtask to focus on the current stage, but also speed up the inference speed, making full use of the advantages of the two-stage method and the one-stage method based on Transformer.
[0241] 3. Design an interaction context content learning module to extract overall and fine-grained context information by mining the multiple relationship contexts in the triples, effectively enhancing the expression ability of interaction features. Through rich context reasoning, this module effectively strengthens the connection between different tasks, promotes the better fusion of interaction information, and improves the accuracy of interaction action classification.
[0242] Embodiment 2: The embodiment of the present invention provides a person interaction detection device based on cascade parallel, which includes an image acquisition module, a feature extraction module, an iterative output module, a query vector module, a human body detection module, an object detection module, an initial interaction module, a relationship construction module, an enhancement module, and a final interaction module.
[0243] The image acquisition module is used to acquire the image to be recognized.
[0244] A feature extraction module, configured to extract image features according to the image to be recognized, and generate position information of the image features.
[0245] An iterative output module, configured to iterate the following steps for a certain number of times or less, and predict a person interaction detection result according to the human feature vector, object feature vector, and interaction feature vector output in the last iteration.
[0246] A query vector module, configured to generate a human query vector and an object query vector according to the image features and a prior bounding box of a person pair, to obtain a query vector group. Wherein, the prior bounding box of the person pair is obtained by random initialization in the first iteration, and is updated according to the human decoded features and object decoded features of the previous iteration in subsequent iterations.
[0247] A human detection module, configured to decode the human decoded features through a human detector according to the image features, the position information, and the human query vector.
[0248] An object detection module, configured to update the object query vector according to the human decoded features, and then decode the object decoded features through an object detector according to the updated object query vector, the image features, and the position information.
[0249] An initial interaction module, configured to input the image features, the human decoded features, and the object decoded features into a classification decoder to obtain initial interaction decoded features.
[0250] A relationship construction module, configured to input the human decoded features, the object decoded features, and the initial interaction decoded features into a multiple relationship construction module to obtain unary relationship features, pairwise relationship features, and ternary relationship features.
[0251] An enhancement module, configured to embed the unary relationship features, the pairwise relationship features, and the image features into the ternary relationship features through an attention module for fusion.
[0252] A final interaction module, configured to enhance the initial interaction decoded features through the fused ternary relationship features to obtain an interaction feature vector.
[0253] Embodiment 3. An embodiment of the present invention provides a person interaction detection device based on cascade parallelism, which includes a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement a person interaction detection method based on cascade parallelism as described in any paragraph of Embodiment 1.
[0254] Embodiment 4. An embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute a method for detecting human interaction based on cascaded parallelism as described in any paragraph of Embodiment 1.
[0255] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0256] In addition, the functional modules in each embodiment of the present invention can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.
[0257] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs. It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0258] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.
[0259] It should be understood that the term "and / or" used herein is only a kind of association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0260] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0261] The "first / second" mentioned in the embodiments is only used to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by the "first / second" can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0262] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for detecting human interaction based on cascade parallelism, characterized in that: Include: Obtain an image to be recognized; Extracting image features according to the image to be identified, and generating position information of the image features; Iteration The following steps are performed, and the human interaction detection result is predicted based on the human feature vector, object feature vector and interaction feature vector outputted in the last iteration; Generate a human query vector and an object query vector according to the image features and the prior bounding box of the person pair, and obtain a query vector group; wherein the prior bounding box of the person pair is obtained by random initialization in the first iteration, and is updated according to the human decoding features and object decoding features of the previous iteration in subsequent iterations; Obtaining a human body decoding feature by decoding through a human body detector according to the image feature, the position information and the human body query vector; updating the object query vector according to the human body decoding feature, and then decoding the object decoding feature by an object detector according to the updated object query vector, the image feature and the position information; Inputting the image features, the human body decoding features and the object decoding features into a classification decoder to obtain initial interactive decoding features; Inputting the human body decoding feature, the object decoding feature and the initial interaction decoding feature into a multiple relationship construction module to obtain unary relationship features, paired relationship features and ternary relationship features; The unary relationship feature, the pairwise relationship feature and the image feature are embedded into the ternary relationship feature for fusion through an attention module; Enhance the initial interactive decoding feature by using the fused ternary relationship feature to obtain an interactive feature vector; Inputting the human body decoding feature, the object decoding feature and the initial interaction decoding feature into a multiple relationship construction module to obtain unary relationship features, paired relationship features and ternary relationship features, specifically including: The human body decoding feature, the object decoding feature and the initial interaction decoding feature are respectively used as unary relationship features; The human body decoding features, the object decoding features and the initial interaction decoding features are combined in pairs in the channel dimension, and then The linear layer performs dimensionality reduction to obtain the pairwise relationship features; The human body decoding feature, the object decoding feature and the initial interaction decoding feature are combined as a ternary relationship feature.
2. The method for detecting human interaction based on cascaded parallelism according to claim 1, characterized in that: The character pair prior bounding box includes a human body prior bounding box and an object prior bounding box; The query generator, human detector, and object detector are all stacked, and the number of stacking layers is layer; According to the image features and the prior bounding box of the person, a human query vector and an object query vector are generated to obtain a query vector group, which specifically includes: According to the human body prior bounding box and the object prior bounding box, obtaining human body region features and object region features from the image features, and merging the two region features according to feature dimensions to obtain human body region features; According to the regional features of the characters, the convolutional layer plus the multi-layer perceptron is added The activation function predicts the key points of the human body and the key points of the object respectively; According to the key points of the human body and the key points of the object, a bilinear interpolation sampling operation is performed on the regional features of the human body to extract local visual features of the human body and local visual features of the object respectively; Aggregating the local visual features of the human body and the local visual features of the object through a fully connected layer to obtain the visual features of the human body and the visual features of the object; Obtaining a position embedding corresponding to each query based on the sinusoidal position encoding of the person's prior bounding box; Obtain an initial query vector of the query generator of the current layer, and fuse the position embedding to obtain a final query vector of the current layer; wherein the initial query vector of the query generator of the first layer uses the human visual features and the object visual features; the initial query vectors of the query generators of the remaining layers need to be obtained by fusing the feature vector output by the detector of the previous layer with the human visual features and the object visual features.
3. The method for detecting human interaction based on cascaded parallelism according to claim 2, characterized in that: The calculation model of the final query vector is: ; ; In the formula, is the layer number, For the The final query vector of the human body at the layer, For the The final query vector of the layer object, For the The initial query vector of the layer human detector, For the The initial query vector of the layer object detector, For human body position embedding, Embedding for object position, For the The feature vector output by the layer human detector, For the The feature vector output by the layer object detector.
4. The method for detecting human interaction based on cascaded parallelism according to claim 1, characterized in that: Both human detectors and object detectors are stacked, and the number of stacked layers is Layer; each layer of human detector and each layer of object detector is a Transformer decoder; The model decoded by the human detector is: ; In the formula, For the The feature vector output by the layer human detector, For the Layer human body detector, is the image feature, for the location information, For the The final query vector of the human body at the layer; The update model for updating the object query vector according to the human body decoding feature is: ; In the formula, For the The updated object query vector of the layer, For the The final query vector of the layer object, For the The feature vector output by the layer human detector; The model for decoding via the object detector is: ; In the formula, For the The feature vector output by the layer object detector, For the Layer object detector, is the image feature, for the location information, For the The updated object query vector for the layer.
5. The method for detecting human interaction based on cascaded parallelism according to claim 1, characterized in that: The number of layers of the interactive classifier stack is layers, and each layer of interactive classifiers contains a classification decoder; The classification decoder model is: ; ; In the formula, is the number of stacked layers, is the layer number, For the Classification decoding of the layer classification decoder output, For the The classification decoder of the layer, is the image feature, for the location information, The character pair decoding features are obtained by fusion and addition of the human decoding features and the object decoding features. The human body detector The feature vector output by the layer decoder, The object detector The feature vector output by the layer decoder, For the Interactive decoding features of layers.
6. The method for detecting human interaction based on cascaded parallelism according to claim 1, characterized in that: The pairwise relationship feature model is: ; ; ; The ternary relationship feature model is: ; In the formula, For the The paired relationship features of the human body decoding features and object decoding features of the layer, For the The paired relationship features of the human body decoding features and the initial interaction decoding features of the layer, For the The paired relationship features of the object decoding features of the layer and the initial interaction decoding features, For the linear layer, The human body detector The feature vector output by the layer decoder, The object detector The feature vector output by the layer decoder, For the The initial interactive decoding features output by the layer classification decoder, For the The ternary relationship characteristics of the layer, Indicates a merge operation; The unary relationship feature, the paired relationship feature and the image feature are embedded into the ternary relationship feature through an attention module for enhancement, specifically including: The contextual content of the unary relation feature is obtained through the self-attention mechanism; ; In the formula, For contextual content, Represents the self-attention mechanism; The context content is embedded into the ternary relationship feature using a cross attention mechanism to obtain a first embedding feature; wherein, ; In the formula, is the first embedding feature, represents the cross attention mechanism, For the The ternary relationship characteristics of the layer; The pairwise relationships of pairwise relationship features are obtained through the self-attention mechanism; among them, ; In the formula, For paired relationships; The pairwise relationship is embedded into the first embedding feature using a cross attention mechanism to obtain a second embedding feature; wherein, ; In the formula, is the second embedded feature; The image feature is embedded into the second embedding feature using a cross attention mechanism to obtain an enhanced ternary relationship feature; wherein, ; In the formula, For the enhanced ternary relationship features, is the image feature; The enhanced ternary relationship feature and the initial interactive decoding feature are converted to obtain the interactive decoding feature, specifically including: pass The linear layer transforms the enhanced ternary relationship features and the initial interactive decoding features; the transformation model is: ; ; In the formula, For the The interaction feature vector of the layer, For the The initial interactive decoding features output by the layer classification decoder, is the conversion factor, Represented as element-wise multiplication, For the enhanced ternary relationship features, for Activation function.
7. A cascaded parallel-based person interaction detection method according to any one of claims 1 to 6, characterized in that: Extracting image features according to the image to be identified and generating location information of the image features specifically includes: Extracting a convolution feature map from the image to be identified by a convolutional neural network module; Generate a fixed position code of the same size as the position information of the feature map according to the convolution feature map; Inputting the position information and the convolution feature map into a Transformer encoder to obtain encoded image features; Iteration The following steps are performed, and the last output is used as the final human interaction detection result, including: According to the feature vector output by the last layer decoder, prediction is performed through the head forward network to obtain the final human interaction detection result; ; ; ; ; In the formula, is the human body bounding box, For prediction The forward network, is the human feature vector output by the last layer of human detector, The human body position information updated for the last iteration, is the object bounding box, For prediction The forward network, is the object feature vector output by the last layer of object detector, The object position information updated for the last iteration, For object categories, For prediction The forward network, For the interaction category, For prediction The forward network, is the interactive decoding feature output by the last layer of interactive classifier, for Activation function, for Activation function.
8. A cascade-based parallel human interaction detection device, characterized in that: Include: An image acquisition module, used for acquiring an image to be identified; A feature extraction module, used to extract image features according to the image to be identified, and generate position information of the image features; Iteration output module, used for iteration The following steps are performed, and the human interaction detection result is predicted based on the human feature vector, object feature vector and interaction feature vector outputted in the last iteration; A query vector module, used to generate a human query vector and an object query vector according to the image features and the prior bounding box of the person pair, so as to obtain a query vector group; wherein the prior bounding box of the person pair is obtained by random initialization in the first iteration, and is updated according to the human decoding features and object decoding features of the previous iteration in subsequent iterations; A human body detection module, configured to obtain human body decoding features by decoding through a human body detector according to the image features, the position information and the human body query vector; An object detection module, used for updating the object query vector according to the human body decoding feature, and then obtaining the object decoding feature by decoding through an object detector according to the updated object query vector, the image feature and the position information; An initial interaction module, used for inputting the image feature, the human body decoding feature and the object decoding feature into a classification decoder to obtain an initial interaction decoding feature; A relationship construction module, used for inputting the human body decoding feature, the object decoding feature and the initial interaction decoding feature into a multiple relationship construction module to obtain a unary relationship feature, a pairwise relationship feature and a ternary relationship feature; An enhancement module, used for embedding the unary relationship feature, the paired relationship feature and the image feature into the ternary relationship feature for fusion through an attention module; A final interaction module, used for enhancing the initial interaction decoding features through the fused ternary relationship features to obtain an interaction feature vector; The relationship building module is used to perform the following steps: The human body decoding feature, the object decoding feature and the initial interaction decoding feature are respectively used as unary relationship features; The human body decoding features, the object decoding features and the initial interaction decoding features are combined in pairs in the channel dimension, and then The linear layer performs dimensionality reduction to obtain the pairwise relationship features; The human body decoding feature, the object decoding feature and the initial interaction decoding feature are combined as a ternary relationship feature.
9. A cascade-based parallel human interaction detection device, characterized in that: It comprises a processor, a memory, and a computer program stored in the memory; the computer program can be executed by the processor to implement a cascaded parallel-based human interaction detection method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a cascade-parallel-based character interaction detection method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Character interaction detection method, device and equipment based on query generator
CN116662587A
Character interaction detection method based on visual language model
CN118212399A