Image action interaction detection method and system based on attention model, and medium
Through the image action interaction detection method based on the visual attention model, the query initialization network is designed using the position information encoding and modulation of the target candidate box, which solves the problem of position context information being ignored in the existing method, and achieves higher detection accuracy and efficiency.
Patent Information
- Application Number
- CN202410116651.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-26
- Publication Date
- 2025-07-29
AI Technical Summary
The existing human-object interaction detection methods have the problem of low detection accuracy when dealing with complex scenes and subtle interactions, mainly because they ignore insufficient position context information and feature representation.
The image action interaction detection method based on the visual attention model is adopted, and the position information of the target candidate box is encoded and modulated, position features are generated, and the query initialization network is designed, and the position context information is used to guide and optimize the person-object interaction detection process.
It improves the accuracy and efficiency of human-object interaction detection in complex scenarios, and improves the model's understanding of complex spatial relationships.
Smart Images

Figure CN120388414A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and machine learning, and particularly to an image action interaction detection method, system, and medium based on an attention model for human-object interaction detection. Background Art
[0002] Human-object interaction detection is an important research direction in the field of computer vision, and its goal is to identify the interaction relationship between humans and objects from images. Traditional methods rely on convolutional neural networks (CNNs) and two-stage detection frameworks, but these methods have limitations in dealing with complex scenes and subtle interactions. Existing human-object interaction detection methods usually ignore the importance of position context information in images, which results in low detection accuracy in some cases. In addition, existing methods also have deficiencies in feature representation and network initialization, which further limits the performance of the detection model.
[0003] After retrieval, Chinese invention application 202210436888.8 discloses a method, device, and electronic device for human interaction detection. This application unifies human instance detection and interaction relationship detection into a human interaction detection model based on a cascaded machine translation network, and combines global context and instance-level information for human interaction reasoning, improving the accuracy of human interaction detection. However, it also does not improve in terms of feature representation and network initialization, limiting the performance of the detection model. Summary of the Invention
[0004] Aiming at the deficiencies in the prior art, the present invention provides an image action interaction detection method, system, and medium based on an attention model, which uses position context information to guide and optimize the process of human-object interaction detection, thereby improving the detection accuracy and efficiency.
[0005] In the first aspect of the present invention, there is provided an image action interaction detection method based on a visual attention model, including:
[0006] Performing object detection on the input image to obtain target candidate boxes, where the target candidate boxes include at least one human candidate box and at least one object candidate box;
[0007] Encoding the position information of the target candidate boxes to obtain a first position encoding, and generating corresponding position features from the first position encoding using a position detection head;
[0008] Modulating the first position encoding using the position information of the target candidate boxes to obtain a second position encoding modulated by human-object combination, where the position relationship between the human and the object is encoded into the visual attention model to predict the interaction action between the human and the object;
[0009] Design a query initialization network using the position feature and the second position encoding, enabling the visual attention model to better understand the interaction type and context;
[0010] Input the image to be detected into a human-object interaction detection model containing the visual attention model to implement image action interaction detection.
[0011] Optionally, the performing object detection on the input image to obtain object candidate boxes includes: performing object detection using an object detection network to obtain object candidate boxes of one or more pairs of human-object combinations.
[0012] Optionally, the encoding the position information of the object candidate boxes to obtain the first position encoding includes: encoding using the position information of the object candidate boxes, and these information are encoded using the ratio of the relative information to the entire image to obtain the first position encoding;
[0013] The position information of the object candidate boxes includes: the relative position of the box center, relative width and height, relative area, aspect ratio of the box, relative distance and direction of the object to the person.
[0014] Optionally, the position information of the object candidate boxes is represented by 18 elements. Take the natural logarithm of each element to obtain another 18 elements, and connect them to form the complete first position encoding representation information.
[0015] Optionally, the modulating the first position encoding using the position information of the object candidate boxes includes: for the first position encoding of the object candidate boxes of a pair of human-object combinations, use the width and height of the object candidate boxes for size modulation to make the first position encoding related to the object candidate boxes and provide position guidance for cross-attention.
[0016] Optionally, the generating corresponding position features from the first position encoding using a position detection head, where: concatenate the first position encodings of the two object candidate boxes of a pair of human-object combinations to obtain the corresponding position features.
[0017] Optionally, the designing a query initialization network using the position feature and the second position encoding includes:
[0018] Provide a query initialization network;
[0019] Input the image features corresponding to the object candidate boxes and the position features into the query initialization network;
[0020] The query initialization network connects the image features and the position features, and after being processed by a multi-layer perceptron, outputs a feature vector containing the image features and the position features;
[0021] Initialize the visual attention model query with the feature vector and the second positional encoding, enabling the visual attention model to better understand the interaction type and context during initialization.
[0022] In a second aspect of the present invention, there is provided an image action interaction detection system based on a visual attention model, comprising:
[0023] A target detection module: performing target detection on an input image to obtain target candidate boxes, where the target candidate boxes include at least one person candidate box and at least one object candidate box;
[0024] A positional feature module: encoding the positional information of the target candidate boxes to obtain a first positional encoding, and using a positional detection head on the first positional encoding to generate corresponding positional features;
[0025] A modulation module: modulating the first positional encoding using the positional information of the target candidate boxes to obtain a second positional encoding modulated by the person-object combination, where the positional relationship between the person and the object is encoded into the visual attention model to predict the interaction action between the person and the object;
[0026] A query initialization module: designing a query initialization network using the positional features and the second positional encoding, enabling the visual attention model to better understand the interaction type and context during initialization;
[0027] An interaction detection module: inputting an image to be detected into a person-object interaction detection model including the visual attention model to implement image action interaction detection.
[0028] In a third aspect of the present invention, there is provided a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned image action interaction detection method based on a visual attention model are implemented.
[0029] In a fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned image action interaction detection method based on a visual attention model are implemented.
[0030] Compared with the prior art, the present invention has at least one of the following beneficial effects:
[0031] The image action interaction detection method and system based on the attention model provided by the present invention generate position features by using the position information of the target candidate boxes, and further perform position encoding with human-object combination modulation, and use the position context information to guide and optimize the process of human-object interaction detection, which helps to improve the model's ability to understand complex spatial relationships and effectively improves the detection accuracy.
[0032] The image action interaction detection method and system based on the attention model provided by the present invention initialize the query of the visual attention model by using the query detection module, which effectively improves the model training efficiency while improving the model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] By reading the following detailed description of the non-restrictive embodiments with reference to the accompanying drawings, other features, objects and advantages of the present invention will become more apparent:
[0034] Figure 1 It is a flowchart of the image action interaction detection method based on the attention model in an embodiment of the present invention;
[0035] Figure 2 It is a module diagram of the image action interaction detection system based on the attention model in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several deformations and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0037] Traditional human-object interaction detection methods have limitations in dealing with complex scenarios and subtle interactions, mainly because they ignore the importance of position context information. In view of this situation, the embodiments of the present invention provide an image action interaction detection method based on a visual attention model, which uses position context information to guide and optimize the recognition process of human-object interaction to improve the accuracy and efficiency of human-object interaction detection in complex scenarios. In this process, the problem that the first-stage detection results are not effectively utilized in the two-stage method of human-object interaction detection is particularly considered. The embodiments of the present invention use the position information of the candidate person boxes and object boxes generated by the object detection in the first stage of human-object interaction detection to initialize the visual attention model in the second stage of interaction detection, which enables the model in the second stage to achieve higher accuracy after training.
[0038] Refer to Figure 1As shown in the figure, based on the above concept, the method for detecting image action interaction based on an attention model in an embodiment of the present invention may include the following steps:
[0039] S100. Perform object detection on the input image to obtain object candidate boxes, where the object candidate boxes include at least one person candidate box and at least one object candidate box;
[0040] S200. Encode the position information of the object candidate boxes to obtain a first position encoding, and use a position detection head to generate corresponding position features for the first position encoding;
[0041] S300. Modulate the first position encoding using the position information of the object candidate boxes to obtain a second position encoding modulated by person-object combination, where the positional relationship between the person and the object is encoded into the visual attention model to predict the interaction actions between the person and the object;
[0042] S400. Design a query initialization network using the position features and the second position encoding, so that the visual attention model can better understand the interaction type and context during initialization;
[0043] S500. Input the image to be detected into a person-object interaction detection model including a visual attention model to implement image action interaction detection.
[0044] In the above embodiment of the present invention, by analyzing the relative positional relationship between the person and the object in the image and introducing position context information, the interaction actions between the person and the object can be understood and predicted more accurately. Among them, a feature representation method that combines the visual features of the image and the position context information is used to improve the richness and effectiveness of the feature representation. At the same time, a query initialization network based on a visual attention model is introduced, so that the visual attention model can better understand the interaction type and context during initialization, effectively improving the detection accuracy.
[0045] In the above embodiment, in S100, object detection is performed on the input image to obtain object candidate boxes. Specifically, any object detection network can be used for object detection to obtain one or more pairs of object candidate boxes of person-object combinations. The object candidate boxes here include the position features of the object candidate boxes and the image features in the object candidate boxes, and these two features are corresponding, that is, paired person candidate boxes and object candidate boxes.
[0046] After obtaining the above object candidate boxes, in S200, the position information of the above object candidate boxes is encoded to obtain a first position encoding. In order to improve the richness and effectiveness of the feature representation, multiple position features are used to encode the spatial relationship between the person and the object. In some possible implementation manners, the encoding of the position information of the object candidate boxes can be performed according to the following operations:
[0047] Encode using the position information of the target candidate box, and these information are encoded using the ratio of the relative information to the entire image to obtain the first position encoding;
[0048] Among them, the above-mentioned position information of the target candidate box includes: the relative position of the box center, relative width and height, relative area, aspect ratio of the box, relative distance and direction of the object to the person.
[0049] In this embodiment, the position information is encoded through the above five types of position information, which can improve the richness and effectiveness of feature representation.
[0050] Exemplarily, in one embodiment, all five types of position information are encoded using the ratio of the relative information to the entire image. For example, the value of the relative area is equal to the ratio of the box area to the area of the entire image. These five types of information can be represented by 18 elements. For the obtained position encoding, take the natural logarithm of each element to obtain another 18 elements, and connect them to form the complete first position encoding representation information.
[0051] Specifically, these five types of information are represented by 4, 4, 3, 2, and 5 elements respectively. The relative position of the candidate box center is obtained by dividing the x and y coordinates of the candidate box by the width and height of the entire picture respectively. In this way, each candidate box has 2 elements (two values between 0 and 1), and a candidate box combination has 4 elements. The relative width and height are the same. For example, divide the width of the candidate box by the width of the entire picture to obtain the relative width. The relative area consists of 3 elements. The first 2 elements are the area ratios of the candidate box relative to the original picture respectively, and the last 1 element is the area ratio of the object box relative to the person box. The aspect ratio of the candidate box is composed of the width of the candidate box divided by the length, and a candidate box combination has 2 elements. The relative distance and direction of the object to the person are composed of 5 elements. The first element is the intersection over union (IoU) of the task box and the object box, and the other 4 elements are composed of the relative position relationship between the object and the person.
[0052] Furthermore, after obtaining the above-mentioned constructed first position encoding, a position detection head is designed to generate the corresponding position features. In some embodiments, the position detection head can be composed of a multi-layer convolutional neural network, and it will output the constructed position features.
[0053] In the above embodiments of the present invention, the first position encoding of the person-object combination is performed using five types of position information, including the relative position of the box center, relative width and height, relative area, aspect ratio of the box, relative distance and direction of the object to the person, etc., so as to more accurately understand the spatial relationship between the person and the object.
[0054] Through the precise encoding of the above five types of location information in the embodiments of the present invention, the interaction relationship between people and objects can be better understood and described. This is crucial for dealing with complex scenarios and interactions, and helps to improve the model's understanding ability of complex spatial relationships.
[0055] In some possible implementation manners, in step S300, the position information of the target candidate box is used to modulate the first position encoding to obtain the second position encoding of the person-object combination modulation, which can be operated as follows: for the position encoding of a single target candidate box pair (a pair of candidate boxes for a person-object combination), the width and height of the target candidate box are used for size modulation to make the first position encoding related to the target candidate box and provide position guidance for cross-attention.
[0056] Furthermore, the position encoding in the above embodiments can concatenate the first position encodings of the two target candidate boxes of a pair of person-object combinations to obtain the corresponding position features.
[0057] In some possible implementation manners, in step S400, in order to effectively improve the detection accuracy, the query initialization network can be designed as follows:
[0058] S401, provide a query initialization network;
[0059] In this step, the query initialization network can adopt an existing network structure, such as a multi-layer perceptron.
[0060] S402, input the image features and position feature information corresponding to the target candidate box into the query initialization network;
[0061] The image features and position feature information corresponding to the target candidate box, that is, the image features and position features determined by the target candidate box obtained by the target detection network in step S100.
[0062] S403, the query initialization network connects the image features and position features, and outputs a feature vector containing the image features and position features after being processed by the multi-layer perceptron of the query initialization network; wherein, the multi-layer perceptron is used to fuse the image features and position features;
[0063] S404, use the feature vector and the second position encoding as the query initialization of the visual attention model, so that the visual attention model can better understand the interaction type and context during initialization.
[0064] In this embodiment, the query initialization network combines image features with location features and uses a second location encoding modulated by the human-object combination to generate a query guided by location information. That is, the image-level information detected by the object detection network is combined with the location information as the input for query initialization. Through this design, not only the accuracy of human-object interaction detection is improved, but also the interaction dynamics in complex scenarios can be effectively processed. These features make the method of the present invention have broad application prospects in the field of computer vision.
[0065] In this embodiment, for the input image, a deep learning-based object detection network is used to perform object detection on it, which is the first stage of the human-object interaction detection process. For the person box and object box obtained in the first stage, the location information of the candidate person box and object box generated by the object detection in the first stage of human-object interaction detection is used to initialize the vision attention-based network in the second-stage interaction detection model, and higher accuracy can be achieved after training. After obtaining the first location encoding of the person box and object box, the location information of the candidate person box and object box generated by the object detection in the first stage of human-object interaction detection is used to perform location features.
[0066] The first location encoding in the vision attention model in the second-stage interaction detection model is modulated to obtain a second location encoding modulated by the human-object combination. Then, for the generated location features and location encoding, a query initialization network is designed to initialize the query part of the vision self-attention model in the interaction detection process. The result of the initialization module is output to the human-object interaction detection model based on the vision attention model. After model training, a successfully trained human-object interaction detection model is obtained, and this human-object interaction detection model can give good interaction predictions.
[0067] Specifically, each of the above parts is further described in detail:
[0068] (1) Object detection network:
[0069] The object detection network DETR based on the vision attention model fine-tuned on the human-object interaction detection dataset is used to perform object detection and generate detection results. Of course, in other embodiments, other object detection networks can also be used, not limited to the object detection network DETR based on the vision attention model, and it can be an object detection network without a vision attention model.
[0070] (2) Location features:
[0071] For a pair of human-object combinations obtained by the object detection network, five types of location feature information are used to encode its location information. All this information is encoded using the ratio of the relative information to the entire image. For example, the value of the relative area is equal to the ratio of the box area to the area of the entire image.
[0072] These five types of information are the relative position of the center of the box, the relative width-to-height ratio, the relative area, the aspect ratio of the box, and the relative distance and direction of the object from the person. These five types of information can be represented using 18 elements. For the obtained position encoding, the natural logarithm is taken for each element to obtain another 18 elements, and they are concatenated to form a complete first position encoding to represent the information.
[0073] For the completed first position encoding, a position detection head is designed to generate the corresponding position features. This detection head consists of a multi-layer convolutional neural network, which will output the completed position features.
[0074] (3) The second position encoding modulated by the person-object combination:
[0075] In this embodiment, it is based on the Transformer structure.
[0076] The definition of the second position encoding uses sine and cosine functions to embed a scalar value, as shown in the following formula.
[0077]
[0078]
[0079] In the formula, is the sine (cosine) encoding that defines a scalar, i is a positive integer that can range from 1 to d / 2, and τ is the temperature parameter.
[0080] The sine (cosine) encoding is a common technique that allows the Transformer-based model to learn position information. Since image information is inherently not in a sequential form, and the Transformer structure is a structure for sequences, when image information is input into the Transformer structure, a position information defined by the cosine function usually needs to be added, so that the Transformer structure can learn the information. This type of embedding can handle different scales, and they are learnable for the model, and the model can adjust them to better express the position patterns that are most useful in a specific problem.
[0081] In this embodiment, the first position encoding is modulated using the target candidate box information. For the first position encoding of a single target candidate box pair (a pair of target candidate boxes for a person-object combination), the width and height of the target candidate box are used for size modulation, so that the position encoding is related to the target candidate box, thereby providing position guidance for cross-attention. For a target candidate box [x, y, w, h], the specific implementation form of the second position encoding is shown in Equation (3):
[0082]
[0083] Among them, PE is the second position encoding of the target candidate box, and wref and href are the reference values of the width and height. wref and href are learnable parameters obtained from the feature f of the target candidate box, as shown in the following formula:
[0084] w_{ref}, h_{ref} = σ(MLP(f)).
[0085] The second position encoding of the above-mentioned person-object pair is completed by concatenating the first position encodings of two target candidate boxes, as shown in the following formula:
[0086] q_p = [PE(b_h), PE(b_o)]
[0087] PE(b_h) and PE(b_o) are the first position encodings of the person box and the object box respectively, and q_p is the position encoding of the query. This operation directly encodes the positional relationship between the person and the object into the attention mechanism of the model.
[0088] (4) Query initialization network design:
[0089] For the image features and position feature information detected by the target detection network, they are input into a pre-designed query initialization network. Specifically, the pre-designed query initialization model includes fully connected layers that separately process the two types of features. After separately processing the image features and position features, these two parts of features are concatenated, and after being processed by a multi-layer perceptron, a feature vector containing the image features and position features is output, and it is used as the query initialization of the attention model. At the same time, the position encoding modulated by the position information of the target candidate box is also used as a part of the query initialization model.
[0090] In summary, the embodiment of the present application proposes a query initialization network by making full use of the position information of the target candidate boxes in the image in the person-object interaction detection, and uses two modules, namely the position feature and the second position encoding modulated by the person-object combination, to deeply explore the use of position information, focusing on using the position context information to guide and optimize the process of person-object interaction detection, which helps to improve the model's ability to understand complex spatial relationships and effectively improves the detection accuracy.
[0091] Based on the same technical concept, in another embodiment, an image action interaction detection system based on a visual attention model is further provided, including: a target detection module 100, a position feature module 200, a modulation module 300, a query initialization module 400, and an interaction detection module 500. Specifically:
[0092] Object Detection Module 100: This module performs object detection on the input image to obtain object candidate boxes, where the object candidate boxes include at least one person candidate box and at least one object candidate box;
[0093] Location Feature Module 200: This module encodes the location information of the object candidate boxes to obtain the first location encoding, and uses a location detection head on the first location encoding to generate corresponding location features;
[0094] Modulation Module 300: This module modulates the first location encoding obtained by the Location Feature Module 200 using the location information of the object candidate boxes obtained by the Object Detection Module 100 to obtain a second location encoding modulated by person-object combination, where the location relationship between the person and the object is encoded into the visual attention model to predict the interaction actions between the person and the object;
[0095] Query Initialization Module 400: This module designs a query initialization network using the location features obtained by the Location Feature Module 200 and the second location encoding obtained by the Modulation Module 300, enabling the visual attention model to better understand the interaction type and context during initialization;
[0096] Interaction Detection Module 500: This module inputs the image to be detected into a person-object interaction detection model containing a visual attention model to implement image action interaction detection.
[0097] In this embodiment, for the image action interaction detection system based on the visual attention model, the technologies implemented by each module can refer to the implementation technologies of the corresponding steps in the above-mentioned image action interaction detection method based on the visual attention model, which will not be elaborated here.
[0098] In another embodiment of the present invention, a computer device is further provided, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned image action interaction detection method based on the visual attention model are implemented.
[0099] Optionally, a memory for storing programs; the memory may include volatile memory (e.g., random-access memory, such as static random-access memory (SRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), etc.); the memory may also include non-volatile memory, such as flash memory. The memory is used to store computer programs (such as application programs and functional modules for implementing the above methods), computer instructions, etc. The above computer programs, computer instructions, etc. can be stored in partitions in one or more memories. And the above computer programs, computer instructions, data, etc. can be called by the processor.
[0100] The above computer programs, computer instructions, etc. can be stored in partitions in one or more memories. And the above computer programs, computer instructions, data, etc. can be called by the processor.
[0101] A processor for executing the computer programs stored in the memory to implement each step in the method described in the above embodiments. For specific details, please refer to the relevant descriptions in the previous method embodiments.
[0102] The processor and the memory can be of an independent structure or an integrated structure integrated together. When the processor and the memory are of an independent structure, the memory and the processor can be coupled and connected through a bus.
[0103] In another embodiment of the present invention, a computer-readable storage medium is further provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the image action interaction detection method based on a visual attention model are implemented.
[0104] Among them, the computer-readable medium includes a computer storage medium and a communication medium, where the communication medium includes any medium facilitating the transmission of a computer program from one place to another. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. Additionally, the ASIC can be located in a user device. Of course, the processor and the storage medium can also exist as discrete components in a communication device.
[0105] The present invention also provides a program product, which includes a computer program stored in a readable storage medium. At least one processor of the server can read the computer program from the readable storage medium, and the execution of the computer program by the at least one processor causes the server to implement the method of any one of the above embodiments of the present invention.
[0106] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the foregoing storage medium includes various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image action interaction detection method based on a visual attention model, characterized in that Including: Performing object detection on the input image to obtain object candidate boxes, where the object candidate boxes include at least one person candidate box and at least one object candidate box; Encoding the position information of the object candidate boxes to obtain a first position encoding, and using a position detection head on the first position encoding to generate corresponding position features; Modulating the first position encoding with the position information of the object candidate boxes to obtain a second position encoding modulated by person-object combination, where the position relationship between the person and the object is encoded into the visual attention model to predict the interaction actions between the person and the object; Designing a query initialization network using the position features and the second position encoding, enabling the visual attention model to better understand the interaction type and context; Inputting the image to be detected into a person-object interaction detection model including the visual attention model to achieve image action interaction detection.
2. The image action interaction detection method based on a visual attention model according to claim 1, wherein The performing object detection on the input image to obtain object candidate boxes includes: Performing object detection using an object detection network to obtain one or more pairs of object candidate boxes of person-object combinations.
3. The image action interaction detection method based on a visual attention model according to claim 1, characterized in that The encoding the position information of the object candidate boxes to obtain a first position encoding includes: Encoding using the position information of the object candidate boxes, and these information are encoded using the ratio of the relative information to the entire image to obtain a first position encoding; The position information of the object candidate boxes includes: the relative position of the box center, relative width and height, relative area, aspect ratio of the box, relative distance and direction of the object to the person.
4. The method for detecting image action interaction based on a visual attention model according to claim 3, characterized in that The position information of the object candidate boxes is represented by 18 elements. Taking the natural logarithm of each element to obtain another 18 elements, and connecting them to form the complete first position encoding representation information.
5. The method for detecting image action interaction based on a visual attention model according to claim 1, wherein The modulating the first position encoding with the position information of the object candidate boxes includes: For the first position encoding of a pair of object candidate boxes of person-object combination, modulating the size using the width and height of the object candidate boxes to make the first position encoding related to the object candidate boxes and provide position guidance for cross attention.
6. The method for detecting image action interaction based on a visual attention model according to claim 1, wherein The using a position detection head on the first position encoding to generate corresponding position features, where: Concatenating the first position encodings of two object candidate boxes of a pair of person-object combination to obtain the corresponding position features.
7. The method for detecting image action interaction based on a visual attention model according to claim 1, wherein The designing a query initialization network using the position features and the second position encoding includes: Providing a query initialization network; Inputting the image features corresponding to the object candidate boxes and the position features into the query initialization network; The query initialization network connects the image features and the position features, and after being processed by a multi-layer perceptron, outputs a feature vector containing image features and position features; Using the feature vector and the second position encoding as the query initialization of the visual attention model, enabling the visual attention model to better understand the interaction type and context during initialization.
8. An image action interaction detection system based on a visual attention model, characterized in that, Including: Object detection module: Performing object detection on the input image to obtain object candidate boxes, where the object candidate boxes include at least one person candidate box and at least one object candidate box; Location Feature Module: Encode the location information of the target candidate box to obtain the first location encoding, and use the location detection head on the first location encoding to generate the corresponding location features; Modulation Module: Modulate the first location encoding using the location information of the target candidate box to obtain the second location encoding modulated by the human-object combination, where the positional relationship between the person and the object is encoded into the visual attention model to predict the interaction actions between the person and the object; Query Initialization Module: Design a query initialization network using the location features and the second location encoding, enabling the visual attention model to better understand the interaction type and context during initialization; Interaction Detection Module: Input the image to be detected into the human-object interaction detection model containing the visual attention model to implement image action interaction detection.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Character interaction detection method and device and electronic equipment
CN114550223A