A method and device for detecting human interaction

By introducing cross-modal calibration and fusion of visual and semantic modal vectors in human interaction detection, the problem of poor correlation between human objects and physical objects is solved, and more accurate interaction detection results are achieved.

CN114596627BActive Publication Date: 2025-10-03ALIBABA (CHINA) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210113567.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-30
Publication Date
2025-10-03
Estimated Expiration
2042-01-30

AI Technical Summary

Technical Problem

In existing human interaction detection methods, the actions of human objects on objects cannot be associated with the objects, resulting in inaccurate detection results that are inconsistent with the real image.

Method used

By obtaining the visual modal vector and semantic modal vector of the image to be detected, inter-modal calibration and fusion are performed to predict the verb category of the human object for the object, including using convolutional neural networks, Transformer encoders and decoders for feature extraction and processing, and introducing the verb semantic features of the object for cross-modal calibration and fusion.

Benefits of technology

The accuracy of human interaction detection is improved, and the interaction between human objects and objects in the image can be detected more accurately, which enhances the authenticity of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114596627B_ABST
    Figure CN114596627B_ABST
Patent Text Reader

Abstract

The present application provides a method for detecting human interaction, comprising: the method provided by the present application comprises: obtaining a visual modal vector of an image to be detected; obtaining a semantic modal vector corresponding to the object object based on the visual vector of the object object; performing inter-modal calibration on the visual modal vector and the semantic modal vector; and predicting the verb category of the human object in the image to be detected with respect to the object object based on the calibrated visual modal vector and the calibrated semantic modal vector. By obtaining the visual modal vector and the semantic modal vector of the image to be detected, and calibrating and fusing the visual modal vector and the semantic modal vector, the method can closely associate the object object in the image to be detected with the action of the human object with respect to the object object, thereby improving the accuracy of human interaction detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method and apparatus for detecting human interaction, a virtual reality device, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the rapid development of computer technology, the use of computer technology to detect images in videos or pictures has been widely applied in a variety of fields, including intelligent robots, live broadcast / short video subject detection, dangerous behavior detection, information detection, and human-computer interaction. Human interaction detection generally involves detecting human and object objects in images, as well as detecting the interaction between people and objects.

[0003] Existing human interaction detection methods typically consist of two parts: human and object detection, and human / object interaction action detection. These two components are performed independently during human interaction detection, often resulting in inaccurate results that are inconsistent with the real image and cause the detected human actions to become inaccurate. Summary of the Invention

[0004] In view of this, the present application provides a method and device for detecting human interaction to solve the technical problems in the prior art that the actions of human objects detected relative to object objects cannot be associated with the object objects, the detection results do not match the real images, and the detection results are inaccurate.

[0005] The present invention provides a method for detecting human interaction, including:

[0006] Obtaining a visual modality vector of the image to be detected; the visual modality vector includes: a visual vector of a person object and a visual vector of an object object;

[0007] Acquire a semantic modal vector corresponding to the object according to the visual vector of the object; the semantic modal vector includes: a verb vector of a candidate verb corresponding to the object;

[0008] Performing inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector;

[0009] The verb category of the person object in the image to be detected with respect to the object object is predicted according to the calibrated visual modality vector and the calibrated semantic modality vector.

[0010] Optionally, the inter-modality calibration of the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector includes: using a channel attention mechanism to perform corresponding calibration on the visual modality vector and the semantic modality vector.

[0011] Optionally, the inter-modal calibration of the visual modal vector and the semantic modal vector to obtain a calibrated visual modal vector and a calibrated semantic modal vector also includes: after the inter-modal calibration of the visual modal vector and the semantic modal vector, intra-modal calibration of the visual modal vector is performed using an information transfer mechanism.

[0012] Optionally, the inter-modal calibration of the visual modal vector and the semantic modal vector to obtain a calibrated visual modal vector and a calibrated semantic modal vector also includes: after the inter-modal calibration of the visual modal vector and the semantic modal vector, intra-modal calibration of the semantic modal vector is performed using an information transfer mechanism.

[0013] Optionally, predicting a verb category of the person object in the image to be detected with respect to the object object based on the calibrated visual modality vector and the calibrated semantic modality vector includes:

[0014] fusing the calibrated visual modality vector and the calibrated semantic modality vector to obtain verb features of the candidate verb;

[0015] The verb category of the person object in the image to be detected for the object object is predicted according to the verb features of the candidate verbs.

[0016] Optionally, the calibrated visual modal vector and the calibrated semantic modal vector are fused to obtain the verb features of the candidate verb, including: using the calibrated visual modal vector and the calibrated semantic modal vector as sequence elements to generate a verb sequence of the candidate verb.

[0017] Optionally, acquiring a semantic modality vector corresponding to the object according to the visual vector of the object includes:

[0018] Obtaining the original vector of the candidate verb corresponding to the object;

[0019] Obtaining a verb conditional probability of the candidate verb corresponding to the object relative to the object;

[0020] According to the original vector of the candidate verb and the verb conditional probability, a semantic modal vector corresponding to the object is obtained.

[0021] The present application also provides a method for detecting human interaction, including:

[0022] Obtaining a visual modality vector of the image to be detected; the visual modality vector includes: a visual vector of a person object and a visual vector of an object object;

[0023] Obtaining the original vector of the candidate verb corresponding to the object, and obtaining the verb conditional probability of the candidate verb corresponding to the object relative to the object;

[0024] Obtaining a semantic modal vector corresponding to the object according to the original vector of the candidate verb and the verb conditional probability; the semantic modal vector includes: a verb vector of the candidate verb corresponding to the object;

[0025] The verb category of the person object with respect to the object object is obtained according to the visual modality vector and the semantic modality vector.

[0026] Optionally, obtaining the original vector of the candidate verb corresponding to the object includes: obtaining the original vector of the candidate verb from a verb vector database according to the visual vector of the object.

[0027] Optionally, obtaining the verb conditional probability of the candidate verb corresponding to the object relative to the object includes: obtaining the verb conditional probability of the candidate verb relative to the object according to the visual vector of the object.

[0028] Optionally, obtaining the semantic modal vector corresponding to the object based on the original vector of the candidate verb and the verb conditional probability includes: taking the product of the original vector of the candidate verb and the verb conditional probability as the semantic modal vector corresponding to the object.

[0029] Optionally, obtaining the verb category of the person object with respect to the object object according to the visual modality vector and the semantic modality vector includes:

[0030] Performing inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector;

[0031] The verb category of the person object in the image to be detected with respect to the object object is predicted according to the calibrated visual modality vector and the calibrated semantic modality vector.

[0032] The embodiment of the present application also provides a person interaction detection device, comprising: a visual modality unit, a semantic modality unit, a calibration unit, and a prediction unit;

[0033] The visual modality unit is used to obtain a visual modality vector of the image to be detected; the visual modality vector includes: a visual vector of a person object and a visual vector of an object object;

[0034] The semantic modality unit is configured to obtain a semantic modality vector corresponding to the object according to the visual vector of the object; the semantic modality vector includes: a verb vector of a candidate verb corresponding to the object;

[0035] The calibration unit is configured to perform inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector;

[0036] The prediction unit is configured to predict a verb category of the person object in the image to be detected for the object object based on the calibrated visual modality vector and the calibrated semantic modality vector.

[0037] The embodiment of the present application further provides a person interaction detection device, comprising: a visual modality unit, a semantic modality unit, and a verb category acquisition unit;

[0038] The visual modality unit is used to obtain a visual modality vector of the image to be detected; the visual modality vector includes: a visual vector of a person object and a visual vector of an object object;

[0039] The semantic modality unit is used to obtain the original vector of the candidate verb corresponding to the object, and obtain the verb conditional probability of the candidate verb corresponding to the object relative to the object;

[0040] The semantic modality unit is further configured to obtain a semantic modality vector corresponding to the object based on the original vector of the candidate verb and the verb conditional probability; the semantic modality vector includes: a verb vector of the candidate verb corresponding to the object;

[0041] The verb category acquisition unit is configured to acquire the verb category of the person object with respect to the object object based on the visual modality vector and the semantic modality vector.

[0042] The present application also provides a virtual reality device, comprising: a memory and a processor; the memory storing a computer instruction set, which, when executed by the processor, performs the following steps:

[0043] Obtaining a visual modality vector of the image to be detected; the visual modality vector includes: a visual vector of a person object and a visual vector of an object object;

[0044] Acquire a semantic modal vector corresponding to the object according to the visual vector of the object; the semantic modal vector includes: a verb vector of a candidate verb corresponding to the object;

[0045] Performing inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector;

[0046] The verb category of the person object in the image to be detected with respect to the object object is predicted according to the calibrated visual modality vector and the calibrated semantic modality vector.

[0047] The present application also provides a virtual reality device, comprising: a memory and a processor; the memory storing a computer instruction set, which, when executed by the processor, performs the following steps:

[0048] Obtaining a visual modality vector of the image to be detected; the visual modality vector includes: a visual vector of a person object and a visual vector of an object object;

[0049] Obtaining the original vector of the candidate verb corresponding to the object, and obtaining the verb conditional probability of the candidate verb corresponding to the object relative to the object;

[0050] Obtaining a semantic modal vector corresponding to the object according to the original vector of the candidate verb and the verb conditional probability; the semantic modal vector includes: a verb vector of the candidate verb corresponding to the object;

[0051] The verb category of the person object with respect to the object object is obtained according to the visual modality vector and the semantic modality vector.

[0052] The embodiment of the present application further provides an electronic device, comprising: a collector, a processor, and a memory;

[0053] The collector is used to collect images to be detected;

[0054] The memory is used to store one or more computer instructions;

[0055] The processor is configured to execute the one or more computer instructions to implement the above method.

[0056] An embodiment of the present application also provides a computer-readable storage medium on which one or more computer instructions are stored, and the instructions are executed by a processor to implement the above method.

[0057] Compared with the prior art, the human interaction detection method provided by the present application includes: obtaining a visual modal vector of an image to be detected; the visual modal vector includes: a visual vector of a human object and a visual vector of an object object; according to the visual vector of the object object, obtaining a semantic modal vector corresponding to the object object; the semantic modal vector includes: a verb vector of a candidate verb corresponding to the object object; performing intermodal calibration on the visual modal vector and the semantic modal vector to obtain a calibrated visual modal vector and a calibrated semantic modal vector; and predicting the verb category of the human object in the image to be detected for the object object based on the calibrated visual modal vector and the calibrated semantic modal vector. By obtaining the visual modal vector (including: a visual vector of a human object and a visual vector of an object object) and the semantic modal vector (i.e., a verb vector of a candidate verb of an object object) of the image to be detected, and calibrating and fusing the visual modal vector and the semantic modal vector, the method can closely associate the object object in the image to be detected with the action of the human object for the object object, thereby improving the accuracy of human interaction detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0059] Figure 1 An application scenario diagram of a method for detecting human interaction provided in an embodiment of the present application;

[0060] Figure 2 An application system diagram of a method for detecting human interaction provided in an embodiment of the present application;

[0061] Figure 3 A flowchart of the character interaction detection method provided in the first embodiment of the present application;

[0062] Figure 4 A schematic diagram of a method for detecting human interaction provided in an embodiment of the present application;

[0063] Figure 5 A flowchart of obtaining the visual modality vector of the image to be detected provided in the first embodiment of the present application;

[0064] Figure 6 A flowchart of the semantic modality vector corresponding to an object provided in the first embodiment of the present application;

[0065] Figure 7A flowchart of a method for detecting human interaction provided in the second embodiment of the present application;

[0066] Figure 8 A flowchart for obtaining the verb category of a person object for an object in an image to be detected provided in the second embodiment of the present application;

[0067] Figure 9 A schematic structural diagram of a person interaction detection device provided in the third embodiment of the present application;

[0068] Figure 10 A schematic structural diagram of a person interaction detection device provided in a fourth embodiment of the present application;

[0069] Figure 11 A schematic structural diagram of a virtual reality device provided in a fifth embodiment of the present application;

[0070] Figure 12 A schematic structural diagram of an electronic device provided in the seventh embodiment of the present application. DETAILED DESCRIPTION

[0071] In order to enable those skilled in the art to better understand the technical solutions of this application, the following clearly and completely describes this application in conjunction with the drawings in the embodiments of this application. However, this application can be implemented in many other ways different from the above description. Therefore, based on the embodiments provided in this application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of this application.

[0072] Among existing methods for detecting human interaction, end-to-end human interaction detection is a widely used approach. This approach is typically designed with two parallel branches: a person / object detection branch, which detects the bounding boxes of people and objects in an image, as well as their categories; and an action detection branch, which infers verbal interaction relationships between people and objects, primarily detecting the actions people take on them. Action detection methods often employ a parallel approach for detecting people and objects, neglecting the "object," a core component of human interaction relationships. In other words, existing methods for detecting human interaction ignore prior knowledge of the action categories associated with the objects involved in human interaction relationships. This decouples the prior relationship between objects and verbs. For example, in a dataset, a person riding a horse has a high probability of occupying all actions for the horse object. However, existing detection methods do not consider the prior probability of the corresponding horse action. Consequently, the detection model assigns only a low probability to the human riding action, decoupling the prior relationship that the person riding a horse has a high probability of occupying all actions for the horse object. Consequently, existing methods for detecting human interaction fail to improve the accuracy of human interaction detection, resulting in discrepancies between the detection results and the real-world image.

[0073] In response to the problems existing in the above-mentioned existing human interaction detection methods, the present application provides a human interaction detection method. While performing visual detection on the image, it introduces the aggregation of verb semantic features of objects in the image, and performs cross-modal calibration and fusion of the visual detection output and semantic aggregation features, so as to obtain detection results that are more consistent with the real image.

[0074] The detection method, detection device, detection system and computer-readable storage medium described in this application are further described in detail below with reference to specific embodiments and drawings.

[0075] Figure 1 This is an application scenario diagram of a method for detecting human interaction provided by an embodiment of the present application. Figure 1 As shown, the character interaction detection method provided in the embodiment of the present application can be applied to the live broadcast / short video main product recommendation scenario. In this scenario, a camera is often set to shoot the interaction between the anchor and the product. Through the character interaction detection method provided in the embodiment of the present application, it is possible to detect relevant information about the anchor, relevant information about the product, and information about the anchor's interactive actions on the product. Through the above-mentioned detection results, the platform can self-generate recommendation information independent of the anchor, and recommend products to users in the form of subtitles, links, etc. In addition, if the above-mentioned detection results reveal unsafe behavior of the anchor, it can also quickly feedback to the platform to avoid adverse effects on users.

[0076] The human interaction detection method provided in the embodiment of the present application can also be applied to various fields such as dangerous behavior detection, information detection, and human-computer interaction.

[0077] Figure 2 This is an application system diagram of a method for detecting human interaction provided by an embodiment of the present application. Figure 2 As shown, the application system includes: a terminal 101 and a server 102. The terminal 101 and the server 102 are communicatively connected via a network. The terminal 101 can be an image acquisition device in various forms, such as a camera, a still camera, etc., and can be one or more. The server 102 can be an independent server, deploying the character interaction detection method provided in this application, or it can be a server group composed of multiple servers, each of which deploys a module of the character interaction detection method provided in this application. For example, a server group may include: a visual server, a semantic server, a calibration server, etc. Of course, the server 102 can also be a cloud server, and the character interaction detection method provided in this application is deployed on a cloud server. The terminal 101 collects the image to be detected, uploads it to the server 102 via the network, and the server 102 performs character interaction detection on the image to be detected.

[0078] The application system of this human interaction detection method can be applied to fields such as dangerous behavior detection. For example, for dangerous behavior detection, the terminal 101 can be set as a depth camera and deployed in various locations such as streets, shopping malls, and residential areas. The terminal 101 will upload the captured video to the server 102 via the network in the form of each frame of image. The server 102 will detect each frame of the image and, if it detects dangerous behavior such as a person holding a knife, it will alert the security management department.

[0079] The first embodiment of the present application provides a method for detecting human interaction.

[0080] Human interaction detection involves detecting the interaction between humans and objects in images or video frames. This includes detecting human objects, objects, and the interaction between humans and objects. For example, human interaction detection for a picture of a person riding a horse involves detecting the person, the horse, and the person's riding action. The output includes the person's bounding box information, the horse's bounding box information, and the verb category corresponding to the interaction between the person and the horse.

[0081] Therefore, human interaction detection is a relatively complex process. It requires not only detecting the objects in the image, but also the actions performed by the people in the image on the objects. Real-world images often do not have a one-to-one relationship; they can contain multiple people and objects, greatly increasing the difficulty of detection.

[0082] This application takes a one-to-two image as an example to explain in detail the human interaction detection method provided in the first embodiment of this application.

[0083] Figure 3 This is a flowchart of the character interaction detection method provided in this embodiment. Figure 4 This is a schematic diagram of the method for detecting human interaction provided by this embodiment. Figure 3 and Figure 4 The human interaction detection method provided by this embodiment is described in detail. The embodiments described below are used to explain the technical solution of this application and are not intended to limit actual use.

[0084] like Figure 3 The method for detecting human interaction provided in this embodiment includes the following steps:

[0085] Step S301 , obtaining a visual modality vector of an image to be detected; the visual modality vector includes: a visual vector of a person object and a visual vector of an object object.

[0086] This step is to detect the human object and the object in the image to be detected from the visual perspective and output a visual modality vector. The visual modality vector includes the visual vector of the human object and the visual vector of the object in the image to be detected.

[0087] The image to be detected refers to an image acquired by an image acquisition device and needs to be detected. It can be an image in the form of a picture or an image output as a video frame. In order to more clearly illustrate the human interaction detection method provided by this embodiment, this embodiment uses a one-to-two image as the image to be detected. Figure 4 The image to be detected shown in the figure includes a man (human object), a cup (first object), and a backpack (second object). The man carries the backpack on his right shoulder and holds the cup in his left hand.

[0088] The height, position, etc. of the person object can be predicted from the visual vector of the person object.

[0089] From the visual vector of the object, we can predict the object's position, size, shape, relative distance from the human object, contact position and contact area with the human object, etc. We can also predict that the category of the first object is a container and the category of the second object is a package.

[0090] The following combination Figure 4 An optional implementation of this step is described in detail.

[0091] Figure 5 This embodiment provides a flowchart for obtaining the visual modality vector of the image to be detected. The specific steps are as follows:

[0092] Step S301-1: Input the image to be detected into the convolutional neural network to extract image features.

[0093] Neural networks (NNs) are composed of neurons and their parameters. They are systems that perform tasks by "learning" from a large number of examples and are typically not programmed with task-specific rules. For example, in image recognition, a neural network can learn the characteristics of a cat by analyzing example images labeled "cat" or "not cat" and use this learning to identify whether other images contain cats. During neural network learning, the neural network is not fed directly with cat characteristics. Instead, it is fed with example images labeled as cats. Through iterative learning, the neural network automatically generates characteristic information representing a cat based on these example images.

[0094] Convolutional Neural Networks (CNNs) are a type of neural network that organizes several neurons into a convolutional layer. Data, starting from the input, propagates sequentially through several convolutional layers through the connections between neurons until it reaches the final output. Convolutional neural networks can also calculate errors based on a specified optimization objective and iteratively update the neural network parameters through backpropagation and gradient descent to optimize the network.

[0095] like Figure 4 As shown in the figure, when an image of a man holding a cup and a backpack is input, the iteratively optimized convolutional neural network can recognize that there is a person, a container, and a bag in the image, and can extract the features of the person, container, and bag.

[0096] Step S301 - 2 : combining the position coding information in the image to be detected with the image features.

[0097] The position coding information may refer to the order of the currently detected image in a video, or may refer to the coordinate position of a pixel point in the image.

[0098] This step combines the position coding information of the image to be detected with the extracted image features. The specific operation is to add the coding sequences of the two.

[0099] Step S301 - 3 : inputting the combined information of the image feature information and the position coding information into an encoder for coding processing to obtain a feature sequence with stronger representation capability.

[0100] The Transformer is an optional encoder that takes the summed feature information of the image to be detected and the positional encoding information and feeds it into its encoder for encoding. Through a series of encoding processes, including self-attention, summation and normalization, and a feedforward neural network, the final output is a feature sequence with stronger representational capabilities.

[0101] Step S301-4: input the encoded feature sequence into a decoder for decoding processing to obtain the visual vector of the person object and the visual vector of the object object.

[0102] The encoded feature sequence can be input into the Transformer for decoding to obtain the visual vector of the person object and the visual vector of the object object.

[0103] The Transformer is a specific form of neural network, consisting of an encoder and a decoder. Typically, the encoder and decoder are stacked together, with each layer having two sublayer connections. The first sublayer connection structure consists of a multi-head self-attention sublayer, a normalization layer, and a residual connection. The second sublayer connection structure consists of a feedforward fully connected sublayer, a normalization layer, and a residual connection.

[0104] Step S301 - 5 : predicting the category of the object through the visual vector of the object.

[0105] like Figure 4 As shown in the figure, an image of a man holding a cup and a backpack undergoes a series of encoding and decoding steps to output the man's visual vector, the first item's visual vector, and the second item's visual vector. The first item's visual vector can be used to predict its category as a container, while the second item's visual vector can be used to predict its category as a bag.

[0106] Through the above steps, the visual modality vector of the image to be detected is obtained from the input image to be detected.

[0107] Step S302: Acquire a semantic modal vector corresponding to the object according to the visual vector of the object; the semantic modal vector includes a verb vector of a candidate verb corresponding to the object.

[0108] This step is to obtain the semantic modal vector corresponding to the object in the image to be detected from the perspective of verb semantics, that is, the verb vector of the candidate verb corresponding to the object.

[0109] The candidate verbs refer to the semantics corresponding to all possible actions that can be performed on the object in the image to be detected. For example, the possible actions that can be performed on a horse include: riding a horse, leading a horse, feeding a horse, beating a horse, etc. The semantics corresponding to these actions include: riding, leading, feeding, beating, etc., so riding, leading, feeding, beating, etc. are the candidate verbs corresponding to horses. For another example: Figure 4 In the example, the man (character object) can perform actions such as taking, grabbing, and throwing on the cup (object object), while the impossible actions include riding, eating, and tearing. Therefore, taking, grabbing, and throwing are the candidate verbs corresponding to the cup (object object).

[0110] The verb vector refers to the multi-dimensional vector corresponding to a verb. Each verb has its corresponding vector. For example, the verb vector of "ride" can be represented as [1, 0, 0, 0, ……, 0], the verb vector of "lead" can be represented as [0, 1, 0, 0, ……, 0], the verb vector of "feed" can be represented as [0, 0, 1, 0, ……, 0], and the verb vector of "beat" can be represented as [0, 0, 0, 1, ……, 0]. The vectors of different verbs cannot be repeated. Therefore, the corresponding verb can be characterized by the verb vector.

[0111] The following combines Figure 4 A detailed description will be given to an optional implementation form of this step.

[0112] Figure 6 It is a flowchart for obtaining the semantic modal vector corresponding to the object in this embodiment.

[0113] The specific steps are as follows:

[0114] Step S302-1: Obtain the original vector of the candidate verb corresponding to the object.

[0115] An optional implementation method includes:

[0116] First, perform neural network mapping on the verb vectors in the verb vector database.

[0117] The verb vector database is a publicly available database of human interaction behaviors, which includes the verb vectors corresponding to almost all verbs and is a collection of verb vectors. The HICO-DET dataset and the V-COCO dataset are relatively comprehensive verb vector databases.

[0118] The purpose of this step is to make the verb vectors in the mapped verb vector database as close as possible to the co-occurrence probability of the verbs.

[0119] Second, obtain the original vector of the candidate verb from the mapped verb vector database according to the visual vector of the object.

[0120] According to the visual vector of the object in the image to be detected, the verb vectors related to the object can be screened out from the above-mentioned mapped verb vector database, that is, the original vectors of the candidate verbs corresponding to the object.

[0121] Such as Figure 4As shown in the figure, an image of a man holding a cup and a backpack is used as an example. The verb vector database includes verb vectors corresponding to verbs such as "take," "hold," "throw," "eat," "drink," "hit," "ride," and "grab." Image feature extraction and analysis show that the visual vector of the second object is a bag. Therefore, verb vectors corresponding to candidate verbs of the bag class are selected from the verb vector database, such as "take," "hold," "throw," and "grab," to form the original vectors of the candidate verbs corresponding to the object.

[0122] Step S302-2: Obtain the verb conditional probability of the candidate verb corresponding to the object relative to the object.

[0123] The verb conditional probability refers to the probability of event A occurring under the condition that event B occurs. In this embodiment, it refers to the probability of all actions that can be performed by the object in the image to be detected.

[0124] like Figure 4 As shown in the image of a man carrying a cup in his backpack, the cup can be subjected to various actions, including: holding, grabbing, throwing, riding, etc. The probability of these actions occurring is: holding 40%, grabbing 20%, throwing 10%, riding 1%. Of course, the cup can be subjected to more than the four actions listed above. In this embodiment, these four actions are used as examples for illustration.

[0125] All possible actions for a bag include: carrying, shouldering, lifting, throwing, etc. The probability of these actions occurring may be: carrying 30%, shouldering 30%, lifting 20%, and throwing 5%. Of course, there are more than the four actions that can be performed on a bag. In this embodiment, the above four actions are used as examples for explanation.

[0126] Obtaining the verb conditional probability is to obtain the verb conditional probability of the candidate verb relative to the object according to the visual vector of the object.

[0127] That is, the verb conditional probability obtained in this step is in one-to-one correspondence with the original vector of the candidate verb obtained in the first step.

[0128] Step S302-3: Obtain the semantic modal vector corresponding to the object according to the original vector of the candidate verb and the verb conditional probability.

[0129] The original verb vector refers to the original state of the verb vector without being processed.

[0130] This step specifically refers to aggregating the original vectors of the candidate verbs corresponding to the object in the image to be detected and the verb conditional probability of the candidate verbs.

[0131] An optional aggregation method is to multiply the original vector of the candidate verb by the verb conditional probability of the candidate verb, that is, to use the product of the original vector of the candidate verb and the verb conditional probability as the semantic modal vector corresponding to the object.

[0132] like Figure 4 As shown, in the image of a man holding a cup and a backpack, the candidate verbs for the first object (container type) are: take, lift, carry, etc., and the candidate verbs for the second object (bag type) are: carry on the back, sling over the shoulder, lift, throw, etc.

[0133] The verb conditional probabilities of the candidate verbs for the first object are: take 30%, lift 20%, carry 10%, and carry 10%; the verb conditional probabilities of the candidate verbs for the second object are: carry on the back 30%, carry over the shoulder 30%, lift 20%, and throw 5%.

[0134] The method of aggregating the original vectors of the candidate verbs corresponding to the first object and the verb conditional probabilities of the candidate verbs is: the original vector of take × 0.3, the original vector of lift × 0.2, the original vector of lift × 0.1, and the original vector of resist × 0.1.

[0135] The method of aggregating the original vectors of the candidate verbs corresponding to the second object and the verb conditional probabilities of the candidate verbs is: the original vector of "carry" × 0.3, the original vector of "sling" × 0.3, the original vector of "lift" × 0.2, and the original vector of "throw" × 0.05.

[0136] Through the above steps, the verb vector of the candidate verb corresponding to the object in the image to be detected can be obtained.

[0137] Step S303 : performing inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector.

[0138] This step is to perform inter-modality calibration on the visual modality vector obtained in step S301 and the semantic modality vector obtained in step S302.

[0139] The calibration provided in this embodiment includes: using a channel attention mechanism to calibrate the visual modality vector and the semantic modality vector. Specifically:

[0140] Assume that the number of features of the visual modality vector and the semantic modality vector is K, and the number of channels of each feature is D.

[0141] Since the semantic modality vector obtained in step S302 is guided by the visual modality vector obtained in step S301, the visual modality vector and the semantic modality vector should be in one-to-one correspondence. A channel attention mechanism is used to calibrate the corresponding channels of the visual modality vector and the semantic modality vector one-to-one.

[0142] For example, first, the channel attention size (D) is calculated using the semantic modality vector. Then, the channel attention size calculated using the semantic modality vector is used to adjust the channel attention size (D) of the visual modality vector. In this way, all features of the K pairs of visual modality vectors and semantic modality vectors are calibrated one by one. Similarly, the semantic modality vector can also be calibrated using the visual modality vector.

[0143] The calibration provided in this embodiment further includes: after performing inter-modality calibration on the visual modality vector and the semantic modality vector, performing intra-modality calibration on the visual modality vector using an information transfer mechanism. Specifically:

[0144] By using the Transformer encoder to transfer information between the K features of the visual modality vector, a visual modality vector with K features with stronger representation capabilities can be obtained.

[0145] The calibration provided in this embodiment further includes: after performing inter-modality calibration on the visual modality vector and the semantic modality vector, performing intra-modality calibration on the semantic modality vector using an information transfer mechanism. Specifically:

[0146] By using the Transformer encoder to transfer information between the K features of the semantic modal vector, a semantic modal vector with K features with stronger representation capabilities can be obtained.

[0147] The above provides a method for performing inter-modality calibration and intra-modality calibration on visual modality vectors and semantic modality vectors.

[0148] Step S304 : predicting the verb category of the person object in the image to be detected with respect to the object object according to the calibrated visual modality vector and the calibrated semantic modality vector.

[0149] This step includes: fusing the calibrated visual modality vector with the calibrated semantic modality vector to obtain verb features of the candidate verb.

[0150] The fusion refers to merging the calibrated visual modality vector and the calibrated semantic modality vector, which can be achieved through a feature fusion function.

[0151] That is, the fusion is to use the calibrated visual modality vector and the calibrated semantic modality vector as sequence elements to generate a verb sequence of the candidate verb.

[0152] This step also includes: predicting the verb category of the person object in the image to be detected for the object object based on the verb features of the candidate verb.

[0153] An optional prediction method is to use a feed-forward neural network (FFN) to predict the verb category of the person object for the object object in the image to be detected based on the verb features of the candidate verbs.

[0154] like Figure 4 As shown, using a feedforward neural network based on the verb features of the candidate verbs, it can be predicted that in the image to be detected, the verb category of the person object for the first object is "take", and the verb category of the person object for the second object is "back".

[0155] The above is an implementation of the character interaction detection method provided in the first embodiment of the present application. Figure 4 As shown, the above method can detect an image of a man holding a cup and a backpack, and the final detection results include: "man holding a cup" and "man carrying a backpack." The detection results also include: a feature description of the person object (e.g., man, height approximately 170 cm, short hair), a feature description of the first object object (e.g., cup, covered, cylindrical, approximately 10 cm high), and a feature description of the second object object (e.g., backpack, blue, rectangular). The detection results also include: the positional relationship between the first object object and the person object (the cup is held in front of the person), and the positional relationship between the second object object and the person object (the bag is carried behind the person). Of course, the detection results can also include other information, which will not be explained in detail here.

[0156] The second embodiment of the present application provides another method for detecting human interaction.

[0157] Figure 7 This is a flow chart of the character interaction detection method provided by this embodiment. Figure 7 and Figure 4 The human interaction detection method provided by this embodiment is described in detail. The embodiments described below are used to explain the technical solution of this application and are not intended to limit actual use.

[0158] Step S701 , obtaining a visual modality vector of an image to be detected; the visual modality vector includes: a visual vector of a person object and a visual vector of an object object.

[0159] The image to be detected refers to an image acquired by an image acquisition device and needs to be detected. It can be an image in the form of a picture or an image output as a video frame. In order to more clearly illustrate the human interaction detection method provided by this embodiment, this embodiment uses a one-to-two image as the image to be detected. Figure 4The image to be detected shown in the figure includes a man (human object), a cup (first object), and a backpack (second object). The man carries the backpack on his right shoulder and holds the cup in his left hand.

[0160] The height, position, etc. of the person object can be predicted from the visual vector of the person object.

[0161] From the visual vector of the object, we can predict the object's position, size, shape, relative distance from the human object, contact position and contact area with the human object, etc. We can also predict that the category of the first object is a container and the category of the second object is a package.

[0162] The following combination Figure 4 An optional implementation of this step is described in detail.

[0163] First, the image to be detected is input into the convolutional neural network to extract image features.

[0164] Neural networks (NNs) are composed of neurons and their parameters. They are systems that perform tasks by "learning" from a large number of examples and are typically not programmed with task-specific rules. For example, in image recognition, a neural network can learn the characteristics of a cat by analyzing example images labeled "cat" or "not cat" and use this learning to identify whether other images contain cats. During neural network learning, the neural network is not fed directly with cat characteristics. Instead, it is fed with example images labeled as cats. Through iterative learning, the neural network automatically generates characteristic information representing a cat based on these example images.

[0165] Convolutional Neural Networks (CNNs) are a type of neural network that organizes several neurons into a convolutional layer. Data, starting from the input, propagates sequentially through several convolutional layers through the connections between neurons until it reaches the final output. Convolutional neural networks can also calculate errors based on a specified optimization objective and iteratively update the neural network parameters through backpropagation and gradient descent to optimize the network.

[0166] like Figure 4 As shown in the figure, when an image of a man holding a cup and a backpack is input, the iteratively optimized convolutional neural network can recognize that there is a person, a container, and a bag in the image, and can extract the features of the person, container, and bag.

[0167] Second, the position coding information in the image to be detected is combined with the image features.

[0168] The position coding information may refer to the order of the currently detected image in a video, or may refer to the coordinate position of a pixel point in the image.

[0169] This step combines the position coding information of the image to be detected with the extracted image features. The specific operation is to add the coding sequences of the two.

[0170] Third, the combined information of image feature information and position coding information is input into the encoder for encoding processing to obtain a feature sequence with stronger representation ability.

[0171] The Transformer is an optional encoder that takes the summed feature information of the image to be detected and the positional encoding information and feeds it into its encoder for encoding. Through a series of encoding processes, including self-attention, summation and normalization, and a feedforward neural network, it ultimately outputs a feature sequence with enhanced representational capabilities.

[0172] Fourth, the encoded feature sequence is input into the decoder for decoding to obtain the visual vector of the person object and the visual vector of the object object;

[0173] The encoded feature sequence can be input into the Transformer for decoding to obtain the visual vector of the person object and the visual vector of the object object.

[0174] The Transformer is a specific form of neural network, consisting of an encoder and a decoder. Typically, the encoder and decoder are stacked together, with each layer having two sublayer connections. The first sublayer connection structure consists of a multi-head self-attention sublayer, a normalization layer, and a residual connection. The second sublayer connection structure consists of a feedforward fully connected sublayer, a normalization layer, and a residual connection.

[0175] Fifth, the category of the object is predicted through the visual vector of the object.

[0176] like Figure 4 As shown in the figure, an image of a man holding a cup and a backpack undergoes a series of encoding and decoding steps to output the man's visual vector, the first item's visual vector, and the second item's visual vector. The first item's visual vector can be used to predict its category as a container, while the second item's visual vector can be used to predict its category as a bag.

[0177] Through the above steps, the visual modality vector of the image to be detected is obtained from the input image to be detected.

[0178] Step S702 : obtaining the original vector of the candidate verb corresponding to the object, and obtaining the verb conditional probability of the candidate verb corresponding to the object relative to the object.

[0179] The candidate verbs refer to the semantics corresponding to all possible actions that can be performed on the object in the image to be detected. For example, the possible actions that can be performed on a horse include: riding a horse, leading a horse, feeding a horse, beating a horse, etc. The semantics corresponding to these actions include: riding, leading, feeding, beating, etc., so riding, leading, feeding, beating, etc. are the candidate verbs corresponding to horses. For another example: Figure 4 In the example, the actions that a man (person object) can perform on a bag (object object) include carrying on his back, slinging over his shoulder, lifting, and throwing, while the actions that he cannot perform include eating and drinking, so carrying on his back, slinging over his shoulder, lifting, and throwing are the candidate verbs corresponding to the bag (object object).

[0180] The raw vector of a verb refers to the original, unprocessed state of the verb vector. Each verb has its own corresponding vector. For example, the verb vector for "back" can be represented as [2, 0, 0, 0, ..., 0], the verb vector for "carry" can be represented as [0, 2, 0, 0, ..., 0], the verb vector for "lift" can be represented as [0, 0, 2, 0, ..., 0], and the verb vector for "throw" can be represented as [0, 0, 0, 2, ..., 0]. The vectors of different verbs are unlikely to be repeated, so the corresponding verbs can be represented by verb vectors.

[0181] This step obtains the original vector of the candidate verb, specifically, obtains the original vector of the candidate verb corresponding to the object in the image to be detected.

[0182] An optional implementation includes:

[0183] First, the verb vectors in the verb vector database are mapped to a neural network;

[0184] The Verb Vector Database is a public database of human interaction behaviors, including verb vectors for almost all verbs. The HICO-DET dataset and the V-COCO dataset are both good candidates for this type of database.

[0185] The purpose of this step is to make the verb vectors in the mapped verb vector database as close as possible to the co-occurrence probability of the verb.

[0186] Second, the original vector of the candidate verb is obtained from the mapped verb vector database according to the visual vector of the object.

[0187] According to the visual vector of the object in the image to be detected, the verb vector related to the object can be screened out from the above-mentioned mapped verb vector database, that is, the original vector of the candidate verb corresponding to the object.

[0188] like Figure 4 As shown in the figure, an image of a man holding a cup and a backpack is used as an example. The verb vector database includes verb vectors corresponding to verbs such as "take," "hold," "throw," "eat," "drink," "hit," "ride," and "grab." Image feature extraction and analysis show that the visual vector of the second object is a bag. Therefore, verb vectors corresponding to candidate verbs of the bag class are selected from the verb vector database, such as "take," "hold," "throw," and "grab," to form the original vectors of the candidate verbs corresponding to the object.

[0189] The verb conditional probability refers to the probability of event A occurring under the condition that event B occurs. In this embodiment, it refers to the probability of all actions that can be performed by the object in the image to be detected.

[0190] like Figure 4 As shown in the image of a man carrying a cup in his backpack, the cup can be subjected to various actions, including: holding, grabbing, throwing, riding, etc. The probability of these actions occurring is: holding 40%, grabbing 20%, throwing 10%, riding 1%. Of course, the cup can be subjected to more than the four actions listed above. In this embodiment, these four actions are used as examples for illustration.

[0191] All possible actions for a bag include: carrying, shouldering, lifting, and throwing. The probability of these actions occurring is: carrying 30%, shouldering 30%, lifting 20%, and throwing 5%. Of course, there are more than the four actions that can be performed on a bag. In this embodiment, the above four actions are used as examples for explanation.

[0192] Obtaining the verb conditional probability is to obtain the verb conditional probability of the candidate verb relative to the object according to the visual vector of the object.

[0193] That is to say, the verb conditional probability obtained in this step corresponds one-to-one to the original vector of the candidate verb. As shown in the following table:

[0194]

[0195] Step S703 , obtaining a semantic modal vector corresponding to the object according to the original vector of the candidate verb and the verb conditional probability; the semantic modal vector includes: a verb vector of the candidate verb corresponding to the object.

[0196] This step specifically refers to aggregating the original vectors of the candidate verbs corresponding to the object in the image to be detected and the verb conditional probability of the candidate verbs.

[0197] An optional aggregation method is to multiply the original vector of the candidate verb by the verb conditional probability, that is, to use the product of the original vector of the candidate verb and the verb conditional probability as the semantic modality vector corresponding to the object.

[0198] like Figure 4 As shown, in the image of a man holding a cup and a backpack, the candidate verbs for the first object (container type) are: take, lift, carry, etc., and the candidate verbs for the second object (bag type) are: carry on the back, sling over the shoulder, lift, throw, etc.

[0199] The verb conditional probabilities of the candidate verbs for the first object are: take 30%, lift 20%, carry 10%, and carry 10%; the verb conditional probabilities of the candidate verbs for the second object are: carry on the back 30%, carry over the shoulder 30%, lift 20%, and throw 5%.

[0200] The method of aggregating the original vectors of the candidate verbs corresponding to the first object and the verb conditional probabilities of the candidate verbs is: the original vector of "take" × 0.3, the original vector of "lift" × 0.2, the original vector of "lift" × 0.1, and the original vector of "resist" × 0.1. As shown in the following table:

[0201]

[0202] The method for aggregating the original vectors of the candidate verbs corresponding to the second object and the verb conditional probabilities of the candidate verbs is: the original vector of "carry" × 0.3, the original vector of "carry" × 0.3, the original vector of "lift" × 0.2, and the original vector of "throw" × 0.05. As shown in the following table:

[0203]

[0204] Through the above steps, the verb vector of the candidate verb corresponding to the object in the image to be detected can be obtained.

[0205] Step S704 : obtaining the verb category of the person object with respect to the object object according to the visual modality vector and the semantic modality vector.

[0206] Figure 8 The flowchart provided in this embodiment for obtaining the verb category of the person object for the object object in the image to be detected has the following specific steps:

[0207] Step S704 - 1 : performing inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector.

[0208] The calibration provided in this embodiment includes: using a channel attention mechanism to calibrate the visual modality vector and the semantic modality vector. Specifically:

[0209] Assume that the number of features of the visual modality vector and the semantic modality vector is K, and the number of channels of each feature is D.

[0210] Since the semantic modality vector is guided by the visual modality vector, the visual modality vector and the semantic modality vector should be in one-to-one correspondence. A channel attention mechanism is used to calibrate the corresponding channels of the visual modality vector and the semantic modality vector one-to-one.

[0211] For example, first, the channel attention size (D) is calculated using the semantic modality vector. Then, the channel attention size calculated using the semantic modality vector is used to adjust the channel attention size (D) of the visual modality vector. In this way, all features of the K pairs of visual modality vectors and semantic modality vectors are calibrated one by one. Similarly, the semantic modality vector can also be calibrated using the visual modality vector.

[0212] The calibration provided in this embodiment further includes: after performing inter-modality calibration on the visual modality vector and the semantic modality vector, performing intra-modality calibration on the visual modality vector using an information transfer mechanism. Specifically:

[0213] By using the Transformer encoder to transfer information between the K features of the visual modality vector, a visual modality vector with K features with stronger representation capabilities can be obtained.

[0214] The calibration provided in this embodiment further includes: after performing inter-modality calibration on the visual modality vector and the semantic modality vector, performing intra-modality calibration on the semantic modality vector using an information transfer mechanism. Specifically:

[0215] By using the Transformer encoder to transfer information between the K features of the semantic modal vector, a semantic modal vector with K features with stronger representation capabilities can be obtained.

[0216] The above provides a method for performing inter-modality calibration and intra-modality calibration on visual modality vectors and semantic modality vectors.

[0217] Step S704 - 2 : fusing the calibrated visual modality vector with the calibrated semantic modality vector to obtain verb features of the candidate verb.

[0218] The fusion refers to merging the calibrated visual modality vector and the calibrated semantic modality vector, which can be achieved through a feature fusion function.

[0219] That is, the fusion is to use the calibrated visual modality vector and the calibrated semantic modality vector as sequence elements to generate a verb sequence of the candidate verb.

[0220] Step S704 - 3 : predicting the verb category of the person object in the image to be detected with respect to the object object according to the verb features of the candidate verb.

[0221] An optional prediction method is to use a feed-forward neural network (FFN) to predict the verb category of the person object for the object object in the image to be detected based on the verb features of the candidate verbs.

[0222] like Figure 4 As shown, using a feedforward neural network based on the verb features of the candidate verbs, it can be predicted that in the image to be detected, the verb category of the person object for the first object is "take", and the verb category of the person object for the second object is "back".

[0223] The above is an implementation of the character interaction detection method provided in the second embodiment of the present application. Figure 4 As shown, the above method can detect an image of a man holding a cup and a backpack, and the final detection results include: "man holding a cup" and "man carrying a backpack." The detection results also include: a feature description of the person object (e.g., man, height approximately 170 cm, short hair), a feature description of the first object object (e.g., cup, covered, cylindrical, approximately 10 cm high), and a feature description of the second object object (e.g., backpack, blue, rectangular). The detection results also include: the positional relationship between the first object object and the person object (the cup is held in front of the person), and the positional relationship between the second object object and the person object (the bag is carried behind the person). Of course, the detection results can also include other information, which will not be explained in detail here.

[0224] The third embodiment of the present application provides a human interaction detection device. Figure 9 This is a schematic diagram of the structure of the human interaction detection device provided in this embodiment.

[0225] like Figure 9 As shown, the human interaction detection device provided by this embodiment includes: a visual modality unit 901, a semantic modality unit 902, a calibration unit 903, and a prediction unit 904;

[0226] The visual modality unit 901 is used to obtain a visual modality vector of the image to be detected; the visual modality vector includes: a visual vector of a person object and a visual vector of an object object;

[0227] The semantic modality unit 902 is configured to obtain a semantic modality vector corresponding to the object according to the visual vector of the object; the semantic modality vector includes: a verb vector of a candidate verb corresponding to the object;

[0228] Optionally, acquiring a semantic modality vector corresponding to the object according to the visual vector of the object includes:

[0229] Obtaining the original vector of the candidate verb corresponding to the object;

[0230] Obtaining a verb conditional probability of the candidate verb corresponding to the object relative to the object;

[0231] According to the original vector of the candidate verb and the verb conditional probability, a semantic modal vector corresponding to the object is obtained.

[0232] The calibration unit 903 is configured to perform inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector;

[0233] Optionally, the inter-modality calibration of the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector includes: using a channel attention mechanism to perform corresponding calibration on the visual modality vector and the semantic modality vector.

[0234] Optionally, the inter-modal calibration of the visual modal vector and the semantic modal vector to obtain a calibrated visual modal vector and a calibrated semantic modal vector also includes: after the inter-modal calibration of the visual modal vector and the semantic modal vector, intra-modal calibration of the visual modal vector is performed using an information transfer mechanism.

[0235] Optionally, the inter-modal calibration of the visual modal vector and the semantic modal vector to obtain a calibrated visual modal vector and a calibrated semantic modal vector also includes: after the inter-modal calibration of the visual modal vector and the semantic modal vector, intra-modal calibration of the semantic modal vector is performed using an information transfer mechanism.

[0236] The prediction unit 904 is configured to predict a verb category of the person object in the image to be detected with respect to the object object based on the calibrated visual modality vector and the calibrated semantic modality vector.

[0237] Optionally, predicting a verb category of the person object in the image to be detected with respect to the object object based on the calibrated visual modality vector and the calibrated semantic modality vector includes:

[0238] fusing the calibrated visual modality vector and the calibrated semantic modality vector to obtain verb features of the candidate verb;

[0239] The verb category of the person object in the image to be detected for the object object is predicted according to the verb features of the candidate verbs.

[0240] Optionally, the calibrated visual modal vector and the calibrated semantic modal vector are fused to obtain the verb features of the candidate verb, including: using the calibrated visual modal vector and the calibrated semantic modal vector as sequence elements to generate a verb sequence of the candidate verb.

[0241] The fourth embodiment of the present application provides a human interaction detection device. Figure 10 This is a schematic diagram of the structure of the human interaction detection device provided in this embodiment.

[0242] like Figure 10 As shown, the human interaction detection device provided by this embodiment includes: a visual modality unit 1001, a semantic modality unit 1002, and a verb category acquisition unit 1003;

[0243] The visual modality unit 1001 is used to obtain a visual modality vector of the image to be detected; the visual modality vector includes: a visual vector of a person object and a visual vector of an object object;

[0244] The semantic modality unit 1002 is configured to obtain the original vector of the candidate verb corresponding to the object, and obtain the verb conditional probability of the candidate verb corresponding to the object relative to the object;

[0245] Optionally, obtaining the original vector of the candidate verb corresponding to the object includes: obtaining the original vector of the candidate verb from a verb vector database according to the visual vector of the object.

[0246] Optionally, obtaining the verb conditional probability of the candidate verb corresponding to the object relative to the object includes: obtaining the verb conditional probability of the candidate verb relative to the object according to the visual vector of the object.

[0247] The semantic modality unit 1002 is further configured to obtain a semantic modality vector corresponding to the object based on the original vector of the candidate verb and the verb conditional probability; the semantic modality vector includes: a verb vector of the candidate verb corresponding to the object;

[0248] Optionally, obtaining the semantic modal vector corresponding to the object based on the original vector of the candidate verb and the verb conditional probability includes: taking the product of the original vector of the candidate verb and the verb conditional probability as the semantic modal vector corresponding to the object.

[0249] The verb category acquiring unit 1003 is configured to acquire the verb category of the person object with respect to the object object according to the visual modality vector and the semantic modality vector.

[0250] Optionally, obtaining the verb category of the person object with respect to the object object according to the visual modality vector and the semantic modality vector includes:

[0251] Performing inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector;

[0252] The verb category of the person object in the image to be detected with respect to the object object is predicted according to the calibrated visual modality vector and the calibrated semantic modality vector.

[0253] The fifth embodiment of the present application provides a virtual reality device. Figure 11 This is a schematic diagram of the structure of the virtual reality device provided in this embodiment.

[0254] like Figure 11 As shown, the virtual reality device provided in this embodiment includes: a memory 1101 and a processor 1102; the memory 1101 stores a computer instruction set, and when executed by the processor 1102, performs the following steps:

[0255] Obtaining a visual modality vector of the image to be detected; the visual modality vector includes: a visual vector of a person object and a visual vector of an object object;

[0256] Acquire a semantic modal vector corresponding to the object according to the visual vector of the object; the semantic modal vector includes: a verb vector of a candidate verb corresponding to the object;

[0257] Performing inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector;

[0258] The verb category of the person object in the image to be detected with respect to the object object is predicted according to the calibrated visual modality vector and the calibrated semantic modality vector.

[0259] Optionally, the inter-modality calibration of the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector includes: using a channel attention mechanism to perform corresponding calibration on the visual modality vector and the semantic modality vector.

[0260] Optionally, the inter-modal calibration of the visual modal vector and the semantic modal vector to obtain a calibrated visual modal vector and a calibrated semantic modal vector also includes: after the inter-modal calibration of the visual modal vector and the semantic modal vector, intra-modal calibration of the visual modal vector is performed using an information transfer mechanism.

[0261] Optionally, the inter-modal calibration of the visual modal vector and the semantic modal vector to obtain a calibrated visual modal vector and a calibrated semantic modal vector also includes: after the inter-modal calibration of the visual modal vector and the semantic modal vector, intra-modal calibration of the semantic modal vector is performed using an information transfer mechanism.

[0262] Optionally, predicting a verb category of the person object in the image to be detected with respect to the object object based on the calibrated visual modality vector and the calibrated semantic modality vector includes:

[0263] fusing the calibrated visual modality vector and the calibrated semantic modality vector to obtain verb features of the candidate verb;

[0264] The verb category of the person object in the image to be detected for the object object is predicted according to the verb features of the candidate verbs.

[0265] Optionally, the calibrated visual modal vector and the calibrated semantic modal vector are fused to obtain the verb features of the candidate verb, including: using the calibrated visual modal vector and the calibrated semantic modal vector as sequence elements to generate a verb sequence of the candidate verb.

[0266] Optionally, acquiring a semantic modality vector corresponding to the object according to the visual vector of the object includes:

[0267] Obtaining the original vector of the candidate verb corresponding to the object;

[0268] Obtaining a verb conditional probability of the candidate verb corresponding to the object relative to the object;

[0269] According to the original vector of the candidate verb and the verb conditional probability, a semantic modal vector corresponding to the object is obtained.

[0270] A sixth embodiment of the present application further provides a virtual reality device, comprising: a memory and a processor; the memory storing a computer instruction set, which, when executed by the processor, performs the following steps:

[0271] Obtaining a visual modality vector of the image to be detected; the visual modality vector includes: a visual vector of a person object and a visual vector of an object object;

[0272] Obtaining the original vector of the candidate verb corresponding to the object, and obtaining the verb conditional probability of the candidate verb corresponding to the object relative to the object;

[0273] Obtaining a semantic modal vector corresponding to the object according to the original vector of the candidate verb and the verb conditional probability; the semantic modal vector includes: a verb vector of the candidate verb corresponding to the object;

[0274] The verb category of the person object with respect to the object object is obtained according to the visual modality vector and the semantic modality vector.

[0275] Optionally, obtaining the original vector of the candidate verb corresponding to the object includes: obtaining the original vector of the candidate verb from a verb vector database according to the visual vector of the object.

[0276] Optionally, obtaining the verb conditional probability of the candidate verb corresponding to the object relative to the object includes: obtaining the verb conditional probability of the candidate verb relative to the object according to the visual vector of the object.

[0277] Optionally, obtaining the semantic modal vector corresponding to the object based on the original vector of the candidate verb and the verb conditional probability includes: taking the product of the original vector of the candidate verb and the verb conditional probability as the semantic modal vector corresponding to the object.

[0278] Optionally, obtaining the verb category of the person object with respect to the object object according to the visual modality vector and the semantic modality vector includes:

[0279] Performing inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector;

[0280] The verb category of the person object in the image to be detected with respect to the object object is predicted according to the calibrated visual modality vector and the calibrated semantic modality vector.

[0281] A seventh embodiment of the present application provides an electronic device. Figure 12 This is a schematic diagram of the structure of the electronic device provided in this embodiment.

[0282] like Figure 12As shown, the electronic device provided by this embodiment includes: a collector 1201, a memory 1202 and a processor 1203.

[0283] The collector 1201 is used to collect images to be detected.

[0284] The memory 1202 is used to store computer instructions for executing the human interaction detection method.

[0285] The processor 1203 is configured to execute computer instructions stored in the memory to perform the methods described in the first and second embodiments of the present application.

[0286] The eighth embodiment of the present application provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed by a processor, they are used to implement the technical solution of any one of the character interaction detection methods in the first and second embodiments of the present application.

[0287] Although the present application is disclosed as above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.

Claims

1. A method for detecting human interaction, characterized in that: include: Obtain the visual modality vector of the image to be detected; The visual modality vector includes: a visual vector of a person object and a visual vector of an object object; Acquire a semantic modal vector corresponding to the object according to the visual vector of the object; the semantic modal vector includes: a verb vector of a candidate verb corresponding to the object; performing inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector, wherein the visual modality vector and the semantic modality vector have a one-to-one correspondence; fusing the calibrated visual modality vector with the calibrated semantic modality vector to obtain verb features of candidate verbs, wherein the candidate verbs are used to represent the semantics corresponding to the set of actions to be performed by the object; The verb category of the person object in the image to be detected for the object object is predicted according to the verb features of the candidate verbs.

2. The method according to claim 1, characterized in that The inter-modality calibration of the visual modality vector and the semantic modality vector to obtain the calibrated visual modality vector and the calibrated semantic modality vector includes: using a channel attention mechanism to perform corresponding calibration on the visual modality vector and the semantic modality vector.

3. The method according to claim 1, characterized in that The inter-modal calibration of the visual modal vector and the semantic modal vector to obtain a calibrated visual modal vector and a calibrated semantic modal vector also includes: after the inter-modal calibration of the visual modal vector and the semantic modal vector, intra-modal calibration of the visual modal vector is performed using an information transmission mechanism.

4. The method according to claim 1, wherein The inter-modal calibration of the visual modal vector and the semantic modal vector to obtain a calibrated visual modal vector and a calibrated semantic modal vector also includes: after the inter-modal calibration of the visual modal vector and the semantic modal vector, intra-modal calibration of the semantic modal vector is performed using an information transfer mechanism.

5. The method according to claim 1, wherein The fusing of the calibrated visual modality vector and the calibrated semantic modality vector to obtain the verb features of the candidate verb includes: using the calibrated visual modality vector and the calibrated semantic modality vector as sequence elements to generate a verb sequence of the candidate verb.

6. The method according to claim 1, characterized in that The acquiring, according to the visual vector of the object, a semantic modality vector corresponding to the object, includes: Obtaining the original vector of the candidate verb corresponding to the object; Obtaining a verb conditional probability of the candidate verb corresponding to the object relative to the object; According to the original vector of the candidate verb and the verb conditional probability, a semantic modal vector corresponding to the object is obtained.

7. A method for detecting human interaction, characterized in that: include: Obtain the visual modality vector of the image to be detected; The visual modality vector includes: a visual vector of a person object and a visual vector of an object object; Obtaining the original vector of the candidate verb corresponding to the object, and obtaining the verb conditional probability of the candidate verb corresponding to the object relative to the object; Obtaining a semantic modal vector corresponding to the object according to the original vector of the candidate verb and the verb conditional probability; the semantic modal vector includes: a verb vector of the candidate verb corresponding to the object; performing inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector, wherein the visual modality vector and the semantic modality vector have a one-to-one correspondence; fusing the calibrated visual modality vector and the calibrated semantic modality vector to obtain verb features of candidate verbs, wherein the candidate verbs are used to represent the semantics corresponding to the set of actions to be performed by the object; The verb category of the person object in the image to be detected for the object object is predicted according to the verb features of the candidate verbs.

8. The method according to claim 7, characterized in that The obtaining of the original vector of the candidate verb corresponding to the object includes: obtaining the original vector of the candidate verb from a verb vector database according to the visual vector of the object.

9. The method according to claim 7, characterized in that The acquiring the verb conditional probability of the candidate verb corresponding to the object relative to the object includes: acquiring the verb conditional probability of the candidate verb relative to the object according to the visual vector of the object.

10. The method according to claim 7, characterized in that The obtaining of the semantic modal vector corresponding to the object according to the original vector of the candidate verb and the verb conditional probability includes: taking the product of the original vector of the candidate verb and the verb conditional probability as the semantic modal vector corresponding to the object.

11. A person interaction detection device, characterized in that: include: Visual modality unit, semantic modality unit, calibration unit, prediction unit; The visual modality unit is used to obtain a visual modality vector of the image to be detected; The visual modality vector includes: a visual vector of a person object and a visual vector of an object object; The semantic modality unit is configured to obtain a semantic modality vector corresponding to the object according to the visual vector of the object; the semantic modality vector includes: a verb vector of a candidate verb corresponding to the object; The calibration unit is configured to perform inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector, wherein the visual modality vector and the semantic modality vector have a one-to-one correspondence; The prediction unit is used to fuse the calibrated visual modality vector with the calibrated semantic modality vector to obtain verb features of candidate verbs, wherein the candidate verbs are used to represent the semantics corresponding to the set of actions to be performed by the object object; and predict the verb category of the person object in the image to be detected for the object object based on the verb features of the candidate verbs.

12. A person interaction detection device, characterized in that: include: Visual modality unit, semantic modality unit, verb category acquisition unit; The visual modality unit is used to obtain a visual modality vector of the image to be detected; The visual modality vector includes: a visual vector of a person object and a visual vector of an object object; The semantic modality unit is used to obtain the original vector of the candidate verb corresponding to the object, and obtain the verb conditional probability of the candidate verb corresponding to the object relative to the object; The semantic modality unit is further configured to obtain a semantic modality vector corresponding to the object based on the original vector of the candidate verb and the verb conditional probability; the semantic modality vector includes: a verb vector of the candidate verb corresponding to the object; The verb category acquisition unit is used to perform inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector, wherein the visual modality vector and the semantic modality vector have a one-to-one correspondence; the calibrated visual modality vector and the calibrated semantic modality vector are fused to obtain verb features of candidate verbs, wherein the candidate verbs are used to represent the semantics corresponding to the set of actions to be performed by the object; and the verb category of the person object in the image to be detected relative to the object object is predicted based on the verb features of the candidate verbs.

13. A virtual reality device, characterized in that: include: memory and processor; The memory stores a computer instruction set, which, when executed by the processor, performs the following steps: Obtain the visual modality vector of the image to be detected; The visual modality vector includes: a visual vector of a person object and a visual vector of an object object; Acquire a semantic modal vector corresponding to the object according to the visual vector of the object; the semantic modal vector includes: a verb vector of a candidate verb corresponding to the object; performing inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector, wherein the visual modality vector and the semantic modality vector have a one-to-one correspondence; fusing the calibrated visual modality vector with the calibrated semantic modality vector to obtain verb features of candidate verbs, wherein the candidate verbs are used to represent the semantics corresponding to the set of actions to be performed by the object; The verb category of the person object in the image to be detected for the object object is predicted according to the verb features of the candidate verbs.

14. A virtual reality device, characterized in that: include: memory and processor; The memory stores a computer instruction set, which, when executed by the processor, performs the following steps: Obtain the visual modality vector of the image to be detected; The visual modality vector includes: a visual vector of a person object and a visual vector of an object object; Obtaining the original vector of the candidate verb corresponding to the object, and obtaining the verb conditional probability of the candidate verb corresponding to the object relative to the object; Obtaining a semantic modal vector corresponding to the object according to the original vector of the candidate verb and the verb conditional probability; the semantic modal vector includes: a verb vector of the candidate verb corresponding to the object; Obtaining a verb category of the person object with respect to the object object according to the visual modality vector and the semantic modality vector; performing inter-modality calibration on the visual modality vector and the semantic modality vector to obtain a calibrated visual modality vector and a calibrated semantic modality vector, wherein the visual modality vector and the semantic modality vector have a one-to-one correspondence; fusing the calibrated visual modality vector and the calibrated semantic modality vector to obtain verb features of candidate verbs, wherein the candidate verbs are used to represent the semantics corresponding to the set of actions to be performed by the object; The verb category of the person object in the image to be detected for the object object is predicted according to the verb features of the candidate verbs.

15. An electronic device, characterized in that: include: Collector, processor and memory; The collector is used to collect images to be detected; The memory is used to store one or more computer instructions; The processor is configured to execute the one or more computer instructions to implement the method according to any one of claims 1 to 10.

16. A computer-readable storage medium having one or more computer instructions stored thereon, characterized in that: The instruction is executed by a processor to implement the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Self-learning event extraction method and application thereof

    CN111881258A