Machine learning model training method and apparatus, interaction relationship detection method and apparatus

Through the machine learning model training method, the self-attention and cross-attention models are used to process the pair of characters and objects in the image, solving the problem that the existing HOI detection methods cannot be trained end-to-end and learn difficult, and achieving efficient and accurate interaction relationship detection.

CN114821634BActive Publication Date: 2025-06-17JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210324555.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-06-17
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

The existing HOI detection methods have problems such as inability to implement end-to-end training, network learning is difficult and slow convergence.

Method used

Using the machine learning model training method, by obtaining pairs of characters and objects in the sample image, using the first machine learning model to process the feature vector and select interactive proposals, combining self-attention and cross-attention models for attention processing, determining the interaction probability value, and training the model through loss function.

Benefits of technology

It improves the accuracy and efficiency of HOI detection, reduces the difficulty of learning, and realizes end-to-end interactive relationship detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821634B_ABST
    Figure CN114821634B_ABST
Patent Text Reader

Abstract

The present disclosure provides a machine learning model training method and apparatus, and an interaction relationship detection method and apparatus, relating to the field of artificial intelligence. The machine learning model training method includes: obtaining all pairs of person objects in a sample image; using a first machine learning model to obtain a first interaction probability value for each pair of person objects; determining a first loss function according to the first interaction probability value of each pair of person objects; selecting a predetermined number of interaction proposals from all pairs of person objects, where the predetermined number of interaction proposals are the top predetermined number of pairs of person objects in all pairs of person objects sorted in descending order of the first interaction probability value; using a second machine learning model to obtain a second interaction probability value for each interaction proposal; determining a second loss function according to the second interaction probability value of each interaction proposal; training the first machine learning model and the second machine learning model according to a first target loss function determined by the first loss function and the second loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and particularly to a method and device for training a machine learning model, and a method and device for detecting interaction relationships. Background Art

[0002] The purpose of HOI (Human Object Interaction) detection is to locate pairs of human objects with interaction relationships in an image, identify specific interaction categories, and finally output a triple in the form of (human, object, interaction).

[0003] Existing HOI detection methods can be roughly divided into two types. The first method uses an indirect prediction method to achieve HOI detection by solving some proxy regression or classification problems. The second method models HOI detection as a direct set prediction problem, so that the prediction result can be optimized "end-to-end". Summary of the Invention

[0004] The inventors noticed that the first method mentioned above needs to merge similar prediction results and requires a manually designed matching method to associate "interaction" with its associated "person" and "object". Therefore, the overall framework cannot achieve "end-to-end" training and is prone to suboptimal solutions. The interaction query information of human objects used in the second method is a learnable parameter initialized randomly, so the network learning is very difficult and the convergence is slow.

[0005] Accordingly, the present disclosure provides a machine learning model training solution that can obtain accurate interaction relationship detection results.

[0006] According to a first aspect of an embodiment of the present disclosure, there is provided a method for training a machine learning model, including: obtaining all pairs of human objects in a sample image; using a first machine learning model to process the constructed feature vectors of each pair of human objects to obtain a first interaction probability value for each pair of human objects; determining a first loss function according to the first interaction probability value of each pair of human objects and the interaction annotation result; selecting a predetermined number of interaction proposals from all pairs of human objects, where the predetermined number of interaction proposals are the first predetermined number of pairs of human objects sorted from large to small according to the first interaction probability value among all pairs of human objects; using a second machine learning model to perform attention processing on the predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information to obtain a second interaction probability value for each interaction proposal; determining a second loss function according to the second interaction probability value of each interaction proposal and the interaction annotation result; determining a first target loss function according to the first loss function and the second loss function; training the first machine learning model and the second machine learning model using the first target loss function.

[0007] In some embodiments, the attention processing of the predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information by using the second machine learning model includes: performing self-attention processing on the predetermined number of interaction proposals by using the self-attention model in the second machine learning model to obtain intermediate features of the first interaction proposals; and performing cross-attention processing on the intermediate features of the first interaction proposals, the image feature map of the sample image, and the position encoding information by using the cross-attention model in the second machine learning model to obtain the interaction probability values of each interaction proposal.

[0008] In some embodiments, during the self-attention processing, the attention weight values of the self-attention model are associated with the predetermined number of interaction proposals; during the cross-attention processing, the attention weight values of the cross-attention model are associated with the intermediate features of the first interaction proposals, the image feature map of the sample image, and the position encoding information.

[0009] In some embodiments, the processing of the constructed feature vectors of each pair of person objects by using the first machine learning model includes: detecting the first intersection over union ratio of the person bounding box and the corresponding person annotation bounding box in the i-th pair of person objects, where 1 ≤ i ≤ N and N is the total number of pairs of person objects; detecting the second intersection over union ratio of the object bounding box and the corresponding object annotation bounding box in the i-th pair of person objects; if both the first intersection over union ratio and the second intersection over union ratio are greater than a preset threshold, then taking the i-th pair of person objects as a positive sample; if at least one of the first intersection over union ratio and the second intersection over union ratio is not greater than the preset threshold, then taking the i-th pair of person objects as a negative sample; and processing the positive samples and the negative samples by using the first machine learning model to obtain the first interaction probability values of each pair of person objects.

[0010] In some embodiments, the processing of the positive samples and the negative samples by using the first machine learning model includes: sampling the negative samples by using a hard negative mining strategy to obtain a negative sample sampling result; and processing the positive samples and the negative sample sampling result by using the first machine learning model to obtain the first interaction probability values of each pair of person objects.

[0011] In some embodiments, the constructed feature vectors of each pair of person objects include the appearance features, spatial features, and semantic features of each pair of person objects.

[0012] In some embodiments, the first target loss function is a weighted sum of the first loss function and the second loss function.

[0013] In some embodiments, obtaining all pairs of human objects in the sample image includes: processing the sample image to obtain an image feature map and position encoding information of the sample image; detecting all humans and all objects in the sample image by using the image feature map and the position encoding information; and generating all pairs of human objects by using the all humans and the all objects.

[0014] In some embodiments, generating all pairs of human objects by using the all humans and the all objects includes: combining any two humans among the all humans, and combining any one human among the all humans with any one object among the all objects to generate the all pairs of human objects.

[0015] In some embodiments, the above method further includes: determining a semantic dependency relationship between any two of the predetermined number of interaction proposals; performing attention processing on the predetermined number of interaction proposals, the semantic dependency relationship, the image feature map and the position encoding information of the sample image by using a second machine learning model to obtain a third interaction probability value for each interaction proposal; determining a third loss function according to the third interaction probability value of each interaction proposal and the interaction annotation result; determining a second target loss function according to the first loss function and the third loss function; and training the first machine learning model and the second machine learning model by using the second target loss function.

[0016] In some embodiments, performing attention processing on the predetermined number of interaction proposals, the semantic dependency relationship, the image feature map and the position encoding information of the sample image by using a second machine learning model includes: performing self-attention processing on the predetermined number of interaction proposals and the semantic dependency relationship by using a self-attention model in the second machine learning model to obtain intermediate features of the second interaction proposals; and performing cross-attention processing on the intermediate features of the second interaction proposals, the image feature map of the sample image and the position encoding information by using a cross-attention model in the second machine learning model to obtain the third interaction probability value for each interaction proposal.

[0017] In some embodiments, during the self-attention processing, the attention weight value of the self-attention model is associated with the predetermined number of interaction proposals and the semantic dependency relationship; and during the cross-attention processing, the attention weight value of the cross-attention model is associated with the intermediate features of the second interaction proposals, the image feature map of the sample image and the position encoding information.

[0018] In some embodiments, the semantic dependency relationships include separated dependency relationships, dependency relationships with the same person, dependency relationships with the same object, cascaded dependency relationships, reverse cascaded dependency relationships, or dependency relationships of the same pair.

[0019] In some embodiments, the second target loss function is a weighted sum of the first loss function and the third loss function.

[0020] In some embodiments, the above method further includes: after determining the semantic dependency relationship between any two of the predetermined number of interaction proposals, determining the spatial structure information of each interaction proposal; using a second machine learning model to perform attention processing on the predetermined number of interaction proposals, the semantic dependency relationship, the spatial structure information, the image feature map of the sample image, and the position encoding information to obtain a fourth interaction probability value for each interaction proposal; determining a fourth loss function according to the fourth interaction probability value of each interaction proposal and the interaction annotation result; determining a third target loss function according to the first loss function and the fourth loss function; and training the first machine learning model and the second machine learning model using the third target loss function.

[0021] In some embodiments, the using the second machine learning model to perform attention processing on the predetermined number of interaction proposals, the semantic dependency relationship, the spatial structure information, the image feature map of the sample image, and the position encoding information includes: using the self-attention model in the second machine learning model to perform self-attention processing on the predetermined number of interaction proposals and the semantic dependency relationship to obtain intermediate features of the third interaction proposal; and using the cross-attention model in the second machine learning model to perform cross-attention processing on the intermediate features of the third interaction proposal, the spatial structure information, the image feature map of the sample image, and the position encoding information to obtain a third interaction probability value for each interaction proposal.

[0022] In some embodiments, during the self-attention processing, the attention weight value of the self-attention model is associated with the predetermined number of interaction proposals and the semantic dependency relationship; during the cross-attention processing, the attention weight value of the cross-attention model is associated with the intermediate features of the second interaction proposal, the spatial structure information, the image feature map of the sample image, and the position encoding information.

[0023] In some embodiments, the spatial structure information includes the background area identifier of the interaction proposal, the union area identifier of the areas where the people and objects are located in the interaction proposal, the area identifier of the area where the people are located in the interaction proposal, the area identifier of the area where the interaction proposal is located, and the intersection area identifier of the areas where the people and objects are located in the interaction proposal.

[0024] In some embodiments, the third target loss function is a weighted sum of the first loss function and the fourth loss function.

[0025] According to a second aspect of the embodiments of the present disclosure, there is provided a machine learning model training device, including: a first training module configured to obtain all person-object pairs in a sample image; a second training module configured to process the constructed feature vectors of each person-object pair by using a first machine learning model to obtain a first interaction probability value of each person-object pair, determine a first loss function according to the first interaction probability value of each person-object pair and the interaction annotation result, and select a predetermined number of interaction proposals from all the person-object pairs, where the predetermined number of interaction proposals are the first predetermined number of person-object pairs sorted from largest to smallest in terms of the first interaction probability value among all the person-object pairs; a third training module configured to perform attention processing on the predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information by using a second machine learning model to obtain a second interaction probability value of each interaction proposal, and determine a second loss function according to the second interaction probability value of each interaction proposal and the interaction annotation result; a fourth training module configured to determine a first target loss function according to the first loss function and the second loss function, and use the first target loss function to train the first machine learning model and the second machine learning model.

[0026] According to a third aspect of the embodiments of the present disclosure, there is provided a machine learning model training device, including: a memory configured to store instructions; a processor coupled to the memory, and the processor is configured to execute the method as described in any one of the above embodiments based on the instructions stored in the memory.

[0027] According to a fourth aspect of the embodiments of the present disclosure, there is provided an interaction relationship detection method, including: obtaining all pairs of human object in the image to be processed; inputting the constructed feature vectors of each pair of human objects into a first machine learning model, so that the first machine learning module outputs a first interaction probability value for each pair of human objects, where the first machine learning model is trained by using the machine learning model training method described in any of the above embodiments; selecting a predetermined number of interaction proposals from all pairs of human objects, where the predetermined number of interaction proposals are the top predetermined number of pairs of human objects sorted from largest to smallest in terms of the first interaction probability value among all pairs of human objects; inputting the predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information into a second machine learning model, so that the second machine learning model outputs the interaction relationship of each interaction proposal, where the second machine learning model is trained by using the machine learning model training method described in any of the above embodiments.

[0028] In some embodiments, the above method further includes: determining the semantic dependency relationship between any two interaction proposals among the predetermined number of interaction proposals; inputting the predetermined number of interaction proposals, the semantic dependency relationship, the image feature map of the sample image, and the position encoding information into the second machine learning model, so that the second machine learning model outputs the interaction relationship of each interaction proposal.

[0029] In some embodiments, the above method further includes: after determining the semantic dependency relationship between any two interaction proposals among the predetermined number of interaction proposals, determining the spatial structure information of each interaction proposal; inputting the predetermined number of interaction proposals, the semantic dependency relationship, the spatial structure information, the image feature map of the sample image, and the position encoding information into the second machine learning model, so that the second machine learning model outputs the interaction relationship of each interaction proposal.

[0030] According to a fifth aspect of the embodiments of the present disclosure, there is provided an interaction relationship detection device, including: a first processing module configured to obtain all pairs of person objects in an image to be processed; a second processing module configured to input the constructed feature vectors of each pair of person objects into a first machine learning model, so that the first machine learning module outputs a first interaction probability value for each pair of person objects, where the first machine learning model is trained by using the machine learning model training method described in any of the above embodiments, and select a predetermined number of interaction proposals from all the pairs of person objects, where the predetermined number of interaction proposals are the top predetermined number of pairs of person objects sorted from large to small in terms of the first interaction probability value among all the pairs of person objects; a third processing module configured to input the predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information into a second machine learning model, so that the second machine learning model outputs the interaction relationship of each interaction proposal, where the second machine learning model is trained by using the machine learning model training method described in any of the above embodiments.

[0031] According to a sixth aspect of the embodiments of the present disclosure, there is provided an interaction relationship detection device, including: a memory configured to store instructions; a processor coupled to the memory, and the processor is configured to execute based on the instructions stored in the memory to implement the method described in any of the above embodiments.

[0032] According to a seventh aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the method described in any of the above embodiments is implemented.

[0033] Through the following detailed description of the exemplary embodiments of the present disclosure with reference to the accompanying drawings, other features and advantages of the present disclosure will become clear. Description of the Drawings

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 It is a schematic flow chart of a machine learning model training method according to an embodiment of the present disclosure;

[0036] Figure 2 It is a schematic diagram of a sample image according to an embodiment of the present disclosure;

[0037] Figure 3 It is a schematic diagram of a training network architecture according to an embodiment of the present disclosure;

[0038] Figure 4 Schematic flowchart of a machine learning model training method according to another embodiment of the present disclosure;

[0039] Figures 5A to 5F Schematic diagram of semantic dependency relationships according to some embodiments of the present disclosure;

[0040] Figure 6 Schematic diagram of a training network architecture according to another embodiment of the present disclosure;

[0041] Figure 7 Schematic flowchart of a machine learning model training method according to still another embodiment of the present disclosure;

[0042] Figure 8 Schematic diagram of a spatial structure according to an embodiment of the present disclosure;

[0043] Figure 9 Schematic diagram of a training network architecture according to still another embodiment of the present disclosure;

[0044] Figure 10 Schematic diagram of the structure of a machine learning model training apparatus according to an embodiment of the present disclosure;

[0045] Figure 11 Schematic diagram of the structure of a machine learning model training apparatus according to another embodiment of the present disclosure;

[0046] Figure 12 Schematic flowchart of an interaction relationship detection method according to an embodiment of the present disclosure;

[0047] Figure 13 Schematic flowchart of an interaction relationship detection method according to another embodiment of the present disclosure;

[0048] Figure 14 Schematic flowchart of an interaction relationship detection method according to still another embodiment of the present disclosure;

[0049] Figure 15 Schematic diagram of the structure of an interaction relationship detection apparatus according to an embodiment of the present disclosure;

[0050] Figure 16 Schematic diagram of the structure of an interaction relationship detection apparatus according to another embodiment of the present disclosure. Detailed implementation manners

[0051] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and in no way restricts the present disclosure and its application or use. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.

[0052] Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.

[0053] At the same time, it should be understood that, for the sake of convenience of description, the sizes of the various parts shown in the drawings are not drawn in actual proportional relationship.

[0054] For technologies, methods and devices known to those of ordinary skill in the relevant art, they may not be discussed in detail, but where appropriate, the said technologies, methods and devices should be regarded as part of the authorization specification.

[0055] In all the examples shown and discussed here, any specific value should be construed as merely exemplary, rather than as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0056] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0057] Figure 1 It is a schematic flowchart of a machine learning model training method according to an embodiment of the present disclosure. In some embodiments, the following machine learning model training method is executed by a machine learning model training device.

[0058] In step 101, obtain all pairs of human objects in the sample image.

[0059] In some embodiments, by processing the sample image to obtain an image feature map and position encoding information of the sample image, and using the image feature map and position encoding information to detect all the people and all the items in the sample image.

[0060] It should be noted that since how to detect all the people and all the items in the sample image is not the inventive point of the present disclosure, it will not be described in detail here.

[0061] Next, all person-object pairs are generated using all the persons and all the objects.

[0062] For example, any two persons among all the persons are combined, and any one person among all the persons is combined with any one object among all the objects to generate all person-object pairs.

[0063] As Figure 2 shown, there are two persons playing baseball in the image. Among them, person 1 on the right side of the image is holding a baseball bat 7, person 2 on the left side of the image is wearing a glove 8, there are also 4 spectators (i.e., persons 3 - 6) above the image, and in addition, there is a baseball bat 9 placed on the ground. For example, through the above processing, Figure 2 all the person-object pairs included are shown in Table 1.

[0064]

[0065]

[0066] Table 1

[0067] Return Figure 1 . In step 102, the first machine learning model is used to process the constructed feature vectors of each person-object pair to obtain the first interaction probability value of each person-object pair.

[0068] In some embodiments, the constructed feature vectors of each person-object pair include the appearance features, spatial features, and semantic features of each person-object pair.

[0069] For example, the appearance features are directly spliced using the corresponding feature vectors of the person and the object (i.e., the 256-dimensional features before the output layer). Let the center coordinates of the normalized person bounding box and the object bounding box be and respectively, then the spatial features are represented as a set of geometric quantities, i.e., [dx, dy, dis, arctan(dy / dx), A h , A o , I, U], where A h , A o , I, U represent the areas of the person, the object, and the intersection and union regions of the two respectively. The semantic features are obtained by using the embedding layer to embed the names of the interacting object categories into 300-dimensional vectors.

[0070] In some embodiments, the Intersection over Union (IOU) of the person bounding box and the corresponding person annotation bounding box in the i-th pair of person objects is detected, where 1 ≤ i ≤ N and N is the total number of pairs of person objects. The Intersection over Union of the object bounding box and the corresponding object annotation bounding box in the i-th pair of person objects is detected. If both the first Intersection over Union and the second Intersection over Union are greater than a preset threshold (e.g., 0.5), the i-th pair of person objects is taken as a positive sample. If at least one of the first Intersection over Union and the second Intersection over Union is not greater than the preset threshold, the i-th pair of person objects is taken as a negative sample. The first machine learning model is used to process the positive samples and negative samples to obtain the first interaction probability value for each pair of person objects.

[0071] In some embodiments, to improve the processing effect, the hard mining strategy is used to sample the negative samples to obtain the negative sample sampling result. Further, the first machine learning model is used to process the positive samples and the negative sample sampling result to obtain the first interaction probability value for each pair of person objects.

[0072] For example, according to the above processing, Figure 2 the interaction probability values of the pairs of person objects included in are shown in Table 2.

[0073] Person object pair Interaction probability value 1 Person 2 - Glove 8 0.9 2 Person 1 - Bat 7 0.9 3 Person 2 - Bat 9 0.7 4 Person 1 - Bat 9 0.6 5 Person 3 - Bat 7 0.1 … … …

[0074] Table 2

[0075] In step 103, a first loss function is determined according to the first interaction probability value and the interaction annotation result of each pair of person objects.

[0076] In some embodiments, the first loss function is FL (focal loss function), as shown in formula (1).

[0077]

[0078] where FL is the focal loss function, is the interaction probability value of the i-th pair of person objects, and z i is the interaction annotation result, and z i ∈ {0, 1}.

[0079] In step 104, a predetermined number of interaction proposals are selected from all pairs of person objects, where the predetermined number of interaction proposals are the top predetermined number of pairs of person objects in all pairs of person objects sorted by the first interaction probability value from large to small.

[0080] It should be noted that if the interaction probability value of a certain pair of character objects is low, it indicates that the possibility of an interaction relationship between the characters and objects in the pair of character objects is low. Only using the pairs of character objects with high interaction probability values as interaction proposals can effectively improve the processing efficiency.

[0081] In step 105, the second machine learning model is used to perform attention processing on a predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information, so as to obtain the second interaction probability value of each interaction proposal.

[0082] In some embodiments, the self-attention model in the second machine learning model is used to perform self-attention processing on a predetermined number of interaction proposals to obtain intermediate features of the first interaction proposals.

[0083] During the self-attention processing, the attention weight values of the self-attention model are associated with a predetermined number of interaction proposals.

[0084] For example, the attention weight values of the self-attention model are as shown in formula (2).

[0085]

[0086] where q=(q1,…,q m ) are m interaction proposals input to the self-attention model, the corresponding key values k=(k1,…,k n ), d key represents the dimension of the key vector, and W q and W k are learnable embedding matrices.

[0087] Next, the cross-attention model in the second machine learning model is used to perform cross-attention processing on the intermediate features of the first interaction proposals, the image feature map of the sample image, and the position encoding information, so as to obtain the interaction probability value of each interaction proposal.

[0088] During the cross-attention processing, the attention weight values of the cross-attention model are associated with the intermediate features of the first interaction proposals, the image feature map of the sample image, and the position encoding information.

[0089] For example, the attention weight values of the cross-attention model are as shown in formula (3).

[0090]

[0091] where are the intermediate features of the first interaction proposals, and x=(x1,…,x n) is the image feature map of the sample image, pos j is the corresponding position encoding information, and are learnable embedding matrices.

[0092] In step 106, a second loss function is determined according to the second interaction probability value and the interaction annotation result of each interaction proposal.

[0093] In some embodiments, the second loss function is the FL loss function, as shown in formula (4).

[0094]

[0095] Where C is the total number of interaction categories, y ic ∈ {0, 1} indicates that the i-th interaction proposal has the c-th type of interaction, is the probability that the c-th type of interaction exists predicted.

[0096] In step 107, a first target loss function is determined according to the first loss function and the second loss function.

[0097] In some embodiments, the first target loss function is the weighted sum of the first loss function and the second loss function.

[0098] For example, the first target loss function is as shown in formula (5).

[0099] L = L1 + L2 (5)

[0100] In step 108, the first machine learning model and the second machine learning model are trained using the first target loss function.

[0101] In the machine learning model training method provided in the above embodiments of the present disclosure, by selecting the pairs of person objects that may have interactions from all the pairs of person objects included in the sample image, that is, the interaction proposals, and then giving the interaction proposals for interaction category prediction, the learning difficulty is effectively reduced, and it helps to obtain accurate interaction relationship detection results.

[0102] Figure 3 is a schematic diagram of the training network architecture of an embodiment of the present disclosure. As Figure 3 shown, the sample image is input into the detection model to detect all the persons and all the items in the sample image. At the same time, the image feature map and the position encoding information of the sample image are provided to the cross-attention model in the second machine learning model.

[0103] Next, all person-object pairs are generated using all the persons and all the objects, and the construction feature vectors of each person-object pair are obtained using a characterization construction model. An interaction prediction model (i.e., the first machine learning model) processes the construction feature vectors of each person-object pair to obtain the interaction probability value of each person-object pair, and determines the first loss function based on the interaction probability value and the interaction annotation result of each person-object pair.

[0104] All the person-object pairs are sorted in descending order of the interaction probability value, and the first predetermined number of person-object pairs are used as interaction proposals. The self-attention model in the second machine learning model performs self-attention processing on the predetermined number of interaction proposals to obtain intermediate interaction proposal features. The cross-attention model in the second machine learning model performs cross-attention processing on the intermediate interaction proposal features, the image feature map of the sample image, and the position encoding information to obtain the interaction probability value of each interaction proposal.

[0105] Next, a second loss function is determined based on the interaction probability value and the interaction annotation result of each interaction proposal. A first target loss function is determined based on the first loss function and the second loss function, and the first machine learning model and the second machine learning model are trained using the first target loss function.

[0106] Figure 4 It is a schematic flowchart of a machine learning model training method according to another embodiment of the present disclosure. In some embodiments, the following machine learning model training method is executed by a machine learning model training device.

[0107] In step 401, all person-object pairs in the sample image are obtained.

[0108] In some embodiments, by processing the sample image to obtain the image feature map and the position encoding information of the sample image, all the persons and all the objects in the sample image are detected using the image feature map and the position encoding information.

[0109] Next, all person-object pairs are generated using all the persons and all the objects.

[0110] For example, any two persons among all the persons are combined, and any one person among all the persons is combined with any one object among all the objects to generate all person-object pairs.

[0111] In step 402, the first machine learning model processes the construction feature vectors of each person-object pair to obtain the first interaction probability value of each person-object pair.

[0112] In some embodiments, the construction feature vector of each person-object pair includes the appearance feature, the spatial feature, and the semantic feature of each person-object pair.

[0113] In step 403, a first loss function is determined according to the first interaction probability value and the interaction annotation result of each pair of character objects.

[0114] In some embodiments, the first loss function is as shown in the above formula (1).

[0115] In step 404, a predetermined number of interaction proposals are selected from all pairs of character objects, where the predetermined number of interaction proposals are the top predetermined number of pairs of character objects sorted from largest to smallest in terms of the first interaction probability value among all pairs of character objects.

[0116] In step 405, the semantic dependency relationship between any two interaction proposals among the predetermined number of interaction proposals is determined.

[0117] In some embodiments, each interaction proposal is used as a graph node, and any two nodes are connected by an edge to construct a complete graph centered on interactions. In this graph, the semantic dependency relationship of interaction proposal HOI(I1) relative to HOI(I2) is expressed as <HOI(I2)→HOI(I1)>, and this semantic dependency relationship is determined by the characters and objects of HOI(I1) and HOI(I2).

[0118] For example, if HOI(I1) and HOI(I2) do not share characters and objects, this semantic dependency relationship is called a disjunctive dependency relationship, as Figure 5A shown.

[0119] If HOI(I1) and HOI(I2) share characters, this semantic dependency relationship is called a same - human dependency relationship, as Figure 5B shown.

[0120] If HOI(I1) and HOI(I2) share objects, this semantic dependency relationship is called a same - object dependency relationship, as Figure 5C shown.

[0121] If the character of HOI(I1) is the object of HOI(I2), this semantic dependency relationship is called a series - opposing dependency relationship, as Figure 5D shown.

[0122] If the object of HOI(I1) is the character of HOI(I2), this semantic dependency relationship is called a series dependency relationship, as Figure 5E shown.

[0123] If the characters and objects of HOI(I1) and HOI(I2) are the same, then this semantic dependency relationship is called the dependency relationship of the same pair, such as Figure 5F as shown.

[0124] Return Figure 4 . In step 406, the second machine learning model is used to perform attention processing on a predetermined number of interaction proposals, semantic dependency relationships, the image feature map of the sample image, and position encoding information to obtain the third interaction probability value of each interaction proposal.

[0125] In some embodiments, the self-attention model in the second machine learning model is used to perform self-attention processing on a predetermined number of interaction proposals and semantic dependency relationships to obtain intermediate features of the second interaction proposal.

[0126] During the self-attention processing, the attention weight value of the self-attention model is associated with a predetermined number of interaction proposals and semantic dependency relationships.

[0127] For example, the attention weight value of the self-attention model is as shown in formula (6).

[0128]

[0129] where q=(q1,…,q m ) are m interaction proposals input to the self-attention model, d ij is the corresponding semantic dependency relationship, d key represents the dimension of the key vector, W q and W k are learnable embedding matrices, E dep is the embedding matrix of the semantic dependency category, and ψ is an encoding function used to encode the semantic dependency between interactions, implemented by a two-layer perceptron, for example.

[0130] Next, the cross-attention model in the second machine learning model is used to perform cross-attention processing on the intermediate features of the second interaction proposal, the image feature map of the sample image, and position encoding information to obtain the third interaction probability value of each interaction proposal.

[0131] During the cross-attention processing, the attention weight value of the cross-attention model is associated with the intermediate features of the second interaction proposal, the image feature map of the sample image, and position encoding information.

[0132] For example, the attention weight value of the cross-attention model is as shown in the above formula (3).

[0133] In step 407, the third loss function is determined according to the third interaction probability value of each interaction proposal and the interaction annotation result.

[0134] In some embodiments, the third loss function L3 is the FL loss function, and the expression form of the above formula (4) can be adopted.

[0135] In step 408, the second target loss function is determined according to the first loss function and the third loss function.

[0136] In some embodiments, the second target loss function is the weighted sum of the first loss function and the third loss function.

[0137] For example, the second target loss function is as shown in formula (7).

[0138] L = L1 + L3 (7)

[0139] In step 409, the first machine learning model and the second machine learning model are trained using the second target loss function.

[0140] Figure 6 It is a schematic diagram of the training network architecture of an embodiment of the present disclosure. As Figure 6 shown, the sample image is input into the detection model to detect all the people and all the objects in the sample image. At the same time, the image feature map and the position encoding information of the sample image are provided to the cross-attention model in the second machine learning model.

[0141] Next, all person-object pairs are generated using all the people and all the objects, and the construction feature vectors of each person-object pair are obtained using the characterization construction model. The interaction prediction model (i.e., the first machine learning model) processes the construction feature vectors of each person-object pair to obtain the interaction probability value of each person-object pair, and determines the first loss function according to the interaction probability value of each person-object pair and the interaction annotation result.

[0142] All the person-object pairs are arranged in descending order of the interaction probability value, and the first predetermined number of person-object pairs are used as interaction proposals. The semantic dependency relationship between any two of the first predetermined number of interaction proposals is determined using the inter-interaction semantic structure model. The self-attention processing is performed on the first predetermined number of interaction proposals and the corresponding semantic dependency relationship using the self-attention model in the second machine learning model to obtain the intermediate features of the interaction proposals. The cross-attention processing is performed on the intermediate features of the interaction proposals, the image feature map of the sample image, and the position encoding information using the cross-attention model in the second machine learning model to obtain the interaction probability value of each interaction proposal.

[0143] Next, determine the third loss function according to the interaction probability value and interaction annotation result of each interaction proposal. Determine the second objective loss function according to the first loss function and the third loss function, and use the second objective loss function to train the first machine learning model and the second machine learning model. By utilizing semantic dependency relationships, the prediction performance of interaction relationships can be improved.

[0144] Figure 7 Schematic diagram of the process of the machine learning model training method according to another embodiment of the present disclosure. In some embodiments, the following machine learning model training method is executed by a machine learning model training device.

[0145] In step 701, obtain all pairs of person objects in the sample image.

[0146] In some embodiments, by processing the sample image to obtain the image feature map and position encoding information of the sample image, detect all persons and all items in the sample image using the image feature map and position encoding information.

[0147] Next, generate all pairs of person objects using all persons and all items.

[0148] For example, combine any two persons among all persons, and combine any one person among all persons with any one item among all items to generate all pairs of person objects.

[0149] In step 702, process the construction feature vector of each pair of person objects using the first machine learning model to obtain the first interaction probability value of each pair of person objects.

[0150] In some embodiments, the construction feature vector of each pair of person objects includes the appearance feature, spatial feature, and semantic feature of each pair of person objects.

[0151] In step 703, determine the first loss function according to the first interaction probability value and interaction annotation result of each pair of person objects.

[0152] In some embodiments, the first loss function is as shown in the above formula (1).

[0153] In step 704, select a predetermined number of interaction proposals from all pairs of person objects, where the predetermined number of interaction proposals are the top predetermined number of pairs of person objects sorted from large to small according to the first interaction probability value among all pairs of person objects.

[0154] In step 705, determine the semantic dependency relationship between any two interaction proposals among the predetermined number of interaction proposals.

[0155] In some embodiments, semantic dependency relationships include separated dependency relationships, dependency relationships with the same person, dependency relationships with the same object, dependency relationships with reverse concatenation, dependency relationships with concatenation, and dependency relationships of the same pair.

[0156] In step 706, the spatial structure information of each interaction proposal is determined.

[0157] In some embodiments, the spatial structure information includes the identification of the background area (background) of the interaction proposal, the identification of the union area (union) of the areas where the person and the object are located in the interaction proposal, the identification of the area where the person is located (human) in the interaction proposal, the identification of the area where the object is located (object) in the interaction proposal, and the identification of the intersection area (intersection) of the areas where the person and the object are located in the interaction proposal.

[0158] As Figure 8 described, area 80 is the background area, area 81 is the area where the person is located, area 82 is the area where the object is located, area 83 is the intersection area of area 81 and area 82, and area 84 is the union area of area 81 and area 82.

[0159] Return Figure 7 . In step 707, the second machine learning model performs attention processing on a predetermined number of interaction proposals, semantic dependency relationships, spatial structure information, the image feature map of the sample image, and position encoding information to obtain the fourth interaction probability value of each interaction proposal.

[0160] In some embodiments, the self-attention model in the second machine learning model performs self-attention processing on a predetermined number of interaction proposals and semantic dependency relationships to obtain the intermediate features of the third interaction proposal.

[0161] During the self-attention processing, the attention weight values of the self-attention model are associated with a predetermined number of interaction proposals and semantic dependency relationships.

[0162] For example, the attention weight values of the self-attention model are as shown in the above formula (6).

[0163] Next, the cross-attention model in the second machine learning model performs cross-attention processing on the intermediate features of the third interaction proposal, spatial structure information, the image feature map of the sample image, and position encoding information to obtain the third interaction probability value of each interaction proposal.

[0164] During the cross-attention processing, the attention weight values of the cross-attention model are associated with the intermediate features of the second interaction proposal, spatial structure information, the image feature map of the sample image, and position encoding information.

[0165] For example, the attention weight value of the cross-attention model is as shown in formula (8).

[0166]

[0167] Among them, is the intermediate feature of the second interaction proposal, x = (x1,..., x n ) is the image feature map of the sample image, pos j is the corresponding position encoding information, and are learnable embedding matrices, l ij is the corresponding spatial structure information, E lay is the embedding matrix of the spatial layout category, and ψ is an encoding function for encoding the spatial structure within the interaction, implemented by, for example, a two-layer perceptron.

[0168] In step 708, determine the fourth loss function according to the fourth interaction probability value and the interaction annotation result of each interaction proposal.

[0169] In some embodiments, the third loss function L4 is the FL loss function, and can adopt the expression form of the above formula (4).

[0170] In step 709, determine the third target loss function according to the first loss function and the fourth loss function.

[0171] In some embodiments, the third target loss function is the weighted sum of the first loss function and the fourth loss function.

[0172] For example, the second target loss function is as shown in formula (9).

[0173] L = L1 + L4 (9)

[0174] In step 710, train the first machine learning model and the second machine learning model using the third target loss function.

[0175] Figure 9 This is a schematic diagram of the training network architecture of another embodiment of the present disclosure. As Figure 9 shown, input the sample image into the detection model to detect all the people and all the objects in the sample image. At the same time, provide the image feature map and the position encoding information of the sample image to the cross-attention model in the second machine learning model.

[0176] Next, all pairs of character objects are generated using all characters and all items, and the construction feature vectors of each pair of character objects are obtained using the representation construction model. The interaction prediction model (i.e., the first machine learning model) processes the construction feature vectors of each pair of character objects to obtain the interaction probability values of each pair of character objects, and determines the first loss function based on the interaction probability values and interaction annotation results of each pair of character objects.

[0177] All pairs of character objects are sorted in descending order of interaction probability values, and the first predetermined number of pairs of character objects are used as interaction proposals. The semantic dependency relationship between any two of the predetermined number of interaction proposals is determined using the inter-interaction semantic structure model, and the spatial structure information of each interaction proposal is determined using the intra-interaction spatial structure model. The self-attention model in the second machine learning model performs self-attention processing on the predetermined number of interaction proposals and the corresponding semantic dependency relationships to obtain intermediate interaction proposal features. The cross-attention model in the second machine learning model performs cross-attention processing on the intermediate interaction proposal features, the corresponding spatial structure information, the image feature map of the sample image, and the position encoding information to obtain the interaction probability values of each interaction proposal.

[0178] Next, a fourth loss function is determined based on the interaction probability values and interaction annotation results of each interaction proposal. A third target loss function is determined based on the first loss function and the fourth loss function, and the first machine learning model and the second machine learning model are trained using the third target loss function. By utilizing the spatial structure information, the prediction performance of the interaction relationship can be further improved.

[0179] Figure 10 The structural schematic diagram of a machine learning model training device according to an embodiment of the present disclosure. As Figure 10 shown, the machine learning model training device includes a first training module 1001, a second training module 1002, a third training module 1003, and a fourth training module 1004.

[0180] The first training module 1001 is configured to obtain all pairs of character objects in the sample image.

[0181] In some embodiments, the first training module 1001 processes the sample image to obtain the image feature map and position encoding information of the sample image, and detects all characters and all items in the sample image using the image feature map and position encoding information.

[0182] The second training module 1002 is configured to process the constructed feature vectors of each pair of person objects by using a first machine learning model to obtain a first interaction probability value for each pair of person objects, determine a first loss function according to the first interaction probability value of each pair of person objects and the interaction annotation result, and select a predetermined number of interaction proposals from all pairs of person objects, where the predetermined number of interaction proposals are the first predetermined number of pairs of person objects sorted from largest to smallest in terms of the first interaction probability value among all pairs of person objects.

[0183] In some embodiments, the constructed feature vectors of each pair of person objects include the appearance features, spatial features, and semantic features of each pair of person objects.

[0184] In some embodiments, the second training module 1002 detects the first intersection over union ratio of the person bounding box and the corresponding person annotation bounding box in the i-th pair of person objects, where 1 ≤ i ≤ N and N is the total number of pairs of person objects. The second intersection over union ratio of the object bounding box and the corresponding object annotation bounding box in the i-th pair of person objects is detected. If both the first intersection over union ratio and the second intersection over union ratio are greater than a preset threshold (for example, 0.5), the i-th pair of person objects is taken as a positive sample. If at least one of the first intersection over union ratio or the second intersection over union ratio is not greater than the preset threshold, the i-th pair of person objects is taken as a negative sample. The first machine learning model is used to process the positive samples and negative samples to obtain a first interaction probability value for each pair of person objects.

[0185] In some embodiments, in order to improve the processing effect, the hard mining strategy is used to sample the negative samples to obtain a negative sample sampling result. Furthermore, the first machine learning model is used to process the positive samples and the negative sample sampling result to obtain a first interaction probability value for each pair of person objects.

[0186] In some embodiments, the first loss function is as shown in the above formula (1).

[0187] The third training module 1003 is configured to perform attention processing on a predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information by using a second machine learning model to obtain a second interaction probability value for each interaction proposal, and determine a second loss function according to the second interaction probability value of each interaction proposal and the interaction annotation result.

[0188] In some embodiments, the third training module 1003 performs self-attention processing on a predetermined number of interaction proposals by using the self-attention model in the second machine learning model to obtain an intermediate feature of the first interaction proposal.

[0189] During the self-attention processing, the attention weight values of the self-attention model are associated with a predetermined number of interaction proposals.

[0190] For example, the attention weight values of the self-attention model are as shown in the above formula (2).

[0191] Next, the third training module 1003 performs cross-attention processing on the intermediate features of the first interaction proposal, the image feature map of the sample image, and the position encoding information using the cross-attention model in the second machine learning model to obtain the interaction probability value of each interaction proposal.

[0192] During the cross-attention processing, the attention weight values of the cross-attention model are associated with the intermediate features of the first interaction proposal, the image feature map of the sample image, and the position encoding information.

[0193] For example, the attention weight values of the cross-attention model are as shown in the above formula (3).

[0194] In some embodiments, the second loss function is as shown in the above formula (4).

[0195] The fourth training module 1004 is configured to determine a first target loss function according to the first loss function and the second loss function, and use the first target loss function to train the first machine learning model and the second machine learning model.

[0196] In some embodiments, the first target loss function is a weighted sum of the first loss function and the second loss function.

[0197] For example, the first target loss function is as shown in the above formula (5).

[0198] In some embodiments, the third training module 1003 determines the semantic dependency relationship between any two interaction proposals among a predetermined number of interaction proposals.

[0199] The fourth training module 1004 performs attention processing on a predetermined number of interaction proposals, the semantic dependency relationship, the image feature map of the sample image, and the position encoding information using the second machine learning model to obtain the third interaction probability value of each interaction proposal.

[0200] In some embodiments, the fourth training module 1004 performs self-attention processing on a predetermined number of interaction proposals and the semantic dependency relationship using the self-attention model in the second machine learning model to obtain the intermediate features of the second interaction proposal.

[0201] During the self-attention processing, the attention weight values of the self-attention model are associated with a predetermined number of interaction proposals and the semantic dependency relationship.

[0202] For example, the attention weight values of the self-attention model are as shown in the above formula (6).

[0203] Next, the fourth training module 1004 performs cross-attention processing on the intermediate features of the second interaction proposal, the image feature map of the sample image, and the position encoding information using the cross-attention model in the second machine learning model to obtain the third interaction probability value for each interaction proposal.

[0204] During the cross-attention processing, the attention weight values of the cross-attention model are associated with the intermediate features of the second interaction proposal, the image feature map of the sample image, and the position encoding information.

[0205] For example, the attention weight values of the cross-attention model are as shown in the above formula (3).

[0206] Next, the fourth training module 1004 determines the third loss function according to the third interaction probability value and the interaction annotation result of each interaction proposal. The second target loss function is determined according to the first loss function and the third loss function. The first machine learning model and the second machine learning model are trained using the second target loss function.

[0207] In some embodiments, the third loss function adopts the expression form of the above formula (4).

[0208] In some embodiments, the second target loss function is the weighted sum of the first loss function and the third loss function.

[0209] For example, the second target loss function is as shown in the above formula (7).

[0210] In some embodiments, the third training module 1003 also determines the spatial structure information of each interaction proposal. For example, the spatial structure information includes the background region identifier of the interaction proposal, the union region identifier of the region where the person is located and the region where the object is located in the interaction proposal, the region identifier of the region where the person is located in the interaction proposal, the region identifier of the region where the interaction proposal is located, and the intersection region identifier of the region where the person is located and the region where the object is located in the interaction proposal.

[0211] The fourth training module 1004 performs attention processing on a predetermined number of interaction proposals, semantic dependency relationships, spatial structure information, the image feature map of the sample image, and the position encoding information using the second machine learning model to obtain the fourth interaction probability value for each interaction proposal.

[0212] In some embodiments, the fourth training module 1004 performs self-attention processing on a predetermined number of interaction proposals and semantic dependency relationships using the self-attention model in the second machine learning model to obtain the intermediate features of the third interaction proposal.

[0213] During the self-attention processing, the attention weight values of the self-attention model are associated with a predetermined number of interaction proposals and semantic dependencies.

[0214] For example, the attention weight values of the self-attention model are as shown in the above formula (6).

[0215] Next, the fourth training module 1004 performs cross-attention processing on the intermediate features of the third interaction proposal, the spatial structure information, the image feature map of the sample image, and the position encoding information using the cross-attention model in the second machine learning model to obtain the third interaction probability value of each interaction proposal.

[0216] During the cross-attention processing, the attention weight values of the cross-attention model are associated with the intermediate features of the second interaction proposal, the spatial structure information, the image feature map of the sample image, and the position encoding information.

[0217] For example, the attention weight values of the cross-attention model are as shown in the above formula (8).

[0218] Next, the fourth training module 1004 determines the fourth loss function according to the fourth interaction probability value and the interaction annotation result of each interaction proposal. The third target loss function is determined according to the first loss function and the fourth loss function, and the first machine learning model and the second machine learning model are trained using the third target loss function

[0219] In some embodiments, the third loss function L4 adopts the expression form of the above formula (4).

[0220] In some embodiments, the third target loss function is a weighted sum of the first loss function and the fourth loss function.

[0221] For example, the second target loss function is as shown in the above formula (9).

[0222] Figure 11 It is a schematic structural diagram of a machine learning model training device according to another embodiment of the present disclosure. As Figure 11 shown, the machine learning model training device includes a memory 111 and a processor 112.

[0223] The memory 111 is used to store instructions, the processor 112 is coupled to the memory 111, and the processor 112 is configured to execute the methods related to any one of the embodiments as Figure 1 , Figure 4 and Figure 7 involved.

[0224] As Figure 11As shown, the machine learning model training device further includes a communication interface 113 for information interaction with other devices. Meanwhile, the machine learning model training device further includes a bus 114, through which the processor 112, the communication interface 113, and the memory 111 complete communication with each other.

[0225] The memory 111 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory. The memory 111 may also be a memory array. The memory 111 may also be partitioned, and the partitions may be combined into virtual volumes according to certain rules.

[0226] In addition, the processor 112 may be a central processing unit CPU, or may be an application specific integrated circuit ASIC, or may be one or more integrated circuits configured to implement the embodiments of the present disclosure.

[0227] The present disclosure also relates to a computer-readable storage medium, in which computer instructions are stored, and when the instructions are executed by a processor, the methods related to any one of the embodiments in Figure 1 、 Figure 4 and Figure 7 are implemented.

[0228] Figure 12 is a schematic flowchart of an interaction relationship detection method according to an embodiment of the present disclosure. In some embodiments, the following interaction relationship detection method is executed by an interaction information detection device.

[0229] In step 1201, all pairs of human objects in the image to be processed are obtained.

[0230] In step 1202, the constructed feature vectors of each pair of human objects are input into a first machine learning model, so that the first machine learning module outputs a first interaction probability value for each pair of human objects, where the first machine learning model is trained by using the machine learning model training method according to any one of the embodiments in Figure 1 、 Figure 4 or Figure 7 .

[0231] In step 1203, a predetermined number of interaction proposals are selected from all pairs of human objects, where the predetermined number of interaction proposals are the top predetermined number of pairs of human objects sorted in descending order of the first interaction probability value among all pairs of human objects.

[0232] In step 1204, the predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information are input into a second machine learning model, so that the second machine learning model outputs the interaction relationship of each interaction proposal, where the second machine learning model uses Figure 1It is trained by the machine learning model training method of any one of the embodiments.

[0233] Figure 13 FIG. 13 is a schematic flowchart of an interaction relationship detection method according to an embodiment of the present disclosure. In some embodiments, the following interaction relationship detection method is executed by an interaction information detection device.

[0234] In step 1301, all pairs of human object pairs in the image to be processed are obtained.

[0235] In step 1302, the constructed feature vectors of each pair of human object pairs are input into a first machine learning model, so that the first machine learning module outputs a first interaction probability value for each pair of human object pairs, where the first machine learning model uses Figure 1 , Figure 4 or Figure 7 It is trained by the machine learning model training method of any one of the embodiments.

[0236] In step 1303, a predetermined number of interaction proposals are selected from all pairs of human object pairs, where the predetermined number of interaction proposals are the top predetermined number of pairs of human object pairs sorted from largest to smallest in the first interaction probability value among all pairs of human object pairs.

[0237] In step 1304, the semantic dependency relationship between any two interaction proposals among the predetermined number of interaction proposals is determined.

[0238] In step 1305, the predetermined number of interaction proposals, the semantic dependency relationship, the image feature map of the sample image, and the position encoding information are input into a second machine learning model, so that the second machine learning model outputs the interaction relationship of each interaction proposal, where the second machine learning model uses Figure 4 It is trained by the machine learning model training method of any one of the embodiments.

[0239] Figure 14 FIG. 14 is a schematic flowchart of an interaction relationship detection method according to an embodiment of the present disclosure. In some embodiments, the following interaction relationship detection method is executed by an interaction information detection device.

[0240] In step 1401, all pairs of human object pairs in the image to be processed are obtained.

[0241] In step 1402, the constructed feature vectors of each pair of human object pairs are input into a first machine learning model, so that the first machine learning module outputs a first interaction probability value for each pair of human object pairs, where the first machine learning model uses Figure 1 , Figure 4 or Figure 7 It is trained by the machine learning model training method of any one of the embodiments.

[0242] In step 1403, a predetermined number of interaction proposals are selected from all pairs of person objects, where the predetermined number of interaction proposals are the top predetermined number of pairs of person objects sorted in descending order of the first interaction probability value among all pairs of person objects.

[0243] In step 1404, the semantic dependency relationship between any two interaction proposals among the predetermined number of interaction proposals is determined.

[0244] In step 1405, the spatial structure information of each interaction proposal is determined.

[0245] In step 1406, the predetermined number of interaction proposals, the semantic dependency relationship, the spatial structure information, the image feature map of the sample image, and the position encoding information are input into the second machine learning model so that the second machine learning model outputs the interaction relationship of each interaction proposal, where the second machine learning model is trained using Figure 7 the machine learning model training method of any one of the

[0246] Figure 15 is a schematic structural diagram of an interaction relationship detection device according to an embodiment of the present disclosure. As Figure 15 shown, the interaction relationship detection device includes a first processing module 151, a second processing module 152, and a third processing module 153.

[0247] The first processing module 151 is configured to obtain all pairs of person objects in the image to be processed.

[0248] The second processing module 152 is configured to input the constructed feature vector of each pair of person objects into the first machine learning model so that the first machine learning module outputs the first interaction probability value of each pair of person objects, where the first machine learning model is trained using Figure 1 , Figure 4 or Figure 7 the machine learning model training method of any one of the

[0249] The third processing module 153 is configured to input the predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information into the second machine learning model so that the second machine learning model outputs the interaction relationship of each interaction proposal, where the second machine learning model is trained using Figure 1 the machine learning model training method of any one of the

[0250] In some embodiments, the second processing module 152 determines the semantic dependency relationship between any two interaction proposals among the predetermined number of interaction proposals.

[0251] The third processing module 153 inputs a predetermined number of interaction proposals, semantic dependency relationships, the image feature map of the sample image, and position encoding information into the second machine learning model, so that the second machine learning model outputs the interaction relationship of each interaction proposal, where the second machine learning model is trained by using Figure 4 the machine learning model training method of any one of the embodiments in

[0252] In some embodiments, the second processing module 152 determines the spatial structure information of each interaction proposal.

[0253] The third processing module 153 inputs a predetermined number of interaction proposals, semantic dependency relationships, spatial structure information, the image feature map of the sample image, and position encoding information into the second machine learning model, so that the second machine learning model outputs the interaction relationship of each interaction proposal, where the second machine learning model is trained by using Figure 7 the machine learning model training method of any one of the embodiments in

[0254] Figure 16 FIG. is a schematic structural diagram of an interaction relationship detection device according to an embodiment of the present disclosure. As Figure 16 shown, the interaction relationship detection device includes a memory 161, a processor 162, a communication interface 163, and a bus 164. Figure 16 Different from Figure 11 is that in the embodiment shown in Figure 16 the processor 162 is configured to execute the methods involved in any one of the embodiments as described in Figures 12 - 14 based on the instructions stored in the memory.

[0255] The present disclosure also relates to a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the methods involved in any one of the embodiments as described in Figures 12 - 14 are implemented.

[0256] In some embodiments, the functional unit modules described above can be implemented as a general-purpose processor, a programmable logic controller (PLC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any suitable combination thereof for executing the functions described in the present disclosure.

[0257] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, or the like.

[0258] The description of the present disclosure is given for purposes of illustration and description, and is not intended to be exhaustive or to limit the present disclosure to the disclosed form. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are chosen and described in order to best explain the principles of the present disclosure and its practical application, and to enable those of ordinary skill in the art to understand the present disclosure so as to design various embodiments with various modifications suitable for a particular purpose.

Claims

1. A method for training a machine learning model, comprising: Obtain all pairs of human objects in the sample image; Process the constructed feature vectors of each pair of human objects using a first machine learning model to obtain a first interaction probability value for each pair of human objects; Determine a first loss function based on the first interaction probability value of each pair of human objects and the interaction annotation result; Select a predetermined number of interaction proposals from all the pairs of human objects, where the predetermined number of interaction proposals are the top predetermined number of pairs of human objects sorted by the first interaction probability value from largest to smallest among all the pairs of human objects; Perform attention processing on the predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information using a second machine learning model to obtain a second interaction probability value for each interaction proposal; Determine a second loss function based on the second interaction probability value of each interaction proposal and the interaction annotation result; Determine a first target loss function based on the first loss function and the second loss function; Train the first machine learning model and the second machine learning model using the first target loss function.

2. The method according to claim 1, wherein, The performing attention processing on the predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information using the second machine learning model includes: Perform self-attention processing on the predetermined number of interaction proposals using the self-attention model in the second machine learning model to obtain intermediate features of the first interaction proposals; Perform cross-attention processing on the intermediate features of the first interaction proposals, the image feature map of the sample image, and the position encoding information using the cross-attention model in the second machine learning model to obtain the interaction probability value of each interaction proposal.

3. The method according to claim 2, wherein, During the self-attention processing, the attention weight values of the self-attention model are associated with the predetermined number of interaction proposals; During the cross-attention processing, the attention weight values of the cross-attention model are associated with the intermediate features of the first interaction proposals, the image feature map of the sample image, and the position encoding information.

4. The method according to claim 1, wherein, The processing the constructed feature vectors of each pair of human objects using the first machine learning model includes: Detect the first intersection over union ratio of the human bounding box and the corresponding human annotation bounding box in the i-th pair of human objects, where 1 ≤ i ≤ N and N is the total number of pairs of human objects; Detect the second intersection over union ratio of the object bounding box and the corresponding object annotation bounding box in the i-th pair of human objects; If both the first intersection over union ratio and the second intersection over union ratio are greater than a preset threshold, then take the i-th pair of human objects as a positive sample; If at least one of the first intersection over union ratio or the second intersection over union ratio is not greater than the preset threshold, then take the i-th pair of human objects as a negative sample; Process the positive samples and the negative samples using the first machine learning model to obtain a first interaction probability value for each pair of human objects.

5. The method according to claim 4, wherein, The processing the positive samples and the negative samples using the first machine learning model includes: Sample the negative samples using a hard negative mining strategy to obtain a negative sample sampling result; Process the sampling results of the positive samples and the negative samples by using the first machine learning model to obtain the first interaction probability value of each pair of person objects.

6. The method according to claim 4, wherein, The constructed feature vector of each pair of person objects includes the appearance feature, spatial feature, and semantic feature of each pair of person objects.

7. The method according to claim 1, wherein, The first objective loss function is the weighted sum of the first loss function and the second loss function.

8. The method according to claim 1, wherein, The obtaining of all pairs of person objects in the sample image includes: Process the sample image to obtain the image feature map and position encoding information of the sample image; Detect all persons and all items in the sample image by using the image feature map and the position encoding information; Generate all pairs of person objects by using all the persons and all the items.

9. The method according to claim 8, wherein, The generating of all pairs of person objects by using all the persons and all the items includes: Combine any two persons among all the persons, and combine any one person among all the persons with any one item among all the items to generate all pairs of person objects.

10. The method according to any one of claims 1-9 further comprises: Determine the semantic dependency relationship between any two interaction proposals among the predetermined number of interaction proposals; Perform attention processing on the predetermined number of interaction proposals, the semantic dependency relationship, the image feature map and position encoding information of the sample image by using a second machine learning model to obtain the third interaction probability value of each interaction proposal; Determine the third loss function according to the third interaction probability value of each interaction proposal and the interaction annotation result; Determine the second objective loss function according to the first loss function and the third loss function; Train the first machine learning model and the second machine learning model by using the second objective loss function.

11. The method according to claim 10, wherein The performing of attention processing on the predetermined number of interaction proposals, the semantic dependency relationship, the image feature map and position encoding information of the sample image by using a second machine learning model includes: Perform self-attention processing on the predetermined number of interaction proposals and the semantic dependency relationship by using the self-attention model in the second machine learning model to obtain the intermediate feature of the second interaction proposal; Perform cross-attention processing on the intermediate feature of the second interaction proposal, the image feature map of the sample image, and the position encoding information by using the cross-attention model in the second machine learning model to obtain the third interaction probability value of each interaction proposal.

12. The method according to claim 11, wherein During the self-attention processing, the attention weight value of the self-attention model is associated with the predetermined number of interaction proposals and the semantic dependency relationship; During the cross-attention processing, the attention weight value of the cross-attention model is associated with the intermediate feature of the second interaction proposal, the image feature map of the sample image, and the position encoding information.

13. The method according to claim 10, wherein The semantic dependency relationship includes a separated dependency relationship, a dependency relationship with the same person, a dependency relationship with the same object, a serial dependency relationship, a reverse serial dependency relationship, or a same-pair dependency relationship.

14. The method according to claim 10, wherein The second objective loss function is the weighted sum of the first loss function and the third loss function.

15. The method according to claim 11 further comprises: After determining the semantic dependency relationship between any two of the predetermined number of interaction proposals, determine the spatial structure information of each interaction proposal; Use a second machine learning model to perform attention processing on the predetermined number of interaction proposals, the semantic dependency relationship, the spatial structure information, the image feature map of the sample image, and the position encoding information to obtain a fourth interaction probability value for each interaction proposal; Determine a fourth loss function according to the fourth interaction probability value of each interaction proposal and the interaction annotation result; Determine a third target loss function according to the first loss function and the fourth loss function; Use the third target loss function to train the first machine learning model and the second machine learning model.

16. The method according to claim 15, wherein The using the second machine learning model to perform attention processing on the predetermined number of interaction proposals, the semantic dependency relationship, the spatial structure information, the image feature map of the sample image, and the position encoding information includes: Use the self-attention model in the second machine learning model to perform self-attention processing on the predetermined number of interaction proposals and the semantic dependency relationship to obtain intermediate features of the third interaction proposal; Use the cross-attention model in the second machine learning model to perform cross-attention processing on the intermediate features of the third interaction proposal, the spatial structure information, the image feature map of the sample image, and the position encoding information to obtain a third interaction probability value for each interaction proposal.

17. The method according to claim 16, wherein During the self-attention processing, the attention weight value of the self-attention model is associated with the predetermined number of interaction proposals and the semantic dependency relationship; During the cross-attention processing, the attention weight value of the cross-attention model is associated with the intermediate features of the second interaction proposal, the spatial structure information, the image feature map of the sample image, and the position encoding information.

18. The method according to claim 15, wherein The spatial structure information includes the background area identifier of the interaction proposal, the union area identifier of the area where the person and the object are located in the interaction proposal, the area identifier of the area where the person is located in the interaction proposal, the area identifier of the area where the object is located in the interaction proposal, and the intersection area identifier of the area where the person and the object are located in the interaction proposal.

19. The method according to claim 15, wherein The third target loss function is the weighted sum of the first loss function and the fourth loss function.

20. A machine learning model training apparatus, comprising: A first training module, configured to obtain all pairs of human objects in the sample image; A second training module, configured to use a first machine learning model to process the constructed feature vectors of each pair of human objects to obtain a first interaction probability value for each pair of human objects, determine a first loss function according to the first interaction probability value of each pair of human objects and the interaction annotation result, and select a predetermined number of interaction proposals from all pairs of human objects, where the predetermined number of interaction proposals are the first predetermined number of pairs of human objects sorted from largest to smallest in terms of the first interaction probability value among all pairs of human objects; The third training module is configured to perform attention processing on the predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information by using a second machine learning model to obtain a second interaction probability value for each interaction proposal, and determine a second loss function according to the second interaction probability value of each interaction proposal and the interaction annotation result; The fourth training module is configured to determine a first target loss function according to the first loss function and the second loss function, and use the first target loss function to train the first machine learning model and the second machine learning model.

21. A machine learning model training device, comprising: The memory is configured to store instructions; The processor is coupled to the memory and is configured to execute the method according to any one of claims 1-19 based on the instructions stored in the memory.

22. An interaction relationship detection method, comprising: Obtain all pairs of human object in the image to be processed; Input the constructed feature vectors of each pair of human objects into a first machine learning model, so that the first machine learning model outputs a first interaction probability value for each pair of human objects, where the first machine learning model is trained by using the machine learning model training method according to any one of claims 1-19; Select a predetermined number of interaction proposals from all the pairs of human objects, where the predetermined number of interaction proposals are the top predetermined number of pairs of human objects sorted from largest to smallest in terms of the first interaction probability value among all the pairs of human objects; Input the predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information into a second machine learning model, so that the second machine learning model outputs the interaction relationship of each interaction proposal, where the second machine learning model is trained by using the machine learning model training method according to any one of claims 1-19.

23. The method according to claim 22, further comprising: Determine the semantic dependency relationship between any two interaction proposals among the predetermined number of interaction proposals; Input the predetermined number of interaction proposals, the semantic dependency relationship, the image feature map of the sample image, and the position encoding information into the second machine learning model, so that the second machine learning model outputs the interaction relationship of each interaction proposal.

24. The method according to claim 23, further comprising: After determining the semantic dependency relationship between any two interaction proposals among the predetermined number of interaction proposals, determine the spatial structure information of each interaction proposal; Input the predetermined number of interaction proposals, the semantic dependency relationship, the spatial structure information, the image feature map of the sample image, and the position encoding information into the second machine learning model, so that the second machine learning model outputs the interaction relationship of each interaction proposal.

25. An interaction relationship detection device, comprising: The first processing module is configured to obtain all pairs of human object in the image to be processed; A second processing module, configured to input the constructed feature vectors of each pair of character objects into a first machine learning model, so that the first machine learning model outputs a first interaction probability value for each pair of character objects, where the first machine learning model is trained by using the machine learning model training method described in any one of claims 1-19, and select a predetermined number of interaction proposals from all the pairs of character objects, where the predetermined number of interaction proposals are the top predetermined number of pairs of character objects in the all pairs of character objects sorted from large to small according to the first interaction probability value; A third processing module, configured to input the predetermined number of interaction proposals, the image feature map of the sample image, and the position encoding information into a second machine learning model, so that the second machine learning model outputs the interaction relationship of each interaction proposal, where the second machine learning model is trained by using the machine learning model training method described in any one of claims 1-19.

26. An interaction relationship detection device, comprising: A memory, configured to store instructions; A processor, coupled to the memory, the processor is configured to execute the method described in any one of claims 22-24 based on the instructions stored in the memory.

27. A computer-readable storage medium, wherein, A computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the method described in any one of claims 1-19, 22-24 is implemented.

Citation Information

Patent Citations

  • Article recommendation method based on multi-heterogeneous graph neural network

    CN114065048A

  • Character interaction detection equipment, method and device and readable storage medium

    CN114170623A