Person interaction detection method, model training method and device
By using self-attention and cross-attention mechanisms in the decoder to decouple the features of character interaction relationships, the problem of insufficient accuracy in character interaction detection models is solved, achieving more efficient model training and more accurate detection results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2023-12-20
- Publication Date
- 2026-07-21
AI Technical Summary
In existing technologies, the accuracy of human interaction detection models in recognizing human interaction relationships in images is insufficient, resulting in poor detection results.
When using a decoder to perform image feature fusion processing based on an initial query matrix, a query vector is used to extract only one feature from a set of human interaction relationships. The features are decoupled through self-attention and cross-attention mechanisms to improve detection accuracy.
By decoupling the various features of human interaction relationships, the accuracy of human interaction detection in images and the efficiency of model training are improved.
Smart Images

Figure CN117743617B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, specifically to computer vision, deep learning, large model and other technical fields, and can be applied to smart city and other scenarios; in particular, it relates to a human interaction detection method, model training method and device. Background Technology
[0002] Human interaction detection is the process of locating people and objects in an image and detecting the interaction between them.
[0003] Accurately identifying the interaction relationships between people in an image is a problem that urgently needs to be solved. Summary of the Invention
[0004] This disclosure provides a method, model training method, and apparatus for detecting human interaction, so as to accurately identify the human interaction relationships in an image.
[0005] According to a first aspect of this disclosure, a method for detecting character interaction is provided, wherein the method includes:
[0006] Extract image features from the image to be detected;
[0007] Obtain an initial query matrix; wherein the initial query matrix includes multiple query sets; the query set is a parameter set used to extract character interaction relationships, the query set includes at least one query vector, the query vector is used to extract features in the character interaction relationships corresponding to the query vector; the character interaction relationship is the interaction relationship between a person and an object;
[0008] Based on the decoder, feature fusion processing is performed on the image features and the initial query matrix to determine the detection result corresponding to the image to be detected; wherein, the detection result represents the interaction relationship between people in the image to be detected.
[0009] According to a second aspect of this disclosure, a method for extracting image features from an image to be trained is provided; wherein the image to be trained has a first human interaction relationship;
[0010] Obtain a query matrix to be trained; wherein, the query matrix to be trained includes multiple query sets; the query set is a parameter set used to extract human interaction relationships, the query set includes at least one query vector, the query vector is used to extract features in human interaction relationships corresponding to the query vector; the human interaction relationship is the interaction relationship between a person and an object;
[0011] Based on the initial decoder, feature fusion processing is performed on the image features of the image to be trained and the query matrix to be trained to obtain the second person interaction relationship corresponding to the image to be trained;
[0012] Based on the interaction relationship between the first character and the second character, the query matrix to be trained and the initial decoder are modified to obtain the trained decoder and the initial query matrix.
[0013] According to a third aspect of this disclosure, a human interaction detection device is provided, wherein the device comprises:
[0014] The first extraction unit is used to extract image features from the image to be detected;
[0015] The first acquisition unit is used to acquire an initial query matrix; wherein, the initial query matrix includes multiple query sets; the query set is a parameter set for extracting character interaction relationships, the query set includes at least one query vector, the query vector is used to extract features in the character interaction relationships corresponding to the query vector; the character interaction relationship is the interaction relationship between a person and an object;
[0016] The first processing unit is used to perform feature fusion processing on the image features and the initial query matrix based on the decoder to determine the detection result corresponding to the image to be detected; wherein the detection result represents the interaction relationship between people in the image to be detected.
[0017] According to a fourth aspect of this disclosure, a model training apparatus is provided, wherein the apparatus comprises:
[0018] The second extraction unit is used to extract image features from the image to be trained; wherein the image to be trained has a first human interaction relationship;
[0019] The second acquisition unit is used to acquire a query matrix to be trained; wherein, the query matrix to be trained includes multiple query sets; the query set is a parameter set for extracting human interaction relationships, the query set includes at least one query vector, the query vector is used to extract features in the human interaction relationships corresponding to the query vector; the human interaction relationship is the interaction relationship between a person and an object;
[0020] The second processing unit is used to perform feature fusion processing on the image features of the image to be trained and the query matrix to be trained based on the initial decoder, so as to obtain the second person interaction relationship corresponding to the image to be trained.
[0021] The correction unit is used to correct the query matrix to be trained and the initial decoder according to the first character interaction relationship and the second character interaction relationship, so as to obtain the trained decoder and the initial query matrix.
[0022] According to a fifth aspect of this disclosure, an electronic device is provided, comprising:
[0023] At least one processor; and
[0024] A memory communicatively connected to the at least one processor; wherein,
[0025] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to perform the method described in the first aspect, or enable the at least one processor to perform the method described in the second aspect.
[0026] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method of the first aspect, or the computer instructions are configured to cause the computer to perform the method of the second aspect.
[0027] According to a seventh aspect of this disclosure, a computer program product is provided, the computer program product comprising: a computer program stored in a readable storage medium, wherein at least one processor of an electronic device can read the computer program from the readable storage medium, the at least one processor executing the computer program causes the electronic device to perform the method of the first aspect, or the at least one processor executing the computer program causes the electronic device to perform the method of the second aspect.
[0028] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0029] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0030] Figure 1 A flowchart illustrating a human interaction detection method provided in an embodiment of this disclosure;
[0031] Figure 2 A flowchart illustrating the second human interaction detection method provided in this embodiment of the present disclosure;
[0032] Figure 3 A schematic diagram of a model structure provided in an embodiment of this disclosure;
[0033] Figure 4 A schematic flowchart of a model training method provided in an embodiment of this disclosure;
[0034] Figure 5 This is a schematic diagram of the structure of a human interaction detection device provided in an embodiment of the present disclosure;
[0035] Figure 6 A schematic diagram of the structure of another human interaction detection device provided in this embodiment of the present disclosure;
[0036] Figure 7 This is a schematic diagram of the structure of a model training device provided in an embodiment of the present disclosure;
[0037] Figure 8 This is a schematic diagram of the structure of the second model training device provided in the embodiments of this disclosure;
[0038] Figure 9 A schematic diagram of an electronic device provided in this disclosure;
[0039] Figure 10 This is a block diagram of an electronic device used to implement the human interaction detection method or model training method of the embodiments of this disclosure. Detailed Implementation
[0040] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0041] In related technologies, the detection of interactions between people and objects can be performed using a DER detector for image detection. Specifically, firstly, features are extracted from the image to be detected, obtaining the corresponding global features. When the decoder in the DER model processes these global features, it fuses them based on multiple received query vectors to determine the interaction relationships between people in the image. It should be noted that in this fusion process, one query vector is used to simultaneously predict all features corresponding to a set of interactions between people and objects. Furthermore, determining all features corresponding to a set of interactions using only one set of query vectors can increase the difficulty of model training, leading to poorer prediction results.
[0042] To avoid at least one of the aforementioned technical problems, the inventors of this disclosure have creatively arrived at the inventive concept of this disclosure: when the decoder performs image feature fusion processing based on the initial query matrix, a query vector is only used to extract one feature from a set of character interaction relationships, that is, to decouple the various features corresponding to the character interaction relationships, and each query vector only needs to focus on the one feature corresponding to itself, so as to make the image detection results more accurate.
[0043] This disclosure provides a method for detecting human interaction, a method for training a model, and an apparatus for application in the fields of computer vision, deep learning, and large models within the field of artificial intelligence.
[0044] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0045] Figure 1 This is a flowchart illustrating a human interaction detection method provided in an embodiment of the present disclosure, wherein the method includes:
[0046] S101. Extract image features from the image to be detected.
[0047] For example, the execution subject in this embodiment can be a human interaction detection method (hereinafter referred to as a detection device). The detection device can be a server (such as a local server or a cloud server), a computer, a processor, a chip, etc. This embodiment does not limit the scope.
[0048] In this embodiment, the human interaction relationship is specifically used to characterize the interaction relationship between people and objects. When performing human interaction detection on the image to be detected, the image features corresponding to the image to be detected can be extracted first.
[0049] It should be noted that this embodiment does not impose specific restrictions on the image feature extraction method. The feature extraction operators provided in related technologies or the model structures provided in related technologies can be used to obtain image features that describe the entire image to be detected.
[0050] S102. Obtain the initial query matrix; wherein, the initial query matrix includes multiple query sets; the query set is a parameter set used to extract the interaction relationship between people, and the query set includes at least one query vector, which is used to extract the feature corresponding to the query vector in the interaction relationship between people; the interaction relationship between people is the interaction relationship between people and objects.
[0051] For example, the initial query matrix in this embodiment can be understood as a set of parameters (or queries) used to extract the interaction relationships between people in an image. Furthermore, the initial query matrix includes multiple query sets. In practical applications, one query set is used to extract the interaction relationships between a group of people and objects in an image. Moreover, since the interaction relationships between people have multiple features in practical applications, the query set can also contain multiple query vectors, each used to extract features from the interaction relationships between people corresponding to that query vector.
[0052] It should be noted that the initial query matrix in this embodiment is obtained during the training of the initial decoder based on the image to be trained and the human interaction relationship of the image to be trained, so that the query vector obtained by training can be used to extract a feature of the human interaction relationship.
[0053] S103. Based on the decoder, perform feature fusion processing on the image features and the initial query matrix to determine the detection result corresponding to the image to be detected; wherein, the detection result represents the interaction relationship between people in the image to be detected.
[0054] For example, after obtaining the initial query matrix, the initial query matrix and the image features corresponding to the image to be detected can be input into the decoder, and the decoder performs feature fusion processing on the image features based on the initial query matrix to obtain the human interaction relationship corresponding to the image to be detected.
[0055] In one example, when the decoder performs feature fusion processing on the initial query matrix and image features, it can perform feature fusion based on the correlation between features in the image features, based on the initial query matrix, to obtain the updated query matrix. Then, each query vector in the initial query matrix can be updated to the feature corresponding to each parameter based on the correlation fusion.
[0056] It is understood that in this embodiment, by setting the initial query matrix as described above, the various features contained in the character interaction relationship can be decoupled. That is, each query vector in the initial query matrix is used to extract a feature in the character interaction relationship, so as to improve the accuracy of the model detection results.
[0057] In one example, the query set includes at least a first query vector, a second query vector, and a third query vector; wherein, the first query vector is used to extract a first feature in the interaction relationship between people, and the first feature is a feature used to indicate the location information of people; the second query vector is used to extract a second feature in the interaction relationship between people, and the second feature is a feature used to indicate the location information and category information of objects; the third query vector is used to extract a third feature in the interaction relationship between people, and the third feature is a feature used to indicate the interaction action between people and objects.
[0058] For example, in this embodiment, the features in the human interaction relationship can be specifically divided into a first feature, a second feature, and a third feature. The first feature describes the position information of the person in the image during the human interaction relationship. The second feature describes the position of an object in the image during the human interaction relationship, as well as the category of the object. The third feature can be understood as the feature representing the interactive actions performed by the person on the object during the human interaction relationship. When the features of the human interaction relationship are refined into the above three features, furthermore, three query vectors (i.e., the first query vector, the second query vector, and the third query vector) can be set in the query set included in the initial query matrix. This allows for the extraction of the human interaction relationships corresponding to each group of people from the image to be detected through the setting of the initial query matrix, and the acquisition of the position of the pairs of people in the image to be detected, as well as the object category and action category interacted by the person.
[0059] Figure 2 This is a flowchart illustrating a second human interaction detection method provided in this disclosure embodiment. The method includes the following steps:
[0060] S201. Based on the convolutional neural network layer, feature extraction processing is performed on the image to be detected to obtain the feature map information of the image to be detected; the feature map information is used to characterize the local features of the image.
[0061] For example, in this embodiment, when extracting the image features corresponding to the image to be detected, the image to be detected can first be subjected to convolutional sampling processing through a convolutional neural network layer in order to obtain the local features corresponding to the image to be detected, namely the feature map information mentioned above.
[0062] S202. Based on the encoder, feature extraction is performed on the feature map information to obtain the image features of the image to be detected.
[0063] For example, in this embodiment, after obtaining the feature map information corresponding to the image to be detected, the feature map information can be input into a pre-trained encoder so that the obtained feature map information can be processed by the encoder to obtain the image features corresponding to the image to be detected.
[0064] In one example, the encoder described above can process the feature map information based on a multi-head self-attention mechanism to obtain the global image features corresponding to the image to be detected.
[0065] It should be noted that the specific structure of the encoder in this embodiment can refer to the specific structure of the encoder corresponding to the DETR model in related technologies. For example, the encoder may include multiple cascaded coding layers, and each coding layer includes a multi-head self-attention layer, a residual and normalization layer, and a feedforward neural network layer.
[0066] It is understood that in this embodiment, image features can be extracted from the image to be detected using convolutional neural network layers and an encoder, so as to obtain feature information representing the global features of the image to be detected. This information can then be combined with the global features of the image to be detected for feature fusion processing during the subsequent decoding process of the decoder, thereby improving the accuracy of the detection results.
[0067] S203. Obtain the initial query matrix; wherein, the initial query matrix includes multiple query sets; the query set is a set of parameters used to extract the interaction relationship between people, and the query set includes at least one query vector, which is used to extract the features corresponding to the query vector in the interaction relationship between people; the interaction relationship between people is the interaction relationship between people and objects.
[0068] For example, the technical principle of step S203 can be found in step S102, and will not be repeated here.
[0069] S204. Based on the first self-attention layer in the decoder, determine the intra-group relevance information of the query set, and update the vectors in the query set according to the intra-group relevance information to obtain the first set; the intra-group relevance information represents the relevance between the query vectors contained in the query set; wherein, the query set includes multiple query vectors.
[0070] For example, in this embodiment, when performing feature fusion processing on the initial query matrix and image features based on the decoder, the first self-attention layer contained in the decoder can be used to perform self-attention processing on each query set in the initial query matrix, so as to update the query vector contained in the query set according to the self-attention mechanism described above.
[0071] Specifically, by performing self-attention processing on each query set, the relevance between the query vectors contained in the query set can be determined, i.e., the intra-group relevance information mentioned above.
[0072] In one example, when performing self-attention processing on a query set, for each query vector in the query set, a relevance score can be determined between the query vector and the query set it belongs to, and the query vector is updated based on the obtained relevance score. Then, by updating each query set in the above manner, a first set corresponding to each query set is obtained.
[0073] In one example, when the query set includes a first query vector, a second query vector, and a third query vector, after self-attention processing, the resulting first set includes three updated query vectors, and the three updated query vectors correspond one-to-one with each query vector in the query set.
[0074] It is understandable that since a query set is used to extract the interaction relationships between a group of people and objects, by performing self-attention processing on multiple query vectors contained in the query set, the correlation between multiple query vectors in the query set can be established. In the subsequent feature fusion process of image features, the correlation between various features in a group of human interaction relationships can be fully combined to perform feature fusion processing, thereby improving the accuracy of subsequent human interaction detection results.
[0075] In one example, after step S204, the following steps may also be included: normalizing each first set based on the normalization layer in the decoder to obtain the processed first set.
[0076] For example, in this embodiment, a normalization layer may also be provided in the decoder. Furthermore, the normalization layer can be used to normalize the results output by the first self-attention layer in the decoder (i.e., the aforementioned first sets) to reduce the difficulty of subsequent data processing.
[0077] S205. Based on the second self-attention layer of the decoder, determine the inter-group correlation information corresponding to the first set; and update the first set according to the inter-group correlation information to obtain the second set; wherein, the inter-group correlation information represents the correlation between the first set and the first query matrix; the first query matrix is composed of each first set.
[0078] For example, in this embodiment, when determining the detection result, the inter-group correlation information corresponding to each first set is first determined based on the multiple first sets obtained in step S204. Specifically, the inter-group correlation information can be used to characterize the correlation between the first sets.
[0079] In one example, when determining the inter-group correlation information corresponding to the first set, the inter-group correlation information can be determined by performing a matrix dot product operation between the first set and the first query matrix composed of each first set. Then, the first set is updated based on the inter-group correlation information corresponding to each first set, resulting in a second set corresponding to each first set. For example, the inter-group correlation information can be directly used as the updated first set.
[0080] In one example, the step S205, "determining the inter-group correlation information corresponding to the first set based on the second self-attention layer of the decoder," includes the following steps:
[0081] Based on the second self-attention layer, the relevance result of the fourth query vector in the first set is determined; the relevance result represents the relevance between the fourth query vector and the fourth query vector in each of the first sets; the fourth query vector is the result of updating the query vector in the query set for extracting target features based on the intra-group relevance information; the target feature is any one of the features used to indicate the interaction between people and objects, the features used to indicate the location and category information of objects, and the features used to indicate the location information of people; the relevance result corresponding to each fourth query vector in the first set is determined as the inter-group relevance information corresponding to the first set.
[0082] For example, in this embodiment, the first set may include a fourth query vector, which is obtained by updating the query vector used to extract target features in the query set through step S203. That is, the fourth query vector can also be regarded as a query vector used to extract target features. When performing inter-group correlation information, the fourth query vectors included in each first set can be correlated with the fourth query vector in the current first set to obtain the correlation result corresponding to the fourth query vector in the current first set.
[0083] For example, when the query set includes a first query vector, a second query vector, and a third query vector, after performing intra-group correlation processing (i.e., step S203 above), the updated query vector 1 corresponding to the first query vector, the updated query vector 2 corresponding to the second query vector, and the updated query vector 3 corresponding to the third query vector are obtained. Query vector 1, query vector 2, and query vector 3 form the first set corresponding to the query set.
[0084] When determining the relevance results corresponding to the first set, inter-group correlation analysis can be performed between query vector 1 in the current first set and query vector 1 contained in each of the other first sets to obtain the relevance results corresponding to query vector 1 in the current first set. Similarly, for query vector 2 in the current first set, component correlation analysis also needs to be performed between query vector 2 contained in each of the other second sets to obtain the relevance results corresponding to query vector 2 in the current first set. The calculation method for the relevance results of query vector 3 is similar to the above process and will not be repeated here.
[0085] It is understandable that by combining query vectors used to extract the same feature from each first set for inter-group correlation calculation (e.g., combining parameters of features used to extract location information of people from different first sets), the same feature in different pairs of people (i.e., people and objects with interactive relationships) can be combined for feature extraction when the decoder performs feature fusion on the image features. This allows for feature extraction by combining the corresponding information in the entire image, thereby improving the accuracy of feature extraction.
[0086] S206. Based on each second set and image features, determine the detection result corresponding to the image to be detected. The detection result represents the interaction relationships between people in the image to be detected.
[0087] For example, after obtaining the second set after intra-group correlation analysis and inter-group correlation processing (i.e., steps S203 and S204), the image features can be correlated and fused according to the obtained second sets in order to determine the human interaction features in the final image to be detected.
[0088] For example, when determining the detection result based on each second set and image features, matrix similarity can be calculated based on each second set and image features, and each feature corresponding to each person pair in the updated fusion can be obtained based on the calculation result and image features. Then, prediction processing is performed based on each obtained feature to obtain the final detection result.
[0089] It is understood that in this embodiment, by combining the analysis of intra-group correlation and inter-group correlation, the features of the same person within the same group and the features of different people within the same group can be fully integrated to perform image feature fusion processing, so as to improve the accuracy of the detection results.
[0090] In one example, step S206 can be implemented as follows: based on the cross-attention layer of the decoder, cross-attention processing is performed on each second set and image features to obtain a second query matrix; the second query matrix includes the third set corresponding to each second set; the third set includes features in the interaction relationship between people; based on the feedforward neural network layer of the decoder, the second query matrix is processed to obtain the detection result corresponding to the image to be detected.
[0091] For example, in this embodiment, after obtaining each second set and image feature, a cross-attention mechanism can be used to fuse each second set and image feature to update each second set, thereby obtaining a third set used to indicate the interaction relationship between each pair of people in the image under test. Specifically, the third set includes each feature in the interaction relationship.
[0092] For example, if the query set includes a first query vector, a second query vector, and a third query vector, then the third set obtained after processing will also correspond to the first feature, the second feature, and the third feature in the character interaction relationship.
[0093] Furthermore, after obtaining the aforementioned third set, the third set can be input into the feedforward neural network layer in the decoder, and the detection result can be predicted based on the feedforward neural network layer to obtain the interaction relationship between people in the image to be detected.
[0094] It is understood that in this embodiment, by combining the cross-attention mechanism and the feedforward neural network, feature fusion and result prediction are performed on image features and the updated query vectors (i.e., the aforementioned second sets) in order to determine the interaction relationship between people in the image to be detected.
[0095] Figure 3 This is a schematic diagram of a model structure provided for an embodiment of this disclosure. For example... Figure 3 As shown in the figure, the model includes convolutional layers, an encoder, and a decoder; the model provided in this embodiment is used for human interaction detection. The decoder includes multiple decoding units connected in series and N feedforward neural network layers. Each decoding unit includes a first self-attention layer, a second self-attention layer, and a cross-attention layer connected in series. The first decoding unit receives the initial query matrix (which includes N query vectors, where N is a positive integer). The principles corresponding to each layer in the decoding unit can be found in [reference needed]. Figure 2 The descriptions in the illustrated embodiments will not be repeated here. The last decoding unit in the decoder can output N features, and each of these features is input into its corresponding feedforward neural network layer, so as to determine the information corresponding to a feature in the human-human interaction relationship based on a feature (e.g., the location information of the person, or the location information of the object, the person's interaction with the object, the category information of the object, etc.). In one possible implementation, a residual network and a normalization layer can be set after the first self-attention layer, the second self-attention layer, and the cross-attention layer of the decoding unit. For example, a residual network and a normalization layer can be set between the first self-attention layer and the second self-attention layer, and a residual network and a normalization layer can be set between the second self-attention layer and the cross-attention layer. It should be noted that the specific principles of the residual network and the normalization layer can be found in the descriptions in related technologies, and will not be repeated here.
[0096] Figure 4 This is a flowchart illustrating a model training method provided in an embodiment of the present disclosure. The method includes the following steps:
[0097] S401. Extract image features from the image to be trained; wherein the image to be trained has a first person interaction relationship.
[0098] For example, the training method provided in this embodiment is used to train a model capable of detecting human interactions. Specifically, it is first necessary to extract image features from the image to be trained to obtain the image features corresponding to the image to be trained. Furthermore, the image to be trained in this embodiment has a first human interaction relationship, which can be regarded as a label corresponding to the image to be trained, used to indicate the interaction information between people and objects contained in the image to be trained.
[0099] It should be noted that the image feature extraction method in this embodiment can be described in parameter step S101, and will not be repeated here.
[0100] S402. Obtain the query matrix to be trained; wherein, the query matrix to be trained includes multiple query sets; the query set is a set of parameters used to extract the interaction relationship between people, and the query set includes at least one query vector, which is used to extract the features corresponding to the query vector in the interaction relationship between people; the interaction relationship between people is the interaction relationship between people and objects.
[0101] For example, in this embodiment, before model training, a query matrix to be trained can be initialized and generated, and this query matrix contains multiple query sets. A query set is used to extract the interaction relationships between a group of people (people and objects) in an image. Furthermore, the number of features specifically corresponding to the interaction relationships corresponds to the number of query vectors contained in the query set, so that subsequent query vectors can be used to extract the features corresponding to the query vectors in the interaction relationships.
[0102] S403. Based on the initial decoder, perform feature fusion processing on the image features of the image to be trained and the query matrix to be trained to obtain the second person interaction relationship corresponding to the image to be trained.
[0103] For example, after obtaining the query matrix to be trained, the query matrix to be trained and the image features of the image to be trained can be input into the initial decoder so that the initial decoder can predict the second person interaction relationship corresponding to the image to be detected.
[0104] It should be noted that the specific principle of step S403 can be found in step S103, and will not be repeated in this embodiment.
[0105] S404. Based on the interaction relationship between the first character and the second character, the query matrix to be trained and the initial decoder are modified to obtain the trained decoder and the initial query matrix.
[0106] For example, in this embodiment, after obtaining the first person interaction relationship and the predicted second person interaction relationship, the parameters of the query matrix to be trained and the initial decoder can be modified and corrected according to the loss function constructed by the first person interaction relationship and the second person interaction relationship, so as to obtain the initial query matrix and decoder required for subsequent model use.
[0107] It is understood that in this embodiment, by decoupling the various features in the character interaction relationship, that is, extracting the features in the character interaction relationship corresponding to the query vector from a query vector in the query matrix, the method in this embodiment can reduce the difficulty of model training and improve the efficiency of model training compared to the training method of using a query vector to extract all features in the character interaction relationship.
[0108] In one example, at least one query vector includes a first query vector, a second query vector, and a third query vector;
[0109] The first query vector is used to extract the first feature in the interaction relationship between people, which is a feature used to indicate the location information of people; the second query vector is used to extract the second feature in the interaction relationship between people, which is a feature used to indicate the location information and category information of objects; and the third query vector is used to extract the third feature in the interaction relationship between people, which is a feature used to indicate the interactive actions between people and objects.
[0110] In one example, based on the initial decoder, feature fusion processing is performed on the image features of the image to be trained and the query matrix to be trained to obtain the second person interaction relationship corresponding to the image to be trained, including:
[0111] Based on the first self-attention layer in the initial decoder, the intra-group relevance information of the query set is determined, and the vectors in the query set are updated according to the intra-group relevance information to obtain the first set; the intra-group relevance information represents the relevance between the query vectors contained in the query set.
[0112] Based on the second self-attention layer of the initial decoder, the inter-group correlation information corresponding to the first set is determined; and the first set is updated according to the inter-group correlation information to obtain the second set; wherein, the inter-group correlation information represents the correlation between the first set and the first query matrix; the first query matrix is composed of each first set;
[0113] Based on each second set and image features, the detection result corresponding to the image to be trained is determined.
[0114] In one example, based on each second set and image features, the detection result corresponding to the image to be trained is determined, including:
[0115] Based on the cross-attention layer of the initial decoder, cross-attention processing is performed on each second set and image features to obtain the second query matrix; the second query matrix includes the third set corresponding to each second set; the third set includes features in the interaction relationship between characters;
[0116] Based on the feedforward neural network layer of the initial decoder, the second query matrix is processed to obtain the detection result corresponding to the image to be trained.
[0117] In one example, based on the second self-attention layer of the initial decoder, the inter-group correlation information corresponding to the first set is determined, including:
[0118] Based on the second self-attention layer, the relevance result of the fourth query vector in the first set is determined; the relevance result represents the relevance between the fourth query vector and the fourth query vector in each of the first sets; the fourth query vector is the result of updating the query vector in the query set for extracting target features based on the intra-group relevance information; the target feature is any one of the features used to indicate the interaction between people and objects, features used to indicate the location and category information of objects, and features used to indicate the location information of people.
[0119] The relevance results corresponding to each fourth query vector in the first set are determined as the inter-group relevance information corresponding to the first set.
[0120] In one example, the method also includes:
[0121] Based on the normalization layer in the initial decoder, each first set is normalized to obtain the processed first set.
[0122] In one example, image features are extracted from the image to be trained, including:
[0123] Based on the convolutional neural network layer, feature extraction processing is performed on the image to be trained to obtain the feature map information of the image to be trained; the feature map information is used to represent the local features of the image;
[0124] Based on the encoder, feature extraction is performed on the feature map information to obtain the image features of the image to be trained.
[0125] The method provided in this embodiment is similar to the one described above. Figure 1-2 The technical principles shown in the Chinese embodiments are similar and will not be repeated here.
[0126] Figure 5 This is a schematic diagram of the structure of a human interaction detection device provided in an embodiment of the present disclosure, wherein the human interaction detection device 500 includes:
[0127] The first extraction unit 501 is used to extract image features of the image to be detected;
[0128] The first acquisition unit 502 is used to acquire an initial query matrix; wherein, the initial query matrix includes multiple query sets; the query set is a parameter set used to extract the interaction relationship between people, and the query set includes at least one query vector, which is used to extract the feature corresponding to the query vector in the interaction relationship between people; the interaction relationship between people is the interaction relationship between people and objects;
[0129] The first processing unit 503 is used to perform feature fusion processing on image features and an initial query matrix based on the decoder to determine the detection result corresponding to the image to be detected; wherein the detection result represents the interaction relationship between people in the image to be detected.
[0130] The apparatus provided in this embodiment is used to implement the technical solution provided by the above method. Its implementation principle and technical effect are similar, and will not be described again.
[0131] Figure 6 This is a schematic diagram of another human interaction detection device provided in an embodiment of the present disclosure, wherein the human interaction detection device 600 includes:
[0132] The first extraction unit 601 is used to extract image features of the image to be detected;
[0133] The first acquisition unit 602 is used to acquire an initial query matrix; wherein, the initial query matrix includes multiple query sets; the query set is a parameter set used to extract the interaction relationship between people, and the query set includes at least one query vector, which is used to extract the feature corresponding to the query vector in the interaction relationship between people; the interaction relationship between people is the interaction relationship between people and objects;
[0134] The first processing unit 603 is used to perform feature fusion processing on image features and an initial query matrix based on the decoder to determine the detection result corresponding to the image to be detected; wherein the detection result represents the interaction relationship between people in the image to be detected.
[0135] In one example, at least one query vector includes a first query vector, a second query vector, and a third query vector;
[0136] The first query vector is used to extract the first feature in the interaction relationship between people, which is a feature used to indicate the location information of people; the second query vector is used to extract the second feature in the interaction relationship between people, which is a feature used to indicate the location information and category information of objects; and the third query vector is used to extract the third feature in the interaction relationship between people, which is a feature used to indicate the interactive actions between people and objects.
[0137] In one example, the first processing unit 603 includes:
[0138] The first determining module 6031 is used to determine the intra-group correlation information of the query set based on the first self-attention layer in the decoder;
[0139] The first update module 6032 is used to update the vectors in the query set according to the intra-group correlation information to obtain the first set; the intra-group correlation information represents the correlation between the query vectors contained in the query set;
[0140] The second determining module 6033 is used to determine the inter-group correlation information corresponding to the first set based on the second self-attention layer of the decoder.
[0141] The second update module 6034 is used to update the first set according to the inter-group correlation information to obtain the second set; wherein, the inter-group correlation information represents the correlation between the first set and the first query matrix; the first query matrix is composed of each first set;
[0142] The third determining module 6035 is used to determine the detection result corresponding to the image to be detected based on each second set and image features.
[0143] In one example, the third determining module 6035 includes:
[0144] The first processing submodule is used to perform cross-attention processing on each second set and image features based on the decoder's cross-attention layer to obtain a second query matrix; the second query matrix includes a third set corresponding to each second set; the third set includes features in the interaction relationship between characters;
[0145] The second processing submodule is used to process the second query matrix based on the feedforward neural network layer of the decoder to obtain the detection result corresponding to the image to be detected.
[0146] In one example, the second determining module 6033 includes:
[0147] The first determining submodule is used to determine the relevance result of the fourth query vector in the first set based on the second self-attention layer; the relevance result represents the relevance between the fourth query vector and the fourth query vector in each of the first sets; the fourth query vector is the result of updating the query vector in the query set for extracting target features based on the intra-group relevance information; the target feature is any one of the following: features for indicating the interaction between people and objects, features for indicating the location and category information of objects, and features for indicating the location information of people;
[0148] The second determining submodule is used to determine the relevance results corresponding to each fourth query vector in the first set as the inter-group relevance information corresponding to the first set.
[0149] In one example, the device also includes:
[0150] The first processing module is used to normalize each first set based on the normalization layer in the decoder to obtain the processed first set.
[0151] In one example, the first extraction unit 601 includes:
[0152] The second processing module 6011 is used to perform feature extraction processing on the image to be detected based on the convolutional neural network layer to obtain feature map information of the image to be detected; the feature map information is used to characterize the local features of the image.
[0153] The first extraction module 6012 is used to extract features from the feature map information based on the encoder to obtain the image features of the image to be detected.
[0154] The apparatus provided in this embodiment is used to implement the technical solution provided by the above method. Its implementation principle and technical effect are similar, and will not be described again.
[0155] Figure 7 This is a schematic diagram of a model training device provided in an embodiment of the present disclosure, wherein the model training device 700 includes:
[0156] The second extraction unit 701 is used to extract image features of the image to be trained; wherein the image to be trained has a first person interaction relationship;
[0157] The second acquisition unit 702 is used to acquire the query matrix to be trained; wherein, the query matrix to be trained includes multiple sets of query sets; the query set is a set of parameters used to extract the interaction relationship between people, and the query set includes at least one query vector, which is used to extract the feature corresponding to the query vector in the interaction relationship between people; the interaction relationship between people is the interaction relationship between people and objects;
[0158] The second processing unit 703 is used to perform feature fusion processing on the image features of the image to be trained and the query matrix to be trained based on the initial decoder, so as to obtain the second person interaction relationship corresponding to the image to be trained.
[0159] The correction unit 704 is used to correct the query matrix to be trained and the initial decoder based on the interaction relationship between the first character and the second character, so as to obtain the trained decoder and the initial query matrix.
[0160] The apparatus provided in this embodiment is used to implement the technical solution provided by the above method. Its implementation principle and technical effect are similar, and will not be described again.
[0161] Figure 8 This is a schematic diagram of the structure of a second model training device provided in an embodiment of the present disclosure, wherein the model training device 800 includes:
[0162] The second extraction unit 801 is used to extract image features of the image to be trained; wherein the image to be trained has a first person interaction relationship;
[0163] The second acquisition unit 802 is used to acquire the query matrix to be trained; wherein, the query matrix to be trained includes multiple sets of query sets; the query set is a set of parameters used to extract the interaction relationship between people, and the query set includes at least one query vector, which is used to extract the features in the interaction relationship between people and objects; the interaction relationship between people and objects is the interaction relationship between people and objects.
[0164] The second processing unit 803 is used to perform feature fusion processing on the image features of the image to be trained and the query matrix to be trained based on the initial decoder, so as to obtain the second person interaction relationship corresponding to the image to be trained.
[0165] The correction unit 804 is used to correct the query matrix to be trained and the initial decoder based on the interaction relationship between the first character and the second character, so as to obtain the trained decoder and the initial query matrix.
[0166] In one example, at least one query vector includes a first query vector, a second query vector, and a third query vector;
[0167] The first query vector is used to extract the first feature in the interaction relationship between people, which is a feature used to indicate the location information of people; the second query vector is used to extract the second feature in the interaction relationship between people, which is a feature used to indicate the location information and category information of objects; and the third query vector is used to extract the third feature in the interaction relationship between people, which is a feature used to indicate the interactive actions between people and objects.
[0168] In one example, the second processing unit 803 includes:
[0169] The fourth determining module 8031 is used to determine the intra-group correlation information of the query set based on the first self-attention layer in the initial decoder;
[0170] The third update module 8032 is used to update the vectors in the query set according to the intra-group correlation information to obtain the first set; the intra-group correlation information represents the correlation between the query vectors contained in the query set;
[0171] The fifth determining module 8033 is used to determine the inter-group correlation information corresponding to the first set based on the second self-attention layer of the initial decoder;
[0172] The fourth update module 8034 is used to update the first set according to the inter-group correlation information to obtain the second set; wherein, the inter-group correlation information represents the correlation between the first set and the first query matrix; the first query matrix is composed of each first set;
[0173] The sixth determining module 8035 is used to determine the detection result corresponding to the image to be trained based on each second set and image features.
[0174] In one example, the sixth determining module 8035 includes:
[0175] The third processing submodule is used to perform cross-attention processing on each second set and image features based on the cross-attention layer of the initial decoder to obtain the second query matrix; the second query matrix includes the third set corresponding to each second set; the third set includes features in the interaction relationship between characters;
[0176] The fourth processing submodule is used to process the second query matrix based on the feedforward neural network layer of the initial decoder to obtain the detection result corresponding to the image to be trained.
[0177] In one example, the fifth determining module 8033 includes:
[0178] The third determining submodule is used to determine the relevance result of the fourth query vector in the first set based on the second self-attention layer; the relevance result represents the relevance between the fourth query vector and the fourth query vector in each of the first sets; the fourth query vector is the result of updating the query vectors in the query set used to extract target features based on the intra-group relevance information; the target feature is any one of the features used to indicate the interaction between people and objects, features used to indicate the location and category information of objects, and features used to indicate the location information of people;
[0179] The fourth determination submodule is used to determine the relevance results corresponding to each fourth query vector in the first set as the inter-group relevance information corresponding to the first set.
[0180] In one example, the device also includes:
[0181] The third processing module is used to normalize each first set based on the normalization layer in the initial decoder to obtain the processed first set.
[0182] In one example, the second extraction unit 801 includes:
[0183] The fourth processing module 8011 is used to perform feature extraction processing on the training image based on the convolutional neural network layer to obtain the feature map information of the training image; the feature map information is used to represent the local features of the image;
[0184] The second extraction module 8012 is used to extract features from the feature map information based on the encoder to obtain the image features of the image to be trained.
[0185] The apparatus provided in this embodiment is used to implement the technical solution provided by the above method. Its implementation principle and technical effect are similar, and will not be described again.
[0186] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0187] This disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method provided in any of the above embodiments.
[0188] Figure 9 This is a schematic diagram of an electronic device provided in this disclosure, such as... Figure 9 As shown, the electronic device 900 in this disclosure may include a processor 901 and a memory 902.
[0189] Memory 902 is used to store programs. Memory 902 may include volatile memory, such as random-access memory (RAM), such as static random-access memory (SRAM), double data rate synchronous dynamic random-access memory (DDR SDRAM), etc.; memory may also include non-volatile memory, such as flash memory. Memory 902 is used to store computer programs (such as application programs, functional modules, etc. that implement the above methods), computer instructions, etc. The computer programs, computer instructions, etc., can be partitioned and stored in one or more memories 902. Furthermore, the computer programs, computer instructions, data, etc., can be accessed by processor 901.
[0190] The aforementioned computer programs and instructions can be stored in one or more partitions of memory 902. Furthermore, the aforementioned computer programs and instructions can be invoked by processor 901.
[0191] The processor 901 is configured to execute the computer program stored in the memory 902 to implement the various steps in the methods described in the above embodiments.
[0192] For details, please refer to the relevant descriptions in the preceding method embodiments.
[0193] The processor 901 and the memory 902 can be independent structures or integrated structures. When the processor 901 and the memory 902 are independent structures, the memory 902 and the processor 901 can be coupled together via bus 903.
[0194] The electronic device in this embodiment can execute the technical solution in the above method. Its specific implementation process and technical principle are the same, and will not be repeated here.
[0195] This disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods provided in any of the above embodiments.
[0196] According to embodiments of this disclosure, this disclosure also provides a computer program product comprising: a computer program stored in a readable storage medium, at least one processor of an electronic device being able to read the computer program from the readable storage medium, and the at least one processor executing the computer program causing the electronic device to perform the scheme provided in any of the above embodiments.
[0197] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0198] like Figure 10As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0199] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0200] The computing unit 1001 can be a variety of general-purpose and / or special-purpose processing units with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as human interaction detection methods or model training methods. For example, in some embodiments, the human interaction detection method, or the model training method, may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the human interaction detection method or model training method described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured in any other suitable manner (e.g., by means of firmware) to perform a human interaction detection method or a model training method.
[0201] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0202] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0203] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0204] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0205] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0206] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0207] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0208] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for detecting human interaction, wherein, The method includes: Extract image features from the image to be detected; Obtain an initial query matrix; wherein the initial query matrix includes multiple query sets; the query set is a parameter set used to extract character interaction relationships, the query set includes at least one query vector, the query vector is used to extract features in the character interaction relationships corresponding to the query vector; the character interaction relationship is the interaction relationship between a person and an object; Based on the decoder, feature fusion processing is performed on the image features and the initial query matrix to determine the detection result corresponding to the image to be detected; wherein, the detection result represents the interaction relationship between people in the image to be detected; Based on the decoder, feature fusion processing is performed on the image features and the initial query matrix to determine the detection result corresponding to the image to be detected, including: Based on the first self-attention layer in the decoder, the intra-group relevance information of the query set is determined, and the vectors in the query set are updated according to the intra-group relevance information to obtain the first set; the intra-group relevance information represents the relevance between the query vectors contained in the query set. Based on the second self-attention layer of the decoder, the inter-group correlation information corresponding to the first set is determined; and the first set is updated according to the inter-group correlation information to obtain the second set; wherein, the inter-group correlation information characterizes the correlation between the first set and the first query matrix; the first query matrix is composed of each first set; Based on each of the second sets and the image features, the detection result corresponding to the image to be detected is determined.
2. The method according to claim 1, wherein, The at least one query vector includes a first query vector, a second query vector, and a third query vector; The first query vector is used to extract a first feature from the interaction between people, and the first feature is a feature used to indicate the location information of people; the second query vector is used to extract a second feature from the interaction between people, and the second feature is a feature used to indicate the location information and category information of objects; the third query vector is used to extract a third feature from the interaction between people, and the third feature is a feature used to indicate the interactive actions between people and objects.
3. The method according to claim 1, wherein, Based on each of the second sets and the image features, the detection result corresponding to the image to be detected is determined, including: Based on the cross-attention layer of the decoder, cross-attention processing is performed on each of the second sets and the image features to obtain a second query matrix; the second query matrix includes a third set corresponding to each of the second sets; the third set includes features in the interaction relationship between characters; Based on the feedforward neural network layer of the decoder, the second query matrix is processed to obtain the detection result corresponding to the image to be detected.
4. The method according to claim 1, wherein, Based on the second self-attention layer of the decoder, the inter-group correlation information corresponding to the first set is determined, including: Based on the second self-attention layer, the relevance result of the fourth query vector in the first set is determined; the relevance result characterizes the relevance between the fourth query vector and the fourth query vector in each first set; the fourth query vector is the result of updating the query vector in the query set for extracting target features based on the intra-group relevance information; the target feature is any one of the following: features for indicating the interaction between a person and an object, features for indicating the location and category information of an object, and features for indicating the location information of a person; The relevance results corresponding to each fourth query vector in the first set are determined as the inter-group relevance information corresponding to the first set.
5. The method according to claim 1, further comprising: Based on the normalization layer in the decoder, each of the first sets is normalized to obtain the processed first set.
6. The method according to any one of claims 1-5, wherein, Extracting image features from the image to be detected, including: Based on the convolutional neural network layer, feature extraction processing is performed on the image to be detected to obtain feature map information of the image to be detected; the feature map information is used to characterize the local features of the image; Based on the encoder, feature extraction is performed on the feature map information to obtain the image features of the image to be detected.
7. A model training method, wherein, The method includes: Extract image features from the image to be trained; wherein the image to be trained has a first person interaction relationship; Obtain a query matrix to be trained; wherein, the query matrix to be trained includes multiple query sets; the query set is a parameter set used to extract human interaction relationships, the query set includes at least one query vector, the query vector is used to extract features in human interaction relationships corresponding to the query vector; the human interaction relationship is the interaction relationship between a person and an object; Based on the initial decoder, feature fusion processing is performed on the image features of the image to be trained and the query matrix to be trained to obtain the second person interaction relationship corresponding to the image to be trained; Based on the interaction relationship between the first character and the interaction relationship between the second character, the query matrix to be trained and the initial decoder are corrected to obtain the trained decoder and the initial query matrix. Based on the initial decoder, feature fusion processing is performed on the image features of the image to be trained and the query matrix to be trained to obtain the second person interaction relationship corresponding to the image to be trained, including: Based on the first self-attention layer in the initial decoder, the intra-group relevance information of the query set is determined, and the vectors in the query set are updated according to the intra-group relevance information to obtain a first set; the intra-group relevance information represents the relevance between the query vectors contained in the query set. Based on the second self-attention layer of the initial decoder, the inter-group correlation information corresponding to the first set is determined; and the first set is updated according to the inter-group correlation information to obtain the second set; wherein, the inter-group correlation information characterizes the correlation between the first set and the first query matrix; the first query matrix is composed of each first set; Based on each of the second sets and the image features, the detection result corresponding to the image to be trained is determined.
8. The method according to claim 7, wherein, The at least one query vector includes a first query vector, a second query vector, and a third query vector; The first query vector is used to extract a first feature from the interaction between people, and the first feature is a feature used to indicate the location information of people; the second query vector is used to extract a second feature from the interaction between people, and the second feature is a feature used to indicate the location information and category information of objects; the third query vector is used to extract a third feature from the interaction between people, and the third feature is a feature used to indicate the interactive actions between people and objects.
9. The method according to claim 7, wherein, Based on each of the second sets and the image features, the detection result corresponding to the image to be trained is determined, including: Based on the cross-attention layer of the initial decoder, cross-attention processing is performed on each of the second sets and the image features to obtain a second query matrix; the second query matrix includes a third set corresponding to each of the second sets; the third set includes features in the character interaction relationship; Based on the feedforward neural network layer of the initial decoder, the second query matrix is processed to obtain the detection result corresponding to the image to be trained.
10. The method according to claim 7, wherein, Based on the second self-attention layer of the initial decoder, the inter-group correlation information corresponding to the first set is determined, including: Based on the second self-attention layer, the relevance result of the fourth query vector in the first set is determined; the relevance result characterizes the relevance between the fourth query vector and the fourth query vector in each first set; the fourth query vector is the result of updating the query vector in the query set for extracting target features based on the intra-group relevance information; the target feature is any one of the following: features for indicating the interaction between a person and an object, features for indicating the location and category information of an object, and features for indicating the location information of a person; The relevance results corresponding to each fourth query vector in the first set are determined as the inter-group relevance information corresponding to the first set.
11. The method according to claim 7, further comprising: Based on the normalization layer in the initial decoder, each of the first sets is normalized to obtain the processed first set.
12. The method according to any one of claims 7-11, wherein, Extracting image features corresponding to the image to be trained includes: Based on the convolutional neural network layer, feature extraction processing is performed on the image to be trained to obtain feature map information of the image to be trained; the feature map information is used to characterize the local features of the image; Based on the encoder, feature extraction is performed on the feature map information to obtain the image features of the image to be trained.
13. A human interaction detection device, wherein, The device includes: The first extraction unit is used to extract image features from the image to be detected; The first acquisition unit is used to acquire an initial query matrix; wherein, the initial query matrix includes multiple query sets; the query set is a parameter set for extracting character interaction relationships, the query set includes at least one query vector, the query vector is used to extract features in the character interaction relationships corresponding to the query vector; the character interaction relationship is the interaction relationship between a person and an object; The first processing unit is configured to perform feature fusion processing on the image features and the initial query matrix based on the decoder to determine the detection result corresponding to the image to be detected; wherein the detection result represents the interaction relationship between people in the image to be detected; The first processing unit includes: The first determining module is used to determine the intra-group relevance information of the query set based on the first self-attention layer in the decoder; The first update module is used to update the vectors in the query set according to the intra-group correlation information to obtain a first set; the intra-group correlation information represents the correlation between the query vectors contained in the query set. The second determining module is used to determine the inter-group correlation information corresponding to the first set based on the second self-attention layer of the decoder; The second update module is used to update the first set according to the inter-group correlation information to obtain a second set; wherein the inter-group correlation information represents the correlation between the first set and the first query matrix; the first query matrix is composed of each first set; The third determining module is used to determine the detection result corresponding to the image to be detected based on each of the second sets and the image features.
14. The apparatus according to claim 13, wherein, The at least one query vector includes a first query vector, a second query vector, and a third query vector; The first query vector is used to extract a first feature from the interaction between people, and the first feature is a feature used to indicate the location information of people; the second query vector is used to extract a second feature from the interaction between people, and the second feature is a feature used to indicate the location information and category information of objects; the third query vector is used to extract a third feature from the interaction between people, and the third feature is a feature used to indicate the interactive actions between people and objects.
15. The apparatus according to claim 13, wherein, The third determining module includes: The first processing submodule is used to perform cross-attention processing on each of the second sets and the image features based on the cross-attention layer of the decoder to obtain a second query matrix; the second query matrix includes a third set corresponding to each of the second sets; the third set includes features in the interaction relationship between the characters; The second processing submodule is used to process the second query matrix based on the feedforward neural network layer of the decoder to obtain the detection result corresponding to the image to be detected.
16. The apparatus according to claim 13, wherein, The second determining module includes: The first determining submodule is used to determine the relevance result of the fourth query vector in the first set based on the second self-attention layer; the relevance result represents the relevance between the fourth query vector and the fourth query vector in each first set; the fourth query vector is the result of updating the query vector in the query set for extracting target features based on the intra-group relevance information; the target feature is any one of the following: features for indicating the interaction between a person and an object, features for indicating the location and category information of an object, and features for indicating the location information of a person; The second determining submodule is used to determine the relevance results corresponding to each fourth query vector in the first set as the inter-group relevance information corresponding to the first set.
17. The apparatus of claim 13, further comprising: The first processing module is used to perform normalization processing on each of the first sets based on the normalization layer in the decoder to obtain the processed first set.
18. The apparatus according to any one of claims 13-17, wherein, The first extraction unit includes: The second processing module is used to perform feature extraction processing on the image to be detected based on the convolutional neural network layer to obtain feature map information of the image to be detected; the feature map information is used to characterize the local features of the image; The first extraction module is used to extract features from the feature map information based on the encoder to obtain the image features of the image to be detected.
19. A model training device, wherein, The device includes: The second extraction unit is used to extract image features from the image to be trained; wherein the image to be trained has a first human interaction relationship; The second acquisition unit is used to acquire a query matrix to be trained; wherein, the query matrix to be trained includes multiple query sets; the query set is a parameter set for extracting human interaction relationships, the query set includes at least one query vector, the query vector is used to extract features in human interaction relationships corresponding to the query vector; the human interaction relationship is the interaction relationship between a person and an object; The second processing unit is used to perform feature fusion processing on the image features of the image to be trained and the query matrix to be trained based on the initial decoder, so as to obtain the second person interaction relationship corresponding to the image to be trained. The correction unit is used to correct the query matrix to be trained and the initial decoder according to the first character interaction relationship and the second character interaction relationship, so as to obtain the trained decoder and the initial query matrix. The second processing unit includes: The fourth determining module is used to determine the intra-group relevance information of the query set based on the first self-attention layer in the initial decoder; The third update module is used to update the vectors in the query set according to the intra-group correlation information to obtain the first set; the intra-group correlation information represents the correlation between the query vectors contained in the query set. The fifth determining module is used to determine the inter-group correlation information corresponding to the first set based on the second self-attention layer of the initial decoder; The fourth update module is used to update the first set according to the inter-group correlation information to obtain the second set; wherein the inter-group correlation information represents the correlation between the first set and the first query matrix; the first query matrix is composed of each first set; The sixth determining module is used to determine the detection result corresponding to the image to be trained based on each of the second sets and the image features.
20. The apparatus according to claim 19, wherein, The at least one query vector includes a first query vector, a second query vector, and a third query vector; The first query vector is used to extract a first feature from the interaction between people, and the first feature is a feature used to indicate the location information of people; the second query vector is used to extract a second feature from the interaction between people, and the second feature is a feature used to indicate the location information and category information of objects; the third query vector is used to extract a third feature from the interaction between people, and the third feature is a feature used to indicate the interactive actions between people and objects.
21. The apparatus according to claim 19, wherein, The sixth determining module includes: The third processing submodule is used to perform cross-attention processing on each of the second sets and the image features based on the cross-attention layer of the initial decoder to obtain a second query matrix; the second query matrix includes a third set corresponding to each of the second sets; the third set includes features in the character interaction relationship; The fourth processing submodule is used to process the second query matrix based on the feedforward neural network layer of the initial decoder to obtain the detection result corresponding to the image to be trained.
22. The apparatus according to claim 19, wherein, The fifth determining module includes: The third determining submodule is used to determine the relevance result of the fourth query vector in the first set based on the second self-attention layer; the relevance result represents the relevance between the fourth query vector and the fourth query vector in each first set; the fourth query vector is the result of updating the query vector in the query set for extracting target features based on the intra-group relevance information; the target feature is any one of the following: features for indicating the interaction between a person and an object, features for indicating the location and category information of an object, and features for indicating the location information of a person; The fourth determination submodule is used to determine the relevance results corresponding to each fourth query vector in the first set as the inter-group relevance information corresponding to the first set.
23. The apparatus of claim 19, further comprising: The third processing module is used to perform normalization processing on each of the first sets based on the normalization layer in the initial decoder to obtain the processed first set.
24. The apparatus according to any one of claims 19-23, wherein, The second extraction unit includes: The fourth processing module is used to perform feature extraction processing on the image to be trained based on the convolutional neural network layer to obtain the feature map information of the image to be trained; the feature map information is used to characterize the local features of the image; The second extraction module is used to extract features from the feature map information based on the encoder to obtain the image features of the image to be trained.
25. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-12.
26. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-12.
27. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-12.