Fine-grained HOI detection method and system based on human-object action correlation mining

By constructing a character action correlation mining model, using a convolutional neural network and a Transformer encoder, combining interactive discriminator and adaptive loss function, HOI-class correlation mining is optimized, which solves the problem of high model complexity in the existing technology and realizes efficient identification of fine-grained HOI-class.

CN120375072APending Publication Date: 2025-07-25SHANDONG UNIV OF FINANCE & ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510470269.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to improve the recognition ability of fine-grained human-object interaction behavior detection without increasing the complexity of the model, and the model structure is complex and memory and time-consuming.

Method used

By constructing a character action correlation mining model, using convolutional neural network to extract features, combining Transformer encoder and decoder, design interactive discriminator and adaptive loss function, optimize HOI-class correlation mining, and improve HOI-class recognition capabilities.

Benefits of technology

Without increasing the complexity of the model, the recognition capability of fine-grained HOIs is significantly improved and the detection performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375072A_ABST
    Figure CN120375072A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of human-object interaction behavior detection and recognition, in particular to a fine-grained HOI detection method and system based on human-object action correlation mining, and the method comprises the steps: extracting image features, and transmitting the extracted image features and a position code Pos into a decoder; a human-object decoder outputs a human detection frame, an object category, an object detection frame and a human-object pair interaction score, an interaction behavior decoder outputs a specific interaction behavior category, and the human-object pair interaction score and the interaction behavior category output by the model are utilized to mine a correlation relationship between HOI categories so as to find the HOI categories which are easy to classify by mistake; then, a self-adaptive loss function is designed based on the correlation between HOI classes to guide the model to pay more attention to the HOI classes prone to error classification, and finally a character interaction pair with the highest interaction score is calculated, so that the recognition capability of the fine-grained HOIs classes is improved as much as possible on the premise that the complexity of the model is not increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-object interaction behavior detection and recognition, and particularly to a fine-grained HOI detection method and system based on mining the correlation between human and object actions. Background Art

[0002] In real life, humans usually play the main role. Therefore, it is particularly important to understand the relationship between humans and the surrounding environment. Human-Object Interaction (HOI) detection in the surveillance scenario can achieve refined scene understanding. By detecting the interaction behaviors between humans and objects in the surveillance scenario, abnormal behaviors can be investigated, tracked, and warned, reducing the occurrence of potential dangers and providing key support for the implementation of an intelligent surveillance system. Among them, HOI detection aims to detect humans and objects and infer the interaction relationship between humans and objects.

[0003] In real natural scenarios, the differences between some human-object interaction behaviors are very small, making the detection of these fine-grained HOIs very difficult. Existing methods extract discriminative feature representations by designing complex model structures, and the effect has been improved to a certain extent. However, the model structure is complex and very memory- and time-consuming. Therefore, how to improve the recognition ability of fine-grained HOIs classes as much as possible without increasing the model complexity is an important and challenging problem. Summary of the Invention

[0004] To solve the technical problems existing in the above background art, the present invention provides a fine-grained HOI detection method and system based on mining the correlation between human and object actions, which can improve the recognition ability of fine-grained HOIs classes.

[0005] To implement the above technical solution, in the first aspect, the present invention provides a fine-grained HOI detection method based on mining the correlation between human and object actions, including the following steps: Step 1: Construct a model for mining the correlation between human and object actions; Step 2: Train the constructed model for mining the correlation between human and object actions; Step 3: Input an image into the trained model for mining the correlation between human and object actions to obtain human detection boxes, object detection boxes, categories, and action categories; The said Step 2 includes: Obtain a training image sample set, and label the training image sample set to generate a detection benchmark dataset GT labeled with actual human detection boxes, actual object detection boxes, actual object categories, actual human matching pairs, and actual human interaction action categories; Convert the obtained training image sample set to a unified image format by performing image format conversion; Input the training image sample set converted to a unified format into a CNN-based feature extractor to extract image features and the position encoding Pos; Input the image features and the position encoding Pos into an encoder to obtain the context information Xs; Input the obtained context information Xs, the position encoding Pos, and the learnable query vector Q into a decoder to obtain human detection boxes, object detection boxes, object categories, and human interaction behavior categories; The interactivity discriminator judges the multiple person-object pairings generated by the human decoder based on binary classification for each person-object pairing in the generated pairings to determine whether there is an interaction, and calculate the interactivity score for each person-object pairing ; Based on the interactivity score for each person-object pairing and the HOI category for each person-object pairing calculate the HOI class correlation, and generate an interaction behavior adaptive function based on the calculated HOI class correlation; Update the HOICM model parameters according to the interaction behavior adaptive loss function.

[0006] Further, the step of inputting the obtained context information Xs, the position encoding Pos, and the learnable query vector Q into a decoder to obtain human detection boxes, object detection boxes, object categories, and human interaction behavior categories includes: Randomly initialize the learnable query vector Q; The human decoder decodes using the context information Xs , the position encoding Pos, and the learnable query vector Q as inputs, and outputs human detection boxes , object categories , and object detection boxes , and based on the human detection boxes and the object detection boxes , generate multiple person-object pairings ; Use the multiple person-object pairings generated by the human decoder as the initial query vector Q, and together with the context information Xs and the encoding positions as inputs for decoding to output the class scores for each person-object pairing ; and predict the HOI category for each person-object pair based on the category scores.

[0007] Further, calculating the HOI class correlation based on the interactivity score for each person-object pair and the HOI category of each person-object matching pair includes: Based on the actual person interaction behavior class and the predicted HOI class of each person-object pair, calculate the frequency matrix of the category error for each person-object pair; Perform row normalization on the predicted frequency matrix to obtain the correlation matrix of the probability of misidentifying the person-object pair category .

[0008] Further, the interactivity score based on each person-object pair and the HOI category of each person-object matching pair includes: According to the actual person interaction behavior class in the detection data benchmark dataset GT and the obtained person interaction behavior category, obtain the following frequency matrix.

[0009]

[0010] Further, performing row normalization on the predicted frequency matrix to obtain the correlation matrix of the probability of misidentifying the person-object pair category : Perform row normalization on the predicted frequency matrix, that is, divide the value of each row in the predicted frequency matrix by the absolute value of the sum of the squares of all elements in each row, so that the sum of the squares of the elements in each row of the matrix is 1; And set the diagonal elements to 1.

[0011] Further, the interactive behavior adaptive loss function is designed as follows: ; where the total loss is composed of the person and object detection box regression loss , the IoU loss , the interactivity loss , the object category loss , and the interactive behavior adaptive loss , that is: ; where the hyperparameters , , , , represent the weights of each loss function respectively.

[0012] Further, the human-object action correlation mining model includes: a feature extractor based on a convolutional neural network, a Transformer encoder, a decoder, an interaction discriminator, an HOI class correlation mining module, and an adaptive HOI learning model.

[0013] Further, the decoder includes: a human decoder and an interaction behavior decoder; wherein, the human decoder and the interaction behavior decoder are cascaded.

[0014] The human decoder includes N Transformer decoder layers and three feed-forward neural network heads; wherein, each Transformer decoder consists of a self-attention module and a multi-head collaborative attention module.

[0015] In a second aspect, the present invention provides a computer-readable storage medium, which includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned fine-grained HOI detection method based on human-object action correlation mining.

[0016] The beneficial effects of the present invention are as follows: The present invention extracts image features and sends the extracted image features and the position encoding Pos into the decoder together; the human-object decoder outputs a human detection box, an object category, an object detection box, and a human-object pair interaction score, and the interaction behavior decoder outputs a specific interaction behavior category, and uses the human-object pair interaction score and the interaction behavior category output by the model to mine the correlation relationship between HOI classes to discover HOI classes that are prone to misclassification; then, an adaptive loss function is designed based on the correlation relationship between HOI classes to guide the model to pay more attention to HOI classes that are prone to misclassification, and finally, the human interaction pair with the highest interaction score is calculated, which helps to improve the recognition ability of fine-grained HOIs classes as much as possible without increasing the model complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings forming a part of this invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0018] Figure 1 It is a flowchart of a fine-grained HOI detection method based on human-object action correlation mining of the present invention.

[0019] Figure 2 It is a framework diagram of the HOICM model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0021] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, each technical and scientific term used in this embodiment has the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0022] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0023] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relational terms determined for the convenience of describing the structural relationship of each component or element of the present invention and do not specifically refer to any component or element of the present invention and should not be construed as a limitation to the present invention.

[0024] In the present invention, terms such as "fixed connection", "connected", "connected to" should be understood in a broad sense, which may mean a fixed connection, an integral connection or a detachable connection; it may be directly connected or indirectly connected through an intermediate medium. For those related scientific research or technical personnel in the field, the specific meanings of the above terms in the present invention can be determined according to specific circumstances and should not be construed as a limitation to the present invention.

[0025] Example 1: As Figure 1 shown, this embodiment provides a fine-grained HOI detection method based on human-object action correlation mining, including the following steps: (1) Construct a human-object action correlation mining model. [A1] Specifically, the human-object action correlation mining model includes: (1) A feature extractor based on a convolutional neural network (CNN), which is used to extract features from the input two-dimensional image, obtain initial features, flatten the initial features to obtain a one-dimensional feature vector, and record the encoding position Pos of the one-dimensional feature vector; ; where R represents a real number, indicating that the value of Pos is a real value; H represents the image height, that is, the ordinate value of the image, usually the number of pixels from the top to the bottom of the image; W (Width): represents the width of the image, that is, the abscissa value of the image; usually the number of pixels from the left to the right of the image; C represents the number of channels.

[0026] (2)A Transformer encoder, which is used to encode one-dimensional feature vectors and encoding positions to obtain rich context information Xs. It should be noted that the context information Xs is the visual feature.

[0027] (3)A decoder, which is used to decode the context information Xs to obtain human detection boxes, object detection boxes, object categories, and human-object interaction (HOI) categories.

[0028] The decoder includes: a human-object decoder and an interaction behavior decoder; among them, the human-object decoder and the interaction behavior decoder are cascaded, and the human-object decoder is connected to three feed-forward neural networks.

[0029] A. The human-object decoder decodes with the context information Xs, position encoding Pos, and learnable query vector Q as inputs, and respectively outputs the human detection box , object category and object detection box through the three feed-forward neural networks connected to it, and combines the human detection box and object detection box to generate multiple human-object pairings ; The human-object pairing can be expressed as: ; represents the k-th detected human-object matching pair, and m represents the number of matching pairs.

[0030] For the sake of convenience of explanation, the human-object decoder is denoted as the human-object decoder .

[0031] More specifically, the human-object decoder includes N Transformer decoder layers and three feed-forward neural network (FFN) head parts. Among them, each Transformer decoder consists of a self-attention module and a multi-head collaborative attention module. Specifically, in the forward propagation process, the learnable query vector Q , context information Xs, and position encoding Pos are input into the Transformer decoder. Among them, the context information Xs and position encoding Pos are input into the Transformer decoder as key-value pairs. Each Transformer decoder first applies a self-attention module to all query vectors Q, and then performs a multi-head collaborative attention operation between the query vectors Q and the key-value pairs, and finally outputs a set of updated query vectors Q1.

[0032] For the three FFN head parts, they are respectively used to obtain the human detection box , object category and object detection bounding boxes and combine the human detection bounding boxes and object detection bounding boxes to generate multiple human-object matching pairs Among them, the FFN head part corresponding to the human detection bounding box and the object detection bounding box consists of three linear layers and a ReLU function, while the FFN head part for identifying the object category consists of one linear layer.

[0033] Among them, it should be noted that the learnable query vector Q is initialized manually, and its size is ; denotes the dimension of the query vector Q, which is the same as the dimension of the decoder hidden layer, and is set to 256 here; Nd denotes the number of query vectors Q, which is fixed at 64 for the HICO-DET dataset. Each query vector Q is a 256-dimensional vector corresponding to a potential target prediction.

[0034] B. Interaction behavior decoder, which is used to use the multiple human-object pairings generated by the human decoder as the initial query vector Q, and together with the context information Xs and the encoded position Pos as inputs for decoding, and output the category scores of each human-object pairing through the FFN head; and predict the HOI class of each human-object pairing based on the category scores. ; and predict the HOI class of each human-object pairing based on the category scores.

[0035] Among them, the specific calculation formula of the category score is as follows: ; Among them, is the category score; is the interaction behavior decoder, and each human-object matching pair , , represents the number of HOI classes.

[0036] Predict the HOI class of each human-object pairing based on the category scores, and the specific formula is as follows: Specifically, the class ; Among them, j is the HOI class, and the class with the largest score is taken as the HOI class of the human-object pairing.

[0037] (4) Interaction discriminator, which is used to judge whether there is an interaction in each human-object pairing in the multiple human-object pairings generated by the human-object decoder and calculate the interaction score of each human-object pairing .

[0038] The specific calculation formula is as follows:

[0039] Among them, is the interactivity score, is the interactivity discriminator.

[0040] (5) The HOI class correlation mining module is used to calculate the HOI class correlation based on the interactivity score and the class score.

[0041] (6) The adaptive HOI learning model is used to perform adaptive learning to update the model parameters.

[0042] (2) Train the constructed human action correlation mining model, as Figure 2 shown.

[0043] S1: Obtain the training image sample set and annotate the training image sample set to generate the detection benchmark dataset GT annotated with actual human detection boxes, actual object detection boxes, actual object categories, actual human-object matching pairs, and actual human-object interaction behavior classes.

[0044] Among them, the actual human-object interaction behavior classes are: , where represents the th class of actions, represents the index value.

[0045] S2: Convert the obtained training image sample set into an image format to convert it into a unified image format.

[0046] For example, convert the obtained training image sample set into an RGB format image of a unified size, and the unified size includes the width, height, and number of channels of the image.

[0047] S3: Input the training image sample set converted into a unified format into the CNN-based feature extractor to extract image features and position encoding Pos.

[0048] Specifically, input the training image sample set in a unified format into the CNN-based feature extractor for feature extraction to obtain initial features, flatten the initial features to obtain flattened one-dimensional features, and record the encoding positions Pos of the one-dimensional feature vectors; S4: Input the image features and the position encoding Pos into the encoder to obtain the context information Xs.

[0049] S5: Input the obtained context information Xs, the position encoding Pos, and the learnable query vector Q into the decoder to obtain the human detection box, the object detection box, the object category, and the human-object interaction (HOI) category.

[0050] Specifically, it includes the following steps: S5-1: Randomly initialize the learnable query vector Q; S5-2: The person-object decoder decodes with the context information Xs, the position encoding Pos, and the learnable query vector Q as inputs, and respectively obtains the person detection box , the object category , and the object detection box through three feed-forward neural networks connected thereto, and generates multiple person-object pairs based on the person detection box and the object detection box ; Specifically, it is expressed as follows:

[0051] Among them, based on the person detection box and the object detection box , multiple person-object pairs are generated as follows: ; represents the k-th person-object matching pair detected, and m represents the number of matching pairs.

[0052] S5-3: Use the multiple person-object pairs generated by the person decoder as the initial query vector Q, and decode with the context information Xs and the encoded position Pos as inputs, and output the class score of each person-object pair through the FFN head; and predict the HOI class of each person-object pair based on the class score.

[0053] Among them, the specific calculation formula of the class score is as follows: ; Among them, is the class score; is the interaction behavior decoder, and each person-object matching pair , , represents the number of HOI classes.

[0054] Predict the HOI class of each person-object pair based on the class score, and the specific formula is as follows: ; Among them, j is the HOI class, and the class with the largest score is taken as the HOI class of the person-object pair.

[0055] Among them, according to the output The correlation relationships between HOI classes can be obtained. For example, if the actual label is person-object pair of class is mostly recognized as class, then it can be considered that HOI classes and are very similar and easily confused behaviors.

[0056] S6: The interactivity discriminator judges whether there is an interaction for each person-object pair generated by the person-object decoder based on binary classification in each person-object pair and calculates the interactivity score for each person-object pair The specific calculation formula is as follows:

[0057] ; ; where is the interactivity score, is the interactivity discriminator.

[0058] S7: Calculate the HOI class correlation based on the interactivity score of each person-object pair and the HOI class of each person-object matching pair The specific steps are as follows:

[0059] S7-1: Initialize the frequency matrix , ; ; S7-2: Calculate the frequency matrix of the category error for each person-object pair based on the actual person-object interaction behavior class and the predicted HOI class of each person-object pair.

[0060] Among them, the frequency matrix of the category error is expressed as: ; where represents the element in the i-th row and j-th column of the matrix , The number of rows and columns in represent the number of HOI classes.

[0061] S7-3: According to the frequency that the actual person-object interaction behavior class in the detection data benchmark dataset GT is correctly predicted as and the frequency that it is wrongly predicted as , obtain the element in the frequency matrix The calculation formula is as follows: .

[0062] S7-4: Initialize relevant matrices , and ; S7-5: Perform row normalization on the prediction frequency matrix to obtain the relevant matrix of the misrecognition probability of the person-object pairing category .

[0063] Specifically, divide the values of each row of the frequency matrix by the absolute value of the sum of the squares of all elements in each row, so that the sum of the squares of the elements in each row of the matrix is 1. The specific formula is as follows:

[0064] Then make the diagonal elements equal to 1. The specific formula is as follows: ; where represents the element in the i-th row and j-th column of the matrix , that is, the correlation between the HOI class and itself is 1. Among them, the number of rows and columns of represents the probability that the sample belongs to but is recognized as .

[0065] S7-6: Generate an interactive behavior adaptive loss function according to the interactivity discrimination score and the HOI class correlation matrix .

[0066] Specifically, design an interactive behavior adaptive loss function according to the interactivity discrimination score and the HOI class correlation matrix to dynamically adjust the attention degree to difficult HOI classes. In other words, for the person-object candidate pairs with low interactivity discrimination scores, their influence on the loss function should be reduced; for the HOI classes with high correlations, the penalty for misclassified categories should be increased. Therefore, the interactive behavior adaptive loss function is designed as:

[0067] where is the balance factor, set to the ratio of positive samples to negative samples; is the adjustable factor, set to 2; is the label, corresponding to 0 and 1 in binary classification; the total loss is composed of the regression loss of the person and object detection boxes , the IoU loss , the interactivity loss , the object category loss Interactive behavior adaptive loss It consists of, that is:

[0068] Among them, the hyperparameters , , , , respectively represent the weights of each loss function.

[0069] S8: Update the HOICM model parameters according to the interactive behavior adaptive loss function, where the model parameters include: the parameters in the person-object decoder, the interactive behavior decoder, and the interactivity discriminator.

[0070] Specifically, update the HOICM model parameters according to the above adaptive function. After each epoch of training, that is, after all samples in the training set have been learned, update the HOI class correlation matrix , and update the model's parameters and loss function to guide the model to adaptively focus on difficult HOI classes.

[0071] (3) Input the image into the trained human action correlation mining model to obtain the human detection box, object detection box and category, and behavior category.

[0072] Inference: Specifically, according to the person-object decoder and the interactive behavior decoder , the human detection box , object category , object detection box , interactivity score , category score can be obtained, and the person-object pair with the highest interactivity score is output.

[0073]

[0074] Therefore, the predicted output can be obtained as , that is, the model outputs the person-object matching pair with the highest score.

[0075] Specifically, for the person-object interaction behavior detection task, input an image, and the HOICM model finally outputs <human detection box, object detection box and category, behavior category>. For example, if there is a person riding a horse in an image, the model will output, <the position of the person, the position and category of the horse, the behavior of riding>.

[0076] Verification results of this method: To verify the effectiveness of the HOICM method, human-object interaction behavior detection evaluation was carried out on the publicly available HICO-DET dataset. Consistent with existing HOI detection methods, the mean average precision (mAP) was used to evaluate the performance of the model. A correctly detected human-object interaction behavior HOI should have the correct behavior category, and the intersection over union (IoU) value between the predicted bounding boxes of the person and the object and their ground truth boxes should be greater than 0.5.

[0077] Table 1 Comparison of detection performance of different methods on the HICO-DET dataset

[0078] On the HICO-DET dataset, results of human-object interaction detection were obtained for three different sets: the full set, the rare set, and the non-rare set, under the default settings. As can be seen from Table 1, the HOICM method showed better performance compared to other methods. This indicates that the proposed method of discovering confusing HOI classes by analyzing the correlation of HOIs and guiding the model to pay more attention to the learning of difficult HOI classes can effectively improve the detection performance of the method.

[0079] Example 2: This example provides a computer-readable storage medium. The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the fine-grained HOI detection method based on human-object action correlation mining described in Example 1.

[0080] For the same and similar parts between the various embodiments in this specification, reference can be made to each other. In particular, for the terminal embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the method embodiments.

[0081] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the systems or units can be in electrical, mechanical or other forms.

[0082] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0083] In addition, it should be noted that the flowchart in the accompanying drawings shows the method of the embodiments of the present disclosure. In the corresponding descriptions in the flowchart or block diagram in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in an order different from that disclosed in the description. Sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and sometimes they can also be executed in the reverse order, which can depend on the functions involved. Each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0084] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A fine-grained HOI detection method based on the mining of human-object action correlation, characterized in that It includes the following steps: Step 1: Construct a model for mining the correlation of human actions; Step 2: Train the constructed model for mining the correlation of human actions; Step 3: Input the image into the trained model for mining the correlation of human actions to obtain human detection boxes, object detection boxes, categories, and action categories; The said Step 2 includes: Obtain a training image sample set and annotate the training image sample set to generate a detection benchmark dataset GT labeled with actual human detection boxes, actual object detection boxes, actual object categories, actual human matching pairs, and actual human interaction behavior categories; Convert the obtained training image sample set in image format to a unified image format; Input the training image sample set converted to a unified format into a CNN-based feature extractor to extract image features and position encoding Pos; Input the image features and the position encoding Pos into an encoder to obtain context information Xs; Input the obtained context information Xs, the position encoding Pos, and a learnable query vector Q into a decoder to obtain human detection boxes, object detection boxes, object categories, and human-object interaction behavior categories; The interactive discriminator determines multiple person-object pairs generated by the person decoder based on binary classification, and for each person-object pair in it, determines whether there is an interaction, and calculates the interaction score of each person-object pair; ​ Based on the interaction score for each person-object pairing and each person-object matching pair calculate the HOI class correlation, and generate an interaction behavior adaptation function based on the calculated HOI class correlation; Update the HOICM model parameters according to the interaction behavior adaptive loss function.

2. The fine-grained HOI detection method based on mining the correlation of human-object actions according to claim 1, wherein The step of inputting the obtained context information Xs, the position encoding Pos, and the learnable query vector Q into a decoder to obtain human detection boxes, object detection boxes, object categories, and human-object interaction behavior categories includes: Randomly initialize the learnable query vector Q; The human-object decoder decodes with the context information Xs, the position encoding Pos, and the learnable query vector Q as inputs, and outputs the human detection box , the object category , and the object detection box , and based on the human detection box and the object detection box , generates multiple human-object pairs ; Multiple person-object pairings generated by the person-object decoder are used as the initial query vector Q, and decoded with the context information Xs and the encoding positions as inputs to output the class scores of each person-object pairing ; and predict the HOI class of each person-object pairing based on the class scores.

3. The fine-grained HOI detection method based on human-object action correlation mining according to claim 2, characterized in that The calculation of the HOI class correlation based on the interaction score for each human-object pair and the HOI class for each human-object matching pair includes: Based on the actual human-object interaction behavior class and the predicted HOI class for each human-object pair, calculate the frequency matrix of class errors for each human-object pair; Perform row normalization on the prediction frequency matrix to obtain a correlation matrix of the misidentification probability of the person-object pairing category .

4. The fine-grained HOI detection method based on human-object action correlation mining according to claim 3, wherein Based on the actual human-object interaction behavior class and the predicted HOI class for each human-object pair, calculate the frequency matrix of class errors for each human-object pair includes: Initialize the frequency matrix , ; Calculate the frequency matrix of the class error for each person-object pair based on the actual human-object interaction behavior classes and the predicted HOI classes for each person-object pair: ; Among them, represents a matrix the element in the i-th row and j-th column of the number of rows and columns in represents the number of HOI categories; According to the actual human-object interaction behavior classes in the detection data benchmark dataset GT The frequency correctly predicted as and the frequency wrongly predicted as are used to obtain the frequency matrix The element in it is calculated as follows: 。 5. The fine-grained HOI detection method based on human-object action correlation mining according to claim 4, wherein Performing row normalization on the prediction frequency matrix to obtain a correlation matrix of the misrecognition probability of the individual-object pairing category comprising the following steps: Perform row normalization on the predicted frequency matrix, that is, divide the values in each row of the predicted frequency matrix by the absolute value of the sum of the squares of all elements in each row, so that the sum of the squares of the elements in each row of the matrix is 1; And set the diagonal elements to 1.

6. The fine-grained HOI detection method based on human-object action correlation mining according to claim 5, characterized in that, The interaction behavior adaptive loss function is designed as follows: ; Among them, the total loss is composed of the regression loss of human and object detection boxes , the IoU loss , the interaction loss , the object category loss , and the interaction behavior adaptive loss , that is: ; Among them, the hyperparameters , , , , respectively represent the weights of each loss function.

7. The fine-grained HOI detection method based on human-object action correlation mining according to claim 1, characterized in that, The model for mining the correlation of human actions includes: a CNN-based feature extractor, a Transformer encoder, a decoder, an interaction discriminator, an HOI class correlation mining module, and an adaptive HOI learning model.

8. The fine-grained HOI detection method based on human-object action correlation mining according to claim 7, wherein The decoder includes: a human-object decoder and an interaction behavior decoder; among them, the human-object decoder and the interaction behavior decoder are cascaded; The human-object decoder includes N Transformer decoder layers and three feed-forward neural network heads; among them, each Transformer decoder consists of a self-attention module and a multi-head collaborative attention module.

9. A computer-readable storage medium, characterized in that, A computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the fine-grained HOI detection method based on human-object action correlation mining according to any one of claims 1 to 8.