Classifier recalibration method for Transform architecture human-object interaction detection model
By classifying the human-object interaction detection model of the Transformer architecture, the prototype supervised comparison learning and difficult-negative feature mining technology are used to solve the wrong prediction problem caused by long-tail deviation, and the recognition accuracy and efficiency of the model are improved.
Patent Information
- Application Number
- CN202510168635.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-11
AI Technical Summary
The existing Transformer-based human-object interaction detection model has long-tail deviations under the long-tail distribution, resulting in incorrect predictions. The existing methods have failed to effectively alleviate this problem.
The prototype supervised comparison learning method is used to recalibrate the human-object interaction detection classifier. By extracting positive and negative features and separating difficult-negative features from negative features, the category prototype is used for optimization calibration, and combined with the image adaptive difficult-negative feature mining technology, the classifier is efficiently calibrated.
It significantly improves the recognition accuracy of the model, especially for tail action categories with very small samples, and improves the accuracy and efficiency of HOI detection.
Smart Images

Figure CN120298742A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of human-object interaction detection, and in particular to a method for recalibrating a classifier for a human-object interaction detection model based on the Transformer architecture. Background Art
[0002] As an important direction in the field of artificial intelligence, computer vision technology is booming at an astonishing speed and is having a profound impact on our real life. In recent years, the computer's understanding of the visual world is no longer satisfied with the instance level, so the reasoning of the relationships between instances has attracted considerable research interest. This type of task not only cares about various objects in the scene, but pays more attention to the abstract relationships between various types of objects in the scene. The human-object interaction detection task is one of the important basic tasks in this research direction, which further promotes the computer vision's understanding ability and influence degree of the real world and is widely applied in fields such as information retrieval, real-time monitoring, and media entertainment.
[0003] The human-object interaction detection task aims to explore the complex relationships between humans and various objects. These relationships usually manifest as <human, object, interaction> triples. Existing human-object interaction detection models can be roughly divided into two categories: two-stage methods and single-stage methods. Two-stage methods rely on pre-trained object detectors to identify all humans and objects, then pair them one by one, and predict the human-object interaction category for each pair. In contrast, single-stage methods predict human-object interaction triples in a unified manner. This method avoids the enumeration of human-object pairs, thus achieving more efficient reasoning. In recent years, inspired by the detection Transformer, single-stage methods regard human-object interaction detection as a set prediction problem. Due to their excellent accuracy and efficiency, these methods have become the mainstream methods in the current research on human-object interaction detection.
[0004] Recently, most of the research on Transformer-based human-object interaction detection models has focused on optimizing the feature extraction process and improving the Transformer encoding and decoding structures. For example, some research improves the feature quality by introducing high-quality image features, interaction-related language prior information, body part features, or multi-scale features. Other research designs more specialized encoders and decoders to better adapt to the characteristics of the human-object interaction detection task. However, few studies focus on optimizing the HOI (Human-Object Interaction) classifier, and our method provides a valuable supplement to the research in this field, and this method can be effectively compatible with other works. Summary of the Invention
[0005] The purpose of the present disclosure is to overcome the above-mentioned defects existing in the prior art, and to provide a classifier recalibration method for a human-object interaction detection model based on the Transformer architecture, which can explicitly recalibrate the human-object interaction detection classifier for the mispredictions caused by the long-tail bias, alleviate the long-tail bias of the model, and improve the recognition accuracy of the model.
[0006] In a first aspect, the present case proposes a classifier recalibration method for a human-object interaction detection model based on the Transformer architecture. The human-object interaction detection model includes a backbone network, a Transformer encoder-decoder, and a learnable human-object interaction classifier. The method includes the following steps: training the human-object interaction detection model using a training set of images to be recognized, and freezing the parameters of the trained model; for all the images to be recognized in the training set, using the backbone network and the Transformer encoder-decoder with frozen parameters to extract the human-object interaction instance features in the images to be recognized, classifying the instance features into positive and negative features, and further separating the hard negative features from the negative features; taking the weight of the human-object interaction classifier on class c as the class prototype p c , using the form of prototype supervised contrastive learning to calculate the loss function for each class c, and performing backpropagation independently by class to optimize and calibrate the class prototype p c ; the loss function L of the prototype supervised contrastive learning c is calculated as follows:
[0007]
[0008] In the formula: is the i-th positive feature of the human-object interaction detection model on class c, is the i-th hard negative feature of the human-object interaction detection model on class c, sim is the cosine similarity, τ is the contrastive learning temperature coefficient, N + and N - are the numbers of positive and negative features in the entire training set.
[0009] In an implementation of the above technical solution, the step of classifying the instance features into positive and negative features includes: through a human-object interaction matching algorithm based on Hungarian matching, matching the true label and the prediction result Y one by one, where is the unique object label, is the human-object interaction label in multi-hot form; calculating the total matching cost, and weighted selecting the prediction result with the lowest total matching cost as the most accurate prediction corresponding to the true label, denoted as N gtRepresents the number of human-object interaction instances in the image; the feature vector of the human-object interaction instance corresponding to the most accurate prediction Y* As positive features, the rest are negative features.
[0010] In an implementation of the above technical solution, by adopting a difficult negative feature mining mechanism for picture adaptation, difficult negative features are separated from negative features. The steps include: dividing the feature vector of the human-object interaction instance into border matching features and border non-matching features; for the border matching features, according to the multi-hot form human-object interaction label corresponding to the border it matches, dividing the label category and non-label category corresponding to this feature, and then for the non-label category, taking the category with the human-object interaction classification score vector corresponding to the comparison feature greater than the first threshold as the difficult negative feature; for the border non-matching features, taking the category with the human-object interaction classification score vector corresponding to the comparison feature greater than the set value as the difficult negative feature.
[0011] In an implementation of the above technical solution, the first threshold is a dynamic threshold, and the dynamic threshold is the minimum positive sample score. The calculation steps of the minimum positive sample score include: for the most accurate prediction Y * Among them, the human-object interaction classification score vector i n On the correct category label The scores are statistically analyzed, and the smallest of all the above scores in Y* is determined as the minimum positive sample score.
[0012] In an implementation of the above technical solution, dividing the feature vector of the human-object interaction instance into border matching features and border non-matching features is achieved through matching. The matching steps include: for the human-object interaction label with an object, matching the true label And all the prediction results of this picture One by one, calculate the intersection over union of the label human body border And the predicted human body side cabinet Calculate the intersection over union of the label object border And the predicted object border If both intersection over unions are greater than the second threshold, and the label object category And the predicted object category o n Are the same, then it is considered that the instance feature corresponding to this prediction result is a border matching feature, otherwise it is a border non-matching feature; for the human-object interaction label without an object, match the true label And all the prediction results of this picture One by one, calculate the label human body border And the predicted human body side cabinet The Intersection over Union. If the Intersection over Union is greater than the third threshold, it is considered that the instance feature corresponding to the prediction result is a bounding box matching feature, otherwise it is a bounding box non-matching feature.
[0013] In one implementation of the above technical solution, the second threshold is 0.5 and the third threshold is 0.5.
[0014] In one implementation of the above technical solution, the backbone network is a ResNet-50 network pre-trained on the COCO dataset and fine-tuned on the HICO-DET dataset.
[0015] In one implementation of the above technical solution, the types of hard negative features include: interaction error, action error, and matching error.
[0016] In a second aspect, the present disclosure proposes a classifier recalibration system for a Transformer architecture human-object interaction detection model. The human-object interaction detection model includes a backbone network, a Transformer encoder-decoder, and a learnable human-object interaction classifier. The system includes a model training module, a feature partitioning module, and an optimization calibration module. Among them: The model training module is configured to train the human-object interaction detection model using a training set of images to be recognized, and freeze the parameters of the trained model. The feature partitioning module is configured to extract human-object interaction instance features in the images to be recognized in the training set using the backbone network and the Transformer encoder-decoder with frozen parameters, divide the instance features into positive and negative features, and further divide hard negative features from the negative features. The optimization calibration module is configured to use the weight of the human-object interaction classifier in class c as the class prototype p c , use the form of prototype supervised contrast learning to calculate the loss function for each class c, and perform backpropagation independently for each class to optimize and calibrate the class prototype p c The loss function L of the prototype supervised contrast learning c is calculated as follows:
[0017]
[0018] In the formula: is the i-th positive feature of the human-object interaction detection model in class c, is the i-th hard negative feature of the human-object interaction detection model in class c, sim is the cosine similarity, τ is the contrast learning temperature coefficient, N + and N- are the numbers of positive and negative features in the entire training set.
[0019] In a third aspect, the present disclosure provides a computer-readable storage medium storing a computer program that can be loaded and executed by a processor to perform any of the above methods.
[0020] Compared with the prior art, the present disclosure has the following advantages: 1) The present disclosure innovatively uses prototype supervised contrastive learning to recalibrate the human-object interaction detection classifier, mine the mispredictions of existing models, and calibrate them specifically, improving the accuracy of human-object interaction detection, especially for the tail action categories with extremely small sample sizes; 2) The present disclosure designs an image adaptive hard negative feature mining technique that can accurately screen out the most valuable hard negative samples from a large number of negative samples for the Transformer architecture and the human-object interaction detection task; 3) The present disclosure innovatively applies the feature storage pool technology in unsupervised contrastive learning to supervised contrastive learning and combines it with hard negative feature mining, ensuring both the high quality of negative samples and the sufficiency of the number of samples in contrastive learning, and realizing the global recalibration of the classifier for the training set. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 and one The schematic workflow diagram in a certain implementation manner.
[0023] Figure 2 and one The schematic workflow diagram of the hard negative feature mining stage of steps S1-S4 in a certain implementation manner.
[0024] Figure 3 and one The schematic workflow diagram of the global prototype contrastive learning stage of step S5 in a certain implementation manner. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The current Human-Object Interaction (HOI) detection datasets have a significant long-tailed distribution problem, which leads to long-tailed bias in data-driven deep learning classifiers, causing the model to assign higher scores to the head classes. To address this issue, existing solutions can be roughly divided into data-oriented and model-oriented methods. Data-oriented methods focus on dataset enhancement and expansion, while model-oriented methods enhance the generalization ability of the model by leveraging multi-modal pre-trained models. These methods have significantly improved the detection performance of rare HOI classes by introducing more data or more powerful models. However, these methods fail to specifically address the mispredictions caused by long-tailed bias. On the other hand, due to memory limitations, these methods cannot eliminate long-tailed bias globally in the training set.
[0026] Based on this, this case proposes a classifier recalibration method for the Transformer architecture-based Human-Object Interaction detection model. By recalibrating the Human-Object Interaction detection classifier, the long-tailed bias of the model is alleviated, and the recognition accuracy of the model is improved. The Human-Object Interaction detection model includes a backbone network, a Transformer encoder-decoder, and a learnable Human-Object Interaction classifier. The steps of this method are as follows:
[0027] Use the training set of images to be recognized to train the Human-Object Interaction detection model, and freeze the parameters of the trained model;
[0028] For all images to be recognized in the training set, use the backbone network and the Transformer encoder-decoder with frozen parameters to extract the Human-Object Interaction instance features in the images to be recognized. Divide the instance features into positive and negative features, and further divide the difficult negative features from the negative features;
[0029] Take the weight of the Human-Object Interaction classifier on class c as the class prototype p c , and in the form of prototype supervised contrast learning, calculate the loss function for each class c, and perform backpropagation independently by class to optimize and calibrate the class prototype p c ;
[0030] The loss function L of the prototype supervised contrast learning c is calculated as follows:
[0031]
[0032] In the formula: is the i-th positive feature of the Human-Object Interaction detection model on class c, is the i-th difficult negative feature of the Human-Object Interaction detection model on class c, sim is the cosine similarity, τ is the contrast learning temperature coefficient, N + and N- are the numbers of positive and negative features in the entire training set.
[0033] Next, in conjunction with the accompanying drawings, how the technical solution of this case is implemented will be clearly and completely described.
[0034] (I) Classifier recalibration method
[0035] First, use the training set of images to be recognized to train the human-object interaction detection model, and freeze the parameters of the trained model.
[0036] Use any image backbone network to extract global image features, and an encoder-decoder based on the detection Transformer architecture with any structure to perform set prediction on human-object interaction instances in the image. In this embodiment, three influential cash models, namely GEN-VLKT, H0ICLIP, and DP-HOI, are selected as the base models for experiments, where N q = 64, D q = 512.
[0037] In one implementation, the backbone network is a ResNet-50 network pre-trained on the COCO dataset and fine-tuned on the HICO-DET dataset.
[0038] Before calibrating the classifier, use the training set of images to be recognized to train the human-object interaction detection model, and freeze the parameters of the trained model.
[0039] Next, refer to Figure 1 , and perform recalibration on the classifier of the Transformer architecture human-object interaction detection model, including the following steps:
[0040] S1. Through the backbone network and the Transformer encoder-decoder, extract a set of human-object interaction instance features in the image to be recognized.
[0041] The backbone network and the Transformer encoder-decoder used in this step are the backbone network and the Transformer encoder-decoder with frozen parameters. The processed image is an image in the training set of images to be recognized for training the human-object interaction detection model.
[0042] Specifically, use the backbone network with frozen parameters to extract global image features; use an encoder-decoder based on the detection Transformer architecture with any structure to perform set prediction on human-object interaction instances in the image, and obtain a set of human-object interaction instance feature vectors Q hoi , whose dimension is N q ×D q , where N q and D q respectively represent the number of query vectors and the feature dimension.
[0043] Human-object interaction instance feature vector Q hoi Generate human-object pair prediction results through a set of feedforward neural networks Including the human body bounding box Object bounding box Unique object category o n And the human-object interaction classification score vector i n .
[0044] S2. Through the human-object interaction matching algorithm based on Hungarian matching, classify the instance features into positive and negative features, and store the positive features.
[0045] Specifically, use the human-object interaction matching algorithm based on Hungarian matching to match the ground truth and the prediction result Y one by one, where is the human body bounding box label, is the object bounding box label, is the unique object label, is the human-object interaction label in multi-hot form, and N gt represents the number of human-object interaction instances in the image. Calculate the total matching cost, and weighted select the prediction result with the lowest total matching cost as the most accurate prediction corresponding to the ground truth, denoted as
[0046] Take the human-object interaction instance feature vector corresponding to the most accurate prediction Y* as the positive feature and store it in the positive feature storage pool corresponding to the human-object interaction label category included.
[0047] Take the human-object interaction classification score vector i * in the most accurate prediction Y n and count the scores on the correct class label . Take the smallest of all the above scores in Y* and call it the "minimum positive sample score" of this picture, and store this score for future use.
[0048] S3. By adopting a picture-adaptive hard negative feature mining mechanism, separate the hard negative features from the negative features and store them.
[0049] Specifically, according to the different human-object interaction categories, adopt the following two matching methods to divide the human-object interaction instance feature vector Q hoi into the bounding box matching feature Q mat and the bounding box non-matching feature Q unm .
[0050] For human-object interaction labels with objects such as "eating pizza" and "riding a horse", the matching method is: match the ground truth All prediction results of this image Match one by one, and calculate the labeled human body bounding box and the predicted human body bounding box to calculate the Intersection over Union (IoU), calculate the labeled object bounding box and the predicted object bounding box of the IoU. If both IoUs are greater than 0.5, and the labeled object category and the predicted object category o n are the same, then it is considered that the instance feature corresponding to this prediction result is the bounding box matching feature Q mat , otherwise it is the bounding box mismatch feature Q unm .
[0051] For the special person-object interaction labels such as "sitting" and "standing" without objects in the V-COCO dataset, the matching method is as follows: Match the ground truth label and all prediction results of this image one by one, and calculate the labeled human body bounding box and the predicted human body bounding box to calculate the Intersection over Union (IoU). If the IoU is greater than 0.5, then it is considered that the instance feature corresponding to this prediction result is the bounding box matching feature Q mat , otherwise it is the bounding box mismatch feature Q unm .
[0052] For the bounding box matching feature Q mat , first divide the label category and non-label category corresponding to this feature according to the multi-hot form person-object interaction label corresponding to the bounding box it matches; in the label category, these features are regarded as "redundant positive features" and ignored; while in the non-label category, by comparing the person-object interaction classification score vector i n corresponding to the feature, whether the scores in these categories are greater than the "minimum positive sample score" to determine whether this feature is a "difficult negative feature" in these categories; the "difficult negative features" are stored in the difficult negative feature storage pool corresponding to these categories, and the remaining simple negative features are ignored. Examples of the difficult negative features in these bounding box matching features are Figure 2 shown as "interaction error". The feature of this "eating pizza" instance is a difficult negative feature relative to the "cutting pizza" category; while for the bounding box matching features in the special interaction categories without objects, its examples are Figure 2 shown as "action error" in. The feature of this "eating pizza" instance is a difficult negative feature relative to the "sitting" category.
[0053] For the bounding box mismatch feature Qunm , only compare the human-object interaction classification score vectors i corresponding to the features n , whether the scores on all categories are greater than the "minimum positive sample score" to determine whether the feature is a "difficult negative feature" on these categories; the "difficult negative features" are stored in the difficult negative feature storage pools corresponding to these categories, while the remaining simple negative features are ignored. Examples of these difficult negative features are Figure 2 shown as "matching error" in
[0054] For the judgment of "difficult negative features", this case uses a method for dividing difficult negative samples adapted to pictures. Specifically, the basis for dividing difficult negative samples and simple negative samples is a classification score threshold adapted to pictures, that is, the minimum positive sample score in this picture. Through such threshold setting, it is possible to avoid the problem that due to the high clustering of instance features in each picture, the average classification score of these features fluctuates with different pictures, making it difficult to determine the global threshold for dividing difficult negative samples.
[0055] S4. Repeat the above three processes until all positive features and difficult negative features of all images in the training set are stored in the feature storage pool by category.
[0056] Use the frozen model to perform human-object interaction detection on the entire training set according to any batch size, and the positive features and difficult negative features of each category in the global training set can be obtained. Among them, the former represents the features of the instances that are indeed of this category in the picture, while the latter represents the features of the instances that are not of this category but are misclassified as this category in the picture.
[0057] S5. Use the instance features in the feature storage pool to re-calibrate the weights of the human-object interaction classifier independently for each category in the form of prototype supervised contrast learning.
[0058] See Figure 3 , first share the parameters of the learnable human-object interaction detection classifier weights with the set of "category prototypes". Specifically, the weight vector of the classifier for each category is the category prototype p of this category c .
[0059] Then, find all positive features of a certain category from the feature storage pool to form positive sample pairs with the category prototype of this category; find all difficult negative features of a certain category from the feature storage pool to form negative sample pairs with the category prototype of this category.
[0060] Specifically, due to the unique characteristics of the HOI task, we merge the instance features of HOI categories with the same verb but different object labels as the instance features of each HOI category. This is based on our observation that HOI categories with the same verb have similar features; in this embodiment, to ensure that the video memory overhead during the calculation process is not too large, we only take this optimization measure for the V-COCO dataset.
[0061] Next, calculate the cosine similarity of positive and negative sample pairs, and calculate the loss function of each category according to the loss function form of prototype supervised contrast learning:
[0062]
[0063] where sim is the cosine similarity, τ is the contrast learning temperature coefficient, N + and N- are the numbers of positive and negative features in the entire training set; in this embodiment, τ = 1.0 is set.
[0064] Then, for each category, calculate this loss function and perform backpropagation to optimize the category prototype p c and then calibrate the classifier.
[0065] Different from traditional prototype supervised contrast learning, the prototype used in step S5 shares weights with the human-object interaction detection classifier; can use the global frozen features stored in the feature storage pool; can explicitly target difficult negative features; and can achieve category-independent backpropagation. The above features enable step S5 to only recalibrate the classifier of the model, ensuring the efficiency of the method.
[0066] (2) Experimental verification
[0067] To verify the actual performance of the present invention in the human-object interaction detection task, experiments were conducted on two commonly used public datasets, HICO-DET and V-COCO. Referring to previous work, we used the mean average precision (mAP) as the evaluation metric. Specifically, when the intersection over union (IoU) of the predicted human and object bounding boxes with their respective ground truth boxes is greater than 0.5, and the predicted action and object categories are correct, the detection result is regarded as a true positive. For HICO-DET, we evaluated the performance in two scenarios: the default scenario considered all categories, and the known object scenario only tested the interactions with objects present in the training set; it should be noted that in this dataset, categories with fewer than 10 training instances among 138 training instances were classified as "rare classes" (rarities). Therefore, we reported the mAP scores for all categories, rare categories, and non-rare categories. For V-COCO, we evaluated the performance in two scenarios: in scenario 1 (S1), the model was required to report the situation without objects, while in scenario 2 (S2), such situations were ignored.
[0068] Comparison of mAP (%) on the HICO-DET test set
[0069]
[0070]
[0071] Table 1 above shows the comparison of the present invention with other methods in terms of the mAP metric on the HICO-DET test set. Among them, "*" represents the known object scenario; the gray background represents the base model cited in the three embodiments of the present invention; "+" represents the result after the improvement of the present invention on this base model; bold and underlined represent the best and sub-best performances. It can be seen from Table 1 that the present invention has achieved obvious performance improvement on the influential advanced models in the past three years, reaching the best performance on the models of the same scale in these three years. Taking the embodiment based on the latest DP HOI model as the base model, its average mAP has increased by 1.62 in all categories and by 3.01 in rare categories, which proves the effectiveness of the present invention and its ability to remove long-tail bias.
[0072] Table 2 shows the comparison of the present invention with other methods in terms of the mAP metric on the V-COCO test set.
[0073] Table 2 Comparison of mAP (%) on the V-COCO test set
[0074]
[0075] According to Table 2, it can be seen that the effect of the present invention on the V-COCO dataset is not as obvious as that on HICO-DET. This is because the detection difficulty on this dataset is relatively small, so fewer difficult negative features are extracted, and the calibration effect on the classifier is limited. However, it still achieves a stable increase in performance.
[0076] Generally speaking, the present invention proposes a novel classifier recalibration method and can significantly improve the performance of the one-stage human-object interaction detection model based on Transformer, especially the performance on rare categories. The present invention is innovative, efficient, and compatible with other mainstream methods in the field of human-object interaction detection.
[0077] Through the description of the above embodiments, those skilled in the art can clearly understand that according to the method of the present disclosure, a classifier recalibration system for a human-object interaction detection model based on the Transformer architecture can be implemented. The Transformer-based human-object interaction detection model includes a backbone network, a Transformer encoder-decoder, and a learnable human-object interaction classifier. The system includes a model training module, a feature division module, and an optimization and calibration module. Among them: The model training module is configured to train the Transformer-based human-object interaction detection model using a training set of images to be recognized, and freeze the parameters of the trained model. The feature division module is configured to extract human-object interaction instance features in the images to be recognized in the training set using the backbone network and the Transformer encoder-decoder with frozen parameters, divide the instance features into positive and negative features, and further divide difficult negative features from the negative features. The optimization and calibration module is configured to use the weight of the human-object interaction classifier on class c as the class prototype p c , and in the form of prototype supervised contrast learning, calculate the loss function for each class c and perform backpropagation independently by class to optimize and calibrate the class prototype p c . The loss function L c of the prototype supervised contrast learning is calculated as follows:
[0078]
[0079] In the formula: is the i-th positive feature of the human-object interaction detection model on class c, is the i-th difficult negative feature of the human-object interaction detection model on class c, sim is the cosine similarity, T is the contrast learning temperature coefficient, and N+ and N- are the numbers of positive and negative features in the entire training set.
[0080] Through the description of the above embodiments, those skilled in the art can clearly understand that the method and system of the present disclosure can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, in more cases for the present disclosure, software program implementation is a better implementation method.
[0081] Although the embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, the present disclosure is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present disclosure, and all of these fall within the scope of protection of the present disclosure.
Claims
1. A classifier recalibration method for a human-object interaction detection model based on the Transformer architecture, the human-object interaction detection model comprising a backbone network, a Transformer encoder-decoder, and a learnable human-object interaction classifier, wherein, The method includes the following steps: Training a human-object interaction detection model using a training set of images to be recognized, and freezing the parameters of the trained model; For all the images to be recognized in the training set, using the backbone network and the Transformer encoder-decoder with frozen parameters to extract the human-object interaction instance features in the images to be recognized, classifying the instance features into positive and negative features, and further separating hard negative features from the negative features; Use the weight of the human-object interaction classifier on class c as the class prototype p c , in the form of prototype supervised contrastive learning, calculate the loss function for each class c, and perform backpropagation independently by class to optimize and calibrate the class prototype p c ; The loss function L of the prototype supervised contrastive learning c is calculated as follows: Wherein: is the i-th positive feature of the human-object interaction detection model on class c, is the i-th hard negative feature of the human-object interaction detection model on class c, sim is the cosine similarity, τ is the contrastive learning temperature coefficient, N + and N - are the numbers of positive and negative features on the entire training set.
2. The method according to claim 1, wherein The step of classifying the instance features into positive and negative features includes: Through the human-object interaction matching algorithm based on Hungarian matching, the real labels are matched one by one with the prediction results Y; Calculating the total matching cost, and weighted selecting the prediction result with the lowest total matching cost as the most accurate prediction for the corresponding ground truth label; The most accurate prediction of Y * The corresponding human-object interaction instance feature vector As positive features, the rest are negative features.
3. The method according to claim 2, characterized in that, By adopting a hard negative feature mining mechanism adapted to the picture, separating hard negative features from the negative features, the steps include: Dividing the human-object interaction instance feature vector into bounding box matching features and bounding box non-matching features; For the bounding box matching features, according to the multi-hot form human-object interaction label corresponding to the bounding box they match, dividing the label categories and non-label categories corresponding to the features, and then for the non-label categories, taking the categories with the human-object interaction classification score vector corresponding to the contrast features greater than the first threshold as hard negative features; For the bounding box non-matching features, taking the categories with the human-object interaction classification score vector corresponding to the contrast features greater than the set value as hard negative features.
4. The method according to claim 3, wherein The first threshold is a dynamic threshold, and the dynamic threshold is the minimum positive sample score. The calculation steps of the minimum positive sample score include: The one that most accurately predicts Y * Among them, the human-object interaction classification score vector i n On the correct class label The scores are statistically counted, and Y * The smallest of all the above scores in is determined as the minimum positive sample score.
5. The method according to claim 1, wherein The division of the human-object interaction instance feature vector into bounding box matching features and bounding box non-matching features is achieved through matching. The matching steps include: For the human-object interaction label with an object, match the true label with all the prediction results Y of this picture, and calculate the label human body bounding box and the predicted human body bounding box to calculate the Intersection over Union (IoU), calculate the label object bounding box and the predicted object bounding box to calculate the IoU. If both IoUs are greater than the second threshold, and the label object category is the same as the predicted object category o n then it is considered that the instance feature corresponding to this prediction result is the bounding box matching feature, otherwise it is the bounding box non-matching feature; For the person-object interaction label without an object, match the true label with all the prediction results Y of this picture, and calculate the label human body bounding box and the predicted human body bounding box of the intersection over union (IoU). If the IoU is greater than the third threshold, it is considered that the instance feature corresponding to the prediction result is a bounding box matching feature, otherwise it is a bounding box non-matching feature.
6. The method according to claim 5, characterized in that The second threshold is 0.5, and the third threshold is 0.
5.
7. The method according to claim 1, wherein The backbone network is a ResNet-50 network pre-trained on the COCO dataset and fine-tuned on the HICO-DET dataset.
8. The method according to claim 1, characterized in that, The types of the hard negative features include: interaction error, action error, and matching error.
9. A classifier recalibration system for a human-object interaction detection model based on the Transformer architecture, the human-object interaction detection model including a backbone network, a Transformer encoder-decoder, and a learnable human-object interaction classifier, characterized in that, The system includes a model training module, a feature division module, and an optimization and calibration module; wherein: The model training module is configured to train a human-object interaction detection model using a training set of images to be recognized, and freeze the parameters of the trained model; The feature division module is configured to, for all the images to be recognized in the training set, use the backbone network and the Transformer encoder-decoder with frozen parameters to extract the human-object interaction instance features in the images to be recognized, classify the instance features into positive and negative features, and further separate hard negative features from the negative features; An optimization and calibration module, configured to use the weight of the human-object interaction classifier on class c as the class prototype p c , in the form of prototype supervised contrastive learning, calculates the loss function for each class c and performs backpropagation independently by class to optimize and calibrate the class prototype p c ; the loss function L of the prototype supervised contrastive learning c is calculated as follows: Wherein: is the i-th positive feature of the human-object interaction detection model on class c, is the i-th hard negative feature of the human-object interaction detection model on class c, sim is the cosine similarity, T is the contrastive learning temperature coefficient, N + and N - are the numbers of positive and negative features on the entire training set.
10. A computer-readable storage medium, characterized in that: There is a computer program stored that can be loaded and executed by a processor to perform any one of the methods as claimed in claims 1 to 8.