Training and Authentication Methods and Devices for Object Detection Models

By using the score difference between the scores output by the teacher model for each anchor box and the score difference between the student model and the student model as a soft label, the student model is trained, and the problem of low detection accuracy of the student model in the application scenario of object detection is solved, achieving higher detection accuracy.

CN114067401BActive Publication Date: 2025-05-27SHANGHAI SENSETIME INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111364379.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-17
Publication Date
2025-05-27
Estimated Expiration
2041-11-17

AI Technical Summary

Technical Problem

The student model trained based on soft labels has low detection accuracy in object detection application scenarios.

Method used

The student model is trained by using the difference between the second score value output by the teacher model for each anchor box and the first score value output by the student model for each anchor box as a soft label, so that the student model can learn how the teacher model scores each anchor box.

Benefits of technology

Improves the accuracy of object detection in the student model, allowing it to more accurately detect target objects in the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067401B_ABST
    Figure CN114067401B_ABST
Patent Text Reader

Abstract

An embodiment of the present disclosure provides a method and apparatus for training and authenticating a target detection model. The difference between the second scores output by the teacher model for each anchor box and the first scores output by the student model for each anchor box is used as a soft label to train the student model, so that the student model can learn the way the teacher model scores each anchor box, thereby improving the target detection accuracy of the student model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular, to a method and apparatus for training a target detection model and identity verification. Background Art

[0002] Knowledge Distillation (KD) refers to using the information obtained by a trained teacher model as soft labels to train a student model. In the target detection application scenario, soft labels generally refer to the probabilities that the target objects in a sample image belong to various categories. However, the target detection accuracy of the student model trained based on the above soft labels is relatively low. Summary of the Invention

[0003] In a first aspect, an embodiment of the present disclosure provides a method for training a target detection model. The method includes: inputting a sample image set into a first target detection model and a trained second target detection model respectively, where the first target detection model is a student model and the second target detection model is a teacher model, and each sample image in the sample image set includes a sample object of one category; for a plurality of predetermined anchor boxes, obtaining a first score for each anchor box among the plurality of anchor boxes determined by the first target detection model for each sample image, and obtaining a second score for each anchor box determined by the second target detection model for each sample image, where the score of an anchor box is used to represent the confidence level of classifying the object in the anchor box as the true category of the sample object; training the first target detection model based on the difference between the first score and the second score of each anchor box.

[0004] In some embodiments, the obtaining a first score for each anchor box among the plurality of anchor boxes determined by the first target detection model for each sample image, and obtaining a second score for each anchor box determined by the second target detection model for each sample image includes: for each sample image, filtering out the anchor boxes corresponding to the background region of the sample image from the plurality of anchor boxes, where the background region is the region in the sample image other than the sample object; obtaining a first score for each of the filtered plurality of anchor boxes determined by the first target detection model for the sample image, and obtaining a second score for each of the filtered anchor boxes determined by the second target detection model for each sample image.

[0005] In some embodiments, training the first object detection model based on the difference between the first scores and the second scores for each anchor box includes: obtaining a first probability distribution of each first score and a second probability distribution of each second score; establishing a loss function based on the difference between the first probability distribution and the second probability distribution; and training the first object detection model based on the loss function.

[0006] In some embodiments, obtaining a first probability distribution of each first score and a second probability distribution of each second score includes: performing an exponential operation on each target score to obtain an exponential score; summing the exponential scores to obtain a total score; and determining a target probability distribution of the target score based on the quotient of each target score and the total score; where the target score is the first score and the target probability distribution is the first probability distribution; or the target score is the second score and the target probability distribution is the second probability distribution.

[0007] In some embodiments, the sample image set includes multiple subsets, and each sample image in each subset includes a sample object of one category; training the first object detection model based on the difference between the first scores and the second scores for each anchor box includes: determining a total difference corresponding to the subset based on the difference between the first scores and the second scores for each anchor box determined for each sample image in a subset; determining a total difference corresponding to the sample image set based on the total differences corresponding to each subset; and training the first object detection model based on the total difference corresponding to the sample image set.

[0008] In some embodiments, the sample image set includes at least one of the following sample images: a first sample image with a brightness value outside a preset brightness range; a second sample image including a target object with a size outside a preset size range; a third sample image including an occluded target object; a fourth sample image in which the difference between the pixel values of the target object included and the background area is less than a preset pixel difference.

[0009] In a second aspect, an embodiment of the present disclosure provides an identity verification method, and the method includes: obtaining a face image of a target object; performing face detection on the face image through a pre-trained face detection model to obtain a face detection result; the face detection model is trained based on the training method of the object detection model according to any embodiment of the present disclosure; segmenting a face region image from the face image based on the face detection result; performing face recognition on the face region image to obtain a face recognition result; and performing identity verification on the target object based on the face recognition result.

[0010] Thirdly, an embodiment of the present disclosure provides a training device for a target detection model. The device includes: an input module configured to input a sample image set into a first target detection model and a trained second target detection model respectively. The first target detection model is a student model, and the second target detection model is a teacher model. Each sample image in the sample image set includes a sample object of one category; an acquisition module configured to, for a plurality of pre-determined anchor boxes, acquire a first score of each of the plurality of anchor boxes determined by the first target detection model for each sample image, and acquire a second score of each of the anchor boxes determined by the second target detection model for each sample image. The score of an anchor box is used to represent the confidence level of classifying the object within the anchor box as the true category of the sample object; a training module configured to train the first target detection model based on the difference between the first score and the second score of each anchor box.

[0011] In some embodiments, the acquisition module is configured to: for each sample image, filter out the anchor boxes corresponding to the background region of the sample image from the plurality of anchor boxes, where the background region is the region other than the sample object in the sample image; acquire the first score of each of the filtered plurality of anchor boxes determined by the first target detection model for the sample image, and acquire the second score of each of the filtered anchor boxes determined by the second target detection model for each sample image.

[0012] In some embodiments, the training module is configured to: acquire a first probability distribution of each first score and a second probability distribution of each second score; establish a loss function based on the difference between the first probability distribution and the second probability distribution; and train the first target detection model based on the loss function.

[0013] In some embodiments, the training module is configured to: perform an exponential operation on each target score to obtain an exponential score; sum up the exponential scores to obtain a total score; and determine a target probability distribution of the target score based on the quotient of each target score and the total score; where the target score is the first score and the target probability distribution is the first probability distribution; or the target score is the second score and the target probability distribution is the second probability distribution.

[0014] In some embodiments, the sample image set includes a plurality of subsets, and the sample images in each subset include a sample object of one category; the training module is configured to: determine a total difference corresponding to the subset based on the difference between the first score and the second score of each anchor box determined for each sample image within a subset; determine a total difference corresponding to the sample image set based on the total differences corresponding to each subset; and train the first target detection model based on the total difference corresponding to the sample image set.

[0015] In some embodiments, the sample image set includes at least one of the following sample images: a first sample image with a luminance value outside a preset luminance range; a second sample image including a target object with a size outside a preset size range; a third sample image including an occluded target object; a fourth sample image in which the difference between the pixel values of the included target object and the background area is less than a preset pixel difference.

[0016] Fourthly, an embodiment of the present disclosure provides an authentication device, including: an image acquisition module for acquiring a face image of a target object; a face detection module for performing face detection on the face image through a pre-trained face detection model to obtain a face detection result, where the face detection model is trained by the training device of the object detection model according to any embodiment of the present disclosure; a segmentation module for segmenting a face region image from the face image based on the face detection result; a face recognition module for performing face recognition on the face region image to obtain a face recognition result; and an authentication module for authenticating the identity of the target object based on the face recognition result.

[0017] Fifthly, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method according to any embodiment.

[0018] Sixthly, an embodiment of the present disclosure provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method according to any embodiment.

[0019] The embodiment of the present disclosure uses the difference between the second score output by the teacher model for each anchor box and the first score output by the student model for each anchor box as a soft label to train the student model, so that the student model can learn the way the teacher model scores each anchor box, thereby improving the object detection accuracy of the student model.

[0020] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings herein are incorporated into the specification and form a part of the specification. These drawings show embodiments consistent with the present disclosure and are used together with the specification to explain the technical solutions of the present disclosure.

[0022] Figure 1 It is a comparison schematic diagram of the teacher model and the student model according to the embodiment of the present disclosure.

[0023] Figure 2 It is a flowchart of a method for training an object detection model according to an embodiment of the present disclosure.

[0024] Figure 3 It is a schematic diagram of an anchor box according to an embodiment of the present disclosure.

[0025] Figure 4A It is a schematic diagram of a soft label in the related art.

[0026] Figure 4B It is a schematic diagram of a soft label according to an embodiment of the present disclosure.

[0027] Figure 5A It is a schematic diagram of the detection result of a simple sample according to an embodiment of the present disclosure.

[0028] Figure 5B It is a schematic diagram of the detection result of a difficult sample according to an embodiment of the present disclosure.

[0029] Figure 6 It is a flowchart of a method for training an object detection model according to another embodiment of the present disclosure.

[0030] Figure 7 It is a schematic diagram of the feature difference and prediction difference according to an embodiment of the present disclosure.

[0031] Figure 8 It is a flowchart of an authentication method according to an embodiment of the present disclosure.

[0032] Figure 9 It is a block diagram of a training device for an object detection model according to an embodiment of the present disclosure.

[0033] Figure 10 It is a block diagram of a training device for an object detection model according to another embodiment of the present disclosure.

[0034] Figure 11 It is a block diagram of an authentication device according to an embodiment of the present disclosure.

[0035] Figure 12 It is a schematic structural diagram of a computer device according to an embodiment of the present disclosure. Detailed implementation manners

[0036] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0037] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. The singular forms "a", "the", and "said" used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. Additionally, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality.

[0038] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0039] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of this disclosure and to make the above-mentioned objects, features, and advantages of the embodiments of this disclosure more obvious and understandable, the technical solutions in the embodiments of this disclosure will be further described in detail below with reference to the accompanying drawings.

[0040] In some embodiments, the knowledge distillation method can be adopted to train the object detection model. That is, a trained object detection model is used as the teacher model, and another object detection model to be trained is used as the student model, and the information obtained by the teacher model is passed to the student model as soft labels, so that the student model is trained based on the passed soft labels. Among them, the scale of the teacher model is generally larger than that of the student model. For example, at least one of the number of network layers, the number of parameters, the quantization bit width, and the memory resources occupied during inference of the teacher model is greater than that of the student model. Correspondingly, the object detection accuracy of the teacher model is generally also greater than that of the student model.

[0041] A target detection model is generally a multi-task model, including an identification task and a regression task. Among them, the identification task is used to identify the category of the target object in the image. In the related art, multiple anchor boxes with preset sizes and preset positions are generally predefined, and the probability of the object in each anchor box belonging to a category is determined through an identification algorithm. For example, for a certain anchor box A, the probability of the object in the anchor box A belonging to each of the common multiple categories can be determined. Suppose the common multiple categories include "cat", "dog", and "elephant", then the probability of the object in the anchor box A belonging to the category "cat", the probability of belonging to the category "dog", and the probability of belonging to the category "elephant" can be determined respectively. The regression task is used to learn the offset parameters of each anchor box, and the offset parameters are used to represent the offset between the actual position of the anchor box and the preset position of the anchor box. Through the above two tasks, the position and category of the target object in the image can be determined. For example, the category corresponding to the maximum probability value among the probability values corresponding to all anchor boxes in each category can be determined as the category of the target object. Suppose the predefined anchor boxes include anchor box A1 and anchor box A2, where the probabilities of the objects in anchor box A1 belonging to "cat", "dog", and "elephant" are {0.23, 0.20, 0.10} respectively, and the probabilities of the objects in anchor box A2 belonging to "cat", "dog", and "elephant" are {0.45, 0.30, 0.02} respectively, then the category "cat" corresponding to the maximum probability value 0.45 can be determined as the category of the target object, and the actual position of anchor box A2 (preset position + offset parameter) can be determined as the position of the target object.

[0042] However, the target detection accuracy of the student model trained with the above category probabilities as soft labels is relatively low. To solve the above problem, the present disclosure makes a comparison between the teacher model and the student model. Refer to Figure 1 , which is a comparison schematic diagram of the teacher model and the student model. When performing target detection, both the teacher model and the student model will first extract features from the input image to obtain the features of the input image, and then perform classification and regression based on the features of the input image. Among them, the features used to perform the classification task are denoted as m1, and the features used to perform the regression task are denoted as m2. Different head information can be used to identify the features corresponding to different tasks. Based on the features m1 and m2, prediction results can be obtained, including the category prediction result and the prediction result of the offset parameter of the anchor box. Further, the optimal one is selected from the prediction results as the final output result.

[0043] In the above process, there are at least the following differences between the teacher model and the student model:

[0044] (1) Feature difference, that is, there is a difference between the features extracted by the teacher model and the features extracted by the student model.

[0045] (2) Prediction differences, that is, there may be differences in the bias parameters and categories output by the teacher model and the student model for the same anchor box.

[0046] (3) Result differences, that is, there may be differences in the categories and / or positions finally determined by the teacher model and the student model for the target object.

[0047] By studying the differences in the above aspects, the present disclosure finds that the prediction difference is an important reason for the result difference. That is, precisely because the categories and bias parameters determined by the teacher model and the student model for each anchor box may be different, the categories and / or positions finally detected by the student model are different from those of the teacher model.

[0048] Based on the above findings, the present disclosure provides a training method for an object detection model, as Figure 2 shown, the method includes:

[0049] Step 201: Input the sample image set into the first object detection model and the trained second object detection model respectively. The first object detection model is a student model, and the second object detection model is a teacher model. Each sample image in the sample image set includes a sample object of one category;

[0050] Step 202: For a plurality of pre-determined anchor boxes, obtain the first score of each anchor box in the plurality of anchor boxes determined by the first object detection model for each sample image, and obtain the second score of each anchor box determined by the second object detection model for each sample image. The score of an anchor box is used to represent the confidence level of classifying the object within the anchor box as the true category of the sample object;

[0051] Step 203: Train the first object detection model based on the difference between the first score and the second score of each anchor box.

[0052] The method of the embodiment of the present disclosure can be executed by a server. The first object detection model and the trained second object detection model can be deployed to the server first, and the first object detection model is trained by the server. The trained first object detection model can be deployed to a resource-constrained device such as a mobile terminal (such as a mobile phone, a tablet computer, etc.) or an edge device to run. Or, the method of the embodiment of the present disclosure can also be executed by a resource-constrained device such as a mobile terminal. The trained second object detection model can be deployed to the server, and the first object detection model can be deployed to the mobile terminal. Then, the server sends the information required for training (such as the above-mentioned second feature map and the second prediction result) to the mobile terminal, and then the mobile terminal trains the first object detection model.

[0053] In step 201, the sample image set may include one or more sample images. Each sample image may include one or more sample objects of at least one category. Among them, one sample image may include one or more sample objects. The sample objects in the same sample image may be of the same category or different categories, and the categories of the sample objects can be pre-annotated. The sample objects can be objects of any category. For example, living bodies such as people and animals, or non-living bodies such as tables, chairs, and buildings. The categories of the sample objects included in the sample image can be set according to the actual application scenario. For example, in the vehicle violation detection scenario, the images of various vehicles can be determined as sample images. In some embodiments, the sample image set includes multiple subsets, each subset includes at least one sample image, and the sample images in different subsets include sample objects of different categories. For example, the sample objects included in the sample images in subset C1 are vehicles, and the sample objects included in the sample images in subset C2 are people, and so on. In order to enable the trained object detection model to detect multiple objects, as many sample images including various different categories of sample objects as possible can be added to the sample image set.

[0054] Each sample image in the sample image set can be respectively input into the first object detection model and the second object detection model. Among them, the first object detection model is the model to be trained, that is, the student model in knowledge distillation. The second object detection model is a trained model, that is, the teacher model in knowledge distillation.

[0055] In step 202, multiple anchor boxes can be determined in advance. Figure 3 Shows the case where the number of pre-determined anchor boxes is 4. Among them, A1 to A4 each represent an anchor box. It can be seen that different anchor boxes can have different sizes. The number, size, and position of these anchor boxes are predefined. When training the object detection model, only the offset parameters of each anchor box need to be adjusted.

[0056] Different from the related art, in the embodiments of the present disclosure, instead of determining the probability that an object in an anchor box belongs to each category and making a local comparison of the probabilities of the same anchor box under each category, the confidence that an object in an anchor box belongs to the true category of the sample object is determined, and a global comparison of the confidences of each anchor box belonging to the true category is made. Among them, the confidence is represented by a score. Figure 4A and Figure 4B respectively show the local comparison method and the global comparison method. In Figure 4AIn the illustrated embodiment, the teacher model obtains the probability values corresponding to each category for the same anchor box. For example, the probabilities of anchor box A1 corresponding to "giraffe", "banana", and "toilet" are 0.28, 0.03, and 0.02 respectively, the probabilities of anchor box A2 corresponding to "giraffe", "banana", and "toilet" are 0.33, 0.08, and 0.01 respectively, and the probabilities of anchor box A3 corresponding to "giraffe", "banana", and "toilet" are 0.45, 0.03, and 0.12 respectively. Then, the probability values for each of the above categories can be used as soft labels and passed to the student model so that the student model can learn various information generated by the teacher model for anchor box A1, anchor box A2, and anchor box A3 respectively. However, in this way, the soft labels do not include the mutual relationship between the information of different anchor boxes.

[0057] In Figure 4B the illustrated embodiment, since the true category of the sample image is "giraffe", the teacher model does not need to pass the probabilities corresponding to the categories "banana" and "toilet" to the student model, but directly passes the probabilities (i.e., confidence levels) of each anchor box corresponding to the true category "giraffe" to the student model. Among them, the true category can be obtained by annotating the sample image. In this way, the student model can learn the distribution relationship of the confidence levels of each anchor box under the true category.

[0058] In some embodiments, in addition to the sample object, a sample image usually may also include some background regions. For example, Figure 4B in the illustrated sample image, in addition to the sample object of the giraffe, it also includes background regions such as grassland and trees. Among the pre-determined multiple anchor boxes, some of the objects in the anchor boxes may belong to the background regions, and these background regions are less helpful for detecting the target object. Therefore, for each sample image, the anchor boxes corresponding to the background regions of the sample image can also be filtered out from the multiple anchor boxes. Among them, other regions in the image except the region where the sample object is located can be determined as the background regions. If the position difference between an anchor box and the position of the true anchor box of the pre-calibrated target object is greater than the preset difference threshold, then it is determined that the anchor box is the anchor box corresponding to the background region. The position difference between two anchor boxes can be determined based on the intersection over union of the two anchor boxes. The smaller the intersection over union, the greater the position difference. For example, in the case where the intersection over union between an anchor box and the true anchor box is less than 0.5, it can be determined that the anchor box is the anchor box corresponding to the background region. Then, the first score of each anchor box in the filtered multiple anchor boxes determined by the first target detection model for the sample image can be obtained, and the second score of each of the filtered anchor boxes determined by the second target detection model for each sample image can be obtained. Each of the filtered anchor boxes is related to the sample object in the sample image and can be called the positive anchor box of the sample object.

[0059] Since the positions of the sample objects in different sample images may vary, the multiple filtered anchor boxes determined for each sample image may be different. For example, assume that the multiple pre-determined anchor boxes include anchor boxes A1 to A10. The filtered anchor boxes corresponding to sample image 1 may include anchor boxes A1 and A3, and the filtered anchor boxes corresponding to sample image 2 may include anchor boxes A3, A5, and A8. Thus, for sample image 1, the difference between the first score and the second score of anchor box A1 can be determined, and the difference between the first score and the second score of anchor box A3 can be determined. For sample image 2, the difference between the first score and the second score of anchor box A3 can be determined, the difference between the first score and the second score of anchor box A5 can be determined, and the difference between the first score and the second score of anchor box A8 can be determined.

[0060] In step 203, the difference between the first score and the second score of each anchor box can be obtained. Since the information output by the object detection model for each anchor box includes the probability of the anchor box under each of multiple categories, as Figure 4A shown. Therefore, only the probability corresponding to the true category of the sample object can be extracted from the probabilities under each category, as Figure 4B shown.

[0061] Then, the first object detection model can be trained based on the difference between the first score and the second score of each anchor box. Since the sample image set may include multiple subsets, and the sample images in different subsets include sample objects of different categories (each category of sample object is called an instance). Therefore, the total difference corresponding to the subset can be determined first based on the difference between the first score and the second score of each anchor box determined for each sample image within each subset, then the total difference corresponding to the sample image set can be determined based on the total differences corresponding to each subset, and then the first object detection model can be trained based on the total difference corresponding to the sample image set.

[0062] For example, the sample image set includes subset 1 and subset 2. Among them, the sample images in subset 1 include sample objects of the category "cat", and the sample images in subset 2 include sample objects of the category "dog". For each sample image in subset 1, the total difference between the first score and the second score of each anchor box corresponding to the category "cat" can be determined. For each sample image in subset 2, the total difference between the first score and the second score of each anchor box corresponding to the category "dog" can be determined. Then, the total difference corresponding to the sample image set is determined based on the total difference corresponding to subset 1 and the total difference corresponding to subset 2, and the first object detection model is trained based on the total difference corresponding to the sample image set.

[0063] In some embodiments, the first probability distribution of each first score and the second probability distribution of each second score can be obtained; a loss function is established based on the difference between the first probability distribution and the second probability distribution; and the first object detection model is trained based on the loss function.

[0064] Specifically, an exponential operation can be performed on each target score to obtain an exponential score; the exponential scores are summed to obtain a total score; and the target probability distribution of the target score is determined based on the quotient of each target score and the total score; where the target score is the first score and the target probability distribution is the first probability distribution; or the target score is the second score and the target probability distribution is the second probability distribution.

[0065] For a specific instance j, its N positive anchor boxes are respectively denoted as where i ∈ {1,, N}, and the first scores and second scores corresponding to these N positive anchor boxes under the target category (i.e., the true category of the sample object) are respectively denoted as and Then the first probability distribution and the second probability distribution can be respectively denoted as:

[0066]

[0067] The above first probability distribution is the normalized first score, and the second probability distribution is the normalized second score. The difference between the two probability distributions can be determined by the KL divergence between these two probability distributions, and the KL divergence can be directly determined as the loss function Loss RM , specifically as follows:

[0068]

[0069] where M is the total number of instances, that is, the total number of subsets. Those skilled in the art can understand that the above KL divergence is only an optional way to measure the difference between two probability distributions. In other embodiments, other ways can also be used to measure the difference between two probability distributions. For example, the L2 loss function can be used.

[0070] In some embodiments, in addition to the above loss function, the loss function may also include the loss function corresponding to the object detection task. For example, the L1 loss function or the intersection over union loss function can be used for the regression task, and the cross-entropy loss function can be used for the recognition task. The final loss function for training the first object detection model can be obtained by summing up each loss function.

[0071] By comparison, it is found that the difference between the output results of the student model and the teacher model for simple samples is small, while the difference between the output results for difficult samples is large. See Figure 5A and Figure 5B , it can be seen that for the Figure 5A simple samples, the anchor boxes of the target objects determined by the student model are consistent with those determined by the teacher model, and both match the true bounding box of the target object. For the Figure 5B simple samples, the anchor boxes of the target objects determined by the student model are inconsistent with those determined by the teacher model. Among them, the anchor boxes of the target objects determined by the teacher model match the true bounding box of the target object, while the anchor boxes of the target objects determined by the student model do not match the true bounding box of the target object. Therefore, in order to improve the training effect, difficult samples can be added to the sample image set. When the sample image set includes multiple subsets, each subset can include difficult samples.

[0072] In some embodiments, the difficult samples include first sample images with brightness values outside a preset brightness range. Generally speaking, when the brightness value of the sample image is within the preset range, the probability of accurately detecting the target object in the sample image is relatively high. Therefore, the first sample images with brightness values outside the preset brightness range can be determined as difficult samples. In some embodiments, the brightness value being outside the preset brightness range may be that the brightness value is less than a preset brightness threshold. In some embodiments, the relationship between the brightness value and the detection accuracy can be determined in advance, and the preset brightness threshold can be determined based on this relationship.

[0073] In some embodiments, the difficult samples include second sample images of target objects with sizes outside a preset size range. The size of the target object in the sample image also affects the detection accuracy. For example, when the size of the target object is too small, it is difficult to extract enough feature information to determine the category of the target object. Therefore, in some embodiments, the size being outside the preset size range may be that the size is less than a preset size threshold. For example, the ratio of the number of pixels occupied by the target object in the sample image to the total number of pixels in the sample image can be determined as the size of the target object in the sample image. If the ratio is less than a preset ratio threshold, it is determined that the size of the target object is less than the preset size threshold.

[0074] When the target object is occluded, it may lead to incorrect recognition results. Therefore, in some embodiments, the difficult samples include third sample images of occluded target objects. When the pixel values of the target object and the background area are relatively close, it may also lead to incorrect recognition results. Therefore, in some embodiments, the difficult samples include fourth sample images in which the difference between the pixel values of the target object and the background area is less than a preset pixel difference.

[0075] In addition to the above-listed cases, difficult samples may also include other types of sample images, which can be specifically determined according to the actual situation and will not be listed one by one here.

[0076] In some embodiments, after training the first object detection model, the trained first object detection model can also be used to perform object detection on the image to be processed. In the embodiments of the present disclosure, the difference between the second score output by the teacher model for each anchor box and the first score output by the student model for each anchor box is used as a soft label to train the student model, so that the student model can learn the way the teacher model scores each anchor box, thereby improving the object detection accuracy of the student model.

[0077] In some embodiments, when performing knowledge distillation, generally the student model is made to learn the feature extraction ability of the teacher model so that the student model and the teacher model can extract as many identical features as possible. This process can be achieved by obtaining the difference between the features extracted by the student model and the features extracted by the teacher model and establishing a loss function based on the difference between the features to train the student model. However, the present disclosure finds that making the two models extract the same features does not necessarily enable the two models to obtain the same object detection results. This is because not all the feature maps output by each network layer in the student model and the teacher model are used for detecting the target object, but only the feature maps with smaller sizes are used for detecting the target object. However, during the training process, even if the feature maps output by a network layer are not used for detecting the target object, a large cost is still incurred to train the parameters of these network layers.

[0078] See Figure 6 , assuming that the teacher model and the student model will output prediction results on the 4th, 5th, and 6th network layers, the difference between the prediction results on the 4th, 5th, and 6th network layers can be obtained based on the prediction results of the two models, and the difference between the features output by the two models on the 4th, 5th, and 6th network layers can be obtained based on the feature maps output by the two models. Among them, P stu and F stu respectively represent the prediction results and the feature maps output by the student model, and P tea and F tea respectively represent the prediction results and the feature maps output by the teacher model. Among them, the sizes of the feature maps and the prediction results are the same. For the feature maps and the prediction results of the same region (e.g., Figure 6 the circular region in), there may be a situation where the difference between the prediction results of this region does not match the difference between the features. For example, the difference between the features in this region is small but the difference between the prediction results is large, or the difference between the features in this region is large but the difference between the prediction results is small.

[0079] To solve the above problems, an embodiment of the present disclosure further provides a method for training a target detection model. Refer to Figure 7 , the method includes:

[0080] Step 701: Input the sample images into the first target detection model and the trained second target detection model respectively. The first target detection model is a student model, and the second target detection model is a teacher model;

[0081] Step 702: Obtain the feature difference between the first feature maps output by each network layer of the first target detection model and the second feature maps output by the corresponding network layer of the second target detection model;

[0082] Step 703: Obtain the prediction difference between the first prediction result maps output by each network layer of the first target detection model and the second prediction results output by the corresponding network layer of the second target detection model. The prediction result output by a network layer is used to represent the category of each point on the feature map output by the network layer;

[0083] Step 704: Perform weighted processing on the feature difference corresponding to each network layer based on the prediction difference corresponding to each network layer to obtain the weighted feature difference of each network layer;

[0084] Step 705: Train the first target detection model based on the weighted feature differences of each network layer.

[0085] In step 701, at least one sample image in the sample image set can be input into the first target detection model and the trained second target detection model respectively. The sample image set adopted in this embodiment and Figure 2 the sample image set in the method embodiment shown can be the same image set, or can also be different subsets of the same sample image set.

[0086] In step 702, for each sample image in the sample image set, each network layer in the multiple network layers of the first target detection model can output the first feature map of the sample image, and each network layer in the multiple network layers of the second target detection model can output the second feature map of the sample image. Among them, the feature maps output by different network layers can have different sizes. In some embodiments, the feature maps with sizes smaller than the preset size threshold can be used to detect target objects, and the feature maps with sizes greater than or equal to the preset size threshold can be only used to detect background regions.

[0087] In some embodiments, the feature difference of a network layer can be determined based on the sum of the feature differences of each channel on the network layer. That is, for each network layer, the feature difference of each channel between the first feature map output by the network layer of the first target detection model and the second feature map output by the network layer of the second target detection model can be obtained; the feature differences of each channel are summed to obtain the feature difference of the network layer.

[0088] Assume F stu represents the feature map output by a network layer of the first target detection model, and F tea respectively represent the feature maps output by the same network layer of the second target detection model, then the feature difference F dif can be denoted as:

[0089]

[0090] where Q is a positive integer representing the number of channels of the feature map output by the network layer. The number of channels of the feature maps output by different network layers can be the same or different.

[0091] In step 703, for each sample image in the sample image set, each network layer in the first target detection model can output the first prediction result of the sample image, and each network layer in the second target detection model can output the second prediction result of the sample image. The prediction result of a network layer can have the same size as the feature map of the network layer.

[0092] The prediction difference of a network layer can be determined based on the sum of the prediction differences of each category output by the network layer. That is, for each network layer, the first prediction results corresponding to each category output by the network layer of the first target detection model and the second prediction results corresponding to each category output by the network layer of the second target detection model can be obtained; the differences between the first prediction results and the second prediction results corresponding to each category are summed to obtain the prediction difference of the network layer.

[0093] Assume P stu represents the first prediction result output by a network layer of the first target detection model, and P tea represents the second prediction result output by the same network layer of the second target detection model, then the prediction difference P dif can be denoted as:

[0094]

[0095] where C represents the total number of categories.

[0096] In step 704, the feature differences corresponding to each network layer can be weighted based on the prediction differences corresponding to each network layer. In practical applications, when the size of the feature map output by a network layer exceeds a certain threshold, this network layer generally only detects the background area in the sample image. Since the detection difficulty of the background area is relatively low, generally, the prediction results of the teacher model and the student model for the background area are the same, that is, the prediction differences between the two are relatively small. For the target object, the prediction results of the two models may be quite different. Through the weighting process in this step, smaller weights can be assigned to the feature differences of the network layers that do not detect the target object, and larger weights can be assigned to the feature differences of the network layers that detect the target object, so that the final prediction result of the model is positively correlated with the feature differences. In this way, the target detection accuracy of the student model can be improved by learning the feature extraction ability of the teacher model.

[0097] In some embodiments, the weighted feature differences of each network layer can be obtained by multiplying and adding the prediction differences corresponding to each network layer and the feature differences corresponding to each network layer pixel by pixel. That is, for a network layer, the element in the u-th row and v-th column of the feature differences of this network layer is multiplied by the element in the u-th row and v-th column of the prediction differences of this network layer to obtain the multiplication result in the u-th row and v-th column. Then, the multiplication results corresponding to the elements in each row and each column are summed to obtain the weighted feature differences of this network layer.

[0098] In step 705, the weighted feature differences of each network layer can be summed to obtain the total feature difference between the first target detection model and the second target detection model; a loss function is established based on the total feature difference, and the first target detection model is trained based on the loss function. The loss function Loss PFI can be denoted as:

[0099]

[0100] where H l and W l respectively represent the height and width of the feature map output by the l-th network layer, represents the operation of multiplying and adding pixel by pixel.

[0101] In some embodiments, the sum of multiple loss functions can also be determined as the total loss function, and the first target detection model is trained using the total loss function. The total loss function Loss can be denoted as:

[0102] Loss = Loss task + αLoss RM + βLoss PFI

[0103] Among them, α and β represent weights used to limit each loss function within the same scale range. In some embodiments, α and β are 4 and 1.5 respectively. Loss task Loss represents the loss function adopted for the object detection task, and different loss functions can be adopted for the object detection task in different application scenarios. For example, the L1 loss function or the intersection over union loss function can be adopted for the regression task in the object detection task, and the cross-entropy loss function can be adopted for the recognition task in the object detection task. Loss RM can be obtained based on the difference between the first score and the second score of each anchor box.

[0104] In the embodiments of the present disclosure, by using the prediction difference between two models to weight the feature difference between the two models to obtain a weighted feature difference, and then training the first object detection model based on the weighted feature difference, the feature difference between the two models is positively correlated with the prediction difference. Therefore, the object detection accuracy of the first object detection model can be improved by enabling the first object detection model to learn the feature extraction ability of the second object detection model.

[0105] In some embodiments, the first object detection model that has been trained can also be used to perform object detection on the image to be processed.

[0106] See Figure 8 , the embodiments of the present disclosure also provide an authentication method, and the method includes:

[0107] Step 801: Obtain a face image of the target object;

[0108] Step 802: Perform face detection on the face image through a pre-trained face detection model to obtain a face detection result; the face detection model is trained based on the training method of the object detection model according to any embodiment of the present disclosure;

[0109] Step 803: Segment a face region image from the face image based on the face detection result;

[0110] Step 804: Perform face recognition on the face region image to obtain a face recognition result;

[0111] Step 805: Authenticate the identity of the target object based on the face recognition result.

[0112] The above method can be executed by resource-constrained devices such as mobile terminals or edge devices. The method of the embodiments of the present disclosure can be used in application scenarios such as mobile terminal account management, access control systems, bank account opening, and security. Taking the application scenario of mobile terminal account management as an example, the target object is the user of the mobile terminal. The user can collect their own face image through the camera on the mobile terminal, and the mobile terminal can be pre-deployed with a face detection model. The camera can input the collected face image into the face detection model to obtain a face detection result, which may include information such as the size and position of the face in the face image. Then, based on the above information, the face region image can be segmented from the face image, and the segmented face region image can be further input into a face recognition model for recognition. The face recognition model can be deployed on the mobile terminal or on the server. In the case where the face recognition model is deployed on the server, the mobile terminal can transmit the face region image to the server, and the server can input the face region image into the face recognition model for recognition to obtain a face recognition result. The face recognition result can be used to represent whether the target object is or is not a specific object, and the specific object can be the owner of the account to be logged in (for example, a bank account, a WeChat account, etc.). When performing identity verification, if the face recognition result represents that the target object is a specific object, it is determined that the identity verification is passed; otherwise, it is determined that the identity verification fails.

[0113] Since the face detection model adopted in the above identity verification method is trained by the training method of the target detection model described in any of the foregoing embodiments, the accuracy of target detection can be improved, and thus the accuracy of identity verification can be improved.

[0114] In addition to the above application scenarios, the first target detection model trained by the training method of the target detection model described in any of the foregoing embodiments can also be used in application scenarios such as subway security inspection. In this application scenario, an image of the item to be detected can be obtained through an X-ray machine, and the image of the item can be input into the trained first target detection model for detection to obtain a detection result, which may include the size and position information of the item in the image. Then, the item can be identified based on the detection result to obtain the category of the item. If the category of the item belongs to a specified category, an alarm message is output. Among them, the specified category can be the category corresponding to dangerous goods such as flammable and explosive, for example, controlled knives, guns, gasoline, etc.

[0115] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0116] See Figure 9, an embodiment of the present disclosure further provides a training device for a target detection model, the device includes:

[0117] An input module 901, configured to input a sample image set into a first target detection model and a trained second target detection model respectively, where the first target detection model is a student model, the second target detection model is a teacher model, and each sample image in the sample image set includes a sample object of one category;

[0118] An acquisition module 902, configured to, for a plurality of pre-determined anchor boxes, acquire a first score of each anchor box in the plurality of anchor boxes determined by the first target detection model for each sample image, and acquire a second score of each anchor box determined by the second target detection model for each sample image. The score of an anchor box is used to represent the confidence level of classifying the object in the anchor box as the true category of the sample object;

[0119] A training module 903, configured to train the first target detection model based on the difference between the first score and the second score of each anchor box.

[0120] See Figure 10 , an embodiment of the present disclosure further provides a training device for a target detection model, the device includes:

[0121] An input module 1001, configured to input a sample image into a first target detection model and a trained second target detection model respectively, where the first target detection model is a student model, and the second target detection model is a teacher model;

[0122] A first acquisition module 1002, configured to acquire the feature difference between the first feature map output by each network layer of the first target detection model and the second feature map output by the corresponding network layer of the second target detection model;

[0123] A second acquisition module 1003, configured to acquire the prediction difference between the first prediction result map output by each network layer of the first target detection model and the second prediction result output by the corresponding network layer of the second target detection model. The prediction result output by a network layer is used to represent the category of each point on the feature map output by the network layer;

[0124] A weighting module 1004, configured to perform a weighting process on the feature difference corresponding to each network layer based on the prediction difference corresponding to each network layer to obtain the weighted feature difference of each network layer;

[0125] A training module 1005, configured to train the first target detection model based on the weighted feature difference of each network layer.

[0126] See Figure 11, embodiments of the present disclosure further provide an identity authentication device, the device includes:

[0127] A third acquisition module 1101, configured to acquire a face image of a target object;

[0128] A face detection module 1102, configured to perform face detection on the face image through a pre-trained face detection model to obtain a face detection result; the face detection model is trained by a training device of the object detection model according to any embodiment of the present disclosure;

[0129] A segmentation module 1103, configured to segment a face area from the face image based on the face detection result;

[0130] A face recognition module 1104, configured to perform face recognition on the face area to obtain a face recognition result;

[0131] An identity authentication module 1105, configured to perform identity authentication on the target object based on the face recognition result.

[0132] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0133] Embodiments of this specification further provide a computer device, which at least includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method described in any of the foregoing embodiments.

[0134] Figure 12 Fig. shows a more specific schematic diagram of the hardware structure of a computing device provided by the embodiments of this specification. The device may include: a processor 1201, a memory 1202, an input / output interface 1203, a communication interface 1204, and a bus 1205. Wherein, the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204 are communicatively connected to each other inside the device through the bus 1205.

[0135] The processor 1201 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification. The processor 1201 may further include a graphics card, and the graphics card may be an Nvidia titan X graphics card or a 1080Ti graphics card, etc.

[0136] The memory 1202 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1202 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1202 and are called and executed by the processor 1201.

[0137] The input / output interface 1203 is used to connect to an input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0138] The communication interface 1204 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).

[0139] The bus 1205 includes a path for transmitting information between various components of the device (such as the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204).

[0140] It should be noted that although the above device only shows the processor 1201, the memory 1202, the input / output interface 1203, the communication interface 1204, and the bus 1205, in the specific implementation process, the device may further include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and do not necessarily include all the components shown in the figure.

[0141] Embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the methods described in any of the foregoing embodiments are implemented.

[0142] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVDs) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0143] From the description of the above embodiments, those skilled in the art can clearly understand that the embodiments of this specification can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the embodiments of this specification, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of this specification.

[0144] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.

[0145] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description in the method embodiments. The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated. When implementing the solutions of the embodiments of this specification, the functions of the modules can be realized in one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solutions of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0146] The above is only the specific implementation manner of the embodiments of this specification. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the embodiments of this specification, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of this specification.

Claims

1. A training method for a target detection model, characterized in that, the method comprises: Inputting a sample image set into a first target detection model and a trained second target detection model respectively. The first target detection model is a student model, and the second target detection model is a teacher model. Each sample image in the sample image set includes a sample object of one category; For a plurality of pre-determined anchor boxes, obtaining a first score for each anchor box among the plurality of anchor boxes determined by the first target detection model for each sample image, and obtaining a second score for each anchor box determined by the second target detection model for each sample image. The score of an anchor box is used to characterize the confidence level of classifying the object within the anchor box as the true category of the sample object; Training the first target detection model based on the difference between the first score and the second score of each anchor box.

2. The method according to claim 1, characterized in that, the obtaining a first score for each anchor box among the plurality of anchor boxes determined by the first target detection model for each sample image, and obtaining a second score for each anchor box determined by the second target detection model for each sample image, includes: For each sample image, filtering out the anchor boxes corresponding to the background region of the sample image from the plurality of anchor boxes. The background region is the region in the sample image other than the sample object; Obtaining a first score for each of the filtered plurality of anchor boxes determined by the first target detection model for the sample image, and obtaining a second score for each of the filtered anchor boxes determined by the second target detection model for each sample image.

3. The method according to claim 1, characterized in that, the training the first target detection model based on the difference between the first score and the second score of each anchor box, includes: Obtaining a first probability distribution of each first score and a second probability distribution of each second score; Establishing a loss function based on the difference between the first probability distribution and the second probability distribution; Training the first target detection model based on the loss function.

4. The method according to claim 3, characterized in that, the obtaining a first probability distribution of each first score and a second probability distribution of each second score, includes: Performing an exponential operation on each target score to obtain an exponential score; Summing up the exponential scores to obtain a total score; Determining a target probability distribution of the target score based on the quotient of each target score and the total score; wherein, the target score is the first score, and the target probability distribution is the first probability distribution; or the target score is the second score, and the target probability distribution is the second probability distribution.

5. The method according to claim 1, characterized in that, the sample image set includes a plurality of subsets, and the sample images in each subset include a sample object of one category; the training the first target detection model based on the difference between the first score and the second score of each anchor box, includes: Determine the total difference corresponding to the subset based on the difference between the first score and the second score of each anchor box determined for each sample image within a subset; Determine the total difference corresponding to the sample image set based on the total differences corresponding to each subset; Train the first object detection model based on the total difference corresponding to the sample image set.

6. The method according to any one of claims 1 to 5, characterized in that the sample image set includes at least one of the following sample images: a first sample image with a brightness value outside a preset brightness range; a second sample image including a target object with a size outside a preset size range; a third sample image including an occluded target object; a fourth sample image in which the difference between the pixel values of the target object included and the background area is less than a preset pixel difference.

7. An authentication method, characterized in that the method includes: obtain a face image of a target object; perform face detection on the face image through a pre-trained face detection model to obtain a face detection result; the face detection model is trained based on the method according to any one of claims 1 to 6; segment a face region image from the face image based on the face detection result; perform face recognition on the face region image to obtain a face recognition result; authenticate the target object based on the face recognition result.

8. A training device for an object detection model, characterized in that the device includes: an input module for inputting a sample image set into a first object detection model and a trained second object detection model respectively, the first object detection model being a student model, the second object detection model being a teacher model, and each sample image in the sample image set including a sample object of one category; an acquisition module for, for a plurality of pre-determined anchor boxes, acquiring a first score of each anchor box in the plurality of anchor boxes determined by the first object detection model for each sample image, and acquiring a second score of each anchor box determined by the second object detection model for each sample image, the score of an anchor box being used to represent the confidence level of classifying the object within the anchor box as the true category of the sample object; a training module for training the first object detection model based on the difference between the first score and the second score of each anchor box.

9. An authentication device, characterized in that the device includes: a third acquisition module for acquiring a face image of a target object; a face detection module for performing face detection on the face image through a pre-trained face detection model to obtain a face detection result; the face detection model is trained based on the device according to claim 8; a segmentation module for segmenting a face region image from the face image based on the face detection result; a face recognition module for performing face recognition on the face region image to obtain a face recognition result; an authentication module for authenticating the target object based on the face recognition result.

10. A computer-readable storage medium, on which a computer program is stored, characterized in that When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 7.

11. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein: when the processor executes the computer program, it implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Binocular image-based model training method and device and data processing equipment

    CN112396073A

  • Image processing method and electronic device

    IN201914005103A