Difficulty sample mining method and device, terminal equipment and storage medium
By calculating the prediction differences between different detection heads in the object detection model, the difficulty value of the sample is determined, so as to accurately dig out the difficult samples, solving the problem of mining high-quality samples from unlabeled sample images, and improving the performance of the object detection model.
Patent Information
- Application Number
- CN202311650376.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2025-06-13
AI Technical Summary
How to accurately mine difficult samples from large numbers of unlabeled sample images to improve the performance of the object detection model and save manual labeling costs.
By obtaining the current sample image of multiple frames and inputting each frame of image into the constructed object detection model for processing, the predicted difference is calculated using the object detection results output by at least two detection heads, determining the difficulty value of the sample, and finally selecting the difficult sample.
It realizes accurate mining of difficult samples from a large number of unlabeled sample images, improves the performance of the object detection model, and saves manual labeling costs.
Smart Images

Figure CN120147767A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of object detection, and in particular, to a method, device, terminal device, and storage medium for mining difficult samples. Background Art
[0002] With the development of camera technology, the acquisition cost of image data is getting lower and lower, which makes it easier to obtain a large number of unlabeled sample images. However, this also poses a huge challenge to manually screening and annotating sample images. How to mine high-quality samples from a large number of unlabeled sample images that can improve the object detection model plays a crucial role in saving the manual annotation cost and improving the model performance. Generally speaking, the above high-quality samples are sample images that are difficult for the object detection model to predict, so they can be called difficult samples. Thus, how to accurately mine difficult samples from a large number of unlabeled sample images has become a technical problem that those skilled in the art need to consider. Summary of the Invention
[0003] In view of this, the embodiments of this application provide a method, device, terminal device, and storage medium for mining difficult samples, which can accurately mine difficult samples from a large number of unlabeled sample images.
[0004] The first aspect of the embodiments of this application provides a method for mining difficult samples, including:
[0005] Obtain multiple frames of current sample images;
[0006] For each frame of the current sample image, input the current sample image into the constructed object detection model for processing, and through at least two detection heads of the object detection model, respectively output at least two object detection results of the current sample image; calculate at least two prediction differences of the at least two detection heads for the current sample image according to the at least two object detection results; determine the difficulty value of the current sample image according to the prediction differences;
[0007] Select difficult samples from multiple frames of current sample images according to the difficulty value of each frame of the current sample image.
[0008] In the embodiments of the present application, first, multiple frames of current sample images are obtained, and then each frame of the current sample image is input into the constructed target detection model for processing; since the target detection model has at least two detection heads, and each detection head can output a target detection result, at least two target detection results can be obtained for each frame of the current sample image; thereafter, according to the at least two target detection results corresponding to each frame of the current sample image, the prediction differences of different detection heads for each frame of the current sample image can be calculated respectively, and the difficulty value of each frame of the current sample image can be determined according to the prediction differences; finally, according to the difficulty value of each frame of the current sample image, difficult samples can be selected from the multiple frames of current sample images. For example, a part of the current sample images with the highest difficulty value can be selected as difficult samples. The above process can accurately evaluate the difficulty of the target detection model in predicting the current sample image by using the prediction differences of different detection heads for the current sample image, and finally accurately mine difficult samples from a large number of unlabeled sample images. For example, if the prediction difference is larger, it means that the prediction results of different detection heads of the target detection model for the current sample image are more different, that is, it can be determined that the difficulty of the target detection model in predicting the current sample image is also greater.
[0009] In one implementation manner of the embodiments of the present application, the at least two detection heads include a first detection head and a second detection head, and the at least two target detection results include a first target detection result output by the first detection head and a second target detection result output by the second detection head; calculating the prediction differences of the at least two detection heads for the current sample image according to the at least two target detection results includes:
[0010] Calculating a regression difference according to the difference between the positions of each target prediction box in the first target detection result and the positions of each target prediction box in the second target detection result;
[0011] Calculating a classification difference according to the difference between the predicted values of each target category in the first target detection result and the predicted values of each target category in the second target detection result;
[0012] Calculating the prediction differences of the first detection head and the second detection head for the current sample image according to the regression difference and the classification difference.
[0013] In one implementation manner of the embodiments of the present application, the prediction differences of the at least two detection heads for the current sample image include the prediction differences between every two detection heads among the at least two detection heads for the current sample image; determining the difficulty value of the current sample image according to the prediction differences includes:
[0014] Calculating the average value of the prediction differences between every two detection heads for the current sample image;
[0015] Determine the average value as the difficulty value of the current sample image.
[0016] In an implementation manner of the embodiment of the present application, after selecting difficult samples from multiple frames of current sample images according to the difficulty values of each frame of the current sample images, it further includes:
[0017] Merge the labeled difficult samples into the training dataset of the target detection model;
[0018] Use the merged training dataset to optimize the training of the target detection model;
[0019] If the performance of the target detection model after the optimized training does not meet the set requirements, return to execute the steps of obtaining multiple frames of current sample images and subsequent steps.
[0020] In an implementation manner of the embodiment of the present application, using the merged training dataset to optimize the training of the target detection model includes:
[0021] Calculate the difference loss according to the prediction difference;
[0022] Add the difference loss to the initial loss function of the target detection model to obtain an updated loss function;
[0023] Use the merged training dataset and based on the updated loss function to optimize the training of the target detection model.
[0024] In an implementation manner of the embodiment of the present application, the training dataset is obtained through the following method:
[0025] Extract a first number of initial sample images from the sample data pool as the initial dataset;
[0026] According to the sample categories included in each frame of the initial sample images in the initial dataset, determine the target sample category with the least number of corresponding initial sample images;
[0027] Extract a second number of initial sample images including the target sample category from the sample data pool and add them to the initial dataset;
[0028] If the number of initial sample images included in the initial dataset reaches the set threshold, determine the initial dataset as the training dataset, otherwise return to execute the step of determining the target sample category with the least number of corresponding initial sample images according to the sample categories included in each frame of the initial sample images in the initial dataset and subsequent steps.
[0029] In an implementation manner of the embodiment of the present application, after using the merged training dataset to optimize the training of the target detection model, it further includes:
[0030] If the performance of the optimized trained object detection model meets the set requirements, all the selected hard samples are merged as the output hard sample set.
[0031] In one implementation manner of the embodiment of the present application, hard samples are selected from multiple frames of current sample images according to the hardness value of each frame of the current sample image, including:
[0032] Select the third quantity of current sample images with the highest hardness value from multiple frames of current sample images as hard samples.
[0033] The second aspect of the embodiment of the present application provides a hard sample mining device, including:
[0034] A current sample acquisition module, configured to acquire multiple frames of current sample images;
[0035] A hardness determination module, configured to input each frame of the current sample image into the constructed object detection model for processing, and respectively output at least two object detection results of the current sample image through at least two detection heads of the object detection model; calculate at least two prediction differences of the detection heads for the current sample image according to the at least two object detection results; and determine the hardness value of the current sample image according to the prediction differences.
[0036] A hard sample selection module, configured to select hard samples from multiple frames of current sample images according to the hardness value of each frame of the current sample image.
[0037] The third aspect of the embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the hard sample mining method provided in the first aspect of the embodiment of the present application is implemented.
[0038] The fourth aspect of the embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the hard sample mining method provided in the first aspect of the embodiment of the present application is implemented.
[0039] The fifth aspect of the embodiment of the present application provides a computer program product, which when running on a terminal device, enables the terminal device to execute the hard sample mining method provided in the first aspect of the embodiment of the present application.
[0040] It can be understood that the beneficial effects of the above second aspect to the fifth aspect can refer to the relevant descriptions in the above first aspect, and will not be repeated here. Description of the Drawings
[0041] Figure 1 It is a flowchart of a method for mining difficult samples provided by an embodiment of the present application;
[0042] Figure 2 It is a schematic structural diagram of a model based on differential learning provided by an embodiment of the present application;
[0043] Figure 3 It is a schematic diagram of the training process of an object detection model provided by an embodiment of the present application;
[0044] Figure 4 It is a schematic structural diagram of a device for mining difficult samples provided by an embodiment of the present application;
[0045] Figure 5 It is a schematic diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0046] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application. Additionally, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0047] When the camera operates normally, a large number of unlabeled sample images will be generated, but not all sample images have a positive impact on subsequent model optimization. Therefore, how to accurately mine difficult samples from a large number of unlabeled sample images has become a technical problem that needs to be considered by those skilled in the art. In view of this, the embodiments of the present application provide a method, device, terminal device, and storage medium for mining difficult samples, which can accurately mine difficult samples from a large number of unlabeled sample images. For more specific technical implementation details of the embodiments of the present application, please refer to the method embodiments described below.
[0048] It should be understood that the execution subject of each method embodiment of the present application is various types of terminal devices or servers. For example, it can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a large-screen TV, and so on. The specific types of the terminal devices and servers are not limited in the embodiments of the present application.
[0049] Please refer to Figure 1 , which shows a method for mining difficult samples provided by an embodiment of the present application, including:
[0050] 101. Obtain multiple frames of current sample images;
[0051] First, obtain multiple frames of current sample images. These current sample images are unlabeled images. Through the method for mining difficult samples provided by the embodiment of the present application, difficult samples can be more accurately mined from these current sample images. In actual operation, image data in the application scenario can be obtained by using data acquisition software or the method of frame extraction at video intervals as the current sample images.
[0052] 102. For each frame of the current sample image, input the current sample image into the constructed target detection model for processing. Through at least two detection heads of the target detection model, at least two target detection results of the current sample image are respectively output; according to the at least two target detection results, calculate at least two prediction differences of the detection heads for the current sample image; according to the prediction differences, determine the difficulty value of the current sample image;
[0053] In the embodiment of the present application, a target detection model is pre-constructed, which has at least two detection heads. Each frame of the current sample image is input into the target detection model for processing. By using the target detection results output by different detection heads, the prediction differences of different detection heads for each frame of the current sample image can be calculated. Finally, the difficulty values of each frame of the current sample image are determined respectively by using the prediction differences. For example, assume that the target detection model has n detection heads. For the current sample image 1, it is input into the target detection model for processing, and each detection head can output a target detection result, that is, n target detection results are obtained. Using these n target detection results, the prediction differences of these n detection heads for the current sample image 1 can be calculated, and the difficulty value of the current sample image 1 is determined according to the prediction differences. And so on, the current sample image 2, the current sample image 3... can all determine their respective difficulty values in the same way as the current sample image 1, so as to determine the difficulty values of all current sample images.
[0054] As an example, the target detection model can be a difference learning model based on yolov5s as the basic framework. Two additional detection heads are added to learn the differences between the output results of each detection head, and on the basis of the original yolov5 loss function, a difference loss is added to optimize the model weights. As Figure 2 shown, it is a schematic diagram of a model structure based on difference learning provided by the embodiment of the present application, where H0 is the original detection head, and H1 and H2 are the two added detection heads.
[0055] The detection head H0 uses the original yolov5 loss function: L = L loc +L cls +L obj . As for the detection head H1 and the detection head H2, an improved loss function is adopted. Specifically, a difference loss is added on the basis of the original yolov5 loss function, and the expression is as follows:
[0056] L = L loc +L cls +L obj -L dis
[0057] Among them, L loc represents the localization loss, L cls represents the classification loss, L obj represents the object loss, and L dis represents the difference loss.
[0058] The localization loss L loc represents the error between the predicted box and the calibrated box, and is defined as follows:
[0059]
[0060] Among them, S is the size of the feature map, and B is the number of bounding boxes predicted by each cell. Indicates whether there is an object in the cell at the i-th row and j-th column. ^ represents the predicted value, and λ coord is a weight coefficient. x and y represent the center coordinates of the box, w represents the width of the box, and h represents the height of the box.
[0061] Classification loss L cls Is used to calculate whether the anchor box and the corresponding calibrated classification are correct, and is defined as follows:
[0062]
[0063] Among them, c is the class index. Indicates whether the actual class c exists in the cell at the i-th row and j-th column.
[0064] Object loss L obj Is used to calculate the confidence of the network, and is defined as follows:
[0065]
[0066] Among them, λ obj is a weight coefficient for the confidence task.
[0067] Difference loss L dis Is used to calculate the prediction difference between each prediction head, and is defined as follows:
[0068]
[0069] Among them, p represents the prediction result output by the detection head H0, and p 1 represents the prediction result output by the detection head H1, and p 2 represents the prediction result output by the detection head H2.
[0070] The prediction difference consists of a regression difference and a classification difference:
[0071] d(p i , p j ) = d 1 (p i , p j ) + d 2 (p i , p j )
[0072] Among them, d(p i , p j ) represents the prediction difference between the detection head i and the detection head j, and d 1 (p i , p j ) represents the regression difference, and d2 (p i ,p j ) represents the classification difference.
[0073] The regression difference is defined as follows:
[0074]
[0075] Where N represents the number of prediction boxes, and the regression difference can be used to represent the positional difference between the prediction boxes output by detection head i and the prediction boxes output by detection head j.
[0076] The classification difference is defined as follows:
[0077]
[0078] Where C represents the total number of categories, represents the predicted value of detection head i for the actual category c, represents the predicted value of detection head j for the actual category c, and the classification difference can be used to represent the category difference between the prediction results output by detection head i and the prediction results output by detection head j.
[0079] In one implementation of the embodiments of the present application, at least two detection heads include a first detection head and a second detection head, and at least two target detection results include a first target detection result output by the first detection head and a second target detection result output by the second detection head; according to the at least two target detection results, the prediction difference of the at least two detection heads for the current sample image is calculated, including:
[0080] (1) Calculate the regression difference according to the difference between the positions of the respective target prediction boxes in the first target detection result and the positions of the respective target prediction boxes in the second target detection result;
[0081] (2) Calculate the classification difference according to the difference between the predicted values of the respective target categories in the first target detection result and the predicted values of the respective target categories in the second target detection result;
[0082] (3) Calculate the prediction difference of the first detection head and the second detection head for the current sample image according to the regression difference and the classification difference.
[0083] As described above, the prediction difference consists of a regression difference and a classification difference. Therefore, when calculating the prediction difference of the current sample image by different detection heads based on the object detection results of different detection heads, the regression difference and the classification difference can be calculated first according to each object detection result, and then the regression difference and the classification difference can be superimposed to obtain the prediction difference. Specifically, the detection head H1 described above can be regarded as the first detection head, and the detection head H2 can be regarded as the second detection head. According to the definition of the regression difference, the regression difference is calculated based on the difference between the positions of each object prediction box in the object detection result output by the first detection head and the positions of each object prediction box in the object detection result output by the second detection head. According to the definition of the classification difference, the classification difference is calculated based on the difference between the predicted values of each object category in the object detection result output by the first detection head and the predicted values of each object category in the object detection result output by the second detection head. Finally, the regression difference and the classification difference are superimposed to obtain the prediction difference of the first detection head and the second detection head for the current sample image.
[0084] In an implementation manner of the embodiment of the present application, the prediction differences of at least two detection heads for the current sample image include the prediction differences between every two detection heads among the at least two detection heads for the current sample image; determining the difficulty value of the current sample image according to the prediction difference includes:
[0085] (1) Calculate the average value of the prediction differences between every two detection heads for the current sample image;
[0086] (2) Determine the average value as the difficulty value of the current sample image.
[0087] Obviously, if the number of detection heads exceeds two, a prediction difference for the current sample image can be calculated for any two of these detection heads, that is, multiple prediction differences will be calculated. The average value of these prediction differences can be determined as the difficulty value of the current sample image. For example, assume there are 3 detection heads, namely H0, H1, and H2. Then calculate the prediction difference 1 between H0 and H1 for the current sample image, the prediction difference 2 between H1 and H2 for the current sample image, and the prediction difference 3 between H0 and H2 for the current sample image, and then calculate the average value of the prediction differences 1, 2, and 3 as the difficulty value of the current sample image.
[0088] 103. Select difficult samples from multiple frames of current sample images according to the difficulty value of each frame of the current sample image.
[0089] After calculating the difficulty value of each frame of the current sample image in the manner of step 102, difficult samples can be selected from the multiple frames of the current sample images obtained in step 101 according to the difficulty value of each frame of the current sample image. Generally speaking, the higher the difficulty value of a sample image, the greater the difficulty for the object detection model to predict the sample image. Therefore, those sample images with higher difficulty values can be selected as difficult samples.
[0090] In an implementation manner of the embodiment of the present application, the method for selecting difficult samples from the multiple frames of the current sample images according to the difficulty value of each frame of the current sample image includes:
[0091] Select the third number of current sample images with the highest difficulty value from the multiple frames of the current sample images as difficult samples.
[0092] In actual operation, the number of selected difficult samples, that is, the third number, can be set. Assuming the third number is 10, then 10 frames of images with the highest difficulty value can be selected from the multiple frames of the current sample images as difficult samples, and so on.
[0093] In an implementation manner of the embodiment of the present application, the training data set of the object detection model is obtained through the following method:
[0094] (1) Extract the first number of initial sample images from the sample data pool as the initial data set;
[0095] (2) Determine the target sample category with the least number of corresponding initial sample images according to the sample category included in each frame of the initial sample images in the initial data set;
[0096] (3) Extract the second number of initial sample images including the target sample category from the sample data pool and add them to the initial data set;
[0097] (4) If the number of initial sample images included in the initial data set reaches the set threshold, determine the initial data set as the training data set; otherwise, return to execute the step of determining the target sample category with the least number of corresponding initial sample images according to the sample category included in each frame of the initial sample images in the initial data set and subsequent steps.
[0098] When constructing the initial dataset for training the object detection model, considering that the sample data of some categories is relatively scarce in the specified scenario, the data is selected from the sample data pool frame by frame, so that the finally constructed initial dataset can ensure the data balance of various categories as much as possible. Among them, the sample data pool stores a large number of sample images with labeled sample categories. Specifically, first, a first number of initial sample images are extracted from the sample data pool as the initial dataset, and the random extraction method can be used here. Then, according to the sample categories included in each frame of the initial sample images in the initial dataset, the target sample category with the least number of corresponding initial sample images is determined, that is, the number of images of each sample category included in the initial dataset is counted respectively, and the sample category with the least number of images is found, denoted as CLS min . Then, a second number of initial sample images including the target sample category are extracted from the sample data pool and added to the initial dataset. For example, some sample images including the CLS min category are randomly selected from the sample pool and added to the initial dataset, and the category quantity in the initial dataset is updated. If the number of sample images of the CLS min category has reached the specified threshold, the data selection for this category will no longer be performed. Finally, it is judged whether the number of initial sample images included in the initial dataset reaches the set threshold. If it reaches, the data selection is stopped, and the current initial dataset is determined as the training dataset of the object detection model. In addition, the initial dataset can also be divided into a training dataset, a validation dataset, and a test dataset according to a certain ratio, and the quantity ratio of the sample images included in these 3 datasets is about 8:1:1. If the number of initial sample images included in the initial dataset has not reached the set threshold, return to execute the step of determining the target sample category with the least number of corresponding initial sample images according to the sample categories included in each frame of the initial sample images in the initial dataset, that is, re-determine the sample category with the least number of corresponding sample images in the current initial dataset, and continue to extract sample images for this sample category, and so on in a loop until the number of initial sample images included in the initial dataset reaches the set threshold.
[0099] In an implementation manner of the embodiment of the present application, after selecting difficult samples from multiple frames of current sample images according to the difficulty values of each frame of the current sample images, it further includes:
[0100] (1) Merging the labeled difficult samples into the training dataset of the object detection model;
[0101] (2) Using the merged training dataset to optimize the training of the object detection model;
[0102] (3) If the performance of the optimized trained object detection model does not meet the set requirements, return to execute the steps of obtaining multiple frames of current sample images and subsequent steps.
[0103] After selecting difficult samples from multiple frames of current sample images, each difficult sample can be labeled by manual annotation or machine annotation, and the labeled difficult samples are merged into the training dataset of the object detection model. Then, the object detection model is optimized and trained using the merged training dataset. This process is equivalent to adding the selected difficult samples to the training dataset of the object detection model after annotation, and continuing to optimize and train the object detection model to improve the model performance. Then, test the model performance. If the performance of the optimized trained object detection model meets the set requirements, the training process can be ended and the object detection model can be released. On the contrary, if the performance of the optimized trained object detection model does not meet the set requirements, continuous training is required. At this time, return to execute the step of obtaining multiple frames of current sample images described in step 101, which is equivalent to starting the next round of difficult sample mining and model training process. For example, when executing the first round of the process, obtain multiple frames of current sample images in the first round, select the difficult samples in the first round according to the method described above, and then optimize and train the object detection model based on the difficult samples in the first round. If the performance of the object detection model meets the requirements after the first round of training, end the training and release the model. If the performance of the object detection model does not meet the requirements after the first round of training, enter the second round of the process. When executing the second round of the process, obtain multiple frames of current sample images in the second round. It should be noted that the multiple frames of current sample images obtained in the second round are different from those obtained in the first round. Then, select the difficult samples in the second round according to the method described above, and optimize and train the object detection model based on the difficult samples in the second round... and so on, until an object detection model with performance meeting the requirements is obtained. Assume that the initial dataset described above is Then As the initial training dataset, when subsequently optimizing and training the model, the difficult samples mined in the jth time and the current training dataset are merged to obtain a new training dataset and based on train the object detection model, and so on.
[0104] In an implementation manner of the embodiment of the present application, after optimizing and training the object detection model using the merged training dataset, it further includes:
[0105] If the performance of the optimized trained object detection model meets the set requirements, merge all the selected difficult samples as the output difficult sample set.
[0106] After several rounds of difficult sample mining and model training processes, the performance of the object detection model will reach the set requirements. At this time, the difficult samples selected and labeled in each round can be merged, and finally a labeled difficult sample set with the expected quantity can be obtained for output. At the same time, during the multiple iterative optimization processes of the model, the effectiveness of the mined difficult samples can also be proved. For example, assume that after j rounds of difficult sample mining, the difficult samples can be merged to obtain the finally output difficult sample set
[0107] In an implementation manner of the embodiment of the present application, the object detection model is optimized and trained by using the merged training data set, including:
[0108] (1) Calculate the difference loss according to the prediction difference;
[0109] (2) Add the difference loss to the initial loss function of the object detection model to obtain an updated loss function;
[0110] (3) Optimize and train the object detection model by using the merged training data set and based on the updated loss function.
[0111] When optimizing and training the object detection model by using the merged training data set, the difference loss can be first calculated according to the prediction differences of the respective detection heads of the object detection model according to the definition of the difference loss described above. Then, the difference loss is added to the initial loss function of the object detection model to obtain an updated loss function, such as the loss function L = L loc +L cls +L obj -L dis . Finally, the object detection model is optimized and trained by using the merged training data set and based on the updated loss function.
[0112] Such as Figure 3As shown in the figure, it is a schematic diagram of the training process of a target detection model provided by an embodiment of the present application. First, a data acquisition process is carried out, that is, an initial data set for model training and multiple frames of current sample images are obtained, where the initial data set serves as the initial training data set. The data mining process refers to mining difficult samples from multiple frames of current sample images according to the method described above. Data annotation refers to annotating the mined difficult samples, and data screening refers to further screening the difficult samples during the annotation process. Data division and merging refer to merging the annotated difficult samples with the training data set to obtain an updated training data set. Model training refers to optimizing and training the target detection model according to the updated training data set. Model testing refers to testing whether the performance of the trained target detection model meets the standard. If the performance meets the standard, the model release process is carried out; if the performance does not meet the standard, the data mining step is returned, that is, the next round of difficult sample mining and model training process is started.
[0113] In an embodiment of the present application, first, multiple frames of current sample images are obtained, and then each frame of the current sample images is input into the constructed target detection model for processing; since the target detection model has at least two detection heads, and each detection head can output a target detection result, at least two target detection results can be obtained for each frame of the current sample images; then, according to the at least two target detection results corresponding to each frame of the current sample images, the prediction differences of different detection heads for each frame of the current sample images can be calculated respectively, and the difficulty value of each frame of the current sample images can be determined respectively according to the prediction differences; finally, according to the difficulty value of each frame of the current sample images, difficult samples can be selected from the multiple frames of current sample images, for example, the part of the current sample images with the highest difficulty value can be selected as the difficult samples. The above process can accurately evaluate the difficulty of the target detection model in predicting the current sample images by using the prediction differences of different detection heads for the current sample images, and finally realize accurately mining difficult samples from a large number of unlabeled sample images.
[0114] It should be understood that the magnitudes of the sequence numbers of the steps in the above various embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0115] The above mainly describes a method for mining difficult samples. Next, a device for mining difficult samples will be described.
[0116] Please refer to Figure 4 , an embodiment of a device for mining difficult samples in an embodiment of the present application includes:
[0117] A current sample acquisition module 401, configured to acquire multiple frames of current sample images;
[0118] A difficulty determination module 402, configured to input a current sample image into a constructed object detection model for each frame of the current sample image, and output at least two object detection results of the current sample image through at least two detection heads of the object detection model; calculate at least two prediction differences of the at least two detection heads for the current sample image according to the at least two object detection results; and determine a difficulty value of the current sample image according to the prediction differences.
[0119] A difficult sample selection module 403, configured to select difficult samples from multiple frames of current sample images according to the difficulty values of each frame of the current sample image.
[0120] In an implementation manner of the embodiment of the present application, the at least two detection heads include a first detection head and a second detection head, and the at least two object detection results include a first object detection result output by the first detection head and a second object detection result output by the second detection head; the difficulty determination module includes:
[0121] A regression difference calculation unit, configured to calculate a regression difference according to the difference between the positions of the object prediction boxes in the first object detection result and the positions of the object prediction boxes in the second object detection result.
[0122] A classification difference calculation unit, configured to calculate a classification difference according to the difference between the predicted values of the object categories in the first object detection result and the predicted values of the object categories in the second object detection result.
[0123] A prediction difference determination unit, configured to calculate a prediction difference between the first detection head and the second detection head for the current sample image according to the regression difference and the classification difference.
[0124] In an implementation manner of the embodiment of the present application, the prediction differences of the at least two detection heads for the current sample image include the prediction differences between every two of the at least two detection heads for the current sample image; the difficulty determination module includes:
[0125] An average value calculation unit, configured to calculate an average value of the prediction differences between every two of the detection heads for the current sample image.
[0126] A difficulty determination unit, configured to determine the average value as the difficulty value of the current sample image.
[0127] In an implementation manner of the embodiment of the present application, the difficult sample mining device further includes:
[0128] A data merging module, configured to merge the labeled difficult samples into a training data set of the object detection model.
[0129] A model training module, configured to optimize and train an object detection model by using the merged training data set;
[0130] A performance testing module, configured to, if the performance of the optimized and trained object detection model does not meet the set requirements, return to execute the steps of obtaining multiple frames of current sample images and subsequent steps.
[0131] In an implementation manner of the embodiment of the present application, the model training module includes:
[0132] A difference loss calculation unit, configured to calculate a difference loss according to a prediction difference;
[0133] A loss function update unit, configured to add the difference loss to an initial loss function of the object detection model to obtain an updated loss function;
[0134] A model training unit, configured to optimize and train the object detection model by using the merged training data set and based on the updated loss function.
[0135] In an implementation manner of the embodiment of the present application, the difficult sample mining device further includes:
[0136] A first image extraction module, configured to extract a first number of initial sample images from a sample data pool as an initial data set;
[0137] A sample category determination module, configured to determine a target sample category with the least number of corresponding initial sample images according to the sample categories included in each frame of initial sample image in the initial data set;
[0138] A second image extraction module, configured to extract a second number of initial sample images including the target sample category from the sample data pool and add them to the initial data set;
[0139] A training data set determination module, configured to, if the number of initial sample images included in the initial data set reaches a set threshold, determine the initial data set as the training data set; otherwise, return to execute the steps of determining the target sample category with the least number of corresponding initial sample images according to the sample categories included in each frame of initial sample image in the initial data set and subsequent steps.
[0140] In an implementation manner of the embodiment of the present application, the difficult sample mining device further includes:
[0141] A difficult sample merging module, configured to, if the performance of the optimized and trained object detection model meets the set requirements, merge all selected difficult samples as an output difficult sample set.
[0142] In an implementation manner of the embodiment of the present application, the difficult sample selection module includes:
[0143] A difficult sample selection unit is configured to select a third quantity of current sample images with the highest difficulty value from multiple frames of current sample images as difficult samples.
[0144] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the method for mining difficult samples as described in any of the above embodiments.
[0145] An embodiment of the present application further provides a computer program product, which when running on a terminal device, causes the terminal device to execute the method for mining difficult samples as described in any of the above embodiments.
[0146] Figure 5 is a schematic diagram of a terminal device provided by an embodiment of the present application. As Figure 5 shown, the terminal device 5 in this embodiment includes: a processor 50, a memory 51, and a computer program 52 stored in the memory 51 and executable on the processor 50. When the processor 50 executes the computer program 52, it implements the steps in the embodiments of the above various methods for mining difficult samples, such as Figure 1 the steps 101 to 103 shown. Alternatively, when the processor 50 executes the computer program 52, it implements the functions of each module / unit in the above device embodiments, such as Figure 4 the functions of the modules 401 to 403 shown.
[0147] The computer program 52 may be divided into one or more modules / units, which are stored in the memory 51 and executed by the processor 50 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 52 in the terminal device 5.
[0148] The so-called processor 50 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0149] The memory 51 may be an internal storage unit of the terminal device 5, such as a hard disk or memory of the terminal device 5. The memory 51 may also be an external storage device of the terminal device 5, such as a plug-in hard disk equipped on the terminal device 5, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 51 may also include both the internal storage unit and the external storage device of the terminal device 5. The memory 51 is used to store the computer program and other programs and data required by the terminal device. The memory 51 may also be used to temporarily store the data that has been output or will be output.
[0150] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.
[0151] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.
[0152] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0153] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0154] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0155] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present application.
[0156] In addition, each functional unit in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0157] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0158] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for mining difficult samples, characterized in that, comprising: Obtaining multiple frames of current sample images; For each frame of the current sample image, inputting the current sample image into a pre-constructed object detection model for processing, and respectively outputting at least two object detection results of the current sample image through at least two detection heads of the object detection model; calculating the prediction difference of the at least two detection heads for the current sample image according to the at least two object detection results; Determining the difficulty value of the current sample image according to the prediction difference; Selecting difficult samples from the multiple frames of current sample images according to the difficulty values of each frame of the current sample image.
2. The method according to claim 1, characterized in that, the at least two detection heads include a first detection head and a second detection head, and the at least two object detection results include a first object detection result output through the first detection head and a second object detection result output through the second detection head; the calculating the prediction difference of the at least two detection heads for the current sample image according to the at least two object detection results includes: Calculating a regression difference according to the difference between the positions of each object prediction box in the first object detection result and the positions of each object prediction box in the second object detection result; Calculating a classification difference according to the difference between the predicted values of each object category in the first object detection result and the predicted values of each object category in the second object detection result; Calculating the prediction difference of the first detection head and the second detection head for the current sample image according to the regression difference and the classification difference.
3. The method according to claim 1, characterized in that, the prediction difference of the at least two detection heads for the current sample image includes the prediction difference between every two detection heads among the at least two detection heads for the current sample image; the determining the difficulty value of the current sample image according to the prediction difference includes: Calculating the average value of the prediction differences between every two detection heads for the current sample image; Determining the average value as the difficulty value of the current sample image.
4. The method according to claim 1, characterized in that, after selecting difficult samples from the multiple frames of current sample images according to the difficulty values of each frame of the current sample image, further comprising: Merging the labeled difficult samples into the training data set of the object detection model; Using the merged training data set to optimize the training of the object detection model; If the performance of the object detection model after the optimized training does not meet the set requirements, return to execute the step of obtaining multiple frames of current sample images and subsequent steps.
5. The method according to claim 4, characterized in that, the using the merged training data set to optimize the training of the object detection model includes: Calculating a difference loss according to the prediction difference; Adding the difference loss to the initial loss function of the object detection model to obtain an updated loss function; Optimize and train the object detection model by using the merged training data set and based on the updated loss function.
6. The method according to claim 4, wherein, the training data set is obtained by the following method: Extract a first number of initial sample images from the sample data pool as the initial data set; Determine the target sample category with the least number of corresponding initial sample images according to the sample categories included in each frame of the initial sample images in the initial data set; Extract a second number of initial sample images including the target sample category from the sample data pool and add them to the initial data set; If the number of initial sample images included in the initial data set reaches the set threshold, determine the initial data set as the training data set, otherwise return to execute the step of determining the target sample category with the least number of corresponding initial sample images according to the sample categories included in each frame of the initial sample images in the initial data set and subsequent steps.
7. The method according to claim 4, wherein, after using the merged training data set to optimize and train the object detection model, further includes: If the performance of the optimized and trained object detection model meets the set requirements, merge all the selected difficult samples as the output difficult sample set.
8. The method according to any one of claims 1 to 7, wherein, selecting difficult samples from the multiple frames of current sample images according to the difficulty value of each frame of the current sample image includes: Selecting a third number of current sample images with the highest difficulty value from the multiple frames of current sample images as the difficult samples.
9. A device for mining difficult samples, wherein, includes: A current sample acquisition module for acquiring multiple frames of current sample images; A difficulty determination module for inputting each frame of the current sample image into a pre-constructed object detection model for processing, and respectively outputting at least two object detection results of the current sample image through at least two detection heads of the object detection model; calculating the prediction difference of the at least two detection heads for the current sample image according to the at least two object detection results; determining the difficulty value of the current sample image according to the prediction difference; A difficult sample selection module for selecting difficult samples from the multiple frames of current sample images according to the difficulty value of each frame of the current sample image.
10. A terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, when the processor executes the computer program, it implements the method for mining difficult samples according to any one of claims 1 to 8.
11. A computer-readable storage medium storing a computer program, wherein, when the computer program is executed by a processor, it implements the method for mining difficult samples according to any one of claims 1 to 8.
Citation Information
Cited By
Unsupervised difficult sample mining method and device, electronic equipment and readable storage medium
CN121279491A
Unsupervised difficult sample mining method and device, electronic equipment and readable storage medium
CN121279491B