Method for determining difficulty level of image sample, electronic equipment and readable storage medium

Through the combination of multiple iterative detection and different performance models, the problem of incomplete mining of difficult example samples is solved, and the quantification of difficult example levels and the generalization ability of model training is improved.

CN120495713APending Publication Date: 2025-08-15ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510329982.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

It is difficult to fully explore and classify difficult samples in the existing technology, resulting in limited generalization capabilities of network models.

Method used

Through multiple iteration detection, the image samples are detected with different performances with object detection models, the difficulty level is determined based on the number of iterations, and the hierarchical difficult sample samples are provided.

Benefits of technology

It improves the accuracy of difficult-case detection results and generalization ability of model training, quantifies the difficult-case level of image samples, and provides more comprehensive data support for subsequent model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495713A_ABST
    Figure CN120495713A_ABST
Patent Text Reader

Abstract

The invention discloses an image sample difficulty level determination method, electronic equipment and a readable storage medium, and the method comprises the steps: carrying out the difficulty detection processing of an image sample according to a preset iteration detection mode, obtaining a difficulty detection result of each iteration, and enabling the difficulty detection result to comprise that the image sample is a difficulty sample or the image sample is a common sample; and determining the difficulty level of the image sample according to the target number of iterations of the image sample as the difficulty sample, wherein the target number of iterations is in direct proportion to the difficulty level. Therefore, through multiple iterations, the accuracy of the difficult case detection result of the image sample can be improved, the difficult case level of the image sample can be quantified, a hierarchical image sample is provided for subsequent model training, and the generalization ability and accuracy of model training are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a method for determining the difficulty level of an image sample, an electronic device, and a computer-readable storage medium. Background Art

[0002] Object detection aims to accurately identify and locate specific targets in images or videos, providing foundational data for subsequent tasks. Using computer vision technology to retrieve specific targets from massive amounts of data requires network models to learn from a large number of data samples to achieve autonomous recognition capabilities. The quality of these data samples directly impacts the accuracy and generalization capabilities of the network models.

[0003] To ensure greater robustness and generalization of trained models, it is necessary to mine difficult examples. Currently, this mining is typically done based on samples where the prediction model incorrectly predicts them. However, this approach is prone to missing difficult examples and fails to classify them into different levels. This makes it impossible to train the network model in a targeted manner, which in turn affects the network's generalization capabilities. Summary of the Invention

[0004] The main technical problem solved by this application is to provide a method for determining the difficulty level of image samples, an electronic device and a computer-readable storage medium, which can obtain more comprehensive difficulty example data.

[0005] In order to solve the above technical problems, a technical solution adopted in the present application is: a method for determining the difficulty level of an image sample is provided, the method comprising: performing difficulty detection processing on the image sample according to a preset iterative detection method, and obtaining a difficulty detection result for each iteration, the difficult example detection result including that the image sample is a difficult example sample or the image sample is an ordinary sample; determining the difficulty level of the image sample according to the target number of iterations for the image sample to be a difficult example sample, the target number of iterations being proportional to the difficulty level.

[0006] To solve the above technical problems, another technical solution adopted in this application is: to provide an electronic device, including a memory and a processor, wherein the memory stores program instructions, and the processor calls the program instructions from the memory to execute the above-mentioned method for determining the difficulty level of image samples.

[0007] In order to solve the above technical problems, another technical solution adopted in the present application is: providing a computer-readable storage medium including program data stored therein, and the program data is used to implement the above-mentioned method for determining the difficulty level of image samples when executed by a processor.

[0008] The beneficial effects of the present application are as follows: The method for determining the difficulty level of an image sample in an embodiment of the present application performs difficulty detection on the image sample according to a preset iterative detection method, obtaining a difficulty detection result for each iteration, wherein the difficulty detection result includes whether the image sample is a difficult sample or an ordinary sample; the difficulty level of the image sample is determined based on the target number of iterations for the image sample to be a difficult sample, where the target number of iterations is proportional to the difficulty level. Thus, through multiple iterations, not only can the accuracy of the difficulty detection result of the image sample be improved, but the difficulty level of the image sample can also be quantified, providing hierarchical image samples for subsequent model training, thereby improving the generalization ability and accuracy of model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which:

[0010] Figure 1 is a flowchart of an exemplary embodiment of a method for determining a difficulty level of an example illustrated in the present application;

[0011] Figure 2 A schematic diagram of the target detection model with different iteration sequences shown in this application;

[0012] Figure 3 A schematic diagram of an object detection model with different pruning degrees shown in this application;

[0013] Figure 4 A flowchart of an exemplary embodiment of a process for obtaining difficult example detection results shown in this application;

[0014] Figure 5 A flowchart of an exemplary embodiment of a difficult sample acquisition process shown in this application;

[0015] Figure 6 1 is a schematic diagram of a specific flow chart of an exemplary embodiment of a method for determining a difficulty level of an example illustrated in the present application;

[0016] Figure 7 is a structural diagram of an exemplary embodiment of a device for determining a difficulty level of an example shown in the present application;

[0017] Figure 8 This is a structural diagram of an embodiment of an electronic device provided by the present application;

[0018] Figure 9 It is a structural diagram of an embodiment of a computer-readable storage medium provided by this application. DETAILED DESCRIPTION

[0019] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It will be understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only some, rather than all, structures related to the present application are shown in the drawings. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0020] First, it's important to note that in practical applications, object detection still presents numerous challenges. Factors such as diverse object shapes, large scale variations, severe occlusions, and complex lighting conditions can lead to false or missed detections. Therefore, finding abundant, difficult examples within massive amounts of data is crucial for improving model performance.

[0021] At present, the method for mining difficult data is usually to input adversarial samples or enhanced samples into the prediction model, and to mine difficult data by analyzing the output results of the prediction model. However, the above methods only use a single prediction to obtain difficult data, and cannot mine more abundant difficult samples at a deeper level, nor can they know the difficulty level of difficult samples. Based on this, an embodiment of the present application provides a method for determining the difficulty level of an image sample, which may be referred to as the difficulty level determination method below. By iteratively detecting the image sample and determining the difficulty level of the image sample based on the difficulty detection result obtained by the iterative detection, more difficult samples can be retrieved during multiple iterations, and the image samples can be graded for difficulty.

[0022] For details, please refer to Figure 1 , Figure 1 It is a flowchart of an exemplary embodiment of the method for determining the difficulty level of an example shown in the present application.

[0023] The execution subject of the difficulty level determination method may be a terminal device or a server or other processing device, wherein the terminal device may be a user equipment (UE), a computer, a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The execution subject of the difficulty level determination method may also be a difficulty level determination device. In some possible implementations, the difficulty level determination method may be implemented by a processor calling computer-readable instructions stored in a memory.

[0024] In the embodiment of the present application, the difficulty level determination device is used as the execution subject for description.

[0025] Specifically, the method for determining the difficulty level of an example in this embodiment includes the following steps:

[0026] S110: Performing hard example detection processing on the image sample according to a preset iterative detection method to obtain a hard example detection result for each iteration, where the hard example detection result includes whether the image sample is a hard example sample or an ordinary sample.

[0027] The preset iterative detection method may be manually pre-set with iteration rules. In some embodiments, the preset iterative detection method may be to process the image sample hard example detection according to a preset number of iterations, obtaining a hard example detection result for each iteration. In other embodiments, the preset iterative detection method may be to determine whether an iteration is complete based on the accuracy of the target detection model, and the iteration is complete when the accuracy of the target detection model reaches a preset accuracy.

[0028] During each iteration, the same target detection model can be used to iteratively optimize the image sample. After each iteration, the target detection model is optimized and hard case detection is performed using the optimized target detection model. The hard case detection result for each iteration is obtained by comparing the target detection model with the actual detection result. Thus, by performing hard case detection on the image sample using target detection models with different target detection accuracies, the difficulty level of the image sample can be evaluated. For example, if the output result of an image sample after detection using a target detection model with higher target detection accuracy still differs from the actual result, it indicates that the difficulty level of the image sample is also higher. In other embodiments, different target detection models can be used to iteratively optimize the image sample, and the hard case detection result for each iteration is obtained by comparing the target detection model with the actual detection result. In other embodiments, multiple target detection models can be used to perform target detection on the image sample, and the target detection results of each target detection model can be compared to determine the hard case detection result of the image sample. There is no specific limitation on the method of determining the hard example detection results here. As long as each iteration process uses a target detection model with different performance to perform hard example detection on the image samples, the performance of the target detection model can be better as the number of iterations increases.

[0029] The image sample can be an image selected from the image sample set. For example, it can be any image in the image sample set, or it can be the image with the best image quality in the image sample set, and the embodiments of the present application do not limit this. The difficult example level determination device selects image samples from the image sample set to perform difficult example detection processing, thereby obtaining difficult example samples from the image sample set. In some embodiments, before performing difficult example detection processing on each image sample in the image sample set, low-quality images in the image sample set can be removed to obtain a target image sample set, and then image samples from the target image sample set are traversed to perform difficult example detection processing to obtain difficult example detection results for each image sample.

[0030] The hard example detection result is used to show whether the image sample is a hard example sample. Exemplarily, the hard example detection process can be performed on the image sample through a preset iterative detection method. In each iterative detection, the hard example detection result of the image sample is determined based on the correctness of the target detection result of the image sample. The hard example detection result includes whether the image sample is a hard example sample or the image sample is a common sample. Hard example samples represent samples that are difficult for a general target detection model to detect correctly, and common samples represent samples that can be easily detected correctly by a general target detection model. Exemplarily, when the target detection result of the image sample indicates that the target detection result is different from the true result, the hard example detection result of the image sample is determined to be a hard example sample; when the target detection result of the image sample indicates that there is no difference between the target detection result and the true result, the hard example detection result of the image sample is determined to be a common sample. In other embodiments, if there is no true label for the image sample, multiple target detection models can be used to perform target detection on the image sample, and whether the image sample is a hard example sample is determined based on the comparison results between the target detection results of each target detection model.

[0031] S120: Determine the difficulty level of the image sample according to the target number of iterations for the image sample to be a difficult sample, where the target number of iterations is proportional to the difficulty level.

[0032] The target number of iterations refers to the number of image samples detected as hard examples in a preset iterative detection method. For example, the hard example level determination device can count the number of hard example detection results of image samples in each iteration to obtain the target number of iterations.

[0033] The difficulty level is proportional to the number of target iterations. The more target iterations there are, the higher the difficulty level. The difficulty level can be expressed as a numerical value. The higher the numerical value, the higher the difficulty level, that is, the greater the difficulty of detecting the image sample. As an example, the difficulty level of each image sample is shown in Table 1, where Figure 3 and Figure 6The difficulty level of is the highest. The reasons for the false detection are that the stone pillar is mistakenly detected as the head and the shadow is mistakenly detected as the head and shoulders. After determining the reasons for the false detection, similar image samples can be found from the massive data as training samples.

[0034] Examples of identified hard examples Difficulty level Target missed detection Figure 1 2 Wrong target type Figure 2 2 The stone pillar was mistakenly detected as a human head Figure 3 5 Signs mistakenly detected as pedestrians Figure 4 4 The back of the head was mistakenly detected as a human face Figure 5 3 Shadow misidentified as head and shoulders Figure 6 5 Animals mistakenly detected as human heads Figure 7 4 Target frame is inaccurate Figure 8 1 …… ……

[0035] Table 1

[0036] After the hard example level determination device obtains the target iteration number for detecting an image sample as a hard example sample, the hard example level is determined according to the target iteration number. As an example, the target iteration number can be determined as the value of the hard example level.

[0037] As can be seen, the method for determining the difficulty level of image samples in the embodiment of the present application performs difficulty detection on image samples according to a preset iterative detection method, obtaining a difficulty detection result for each iteration, wherein the difficulty detection result includes whether the image sample is a difficult sample or an ordinary sample. The difficulty level of the image sample is determined based on the target number of iterations for the image sample to be a difficult sample, where the target number of iterations is proportional to the difficulty level. Thus, through multiple iterations, not only can the accuracy of the difficulty detection results of image samples be improved, but the difficulty level of image samples can also be quantified, providing hierarchical image samples for subsequent model training, thereby improving the generalization ability and accuracy of model training.

[0038] In some embodiments, step S110 may further include: performing a first hard example detection process on the image sample using each target detection model to obtain a hard example detection result of a first iteration; optimizing each target detection model based on the hard example detection result of the first iteration to obtain each optimized target detection model; and performing a second hard example detection process on the image sample using each optimized target detection model to obtain a hard example detection result of a second iteration. Based on the optimized target detection models, hard example samples with a higher difficulty level can be obtained, thereby achieving a hierarchical difficulty level for the image samples.

[0039] Before performing difficult case detection processing, it is necessary to select various target detection models, each of which can have different performance. For example, a preset target detection model set is obtained, the preset target detection model set including at least two preset target detection models; and each target detection model is selected from the preset target detection model set based on the target detection accuracy of each preset target detection model in the preset target detection model set.

[0040] The preset target detection model set can be a set of target detection models with different target detection performance. In some embodiments, the preset target detection model set can be obtained during the training process. Specifically, the training sample set is divided according to a preset ratio to obtain a first training sample set and a second training sample set; the first training sample set is used to perform model training on the target detection model to be trained to obtain a target detection model after model training; in the process of training the target detection model after model training using the second training sample set, the target detection model trained by each second training sample is obtained; and the target detection models trained by each second training sample are combined into a preset target detection model set. Among them, the preset ratio can be 7:3, that is, the training samples in the training sample set are divided into 30% and 70%, and 70% of the first training sample set is first used to train the target detection model to be trained, and then 30% of the second training sample set is used to train the target detection model after model training. It should be noted that the same target detection model has different detection performance at different stages during training. Compared with the early stage, the performance of the target detection model trained in the later stage will be more stable, and it is easier for ordinary samples to learn stable results, showing stable detection results at all stages. Difficult samples are prone to inconsistency in the convergence process of a target detection model. As an example, you can refer to Figure 2 During the training process, the first 70% of the iterations of the same object detection model showed unstable results, while the last 30% of the iterations showed relatively stable results. Therefore, after optimizing the object detection model with each second training sample, each optimized object detection model is retained and added to the preset object detection model set, allowing the obtained object detection model to detect difficult samples more accurately.

[0041] In other embodiments, a preset target detection model set can be formed based on the specifications of the target detection model, and then target detection models of different specifications can be obtained from the preset target detection model set. The specifications include the model parameter amount and computational amount of the target detection model. Low specifications mean that the model has fewer parameters and a smaller computational amount, and it is easy for the model to show poor learning due to underfitting for difficult samples or noise samples; high specifications mean that the model has more parameters and a larger computational amount, and it is easy for the model to show memory due to overfitting, and show differences for difficult samples. For example, for the same image sample, a low-specification target detection model is prone to miss detection of edge truncated targets, while a high-specification target detection model can maintain a better effect on edge target detection.

[0042] In other embodiments, a preset target detection model set can be formed based on the quantization degree of the target detection model, and then target detection models with different quantization degrees can be obtained from the preset target detection model set. Quantization is an acceleration and model compression technology, which mainly converts floating-point operations with high memory usage and low computing speed into floating-point or integer operations with low memory usage and high computing speed. As an example, in a general 64-bit system, a standard floating-point number occupies 32 bits (bits), which can be converted to 16 bits, 8 bits or even 4 bits after quantization, corresponding to fp16, int8, and int4 respectively. Therefore, the size of the target detection model in memory can be reduced to nearly 50%, 25%, and 12.5% through quantization technology, and its computing speed will also increase significantly, but some accuracy will be sacrificed. For the same unquantized object detection model, output results will vary under different quantization accuracies. Common samples maintain stable output due to multiple redundant features, while difficult samples, due to poor feature redundancy (for example, multiple required features are essential), are prone to unstable output due to the uncertainty caused by quantization. Partial precision differences are the root cause of inconsistent output results. Therefore, we can cleverly use the precision error introduced by quantization as a form of noise. We can select fp16, int8, and int4 from the same object detection model for difficult example detection and determine which targets the object detection model is not robust to.

[0043] In other embodiments, a preset target detection model set can be determined based on the pruning degree of the target detection model, and then target detection models with different pruning degrees can be obtained from the preset target detection model set. Pruning and quantization have the same principle. Quantization increases the instability of distributed features, while pruning directly discards some distributed features. As an example, please refer to Figure 3 The image obtains four layers of features through convolution 1, while only two layers of features are obtained through pruning convolution 1. Compared to convolution, pruned convolution discards some unimportant features through feature evaluation, thereby reducing some computation and memory usage, improving computation speed, but also sacrificing some accuracy. Therefore, the accuracy loss introduced by pruning can be used as a means of noise injection to obtain less robust targets.

[0044] Each target detection model can be obtained by any one of the above four methods or a combination thereof. In some embodiments, each target detection model can be a target detection model of the same category. In other embodiments, the target detection model can be a target detection sub-model of multiple categories, for example, it can include at least two first-category target detection sub-models and at least one second-category target detection sub-model, wherein each first-category target detection sub-model can be obtained by any one of the above four methods or a combination thereof. The implementation method in which the target detection model includes target detection sub-models of multiple categories is explained in the following embodiments and will not be repeated here.

[0045] After obtaining each target detection model, each target detection model can be used to perform a first hard example detection process on the image sample to obtain a first iteration of hard example detection results. In some embodiments, the image sample can be input into each target detection model for target detection processing, and the target detection results output by each target detection model can be obtained; the first iteration of hard example detection results for the image sample can be determined based on the comparison results between the target detection results output by each target detection model. In this way, it is possible to determine whether an image sample is a hard example by comparing the output results of multiple target detection models without the need for image annotation.

[0046] Target detection is the process of detecting a target object of interest from an image sample. The target detection result includes the target area of the target object and the corresponding target type. Exemplarily, the target detection model has a target detection function, and the image samples are respectively input into each target detection model with a target detection function to obtain the target detection results output by each target detection model. Among them, each target detection model can be a target detection model with different target detection performance, and there are differences in the performance of difficult samples. The same image sample is detected by each target detection model with different target detection performance, and whether the image sample is a difficult sample can be determined based on the target detection results. As an example, the target detection results output by each target detection model are compared pairwise to determine whether the target detection results are consistent. If there is a difference, the image sample is determined to be a difficult sample.

[0047] Furthermore, the method of performing a second difficult example detection process on the image sample using each optimized target detection model to obtain the difficult example detection result of the second iteration can also be to determine the difficult example detection result of the image sample in the second iteration by comparing the target detection results output by each optimized target detection model. If the target detection results output by each optimized target detection model are consistent, the difficult example detection result is determined as a common sample. If the target detection results output by each optimized target detection model are inconsistent, the difficult example detection result is determined as a difficult example sample. Similarly, the difficult example detection result of each iterative process of the image sample can be determined by comparing the target detection results of multiple target detection models.

[0048] Specifically, when obtaining the comparison result between the target detection results of each target detection model, it can be obtained by comparing the target areas in each target detection result. The target area can also be called a target box, which is the area obtained by the target detection model by selecting the target object of interest through a closed shape box. The difficult example level determination device calculates the intersection and union ratio of the target area output by each target detection model and the target area output by other target detection models, and obtains multiple area intersection and union ratios; determines the comparison result between the target detection results output by each target detection model according to the size relationship between the multiple area intersection and union ratios and the preset intersection and union ratio threshold; in response to the comparison result indicating that the target detection results output by each target detection model are different, the difficult example detection result of the first iteration is determined to be a difficult example sample. Thus, the area overlap of each target area is obtained by the intersection and union calculation, thereby obtaining the comparison result between the target detection results output by each target detection model.

[0049] Specifically, the difficulty level determination device traverses each target area output by each target detection model, and then calculates the intersection-and-union ratio of the currently traversed target area with the target areas output by other target detection models to obtain multiple regional intersection-and-union ratios. According to the regional intersection-and-union ratios, it is determined whether there are areas in other target detection models that belong to the same target as the currently traversed target area, wherein if the regional intersection-and-union ratio between the target area output by other target detection models and the currently traversed target area is greater than a preset intersection-and-union ratio threshold, it is considered that the two belong to the same target. For example, if there are ten target detection models, for the target area output by one of the target detection models, the intersection-and-union ratio of each target area output by the target detection model is calculated with the target areas output by the other nine target detection models to determine whether each target area output by the target detection model also exists in the target detection results output by other target detection models. If the target areas output by each target detection model exist in other target detection models, that is, the target areas output by each target detection model are the same, then the comparison result is determined to be the target detection results output by each target detection model are the same; in response to the comparison result characterizing that the target detection results output by each target detection model are the same, then the difficult example detection result of the first iteration is determined to be a common sample; if there is a target area output by a target detection model that does not exist in other target detection models, or is inconsistent with the target area output by other target detection models, then the comparison result is determined to be the target detection results output by each target detection model are different; in response to the comparison result characterizing that the target detection results output by each target detection model are different, then the difficult example detection result of the first iteration is determined to be a difficult example sample.

[0050] In other embodiments, after the difficulty level determination device determines that the target areas output by each target detection model belong to the same target, it determines whether the target types corresponding to each target area are the same. If they are the same, the comparison result is determined to be that the target detection results output by each target detection model are the same; if they are different, the comparison result is determined to be that the target detection results output by each target detection model are different. That is, only when the region intersection-over-union ratio between the target areas output by each target detection model is greater than the preset intersection-over-union ratio threshold, and the target types of the corresponding target areas are the same, the comparison result is determined to be that the target detection results output by each target detection model are the same. In this way, image samples such as missed detection, false detection, inconsistent target type, or inaccurate target area between each target detection model can be obtained, and the union of all image samples is composed of multiple target detection difference data.

[0051] Furthermore, an embodiment of a target detection model including a plurality of target detection sub-models of different categories is described as follows: the target detection model includes at least two first-category target detection sub-models and at least one second-category target detection sub-model, and image samples are respectively input into each first-category target detection sub-model for target detection processing to obtain target detection results output by each first-category target detection sub-model; in response to the target detection results output by each first-category target detection sub-model having the same representation, the image samples are input into each second-category target detection sub-model for target detection processing to obtain target detection results output by each second-category target detection sub-model; the difficult example detection results of the first iteration are determined based on the target detection results output by each first-category target detection sub-model and the target detection results output by each second-category target detection sub-model. Thus, target detection is performed by target detection sub-models of different categories, which can solve the current problem that a single target detection model cannot recall all difficult example samples when screening difficult example samples, and cannot mine data with a higher level of difficulty.

[0052] Among them, the first category target detection sub-model and the second category target detection sub-model can be selected from the multi-target detection model (object detection) and the multi-target classification model (object classification). There are two main differences between the multi-target detection model and the multi-target classification model: first, the multi-target detection model uses full-image scaling, which loses more detailed information about distant and small targets, while the multi-target classification model cuts out small images for recognition, which loses less information than the multi-target detection model; second, the multi-target detection model performs real-time full-image detection on all image samples, and the algorithm complexity of the NMS (Non-Maximum Suppression) stage is O(n times the number of targets n). 2 ), so the performance of the multi-target detection model is greatly squeezed, and the multi-target classification model only classifies and detects the required targets, so the model can be relatively large.

[0053] The device for determining the difficulty level first inputs the image sample into each first-category object detection sub-model for object detection processing, thereby obtaining an object detection result output by each first-category object detection sub-model. The selection of each first-category object detection sub-model may refer to the above-described embodiment for selecting an object detection model. For example, each first-category object detection sub-model may be selected from a set of preset first-category object detection sub-models based on the training level, specification, quantization level, or pruning level of the preset first-category object detection sub-model.

[0054] After obtaining the target detection results output by each first-category target detection sub-model, the target detection results include the target area and target type of the target object. The overlap between the target areas output by each first-category target detection sub-model is compared pairwise. For example, the overlap between the target areas output by any two first-category target detection sub-models can be calculated by using the intersection-over-union (IoU) ratio. If the IoU value between the target areas output by any two first-category target detection sub-models is greater than a preset IoU threshold, the target types of the corresponding target areas are compared to determine whether they are consistent. This allows determining whether the target detection results output by each first-category target detection sub-model are consistent. Furthermore, the difference data between each first-category target detection sub-model can be further refined into target omission, target misdetection, inaccurate target area, and incorrect target type using the IoU value and target type. Among them, missed target detection refers to the failure to detect relevant targets or target parts; false target detection includes false detection of stationary targets, such as stone pillars, decorations, signboards, garbage dumps, debris piles and other stationary targets; false face detection, such as billboards, necks, hands, wheels, back of the head, bald heads, etc. are falsely detected as faces; false shadow detection, such as shadows are falsely detected as heads and shoulders, and animals are falsely detected; inaccurate target area includes target areas that are too large or too small; target type error refers to the incorrect detection of the type of target object, such as pedestrians are falsely detected as non-motor vehicles, non-motor vehicles are falsely detected as motor vehicles, etc.

[0055] If the target detection results output by each first-category target detection sub-model have the same representation, that is, they may be correct data, but because the target detection sub-models of the same category have certain similarities, they may miss difficult samples that cannot be detected by multiple first-category target detection sub-models. Therefore, in order to find more difficult samples, at least one second-category target detection sub-model is used to help recall more difficult samples and obtain more accurate difficult detection results. In the embodiment of the present application, after the target detection results output by each first-category target detection sub-model have the same representation, each second-category target detection sub-model is used to perform target detection processing on the image sample to obtain the target detection results output by each second-category target detection sub-model, and determine the difficult detection results of the first iteration based on the target detection results output by each first-category target detection sub-model and the target detection results output by each second-category target detection sub-model.

[0056] Among them, in response to the target detection results output by each first-category target detection sub-model being the same as the target detection result representation results output by each second-category target detection sub-model, the difficult example detection results of the first iteration are determined as common samples; in response to the target detection results output by each first-category target detection sub-model being different from the target detection result representation results output by each second-category target detection sub-model, the difficult example detection results of the first iteration are determined as difficult example samples. Thus, screening difficult example samples based on the consistency between the target detection results output by different categories of target detection sub-models can not only reduce the data annotation requirements, but also obtain more accurate difficult example detection results and recall more difficult example samples. In other embodiments, the difficult example level determination device determines whether the target detection results output by each second-category target detection sub-model are the same; if so, the difficult example detection results of the first iteration are determined as common samples; otherwise, the difficult example detection results of the first iteration are determined as difficult example samples.

[0057] It should be noted that the difficult example detection results of each iterative process can be obtained by referring to the above embodiments.

[0058] In order to elaborate on the process of obtaining the difficult case detection results in the embodiment of this application, Figure 4 The flowchart shown further illustrates this, as follows:

[0059] The image samples are input into each first-category target detection sub-model for target detection processing, and the target detection results output by each first-category target detection sub-model are obtained; the target detection results output by each first-category target detection sub-model are compared in pairs, for example, the target detection results of OD1 (object detection1, multi-target detection model 1) (including the coordinates of the target area and the target type, etc.) are compared with those of OD2 (object detection model 2). detection2, multi-target detection model 2) are compared with the target detection results (including the coordinates of the target area and the target type, etc.) of the target area; it is determined whether the target area overlap between the two is greater than the preset intersection-over-union ratio threshold; if not, it is determined that it is a target missed detection, a target misdetection or an inaccurate target area, and if so, it is determined whether the target type of the corresponding target area is consistent, and if not, it is determined that the target type is wrong, and if so, it is determined to be possible correct data; for possible correct data, each second category target detection sub-model is used to perform target detection processing on the image sample to obtain the target detection result output by each second category target detection sub-model, and the target detection result includes the target type; it is determined whether the target type output by each first category target detection sub-model is consistent with the target type output by each second category target detection sub-model, and if so, it is determined to be a simple sample, and if not, it is determined to be a difficult sample.

[0060] Furthermore, after obtaining the hard example detection results in each iteration, the hard example detection results can be used to optimize each target detection model. Specifically, if the hard example detection result indicates that an image sample is a hard example sample, the image sample is manually labeled, and the manually labeled image samples are used to optimize each target detection model, thereby obtaining each optimized target detection model; then, each optimized target detection model is used to locate image samples with a higher difficulty level. This process is repeated, ultimately resulting in a hierarchical hard example sample, a series of target detection models for hard example samples, and a set of labeled image samples to be processed.

[0061] In some embodiments, the difficulty level determination device can utilize AdaBoost (Adaptive Boosting) to gradually improve the learning effect of each target detection model by changing the composition of image samples. For example, image samples are weighted based on the difficulty detection results of the first iteration to obtain weighted image samples; the weighted image samples are then input into each optimized target detection model for a second difficulty detection process, resulting in the difficulty detection results of the second iteration. As image samples are continuously weighted, the generalization and accuracy of the target detection model will continue to improve.

[0062] Furthermore, the device for determining the difficulty level performs a second weighting process on the weighted image samples according to the second iterative difficulty detection result to obtain a second weighted image sample, and simultaneously performs a second optimization process on each optimized object detection model according to the second iterative difficulty detection result to obtain each second optimized object detection model; the device inputs the second weighted image samples into each second optimized object detection model to perform a third difficult example detection process to obtain a third iterative difficult example detection result, and so on.

[0063] It can be seen that the weighted processing in the embodiment of the present application is completed on the basis of the image samples after weighted processing in the previous iteration process. Therefore, as the number of historical iterations in which the image samples are detected as difficult samples increases, their weighted weights will become larger and larger, thereby gradually increasing the emphasis on difficult samples, improving the learning ability of the target detection model, and gradually reducing the possibility of image samples being misdetected or missed. The difficulty level of the image samples can be quantified through the weighted weight of the image samples (number of oversampling times).

[0064] It should be noted that the image sample is one of the target image sample set. During each iteration, all image samples in the target image sample set are subjected to hard example detection processing to obtain the hard example detection results of each image sample in different iterations. Then, the corresponding image samples are weighted according to the hard example detection results of each image sample to obtain each weighted image sample. Therefore, as the number of iterations increases, the weight of the hard example samples in the target image sample set will become larger and larger, while the weight of the common samples will become smaller and smaller. As an example, please refer to Figure 5 During the i+1th iteration, the difficulty level determination device obtains image samples and full data with a difficulty level below i, wherein the image samples below level i are image samples after weighted processing during the i-th iteration, and the full data are other image samples in the target image sample set except the image samples below level i, which can be considered as ordinary samples; each target detection model is optimized according to the image samples with a difficulty level below i to obtain each optimized target detection model; the image samples and full data with a difficulty level below i are input into each optimized target detection model for difficult example detection processing to obtain the difficult example detection result of each image sample, and the image samples with a difficulty level i+1 are searched from the image samples with a difficulty level below i, thereby obtaining image samples and full data with a difficulty level i+1 or less. Therefore, each target detection model that has been optimized for difficulty level below i is used to search for image samples with difficulty level i+1, indicating that the embodiment of the present application infinitely weights the prediction results of the new learner when facing difficult samples, and discards the prediction results of each target detection model in the previous iteration step.

[0065] In order to elaborate on the process of obtaining the difficult case detection results in the embodiment of this application, Figure 6 The flowchart shown further illustrates this, as follows:

[0066] The embodiment of the present application takes the example that each target detection model includes a first category target detection sub-model and a second category target detection sub-model, the first category target detection sub-model is a multi-target detection model, and the second category target detection sub-model is a multi-target classification model. It mainly includes four parts: image quality assessment stage, multi-target detection model stage, multi-target classification model stage and AdaBoost iterative difficult example stage.

[0067] Image quality assessment: Initialize the number of iterations, AdaBoost_i, to 0, and obtain a set of image samples. Because image blur, brightness distortion, color distortion, and other factors can make corresponding images difficult to recognize or classify, severely impacting model training and final performance, an image quality assessment algorithm can be used to assess the quality of each image sample in the set and filter out low-quality samples. There are two approaches to image quality assessment: Full-Reference Image Quality Assessment (FR-IQA), which requires the original image as a reference, such as Peak Signal to Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM); and No-Reference Image Quality Assessment (NR-IQA), which does not require the original image as a reference, such as BRISQUE (Blind / Referenceless Image Spatial Quality Evaluator) and NIQE (Natural Image Quality Evaluator). The non-reference image quality assessment method has low computational complexity and is applicable to various image processing tasks. Because the present embodiment requires processing massive amounts of data for cleaning, it uses a non-reference image quality assessment method to filter out low-quality image samples. The present embodiment utilizes a trained image quality assessment algorithm model to assess image quality and then filters out low-quality image samples.

[0068] Multi-target detection model stage: Within each AdaBoost iteration, at least two multi-target detection models are used to locate the difference samples between the models. There are four ways to obtain the at least two multi-target detection models: 1. Using different models from the last 30% of iterations of the same target detection model training process; 2. Using multi-target detection models of different specifications; 3. Using multi-target detection models of different quantization levels; 4. Using multi-target detection models of different pruning levels. After inputting the image samples into each multi-target detection model, the target detection results output by each multi-target detection model are obtained; if the image samples already have target area labels, the target detection results output by each multi-target detection model are compared with the area labels; if the image samples do not have target area labels, the target areas in any two target detection results are compared to see whether there are differences. If so, they are determined to be difference image samples, and the difficult example detection results are determined as difficult example samples; otherwise, they are determined to be indifference image samples; then, for indifference image samples, it is determined whether the corresponding image samples have classification labels (such as pedestrian, motor vehicle and non-motor vehicle labels). If so, the target detection results output by each multi-target detection model are directly compared with the classification labels. If not, the indifference image samples are input into each multi-target classification model.

[0069] Multi-target classification model stage: The device for determining the difficulty level inputs the indifferent image samples into each multi-target classification model to obtain the target detection results output by each multi-target classification model; compares whether the target type in the target detection results output by each multi-target classification model is consistent with the target type in the target detection results output by each multi-target detection model; if not, determines that the corresponding image sample is a difference image sample, and determines the difficult example detection result as a difficult example sample. For example, each multi-target detection model outputs a target area corresponding to a non-motor vehicle, but the human non-attribute classification model (pedestrian, non-motor vehicle attribute classification model) identifies the target area as a pedestrian; each multi-target detection model outputs a target area corresponding to a motor vehicle, but the motor vehicle attribute classification model identifies the target area as a tricycle.

[0070] Therefore, through the multi-target detection model and the multi-target classification model, the difficult samples in this AdaBoost iteration step are mainly divided into four categories: target missed detection, D miss , target type error D type , target false detection D fake The case where the target frame is not accurate D location .

[0071] AdaBoost iterative difficulty phase: Within this iteration, AdaBoost_i, the difficulty level of the difference image samples is increased by 1. The difference image samples are manually reviewed and annotated, and the multi-target detection model and multi-target classification model are fine-tuned using the difference image samples. At the same time, each image sample is weighted according to its difficulty detection result, and the weighted image samples are used for the next iteration. A check is made to determine whether the current number of iterations, AdaBoost_i, has reached the preset number of iterations, N. If so, the process ends and the difficulty level of each image sample is output. Otherwise, the weighted image samples are input into the fine-tuned multi-target detection model and multi-target classification model to continue the iterative process.

[0072] See also Figure 7 , Figure 7 is a schematic diagram of the structure of an exemplary embodiment of a device for determining a difficulty level illustrated in this application. The device 700 includes an iteration module 710 and a determination module 720. The iteration module 710 is configured to perform difficulty detection on an image sample according to a preset iterative detection method, obtaining a difficulty detection result for each iteration, wherein the difficulty detection result includes whether the image sample is a difficult sample or a common sample. The determination module 720 is configured to determine the difficulty level of the image sample based on a target number of iterations for determining that the image sample is a difficult sample, where the target number of iterations is proportional to the difficulty level.

[0073] In the above scheme, the difficulty level determination device performs difficulty detection on image samples according to a preset iterative detection method, obtaining a difficulty detection result for each iteration, which may indicate whether the image sample is a difficult example or a common example. The difficulty level of the image sample is determined based on the target number of iterations for the image sample to be a difficult example, with the target number of iterations being proportional to the difficulty level. This, through multiple iterations, not only improves the accuracy of the difficulty detection results for image samples but also quantifies the difficulty level of the image samples, providing hierarchical image samples for subsequent model training and improving the generalization and accuracy of model training.

[0074] Among them, the functions of each module can be found in the embodiment of the method for determining the difficulty level, and will not be repeated here.

[0075] In order to implement the method for determining the difficulty level of the above embodiment, this application proposes another electronic device, which can be found in detail. Figure 8 , Figure 8 It is a structural diagram of an embodiment of an electronic device provided by this application.

[0076] The electronic device 800 includes a memory 810 and a processor 820 , wherein the memory 810 and the processor 820 are coupled.

[0077] The memory 810 is used to store program data, and the processor 820 is used to execute the program data to implement the method for determining the difficulty level of the above embodiment.

[0078] In this embodiment, the processor 820 may also be referred to as a CPU (Central Processing Unit). The processor 820 may be an integrated circuit chip having signal processing capabilities. The processor 820 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor, or the processor 820 may be any conventional processor.

[0079] This application also provides a computer-readable storage medium, such as Figure 9 As shown, the computer-readable storage medium 900 is used to store program data 910. When the program data 910 is executed by the processor, it is used to implement the method for determining the difficulty level of the example in the method embodiment of the present application.

[0080] The method involved in the embodiment of the method for determining the difficulty level of the present application, when implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a device, such as a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0081] The above description is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for determining the difficulty level of an image sample, characterized in that: The method for determining the difficulty level of the image sample includes: Performing hard example detection processing on the image sample according to a preset iterative detection method to obtain a hard example detection result for each iteration, wherein the hard example detection result includes whether the image sample is a hard example sample or the image sample is a common sample; The difficulty level of the image sample is determined according to a target number of iterations for the image sample to be a difficult sample, wherein the target number of iterations is proportional to the difficulty level.

2. The method for determining the difficulty level of an image sample according to claim 1, wherein: The step of performing hard example detection processing on the image samples according to the preset iterative detection method to obtain the hard example detection result of each iteration includes: Performing a first hard example detection process on the image sample using each target detection model to obtain a hard example detection result of a first iteration; Optimizing each target detection model according to the hard example detection result of the first iteration to obtain each optimized target detection model; The optimized target detection models are used to perform a second difficult example detection process on the image samples to obtain a second iterative difficult example detection result.

3. The method for determining the difficulty level of an image sample according to claim 2, wherein: The step of performing a first hard example detection process on the image sample using each target detection model to obtain a hard example detection result of a first iteration includes: Inputting the image samples into each target detection model to perform target detection processing, and obtaining target detection results output by each target detection model; The hard example detection result of the first iteration of the image sample is determined according to the comparison result between the target detection results output by each target detection model.

4. The method for determining the difficulty level of an image sample according to claim 3, wherein: The target detection result includes at least one target area, and the step of determining the hard-example detection result of the first iteration of the image sample based on the comparison result between the target detection results output by each target detection model includes: The target area output by each target detection model is calculated with the target area output by other target detection models to obtain multiple area intersection and union ratios; Determine a comparison result between the target detection results output by each target detection model based on a size relationship between the IoU values of the multiple regions and a preset IoU threshold; In response to the comparison result indicating that the target detection results output by each target detection model are different, the difficult example detection result of the first iteration is determined to be a difficult example sample.

5. The method for determining the difficulty level of an image sample according to claim 2, wherein: The step of performing a second hard example detection process on the image sample using each optimized target detection model to obtain a hard example detection result of a second iteration includes: performing weighted processing on the image samples according to the hard example detection result of the first iteration to obtain weighted image samples; The weighted image samples are input into the optimized target detection models to perform a second difficult example detection process, and obtain a second iterative difficult example detection result.

6. The method for determining the difficulty level of an image sample according to claim 2, wherein: The object detection model includes at least two first-category object detection sub-models and at least one second-category object detection sub-model. The step of performing a first hard example detection process on the image sample using each object detection model to obtain a first iterative hard example detection result includes: Inputting the image samples into each first category target detection sub-model for target detection processing, thereby obtaining target detection results output by each first category target detection sub-model; In response to the target detection results output by the first-category target detection sub-models having the same representation, inputting the image sample into the second-category target detection sub-models for target detection processing to obtain target detection results output by the second-category target detection sub-models; The difficult example detection result of the first iteration is determined based on the target detection results output by each first-category target detection sub-model and the target detection results output by each second-category target detection sub-model.

7. The method for determining the difficulty level of an image sample according to claim 6, wherein: The step of determining the hard example detection result of the first iteration based on the target detection results output by each first category target detection sub-model and the target detection results output by each second category target detection sub-model includes: In response to the target detection results output by each first-category target detection sub-model being identical to the target detection result representation results output by each second-category target detection sub-model, determining the hard example detection result of the first iteration as a common sample; In response to the target detection results output by each first-category target detection sub-model being different from the target detection result representation results output by each second-category target detection sub-model, the hard example detection result of the first iteration is determined as a hard example sample.

8. The method for determining the difficulty level of an image sample according to claim 2, wherein: Before the step of performing a first hard example detection process on the image sample using each object detection model to obtain a hard example detection result of a first iteration, the method further includes: Obtaining a preset target detection model set, wherein the preset target detection model set includes at least two preset target detection models; Each target detection model is selected from the preset target detection model set according to the target detection accuracy of each preset target detection model in the preset target detection model set.

9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores program instructions, and the processor calls the program instructions from the memory to execute the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that include: Program data is stored, and when the program data is executed by a processor, it is used to implement the method according to any one of claims 1 to 8.