Target recognition model training device, target recognition device and method thereof
Patent Information
- Application Number
- CN202610969521.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-22
AI Technical Summary
然而,受光照变化、天气干扰(如雨、雾)、遮挡物干扰等环境因素影响,红绿灯在物理成像层面会出现颜色混淆,导致红灯与黄灯之间、或黄灯与绿灯之间易出现视觉混淆
[0012]本公开实施例提供的技术方案,对样本进行混淆类型标注后,在目标识别模型训练阶段通过类别标签、混淆类型标签及置信度标签构建的损失函数对初始识别模型进行迭代训练,使训练过程中同时兼顾分类精度和置信度准确度,使训练后的目标识别模型输出的置信度概率分布更准确,提升了目标识别模型对目标识别的可靠性,从而避免目标识别模型在视觉混淆场景下输出高置信度的错误类别的情况。
Smart Images

Figure CN122799082A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer vision technology, specifically to a target recognition model training device, a target recognition device, and a method thereof. Background Technology
[0002] In scenarios such as intelligent transportation and autonomous driving, traffic light recognition is a core visual task to ensure driving safety, and the accuracy of traffic light recognition directly affects traffic safety. However, due to environmental factors such as changes in lighting, weather interference (such as rain and fog), and obstructions, color confusion can occur at the physical imaging level of traffic lights, leading to visual confusion between red and yellow lights, or between yellow and green lights.
[0003] How to avoid color confusion and improve recognition accuracy in traffic light recognition scenarios is a technical problem that urgently needs to be solved. Summary of the Invention
[0004] To address the aforementioned technical issues, this disclosure provides a target recognition model training device and method, an electronic device, and a storage medium to improve the accuracy of target recognition models in visually confusing scenarios.
[0005] The first aspect of this disclosure provides a target recognition model training apparatus, including a processor configured to: acquire a training sample set; wherein the training sample set includes multiple training samples, the training samples including sample images and category labels of the target to be recognized in the sample images, confusion type labels corresponding to the category labels, and confidence labels corresponding to the category labels; input the training samples into an initial recognition model to obtain the probability distribution of the target to be recognized belonging to each candidate category; determine the prediction confidence of the predicted category based on the probability distribution; construct a target loss function based on the category labels, confusion type labels, confidence labels, predicted categories, and prediction confidence; and train the initial recognition model based on the target loss function to obtain a target recognition model.
[0006] A second aspect of this disclosure provides a target recognition apparatus, including a processor configured to: acquire an image to be recognized; input the image to be recognized into a target recognition model to obtain a probability distribution of each candidate category of the image to be recognized output by the target recognition model; wherein the target recognition model is trained based on the aforementioned target recognition model training apparatus; and determine the target category and target prediction confidence of the image to be recognized based on the probability distribution.
[0007] A third aspect of this disclosure provides a method for training a target recognition model, comprising: acquiring a training sample set; wherein the training sample set includes multiple training samples, the training samples including sample images and category labels of the target to be recognized in the sample images, confusion type labels corresponding to the category labels, and confidence labels corresponding to the category labels; inputting the training samples into an initial recognition model to obtain the probability distribution of the target to be recognized belonging to each candidate category; determining the prediction confidence of the predicted category based on the probability distribution; constructing a target loss function based on the category labels, confusion type labels, confidence labels, predicted categories, and prediction confidence; and training the initial recognition model based on the target loss function to obtain a target recognition model.
[0008] A fourth aspect of this disclosure provides a target recognition method, comprising: acquiring an image to be recognized; inputting the image to be recognized into a target recognition model to obtain a probability distribution of each candidate category of the image to be recognized output by the target recognition model; wherein the target recognition model is trained based on the target recognition model training method described above; and determining the target category and target prediction confidence of the image to be recognized based on the probability distribution.
[0009] A fifth aspect of this disclosure provides a computer-readable storage medium storing a computer program that is executed by a processor to perform a target recognition model training method of the third aspect of this disclosure, or to perform a target recognition method of the fourth aspect of this disclosure.
[0010] A sixth aspect of this disclosure provides an electronic device comprising: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the target recognition model training method of the third aspect of this disclosure, or to execute the target recognition method of the fourth aspect of this disclosure.
[0011] The seventh aspect of this disclosure provides a computer program product that, when executed by an instruction processor, performs either the target recognition model training method of the third aspect of this disclosure or the target recognition method of the fourth aspect of this disclosure.
[0012] The technical solution provided in this disclosure, after labeling samples with confusion types, iterates the initial recognition model during the target recognition model training stage by constructing a loss function based on category labels, confusion type labels, and confidence labels. This allows the training process to simultaneously consider classification accuracy and confidence accuracy, making the confidence probability distribution output by the trained target recognition model more accurate and improving the reliability of the target recognition model in target recognition. This avoids the situation where the target recognition model outputs a high-confidence incorrect category in visual confusion scenarios. Attached Figure Description
[0013] Figure 1 This is a structural diagram of a target recognition model training device provided in an embodiment of the present disclosure.
[0014] Figure 2 This is a schematic diagram of the structure of a second target recognition model training device provided in an embodiment of this disclosure.
[0015] Figure 3 This is a structural diagram of a target recognition device provided in an embodiment of the present disclosure.
[0016] Figure 4 This is a flowchart illustrating a target recognition model training method provided in an embodiment of the present disclosure.
[0017] Figure 5 This is a flowchart illustrating a target recognition method provided in an embodiment of the present disclosure.
[0018] Figure 6 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation
[0019] To explain this disclosure, exemplary embodiments of the disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the disclosure, and not all of them. It should be understood that the disclosure is not limited to exemplary embodiments.
[0020] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0021] Application Overview In intelligent transportation and autonomous driving scenarios, environmental perception is fundamental for vehicles to achieve autonomous decision-making and planning. Among these, accurate recognition of traffic lights, especially real-time and accurate identification of traffic light colors (red, yellow, and green), is one of the core perception tasks for ensuring safe passage and improving traffic efficiency. Incorrect traffic light recognition could lead to vehicles mistakenly proceeding through a red light or braking unnecessarily during a green light, potentially causing serious accidents. Therefore, accurate traffic light recognition is crucial in autonomous driving systems.
[0022] Deep learning-based traffic light color recognition methods have been widely applied. These methods typically employ convolutional neural networks as the core for feature extraction and classification, training the model using image data labeled with traffic light colors to identify color categories. However, environmental factors such as changes in lighting, weather conditions, or obstructions can cause color confusion at the physical imaging level. For example, red lights tend to appear yellowish under strong backlighting or smog, leading to visual confusion between red and yellow lights; under overexposure or rain / fog scattering, green and yellow colors are weakened, resulting in a whitish or yellowish hue, further causing visual confusion between yellow and green lights. When color confusion occurs, the traffic light color recognition model trained using deep learning may output incorrect recognition results, further affecting the decisions of downstream decision-making modules and even posing significant safety hazards.
[0023] In related technologies, models are optimized using loss functions such as cross-entropy loss, which are trained solely to minimize classification errors. This may lead to the model easily outputting errors on samples with color confusion.
[0024] Therefore, how to effectively avoid high-confidence erroneous predictions by target recognition models on easily confused samples in visually confusing scenarios has become an urgent problem to be solved.
[0025] This disclosure provides a target recognition model training device, a target recognition device, and a method thereof, which effectively reduces the high-confidence misclassification of the target recognition model in visually confusing scenarios and improves the reliability of the target recognition results.
[0026] The target recognition model training device and method, as well as the target recognition device and method, provided in the embodiments of this disclosure are described in detail below with reference to the accompanying drawings.
[0027] Exemplary device This disclosure provides a target recognition model training device, which can be an electronic device with data processing capabilities. The electronic device includes, but is not limited to, in-vehicle computing devices, roadside computing devices, cloud servers, controllers, and edge computing nodes communicatively connected to the aforementioned devices. The electronic device is used to execute the training process of the target recognition model and to optimize the parameters of the target recognition model based on training samples.
[0028] In some embodiments, the in-vehicle computing device may be deployed in the vehicle, for example, integrated into the domain controller of an autonomous vehicle, a connected vehicle, or an Advanced Driver Assistance System (ADAS), to participate in the training process of the target recognition model.
[0029] Roadside computing devices can be roadside infrastructure, such as roadside units (RSUs) or other roadside edge computing nodes, used to participate in the training process of target recognition models.
[0030] Cloud servers can establish communication connections with vehicle-mounted computing devices, roadside computing devices, or edge computing nodes through communication networks to collect training data and perform centralized or distributed training of target recognition models.
[0031] The controller can be deployed in the robot system to perform local training of the target recognition model or participate in distributed training.
[0032] It should be understood that the devices listed above are merely examples and are not intended to be limiting.
[0033] like Figure 1 As shown, the target recognition model training device 100 includes a processor 101 and a memory 102.
[0034] The processor 101 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other types of processing chips. The processor 101 executes computer program instructions to train the initial recognition model to obtain a target recognition model.
[0035] The memory 102 is used to store computer program instructions and data, and the data may include training samples, images to be processed, and / or intermediate data. The memory 102 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, read-only memory, flash memory, or optical disk.
[0036] The processor 101 is configured to: acquire a training sample set; wherein the training sample set includes multiple training samples, and the training samples include sample images and the category labels of the target to be identified in the sample images, the confusion type labels corresponding to the category labels, and the confidence labels corresponding to the category labels; input the training samples into an initial recognition model to obtain the probability distribution of the target to be identified belonging to each candidate category; determine the prediction confidence of the predicted category based on the probability distribution; construct a target loss function based on the category labels, confusion type labels, confidence labels, predicted categories, and prediction confidence; and train the initial recognition model based on the target loss function to obtain a target recognition model.
[0037] The training sample set includes multiple training samples. Each training sample includes a sample image, along with its category label, confusion type label, and corresponding confidence label. The sample images can be in RGB, RAW, or YUV format, etc. Each sample image may contain one or more targets to be identified.
[0038] The training sample set can be pre-stored in memory 102, and the processor 101 can read the training sample set from memory 102.
[0039] The processor 101 can also obtain the training sample set required for this embodiment from other devices according to instructions or configuration information. This disclosure does not limit the specific method of obtaining the training sample set.
[0040] Based on the task type of the initial recognition model, sample images of different scenes can be obtained to construct a training sample set.
[0041] In this embodiment of the disclosure, the initial recognition model can be a deep learning network used to complete the target category recognition, such as a lightweight classification network, a deep convolutional neural network, or a Transformer-based visual classification network.
[0042] In a specific embodiment, the initial recognition model can be used to identify traffic light colors, traffic signs, or obstacles. The following example, using traffic light recognition, illustrates how the training sample set is constructed.
[0043] The sample images in the training samples can be collected from a variety of different scenes. These different scenes can include: different light intensities, different weather conditions, different times of day, and different occlusion scenarios.
[0044] Different lighting intensities: including but not limited to strong light, weak light, and backlighting scenarios. Among them, strong light scenarios may cause the sample image to be overexposed, weak light scenarios may cause the image to be insufficiently bright, and backlighting scenarios may cause the traffic light subject to be dark or produce a halo effect.
[0045] Different weather conditions: including but not limited to sunny days, rainy days, and foggy days. In rainy days, light is easily scattered by raindrops, causing a whitening effect. In foggy days, image contrast decreases, which may cause color distortion of traffic lights.
[0046] Different time periods: including but not limited to daytime and nighttime scenarios. In nighttime scenarios, traffic lights and surrounding ambient light sources mix, increasing interference factors for color recognition.
[0047] Different degrees of occlusion: including but not limited to partial occlusion and long-distance blurring scenarios. In partially occluded scenarios, the traffic light area is partially obscured by objects such as tree branches and vehicles. In long-distance blurring scenarios, the traffic light occupies fewer pixels in the image, resulting in loss of detail information.
[0048] By covering the above-mentioned multiple scenarios, the training sample set is made diverse and practical, so as to adapt to the complexity of autonomous driving scenarios and improve the generalization ability and robustness of the target recognition model obtained after training the initial recognition model in different environments.
[0049] The category label of the target to be identified in the training samples is used to identify the true category of the target to be identified in the sample image. The category label will be used as the ground truth during the initial training phase of the recognition model.
[0050] Category labels are set according to the type of target to be identified. Taking a traffic light as an example, the category label can be any one of red, yellow, or green lights. Taking a traffic sign as an example, the category label can be any one of various traffic signs such as speed limit signs, stop signs, and no-entry signs. Taking an obstacle as an example, the category label can be any one of different types of obstacles such as pedestrians, vehicles, and non-motorized vehicles. Category labels can be represented using one-hot encoding, where true categories are represented by 1, and other non-true categories by 0.
[0051] The confusion type label indicates the difficulty of identifying the target category in the sample image. The confusion type label can identify whether there is type confusion in the target category in the sample image, and if so, the specific type of confusion.
[0052] In some embodiments, the obfuscation type label may include no obfuscation and at least one obfuscation type. Taking an obfuscation type label including no obfuscation and two obfuscation types as an example, the type of the target to be identified includes type 1, type 2, and type 3, where type 1 and type 2 are easily obfuscated, type 3 and type 2 are easily obfuscated, and type 1 and type 3 are not easily obfuscated. In this case, the obfuscation type label may include: no obfuscation, type 1 and type 2 obfuscated, and type 3 and type 2 obfuscated. Taking color recognition as an example, the obfuscation type label includes: no obfuscation, color 1 and color 2 obfuscated, and color 3 and color 2 obfuscated. For a traffic light color recognition scenario, color 1 can be red, color 2 can be yellow, and color 3 can be green. Alternatively, if the target to be identified is an obstacle, the obfuscation type label includes: no obfuscation, type 1 and type 2 obfuscation, and type 3 and type 2 obfuscation. Type 1 is a cat, type 2 is a dog, and type 3 is a wolf.
[0053] It should be noted that the number of obfuscation types in the obfuscation type label can be flexibly determined based on the total number of categories of the target to be identified and the number of category pairs that are prone to confusion. No specific limit is set here.
[0054] In a specific embodiment, taking the identification of traffic light colors as an example, the confusion type labels include three types: no color confusion, red-yellow confusion, or yellow-green confusion. No color confusion indicates that the red, yellow, or green features of the traffic light image area in the sample image are clear, with obvious visual color differences, enabling accurate determination of the traffic light's color category. Red-yellow confusion indicates that due to environmental factors such as changes in lighting, weather interference, or obstructions, red and yellow in the sample image appear similar or indistinguishable visually (e.g., color, brightness). Yellow-green confusion indicates that due to environmental factors such as changes in lighting, weather interference, or obstructions, yellow and green in the sample image appear similar or indistinguishable visually (e.g., color, brightness).
[0055] The confidence labels in the training samples represent the degree of certainty regarding the category label. In scenarios where target types are easily confused, even human annotators may have ambiguity about the true category of the target to be identified in some sample images. The confidence labels are a quantitative record of this ambiguity. Confidence labels can be divided into multiple levels, such as high and low, or more levels can be divided according to actual needs, such as high, medium, and low. This disclosure does not limit the specific granularity of the confidence label division.
[0056] In a specific embodiment, taking the identification of traffic light colors as an example, the confidence label is used to represent the category label of the sample image obtained after color category annotation. This category label corresponds to the true color of the traffic light in the sample image, which can be any one of red, yellow, or green. Based on the degree of certainty of the category label of the sample image by human vision, a corresponding confidence label is set for the category label of the sample image. In some examples, the confidence label corresponding to the category label can be divided into three levels: high, medium, and low.
[0057] A high confidence level indicates that the human eye can readily identify the color of traffic lights in sample images when determining their category. For example, in sample images taken under clear, well-lit, and unobstructed conditions, the color of the traffic lights is clear, and the identification result is unique. A medium confidence level indicates that determining the category of traffic lights requires contextual information (such as adjacent frames and the physical location of the traffic lights) to assist in judging their color. For example, in images taken at dusk or in backlight, the red light appears yellowish, creating a visual bias. While this may lead to a category determination, environmental interference introduces uncertainty. A low confidence level indicates significant uncertainty in determining the color category, making it difficult to clearly distinguish the color of traffic lights in images. For example, in images taken during dense fog or heavy rain, the color characteristics of traffic lights are greatly weakened, and the boundaries between red / yellow and yellow / green categories are blurred, making accurate determination impossible.
[0058] It should be noted that training samples can also include the location of the target to be identified in the sample image. For example, when annotating the sample image, the location information of the region where the target to be identified is located can also be annotated. For example, a bounding box can be used to record the coordinates of the upper left and lower right corners of the region where the target to be identified is located, or the coordinates of the center point of the region where the target to be identified is located, as well as the width and height of the region. The target to be identified can be a traffic light, an obstacle, or something else.
[0059] In an optional embodiment, the training sample set can be divided into a training set, a validation set, and a test set according to a preset ratio, which are used for training, validating, and testing the initial recognition model, respectively. For example, the preset ratio can be 7:2:1, that is, the training set accounts for 70% of the total data, the validation set accounts for 20%, and the test set accounts for 10%. Of course, it is understood that those skilled in the art can adjust the above preset ratio according to the actual amount of data and task requirements, such as 6:2:2 or 8:1:1, etc., and this embodiment does not limit this.
[0060] It is understandable that for other classification and recognition-related tasks such as traffic sign recognition and obstacle classification, the construction method of the training sample set corresponding to the task of recognizing the color of traffic lights can be referenced, and the training sample set of the task can be constructed based on the same principle.
[0061] The training samples are input into the initial recognition model, which outputs the probability distribution of the target belonging to each candidate category. Each candidate category has the same label as the category in the training samples. For any candidate category, the higher the probability of that candidate category output by the initial recognition model, the greater the likelihood that the target belongs to that candidate category; conversely, the lower the probability of that candidate category output by the initial recognition model, the lower the likelihood that the target belongs to that candidate category.
[0062] When determining the prediction confidence of a prediction category based on a probability distribution, the candidate category corresponding to the highest probability in the probability distribution is determined as the prediction category, and this highest probability is the prediction confidence of the prediction category.
[0063] After obtaining the prediction confidence score, a target loss function is constructed based on the class label, confusion type label, confidence score label, predicted class, and prediction confidence score. The initial recognition model is then trained using this target loss function; that is, the loss value of the initial recognition model is calculated based on the target loss function, and the network parameters of the initial recognition model are updated using the backpropagation algorithm based on this loss value. This process of inputting training samples into the model and updating the network parameters is repeated until the model converges, resulting in the target recognition model.
[0064] It should be noted that the embodiments disclosed herein do not limit the form of the target loss function, as long as it can simultaneously constrain the probability distribution of the model's classification accuracy and output by combining the class label, confusion type label, and confidence label. The target loss function can be obtained by combining different types of loss functions. For example, the loss function includes an accuracy loss function and a confidence correction loss function.
[0065] The target recognition model training device provided in this embodiment annotates the samples with confusion types, and then iteratively trains the initial recognition model during the target recognition model training stage using a loss function constructed from category labels, confusion type labels, and confidence labels. This allows the training process to simultaneously consider classification accuracy and confidence accuracy, making the confidence probability distribution output by the trained target recognition model more accurate and improving the reliability of the target recognition model in target recognition. This avoids the situation where the target recognition model outputs a high-confidence incorrect category in visual confusion scenarios.
[0066] In some embodiments, the processor 101 executes instructions to construct a target loss function based on the class label, confusion type label, confidence label, predicted class, and predicted confidence, and is further configured to: construct an accuracy loss function based on the class label and the predicted confidence of the predicted class; construct a confidence correction loss function based on the confusion type label, confidence label, and probability distribution; and determine the target loss function based on the accuracy loss function and the confidence correction loss function.
[0067] In this embodiment, the accuracy loss function is used to constrain the model's output accuracy and improve the model's classification accuracy. The accuracy loss function can be constructed based on the difference between the class label and the predicted class output by the model. The confidence correction loss function is used to constrain the predicted confidence of the model's output, making it consistent with the confidence label. The confidence correction loss function can be constructed based on the difference between the confidence label and the predicted confidence of the model's output. The target loss function can be a weighted combination of the accuracy loss function and the confidence correction loss function, and the specific form of the target loss function can be set according to actual training needs. Therefore, the processor 101 can simultaneously introduce the accuracy loss function and the confidence correction loss function into the target loss function, enabling the model to simultaneously optimize classification accuracy and confidence reliability during the training phase. This reduces the initial recognition model from outputting high-confidence incorrect classes in scenarios where target types are easily confused, thereby improving the model's recognition accuracy.
[0068] In some embodiments, the accuracy loss function constructed based on the class labels and the probability distribution of multiple candidate classes may include a focal loss function. The focal loss function adds dynamic weights to the standard cross-entropy loss function, causing the model to pay less attention to easy samples and more attention to difficult or easily confused samples, thereby mitigating the problem of sample or class imbalance.
[0069] In a specific embodiment, the training samples are divided into red light samples, green light samples, and yellow light samples according to the color of the traffic lights in the sample images. Red light samples refer to training samples where the category label of the traffic lights in the sample image is marked in red; green light samples refer to training samples where the category label of the traffic lights in the sample image is marked in green; and yellow light samples refer to training samples where the category label of the traffic lights in the sample image is marked in yellow. When identifying the color of traffic lights, there are two main problems: 1. The number of samples for different color categories may be unbalanced. For example, there may be more red light samples and fewer yellow light samples; or more yellow light samples and fewer red light samples. 2. The difficulty of identifying the color category of traffic lights varies in different sample images. For example, the difficulty of identifying the color category of traffic lights in sample images with confusing colors is higher than the difficulty of identification in images with clear images.
[0070] In related technologies, the cross-entropy loss function does not distinguish between all training samples, applies equal weights to the loss contribution of each training sample, and does not consider the sample category, making it difficult to balance the number of samples in different categories. For example, when the number of samples in different color categories is unbalanced, the gradient of the loss function will skew towards the color category with a large number of training samples. In addition, it is difficult to focus on training on difficult samples (training samples with blurred visual features and high discrimination difficulty). For example, applying equal supervision to easy samples with high confidence and difficult samples with low confidence makes it impossible for the model to accurately identify difficult samples with blurred features and easy confusion.
[0071] In this embodiment, a dynamic weight term is added to the cross-entropy loss function, including: introducing class weights when the number of training samples of different classes is imbalanced. By applying different loss weights to different categories, the difference in the number of training samples for each category can be balanced; a modulation factor can also be introduced. By automatically reducing the loss weight of easily classified training samples, the model training focuses on difficult samples. The accuracy loss function constructed based on this method allows the initial recognition model to pay more attention to learning easily confused difficult samples, improving the model's recognition accuracy in complex scenarios.
[0072] The accuracy loss function can be constructed based on the class labels, the probability distribution of multiple candidate classes predicted by the model, the class weights used to balance class imbalance, and the focus parameters used to focus on training hard examples.
[0073] The formula for calculating the accuracy loss function is shown in Formula (1) below.
[0074] Formula (1): .
[0075] In formula (1), The category index corresponding to the category label in the sample image. The category corresponding to the category label Category weights, The probability value corresponding to the category output by the initial recognition model; This is a focusing parameter used to reduce the weight of easy examples and focus training on difficult examples. The value of can be determined based on the actual situation, for example... Option 2 is acceptable. The value can be adaptively adjusted according to the category distribution of the training sample set. Specifically, it can be determined by referring to the conventional category weight setting method.
[0076] It should be noted that, for traffic light color recognition scenarios, easily confused samples include training samples in color confusion scenarios such as red-yellow confusion and yellow-green confusion; the category label can include red light, yellow light, or green light. =1 represents a red light. =2 represents a yellow light. =3 represents a green light. Of course, the physical meaning of each parameter in the above accuracy loss function can be adaptively adjusted based on the type of target recognition task.
[0077] Processor 101 uses a focus loss function to constrain the classification accuracy of the model. By introducing class weights and focus parameters, it effectively alleviates the problem of imbalanced samples and focuses on training difficult samples, thereby improving the target recognition accuracy of the model in visually confusing scenarios and avoiding high confidence in outputting incorrect categories.
[0078] In some embodiments, the processor 101 executes instructions to construct a confidence correction loss function based on the confusion type label, the confidence label, and the probability distribution, and is further configured to: perform label smoothing on the confidence label based on the confidence label and the confusion type label to obtain a smoothed confidence label; construct a first loss component based on the smoothed confidence label and the probability distribution; and construct the confidence correction loss function based on the first loss component.
[0079] In scenarios where the categories of the target to be identified are easily confused, models trained based on traditional hard labels (one-hot labels) set the values corresponding to the true categories of the training samples to 1, and the values corresponding to the other categories to 0. The model training process uses the labeling results of the hard labels as the fitting target.
[0080] For example, traffic light category labels have three possible outcomes: red, yellow, and green. Therefore, a set of three numbers is used to correspond to these three categories, with the position order fixed as [red, yellow, green]. In the vector [1, 0, 0], the first position is 1, and the other two positions are 0, representing a red light for this training sample; in the vector [0, 1, 0], the second position is 1, and the other two positions are 0, representing a yellow light for this training sample; and in the vector [0, 0, 1], the third position is 1, and the other two positions are 0, representing a green light for this training sample. The rule is: write 1 in the position corresponding to the actual category of the training sample, and write 0 in all other category positions. This is the one-hot hard labeling rule.
[0081] In scenarios where the types of targets to be identified are easily confused, even with manual annotation, various environmental factors may lead to ambiguity regarding the target category in some sample images. For example, in the scenario of recognizing the color of traffic lights, if a sample image's light color appears to be between red and yellow due to lighting or occlusion, based on the aforementioned one-hot hard labeling rules, it may be difficult to uniquely determine its true category during annotation. However, hard labels require definitive labeling as one of red, yellow, or green, which may result in the annotation results failing to accurately reflect the true visual state of the sample image. This forces the model to learn a clear judgment boundary for ambiguous or easily confused samples. This approach may cause the model to output high confidence in the wrong category when faced with easily confused samples. To address this, embodiments of this disclosure introduce label smoothing processing, converting hard labels (e.g., [1, 0, 0]) into smooth confidence labels (e.g., [0.9, 0.05, 0.05]) that reflect annotation uncertainty, to simulate the category uncertainty that may exist during the annotation process. Supervised training of the model is then conducted using smooth confidence labels, improving the model's output accuracy.
[0082] Label smoothing refers to smoothing the probability distribution of confidence labels to reduce overfitting to a single class during model training, thereby improving the model's generalization ability and robustness. Specifically, label smoothing is performed on the confidence labels corresponding to training samples to adjust the label values corresponding to the target class from a first value to a second value less than the first value, and adjust the label values corresponding to the remaining classes to a third value greater than zero, thus generating smooth confidence labels; the target recognition model is then trained based on these smooth confidence labels.
[0083] The process of converting hard labels into smooth confidence labels by label smoothing of the confidence labels of sample images can be executed by processor 101. Processor 101 is configured to: obtain the confusion type label and confidence label of the sample image; keep the confidence label unchanged if the confusion type label is unconfined; and perform label smoothing on each confidence label according to the different confidence labels in the sample if the confusion type label is easily confused.
[0084] In a specific embodiment, taking the traffic light color recognition scenario as an example, the confusion type labels include: red-yellow confusion, yellow-green confusion, and no confusion. Therefore, the aforementioned easily confused categories can include red-yellow confusion or yellow-green confusion. The confidence level label is high, medium, or low.
[0085] In one example, label smoothing includes: when the confusion type label is easily confused between the first color category and the second color category, if the category label is the first color category and the confidence label corresponding to the category label is high, then a smoothed confidence label is generated based on the level information of the confidence label. This smoothed confidence label contains the label value distribution corresponding to each category. The label value corresponding to the category refers to: constructing a one-dimensional vector for multi-category partitioning, where each position in the vector uniquely corresponds to a category, and the value filled in at each position is the label value corresponding to that category. This vector represents the weighted distribution of training samples to each category.
[0086] The process of generating smooth confidence labels and configuring the distribution of label values for each category is as follows: The initial label value for the first color category is a first value (1 in the case of hard labels), and this first value is adjusted to a second value that is less than the first value; the label value for the second color category is adjusted to a third value that is greater than 0, and the label values for the remaining color categories are set to 0. The one-dimensional vector formed by the transformed label values of all color categories is determined as the smooth confidence label.
[0087] When the confidence level is high, the second value is very close to the first value; for example, the first value is 1, the second value is 0.95, and the third value is 0.05. If the confidence level of the above category label is medium, the difference between the second and first values will be larger than when the confidence level of the category label is high; for example, the first value is 1, the second value is 0.75, and the third value is 0.25. If the confidence level is low, the difference between the second and first values can be further increased compared to when the confidence level of the category label is medium. For example, the second value can be adjusted to 0.6, and the third value can be adjusted to 0.4, and so on. The magnitude of the difference between the second and first values can be adaptively adjusted according to the level of the confidence level. The smoothed confidence level label obtained after this adjustment can accurately reflect the category uncertainty inherent in the training sample itself, match the actual judgment result of the annotator on the sample, and allow the model to learn the judgment logic that conforms to the actual scenario.
[0088] The following example uses traffic light color recognition as an illustration, and Table 1 shows examples of smooth confidence label values corresponding to different confidence labels.
[0089] Table 1
[0090] In Table 1 above, P1, P2, P3, and P4 satisfy the following conditions: P1 > P2 and P3 > P4, meaning the label value of the category corresponding to the category label is greater than the label value of the category that is easily confused with that category; P1 > P3 and P2 < P4, meaning the lower the confidence level, the lower the label value of the category corresponding to the category label, and the higher the label value of the category that is easily confused with that category; P1 + P2 = 1 and P3 + P4 = 1, meaning the sum of the label values of the two easily confused categories is 1, and the label values of the other categories are fixed at 0, which conforms to the characteristics of probability distribution and can intuitively match the actual judgment tendency of the target category to be identified during the category labeling process.
[0091] In some embodiments, constructing a first loss component based on the probability distribution of smoothed confidence labels and model output may include: constructing a cross-entropy loss function based on smoothed confidence labels and probability distribution.
[0092] When constructing the confidence correction loss function based on the first loss component, the cross-entropy loss function constructed based on the smoothed confidence label and probability distribution can be determined as the confidence correction loss function.
[0093] After generating smoothed confidence labels based on the confidence labels of the sample images, the cross-entropy loss is calculated by comparing the smoothed confidence labels with the probability values of each category output by the model. Cross-entropy loss This is used to constrain the probability values of each category output by the recognition model for a sample image, so that the probability values of each category are consistent with the label values of each type of the sample image.
[0094] Cross-entropy loss The calculation formula is shown in formula (2) below.
[0095] Formula (2): .
[0096] In formula (2), As a category, Each corresponds to a different category. To smooth the label values of class c in the confidence labels, This represents the probability value for class c predicted by the model.
[0097] By minimizing the above cross-entropy loss The target recognition model's output category probability distribution is fitted to the label value distribution represented by the smooth confidence labels. As the confidence level of the training samples' confidence labels decreases, the label value of the category corresponding to the category label gradually decreases, while the label value of easily confused categories gradually increases. This allows the model to learn the degree of category uncertainty of different samples during training, and no longer forcibly fits the deterministic hard labels. This avoids the model outputting excessively high prediction confidence for easily confused categories on easily confused samples, thereby preventing the easily confused categories from being misclassified as the true categories in post-processing and improving the accuracy of the model output.
[0098] In some embodiments, when constructing the confidence-corrected loss function based on the first loss component, a second loss component can also be constructed based on the confidence label and the predicted confidence; then, the confidence-corrected loss function is constructed based on the first loss component and the second loss component.
[0099] The second loss component can be a distribution consistency constraint loss function constructed based on confidence labels and predicted confidence, to characterize the distribution difference between predicted confidence and confidence labels in the same training batch, thereby constraining the predicted confidence of the model output in the dimension of training batch, so that the statistical distribution of predicted confidence in a training batch output by the model approaches the distribution of manually labeled confidence labels.
[0100] Specifically, the first loss component (cross-entropy loss) only constrains the probability values of each category in a single sample image to fit their corresponding smooth confidence labels, without considering the overall consistency of the confidence distribution among different sample images within a training batch. Specifically, within a single training batch, each level of the confidence label has an initial proportion distribution, while the high, medium, and low confidence samples predicted by the model will form a different proportion distribution than the initial distribution, and the two proportion distributions are prone to significant deviation. For example, in a certain training batch, samples with high, medium, and low confidence labels each account for 1 / 3 of the total samples in that batch, but more than 1 / 2 of the samples in that batch are judged as high-confidence samples by the model, causing the proportion of high-confidence samples predicted by the model to be much higher than the proportion of samples with high confidence labels, interfering with the model's optimization direction and hindering stable training. To solve this problem, embodiments of this disclosure further introduce a second loss component to constrain the consistency between the probability value distribution of each predicted category in all sample images of the current training batch and the corresponding smooth confidence label distribution.
[0101] In some embodiments, the second loss component construction process is as follows: processor 101 executes instructions to construct a second loss component based on confidence labels and predicted confidence, including: identifying multiple target training samples in the same training batch whose confusion type label is the target confusion type; determining the confidence label probability distribution of multiple target training samples in the training batch based on the confidence labels of the category labels of the multiple target training samples; determining the weight of each confidence level of each target training sample based on a preset threshold and the predicted confidence of the multiple target training samples; determining the confidence prediction distribution of each confidence level based on the weight of each confidence level of each target training sample; and constructing the second loss component based on the confidence label probability distribution and the confidence prediction distribution.
[0102] The target confusion type is a pre-defined combination of easily confused categories. For example, in the scenario of traffic light color recognition, confusion types are divided into three categories: no confusion, red-yellow confusion, and yellow-green confusion. In some examples, because red-yellow confusion is prone to causing serious recognition misjudgments at the vehicle driving level, it can be designated as the target confusion type. In some examples, in addition to red-yellow confusion, yellow-green confusion can also cause autonomous driving perception misjudgments; therefore, yellow-green confusion can also be classified together with red-yellow confusion as a target confusion type.
[0103] In a specific embodiment, if the confidence labels include three categories: high, medium, and low, for multiple target training samples in the same training batch whose confusion type label is the target confusion type, the target training samples with high confidence labels among the multiple target training samples are determined as the first training sample, the target training samples with medium confidence labels among the multiple target training samples are determined as the second training sample, and the target training samples with low confidence labels among the multiple target training samples are determined as the third training sample. The first number of the first training sample, the second number of the second training sample, and the third number of the third training sample among the multiple target training samples are counted. Then, the first ratio of the first number to the total number of multiple target training samples, the second ratio of the second number to the total number, and the third ratio of the third number to the total number are obtained (the sum of the first ratio, the second ratio, and the third ratio is 1).
[0104] Then, after obtaining the first, second, and third ratios, a first threshold is pre-set to divide the target training sample's prediction confidence into high and medium confidence levels, and a second threshold is set to divide the target training sample's prediction confidence into medium and low confidence levels. Based on the first threshold and the target training sample's prediction confidence, a weight representing the target training sample approaching high confidence is generated; based on the second threshold and the target training sample's prediction confidence, a weight representing the target training sample approaching low confidence is generated. For any target training sample, the sum of the weights representing the target training sample approaching high confidence, the weights representing the target training sample approaching medium confidence, and the weights representing the target training sample approaching low confidence is 1. Therefore, after obtaining the weights representing the target training sample approaching high and low confidence, the weight representing the target training sample approaching medium confidence can be calculated. The above weights are calculated only based on the target training sample's prediction confidence, without using the target training sample's confidence label in the calculation. The first and second thresholds shall be set by those skilled in the art based on the actual situation.
[0105] To determine the confidence prediction distribution for each confidence level based on the weights of each confidence level for each target training sample, the weights of all target training samples approaching high confidence are averaged to obtain the mean of high confidence in that training batch. Similarly, the weights of all target training samples approaching medium confidence are averaged to obtain the mean of medium confidence in that training batch. The weights of all target training samples approaching low confidence are averaged to obtain the mean of low confidence in that training batch. These three means together constitute the confidence prediction distribution for each confidence level in that training batch.
[0106] Finally, a second loss component is constructed based on the confidence label probability distribution and the confidence prediction distribution within the same training batch. The KL divergence can be used to calculate the distributional difference between the confidence label probability distribution and the confidence prediction distribution, and the calculated KL divergence value is used as the second loss component.
[0107] In a specific embodiment, taking the traffic light color recognition scenario as an example, the second loss component construction process is as follows.
[0108] The first step is to identify multiple target samples in the current training batch that are labeled as red-yellow confusion. Then, count the proportion of the first sample with the highest confidence label among these multiple target samples out of the total number of samples in the multiple target samples. The proportion of the second sample in the confidence-labeled sample to the total number of samples in the multiple target samples. The proportion of the number of third samples with low confidence labels to the total number of samples in the multiple target samples. The confidence label probability distribution of this training batch is obtained. , Satisfy normalization constraints: Among them, the confidence label probability distribution The confidence label is statistically obtained based on the confidence labels of samples in the current training batch and is used as a supervisory signal to constrain the output of the model.
[0109] The second step involves setting the first and second thresholds to 0.75 and 0.5, respectively. Based on the relationship between the predicted confidence of the target training sample and the first and second thresholds, the confidence level is determined. A predicted confidence greater than the first threshold is considered high confidence; a predicted confidence less than or equal to the first threshold but greater than the second threshold is considered medium confidence; and a predicted confidence less than or equal to the second threshold is considered low confidence. For each sample image within the training batch, after processing by the initial recognition model, the probability distribution of each category is then used to determine the maximum probability in that distribution. The Sigmoid differentiable activation function is used to maximize the probability of the sample image. The weights of the sample image are mapped to the three levels of low confidence, medium confidence, and high confidence, respectively. This mapping method avoids the gradient breakage problem caused by the traditional hard threshold segmentation method, and makes the confidence level allocation operation differentiable throughout the process, allowing it to participate in end-to-end gradient backpropagation training with the model. The mapping process is shown in the following formula (3).
[0110] Formula (3): .
[0111] In formula (3), To maximize the probability of the sample The weights with low confidence are obtained by subtracting from the second threshold and then applying a Sigmoid mapping. To maximize the probability of the sample The high-confidence weights are obtained by subtracting the first threshold and then mapping the result to a Sigmoid function. The weights are for the medium confidence level; It is a differentiable activation function of the sigmoid class. This is the steepness coefficient, used to control the smoothness of the mapping curve; a value of 15 can be selected. This is a second threshold (e.g., 0.5) used to distinguish between low and medium confidence levels. The first threshold (e.g., 0.75) is used to distinguish between medium and high confidence levels.
[0112] The third step is to construct the confidence prediction distribution for the current training batch. Based on the calculation results of the second step, the low-confidence weights of all target training samples within the training batch are averaged to obtain the biased mean of low confidence for that training batch; the medium-confidence weights of all target training samples within the training batch are averaged to obtain the biased mean of medium confidence for that training batch; and the high-confidence weights of all target training samples within the training batch are averaged to obtain the biased mean of high confidence for that training batch. The biased means corresponding to low confidence, medium confidence, and high confidence respectively collectively constitute the confidence prediction distribution of that training batch. .
[0113] The calculation method is shown in the following formula (4).
[0114] Formula (4): .
[0115] In formula (4), N is the number of target samples in the current batch, and N is a positive integer; This represents the propensity mean for low confidence levels in the current training batch. This represents the propensity mean of the median confidence level for this training batch. This represents the propensity mean with high confidence for this training batch.
[0116] Step 4: Calculate the KL divergence loss for the current training batch. This is achieved by minimizing the relative entropy between the built-in confidence label probability distribution of the current training batch and the confidence prediction distribution of the model within the current training batch, thus constraining the consistency between the global prediction confidence distribution of the model and the sample confidence label distribution. The calculation method is shown in the following formula (5).
[0117] Formula (5): .
[0118] In formula (5), KL divergence constraints are used to calculate the confidence label probability distribution. With prediction confidence distribution The differences between them; k is the index of confidence level, representing low confidence (low), medium confidence (mid), and high confidence (high), respectively. Let be the probability value of the k-th level in the confidence label probability distribution. To predict the probability value of the k-th level in the confidence distribution.
[0119] Based on the cross-entropy loss obtained above KL divergence constraint Confidence-corrected loss function The calculation method is shown in the following formula (6).
[0120] Formula (6): .
[0121] In formula (6), These are the weighting coefficients for the KL divergence constraint, used to adjust the strength of the KL divergence constraint. It can be adaptively adjusted according to the model training effect (the value range is 0.1~1.0).
[0122] When constructing a confidence correction loss function based on the first and second loss components, the first and second loss components can be weighted and combined to obtain the final confidence correction loss function. The first loss component constrains the confidence distribution of the initial recognition model on a single sample, while the second loss component constrains the distribution of predicted confidence in the same training batch to be consistent with the confidence label distribution of that batch. This avoids the probability shift problem where the initial recognition model outputs an overall higher predicted probability after processing a large number of simple samples and an overall lower predicted probability after processing difficult samples. Global calibration of predicted confidence is achieved from both the single-sample and batch dimensions.
[0123] In some embodiments, the processor 101 executes instructions to determine a target loss function based on a precision loss function and a confidence correction loss function, and is further configured to: determine a first weighting coefficient of the precision loss function and a second weighting coefficient of the confidence correction loss function based on a confusion type label, a confidence label, a class label, and a predicted class; and determine the target loss function by performing a weighted sum of the precision loss function and the confidence correction loss function based on the first weighting coefficient and the second weighting coefficient.
[0124] For the precision loss function and confidence-corrected loss function The weighted combination yields the total loss function of the initial recognition model. The calculation method is shown in the following formula (7).
[0125] Formula (7): .
[0126] In formula (7), Weighting for accuracy loss. To adjust the loss weights for confidence levels, .
[0127] The first weighting coefficient can be the precision loss weight. The second weighting coefficient can be the confidence-corrected loss weight. .in, , Instead of globally fixed parameters, for each training sample, a corresponding set of parameters is dynamically configured independently based on the scene to which the sample belongs. , Values.
[0128] In some embodiments, the processor 101 executes instructions to determine a first weight coefficient of the accuracy loss function and a second weight coefficient of the confidence correction loss function based on the confusion type label, confidence label, category label, and predicted category. This is further configured to: adjust the second weight coefficient to be greater than the first weight coefficient when the confusion type label indicates that the sample image belongs to the target confusion type, the confidence label does not exceed a preset level, and the category label is inconsistent with the predicted category; and adjust the first weight coefficient to be greater than the second weight coefficient when the confusion type label indicates that the sample image does not belong to the target confusion type, the confidence label exceeds a preset level, or the category label is consistent with the predicted category.
[0129] "Confidence label not exceeding the preset level" means that the confidence label indicates a level that is either the preset level or lower. For example, if the confidence label levels include high, medium, and low, and the preset level is medium, then "confidence label not exceeding the preset level" means that the confidence label is medium or low; "confidence label exceeding the preset level" means that the confidence label is high.
[0130] The target confusion type refers to the visual confusion present in the sample image, and it belongs to a preset confusion type. The preset confusion type can be set according to the actual situation. For example, to address types that are easily confused by the target's category, type 1 and type 2 confusion can be set as the target confusion type; similarly, to address types that are easily confused by the target's color, color 1 and color 2 confusion can be set as the target confusion type. Specifically, taking traffic light color recognition as an example, the target confusion type could be red-yellow confusion.
[0131] Taking the confidence level as having three levels—high, medium, and low—and setting the default level as medium, if the confusion type label indicates that the sample image belongs to the target confusion type, the confidence label is medium or low, and the category label is inconsistent with the predicted category, then the sample image is a sample that is easily confused and misclassified. In this case, the confidence correction has a higher priority, that is, the constraint effect of the confidence correction loss is strengthened, and the second weight coefficient is adjusted to be greater than the first weight coefficient. Conversely, if the confusion type label indicates that the sample image does not belong to the target confusion type, the confidence label is high, or the category label is consistent with the predicted category, then the sample is not a sample that is easily confused and misclassified. In this case, the constraint effect of the accuracy loss can be strengthened, and the first weight coefficient is adjusted to be greater than the second weight coefficient.
[0132] In this embodiment, when the confusion type label indicates that the sample image does not belong to the target confusion type, it means that the sample image itself is not a sample that is easily confused and misclassified. Confidence correction is no longer the main concern, and the priority is to constrain the accuracy loss to optimize the classification accuracy. When the confidence label is high, it means that the true category of the sample image is highly recognizable, and the model output probability itself matches the actual classification accuracy well. Confidence correction is not the focus of optimization, and the priority is to constrain the accuracy loss to correct the classification result. When the category label is consistent with the predicted category, it means that there is no misclassification due to type confusion. In this case, the priority is to ensure accuracy, and the confidence correction constraint is placed in a secondary position.
[0133] It should be noted that the first and second weighting coefficients mentioned above are not fixed, but are dynamically adjusted based on the category label of each sample image in each training batch, the predicted category output by the initial recognition model, and the confidence level of the predicted category. During the training of each batch, the processor 101 iterates through each sample image within the batch, determines the first weight of the accuracy loss function and the second weight of the confidence correction loss function corresponding to that sample according to the aforementioned weighting adjustment method, and then calculates the weighted sum of the losses of all samples in the batch to obtain the total loss for that batch. This sample-by-sample dynamic weighting strategy enables the initial recognition model to adaptively balance the relationship between recognition accuracy and confidence correction in different scenarios.
[0134] In this embodiment of the disclosure, by determining a first weighting coefficient and a second weighting coefficient, and by performing a weighted summation of the accuracy loss function and the confidence correction loss function based on the first weighting coefficient and the second weighting coefficient, a target loss function is determined. This enables differentiated loss constraints for training samples with different types of confusion: for training samples without confusion, the weight of the confidence correction loss can be reduced, focusing more on the model's learning of category recognition accuracy; for training samples with confusion, the weight of the confidence correction loss is increased, strengthening the correction effect on the model. Thus, without sacrificing the sample recognition accuracy, the problem of an overall high model probability on easily confused samples is effectively reduced.
[0135] Taking the scenario of traffic light color recognition, where the target confusion type is red-yellow confusion or yellow-green confusion, as an example, this paper explains in detail how to adjust the accuracy loss weight and confidence correction loss weight: First, obtain the class label, confusion type label, confidence label, and consistency between the predicted class and the class label for the sample image. If the confusion type label is red-yellow confusion or yellow-green confusion, the confidence label is medium or low, and the predicted class is inconsistent with the class label, adjust the confidence correction loss weight and the accuracy loss weight so that the confidence correction loss weight is greater than the accuracy loss weight. For example, set... =0.3, =0.7. When the confusion type label is not red-yellow or yellow-green confusion, the confidence label is high, or the predicted category matches the category label, adjust the confidence correction loss weight and the accuracy loss weight so that the confidence correction loss weight is less than the accuracy loss weight. For example, =0.7, =0.3. Of course, the weight coefficients can be adjusted adaptively based on the actual training effect of the model, so as to improve the reliability of the target recognition model's output confidence while ensuring recognition accuracy.
[0136] In an optional embodiment, after obtaining the total loss function, the processor 101 is configured to: correct the prediction confidence of the initial recognition model output based on the total loss function and the sample training set, to obtain a trained target recognition model.
[0137] In a specific embodiment, during the training of the initial recognition model, the processor 101 uses a validation set to monitor the accuracy of the predicted class and the reliability of the predicted confidence in real time. In some examples, the reliability of the predicted confidence can be evaluated using the Expected Calibration Error (ECE). The lower the ECE value, the higher the consistency between the predicted confidence output by the initial recognition model and the accuracy of the predicted class.
[0138] In this embodiment of the disclosure, when calculating the total loss of a single training sample, the first weight of the accuracy loss function and the second weight of the confidence correction loss function are adaptively configured based on the matching of the confusion type label, confidence label, category label and prediction result. This makes the confidence correction focus on training samples that are prone to confusion, while other training samples focus on classification accuracy, thus taking into account both classification accuracy and confidence calibration effect.
[0139] The following uses the identification of traffic light colors, with the initial identification model being the initial color identification model, as an example to illustrate the target recognition model training device of this disclosure through a specific embodiment.
[0140] First, a pre-constructed training sample set is read or received. This training sample set includes multiple sample images, each labeled with a color category label (representing the color class), a color confusion type label (representing the color confusion type), and a color confidence label (representing the confidence level for the labeled color class). The sample images from the training sample set are then input into an initial color recognition model to obtain the color category probability distribution information output by the initial color recognition model (corresponding to the probability distribution of the target to be recognized belonging to each candidate category). Based on the color category probability distribution information, the predicted confidence level of the color recognition result is determined, and the color recognition result is the color category probability score. The color with the highest probability in the information is identified. Based on the color category label, color confusion type label, color confidence label, color recognition result, and the predicted confidence of the color recognition result of the sample image, an accuracy loss function and a confidence correction loss function are constructed, and the accuracy loss weight and confidence correction loss weight are determined. The total loss is calculated based on the accuracy loss function, the confidence correction loss function, the accuracy loss weight, and the confidence correction loss weight. The model parameters of the initial color recognition model are updated based on the total loss function. While optimizing the color classification accuracy, the confidence of each candidate category output by the initial color recognition model is corrected to obtain the trained color recognition model.
[0141] The initial color recognition model in this disclosure can be any model that can recognize colors. This disclosure only provides examples through the model structure described below. It is understood that the initial color recognition model in this disclosure can also be other model structures besides the model structure shown below, and this disclosure does not limit it in this regard.
[0142] In some examples, for traffic light color recognition scenarios, the initial recognition model in this disclosure embodiment may also be referred to as an initial color recognition model. In a possible implementation, such as Figure 2 As shown, Figure 2 This is a schematic diagram of the network structure of an initial color recognition model provided in this disclosure. The initial color recognition model includes a feature extraction module and a color classification module.
[0143] The feature extraction module is used to extract feature information of the traffic light region in the input sample image. The feature information includes deep semantic features and shallow texture features. The feature extraction module can include ResNet50 network, VGG16 network, EfficientNet, or other convolutional neural network structures with multi-scale feature extraction capabilities.
[0144] The color classification module is used to transform and fuse the color category space dimension based on the feature information output by the feature extraction module, so as to output the color category, such as red, yellow or green. The color classification module includes a linear network connection layer (i.e., a fully connected layer) and a normalization layer.
[0145] Specifically, the linear network connection layer is used to perform dimensionality transformation and feature fusion on the feature information, mapping the feature information to raw scores corresponding to each color category. The magnitude of the raw score reflects the model's relative bias towards each category; a larger value indicates stronger evidence supporting that category. The normalization layer is used to convert the raw scores into probability distributions for each color category. This probability distribution can be represented as a three-dimensional probability vector. , , , These represent the probability values for a traffic light to be classified as red, yellow, or green, respectively. , , , Based on the probability distribution of each color category, the color category corresponding to the maximum probability value is selected, and this color category is output as the color recognition result. For example, if... If so, a red light will be output.
[0146] The embodiments disclosed herein have the following beneficial effects: (i) Dataset annotation is more targeted This disclosure provides precise supervision information for model confidence correction by additionally labeling confusion type and confidence level during the dataset construction phase, thus solving the problem that existing datasets cannot support confidence correction training. Based on the dataset provided by this disclosure, the model can more accurately identify targets in real-world driving scenarios where complex factors such as lighting changes and weather interference cause confusion in target category colors or similar visual features.
[0147] (ii) The confidence level correction effect is significant. The target loss function constructed based on the confidence correction loss function provided in this disclosure can effectively constrain the model's confidence output in scenarios where target categories are easily confused. This avoids the model outputting high-confidence incorrect categories in visually confusing scenarios, ensuring that the model's output confidence matches the actual recognition accuracy and significantly improving the model's reliability in complex scenarios. Taking the traffic light recognition scenario as an example, the confidence correction loss function can effectively constrain the model's confidence output in scenarios such as red and yellow confusion, preventing the model from outputting high-confidence incorrect categories when colors are confused.
[0148] (III) Dual Improvement in Recognition Accuracy and Confidence Reliability This disclosure employs a dynamically weighted total loss function, which adaptively adjusts the weight coefficients of the accuracy loss function and the confidence correction loss function according to different scenarios. For clear samples or samples correctly classified by the model, the accuracy loss dominates, ensuring high recognition accuracy. For samples with visually confusing features and low labeling certainty, the confidence correction loss dominates, constraining the model's confidence output. Through this dynamic weighting strategy, this disclosure achieves a dual improvement in recognition accuracy and confidence reliability, making it particularly suitable for applications with extremely high safety requirements, such as intelligent driving and assisted driving. Taking traffic light recognition as an example, this disclosure can effectively constrain the model's confidence output in scenarios such as red-yellow confusion, preventing the model from outputting high-confidence incorrect categories when red and yellow are confused. This prevents dangerous driving behaviors such as running red lights and wrong parking caused by misidentification. While ensuring the accuracy of traffic light color recognition, it effectively improves the reliability of the model's output confidence, providing more reliable recognition results for autonomous driving decision-making and further ensuring driving safety.
[0149] Figure 3 This is a schematic diagram of the structure of a target recognition device provided in an embodiment of this disclosure. Figure 3 As shown, the target recognition device 300 includes a processor 301 and a memory 302.
[0150] The processor 301 may be, for example, a CPU, GPU, NPU, ASIC, FPGA or other types of processing chip. The processor 301 executes computer program instructions to acquire the image to be recognized and inputs it into the target recognition model. Based on the probability distribution of each category output by the target recognition model, the target category and prediction confidence are determined.
[0151] Memory 302 is used to store computer program instructions and data. Memory 102 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, read-only memory, flash memory, or optical disk.
[0152] The processor 301 is configured to acquire an image to be recognized; input the image to be recognized into a target recognition model to obtain the probability distribution of each candidate category of the image to be recognized output by the target recognition model; wherein the target recognition model is trained based on the aforementioned target recognition model training device.
[0153] The image to be identified can include images from any classification task. For example, in an autonomous driving scenario, the image to be identified could be a road scene image including traffic lights captured by an onboard camera, where the target category is the color category of the traffic lights, and the target prediction confidence is the confidence level of the color category. Similarly, in an obstacle detection scenario for autonomous driving, the image to be identified could be a road scene image captured by an onboard camera, where the target category is the obstacle category, and the target prediction confidence is the confidence level of the obstacle category output by the model.
[0154] In some embodiments, the processor 301, through the target recognition model trained with confidence correction described above, can not only improve the accuracy of easily confused category recognition and reduce the probability of misjudgment of confused categories, but also output a confidence value that matches the actual recognition accuracy of the model, thereby improving the accuracy of the model's own classification and the accuracy of the confidence value. For example, after obtaining the traffic light color recognition result and the corresponding confidence value, the processor 301 will send this information to the vehicle's decision control module, which will then combine it with other perception information to determine the vehicle's passage strategy. When the confidence value of the recognition result is high, the vehicle will be controlled to pass directly according to the recognition result; when the confidence value of the recognition result is low, a safety reminder for the vehicle will be triggered.
[0155] Based on the target recognition device provided in this disclosure, the target recognition model trained by confidence correction has higher prediction confidence accuracy and can effectively suppress confidence errors. For easily confused categories with similar colors and overlapping features, it can reduce the probability of category misclassification, improve classification accuracy, optimize the model's own classification and discrimination capabilities, and improve the overall robustness of the model.
[0156] Exemplary methods Figure 4 This is a schematic flowchart illustrating a target recognition model training method provided in an exemplary embodiment of this disclosure. This method can be applied to the aforementioned target recognition device or electronic device. Figure 4 As shown, the method includes the following steps S410~S440.
[0157] S410: Obtain the training sample set; wherein, the training sample set includes multiple training samples, and the training samples include sample images and the category labels of the target to be identified in the sample images, the confusion type labels corresponding to the category labels, and the confidence labels corresponding to the category labels.
[0158] S420: Input the training samples into the initial recognition model to obtain the probability distribution of the target to be recognized belonging to each candidate category.
[0159] S430: Determine the prediction confidence level for the prediction category based on the probability distribution.
[0160] S440: Construct a target loss function based on category label, confusion type label, confidence label, predicted category and predicted confidence. Train the initial recognition model based on the target loss function to obtain the target recognition model.
[0161] The target recognition model training method provided in this disclosure, after labeling the samples with confusion types, iterates the initial recognition model during the target recognition model training stage by constructing a loss function based on category labels, confusion type labels, and confidence labels. This allows the training process to simultaneously consider classification accuracy and confidence accuracy, making the confidence probability distribution output by the trained target recognition model more accurate and improving the reliability of the target recognition model in target recognition. This avoids the situation where the target recognition model outputs a high-confidence incorrect category in visual confusion scenarios.
[0162] In some embodiments, a target loss function is constructed based on the category label, confusion type label, confidence label, predicted category, and predicted confidence, including: constructing an accuracy loss function based on the category label and the predicted confidence of the predicted category; constructing a confidence correction loss function based on the confusion type label, confidence label, and probability distribution; and determining the target loss function based on the accuracy loss function and the confidence correction loss function.
[0163] In some embodiments, the accuracy loss function is constructed by using a focus loss function based on the probability distribution of the category label and multiple candidate categories.
[0164] In some embodiments, a confidence correction loss function is constructed based on the confusion type label, the confidence label, and the probability distribution, including: performing label smoothing on the confidence label based on the confidence label and the confusion type label to obtain a smoothed confidence label; constructing a first loss component based on the probability distribution of the smoothed confidence label and the category; and constructing the confidence correction loss function based on the first loss component.
[0165] In some embodiments, constructing a confidence-corrected loss function based on a first loss component includes: constructing a second loss component based on confidence labels and predicted confidence; and constructing a confidence-corrected loss function based on the first and second loss components.
[0166] In some embodiments, constructing a second loss component based on confidence labels and predicted confidence includes: identifying multiple target training samples in the same training batch whose confusion type label is the target confusion type; determining the confidence label probability distribution of the multiple target training samples in the training batch based on the confidence labels of the category labels of the multiple target training samples; determining the weight of each confidence level of each target training sample based on a preset threshold and the predicted confidence of the multiple target training samples; determining the confidence prediction distribution of each confidence level based on the weight of each confidence level of each target training sample; and constructing a second loss component based on the confidence label probability distribution and the confidence prediction distribution.
[0167] In some embodiments, determining a target loss function based on an accuracy loss function and a confidence correction loss function includes: determining a first weight coefficient corresponding to the accuracy loss function and a second weight coefficient corresponding to the confidence correction loss function based on a confusion type label, a confidence label, a category label, and a predicted category; and determining the target loss function based on the first weight coefficient, the second weight coefficient, the accuracy loss function, and the confidence correction loss function.
[0168] In some embodiments, determining a first weight coefficient corresponding to the accuracy loss function and a second weight coefficient corresponding to the confidence correction loss function based on the confusion type label, confidence label, category label, and predicted category includes: adjusting the second weight coefficient to be greater than the first weight coefficient when the confusion type label indicates that the sample image belongs to the target confusion type, the confidence label does not exceed a preset level, and the category label is inconsistent with the predicted category; adjusting the first weight coefficient to be greater than the second weight coefficient when the confusion type label indicates that the sample image does not belong to the target confusion type, the confidence label exceeds a preset level, or the category label is consistent with the predicted category.
[0169] Figure 5 This is a flowchart illustrating a target recognition method provided in an exemplary embodiment of this disclosure. Figure 5 As shown, the method includes the following steps S510~S530.
[0170] S510: Acquire the image to be recognized.
[0171] S520: Input the image to be recognized into the target recognition model to obtain the probability distribution of each candidate category of the image to be recognized output by the target recognition model; wherein, the target recognition model is trained based on the above target recognition model training method.
[0172] S530: Based on the probability distribution, determine the target category and target prediction confidence of the image to be identified.
[0173] Based on the target recognition method provided in this disclosure, the target recognition model trained by confidence correction has higher prediction confidence accuracy, which can effectively suppress the problem of confidence classification error. For easily confused categories with similar colors and overlapping features, it can reduce the probability of misclassification, improve the classification accuracy, optimize the model's own classification and discrimination ability, and improve the overall recognition robustness of the model.
[0174] Exemplary electronic devices Figure 6 A structural diagram of an electronic device provided in an embodiment of this disclosure includes at least one processor 111 and a memory 112.
[0175] The processor 111 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 11 to perform desired functions.
[0176] The memory 112 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 111 may execute one or more computer program instructions to implement the target recognition model training method or target recognition method and / or other desired functions of the various embodiments of this disclosure described above.
[0177] In one example, the electronic device 11 may also include an input device 113 and an output device 114, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0178] The input device 113 may also include, for example, a keyboard, a mouse, etc.
[0179] The output device 114 can output various information to the outside, including, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0180] Of course, for the sake of simplicity, Figure 6 Only some of the components of the electronic device 11 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 11 may include any other suitable components depending on the specific application.
[0181] Exemplary computer program products and computer-readable storage media In addition to the methods and devices described above, embodiments of this disclosure may also provide a computer program product, including computer program instructions, which, when executed by a processor, cause the processor to perform the steps in the target recognition model training method or target recognition method of the various embodiments of this disclosure described in the "Exemplary Methods" section above.
[0182] Computer program products can be written in any combination of one or more programming languages to perform the operations of embodiments of this disclosure. These programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0183] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the target recognition model training method or target recognition method of the various embodiments of this disclosure described in the "Exemplary Methods" section above.
[0184] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may include, but is not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0185] The basic principles of this disclosure have been described above with reference to specific embodiments. However, the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0186] Various modifications and variations can be made to this disclosure without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, this disclosure is also intended to include such modifications and variations.
Claims
1. A target recognition model training device, comprising a processor, the processor being configured to: Obtain the training sample set; among which, The training sample set includes multiple training samples, each training sample including a sample image and a category label of the target to be identified in the sample image, a confusion type label corresponding to the category label, and a confidence label corresponding to the category label; The training samples are input into the initial recognition model to obtain the probability distribution of the target to be recognized belonging to each candidate category; Based on the probability distribution, the prediction confidence level of the predicted category is determined; A target loss function is constructed based on the category label, the confusion type label, the confidence label, the predicted category, and the predicted confidence. The initial recognition model is then trained based on the target loss function to obtain the target recognition model.
2. The apparatus according to claim 1, wherein, The processor is further configured to: Based on the category label and the prediction confidence of the predicted category, an accuracy loss function is constructed; Based on the confusion type label, the confidence label, and the probability distribution, a confidence correction loss function is constructed; The target loss function is determined based on the accuracy loss function and the confidence correction loss function.
3. The apparatus according to claim 2, wherein, The processor is further configured to: Based on the confidence label and the obfuscation type label, the confidence label is smoothed to obtain a smoothed confidence label; Based on the smoothed confidence labels and the probability distribution, a first loss component is constructed; Based on the first loss component, the confidence correction loss function is constructed.
4. The apparatus according to claim 3, wherein, The processor is further configured to: Based on the confidence label and the predicted confidence, a second loss component is constructed; The confidence correction loss function is constructed based on the first loss component and the second loss component.
5. The apparatus according to claim 4, wherein, The processor is further configured to: A plurality of target training samples in the same training batch are identified as having the target confusion type label. Based on the confidence labels of the category labels of the plurality of target training samples, the probability distribution of the confidence labels of the plurality of target training samples in the training batch is determined. Based on a preset threshold and the prediction confidence of multiple target training samples, the weight of each prediction confidence level of each target training sample is determined. Based on the weights of each prediction confidence level of each target training sample, the confidence prediction distribution of each prediction confidence level is determined; A second loss component is constructed based on the confidence label probability distribution and the confidence prediction distribution.
6. The apparatus according to any one of claims 2-5, wherein, The processor is further configured to: Based on the confusion type label, the confidence label, the category label, and the predicted category, determine the first weight coefficient corresponding to the accuracy loss function and the second weight coefficient corresponding to the confidence correction loss function; The target loss function is determined based on the first weighting coefficient, the second weighting coefficient, the accuracy loss function, and the confidence correction loss function.
7. The apparatus according to claim 6, wherein, The processor is further configured to: If the confusion type label indicates that the sample image belongs to the target confusion type, the confidence label does not exceed the preset level, and the category label is inconsistent with the predicted category, the second weight coefficient is adjusted to be greater than the first weight coefficient. If the confusion type label indicates that the sample image does not belong to the target confusion type, the confidence label exceeds the preset level, or the category label is consistent with the predicted category, the first weight coefficient is adjusted to be greater than the second weight coefficient.
8. A target recognition device, comprising a processor, the processor being configured to: Acquire the image to be recognized; The image to be identified is input into a target recognition model to obtain the probability distribution of each candidate category of the image to be identified, output by the target recognition model; wherein, The target recognition model is trained based on the target recognition model training device according to any one of claims 1-6; Based on the probability distribution, the target category and target prediction confidence of the image to be identified are determined.
9. A method for training a target recognition model, comprising: Obtain a training sample set; wherein the training sample set includes multiple training samples, and the training samples include sample images and category labels of the target to be identified in the sample images, confusion type labels corresponding to the category labels, and confidence labels corresponding to the category labels; The training samples are input into the initial recognition model to obtain the probability distribution of the target to be recognized belonging to each candidate category; Based on the probability distribution, the prediction confidence level of the predicted category is determined; A target loss function is constructed based on the category label, the confusion type label, the confidence label, the predicted category, and the predicted confidence. The initial recognition model is then trained based on the target loss function to obtain the target recognition model.
10. A target recognition method, comprising: Acquire the image to be recognized; The image to be identified is input into the target recognition model to obtain the probability distribution of each candidate category of the image to be identified output by the target recognition model; wherein, the target recognition model is trained based on the target recognition model training method of claim 9; Based on the probability distribution, the target category and target prediction confidence of the image to be identified are determined.
11. A computer-readable storage medium storing a computer program for executing the target recognition model training method of claim 9 or the target recognition method of claim 10.
12. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the target recognition model training method of claim 9 or the target recognition method of claim 10.