Object recognition method, apparatus, vehicle, and storage medium
The object recognition model trained with single-stage gradient loss and regression loss solves the problem of poor object recognition performance in existing technologies, improves recognition accuracy, reduces costs, and enhances the safety of autonomous driving.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU AUTOMOBILE GROUP CO LTD
- Filing Date
- 2023-02-21
- Publication Date
- 2026-07-31
AI Technical Summary
Existing object recognition models perform poorly in autonomous driving, resulting in low accuracy in identifying target objects and posing safety hazards. Furthermore, ultrasonic detection is greatly affected by weather conditions and is costly, while lidar detection is expensive.
The object recognition model is trained using single-stage gradient loss and regression loss. The single-stage gradient loss in each iteration represents the accuracy of the model on the target sample, and the regression loss represents the fit of the object recognition. The model is trained by combining multi-scale feature map information to improve the recognition accuracy.
It improves the recognition effect and accuracy of the object recognition model, reduces costs, reduces security risks, and enhances the reliability of object recognition results.
Smart Images

Figure CN116189128B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle control technology, and more specifically, to an object recognition method, apparatus, vehicle, and computer-readable storage medium. Background Technology
[0002] With the development of automotive electronics and network technology, object detection is becoming increasingly important and is gradually becoming a crucial part of environmental perception in autonomous driving systems.
[0003] Currently, object detection can be vision-based. First, bounding boxes (used to select target objects within the sample images) are labeled on color sample images, and an initial model is trained using these labeled color images to obtain an object recognition model. During autonomous driving, color images of the vehicle's driving environment are acquired. These color images are processed by the object recognition model to obtain the probability that the color images contain target objects. Then, based on this probability, the objects included in the driving environment are determined.
[0004] However, the object recognition models obtained using existing methods have poor recognition performance, resulting in a low accuracy in probabilities that the obtained color images include the target objects, which poses a safety hazard to vehicles during autonomous driving. Summary of the Invention
[0005] In view of this, this application proposes an object recognition method, apparatus, vehicle, and computer-readable storage medium to solve the above problems.
[0006] In a first aspect, embodiments of this application provide an object recognition method, the method comprising: acquiring a captured image of a target region; inputting the captured image into an object recognition model to obtain a predicted probability of the target object; the predicted probability refers to the probability that the captured image includes the target object; the object recognition model is trained using single-stage gradient loss values and regression loss values in each iteration, wherein the single-stage gradient loss value in each iteration characterizes the accuracy of the training model in that iteration in predicting the target sample, and the regression loss value in each iteration characterizes the degree of fitting of the training model in that iteration to the location of the target object recognition; the single-stage gradient loss value in each iteration is determined by the degree of overlap between the predicted bounding box and the ground truth bounding box of the sample image in that iteration; the predicted bounding box of the sample image in each iteration is determined by the training model in that iteration, wherein the predicted bounding box of the sample image in each iteration refers to the selection box predicted by the training model in that iteration for selecting the target object, and the ground truth bounding box refers to the labeled selection box for selecting the target object; and determining the object recognition result for the target region based on the predicted probability.
[0007] Secondly, embodiments of this application provide an object recognition device, comprising: an acquisition module for acquiring a captured image of a target region; a prediction module for inputting the captured image into an object recognition model to obtain a predicted probability of the target object; the predicted probability refers to the probability that the captured image includes the target object; the object recognition model is trained using single-stage gradient loss values and regression loss values in each iteration, wherein the single-stage gradient loss value in each iteration characterizes the accuracy of the training model in that iteration in predicting the target sample, and the regression loss value in each iteration characterizes the degree of fitting of the training model in that iteration to the location of the target object recognition; the single-stage gradient loss value in each iteration is determined by the degree of overlap between the predicted bounding box and the ground truth bounding box corresponding to the sample image in that iteration; the predicted bounding box corresponding to the sample image in each iteration is determined by the training model in that iteration, and the predicted bounding box refers to the selection box predicted by the training model in that iteration for selecting the target object, and the ground truth bounding box refers to the labeled selection box for selecting the target object; and a result determination module for determining the object recognition result for the target region based on the predicted probability.
[0008] Thirdly, embodiments of this application provide a vehicle, the vehicle comprising:
[0009] One or more processors;
[0010] Memory;
[0011] One or more applications, wherein the applications are stored in memory and configured to be executed by one or more processors, and the applications are configured to perform the methods of the first aspect described above.
[0012] Fourthly, embodiments of this application provide a computer-readable storage medium storing program code, which can be invoked by a processor to execute the method described in the first aspect.
[0013] This application provides an object recognition method, device, vehicle, and computer-readable storage medium. In this application, the object recognition model is trained by using a single-stage gradient loss value and a regression loss value to train the model during the iterative process. The single-stage gradient loss value represents the attention the model pays to the target sample during the iterative process, and the regression loss value represents the accuracy of the model in recognizing the target object during the iterative process. This results in a better recognition effect of the trained object recognition model, which in turn leads to a higher accuracy in predicting the probability of the target object determined by the object recognition model. Consequently, the object recognition result is more accurate, greatly reducing the occurrence of vehicle safety hazards caused by inaccurate object recognition results.
[0014] These or other aspects of this application will become more apparent from the description of the following embodiments. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 A schematic diagram of a vehicle hardware environment applicable to embodiments of this application is shown.
[0017] Figure 2 A flowchart of an object recognition method according to an embodiment of this application is shown.
[0018] Figure 3 A flowchart of an object recognition method according to yet another embodiment of this application is shown.
[0019] Figure 4 A flowchart of an object recognition method according to another embodiment of this application is shown.
[0020] Figure 5 A schematic diagram illustrating the training process of an object recognition model according to an embodiment of this application is shown;
[0021] Figure 6 A structural block diagram of an object recognition device according to an embodiment of this application is shown.
[0022] Figure 7 This invention provides a structural block diagram of a vehicle according to one embodiment of the present application.
[0023] Figure 8 A structural block diagram of a computer-readable storage medium provided in an embodiment of this application is shown. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0025] Currently, object recognition results can also be obtained through ultrasonic detection. Ultrasonic detection utilizes ultrasound as a signal source to transmit information. An ultrasonic sensor emits ultrasound waves through a transmitter. When the ultrasound waves encounter an obstacle, some are reflected back. A receiver picks up the reflected ultrasound waves, and using a timer, the ultrasonic sensor calculates the travel time of the reflected ultrasound waves in the air. Based on the distance the ultrasound waves travel in the air, the distance between the obstacle and the ultrasonic sensor is calculated, thus enabling obstacle detection and object recognition.
[0026] However, the transmission speed of ultrasound is easily affected by weather and temperature changes, which can lead to increased measurement errors. At the same time, when measuring at long distances, the echo signal of ultrasound is weak, making it unsuitable for obstacle detection and tracking at slightly longer distances, resulting in a lower accuracy of object recognition results obtained by ultrasound detection.
[0027] In addition, object recognition results can now also be obtained through lidar detection. Lidar detection involves emitting a detection signal (laser beam) at the target, then comparing the received signal reflected back from the target (target echo) with the emitted signal. After appropriate processing, relevant information about the target can be obtained, such as parameters like target distance, azimuth, altitude, speed, attitude, and even shape, thereby enabling target detection, tracking, and identification.
[0028] Obstacle detection methods based on lidar are unaffected by light conditions and have excellent environmental adaptability, exhibiting significant advantages in key performance aspects such as detection accuracy, detection range, and stability. However, lidar is expensive, resulting in high costs and limited practicality for lidar-based detection.
[0029] In view of this, the inventors have proposed the object recognition method, device, vehicle, and computer-readable storage medium of this application. The single-stage gradient loss value represents the attention of the training model to the target sample during the iteration process, and the regression loss value represents the accuracy of the training model in recognizing the target object during the iteration process. This makes the object recognition model obtained by training perform well, thereby making the prediction probability accuracy of the target object determined by the object recognition model higher, and thus making the object recognition result more accurate. This greatly reduces the occurrence of vehicle safety hazards caused by inaccurate object recognition results.
[0030] Meanwhile, the training process of the object recognition model is relatively inexpensive, which makes the object recognition method of this application less expensive and improves its practicality.
[0031] Reference Figure 1 , Figure 1A schematic diagram of a vehicle hardware environment applicable to an embodiment of this application is shown. The vehicle 100 includes an autonomous driving system 110. The autonomous driving system 110 can have a variety of built-in autonomous driving functions. The autonomous driving system 110 controls the vehicle to drive autonomously according to the built-in autonomous driving functions. The autonomous driving functions may include, for example, automatic lane changing function, automatic overtaking function, and automatic parking function.
[0032] The autonomous driving system 110 may include an on-board data acquisition device 111, one or more (only one is shown in the figure) processors 112 and memory 113.
[0033] The vehicle-mounted acquisition device 111 is used to capture images of a target area during vehicle operation. The target area can refer to the area surrounding the vehicle during operation (or the driving environment during operation), and the vehicle-mounted acquisition device 111 can be, for example, a high-definition camera.
[0034] The processor 112 may be a microcontroller unit (MCU) with a built-in memory 113 containing a program that can execute the contents of the following embodiments, and the processor 112 can execute the program stored in the memory 113.
[0035] The processor 112 may include one or more processors. The processor 112 connects to various parts of the vehicle 100 through various interfaces and lines, and performs various functions of the vehicle 10 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 113, and calling data stored in the memory 113.
[0036] Memory 113 may include random access memory (RAM) or read-only memory (ROM). Memory 15 may be used to store instructions, programs, code, code sets, or instruction sets. Memory 15 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), instructions for implementing the various method embodiments described below, etc.
[0037] Reference Figure 2 , Figure 2 A flowchart of an object recognition method according to an embodiment of this application is shown. The method is used for vehicles and includes:
[0038] S110, Acquire images of the target area.
[0039] In this embodiment, the target area can refer to the area surrounding the vehicle during its driving process. The vehicle can be equipped with a camera (e.g., a high-definition camera) to capture images of the target area. These images can be RGB images. The camera can use an external ISP (Image Signal Processor) to meet the wide dynamic range requirements of HDR (High Dynamic Range Imaging), thereby enabling the machine to effectively recognize the image content.
[0040] S120. Input the captured image into the object recognition model to obtain the predicted probability of the target object; the predicted probability refers to the probability that the captured image includes the target object; the object recognition model is trained by the single-stage gradient loss value and regression loss value of each iteration process. The single-stage gradient loss value of each iteration process represents the accuracy of the training model in that iteration process in predicting the target sample, and the regression loss value of each iteration process represents the degree of fitting of the training model in that iteration process to the position of the target object recognition. The single-stage gradient loss value of each iteration process is determined by the degree of overlap between the predicted box and the ground truth box corresponding to the sample image in that iteration process. The predicted box corresponding to the sample image in each iteration process is determined by the training model in that iteration process. The predicted box corresponding to the sample image in each iteration process refers to the selection box predicted by the training model in that iteration process for selecting the target object, and the ground truth box refers to the labeled selection box for selecting the target object.
[0041] The target object can refer to an object of a certain category, such as people; the target object can also include multiple target objects, each of which refers to an object of a certain category, such as cars and people.
[0042] Target samples can refer to difficult samples, which are images that contain a target object but whose object is not obvious. When the model to be trained during the iterative process (i.e., the iterative training process) identifies such images, it is difficult to accurately determine whether they contain the target object. The parameters of the models to be trained differ in different iterative processes, but their structure is the same. The model with initialized parameters can be designated as the initial model and used as the model to be trained in the first iteration. The model to be trained after each iteration can be used as the model to be trained in the next iteration. Sample images can refer to images used to train the model; sample images may or may not contain the target object.
[0043] An object recognition model is used to identify whether a captured image contains a target object. It can be obtained by iteratively training an initial model using single-stage gradient loss and regression loss. The initial model can be a general neural network model, such as VGG or ResNet.
[0044] In each iteration, the predicted bounding box is the selection box predicted by the training model of the current iteration after processing the sample image (the predicted bounding box of the same sample image may be different in different iterations); the ground truth box is the labeled selection box, which can be labeled by an electronic device or manually labeled by the user. The selection box (ground truth box and predicted bounding box) can be a rectangle, and the selection box is used to select the target object.
[0045] The single-stage gradient loss value in each iteration represents the accuracy of the training model's prediction of the target sample in that iteration. Simultaneously, the regression loss value in each iteration represents the degree of fit of the training model to the target object's location. Combining the single-stage gradient loss value and the regression loss value, the final loss value for that iteration is obtained (this final loss value can be obtained by summing the single-stage gradient loss value and the regression loss value; this summation can be a weighted summation or a direct summation). The training model for that iteration is then trained using this final loss value, resulting in the trained model for that iteration. This trained model serves as the training model for the next iteration. After the final iteration, the trained model for the final iteration is used as the object recognition model. The training model for the first iteration can refer to the initial model.
[0046] The object recognition model trained according to this method has a high accuracy rate in recognizing target objects and a strong ability to recognize difficult-to-recognize images, thus achieving the effect of improving the recognition ability and accuracy of the object recognition model.
[0047] For each iteration, feature information of the sample images can be extracted from the model to be trained in that iteration, and then regression calculation can be performed on the extracted feature information of the sample images to obtain the regression loss value of that iteration.
[0048] Meanwhile, for each iteration, the feature information of the sample image can be classified to obtain the prediction probability and prediction box corresponding to the iteration, and the ground truth box of the sample image can be obtained. The single-stage gradient loss value of the iteration is determined based on the degree of overlap between the prediction box and the ground truth box corresponding to the iteration (in this application, the degree of overlap can be characterized by IOU, which is short for Intersection over Union).
[0049] After obtaining the captured image, it is input into the object recognition model to obtain the predicted probability output by the model. The predicted probability output by the object recognition model includes the individual predicted probability for each target object. For example, if the target objects include people and cars, the predicted probability output by the object recognition model may include a predicted probability of 0.56 for cars and a predicted probability of 0.8 for people.
[0050] S130. Based on the predicted probability, determine the object recognition result for the target area.
[0051] Based on the predicted probabilities of each target object, the object recognition result of the target region is determined. The object recognition result may include: the target region includes the target object or the target region does not include the target object.
[0052] A probability threshold can be set for each target object (the probability thresholds for different target objects can be the same or different); when the predicted probability of the target object reaches the corresponding probability threshold, it is determined that the captured image includes the target object, and thus the object recognition result of the target region including the target object is obtained; when the predicted probability of the target object does not reach the corresponding probability threshold, it is determined that the captured image does not include the target object, and thus the object recognition result of the target region not including the target object is obtained.
[0053] In this embodiment, the object recognition model is trained by using single-stage gradient loss and regression loss values to train the model to be trained during the iterative process. The single-stage gradient loss value represents the attention the model to be trained pays to the target sample during the iterative process, and the regression loss value represents the accuracy of the model to be trained in identifying the target object during the iterative process. This results in a better recognition effect of the trained object recognition model, which in turn leads to a higher accuracy in predicting the probability of the target object determined by the object recognition model. Consequently, the object recognition result is more accurate, greatly reducing the occurrence of vehicle safety hazards caused by inaccurate object recognition results.
[0054] Reference Figure 3 , Figure 3 A flowchart of an object recognition method according to another embodiment of this application is shown. The method is used for vehicles and includes:
[0055] S210. Obtain sample images of the target object; determine the multi-scale feature map information of the Nth iteration process of the sample image using the model to be trained in the Nth iteration process.
[0056] In this application, N is an integer greater than 1. A sample image can refer to a color image (RGB image) used as a training sample. There can be multiple sample images. A sample image may include the target object or may not include the target object.
[0057] For the Nth iteration, features can be extracted from each sample image using the model to be trained in the Nth iteration, resulting in multi-scale feature map information for each sample image in the Nth iteration.
[0058] In one implementation, the multi-scale feature map information for the Nth iteration of a sample image is determined using the training model in the Nth iteration process. This includes: extracting features from the sample image using a pre-trained backbone network (e.g., the backbone of RetinaNet) in the training model in the Nth iteration process to obtain image features for the Nth iteration process; and extracting features from the image features in the Nth iteration process using a feature pyramid (e.g., Featured image pyramid, Single feature map, Pyramidal feature hierarchy, or Feature Pyramid Network) in the training model in the Nth iteration process to obtain multi-scale feature map information for the Nth iteration process. That is, in this embodiment, the training model in the Nth iteration process includes a backbone network and a feature pyramid, and the multi-scale feature map information of a sample image can refer to feature maps at multiple scales corresponding to that sample image.
[0059] The backbone network needs to have a certain image feature extraction capability, so it is necessary to import pre-trained weights. By using a pre-trained backbone network with prior knowledge, image features can be detected quickly, effectively improving the training effect.
[0060] Typically, feature map extraction in Convolutional Neural Networks (CNNs) is top-down. However, the feature pyramid is designed with a bottom-up structure. Simply put, it involves upsampling multiple times and fusing top-down features into a single convolution to obtain the feature map. Bottom-up features have lower semantic information, but because they are sampled less, they have stronger localization information, thus enhancing the model's localization ability.
[0061] In some possible implementations, the image features of the Nth iteration process are obtained by extracting features from the sample images through the pre-trained backbone network in the model to be trained in the Nth iteration process. This includes: resizing multiple sample images to obtain multiple intermediate sample images of the same size; normalizing the pixel values of each pixel in each intermediate sample image to obtain the input sample image corresponding to each intermediate sample image; and extracting features from each input sample image through the pre-trained backbone network in the model to be trained in the Nth iteration process to obtain the image features of the Nth iteration process corresponding to each input sample image.
[0062] The multiple sample images used as training samples may have different sizes, so it is necessary to resize these multiple sample images to obtain multiple intermediate sample images of the same size. For example, the size of an intermediate sample image can be 512×512.
[0063] The pixel value of each pixel in the intermediate sample image can be divided by 255 to normalize the pixel value of each pixel in the intermediate sample image, so as to obtain the input sample image (the pixel value of the input sample image is in [0,1]). Then, the input image is input into the training model in the Nth iteration process to obtain the image features of the Nth iteration process extracted by the backbone network.
[0064] Normalization reduces the size of data values, making it easier for the model training process to converge.
[0065] S220. Perform regression processing on the multi-scale feature map information of the Nth iteration process to obtain the regression loss value of the Nth iteration process.
[0066] After obtaining the multi-scale feature map information of the Nth iteration of the sample image, regression processing is performed on the multi-scale feature map information of the Nth iteration (e.g., by calculating using a regression loss function) to obtain the regression loss value of the Nth iteration. The regression loss function can be, for example, the mean squared error loss function, the mean absolute error loss function, the Huber loss function, the Log-Cosh loss function, and the quantile loss function.
[0067] Understandably, for each iteration, the multi-scale feature map information of that iteration is determined according to the method in S210, and the regression loss value of that iteration is determined based on the multi-scale feature map information of that iteration and the method in S220.
[0068] S230. Determine the single-stage gradient loss value of the Nth iteration process; train the model to be trained in the Nth iteration process using the single-stage gradient loss value and the regression loss value of the Nth iteration process to obtain the object recognition model.
[0069] S240. Acquire a captured image of the target area; input the captured image into the object recognition model to obtain the predicted probability of the target object; determine the object recognition result for the target area based on the predicted probability.
[0070] The descriptions of S230-S240 are the same as those of S110-S130 above, and will not be repeated here.
[0071] In this embodiment, by using training samples and the model to be trained in the iterative process, high-accuracy multi-scale feature map information can be obtained. Based on the high-accuracy multi-scale feature map information, a high-accuracy regression loss value can be determined, thereby making the object recognition model trained based on the regression loss value more accurate, and thus improving the accuracy of object recognition results of target regions obtained by the object recognition model.
[0072] Reference Figure 4 , Figure 4 A flowchart of an object recognition method according to another embodiment of this application is shown. The method is used for vehicles and includes:
[0073] S310. Obtain the label, ground truth box, sample prediction probability of the Nth iteration process, and prediction box of the Nth iteration process corresponding to the sample image; determine the multi-scale feature map information of the Nth iteration process for the sample image through the training model of the Nth iteration process.
[0074] The description of S310 is the same as that of S210 above, and will not be repeated here.
[0075] The label represents whether the sample image includes the target object or not. The sample prediction probability in the Nth iteration represents the probability that the model to be trained in the Nth iteration predicts that the sample image includes the target object. N is an integer greater than 1.
[0076] The sample image includes the target object, and its corresponding label can be 1; the sample image does not include the target object, and its corresponding label can be 0.
[0077] S320. Perform convolution operation on the multi-scale feature map information of the Nth iteration process to obtain the sample prediction probability and prediction box of the Nth iteration process of the sample image.
[0078] Among them, the sample prediction probability of the Nth iteration process represents the probability that the training model predicts that the sample image includes the target object in the Nth iteration process; the prediction box of the Nth iteration process refers to the selection box predicted by the training model in the Nth iteration process.
[0079] The multi-scale feature map information from the Nth iteration can be input into a classifier constructed from a convolutional network to perform convolution operations on the multi-scale feature map information from the Nth iteration, thereby obtaining the sample prediction probability and prediction box of the sample image in the Nth iteration. The prediction box in the Nth iteration can also include the center point width and height information of the selection box predicted by the model to be trained in the Nth iteration.
[0080] It is understandable that, for each iteration, the multi-scale feature map information of that iteration is determined according to the method of S310, and the sample prediction probability and prediction box of that iteration are determined based on the multi-scale feature map information of that iteration and the method of S320.
[0081] S330. Based on the label and the sample prediction probability of the Nth iteration process, determine the gradient value of the sample image in the Nth iteration process; determine the gradient interval of the Nth iteration process based on the gradient value of the Nth iteration process, and divide the gradient interval of the Nth iteration process to obtain multiple sub-intervals of the Nth iteration process.
[0082] The probability can be predicted based on the labels and the samples from the Nth iteration. The cross-entropy loss value for the Nth iteration can then be calculated using Formula 1, as follows:
[0083]
[0084] Among them, Loss CE (p,p * ) is the cross-entropy loss value of the Nth iteration of the calculated sample image, and p is the prediction probability of the Nth iteration corresponding to the sample image. * The label for the sample image is either 1 or 0.
[0085] After obtaining the cross-entropy loss value, we take the derivative of the cross-entropy loss value according to Formula 2 to obtain the gradient value of the Nth iteration. Formula 2 is as follows:
[0086]
[0087] Where g is the gradient value of the Nth iteration.
[0088] The gradient value of the Nth iteration is usually within [0,1]. The interval formed by the minimum and maximum gradient values of the Nth iteration (that is, the gradient interval of the Nth iteration) can be evenly divided to obtain multiple sub-intervals (e.g., 30 sub-intervals) of the Nth iteration.
[0089] S340. Determine the degree of overlap between the prediction box and the truth box in the Nth iteration process, and use it as the overlap of the Nth iteration process; determine the single-stage gradient loss value of the Nth iteration process based on the multiple sub-intervals of the Nth iteration process and the overlap of the Nth iteration process.
[0090] The predicted bounding box and ground truth bounding box of each sample image in the Nth iteration process can be obtained. Based on the area of the predicted bounding box, the area of the ground truth bounding box, and the area of their overlap in the Nth iteration process, the overlap degree of the Nth iteration process is determined (in this application, the overlap degree can be characterized by IOU, which stands for Intersection over Union). This overlap degree is then used as the single-stage gradient loss value for the Nth iteration process. Finally, based on the multiple sub-intervals of the Nth iteration process and the overlap degree of the Nth iteration process, the single-stage gradient loss value for the Nth iteration process is determined.
[0091] Specifically, the single-stage gradient loss value of the Nth iteration process is determined based on multiple sub-intervals and the overlap of the Nth iteration process. This includes: counting the total number of sample images and the number of sample images in each sub-interval of the Nth iteration process; and determining the single-stage gradient loss value of the Nth iteration process based on the total number of sample images, the gradient value of each sample image in the Nth iteration process, the overlap of each sample image in the Nth iteration process, and the number of sample images corresponding to each sub-interval of the Nth iteration process.
[0092] Formula 3 can be used to determine the single-stage gradient loss value of the Nth iteration process, based on the total number of sample images, the gradient value of each sample image in the Nth iteration process, the overlap of each sample image in the Nth iteration process, and the number of sample images corresponding to each sub-interval in the Nth iteration process. Formula 3 is as follows:
[0093]
[0094] Among them, Loss IOU- Let iou be the single-stage gradient loss value of the Nth iteration process. i Let g be the overlap between the predicted bounding box and the ground truth bounding box for the Nth iteration corresponding to the i-th sample image. iIt is the gradient value of the Nth iteration corresponding to the i-th sample image, where N is the total number of sample images, gd(g i ) is the number of sample images in the sub-interval where the i-th sample image is located in the N-th iteration (that is, the total number of sample images corresponding to this sub-interval).
[0095] S350. The model to be trained in the Nth iteration process is trained using the single-stage gradient loss value and regression loss value of the Nth iteration process to obtain the object recognition model.
[0096] S360. Acquire a captured image of the target area; input the captured image into the object recognition model to obtain the predicted probability of the target object; determine the object recognition result for the target area based on the predicted probability.
[0097] The descriptions of S350-S360 are the same as those of S110-S130 above, and will not be repeated here.
[0098] In this embodiment, the training process of the object recognition model is as follows: Figure 5 As shown, images can be captured by a camera, and the captured images can be annotated with truth boxes and labels to obtain sample images. The sample images are then preprocessed: the sample images are resized to obtain intermediate sample images, and the pixel values of each pixel in the intermediate sample images after unification are normalized to obtain the input sample images.
[0099] For each iteration, the input sample image is fed into the training model for that iteration. The model's backbone network and feature pyramid are used for processing to obtain multi-scale feature map information for that iteration. Then, regression processing is performed on this multi-scale feature map information to obtain the regression loss value for that iteration. Simultaneously, convolution operations are performed on the multi-scale feature map information to obtain the sample prediction probability and prediction box for that iteration.
[0100] Then, the overlap is calculated based on the predicted bounding box and the ground truth bounding box of this iteration, and the cross-entropy loss value is calculated based on the sample prediction probability and the label. Multiple sub-intervals are obtained through the cross-entropy loss value. Based on the multiple sub-intervals and the overlap, the single-stage gradient loss value of this iteration is obtained. The final loss value of this iteration is determined by the single-stage gradient loss value and the regression loss value of this iteration.
[0101] Finally, the model to be trained in this iteration is iteratively trained using the final loss value of this iteration until the last iteration is completed, resulting in the object recognition model.
[0102] By utilizing the gradient distribution of sample images, the model suppresses easily trained sample images and outlier sample images, thereby increasing the model's focus on target samples. Simultaneously, by incorporating sample overlap, the model's training is concentrated on classifying boxes with relatively accurate predictions, enabling the model to find suitable target samples for training, reducing overfitting, and improving the model's generalization ability for autonomous driving scenarios.
[0103] In this embodiment, the single-stage gradient loss value of the Nth iteration process is determined based on the total number of sample images, the gradient value of each sample image in the Nth iteration process, the overlap of each sample image in the Nth iteration process, and the number of sample images corresponding to each sub-interval in the Nth iteration process. The single-stage gradient loss value of the Nth iteration process can accurately characterize the attention of the training model to the target sample in the Nth iteration process, thereby making the object recognition model trained based on the single-stage gradient loss value have a better recognition effect, and thus improving the accuracy of object recognition results in the target region.
[0104] Reference Figure 6 , Figure 6 This diagram illustrates a structural block diagram of an object recognition device according to an embodiment of this application. The device 800 is used in a vehicle and includes:
[0105] The acquisition module 810 is used to acquire images of the target area.
[0106] The prediction module 820 is used to input the captured image into the object recognition model to obtain the predicted probability of the target object. The predicted probability refers to the probability that the captured image includes the target object. The object recognition model is trained by the single-stage gradient loss value and regression loss value of each iteration. The single-stage gradient loss value of each iteration represents the accuracy of the training model in predicting the target sample in that iteration. The regression loss value of each iteration represents the degree of fit of the training model in the location of the target object in that iteration. The single-stage gradient loss value of each iteration is determined by the degree of overlap between the predicted box and the ground truth box of the sample image in that iteration. The predicted box of the sample image in each iteration is determined by the training model in that iteration. The predicted box of the sample image in each iteration refers to the selection box predicted by the training model in that iteration for selecting the target object. The ground truth box refers to the labeled selection box for selecting the target object.
[0107] The result determination module 830 is used to determine the object recognition result for the target area based on the predicted probability.
[0108] Furthermore, the device also includes a first acquisition module, used to acquire the label, ground truth box, sample prediction probability of the Nth iteration, and prediction box of the Nth iteration corresponding to the sample image; the label represents whether the sample image includes the target object or not, and the sample prediction probability of the Nth iteration represents the probability that the training model predicts the sample image includes the target object in the Nth iteration, where N is an integer greater than 1; the gradient value of the sample image in the Nth iteration is determined according to the label and the sample prediction probability of the Nth iteration; the gradient interval of the Nth iteration is determined according to the gradient value of the Nth iteration, and the gradient interval of the Nth iteration is divided to obtain multiple sub-intervals of the Nth iteration; the degree of overlap between the prediction box and the ground truth box in the Nth iteration is determined as the overlap of the Nth iteration; the single-stage gradient loss value of the Nth iteration is determined according to the multiple sub-intervals of the Nth iteration and the overlap of the Nth iteration.
[0109] Furthermore, the first acquisition module is also used to count the total number of sample images and the number of sample images in each sub-interval of the Nth iteration process; and to determine the single-stage gradient loss value of the Nth iteration process based on the total number of sample images, the gradient value of each sample image in the Nth iteration process, the overlap of each sample image in the Nth iteration process, and the number of sample images corresponding to each sub-interval of the Nth iteration process.
[0110] Furthermore, the device also includes a second acquisition module, used to determine the multi-scale feature map information of the Nth iteration process of the sample image through the training model of the Nth iteration process; and to perform convolution operation on the multi-scale feature map information of the Nth iteration process to obtain the sample prediction probability and prediction box of the Nth iteration process of the sample image.
[0111] Furthermore, the device also includes a third acquisition module for acquiring sample images of the target object; determining multi-scale feature map information of the Nth iteration process of the sample image through the training model of the Nth iteration process, where N is an integer greater than 1; and performing regression processing on the multi-scale feature map information of the Nth iteration process to obtain the regression loss value of the Nth iteration process.
[0112] Furthermore, the device also includes a fourth acquisition module, which is used to extract features from the sample image through the pre-trained backbone network in the training model of the Nth iteration process to obtain the image features of the Nth iteration process; and to extract features from the image features of the Nth iteration process through the feature pyramid in the training model of the Nth iteration process to obtain the multi-scale feature map information of the Nth iteration process.
[0113] Furthermore, the fourth acquisition module is also used to resize multiple sample images to obtain multiple intermediate sample images of the same size; to normalize the pixel values of each pixel in each intermediate sample image to obtain the input sample image corresponding to each intermediate sample image; and to extract features from each input sample image through the pre-trained backbone network in the training model of the Nth iteration process to obtain the image features of the Nth iteration process corresponding to each input sample image.
[0114] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0115] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0116] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0117] Please see Figure 7 The diagram illustrates a structural block diagram of a vehicle 900 provided in an embodiment of this application. The vehicle 900 in this application may include one or more of the following components: a processor 910, a memory 920, and one or more application programs. The one or more application programs may be stored in the memory 920 and configured to be executed by one or more processors 910. The one or more programs are configured to perform the methods described in the foregoing method embodiments.
[0118] The processor 910 may include one or more processing cores. The processor 910 connects to various parts of the vehicle 900 via various interfaces and lines, and performs various functions and processes data of the vehicle 900 by running or executing instructions, programs, code sets, or instruction sets stored in the memory 920, and by calling data stored in the memory 920. Optionally, the processor 910 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 910 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 910 and may be implemented separately through a communication chip.
[0119] The memory 920 may include random access memory (RAM) or read-only memory (ROM). The memory 920 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 920 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the vehicle 900 during use (such as phonebooks, audio and video data, chat log data, etc.).
[0120] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0121] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0122] refer to Figure 8 , Figure 8A structural block diagram of a computer-readable storage medium provided in an embodiment of this application is shown. The computer-readable storage medium 1100 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0123] The computer-readable storage medium 1100 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 1100 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 1100 has storage space for program code 1110 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 1110 may, for example, be compressed in a suitable form.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method of object recognition, characterized by, The method includes: Acquire images of the target area; The captured image is input into an object recognition model to obtain the predicted probability of the target object. The predicted probability refers to the probability that the captured image includes the target object. The object recognition model is trained using the single-stage gradient loss value and regression loss value of each iteration. The single-stage gradient loss value of each iteration represents the accuracy of the model's prediction of the target sample in that iteration, and the regression loss value of each iteration represents the degree of fit of the model's position to the target object in that iteration. The process of obtaining the single-stage gradient loss value of each iteration includes: obtaining the label, ground truth box, sample prediction probability of the Nth iteration, and prediction box of the Nth iteration. The label represents whether the sample image includes the target object or not. The sample prediction probability of the Nth iteration represents... In the Nth iteration, the training model predicts the probability that the sample image includes the target object, where N is an integer greater than 1; the predicted bounding box of the sample image in each iteration refers to the selection box predicted by the training model in that iteration for selecting the target object, and the ground truth box refers to the labeled selection box for selecting the target object; based on the label and the sample prediction probability in the Nth iteration, the gradient value of the sample image in the Nth iteration is determined; the gradient interval of the Nth iteration is divided to obtain multiple sub-intervals of the Nth iteration; the degree of overlap between the predicted bounding box and the ground truth box in the Nth iteration is determined as the overlap degree of the Nth iteration; based on the multiple sub-intervals of the Nth iteration and the overlap degree of the Nth iteration, the single-stage gradient loss value of the Nth iteration is determined; Based on the predicted probability, the object recognition result for the target region is determined.
2. The method of claim 1, wherein, The step of determining the single-stage gradient loss value of the Nth iteration process based on multiple sub-intervals of the Nth iteration process and the overlap of the Nth iteration process includes: The total number of sample images and the number of sample images in each sub-interval of the Nth iteration process are counted. The single-stage gradient loss value of the Nth iteration process is determined based on the total number of sample images, the gradient value of each sample image in the Nth iteration process, the overlap of each sample image in the Nth iteration process, and the number of sample images corresponding to each sub-interval in the Nth iteration process.
3. The method of claim 1, wherein, The methods for obtaining the sample prediction probability and prediction box in the Nth iteration process include: The multi-scale feature map information for the Nth iteration of the sample image is determined by the model to be trained in the Nth iteration process. Convolution operations are performed on the multi-scale feature map information of the Nth iteration process to obtain the sample prediction probability and prediction box of the sample image in the Nth iteration process.
4. The method of claim 1, wherein, The methods for obtaining the regression loss value in each iteration include: Obtain a sample image of the target object; The multi-scale feature map information for the Nth iteration of the sample image is determined by the model to be trained in the Nth iteration process, where N is an integer greater than 1. The multi-scale feature map information of the Nth iteration process is subjected to regression processing to obtain the regression loss value of the Nth iteration process.
5. The method according to claim 3 or 4, characterized in that, The process of determining the multi-scale feature map information for the sample image in the Nth iteration of the training model includes: The sample image is extracted by pre-training the backbone network in the training model during the Nth iteration process to obtain the image features of the Nth iteration process. Feature extraction is performed on the image features of the Nth iteration process using the feature pyramid in the model to be trained in the Nth iteration process, to obtain the multi-scale feature map information of the Nth iteration process.
6. The method of claim 5, wherein, The process of extracting features from the sample image using the pre-trained backbone network in the training model during the Nth iteration to obtain the image features for the Nth iteration includes: The size of multiple sample images is changed to obtain multiple intermediate sample images of the same size; The pixel values of each pixel in each intermediate sample image are normalized to obtain the input sample image corresponding to each intermediate sample image. By using the pre-trained backbone network in the training model during the Nth iteration, feature extraction is performed on each input sample image to obtain the image features of the Nth iteration corresponding to each input sample image.
7. An object recognition apparatus characterized by comprising: The device includes: The acquisition module is used to acquire images of the target area. A prediction module is used to input the captured image into an object recognition model to obtain the predicted probability of the target object. The predicted probability refers to the probability that the captured image includes the target object. The object recognition model is trained using single-stage gradient loss and regression loss values in each iteration. The single-stage gradient loss value in each iteration represents the accuracy of the model's prediction of the target sample in that iteration, and the regression loss value in each iteration represents the degree of fit of the model's position to the target object in that iteration. The process of obtaining the single-stage gradient loss value in each iteration includes: obtaining the label, ground truth box, sample prediction probability in the Nth iteration, and prediction box in the Nth iteration for the sample image. The label indicates that the sample image includes the target object or the sample... The image does not include the target object. The sample prediction probability of the Nth iteration represents the probability that the training model in the Nth iteration predicts that the sample image includes the target object, where N is an integer greater than 1. The prediction box corresponding to the sample image in each iteration refers to the selection box predicted by the training model in that iteration for selecting the target object, and the ground truth box refers to the labeled selection box for selecting the target object. Based on the label and the sample prediction probability of the Nth iteration, the gradient value of the sample image in the Nth iteration is determined. The gradient interval of the Nth iteration is divided to obtain multiple sub-intervals of the Nth iteration. The degree of overlap between the prediction box and the ground truth box in the Nth iteration is determined as the overlap of the Nth iteration. Based on the multiple sub-intervals of the Nth iteration and the overlap of the Nth iteration, the single-stage gradient loss value of the Nth iteration is determined. The result determination module is used to determine the object recognition result for the target area based on the predicted probability.
8. A vehicle characterized by comprising: The vehicles include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the method as described in any one of claims 1-6.
9. A computer readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1-6.