Method, device and electronic device for generating autonomous driving perception model

By fusing the prediction results and labels of the teacher perception model and the initial autonomous driving perception model, a lightweight and performance-enhanced target autonomous driving perception model is generated, which solves the problem of insufficient vehicle-side deployment performance and achieves efficient model training and lightweight deployment.

CN115984791BActive Publication Date: 2025-10-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211646569.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-19
Publication Date
2025-10-03
Estimated Expiration
2042-12-19

AI Technical Summary

Technical Problem

The existing autonomous driving perception models have limited room for performance improvement when deployed on the vehicle side, making it difficult to meet the requirements of lightweight and high efficiency.

Method used

By obtaining a sample training dataset, the prediction results of the teacher perception model and the initial autonomous driving perception model are fused with the labels, and the initial model is corrected to generate the target autonomous driving perception model, realizing knowledge distillation and improving model performance.

Benefits of technology

The generated target autonomous driving perception model is lightweight and has improved performance, saving vehicle-side storage space and computing resources, and improving training speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984791B_ABST
    Figure CN115984791B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, device and electronic device for generating an autonomous driving perception model, which relates to the field of computer technology, and in particular to the field of artificial intelligence technology such as autonomous driving and deep learning. The method comprises: inputting a lane image into a teacher perception model and an initial autonomous driving perception model respectively to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model; fusing the first prediction result and the first label corresponding to the lane image to obtain a second label corresponding to the lane image; and correcting the initial autonomous driving perception model according to the difference between the second label and the second prediction result to obtain a target autonomous driving perception model. In this way, the initial autonomous driving perception model learns the knowledge learned by the teacher perception model through the second label, which not only obtains a lightweight target autonomous driving perception model, but also improves the performance of the generated target autonomous driving perception model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, in particular to the field of artificial intelligence technologies such as autonomous driving and deep learning, and specifically to a method, device, and electronic device for generating an autonomous driving perception model. Background Art

[0002] With the continuous development and improvement of artificial intelligence technology, it has played an extremely important role in various fields related to human daily life. For example, artificial intelligence has made significant progress in the field of autonomous driving.

[0003] Currently, autonomous driving perception models are often very lightweight due to their need for on-board deployment, leaving significant room for performance improvement. Therefore, improving the performance of on-board models has become a key research topic. Summary of the Invention

[0004] The present disclosure provides a method, device, and electronic device for generating an autonomous driving perception model.

[0005] According to a first aspect of the present disclosure, a method for generating an autonomous driving perception model is provided, comprising:

[0006] Obtain a sample training data set, wherein the sample training data set includes lane images and first labels corresponding to the lane images under the type of task to be trained;

[0007] Inputting the lane image into the teacher perception model and the initial autonomous driving perception model respectively to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model;

[0008] Fusing the first prediction result and the first label to obtain a second label corresponding to the lane image;

[0009] According to the difference between the second label and the second prediction result, the initial autonomous driving perception model is corrected to obtain a target autonomous driving perception model.

[0010] According to a second aspect of the present disclosure, a device for generating an autonomous driving perception model is provided, comprising:

[0011] A first acquisition module is configured to acquire a sample training data set, wherein the sample training data set includes a lane image and a first label corresponding to the lane image under a task type to be trained;

[0012] a second acquisition module, configured to input the lane image into the teacher perception model and the initial autonomous driving perception model, respectively, to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model;

[0013] a third acquisition module, configured to fuse the first prediction result and the first label to obtain a second label corresponding to the lane image;

[0014] A fourth acquisition module is used to correct the initial autonomous driving perception model according to the difference between the second label and the second prediction result to obtain a target autonomous driving perception model.

[0015] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0016] at least one processor; and

[0017] a memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for generating an autonomous driving perception model as described in the first aspect.

[0019] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method for generating an autonomous driving perception model as described in the first aspect.

[0020] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement the steps of the method for generating an autonomous driving perception model as described in the first aspect.

[0021] The method, device, and electronic device for generating an autonomous driving perception model provided by the present disclosure have the following beneficial effects:

[0022] In the disclosed embodiment, a sample training dataset is first obtained, wherein the sample training dataset includes a lane image and a first label corresponding to the lane image under the type of task to be trained. The lane image is then input into a teacher perception model and an initial autonomous driving perception model, respectively, to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model. The first prediction result and the first label are then fused to obtain a second label corresponding to the lane image. Finally, based on the difference between the second label and the second prediction result, the initial autonomous driving perception model is corrected to obtain a target autonomous driving perception model. Thus, by fusing the first label with the first prediction result output by the teacher perception model to obtain a second label for training the initial autonomous driving perception model, the initial autonomous driving perception model learns the knowledge learned by the teacher perception model through the second label, thereby not only obtaining a lightweight target autonomous driving perception model but also improving the performance of the generated target autonomous driving perception model.

[0023] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0025] Figure 1 This is a flowchart of a method for generating an autonomous driving perception model according to an embodiment of the present disclosure;

[0026] Figure 2 is a flowchart of a method for generating an autonomous driving perception model according to another embodiment of the present disclosure;

[0027] Figure 3 is a flowchart of a method for generating an autonomous driving perception model according to another embodiment of the present disclosure;

[0028] Figure 4 1 is a schematic structural diagram of a device for generating an autonomous driving perception model according to an embodiment of the present disclosure;

[0029] Figure 5 This is a block diagram of an electronic device used to implement the method for generating an autonomous driving perception model according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0030] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0031] The embodiments of the present disclosure relate to the fields of artificial intelligence technologies such as computer vision and deep learning.

[0032] Artificial Intelligence (AI) is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence.

[0033] Deep learning involves learning the inherent patterns and representational hierarchies of sample data. The information gained from this learning process is highly helpful in interpreting data such as text, images, and sounds. The ultimate goal of deep learning is to enable machines to have the same analytical and learning capabilities as humans, enabling them to recognize data such as text, images, and sounds.

[0034] Autonomous driving generally refers to autonomous driving systems. These systems utilize advanced communications, computing, networking, and control technologies to achieve real-time, continuous control of trains. Key technologies associated with autonomous driving systems include environmental perception, logical reasoning and decision-making, motion control, and processor performance.

[0035] The following describes the method, device, and electronic device for generating an autonomous driving perception model according to an embodiment of the present disclosure with reference to the accompanying drawings.

[0036] It should be noted that the executor of the method for generating an autonomous driving perception model in this embodiment is a device for generating an autonomous driving perception model. The device can be implemented by software and / or hardware, and the device can be configured in an electronic device. The electronic device may include but is not limited to a terminal, a server, etc.

[0037] Figure 1 This is a flow chart of a method for generating an autonomous driving perception model according to an embodiment of the present disclosure.

[0038] like Figure 1 As shown, the method for generating the autonomous driving perception model includes:

[0039] S101: Obtain a sample training data set, wherein the sample training data set includes lane images and first labels corresponding to the lane images under the type of task to be trained.

[0040] The sample training data set may include a large number of lane images and a first label corresponding to each lane image under each task type to be trained.

[0041] For example, if the training task is detecting yellow painted lane lines, the first label corresponding to the pixels of the yellow painted lane lines in the lane image is 1, and the label corresponding to the pixels of non-yellow painted lane lines in the lane image is 0. Alternatively, if the training task is detecting white painted lane lines, the first label corresponding to the pixels of the white painted lane lines in the lane image is 1, and the label corresponding to the pixels of non-white painted lane lines in the lane image is 0.

[0042] S102: Input the lane image into the teacher perception model and the initial autonomous driving perception model respectively to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model.

[0043] The teacher perception model can be a pre-trained large model for autonomous driving perception. The initial autonomous driving perception model can be a lightweight, untrained small model for autonomous driving sensors. It should be noted that the teacher perception model has a more complex structure, a larger amount of data, and higher accuracy. The initial autonomous driving perception model has a simpler structure and smaller amount of data than the teacher perception model.

[0044] Optionally, the initial parameters of the initial autonomous driving perception model can be determined based on the parameters of the teacher perception model.

[0045] The first prediction result may include the probability of each pixel pair in the sample image belonging to each category predicted by the teacher perception model. The second prediction result may include the probability of each pixel in the sample image belonging to each category predicted by the initial autonomous driving perception model.

[0046] Optionally, the teacher perception model may include a branch structure corresponding to multiple perception tasks. For example, the multiple perception tasks may include a perception task for white painted lane lines, a perception task for yellow painted lines, a perception task for free space, and so on. Similarly, the initial autonomous driving perception model may also include a score structure corresponding to multiple perception tasks. This disclosure is not limited to this.

[0047] S103: Fusing the first prediction result and the first label to obtain a second label corresponding to the lane image.

[0048] Optionally, corresponding weights may be set for the first prediction result and the first label respectively, and then the first prediction result and the second prediction result may be fused based on the weights corresponding to the first prediction result and the first label respectively to obtain the second label corresponding to the lane line.

[0049] For example, if the weight corresponding to the first prediction result is set to 0.5 and the weight corresponding to the first label is set to 0.5, then the second label can be determined to be 0.5*first prediction result+0.5*first label.

[0050] Alternatively, the weight corresponding to the first prediction result can be determined based on the overall loss value corresponding to the first prediction result. The smaller the overall loss value, the greater the weight corresponding to the first prediction result; the larger the overall loss value, the smaller the weight corresponding to the first prediction result. After determining the weight corresponding to the first prediction result, the weight corresponding to the first label is determined based on the sum of the weights corresponding to the first prediction result and the first label being 1.

[0051] S104: According to the difference between the second label and the second prediction result, the initial autonomous driving perception model is corrected to obtain a target autonomous driving perception model.

[0052] In the disclosed embodiment, since the second label fuses the first prediction result and the first label, the loss of the second prediction result is determined based on the difference between the second label and the second prediction result, and then the initial autonomous driving perception model is corrected according to the loss. Therefore, the knowledge learned by the teacher perception model can be transferred to the initial autonomous driving perception model through the second label that fuses the first prediction result, so that the initial autonomous driving perception model can quickly learn the knowledge learned by the teacher perception model, thereby improving the training speed and performance of the initial autonomous driving perception model.

[0053] Optionally, a cross entropy loss function may be used to calculate the difference between the second label and the second prediction result.

[0054] In the disclosed embodiment, the target autonomous driving perception model is obtained by distilling and learning the teacher perception model. The target autonomous driving perception model is small in size, high in accuracy, and requires fewer computing resources than the teacher perception model. Therefore, the trained target autonomous driving perception model is deployed on the vehicle side, which can save storage space and computing resources on the vehicle side.

[0055] In the disclosed embodiment, a sample training dataset is first obtained, wherein the sample training dataset includes a lane image and a first label corresponding to the lane image under the type of task to be trained. The lane image is then input into a teacher perception model and an initial autonomous driving perception model, respectively, to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model. The first prediction result and the first label are then fused to obtain a second label corresponding to the lane image. Finally, based on the difference between the second label and the second prediction result, the initial autonomous driving perception model is corrected to obtain a target autonomous driving perception model. Thus, by fusing the first label with the first prediction result output by the teacher perception model to obtain a second label for training the initial autonomous driving perception model, the initial autonomous driving perception model learns the knowledge learned by the teacher perception model through the second label, thereby not only obtaining a lightweight target autonomous driving perception model but also improving the performance of the generated target autonomous driving perception model.

[0056] Figure 2 is a flow chart of a method for generating an autonomous driving perception model according to another embodiment of the present disclosure; Figure 2 As shown, the method for generating the autonomous driving perception model includes:

[0057] S201: Obtain a sample training data set, wherein the sample training data set includes lane images and first labels corresponding to the lane images under the type of task to be trained.

[0058] S202: Input the lane image into the teacher perception model and the initial autonomous driving perception model respectively to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model.

[0059] The specific implementation of step S201 and step S202 can refer to the detailed description in other embodiments of the present disclosure, and will not be described in detail here.

[0060] S203: Determine a first weight factor corresponding to the first prediction result of each pixel in the lane image according to a difference between the first prediction result of each pixel in the lane image and the first label.

[0061] In the disclosed embodiment, a first weighting factor corresponding to the first prediction result for each pixel in the lane image can be determined based on the difference between the first prediction result and the first label for each pixel. A larger difference corresponds to a smaller first weighting factor, and a smaller difference corresponds to a larger first weighting factor. This allows the initial autonomous driving perception model to better learn the knowledge from the teacher perception model, which has more accurate prediction results. This improves the training efficiency and performance of the initial autonomous driving perception model.

[0062] Optionally, a target loss value corresponding to the first prediction result for each pixel in the lane image can be determined based on the difference between the first prediction result and the first label for each pixel in the lane image. Subsequently, a first weight factor corresponding to the first prediction result for each pixel in the lane image can be determined based on the mapping relationship between the loss value and the weight factor and the target loss value corresponding to the first prediction result for each pixel in the lane image. Thus, the target loss value can accurately reflect the accuracy of the first prediction result of each pixel by the teacher perception model, thereby improving the accuracy of the determined first weight factor.

[0063] In the mapping relationship between loss values ​​and weight factors, larger loss values ​​indicate less accurate predictions and smaller corresponding weight factors; smaller loss values ​​indicate more accurate predictions and larger corresponding weight factors. This allows the initial autonomous driving perception model to better learn from the more accurate predictions of the teacher perception model.

[0064] Optionally, a cross entropy loss function may be used to calculate the difference between the first prediction result and the first label for each pixel in the lane image, and the corresponding target loss value.

[0065] S204: Determine a second weight factor corresponding to each pixel in the lane image according to the type of the task to be trained.

[0066] In the disclosed embodiment, the second weighting factor corresponding to each pixel in the lane image varies depending on the training task type. Optionally, the second weighting factor corresponding to each pixel in the training task type may be determined based on whether the first label corresponding to the training task type contains subjective bias from the annotator.

[0067] Specifically, if the first label corresponding to the task type to be trained is based on paint lines and does not contain subjective tendencies of the annotator, such as the detection of white painted lane lines, the detection of yellow painted lines, etc., then the corresponding second weight factor can be 1, 0.9, etc. If the first label corresponding to the task type to be trained lacks the support of paint lines, it often carries the subjective tendencies of the annotator, such as the free space detection training task type. For the task type to be processed with subjective noise in the first label, the corresponding second weight factor can be appropriately increased, for example, the second weight factor can be 1.2, 1.1, etc. This disclosure is not limited to this.

[0068] Therefore, when there is a subjective tendency of the standard person in the first label corresponding to the task type to be trained, the second weight factor can be appropriately increased, thereby increasing the target weight of the first prediction result when generating the second label, so that the initial autonomous driving perception model can learn more about the first prediction result of the teacher perception model for the sample image under the task type to be trained, thereby improving the training efficiency and performance of the initial autonomous driving perception model.

[0069] Optionally, based on the first label, the positive samples and negative samples contained in the lane image can be determined, and then based on the type of task to be trained, the weight factor mapping table can be queried to obtain the first value corresponding to the positive sample of the task type to be trained and the second value corresponding to the negative sample, wherein the first value is greater than the second value. Finally, it is determined that the second weight factor corresponding to each pixel point in the positive sample of the lane image is the first value, and the second weight factor corresponding to each pixel point in the negative sample of the lane image is the second value.

[0070] The weight factor mapping table may be pre-generated and include weight factors (ie, first values) corresponding to positive samples of each task type to be trained and weight factors (ie, second values) corresponding to negative samples.

[0071] In the disclosed embodiment, different second weight factors can be set for positive samples and negative samples in the sample image, and the second weight factor corresponding to the positive sample is greater than the second weight factor corresponding to the negative sample, so that the initial autonomous driving perception model can learn more about the first prediction results of the teacher perception model for the positive samples in the sample image during the training process.

[0072] S205: Obtain a first initial weight corresponding to the first prediction result of each pixel in the lane image, and a second initial weight corresponding to the first label.

[0073] The first initial weight and the second initial weight may be pre-set. The first prediction result of each pixel in the lane image corresponds to the same first initial weight, and the first label of each pixel in the lane image corresponds to the same second initial weight.

[0074] The sum of the first initial weight and the second initial weight may be 1. The first initial weight and the second initial weight may be the same or different, which is not limited in this disclosure.

[0075] S206: Determine the product of the first initial weight, the first weight factor, and the second weight factor corresponding to the first prediction result of each pixel point in the lane image as the target weight corresponding to the first prediction result of each pixel point in the lane image.

[0076] In the embodiment of the present disclosure, after determining the first initial weight, first weight factor and second weight factor corresponding to the first prediction result of each pixel point in the lane image, the target weight corresponding to the first prediction result of each pixel point can be accurately determined, thereby accurately determining the weight of the knowledge distillation of the first prediction result predicted by the teacher perception model by the initial autonomous driving perception model, thereby enabling the initial autonomous driving perception model to better learn the knowledge of the teacher model.

[0077] S207: Based on the second initial weight corresponding to the first label of each pixel in the lane image and the target weight corresponding to the first prediction result, the first prediction result and the first label of each pixel in the lane image are fused to obtain the second label corresponding to each pixel in the lane image.

[0078] In an embodiment of the present disclosure, after determining the second initial weight corresponding to the first label of each pixel point in the lane image and the target weight corresponding to the first prediction result, the weighted sum corresponding to each pixel point can be determined as the initial second label corresponding to each pixel point in the lane image, and then the initial second label corresponding to each pixel point is normalized to determine the second label corresponding to each pixel point in the lane image.

[0079] S208: Generate a second label corresponding to the lane image based on the second label corresponding to each pixel in the lane image.

[0080] It is understandable that after determining the second label corresponding to each pixel in the lane image, the second label corresponding to each pixel can be combined to obtain the second label corresponding to the lane image.

[0081] S209: According to the difference between the second label and the second prediction result, the initial autonomous driving perception model is corrected to obtain a target autonomous driving perception model.

[0082] In an embodiment of the present disclosure, the lane images in the sample training data set are first input into the teacher perception model and the initial autonomous driving perception model respectively to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model. Then, based on the difference between the first prediction result and the first label of each pixel in the lane image, the first weight factor corresponding to the first prediction result of each pixel in the lane image is determined. According to the type of task to be trained, the second weight factor corresponding to each pixel in the lane image is determined. Then, based on the product of the first initial weight, the first weight factor and the second weight factor corresponding to the first prediction result of each pixel, the target weight corresponding to the first prediction result of each pixel in the lane image is determined. Based on the second initial weight and the target weight corresponding to the first label, the first prediction result and the first label corresponding to each pixel are fused to obtain the second label corresponding to the lane image. Finally, based on the difference between the second label and the second prediction result, the initial autonomous driving perception model is corrected to obtain the target autonomous driving perception model. Therefore, the target weight corresponding to the first prediction result of each pixel point in the sample image can be determined according to the accuracy of the first prediction result corresponding to each pixel point in the sample image and the type of task to be trained, and then based on the target weight, the first label is fused with the first prediction result, so that the initial autonomous driving perception model can better learn the knowledge learned by the teacher perception model based on the second label, thereby further improving the performance of the acquired target autonomous driving perception model.

[0083] Figure 3 is a flow chart of a method for generating an autonomous driving perception model according to another embodiment of the present disclosure; Figure 3 As shown, the method for generating the autonomous driving perception model includes:

[0084] S301: Obtain a sample training data set, wherein the sample training data set includes lane images and first labels corresponding to the lane images under the type of task to be trained.

[0085] S302: Input the lane image into the teacher perception model and the initial autonomous driving perception model respectively to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model.

[0086] S303: Determine a first weight factor corresponding to the first prediction result of each pixel in the lane image based on a difference between the first prediction result of each pixel in the lane image and the first label.

[0087] S304: Determine a second weight factor corresponding to each pixel in the lane image according to the type of the task to be trained.

[0088] S305: Obtain a first initial weight corresponding to the first prediction result of each pixel in the lane image, and a second initial weight corresponding to the first label.

[0089] S306: Determine the product of the first initial weight, the first weight factor, and the second weight factor corresponding to the first prediction result of each pixel point in the lane image as the target weight corresponding to the first prediction result of each pixel point in the lane image.

[0090] The specific implementation of step S301 and step S306 can refer to the detailed description in other embodiments of the present disclosure, and will not be described in detail here.

[0091] S307: Determine the ratio between the number of pixels corresponding to the positive samples in the lane image and the total number of pixels in the lane image.

[0092] S308: Determine a third weighting factor corresponding to each pixel in the lane image according to the ratio.

[0093] It should be noted that the smaller the ratio, the smaller the proportion of positive samples in the sample image. It is necessary to increase the third weight factor corresponding to the pixels of the positive samples, so that the initial autonomous driving perception model can better learn the first prediction results of the teacher perception model for the positive samples in the sample image during the training process.

[0094] Optionally, when the ratio is less than the first threshold and greater than the second threshold, the third weight factor corresponding to each pixel point in the positive sample of the lane image is determined to be a third value, and the third weight factor corresponding to each pixel point in the negative sample of the lane image is determined to be a fourth value, wherein the third value is greater than the fourth value.

[0095] For example, the first threshold may be 10%, 15%, etc. The second threshold may be 5%, 8%, etc. This disclosure does not limit this.

[0096] In the disclosed embodiment, the third weight factor corresponding to the positive sample is greater than the third weight factor corresponding to the negative sample, thereby appropriately increasing the target weight corresponding to the first prediction result of the positive sample in the sample image, further enabling the initial autonomous driving perception model to better learn the first prediction result of the teacher perception model for the positive sample in the sample image during the training process.

[0097] Alternatively, when the ratio is less than or equal to the second threshold, the third weight factor corresponding to each pixel point in the positive sample of the lane image is determined to be the fifth value, and the third weight factor corresponding to each pixel point in the negative sample of the lane image is 0, where the fifth value is greater than the third value.

[0098] It should be noted that when the ratio is less than or equal to the second threshold, it means that the proportion of positive samples in the sample image is very small. At this time, the third weight factor corresponding to each pixel point in the negative sample of the lane image can be set to 0, so that the target weight corresponding to the first prediction result of each pixel point in the negative sample is 0, thereby achieving the purpose of only distilling positive samples, so that the initial autonomous driving perception model only learns the first prediction result of the teacher perception model for the positive samples in the sample image during the training process.

[0099] S309: Update the target weight corresponding to the first prediction result of each pixel in the lane image according to the product of the first weight factor, the second weight factor, the third weight factor and the first initial weight corresponding to each pixel in the lane image.

[0100] In the disclosed embodiment, the third weight factor corresponding to each pixel in the lane image can be further determined based on the ratio between the number of pixel points corresponding to the positive samples in the lane image and the total number of pixel points in the lane image, and the target weight can be updated according to the third weight factor, so that the determined target weight is more accurate.

[0101] S310: Based on the second initial weight corresponding to the first label of each pixel in the lane image and the target weight corresponding to the first prediction result, the first prediction result and the first label of each pixel in the lane image are fused to obtain the second label corresponding to each pixel in the lane image.

[0102] S311: Generate a second label corresponding to the lane image based on the second label corresponding to each pixel in the lane image.

[0103] S312: According to the difference between the second label and the second prediction result, the initial autonomous driving perception model is corrected to obtain a target autonomous driving perception model.

[0104] In an embodiment of the present disclosure, lane images in a sample training dataset are first input into a teacher perception model and an initial autonomous driving perception model, respectively, to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model. A first weight factor corresponding to the first prediction result of each pixel in the lane image is then determined based on the difference between the first prediction result and the first label of each pixel in the lane image. A second weight factor corresponding to each pixel in the lane image is determined based on the type of task to be trained. A third weight factor corresponding to each pixel in the lane image is determined based on the ratio between the number of pixels corresponding to positive samples in the lane image and the total number of pixels in the lane image. A target weight corresponding to the first prediction result of each pixel in the lane image is then determined based on the product of the first initial weight, the first weight factor, the second weight factor, and the third weight factor corresponding to the first prediction result of each pixel. The first prediction result and the first label corresponding to each pixel are fused based on the second initial weight and the target weight corresponding to the first label to obtain a second label corresponding to the lane image. Finally, the initial autonomous driving perception model is modified based on the difference between the second label and the second prediction result to obtain a target autonomous driving perception model. Therefore, based on the accuracy of the first prediction result corresponding to each pixel point in the sample image, the type of task to be trained and the proportion of positive samples in the sample image, the target weight corresponding to the first prediction result of each pixel point can be further accurately determined, so that the initial autonomous driving perception model and the second label can better learn the knowledge learned by the teacher perception model, thereby further improving the performance of the acquired target autonomous driving perception model.

[0105] Figure 4 is a structural diagram of a device for generating an autonomous driving perception model according to an embodiment of the present disclosure; Figure 4 As shown, the device 400 for generating the autonomous driving perception model includes:

[0106] A first acquisition module 410 is configured to acquire a sample training dataset, wherein the sample training dataset includes lane images and first labels corresponding to the lane images under the type of task to be trained;

[0107] A second acquisition module 420 is configured to input the lane image into the teacher perception model and the initial autonomous driving perception model, respectively, to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model;

[0108] A third acquisition module 430 is configured to fuse the first prediction result and the first label to obtain a second label corresponding to the lane image;

[0109] The fourth acquisition module 440 is used to correct the initial autonomous driving perception model according to the difference between the second label and the second prediction result to obtain the target autonomous driving perception model.

[0110] In some embodiments of the present disclosure, the third acquisition module 430 includes:

[0111] a first determining unit, configured to determine a first weight factor corresponding to the first prediction result for each pixel in the lane image based on a difference between the first prediction result for each pixel in the lane image and the first label;

[0112] A second determining unit is used to determine a second weight factor corresponding to each pixel in the lane image according to the type of the task to be trained;

[0113] A first acquisition unit is used to acquire a first initial weight corresponding to the first prediction result of each pixel in the lane image, and a second initial weight corresponding to the first label;

[0114] a third determining unit, configured to determine a target weight corresponding to the first prediction result for each pixel point in the lane image by multiplying the first initial weight, the first weight factor, and the second weight factor corresponding to the first prediction result for each pixel point in the lane image;

[0115] a second acquisition unit, configured to fuse the first prediction result and the first label of each pixel in the lane image based on a second initial weight corresponding to the first label of each pixel in the lane image and a target weight corresponding to the first prediction result, so as to obtain a second label corresponding to each pixel in the lane image;

[0116] A generating unit is configured to generate a second label corresponding to the lane image based on the second label corresponding to each pixel in the lane image.

[0117] In some embodiments of the present disclosure, the first determining unit is specifically configured to:

[0118] determining a target loss value corresponding to the first prediction result for each pixel in the lane image based on a difference between the first prediction result and the first label for each pixel in the lane image;

[0119] According to the mapping relationship between the loss value and the weight factor, and the target loss value corresponding to the first prediction result of each pixel point in the lane image, the first weight factor corresponding to the first prediction result of each pixel point in the lane image is determined.

[0120] In some embodiments of the present disclosure, the second determining unit is configured to:

[0121] Based on the first label, determining positive samples and negative samples contained in the lane image;

[0122] Based on the type of the task to be trained, query the weight factor mapping table to obtain a first value corresponding to the positive sample of the task type to be trained and a second value corresponding to the negative sample, wherein the first value is greater than the second value;

[0123] It is determined that the second weight factor corresponding to each pixel in the positive sample of the lane image is a first value, and the second weight factor corresponding to each pixel in the negative sample of the lane image is a second value.

[0124] In some embodiments of the present disclosure, the third acquisition module further includes:

[0125] a fourth determining unit, configured to determine a ratio between the number of pixels corresponding to the positive samples in the lane image and the total number of pixels in the lane image;

[0126] a fifth determining unit, configured to determine a third weighting factor corresponding to each pixel in the lane image according to the ratio;

[0127] An updating unit is used to update the target weight corresponding to the first prediction result of each pixel point in the lane image according to the product of the first weight factor, the second weight factor, the third weight factor and the first initial weight corresponding to each pixel point in the lane image.

[0128] In some embodiments of the present disclosure, the fifth determining unit is specifically configured to:

[0129] When the ratio is less than the first threshold and greater than the second threshold, determining that the third weight factor corresponding to each pixel in the positive sample of the lane image is a third value, and the third weight factor corresponding to each pixel in the negative sample of the lane image is a fourth value, wherein the third value is greater than the fourth value; or

[0130] When the ratio is less than or equal to the second threshold, the third weight factor corresponding to each pixel in the positive sample of the lane image is determined to be the fifth value, and the third weight factor corresponding to each pixel in the negative sample of the lane image is 0, where the fifth value is greater than the third value.

[0131] It should be noted that the above explanation of the method for generating an autonomous driving perception model is also applicable to the device for generating an autonomous driving perception model in this embodiment and will not be repeated here.

[0132] In the disclosed embodiment, a sample training dataset is first obtained, wherein the sample training dataset includes a lane image and a first label corresponding to the lane image under the type of task to be trained. The lane image is then input into a teacher perception model and an initial autonomous driving perception model, respectively, to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model. The first prediction result and the first label are then fused to obtain a second label corresponding to the lane image. Finally, based on the difference between the second label and the second prediction result, the initial autonomous driving perception model is corrected to obtain a target autonomous driving perception model. Thus, by fusing the first label with the first prediction result output by the teacher perception model to obtain a second label for training the initial autonomous driving perception model, the initial autonomous driving perception model learns the knowledge learned by the teacher perception model through the second label, thereby not only obtaining a lightweight target autonomous driving perception model but also improving the performance of the generated target autonomous driving perception model.

[0133] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0134] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0135] like Figure 5 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0136] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0137] The computing unit 501 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the method for generating an autonomous driving perception model. For example, in some embodiments, the method for generating an autonomous driving perception model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method for generating an autonomous driving perception model described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the method for generating an autonomous driving perception model by any other suitable means (e.g., via firmware).

[0138] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0139] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0140] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0141] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0142] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0143] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.

[0144] In this embodiment, a sample training dataset is first obtained, wherein the sample training dataset includes lane images and first labels corresponding to the lane images for the type of task to be trained. The lane images are then input into a teacher perception model and an initial autonomous driving perception model, respectively, to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model. The first prediction result and the first label are then fused to obtain a second label corresponding to the lane image. Finally, based on the difference between the second label and the second prediction result, the initial autonomous driving perception model is modified to obtain a target autonomous driving perception model. Thus, by fusing the first label with the first prediction result output by the teacher perception model to obtain a second label for training the initial autonomous driving perception model, the initial autonomous driving perception model learns the knowledge learned by the teacher perception model through the second label, thereby not only obtaining a lightweight target autonomous driving perception model but also improving the performance of the generated target autonomous driving perception model.

[0145] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0146] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present disclosure, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In the description of the present disclosure, the words "if" and "if" used can be interpreted as "at the time of" or "when" or "in response to a determination" or "under the circumstances of".

[0147] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for generating an autonomous driving perception model, comprising: Obtain a sample training data set, wherein the sample training data set includes lane images and first labels corresponding to the lane images under the type of task to be trained; Inputting the lane image into the teacher perception model and the initial autonomous driving perception model respectively to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model; Fusing the first prediction result and the first label to obtain a second label corresponding to the lane image; According to the difference between the second label and the second prediction result, the initial autonomous driving perception model is modified to obtain a target autonomous driving perception model; wherein fusing the first prediction result and the first label to obtain the second label includes: determining a first weight factor corresponding to the first prediction result for each pixel in the lane image based on a difference between the first prediction result for each pixel in the lane image and the first label; Determining a second weight factor corresponding to each pixel in the lane image according to the type of the task to be trained; Obtaining a first initial weight corresponding to the first prediction result of each pixel in the lane image and a second initial weight corresponding to the first label; Determine the target weight corresponding to the first prediction result for each pixel in the lane image by multiplying the first initial weight, the first weight factor, and the second weight factor corresponding to the first prediction result for each pixel in the lane image; fusing the first prediction result and the first label for each pixel in the lane image based on the second initial weight corresponding to the first label of each pixel in the lane image and the target weight corresponding to the first prediction result to obtain a second label corresponding to each pixel in the lane image; Based on the second label corresponding to each pixel in the lane image, the second label corresponding to the lane image is generated.

2. The method according to claim 1, wherein The determining, based on a difference between the first prediction result of each pixel in the lane image and the first label, a first weight factor corresponding to the first prediction result of each pixel in the lane image includes: determining a target loss value corresponding to the first prediction result for each pixel in the lane image based on a difference between the first prediction result and the first label for each pixel in the lane image; According to the mapping relationship between the loss value and the weight factor, and the target loss value corresponding to the first prediction result of each pixel point in the lane image, the first weight factor corresponding to the first prediction result of each pixel point in the lane image is determined.

3. The method according to claim 1, wherein Determining, according to the type of the task to be trained, a second weight factor corresponding to each pixel in the lane image includes: Determining positive samples and negative samples contained in the lane image based on the first label; Based on the task type to be trained, query a weight factor mapping table to obtain a first value corresponding to a positive sample of the task type to be trained and a second value corresponding to a negative sample, wherein the first value is greater than the second value; Determine that a second weight factor corresponding to each pixel in the positive sample of the lane image is the first value, and a second weight factor corresponding to each pixel in the negative sample of the lane image is the second value.

4. The method according to claim 3, wherein: Also includes: Determining a ratio between the number of pixels corresponding to the positive samples in the lane image and the total number of pixels in the lane image; determining a third weight factor corresponding to each pixel in the lane image according to the ratio; The target weight corresponding to the first prediction result of each pixel point in the lane image is updated according to the product of the first weight factor, the second weight factor, the third weight factor and the first initial weight corresponding to each pixel point in the lane image.

5. The method according to claim 4, wherein Determining a third weight factor corresponding to each pixel in the lane image according to the ratio includes: When the ratio is less than the first threshold and greater than the second threshold, determining that the third weight factor corresponding to each pixel in the positive sample of the lane image is a third value, and the third weight factor corresponding to each pixel in the negative sample of the lane image is a fourth value, wherein the third value is greater than the fourth value; or When the ratio is less than or equal to the second threshold, the third weight factor corresponding to each pixel in the positive sample of the lane image is determined to be a fifth value, and the third weight factor corresponding to each pixel in the negative sample of the lane image is 0, wherein the fifth value is greater than the third value.

6. A device for generating an autonomous driving perception model, comprising: A first acquisition module is configured to acquire a sample training data set, wherein the sample training data set includes a lane image and a first label corresponding to the lane image under a task type to be trained; a second acquisition module, configured to input the lane image into the teacher perception model and the initial autonomous driving perception model, respectively, to obtain a first prediction result output by the teacher perception model and a second prediction result output by the initial autonomous driving perception model; a third acquisition module, configured to fuse the first prediction result and the first label to obtain a second label corresponding to the lane image; a fourth acquisition module, configured to modify the initial autonomous driving perception model according to a difference between the second label and the second prediction result to obtain a target autonomous driving perception model; Wherein, the third acquisition module includes: a first determining unit, configured to determine a first weight factor corresponding to the first prediction result for each pixel in the lane image based on a difference between the first prediction result for each pixel in the lane image and the first label; A second determining unit is configured to determine a second weight factor corresponding to each pixel in the lane image according to the type of the task to be trained; A first acquisition unit is configured to acquire a first initial weight corresponding to a first prediction result of each pixel in the lane image, and a second initial weight corresponding to the first label; a third determining unit, configured to determine, by multiplying the first initial weight, the first weight factor, and the second weight factor corresponding to the first prediction result for each pixel in the lane image, a target weight corresponding to the first prediction result for each pixel in the lane image; a second acquiring unit, configured to fuse the first prediction result and the first label for each pixel in the lane image based on the second initial weight corresponding to the first label of each pixel in the lane image and the target weight corresponding to the first prediction result, so as to acquire a second label corresponding to each pixel in the lane image; A generating unit is configured to generate the second label corresponding to the lane image based on the second label corresponding to each pixel in the lane image.

7. The device according to claim 6, wherein The first determining unit is specifically configured to: determining a target loss value corresponding to the first prediction result for each pixel in the lane image based on a difference between the first prediction result and the first label for each pixel in the lane image; According to the mapping relationship between the loss value and the weight factor, and the target loss value corresponding to the first prediction result of each pixel point in the lane image, the first weight factor corresponding to the first prediction result of each pixel point in the lane image is determined.

8. The device according to claim 6, wherein The second determining unit is configured to: Determining positive samples and negative samples contained in the lane image based on the first label; Based on the task type to be trained, query a weight factor mapping table to obtain a first value corresponding to a positive sample of the task type to be trained and a second value corresponding to a negative sample, wherein the first value is greater than the second value; Determine that a second weight factor corresponding to each pixel in the positive sample of the lane image is the first value, and a second weight factor corresponding to each pixel in the negative sample of the lane image is the second value.

9. The device according to claim 8, wherein The third acquisition module further includes: a fourth determining unit, configured to determine a ratio between the number of pixels corresponding to the positive samples in the lane image and the total number of pixels in the lane image; a fifth determining unit, configured to determine, based on the ratio, a third weighting factor corresponding to each pixel in the lane image; An updating unit is used to update the target weight corresponding to the first prediction result of each pixel point in the lane image according to the product of the first weight factor, the second weight factor, the third weight factor and the first initial weight corresponding to each pixel point in the lane image.

10. The device according to claim 9, wherein The fifth determining unit is specifically configured to: When the ratio is less than the first threshold and greater than the second threshold, determining that the third weight factor corresponding to each pixel in the positive sample of the lane image is a third value, and the third weight factor corresponding to each pixel in the negative sample of the lane image is a fourth value, wherein the third value is greater than the fourth value; or When the ratio is less than or equal to the second threshold, the third weight factor corresponding to each pixel in the positive sample of the lane image is determined to be a fifth value, and the third weight factor corresponding to each pixel in the negative sample of the lane image is 0, wherein the fifth value is greater than the third value.

11. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 5.

13. A computer program product comprising computer instructions, wherein when the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Image classification method and device, readable storage medium and terminal equipment

    CN110147456A

  • Method, device and equipment for training model, medium and product

    CN113392984A