Automatic driving algorithm optimization method and device, and computer program product
By collecting and optimizing virtual driving instructions for the autonomous driving system in the same driving scenario, the problem that autonomous driving algorithms are difficult to adapt to individual differences between different drivers is solved, and safety and comfort are improved.
Patent Information
- Application Number
- CN202510467658.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-22
AI Technical Summary
Existing autonomous driving algorithms are difficult to adapt to individual differences between different drivers, resulting in insufficient safety and comfort, especially learning an unsafe driving style may pose risks.
By collecting the differential data of the actual driving instructions of the human driver and the virtual driving instructions of the autonomous driving system in the same driving scenario, the server divides the virtual driving instructions into three types: conservative, moderate and radical, selects the moderate type as a positive sample for optimization training, uses vehicle perception data for supervision and learning, and optimizes the autonomous driving algorithm.
The continuous iterative optimization of the autonomous driving algorithm has been realized, making it closer to most people's driving habits, improving safety and comfort, and reducing the risks caused by learning unsafe styles.
Smart Images

Figure CN120354607A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and particularly relates to a method and device for optimizing an autonomous driving algorithm and a computer program product. Background Art
[0002] As an important means for the continuous iterative upgrade of an autonomous driving model algorithm based on a data closed-loop method, the shadow mode provides strong support for the development of autonomous driving technology through the real-time data collection and analysis of the differences between human driving behaviors.
[0003] Currently, the shadow mode usually collects both driver behavior data and autonomous driving system behavior data at the same time, and optimizes them through least squares fitting. The least squares fitting method can find an optimal model to minimize the error between the predicted value (autonomous driving system behavior) and the actual value (driver behavior), so that the driving style of the autonomous driving system is as close as possible to that of the driver. However, the driving styles of drivers vary, and the data of each driver needs to be fitted separately. Moreover, not all driving styles of drivers are safe and comfortable. If the autonomous driving system learns an unsafe driving style, it may bring safety risks during actual driving. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and device for optimizing an autonomous driving algorithm and a computer program product, so as to achieve the continuous iterative optimization of the autonomous driving algorithm, and make the optimized autonomous driving algorithm more suitable for the driving habits of most people, meeting the requirements of driving safety and comfort.
[0005] To achieve the above purpose, according to the first aspect of the present invention, a method for optimizing an autonomous driving algorithm is provided. The method includes:
[0006] During the process of a human driver manually driving a vehicle, run the autonomous driving system;
[0007] Obtain the actual driving instructions of the human driver and the virtual driving instructions output by the autonomous driving system based on vehicle perception data in the same driving scenario;
[0008] When there is a difference between the actual driving instructions and the virtual driving instructions, collect the actual driving instructions, vehicle perception data, and virtual driving instructions, and upload them to the server, so that the server divides the virtual driving instructions into three types: conservative, moderate, and aggressive according to the difference between the actual driving instructions and the virtual driving instructions, select at least one type of virtual driving instruction as a positive sample, at least one type of virtual driving instruction as a negative sample, and optimize and train the autonomous driving algorithm according to the vehicle perception data, positive samples, and negative samples; the positive samples include at least moderate-type virtual driving instructions.
[0009] According to a second aspect of the present invention, there is provided a method for optimizing an autonomous driving algorithm, the method comprising:
[0010] Obtain the human-machine driving behavior difference data uploaded by the vehicle; wherein, the human-machine driving behavior difference data includes the actual driving instructions and virtual driving instructions that are different in the same driving scenario and the vehicle perception data corresponding to the driving scenario, the actual driving instructions are the driving instructions of a human driver, and the virtual driving instructions are the virtual driving instructions output by the autonomous driving algorithm based on the vehicle perception data;
[0011] According to the difference between the actual driving instruction and the virtual driving instruction, divide the virtual driving instruction into three types: conservative, moderate, and aggressive;
[0012] Select at least one type of virtual driving instruction as a positive sample and at least one type of virtual driving instruction as a negative sample; the positive sample includes at least the moderate type of virtual driving instruction;
[0013] Use the vehicle perception data as the input of the autonomous driving algorithm, and use the positive sample and the negative sample as supervision signals to guide the optimization training of the autonomous driving algorithm.
[0014] According to a third aspect of the present invention, there is provided an apparatus for optimizing an autonomous driving algorithm, including a module for executing the method described in the first aspect or the second aspect of the present invention.
[0015] According to a fourth aspect of the present invention, there is provided an apparatus for optimizing an autonomous driving algorithm, including:
[0016] A communication interface for communicating with other electronic devices;
[0017] A memory for storing computer program instructions;
[0018] A processor for executing the computer program instructions to support the apparatus to implement the method described in the first aspect or the second aspect of the present invention.
[0019] According to a fifth aspect of the present invention, there is provided a computer program product, including computer program instructions, and the computer program instructions instruct a computer device to execute the operations corresponding to the method described in the first aspect or the second aspect of the present invention.
[0020] An autonomous driving algorithm optimization method, apparatus, and computer program product proposed by the present invention have the following beneficial effects:
[0021] During the driving process of the vehicle, the vehicle end collects the data of the differences in human driving behaviors and uploads it to the cloud server. Based on the data of the differences in human driving behaviors of the vehicle, through the differential analysis of the virtual driving instructions and the actual driving instructions, the virtual driving instructions are divided into three types: conservative, moderate, and aggressive, and appropriate positive and negative samples are selected for optimization training, which can make the autonomous driving algorithm closer to the safe and comfortable driving habits of most people and reduce the risks brought by learning unsafe driving styles. In summary, the present invention can achieve the continuous iterative optimization of the autonomous driving algorithm, and compared with the traditional shadow mode, the present invention can make the optimized autonomous driving algorithm more suitable for the driving habits of most people and meet the requirements of driving safety and comfort. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a flowchart of an autonomous driving algorithm optimization method in Embodiment 1 of the present invention.
[0024] Figure 2 It is a schematic diagram of the differences in human and machine driving behaviors.
[0025] Figure 3 It is a schematic diagram of the distribution of driving style data.
[0026] Figure 4 It is a flowchart of an autonomous driving algorithm optimization method in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The detailed description of the drawings is intended to be an illustration of the current embodiments of the present invention, rather than an indication of the only form in which the present invention can be implemented. It should be understood that the same or equivalent functions can be achieved by different embodiments intended to be included within the spirit and scope of the present invention.
[0028] Refer to Figure 1 , Embodiment 1 of the present invention proposes an autonomous driving algorithm optimization method, which is applied to the vehicle end. The method includes:
[0029] Step S11, during the process of a human driver manually driving the vehicle, run the autonomous driving system;
[0030] Specifically, when a human driver manually drives a vehicle, the autonomous driving system runs in the background, but the autonomous driving system does not actually control the vehicle. Instead, it only outputs virtual driving instructions for differential comparison with the actual driving instructions of the human driver.
[0031] Step S12: Obtain the actual driving instructions of the human driver and the virtual driving instructions output by the autonomous driving system based on vehicle perception data in the same driving scenario.
[0032] Specifically, vehicle perception data refers to the perception data obtained by the vehicle's sensing system for detecting the current scene environment. The sensing system includes, but is not limited to, cameras, lidar, ultrasonic sensors, etc. At each moment, there are both the actual driving instructions of the human driver and the virtual driving instructions of the autonomous driving system / algorithm.
[0033] Step S13: When there is a difference between the actual driving instructions and the virtual driving instructions, collect the actual driving instructions, vehicle perception data, and virtual driving instructions and upload them to the server, so that the server classifies the virtual driving instructions into three types: conservative, moderate, and aggressive based on the difference between the actual driving instructions and the virtual driving instructions, select at least one type of virtual driving instruction as the positive sample, at least one type of virtual driving instruction as the negative sample, and optimize and train the autonomous driving algorithm based on the vehicle perception data, positive samples, and negative samples; the positive samples include at least moderate-type virtual driving instructions.
[0034] Specifically, the vehicle continuously monitors and compares the actual driving instructions and the virtual driving instructions to determine whether there is a significant difference between them. For example Figure 2 as shown, when facing an obstacle vehicle, the human driver chooses to detour, while the autonomous driving system chooses to decelerate and avoid. Once a command difference is detected, data collection is immediately triggered. The collected content includes the actual driving instructions, vehicle perception data, and virtual driving instructions. The actual driving instructions are the real operation instructions issued by the human driver, such as throttle, brake, steering, etc.; the vehicle perception data is the environmental data collected by the vehicle sensors in real time, such as road conditions, traffic signals, the behavior of surrounding vehicles, etc.; the virtual driving instructions are the simulated driving instructions generated by the autonomous driving system based on the vehicle perception data; finally, through the communication interface of the vehicle, the collected data is transmitted to the remote server safely and efficiently to ensure the integrity and real-time nature of the data transmission.
[0035] After the server receives the data uploaded by the vehicle, it first analyzes the differences between the actual driving instructions and the virtual driving instructions. Based on these differences, the virtual driving instructions are divided into three types: conservative, moderate, and aggressive. This classification is based on a comprehensive assessment of driving style, risk preference, and operating habits. Specifically, the analysis of the differences between the actual driving instructions and the virtual driving instructions can follow the following classification rules, such as:
[0036] Conservative type: The virtual driving instructions are significantly different from the actual driving instructions and tend to be safer and slower in operation.
[0037] Moderate type: The virtual driving instructions are slightly different from the actual driving instructions, and the operating style is similar to that of a human driver.
[0038] Aggressive type: The virtual driving instructions are significantly different from the actual driving instructions and tend to be faster and more intense in operation.
[0039] In addition, human drivers with rich driving experience can also conduct the analysis of the differences and the classification between the actual driving instructions and the virtual driving instructions.
[0040] Such as Figure 3 shown, for a wide range of driving populations, the driving behavior styles should follow a normal distribution pattern, that is, the populations with "conservative" and "aggressive" driving styles are in the minority, and the driving styles of the vast majority of people are in the "suitable" range.
[0041] From the classified virtual driving instructions, select at least one type of instruction as the positive sample. The positive sample represents the driving style that the autonomous driving system needs to learn and imitate, and the positive sample includes at least the virtual driving instructions of the moderate type. At the same time, select at least one type of instruction as the negative sample. The negative sample represents the driving style that the autonomous driving system needs to avoid or correct. Generally speaking, to construct an autonomous driving algorithm suitable for the driving and riding experience of most people, the algorithm data labeled as "suitable" can be screened out as the positive sample, while the algorithm data of "conservative" and "aggressive" are both used as negative samples. If the expectation for the autonomous driving algorithm is a more "aggressive" style, the data with the "aggressive" label can be removed from the negative sample and added to the positive sample, which can make the driving behavior of the autonomous driving algorithm more "aggressive".
[0042] Using the uploaded vehicle perception data (model input), selected positive and negative samples (supervision signals), an optimized training set is constructed to perform optimized training on the autonomous driving algorithm under supervised learning. Through iterative training, the driving style of the autonomous driving algorithm approaches the positive samples and moves away from the negative samples, while ensuring driving safety and comfort. For the newly obtained autonomous driving algorithm model from the training, compared with the original algorithm, the driving behavior instructions output by it will likely be evaluated as "suitable", and the likelihood of the driving behaviors evaluated as "conservative" and "aggressive" will be reduced. The server feeds back the optimized autonomous driving algorithm to the vehicle side, and the vehicle side updates the algorithm, and continues to collect data, detect differences, and upload to the server during subsequent driving, forming a closed-loop iterative optimization process. Through this detailed step, the method of this embodiment can achieve continuous optimization of the autonomous driving algorithm, making the autonomous driving system more intelligent, safe, and in line with the driving habits of human drivers.
[0043] In some embodiments, the method described in Embodiment 1 includes:
[0044] When there is a difference between the actual driving instruction and the virtual driving instruction, suspend the operation of the autonomous driving system. After completing the collection of the actual driving instruction, vehicle perception data, and virtual driving instruction, resume the operation of the autonomous driving system.
[0045] Specifically, suspending the operation of the autonomous driving system is to avoid continuously collecting data due to the continuous inconsistency between the algorithm and human driving behavior during this period; after the data collection is completed, resume the operation of the background autonomous driving system, that is, return to the state of step S11 again, so that during the vehicle driving on the road, it can continuously collect human and algorithm driving behavior data.
[0046] In some embodiments, the actual driving instruction and the virtual driving instruction with differences in the same driving scenario described in Embodiment 1 include: in the same driving scenario, the actual driving instruction and the virtual driving instruction with inconsistent behaviors, inconsistent degrees, or inconsistent timings;
[0047] The inconsistent behaviors include longitudinal or lateral behavior inconsistencies; specifically, the longitudinal behavior refers to the operation instructions for the accelerator and brake, and the lateral behavior refers to the operation instructions for the steering wheel; for example, when a human driver steps on the accelerator to accelerate and turns the steering wheel to the left at the same time, when the virtual driving instruction output by the autonomous driving system is also to accelerate and turn left, it is considered that the two behaviors are consistent, otherwise they are inconsistent;
[0048] The inconsistency in degree includes that the control parameter error of the driving instruction is greater than a preset threshold. Specifically, when |virtual driving instruction - actual driving instruction| / |actual driving instruction| > preset threshold (for example, 50%), there is an inconsistency in degree. For example, if the human driver's steering wheel angle instruction is 20°, the autonomous driving system's steering wheel angle instruction within the range of 10° - 30° is considered consistent with the driver's operation, while if the autonomous driving system's steering wheel angle instruction is 50°, it is considered inconsistent with the driver's operation.
[0049] The inconsistency in timing includes that the execution time error of the driving instruction is greater than a preset threshold. Specifically, when the time difference between the execution time points of the virtual driving instruction and the actual driving instruction exceeds the preset threshold (for example, 3 seconds), there is an inconsistency in timing, that is, when |virtual driving instruction time point - actual driving instruction time point| > 3S.
[0050] In some embodiments, the vehicle perception data described in Embodiment 1 includes: the vehicle perception data in the time period between the preset time before and the preset time after the corresponding moment when there are behavioral inconsistencies, degree inconsistencies, or timing inconsistencies between the actual driving instruction and the virtual driving instruction.
[0051] Specifically, the preset time before and the preset time after can be 15 seconds, that is, the data collection time period can be the environmental perception data detected by the vehicle sensing system within the 30 - second time period between 15 seconds before and 15 seconds after the moment when it is determined that the driving instructions of the autonomous driving system and the human driver are inconsistent.
[0052] Refer to Figure 4 , Embodiment 2 of the present invention proposes an autonomous driving algorithm optimization method, which is applied to a cloud server. The method includes:
[0053] Step S21, obtaining the human - machine driving behavior difference data uploaded by the vehicle; wherein, the human - machine driving behavior difference data includes the actual driving instruction and the virtual driving instruction with differences in the same driving scenario and the vehicle perception data corresponding to the driving scenario. The actual driving instruction is the driving instruction of the human driver, and the virtual driving instruction is the virtual driving instruction output by the autonomous driving algorithm based on the vehicle perception data.
[0054] Specifically, when a human driver manually drives a vehicle, the autonomous driving system runs in the background, but the autonomous driving system does not actually control the vehicle. Instead, it only outputs virtual driving instructions for differential comparison with the actual driving instructions of the human driver. Vehicle perception data refers to the perception data obtained by the vehicle's sensing system for detecting the current scene environment. The sensing system includes, but is not limited to, cameras, lidar, ultrasonic sensors, etc. At each moment, there are both the actual driving instructions of the human driver and the virtual driving instructions of the autonomous driving system / algorithm. The vehicle continuously monitors and compares the actual driving instructions with the virtual driving instructions to determine whether there are significant differences between the two. For example Figure 2 As shown, when facing an obstacle vehicle, the human driver chooses to detour, while the autonomous driving system chooses to decelerate and avoid. Once a difference in instructions is detected, data collection is immediately triggered. The collected content includes actual driving instructions, vehicle perception data, and virtual driving instructions. The actual driving instructions are the real operation instructions issued by the human driver, such as throttle, brake, steering, etc.; the vehicle perception data is the environmental data collected by the vehicle sensors in real time, such as road conditions, traffic signals, the behaviors of surrounding vehicles, etc.; the virtual driving instructions are the simulated driving instructions generated by the autonomous driving system based on the vehicle perception data; finally, through the communication interface of the vehicle, the collected data is securely and efficiently transmitted to the remote server to ensure the integrity and real-time nature of data transmission.
[0055] Step S22: According to the difference between the actual driving instruction and the virtual driving instruction, divide the virtual driving instruction into three types: conservative, moderate, and aggressive.
[0056] Specifically, after the server receives the data uploaded by the vehicle, it first analyzes the difference between the actual driving instruction and the virtual driving instruction. According to the difference, the virtual driving instruction is divided into three types: conservative, moderate, and aggressive. This classification is based on a comprehensive evaluation of driving styles, risk preferences, and operating habits. Specifically, the analysis of the difference between the actual driving instruction and the virtual driving instruction can follow the following classification rules, such as:
[0057] Conservative type: The difference between the virtual driving instruction and the actual driving instruction is large, and it tends to be safer and slower in operation.
[0058] Moderate type: The difference between the virtual driving instruction and the actual driving instruction is small, and the operating style is close to that of the human driver.
[0059] Aggressive type: The difference between the virtual driving instruction and the actual driving instruction is large, and it tends to be faster and more intense in operation.
[0060] In addition, human drivers with rich driving experience can also conduct the analysis of the difference and type division between the actual driving instruction and the virtual driving instruction.
[0061] As Figure 3 shown, for a wide range of drivers, the driving behavior styles should follow a normal distribution pattern, that is, the number of people with "conservative" and "aggressive" driving styles is small, and the driving styles of the vast majority of people fall within the "suitable" range.
[0062] Step S23: Select at least one type of virtual driving instruction as the positive sample and at least one type of virtual driving instruction as the negative sample; the positive sample includes at least the moderate type of virtual driving instruction.
[0063] Specifically, from the classified virtual driving instructions, select at least one type of instruction as the positive sample, where the positive sample represents the driving style that the autonomous driving system needs to learn and imitate, and the positive sample includes at least the moderate type of virtual driving instruction; at the same time, select at least one type of instruction as the negative sample, where the negative sample represents the driving style that the autonomous driving system needs to avoid or correct.
[0064] Generally speaking, to build an autonomous driving algorithm suitable for the driving and riding experiences of most people, the algorithm data labeled as "suitable" can be selected as the positive sample, while the algorithm data of "conservative" and "aggressive" are both used as negative samples. If the expectation for the autonomous driving algorithm is a more "aggressive" style, the data with the "aggressive" label can be removed from the negative sample and added to the positive sample, which can make the driving behavior of the autonomous driving algorithm more "aggressive".
[0065] Step S24: Use the vehicle perception data as the input of the autonomous driving algorithm, and use the positive sample and negative sample as the supervision signals to guide the optimization training of the autonomous driving algorithm.
[0066] Specifically, using the uploaded vehicle perception data (model input), the selected positive sample and negative sample (supervision signals), an optimization training set is constructed to perform optimization training on the autonomous driving algorithm under supervised learning. Through iterative training, the driving style of the autonomous driving algorithm approaches the positive sample and moves away from the negative sample, while ensuring driving safety and comfort. The newly trained autonomous driving algorithm model, compared with the original algorithm, is likely to have its output driving behavior instructions evaluated as "suitable", while the likelihood of those driving behaviors evaluated as "conservative" and "aggressive" will decrease. The server feeds back the optimized autonomous driving algorithm to the vehicle side, and the vehicle side updates the algorithm, and continues to collect data, detect differences, and upload to the server during subsequent driving processes, forming a closed-loop iterative optimization process. Through this detailed step, the method of this embodiment can achieve continuous optimization of the autonomous driving algorithm, making the autonomous driving system more intelligent, safe, and in line with the driving habits of human drivers.
[0067] In some embodiments, the actual driving instructions and virtual driving instructions that are different in the same driving scenario described in Embodiment 2 include: actual driving instructions and virtual driving instructions with inconsistent behaviors, inconsistent degrees, or inconsistent timings in the same driving scenario;
[0068] The inconsistent behaviors include longitudinal behaviors or lateral behaviors; specifically, the longitudinal behaviors refer to operation instructions for the accelerator and brake, and the lateral behaviors refer to operation instructions for the steering wheel; for example, when a human driver steps on the accelerator to accelerate and at the same time turns the steering wheel to the left, when the virtual driving instructions output by the automatic driving system are also to accelerate and turn left, it is considered that the two behaviors are consistent, otherwise they are inconsistent;
[0069] The inconsistent degrees include that the control parameter error of the driving instructions is greater than a preset threshold; specifically, when |virtual driving instruction - actual driving instruction| / |actual driving instruction| > preset threshold (for example, 50%), the degree is inconsistent. For example, when the human driver's steering wheel steering angle instruction is 20°, the automatic driving system's steering wheel steering angle instruction within the range of 10° to 30° is considered consistent with the driver's operation, while if the automatic driving system's steering wheel steering angle instruction is 50°, it is considered inconsistent with the driver's operation.
[0070] The inconsistent timings include that the execution time error of the driving instructions is greater than a preset threshold. Specifically, if the time difference between the execution time points of the virtual driving instruction and the actual driving instruction exceeds the preset threshold (for example, 3 seconds), the timing is inconsistent, that is, when |virtual driving instruction time point - actual driving instruction time point| > 3S.
[0071] In some embodiments, the vehicle perception data corresponding to the driving scenario described in Embodiment 2 includes: vehicle perception data in the time period between the pre-preset time and the post-preset time corresponding to the moment when there are inconsistent behaviors, inconsistent degrees, or inconsistent timings between the actual driving instruction and the virtual driving instruction.
[0072] Specifically, the pre-preset time and the post-preset time can be 15 seconds, that is, the data collection time period can be the environmental perception data detected by the vehicle sensing system within the 30-second time period between 15 seconds before and 15 seconds after the moment when it is determined that the driving instructions of the automatic driving system and the human driver are inconsistent.
[0073] Embodiment 3 of the present invention proposes an automatic driving algorithm optimization device, including a module for executing the method described in Embodiment 1; the module can be implemented in a software manner, a hardware manner, or a combination of software and hardware manners. It should be noted that the device provided in Embodiment 3 can be used to execute the method described in Embodiment 1 above. Therefore, the content not detailed in this Embodiment 3 can be obtained by referring to the content of the method in Embodiment 1 above, and thus will not be elaborated here.
[0074] Embodiment 4 of the present invention provides an automatic driving algorithm optimization device, including a module for executing the method described in Embodiment 2; the module can be implemented in a software manner, a hardware manner, or a combination of software and hardware. It should be noted that the device provided in Embodiment 4 can be used to execute the method described in Embodiment 2 above. Therefore, the content not detailed in Embodiment 4 can be obtained by referring to the content of the method in Embodiment 2 above, so it will not be elaborated here.
[0075] Embodiment 5 of the present invention provides an automatic driving algorithm optimization device, including:
[0076] A communication interface for communicating with other electronic devices;
[0077] A memory for storing computer program instructions;
[0078] A processor for executing the computer program instructions to support the device to implement the method according to Embodiment 1.
[0079] In Embodiment 5, the memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store operating devices, application programs required for at least one function, etc., and the data storage area can store relevant data, etc. In addition, the memory can be a high-speed random access memory, or a non-volatile memory, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc., or the memory can also be other volatile solid-state storage devices.
[0080] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor can also be any conventional processor. The processor is the control center of the device, and uses various interfaces and lines to connect each part of the device described in Embodiment 5.
[0081] Embodiment 6 of the present invention provides an automatic driving algorithm optimization device, including:
[0082] A communication interface for communicating with other electronic devices;
[0083] A memory for storing computer program instructions;
[0084] A processor for executing the computer program instructions to support the apparatus to implement the method according to Embodiment 2.
[0085] In Embodiment 6, the memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store operating devices, application programs required for at least one function, etc., and the data storage area can store related data, etc. In addition, the memory can be a high-speed random access memory, or a non-volatile memory, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc., or the memory can also be other volatile solid-state storage devices.
[0086] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor can also be any conventional processor. The processor is the control center of the apparatus, and uses various interfaces and lines to connect each part of the apparatus according to Embodiment 6.
[0087] Embodiment 7 of the present invention proposes a computer program product, including computer program instructions, and the computer program instructions instruct a computer device to execute operations corresponding to the method according to Embodiment 1 of the present invention; specifically, the computer program product includes a series of computer program instructions, and these computer program instructions are codes written in the computer program, which define how to execute specific operations. These computer program instructions are designed to be loaded onto a computer device and guide the device to execute specific operations, and these operations refer to each step in the autonomous driving algorithm optimization method described in Embodiment 1 above. In this way, the computer program product of Embodiment 7 provides a complete software solution, which can run on various computer devices and implement the autonomous driving algorithm optimization method of Embodiment 1.
[0088] Embodiment 8 of the present invention provides a computer program product, including computer program instructions, and the computer program instructions direct a computer device to perform operations corresponding to the method described in Embodiment 2 of the present invention; specifically, the computer program product includes a series of computer program instructions, which are codes written in a computer program, define how to perform specific operations, are designed to be loaded onto a computer device, and guide the device to perform specific operations, which refer to the respective steps in the automatic driving algorithm optimization method described in Embodiment 2 above. In this way, the computer program product of Embodiment 8 provides a complete software solution, which can run on various computer devices to implement the automatic driving algorithm optimization method of Embodiment 2 above.
[0089] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the disclosed embodiments.
Claims
1. An optimization method for an autonomous driving algorithm, characterized in that The method includes: During the process of a human driver manually driving a vehicle, running an automatic driving system; Obtaining the actual driving instructions of the human driver and the virtual driving instructions output by the automatic driving system based on vehicle perception data decision-making in the same driving scenario; When there are differences between the actual driving instructions and the virtual driving instructions, collecting the actual driving instructions, vehicle perception data, and virtual driving instructions, and uploading them to the server, so that the server divides the virtual driving instructions into three types: conservative, moderate, and aggressive according to the differences between the actual driving instructions and the virtual driving instructions, selects at least one type of virtual driving instruction as a positive sample, at least one type of virtual driving instruction as a negative sample, and optimizes and trains the automatic driving algorithm according to the vehicle perception data, positive samples, and negative samples; the positive samples include at least moderate-type virtual driving instructions.
2. The method according to claim 1, characterized in that The method includes: When there are differences between the actual driving instructions and the virtual driving instructions, suspend the operation of the automatic driving system, and resume the operation of the automatic driving system after completing the collection of the actual driving instructions, vehicle perception data, and virtual driving instructions.
3. The method according to claim 1, wherein The actual driving instructions and virtual driving instructions with differences in the same driving scenario include: actual driving instructions and virtual driving instructions with inconsistent behaviors, inconsistent degrees, or inconsistent timings in the same driving scenario; The inconsistent behaviors include longitudinal behavior or lateral behavior inconsistency; The inconsistent degree includes that the control parameter error of the driving instruction is greater than a preset threshold; The inconsistent timing includes that the execution time error of the driving instruction is greater than a preset threshold.
4. The method according to claim 3, characterized in that, The vehicle perception data includes: vehicle perception data in the time period between the preset time before and the preset time after the corresponding moment when there are inconsistent behaviors, inconsistent degrees, or inconsistent timings between the actual driving instructions and the virtual driving instructions.
5. A method for optimizing an autonomous driving algorithm, characterized in that, The method includes: Obtaining the human-machine driving behavior difference data uploaded by the vehicle; wherein, the human-machine driving behavior difference data includes the actual driving instructions and virtual driving instructions with differences in the same driving scenario and the vehicle perception data corresponding to the driving scenario, the actual driving instructions are the driving instructions of the human driver, and the virtual driving instructions are the virtual driving instructions output by the automatic driving algorithm based on the vehicle perception data decision-making; Dividing the virtual driving instructions into three types: conservative, moderate, and aggressive according to the differences between the actual driving instructions and the virtual driving instructions; Selecting at least one type of virtual driving instruction as a positive sample and at least one type of virtual driving instruction as a negative sample; the positive samples include at least moderate-type virtual driving instructions; Taking the vehicle perception data as the input of the automatic driving algorithm, and taking the positive samples and negative samples as supervision signals to guide the optimization and training of the automatic driving algorithm.
6. The method according to claim 5, characterized in that The actual driving instructions and virtual driving instructions with differences in the same driving scenario include: actual driving instructions and virtual driving instructions with inconsistent behaviors, inconsistent degrees, or inconsistent timings in the same driving scenario; The inconsistent behaviors include longitudinal behavior or lateral behavior inconsistency; The inconsistency in degree includes that the control parameter error of the driving instruction is greater than a preset threshold value; The inconsistency in timing includes that the execution time error of the driving instruction is greater than a preset threshold value.
7. The method according to claim 6, characterized in that, The vehicle perception data corresponding to the driving scenario includes: the vehicle perception data in the time period between the preset time before and the preset time after the corresponding moment when there is a behavior inconsistency, a degree inconsistency or a timing inconsistency between the actual driving instruction and the virtual driving instruction.
8. An automatic driving algorithm optimization device, characterized in that, It includes a module for executing the method according to any one of claims 1 to 7.
9. An automatic driving algorithm optimization device, characterized in that, It includes: A communication interface for communicating with other electronic devices; A memory for storing computer program instructions; A processor for executing the computer program instructions to support the device to implement the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that, It includes computer program instructions, and the computer program instructions instruct a computer device to execute the operations corresponding to the method according to any one of claims 1 to 7.