Optimization method and system for improving target-level fusion effect of sensor, medium and equipment

By determining the target-level fusion key scenario library and adaptively setting the Kalman gain coefficient using reinforcement learning network, the problem of unsatisfactory target-level fusion effect in the existing technology is solved, and an efficient and adaptive target-level fusion effect is achieved, reducing development costs.

CN120086791APending Publication Date: 2025-06-03TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510135608.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing target-level fusion technology is not ideal in multi-sensor fusion, resulting in target recognition bias, affecting the accuracy of the decision algorithm and the smoothness of the control algorithm, and the cost of manual parameter adjustment is high.

Method used

By determining the target-level fusion key scenario library, collecting core perceptual target perception results, forming a data set, and using reinforcement learning network to learn Kalman gain coefficients of different sensors, adaptively setting the best fusion parameters.

Benefits of technology

The effect of target-level fusion of sensors is improved, the development cost is reduced, large-scale manual parameter adjustment in different scenarios is avoided, and the better target-level fusion effect is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086791A_ABST
    Figure CN120086791A_ABST
Patent Text Reader

Abstract

The invention relates to the field of driving assistance systems and unmanned driving systems, and discloses an optimization method, system, medium and equipment for improving the target-level fusion effect of a sensor, and the method comprises the steps: determining a target-level fusion key scene library based on feedback information and laws and regulations through which a vehicle model needs to pass; acquiring a sensing result of a core sensing target in the key scene library, and performing data processing on the sensing result to form a data set; kalman gain coefficients of different sensors are learned according to the data set, a deployment method of reinforcement learning training results is determined according to vehicle-mounted computing power, and a final Kalman gain coefficient is obtained. According to the invention, the target level fusion effect in different scenes can be optimized, and the development cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of advanced driver assistance systems and driverless systems, and particularly to an efficient optimization method, system, medium and device for improving the target-level fusion effect of sensors. Background Art

[0002] Target-level fusion is one of the important perception solutions for advanced driver assistance systems and driverless systems, and it is widely used on low-computing-power platforms. Existing mass-produced target-level fusion generally updates the track based on Kalman filtering or its improved strategy, and determines the proportion of different sensors in the track update through the Kalman gain coefficient. When the sensing effect of the sensor is good, the Kalman gain coefficient of the sensor is appropriately increased, and when the sensing effect of the sensor is poor, the Kalman gain coefficient of the sensor is appropriately decreased.

[0003] Due to the unsatisfactory fusion result of multiple sensors, it will lead to target recognition deviation, affect the accuracy of the decision-making algorithm and the smoothness of the control algorithm, and even endanger the safety of the vehicle. Moreover, due to the complex and changeable scenarios and various target types involved in sensing, large-scale manual parameter tuning for different scenarios has high costs in terms of time and resources. Summary of the Invention

[0004] Aiming at the above problems, the purpose of the present invention is to provide an optimization method, system, medium and device for improving the target-level fusion effect of sensors, which can optimize the target-level fusion effect in different scenarios and reduce the development cost.

[0005] To achieve the above purpose, in the first aspect, the technical solution adopted by the present invention is: an optimization method for improving the target-level fusion effect of sensors, which includes: determining a target-level fusion key scenario library based on feedback information and regulations that the vehicle model needs to pass; collecting the perception results of core perception targets in the key scenario library, and performing data processing on the perception results to form a data set; learning the Kalman gain coefficients of different sensors for the data set, and determining the deployment method of the reinforcement learning training results according to the vehicle-mounted computing power to obtain the final Kalman gain coefficient.

[0006] Further, determining a target-level fusion key scenario library based on feedback information and regulations that the vehicle model needs to pass includes: Based on feedback information, determining the key scenarios based on experience as experience scenarios; According to the regulations that the vehicle model needs to pass, selecting the key scenarios under this regulation as regulation scenarios; Taking the maximum of the experience scenarios and the regulation scenarios to obtain the target-level fusion key scenarios; Selecting a core perception target for each target-level fusion key scenario; A sample in the key scenario library of object-level fusion consists of a key scenario at the object level and a core perception object in this key scenario.

[0007] Furthermore, the key scenarios in the experience scenarios include sudden braking of the vehicle in front, pedestrians crossing and quickly stopping, and motorcycles quickly changing lanes. The key scenarios in the regulation scenarios include oncoming vehicles and following vehicles for the emergency lane keeping function, the target vehicle going straight when the ego vehicle is turning for the automatic emergency braking function, children "popping out", bicycles crossing while blocked, pedestrians crossing when the ego vehicle is turning, and motorcycles going straight when the ego vehicle is turning.

[0008] Furthermore, collect the perception results of the core perception object in the key scenario library, and perform data processing on the perception results to form a data set, including: According to the set sensor scheme of the vehicle model project, collect the perception results of different sensors in the key scenario library for the core perception object. Collect the ground truth data of the perception results of the core perception object in the key scenario library. Stitch the perception data of the same sensor for the core perception object in different scenarios. At the same time, stitch the ground truth data of the core perception object in the same time dimension. The stitched perception data and the stitched ground truth data together form a data set.

[0009] Furthermore, collect the ground truth data of the perception results of the core perception object in the key scenario library, including: using an inertial navigation differential positioning system or lidar to collect the ground truth data of the perception results of the core perception object in the key scenario library.

[0010] Furthermore, learn the Kalman gain coefficients of different sensors for the data set, including: Select state variables according to the types of targets and tracks. For the selected state variables, multiply the square of the difference between the track and the ground truth by a weight coefficient as the reward function. The action is selected as the Kalman gain coefficient corresponding to the track update, and each sensor has one action; the action space is selected in two forms, discrete or continuous, and the range is between [0,1]. Different forms of action spaces correspond to different selections of reinforcement learning methods. Two strategies are adopted for the training of reinforcement learning. The first strategy assumes that the Kalman gain coefficient remains unchanged within one round, and the second strategy assumes that the Kalman gain coefficient is dynamically changing within one round.

[0011] Furthermore, determine the deployment method of the reinforcement learning training results according to the in-vehicle computing power to obtain the final Kalman gain coefficient, including: When the training result fitted by the integrated neural network is applied online, if the in-vehicle computing power load exceeds the set value or the maximum threshold required by the customer, the first strategy is adopted for the training result, and each sensor adopts a fixed Kalman gain coefficient; otherwise, the second strategy is adopted.

[0012] In a second aspect, the technical solution adopted by the present invention is: an optimization system for improving the target-level fusion effect of sensors, which includes: a key scenario library determination module for determining a target-level fusion key scenario library based on feedback information and regulations that the vehicle model needs to pass; a data acquisition and processing module for acquiring the perception results of core perception targets in the key scenario library and performing data processing on the perception results to form a data set; a Kalman gain coefficient determination module for learning the Kalman gain coefficients of different sensors for the data set, determining the deployment method of the reinforcement learning training result according to the in-vehicle computing power, and obtaining the final Kalman gain coefficient.

[0013] In a third aspect, the technical solution adopted by the present invention is: a computer-readable storage medium storing one or more programs, where the one or more programs include instructions that, when executed by a computing device, cause the computing device to execute any one of the above methods.

[0014] In a fourth aspect, the technical solution adopted by the present invention is: a computing device, which includes: one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the above methods.

[0015] Due to the above technical solutions adopted by the present invention, it has the following advantages: The present invention can adaptively set the optimal fusion parameters by the reinforcement learning network according to the current scenario, so that a better target-level fusion effect can be obtained. By network training, large-scale manual parameter adjustment in different scenarios is avoided, the development cost is reduced, and it is easy to achieve mass production and promotion. Description of the Drawings

[0016] Figure 1 is a flowchart of an optimization method for improving the target-level fusion effect of sensors in an embodiment of the present invention; Figure 2 is a schematic diagram of the input and output of a neural network for fitting the reinforcement learning training result in an embodiment of the present invention. Detailed Embodiments

[0017] In view of the problems that there are various types of existing sensors, the sensing characteristics of the same type of sensors also vary, the adjustment of the Kalman gain coefficient is highly subjective, which limits the effect of target-level fusion. Also, since the sensing effect of the sensor is not fixed but related to the actual scenario, and the strength relationship of the sensing capabilities between different sensors may change with the distance and motion state of the sensing target, setting a fixed Kalman gain coefficient will result in the fusion effect not reaching the optimum. The weakened fusion result may lead to target recognition deviation, thus affecting the accuracy of the decision-making algorithm and the smoothness of the control algorithm, and even endangering the safety of the vehicle in severe cases. At the same time, due to the complex and changeable scenarios and various target types involved in sensing, large-scale manual parameter tuning for different scenarios incurs high costs in terms of time and resources. Therefore, the present invention proposes an optimization method, system, medium and device for improving the sensor target-level fusion effect. The efficient and adaptive optimization method of the present invention can effectively enhance the sensor target-level fusion effect.

[0018] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention fall within the scope of protection of the present invention.

[0019] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0020] In an embodiment of the present invention, as Figure 1 shown, there is provided an efficient optimization method for improving the sensor target-level fusion effect. In this embodiment, the method includes the following steps: 1) Determine the key scenario library for target-level fusion based on the feedback information and the regulations that the vehicle model needs to comply with; 2) Collect the sensing results of the core sensing targets in the key scenario library, and perform data processing on the sensing results to form a data set; 3) Learn the Kalman gain coefficients of different sensors for the data set, determine the deployment method of the reinforcement learning training results according to the in-vehicle computing power, and obtain the final Kalman gain coefficient.

[0021] In step 1) above, based on the feedback information and the regulations that the vehicle model needs to pass, determine the target-level fusion key scenario library, including the following steps: 1.1) Based on the feedback information, determine the key scenarios based on experience as experience scenarios.

[0022] In this embodiment, the feedback information includes feedback from developers, testers, and users.

[0023] In this embodiment, the key scenarios in the experience scenarios include sudden braking of the vehicle ahead, rapid stop of a pedestrian crossing, and rapid lane change of a motorcycle, etc.

[0024] 1.2) According to the regulations that the vehicle model needs to pass, select the key scenarios under this regulation as regulation scenarios.

[0025] In this embodiment, taking CN CAP2024 as an example, the key scenarios in the regulation scenarios include oncoming vehicles and following vehicles for the emergency lane keeping function, the target vehicle going straight when the host vehicle is turning for the automatic emergency braking function, a child "popping out suddenly", a bicycle crossing with occlusion, a pedestrian crossing when the host vehicle is turning, and a motorcycle going straight when the host vehicle is turning, etc.

[0026] 1.3) Take the maximum of the experience scenarios and the regulation scenarios to obtain the target-level fusion key scenarios.

[0027] 1.4) Select a core perception target for each target-level fusion key scenario.

[0028] For example, the target vehicle ahead in the sudden braking of the vehicle ahead scenario, the motorcycle target in the rapid lane change of a motorcycle scenario, the child target in the "child popping out suddenly" scenario, and the pedestrian target in the pedestrian crossing when the host vehicle is turning scenario.

[0029] 1.5) A sample in the target-level fusion key scenario library is composed of a target-level fusion key scenario and a core perception target under this key scenario.

[0030] In step 2) above, collect the perception results of the core perception targets in the key scenario library and perform data processing on the perception results to form a data set, including the following steps: 2.1) According to the set sensor scheme of the vehicle model project, collect the perception results of different sensors on the core perception targets in the key scenario library.

[0031] 2.2) Collect the ground truth data of the perception results of the core perception targets in the key scenario library; In this embodiment, to collect the ground truth data of the perception results of the core perception targets in the key scenario library, various means can be used to measure the ground truth data. For example, it can include using an inertial navigation differential positioning system or lidar, etc., to collect the ground truth data of the perception results of the core perception targets in the key scenario library.

[0032] 2.3) Concatenate the perception data of the same sensor for the core perception target in different scenarios. At the same time, concatenate the ground truth data of the core perception target in the same time dimension. The concatenated perception data and the concatenated ground truth data together form a dataset.

[0033] In step 3) above, learning the Kalman gain coefficients of different sensors for the dataset includes the following steps: 3.1) Select state variables according to the types of targets and trajectories.

[0034] In this embodiment, for example, select state variables according to the types of targets and trajectories. For vehicle targets and trajectories, the selected state variables are six state variables: lateral distance, longitudinal distance, lateral speed, longitudinal speed, lateral acceleration, and longitudinal acceleration. Five state variables excluding lateral acceleration can also be selected. For pedestrian targets and trajectories, the selected state variables are generally lateral distance, longitudinal distance, lateral speed, and longitudinal speed. Here, the target refers to the target recognized by the sensor, such as vehicles, pedestrians, cyclists, traffic cones, etc.

[0035] 3.2) For the selected state variables, multiply the square of the difference between the trajectory and the ground truth by a weight coefficient as the reward function.

[0036] In this embodiment, for the selected five or six state variables, multiply the square of the difference between the trajectory and the ground truth by a weight coefficient as the reward function. If you want a certain state variable to be more important, you can appropriately increase the corresponding weight coefficient, and vice versa.

[0037] 3.3) The action is to update the Kalman gain coefficient corresponding to the trajectory, and each sensor has one action; the action space is selected in two forms: discrete or continuous, and the range is between [0,1]; different forms of action spaces correspond to different choices of reinforcement learning methods.

[0038] 3.4) Two strategies are adopted for the training of reinforcement learning. The first strategy assumes that the Kalman gain coefficient remains unchanged within one round, and the second strategy assumes that the Kalman gain coefficient changes dynamically within one round.

[0039] In step 3) above, determine the deployment method of the reinforcement learning training result according to the in-vehicle computing power to obtain the final Kalman gain coefficient. Specifically: When the training result fitted by the integrated neural network is applied online, if the in-vehicle computing power load exceeds the set value (for example, the set value is 80%) or the maximum threshold required by the customer, the training result adopts the first strategy, and each sensor adopts a fixed Kalman gain coefficient; otherwise, the second strategy is adopted.

[0040] For the first strategy for reinforcement learning training, since the deployment of its training results is relatively simple and only the optimal training results are required, a fixed Kalman gain coefficient for each sensor can be selected. For the second strategy for reinforcement learning training, it is necessary to use a neural network to fit the training results to achieve online application, and adaptively set the Kalman gain coefficient according to the current state of the host vehicle and the target.

[0041] In this embodiment, as Figure 2 shown, the neural network adopts a feasible input-output structure, and the input can be a combination of all or part of the parameters of the host vehicle speed, host vehicle acceleration, host vehicle steering wheel angle, target lateral distance, target longitudinal distance, target lateral speed, target longitudinal speed, target lateral acceleration, and target longitudinal acceleration.

[0042] After the above steps are completed, a step of in-vehicle verification is also included. After the in-vehicle deployment of the training results, an in-vehicle verification of a set mileage is carried out to ensure the stable and reliable integration of the production target level.

[0043] In the above embodiments, after the in-vehicle verification step, a step of software release is also included. After the in-vehicle verification is completed, the target-level fusion software adapted to the sensor scheme of this vehicle model is released for subsequent path planning and control algorithms to use.

[0044] In an embodiment of the present invention, an optimization system for improving the sensor target-level fusion effect is provided, which includes: A key scenario library determination module, which determines the target-level fusion key scenario library based on the feedback information and the regulations that the vehicle model needs to pass; A data collection and processing module, which collects the perception results of the core perception targets in the key scenario library and processes the perception results to form a data set; A Kalman gain coefficient determination module, which learns the Kalman gain coefficients of different sensors for the data set, determines the deployment method of the reinforcement learning training results according to the on-vehicle computing power, and obtains the final Kalman gain coefficient.

[0045] In the above embodiments, determining the target-level fusion key scenario library based on the feedback information and the regulations that the vehicle model needs to pass includes: Based on the feedback information, determining the key scenarios based on experience as the experience scenarios; According to the regulations that the vehicle model needs to pass, selecting the key scenarios under this regulation as the regulation scenarios; Taking the maximum of the experience scenarios and the regulation scenarios to obtain the target-level fusion key scenarios; Selecting a core perception target for each target-level fusion key scenario; A sample in the target-level fusion key scenario library is composed of a target-level fusion key scenario and a core perception target under this key scenario.

[0046] In this embodiment, the key scenarios in the experience scenario include sudden braking of the vehicle ahead, pedestrians crossing the road and quickly stopping, and motorcycles quickly changing lanes.

[0047] In this embodiment, the key scenarios in the regulation scenario include oncoming vehicles and rear vehicles for the emergency lane keeping function, the target vehicle going straight when the host vehicle turns for the automatic emergency braking function, children "popping out", bicycles crossing the road blocked, pedestrians crossing the road when the host vehicle turns, and motorcycles going straight when the host vehicle turns.

[0048] In the above embodiment, the perception results of the core perception targets in the key scenario library are collected, and the perception results are processed to form a data set, including: According to the set sensor scheme of the vehicle model project, the perception results of different sensors for the core perception targets in the key scenario library are collected; The true value data of the perception results of the core perception targets in the key scenario library are collected; The perception data of the same sensor for the core perception targets in different scenarios are spliced. At the same time, the true value data of the core perception targets in the same time dimension are spliced. The spliced perception data and the spliced true value data together form a data set.

[0049] In this embodiment, the true value data of the perception results of the core perception targets in the key scenario library are collected, including: using an inertial navigation differential positioning system or lidar to collect the true value data of the perception results of the core perception targets in the key scenario library.

[0050] In the above embodiment, learning the Kalman gain coefficients of different sensors for the data set includes: Select state variables according to the types of targets and tracks; For the selected state variables, the square of the difference between the track and the true value is multiplied by a weight coefficient as the reward function; The action is selected as the Kalman gain coefficient corresponding to the track update, and each sensor has one action; the action space is selected in two forms, discrete or continuous, and the ranges are both between [0,1]. Different forms of action spaces correspond to different selection methods of reinforcement learning; Two strategies are adopted for the training of reinforcement learning. The first strategy is that the Kalman gain coefficient remains unchanged within one round, and the second strategy is that the Kalman gain coefficient is dynamically changed within one round.

[0051] In the above embodiment, determining the deployment method of the reinforcement learning training results according to the in-vehicle computing power to obtain the final Kalman gain coefficient includes: When the training result fitted by the integrated neural network is applied online, if the in-vehicle computing power load exceeds the set value or the maximum threshold required by the customer, the first strategy is adopted for the training result, and each sensor adopts a fixed Kalman gain coefficient; otherwise, the second strategy is adopted.

[0052] The system provided in this embodiment is used to execute the above method embodiments. For the specific process and detailed content, please refer to the above embodiments and will not be elaborated here.

[0053] In an embodiment of the invention, a computing device is provided. The computing device may be a terminal, and it may include: a processor, a communications interface, a memory, a display screen, and an input device. Among them, the processor, the communications interface, and the memory complete their mutual communication through a communication bus. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. When the computer program is executed by the processor, it is used to implement the methods in the above embodiments; the internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communications interface is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen may be a liquid crystal display screen or an electronic ink display screen. The input device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computing device, or an external keyboard, touchpad, or mouse, etc. The processor can call the logical instructions in the memory.

[0054] In addition, when the logical instructions in the above memory are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. And the foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes.

[0055] In an embodiment of the present invention, a computer program product is provided. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above method embodiments.

[0056] In an embodiment of the present invention, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores server instructions, and the computer instructions cause the computer to execute the methods provided in the above embodiments.

[0057] For a computer-readable storage medium provided in the above embodiments, its implementation principle and technical effects are similar to those of the above method embodiments, and will not be elaborated here.

[0058] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a plurality of blocks.

[0059] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the specified functions in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a plurality of blocks.

[0060] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a plurality of blocks.

[0061] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An optimization method for improving sensor target level fusion effect, characterized in that: include: Determine the target-level fusion key scenario library based on feedback information and the regulations that the vehicle model needs to pass; Collect the perception results of the core perception targets in the key scene library, and process the perception results to form a data set; The Kalman gain coefficients of different sensors are learned for the data set, and the deployment method of the reinforcement learning training results is determined according to the on-board computing power to obtain the final Kalman gain coefficients.

2. The optimization method for improving sensor target level fusion effect as claimed in claim 1, characterized in that: Based on the feedback information and the regulations that the vehicle model needs to pass, determine the target-level fusion key scenario library, including: Based on the feedback information, key scenarios based on experience are determined as experience scenarios; According to the regulations that the vehicle model needs to pass, select the key scenarios under the regulations as the regulatory scenarios; Maximize the experience scenario and the regulatory scenario to obtain the target-level fusion key scenario; A core perception target is selected for each target-level fusion key scenario; A sample in the target-level fusion key scene library consists of a target-level fusion key scene and a core perception target under the key scene.

3. The optimization method for improving sensor target level fusion effect as claimed in claim 2, characterized in that: Key scenarios in the experience scenarios include sudden braking of the vehicle ahead, quick stopping of pedestrians crossing the road, and quick lane switching of motorcycles; The key scenarios in the regulatory scenarios include oncoming and rearward vehicles with the emergency lane keeping function, the target vehicle going straight when the vehicle is turning with the automatic emergency braking function, a child "peeking out", an obstructed bicycle crossing the road, pedestrians crossing the road when the vehicle is turning, and a motorcycle going straight when the vehicle is turning.

4. The optimization method for improving sensor target level fusion effect as claimed in claim 1, characterized in that: Collect the perception results of the core perception targets in the key scene library, and process the perception results to form a data set, including: According to the sensor scheme set for the vehicle model project, collect the perception results of different sensors on the core perception targets in the key scene library; Collect the true value data of the perception results of the core perception targets in the key scene library; The perception data of the core perception target of the same sensor in different scenarios are spliced, and at the same time, the true value data of the core perception target in the same time dimension are spliced, and the spliced ​​perception data and the spliced ​​true value data are combined into a data set.

5. The optimization method for improving sensor target level fusion effect as claimed in claim 4, characterized in that: Collecting the true value data of the perception results of the core perception targets in the key scene library, including: using an inertial navigation differential positioning system or a lidar to collect the true value data of the perception results of the core perception targets in the key scene library.

6. The optimization method for improving sensor target level fusion effect as claimed in claim 1, characterized in that: Learn the Kalman gain coefficients of different sensors for the dataset, including: Select state variables according to the type of target and track; For the selected state variable, the square of the difference between the track and the true value multiplied by the weight coefficient is used as the reward function; The action selection is the Kalman gain coefficient corresponding to the track update, and each sensor has one action; the action space can be discrete or continuous, and the range is between [0,1]. Different forms of action space correspond to different reinforcement learning method selections; Reinforcement learning training adopts two strategies. The first strategy assumes that the Kalman gain coefficient within a round remains unchanged, and the second strategy assumes that the Kalman gain coefficient within a round changes dynamically.

7. The optimization method for improving sensor target level fusion effect as claimed in claim 6, characterized in that: Determine the deployment method of the reinforcement learning training results based on the vehicle computing power to obtain the final Kalman gain coefficient, including: When the training results of the integrated neural network fitting are applied online, if the on-board computing load exceeds the set value or the maximum threshold required by the customer, the training results adopt the first strategy, and each sensor uses a fixed Kalman gain coefficient; otherwise, the second strategy is adopted.

8. An optimization system for improving sensor target level fusion effect, characterized in that: include: The key scenario library determination module determines the target-level fusion key scenario library based on feedback information and the regulations that the vehicle model needs to pass; The data collection and processing module collects the perception results of the core perception targets in the key scene library and processes the perception results to form a data set; The Kalman gain coefficient determination module learns the Kalman gain coefficients of different sensors for the data set, determines the deployment method of the reinforcement learning training results based on the on-board computing power, and obtains the final Kalman gain coefficient.

9. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions, which, when executed by a computing device, cause the computing device to perform any one of the methods of claims 1 to 7.

10. A computing device, characterized in that include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the methods described in claims 1 to 7.