Driving track generation method and device, equipment and storage medium
By fusing data from multiple sensors to generate a risk semantic map, the problem of insufficient ability of existing technologies to handle complex scenarios and low safety of the generated driving trajectories is solved, achieving a more human-like and comfortable driving experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-19
- Publication Date
- 2026-04-03
AI Technical Summary
Existing driving trajectory generation methods are insufficient in handling complex, dynamic, and uncertain traffic environments, resulting in low safety, especially in situations where sensors are obstructed or in adverse weather conditions, making it difficult to make reasonable avoidance maneuvers.
By fusing data from multiple sensors, a risk semantic map is generated, which includes environmental information and risk levels. A decision-making and planning network combining reinforcement learning and imitation learning is used to generate a driving trajectory that comprehensively considers safety, comfort, and efficiency.
It enhances the driving ability and safety in complex environments. Through the fusion of data from multiple sensors and autonomous solutions, it improves the ability to cope with complex scenarios and safety, and provides a more human-like and comfortable driving experience.
Smart Images

Figure CN121783176A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicles, and more particularly to a method, apparatus, device, and storage medium for generating driving trajectories. Background Technology
[0002] With the rapid development of intelligent driving technology, trajectory planning in intelligent driving has become a key link in ensuring safe and comfortable vehicle driving. Existing trajectory planning methods usually make path decisions based on current environmental information, such as vehicles, pedestrians, obstacles, and lane lines, and plan future trajectories by predicting the movement trends of surrounding traffic participants.
[0003] However, traditional simple probability models usually separate environmental information from risk assessment, making it impossible to conduct comprehensive risk assessment in dynamic environments. When faced with situations where sensor obstruction or severe weather leads to blurred environmental perception, they cannot make reasonable avoidance of high-risk areas during trajectory generation, thus making it difficult to adapt to complex, dynamic, and uncertain real traffic environments.
[0004] Therefore, existing technologies suffer from insufficient ability to handle complex scenarios and low safety due to the generated driving trajectories.
[0005] Application content The purpose of this application is to provide a driving trajectory generation method, apparatus, device, and storage medium to improve the ability of the generated driving trajectory to cope with complex scenarios and enhance its safety.
[0006] Firstly, this application provides a method for generating driving trajectories, including: Obtain an environmental information map containing current and historical environmental information; Based on current and historical environmental information, a predicted environmental information map for future time periods is generated, along with a set of confidence scores corresponding to each region of the predicted environmental information map. The scores in the set of confidence scores represent the degree of confidence in the environmental characteristics of the region. By fusing the environmental information map, the predicted environmental information map, and the confidence score set, a risk semantic map is obtained, which contains environmental information and the corresponding risk level for each region. Based on the risk semantic map, a driving trajectory is generated.
[0007] Furthermore, an environmental information map containing current and historical environmental information is obtained, including: The current environmental features collected by multiple sensors are fused to obtain the current environmental information, where the current environmental features include the environmental features collected by the sensors at the current moment. By fusing historical environmental features collected by sensors, historical environmental information is obtained, which includes historical environmental features collected by sensors at past moments. By integrating current and historical environmental information, an environmental information map is obtained.
[0008] Furthermore, the current environmental features collected by multiple sensors are fused to obtain current environmental information, including: Based on the attention mechanism in the cross-modal Transformer, the current environmental features collected by different sensors in their respective coordinate systems are mapped to a unified coordinate system. The sensors include camera equipment, radar and laser point clouds, and millimeter-wave radar point traces. By fusing historical environmental features collected by sensors, historical environmental information is obtained, including: Based on the attention mechanism in cross-modal Transformer, historical environmental features collected by different sensors in their respective coordinate systems are mapped to a unified coordinate system. By integrating current and historical environmental information, an environmental information map is obtained, including: By fusing current and historical environmental information using recurrent neural networks and / or Transformer networks, an environmental information map is obtained, which consists of a rasterized environmental feature map and a dynamic target structured list.
[0009] Furthermore, based on current and historical environmental information, a predicted environmental information map for future time periods is generated, along with a set of confidence scores corresponding to each region of the predicted environmental information map, including: Based on a self-supervised learning head, a predicted environmental information map for future time periods is generated according to current and historical environmental information. The predicted environmental information map includes the predicted locations of static obstacles, passable areas, and dynamic objects. The confidence level of each region in the predicted environment information map is quantified to obtain a set of confidence scores corresponding to each region in the predicted environment information map, wherein each region includes at least one pixel.
[0010] Furthermore, based on the risk semantic map, a driving trajectory is generated, including: The risk semantic map is input into a decision-making and planning network that combines reinforcement learning and imitation learning. The reward function in the decision-making and planning network is to punish based on the score, and the score is directly proportional to the degree of punishment. The decision-making and planning network outputs a driving trajectory that meets preset driving conditions, including safety, comfort, compliance with traffic rules, and efficiency.
[0011] Furthermore, the reward function R satisfies: R = Rsafe + α * Rcomfort + β * Rprogress - γ * Runcertainty; Where Rsafe is the safety score, α is the comfort weight coefficient, Rcomfort is the comfort score, β is the feasibility weight coefficient, Rprogress is the feasibility score, γ is the confidence weight coefficient, and Runcertainty is the confidence score.
[0012] Furthermore, before generating the driving trajectory based on the risk semantic map, the following steps are included: Based on multiple historical user driving data, the decision planning network is pre-trained to learn a driving style that meets the target driving conditions, where the target driving conditions are at least one of the preset driving conditions input by the user. After generating the driving trajectory based on the risk semantic map, it includes: Based on driving style and driving trajectory, driving instructions that meet the target driving conditions are generated. These driving instructions include accelerator, brake, and steering. Secondly, this application also provides a driving trajectory generation device, comprising: The acquisition module is used to acquire an environmental information map containing current and historical environmental information. The first generation module is used to generate a predicted environmental information map for a future time period based on current environmental information and historical environmental information, as well as a set of confidence scores corresponding to each region of the predicted environmental information map, wherein the scores in the set of confidence scores represent the degree of confidence in the environmental characteristics of the region. The fusion module is used to fuse the environmental information map, the predicted environmental information map, and the confidence score set to obtain a risk semantic map, which contains environmental information and the corresponding risk level for each region. The second generation module is used to generate driving trajectories based on the risk semantic map.
[0013] Thirdly, this application also provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is used to execute the driving trajectory generation method described above by running instructions in memory.
[0014] Fourthly, this application also provides a computer storage medium storing instructions that, when executed, implement the above-described driving trajectory generation method.
[0015] This application embodiment obtains an environmental information map containing current and historical environmental information; based on the current and historical environmental information, it generates a predicted environmental information map for future time periods, and a confidence score set corresponding to each region of the predicted environmental information map, where the scores in the confidence score set represent the degree of confidence in the environmental characteristics of a region; it fuses the environmental information map, the predicted environmental information map, and the confidence score set to obtain a risk semantic map, where the risk semantic map contains environmental information and the corresponding risk level for each region; and it generates a driving trajectory based on the risk semantic map. This solves the technical problem of insufficient ability to handle complex scenarios and low safety in existing technologies for generated driving trajectories, achieving the technical effect of improving the ability and safety of generated driving trajectories in handling complex scenarios.
[0016] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a driving trajectory generation method provided in this application embodiment; Figure 2 This application provides a schematic diagram of the hardware and logic architecture of a driving trajectory generation system. Figure 3 A schematic diagram illustrating another driving trajectory generation method provided in an embodiment of this application; Figure 4 A structural diagram of a driving trajectory generation device provided in an embodiment of this application; Figure 5 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0018] To facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. For example, the first threshold and the second threshold are only used to distinguish different thresholds and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0019] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0020] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, a combination of a and b, a combination of a and c, a combination of b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0021] With the development of artificial intelligence technology, Level 2 / Level 3 autonomous driving systems have begun large-scale commercial use. However, existing autonomous driving systems still face many challenges in the decision-making and planning module, as follows: Perception relies heavily on high-precision maps: Most systems heavily depend on pre-recorded high-precision prior maps for positioning and decision-making. In areas where high-precision maps are not covered, are outdated, or elements have changed (such as road construction or traffic sign updates), system performance will significantly degrade or even crash.
[0022] The multimodal fusion layer is shallow: existing fusion schemes are mostly pre-fusion or post-fusion, which makes it difficult to achieve deep and complementary spatiotemporal alignment at the feature level. This leads to a decrease in perception reliability when a single sensor is limited (such as camera glare or the performance degradation of lidar in rain and fog).
[0023] Poor generalization ability of decision-making modules: Decision-making systems based on rules or pure imitation learning lack a deep understanding of the physical world and have difficulty handling scenarios that have not appeared in the training set, such as handling foreign objects falling from the vehicle in front or the passage of emergency rescue vehicles.
[0024] Lack of quantification and utilization of uncertainty: Existing systems typically output a definite perception or decision result, but cannot assess the reliability of that result. When faced with situations that cause perception ambiguity, such as obstruction or severe weather, the system cannot perceive the level of "uncertainty," thus failing to make more conservative and safer strategy choices.
[0025] Therefore, the applicant is committed to an autonomous driving solution that can achieve deep multimodal fusion without relying on high-precision maps, and can autonomously assess and utilize uncertainties to make robust decisions.
[0026] To address the technical problems of insufficient ability to handle complex scenarios and low safety in existing technologies for generating driving trajectories, embodiments of this application provide a driving trajectory generation method, such as... Figure 1 As shown: Figure 1 A flowchart of a driving trajectory generation method provided in this application embodiment includes: S101: Obtain an environmental information map containing current and historical environmental information; Specifically, deep neural networks are used to extract raw features from camera images, LiDAR point clouds, and millimeter-wave radar traces at the current moment and several past moments. A spatiotemporal fusion module can be designed, which includes: a spatial alignment unit, which maps features from different sensors in their respective coordinate systems to a unified vehicle coordinate system through calibration parameters and an attention mechanism; and a temporal alignment unit, which uses a recurrent neural network (RNN) or Transformer network to fuse temporal features and capture the motion trends of dynamic targets. The output is a rasterized feature map of the environment rich in spatiotemporal context information and a structured list of dynamic targets.
[0027] S102; Based on current environmental information and historical environmental information, generate a predicted environmental information map for a future time period, and a confidence score set corresponding to each region of the predicted environmental information map, wherein the scores in the confidence score set represent the degree of confidence in the environmental characteristics of the region. Specifically, a self-supervised learning head is introduced, working in parallel or cascaded with the perception network in S101. The task of this learning head is: future scene prediction, using current and past multimodal fusion features as input to predict the environmental BEV feature map within a short future timeframe (e.g., 3-5 seconds), including the possible locations of static obstacles, passable areas, and dynamic targets. Uncertainty quantification involves generating an uncertainty confidence score for each pixel or region of the prediction result during the prediction process. For example, the uncertainty score will significantly increase in areas where sensors are obstructed or in adverse weather conditions.
[0028] S103: The environmental information map, the predicted environmental information map, and the confidence score set are fused to obtain a risk semantic map, which contains environmental information and the corresponding risk level for each region. Specifically, the current environment feature map output by S101 is concatenated and fused with the future scene feature map predicted by S102. Simultaneously, the uncertainty confidence score generated by S102 is used as an independent channel and overlaid onto the fused feature map to generate a risk-aware semantic map. This map not only contains geometric and semantic information about "what is there," but also cognitive information about "how reliable the information there is" and "where there might be future dangers."
[0029] S104: Generate driving trajectory based on risk semantic map.
[0030] Specifically, the risk semantic map generated by S103 is input into a decision-making and planning network combining reinforcement learning (RL) and imitation learning (IL). The network's reward function is explicitly designed to include an uncertainty penalty term. The intelligent system is penalized for entering high-uncertainty areas, thus tending to choose more conservative and safer strategies. Simultaneously, the network is pre-trained using a large amount of human driving data (IL) to learn a smooth and comfortable driving style. Ultimately, the network outputs an optimal future trajectory that comprehensively considers safety, comfort, traffic rules, and destination guidance. The trajectory planned by S104 is then converted into specific control commands such as accelerator, brake, and steering, which are executed by the vehicle's drive-by-wire system.
[0031] This method involves acquiring an environmental information map containing current and historical environmental information; generating a predicted environmental information map for future time periods based on the current and historical environmental information, and a confidence score set corresponding to each region of the predicted environmental information map, where the scores in the confidence score set represent the degree of confidence in the environmental characteristics of a region; fusing the environmental information map, the predicted environmental information map, and the confidence score set to obtain a risk semantic map, which contains environmental information and the corresponding risk level for each region; and generating a driving trajectory based on the risk semantic map. This approach addresses the technical problems of insufficient capability and low safety of driving trajectories generated in complex scenarios in existing technologies, achieving the technical effect of improving the capability and safety of generated driving trajectories in complex scenarios.
[0032] In an optional embodiment, S101: Obtaining an environmental information map containing current environmental information and historical environmental information, including: S1011: The current environmental features collected by multiple sensors are fused to obtain current environmental information, wherein the current environmental features include the environmental features collected by the sensors at the current moment; S1012: The historical environmental features collected by the sensors are fused to obtain historical environmental information, wherein the historical environmental features include the historical environmental features collected by the sensors at past moments; S1013: Integrate current environmental information and historical environmental information to obtain an environmental information map.
[0033] In an optional embodiment, S1011: The current environmental features collected by multiple sensors are fused to obtain current environmental information, including: S10111: Based on the attention mechanism in cross-modal Transformer, the current environmental features collected by different sensors in their respective coordinate systems are mapped to a unified coordinate system. The sensors include camera equipment, radar and laser point clouds, and millimeter-wave radar point traces. S10112: Fuse historical environmental features collected by sensors to obtain historical environmental information, including: S10113: Based on the attention mechanism in cross-modal Transformer, historical environmental features collected by different sensors in their respective coordinate systems are mapped to a unified coordinate system; In an optional embodiment, S1013: The current environmental information and historical environmental information are fused to obtain an environmental information map, including: S10131: Based on recurrent neural networks and / or Transformer networks, current environmental information and historical environmental information are fused to obtain an environmental information map, wherein the environmental information map consists of an environmental rasterized feature map and a dynamic target structured list.
[0034] Furthermore, S102: Based on current environmental information and historical environmental information, generate a predicted environmental information map for future time periods, and a set of confidence scores corresponding to each region of the predicted environmental information map, including: S1021: Based on a self-supervised learning head, a predicted environmental information map for a future time period is generated according to the current environmental information and historical environmental information. The predicted environmental information map includes the predicted locations of static obstacles, passable areas, and dynamic objects. S1022: Quantify the confidence level of each region of the predicted environment information map to obtain a set of confidence scores corresponding to each region of the predicted environment information map, wherein each region includes at least one pixel.
[0035] In an optional embodiment, S104: Generate a driving trajectory based on the risk semantic map, including: S1041: Input the risk semantic map into a decision planning network that combines reinforcement learning and imitation learning. The reward function in the decision planning network is to punish according to the score, and the score is proportional to the degree of punishment. S1042: The decision-making and planning network outputs a driving trajectory that meets preset driving conditions, including safety, comfort, compliance with traffic rules, and efficiency.
[0036] Optionally, the reward function R satisfies: R = Rsafe + α * Rcomfort + β * Rprogress - γ * Runcertainty; Where Rsafe is the safety score, α is the comfort weight coefficient, Rcomfort is the comfort score, β is the feasibility weight coefficient, Rprogress is the feasibility score, γ is the confidence weight coefficient, and Runcertainty is the confidence score.
[0037] In an optional embodiment, S104: Before generating the driving trajectory based on the risk semantic map, the following steps are included: S110: Based on multiple historical user driving data, the decision planning network is pre-trained to learn a driving style that meets the target driving conditions, wherein the target driving conditions are at least one of the preset driving conditions input by the user. S104: After generating the driving trajectory based on the risk semantic map, it includes: S105: Generate driving instructions that meet the target driving conditions based on driving style and driving trajectory. The driving instructions include accelerator, brake, and steering. In one exemplary embodiment, this application provides a driving trajectory generation method, such as... Figure 2 and Figure 3 As shown, Figure 2 This application provides a schematic diagram of the hardware and logic architecture of a driving trajectory generation system, which is a system that runs the driving trajectory generation method. Figure 3 This is a schematic diagram of another driving trajectory generation method provided in an embodiment of this application. After the driving trajectory generation system is started, the sensor module continuously collects data, and the specific steps of the data processing flow are as follows: Step 1. Perception and Fusion: Camera images are transformed to BEV space using methods such as LSS to generate feature maps F_cam; LiDAR point clouds are used to generate voxel feature maps F_lidar using a PointPillar network. The fusion network employs a cross-modal Transformer to calculate the attention weights between F_cam and F_lidar, performs adaptive fusion, and outputs F_multi.
[0038] Step 2. Prediction and Uncertainty Estimation: A prediction network with a U-Net structure is used. Taking the F_multi of the current frame and the previous T frames as input, it outputs the predicted feature map F_pred^{t+k} for the next K frames and the corresponding uncertainty map U_map. The loss function is the Huber loss between the predicted map and the actual future data. The uncertainty is obtained through the log-variance term of the additional output of the prediction network.
[0039] Step 3. Decision Planning: Concatenate F_multi^t with F_pred^{t+1} to F_pred^{t+3} along the channel dimension and incorporate them into U_map to form R_map. The decision network uses a CNN-based encoder to extract features from R_map, followed by an MLP policy header and value function header. The reward function is designed as follows: `R = R_safe + α * R_comfort + β * R_progress - γ * R_uncertainty` Here, `R_uncertainty` is the average value of U_map within the area covered by the planned trajectory. The network is first trained with millions of steps of the PPO algorithm in simulation environments such as CARLA, and then fine-tuned under supervision by loading human driving data.
[0040] Step 4. Control: Use an LQR controller to track the optimal trajectory generated by S5 and output control commands.
[0041] The embodiments of this application have the following technical effects: Enhanced robustness: Through multimodal deep fusion and self-supervised future prediction, the system's ability to cope with scenarios such as single sensor timeliness, severe weather, and lack of high-precision maps is significantly improved. Uncertainty estimation provides the system with "self-awareness," enabling it to adopt safer degradation strategies when confidence is insufficient. Better generalization ability: Self-supervised learning tasks do not require a large amount of manually labeled data; they can utilize massive amounts of real driving data for pre-training and online learning, thereby better handling unseen long-tail scenarios. More human-like and comfortable experience: The decision-making process comprehensively considers uncertainty risks and human driving styles, resulting in smoother and more natural trajectories while ensuring safety, improving passenger comfort and trust. End-to-end optimization: The entire system, from perception to decision-making, can be jointly trained and optimized, avoiding the problem of error accumulation in traditional modular systems.
[0042] Based on the same concept, this application also provides a driving trajectory generation device, please refer to... Figure 4 , Figure 4 A structural diagram of a driving trajectory generation device provided in this application embodiment includes: The acquisition module 201 is used to acquire an environmental information map containing current environmental information and historical environmental information; The first generation module 202 is used to generate a predicted environmental information map for a future time period based on current environmental information and historical environmental information, as well as a set of confidence scores corresponding to each region of the predicted environmental information map, wherein the scores in the set of confidence scores represent the degree of confidence in the environmental characteristics of the region. The fusion module 203 is used to fuse the environmental information map, the predicted environmental information map, and the confidence score set to obtain a risk semantic map, wherein the risk semantic map contains environmental information and the corresponding risk level of each region; The second generation module 204 is used to generate driving trajectories based on the risk semantic map.
[0043] This method involves acquiring an environmental information map containing current and historical environmental information; generating a predicted environmental information map for future time periods based on the current and historical environmental information, and a confidence score set corresponding to each region of the predicted environmental information map, where the scores in the confidence score set represent the degree of confidence in the environmental characteristics of a region; fusing the environmental information map, the predicted environmental information map, and the confidence score set to obtain a risk semantic map, which contains environmental information and the corresponding risk level for each region; and generating a driving trajectory based on the risk semantic map. This approach addresses the technical problems of insufficient capability and low safety of driving trajectories generated in complex scenarios in existing technologies, achieving the technical effect of improving the capability and safety of generated driving trajectories in complex scenarios.
[0044] Furthermore, the acquisition module 201 includes a first fusion unit, a second fusion unit, and a third fusion unit.
[0045] The first fusion unit is used to fuse the current environmental features collected by multiple sensors to obtain current environmental information, wherein the current environmental features include the environmental features collected by the sensors at the current moment; The second fusion unit is used to fuse the historical environmental features collected by the sensors to obtain historical environmental information, wherein the historical environmental features include the historical environmental features collected by the sensors at past moments. The third fusion unit is used to fuse current environmental information and historical environmental information to obtain an environmental information map.
[0046] Furthermore, the first fusion unit includes a first mapping component.
[0047] The first mapping component is used to map the current environmental features collected by different sensors in their respective coordinate systems to a unified coordinate system based on the attention mechanism in the cross-modal Transformer. The sensors include camera devices, radar and laser point clouds, and millimeter-wave radar traces. Furthermore, the second fusion unit includes a second mapping component.
[0048] The second mapping component is used to map the historical environmental features collected by different sensors in their respective coordinate systems to a unified coordinate system based on the attention mechanism in the cross-modal Transformer. Furthermore, the third fusion unit includes fusion components.
[0049] The fusion component is used to fuse current environmental information and historical environmental information based on recurrent neural networks and / or Transformer networks to obtain an environmental information map, which consists of an environmental rasterized feature map and a dynamic target structured list.
[0050] Furthermore, the first generation module 202 includes a generation unit and a quantization unit. The generation unit is used to generate a predicted environmental information map for a future time period based on a self-supervised learning head, according to current environmental information and historical environmental information. The predicted environmental information map includes the predicted locations of static obstacles, passable areas, and dynamic objects. A quantization unit is used to quantify the confidence level of each region of the predicted environment information map, thereby obtaining a set of confidence scores corresponding to each region of the predicted environment information map, wherein each region includes at least one pixel.
[0051] Furthermore, the second generation module 204 includes an input unit and an output unit.
[0052] The input unit is used to input the risk semantic map into a decision-making and planning network that combines reinforcement learning and imitation learning. The reward function in the decision-making and planning network is to punish according to the score, and the score is proportional to the degree of punishment. The output unit is used to output a driving trajectory that meets preset driving conditions from the decision planning network. These preset driving conditions include safety, comfort, compliance with traffic rules, and efficiency.
[0053] Optionally, the reward function R satisfies: R = Rsafe + α * Rcomfort + β * Rprogress - γ * Runcertainty; Where Rsafe is the safety score, α is the comfort weight coefficient, Rcomfort is the comfort score, β is the feasibility weight coefficient, Rprogress is the feasibility score, γ is the confidence weight coefficient, and Runcertainty is the confidence score.
[0054] Furthermore, the driving trajectory generation device also includes a learning module and a third generation module.
[0055] The learning module is used to pre-train the decision planning network based on multiple historical user driving data before generating the driving trajectory according to the risk semantic map, and learn the driving style that meets the target driving conditions. The target driving conditions are at least one of the preset driving conditions input by the user. The third generation module is used to generate a driving trajectory based on the risk semantic map, and then generate driving instructions that meet the target driving conditions based on the driving style and driving trajectory. The driving instructions include accelerator, brake, and steering. This application also provides an electronic device, please refer to... Figure 5 , Figure 5 This is a structural diagram of an electronic device provided in an embodiment of this application.
[0056] like Figure 5 As shown, the electronic device 400 includes a processor 410.
[0057] like Figure 5 As shown, the processor 410 described above can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program in this application.
[0058] like Figure 5 As shown, the electronic device 400 may further include a communication line 440. The communication line 440 may include a path for transmitting information between the components.
[0059] Optional, such as Figure 5 As shown, the above-described electronic device may further include a communication interface 420. There may be one or more communication interfaces 420. The communication interface 420 may use any transceiver-like device for communicating with other devices or communication networks.
[0060] Optional, such as Figure 5As shown, the electronic device may further include a memory 430. The memory 430 stores computer execution instructions for implementing the scheme of this application, and its execution is controlled by a processor. The processor executes the computer execution instructions stored in the memory, thereby implementing the driving trajectory generation method provided in the embodiments of this application.
[0061] like Figure 5 As shown, memory 430 can be read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Memory 430 can exist independently and be connected to processor 410 via communication line 440. Memory 430 can also be integrated with processor 410.
[0062] Optionally, the computer execution instructions in the embodiments of this application may also be referred to as application code, and the embodiments of this application do not specifically limit this.
[0063] In a specific implementation, as one example, such as Figure 5 As shown, processor 410 may include one or more CPUs, such as Figure 5 CPU0 and CPU1 in the CPU.
[0064] In a specific implementation, as one example, such as Figure 5 As shown, the terminal device may include multiple processors, such as Figure 5 The first processor 4101 and the second processor 4102 are included. Each of these processors can be a single-core processor or a multi-core processor.
[0065] The methods disclosed in the embodiments of this application can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied as execution by a hardware decoding processor, or as a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above driving trajectory generation method.
[0066] This application also provides a computer-readable storage medium storing instructions that, when executed, implement the functions performed by the terminal device in the above embodiments.
[0067] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD).
[0068] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.
[0069] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.
Claims
1. A method for generating driving trajectories, characterized in that, include: Obtain an environmental information map containing current and historical environmental information; Based on current and historical environmental information, a predicted environmental information map for a future time period is generated, along with a set of confidence scores corresponding to each region of the predicted environmental information map, wherein the scores in the set of confidence scores represent the degree of confidence in the environmental characteristics of the region. The environmental information map, the predicted environmental information map, and the confidence score set are fused to obtain a risk semantic map, wherein the risk semantic map contains environmental information and the corresponding risk level for each region; The driving trajectory is generated based on the risk semantic map.
2. The method according to claim 1, characterized in that, Obtain an environmental information map containing current and historical environmental information, including: The current environmental features collected by multiple sensors are fused to obtain current environmental information, wherein the current environmental features include the environmental features collected by the sensors at the current moment; The historical environmental features collected by the sensors are fused to obtain historical environmental information, wherein the historical environmental features include historical environmental features collected by the sensors at past times; By fusing the current environmental information with the historical environmental information, an environmental information map is obtained.
3. The method according to claim 2, characterized in that, By fusing current environmental features collected from multiple sensors, current environmental information is obtained, including: Based on the attention mechanism in the cross-modal Transformer, the current environmental features collected by different sensors in their respective coordinate systems are mapped to a unified coordinate system. The sensors include camera devices, radar and laser point clouds, and millimeter-wave radar traces. The historical environmental features collected by the sensors are fused to obtain historical environmental information, including: Based on the attention mechanism in the cross-modal Transformer, the historical environmental features collected by different sensors in their respective coordinate systems are mapped to a unified coordinate system; By fusing the current environmental information and the historical environmental information, an environmental information map is obtained, including: The current environmental information and the historical environmental information are fused based on recurrent neural networks and / or Transformer networks to obtain an environmental information map, wherein the environmental information map consists of an environmental rasterized feature map and a dynamic target structured list.
4. The method according to claim 1, characterized in that, Based on current and historical environmental information, a predicted environmental information map for future time periods is generated, along with a set of confidence scores corresponding to each region of the predicted environmental information map, including: Based on a self-supervised learning head, a predicted environmental information map for a future time period is generated according to current and historical environmental information. The predicted environmental information map includes the predicted locations of static obstacles, passable areas, and dynamic objects. The confidence level of each region of the predicted environment information map is quantified to obtain a confidence score set corresponding to each region of the predicted environment information map, wherein each region includes at least one pixel.
5. The method according to claim 4, characterized in that, Based on the risk semantic map, the driving trajectory is generated, including: The risk semantic map is input into a decision planning network that combines reinforcement learning and imitation learning, wherein the reward function in the decision planning network is to punish according to the score, and the score is proportional to the degree of punishment; The decision planning network outputs a driving trajectory that meets preset driving conditions, including safety, comfort, compliance with traffic rules, and efficiency.
6. The method according to claim 5, characterized in that, The reward function R satisfies: R = Rsafe + α * Rcomfort + β * Rprogress - γ * Runcertainty; Where Rsafe is the safety score, α is the comfort weight coefficient, Rcomfort is the comfort score, β is the feasibility weight coefficient, Rprogress is the feasibility score, γ is the confidence weight coefficient, and Runcertainty is the confidence score.
7. The method according to claim 5, characterized in that, Before generating the driving trajectory based on the risk semantic map, the process includes: Based on multiple historical user driving data, the decision planning network is pre-trained to learn a driving style that meets the target driving conditions, wherein the target driving conditions are at least one of the preset driving conditions input by the user. After generating the driving trajectory based on the risk semantic map, the process includes: Based on the driving style and the driving trajectory, driving instructions that meet the target driving conditions are generated, wherein the driving instructions include accelerator, brake, and steering.
8. A driving trajectory generation device, characterized in that, include: The acquisition module is used to acquire an environmental information map containing current and historical environmental information. The first generation module is used to generate a predicted environmental information map for a future time period based on current environmental information and historical environmental information, as well as a set of confidence scores corresponding to each region of the predicted environmental information map, wherein the scores in the set of confidence scores represent the degree of confidence in the environmental characteristics of the region. The fusion module is used to fuse the environmental information map, the predicted environmental information map, and the confidence score set to obtain a risk semantic map, wherein the risk semantic map contains environmental information and the corresponding risk level for each region; The second generation module is used to generate the driving trajectory based on the risk semantic map.
9. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the driving trajectory generation method according to any one of claims 1 to 7 by running instructions in the memory.
10. A computer storage medium, characterized in that, The computer storage medium stores instructions that, when executed, implement the driving trajectory generation method according to any one of claims 1 to 7.