Vehicle control method, system, and electronic device
The vehicle control system, through multi-module collaborative processing, achieves a complete closed loop from perception to decision-making, improving the safety and reliability of assisted driving and solving the shortcomings of environmental perception and decision-making planning in existing technologies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NULLMAX INC
- Filing Date
- 2026-03-24
- Publication Date
- 2026-05-19
AI Technical Summary
Existing driver assistance technologies are insufficient in terms of safety and reliability, making it difficult to achieve high-precision environmental perception and decision-making.
The vehicle control system employs multi-module collaborative processing, including an end-to-end model module, an environment model module, an obstacle post-fusion module, an obstacle prediction post-processing module, and a decision planning module. Through multi-modal information fusion and a clearly defined information processing flow, it improves the accuracy and consistency of perception, fusion, environmental modeling, and decision planning.
It improves the safety and reliability of assisted driving, ensures safe, efficient and compliant driving behavior in complex traffic environments, reduces information distortion and error accumulation, and enhances the overall perception accuracy and real-time performance of the system.
Smart Images

Figure CN121893993B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of driver assistance technology, and in particular to a vehicle control method, system and electronic device. Background Technology
[0002] Assisted driving (ADAS) is a collective term for technologies that use onboard sensors, cameras, radar (including millimeter-wave radar and lidar), ultrasonic sensors, and other devices to perceive the vehicle's driving environment in real time and assist the driver in completing some or all driving tasks. It encompasses driving functions such as lane keeping assist, automatic parking, and adaptive cruise control, aiming to improve driving comfort, safety, and intelligence. With the continuous development of ADAS technologies, they have become one of the core functions of intelligent vehicles. Therefore, how to achieve safer and more reliable ADAS is a direction of ongoing exploration and research in this field. Summary of the Invention
[0003] The purpose of this application is to address the technical problem of how to achieve safer and more reliable driver assistance.
[0004] To address the aforementioned technical problems, in a first aspect, this application discloses a vehicle control method applied to a vehicle control system for assisted driving control of a vehicle. The system includes an end-to-end model module, an environment model module, an obstacle post-fusion module, an obstacle prediction post-processing module, and a decision planning module. The method includes: the end-to-end model module obtaining corresponding perception result information based on the vehicle's driving perception information input to the end-to-end model module. The perception result information includes road feature identification information corresponding to the vehicle's driving environment, obstacle perception information corresponding to obstacles around the vehicle, and the vehicle's self-planned trajectory information. The road feature identification information is output to the environment model module, the obstacle perception information is output to the obstacle post-fusion module, and the self-planned trajectory information is output to the decision planning module. The obstacle post-fusion module obtains corresponding obstacle post-fusion information based on the obstacle perception information and obstacle post-fusion reference information, and outputs the obstacle post-fusion information to the environment model module and the obstacle prediction post-processing module, respectively. The obstacle prediction post-processing module and the environment model module obtain corresponding driving environment identification information based on road feature information, obstacle post-fusion information, and driving environment reference information, and output the driving environment identification information to the obstacle prediction post-processing module and the decision planning module respectively; the obstacle prediction post-processing module obtains corresponding obstacle prediction trajectory information based on obstacle post-fusion information, driving environment identification information, and prediction post-processing reference information, and outputs the obstacle prediction trajectory information to the decision planning module; the decision planning module obtains corresponding driving control information based on the vehicle's planned trajectory information, obstacle prediction trajectory information, driving environment identification information, and decision planning reference information for vehicle driving control.
[0005] Using the above method, the end-to-end model module comprehensively analyzes the driving perception information, outputting the obtained road feature identification information, obstacle perception information, and vehicle planned trajectory information to their respective modules. This achieves the separation and preprocessing of driving perception information, providing high-quality basic data for subsequent environmental modeling, obstacle fusion, and decision planning, improving the accuracy and real-time performance of overall driving perception, thereby enhancing the safety and reliability of assisted driving. The obstacle post-fusion module fuses obstacle perception information and obstacle post-fusion reference information to generate more reliable obstacle fusion results, which are simultaneously transmitted to the environment model module and the obstacle prediction post-processing module. This ensures the consistency of data used in environmental modeling and obstacle prediction, reducing information distortion and error accumulation, thus improving the safety and reliability of assisted driving. The environment model module integrates road feature information, obstacle post-fusion information, and driving environment reference information to construct accurate driving environment identifiers, providing a unified and reliable environmental description for subsequent obstacle prediction and decision planning. This enables obstacle prediction and decision planning to work collaboratively based on the same environment, avoiding decision conflicts caused by environmental perception biases, further ensuring the safety and reliability of assisted driving. The obstacle prediction post-processing module integrates obstacle post-fusion information, driving environment labeling information, and prediction post-processing reference information to more accurately predict the future trajectory of surrounding obstacles, thereby improving the safety and reliability of assisted driving. The decision-making and planning module combines the vehicle's planned trajectory information, obstacle prediction trajectory information, driving environment labeling information, and decision-making and planning reference information to generate optimized driving control information. Based on this driving control information, the module controls the vehicle's movement, ensuring that the vehicle can make safe, efficient, and compliant driving behaviors in complex and changing traffic environments. Ultimately, this achieves safer and more reliable assisted driving control, thus improving the safety and reliability of assisted driving.
[0006] In summary, this solution, through multi-module collaborative processing, with clear division of labor among modules and progressive refinement and interaction of information, forms a complete closed loop from perception, fusion, environmental modeling, trajectory prediction to decision planning, thereby improving the precision and accuracy of information processing and enhancing the safety and reliability of assisted driving.
[0007] In one possible implementation of the first aspect above, the vehicle control system further includes a pre-processing module, and the method further includes: the pre-processing module performs feature extraction processing based on the original driving-related information input to the pre-processing module to obtain driving perception information corresponding to the original driving-related information, and outputs the driving perception information to the end-to-end model module.
[0008] By employing the above method, feature extraction of raw driving-related information through a pre-processing module can effectively filter out noise and redundant information in the raw data, extract key features highly relevant to driving decisions, and generate high-quality driving perception information. This provides accurate input for subsequent end-to-end model modules, thereby improving the perception efficiency and accuracy of the entire system and laying a solid data foundation for safe and reliable assisted driving.
[0009] In one possible implementation of the first aspect above, the end-to-end model module includes an end-to-end model, which is a multimodal large language model. The end-to-end model module obtains corresponding perception result information based on the driving perception information corresponding to the vehicle input to the end-to-end model module, including: performing multimodal feature extraction processing and feature fusion processing based on the driving perception information to obtain the corresponding perception result information.
[0010] Using the above method, the end-to-end model module employs a multimodal large language model to extract and fuse multimodal features from driving perception information. This deeply integrates information from different sources such as vision, radar, and navigation, enabling a more comprehensive understanding of complex traffic scenarios. This allows for more accurate identification of road features, obstacle states, and the vehicle's planned trajectory, improving the accuracy of perception results and providing a more reliable basis for subsequent environmental modeling and decision-making. Ultimately, this enhances the safety and reliability of assisted driving.
[0011] In one possible implementation of the first aspect above, the obstacle perception information includes information about visually perceived obstacles, the obstacle post-fusion reference information includes information about obstacles to be fused, the obstacles to be fused include non-visually perceived obstacles, and the obstacle post-fusion module obtains corresponding obstacle post-fusion information based on the obstacle perception information and the obstacle post-fusion reference information, including: fusing the visually perceived obstacles and the obstacles to be fused based on the obstacle perception information and the obstacle post-fusion reference information to obtain obstacle post-fusion information.
[0012] By employing the above method, visually perceived obstacles and non-visually perceived obstacles such as radar obstacles are fused together, the advantages of different sensors can be fully utilized to overcome the limitations of a single sensor in a specific environment. This generates more complete and accurate obstacle fusion information, thereby improving the system's reliability and accuracy in perceiving surrounding obstacles and enhancing the safety and reliability of assisted driving.
[0013] In one possible implementation of the first aspect above, the environment model module obtains corresponding driving environment identification information based on road feature information, obstacle post-fusion information, and driving environment reference information, including: performing local mapping processing and driving environment topology extraction processing on the road feature information, obstacle post-fusion information, and driving environment reference information to obtain environmental constraint information corresponding to the driving environment in which the vehicle is located; obtaining driving reference line information corresponding to the vehicle based on the environmental constraint information; and obtaining driving environment identification information based on the driving reference line information.
[0014] Using the above method, based on road feature information, obstacle fusion information, and driving environment reference information, local mapping and topology extraction of the driving environment are performed. Through these processes, a precise geometric map of the current driving environment can be constructed, and its road topology can be understood. This allows for the acquisition of environmental constraints and driving reference lines, enabling the vehicle to accurately perceive key environmental elements such as lanes and intersections. This provides clear environmental boundaries for obstacle prediction and autonomous vehicle decision-making, thereby improving the rationality of path planning and the smoothness of the driving trajectory, ensuring the safety of assisted driving. This ultimately enhances the safety and reliability of assisted driving.
[0015] In one possible implementation of the first aspect described above, the prediction post-processing module obtains the corresponding obstacle prediction trajectory information based on the obstacle post-fusion information, driving environment identification information, and prediction post-processing reference information. This includes: obtaining the corresponding model intent and rule intent based on a preset matching strategy, according to the obstacle post-fusion information, driving environment identification information, and prediction post-processing reference information; and performing matching processing on the model intent and rule intent to determine whether the model intent and rule intent match successfully. If the match is successful, smoothing and extending processing is performed on the initial obstacle prediction trajectory information corresponding to the obstacle post-fusion information to generate obstacle prediction trajectory information. If the match is unsuccessful, obstacle prediction trajectory information is generated based on a kinematic model.
[0016] Using the above method, based on a preset matching strategy, the corresponding model intent and rule intent are obtained according to obstacle post-fusion information, driving environment identification information, and prediction post-processing reference information. By matching the model intent with the rule intent, while fusing data to drive prediction, prior knowledge such as traffic rules is introduced for constraint verification. When a match is successful, the trajectory is smoothly extended to ensure the continuity and rationality of the prediction; when a match fails, a conservative kinematic model is used as a safe alternative to ensure the basic reliability of the prediction. This scheme improves the accuracy of obstacle behavior prediction and effectively prevents potential collision risks, thereby enhancing the safety and reliability of assisted driving.
[0017] In one possible implementation of the first aspect mentioned above, the decision-making and planning module includes a decision-making submodule and a planning submodule. The decision-making and planning module obtains corresponding driving control information based on the vehicle's planned trajectory information, obstacle prediction trajectory information, driving environment identification information, and decision-making and planning reference information. This includes: the decision-making submodule obtaining corresponding driving behavior decision information as first driving control information based on obstacle prediction trajectory information, driving environment identification information, and decision-making and planning reference information, and outputting the driving behavior decision information to the planning submodule; and the planning submodule obtaining corresponding initial vehicle control trajectory information based on the vehicle's planned trajectory information, driving environment identification information, and decision-making and planning reference information, and dynamically correcting the initial vehicle control trajectory information based on obstacle prediction trajectory information and driving behavior decision information to obtain corresponding target vehicle control trajectory information as second driving control information.
[0018] Using the above method, the decision-making and planning module is divided into a decision-making sub-module and a planning sub-module. The decision-making sub-module generates corresponding driving behavior decision information, i.e., the first driving control information, based on obstacle prediction information, driving environment identification information, and decision-making and planning reference information, providing guidance for decision-making and planning. The planning sub-module generates corresponding initial vehicle control trajectory information based on the vehicle's planned trajectory information, driving environment identification information, and decision-making and planning reference information, and dynamically corrects it based on obstacle prediction trajectory information and driving behavior decision information to generate the final target vehicle control trajectory, i.e., the second driving control information. This hierarchical architecture enables decision-making to consider traffic rules and scenario constraints globally, and planning to flexibly respond to dynamic obstacles. It ensures the rationality and safety of driving strategies while improving the real-time performance and accuracy of trajectory planning, thereby effectively improving the reliability and safety of assisted driving.
[0019] In one possible implementation of the first aspect mentioned above, the driving perception information includes: vehicle driving status information, driving environment information corresponding to the driving environment in which the vehicle is located, and navigation information and positioning information corresponding to the vehicle.
[0020] In one possible implementation of the first aspect described above, the obstacle to be fused is a radar obstacle.
[0021] In one possible implementation of the first aspect above, the obstacle post-fusion information includes: fused obstacle information.
[0022] In one possible implementation of the first aspect above, the driving environment reference information includes: traffic safety constraint information, general obstacle information, and navigation and positioning information corresponding to the driving environment in which the vehicle is located.
[0023] In one possible implementation of the first aspect mentioned above, the driving environment identification information includes a local map of the driving environment, a driving environment topology, lane information of the vehicle in the local map of the driving environment, the location navigation information of the target lane in the local map of the driving environment, traffic flow information, traffic scene safety constraint information, obstacle information, and driving speed constraint information.
[0024] In one possible implementation of the first aspect above, the prediction post-processing reference information includes: the vehicle's corresponding positioning information.
[0025] In one possible implementation of the first aspect mentioned above, the decision planning reference information includes: vehicle-specific dynamics constraints, traffic rule constraints, driving mode preference information, speed limits and navigation guidance information, and vehicle parameters and load status information.
[0026] In one possible implementation of the first aspect mentioned above, driving behavior decision information includes lane keeping, lane changing, ramp merging and merging, and intersection turning decisions.
[0027] By employing the above methods, the system incorporates vehicle driving status, driving environment, navigation, and positioning information into the driving perception information, enabling it to comprehensively perceive the driving situation from multiple dimensions, laying a solid foundation for subsequent processing. By fusing radar obstacles as obstacles to be fused with visual perception, the system achieves complementary advantages of multiple sensors, significantly enhancing the accuracy of perception. By introducing traffic safety constraint information, general obstacle information, and the vehicle's corresponding navigation and positioning information as driving environment references, the environmental model can construct high-precision environmental markings that conform to traffic rules and accurately reflect actual roads. Using positioning information as a reference for prediction post-processing allows obstacle trajectory prediction to be accurately analyzed in conjunction with the vehicle's position, improving the ability to predict potential risks. By incorporating multi-dimensional information into the driving environment marking information, the system provides detailed environmental awareness for decision-making and planning. By covering driving behaviors such as lane keeping, lane changing, ramp merging and merging, and intersection turning, the system ensures that the decision-making submodule can cope with diverse traffic scenarios, thereby improving the reliability of the assisted driving system throughout the entire chain from perception to decision-making. This effectively enhances the reliability and safety of assisted driving.
[0028] Secondly, this application discloses a vehicle control system, which includes an end-to-end model module, an environment model module, an obstacle post-fusion module, an obstacle prediction post-processing module, and a decision planning module. The end-to-end model module is used to obtain corresponding perception result information based on the vehicle's driving perception information input to the end-to-end model module. The perception result information includes road feature identification information corresponding to the vehicle's driving environment, obstacle perception information corresponding to obstacles around the vehicle, and the vehicle's self-planned trajectory information. The road feature identification information is output to the environment model module, the obstacle perception information is output to the obstacle post-fusion module, and the self-planned trajectory information is output to the decision planning module. The obstacle post-fusion module is used to obtain... After obtaining the corresponding obstacle fusion information, the obstacle fusion information is output to the environment model module and the obstacle prediction post-processing module respectively. The environment model module is used to obtain the corresponding driving environment identification information based on road feature information, obstacle fusion information, and driving environment reference information, and outputs the driving environment identification information to the obstacle prediction post-processing module and the decision planning module respectively. The obstacle prediction post-processing module is used to obtain the corresponding obstacle prediction trajectory information based on the obstacle fusion information, driving environment identification information, and prediction post-processing reference information, and outputs the obstacle prediction trajectory information to the decision planning module. The decision planning module is used to obtain the corresponding driving control information based on the vehicle's planned trajectory information, obstacle prediction trajectory information, driving environment identification information, and decision planning reference information for vehicle driving control.
[0029] Thirdly, embodiments of this application disclose an electronic device that includes the aforementioned vehicle control system for implementing the aforementioned vehicle control method.
[0030] The relevant beneficial effects of the second and third aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0031] To more clearly illustrate the technical solution of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below.
[0032] Figure 1 A schematic diagram of a vehicle control system provided in an embodiment of this application;
[0033] Figure 2 Another structural schematic diagram of the vehicle control system provided in the embodiments of this application;
[0034] Figure 3 A schematic flowchart of a vehicle control method provided in an embodiment of this application;
[0035] Figure 4Another schematic flowchart of the vehicle control method provided in the embodiments of this application;
[0036] Figure 5 A schematic diagram illustrating the principle of an end-to-end large model processing procedure provided in an embodiment of this application;
[0037] Figure 6 This is a schematic diagram illustrating the principle of lane sequence filling provided in an embodiment of this application;
[0038] Figure 7 This is a schematic diagram illustrating the principle of smoothing lane sequence centerline interpolation provided in an embodiment of this application.
[0039] Figure 8 A schematic diagram illustrating the principle of the environmental model module processing procedure provided in this application embodiment;
[0040] Figure 9 A schematic diagram illustrating the processing procedure of the obstacle prediction post-processing module provided in an embodiment of this application;
[0041] Figure 10 This is a schematic diagram illustrating the principle of a vehicle control method based on a vehicle control system, as provided in an embodiment of this application. Detailed Implementation
[0042] To achieve safer and more reliable driver assistance, such as Figure 1 As shown, an embodiment of this application discloses a vehicle control system, which includes an end-to-end model module, an environment model module, an obstacle post-fusion module, an obstacle prediction post-processing module, and a decision planning module.
[0043] The end-to-end model module is used to obtain corresponding perception result information based on the driving perception information of the vehicle input to the end-to-end model module. The perception result information includes road feature identification information corresponding to the driving environment of the vehicle, obstacle perception information corresponding to the obstacles around the vehicle, and the vehicle's self-planned trajectory information. The road feature identification information is output to the environment model module, the obstacle perception information is output to the obstacle post-fusion module, and the self-planned trajectory information is output to the decision planning module.
[0044] The obstacle post-fusion module is used to obtain the corresponding obstacle post-fusion information based on the obstacle perception information and the obstacle post-fusion reference information, and outputs the obstacle post-fusion information to the environment model module and the obstacle prediction post-processing module respectively.
[0045] The environment model module is used to obtain corresponding driving environment identification information based on road feature information, obstacle post-fusion information, and driving environment reference information, and outputs the driving environment identification information to the obstacle prediction post-processing module and the decision planning module respectively.
[0046] The obstacle prediction post-processing module is used to obtain the corresponding obstacle prediction trajectory information based on the obstacle post-fusion information, driving environment identification information, and prediction post-processing reference information, and output the obstacle prediction trajectory information to the decision planning module.
[0047] The decision planning module is used to obtain corresponding driving control information based on the vehicle's planned trajectory information, obstacle prediction trajectory information, driving environment identification information, and decision planning reference information, for use in vehicle driving control.
[0048] The end-to-end model module includes the end-to-end model, which is the end-to-end (E2E) large model, and can be simply referred to as the end-to-end model. Additionally, the environment model module can be simply referred to as the environment model, the obstacle post-fusion module can be simply referred to as obstacle post-fusion, and the obstacle prediction post-processing module can be simply referred to as obstacle prediction post-processing.
[0049] Furthermore, such as Figure 2 As shown, the vehicle control system may also include a pre-processing module, which is used to perform feature extraction processing based on the original driving-related information input to the pre-processing module to obtain driving perception information corresponding to the original driving-related information, and output the driving perception information to the end-to-end model module.
[0050] like Figure 3 As shown, an embodiment of this application discloses a vehicle control method applied to the aforementioned vehicle control system for assisted driving control of a vehicle. The method includes:
[0051] In step S100, the end-to-end model module obtains corresponding perception result information based on the driving perception information of the vehicle input to the end-to-end model module. The perception result information includes road feature identification information corresponding to the driving environment of the vehicle, obstacle perception information corresponding to obstacles around the vehicle, and vehicle planning trajectory information. The road feature identification information is output to the environment model module, the obstacle perception information is output to the obstacle fusion module, and the vehicle planning trajectory information is output to the decision planning module.
[0052] The driving perception information can include, for example, vehicle driving status information, driving environment information corresponding to the vehicle's current driving environment, and corresponding navigation and positioning information. Vehicle driving status information mainly refers to information related to the vehicle's current kinematic and dynamic state. For example, this could be ego motion information, including vehicle speed, acceleration, steering wheel angle, yaw rate, and vehicle bus status, etc., reflecting the vehicle's real-time motion. Driving environment information mainly refers to raw data or lightweight features of the surrounding environment acquired through sensors such as cameras and radar. Examples include images, such as raw images captured by onboard cameras, bird's-eye-view (BEV) features including lane lines, stop lines, crosswalks, road arrows, merging and diverging points, obstacles, and drivable areas, as well as environmental information such as traffic light status, stop line positions, and traffic signs. The corresponding navigation information mainly refers to guidance information provided by global path planning, such as navigation information including path shape, target lane, and next intersection behavior. The vehicle's positioning information mainly refers to the vehicle's precise position and attitude in a global or local coordinate system. For example, positioning information may include Global Positioning System (GPS) coordinates and heading angle.
[0053] Furthermore, the input to the end-to-end model module can also include initialization information to guide the end-to-end model in decoding specific output tasks from the fused multimodal features, such as spatial paths, temporal trajectories, and perception results. The initialization information does not reflect any real physical quantities or external states. Examples include spatial path initialization information (pathquery), temporal trajectory initialization information (trajectory query), perception output initialization information (perception query), and text description initialization information (description query).
[0054] Perception result information refers to the output of the end-to-end model module after performing deep processing on driving perception information, such as feature extraction and feature fusion. It includes structured and semantic environmental understanding and the vehicle's future trajectory. This includes road feature identification information, obstacle perception information, and vehicle planned trajectory information. Road feature identification information describes the road geometry and attributes of the vehicle's driving area, such as lane lines and related features. Obstacle perception information identifies and quantifies information about surrounding dynamic / static obstacles, such as obstacle categories, locations, motion states, and origin-destination prediction (OD) trajectories. Vehicle planned trajectory information refers to the vehicle's future driving path and temporal trajectory generated by the end-to-end model module, such as spatial paths, temporal trajectories, perception outputs, and text descriptions.
[0055] Furthermore, as mentioned earlier, the input driving perception information for the end-to-end model module is generated by the aforementioned pre-processing module. The aforementioned raw driving-related information refers to unprocessed, raw data directly collected from vehicle sensors and communication buses, which may include diverse information such as visual perception data, radar perception data, lidar data, navigation data, positioning data, and vehicle status data. This raw driving-related information is input to the pre-processing module, which is responsible for feature extraction, encoding, and formatting of the multi-source, heterogeneous raw driving-related information, converting it into unified, high-dimensional driving perception information for use by the end-to-end model module (i.e., the end-to-end large model). This pre-processing module can be, for example, a tokenization module or other types of pre-processing modules.
[0056] In step S200, the obstacle post-fusion module obtains the corresponding obstacle post-fusion information based on the obstacle perception information and the obstacle post-fusion reference information, and outputs the obstacle post-fusion information to the environment model module and the obstacle prediction post-processing module respectively.
[0057] The obstacle post-fusion reference information refers to the auxiliary information input to the obstacle post-fusion module, used to fuse with visually perceived obstacle information to generate more complete and accurate obstacle post-fusion information. The obstacle post-fusion reference information may include obstacle data detected by non-visual sensors (such as millimeter-wave radar and lidar), as well as other necessary fusion parameters, such as sensor calibration parameters and time synchronization information. The obstacle post-fusion information refers to the fused obstacle information, including, for example, the fusion results of visual perception and non-visual sensors, such as obstacle post-fusion information.
[0058] In step S300, the environment model module obtains the corresponding driving environment identification information based on road feature information, obstacle post-fusion information, and driving environment reference information, and outputs the driving environment identification information to the obstacle prediction post-processing module and the decision planning module respectively.
[0059] Among them, driving environment reference information refers to the auxiliary information input into the environment model module, which is used to combine road feature information and obstacle fusion information to jointly construct an accurate driving environment label. Driving environment reference information may include, for example, traffic safety constraint information corresponding to the vehicle's driving environment, such as traffic light signal information, traffic sign information, etc., general obstacle information, and vehicle-specific navigation information such as standard definition (SD) navigation, positioning information such as GPS dead reckoning (GPS-DR) positioning, odometer information, etc.
[0060] Driving environment identification information refers to the output of the environment model module. It is a comprehensive digital description of the vehicle's current driving environment, providing static boundaries and road topology for downstream decision-making and planning. Driving environment identification information can include, for example, environmental information such as a local driving environment map, driving environment topology, lane information of the vehicle in the local driving environment map, the location navigation score of the target lane in the local driving environment map, traffic flow information, traffic scene safety constraints, obstacle information, driving speed constraints, etc.
[0061] In step S400, the obstacle prediction post-processing module obtains the corresponding obstacle prediction trajectory information based on the obstacle post-fusion information, driving environment identification information, and prediction post-processing reference information, and outputs the obstacle prediction trajectory information to the decision planning module.
[0062] The prediction post-processing reference information refers to the auxiliary information input into the obstacle prediction post-processing module. It is used to correct and extend the initial predicted trajectory of the obstacle by combining the obstacle post-fusion information and driving environment identification information. The prediction post-processing reference information may include, for example, the vehicle's corresponding positioning information.
[0063] Obstacle prediction trajectory information refers to the output of the obstacle prediction post-processing module, which describes the expected movement trajectory of surrounding dynamic obstacles over a period of time, providing dynamic interactive input for decision-making and planning. This information typically includes, for example, a time series of trajectory points, with each point carrying attributes such as position, velocity, orientation angle, and confidence level, such as an 8-second obstacle prediction trajectory.
[0064] In step S500, the decision planning module obtains corresponding driving control information based on the vehicle's planned trajectory information, obstacle prediction trajectory information, driving environment identification information, and decision planning reference information, for use in vehicle driving control.
[0065] Among them, decision-making and planning reference information refers to auxiliary information input into the decision-making and planning module, which is used to consider factors such as the vehicle's own capabilities, regulatory constraints, and driving preferences when generating final driving control information. This information may include, for example, the vehicle's corresponding dynamics constraints, traffic rule constraints, driving mode preference information, legal speed limits and navigation guidance information, as well as vehicle parameters and load status information, such as vehicle dynamics parameters and traffic rules.
[0066] Driving control information refers to the final output of the decision-making and planning module, used to directly control the vehicle's actuators, such as steering, driving, and braking, to achieve safe and reliable assisted driving. This information can be, for example, a sequence of trajectory points or low-level control commands.
[0067] Using the above method, the end-to-end model module comprehensively analyzes the driving perception information, outputting the obtained road feature identification information, obstacle perception information, and vehicle planned trajectory information to their respective modules. This achieves the separation and preprocessing of driving perception information, providing high-quality basic data for subsequent environmental modeling, obstacle fusion, and decision planning, improving the accuracy and real-time performance of overall driving perception, thereby enhancing the safety and reliability of assisted driving. The obstacle post-fusion module fuses obstacle perception information and obstacle post-fusion reference information to generate more reliable obstacle fusion results, which are simultaneously transmitted to the environment model module and the obstacle prediction post-processing module. This ensures the consistency of data used in environmental modeling and obstacle prediction, reducing information distortion and error accumulation, thus improving the safety and reliability of assisted driving. The environment model module integrates road feature information, obstacle post-fusion information, and driving environment reference information to construct accurate driving environment identifiers, providing a unified and reliable environmental description for subsequent obstacle prediction and decision planning. This enables obstacle prediction and decision planning to work collaboratively based on the same environment, avoiding decision conflicts caused by environmental perception biases, further ensuring the safety and reliability of assisted driving. The obstacle prediction post-processing module integrates obstacle post-fusion information, driving environment labeling information, and prediction post-processing reference information to more accurately predict the future trajectory of surrounding obstacles, thereby improving the safety and reliability of assisted driving. The decision-making and planning module combines the vehicle's planned trajectory information, obstacle prediction trajectory information, driving environment labeling information, and decision-making and planning reference information to generate optimized driving control information. This ensures that the vehicle can make safe, efficient, and compliant driving behaviors in complex and changing traffic environments, ultimately achieving safer and more reliable assisted driving control, thus improving the safety and reliability of assisted driving.
[0068] In summary, this solution, through multi-module collaborative processing, with clear division of labor among modules and progressive refinement and interaction of information, forms a complete closed loop from perception, fusion, environmental modeling, trajectory prediction to decision planning, thereby improving the precision and accuracy of information processing and enhancing the safety and reliability of assisted driving.
[0069] In one possible implementation of the above method, the end-to-end model module includes an end-to-end model, which is a multimodal large language model. The end-to-end model module obtains corresponding perception result information based on the driving perception information corresponding to the vehicle input to the end-to-end model module, including: performing multimodal feature extraction processing and feature fusion processing based on the driving perception information to obtain the corresponding perception result information.
[0070] The processing of the end-to-end model module can be as follows: The end-to-end model of the end-to-end model module first performs modality-specific feature extraction on multi-source driving perception information, such as vehicle driving status information, environmental image information, navigation instructions and positioning data, for example, extracting image spatial features through a visual encoder, extracting kinematic features through a multilayer perceptron, and extracting navigation semantic features through an embedding layer; then, these heterogeneous features are mapped to a unified high-dimensional feature space, and deep fusion is performed using a self-attention and cross-attention mechanism based on a transformer, so that information from different modalities can fully interact and complement each other; finally, the fused features are fed into multiple task-specific decoding heads, and are decoded in parallel by the end-to-end large model to generate structured perception result information, such as road feature identification, such as lane line type and location, obstacle perception information, such as obstacle target category and motion state, and vehicle planned trajectory, such as spatial path and temporal trajectory.
[0071] Using the above method, the end-to-end model module employs a multimodal large language model to extract and fuse multimodal features from driving perception information. This deeply integrates information from different sources such as vision, radar, and navigation, enabling a more comprehensive understanding of complex traffic scenarios. This allows for more accurate identification of road features, obstacle states, and the vehicle's planned trajectory, improving the accuracy of perception results and providing a more reliable basis for subsequent environmental modeling and decision-making. Ultimately, this enhances the safety and reliability of assisted driving.
[0072] In one possible implementation of the above method, obstacle perception information includes information about visually perceived obstacles, obstacle post-fusion reference information includes information about obstacles to be fused, obstacles to be fused include non-visually perceived obstacles, and the obstacle post-fusion module obtains corresponding obstacle post-fusion information based on the obstacle perception information and the obstacle post-fusion reference information, including: fusing the visually perceived obstacles and the obstacles to be fused based on the obstacle perception information and the obstacle post-fusion reference information to obtain obstacle post-fusion information.
[0073] The obstacle to be merged is a radar obstacle. Of course, the obstacle to be merged can also be other types of obstacles.
[0074] Furthermore, the processing procedure of the obstacle post-fusion module can be as follows: First, the obstacle post-fusion module performs spatiotemporal synchronization and coordinate system unification on the input visual perception obstacle information, such as obstacle category, position, and velocity, and radar obstacle information, such as distance, radial velocity, and reflection intensity, to ensure that the two are aligned under the same spatiotemporal reference. Then, it matches the visual target and the radar target through a data association algorithm to identify different sensor observation results of the same obstacle. Finally, it uses a fusion algorithm to perform state estimation and fusion on the matched multi-source observation data, combining the semantic information of vision with the precise ranging and velocity measurement advantages of radar to generate obstacle post-fusion information that includes obstacle category, precise position, motion state, and confidence level.
[0075] By employing the above method, visually perceived obstacle information is fused with non-visually perceived obstacle information, such as radar obstacles, which is the obstacle to be fused. This fully leverages the complementary advantages of different sensors, overcomes the limitations of a single sensor in a specific environment, and generates more complete and accurate obstacle fusion information. This improves the system's reliability and accuracy in perceiving surrounding obstacles, thereby enhancing the safety and reliability of assisted driving.
[0076] In one possible implementation of the above method, the environment model module obtains corresponding driving environment identification information based on road feature information, obstacle post-fusion information, and driving environment reference information. This includes: performing local mapping processing and driving environment topology extraction processing on the road feature information, obstacle post-fusion information, and driving environment reference information to obtain environmental constraint information corresponding to the driving environment in which the vehicle is located; obtaining driving reference line information corresponding to the vehicle based on the environmental constraint information; and obtaining driving environment identification information based on the driving reference line information.
[0077] Among them, the driving reference line information can be simply referred to as the reference line.
[0078] Furthermore, the processing procedure of the environment model module can be as follows: First, the environment model module performs local mapping based on road feature information such as lane lines, obstacle fusion information, and driving environment reference information such as navigation, positioning, and traffic lights to construct the road geometry within the vehicle's surrounding area; then, through topology extraction processing, it identifies lane grouping, merging and diverging relationships, and road connectivity to obtain environmental constraint information containing lane relationships and traffic rules; based on this constraint information, it further extracts reference lines for vehicle travel, such as lane center lines, and performs interpolation and smoothing processing on them; finally, it integrates the above information to generate driving environment identification information.
[0079] Using the above method, based on road feature information, obstacle fusion information, and driving environment reference information, local mapping and topology extraction of the driving environment are performed. Through these processes, a precise geometric map of the current driving environment can be constructed, and its road topology can be understood. This allows for the acquisition of environmental constraints and driving reference lines, enabling the vehicle to accurately perceive key environmental elements such as lanes and intersections. This provides clear environmental boundaries for obstacle prediction and autonomous vehicle decision-making, thereby improving the rationality of path planning and the smoothness of the driving trajectory, ensuring the safety of assisted driving. This ultimately enhances the safety and reliability of assisted driving.
[0080] In one possible implementation of the above method, the prediction post-processing module obtains the corresponding obstacle prediction trajectory information based on obstacle post-fusion information, driving environment identification information, and prediction post-processing reference information. This includes: based on a preset matching strategy, obtaining the corresponding model intent and rule intent based on obstacle post-fusion information, driving environment identification information, and prediction post-processing reference information, and performing matching processing on the model intent and rule intent to determine whether the model intent and rule intent match successfully; if the match is successful, smoothing and extending the initial obstacle prediction trajectory information corresponding to the obstacle post-fusion information is performed to generate obstacle prediction trajectory information; if the match is unsuccessful, obstacle prediction trajectory information is generated based on a kinematic model.
[0081] The processing procedure of the prediction post-processing module can be as follows: First, based on a preset matching strategy, the prediction post-processing module comprehensively utilizes obstacle fusion information such as the obstacle's position, speed, category, and motion state; driving environment identification information such as lane topology, traffic rules, and road boundaries; and prediction post-processing reference information such as the vehicle's positioning and motion state. Through parallel reasoning, it generates model intent and rule intent respectively. The model intent is generated by a data-driven behavior prediction model, reflecting the statistical behavior trend of obstacles in historical trajectories and contextual scenarios. The rule intent is constructed based on traffic rules, road geometric constraints, and scene common sense, reflecting the normative behaviors that obstacles should follow in specific environments, such as driving within lanes and yielding at intersections. Subsequently, the system performs consistency matching and confidence assessment on the two intentions: if the model intention and the rule intention are consistent in key spatiotemporal constraints, such as both predicting that the obstacle will keep going straight in the lane, the match is considered successful. At this time, the initial predicted trajectory of the obstacle is smoothed and temporally extended to generate a final predicted trajectory that conforms to environmental constraints and is continuous and smooth; if there is a significant conflict between the two, such as the model intention predicting that the obstacle will cross the solid line to change lanes, while the rule intention determines that the behavior is illegal, the match is considered to have failed. The system triggers a safety redundancy mechanism and degenerates into generating a conservative predicted trajectory that conforms to physical laws based on the kinematic model.
[0082] Using the above method, based on a preset matching strategy, the corresponding model intent and rule intent are obtained according to obstacle post-fusion information, driving environment identification information, and prediction post-processing reference information. By matching the model intent with the rule intent, while fusing data to drive prediction, prior knowledge such as traffic rules is introduced for constraint verification. When a match is successful, the trajectory is smoothly extended to ensure the continuity and rationality of the prediction; when a match fails, a conservative kinematic model is used as a safe alternative to ensure the basic reliability of the prediction. This scheme improves the accuracy of obstacle behavior prediction and effectively prevents potential collision risks, thereby enhancing the safety and reliability of assisted driving.
[0083] In one possible implementation of the above method, the decision-making and planning module includes a decision-making submodule and a planning submodule. The decision-making and planning module obtains corresponding driving control information based on the vehicle's planned trajectory information, obstacle prediction trajectory information, driving environment identification information, and decision-making and planning reference information. This includes: the decision-making submodule obtaining corresponding driving behavior decision information as first driving control information based on obstacle prediction trajectory information, driving environment identification information, and decision-making and planning reference information, and outputting the driving behavior decision information to the planning submodule; and the planning submodule obtaining corresponding initial vehicle control trajectory information based on the vehicle's planned trajectory information, driving environment identification information, and decision-making and planning reference information, and dynamically correcting the initial vehicle control trajectory information based on obstacle prediction trajectory information and driving behavior decision information to obtain corresponding target vehicle control trajectory information as second driving control information.
[0084] Among them, driving behavior decision information includes lane keeping, lane changing, ramp merging and merging, and intersection turning decisions.
[0085] Furthermore, the processing procedure of the decision planning module can be as follows: The decision submodule of the decision planning module first uses obstacle prediction trajectory information, driving environment identification information, and decision planning reference information to analyze the current scene through behavioral decision algorithms, such as rule-based logical judgment, to generate driving behavior decision information as the first driving control information, such as instructions to maintain the current lane cruise, change lanes to the left to overtake, prepare to exit the ramp or stop at a red light, etc., and outputs the generated driving behavior decision information to the planning submodule; The planning submodule then generates an initial vehicle control trajectory that conforms to the static boundary based on the driving behavior decision information, the vehicle's planned trajectory information, driving environment identification information, and decision planning reference information, and then dynamically corrects and locally optimizes the initial trajectory by combining the real-time updated obstacle prediction trajectory information, and finally outputs the target vehicle control trajectory that takes into account safety, comfort, and rule constraints as the second driving control information for real-time vehicle control.
[0086] Using the above method, the decision-making and planning module is divided into a decision-making sub-module and a planning sub-module. The decision-making sub-module generates corresponding driving behavior decision information, i.e., the first driving control information, based on obstacle prediction information, driving environment identification information, and decision-making and planning reference information, providing guidance for decision-making and planning. The planning sub-module generates corresponding initial vehicle control trajectory information based on the vehicle's planned trajectory information, driving environment identification information, and decision-making and planning reference information, and dynamically corrects it based on obstacle prediction trajectory information and driving behavior decision information to generate the final target vehicle control trajectory, i.e., the second driving control information. This hierarchical architecture enables decision-making to consider traffic rules and scenario constraints globally, and planning to flexibly respond to dynamic obstacles. It ensures the rationality and safety of driving strategies while improving the real-time performance and accuracy of trajectory planning, thereby effectively improving the reliability and safety of assisted driving.
[0087] The following example illustrates the vehicle control method provided in this application, which utilizes road feature identification information, obstacle perception information, and vehicle planning trajectory information output by an end-to-end large model based on a bird's-eye-view transformer (Bev Transformer). This method combines a new environmental model framework for roads with a spatiotemporal joint decision-making and planning system to achieve real-time vehicle control.
[0088] like Figure 4 As shown, the vehicle control method includes the following steps.
[0089] Step S101: The end-to-end large model encodes driving perception information into high-dimensional features, and performs feature interaction and fusion based on a multimodal architecture, including using a gating network to dynamically allocate expert network weights according to environmental constraints such as lane line type, and finally decodes and outputs road feature identification information, obstacle perception information and vehicle planning trajectory information in parallel, and outputs them to the environment model, obstacle post-fusion module and decision planning module respectively.
[0090] The end-to-end large model differs from the traditional modular model by using a one-model to realize the processing of lane line perception, obstacle prediction trajectory, and vehicle planning trajectory output.
[0091] The inputs to an end-to-end large model may include, for example, BEV features, ego motion information such as speed, acceleration, and steering wheel angle, environmental information such as traffic lights / stop lines, navigation information, and initialization information such as initialization information for spatial paths, initialization information for temporal trajectories, initialization information for perception outputs, and initialization information for text descriptions.
[0092] The end-to-end large model processing can include, for example, the following: The end-to-end large model uses a multimodal (MoE) architecture to encode raw information such as the aforementioned images, navigation, environmental information, vehicle status, and other relevant driving perception information into high-dimensional features. These high-dimensional features interact fully within the Transformer-based end-to-end large model. This includes a gate network that, based on perceived lane lines, turn arrows, and other environmental information, prohibits invalid modes in conjunction with lane line type, and assigns weights to the remaining modes using a normalized exponential function (softmax). Finally, the various expert networks are combined, including lane keeping, left lane changing, right lane changing, overtaking, and detour processing, to obtain trajectory output, perception output, and process description, which serve as perception result information for downstream use.
[0093] Furthermore, a multimodal MoE architecture is adopted. In short, lane-type information (such as solid line, dashed line, solid left and dashed right, dashed left and solid right, etc.) is used as a key input to the gate network. The gate network then distributes tasks to several corresponding expert networks based on the scenario. Each expert network corresponds to a high-level policy distribution, meaning each expert network focuses on one "modality," such as lane keeping, merging to the left, merging to the right, and emergency avoidance. For example, it could be Expert0 representing lane keeping, Expert1 representing left lane change, Expert2 representing right lane change, and Expert3 representing overtaking / obstacle avoidance, among other expert networks with different modalities. The gating network takes into account environmental information, including lane marking features such as lane lines, stop lines, zebra crossings, turn arrows, and lane signs, as well as the status of surrounding obstacles. First, invalid modes are prohibited using a lane marking mask (e.g., if the right side is a solid line, the gate weight for "right lane change expert" is set to 0). Then, a softmax operation is performed on the remaining modes (i.e., valid modes) to assign weights and obtain corresponding gating weights. Finally, the weights are fused, and the outputs of each expert network are combined to obtain the policy and trajectory. During training, constraints, penalties, and balance terms ensure expert specialization and driving compliance, allowing the end-to-end large model to autonomously select the optimal policy within the allowed modes. The final action distribution can be, for example, action distribution = Σ(gate_i * expert_i_output).
[0094] The output of the end-to-end large model may include, for example, road feature identification information (such as lane lines), obstacle perception information, and vehicle planning trajectory information (such as the aforementioned vehicle path (a spatial path with a length of 50m) and vehicle trajectory (a temporal trajectory with a duration of 8s), etc. The lane lines are output to the environment model, the obstacle perception information is output to the obstacle fusion module, and the vehicle planning trajectory information is output to the decision planning module.
[0095] Furthermore, such as Figure 5As shown, the end-to-end large model processing process can be as follows: The end-to-end large model receives input information such as images, navigation, and CAN-bus. The image is a token sequence processed by its corresponding tokenization module, visual-language encoder (VL encoder), and a combination of the VL encoder and NM vision model. The navigation and CAN-bus information are processed by their respective tokenization modules, resulting in a token sequence that the end-to-end large model can process. Simultaneously, the end-to-end large model also receives initialization information such as text descriptions, perception outputs, temporal trajectories, and spatial paths. All these inputs are fed into the end-to-end large model, such as a large multimodal model, for unified processing. Correspondingly, the output includes perception results such as spatial paths, temporal trajectories, perception outputs, and text descriptions, which support the vehicle's driving decisions and environmental perception.
[0096] Step S102: The obstacle post-fusion module receives the obstacle perception information output from the end-to-end large model and fuses it with the obstacle post-fusion reference information to obtain obstacle post-fusion information, which is then output to the environment model and the obstacle prediction post-processing module respectively.
[0097] Step S103: The environmental model integrates road feature information, obstacle post-fusion information, and driving environment reference information, and performs local mapping, topology extraction, lane sequence filling, and reference line generation in sequence. Finally, it outputs driving environment identification information containing local lane line scatter points and reference line scatter points, which are output to the obstacle prediction post-processing module and the decision planning module, respectively.
[0098] Among them, the environmental model is an intermediate link between perception, decision-making, and planning. It aims to construct the road topology within the visible range around the vehicle and use it as a static boundary input to the downstream decision-making and planning (Planner) module.
[0099] The inputs to the environmental model may include, for example, road feature information such as lane lines, obstacle fusion information such as OD obstacle fusion information, driving environment reference information such as navigation information, positioning information (such as odometer (odom), GPS dead reckoning (gps-dr) positioning, etc.), traffic light information (TLI), traffic sign recognition information (TSR), general obstacle information (Occupancy grid (OCC), etc.), and other information.
[0100] The environmental model processing can include, for example, the following steps: First, environmental modeling is completed through local mapping and topology extraction. Topology extraction includes flow control zone calculation, lane grouping, road segmentation, processing of extra-wide lanes and single-sided lanes, merging and branching traffic, and centerline generation. Then, lane sequences are filled to generate connected sequences and their relationships. Further, the centerlines of the lane sequences are interpolated and smoothed to obtain reference lines. Subsequently, static navigation scores are calculated by combining standard definition navigation information and lane connectivity relationships. Based on the scores, the direction of travel is determined, and the lane's direction of travel is linked to traffic lights and the reference line. Finally, speed limits are implemented by integrating navigation information and Traffic Sign Recognition (TSR) information, resulting in a complete lane topology and reference line output.
[0101] The output of the environment model may include, for example, driving environment identification information, i.e., environment information (Env Info), such as lane line scatter points and reference line scatter points after local mapping. This driving environment identification information is then output to the obstacle prediction post-processing module and the decision planning module, respectively.
[0102] Furthermore, such as Figure 6 The diagram illustrates a principle for lane sequence filling. In the diagram, "vehicle" (ego) represents the vehicle itself, and R1-1, R2-1, R2-1', etc., represent lanes or road nodes. Solid lines represent the main reference lines that are already determined and available for vehicle travel within the current moving window; these are the core paths planned. Dashed lines represent alternative reference lines or historical information generated within the moving window, used for path selection in scenarios such as lane changing and obstacle avoidance. The environment model only calculates the road or lane nodes within the moving window range in front of the vehicle, generating the corresponding reference lines.
[0103] like Figure 7The diagram illustrates a principle for smoothing lane sequence centerline interpolation. Three independent reference lines (ref line 1, ref line 2, and ref line 3) are represented by lines of different colors. These reference lines represent multiple driving paths available to the vehicle on the road ahead, such as different lanes, turning routes, or obstacle avoidance alternatives. The environment model selects the optimal reference line from these pre-generated reference lines, or smoothly switches between multiple lines, based on real-time road conditions, navigation instructions, and obstacle information, to achieve safe and comfortable driving decisions.
[0104] Furthermore, such as Figure 8 As shown, the overall framework of the environment model mainly includes seven processing modules: local mapping module, topology extraction module, lane sequence filling, reference line generation, static navigation calculation, traffic light binding reference line, and speed limit.
[0105] Local mapping is used to achieve mapping.
[0106] Topology extraction is used to achieve precise topology, mainly including six processes: guide zone calculation, lane grouping, road section generation, wide lane and single-sided lane processing, merging and branching processing, and centerline generation. Specifically, guide zone calculation mainly involves determining the location, shape, and direction of the guide zone based on SD navigation information and the perceived lane line alignment; lane grouping mainly involves determining road merging and branching based on the curb and guide zone, thereby grouping lane lines; road section segmentation mainly involves dividing lane lines into segments and topology from lane grouping and lane addition / reduction; wide lane and single-sided lane processing mainly involves finding wide lanes and single-sided lanes based on the basic segmentation results and refining the topology; merging and branching processing mainly involves establishing virtual lanes based on the perceived merging and branching points and the relationship between lane lines; and centerline generation mainly involves averaging the left and right side lines.
[0107] Lane sequence filling refers to generating connected sequences and interrelationships based on topological relationships. That is, after topology extraction, by integrating information such as perceived lane lines, guide zones, and curbs, the sequential connections between lanes are established, generating a continuous lane topology structure that can be used for navigation and path planning. The construction of lane sequences provides a structured road network foundation for subsequent reference line generation, navigation path selection, and behavioral decision-making.
[0108] Reference line generation refers to the interpolation and smoothing of the centerline in the Lane Sequence. Specifically, after the lane sequence is constructed, the lane centerline is densely sampled and then smoothed using spline interpolation or filtering algorithms to eliminate perceptual noise and local jitter, ensuring the geometric continuity and reasonable curvature changes of the reference line. The generated reference line serves as a static boundary input to the downstream decision-making and planning module.
[0109] Static navigation score calculation refers to calculating a score by combining standard definition (SD) navigation information and lane topology, and determining the direction of travel based on the score. In other words, it calculates the score for each lane sequence based on navigation instructions (such as straight, left turn, right turn, etc.) and lane topology, and the score determines the lane direction to be prioritized.
[0110] Traffic light reference lines refer to the binding of traffic light information to the corresponding reference lines based on the direction of traffic flow (e.g., left turn, straight ahead, right turn) and the traffic lights in each direction. This provides traffic signal constraints for downstream decision-making and planning modules.
[0111] Speed limit refers to the process of integrating static road speed limit information from navigation maps and dynamic speed limit signs detected by the Traffic Sign Recognition (TSR) system to generate the final speed limit constraint for the current driving scenario. This involves extracting the legal speed limits for each lane or road segment provided by the navigation system, while simultaneously receiving real-time speed limit indicators (such as temporary construction speed limits, variable speed limit signs, etc.) from the TSR output, and fusing the two types of information through spatiotemporal alignment and confidence assessment. During the fusion process, the system performs reasonableness verification and smoothing of the speed limit based on the road environment and vehicle status. The generated speed limit value or speed curve is then bound to the corresponding reference line or lane sequence, serving as the speed constraint input for the downstream decision-making and planning module, ensuring that the vehicle achieves safe and comfortable autonomous driving while adhering to traffic rules.
[0112] Therefore, as Figure 8As shown, the processing of the environment model module can be as follows: GPS dead reckoning (GPS-DR) positioning, odometry (ODM), standard definition (SD) navigation, shape points and instructions, lane lines, stop lines, zebra crossings, road arrows, merging and branching points, etc., are localized and then subjected to precise topology (Topo exact), which involves sequentially performing guide zone identification, lane line grouping, road section generation, wide lane and single-sided lane processing, merging and branching processing, and centerline generation, etc., to output scatter points. The scatter points are then filled with lane sequences to output a lane sequence list. After static navigation sub-processing, the position of the target lane on the local map is output, and the navigation sub-processing is smoothed to generate a reference line (Ref Line), or the lane sequence list is directly smoothed to generate a reference line. At the same time, obstacle back fusion is processed by traffic flow and then smoothed to generate a reference line. Traffic lights and traffic signs undergo traffic light recognition (including countdown) processing, navigation perception matching based on standard definition navigation traffic lights, traffic light binding based on generated reference lines, and inference of traffic lights with environmental information. The output reference lines are then smoothed to generate reference lines with traffic light status (Ref line with TL status), which are then smoothed further. General obstacle information is also smoothed to generate reference lines. Furthermore, areas where traffic participants may be present undergo occlusion scene recognition processing, and then combined with traffic speed limit signs from traffic lights and traffic signs, as well as navigation speed limits from standard definition navigation, to complete scattered speed limits for reference lines (Ref Line), which are then smoothed to generate reference lines. The smoothed reference lines generated above are the final output reference lines. Finally, the generated reference lines are combined with the vehicle's drivable area to output environmental information.
[0113] Step S104: The obstacle prediction post-processing module, based on the obstacle post-fusion information, driving environment identification information, and prediction post-processing reference information, determines the obstacle prediction trajectory by matching the model intent with the rule intent. If the match is successful, the initial trajectory of the obstacle is smoothly extended to generate the predicted trajectory; otherwise, the predicted trajectory is generated based on the kinematic model. Finally, the obstacle prediction trajectory information is output to the decision planning module.
[0114] The inputs to the obstacle prediction post-processing module may include, for example, obstacle post-fusion information such as OD obstacle post-fusion information, driving environment identification information such as environmental information, and prediction post-processing reference information such as positioning information.
[0115] The obstacle prediction post-processing module's process may include, for example, obtaining the model intent and rule intent based on a preset matching strategy and performing a match; if the match is successful, smoothing and extending the initial obstacle prediction trajectory information to generate the obstacle prediction trajectory, such as an 8-second obstacle prediction trajectory; if the match is unsuccessful, generating the obstacle prediction trajectory based on the kinematic model. In other words, the core of obstacle prediction post-processing lies in matching the model intent and rule intent. If the match is successful, the OD prediction trajectory output by the end-to-end large model is smoothed and extended to generate the final 8-second trajectory; if the match is unsuccessful, an 8-second trajectory is generated based on the kinematic model and output as dynamic obstacle interaction information to the downstream Planner (decision-making, planning) module.
[0116] The output of the obstacle prediction post-processing module may include, for example, obstacle prediction trajectory information, which is then output to the decision planning module.
[0117] Furthermore, such as Figure 9 As shown, the obstacle prediction post-processing module's processing procedure can be as follows: The obstacle prediction post-processing module takes vehicle localization information, OD obstacle post-fusion information, environmental model information, and a prediction model / NN driver model as inputs. The OD obstacle post-fusion information is output by the obstacle post-fusion module based on the OD predicted trajectory, and the OD predicted trajectory is output from the end-to-end large model and input into the obstacle post-fusion module. This process includes: inputting the vehicle localization information into the prediction model / NN driver model for processing, and obtaining the historical trajectory by accumulating the target information coordinate transformation and trajectory (tracking) of the vehicle localization information and the OD obstacle post-fusion information; and then comparing the historical trajectory with the predicted model / NN driver model... The driver model processes the vehicle's localization information and combines it with the predicted trajectory to generate a merged trajectory through modal merging. The merged trajectory undergoes a curvature orientation feasibility check to generate a curvature-feasible trajectory. Subsequently, the traffic flow wall generation module outputs the traffic flow wall for curb / traffic flow wall interception checks, generating a post-interception trajectory. The post-interception trajectory undergoes model-predicted trajectory intent recognition to generate a model intent. Simultaneously, environmental model information is combined with historical trajectories to complete rule-based target lateral / longitudinal intent recognition, outputting a rule intent. Next, the model intent and rule intent are filtered / matched. If a match is successful, a smooth trajectory is generated through lane adsorption (trajectory correction or generation), trajectory correction, and optimization and smoothing of the predicted trajectory. This smooth trajectory is then extended to 8 seconds using a kinematic model. If a match fails, an 8-second trajectory is directly generated by the kinematic model. Finally, both types of trajectories are output together as the final obstacle prediction trajectory for use by the decision planning module (Planner).
[0118] Step S105: The decision planning module generates driving behavior decisions through the decision submodule based on information such as the self-planned trajectory information, obstacle prediction trajectory information, driving environment identification information, and decision planning reference information. The planning submodule then integrates the self-planned trajectory with the static reference line based on the decision, and makes dynamic corrections by combining the obstacle prediction trajectory. Finally, it outputs the target vehicle control trajectory as driving control information.
[0119] The inputs to the decision planning module may include, for example, vehicle trajectory planning information such as path and trajectory, obstacle prediction trajectory information, driving environment identification information such as environmental information, and decision planning reference information such as vehicle dynamics parameters and traffic rules.
[0120] The decision-making and planning module's processing may include, for example, the decision submodule making driving behavior decisions based on obstacle prediction trajectory information, driving environment signage information, and decision-making and planning reference information, such as lane keeping, lane changing, ramp merging and merging, intersection turning, etc., obtaining driving behavior decision information as the first driving control information, and outputting it to the planning submodule. The planning submodule, based on the vehicle's planned trajectory information, driving environment signage information, and decision-making and planning reference information, merges the reference line with the vehicle's planned trajectory, outputs initial vehicle control trajectory information based on static boundaries, and dynamically corrects the initial vehicle control trajectory information based on obstacle prediction trajectory information and driving behavior decision information. When there is dynamic obstacle interaction, obstacle prediction trajectory correction is considered to obtain target vehicle control trajectory information as the second driving control information.
[0121] The output of the decision planning module may include, for example, the final driving control information, such as the vehicle trajectory, for vehicle driving control.
[0122] Furthermore, such as Figure 10As shown, the corresponding vehicle control method based on the vehicle control system may include, for example, the output of road feature identification information such as lane lines, stop lines, zebra crossings, road arrows, and merging / branching points from an end-to-end large model, as well as the vehicle's planned trajectory and obstacle perception information. The positioning information (i.e., GPS dead reckoning (GPS-DR) positioning, odometer (ODM)), standard definition (SD) navigation shape points and instructions, common obstacles, traffic lights & traffic signs, and obstacle perception information, after being processed by obstacle post-fusion, are all input into the environment model. The environment model processes the various input information to generate environmental information, which is then input into the obstacle prediction post-processing module and the decision planning module. The obstacle prediction post-processing module generates the OD prediction trajectory based on the obstacle post-fusion information and the environmental information, and inputs the OD prediction trajectory into the decision planning module. The decision planning module receives environmental information output by the environmental model, OD prediction trajectory output by the obstacle prediction post-processing module, vehicle planning trajectory (path, trajectory), positioning information, and vehicle body and chassis signals from the end-to-end large model. It combines the above information and processes it through the decision and planning sub-modules to determine the driving policy. Then, the decision planning module generates the planned trajectory and target speed information to ultimately achieve vehicle control.
[0123] Using the above method, the end-to-end large model accurately outputs road feature identification information, obstacle perception information, and vehicle planning trajectory information; the environment model generates high-precision reference lines and other static boundary information through local mapping and topology extraction; the obstacle post-fusion module deeply fuses visual and radar obstacles to eliminate single sensor errors; the obstacle prediction post-processing module generates reliable obstacle prediction trajectories by matching model intent with rule intent; and the decision planning module combines vehicle planning trajectory information, obstacle prediction trajectory information, driving environment identification information, and decision planning reference information to generate driving behavior decisions in layers and dynamically correct the vehicle control trajectory. In other words, when the aforementioned image information is Bev feature information, the vehicle control method provided in this application can be understood as a method that uses lane lines, obstacle prediction trajectories, and vehicle planning trajectories output by an end-to-end (E2E) model based on Bev Transformer, combined with a new road environment model framework and a spatiotemporal joint decision planning system to achieve real-time vehicle control. The road environment model is built locally based on visual perception results, and topology extraction is performed. The resulting reference line is then provided to the downstream decision-making and planning module as static boundary input. Simultaneously, the obstacle prediction and post-processing module fuses visual and radar obstacle data and corrects the predicted trajectory, providing this data as dynamic obstacle interaction input to the downstream decision-making and planning module. Finally, the spatiotemporal joint decision-making and planning system generates a vehicle control trajectory based on the static boundary reference line, the E2E vehicle planning trajectory, and the dynamic obstacle interaction. This enables safe and stable operation of driving functions such as lane keeping, lane changing, ramp merging, and intersection turning. The solution is adaptable to various complex scenarios, including highway ramps, urban intersections, roundabouts, and construction sites, and possesses mass-production-grade performance. This improves the reliability and safety of assisted driving.
[0124] This application also discloses an electronic device, including the aforementioned vehicle control system, for implementing the aforementioned vehicle control method. This electronic device can be a vehicle, or a server or other device corresponding to the vehicle.
[0125] It should be noted that, in addition to the specific embodiments described above, those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Although the description of this application is presented in conjunction with preferred embodiments, this does not mean that the features of this application are limited to these embodiments. On the contrary, the purpose of describing the application in conjunction with the embodiments is to cover other options or modifications that may be derived based on the claims of this application. To provide a thorough understanding of this application, many specific details are included in the above description, and this application may also be implemented without using these details. Furthermore, to avoid confusion or obscuring the focus of this application, some specific details will be omitted in the description. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.
[0126] It should be noted that in this specification, similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0127] The terms “first”, “second”, etc., are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0128] Although this application has been illustrated and described with reference to certain preferred embodiments, those skilled in the art should understand that the above description is a further detailed explanation of the application in conjunction with specific embodiments, and should not be construed as limiting the specific implementation of the application to these descriptions. Those skilled in the art can make various changes in form and detail, including some simple deductions or substitutions, without departing from the spirit and scope of this application.
Claims
1. A vehicle control method, characterized in that, A vehicle control system applied to assist driving control of a vehicle, the vehicle control system including an end-to-end model module, an environment model module, an obstacle post-fusion module, an obstacle prediction post-processing module, and a decision planning module, the method including: The end-to-end model module obtains corresponding perception result information based on the driving perception information corresponding to the vehicle input into the end-to-end model module. The perception result information includes road feature identification information corresponding to the driving environment in which the vehicle is located, obstacle perception information corresponding to the obstacles around the vehicle, and the vehicle's self-planned trajectory information. The road feature identification information is output to the environment model module, the obstacle perception information is output to the obstacle fusion module, and the self-planned trajectory information is output to the decision planning module. The obstacle post-fusion module obtains corresponding obstacle post-fusion information based on the obstacle perception information and the obstacle post-fusion reference information, and outputs the obstacle post-fusion information to the environment model module and the obstacle prediction post-processing module respectively. The environment model module obtains corresponding driving environment identification information based on the road feature information, the obstacle post-fusion information, and the driving environment reference information, and outputs the driving environment identification information to the obstacle prediction post-processing module and the decision planning module respectively. The obstacle prediction post-processing module obtains the corresponding obstacle prediction trajectory information based on the obstacle post-fusion information, the driving environment identification information, and the prediction post-processing reference information, and outputs the obstacle prediction trajectory information to the decision planning module. The decision planning module obtains corresponding driving control information based on the vehicle's planned trajectory information, the obstacle predicted trajectory information, the driving environment identification information, and the decision planning reference information, for use in vehicle driving control.
2. The method according to claim 1, characterized in that, The vehicle control system further includes a pre-processing module, and the method further includes: The pre-processing module performs feature extraction processing based on the original driving-related information input to the pre-processing module to obtain the driving perception information corresponding to the original driving-related information, and outputs the driving perception information to the end-to-end model module.
3. The method according to claim 2, characterized in that, The end-to-end model module includes an end-to-end model, which is a multimodal large language model. The end-to-end model module obtains corresponding perception result information based on the driving perception information corresponding to the vehicle input into the end-to-end model module, including: Based on the driving perception information, multimodal feature extraction and feature fusion processing are performed to obtain the corresponding perception result information.
4. The method according to claim 3, characterized in that, The obstacle perception information includes information about visually perceived obstacles, and the obstacle post-fusion reference information includes information about obstacles to be fused, including non-visually perceived obstacles. The obstacle post-fusion module obtains corresponding obstacle post-fusion information based on the obstacle perception information and the obstacle post-fusion reference information, including: Based on the obstacle perception information and the obstacle post-fusion reference information, the visually perceived obstacle and the obstacle to be fused are fused to obtain the obstacle post-fusion information.
5. The method according to claim 4, characterized in that, The environment model module obtains corresponding driving environment identification information based on the road feature information, the obstacle post-fusion information, and the driving environment reference information, including: Based on the road feature information, the obstacle fusion information and the driving environment reference information, local mapping and topology extraction of the driving environment are performed to obtain the environmental constraint information corresponding to the driving environment of the vehicle. Based on the environmental constraint information, the driving reference line information corresponding to the vehicle is obtained; The driving environment identification information is obtained based on the driving reference line information.
6. The method according to claim 5, characterized in that, The prediction post-processing module obtains the corresponding obstacle prediction trajectory information based on the obstacle post-fusion information, the driving environment identification information, and the prediction post-processing reference information, including: Based on a preset matching strategy, the corresponding model intent and rule intent are obtained according to the obstacle fusion information, the driving environment identification information and the prediction post-processing reference information. The model intent and the rule intent are then matched to determine whether the model intent and the rule intent are successfully matched. If the match is successful, the initial predicted trajectory information of the obstacle corresponding to the post-fusion information of the obstacle is smoothed and extended to generate the predicted trajectory information of the obstacle. If a match is not found, the predicted trajectory information of the obstacle is generated based on the kinematic model.
7. The method according to claim 6, characterized in that, The decision-making and planning module includes a decision-making submodule and a planning submodule. The decision-making and planning module obtains corresponding driving control information based on the vehicle's planned trajectory information, the obstacle prediction trajectory information, the driving environment identification information, and the decision-making and planning reference information, including: The decision submodule obtains corresponding driving behavior decision information as first driving control information based on the obstacle prediction trajectory information, the driving environment identification information, and the decision planning reference information, and outputs the driving behavior decision information to the planning submodule. The planning submodule obtains the corresponding initial vehicle control trajectory information based on the self-planned trajectory information, the driving environment identification information, and the decision planning reference information. It then dynamically corrects the initial vehicle control trajectory information based on the obstacle prediction trajectory information and the driving behavior decision information to obtain the corresponding target vehicle control trajectory information as the second driving control information.
8. The method according to claim 7, characterized in that, The driving perception information includes: the driving status information of the vehicle, the driving environment information corresponding to the driving environment in which the vehicle is located, and the navigation information and positioning information corresponding to the vehicle. The obstacle to be fused is a radar obstacle; The obstacle fusion information includes: fused obstacle information; The driving environment reference information includes: traffic safety constraint information, general obstacle information, navigation information, and positioning information corresponding to the driving environment in which the vehicle is located; The driving environment identification information includes a local driving environment map, a driving environment topology, lane information of the vehicle in the local driving environment map, the location navigation information of the target lane in the local driving environment map, traffic flow information, traffic scene safety constraint information, obstacle information, and driving speed constraint information. The prediction post-processing reference information includes: the positioning information corresponding to the vehicle; The decision-making and planning reference information includes: the vehicle's corresponding dynamics constraints, traffic rule constraints, driving mode preference information, speed limits and navigation guidance information, and vehicle body parameters and load status information; The driving behavior decision information includes lane keeping, lane changing, ramp merging and merging, and intersection turning decisions.
9. A vehicle control system, characterized in that, The system includes an end-to-end model module, an environment model module, an obstacle post-fusion module, an obstacle prediction post-processing module, and a decision planning module, wherein: The end-to-end model module is used to obtain corresponding perception result information based on the driving perception information corresponding to the vehicle input to the end-to-end model module. The perception result information includes road feature identification information corresponding to the driving environment in which the vehicle is located, obstacle perception information corresponding to the obstacles around the vehicle, and the vehicle's self-planned trajectory information. The road feature identification information is output to the environment model module, the obstacle perception information is output to the obstacle fusion module, and the self-planned trajectory information is output to the decision planning module. The obstacle post-fusion module is used to obtain corresponding obstacle post-fusion information based on the obstacle perception information and the obstacle post-fusion reference information, and output the obstacle post-fusion information to the environment model module and the obstacle prediction post-processing module respectively. The environment model module is used to obtain corresponding driving environment identification information based on the road feature information, the obstacle post-fusion information, and the driving environment reference information, and output the driving environment identification information to the obstacle prediction post-processing module and the decision planning module respectively. The obstacle prediction post-processing module is used to obtain the corresponding obstacle prediction trajectory information based on the obstacle post-fusion information, the driving environment identification information, and the prediction post-processing reference information, and output the obstacle prediction trajectory information to the decision planning module. The decision planning module is used to obtain corresponding driving control information based on the vehicle's planned trajectory information, the obstacle predicted trajectory information, the driving environment identification information, and the decision planning reference information, for use in vehicle driving control.
10. An electronic device, characterized in that, Includes the vehicle control system as described in claim 9, used to implement the vehicle control method as described in any one of claims 1-8.