Path planning method and device based on multimodal data fusion and ai inference collaboration
By employing a path planning method that combines multimodal data fusion with AI inference, and utilizing hierarchical thinking trees to generate and evaluate multiple candidate paths, the system addresses the problem of insufficient path planning in complex dynamic scenarios for intelligent driving systems, thereby improving the system's adaptability and safety in complex traffic situations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-03
- Publication Date
- 2026-04-14
AI Technical Summary
Existing intelligent driving systems perform poorly in path planning in complex and dynamic scenarios, making it difficult to generate multiple candidate decision paths and perform dynamic evaluation.
A path planning method that combines multimodal data fusion with AI reasoning is adopted. By acquiring multimodal data, candidate actions in the current thinking state, and reasoning levels, the method predicts the next thinking state and evaluation results, generates multiple candidate thought paths, and uses hierarchical thinking trees for path reasoning and evaluation.
It improves the adaptability of intelligent driving systems in complex dynamic scenarios, generates multiple path options, and enhances the safety and adaptability of the system in complex traffic scenarios.
Smart Images

Figure CN121632198B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent driving technology, and in particular to a path planning method and apparatus that combines multimodal data fusion and AI reasoning. Background Technology
[0002] In scenarios such as urban open roads, highways, and complex intersections, vehicles need to perceive the surrounding environment in real time through multimodal sensors (such as cameras, lidar, and millimeter-wave radar) to complete environmental modeling, target detection, path planning, and decision control.
[0003] Currently, intelligent driving systems typically employ linear sequential reasoning when planning routes.
[0004] However, this linear sequential reasoning of existing technologies performs poorly in complex dynamic scenarios. Summary of the Invention
[0005] This application provides a path planning method and apparatus that combines multimodal data fusion and AI inference to improve the adaptability of intelligent driving systems in complex dynamic scenarios.
[0006] In a first aspect, embodiments of this application provide a path planning method for multimodal data fusion and AI inference collaboration, including:
[0007] The system acquires multimodal data during vehicle driving, at least two candidate actions in the current thought state, and the reasoning level corresponding to each candidate action. The current thought state is determined based on the driving situation of the vehicle at the current moment, and the reasoning level is used to characterize the amount of computing resources used in path reasoning.
[0008] Based on the multimodal data, at least two candidate actions in the current thinking state, and the reasoning level corresponding to each candidate action, predict the next thinking state and the evaluation result of the next thinking state after the vehicle executes each candidate action.
[0009] Based on the evaluation results of the candidate actions and each next thought state, at least two candidate thought paths are generated.
[0010] Secondly, embodiments of this application provide a path planning device for multimodal data fusion and AI inference collaboration, comprising:
[0011] The data acquisition module is used to acquire multimodal data during vehicle driving, at least two candidate actions in the current thinking state, and the reasoning level corresponding to each candidate action. The current thinking state is determined based on the driving situation of the vehicle at the current moment, and the reasoning level is used to characterize the amount of computing resources used in path reasoning.
[0012] The state reasoning module is used to predict the next state of thought and the evaluation result of the next state of thought after the vehicle executes each candidate action, based on the multimodal data, at least two candidate actions in the current thinking state, and the reasoning level corresponding to each candidate action.
[0013] The path generation module is used to generate at least two candidate thought paths based on the candidate actions and the evaluation results of each next thought state.
[0014] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0015] Fourthly, embodiments of this application provide a vehicle, including a vehicle body and the aforementioned electronic equipment disposed in the vehicle body.
[0016] Fifthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0017] Sixthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0018] The multimodal data fusion and AI inference collaborative path planning method and apparatus provided in this application embodiment can generate paths corresponding to the candidate action at different inference levels by configuring a corresponding inference level for each candidate action, thereby realizing the generation of multiple paths. This can adapt to the needs of highly dynamic traffic scenarios and improve the adaptability in complex dynamic scenarios. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0020] Figure 1 A schematic diagram of the path planning method for multimodal data fusion and AI inference collaboration provided in this application;
[0021] Figure 2 This is a schematic diagram illustrating the generation of multimodal data provided in an embodiment of this application;
[0022] Figure 3This is a schematic diagram of the multimodal data fusion process provided in an embodiment of this application;
[0023] Figure 4 This is a schematic diagram of the millimeter-wave radar enhanced feature generation process provided in the embodiments of this application;
[0024] Figure 5 This is a schematic diagram illustrating the process of converting a bird's-eye view feature map into a word sequence, as provided in an embodiment of this application.
[0025] Figure 6 This is a schematic diagram of the word sequence generation process provided in the embodiments of this application;
[0026] Figure 7 This is a schematic diagram of the path reasoning generation process provided in an embodiment of this application;
[0027] Figure 8 This is a schematic diagram of the intelligent driving decision-making process in an embodiment of this application;
[0028] Figure 9 A flowchart illustrating the path explanation and tracing process provided in this application's embodiments;
[0029] Figure 10 This is an overall flowchart of path reasoning provided in the embodiments of this application;
[0030] Figure 11 This is a schematic diagram of the path reasoning process provided in the embodiments of this application;
[0031] Figure 12 This is a schematic diagram of a path reasoning process provided in another embodiment of this application;
[0032] Figure 13 A schematic diagram of the path planning device for multimodal data fusion and AI inference collaboration provided in this application;
[0033] Figure 14 A schematic diagram of the structure of the electronic device provided in this application.
[0034] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0035] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0036] With the development of intelligent driving technology, multimodal perception systems have gradually become a core component of vehicle perception systems. Multimodal refers to the intelligent driving system integrating the inputs from sensors such as camera images, LiDAR (Light Detection and Ranging), and Radar (Radio Detection and Ranging) to perform comprehensive, multi-scale, and cross-semantic information modeling of the environment, and based on the environment modeling, performing target detection, path planning, and decision control.
[0037] In traditional intelligent driving path decision-making, there are two main approaches: (1) Multimodal fusion model based on bird's eye view (BEV-based). This model achieves unified spatial modeling by uniformly projecting multimodal features onto the bird's eye view (BEV) space. However, this approach lacks a clear reasoning mechanism and usually relies on deep neural networks for end-to-end training, making it difficult to explain the intermediate decision-making process in the path reasoning process. (2) Language reasoning method based on chain-of-thought (CoT). The CoT method, which has emerged in large language models, can decompose complex tasks into multi-step reasoning paths. However, this approach is mainly used for text input and has not yet formed a standardized, hierarchical reasoning framework under multimodal input.
[0038] To address the aforementioned issues, this application provides a path planning method and apparatus that integrates multimodal data fusion and AI reasoning. It innovatively integrates the two methods mentioned above, inheriting the advantages of unified representation in BEV space while introducing a human-like hierarchical thinking mechanism to achieve a complete closed-loop reasoning path with interpretability from multimodal perception to path decision-making.
[0039] The multimodal data fusion and AI inference collaborative path planning method and apparatus provided in this application embodiment can be applied to at least the following scenarios:
[0040] (1) Passenger vehicle autonomous driving: In scenarios such as urban open roads, highways, and complex intersections, vehicles need to perceive the surrounding environment in real time through multi-modal sensors (cameras, lidar, millimeter-wave radar) to complete environmental modeling, target detection, path planning and decision control.
[0041] (2) Intelligent dispatching of commercial vehicles: such as unmanned delivery vehicles and smart bus systems, need to achieve efficient and safe path planning and conflict avoidance in scenarios with dense dynamic traffic participants (pedestrians, non-motorized vehicles and other vehicles).
[0042] (3) Adaptation to extreme environments: For example, in scenarios such as low light at night, rain and fog, and sensor failure, the reliability of the system needs to be ensured through multimodal data complementarity and robust inference mechanisms.
[0043] (4) Smart city traffic management: By integrating vehicle-side and roadside perception data, traffic flow optimization, accident early warning and collaborative control can be achieved.
[0044] The technical solution of this application and how it solves the above-mentioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0045] Figure 1 This is a flowchart illustrating the path planning method for multimodal data fusion and AI inference collaboration provided in this application. This method can be applied to vehicles, specifically to the vehicle's intelligent driving system. Taking the vehicle as the executing entity as an example... Figure 1 As shown, the method specifically includes the following steps:
[0046] Step 110: Obtain multimodal data during vehicle driving, at least two candidate actions in the current thought state, and the inference level corresponding to each candidate action.
[0047] The current thought state is determined based on the vehicle's driving situation at the current moment, and the reasoning level is used to represent the amount of computing resources used during path reasoning.
[0048] Step 120: Based on multimodal data, at least two candidate actions in the current thinking state, and the reasoning level corresponding to each candidate action, predict the next thinking state after the vehicle executes each candidate action and the evaluation result corresponding to each next thinking state.
[0049] Step 130: Based on each next thought state and the evaluation results corresponding to each next thought state, generate at least two candidate thought paths.
[0050] Traditional path generation techniques primarily rely on linear thought chains or fixed-structure networks, making it impossible to generate multiple candidate decision paths and perform dynamic evaluation in complex scenarios. The multimodal data fusion and AI inference-based path planning method provided in this application, by configuring a corresponding inference level for each candidate action, enables the vehicle to generate paths corresponding to that candidate action at different inference levels, achieving multi-path generation. This adapts to the needs of highly dynamic traffic scenarios and improves adaptability in complex dynamic environments.
[0051] Regarding step 110 above, multimodal data can be obtained by sensing the surrounding environment through multimodal sensors (such as cameras, lidar, and millimeter-wave radar) mounted on the vehicle.
[0052] In this embodiment, the current thought state can refer to the traffic environment generated by the vehicle's autonomous driving system based on the past and present times. Based on the current thought state, the autonomous driving system makes inferences about the next thought state, thereby determining the driving strategy that the vehicle should take in the next moment, and performing path planning inference based on the driving strategy taken in the next moment.
[0053] Specifically, the current state of mind can encompass current environmental perception information and the vehicle's historical state and intentions. For example, current environmental perception information could be "the vehicle is in the middle lane, there is a slow vehicle 50 meters ahead, the left lane is empty, and a vehicle is rapidly approaching from the right rear." The vehicle's historical state and intentions can include the current vehicle speed, acceleration, heading angle, and the actions the vehicle has just performed.
[0054] In this embodiment, the reasoning level can refer to the search depth, breadth, and computational granularity of the intelligent driving system's thought process during path reasoning. For example, in simple traffic scenarios, the reasoning level can be selected as "ultra-fast," meaning the autonomous driving system requires relatively little thought and can quickly generate the reasoned path in a short time. In complex traffic scenarios, the reasoning level can be selected as "quality," meaning the autonomous driving system will employ a deeper thinking mode, spending more computation time weighing multiple possible solutions to ensure the safety, rationality, and comfort of the reasoned path.
[0055] In this embodiment, candidate actions can refer to driving actions that the vehicle can take in the current traffic scenario. For example, candidate actions may include at least one of deceleration, acceleration, lane change, stopping, and avoidance.
[0056] By configuring multiple candidate actions, the intelligent driving system can analyze and reason about each candidate action, thereby generating more alternative paths to adapt to complex traffic scenarios.
[0057] To avoid incomplete environmental perception caused by single-modal data and to ensure that the input information of the autonomous driving system is comprehensive and covers geometric structure, semantic texture, and speed dynamics, multimodal data can be acquired through the following steps in some implementations:
[0058] Step (1): Acquire sensor data collected by various types of sensors during vehicle driving, as well as the vehicle's driving status at historical moments.
[0059] Step (2): Generate multimodal data based on sensor data and driving status.
[0060] The different types of sensors can specifically include cameras, lidar, millimeter-wave radar, etc.
[0061] Among them, the vehicle's driving state at a historical moment can be a sequence of the vehicle's historical driving states over a preset time period (e.g., 5 seconds) h={(V t C t )}. V t For vehicle speed, C t The trajectory curvature is used to characterize the dynamic tendency of a vehicle.
[0062] For example, Figure 2 This is a schematic diagram illustrating the generation of multimodal data provided in an embodiment of this application, such as... Figure 2 As shown, it integrates RGB image data collected by the camera, 3D point cloud data collected by the LIDAR sensor, point cloud data collected by the millimeter-wave radar, and historical driving state sequences recorded from historical driving states to form multimodal data.
[0063] Furthermore, in some embodiments, the sensor data includes image data acquired by a camera or camera, laser point cloud data acquired by a lidar, and millimeter-wave radar point cloud data acquired by a millimeter-wave radar.
[0064] LiDAR is used to measure the three-dimensional position and reflection intensity of objects around the vehicle. Millimeter-wave radar is a sensor that emits millimeter waves to detect the distance and speed of objects; it can still work in rain and fog, but lacks altitude information.
[0065] Specifically, taking sensors including cameras, LiDAR, and millimeter-wave radar as an example, at each time step t, the autonomous driving system simultaneously acquires RGB image data captured by the camera. 3D point cloud data output by lidar and point cloud data output by millimeter-wave radar. .
[0066] Among them, O I This represents an RGB image tensor, where the subscript 'I' denotes the image. This represents a three-dimensional tensor with height H, width W, and 3 channels over a real number field. (3D point cloud data O) L Each action (x) i y i , z i r i ), representing the three-dimensional coordinates (x, y) of the point in the vehicle-centered coordinate system. i y i , z i and reflectivity r i Point cloud data O R Each row (x j y j , z j v j In addition to three-dimensional coordinates, it also carries radial velocity information v. j .
[0067] For step 120 above, unified spatial modeling can be performed on multimodal data to improve the quality of multimodal data integration and obtain fused features.
[0068] In this embodiment, taking lane changing and deceleration as candidate actions of the vehicle at the current moment, the autonomous driving system can use various different inference levels to analyze and predict the driving situation the vehicle will face after performing the candidate action of "lane changing" based on the fusion features obtained by multimodal data fusion, thereby determining the next thinking state.
[0069] The reasoning levels can be divided into "Ultra-fast," "Fast," "Balanced," and "Quality" based on the search depth. The autonomous driving system uses the "Ultra-fast" reasoning level to analyze and predict the next thought state after a vehicle performs a lane change, and it can also use the "Fast" and "Balanced" reasoning levels to analyze and predict the next thought state after a vehicle performs a lane change. This allows the system to obtain the next thought state obtained by using different reasoning levels for the same candidate action.
[0070] At the same time, for each candidate action, different levels of reasoning can be used to obtain the corresponding next thought state, thus obtaining multiple next thought states.
[0071] After obtaining the next thought state, each next thought state can be scored using a value evaluation function (covering indicators such as safety, rule constraints, efficiency, comfort, and uncertainty penalties) to obtain an evaluation result. This evaluation result is used to characterize the quality of the vehicle's execution of the corresponding candidate action.
[0072] Among them, the next thought state with a low score can be directly eliminated, while the next thought state with a relatively high score can be retained to generate candidate thought paths.
[0073] Regarding step 130 above, if the vehicle performs a certain candidate action and the corresponding next thought state has a high score, then at least two candidate actions for that next thought state are obtained, and the next thought state after that next thought state is predicted. This process is continued, and candidate thought paths are constructed based on the candidate actions corresponding to each thought state with a high score.
[0074] Here, candidate thought paths refer to the driving paths that the vehicle can choose from in subsequent moments. Furthermore, based on the evaluation results of the next thought state in the candidate thought paths, the optimal path can be determined from each candidate thought path as the actual driving path taken by the vehicle in subsequent moments.
[0075] Traditional methods often rely on inefficient feature stacking or unstructured fusion, making it difficult to achieve collaborative reasoning in a unified space. In some embodiments, to achieve high-precision fusion of different modal data in a unified spatiotemporal coordinate system, alleviate the problems of uneven data distribution and missing radar altitude, and provide stable input for subsequent reasoning, a bird's-eye view feature map can be introduced as a unified representation space. This allows for structured mapping of heterogeneous data such as visual, lidar, and millimeter-wave radar data, achieving high-precision environmental modeling. Specifically, multimodal data is first fused to obtain a bird's-eye view feature map. Then, based on the bird's-eye view feature map, at least two candidate actions in the current thought state, and the reasoning level corresponding to each candidate action, the next thought state after the vehicle executes each candidate action is determined.
[0076] In this embodiment, the bird's-eye view feature map can be understood as a grid map formed with the vehicle as the center. It is a 2D coordinate system viewed from directly above, which can make it easier for the autonomous driving system to make path planning decisions.
[0077] Furthermore, in some embodiments, Figure 3 This is a schematic diagram of the multimodal data fusion process provided in the embodiments of this application, such as... Figure 3 As shown, it includes the following steps:
[0078] Step 310: Extract lidar features from lidar point cloud data and image features from camera image data;
[0079] Step 320: Based on the lidar features, perform feature enhancement on the millimeter-wave radar point cloud data to generate millimeter-wave radar enhanced features;
[0080] Step 330: Project the image features, millimeter-wave radar enhancement features, and lidar features into the grid of the bird's-eye view feature map, and then stitch and fuse them to form the bird's-eye view feature map.
[0081] In this embodiment, taking the fusion of lidar point cloud data, millimeter-wave radar point cloud data and camera image data as an example, lidar point cloud data can provide accurate three-dimensional geometry and distance information, but its understanding of object semantics is weak, the point cloud is sparse, and the initial lidar data is in three-dimensional space.
[0082] While camera image data can provide rich texture, color, and semantic information (such as traffic signs and traffic lights), it lacks depth information, and camera image data is on a 2D image plane.
[0083] Millimeter-wave radar point cloud data serves as a crucial supplement to information, primarily including target velocity, and provides reliable information on target presence and motion even under weather conditions where lidar performance is degraded, such as rain, snow, fog, and dust.
[0084] In this embodiment, in order to reduce the representation gap between various sensors, a bidirectional LiDAR-Radar fusion strategy is adopted. The features of the LiDAR are used to make up for the lack of height information in the millimeter-wave radar point cloud, thereby enhancing the features of the millimeter-wave radar point cloud data. This enables the autonomous driving system to maintain stable environmental modeling capabilities under complex weather, occlusion and low light conditions, so as to facilitate the subsequent construction of a safer and more robust environmental perception system.
[0085] In this embodiment, during feature projection, since the point cloud data of the LiDAR is inherently three-dimensional, it can be directly "flattened" onto the grid plane of the bird's-eye view feature map using methods such as voxelization. For example, each grid cell in the bird's-eye view feature map can encode information such as the point cloud density and average height at that location.
[0086] For millimeter-wave radar point cloud data, after compensating for the lack of height information in millimeter-wave radar point clouds by using lidar features, the three-dimensional features can be projected onto a two-dimensional bird's-eye view grid to generate a bird's-eye view representation aligned with the lidar feature space.
[0087] In addition, for two-dimensional camera image data, in order to solve the problem of "near objects appearing larger and farther objects appearing smaller"—that is, the size and position of the same object in the image change drastically with distance—deep learning technology can be used to predict a series of possible depth distributions for each pixel in the two-dimensional image, thereby "lifting" the two-dimensional image features into a three-dimensional frustum point cloud. Then, all these generated three-dimensional frustum features are projected into the grid of the bird's-eye view feature map through operations such as pooling.
[0088] Furthermore, in some embodiments, Figure 4 This is a schematic diagram of the millimeter-wave radar enhanced feature generation process provided in the embodiments of this application, such as... Figure 4 As shown, it specifically includes the following steps:
[0089] Step 410: Obtain high-dimensional spatial features from the lidar point cloud data;
[0090] Step 420: Use high-dimensional spatial features as a query library and determine the missing height information in the millimeter-wave radar point cloud data through a query function;
[0091] Step 430: Perform feature fusion between the altitude information and the millimeter-wave radar point cloud data to generate millimeter-wave radar enhanced features.
[0092] In this embodiment, the autonomous driving system first extracts the high-dimensional spatial feature tensor F from the lidar point cloud data. L As a "query database", and using the current set of three-dimensional coordinates P of the lidar point cloud. R It is considered a "query index".
[0093] This can be achieved through a pre-trained Query function Q. h Perform a two-way interaction: for each lidar point coordinate in F L The most relevant spatial-semantic context is found in the data, and the "height embedding" of the LiDAR point is predicted accordingly to obtain the height vector.
[0094] All predicted height vectors will be stacked into a matrix. C h N represents the number of embedded channels. r This represents the number of LiDAR points in the current frame. This matrix H is the "high embedding" and can be directly stitched with the original millimeter-wave radar point cloud data for subsequent fusion into the bird's-eye view feature map.
[0095] Specifically, after obtaining the highly embedded matrix H, the autonomous driving system first processes the original millimeter-wave radar features F... R The feature tensor is concatenated with H along the channel dimension to form a new composite feature tensor. This composite tensor is then fed into a lightweight convolutional network Conv(·) for local fusion and dimensionality reduction, and the output is the enhanced millimeter-wave radar enhancement feature F. R + .
[0096] The convolution operation performed by the lightweight convolutional network Conv can fully integrate the height semantics of the original point cloud data from millimeter-wave radar and the LiDAR-completed data while maintaining the spatial structure, providing robust input for subsequent unified bird's-eye view fusion.
[0097] Furthermore, after obtaining the enhanced characteristics F of millimeter-wave radar R + Then, the millimeter-wave radar enhanced signature (FR+) can be sent to Proj. BEV The (·) function projects 3D features onto a uniform 2D bird's-eye view grid, generating a model with the lidar feature F. L Spatially aligned BEV representation. Then F L The projected radar BEV features are concatenated along the channel dimension and then input into a lightweight convolutional network Conv(·) for deep fusion.
[0098] The output of the lightweight convolutional network Conv(·) is the fused bird's-eye view feature map Z, which has a dimension of C. x ×H×W,C x H represents the number of channels, and H×W represents the grid size of the bird's-eye view feature map.
[0099] In the above embodiments, by using a unified BEV space as the central point, the spatiotemporal alignment and uncertainty modeling of heterogeneous sensor data such as vision, lidar, and millimeter-wave radar are completed. This effectively solves the alignment problem of multimodal sensors in the spatial and temporal dimensions, enabling high-precision fusion of different modalities under unified spatiotemporal coordinates. It alleviates the problems of uneven data distribution and missing radar height, improves the robustness and accuracy of multimodal perception fusion, and provides stable input for subsequent path reasoning.
[0100] Furthermore, in some embodiments, after obtaining the bird's-eye view feature map, in order to facilitate direct processing by the subsequent large language model and improve processing efficiency, the bird's-eye view feature map can be further compressed into a token sequence of a preset length. This allows the large language model to determine the next thought state of the vehicle after executing each candidate action based on the token sequence, at least two candidate actions in the current thought state, and the reasoning level corresponding to each candidate action.
[0101] In this embodiment, in the large language model, the token is the basic information unit for language processing. It can be a character, a word, or even a phrase or sentence. The sensor data is compressed into a fixed-length token sequence so that the large language model can "understand" it.
[0102] Among them, large language models can refer to multimodal large language models (MLLMs). They are large-scale pre-trained language models that have the ability to understand and generate multimodal inputs such as text, images, and point clouds. They are usually based on the Transformer architecture and have billions of parameters.
[0103] Transformer is a neural network structure based on self-attention mechanism, which is suitable for sequence modeling.
[0104] In this embodiment, MLLM serves as the core inference engine and can be used for semantic understanding. Additionally, MLLM can also be used for candidate trajectory scoring and generating natural language explanations, which will be further described in subsequent embodiments.
[0105] Furthermore, in some embodiments, Figure 5 This is a schematic diagram of the bird's-eye view feature map to word sequence conversion process provided in the embodiments of this application, such as... Figure 5 As shown, it includes the following steps:
[0106] Step 510: Using the bird's-eye view feature map as input to the convolutional network, obtain the 3D feature tensor output by the convolutional network. The convolutional network is used to extract local features and compress the channel dimensions of the bird's-eye view feature map.
[0107] Step 520: Flatten the three-dimensional feature tensor into a two-dimensional matrix, where each row of the two-dimensional matrix corresponds to a local feature in the bird's-eye view feature map.
[0108] Step 530: Obtain the linear mapping weight matrix, bias vector, and position encoding matrix.
[0109] Step 540: Based on the linear mapping weight matrix, bias vector and position encoding matrix, perform spatial transformation on the two-dimensional matrix to form a word sequence.
[0110] In this embodiment, firstly, the autonomous driving system feeds the fused BEV feature map Z into the convolutional sub-network ConvNet(·) to extract the local spatial context and compress the channel dimension.
[0111] Secondly, the three-dimensional feature tensor output by the convolutional sub-network ConvNet(·) is flattened into a two-dimensional matrix through the Flatten(·) operation, with each row corresponding to a local feature block.
[0112] Next, the flattened two-dimensional matrix and the linear mapping weight matrix W proj ⊤ Multiply and add the bias vector b. proj This completes the alignment from the feature space to the token space.
[0113] In order to preserve spatial location information, a learnable location encoding matrix P can be added.
[0114] Finally, the entire sequence is normalized using LayerNorm(·), outputting a token sequence of a preset length (e.g., 64). d represents the Token dimension, which can be directly input into a large language model for high-level semantic reasoning.
[0115] Figure 6 This is a schematic diagram of the word sequence generation process provided in the embodiments of this application, such as... Figure 6 As shown, it includes the following steps:
[0116] Step 610: Input multimodal sensor data;
[0117] Step 620: Radar signature enhancement;
[0118] Step 630: Unify and fuse features into the bird's-eye view feature map;
[0119] Step 640: Bird's-eye view feature map word sequence embedding;
[0120] Step 650: Output a fixed-length word sequence.
[0121] In the above embodiments, by compressing the bird's-eye view feature map into a sequence of tokens of a preset length, it can be directly used as input to a large language model, so that the large language model can directly perform path reasoning and improve the efficiency of path reasoning generation.
[0122] In some embodiments, the inference level corresponding to a candidate action can be obtained through the following steps:
[0123] Step 1: Obtain the scene complexity of the traffic scene in which the vehicle is located at the current moment;
[0124] Step 2: Determine the reasoning level corresponding to each candidate action based on the complexity of the scenario.
[0125] In this embodiment, the autonomous driving system can continuously monitor the current traffic environment and, based on the current traffic environment, provide the inference level corresponding to the candidate action.
[0126] For example, taking a deserted rural road as an example, if the candidate action is "lane change", the corresponding reasoning level can be "super fast". However, for busy urban roads, if there are many road participants (such as pedestrians, motor vehicles and non-motor vehicles) around the vehicle, if the candidate action is "lane change", the corresponding reasoning level can be "quality" in order to ensure safety.
[0127] Furthermore, in some embodiments, for the autonomous driving system, corresponding recognition logic can be configured to determine the scene complexity, specifically including the following steps:
[0128] Step S1: Extract N key scene features from the traffic scene, where N is a positive integer.
[0129] Step S2: Normalize each key scene feature and obtain the normalization result corresponding to each key scene feature.
[0130] Step S3: Determine the scene complexity based on the normalization result and the preset weight corresponding to each key scene feature.
[0131] In this embodiment, the autonomous driving system extracts N key contextual features C i , as a key scene feature.
[0132] For example, key contextual feature C i This can include target density, velocity variance, road curvature, etc.
[0133] Each key contextual feature is first normalized and then weighted by its corresponding weight α. i Multiply them and then linearly sum them to obtain a single scalar Comp.
[0134] The scalar Comp value reflects the complexity of the scene in real time and is used as a "speed regulator" for the path reasoning tree. The higher the Comp value, the more the autonomous driving system tends to choose a deeper reasoning level to ensure decision accuracy in complex scenes; the lower the Comp value, the more the autonomous driving system tends to reason quickly to save computing power.
[0135] In this embodiment, by utilizing the scenario complexity to configure the reasoning level corresponding to each candidate action, path reasoning can be performed using an appropriate amount of computing resources in various traffic scenarios, thereby improving the utilization rate of computing resources in the autonomous driving system.
[0136] To overcome the shortcomings of traditional single-path reasoning models, this paper supports multi-path exploration and layer-by-layer selection, thereby adapting to complex and uncertain traffic scenarios while preserving traceable logical evidence throughout the process. The path reasoning process is detailed below through several examples.
[0137] In some embodiments, a hierarchical tree-of-thought (HTOT) approach can be used for path reasoning. Specifically, Figure 7 This is a schematic diagram of the path reasoning generation process provided in the embodiments of this application, such as... Figure 7 As shown, it includes the following steps:
[0138] Step 710: Construct a hierarchical mind tree with the current thought state as the root node and the next thought state after the vehicle performs each candidate action as the child node;
[0139] Step 720: Based on the evaluation results of each next thinking state, configure the scores of each sub-node in the hierarchical thinking tree;
[0140] Step 730: Based on the scores of each child node, determine at least two node paths in the hierarchical mind tree;
[0141] Step 740: Based on the candidate actions corresponding to each node in each node path, generate at least two candidate thought paths.
[0142] In this embodiment, the hierarchical mind tree is a reasoning and decision-making method based on a hierarchical structure. Its core idea is to break down the reasoning process into multiple levels. Each level in the hierarchical mind tree contains several child nodes that represent the next thinking state. The hierarchical mind tree converges layer by layer to the optimal reasoning path (i.e., the target path, which the vehicle travels on in the next moment) through a cyclical process of generation, evaluation, and selection.
[0143] In this hierarchical thinking tree, each level can be regarded as a thought state at a certain moment. The root node in the root node level represents the thought state at the current moment, and the child nodes of the next level of the root node represent the next thought state at the next moment. And so on, the child nodes in the Nth level of the hierarchical thinking tree represent the next N thought states at the next N moments.
[0144] In this embodiment of the application, by utilizing the reasoning mechanism of hierarchical thinking tree, candidate thought paths can be dynamically generated, evaluated and screened layer by layer, and a human-like reasoning process of "thinking before and after" can be constructed, which can effectively adapt to various uncertainties in complex traffic scenarios.
[0145] For example, Figure 8 This is a schematic diagram of the intelligent driving decision-making process in an embodiment of this application, such as... Figure 8 As shown, it includes the following steps:
[0146] Step 810: Obtain the current thought state s; Step 820: Monitor the traffic environment; Step 830: Extract key features; Step 840: Perform feature normalization; Step 850: Calculate the scene complexity COMP; Step 860: Determine the inference level L based on the scene complexity COMP; Step 870: Enumerate candidate actions in the candidate action set A. Specifically, this includes acceleration, deceleration, lane changing, avoidance, stopping, and maintaining the current driving state. Step 880: Use the current thought state s, inference level L, and candidate actions as input to the hierarchical transfer function; Step 890: Generate the next thought state s'; Step 8100: Perform scoring, pruning, and expansion; Step 8110: Update the current thought state s.
[0147] In this embodiment, the autonomous driving system views the intelligent driving decision-making process as a gradual evolution within a hierarchical thinking space S. For the current thinking state s at the current moment, it first enumerates possible candidate actions (e.g., deceleration, lane change, parking) from the candidate action set A, and then selects an appropriate search depth {"ultra-fast", "fast", "balanced", "quality"} from the inference level set L based on the scenario complexity. Subsequently, the hierarchical transfer function T... h By combining these three types of inputs, the next thought state s' is generated, realizing a step-by-step progression of states and providing a structured framework for subsequent scoring, pruning, and expansion.
[0148] As mentioned in the above embodiments, each next mental state can be scored by combining a value evaluation function (covering indicators such as security, rule constraints, efficiency, comfort, and uncertainty penalty) to obtain an evaluation result, which can be used to configure the scoring of the child node.
[0149] For example, in the current thought state, there is a candidate action "lane change". The autonomous driving system predicts the next thought state after executing "lane change" and, combined with the value evaluation function, obtains the evaluation result as "collision may occur". In this case, the score of the child node corresponding to the next thought state is a failing score.
[0150] In this embodiment, by introducing a hierarchical thought tree (HToT) structure into the path reasoning step and combining it with value evaluation functions (safety, rule constraints, efficiency, comfort, and uncertainty penalty), a human-like hierarchical reasoning mechanism and a more flexible and robust decision-making process are achieved, enhancing adaptability to complex scenarios. Furthermore, the generation of at least two candidate thought paths breaks through the limitations of traditional single-path or linear thought chain reasoning and can adaptively adjust the reasoning depth, ensuring rapid convergence in simple scenarios and in-depth exploration in complex scenarios.
[0151] Based on the above embodiments, in some embodiments, a pruning operation is performed on the hierarchical thinking tree, specifically referring to deleting target child nodes from the hierarchical thinking tree. The target child node's score is less than or equal to a first score threshold.
[0152] In this embodiment, the first score threshold can refer to the "failing score" mentioned above. That is, when the score of a child node is less than or equal to the "failing score", the node will be removed from the hierarchical mind tree.
[0153] By filtering out some substandard target sub-nodes, the reliability of the final generated candidate thought paths can be improved, thus preventing safety accidents when vehicles use a particular candidate thought path and improving safety.
[0154] In addition, in order to achieve interpretability and traceability of decision results, avoid the "black box output" of traditional end-to-end models, and improve user trust and transparency of autonomous driving system decision-making, the target path can be determined from at least two candidate thought paths; the leaf nodes in the target path are decoded to generate decision explanation text.
[0155] In this context, if the scores of all sub-nodes in the target path are greater than the second score threshold, then this target path is the optimal path, and subsequent vehicles will use this target path. The decision explanation text includes target speed information, driving intention information, the intelligent driving strategy adopted by the vehicle, and a textual explanation of the adopted intelligent driving strategy.
[0156] In this embodiment, the second score threshold is greater than the first score threshold. For example, each child node in the target path is the highest-scoring child node among all child nodes in each layer of the hierarchical mind tree. For instance, the hierarchical mind tree has three layers: the first layer is the root node, the highest-scoring child node in the second layer, and the highest-scoring leaf node in the third layer, which together form the node path, corresponding to the target path.
[0157] In this embodiment, after completing the node search of the hierarchical thinking tree and extracting the leaf nodes in the target path, the autonomous driving system decodes them into a structured decision dictionary (Decision) as the decision explanation text.
[0158] The decision dictionary contains five fields: speed (target vehicle speed), direction (driving intention information, such as going straight, turning left, turning right, etc.), action (specific intelligent driving strategy, such as acceleration, deceleration, lane change, stopping, etc.), rationale (natural language explanation, i.e., the text explanation of why the intelligent driving strategy is adopted), and trajectory (i.e., the three-dimensional waypoint sequence at future time, such as the sequence of 10 two-dimensional waypoints sampled at 0.5-second intervals in the next 5 seconds).
[0159] In addition, in some embodiments, since the hierarchical mind tree records the thought state at each moment, which is equivalent to the path reasoning generation process, in order to further improve the transparency of the path decision of the intelligent driving system, path reasoning process information can be generated based on the hierarchical mind tree after pruning.
[0160] The path reasoning process information includes the number of nodes in the hierarchical thinking tree after pruning, the nodes in each layer, the score of at least one node in each layer, and the evaluation result of the next thinking state corresponding to each node.
[0161] In this embodiment, the entire path reasoning process information can be recorded as T. trace={(T d Scores d , l d )} d=1 D T d The set of thought nodes reserved for the d-th layer (i.e., the nodes of each layer), Scores d Let l be the model score vector corresponding to the d-th layer node, used to score that node. d The scores are given to the top-k nodes in the d-th layer, where D is the maximum depth of the final reasoning tree (i.e., the number of node layers in the hierarchical thinking tree). This path reasoning process information is used to audit, debug, or explain the decision basis to users after the fact for the path decision of the intelligent driving system.
[0162] In this embodiment, the autonomous driving system outputs natural language descriptions of the path reasoning logic and the basis for the final decision at each step while generating path planning results. This can enhance user trust and system transparency, and facilitate subsequent debugging and supervision.
[0163] Figure 9 The path interpretation and tracing flowchart provided for the embodiments of this application is as follows: Figure 9 As shown, it includes the following steps:
[0164] Step 910: Hierarchical mind tree search; Step 920: Extract the leaf node with the highest score in the node path; Step 930: Decode into a structured decision dictionary (i.e., decision explanation text); Step 940: Decision field parsing; Step 950: Path evaluation and filtering; Step 960: Extract the complete reasoning path; Step 970: Construct path reasoning process information T. trace Step 980: Interpretable decision and trajectory output.
[0165] Furthermore, in some embodiments, after generating multiple candidate thought paths, a target path can be determined from among the multiple candidate thought paths, and then the vehicle's driving can be controlled based on the target path to achieve autonomous driving control of the vehicle.
[0166] Compared to traditional end-to-end "black box output," the above embodiments retain the evidence pointers and scoring process in the hierarchical reasoning tree during the output stage and generate natural language explanations based on structured information. This allows each final decision to be traced back to specific evidence and logical basis, outputting a complete chain of reasoning evidence and natural language explanations. This improves the interpretability of the intelligent driving system's decisions and user trust, and meets the needs of subsequent debugging and supervision.
[0167] The overall process of this application is described in detail below through some embodiments.
[0168] Figure 10The overall flowchart of path reasoning provided in the embodiments of this application is as follows: Figure 10 As shown, it includes the following steps: Step 1010: Input of multimodal sensor data; Step 1020: Unified fusion of multimodal data into bird's-eye view feature map; Step 1030: Output of fixed-length word sequence; Step 1040: Hierarchical thinking tree planning; Step 1050: Path evaluation and selection; Step 1060: Interpretability decision and trajectory output.
[0169] Furthermore, Figure 11 The schematic diagram of the path reasoning process provided in this application embodiment specifically includes the following modules: multimodal data acquisition module, data preprocessing module, radar feature enhancement module, BEV fusion module, token encoding module, hierarchical mind tree reasoning module, trajectory evaluation module, and interpretable output module. Figure 11 As shown, the above module can specifically perform the following steps:
[0170] Step 1110: Simultaneously receive RGB images from the vehicle-mounted camera, 3D point cloud data from the LiDAR, and point cloud data from the millimeter-wave radar, and synchronously record historical status information such as the vehicle's speed and trajectory curvature over the past 5 seconds. Step 1110 is executed by the multimodal data acquisition module.
[0171] Step 1120: Perform distortion correction and normalization on the RGB image, and perform spatiotemporal alignment and coordinate transformation on the point clouds of the LiDAR and millimeter-wave radar to obtain multimodal raw data in a unified reference coordinate system. Step 1120 is executed by the data preprocessing module.
[0172] Step 1130: Using a bidirectional LiDAR-Radar fusion strategy, the high-dimensional spatial features extracted by the LiDAR are used as a query library to predict and complete the missing height embedding information for the millimeter-wave radar point cloud. Step 1130 is executed by the radar feature enhancement module.
[0173] Step 1140: Project the enhanced millimeter-wave radar features and lidar features onto a unified bird's-eye view grid, and perform deep fusion using a lightweight convolutional network to obtain a comprehensive BEV feature map. Step 1140 is executed by the BEV fusion module.
[0174] Step 1150: Compress, flatten, and linearly map the synthesized BEV feature map into a fixed-length BEV token sequence, while adding learnable positional encoding for input into the visual language large model. Step 1150 is performed by the token encoding module.
[0175] Step 1160: Driven by the large visual language model, and under the control of the candidate action set and dynamic reasoning depth, multiple hierarchical candidate trajectory ideas are generated. Step 1160 is executed by the hierarchical thought tree reasoning module.
[0176] Step 1170: Based on indicators such as safety, comfort, and efficiency, the candidate thought paths are scored, pruned, and screened to select the optimal candidate thought path. Step 1170 is executed by the trajectory evaluation module.
[0177] Step 1180: Decode the optimal path into structured control commands and natural language interpretation, and output a sequence of reference trajectory points for the next 5 seconds. This step 1180 is performed by the interpretable output module.
[0178] Furthermore, Figure 12 This is a schematic diagram of a path reasoning process provided in another embodiment of this application. The path reasoning process specifically includes the following modules: a multimodal perception module, a spatial feature completion module, a spatial unified representation module, a tokenization and location encoding module, an HToT hierarchical reasoning module, a thought path scoring module, a decision generation module, and a decision recording and traceability module. For example... Figure 12 As shown, the above module can specifically perform the following steps:
[0179] Step 1210: Receive RGB images, LiDAR point clouds, millimeter-wave radar point clouds, and historical driving state sequences, and synchronize the data timestamps. Step 1210 is executed by the multimodal perception module.
[0180] Step 1220: Using a cross-modal query mechanism, the missing altitude information of the millimeter-wave radar is completed using LiDAR features to generate enhanced radar features. Step 1220 is executed by the spatial feature completion module.
[0181] Step 1230: Project the enhanced radar features and LiDAR features onto a unified BEV mesh representation, and then stitch and fuse them along the channel dimension. Step 1230 is performed by the spatial unified representation module.
[0182] Step 1240: Compress the fused BEV feature map into a fixed-length token sequence and overlay spatial location encoding to ensure that spatial structure information is preserved during inference. Step 1240 is performed by the tokenization and location encoding module.
[0183] Step 1250: Guided by the large visual language model, dynamically adjust the number of inference layers according to the scene complexity, and generate and evaluate candidate thought paths layer by layer. Step 1250 is executed by the HToT hierarchical inference module.
[0184] Step 1260: Calculate the scene complexity score using traffic environment features (such as road curvature, target density, and speed variance), and dynamically adjust the inference tree depth and expansion strategy. Step 1260 is executed by the thought path scoring module.
[0185] Step 1270: Decode the target speed, driving intention, control strategy, natural language interpretation, and predicted trajectory from the leaf node of the highest-scoring thought path. Step 1270 is executed by the decision generation module.
[0186] Step 1280: Record the reasoning process, thought process, and scoring vector in their entirety for post-audit and model optimization. This step 1280 is executed by the decision recording and traceability module.
[0187] The above embodiments introduce the entire path reasoning process, and the following advantages of this solution can be deduced: (1) The deep integration of HToT hierarchical thinking tree with multimodal sensors breaks through the bottleneck of "black box output" of traditional end-to-end systems. In complex traffic scenarios, the efficiency and safety of intelligent driving decision-making are significantly improved through dynamic reasoning depth and interpretable paths. (2) The HToT hierarchical thinking tree framework is extended to multimodal sensor data scenarios to solve the representation gap between LiDAR, Radar, and image data in semantic and geometric space. Through BEV feature maps and token sequences, a more comprehensive and adaptable solution is provided for decision-making under different weather, lighting, and traffic density conditions. (3) For urban open road environments with high dynamics and high uncertainty, such as when pedestrians suddenly cross at night, traditional single-modal systems are prone to misjudgment due to visual failure. However, this solution can still output reliable trajectories and interpretations under low illumination by using Radar-LiDAR complementary perception and hierarchical reasoning, which reduces delay and accident risk, and achieves low-cost and highly reliable real-time intelligent driving decision-making.
[0188] Figure 13 A schematic diagram of the path planning device for multimodal data fusion and AI inference collaboration provided in this application is shown below. Figure 13 As shown, the path planning device 1300 provided in this embodiment includes:
[0189] The data acquisition module 1310 is used to acquire multimodal data during vehicle driving, at least two candidate actions in the current thinking state, and the reasoning level corresponding to each candidate action.
[0190] The current thought state is determined based on the vehicle's driving situation at the current moment, and the reasoning level is used to represent the amount of computing resources used during path reasoning.
[0191] The state reasoning module 1320 is used to predict the next state of thought and the evaluation result of the next state of thought after the vehicle executes each candidate action, based on multimodal data, at least two candidate actions in the current thinking state, and the reasoning level corresponding to each candidate action.
[0192] The path generation module 1330 is used to generate at least two candidate thought paths based on the evaluation results of candidate actions and each next thought state.
[0193] In one possible implementation, the data acquisition module can be used to: acquire sensor data collected by various types of sensors during vehicle driving, as well as the vehicle's driving status at historical moments; and then generate multimodal data based on the sensor data and driving status.
[0194] In one possible implementation, the sensor data includes image data, lidar point cloud data, and millimeter-wave radar point cloud data.
[0195] In one possible implementation, the state reasoning module can be used to: fuse multimodal data to obtain a bird's-eye view feature map; and based on the bird's-eye view feature map, at least two candidate actions in the current thinking state, and the reasoning level corresponding to each candidate action, determine the next thinking state of the vehicle after executing each candidate action.
[0196] In one possible implementation, the state reasoning module can be used to: extract lidar features from lidar point cloud data and image features from camera image data; then, based on the lidar features, perform feature enhancement on the millimeter-wave radar point cloud data to generate millimeter-wave radar enhanced features; finally, project the image features, millimeter-wave radar enhanced features, and lidar features onto the grid of the bird's-eye view feature map, and stitch and fuse them to form the bird's-eye view feature map.
[0197] In one possible implementation, the state reasoning module can be used to: acquire high-dimensional spatial features from lidar point cloud data; then use the high-dimensional spatial features as a query library to determine the missing height information in the millimeter-wave radar point cloud data through a query function; and finally fuse the height information with the millimeter-wave radar point cloud data to generate millimeter-wave radar enhanced features.
[0198] In one possible implementation, the state reasoning module can be used to: compress the bird's-eye view feature map into a word sequence of a preset length; and determine the next state of mind of the vehicle after executing each candidate action based on the word sequence, at least two candidate actions in the current state of mind, and the reasoning level corresponding to each candidate action.
[0199] In one possible implementation, the state reasoning module can be used to: take the bird's-eye view feature map as input to the convolutional network and obtain the three-dimensional feature tensor output by the convolutional network; then flatten the three-dimensional feature tensor into a two-dimensional matrix; then obtain the linear mapping weight matrix, bias vector and position encoding matrix; finally, based on the linear mapping weight matrix, bias vector and position encoding matrix, perform spatial transformation on the two-dimensional matrix to form a word sequence.
[0200] The convolutional network is used to extract local features and compress channel dimensions in the bird's-eye view feature map. Each row of the two-dimensional matrix corresponds to each local feature in the bird's-eye view feature map.
[0201] In one possible implementation, the data acquisition module can be used to: acquire the scene complexity of the traffic scene in which the vehicle is currently located; and determine the inference level corresponding to each candidate action based on the scene complexity.
[0202] In one possible implementation, the data acquisition module can be used to: extract N key scene features from the traffic scene; then normalize each key scene feature to obtain the normalization result corresponding to each key scene feature; finally, determine the scene complexity based on the normalization result corresponding to each key scene feature and the preset weight corresponding to each key scene feature. Here, N is a positive integer.
[0203] In one possible implementation, the path generation module can be used to: construct a hierarchical thought tree with the current thought state as the root node and the next thought state after the vehicle performs each candidate action as the child node; then configure the score of each child node in the hierarchical thought tree based on the evaluation result of each next thought state; then determine at least two node paths in the hierarchical thought tree based on the scores of each child node; finally, generate candidate thought paths corresponding to each node path based on the candidate actions corresponding to each node in each node path.
[0204] In one possible implementation, a pruning module is also included to prune the hierarchical mind tree, deleting target child nodes in the hierarchical mind tree whose scores are less than or equal to a first score threshold.
[0205] In one possible implementation, an explanatory text generation module is also included, which is used to determine the target path among at least two candidate thought paths; and to decode the leaf nodes in the target path to generate decision explanatory text.
[0206] Among them, the scores of each sub-node in the node path corresponding to the target path are all greater than the second score threshold. The decision explanation text includes target vehicle speed information, driving intention information, intelligent driving strategy adopted by the vehicle, text explanation of the intelligent driving strategy adopted, and three-dimensional waypoint sequence at future time.
[0207] In one possible implementation, an information reasoning generation module is also included, which generates path reasoning process information based on the hierarchical mind tree after pruning.
[0208] The path reasoning process information includes the number of nodes in the hierarchical thinking tree after pruning, the nodes in each layer, the score of at least one node in each layer, and the evaluation result of the next thinking state corresponding to each node.
[0209] In one possible implementation, a control module is also included for controlling the driving of the vehicle based on the target path.
[0210] In one possible implementation, candidate actions include at least one of acceleration, deceleration, lane change, avoidance, stopping, and maintaining the current driving state.
[0211] The path generation device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0212] Figure 14 A schematic diagram of the structure of the electronic device provided in this application. Figure 14 As shown, the electronic device 1400 provided in this embodiment includes at least one processor 1401 and a memory 1402. Optionally, the electronic device 1400 further includes a communication component 1403. The processor 1401, the memory 1402, and the communication component 1403 are connected via a bus 1404.
[0213] In a specific implementation, at least one processor 1401 executes computer execution instructions stored in memory 1402, causing at least one processor 1401 to perform the above-described method.
[0214] The specific implementation process of processor 1401 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so it will not be repeated here.
[0215] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0216] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0217] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0218] This application also provides a vehicle, which includes a vehicle body and the aforementioned electronic equipment disposed in the vehicle body. The vehicle may be an electric vehicle, a gasoline-powered vehicle, or a hybrid vehicle.
[0219] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0220] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0221] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0222] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0223] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0224] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0225] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0226] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0227] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0228] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A path planning method for multimodal data fusion and AI inference collaboration, characterized in that, include: The system acquires multimodal data during vehicle driving, at least two candidate actions in the current thought state, and the reasoning level corresponding to each candidate action. The current thought state is determined based on the driving situation of the vehicle at the current moment, and the reasoning level is used to characterize the amount of computing resources used in path reasoning. Based on the multimodal data, at least two candidate actions in the current thinking state, and the reasoning level corresponding to each candidate action, predict the next thinking state and the evaluation result of the next thinking state after the vehicle executes each candidate action. Based on the evaluation results of the candidate actions and each next thought state, at least two candidate thought paths are generated.
2. The method according to claim 1, characterized in that, The acquisition of multimodal data during vehicle driving includes: The sensor data collected by various types of sensors during the driving process of the vehicle, as well as the driving status of the vehicle at historical moments, are obtained. The multimodal data is generated based on the sensor data and the driving state.
3. The method according to claim 2, characterized in that, The sensor data includes image data, lidar point cloud data, and millimeter-wave radar point cloud data.
4. The method according to claim 3, characterized in that, The step of predicting the next thought state of the vehicle after executing each candidate action, based on the multimodal data, at least two candidate actions in the current thought state, and the reasoning level corresponding to each candidate action, includes: The multimodal data is fused to obtain a bird's-eye view feature map; Based on the bird's-eye view feature map, at least two candidate actions in the current thinking state, and the reasoning level corresponding to each candidate action, the next thinking state after the vehicle executes each candidate action is determined.
5. The method according to claim 4, characterized in that, The process of fusing the multimodal data to obtain a bird's-eye view feature map includes: Extract lidar features from the lidar point cloud data and image features from the image data; Based on the features of the lidar, feature enhancement is performed on the millimeter-wave radar point cloud data to generate millimeter-wave radar enhanced features; The image features, the millimeter-wave radar enhancement features, and the lidar features are all projected into the grid of the bird's-eye view feature map and then stitched and fused together to form the bird's-eye view feature map.
6. The method according to claim 5, characterized in that, The step of enhancing the millimeter-wave radar point cloud data based on the lidar features to generate millimeter-wave radar enhanced features includes: Obtain high-dimensional spatial features from the lidar point cloud data; Using the high-dimensional spatial features as a query library, the missing height information in the millimeter-wave radar point cloud data is determined by a query function. The height information is fused with the millimeter-wave radar point cloud data to generate the millimeter-wave radar enhanced features.
7. The method according to claim 4, characterized in that, The process of determining the next thought state of the vehicle after executing each candidate action, based on the bird's-eye view feature map, at least two candidate actions in the current thought state, and the reasoning level corresponding to each candidate action, includes: The bird's-eye view feature map is compressed into a word sequence of a preset length; Based on the word sequence, at least two candidate actions in the current thought state, and the reasoning level corresponding to each candidate action, the next thought state after the vehicle executes each candidate action is determined.
8. The method according to claim 7, characterized in that, The process of compressing the bird's-eye view feature map into a word sequence of a preset length includes... Using the bird's-eye view feature map as input to a convolutional network, a three-dimensional feature tensor output by the convolutional network is obtained. The convolutional network is used to extract local features and compress channel dimensions of the bird's-eye view feature map. The three-dimensional feature tensor is flattened into a two-dimensional matrix, and each row of the two-dimensional matrix corresponds to each local feature in the bird's-eye view feature map. Obtain the linear mapping weight matrix, bias vector, and position encoding matrix; Based on the linear mapping weight matrix, the bias vector, and the positional encoding matrix, the two-dimensional matrix is spatially transformed to form the word sequence.
9. The method according to claim 1, characterized in that, Obtain the inference level corresponding to each candidate action, including: Obtain the scene complexity of the traffic scene in which the vehicle is located at the current moment; Based on the complexity of the scenario, determine the inference level corresponding to each candidate action.
10. The method according to claim 9, characterized in that, The process of obtaining the scene complexity of the traffic scene in which the vehicle is currently located includes: Extract N key scene features from the traffic scene, where N is a positive integer; Normalize each key scene feature to obtain the normalization result corresponding to each key scene feature; The scene complexity is determined based on the normalization result corresponding to each key scene feature and the preset weight corresponding to each key scene feature.
11. The method according to claim 1, characterized in that, Based on the evaluation results of the candidate actions and each next thought state, at least two candidate thought paths are generated, including: Using the current thought state as the root node and the next thought state after the vehicle performs each candidate action as the child node, a hierarchical thought tree is constructed. Based on the evaluation results of each next thinking state, configure the score of each sub-node in the hierarchical thinking tree; Based on the scores of each child node, at least two node paths are determined in the hierarchical mind tree. Based on the candidate actions corresponding to each node in each node path, candidate thought paths corresponding to the node paths are generated.
12. The method according to claim 11, characterized in that, The method further includes: The hierarchical thinking tree is pruned to delete the target child node, where the score of the target child node is less than or equal to a first score threshold.
13. The method according to claim 11, characterized in that, The method further includes: The target path is determined from the at least two candidate thought paths, wherein the scores of each sub-node in the node path corresponding to the target path are all greater than the second score threshold. The leaf nodes in the target path are decoded to generate a decision explanation text, which includes target vehicle speed information, driving intention information, the intelligent driving strategy adopted by the vehicle, a textual explanation of the intelligent driving strategy, and a three-dimensional waypoint sequence at future time.
14. The method according to claim 12, characterized in that, After pruning the hierarchical mind tree, the method further includes: Based on the hierarchical mind tree after the pruning operation, path reasoning process information is generated. The path reasoning process information includes the number of node layers in the hierarchical mind tree after the pruning operation, the nodes in each layer, the score of at least one node in each layer, and the evaluation result of the next thinking state corresponding to each node.
15. The method according to claim 13, characterized in that, The method further includes: Based on the target path, control the driving of the vehicle.
16. The method according to any one of claims 1-15, characterized in that, The candidate actions include at least one of acceleration, deceleration, lane change, avoidance, stopping, and maintaining the current driving state.
17. A path planning device for multimodal data fusion and AI inference collaboration, characterized in that, include: The data acquisition module is used to acquire multimodal data during vehicle driving, at least two candidate actions in the current thinking state, and the reasoning level corresponding to each candidate action. The current thinking state is determined based on the driving situation of the vehicle at the current moment, and the reasoning level is used to characterize the amount of computing resources used in path reasoning. The state reasoning module is used to predict the next state of thought and the evaluation result of the next state of thought after the vehicle executes each candidate action, based on the multimodal data, at least two candidate actions in the current thinking state and the reasoning level corresponding to each candidate action. The path generation module is used to generate at least two candidate thought paths based on the candidate actions and the evaluation results of each next thought state.
18. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-16.
19. A vehicle, characterized in that, It includes a vehicle body and an electronic device as described in claim 18 disposed in the vehicle body.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-16.
21. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-16.
Citation Information
Patent Citations
Automatic driving large model framework based on 3D space-time perception and human-like decision-making reasoning
CN120182938A
Autonomous navigation method, system and equipment based on multi-modal large model and medium
CN121230741A