Target vehicle path planning method and device, equipment and medium

By synchronous data collection and three-dimensional spatial semantic modeling of the SpatialLM model, combined with graph neural networks and reinforcement learning, the spatiotemporal synchronization problem of multi-sensor fusion in complex traffic environments is solved, the accuracy and adaptability of vehicle path planning are achieved, and the safety and reliability of intelligent driving are improved.

CN120628137APending Publication Date: 2025-09-12SHANGHAI DONGPU INFORMATION TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510973164.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In existing technologies, multi-sensor fusion solutions lack the accuracy of spatiotemporal synchronization in complex outdoor traffic environments, resulting in information loss or misjudgment. Spatial understanding models have not been widely used in the driving field and are difficult to effectively integrate into the perception-decision-control link.

Method used

By triggering a synchronization mechanism, each sensor collects data at the same time. Cameras, lidar, millimeter-wave radar, GNSS, and IMU are used for feature extraction and fusion. The SpatialLM model is used for three-dimensional spatial semantic modeling. Graph neural networks and reinforcement learning are combined for path planning. Traffic rules and vehicle dynamics models are introduced for fine-tuning, and transfer learning is used for model optimization.

Benefits of technology

It achieves precise fusion of multimodal feature data and three-dimensional spatial semantic modeling, generates a clear spatial structure map, provides intuitive environmental cognition, ensures the scientific nature and adaptability of path planning, and improves the safety and traffic efficiency of vehicles in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120628137A_ABST
    Figure CN120628137A_ABST
Patent Text Reader

Abstract

The invention provides a target vehicle path planning method, apparatus and device, and a medium. The method comprises the steps of triggering a synchronization mechanism to enable each sensor to perform data acquisition at the same moment; performing feature extraction based on the original data acquired by each sensor to obtain multi-modal feature data; fusing the multi-modal feature data to obtain a multi-modal feature vector; carrying out three-dimensional space semantic modeling on the multi-modal feature vector by utilizing a SpatialLM model to obtain a space structure map; performing path planning on the target vehicle through a path planning algorithm based on the spatial structure map and the traffic rule knowledge base meeting the preset requirement; the sensors comprise a camera, a laser radar, a millimeter wave radar, a GNSS and an IMU.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle path planning, and in particular to a target vehicle path planning method, device, equipment and medium. Background Art

[0002] Autonomous driving requires real-time processing of complex environmental information and accurate decision-making. Existing multi-sensor fusion solutions lack temporal and spatial synchronization accuracy. Furthermore, multimodal data fusion presents technical challenges, leading to information loss or misjudgment. In recent years, large-scale spatial understanding models based on the Transformer architecture have achieved breakthroughs in areas such as indoor scene reconstruction and robot navigation. These models can extract semantically informed spatial structures from multimodal inputs and perform reasoning and prediction. However, current spatial understanding models have not yet been widely applied in the driving field, especially in complex outdoor traffic environments. How to effectively integrate such models into the perception-decision-control chain remains a pressing technical challenge. Summary of the Invention

[0003] The main purpose of the present invention is to solve the technical problem in the prior art of how to effectively integrate the spatial understanding model into the perception-decision-control link in a complex outdoor traffic environment.

[0004] A first aspect of the present invention provides a target vehicle path planning method, comprising: Triggering the synchronization mechanism to enable each sensor to collect data at the same time; extracting features based on the raw data collected by each sensor to obtain multimodal feature data; fusing the multimodal feature data to obtain a multimodal feature vector; The SpatialLM model is used to perform three-dimensional spatial semantic modeling on the multimodal feature vector to obtain a spatial structure map; The target vehicle's path is planned using a path planning algorithm based on the spatial structure graph and a traffic rules knowledge base that meets preset requirements; The sensors include: camera, lidar, millimeter wave radar, GNSS and IMU.

[0005] Optionally, in a first implementation of the first aspect of the present invention, the triggering synchronization mechanism to enable each sensor to collect data at the same time includes: When collecting data, each sensor adds a timestamp with a timestamp accuracy that meets the preset requirements; for data with a timestamp deviation, a trigger signal is sent to all sensors through a synchronous trigger device so that each sensor collects data at the same time.

[0006] Optionally, in a second implementation of the first aspect of the present invention, extracting features based on the raw data collected by each sensor to obtain multimodal feature data includes: Based on the image data collected by the camera, the collected image data is preprocessed including denoising, grayscale conversion and normalization to obtain preprocessed image data; the preprocessed image data is extracted with a convolutional neural network; the visual features include: edges, textures, shape features in the image, and target features including vehicles, pedestrians, and traffic signs; LiDAR-based point cloud data is collected and converted into a bird's-eye view or voxel grid format. The spatial geometric features of the point cloud are extracted using a point cloud-specific neural network. These features include point positions, normal vectors, curvature, and object shape and size features represented by point cloud clusters. These point cloud-specific neural networks include PointNet and PointNet++. The millimeter-wave radar collects echo signals, performs filtering, noise removal, and clutter preprocessing on the collected echo signals to obtain preprocessed echo signals. The preprocessed echo signals are converted from time domain signals to frequency domain signals through Fourier transform, and target position and motion parameter features are extracted based on the frequency domain signals. The vehicle's position and speed information is collected based on GNSS, and the collected vehicle position and speed information is decoded and positioned to obtain the vehicle's latitude and longitude, altitude, and speed information; The acceleration, angular velocity and attitude information of the vehicle are collected based on the IMU, and the collected acceleration, angular velocity and attitude information of the vehicle are preprocessed including denoising and zero bias correction to obtain the preprocessed vehicle acceleration, angular velocity and attitude information; The vehicle's latitude and longitude, altitude and speed information are fused with the pre-processed vehicle acceleration, angular velocity and attitude information to obtain the vehicle's attitude and position information; the vehicle's motion state characteristics are extracted based on the vehicle's attitude and position information.

[0007] Optionally, in a third implementation of the first aspect of the present invention, fusing the multimodal feature data to obtain a multimodal feature vector includes: Multimodal feature data are fused by splicing or weighted fusion to obtain a multimodal feature vector; The multimodal feature data is fused by weighted fusion, including: A data fusion weight distribution model is established; the weight distribution model dynamically adjusts the fusion weight of each sensor data according to the performance of different sensors in different environments.

[0008] Optionally, in a fourth implementation of the first aspect of the present invention, performing three-dimensional spatial semantic modeling on the multimodal feature vector using the SpatialLM model to obtain a spatial structure map includes: The SpatialLM model is trained using a large-scale three-dimensional spatial dataset that meets the preset requirements to obtain a pre-trained SpatialLM model; Traffic rule constraints and vehicle dynamics model constraints are introduced to fine-tune the pre-trained SpatialLM model to obtain a fine-tuned SpatialLM model. The fine-tuned SpatialLM model encodes the multimodal feature vectors to generate feature vectors containing spatial location and semantic information. Based on the generated feature vectors containing spatial location and semantic information, the self-attention mechanism is used to obtain the dependencies between different regions in the space and generate a spatial structure map with functional labels. The spatial structure graph with functional labels includes: a spatial structure graph with functional labels including express delivery vehicle passable areas, dangerous areas, and traffic nodes; The spatial structure graph includes nodes and edges; wherein the nodes represent spatial entities, including location coordinates, semantic categories, and attribute information; and the edges represent semantic relationships and spatial connection relationships between spatial entities.

[0009] Optionally, in a fifth implementation of the first aspect of the present invention, the path planning for the target vehicle is performed using a path planning algorithm based on the spatial structure graph and a traffic rule knowledge base that meets preset requirements, including: Construct a graph neural network model and use it to infer spatial structure graphs with functional labels to obtain inference results; wherein the inference results include: extracting implicit relationships between spatial entities and global structural features; Based on reinforcement learning methods, a driving behavior decision model is constructed with the goal of ensuring the safety, efficiency, and compliance of the target vehicle. The state space of the driving behavior decision model includes the vehicle's current position, speed, posture, the spatial structure of the surrounding environment, and relevant rules in a traffic rule knowledge base that meet preset requirements. The action space of the driving behavior decision model includes driving actions such as acceleration, deceleration, steering, lane changing, and parking. Behavioral decisions are made using the constructed driving behavior decision model. Based on the inference results of the graph neural network model and the behavioral decisions of the driving behavior decision model, a path planning algorithm is used to perform path planning on a spatial structure map with functional labels to generate the optimal driving path from the starting point to the end point; at the same time, the generated optimal driving path is dynamically adjusted according to real-time traffic conditions and environmental changes.

[0010] Optionally, in a sixth implementation of the first aspect of the present invention, the method further includes: fine-tuning the SpatialLM model based on actual driving feedback data using a transfer learning method to obtain a fine-tuned SpatialLM model; The actual driving feedback data includes: raw data collected by each sensor, spatial structure map, execution results of behavioral decisions, actual driving status of the vehicle and information on changes in the traffic environment; Construct a SpatialLM model evaluation metric, and evaluate the fine-tuned SpatialLM model based on the constructed SpatialLM model evaluation metric. If the current SpatialLM model evaluation metric meets the preset requirements, stop fine-tuning the SpatialLM model. If the current SpatialLM model evaluation metric does not meet the preset requirements, continue to collect actual driving feedback data, and use the collected actual driving feedback data to continue fine-tuning the SpatialLM model until the evaluation metric of the fine-tuned SpatialLM model meets the preset requirements. The evaluation indicators include: the accuracy of the spatial structure map, the correctness of the behavioral decision and the efficiency of path planning.

[0011] A second aspect of the present invention provides a target vehicle path planning device, characterized by comprising: The feature extraction module is used to trigger the synchronization mechanism so that each sensor collects data at the same time; extract features based on the raw data collected by each sensor to obtain multimodal feature data; and fuse the multimodal feature data to obtain a multimodal feature vector; The spatial structure map construction module is used to use the SpatialLM model to perform three-dimensional spatial semantic modeling on multimodal feature vectors to obtain a spatial structure map; The path planning module is used to plan the path of the target vehicle through a path planning algorithm based on the spatial structure map and the traffic rules knowledge base that meets the preset requirements; The sensors include: camera, lidar, millimeter wave radar, GNSS and IMU.

[0012] Optionally, in a first implementation of the second aspect of the present invention, triggering the synchronization mechanism so that each sensor collects data at the same time includes: When collecting data, each sensor adds a timestamp with a timestamp accuracy that meets the preset requirements; for data with a timestamp deviation, a trigger signal is sent to all sensors through a synchronous trigger device so that each sensor collects data at the same time.

[0013] Optionally, in a second implementation of the second aspect of the present invention, extracting features based on the raw data collected by each sensor to obtain multimodal feature data includes: Based on the image data collected by the camera, the collected image data is preprocessed including denoising, grayscale conversion and normalization to obtain preprocessed image data; the preprocessed image data is extracted with a convolutional neural network; the visual features include: edges, textures, shape features in the image, and target features including vehicles, pedestrians, and traffic signs; LiDAR-based point cloud data is collected and converted into a bird's-eye view or voxel grid format. The spatial geometric features of the point cloud are extracted using a point cloud-specific neural network. These features include point positions, normal vectors, curvature, and object shape and size features represented by point cloud clusters. These point cloud-specific neural networks include PointNet and PointNet++. The millimeter-wave radar collects echo signals, performs filtering, noise removal, and clutter preprocessing on the collected echo signals to obtain preprocessed echo signals. The preprocessed echo signals are converted from time domain signals to frequency domain signals through Fourier transform, and target position and motion parameter features are extracted based on the frequency domain signals. The vehicle's position and speed information is collected based on GNSS, and the collected vehicle position and speed information is decoded and positioned to obtain the vehicle's latitude and longitude, altitude, and speed information; The acceleration, angular velocity and attitude information of the vehicle are collected based on the IMU, and the collected acceleration, angular velocity and attitude information of the vehicle are preprocessed including denoising and zero bias correction to obtain the preprocessed vehicle acceleration, angular velocity and attitude information; The vehicle's latitude and longitude, altitude and speed information are fused with the pre-processed vehicle acceleration, angular velocity and attitude information to obtain the vehicle's attitude and position information; the vehicle's motion state characteristics are extracted based on the vehicle's attitude and position information.

[0014] Optionally, in a third implementation of the second aspect of the present invention, fusing the multimodal feature data to obtain a multimodal feature vector includes: Multimodal feature data are fused by splicing or weighted fusion to obtain a multimodal feature vector; The multimodal feature data is fused by weighted fusion, including: A data fusion weight distribution model is established; the weight distribution model dynamically adjusts the fusion weight of each sensor data according to the performance of different sensors in different environments.

[0015] Optionally, in a fourth implementation of the second aspect of the present invention, the spatial structure map construction module includes: The training submodule is used to train the SpatialLM model using a large-scale three-dimensional spatial dataset that meets the preset requirements to obtain a pre-trained SpatialLM model; The fine-tuning submodule is used to introduce traffic rule constraints and vehicle dynamics model constraints to fine-tune the pre-trained SpatialLM model to obtain a fine-tuned SpatialLM model; The spatial structure map generation submodule is used to encode the multimodal feature vector using the fine-tuned SpatialLM model to generate a feature vector containing spatial position and semantic information. Based on the generated feature vector containing spatial position and semantic information, the self-attention mechanism is used to obtain the dependency relationship between different regions in the space and generate a spatial structure map with functional labels. The spatial structure graph with functional labels includes: a spatial structure graph with functional labels including express delivery vehicle passable areas, dangerous areas, and traffic nodes; The spatial structure graph includes nodes and edges; wherein the nodes represent spatial entities, including location coordinates, semantic categories, and attribute information; and the edges represent semantic relationships and spatial connection relationships between spatial entities.

[0016] Optionally, in a fifth implementation of the second aspect of the present invention, the path planning module includes: The inference result acquisition submodule is used to build a graph neural network model and use the graph neural network model to infer the spatial structure graph with functional labels to obtain inference results; wherein, the inference results include: extracting implicit relationships between spatial entities and global structural features; A behavior decision submodule is used to construct a driving behavior decision model based on reinforcement learning methods, with the driving safety, efficiency, and compliance of the target vehicle as the goals. The state space of the driving behavior decision model includes: the vehicle's current position, speed, posture, the spatial structure map of the surrounding environment, and relevant rules in the traffic rule knowledge base that meet preset requirements; the action space of the driving behavior decision model includes: driving actions such as acceleration, deceleration, steering, lane changing, and parking; and behavioral decisions are made using the constructed driving behavior decision model. The driving path generation submodule is used to make behavioral decisions based on the inference results of the graph neural network model and the driving behavior decision model. It adopts a path planning algorithm to perform path planning on a spatial structure map with functional labels to generate the optimal driving path from the starting point to the end point; at the same time, the generated optimal driving path is dynamically adjusted according to real-time traffic conditions and environmental changes.

[0017] Optionally, in a sixth implementation of the second aspect of the present invention, the apparatus further comprises: a SpatialLM model fine-tuning module, configured to fine-tune the SpatialLM model based on actual driving feedback data using a transfer learning method to obtain a fine-tuned SpatialLM model; The actual driving feedback data includes: raw data collected by each sensor, spatial structure map, execution results of behavioral decisions, actual driving status of the vehicle and information on changes in the traffic environment; Construct a SpatialLM model evaluation metric, and evaluate the fine-tuned SpatialLM model based on the constructed SpatialLM model evaluation metric. If the current SpatialLM model evaluation metric meets the preset requirements, stop fine-tuning the SpatialLM model. If the current SpatialLM model evaluation metric does not meet the preset requirements, continue to collect actual driving feedback data, and use the collected actual driving feedback data to continue fine-tuning the SpatialLM model until the evaluation metric of the fine-tuned SpatialLM model meets the preset requirements. The evaluation indicators include: the accuracy of the spatial structure map, the correctness of the behavioral decision and the efficiency of path planning.

[0018] A third aspect of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the target vehicle path planning method described above.

[0019] The fourth aspect of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the target vehicle path planning method described above when executed by the processor.

[0020] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention triggers a synchronization mechanism to ensure that multiple types of sensors, such as cameras and lidar, collect data at the same time, avoiding data inconsistency caused by time deviations. At the same time, it adds high-precision timestamps to further ensure the timeliness and accuracy of the data. 2. The present invention uses a splicing or weighted fusion method for multimodal feature data, especially a dynamic adjustment of the weight distribution model in weighted fusion. This can fully utilize the advantages of each sensor based on its performance in different environments, so that the fused multimodal feature vector can more accurately and comprehensively reflect the information about the vehicle's surrounding environment; 3. This paper pre-trains the SpatialLM model using a large-scale 3D spatial dataset and fine-tunes it by introducing traffic rules and vehicle dynamics model constraints. This enables accurate 3D spatial semantic modeling of multimodal feature vectors, generating a spatial structure graph containing rich spatial location and semantic information. The nodes and edges in the graph clearly represent spatial entities, their semantics, and spatial relationships, and are labeled with functions, providing intuitive and intelligent environmental awareness for the vehicle. 4. The present invention is based on a spatial structure map and a traffic rules knowledge base, and combines graph neural network model reasoning and reinforcement learning to construct a driving behavior decision model. It comprehensively considers various factors in the vehicle driving process, from extracting implicit relationships between spatial entities to making behavioral decisions with safety, efficiency and compliance as the goals.

[0021] 5. The present invention uses a path planning algorithm to generate the optimal driving path and can dynamically adjust it according to real-time traffic conditions and environmental changes. This not only ensures the scientific nature of path planning, but also improves the adaptability of vehicles in complex and changing environments, effectively avoiding traffic accidents, improving traffic efficiency, and enhancing the safety and reliability of vehicle driving; 6. This invention uses a transfer learning approach to continuously fine-tune the SpatialLM model based on multi-dimensional feedback data from actual driving. It also scientifically evaluates and optimizes the model by constructing an evaluation system that includes indicators such as spatial structure map accuracy, behavioral decision-making accuracy, and path planning efficiency. This iterative optimization mechanism enables the model to continuously adapt to new driving scenarios and environmental changes, continuously improve performance, and ensure the advancement and practicality of the target vehicle path planning method. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings: Figure 1 This is a first flow chart of a target vehicle path planning method provided by an embodiment of the present invention.

[0023] Figure 2 This is a second flow chart of the target vehicle path planning method provided by an embodiment of the present invention.

[0024] Figure 3 This is a third flow chart of the target vehicle path planning method provided by an embodiment of the present invention.

[0025] Figure 4 A schematic structural diagram of a target vehicle path planning device provided in an embodiment of the present invention.

[0026] Figure 5 Another structural schematic diagram of a target vehicle path planning device provided by an embodiment of the present invention.

[0027] Figure 6 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0028] Embodiments of the present invention provide a target vehicle path planning method, apparatus, device, and medium, including: triggering a synchronization mechanism to enable each sensor to collect data at the same time; performing feature extraction based on the raw data collected by each sensor to obtain multimodal feature data; fusing the multimodal feature data to obtain a multimodal feature vector; using the SpatialLM model to perform three-dimensional spatial semantic modeling on the multimodal feature vector to obtain a spatial structure map; and planning the path of the target vehicle using a path planning algorithm based on the spatial structure map and a traffic rule knowledge base that meets preset requirements; wherein the sensors include: a camera, a lidar, a millimeter-wave radar, a GNSS, and an IMU. The present invention solves the technical problem of how to effectively integrate spatial understanding models into the perception-decision-control chain in complex outdoor traffic environments in the prior art.

[0029] The terms "first," "second," "third," "fourth," and so on (if any) in the description and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that shown or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements expressly listed, but may include other steps or elements not expressly listed or inherent to such process, method, product, or apparatus.

[0030] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 , a first embodiment of the target vehicle path planning method in an embodiment of the present invention includes: 101. Triggering a synchronization mechanism so that each sensor collects data at the same time; performing feature extraction based on the raw data collected by each sensor to obtain multimodal feature data; fusing the multimodal feature data to obtain a multimodal feature vector; In this embodiment, each sensor adds an accurate timestamp during data collection, and the timestamp accuracy reaches the microsecond level. For data with timestamp deviations, hardware trigger synchronization is performed using a synchronous clock signal, and a trigger signal is sent to all sensors through a synchronous trigger device, so that each sensor starts data collection at the same time.

[0031] The sensors include: camera, lidar, millimeter wave radar, GNSS and IMU; The feature extraction based on the raw data collected by each sensor to obtain multimodal feature data includes: Convolutional neural networks are used for visual feature extraction. For camera data of different resolutions and frame rates, appropriate CNN architectures, such as ResNet and VGG, are selected. The image is first preprocessed, including denoising, grayscale conversion, and normalization. The image is then fed into the CNN network to extract low-level visual features such as edges, textures, and shapes, as well as high-level semantic features of objects such as vehicles, pedestrians, and traffic signs. For point cloud data collected by LiDAR, point cloud processing algorithms are used for feature extraction. First, the point cloud data is converted into a bird's-eye view or voxel grid format. Then, point cloud-specific neural networks such as PointNet and PointNet++ are used to extract the spatial geometric features of the point cloud, such as point position, normal vector, curvature, etc., as well as the shape and size of the object represented by the point cloud cluster. The millimeter-wave radar echo signal is first filtered to remove noise and clutter. Then, the time-domain signal is converted into a frequency-domain signal through Fourier transform to extract information such as the target's distance, speed, and angle. Radar signal processing algorithms, such as constant false alarm detection, are used to detect the target's position and motion parameters and convert them into coordinate data in a Cartesian coordinate system. GNSS provides the vehicle's position and velocity information, while IMU provides the vehicle's acceleration, angular velocity, and attitude information. GNSS data is decoded and positioned to obtain the vehicle's latitude, longitude, altitude, and speed. IMU data is preprocessed, including denoising and bias correction. GNSS and IMU data are then fused through inertial navigation algorithms, such as the extended Kalman filter (EKF), to obtain the vehicle's precise attitude and position information and extract the vehicle's motion state characteristics, such as acceleration, angular velocity, and heading angle.

[0032] The multimodal data that have undergone synchronization and preliminary feature extraction are fused, including: splicing or weighted fusion of feature vectors from different sensors to form a multimodal feature vector; fusing the detection results of each sensor, and using a voting mechanism or Bayesian network for decision fusion. At the same time, a data fusion weight distribution model is established to dynamically adjust the fusion weight of each sensor data according to the performance of different sensors in different environments.

[0033] 102. Use the SpatialLM model to perform three-dimensional spatial semantic modeling on multimodal feature vectors to obtain a spatial structure map; In this embodiment, a pre-trained SpatialLM model based on the Transformer architecture is used. The model is pre-trained on a large-scale three-dimensional spatial dataset and has learned rich spatial semantic knowledge. The input of the SpatialLM model is a multimodal feature vector, which is processed by a multi-layer Transformer encoder to capture the semantic relationship and spatial structure information between different objects in the space. The output of the SpatialLM model is a three-dimensional spatial semantic representation, which contains the semantic category and attribute information of each spatial position. The fused multimodal feature vector is input into the pre-trained SpatialLM model. The model first encodes the input data to generate a feature vector containing spatial position and semantic information, and then, through the self-attention mechanism, the model The model can focus on the dependencies between different areas in space and realize global semantic modeling of three-dimensional space. During the modeling process, taking into account the particularity of vehicle driving scenarios, the model is adjusted in a targeted manner, such as introducing prior knowledge such as traffic rule constraints and vehicle dynamics models to improve the model's understanding of driving-related spatial semantics; based on the output results of the SpatialLM model, a spatial structure map with functional labels is generated; the definition of functional labels is based on the actual needs of driving scenarios, including areas where express vehicles can pass, dangerous areas, traffic nodes, etc. Each node in the map contains location coordinates, semantic categories, attribute information, and connection relationships with other nodes. It is stored in a graph data structure, with nodes representing spatial entities and edges representing semantic relationships and spatial connection relationships between entities.

[0034] 103. Based on the spatial structure map and the traffic rules knowledge base that meets the preset requirements, the target vehicle is planned through the path planning algorithm; In this embodiment, the traffic rules knowledge base includes all traffic rules and regulations related to vehicle driving, as well as common driving scenarios and response strategies; First, the spatial structure graph is converted into the input format of a graph neural network. Node features contain the semantic categories, attribute information, and location coordinates of spatial entities, while edge features contain the connection and semantic relationships between entities. Graph neural network models such as graph convolutional neural networks and graph attention networks are used to reason on the spatial graph and extract the implicit relationships and global structural features between spatial entities. Through graph neural network reasoning, information such as the possibility of passage, traffic flow distribution, and potential hazards in different spatial areas can be predicted, providing a basis for path planning and behavioral decision-making. A reinforcement learning method is used to establish a driving behavior decision-making model with the goal of ensuring the driving safety, efficiency, and compliance of express delivery vehicles. The state space includes the vehicle's current position, speed, posture, spatial structure map information of the surrounding environment, and relevant rules in the traffic rules knowledge base; the action space includes driving actions such as acceleration, deceleration, steering, lane changing, and parking; the design of the reward function comprehensively considers factors such as driving safety, driving efficiency, and compliance. Through interactive training with the environment, the reinforcement learning model can learn the optimal driving strategy in different scenarios; combining the inference results of the graph neural network and the behavioral decision-making of reinforcement learning, classic path planning algorithms such as the Dijkstra algorithm and the A* algorithm are used to perform path planning on the spatial structure map. During the planning process, factors such as road capacity, speed limit requirements, traffic flow, and special area restrictions are taken into account to generate the optimal driving path from the starting point to the end point. At the same time, the path is dynamically adjusted according to real-time traffic conditions and environmental changes.

[0035] Through multi-link collaborative innovation, this embodiment has built a complete and efficient intelligent path planning system, significantly improving the perception, decision-making and execution capabilities of target vehicles in complex traffic scenarios, and providing strong technical support for the safety, reliability and intelligence level of intelligent driving. See also Figure 2 , a second embodiment of the target vehicle path planning method in an embodiment of the present invention includes: 201. Trigger a synchronization mechanism to enable each sensor to collect data at the same time; extract features based on the raw data collected by each sensor to obtain multimodal feature data; and fuse the multimodal feature data to obtain a multimodal feature vector; wherein the sensors include: a camera, a lidar, a millimeter-wave radar, a GNSS, and an IMU; 202. Use the SpatialLM model to perform three-dimensional spatial semantic modeling on the multimodal feature vector to obtain a spatial structure map; 203. Based on the spatial structure map and the traffic rules knowledge base that meets the preset requirements, the path planning algorithm is used to plan the path of the target vehicle; 204. Using the transfer learning method, the SpatialLM model is fine-tuned based on the actual driving feedback data to obtain a fine-tuned SpatialLM model; In this embodiment, various feedback data are collected in real time, including: raw data collected by sensors, output results of spatial semantic modeling, execution results of behavioral decisions, actual driving status of the vehicle, and information on changes in the traffic environment. Transfer learning is used to fine-tune the pre-trained SpatialLM model using the processed feedback data. During fine-tuning, the model's underlying general feature extraction layer is retained, while the higher-level semantic modeling layer undergoes targeted training to adapt the model to the specific area's road characteristics, traffic regulations, and unique driving scenarios. Based on the type and amount of feedback data, an appropriate fine-tuning strategy is selected, such as full fine-tuning, partial layer fine-tuning, or parameter freezing fine-tuning, to improve training efficiency while ensuring improved model performance. A model evaluation metric system is established, including spatial semantic modeling accuracy, behavioral decision-making accuracy, and path planning efficiency, to evaluate the fine-tuned model. Based on the evaluation results, a decision is made on whether to update the model. If the model performance meets the requirements, it is deployed in the production system. If not, data collection, processing, and fine-tuning are repeated until the model performance reaches the expected target. A model version management mechanism is also established to record the model's update history and performance changes.

[0036] This embodiment fine-tunes the SpatialLM model based on feedback data from actual driving, which can improve the adaptability of the model in specific areas or special scenarios.

[0037] See also Figure 3 , a second embodiment of the target vehicle path planning method in an embodiment of the present invention includes: 301. Trigger a synchronization mechanism so that each sensor collects data at the same time; extract features based on the raw data collected by each sensor to obtain multimodal feature data; and fuse the multimodal feature data to obtain a multimodal feature vector. The sensors include: a camera, a lidar, a millimeter-wave radar, a GNSS, and an IMU. 302. Use the SpatialLM model to perform three-dimensional spatial semantic modeling on the multimodal feature vector to obtain a spatial structure map; 303. Based on the spatial structure map and the traffic rules knowledge base that meets the preset requirements, the path planning algorithm is used to plan the path of the target vehicle; 304. Using the transfer learning method, the SpatialLM model is fine-tuned based on the actual driving feedback data to obtain a fine-tuned SpatialLM model; 305. Present the spatial structure map in a visual form to the target vehicle remote monitoring system to enhance system transparency and trustworthiness. In this embodiment, the target vehicle may be a courier truck. Through a visual interface, the work process and decision-making basis are clearly displayed to remote monitoring personnel, such as the fusion results of multimodal data, the generation process of spatial structure maps, the calculation steps of path planning, etc., to improve the transparency of the system. At the same time, it provides model performance evaluation reports and historical data records to enable operators to understand the reliability and stability of the model and enhance their trust in the system.

[0038] The target vehicle path planning method according to the embodiment of the present invention is described above. The target vehicle path planning device according to the embodiment of the present invention is described below. Figure 4 In one embodiment of the present invention, a target vehicle path planning device includes: The feature extraction module 401 is used to trigger the synchronization mechanism so that each sensor collects data at the same time; extract features based on the raw data collected by each sensor to obtain multimodal feature data; and fuse the multimodal feature data to obtain a multimodal feature vector; In this embodiment, the feature extraction module 401 includes: The synchronous trigger submodule 4011 is used to add a timestamp with a timestamp accuracy that meets the preset requirements when each sensor is collecting data; if there is a deviation in the timestamp, a trigger signal is sent to all sensors through the synchronous trigger device so that all sensors collect data at the same time; The multimodal feature data acquisition submodule 4012 is used to collect raw data based on the camera, lidar, millimeter-wave radar, GNSS, and IMU, and perform feature extraction on the collected raw data to obtain multimodal feature data; In this embodiment, the multimodal feature data acquisition submodule 4012 includes: Based on the image data collected by the camera, the collected image data is preprocessed including denoising, grayscale conversion and normalization to obtain preprocessed image data; the preprocessed image data is extracted with a convolutional neural network; the visual features include: edges, textures, shape features in the image, and target features including vehicles, pedestrians, and traffic signs; LiDAR-based point cloud data is collected and converted into a bird's-eye view or voxel grid format. The spatial geometric features of the point cloud are extracted using a point cloud-specific neural network. These features include point positions, normal vectors, curvature, and object shape and size features represented by point cloud clusters. These point cloud-specific neural networks include PointNet and PointNet++. The millimeter-wave radar collects echo signals, performs filtering, noise removal, and clutter preprocessing on the collected echo signals to obtain preprocessed echo signals. The preprocessed echo signals are converted from time domain signals to frequency domain signals through Fourier transform, and target position and motion parameter features are extracted based on the frequency domain signals. The vehicle's position and speed information is collected based on GNSS, and the collected vehicle position and speed information is decoded and positioned to obtain the vehicle's latitude and longitude, altitude, and speed information; The acceleration, angular velocity and attitude information of the vehicle are collected based on the IMU, and the collected acceleration, angular velocity and attitude information of the vehicle are preprocessed including denoising and zero bias correction to obtain the preprocessed vehicle acceleration, angular velocity and attitude information; The vehicle's latitude and longitude, altitude and speed information are fused with the pre-processed vehicle acceleration, angular velocity and attitude information to obtain the vehicle's attitude and position information; the vehicle's motion state characteristics are extracted based on the vehicle's attitude and position information.

[0039] The fusion submodule 4013 is used to fuse the multimodal feature data by splicing or weighted fusion to obtain a multimodal feature vector; In this embodiment, the multimodal feature data is fused by a weighted fusion method, including: A data fusion weight distribution model is established; the weight distribution model dynamically adjusts the fusion weight of each sensor data according to the performance of different sensors in different environments.

[0040] A spatial structure map construction module 402 is used to perform three-dimensional spatial semantic modeling on the multimodal feature vector using the SpatialLM model to obtain a spatial structure map; In this embodiment, the spatial structure map construction module 402 includes: A training submodule 4021 is used to train the SpatialLM model using a large-scale three-dimensional spatial data set that meets preset requirements to obtain a pre-trained SpatialLM model; A fine-tuning submodule 4022 is used to introduce traffic rule constraints and vehicle dynamics model constraints to fine-tune the pre-trained SpatialLM model to obtain a fine-tuned SpatialLM model; The spatial structure map generation submodule 4023 is used to encode the multimodal feature vector using the fine-tuned SpatialLM model to generate a feature vector containing spatial position and semantic information; based on the generated feature vector containing spatial position and semantic information, the self-attention mechanism is used to obtain the dependency relationship between different regions in the space to generate a spatial structure map with functional labels; The spatial structure graph with functional labels includes: a spatial structure graph with functional labels including express delivery vehicle passable areas, dangerous areas, and traffic nodes; The spatial structure graph includes nodes and edges; wherein the nodes represent spatial entities, including location coordinates, semantic categories, and attribute information; and the edges represent semantic relationships and spatial connection relationships between spatial entities.

[0041] A path planning module 403 is configured to plan a path for a target vehicle using a path planning algorithm based on a spatial structure map and a traffic rule knowledge base that meets preset requirements; In this embodiment, the path planning module 403 includes: The inference result acquisition submodule 4031 is used to build a graph neural network model and use the graph neural network model to infer the spatial structure graph with functional labels to obtain inference results; wherein the inference results include: extracting implicit relationships between spatial entities and global structural features; The behavior decision submodule 4032 is configured to construct a driving behavior decision model based on a reinforcement learning method, with the driving safety, efficiency, and compliance of the target vehicle as the goals. The state space of the driving behavior decision model includes: the vehicle's current position, speed, posture, the spatial structure map of the surrounding environment, and relevant rules in the traffic rule knowledge base that meet preset requirements; the action space of the driving behavior decision model includes: driving actions such as acceleration, deceleration, steering, lane changing, and parking; and behavioral decisions are made using the constructed driving behavior decision model. The driving path generation submodule 4033 is used to make behavioral decisions based on the inference results of the graph neural network model and the driving behavior decision model. It adopts a path planning algorithm to perform path planning on a spatial structure map with functional labels to generate the optimal driving path from the starting point to the end point; at the same time, the generated optimal driving path is dynamically adjusted according to real-time traffic conditions and environmental changes.

[0042] A SpatialLM model fine-tuning module 404 is configured to fine-tune the SpatialLM model based on actual driving feedback data using a transfer learning method to obtain a fine-tuned SpatialLM model. In this embodiment, the fine-tuning SpatialLM model module 404 includes: The actual driving feedback data includes: raw data collected by each sensor, spatial structure map, execution results of behavioral decisions, actual driving status of the vehicle and change information of the traffic environment; Construct a SpatialLM model evaluation metric, and evaluate the fine-tuned SpatialLM model based on the constructed SpatialLM model evaluation metric. If the current SpatialLM model evaluation metric meets the preset requirements, stop fine-tuning the SpatialLM model. If the current SpatialLM model evaluation metric does not meet the preset requirements, continue to collect actual driving feedback data, and use the collected actual driving feedback data to continue fine-tuning the SpatialLM model until the evaluation metric of the fine-tuned SpatialLM model meets the preset requirements. The evaluation indicators include: the accuracy of the spatial structure map, the correctness of the behavioral decision and the efficiency of path planning.

[0043] See also Figure 5 Another embodiment of the target vehicle path planning device in the embodiment of the present invention includes: The feature extraction module 501 is used to trigger the synchronization mechanism so that each sensor collects data at the same time; extract features based on the raw data collected by each sensor to obtain multimodal feature data; and fuse the multimodal feature data to obtain a multimodal feature vector; A spatial structure map construction module 502 is used to perform three-dimensional spatial semantic modeling on the multimodal feature vector using the SpatialLM model to obtain a spatial structure map; A path planning module 503 is used to plan a path for a target vehicle using a path planning algorithm based on a spatial structure map and a traffic rule knowledge base that meets preset requirements; The visualization and human-computer interaction module 504 is used to present the spatial structure map in a visual form to the target vehicle remote monitoring system to enhance the transparency and trust of the system; In this embodiment, the visualization and human-computer interaction module 504 includes: Through a visual interface, the work process and decision-making basis are clearly displayed to remote monitoring personnel, such as the fusion results of multimodal data, the generation process of spatial structure maps, the calculation steps of path planning, etc., to improve the transparency of the system. At the same time, it provides model performance evaluation reports and historical data records to enable operators to understand the reliability and stability of the model and enhance their trust in the system.

[0044] above Figure 4 and Figure 5 The target vehicle path planning device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The electronic device in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0045] Figure 6is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device 700 may vary significantly due to different configurations or performance, and may include one or more processors (central processing units, CPUs) 710 (for example, one or more processors), a memory 720, and one or more storage media 730 (for example, one or more mass storage devices) storing application programs 733 or data 732. The memory 720 and storage medium 730 may be either transient or persistent storage. The program stored in the storage medium 730 may include one or more modules (not shown), each of which may include a series of instruction operations on the electronic device 700. Furthermore, the processor 710 may be configured to communicate with the storage medium 730 to execute the series of instruction operations in the storage medium 730 on the electronic device 700.

[0046] The electronic device 700 may further include one or more power supplies 740, one or more wired or wireless network interfaces 750, one or more input and output interfaces 750, and / or one or more operating systems 731, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be understood by those skilled in the art that Figure 6 The illustrated electronic device structure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0047] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, cause the computer to execute the steps of the target vehicle path planning method.

[0048] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0049] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0050] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A target vehicle path planning method, characterized in that: include: Trigger the synchronization mechanism so that each sensor collects data at the same time; Based on the original data collected by each sensor, feature extraction is performed to obtain multimodal feature data; the multimodal feature data is fused to obtain a multimodal feature vector; The SpatialLM model is used to perform three-dimensional spatial semantic modeling on the multimodal feature vector to obtain a spatial structure map; The target vehicle's path is planned using a path planning algorithm based on the spatial structure graph and a traffic rules knowledge base that meets preset requirements; The sensors include: camera, lidar, millimeter wave radar, GNSS and IMU.

2. The target vehicle path planning method according to claim 1, characterized in that: The trigger synchronization mechanism enables each sensor to collect data at the same time, including: When collecting data, each sensor adds a timestamp with a timestamp accuracy that meets the preset requirements; for data with a timestamp deviation, a trigger signal is sent to all sensors through a synchronous trigger device so that each sensor collects data at the same time.

3. The target vehicle path planning method according to claim 1, characterized in that: The feature extraction based on the raw data collected by each sensor to obtain multimodal feature data includes: Based on the image data collected by the camera, the collected image data is preprocessed including denoising, grayscale conversion and normalization to obtain preprocessed image data; the preprocessed image data is extracted with a convolutional neural network; the visual features include: edges, textures, shape features in the image, and target features including vehicles, pedestrians, and traffic signs; LiDAR-based point cloud data is collected and converted into a bird's-eye view or voxel grid format. The spatial geometric features of the point cloud are extracted using a point cloud-specific neural network. These features include point positions, normal vectors, curvature, and object shape and size features represented by point cloud clusters. These point cloud-specific neural networks include PointNet and PointNet++. The millimeter-wave radar collects echo signals, performs filtering, noise removal, and clutter preprocessing on the collected echo signals to obtain preprocessed echo signals. The preprocessed echo signals are converted from time domain signals to frequency domain signals through Fourier transform, and target position and motion parameter features are extracted based on the frequency domain signals. The vehicle's position and speed information is collected based on GNSS, and the collected vehicle position and speed information is decoded and positioned to obtain the vehicle's latitude and longitude, altitude, and speed information; The acceleration, angular velocity and attitude information of the vehicle are collected based on the IMU, and the collected acceleration, angular velocity and attitude information of the vehicle are preprocessed including denoising and zero bias correction to obtain the preprocessed vehicle acceleration, angular velocity and attitude information; The vehicle's latitude and longitude, altitude and speed information are fused with the pre-processed vehicle acceleration, angular velocity and attitude information to obtain the vehicle's attitude and position information; the vehicle's motion state characteristics are extracted based on the vehicle's attitude and position information.

4. The target vehicle path planning method according to claim 1, characterized in that: The multimodal feature data is fused to obtain a multimodal feature vector. include: Multimodal feature data are fused by splicing or weighted fusion to obtain a multimodal feature vector; The multimodal feature data is fused by weighted fusion, including: A data fusion weight distribution model is established; the weight distribution model dynamically adjusts the fusion weight of each sensor data according to the performance of different sensors in different environments.

5. The target vehicle path planning method according to claim 1, characterized in that: The method of using the SpatialLM model to perform three-dimensional spatial semantic modeling on the multimodal feature vector to obtain a spatial structure map includes: The SpatialLM model is trained using a large-scale three-dimensional spatial dataset that meets the preset requirements to obtain a pre-trained SpatialLM model; Traffic rule constraints and vehicle dynamics model constraints are introduced to fine-tune the pre-trained SpatialLM model to obtain a fine-tuned SpatialLM model. The fine-tuned SpatialLM model encodes the multimodal feature vectors to generate feature vectors containing spatial location and semantic information. Based on the generated feature vectors containing spatial location and semantic information, the self-attention mechanism is used to obtain the dependencies between different regions in the space and generate a spatial structure map with functional labels. The spatial structure graph with functional labels includes: a spatial structure graph with functional labels including express delivery vehicle passable areas, dangerous areas, and traffic nodes; The spatial structure graph includes nodes and edges; wherein the nodes represent spatial entities, including location coordinates, semantic categories, and attribute information; and the edges represent semantic relationships and spatial connection relationships between spatial entities.

6. The target vehicle path planning method according to claim 1, characterized in that: The path planning algorithm is used to plan the path of the target vehicle based on the spatial structure map and the traffic rule knowledge base that meets the preset requirements, including: Construct a graph neural network model and use it to infer the spatial structure graph with functional labels to obtain inference results; wherein the inference results include: extracting implicit relationships between spatial entities and global structural features; Based on reinforcement learning methods, a driving behavior decision model is constructed with the goal of ensuring the safety, efficiency, and compliance of the target vehicle. The state space of the driving behavior decision model includes the vehicle's current position, speed, posture, the spatial structure map of the surrounding environment, and relevant rules in a traffic rule knowledge base that meet preset requirements. The action space of the driving behavior decision model includes driving actions such as acceleration, deceleration, steering, lane changing, and parking. Behavioral decisions are made using the constructed driving behavior decision model. Based on the inference results of the graph neural network model and the behavioral decisions of the driving behavior decision model, a path planning algorithm is used to perform path planning on a spatial structure map with functional labels to generate the optimal driving path from the starting point to the end point; at the same time, the generated optimal driving path is dynamically adjusted according to real-time traffic conditions and environmental changes.

7. The target vehicle path planning method according to claim 1, characterized in that: The method further includes: using a transfer learning method to fine-tune the SpatialLM model based on actual driving feedback data to obtain a fine-tuned SpatialLM model; The actual driving feedback data includes: raw data collected by each sensor, spatial structure map, execution results of behavioral decisions, actual driving status of the vehicle and information on changes in the traffic environment; Construct a SpatialLM model evaluation metric, and evaluate the fine-tuned SpatialLM model based on the constructed SpatialLM model evaluation metric. If the current SpatialLM model evaluation metric meets the preset requirements, stop fine-tuning the SpatialLM model. If the current SpatialLM model evaluation metric does not meet the preset requirements, continue to collect actual driving feedback data, and use the collected actual driving feedback data to continue fine-tuning the SpatialLM model until the evaluation metric of the fine-tuned SpatialLM model meets the preset requirements. The evaluation indicators include: the accuracy of the spatial structure map, the correctness of behavioral decisions, and the efficiency of path planning.

8. A target vehicle path planning device, characterized in that: include: Feature extraction module, used to trigger the synchronization mechanism so that each sensor collects data at the same time; Based on the original data collected by each sensor, feature extraction is performed to obtain multimodal feature data; the multimodal feature data is fused to obtain a multimodal feature vector; The spatial structure map construction module is used to use the SpatialLM model to perform three-dimensional spatial semantic modeling on multimodal feature vectors to obtain a spatial structure map; The path planning module is used to plan the path of the target vehicle through a path planning algorithm based on the spatial structure map and the traffic rules knowledge base that meets the preset requirements; The sensors include: camera, lidar, millimeter wave radar, GNSS and IMU.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the target vehicle path planning method according to any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by a processor, the steps of the target vehicle path planning method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Indoor parking lot vehicle navigation method, equipment and computer program product

    CN121171056A

  • Intelligent driving method and system based on multi-mode memory assistance and vehicle

    CN121224744A

  • Automatic driving path planning method based on multi-sensor fusion

    CN121323672A