Large-language-model-driven space-ground integrated automatic driving intelligent decision-making system and method
Through the integrated world-wide autonomous driving intelligent decision-making system driven by a large language model, combined with low-orbit satellite positioning and vehicle-end sensor data, the cloud-based large language model optimizes decision-making, solving the accuracy and reliability problems of existing autonomous driving technology in complex environments, and achieving efficient and accurate autonomous driving decisions.
Patent Information
- Application Number
- CN202510247767.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
When existing autonomous driving technology faces complex road conditions, emergencies and long-tail events, the sensor field of vision is limited, the accuracy is disturbed, and the learning ability is insufficient, making it difficult to cope with the diversity and uncertainty of the real world.
A large language model-driven intelligent decision-making system for integrated autonomous driving is designed to provide high-precision positioning information through low-orbit satellites, the vehicle side collects and analyzes environmental data in real time, and the cloud uses the large language model to optimize decision-making, forming a tight closed-loop interactive process.
It realizes more comprehensive and accurate environmental perception, improves the accuracy and reliability of autonomous driving decisions, reduces the requirements for vehicle-side sensor configuration, reduces hardware costs and data processing complexity, and improves the overall operating efficiency of the system.
Smart Images

Figure CN120183180A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of intelligent transportation systems and autonomous driving, and particularly to a space-earth integrated autonomous driving intelligent decision-making system and method driven by a large language model. Background Art
[0002] As a revolutionary breakthrough in the transportation field, autonomous driving technology has received extensive attention and made remarkable progress. Despite many achievements, existing methods often expose problems such as insufficient learning ability and poor generalization performance when facing complex road conditions, emergencies, and long-tail events, because the sensor field of view is limited and the accuracy is interfered, making it difficult to cope with the diversity and uncertainty of the real world.
[0003] With the rapid development of artificial intelligence technology, large language models have shown great application potential in the field of autonomous driving decision-making due to their powerful language understanding, reasoning, and generation capabilities. Large language models can process and understand massive text information, have the ability to logically analyze and judge complex scenarios, and can provide rich knowledge support and intelligent guidance for autonomous driving decision-making. However, deploying large language models on the vehicle side to process massive sensor data in real time requires extremely high computing power for in-vehicle computing units, increasing costs and power consumption. At the same time, space-earth integration brings new opportunities for autonomous driving. As a key component of the space-earth integrated network, low-earth orbit satellites can provide high-precision positioning information, have a wide coverage range, are not affected by ground obstacles, and can effectively make up for the limitations of ground sensors in complex environments, providing more comprehensive and accurate spatio-temporal information for autonomous driving vehicles and helping the vehicles to understand their own positions and surrounding situations in real time. In this context, the present invention proposes a space-earth integrated autonomous driving intelligent decision-making system driven by a large language model, which includes a satellite layer, a vehicle side layer, and a cloud layer. By providing wide-area spatio-temporal information through low-earth orbit satellites, the vehicle side collects and analyzes environmental data in real time, and the cloud uses the knowledge reserve and powerful computing power of large language models to optimize decisions, providing an innovative solution to the autonomous driving decision-making dilemma. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, the present invention provides a space-earth integrated autonomous driving intelligent decision-making system driven by a large language model, which gives full play to the advantages of space-earth cooperation, effectively integrates multi-source information, and realizes high-precision and high-reliability autonomous driving decision-making.
[0005] To achieve the above technical objectives, the present invention adopts the following technical solutions:
[0006] The present invention designs a space-earth integrated autonomous driving intelligent decision-making system (hereinafter referred to as the space-earth integrated decision-making system), and the implementation method of the decision-making system includes the following steps:
[0007] Step 1: Construct the satellite layer as the high-altitude information node of the space-ground integrated system. The satellite layer consists of a low-earth orbit satellite constellation;
[0008] Step 2: Construct the vehicle terminal layer as the direct executor of the space-ground integrated system. The vehicle terminal layer is equipped with multiple sensors and has a preliminary decision-making function;
[0009] Step 3: Construct the cloud layer as the central brain of the space-ground integrated system, with supercomputing power and massive data storage capabilities;
[0010] Step 4: Form a tight closed-loop interaction process among the layers constructed in Step 1, Step 2, and Step 3.
[0011] Furthermore, the satellite layer in Step 1 consists of a low-earth orbit satellite constellation and can provide centimeter-level high-precision positioning information.
[0012] The positioning information covers dynamic features and static features:
[0013] In terms of dynamic features, it includes the running speed, acceleration, and change in heading angle of the vehicle. Real-time monitoring of the running speed of the vehicle can intuitively reflect how fast the vehicle is traveling at present. The acceleration feature is closely related to the dynamic changes in the vehicle's power performance and driving state. The change rate of the heading angle reflects the rate of change of the vehicle's driving direction and is crucial for judging the vehicle's turning intention and handling flexibility.
[0014] Static features include the absolute position information and relative position features of the vehicle. With the cooperation of a high-precision map, the vehicle can accurately know its own lane, road section, and the relative position relationship with surrounding landmark buildings and traffic facilities, as well as the spatial position distribution among other surrounding traffic participants.
[0015] The positioning information composed of dynamic features and static features is interrelated and synergistic, jointly providing all-round and multi-level information support for autonomous driving decision-making;
[0016] Furthermore, the vehicle terminal layer in Step 2 is equipped with a variety of sensors, including lidar, cameras, and millimeter-wave radars. These sensors collect environmental data around the vehicle in real time from different dimensions.
[0017] Lidar can accurately construct a three-dimensional point cloud map around the vehicle.
[0018] Cameras capture visual images for identifying targets such as traffic signs, signal lights, pedestrians, and other vehicles.
[0019] Millimeter-wave radars focus on detecting the speed and distance information of target objects.
[0020] The in-vehicle layer incorporates an end-to-end rapid decision-making system that uses an efficient end-to-end model to deeply fuse, analyze, and infer multi-source sensor data. Combining the high-precision positioning information provided by low-orbit satellites, it quickly generates driving decisions suitable for the current complex road conditions, including speed and steering. The end-to-end rapid decision-making system is implemented as a multi-modal fusion Transformer architecture. Specifically:
[0021] The input features include image data I ∈ R H×W×C , lidar point cloud data P1 ∈ R N×C , millimeter-wave radar point cloud data P2 ∈ R N×C and high-precision positioning information M, where H and W are the length and width of the image, C is the number of channels, and N is the number of points in the point cloud.
[0022] For the image data I, it is divided into multiple small blocks, and each small block is regarded as a sequence element F I , which is mapped to a feature vector of a fixed dimension through linear projection, and position encoding is added to retain spatial position information.
[0023] For the features P1 and P2 of the lidar and millimeter-wave radar, they are converted into a sequence of feature vectors F P , and corresponding sensor embeddings are added to distinguish features from different sensor sources.
[0024] For the high-precision positioning information, based on longitude, latitude, heading angle, etc., it is encoded into a position feature sequence F M .
[0025] Subsequently, the above sequences are used as inputs to the multi-modal fusion Transformer. In the Transformer encoder, the multi-head self-attention mechanism is used to interact and fuse the feature sequences of different sensors. The multi-head self-attention mechanism can simultaneously focus on different subspaces of the input sequence, capture the complex correlations between different sensor features, and obtain the fused features. Specifically:
[0026] First, the input sequence is linearly transformed to generate query (Q), key (K), and value (V) matrices, and each head corresponds to an independent weight matrix. By splitting the Q, K, and V matrices according to the number of heads, each head processes a sub-vector of dimension d model / h. When calculating the attention, each head performs a scaled dot-product attention operation in parallel: for the i-th head, its attention score is:
[0027]
[0028] where softmax is the activation function. Subsequently, the fused features are obtained, which can be expressed as:
[0029] output i= score i V i
[0030] Then, the features after encoder fusion are input into the Transformer decoder, and combined with the preset driving goals to generate continuous decision control instructions such as speed and steering angle;
[0031] Furthermore, the cloud layer in step 3 is responsible for managing and maintaining a huge database, providing data support for the vehicle terminal and the satellite terminal. Specifically:
[0032] The cloud storage architecture adopts a distributed file system, which has high scalability, high availability, and high-performance data storage capabilities. These distributed file systems store data dispersedly on multiple nodes, and through the data redundancy backup strategy, ensure that data can still be obtained completely when some nodes fail, guaranteeing the continuous operation of the system.
[0033] For the stored data, it is subdivided into multiple types according to its characteristics and uses: high-precision maps, historical traffic data, real-time road conditions information, and various updated traffic regulation knowledge.
[0034] Among them, the high-precision map data covers detailed information such as road topology, lane line information, and traffic sign positions, and is stored in the form of vector maps for quick query and precise navigation.
[0035] The historical traffic data includes traffic flow, congestion conditions, accident records, etc. at different times and sections, providing references for road condition prediction and driving route planning, and is stored in a time series database according to time series to support efficient time range queries.
[0036] The real-time road conditions information focuses on the dynamic changes of the current road, such as road construction, temporary control, traffic jams caused by emergencies, etc., and is obtained through multiple sources such as real-time reporting from the vehicle terminal and satellite remote sensing monitoring, and is updated in real time in the form of a message queue to ensure the timeliness of the data.
[0037] Various traffic regulations and driving common sense documents are stored in text form for the large language model to retrieve and learn at any time to update its knowledge reserve.
[0038] When the data grows massively, the distributed system can expand by adding storage nodes to meet the ever-increasing data storage requirements;
[0039] Furthermore, the cloud layer in step 3 runs a powerful cloud large language model. By deeply mining and learning the massive data, it continuously optimizes its own knowledge system, provides knowledge distillation for the vehicle terminal model, and improves the decision-making ability of the vehicle terminal model. Specifically:
[0040] The large language model adopts a GPT-based architecture, and its input is a large amount of multi-source data, including a large amount of text knowledge, traffic rules, accident cases, and scene response logic on the one hand, and on the other hand, it integrates the actual driving data uploaded by the vehicle end, including sensor environment data and decision sequences, enabling the model to closely fit the real driving situation.
[0041] In terms of the model training algorithm, a pre-training - fine-tuning paradigm based on the Transformer architecture is adopted.
[0042] In the pre-training stage, large-scale unsupervised corpora are used for pre-training to enable the model to learn general language patterns and semantic understanding capabilities.
[0043] In the fine-tuning stage, combined with specific autonomous driving tasks, actual scenario data is used for fine-tuning to strengthen the model's risk recognition and decision optimization capabilities;
[0044] Furthermore, the cloud layer in step 3 undertakes the tasks of remote monitoring and management of the entire system, collects in real-time the driving data, sensor status, and decision execution feedback information uploaded by the vehicle end, evaluates and warns the vehicle operation status. When the vehicle driving scenario is safe, it complies with the vehicle's decision. When potential risks or abnormal situations are detected, the decision of the large language model is adopted, and optimization instructions, updated model parameters, or emergency disposal plans are pushed to the vehicle end in real-time to assist the vehicle end in correcting the decision. When the system is unable to make a decision in an extremely complex emergency scenario (such as a traffic accident), the ultimate decision of emergency braking and pulling over to the side of the road is initiated;
[0045] Furthermore, the specific content of the closed-loop interaction process in step 4 is as follows:
[0046] The satellite layer continuously transmits high-precision positioning information and macro environment data to the vehicle end and the cloud.
[0047] Based on this information, the vehicle end collects and processes local sensor data in real-time, generates preliminary driving decisions using an end-to-end fast decision-making system, and simultaneously uploads the driving data, environment perception data, and decision-making requirements to the cloud in real-time.
[0048] The cloud combines the data stored in itself and its powerful computing capabilities, comprehensively analyzes the information uploaded by the vehicle end, and evaluates the vehicle operation situation. When the driving scenario is safe, it directly adopts the decision of the vehicle end. When the driving scenario is abnormal, it uses the cloud large language model to provide knowledge enhancement and model optimization suggestions for the vehicle end, and feeds back the processing results to the vehicle end in real-time to assist the vehicle end in correcting the decision.
[0049] The vehicle end further optimizes the driving operation according to the cloud feedback and feeds back the execution result to the cloud again, realizing the continuous learning and dynamic optimization of the system, thus ensuring the safe and efficient operation of autonomous vehicles in various complex scenarios;
[0050] Advantages of the present invention:
[0051] 1) The present invention proposes a space-ground integrated autonomous driving intelligent decision-making system driven by a large language model. By making full use of the high-precision positioning information of low-earth orbit satellites and organically combining it with the rapid decision-making system at the vehicle end on the ground and the large language model in the cloud, the autonomous driving decision-making system can perceive the environment more comprehensively and accurately, and respond to complex road conditions more intelligently and flexibly, thereby improving the accuracy and reliability of decision-making.
[0052] 2) The present invention constructs an end-to-end decision optimization closed-loop system, realizing the full-process automation and intelligence from input information collection, model decision generation to execution result feedback. The system evaluates the decision-making effect in real time according to the actual driving state of the vehicle fed back by the actuator. Once a deviation or anomaly is found, it immediately triggers the re-decision mechanism of the large language model and conducts targeted optimization and adjustment in combination with space-ground integrated information to ensure that the vehicle always travels along the optimal trajectory, comprehensively improving the safety and stability of autonomous driving.
[0053] 3) Supported by high-precision positioning, the present invention effectively reduces the configuration requirements for vehicle-end sensors, which can not only reduce the hardware cost of the vehicle, but also reduce the complexity of vehicle-end data processing, further improving the overall operation efficiency of the system.
[0054] 4) The multi-node parallel data processing design in the cloud layer of the present invention greatly improves the data access speed, reduces latency, and meets the strict real-time requirements of the autonomous driving system. At the same time, the redundant backup mechanism effectively resists risks such as hardware failures and network interruptions, ensuring data security and integrity, and laying a solid foundation for the stable operation of the autonomous driving decision-making system. Description of the Drawings
[0055] Figure 1 is an architecture diagram of a space-ground integrated autonomous driving intelligent decision-making system driven by a large language model provided by the present invention;
[0056] Figure 2 is a cloud distributed file storage architecture diagram provided by the present invention;
[0057] Figure 3 is a vehicle decision-making process flow chart provided by the present invention; Detailed Embodiments
[0058] The present invention will be further described below with reference to the drawings. The drawings are used to provide a further understanding of the embodiments of the present disclosure, and constitute a part of the specification, and are used together with the embodiments of the present disclosure to explain the present disclosure, and do not constitute a limitation to the present disclosure.
[0059] A space-ground integrated autonomous driving intelligent decision-making system driven by a large language model proposed by the present invention, as Figure 1As shown in the figure, it is an architecture diagram of a space-ground integrated autonomous driving intelligent decision-making system driven by a large language model provided by the implementation of this application, mainly including the following steps:
[0060] Step 1: Construct a satellite layer as the high-altitude information node of the space-ground integrated system. The satellite layer consists of a low-earth orbit satellite constellation;
[0061] Furthermore, the satellite layer consists of a low-earth orbit satellite constellation, which can provide centimeter-level high-precision positioning information.
[0062] Among them, the positioning information covers dynamic features and static features:
[0063] In terms of dynamic features, it includes the running speed, acceleration, and heading angle change of the vehicle.
[0064] Real-time monitoring of the running speed of the vehicle can intuitively reflect the current driving speed of the vehicle.
[0065] The acceleration feature is closely related to the dynamic changes in the power performance and driving state of the vehicle.
[0066] The rate of change of the heading angle reflects the rate of change of the driving direction of the vehicle, which is crucial for judging the turning intention and maneuverability of the vehicle.
[0067] The static features include the absolute position information and relative position features of the vehicle. With the cooperation of a high-precision map, the vehicle can accurately know its own lane, road section, and the relative position relationship with surrounding landmark buildings and traffic facilities, as well as the spatial position distribution among other surrounding traffic participants.
[0068] The positioning information composed of dynamic features and static features is interrelated and synergistic, jointly providing all-round and multi-level information support for autonomous driving decision-making;
[0069] Step 2: Construct a vehicle terminal layer as the direct executor of the space-ground integrated system. The vehicle terminal layer is equipped with multiple sensors and has a preliminary decision-making function;
[0070] Furthermore, the vehicle terminal layer is equipped with a variety of sensors, including lidar, cameras, and millimeter-wave radars, which collect environmental data around the vehicle in real time from different dimensions.
[0071] Lidar can accurately construct a three-dimensional point cloud map around the vehicle.
[0072] Cameras capture visual images for identifying targets such as traffic signs, signal lights, pedestrians, and other vehicles.
[0073] Millimeter-wave radars focus on detecting the speed and distance information of target objects.
[0074] An end-to-end fast decision-making system is built into the vehicle end. It uses an efficient end-to-end model to deeply fuse, analyze, and infer multi-source sensor data. Combining the high-precision positioning information provided by low-earth orbit satellites, it quickly generates driving decisions suitable for the current complex road conditions, including speed and steering, etc. The end-to-end fast decision-making system is implemented as a multi-modal fusion Transformer architecture. Specifically:
[0075] The input features include image data I ∈ R H×W×C 、lidar point cloud data P1 ∈ R N×C 、millimeter-wave radar point cloud data P2 ∈ R N×C and high-precision positioning information M, where H and W are the length and width of the image, C is the number of channels, and N is the number of points in the point cloud.
[0076] For the image data I, it is divided into multiple small blocks, and each small block is regarded as a sequence element F I . It is mapped to a feature vector of a fixed dimension through linear projection, and positional encoding is added to retain the spatial position information.
[0077] For the features P1 and P2 of the lidar and millimeter-wave radar, they are converted into a sequence of feature vectors F P and the corresponding sensor embeddings are added to distinguish the features from different sensor sources.
[0078] For the high-precision positioning information, based on longitude, latitude, heading angle, etc., it is encoded into a sequence of position features F M .
[0079] Subsequently, the above sequences are used as the input of the multi-modal fusion Transformer. In the Transformer encoder, the multi-head self-attention mechanism is used to interact and fuse the feature sequences of different sensors. The multi-head self-attention mechanism can simultaneously focus on different subspaces of the input sequence, capture the complex correlations between different sensor features, and obtain the fused features.
[0080] Here is a specific example analysis as follows. When processing a sequence containing vehicle image features, lidar point cloud features, and millimeter-wave radar target features, the model can learn the internal connection between the vehicle appearance features and the vehicle contour measured by the lidar and the vehicle speed detected by the millimeter-wave radar through the attention mechanism, so as to achieve more accurate object perception.
[0081] Specifically:
[0082] First, the input sequence is linearly transformed to generate query (Q), key (K), and value (V) matrices, and processed using the multi-head self-attention mechanism. Each head corresponds to an independent weight matrix. By splitting the Q, K, and V matrices according to the number of heads, each head processes a dimension of d modelsub-vectors per hour, where d model represents the overall dimension of the input, and h represents the number of heads. When calculating attention, each head performs the scaled dot-product attention operation in parallel: for the (i)-th head, its attention score is:
[0083]
[0084] where softmax is the activation function.
[0085] Subsequently, the fused feature output is obtained i , which can be expressed as:
[0086] output i = score i V i
[0087] Then, the fused features of the encoder are input into the Transformer decoder, and combined with the preset driving goals to generate continuous decision control instructions such as speed and steering angle;
[0088] Step 3, construct a cloud layer as the central brain of the space-ground integrated system, with supercomputing power and massive data storage capabilities;
[0089] Furthermore, the cloud layer is responsible for managing and maintaining a huge database, providing data support for the vehicle side and the satellite side. Specifically:
[0090] As shown in the cloud distributed file storage architecture diagram in Figure 2 , the cloud storage architecture adopts a distributed file system, with high scalability, high availability, and high-performance data storage capabilities.
[0091] These distributed file systems store data dispersedly on multiple nodes. Through the data redundancy backup strategy, it is ensured that data can still be obtained completely when some nodes fail, guaranteeing the continuous operation of the system.
[0092] For the stored data, it is classified into multiple types according to its characteristics and uses: high-precision maps, historical traffic data, real-time road conditions information, and various updated traffic regulation knowledge.
[0093] Among them, the high-precision map data covers detailed information such as road topology, lane line information, and traffic sign positions, and is stored in the form of vector maps for quick query and precise navigation.
[0094] The historical traffic data includes traffic flow, congestion conditions, accident records, etc. at different times and sections, providing references for road condition prediction and driving route planning, and is stored in a time series database according to time series to support efficient time range queries.
[0095] Real-time traffic information focuses on the dynamic changes in current roads, such as road construction, temporary controls, traffic jams caused by emergencies, etc. It is obtained through real-time reporting from the vehicle and multiple sources such as satellite remote sensing monitoring, and updated in real time in the form of message queues to ensure the timeliness of the data.
[0096] Various traffic regulations and driving knowledge documents are stored in text form, which makes it easy for the large language model to retrieve and learn at any time and update the knowledge reserve.
[0097] When data grows in large amounts, the distributed system can expand capacity by adding storage nodes to meet the ever-increasing data storage needs;
[0098] Furthermore, the cloud layer runs a powerful cloud language model, which continuously optimizes its own knowledge system through deep mining and learning of massive data, provides knowledge distillation for the vehicle-side model, and improves the decision-making ability of the vehicle-side model. Specifically:
[0099] The large language model adopts a GPT-based architecture, and its input is massive multi-source data.
[0100] On the one hand, it contains massive text knowledge, traffic rules, accident cases, and scenario response logic.
[0101] On the other hand, it integrates the actual driving data uploaded by the vehicle, including sensor environment data and decision sequences, so that the model closely fits the actual driving situation.
[0102] In terms of model training algorithm, the pre-training-fine-tuning paradigm based on the Transformer architecture is adopted.
[0103] The pre-training phase uses large-scale unsupervised corpus for pre-training, allowing the model to learn common language patterns and semantic understanding capabilities.
[0104] During the fine-tuning phase, specific autonomous driving tasks are combined with actual scenario data to fine-tune the model and enhance its risk identification and decision-making optimization capabilities.
[0105] Furthermore, the cloud layer undertakes the remote monitoring and management tasks of the entire system, collects driving data, sensor status and decision-making execution feedback information uploaded by the vehicle in real time, evaluates and warns the vehicle's operating status, and follows the vehicle's decision when the vehicle's driving scene is safe. When potential risks or abnormal situations are found, the decision of the large language model is adopted to push optimization instructions, updated model parameters or emergency response plans to the vehicle in real time to assist the vehicle in correcting its decision. When the system is unable to make a decision in an extremely complex emergency scenario (such as a traffic accident), the ultimate decision of emergency braking and pulling over is initiated;
[0106] Step 4: Form a tight closed-loop interactive process by combining the layers constructed in steps 1, 2, and 3.
[0107] Furthermore, as shown in the vehicle decision-making process flowchart of Figure 3 , the closed-loop interaction of the system and the vehicle decision-making process are as follows:
[0108] The satellite layer continuously transmits high-precision positioning information and macro-environment data to the vehicle terminal and the cloud.
[0109] Based on this information, the vehicle terminal collects and processes local sensor data in real time, generates preliminary driving decisions using the end-to-end fast decision-making system, and simultaneously uploads driving data, environmental perception data, and decision-making requirements to the cloud in real time.
[0110] Combining the data stored in itself and its powerful computing power, the cloud comprehensively analyzes the information uploaded by the vehicle terminal and evaluates the vehicle operation status. When the driving scenario is safe, it directly adopts the decision of the vehicle terminal. When the driving scenario is abnormal, it uses the cloud large language model to provide knowledge enhancement and model optimization suggestions for the vehicle terminal, and feeds back the processing results to the vehicle terminal in real time to assist the vehicle terminal in correcting the decision.
[0111] The vehicle terminal further optimizes the driving operation according to the cloud feedback and feeds back the execution result to the cloud again, realizing the continuous learning and dynamic optimization of the system, so as to ensure the safe and efficient operation of the autonomous vehicle in various complex scenarios.
[0112] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation modes of the present invention, and they are not intended to limit the protection scope of the present invention. Any equivalent modes or changes made without departing from the technology created by the present invention should be included in the protection scope of the present invention.
Claims
1. A large language model-driven, integrated, autonomous driving intelligent decision-making system, characterized by: Including satellite layer, vehicle-side layer and cloud layer; The satellite layer, as a high-altitude information node of the integrated space-ground network, consists of a constellation of low-orbit satellites, which transmits high-precision positioning information to ground vehicles in real time, providing a time and space reference for the vehicles; The vehicle-side layer is the direct executor of the decision-making system. It is equipped with a variety of sensors to collect environmental data around the vehicle in real time. It uses high-precision positioning data provided by low-orbit satellites and the end-to-end rapid decision-making model built into the vehicle-side layer to generate preliminary driving decisions. At the same time, it uploads driving data, environmental perception data, and decision-making requirements to the cloud layer. The cloud layer is responsible for remote monitoring and management of the entire decision-making system, collecting driving data, perception data and decision-making execution feedback information uploaded by the vehicle in real time, evaluating and warning the vehicle's operating status, and following the vehicle's quick decision when the vehicle's driving scenario is safe. When potential risks or abnormal situations are found, a decision model based on a large language model is used to push optimization instructions, updated model parameters or emergency response plans to the vehicle in real time to assist the vehicle in correcting its decisions.
2. The large language model-driven, integrated ground-ground autonomous driving intelligent decision-making system according to claim 1, characterized in that: The high-precision positioning information transmitted by the satellite layer to the ground vehicle includes dynamic features and static features; Dynamic features include the vehicle's running speed, acceleration, and heading angle, which monitor the vehicle's running speed in real time and directly reflect the vehicle's current driving speed. The acceleration characteristics reflect the dynamic changes in the vehicle's power performance and driving state. The heading angle changes reflect the rate of change of the vehicle's driving direction, determine the vehicle's steering intention, and achieve control flexibility. Static features include the absolute and relative position information of the vehicle. Combined with high-precision maps, the vehicle can accurately know the lane and road section it is in, as well as its relative position relationship with surrounding landmarks and traffic facilities, as well as the spatial position distribution of other surrounding traffic participants. The positioning information composed of dynamic features and static features is interrelated and synergistic, jointly providing all-round and multi-level information support for autonomous driving decisions.
3. The large language model-driven, integrated ground-ground autonomous driving intelligent decision-making system according to claim 1, characterized in that: The vehicle-side layer is equipped with a variety of sensors, including: laser radar, camera, millimeter wave radar, which collects environmental data around the vehicle in real time from different dimensions; The laser radar can accurately construct a three-dimensional point cloud map around the vehicle; The camera captures visual images for identifying traffic signs, signal lights, pedestrians and other vehicle targets; The millimeter wave radar is used to detect the speed and distance information of the target object.
4. The large language model-driven, integrated ground-ground autonomous driving intelligent decision-making system according to claim 1, characterized in that: The vehicle-side layer has a built-in end-to-end rapid decision-making system, which uses an end-to-end model to deeply fuse, analyze and reason multi-source sensor data, and combines the high-precision positioning information provided by low-orbit satellites to quickly generate driving decisions that adapt to current complex road conditions, including speed and steering.
5. The large language model-driven integrated ground-ground autonomous driving intelligent decision-making system according to claim 4, characterized in that: The end-to-end fast decision-making system adopts a multi-modal fusion Transformer architecture, as follows: The input features include image data I∈R H×W×C , LiDAR point cloud data P1∈R N×C , millimeter wave radar point cloud data P2∈R N×C and high-precision positioning information M, where H and W are the length and width of the image, C is the number of channels, and N is the number of point clouds; For the image data I, it is divided into multiple small blocks, each of which is regarded as a sequence element F I , map it to a feature vector of fixed dimension by linear projection, and add position encoding to preserve the spatial location information; For the point cloud features P1 and P2 of the laser radar and millimeter wave radar, convert them into feature vector sequences F P , and add corresponding sensor embeddings to distinguish the features from different sensor sources; For high-precision positioning information, based on longitude, latitude and heading angle, it is encoded as a position feature sequence F M ; The above sequence is then used as the input of the multimodal fusion Transformer. In the Transformer encoder, the multi-head self-attention mechanism is used to interactively fuse the feature sequences of different sensors. The multi-head self-attention mechanism can simultaneously focus on different subspaces of the input sequence, capture the complex correlation between the features of different sensors, and obtain the fused features. The encoder fused features are input into the Transformer decoder, and combined with the preset driving goals to generate continuous decision instructions for speed and steering angle.
6. The large language model-driven, integrated ground-ground autonomous driving intelligent decision-making system according to claim 5, characterized in that: The Transformer encoder generates fused features as follows: First, the input sequence is linearly transformed to generate query (Q), key (K), and value (V) matrices. Each head corresponds to an independent weight matrix. By splitting the Q, K, and V matrices according to the number of heads, each head has a processing dimension of d model / h’s subvector, when calculating attention, each head performs the scaled dot product attention operation in parallel: for the i-th head, its attention score is: Among them, softmax is the activation function; Then the fused features are obtained, which can be expressed as: output i =score i V i 。 7. The large language model-driven, integrated ground-ground autonomous driving intelligent decision-making system according to claim 1, characterized in that: The storage architecture of the cloud layer adopts a distributed file system, which stores data in multiple nodes in a dispersed manner. Through a data redundancy backup strategy, it is ensured that when some nodes fail, the data can still be fully obtained, thus ensuring the continuous operation of the system; The stored data is divided into the following categories according to its characteristics and uses: high-precision maps, historical traffic data, real-time traffic information, and various updated traffic regulations knowledge; Among them, high-precision map data covers road topology, lane line information, and traffic sign locations, and is stored in the form of vector graphics to facilitate quick query and accurate navigation; Historical traffic data includes traffic volume, congestion conditions, and accident records at different time periods and road sections, providing a reference for road condition prediction and driving route planning. It is stored in a time series database in time series to support efficient time range queries. Real-time traffic information focuses on the dynamic changes of current roads, such as road construction, temporary control, and traffic jams caused by emergencies. It is obtained through real-time reporting from the vehicle end and satellite remote sensing monitoring and other multi-source channels, and is updated in real time in the form of message queues to ensure the timeliness of the data. Traffic regulations knowledge is stored in text form, which makes it easy for large language models to retrieve and learn at any time and update knowledge reserves.
8. The large language model-driven, integrated ground-ground autonomous driving intelligent decision-making system according to claim 1, characterized in that: The cloud layer includes a large language model, which continuously optimizes its own knowledge system through deep mining and learning of massive data, provides knowledge distillation for the vehicle-side decision model, and improves the decision-making ability of the vehicle-side decision model. The details are as follows: The large language model adopts a GPT-based architecture. Its input is a massive amount of multi-source data, which includes massive text knowledge, traffic rules, accident cases, and scenario response logic. On the other hand, it integrates the actual driving data uploaded by the vehicle, including sensor environment data and decision sequences, so that the model closely fits the actual driving situation. In the training of large language models, a pre-training-fine-tuning paradigm based on the Transformer architecture is adopted. In the pre-training stage, large-scale unsupervised corpus is used for pre-training, allowing the model to learn common language patterns and semantic understanding capabilities. During the fine-tuning phase, specific autonomous driving tasks are combined with actual scenario data for fine-tuning to enhance the model's risk identification and decision-making optimization capabilities.
9. A large language model-driven, ground-ground integrated autonomous driving intelligent decision-making method, characterized in that: The steps include: Step 1: Construct a satellite layer as a high-altitude information node of the integrated space-ground network system. The satellite layer consists of a low-orbit satellite constellation. The high-precision positioning information transmitted to ground vehicles includes dynamic and static features. Dynamic features include the vehicle's running speed, acceleration, and heading angle, which monitor the vehicle's running speed in real time and directly reflect the vehicle's current driving speed. The acceleration characteristics reflect the dynamic changes in the vehicle's power performance and driving state. The heading angle changes reflect the rate of change of the vehicle's driving direction, determine the vehicle's steering intention, and achieve control flexibility. Static features include the absolute and relative position information of the vehicle. Combined with high-precision maps, the vehicle can accurately know the lane and road section it is in, as well as its relative position relationship with surrounding landmarks and traffic facilities, as well as the spatial position distribution of other surrounding traffic participants. The positioning information composed of dynamic features and static features are interrelated and synergistic, providing all-round and multi-level information support for autonomous driving decisions; Step 2: Construct the vehicle-side layer. As the executor of the integrated space-ground network system, the vehicle-side layer is equipped with multiple sensors, including: laser radar, camera, millimeter wave radar, which collects environmental data around the vehicle in real time from different dimensions; the laser radar can accurately construct a three-dimensional point cloud map around the vehicle; The camera captures visual images for identifying traffic signs, signal lights, pedestrians and other vehicle targets; The millimeter wave radar is used to detect the speed and distance information of the target object; The vehicle-side layer has a built-in end-to-end rapid decision-making system, which uses the end-to-end decision-making system to deeply integrate, analyze and reason multi-source sensor data, and combines the high-precision positioning information provided by low-orbit satellites to quickly generate driving decisions that adapt to current complex road conditions, including speed and steering; The end-to-end fast decision-making system adopts a multi-modal fusion Transformer architecture, as follows: The input features include image data I∈R H×W×C , LiDAR point cloud data P1∈R N×C , millimeter wave radar point cloud data P2∈R N×C and high-precision positioning information M, where H and W are the length and width of the image, C is the number of channels, and N is the number of point clouds; For the image data I, it is divided into multiple small blocks, each of which is regarded as a sequence element F I , map it to a feature vector of fixed dimension by linear projection, and add position encoding to preserve the spatial location information; For the point cloud features P1 and P2 of the laser radar and millimeter wave radar, convert them into feature vector sequences F P , and add corresponding sensor embeddings to distinguish the features from different sensor sources; For high-precision positioning information, based on longitude, latitude and heading angle, it is encoded as a position feature sequence F M ; The above sequence is then used as the input of the multimodal fusion Transformer. In the Transformer encoder, the multi-head self-attention mechanism is used to interactively fuse the feature sequences of different sensors. The multi-head self-attention mechanism can simultaneously focus on different subspaces of the input sequence, capture the complex correlation between the features of different sensors, and obtain the fused features. The encoder fused features are input into the Transformer decoder, and combined with the preset driving goals, continuous decision instructions for speed and steering angle are generated; Step 3: Build a cloud layer as the central brain of the integrated space-ground network system, with super computing power and massive data storage capacity; the cloud layer remotely monitors and manages the entire decision-making system, collects driving data, perception data and decision-making execution feedback information uploaded by the vehicle in real time, evaluates and warns the vehicle's operating status, and follows the vehicle's quick decision when the vehicle's driving scene is safe. When potential risks or abnormal situations are found, a decision model based on a large language model is used to push optimization instructions, updated model parameters or emergency response plans to the vehicle in real time to assist the vehicle in correcting its decision; The layers constructed in step 4, step 1, step 2 and step 3 form a tight closed-loop interactive process; the details are as follows: The satellite layer continuously transmits high-precision positioning information and macro-environmental data to the vehicle layer and the cloud layer; Based on this information, the vehicle-side layer collects and processes local sensor data in real time, generates preliminary driving decisions using an end-to-end rapid decision-making system, and uploads driving data, environmental perception data, and decision-making requirements to the cloud in real time; The cloud layer combines its own stored data with powerful computing power to conduct a comprehensive analysis of the information uploaded by the vehicle and evaluate the vehicle's operating conditions. When the driving scene is safe, the vehicle's decision is directly adopted. When the driving scene is abnormal, the cloud-based large language model is used to provide knowledge enhancement and model optimization suggestions for the vehicle, and the processing results are fed back to the vehicle in real time to assist the vehicle in correcting its decision. The vehicle further optimizes driving decisions based on feedback from the cloud and feeds the execution results back to the cloud, achieving continuous learning and dynamic optimization to ensure the safe and efficient operation of autonomous vehicles in various complex scenarios.
10. The large language model-driven, integrated ground-ground autonomous driving intelligent decision-making method according to claim 9, characterized in that: In step 2, the Transformer encoder of the vehicle-side layer generates fused features. The specific method is as follows: First, the input sequence is linearly transformed to generate query (Q), key (K), and value (V) matrices. Each head corresponds to an independent weight matrix. By splitting the Q, K, and V matrices according to the number of heads, each head has a processing dimension of d model / h’s subvector, when calculating attention, each head performs the scaled dot product attention operation in parallel: for the i-th head, its attention score is: Among them, softmax is the activation function; Then the fused features are obtained, which can be expressed as: output i =score i V i ; The storage architecture of the cloud layer in step 3 adopts a distributed file system, which stores data in multiple nodes in a dispersed manner. Through a data redundancy backup strategy, it is ensured that when some nodes fail, the data can still be fully obtained, thereby ensuring the continuous operation of the system; The stored data is divided into the following categories according to its characteristics and uses: high-precision maps, historical traffic data, real-time traffic information, and various updated traffic regulations knowledge; Among them, high-precision map data covers road topology, lane line information, and traffic sign locations, and is stored in the form of vector graphics to facilitate quick query and accurate navigation; Historical traffic data includes traffic volume, congestion conditions, and accident records at different time periods and road sections, providing a reference for road condition prediction and driving route planning. It is stored in a time series database in time series to support efficient time range queries. Real-time traffic information focuses on the dynamic changes of current roads, such as road construction, temporary control, and traffic jams caused by emergencies. It is obtained through real-time reporting from the vehicle end and satellite remote sensing monitoring and other multi-source channels, and is updated in real time in the form of message queues to ensure the timeliness of the data. Traffic regulations knowledge is stored in text form, which makes it easy for large language models to retrieve and learn at any time and update knowledge reserves; The cloud layer in step 3 uses a large language model to continuously optimize its own knowledge system through deep mining and learning of massive data, provide knowledge distillation for the vehicle-side decision model, and improve the decision-making ability of the vehicle-side decision model. The details are as follows: The large language model adopts a GPT-based architecture. Its input is a massive amount of multi-source data, which includes massive text knowledge, traffic rules, accident cases, and scenario response logic. On the other hand, it integrates the actual driving data uploaded by the vehicle, including sensor environment data and decision sequences, so that the model closely fits the actual driving situation. In the training of large language models, a pre-training-fine-tuning paradigm based on the Transformer architecture is adopted; in the pre-training stage, large-scale unsupervised corpus is used for pre-training to allow the model to learn common language patterns and semantic understanding capabilities; in the fine-tuning stage, specific autonomous driving tasks are combined with actual scenario data for fine-tuning to enhance the model's risk identification and decision-making optimization capabilities.
Citation Information
Cited By
Decision-making method, device and equipment for vehicle auxiliary driving, storage medium and product
CN120382908A
Decision-making methods, devices, equipment, storage media and products for vehicle assisted driving
CN120382908B
Vehicle-mounted system modular framework and collaboration method thereof
CN120409536A
Unmanned equipment contact sense sensor device based on large model
CN120663362A
Internet of vehicles resource negotiation allocation method and system based on big language model
CN120785926A