Automatic driving large model framework based on 3D space-time perception and human-like decision-making reasoning
By adopting a large-scale model framework based on 3D spatiotemporal perception and human-like decision-making reasoning in the autonomous driving system, combining dynamic semantic patch embedding, multi-scale chain reasoning and multi-level spatiotemporal semantic adaptive data extraction paradigm, the existing system's insufficient semantic perception ability and uninterpretation of decisions in complex environments is solved, and more efficient and transparent autonomous driving decisions are achieved.
Patent Information
- Application Number
- CN202510645388.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing end-to-end autonomous driving system lacks semantic perception capabilities in a complex and changeable three-dimensional global environment, making it difficult to accurately predict the future status of the vehicle or adapt to isomerized traffic scenarios. At the same time, the decision-making process lacks interpretability, affecting public trust and legal compliance.
The autonomous driving big model framework based on 3D space-time perception and human-like decision-making reasoning is adopted. Through the deep integration of dynamic semantic patch embedding and multi-scale chain reasoning, the three-dimensional space-time semantic inference capabilities are improved, and through the multi-level spatio-temporal semantic adaptive data extraction paradigm based on rules-driven interaction with semantic actions, and the multi-level training matrix architecture that enhances scenario understanding, decision transparency and end-to-end driving task optimization capabilities.
It significantly improves the three-dimensional spatio-temporal semantic reasoning capabilities of the autonomous driving system in complex scenarios, improves the accuracy of dynamic behavior deduction and decision interpretability, and enhances the practicality and safety of the system.
Smart Images

Figure CN120182938A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving, and particularly relates to an autonomous driving large model framework based on 3D spatio-temporal perception and human-like decision-making reasoning. Background Art
[0002] In recent years, with the rapid development of deep learning and foundation model technologies, multi-modal large language models (MLLMs) have shown important application prospects in the field of end-to-end autonomous driving due to their powerful cross-modal semantic reasoning capabilities and generalization performance. End-to-end autonomous driving aims to directly analyze in-vehicle multi-source sensor data (such as multi-view dynamic video streams, lidar point clouds, and bird's-eye views) to generate path planning or low-level control instructions for the vehicle, achieving full-link integration from environmental perception to decision output. Compared with traditional modular systems, it can significantly reduce error transmission and computational redundancy in intermediate links. However, this technology faces multiple technical challenges in actual deployment, and scene semantic understanding and decision interpretability have become key bottlenecks.
[0003] Traditional autonomous driving systems mostly adopt modular small neural network architectures, such as independent perception units (extracting two-dimensional features based on convolutional networks), planning units (generating paths relying on rule-driven methods), and control units (outputting instructions through classical controllers), which operate through interface cascading. This design is functionally limited to specific operating scenarios (such as urban structured roads or highways), and its reasoning ability is limited by the deficiencies in model capacity and data dimensions, making it difficult to handle complex and changing three-dimensional global environments. In addition, most existing end-to-end methods use single-frame front-view images as the main input, lacking the ability to deeply model long-range spatio-temporal relationships, resulting in insufficient semantic perception of three-dimensional dynamic / static scenes and difficulty in accurately predicting the future state of the vehicle or adapting to heterogeneous traffic scenarios.
[0004] The "black box" nature of the decision-making process further exacerbates the practical problems of end-to-end autonomous driving. Although the full-stack integrated design improves system efficiency, its internal reasoning logic lacks interpretability and is difficult to intuitively reveal the decision-making process. This opacity not only weakens the public's trust in autonomous driving systems but also poses challenges in legal compliance and safety certification. Existing research has tried to alleviate this problem in various ways, such as using visual spatial topology maps to display perception results or generating intermediate semantic representations (such as object detection bounding boxes) to explain behavioral intentions, but these methods have a high understanding threshold for users and limited practicality. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide an autonomous driving large model framework based on 3D spatio-temporal perception and human-like decision-making reasoning, aiming to solve the problems raised in the above background art.
[0006] The embodiments of the present invention are implemented as follows. Based on the 3D spatio-temporal perception and humanoid decision-making reasoning-based autonomous driving large model framework, the construction steps of the autonomous driving large model framework are as follows: Step 1: Construction of an autonomous driving vision-language framework for dynamic semantic patch embedding and multi-scale chain reasoning; By designing the AutoSceneNet system, a multi-modal large language model architecture with humanoid multi-scale reasoning ability is constructed. The video-image dynamic semantic patch embedding joint encoder is used to extract features from the full-view scene video stream and the global bird's-eye view, generating dynamic video sequences and global spatial features. Subsequently, through an efficient cross-modal semantic alignment mechanism, the visual features are mapped to a unified semantic representation space, and projected to the high-dimensional language embedding domain using a shared vision-language connector to generate unified visual tokens. After being concatenated with the text tokens encoded by the text encoder, they are input into the large language model backbone for multi-modal reasoning. At the same time, after the output of the shared vision-language connector is averaged pooled, a multi-task perception head is used to generate outputs aligned with the panoramic semantic map, traffic element perception, and visual entity localization labels. Finally, a natural language response is generated through the Q-Former classification decoder, including three-dimensional scene parsing, semantic action inference, behavioral logic interpretation, motion trajectory planning, and driving control instructions; Step 2: A multi-level spatio-temporal semantic adaptive data extraction paradigm based on rule-driven and semantic action interaction; Through a rule-driven approach, the generation of semantic actions is guided by the lateral and longitudinal speed targets reflecting the motion intention of the vehicle itself and the driving direction intention providing directional information. Subsequently, information such as moving bodies and interactors, infrastructure and control bodies, and environmental factors and road conditions in the scene are normalized, and through multi-view image stitching and semantic map construction, hierarchical scene semantic images are generated. These images and semantic actions interact with accurate object detection results to form multi-modal semantic understanding, and finally, semantic action interpretations and adaptive guidance instructions are generated. Based on this, an interpretable end-to-end autonomous driving comprehensive dataset AutoSceneVQA is developed, which provides sufficient semantic support while ensuring that the model can perform tasks in a clear and transparent manner during the training process, providing a cross-modal semantic basis for the model training to support three-dimensional scene semantic perception and long-axis time-series task reasoning; Step 3: Design of a cross-modal task-oriented multi-level training matrix architecture; A multi-level training matrix architecture oriented to cross-modal task is designed, covering multiple key task modules such as cross-modal feature pre-training, 3D scene perception pre-training, end-to-end driving warm-up fine-tuning, and comprehensive fine-tuning. In this architecture, cross-modal feature pre-training jointly learns the deep alignment of visual and language features on multi-source datasets. 3D scene perception pre-training relies on large-scale 3D datasets to further enhance the model's 3D spatial semantic insight ability in complex driving environments by learning semantic structures and dynamic interactions in 3D environments. End-to-end driving warm-up fine-tuning focuses on quickly adapting to driving tasks using pre-trained knowledge to optimize the full link performance from perception to decision-making. Finally, through the comprehensive fine-tuning link, the model's ability in long-time sequence task reasoning and decision-making optimization is further improved. At the same time, in order to further improve the training effect, a loss function is specially designed, which ensures efficient collaborative learning and target guidance between different tasks by adaptively weighted fusion of the optimization objectives of each task module.
[0007] The self-driving large model framework based on 3D spatio-temporal perception and human-like decision-making reasoning provided by the embodiments of the present invention effectively improves the 3D spatio-temporal semantic reasoning ability of the self-driving system in complex scenarios through the deep integration of cross-modal dynamic semantic patch embedding and multi-scale chain reasoning, and realizes the efficient collaborative processing of full-view scene video streams and global bird's-eye views. At the same time, through the multi-level spatio-temporal semantic adaptive data extraction paradigm based on rule-driven and semantic action interaction and the generation of semantic action representations, combined with the multi-level training matrix architecture, the model performs excellently in scene understanding, decision transparency, and end-to-end driving task optimization. Compared with traditional methods, the accuracy of the system in dynamic behavior deduction is improved by about 15%, the decision interpretability is improved by 20%, and better performance is achieved in the path planning and behavior prediction tasks on the nuScenes dataset, providing strong technical support for the global path planning and human-like decision-making of self-driving. Brief Description of the Drawings
[0008] Figure 1 It is the process of the AutoSceneNet inference framework; Figure 2 It is the framework of AutoSceneNet; Figure 3 It is the development process of the multi-level spatio-temporal semantic adaptive data extraction paradigm based on rule-driven and semantic action interaction; Figure 4 It is the multi-level training matrix architecture oriented to cross-modal tasks; Figure 5 It is the hierarchical visualization in complex scenarios. Detailed Embodiments
[0009] In order to make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0010] The following describes the specific implementation of the present invention in detail with reference to specific embodiments.
[0011] As Figure 1 and Figure 2 shown, it is a large-scale autonomous driving model framework based on 3D spatio-temporal perception and humanoid decision-making reasoning provided by an embodiment of the present invention. It systematically constructs an end-to-end multi-modal semantic reasoning and collaborative optimization training framework for autonomous driving. With the AutoSceneNet system as the core, the driving decision-making process is finely deconstructed into a progressive technical link of constructing an autonomous driving vision-language framework with dynamic semantic patch embedding and multi-scale chain reasoning, developing a multi-level spatio-temporal semantic adaptive data extraction paradigm based on rule-driven and semantic action interaction, and optimizing the design of a cross-modal task-oriented multi-level training matrix architecture. Through deep spatio-temporal semantic modeling and humanoid reasoning link simulation, it realizes the full-link mapping from full-view scene perception to accurate control instruction generation, effectively improving the semantic integration efficiency of multi-modal data and the interpretability of decisions, and providing a new technical direction for the fields of end-to-end autonomous driving and multi-modal large language models. The specific technical process of the present invention is as follows: Step 1: Construction of an autonomous driving vision-language framework with dynamic semantic patch embedding and multi-scale chain reasoning; By designing the AutoSceneNet system, it aims to construct a multi-modal large language model architecture with humanoid multi-scale reasoning ability. It uses a video-image dynamic semantic patch embedding joint encoder to extract features from the full-view scene video stream and the global bird's-eye view, generating a dynamic video sequence and global spatial features. Subsequently, through an efficient cross-modal semantic alignment mechanism, the visual features are mapped to a unified semantic representation space, and projected to a high-dimensional language embedding domain using a shared vision-language connector to generate unified visual tokens. After being concatenated with the text tokens encoded by the text encoder, they are input into the backbone of the large language model for multi-modal reasoning. At the same time, after the output of the shared vision-language connector is average-pooled, a multi-task perception head is used to generate outputs aligned with the panoramic semantic map, traffic element perception, and visual entity localization labels. Finally, a natural language response is generated through a Q-Former classification decoder, including 3D scene parsing, semantic action inference, behavioral logic interpretation, motion trajectory planning, and driving control instructions, providing strong technical support for the deep coupling of multi-modal data and the interpretability of humanoid decisions.
[0012] Step 2: Development of a multi-level spatio-temporal semantic adaptive data extraction paradigm based on rule-driven and semantic action interaction; Using a novel data extraction paradigm, aiming to provide cross-modal semantic support for model training to support 3D scene semantic perception and long-axis temporal task reasoning. Through a rule-driven approach, based on the lateral and longitudinal speed targets reflecting the vehicle's motion intention and the driving direction intention providing directional information, it guides the generation of semantic actions. Subsequently, after normalizing information such as moving objects, interactors, infrastructure, control bodies, environmental factors, and road conditions in the scene, through multi-view image stitching and semantic map construction, a hierarchical scene semantic image is generated. These images and semantic actions interact with accurate object detection results to form multi-modal semantic understanding, and finally generate semantic action interpretations and adaptive guidance instructions. Based on this, an end-to-end interpretable autonomous driving comprehensive dataset AutoSceneVQA is developed, which not only provides sufficient semantic support but also ensures that the model can perform tasks in a clear and transparent manner during the training process; Step 3: Design of a multi-level training matrix architecture for cross-modal task orientation; By designing a multi-level training matrix architecture of cross-modal feature pre-training, 3D scene perception pre-training, end-to-end driving warm-up fine-tuning, and comprehensive fine-tuning, it aims to improve the model's perception and reasoning abilities through a task-oriented hierarchical optimization strategy. Among them, cross-modal feature pre-training jointly learns the deep alignment of visual and language features on multi-source datasets. 3D scene perception pre-training relies on large-scale 3D datasets to further enhance the model's 3D spatial semantic insight ability in complex driving environments by learning semantic structures and dynamic interactions in the 3D environment. End-to-end driving warm-up fine-tuning focuses on quickly adapting to driving tasks using pre-trained knowledge to optimize the full link performance from perception to decision-making. Finally, through the comprehensive fine-tuning link, the model's ability in long-time series task reasoning and decision-making optimization is further improved. At the same time, in order to further improve the training effect, a loss function is specifically designed, which ensures efficient collaborative learning and target guidance between different tasks by adaptively weighted fusion of the optimization objectives of each task module.
[0013] As Figure 1 shown, as a preferred embodiment of the present invention, in the said Step 1, the proposed AutoSceneNet system relies on the full-view scene video stream, global bird's-eye view, and multi-type semantic queries as inputs, aiming to efficiently solve complex tasks such as 3D scene semantic deconstruction, dynamic behavior deduction, human-like decision-making logic clarification, and precise control instruction generation in autonomous driving, and constructs a full-link technical system from perception to decision-making. Specifically, the AutoSceneNet system consists of a video-image dynamic semantic patch embedding joint encoder , text encoder , large language model backbone , shared vision-language connector , average pooling layer , multi-task perception head and the Q-Former classification decoder are composed to achieve 3D spatio-temporal perception and human-like decision-making reasoning through dynamic semantic patch embedding. The reasoning process is as follows: Given multi-type semantic queries , full-view dynamic video stream and global bird's-eye view , first encode the multi-type semantic queries into text tokens through the text encoder : ; Then, synchronously extract the dynamic visual semantic representation of the full-view dynamic video stream and the global spatial semantic representation of the global bird's-eye view through the video-image dynamic semantic patch embedding joint encoder : ; Among them, is the dynamic visual semantic representation, is the global spatial semantic representation; Subsequently, splice and and input them into the shared vision-language connector: ; Among them, is the unified vision token; Then, splice the unified vision token with the text token and input them into the large language model backbone for cross-modal semantic reasoning to generate multi-modal semantic reasoning results: ; Among them, is the multi-modal semantic reasoning result; At the same time, input the unified vision token into the average pooling layer and then, through the multi-task perception head , generate outputs of panoramic semantic map, traffic element perception, and visual entity localization label alignment: ; Among them, is the output aligned with the panoramic semantic map, is the output aligned with traffic element perception, is the output aligned with visual entity localization; Finally, through the Q-Former classification decoder Decode the multimodal semantic reasoning results into natural language semantic outputs and semantic classification results: ; Among them, is the natural language semantic output, is the semantic classification result; This system effectively improves the cross-scenario adaptability in complex scenarios by deeply integrating local dynamic information and global spatial information.
[0014] As a preferred embodiment of the present invention, in step 2, in order to support the three-dimensional scene semantic deconstruction and human-like decision-making reasoning ability of the AutoSceneNet system in complex outdoor scenarios, a multi-level spatio-temporal semantic adaptive data extraction paradigm based on rule-driven and semantic action interaction is proposed, and based on this, a comprehensive VQA driving instruction dataset AutoSceneVQA is developed, overcoming the limitations of traditional VQA datasets in terms of modal singularity and task complexity. This dataset is constructed based on the nuScenes open-source driving dataset, deeply integrating full-view dynamic video streams, global bird's-eye views, and multi-round semantic question-and-answer annotations, aiming to support the efficient reasoning of the model in three-dimensional scene understanding and end-to-end driving tasks through multi-modal semantic collaboration and spatio-temporal information coupling. It specifically includes the following steps: Step 2.1: Construction of a hierarchical scene understanding dataset; A hierarchical multi-modal semantic system is proposed, focusing on three core semantic elements in the scene: moving objects and interactors, infrastructure and controllers, and environmental factors and road conditions, to construct a multi-dimensional semantic representation space; among them, moving objects and interactors include the ego vehicle, pedestrians, and other traffic participants, and their dynamic positions and behaviors are detected in real time through sensors (such as lidar, cameras, etc.). Infrastructure and controllers cover fixed facilities such as road signs, traffic lights, and lane lines, which are extracted through visual recognition and sensor inputs. Environmental factors and road conditions include external environmental information such as weather conditions, lighting changes, and road conditions.
[0015] To address the lack of relative azimuth semantics of traffic participants in the nuScenes dataset, a spatial azimuth derivation mechanism is proposed, which calculates the angular difference between the ego vehicle's three-dimensional bounding box coordinates and traffic participants , and then determines the spatial relative azimuth. The function is used to calculate the angular difference between the ego vehicle and traffic participants, and its definition is: given the ego vehicle position and traffic participant position , first calculate the two-dimensional plane projection vectors and , and then calculate the angular difference through the vector angle formula: ; Based on the calculated angle difference , determine the relative orientation of traffic participants: When , the surrounding traffic participants are classified as "Front", when , the surrounding traffic participants are classified as "Front-right", when , the surrounding traffic participants are classified as "Front-left", when , the surrounding traffic participants are classified as "Rear-right", when , the surrounding traffic participants are classified as "Rear-right", and in other angle ranges, the surrounding traffic participants are classified as "Rear".
[0016] Among them, the forward direction of the ego vehicle is defined as 0°, and the clockwise angle is positive; in addition, the two-dimensional dynamic semantic attributes of traffic participants describe their movement trajectories through fixed sentences.
[0017] Based on Figure 3 the development process shown, the construction of the hierarchical scene understanding dataset is as follows: Step 2.1.1: Scene element extraction; Based on the scene information of the nuScenes dataset, extract information such as moving objects and interactors, infrastructure and control bodies, environmental factors and road conditions, and use a hierarchical modeling method to further divide each type of element into different hierarchical structures to form a multi-level semantic representation, providing accurate basic data for subsequent scene analysis and task reasoning.
[0018] Step 2.1.2: Normalization processing; After the extraction of scene elements is completed, enter the normalization processing stage. At this time, all the extracted information such as moving objects, infrastructure, and environmental factors will be uniformly converted into a standardized data format to ensure compatibility and consistency between different sensors and information sources. Specifically, it is uniformly processed through a data fusion algorithm to generate structured data including standardized position information, class labels, and timestamps, thereby ensuring data consistency and accuracy in subsequent steps.
[0019] Step 2.1.3: Multi-view image stitching and semantic map construction; Based on the full - perspective scene video stream, first, each frame is annotated with perspective azimuth and mapped into the global bird's - eye view. Then, combined with the semantic question set, a large - language model is called to perform semantic parsing on the spliced image data, automatically generating interactive semantic query pairs. Finally, the results of these query pairs are incorporated into the image splicing process to construct a scene semantic map with deep - level semantic information. This process not only completes the seamless splicing between different camera perspectives but also strengthens the multi - modal semantic understanding through interactive query pairs, providing solid semantic support for subsequent task reasoning and decision - making.
[0020] Step 2.2: Construction of an interpretable end - to - end driving dataset; Through a rule - driven method, semantic actions are generated based on the lateral and longitudinal speed targets of the ego - vehicle and the driving direction intention. The system uses these dynamic targets and intentions to guide the generation of semantic actions, thus accurately reflecting the motion state and driving intention of the ego - vehicle. Then, the generated semantic actions interact with the hierarchical semantic images. Combining with the accurate object detection results, the system finally generates adaptive guidance instructions. These instructions not only reflect the immediate decision - making process of the system but also provide clear and traceable guidance for subsequent tasks (such as path planning, decision - making, etc.). Finally, based on the generated semantic action interpretations and guidance instructions, an interpretable end - to - end driving dataset is constructed.
[0021] The specific construction process of the interpretable end - to - end driving dataset is as follows: Step 2.2.1: Extract ego - vehicle and scene information; Extract the original motion target information of the ego - vehicle from the nuScenes dataset, including lateral speed target, longitudinal speed target, and driving direction intention, and process it to extract the control signal sequence of the ego - vehicle, forming a historical control signal sequence as semantic context for subsequent vehicle motion control signal prediction. To quantify the extraction quality of the control signal, a signal integrity index is defined , and the calculation formula is: ; where is the number of frames of missing signals, is the total number of frames, The closer it is to 1, the more complete the signal sequence is.
[0022] Step 2.2.2: Rule - driven generation of semantic actions; Based on the dynamic targets and intentions extracted from the ego - vehicle and the scene, semantic actions are generated through a rule - driven decision framework. In this process, the system combines the lateral and longitudinal speed targets with the driving direction intention, and accurately calculates the next action plan of the ego - vehicle by designing the generation rules of semantic action representations and the interpretation templates of human - like decision - making logic.
[0023] Step 2.2.3: Generate adaptive guidance instructions and semantic action interpretations; The system interacts the generated semantic actions with the hierarchical semantic images, and combines the accurate object detection results to further generate adaptive guidance instructions and semantic action interpretations through interactive semantic queries. Through this interactive semantic query process, the system can extract the key semantic information in the scene and make dynamic adjustments according to the generated semantic actions and environmental conditions. At the same time, the GPT-4V model is used to decompose the control signal sequence to generate semantic action representations, so as to ensure that the generated guidance instructions are not only accurate, but also can effectively guide the system to perform task reasoning and decision-making. At the same time, to evaluate semantic consistency, a semantic consistency index is defined , and the calculation formula is: ; where and are the semantic embedding vectors of the question and the answer respectively, is the number of data pairs, The closer it is to 1, the more consistent the semantics of the question and the answer are; the generated semantic pairs and decision logics are corrected through algorithm rules to ensure that they are highly consistent with human driving thinking, thereby improving decision transparency and interpretability; to quantify the interpretability of decision logics, an interpretability index is defined , and the calculation formula is: ; where and are the semantic vectors of the original and corrected decision logics respectively, is the number of decision logics, The closer it is to 1, the more interpretable the corrected decision logic is.
[0024] Semantic action representation is generated through direction semantic inference , longitudinal speed semantic inference and lateral speed semantic inference , the maximum longitudinal acceleration threshold , the minimum longitudinal acceleration threshold , the maximum lateral acceleration threshold , the minimum lateral acceleration threshold , the speed threshold , the displacement threshold and the yaw angle threshold Lateral speed semantic inference : ; Among them, ClsLat The function is defined as: if , then return "fast"; if , then return "slow", otherwise return "medium"; is the lateral acceleration of the host vehicle.
[0025] Longitudinal velocity semantic inference : ; Among them, ClsLong The function is defined as: if , then return "fast-accel"; if , then return "accel", if , then return "decel"; if , then return "fast-decel", otherwise return "steady"; is the longitudinal acceleration of the host vehicle.
[0026] Direction semantic inference : ; Among them, ClsDir The function is defined as: if , then return "idle"; if and , then return "move-straight"; if and and , then return "turn-left"; if and and , then return "turn-right"; if and and , then return "shift-left"; if and and , then return "shift-right"; by default, return "move-straight"; is the lateral displacement, is the yaw angle.
[0027] Semantic action representation : ; Among them, The function combines the lateral velocity semantics , direction semantics and longitudinal velocity semantics into a triple in a fixed format. Through the above derivation, a total of 64 semantic action representations are generated, comprehensively covering the behavior patterns of the ego vehicle in diverse scenarios.
[0028] As a preferred embodiment of the present invention, in step 3, a cross-modal task-oriented multi-level training matrix architecture is designed to enhance the inference ability of the AutoSceneNet system in end-to-end autonomous driving tasks. As Figure 4 shown, it includes cross-modal feature pre-training, 3D scene perception pre-training, end-to-end driving warm-up fine-tuning, and comprehensive fine-tuning, aiming to enhance the model's 3D spatio-temporal perception and human-like decision-making inference ability. The training process is as follows: First, in the cross-modal feature pre-training stage, a multi-modal dataset is used to train the video-image dynamic semantic patch embedding joint encoder and the shared vision-language connector through the loss function . The weights of the text encoder , large language model backbone and Q-Former classification decoder are frozen to align the dynamic visual features, global visual features, and text features: ; Secondly, in the 3D scene perception pre-training stage, the hierarchical scene understanding dataset of AutoSceneVQA is used to optimize the AutoSceneNet system through the loss function to improve its 3D scene parsing ability: ; Then, in the end-to-end driving warm-up fine-tuning stage, the weights of the video-image dynamic semantic patch embedding joint encoder , the shared vision-language connector and the large language model backbone are frozen, and the Q-Former classification decoder is trained using cross-entropy loss to supervise the training process: ; where is the output of the Q-Former classification decoder.
[0029] Finally, in the end-to-end driving comprehensive fine-tuning stage, based on the interpretable end-to-end driving dataset of AutoSceneVQA , through the joint loss function Optimize the AutoSceneNet system to improve the complete link from 3D scene parsing to human-like decision-making generation. The joint loss function consists of two parts: the first part is the cross-entropy loss of the natural language output, and the second part is the mean squared error loss of the classification result, as follows: First, calculate the semantic response of multi-modal reasoning : ; Then, generate the natural language output and the classification result through the Q-Former classification decoder : ; Based on the above outputs, the joint loss function is defined as: ; where and are the true labels of the natural language output and the classification result in the dataset respectively, and and are weight coefficients that respectively adjust the loss contributions of the natural language output and the classification result.
[0030] As a preferred embodiment of the present invention, in order to verify the feasibility of the present invention, its efficiency and robustness in the end-to-end autonomous driving task are comprehensively demonstrated through experiments. The specific composition and effects are as follows: The present invention takes the AutoSceneNet system as the core, combines the comprehensive VQA driving instruction dataset AutoSceneVQA developed based on a multi-level spatio-temporal semantic adaptive data extraction paradigm driven by rules and semantic action interaction, and a cross-modal task-oriented multi-level training matrix architecture, and constructs a complete technical link from multi-modal data processing to human-like decision-making generation. In terms of experimental settings, the AutoSceneNet system takes the full-view scene video stream, global bird's-eye view, and multi-type semantic queries as inputs, and performs cross-modal semantic reasoning through a video-image dynamic semantic patch embedding joint encoder, text encoder, shared vision-language connector, and large language model backbone. The multi-level training matrix adopts a matrix architecture of cross-modal feature pre-training, 3D scene perception pre-training, end-to-end driving warm-up fine-tuning, and end-to-end driving comprehensive fine-tuning. The pre-training stage is completed on 8 A100 (80GB) GPUs (learning rate 1e-4, batch size 128, training for 20 epochs). The fine-tuning stage uses the AdamW optimizer (learning rate 5e-5, batch size 128, training for 5 epochs) to complete the optimization of 3D scene perception pre-training and end-to-end driving fine-tuning on 8 L20 (48GB) GPUs.
[0031] Based on the above experimental settings, first, through comparative experiments on various mainstream multi-modal large language model architectures, the performance of these architectures in end-to-end autonomous driving tasks was systematically evaluated. As shown in Table 1, compared with other mainstream architectures (such as DriveGPT4, RAG-Driver), AutoSceneNet demonstrated effective superiority in semantic action reasoning and behavioral rationality analysis tasks. The semantic understanding ability (measured by the CIDEr score C) was improved by approximately 10%, fully reflecting the technical breakthroughs of the present invention in cross-modal dynamic semantic patch embedding and multi-scale chain reasoning. At the same time, as shown in Table 2, in the motion planning task, the average L2 error of AutoSceneNet was only 0.36 meters, which was significantly lower than baseline methods such as UniAD and VAD, demonstrating its robustness and efficiency in planning human-like driving trajectories and verifying the effectiveness of the multi-level spatio-temporal semantic adaptive data extraction paradigm driven by rules and semantic action interaction.
[0032] Table 1 Comparison of Semantic Action Reasoning and Behavioral Rationality Analysis ; Table 2 Motion Planning Evaluation of Different Architectures on the AutoSceneVQA Dataset ; Subsequently, the present invention conducted an in-depth analysis of the dataset AutoSceneVQA developed based on the proposed data extraction paradigm. As shown in Table 3, compared with other existing autonomous driving datasets, AutoSceneVQA shows significant advantages in multi-view support, three-dimensional information coverage and multi-task capabilities. It not only supports multi-view (6 views) and three-dimensional information, but also comprehensively covers environmental perception, behavioral decision-making, motion planning and vehicle control tasks, making up for the limitations of existing datasets in spatiotemporal dynamics and task complexity, and providing strong support for three-dimensional spatiotemporal perception and human-like decision reasoning.
[0033] Table 3 Comparison of AutoSceneVQA dataset with other autonomous driving datasets ; Finally, the present invention intuitively demonstrates the effect of technical application through visual analysis. Figure 5 As shown in Tables 4 and 5 below, in complex urban intersection scenarios, AutoSceneNet can complete hierarchical scene understanding based on multimodal input, accurately infer driving strategies, generate human-like behavior explanations, and ultimately output safe and reliable control signals, which fully demonstrates the efficient applicability and decision transparency of the present invention in 3D spatiotemporal perception, long-axis task planning, and human-like decision reasoning, providing effective technical support for the practical application of end-to-end autonomous driving and demonstrating the potential for practical industrialization.
[0034] Table 4 Hierarchical scene understanding
[0035] Table 5 Explainable end-to-end driving
[0036] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A large-scale autonomous driving model framework based on 3D spatiotemporal perception and human-like decision reasoning, characterized by: The steps for building the autonomous driving large model framework are as follows: Step 1: Construction of a visual language framework for autonomous driving with dynamic semantic patch embedding and multi-scale chain reasoning; By designing the AutoSceneNet system, a multimodal large language model architecture with human-like multi-scale reasoning capabilities is constructed. The video-image dynamic semantic patch embedding joint encoder is used to extract features from the full-view scene video stream and the global bird's-eye view, generating dynamic video sequences and global spatial features. Subsequently, the visual features are mapped to a unified semantic representation space through a cross-modal semantic alignment mechanism, and projected to a high-dimensional language embedding domain using a shared visual-language connector to generate unified visual tags. After being spliced with the text tags encoded by the text encoder, they are input into the backbone of a large language model for multimodal reasoning. At the same time, after the output of the shared visual-language connector is average-pooled, a multi-task perception head is used to generate outputs aligned with panoramic semantic maps, traffic element perception, and visual entity positioning labels. Finally, a natural language response is generated through the Q-Former classification decoder, including three-dimensional scene analysis, semantic action inference, behavioral logic interpretation, motion trajectory planning, and driving control instructions. Step 2: Multi-level spatiotemporal semantic adaptive data extraction paradigm based on rule-driven and semantic-action interaction; In a rule-driven way, the generation of semantic actions is guided based on the lateral and longitudinal speed targets that reflect the vehicle's motion intention and the driving direction intention that provides directional information. Subsequently, the moving objects and interactors, infrastructure and control objects, as well as environmental factors and road condition information in the scene are normalized, and then multi-view image stitching and semantic graph construction are performed to generate a hierarchical scene semantic image. These images and semantic actions are interacted with the target detection results to form a multi-modal semantic understanding, and finally semantic action interpretation and adaptive guidance instructions are generated. Based on this, an interpretable end-to-end autonomous driving comprehensive dataset AutoSceneVQA is developed to provide a cross-modal semantic foundation for model training that supports three-dimensional scene semantic perception and long-axis temporal task reasoning. Step 3: Design of cross-modal task-oriented multi-level training matrix architecture; A cross-modal task-oriented multi-level training matrix architecture is designed, covering cross-modal feature pre-training, 3D scene perception pre-training, end-to-end driving warm-up fine-tuning and comprehensive fine-tuning; in this architecture, cross-modal feature pre-training jointly learns the deep alignment of visual and language features on multi-source datasets, 3D scene perception pre-training relies on large-scale 3D datasets to learn the semantic structure and dynamic interaction in the 3D environment, and end-to-end driving warm-up fine-tuning uses the pre-trained knowledge to quickly adapt to driving tasks and thus optimize the full-link performance from perception to decision-making; finally, through the comprehensive fine-tuning link, the model's capabilities in long-term task reasoning and decision optimization are improved. At the same time, a loss function is designed that fuses the optimization objectives of each task module through adaptive weighting.
2. The large-scale model framework for autonomous driving based on 3D spatiotemporal perception and human-like decision-making reasoning according to claim 1 is characterized in that: In step 1, the AutoSceneNet system consists of a video-image dynamic semantic patch embedding joint encoder , Text Encoder , Large Language Model Backbone , Shared Vision-Language Connector , average pooling layer , Multi-task perception head and Q-Former classification decoder Composition,aims to achieve 3D spatiotemporal perception and human-like decision reasoning through,dynamic semantic patch embedding; The reasoning process is as follows: Given a multi-type semantic query , full-view scene video stream and a bird's eye view First, the multi-type semantic query is encoded into text tags through a text encoder : ; Next, we use the video-image dynamic semantic patch embedding joint encoder Synchronously extract full-view scene video stream Dynamic visual semantic representation and global bird's-eye view The global spatial semantic representation of: ; in, is a dynamic visual semantic representation, It is a global spatial semantic representation; Afterwards, and To stitch, enter the shared vision-language connector: ; in, To unify visual markings; Then, unify the visual mark With text mark Concatenation, input large language model backbone Perform cross-modal semantic reasoning and generate multimodal semantic reasoning results: ; in, It is the result of multimodal semantic reasoning; At the same time, the visual mark Input average pooling layer Afterwards, through the multi-task perception head , generating outputs of panoramic semantic maps, traffic element perception, and visual entity localization label alignment: ; in, is the output of the panoramic semantic map alignment, The output of traffic element perception alignment, Output for visual entity localization alignment; Finally, the Q-Former classification decoder Decode the multimodal semantic reasoning results into natural language semantic output and semantic classification results: ; in, is the natural language semantic output, The semantic classification result.
3. The large-scale model framework for autonomous driving based on 3D spatiotemporal perception and human-like decision-making reasoning according to claim 2 is characterized in that: In step 2, the construction of the comprehensive VQA driving instruction dataset AutoSceneVQA based on the multi-level spatiotemporal semantic adaptive data extraction paradigm of rule-driven and semantic-action interaction includes the following steps: Step 2.1: Construction of hierarchical scene understanding dataset; Focusing on the three core semantic elements of mobile objects and interactors, infrastructure and control objects, and environmental factors and road conditions in the scene, a multi-dimensional semantic representation space is constructed; among them, mobile objects and interactors include vehicles, pedestrians and other traffic participants, and their dynamic positions and behaviors are detected in real time through sensors; infrastructure and control objects include road signs, traffic lights, and lane lines, which are extracted through visual recognition and sensor input; environmental factors and road conditions include weather conditions, lighting changes and road conditions, which are evaluated and captured through the environmental perception module; Calculate the angle difference between the 3D bounding box coordinates of the vehicle and the traffic participants , and then determine the relative position in space; function It is used to calculate the angle difference between the ego vehicle and the traffic participants, which is defined as: given the ego vehicle position and the location of traffic participants , first calculate the two-dimensional plane projection vectors of the two and , and then use the vector angle formula Calculate the angle difference: ; Based on the calculated angle difference , determine the relative position of traffic participants: when When , the surrounding traffic participants are classified as "Front", when When When , the surrounding traffic participants are classified as "Front-left", when When , the surrounding traffic participants are classified as "Rear-right", when When the angle is 0.0000, the surrounding traffic participants are classified as "Rear-right", and when the angle is in other ranges, the surrounding traffic participants are classified as "Rear"; Among them, the forward direction of the vehicle is defined as 0°, and the clockwise angle is positive; in addition, the two-dimensional dynamic semantic attributes of traffic participants describe their motion trajectories through fixed sentences; The specific process of constructing a hierarchical scene understanding dataset is as follows: Step 2.1.1: scene element extraction; Based on the scene information of the nuScenes dataset, mobile objects and interactors, infrastructure and control objects, environmental factors and road conditions are extracted. Each type of element is further subdivided into different levels of structure using a hierarchical modeling method to form a multi-level semantic representation, providing accurate basic data for subsequent scene analysis and task reasoning; Step 2.1.2: Normalization processing; After the scene elements are extracted, the normalization process begins. At this point, all extracted information about moving objects, infrastructure, and environmental factors will be uniformly converted into a standardized data format. Specifically, the data will be uniformly processed through a data fusion algorithm to generate structured data including standardized location information, category labels, and timestamps, ensuring data consistency and accuracy in subsequent steps. Step 2.1.3: Multi-view image stitching and semantic graph construction; Based on the full-view scene video stream, the viewpoint orientation of each frame is first annotated and mapped to the global bird's-eye view. Then, combined with the semantic question set, the large language model is called to perform semantic analysis on the spliced image data and automatically generate interactive semantic query pairs. Finally, the results of these query pairs are integrated into the image splicing process to construct a scene semantic graph with deep semantic information, which provides semantic support for subsequent task reasoning and decision-making. Step 2.2: Construction of an explainable end-to-end driving dataset; Through a rule-driven approach, semantic actions are generated according to the lateral and longitudinal speed targets and driving direction intentions of the ego vehicle. The system guides the generation of semantic actions through these dynamic targets and intentions, thereby reflecting the motion state and driving intentions of the ego vehicle. Then, the generated semantic actions interact with the hierarchical semantic images, and combined with the precise target detection results, the system finally generates adaptive guidance instructions. Finally, based on the generated semantic action interpretation and guidance instructions, an interpretable end-to-end driving dataset is constructed, providing a semantic basis for three-dimensional scene perception and long-term task reasoning for the training of autonomous driving systems.
4. The large-scale model framework for autonomous driving based on 3D spatiotemporal perception and human-like decision reasoning according to claim 3 is characterized in that: In step 2.2, the specific construction process of the explainable end-to-end driving dataset is as follows: Step 2.2.1: Extract vehicle and scene information; Extract the original motion target information of the vehicle from the nuScenes dataset, including the lateral speed target, longitudinal speed target, and driving direction intention, and process it to extract the control signal sequence of the vehicle to form a historical control signal sequence as a semantic context for subsequent vehicle motion control signal prediction; To quantify the extraction quality of the control signal, define signal integrity indicators , the calculation formula is: ; in, is the number of frames of missing signal, is the total number of frames, The closer it is to 1, the more complete the signal sequence is; Step 2.2.2: Generate semantic actions driven by rules; Based on the dynamic goals and intentions extracted from the ego vehicle and the scene, semantic actions are generated through a rule-driven decision framework. In this process, the system combines the lateral and longitudinal speed goals with the driving direction intention, and calculates the next action plan of the ego vehicle by designing the generation rules of semantic action representation and the interpretation template of human-like decision logic. Step 2.2.3: Generate adaptive guidance instructions and semantic action interpretation; The system interacts the generated semantic actions with the hierarchical semantic images, and combines the target detection results to further generate adaptive guidance instructions and semantic action interpretations through interactive semantic queries. At the same time, the GPT-4V model is used to decompose the control signal sequence to generate semantic action representations. At the same time, in order to evaluate semantic consistency, a semantic consistency index is defined. , the calculation formula is: ; in, and are the semantic embedding vectors of questions and answers respectively, is the number of data pairs, The closer it is to 1, the more consistent the semantics of the question and answer are; the semantic pairs and decision logic generated by the algorithm rules are corrected to ensure that they are highly consistent with human driving thinking; to quantify the interpretability of the decision logic, define the interpretability index , the calculation formula is: ; in, and are the original and revised decision logic semantic vectors, is the number of decision logics, The closer it is to 1, the more explainable the revised decision logic is.
5. The large-scale model framework for autonomous driving based on 3D spatiotemporal perception and human-like decision reasoning according to claim 4 is characterized in that: In step 2.2.2, the semantic action representation The generation of , Longitudinal velocity semantic inference and lateral velocity semantic inference Three-dimensional implementation; first, define the threshold vector, including the maximum threshold of longitudinal acceleration , Minimum threshold of longitudinal acceleration , Maximum threshold of lateral acceleration , minimum threshold of lateral acceleration , speed threshold , displacement threshold and yaw angle threshold , based on which the semantic type is derived: Lateral velocity semantic inference : ; in, ClsLat The function is defined as: , then return "fast"; if , then returns "slow", otherwise returns "medium"; is the lateral acceleration of the vehicle; ; in, ClsLong The function is defined as: , then returns "fast-accel"; if , then returns "accel", if , then return "decel"; if , then returns "fast-decel", otherwise returns "steady"; is the longitudinal acceleration of the vehicle; Direction semantic inference : ; in, ClsDir The function is defined as: , then return "idle"; if and , then return "move-straight"; if and and , then return "turn-left"; if and and , then return "turn-right"; if and and , then returns "shift-left"; if and and , then returns "shift-right"; the default return is "move-straight"; is the lateral displacement, is the yaw angle; Semantic Action Representation : ; in, Function to convert lateral velocity semantics , Direction semantics and longitudinal velocity semantics The three tuples are spliced into a fixed format. Through the above derivation, a total of 64 semantic action representations are generated, which fully cover the behavior patterns of the vehicle in various scenarios.
6. The large-scale model framework for autonomous driving based on 3D spatiotemporal perception and human-like decision-making reasoning according to claim 5 is characterized in that: In step 3, a cross-modal task-oriented multi-level training matrix architecture is designed to improve the reasoning ability of the AutoSceneNet system in end-to-end autonomous driving tasks, including cross-modal feature pre-training, three-dimensional scene perception pre-training, end-to-end driving warm-up fine-tuning, and end-to-end driving comprehensive fine-tuning; The training process is as follows: First, in the cross-modal feature pre-training stage, a multimodal dataset is used , through the loss function Training a joint video-image dynamic semantic patch embedding encoder and shared visual-linguistic connectors , frozen text encoder , Large Language Model Backbone and Q-Former classification decoder The weights of to align dynamic visual features, global visual features and text features: ; Secondly, in the 3D scene perception pre-training stage, we use the hierarchical scene understanding dataset of AutoSceneVQA , through the loss function Optimize the AutoSceneNet system to improve its 3D scene analysis capabilities: ; Then, in the end-to-end driving warm-up fine-tuning stage, the video-image dynamic semantic patch embedding joint encoder is maintained and shared visual-linguistic connectors and large language model backbones Weight freezing, training Q-Former classification decoder , using cross entropy loss Supervised training process: ; in, is the output of the Q-Former classification decoder; Finally, in the end-to-end driving comprehensive fine-tuning stage, the interpretable end-to-end driving dataset based on AutoSceneVQA , through the joint loss function Optimize the AutoSceneNet system and improve the complete link from 3D scene analysis to human-like decision generation; joint loss function It consists of two parts: the first part is the cross entropy loss of natural language output, and the second part is the mean square error loss of the classification result, as follows: First, the semantic response of multimodal reasoning is calculated : ; Then, the Q-Former classification decoder Generate natural language output And the classification results : ; Based on the above output, the joint loss function Defined as: ; in, and The data sets The true labels of the natural language output and classification results in and are weight coefficients, which adjust the loss contribution of natural language output and classification results respectively.
Citation Information
Patent Citations
Logical reasoning problem-oriented anthropomorphic thinking-driven embedding enhancement method
CN119025646A
Equipment operation manual vector knowledge base construction method
CN119128219A
Automatic driving interpretation text determination method based on large visual language model
CN119142366A
Multi-mode instruction following data set construction method special for automatic driving
CN119691692A
Visual question and answer data set construction method for three-dimensional driving scene understanding
CN119721263A
Cited By
Automatic driving end-to-end model self-correction method and device and medium
CN120564157A
Self-correction method and device for end-to-end model of automatic driving and medium
CN120564157B
Human motion generation method and system based on multi-token large language model
CN120597896A
Driver abnormity monitoring system based on large model collaborative decision
CN120635869A
A driver abnormality monitoring system based on large model collaborative decision making
CN120635869B