Method and device for determining end-to-end perception decision regulation and control architecture
Through an end-to-end perception, decision-making and control architecture, and by utilizing large multimodal models and reinforcement learning, the information transmission problem of the layered architecture and the perception deficiencies of the single-modal model are solved, enabling efficient and explainable decision-making capabilities of the autonomous driving system in complex scenarios.
Patent Information
- Application Number
- CN202511130138.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-13
AI Technical Summary
In existing autonomous driving technology, the layered architecture leads to information loss and high system complexity, making it difficult to meet real-time requirements. In addition, the existing end-to-end model relies on single-modal data, has insufficient perception capabilities and is highly unexplainable, making it difficult to perform well in complex scenarios.
By adopting an end-to-end perception, decision-making and control architecture, and through multimodal large model training and combining multimodal data sets, we can gradually achieve the transition from basic traffic signal recognition to dynamic traffic perception and then to driving decision-making. We use cross-modal feature alignment and rule understanding modules to fuse semantic features, and optimize the decision-making module through reinforcement learning to form an end-to-end perception decision-making process.
It significantly improves the real-time and robustness of the autonomous driving system, enhances perception accuracy and decision-making interpretability, and is able to operate stably in complex scenarios, meeting the needs of highly complex driving tasks.
Smart Images

Figure CN120726596A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to a method and device for determining an end-to-end perception, decision-making, and control architecture. Background Art
[0002] Existing autonomous driving technologies primarily utilize a layered architecture, which divides the system into multiple modules, including perception, prediction, planning, and control. The perception module uses sensors such as cameras, lidar, and radar to identify static and dynamic objects in the surrounding environment and generate a map of the environment. The prediction module analyzes the behavior of dynamic objects and understands the movement trends of pedestrians and vehicles. The planning module develops a reasonable path based on perception and prediction information, and the control module executes the planned actions. This layered architecture, with its clear structure, is the mainstream approach in many current autonomous driving solutions.
[0003] However, the independence between modules may lead to information loss or limited expression during the transmission process, making it difficult to optimize the overall system. At the same time, since multiple modules need to work together, the system complexity is high, the design cycle is long, and it will bring certain computing delays, which is not ideal for autonomous driving tasks with high real-time requirements. Summary of the Invention
[0004] The present invention provides a method and device for determining an end-to-end perception, decision-making and control architecture to address the defect that a layered architecture is not ideal for autonomous driving tasks with high real-time requirements.
[0005] The present invention provides a method for determining an end-to-end perception decision-making and control architecture, comprising: A first multimodal dataset constructed based on basic traffic signal images, road map data, and basic driving rules for autonomous driving scenarios is used to train a preset multimodal large model; A second multimodal dataset constructed based on dynamic traffic element image sequence data, complex environmental element image sequence data, and complex driving rules of an autonomous driving scenario is used to train the multimodal large model trained on the first multimodal dataset; A third multimodal dataset constructed based on autonomous driving scene images, autonomous driving targets, and sensor time series data is used to perform end-to-end training on the multimodal large model trained on the second multimodal dataset, thereby obtaining an end-to-end perception, decision-making, and control architecture suitable for autonomous driving scenarios. The autonomous driving scene image includes one or more of basic traffic signals, road map data, dynamic traffic elements and complex environment elements.
[0006] As an embodiment, the multimodal large model includes a cross-modal feature alignment module and a rule understanding module; The cross-modal feature alignment module is used to extract basic semantic features from the basic traffic signal image, the road map data, and the basic driving rules, and map the basic semantic features into the same semantic embedding space to complete semantic feature alignment and semantic feature fusion to obtain basic semantic fusion features; The rule understanding module is used to convert the basic driving rules into a structured semantic representation.
[0007] As an embodiment, the road map data includes environmental map information.
[0008] As an embodiment, the multimodal large model further includes a multimodal perception module and a modal feature cross-decoding module; The multimodal perception module is used to extract the temporal characteristics of the dynamic traffic element image sequence data, the complex environment element image sequence data, and the complex driving rules; The modal feature cross-decoding module is used to interactively decode and globally aggregate the temporal features to generate a unified high-order representation.
[0009] As an embodiment, the multimodal large model further includes a decision module; The decision module is used to encode the autonomous driving scene image, the autonomous driving target and the sensor time series data into a joint representation, input the structured semantic representation, the unified high-order representation and the joint representation into a reinforcement learning model, and generate a decision, which is used to represent a set of autonomous driving control instructions.
[0010] In one embodiment, the third multimodal dataset includes a simulated multimodal data subset and a real multimodal data subset. Correspondingly, the multimodal large model trained on the second multimodal dataset is end-to-end trained to obtain an end-to-end perception, decision-making, and control architecture suitable for autonomous driving scenarios, including: Based on the simulated multimodal data subset, simulation training is performed on the multimodal large model after complex recognition training, and during the simulation training process, parameters of the multimodal large model are optimized based on an end-to-end back-propagation mechanism; Based on the real multimodal data subset, the parameters of the multimodal large model after simulation training are adjusted to obtain the end-to-end perception decision-making and control architecture.
[0011] The present invention also provides an end-to-end perception decision-making and control architecture determination device, comprising: A first training module is used to train a preset multimodal large model based on a first multimodal dataset constructed based on basic traffic signal images, road map data, and basic driving rules of an autonomous driving scenario; A second training module is configured to train the multimodal large model trained on the first multimodal dataset based on a second multimodal dataset constructed based on image sequence data of dynamic traffic elements, image sequence data of complex environmental elements, and complex driving rules of an autonomous driving scenario; A third training module is configured to perform end-to-end training on the multimodal large model trained on the second multimodal dataset based on a third multimodal dataset constructed from autonomous driving scene images, autonomous driving targets, and sensor time series data, thereby obtaining an end-to-end perception, decision-making, and control architecture suitable for autonomous driving scenarios; The autonomous driving scene image includes one or more of basic traffic signals, road map data, dynamic traffic elements and complex environment elements.
[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the computer program, it implements any of the above-described methods for determining an end-to-end perception decision-making and control architecture.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the end-to-end perception decision-making and control architecture determination method as described in any one of the above.
[0014] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described methods for determining an end-to-end perception decision-making and control architecture.
[0015] The present invention provides a method and device for determining an end-to-end perception, decision-making, and regulatory control architecture. A multimodal large model is trained through a phased training strategy, gradually moving from basic element recognition to complex element recognition and then to decision training, thereby enhancing the real-time performance of the model. The end-to-end process avoids the information loss problem in traditional layered architectures, significantly improves response efficiency and anti-interference capabilities, and meets the real-time requirements of autonomous driving tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 This is one of the flow charts of the end-to-end perception decision-making and control architecture determination method provided by the present invention.
[0018] Figure 2This is the second flow chart of the end-to-end perception decision-making and control architecture determination method provided by the present invention.
[0019] Figure 3 It is a structural diagram of the multimodal large model provided by the present invention.
[0020] Figure 4 It is a structural diagram of the end-to-end perception decision-making and control architecture determination device provided by the present invention.
[0021] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0022] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0023] Existing autonomous driving technologies face numerous shortcomings in practical applications. While layered architectures offer the advantage of a clearly structured, multi-module design, the independence between modules can lead to information loss or limited representation during transmission, making overall system optimization difficult. Furthermore, the need for multiple modules to work together increases system complexity, lengthens the design cycle, and introduces computational latency, making them unsuitable for real-time autonomous driving tasks. Furthermore, while end-to-end deep learning approaches can simplify system architecture and directly map inputs to control commands, current mainstream end-to-end models rely on single-modal data (such as camera images) and fail to leverage the rich information from multimodal data. This results in insufficient environmental perception and high levels of uninterpretability, making driving decision logic difficult to verify and generalizing poorly. System performance can also significantly degrade when encountering scenarios not seen in the training data. Furthermore, while multimodal fusion technology has been used to enhance perception capabilities, current approaches primarily focus on the perception layer and fail to fully integrate into the end-to-end driving decision-making process, resulting in challenges in robustness and real-time performance in practical applications. Overall, these shortcomings make it difficult for existing technologies to achieve ideal performance levels in complex and diverse real-world driving scenarios.
[0024] To this end, the present invention provides a method and device for determining an end-to-end perception decision-making and control architecture, which is described in detail below with reference to the accompanying drawings.
[0025] Figure 1 This is one of the flow charts of the method for determining the end-to-end perception decision-making and control architecture provided by the present invention, such as Figure 1 As shown, the present invention provides a method for determining an end-to-end perception decision-making and control architecture, including steps S100-S300.
[0026] Step S100 , training a preset multimodal large model based on a first multimodal dataset constructed based on basic traffic signal images, road map data, and basic driving rules of the autonomous driving scenario.
[0027] Optionally, the road map data includes environmental map information, such as location data and lane distribution related to the location data, etc. The basic traffic signal image is used to represent an image containing traffic signals (such as traffic lights, turn signs, speed limit signs, etc.), and the basic driving rules are used to represent the driving rules set for the basic traffic signal image and the road map data, such as stop at red lights and go at green lights, and drive with caution in locations with high accident rates.
[0028] The first multimodal dataset includes basic traffic signal images in image mode, road map data in spatial mode, and basic driving rules in text mode. It uses image data as the main modality input, and combines high-precision road map data and manually set basic driving rules to enhance the multimodal large model's understanding of traffic signals, rules, and scenarios through a combination of vision, space, and language.
[0029] The preset multimodal large model can be trained using weakly supervised or semi-supervised learning methods, so that the multimodal large model after basic recognition training can accurately classify and semantically label the basic traffic elements in the scene, and generate a preliminary high-level semantic representation of the scene, such as a green light for going straight or a dangerous sharp turn.
[0030] Step S200: Training the multimodal large model trained on the first multimodal dataset based on a second multimodal dataset constructed based on dynamic traffic element image sequence data, complex environment element image sequence data, and complex driving rules of the autonomous driving scene.
[0031] Dynamic traffic elements include pedestrians, vehicles, bicycles, and other dynamic participants on the road. Complex environmental elements include complex road structures (such as intersections and roundabouts) and driving environments (such as inclement weather). Complex driving rules are used to represent driving rules set for image sequence data of dynamic traffic elements and complex environmental elements (for example, slow down to avoid a pedestrian on the left entering the lane). Complex driving rules contain scene information described in language and can serve as supplementary features.
[0032] The multimodal large model trained on the first multimodal data set can be trained using a weakly supervised learning method to achieve complex recognition, real-time detection and identification of the position, category and behavior trend of dynamic traffic participants (such as pedestrians approaching from the right, slowing down and braking), and global semantic graph construction to form a complete dynamic scene understanding capability.
[0033] Step S300: Based on the third multimodal dataset constructed based on the autonomous driving scene images, autonomous driving targets, and sensor time series data, the multimodal large model trained on the second multimodal dataset is end-to-end trained to obtain an end-to-end perception, decision-making, and control architecture suitable for autonomous driving scenarios.
[0034] The autonomous driving scene image includes one or more of basic traffic signals, road map data, dynamic traffic elements and complex environment elements.
[0035] The autonomous driving goal is used to represent the natural language instructions set for the autonomous driving scene image, such as turning left at the next intersection. The sensor time series data includes the time series data collected by the radar sensor, which is used to help the vehicle identify roads, obstacles, pedestrians, etc., and enhance its adaptability to extreme scenarios (such as low visibility and irregular moving objects).
[0036] like Figure 2 As shown, the training phase of the multimodal large model provided by the present invention includes a basic traffic signal recognition training step (i.e., step S100), a road traffic element recognition training step (i.e., step S200), and an end-to-end decision-making training step (i.e., step S300). Through a step-by-step training strategy, an efficient, reliable, and interpretable end-to-end autonomous driving model based on the multimodal large model (i.e., an end-to-end perception, decision-making, and control architecture) is constructed to realize the complete process of autonomous behavior by perceiving the environment, making decisions, and executing control.
[0037] During the application phase, real-time image data collected by high-definition cameras, location information collected by the positioning system, and radar data collected by sensors are input into the end-to-end perception, decision-making, and control architecture. Driving control instructions (such as acceleration, steering, braking, moving forward, changing lanes, and deceleration) output by the end-to-end perception, decision-making, and control architecture are obtained. Real-time driving control decisions can be reliably generated in various complex scenarios, ensuring that the logical path leading to driving behavior has complete link explainability.
[0038] It is understood that this invention effectively improves perception accuracy and scene understanding capabilities by introducing a large multimodal model and deeply integrating multiple sources of information, such as scene information described in language. Furthermore, through a phased training strategy, it gradually optimizes the process from basic traffic signal recognition to dynamic traffic element perception and finally to driving decision-making, significantly enhancing generalization, robustness, and real-time performance. This end-to-end perception, decision-making, and control architecture avoids the information loss inherent in traditional layered architectures through an end-to-end process, significantly improving response efficiency and anti-interference capabilities. The deep integration of multimodal information makes the decision-making process more robust and reasonable in complex scenarios.
[0039] Based on the above embodiment, as an optional embodiment, the multimodal large model includes a cross-modal feature alignment module and a rule understanding module.
[0040] The cross-modal feature alignment module is used to extract the basic semantic features of the basic traffic signal image, the road map data, and the basic driving rules, and map each of the basic semantic features to the same semantic embedding space to complete semantic feature alignment and semantic feature fusion to obtain basic semantic fusion features.
[0041] The rule understanding module is used to convert the basic driving rules into a structured semantic representation.
[0042] The cross-modal feature alignment module mainly includes a multi-channel encoder and an aggregation network. First, each modal data in the multimodal dataset (such as visual information, environmental structure, text rules, etc.) is extracted with basic semantic features through an independent dedicated neural network. The features of all modalities are mapped to the same semantic embedding space. Semantic-level alignment and fusion are achieved through cross-modal attention mechanism, fusion Transformer structure or contrastive learning algorithms, thereby establishing deep semantic connections between different modal information.
[0043] The rule understanding module adopts a Transformer-based deep understanding model or a large language model. After feature fusion, it performs contextual analysis and reasoning on the text rules, expresses the text rules as structured semantic representations, and combines them with multimodal environmental perception features to input them into the decision module, realizing intelligent understanding and application of text rules in complex scenarios.
[0044] It can be understood that the present invention performs basic recognition training on the multimodal large model, that is, trains the cross-modal feature alignment module and the rule understanding module. Different from the independent processing or simple splicing of single modal data in the existing model, the traditional method is difficult to capture the complex semantic dependencies between modalities. The cross-modal feature alignment module and the rule understanding module are used to realize the semantic interaction and scene association understanding of different modal data.
[0045] Based on the above embodiment, as an optional embodiment, the multimodal large model further includes a multimodal perception module and a modal feature cross-decoding module.
[0046] The multimodal perception module is used to extract the respective temporal features of the dynamic traffic element image sequence data, the complex environment element image sequence data, and the complex driving rules.
[0047] The modal feature cross-decoding module is used to interactively decode and globally aggregate the temporal features to generate a unified high-order representation.
[0048] Optionally, the multimodal perception module can also be connected to the cross-modal feature alignment module to transmit the extracted timing features to the cross-modal feature alignment module for feature alignment. The cross-modal feature alignment module transmits the aligned timing features to the rule understanding module, and the rule understanding module converts the timing features into timing features of structured semantic representation. Correspondingly, the modal feature cross-decoding module interactively decodes and globally aggregates the timing features of the structured semantic representation to generate a unified high-order representation.
[0049] Optionally, the multimodal large model also includes a dynamic prediction network, which adopts a time series modeling structure, such as a neural module based on Transformer, temporal convolutional network (TCN) or long short-term memory network (LSTM), to model the temporal dynamics of the time series data in the second multimodal dataset. Through the input of perception data streams at continuous moments, it effectively captures the evolution laws of environmental elements and the behavior of participants (pedestrians, vehicles, etc.), and realizes early perception and state prediction of motion, changes and emergencies in the scene.
[0050] Optionally, the modal feature cross-decoding module includes a heterogeneous information decoder or a gated fusion unit, which uses a cross-attention mechanism to deeply decouple and reintegrate the time series data from different modalities in time and semantic space to make up for the limitations of a single modality in capturing global dynamics and sudden situations.
[0051] It can be understood that the present invention can make up for the limitations of a single modality in capturing global dynamics and sudden situations through a multimodal perception module and a modal feature cross-decoding module, thereby achieving robust perception of dynamic road elements.
[0052] Based on the above embodiment, as an optional embodiment, the multimodal large model further includes a decision module.
[0053] The decision module is used to encode the autonomous driving scene image, the autonomous driving target and the sensor time series data into a joint representation, input the structured semantic representation, the unified high-order representation and the joint representation into a reinforcement learning model, and generate a decision, which is used to represent a set of autonomous driving control instructions.
[0054] It is understandable that the decision-making module is driven by end-to-end reinforcement learning, which can achieve collaborative optimization of different goals (such as efficiency, stability, energy consumption, etc.).
[0055] Based on the above embodiment, as an optional embodiment, the third multimodal data set includes a simulated multimodal data subset and a real multimodal data subset. Correspondingly, the multimodal large model trained on the second multimodal data set is end-to-end trained to obtain an end-to-end perception, decision-making and control architecture suitable for autonomous driving scenarios, including: Based on the simulated multimodal data subset, simulation training is performed on the multimodal large model after complex recognition training. During the simulation training process, the parameters of the multimodal large model are optimized based on an end-to-end back propagation mechanism.
[0056] Based on the real multimodal data subset, the parameters of the multimodal large model after simulation training are adjusted to obtain the end-to-end perception decision-making and control architecture.
[0057] It can be understood that the present invention performs end-to-end reinforcement learning in a simulation environment, optimizes path planning and driving control by continuously adjusting the driving strategy, and then further fine-tunes it in a real data environment to cope with noise and diverse terrain. This is conducive to reliably generating real-time driving control decisions in various complex scenarios and ensuring that the logical path leading to driving behavior has complete link interpretability.
[0058] Figure 3 This is a structural diagram of a preferred embodiment of the multimodal large model provided by the present invention, and the arrows in the figure represent the signal flow.
[0059] In the basic traffic signal recognition training steps, the training process is organically integrated with environmental perception, rule processing, and multimodal feature alignment. End-to-end backpropagation optimizes the parameters of the entire process, enabling the intelligent agent to achieve comprehensive understanding of multimodal information, contextual reasoning, and action planning, greatly improving the generalization ability of large multimodal models and their adaptability to complex environments.
[0060] The loss function expression of the basic traffic signal recognition training step is as follows: L=λ1L cross +λ2L align +λ3L rule ; Where L represents the loss function of the basic traffic signal recognition training step; λ1, λ2, and λ3 represent the weight coefficients of the loss of different subtasks, which are used to balance the contribution of each part to the overall training goal; L cross Represents the main task loss of basic recognition (such as classification, detection, etc.) in multimodal scenarios; L align Indicates the degree of alignment of features of different modalities (such as vision, structure, text, etc.) in the common semantic space constrained by contrastive learning or embedding similarity loss, such as contrastive loss (such as InfoNCE), mean square error (MSE) or triplet loss (Triplet Loss). rule Represents the contextual parsing loss of the rule understanding module, which is used to constrain the model's understanding and reasoning accuracy of text rule expressions (such as the cross-entropy loss of rule prediction).
[0061] When training a large multimodal model, all parameters θ (including the weights and biases of the multimodal encoder, feature aggregation network, cross-modal fusion module, and rule parsing network) are updated synchronously through an end-to-end backpropagation algorithm (such as Adam or SGD optimizer), with the optimization goal of minimizing the total loss L mentioned above. The parameter update expression is as follows: θ←θ-η L / θ Among them, θ represents the set of trainable parameters of the entire model; η represents the learning rate; L / θ represents the gradient of the total loss L with respect to the model parameters θ.
[0062] Through the above-mentioned end-to-end loss combination and parameter optimization process, it is possible to efficiently achieve the unified fusion and generalized modeling of multimodal features, contextual rules and environmental perception information, thereby significantly improving the comprehensive intelligent capabilities of large multimodal models in complex scenarios.
[0063] The road traffic element recognition training step is trained and optimized through end-to-end deep learning, using algorithmic frameworks such as sequence-to-sequence loss functions, dynamic event detection loss, or multi-task loss to achieve accurate perception and intelligent prediction of events and states in complex and changing environments, significantly improving the intelligent agent's global understanding and dynamic adaptability to complex scenarios.
[0064] The decision module consists of an input perceptual encoding unit, a global scene perception module, a policy network, a multi-task optimization component, and an execution output layer. The input perceptual encoding unit encodes the raw image, target information, and real-time dynamic data streams (including multimodal temporal features from sensors, agent and environment states, etc.) into a joint representation, which is then fed into the global scene perception module. This module, using a fusion of attention mechanisms and contextual association modeling, performs a global and dynamic understanding of multiple targets and environmental elements in the scene, outputting high-dimensional features of the current state. The downstream multi-task policy optimization network utilizes reinforcement learning as its foundational framework, typically employing deep reinforcement learning algorithms (such as DQN, DDPG, PPO, SAC, etc.) or their multi-task variants. The policy network makes autonomous decisions about the environment state and outputs a series of candidate actions and their probability distributions, achieving coordinated optimization of various objectives (such as performance, stability, and energy consumption). During training, the policy network parameters are continuously optimized through end-to-end backpropagation using environmental feedback (reward signals), enabling adaptive dynamic adjustments and multi-task trade-offs. The module's overall process involves perception, fusion, global understanding, policy decision-making, execution feedback, and continuous optimization. It enables dynamic prediction of agent behavior and path planning in complex dynamic environments and under multi-objective constraints, while continuously improving its policy robustness and generalization capabilities. Furthermore, to improve the efficiency of multi-objective optimization, mechanisms such as hierarchical reinforcement learning, multi-objective reward function design, and adaptive experience replay can be introduced to enable the policy network to efficiently adapt to complex scenarios and dynamic objectives and output optimal behavior. By introducing a global scene perception module and a multi-task policy optimization network, dynamic prediction of behavior and multi-objective optimized path planning are achieved (this differs from existing end-to-end approaches that do not utilize reinforcement learning and lack dynamic modeling of real-world objectives). Ultimately, an end-to-end perception, decision-making, and control model capable of adapting to complex dynamic environments can be obtained.
[0065] Next, the hardware and software environment of the experiment of the present invention, the data set used for the experiment, the experimental settings and the experimental evaluation indicators are introduced in detail.
[0066] (1) Experimental environment.
[0067] The detailed information of the environment configuration is shown in Table 1.
[0068] Table 1 Experimental environment configuration
[0069] (2) Experimental dataset.
[0070] This paper verifies the proposed method on classic datasets, namely NuScenes and Waymo datasets, and tries multiple experimental settings.
[0071] (3) Experimental setup.
[0072] This paper compares its results with those of traditional models (UniAD, VAD, and BEV-Planner); pure multimodal large models (LLava and Vicuna); and end-to-end models based on multimodal large models (EMMA and Omnidrive). This paper uses metrics provided by nuScenes for overall evaluation, including L2, collision rate, and boundary rate. Experiments show that the end-to-end perception, decision-making, and control architecture proposed in this paper outperform existing models.
[0073] In summary, this invention integrates multimodal data into an end-to-end architecture by introducing a large multimodal model. Combined with weakly supervised learning methods, this method incorporates multi-source information, including scene information described in language, for deep fusion, effectively improving perception accuracy and scene understanding capabilities. Progressively progressing from accurate recognition of basic traffic signals to robust perception of dynamic road elements, and ultimately completing driving decision training, this approach enhances the adaptability, real-time performance, and safety of autonomous driving systems in complex and open environments. It also enhances the interpretability of driving decisions, overcoming the black-box nature of existing end-to-end models and meeting the practical needs of highly complex driving scenarios. Multimodal cross-training enables generalization of multi-source traffic rule data. Leveraging the pre-training capabilities of language models, abstract rules are embedded into the model, deepening traffic semantic understanding. Enriched sensor fusion information enhances the system's adaptability to extreme scenarios (such as low visibility and irregularly moving objects). Time-series dynamic modeling improves predictive capabilities for complex driving scenarios. This end-to-end process avoids the information loss inherent in traditional layered architectures, significantly improving response efficiency and interference immunity. The deep fusion of multimodal information makes the decision-making process more robust and reasonable in complex scenarios.
[0074] The end-to-end perception decision-making and control architecture determination device provided by the present invention is described below. The end-to-end perception decision-making and control architecture determination device described below and the end-to-end perception decision-making and control architecture determination method described above can be referenced to each other.
[0075] Figure 4 This is a schematic diagram of the structure of the end-to-end perception decision-making and control architecture determination device provided by the present invention. Figure 4 As shown, the present invention also provides an end-to-end perception decision-making and control architecture determination device, including the following modules.
[0076] A first training module 410 is configured to train a preset multimodal large model based on a first multimodal dataset constructed based on basic traffic signal images, road map data, and basic driving rules of an autonomous driving scenario; A second training module 420 is configured to train the multimodal large model trained with the first multimodal dataset based on a second multimodal dataset constructed based on image sequence data of dynamic traffic elements, image sequence data of complex environmental elements, and complex driving rules of an autonomous driving scenario; A third training module 430 is configured to perform end-to-end training on the multimodal large model trained on the second multimodal dataset based on a third multimodal dataset constructed from autonomous driving scene images, autonomous driving targets, and sensor time series data, to obtain an end-to-end perception, decision-making, and control architecture suitable for autonomous driving scenarios; The autonomous driving scene image includes one or more of basic traffic signals, road map data, dynamic traffic elements and complex environment elements.
[0077] As an embodiment, the multimodal large model includes a cross-modal feature alignment module and a rule understanding module; The cross-modal feature alignment module is used to extract basic semantic features from the basic traffic signal image, the road map data, and the basic driving rules, and map the basic semantic features into the same semantic embedding space to complete semantic feature alignment and semantic feature fusion to obtain basic semantic fusion features; The rule understanding module is used to convert the basic driving rules into a structured semantic representation.
[0078] As an embodiment, the road map data includes environmental map information.
[0079] As an embodiment, the multimodal large model further includes a multimodal perception module and a modal feature cross-decoding module; The multimodal perception module is used to extract the temporal characteristics of the dynamic traffic element image sequence data, the complex environment element image sequence data, and the complex driving rules; The modal feature cross-decoding module is used to interactively decode and globally aggregate the temporal features to generate a unified high-order representation.
[0080] As an embodiment, the multimodal large model further includes a decision module; The decision module is used to encode the autonomous driving scene image, the autonomous driving target and the sensor time series data into a joint representation, input the structured semantic representation, the unified high-order representation and the joint representation into a reinforcement learning model, and generate a decision, which is used to represent a set of autonomous driving control instructions.
[0081] In one embodiment, the third multimodal dataset includes a simulated multimodal data subset and a real multimodal data subset. Correspondingly, the multimodal large model trained on the second multimodal dataset is end-to-end trained to obtain an end-to-end perception, decision-making, and control architecture suitable for autonomous driving scenarios, including: Based on the simulated multimodal data subset, simulation training is performed on the multimodal large model after complex recognition training, and during the simulation training process, parameters of the multimodal large model are optimized based on an end-to-end back-propagation mechanism; Based on the real multimodal data subset, the parameters of the multimodal large model after simulation training are adjusted to obtain the end-to-end perception decision-making and control architecture.
[0082] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the end-to-end perception decision-making and control architecture determination method, which includes: A first multimodal dataset constructed based on basic traffic signal images, road map data, and basic driving rules for autonomous driving scenarios is used to train a preset multimodal large model; A second multimodal dataset constructed based on dynamic traffic element image sequence data, complex environmental element image sequence data, and complex driving rules of an autonomous driving scenario is used to train the multimodal large model trained on the first multimodal dataset; A third multimodal dataset constructed based on autonomous driving scene images, autonomous driving targets, and sensor time series data is used to perform end-to-end training on the multimodal large model trained on the second multimodal dataset, thereby obtaining an end-to-end perception, decision-making, and control architecture suitable for autonomous driving scenarios. The autonomous driving scene image includes one or more of basic traffic signals, road map data, dynamic traffic elements and complex environment elements.
[0083] Furthermore, the logic instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0084] In another aspect, the present invention further provides a computer program product, comprising a computer program. The computer program may be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the end-to-end perception decision-making and control architecture determination method provided by each of the above methods, the method comprising: A first multimodal dataset constructed based on basic traffic signal images, road map data, and basic driving rules for autonomous driving scenarios is used to train a preset multimodal large model; A second multimodal dataset constructed based on dynamic traffic element image sequence data, complex environmental element image sequence data, and complex driving rules of an autonomous driving scenario is used to train the multimodal large model trained on the first multimodal dataset; A third multimodal dataset constructed based on autonomous driving scene images, autonomous driving targets, and sensor time series data is used to perform end-to-end training on the multimodal large model trained on the second multimodal dataset, thereby obtaining an end-to-end perception, decision-making, and control architecture suitable for autonomous driving scenarios. The autonomous driving scene image includes one or more of basic traffic signals, road map data, dynamic traffic elements and complex environment elements.
[0085] In yet another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program is implemented to perform the end-to-end perception decision-making and control architecture determination method provided by the above methods, the method comprising: A first multimodal dataset constructed based on basic traffic signal images, road map data, and basic driving rules for autonomous driving scenarios is used to train a preset multimodal large model; A second multimodal dataset constructed based on dynamic traffic element image sequence data, complex environmental element image sequence data, and complex driving rules of an autonomous driving scenario is used to train the multimodal large model trained on the first multimodal dataset; A third multimodal dataset constructed based on autonomous driving scene images, autonomous driving targets, and sensor time series data is used to perform end-to-end training on the multimodal large model trained on the second multimodal dataset, thereby obtaining an end-to-end perception, decision-making, and control architecture suitable for autonomous driving scenarios. The autonomous driving scene image includes one or more of basic traffic signals, road map data, dynamic traffic elements and complex environment elements.
[0086] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0087] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for determining an end-to-end perception, decision-making, and regulation control architecture, characterized in that: include: A first multimodal dataset constructed based on basic traffic signal images, road map data, and basic driving rules for autonomous driving scenarios is used to train a preset multimodal large model; A second multimodal dataset constructed based on dynamic traffic element image sequence data, complex environmental element image sequence data, and complex driving rules of an autonomous driving scenario is used to train the multimodal large model trained on the first multimodal dataset; A third multimodal dataset constructed based on autonomous driving scene images, autonomous driving targets, and sensor time series data is used to perform end-to-end training on the multimodal large model trained on the second multimodal dataset, thereby obtaining an end-to-end perception, decision-making, and control architecture suitable for autonomous driving scenarios. The autonomous driving scene image includes one or more of basic traffic signals, road map data, dynamic traffic elements and complex environment elements.
2. The method for determining an end-to-end perception decision-making and control architecture according to claim 1, characterized in that: The multimodal large model includes a cross-modal feature alignment module and a rule understanding module; The cross-modal feature alignment module is used to extract basic semantic features from the basic traffic signal image, the road map data, and the basic driving rules, and map the basic semantic features into the same semantic embedding space to complete semantic feature alignment and semantic feature fusion to obtain basic semantic fusion features; The rule understanding module is used to convert the basic driving rules into a structured semantic representation.
3. The method for determining an end-to-end perception decision-making and control architecture according to claim 2, characterized in that: The road map data includes environmental map information.
4. The method for determining an end-to-end perception decision-making and control architecture according to claim 2, wherein: The multimodal large model also includes a multimodal perception module and a modal feature cross-decoding module; The multimodal perception module is used to extract the temporal characteristics of the dynamic traffic element image sequence data, the complex environment element image sequence data, and the complex driving rules; The modal feature cross-decoding module is used to interactively decode and globally aggregate the temporal features to generate a unified high-order representation.
5. The method for determining an end-to-end perception decision-making and control architecture according to claim 4, characterized in that: The multimodal large model also includes a decision module; The decision module is used to encode the autonomous driving scene image, the autonomous driving target and the sensor time series data into a joint representation, input the structured semantic representation, the unified high-order representation and the joint representation into a reinforcement learning model, and generate a decision, which is used to represent a set of autonomous driving control instructions.
6. The method for determining an end-to-end perception, decision-making, and regulation control architecture according to any one of claims 1 to 5, wherein: The third multimodal data set includes a simulated multimodal data subset and a real multimodal data subset. Correspondingly, the multimodal large model trained on the second multimodal data set is end-to-end trained to obtain an end-to-end perception, decision-making and control architecture suitable for autonomous driving scenarios, including: Based on the simulated multimodal data subset, simulation training is performed on the multimodal large model after complex recognition training, and during the simulation training process, parameters of the multimodal large model are optimized based on an end-to-end back-propagation mechanism; Based on the real multimodal data subset, the parameters of the multimodal large model after simulation training are adjusted to obtain the end-to-end perception decision-making and control architecture.
7. An end-to-end perception decision-making and control architecture determination device, characterized in that: include: A first training module is used to train a preset multimodal large model based on a first multimodal dataset constructed based on basic traffic signal images, road map data, and basic driving rules of an autonomous driving scenario; A second training module is configured to train the multimodal large model trained on the first multimodal dataset based on a second multimodal dataset constructed based on image sequence data of dynamic traffic elements, image sequence data of complex environmental elements, and complex driving rules of an autonomous driving scenario; A third training module is configured to perform end-to-end training on the multimodal large model trained on the second multimodal dataset based on a third multimodal dataset constructed from autonomous driving scene images, autonomous driving targets, and sensor time series data, thereby obtaining an end-to-end perception, decision-making, and control architecture suitable for autonomous driving scenarios; The autonomous driving scene image includes one or more of basic traffic signals, road map data, dynamic traffic elements and complex environment elements.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the end-to-end perception decision-making and control architecture determination method as described in any one of claims 1-6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for determining an end-to-end perception decision-making and control architecture as described in any one of claims 1-6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for determining an end-to-end perception decision-making and control architecture as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Automatic driving model, training method, automatic driving method and vehicle
CN116880462A
Intelligent driving simulation test method and system for internal combustion locomotive
CN120178700A
Cited By
Intelligent driving system model training method and device and vehicle
CN121351012A
A smart driving system model training method and device, and a vehicle
CN121351012B
Two-wheeled vehicle automatic driving road condition decision-making method and device based on near-end strategy optimization
CN121671669A
Traffic perception system and method based on vision measurement
CN121725438A