Driving control method and moving device
By using a multimodal diffusion model to fill sensor data to generate control sub-information, the problem of the lack of rigor in causal logic in autonomous driving systems is solved, thereby improving the reliability and safety of driving control for mobile devices.
Patent Information
- Application Number
- CN202610320721.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-16
- Publication Date
- 2026-06-30
AI Technical Summary
In existing autonomous driving systems, large language models suffer from difficulties in quantifying textual reasoning and poor interpretability of the decision-making process in driving control decisions. This results in a lack of rigor in the causal logic of driving decisions, making it difficult to meet the reliability and safety requirements of autonomous driving control for mobile devices.
A multimodal diffusion model is adopted, which fills the prompt template by acquiring sensor data from multiple angles and generates control sub-information at multiple times. The control sub-information is generated by combining the prompt information and the time sequence, so as to ensure the rigor of the causal logic of driving decisions.
It improves the reliability and safety of mobile device driving control, and ensures the accuracy and consistency of driving decisions through multimodal feature matching and time-series continuous control sub-information generation.
Smart Images

Figure CN122308359A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of mobile device technology, and to, but is not limited to, a driving control method and a mobile device. Background Technology
[0002] With the rapid development of artificial intelligence and intelligent transportation technologies, autonomous driving systems are gradually becoming an important direction for improving traffic efficiency and driving safety. Autonomous driving systems can rely on large language models (LLM) to extract features and make control decision inferences from the raw data collected by sensors, thereby realizing the autonomous driving control of mobile devices.
[0003] In related technologies, the direct application of large language models to driving control decisions is still in the exploratory stage. There are problems such as difficulty in quantifying text reasoning and poor interpretability of the decision-making process, which make it difficult to meet the reliability and safety requirements of autonomous driving control of mobile devices. Summary of the Invention
[0004] The driving control method and mobile device provided in this application can obtain prompt information by filling prompt templates with sensor data through a multimodal diffusion model, and generate multiple control sub-information by combining the prompt information and time sequence, thereby ensuring the causal logic rigor of driving decisions and improving the reliability of mobile device driving.
[0005] The first aspect of this application provides a driving control method, including: The system acquires the target driving task, target driving status data, and target sensor data of the mobile device; wherein the target sensor data includes multiple images of the same location, each image being acquired from a different angle. A target prompt template is generated based on the target driving task and the target driving status data; wherein, the target prompt template includes content to be filled in; The target sensor data and the target prompt template are input into a pre-trained multimodal diffusion model to obtain target driving control information output by the multimodal diffusion model; wherein, the target driving control information includes multiple control sub-information corresponding to multiple time points; the multimodal diffusion model is used to fill the content to be filled in the target prompt template according to the target sensor data to obtain target prompt information, and to generate the multiple control sub-information according to the target prompt information in the chronological order of the multiple time points; The mobile device is controlled to move according to the multiple control sub-information.
[0006] By implementing the above technical solution, the target driving task, target driving state data, and target sensor data containing multiple angle images of the same location are acquired for the mobile device. A target prompt template with content to be filled is generated by combining the target driving task and target driving state data. The target sensor data and the target prompt template are input into a pre-trained multimodal diffusion model. The multimodal diffusion model fills in the content of the target prompt template based on the target sensor data, obtaining target prompt information. Furthermore, by combining the target prompt information and the chronological order of multiple moments, multiple control sub-information corresponding to each moment is generated. The mobile device is controlled based on these multiple control sub-information. This approach enables comprehensive perception of the driving environment based on multi-angle images of the same location, providing complete and accurate environmental data support for driving decisions. Simultaneously, the construction and filling of the target prompt template forms a standardized decision-making reasoning logic, ensuring the rigor of the causal logic of driving decisions. Furthermore, the generation of sequentially continuous control sub-information makes the mobile device's driving control more coherent and precise, improving the reliability and safety of the mobile device's driving control.
[0007] As an optional implementation, in the first aspect of the embodiments of this application, the step of obtaining target prompt information by filling the content to be filled in the target prompt template according to the target sensor data includes: Visual features and text features are extracted from the target sensor data and the target prompt template, respectively; The visual features and text features are aligned and concatenated to generate a fused feature; The target prompt information is obtained by filling the content to be filled in the target prompt template according to the fusion features.
[0008] By implementing the above technical solution, visual features are extracted from target sensor data, and textual features are extracted from target prompt templates. The visual and textual features are first aligned, then concatenated to generate a fused feature. This fused feature is then used to fill in the content to be filled in the target prompt template, resulting in target prompt information. This effectively solves the problem of heterogeneity between visual and textual modal features, achieving accurate matching of multimodal features in the semantic space. Simultaneously, the fused feature integrates visual perception information from the driving scenario with textual guidance information from the prompt template, making the filled target prompt information more closely aligned with the actual driving task and state of the mobile device, thus making it more targeted and reasonable. This provides a decision-making basis for subsequent generation of control sub-information, improves the rigor of driving decisions, and ensures the accuracy and reliability of mobile device driving control.
[0009] As an optional implementation, in a first aspect of this application, the multimodal diffusion model includes a visual encoder, a text encoder, a connector, and a diffusion backbone network; the visual encoder is used to extract the visual features from the target sensor data; the text encoder is used to extract the text features from the target prompt template; the connector is used to align and splice the visual features and the text features to generate the fused features; the diffusion backbone network is used to fill the content to be filled in the target prompt template according to the fused features to obtain the target prompt information, and to generate the multiple control sub-information according to the target prompt information in the chronological order of the multiple time points.
[0010] By implementing the above technical solution, the multimodal diffusion model can include a visual encoder, a text encoder, a connector, and a diffusion backbone network. The visual encoder extracts visual features from target sensor data, the text encoder extracts text features from target prompt templates, the connector aligns and splices visual and text features to generate fused features, and the diffusion backbone network is used to fill in the content to be filled based on the fused features to obtain target prompt information, and to generate multiple control sub-information by combining the target prompt information and the time sequence. The diffusion backbone network completes the prompt information generation and control sub-information output, achieving integrated connection from scene perception to decision reasoning to control output. This avoids decision bias caused by disconnections between links, improves the overall adaptability and robustness of the model, and further ensures the accuracy and reliability of mobile device driving control.
[0011] As an optional implementation, in a first aspect of the embodiments of this application, the target prompt information includes a first reasoning step and a second reasoning step, and the generation of the plurality of control sub-information according to the target prompt information in the chronological order of the plurality of times includes: The first inference step is executed based on the diffusion backbone network, the target prompt information, and the fusion features to generate a time step sequence; wherein, the time step sequence includes multiple mask markers corresponding to the multiple time points; The second inference step is executed based on the diffusion backbone network, the target prompt information, the time step sequence, and the fusion features. The control sub-information corresponding to each mask mark is obtained in the order of the multiple time points to generate the multiple control sub-information.
[0012] By implementing the above technical solution, the target prompt information includes a first inference step and a second inference step. Relying on the diffusion backbone network, the first inference step is executed by combining the target prompt information and fusion features to generate a time-step sequence containing mask markers corresponding to multiple time points. The diffusion backbone network further combines the target prompt information, the time-step sequence, and the fusion features to execute the second inference step, generating control sub-information corresponding to each mask marker according to the chronological order of the multiple time points, ultimately generating multiple control sub-information pieces. This step-by-step inference logic decomposes the generation process of control sub-information, making the decision-making process for generating control sub-information more hierarchical and rigorous, avoiding the temporal disorder and matching bias problems caused by generating control information at multiple time points through a single inference step.
[0013] As an optional implementation, in the first aspect of the embodiments of this application, the control sub-information includes at least one control parameter, each control parameter consisting of at least two parameter parts of different priorities, and the diffusion backbone network generates the control sub-information in the following manner: The diffusion backbone network generates parameter parts of each priority according to the target prompt information, the time step sequence and the fusion features, in a preset priority order; wherein, parameter parts of all control parameters of the same priority are generated in parallel.
[0014] By implementing the above technical solution, the control sub-information includes at least one control parameter, and each control parameter can be composed of at least two parameter parts with different priorities. Based on target prompt information, time step sequences, and fusion features, a diffusion backbone network generates parameter parts of each priority according to a preset priority order, and generates parameter parts of the same priority for all control parameters in parallel. This ensures a clear hierarchical order in the generation process of control parameters, avoiding logical confusion and timing conflicts caused by the mixed generation of multiple control parameters and parameter parts. Simultaneously, the parallel generation of parameter parts of the same priority effectively improves the generation efficiency of control sub-information and reduces decision-making inference latency. The preset priority order ensures the rationality and relevance of control parameter generation, improving the accuracy and reliability of driving control.
[0015] As an optional implementation, in the first aspect of the embodiments of this application, the content to be filled in the target prompt template includes environmental perception content, decision rationale content, and action suggestion content; the environmental perception content, the decision rationale content, and the action suggestion content are filled in parallel based on the target sensor data; the environmental perception content is used to characterize the environmental perception results of the environment in which the mobile device is located; the decision rationale content is used to characterize the reasoning basis for determining the target driving control information based on the environmental perception results; and the action suggestion content is used to indicate the target driving control information generated based on the reasoning basis.
[0016] By implementing the above technical solution, the content to be filled in the target prompt template includes environmental perception content, decision rationale content, and action suggestion content. Environmental perception content represents the perception results of the driving environment, decision rationale content provides the basis for decision deduction, and action suggestion content indicates the final driving control information, so as to realize the structured and standardized filling of the prompt template content, so that the target prompt information has a clear hierarchy of environmental perception, logical deduction and action guidance, and the parallel filling method can improve the content generation efficiency.
[0017] As an optional implementation, in the first aspect of this application, the multimodal diffusion model is obtained by training an initial model based on a training dataset; the training dataset includes reference sensor data, reference prompt templates, reference prompt information, and reference driving control information; the training phase of the multimodal diffusion model includes a first training phase and a second training phase; before inputting the target sensor data and the target prompt template into the pre-trained multimodal diffusion model to obtain the target driving control information output by the multimodal diffusion model, the method further includes: In the first training phase, the initial model is trained based on the reference sensor data, the reference prompt template, and the reference prompt information; In the second training phase, the initial model is trained based on the reference sensor data, the reference prompt template, the reference prompt information, and the reference driving control information.
[0018] By implementing the above technical solution, the initial model learns the basic capabilities of multimodal feature processing, prompt template filling, and prompt information generation in the first training phase. In the second training phase, reference driving control information is further introduced, enabling the model to learn the complete decision-making logic from prompt information reasoning to the final driving control information output. This phased training reduces the learning difficulty of the model, allowing it to gradually master the entire chain of environmental perception, reasoning generation, and control output capabilities. It ensures a high degree of logical consistency between the target prompt information generated by the model and the target driving control information, thereby improving the accuracy and reliability of the multimodal diffusion model's output results.
[0019] As an optional implementation, in a first aspect of the embodiments of this application, the step of training the initial model based on the reference sensor data, the reference prompt template, and the reference prompt information during the first training phase includes: In the first training phase, the reference sensor data and the reference prompt template are input into the initial model to obtain the predicted prompt information output by the initial model; the parameters of the initial model are updated according to the difference between the predicted prompt information and the reference prompt information. In the second training phase, the initial model is trained based on the reference sensor data, the reference prompt template, the reference prompt information, and the reference driving control information, including: The reference sensor data, the reference prompt template, and the reference prompt information are input into the initial model to obtain the predicted driving control information output by the initial model; the parameters of the initial model are updated according to the difference between the predicted driving control information and the reference driving control information.
[0020] By implementing the above technical solution, the first training phase uses reference prompt information as the supervised benchmark. Reference sensor data and a reference prompt template are input into the initial model to obtain predicted prompt information, and the model parameters are updated based on the difference between the two. The second training phase uses reference driving control information as the supervised benchmark. Reference sensor data, a reference prompt template, and reference prompt information are input into the initial model to obtain predicted driving control information, and the model parameters are updated based on the difference between the predicted driving control information and the reference driving control information. This avoids the task coupling and parameter optimization bias problems caused by single-stage training, improving model convergence efficiency and training stability. Simultaneously, the phased supervised learning mechanism provides clear guidance for model parameter updates, improving the inference reliability and decision-making accuracy of the multimodal diffusion model.
[0021] As an optional implementation, in the first aspect of the embodiments of this application, the initial model includes a visual encoder, a text encoder, a connector, and an initial diffusion backbone network; updating the parameters of the initial model includes: Update the parameters of the initial diffusion backbone network.
[0022] By implementing the above technical solution, the initial model consists of a visual encoder, a text encoder, a connector, and an initial diffusion backbone network. The visual encoder, text encoder, and connector can be transferred to the multimodal diffusion model without additional training. Only the parameters of the initial diffusion backbone network are updated, ensuring the accuracy of multimodal feature extraction and fusion, and improving training efficiency.
[0023] A second aspect of this application provides a mobile device including a memory and a processor, the memory storing a computer program executable on the processor, the processor executing the program to implement the method described in the first aspect of this application. Attached Figure Description
[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.
[0025] Figure 1 This is a flowchart illustrating a driving control method disclosed in an embodiment of this application. Figure 1 ; Figure 2 This is a schematic diagram of the inference process of a pre-trained multimodal diffusion model disclosed in an embodiment of this application. Figure 1 ; Figure 3 This is a schematic diagram of the inference process of a pre-trained multimodal diffusion model disclosed in an embodiment of this application. Figure 2 ; Figure 4 A schematic diagram illustrating the generation of multiple control sub-information for the multimodal diffusion model disclosed in this application embodiment; Figure 5 This is a schematic diagram of the inference process of a pre-trained multimodal diffusion model disclosed in an embodiment of this application; Figure 6 This is a flowchart illustrating a driving control method disclosed in an embodiment of this application. Figure 2 ; Figure 7 This is a structural block diagram of a driving control device disclosed in an embodiment of this application; Figure 8 This is a structural block diagram of a mobile device disclosed in an embodiment of this application. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of this application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0028] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0029] It should be noted that the terms "first, second, third" used in the embodiments of this application are used to distinguish similar or different objects and do not represent a specific order of objects. It can be understood that "first, second, third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0030] With the rapid development of artificial intelligence and intelligent transportation technologies, autonomous driving systems are gradually becoming an important direction for improving traffic efficiency and driving safety. Autonomous driving systems typically integrate one or more sensor modules to collect environmental perception data around the mobile device, and then use a central computing unit to complete data fusion, feature analysis, and control decision-making processes.
[0031] Among related technologies, Large Language Models (LLMs), with their superior natural language understanding, contextual semantic association, and complex logical reasoning capabilities, are widely used in the decision-making and reasoning stages of autonomous driving systems, providing guidance for driving control to the underlying actuators of mobile devices. However, the reasoning process of LLMs typically relies on natural language text, lacks explicit reasoning steps, and is difficult to quantify. Furthermore, the decision-making chain is opaque and lacks interpretability, resulting in a lack of rigor in the causal logic of driving decisions, affecting the accuracy of driving control, and making it difficult to meet the reliability and safety requirements of autonomous driving control for mobile devices.
[0032] In view of this, embodiments of this application provide a driving control method, which includes: acquiring a target driving task, target driving state data, and target sensor data of a mobile device; wherein, the target sensor data includes multiple images of the same location, each image being acquired from a different angle; generating a target prompt template based on the target driving task and target driving state data; wherein, the target prompt template includes content to be filled; inputting the target sensor data and the target prompt template into a pre-trained multimodal diffusion model to obtain target driving control information output by the multimodal diffusion model; wherein, the target driving control information includes multiple control sub-information corresponding to multiple time points; the multimodal diffusion model is used to fill the content to be filled in the target prompt template based on the target sensor data to obtain target prompt information, and to generate multiple control sub-information according to the target prompt information in a sequential order of multiple time points; controlling the driving of the mobile device based on the multiple control sub-information. In this embodiment of the application, the prompt information is obtained by filling the prompt template based on sensor data using a multimodal diffusion model, and multiple control sub-information are generated by combining the prompt information and the time sequence, ensuring the causal logic rigor of driving decisions and improving the reliability of mobile device driving.
[0033] The driving control method proposed in this application can be applied to mobile devices, such as vehicles, drones, airplanes, ships, robots, etc. This application does not limit the application to these devices.
[0034] It should be noted that in the exemplary applications of vehicles provided in this application, the vehicles can be implemented as cars, commercial vehicles, special operation vehicles (such as fire trucks, emergency rescue vehicles, police cars, ambulances, etc.), motorcycles, electric bicycles, agricultural and engineering machinery vehicles (such as tractors, harvesters, excavators, bulldozers, etc.), special robot vehicles, etc.
[0035] To make the objectives and technical solutions of this application clearer and more intuitive, the driving control method disclosed in this application will be described in detail below with reference to the accompanying drawings. It should be understood that the execution subject of the embodiments of this application may also be a processor or chip in a mobile device.
[0036] To facilitate the explanation and description of the driving control method provided in the embodiments of this application, the following description mainly uses the application of the control method for a mobile device to a vehicle as an example in the relevant drawings. Please refer to... Figure 1 , Figure 1 This is a flowchart illustrating a driving control method disclosed in an embodiment of this application. Figure 1 ,like Figure 1 The method shown may include the following steps: S101, acquire the target driving task, target driving status data and target sensor data of the mobile device; wherein, the target sensor data includes multiple images of the same location, each image is acquired from a different angle.
[0037] It should be noted that the mobile device in the embodiments of this application can be a vehicle, motorcycle, etc. That is to say, any mobile device can use the driving control method provided in this application.
[0038] In this embodiment, the target driving task is a driving planning or driving operation task that the mobile device needs to perform, used to indicate the driving intention of the mobile device. For example, the target driving task may be driving in the center of the lane, avoiding surrounding obstacles, turning at road intersections, or driving along a designated route. The target driving task may be in the form of structured instructions, text descriptions, code identifiers, etc., and this application does not limit the specific form of the task.
[0039] As an optional implementation, the target driving task is a task generated based on user input data; or, the target driving task is a preset task; or, the target driving task is a task generated based on navigation data.
[0040] Optionally, the target driving task is a task generated based on user input data. This type of task reflects the user's true driving intentions and personalized driving needs. User input data can be input through voice, touch operation, physical control component operation, etc. For example, user input data can be voice commands such as "turn left onto the auxiliary road" or "overtake the slow-moving vehicle ahead," or issuing lane change commands by moving the lane change lever, or planning the driving route by touching the in-vehicle terminal. After obtaining the user input data, the target driving task can be generated based on the user input data by recognition, parsing, and standardization conversion, adapting to the inference requirements of the multimodal diffusion model.
[0041] For example, the input data is the voice command "turn left into the auxiliary road". After the command parsing module recognizes and parses the voice data, it generates a standardized target driving task of "turn left at the left turn sign ahead and enter the auxiliary road, keeping the lane in the center of the auxiliary road".
[0042] Optionally, the target driving task is a preset task, such as a basic driving task pre-set by the vehicle system for regular driving scenarios without navigation or additional user driving instructions. It does not require user input data or navigation data support and can adapt to the general driving needs of mobile devices roaming without navigation.
[0043] For example, preset tasks may include general tasks such as roaming in the center of the lane without navigation, avoiding dynamic and static obstacles in real time, staying away from dedicated driving lanes such as public transportation lanes, autonomously choosing the driving direction at road intersections according to traffic rules, and dynamically adjusting the driving speed according to road speed limits. These tasks may also be tasks determined in advance based on the user's historical driving behavior. This application does not limit them here.
[0044] Optionally, the target driving task is a task generated based on navigation data. Such tasks can be generated based on the starting point, ending point, and driving route planned by the navigation, combined with navigation-related information such as road attributes, traffic signs, and speed limits. It can fit the overall driving planning requirements of the navigation and adapt to driving scenarios where users have a clear driving destination. Based on the navigation data, target driving tasks can be generated such as turning right at the designated intersection ahead along the planned route, entering the highway ramp and accelerating to the prescribed driving speed according to the highway section requirements, and completing merging, diverging, and lane changing at the planned location while maintaining the center of the lane.
[0045] For example, navigation data can be obtained through in-vehicle navigation systems, such as real-time location data obtained through satellite positioning systems, or core route planning data issued by cloud navigation service platforms based on user-defined starting and ending points. This ensures that the target driving task generated based on the navigation data matches real-time road conditions and the planned route. After obtaining the navigation data, the target driving task can be generated by identifying, parsing, and standardizing it. For instance, if the navigation data indicates a right turn at the intersection 200 meters ahead, identifying and parsing the navigation data will reveal keywords such as "200 meters," "intersection," and "right turn." These keywords can then be standardized and converted into the target driving task of "driving along the current lane to the intersection 200 meters ahead, completing the right turn, and entering the corresponding auxiliary road."
[0046] In this embodiment, the target driving state data is a set of parameters reflecting the driving state of the mobile device itself, providing underlying state support for subsequent target prompt template generation and driving decision reasoning. For example, in the scenario where the mobile device is a motor vehicle, the target driving state data includes, but is not limited to, one or more of the following vehicle state parameters: real-time driving speed, current gear status, steering angle, remaining fuel / remaining battery power, driving mode (economy mode, sport mode, etc.), braking status, tire pressure, engine speed, and steering wheel angle deviation. This application does not limit the scope of the data.
[0047] As an optional implementation, target driving status data can be acquired in real time through the sensing module and bus acquisition module of the mobile device. For example, raw status data can be acquired in real time through the vehicle's onboard CAN bus, onboard diagnostic system (OBD), and various body sensors. After preliminary filtering, noise reduction, and standardization, regularized target driving status data is obtained, avoiding the interference of raw data noise with the accuracy of subsequent decision-making and reasoning.
[0048] As an optional implementation method, the target driving state data can be divided into dynamic state data and static state data. Dynamic state data refers to parameters that change in real time during driving, such as real-time vehicle speed, steering angle, and braking status; static state data refers to parameters that remain constant for short periods during driving, such as vehicle rated load, driving mode, and tire specifications. By classifying and organizing the target driving state data, different collection frequencies can be set for different types of state data. For example, a high-frequency collection mode can be used for dynamic state data, and a low-frequency collection mode can be used for static state data. This reduces the overhead of transmitting and processing redundant data and avoids the repeated collection of static parameters that consumes computing resources.
[0049] Optionally, the target driving status data may also include auxiliary status parameters, such as parking status, light on status, and wiper working status. These parameters can be flexibly added or removed according to the actual driving scenario. This application does not limit the types of parameters or the dimensions of data collection for the target driving status data.
[0050] In this embodiment of the application, the target sensor data is the data of the mobile device perceiving the surrounding driving environment, including multiple images of the same location, each image being acquired from a different angle.
[0051] It should be noted that mobile devices can use multiple cameras deployed at different locations to create overlapping coverage of the field of view at the same environmental location, acquiring multiple images of that location from different shooting angles. The target sensor data can be acquired in the form of single-sampling-time acquisition or continuous acquisition at multiple sampling-times, and the target sensor data in both implementation methods can be used as effective input to the multimodal diffusion model to adapt to the model's perception requirements of the driving environment.
[0052] Optionally, when data acquisition is performed using a single sampling moment, the mobile device completes one data acquisition according to a preset sensor data sampling frequency, which is recorded as a single sampling moment. At this sampling moment, multiple cameras maintain synchronized acquisition timing and acquire images of the same environmental location, obtaining multiple images of the mobile device at different acquisition angles at that environmental location.
[0053] For example, within a single sampling moment, the mobile device simultaneously captures five images from different angles of its surrounding environment using five cameras positioned at different locations: the front-view main camera, the left front-view camera, the right front-view camera, the left rear-view camera, and the right rear-view camera. This generates multi-view image data at a single moment. This acquisition method can accurately reconstruct the spatial environmental characteristics of the same location at a specific instant, providing stereoscopic visual perception data for that location to support multimodal diffusion models.
[0054] Optionally, when using multiple sampling times for continuous data acquisition, the mobile device performs the above-mentioned multi-angle image acquisition operation at each of the multiple sampling times, and integrates the multi-angle images acquired at each sampling time to form multi-time, multi-view image data.
[0055] For example, during three consecutive sampling times, images of the road environment in which the mobile device is traveling are captured by five cameras at different angles at each time point, ultimately resulting in image data consisting of 15 images. This acquisition method can capture the environmental change characteristics of the mobile device in a continuous time dimension, such as the driving trajectory of other mobile devices on the road, the state switching of traffic lights, and the movement trend of pedestrians at intersections, providing time-series environmental perception data for multimodal diffusion models.
[0056] It should be noted that the time interval between sampling moments, the number of cameras participating in the acquisition at a single sampling moment, and their angles can be adapted and adjusted by those skilled in the art according to the actual driving scenario of the mobile device; and the acquisition parameters of each camera, such as frame rate, resolution, exposure, white balance, etc., can be calibrated based on the needs of the driving scenario to ensure the consistency of image data acquired at different angles and at different times in the temporal and spatial dimensions, and to avoid image feature deviations caused by differences in hardware parameters. This application does not limit the number of cameras, angles, and acquisition parameters.
[0057] As an optional implementation, in addition to image data collected from multiple angles at the same location, target sensor data may also include radar data, such as radar point cloud data, millimeter-wave radar data, and ultrasonic radar data, or one or more of these, to supplement driving environment perception information from multiple dimensions and improve the environmental perception coverage. Specifically: millimeter-wave radar data can be acquired through millimeter-wave radar deployed on the vehicle body. This data has advantages in long-range, high-refresh-rate speed and distance measurement, and can accurately capture the driving speed, relative distance, and movement trend of surrounding dynamic targets. It is particularly suitable for long-range forward vehicle perception needs in highway scenarios and can complement radar point cloud data for near- and long-range perception. Ultrasonic radar data can be acquired through ultrasonic radar deployed around the vehicle body. This data has short-range, high-precision detection characteristics and can accurately identify the position and outline features of near-range static obstacles and low obstacles around the mobile device, such as curbs, bollards, guardrails, and green belts. It is suitable for near-range perception scenarios such as urban roads, underground garages, and narrow road passages, and can complement radar point cloud data and millimeter-wave radar data to form a comprehensive perception hierarchy covering near, medium, and long ranges.
[0058] The radar data mentioned above can be combined with image data from multiple acquisition perspectives to construct a multimodal perception system for mobile devices in the driving environment. This effectively compensates for the limitations of image data in scenarios such as three-dimensional distance measurement, obstacle contour recognition, and perception in low light and adverse weather conditions, providing more comprehensive and accurate environmental feature inputs for the multimodal diffusion model.
[0059] It should be noted that radar data can be collected and generated by radar sensors mounted on the mobile device. The radar sensors can be deployed in positions such as the front of the vehicle, the front of the roof, and the sides of the vehicle to ensure that the collection range fully covers the driving area around the mobile device. This application does not limit this.
[0060] Optionally, when the target sensor data includes at least two types of perception data, spatiotemporal alignment between different types of data can be achieved through preprocessing operations such as timestamp synchronization and spatial coordinate calibration. This ensures that perception data of different types and different acquisition sources remain consistent in time and space, and ensures that all types of perception data can accurately reflect the state of the same driving environment location, providing data support for subsequent driving information inference.
[0061] As an optional implementation method, the driving control method provided in this application acquires target sensor data only after the target driving task has been determined. This approach, by clearly defining the target driving task before collecting and processing target sensor data, improves data utilization efficiency and the relevance of model inference, while avoiding resource waste and additional computational overhead caused by meaningless collection of irrelevant environmental data.
[0062] S102, Generate a target prompt template based on the target driving task and target driving status data; wherein, the target prompt template includes content to be filled in.
[0063] In this embodiment, the target prompt template is a structured reasoning framework constructed based on the driving intention of the target driving task and the working condition characteristics of the target driving state data. It contains content to be filled in. As the input of the multimodal diffusion model, the target prompt template can standardize the reasoning path of the model, limit the reasoning scope, and ensure the causal logic rigor of driving decisions. This application does not limit the form of the target prompt template.
[0064] Optionally, the target cue template can be a Chain of Thought (CoT) text template with a filled area.
[0065] It's important to note that the Thinking Chain is a text-based reasoning paradigm that simulates step-by-step logical deduction in humans. By breaking down complex decisions into multi-level, sequential steps, it makes the reasoning process explicit. When the target prompt template is a Thinking Chain text template, the template can be divided into multi-level filling areas according to the logical deduction sequence of the driving decision. These filling areas correspond to the content to be filled and are reserved blank areas for subsequent semantic content filling by the multimodal diffusion model based on target sensor data. Through the hierarchical design of the Thinking Chain, the model's reasoning process forms an explicit and traceable logical chain.
[0066] Optionally, the target hint template can be a Directed Acyclic Graph (DAG) structure template with blank nodes.
[0067] It should be noted that a directed acyclic graph (DAG) is a topological structure without closed loops and with a fixed flow direction. It consists of nodes carrying information and directed edges defining the relationships between nodes. The reasoning process can only flow unidirectionally along the directed edges, and there are no logical loops. It is commonly used in scenarios such as complex logical decision-making, multi-branch process deduction, and hierarchical task scheduling. When the target hint template is a DAG structure template, the target hint template builds a driving decision-making reasoning framework based on the nodes and directed edges of the DAG. Blank nodes carry the content to be filled, and directed edges represent the causal relationships and execution order between nodes, strictly limiting the unique direction of reasoning flow. This allows it to adapt to the multi-branch decision-making needs in complex driving scenarios, preserving the complete area of content to be filled while using the acyclic topology to constrain the rigor of the reasoning logic, avoiding logical disorder in the decision-making chain, and improving the accuracy of subsequent driving control information generation.
[0068] As an optional implementation, a target prompt template is generated based on the target driving task and target driving status data, including: Based on the target driving task and target driving status data, the preset initial template is filled in to generate the target prompt template.
[0069] It should be noted that the initial template is a pre-built, standardized template with a general reasoning framework. It internally reserves two types of filling areas: one type is filled using target driving task and target driving state data, and the other type is filled using a multimodal diffusion model based on target sensor data. The type of this initial template is consistent with the type of the target prompt template. That is, if the target prompt template uses a thought chain text template, the initial template is also a thought chain initial template of the corresponding format; if the target prompt template uses a directed acyclic graph structure template, the initial template is also a directed acyclic graph initial template of the corresponding format, ensuring template structure compatibility and conflict-free filling processes. When generating the target prompt template, only the acquired target driving task and target driving state data need to be used as known information and filled into the corresponding filling areas of the initial template. A target prompt template with a structured reasoning framework can be quickly generated without relying on model recognition.
[0070] For example, the initial template is a thought chain initial template, such as "Traffic light status __, vehicles ahead __, pedestrians ahead __, weather __, current driving presents __ risk, current operating status __; please combine driving status and surrounding perception information to execute __". If the target driving task is roaming, and the target driving status data is current speed 30km / h and gear D, fill the above information into the corresponding fill area in the initial template, and the generated target prompt template will be "Traffic light status __, vehicles ahead __, pedestrians ahead __, weather __, current driving presents __ risk, current operating status is current speed 30km / h and gear D; please combine driving status and surrounding perception information to execute the roaming driving task".
[0071] Optionally, the filling areas inside the initial template can be distinguished and marked to ensure that the target driving task and target driving status data are accurately filled into specific filling areas in the initial template.
[0072] It should be noted that character markers, field labels, node attributes, and other methods can be used to differentiate and mark the two types of fill areas within the initial template. For example, pre-defined exclusive symbols and attribute fields can be used to mark fill areas for target driving task and target driving status data; at the same time, fill areas filled based on target sensor data using a multimodal diffusion model can be differentiated and marked. Different markings can avoid problems such as misalignment and content confusion, ensuring the accuracy and reliability of the generated target prompt template.
[0073] As an optional implementation, generating a target prompt template based on the target driving task and target driving state data includes: inputting the target driving task and target driving state data into a pre-trained language model to obtain the target prompt template output by the language model.
[0074] It should be noted that the pre-trained language model can be a lightweight inference model obtained by supervised training of an initial language model based on a sample dataset. The initial language model can be a small language model based on the Transformer architecture (such as lightweight models like DistilBERT or MobileBERT). The sample dataset includes sample driving tasks, sample driving state data, and sample prompt templates. Through iterative training, the mapping relationship between driving tasks, driving states, and prompt templates is fitted, enabling the language model to output adapted, structured prompt templates based on the input task and state data.
[0075] S103, Input the target sensor data and target prompt template into the pre-trained multimodal diffusion model to obtain the target driving control information output by the multimodal diffusion model; wherein, the target driving control information includes multiple control sub-information corresponding to multiple time points; the multimodal diffusion model is used to fill the content to be filled in the target prompt template according to the target sensor data to obtain the target prompt information, and to generate multiple control sub-information according to the target prompt information in the order of multiple time points.
[0076] In this embodiment, the multimodal diffusion model is a network model built on a diffusion backbone network, capable of cross-modal data processing. It can process target sensor data and target prompt templates from different modalities, and output target driving control information adapted to the current driving scenario through a progressive denoising generation paradigm. Specifically, the multimodal diffusion model is used to fill in the content to be filled in the target prompt template based on the target sensor data to obtain target prompt information, and to generate multiple control sub-information based on the target prompt information in a sequential order across multiple time points.
[0077] It should be noted that the target driving control information is used to control the mobile device to drive according to the target driving task. The target driving control information can be at least one of the target driving control command and the target driving trajectory information. The target driving control command may include, but is not limited to, driving speed command, steering angle command, etc., used to directly control the driving operation of the mobile device; while the target driving trajectory information may include the coordinates of the driving path and driving posture within a preset time period in the future, used to guide the mobile device to drive along the preset trajectory.
[0078] It is understood that both the target driving control command and the target driving trajectory information serve the mobile device to complete the target driving task. The two are related and can be converted into each other through a preset conversion algorithm. That is, the target driving control command can be derived and generated based on the target driving trajectory information, and the corresponding target driving trajectory information can be fitted and generated based on the target driving control command. They can be flexibly selected or combined according to the driving scenario and control requirements. This application does not limit the type of target driving control information.
[0079] In this embodiment, the target driving control information includes multiple control sub-information corresponding to multiple time points. The control sub-information is a time-sequential decomposition unit of the target driving control information. Each control sub-information corresponds to the driving control parameters at a certain time point. Multiple control sub-information are arranged sequentially in chronological order to form a complete time-sequential control link, ensuring that the mobile device can smoothly and continuously complete the target driving task.
[0080] For example, when the target driving control information is a target driving control command, the control sub-information is the driving control sub-command corresponding to a certain moment. For instance, the number of control sub-information is set to 3, namely: the driving control sub-command "vehicle speed A1km / h, steering angle B1°" corresponding to moment 1, the driving control sub-command "vehicle speed A2km / h, steering angle B2°" corresponding to moment 2, and the driving control sub-command "vehicle speed A3km / h, steering angle B3°" corresponding to moment 3. When the target driving control information is target driving trajectory information, the control sub-information is the driving trajectory sub-information corresponding to a certain moment. For instance, the number of control sub-information is set to 3, namely: the driving trajectory sub-information "position coordinates (X1,Y1)" corresponding to moment 1, the driving trajectory sub-information "position coordinates (X2,Y2)" corresponding to moment 2, and the driving trajectory sub-information "position coordinates (X3,Y3)" corresponding to moment 3.
[0081] As an optional embodiment, the content to be filled in the target prompt template includes environmental perception content, decision rationale content, and action suggestion content; the environmental perception content, decision rationale content, and action suggestion content are filled in parallel based on target sensor data; the environmental perception content is used to characterize the environmental perception results of the environment in which the mobile device is located; the decision rationale content is used to characterize the reasoning basis for determining the target driving control information based on the environmental perception results; and the action suggestion content is used to indicate the target driving control information generated based on the reasoning basis.
[0082] It should be noted that the multimodal diffusion model can simultaneously complete the parsing and filling of multiple types of content to be filled using a parallel filling method, reducing the timing delay caused by step-by-step filling and improving the model's inference efficiency. By setting multiple dimensions of content to be filled, such as environmental perception, decision rationale, and action suggestions, the rigor of the model's inference is improved.
[0083] For example, the unfilled environmental perception content could be "Traffic light status is __ lights, __ vehicles ahead, __ pedestrians ahead, __ weather, __ road condition", while the filled environmental perception content could be "Traffic light status is red, 2 vehicles ahead, 3 pedestrians ahead, cloudy weather, dry road condition". The unfilled decision rationale could be "Based on the current environmental perception results and the vehicle's driving conditions, it is necessary to follow traffic rules __ and yield to __ to ensure driving safety", while the filled decision rationale could be "Based on the current environmental perception results and the vehicle's driving conditions, it is necessary to follow..." "Follow the traffic rules of stopping at red lights and yield to pedestrians and vehicles ahead at intersections to ensure driving safety"; the unfilled action suggestion content can be "Based on the target prompt information, generate the time step sequence corresponding to the target driving control information, with the sequence format being __; fill the time step sequence according to the rule of integers first and decimals second to obtain __", and the filled action suggestion content can be "Based on the target prompt information, generate the time step sequence corresponding to the target driving control information, with the sequence format being (X1, Y1) and (X2, Y2); fill the time step sequence according to the rule of integers first and decimals second to obtain continuous time-series target driving trajectory information".
[0084] Optionally, the content to be filled in the target prompt template may also include risk assessment content, which is used to characterize the safety hazards, conflict risks and risk levels in the driving environment around the mobile device.
[0085] For example, the unfilled risk assessment content can be "Current driving presents __ risks", while the filled risk assessment content can be "Current driving presents risks of vehicles approaching from behind and cross-roading at intersections".
[0086] It should be noted that the types of content to be filled, template formats, and filling examples mentioned above are only illustrative. The content to be filled in the target prompt template can be flexibly added, deleted, or adjusted according to the driving scenario and reasoning needs. This application does not impose any restrictions on this.
[0087] As an optional implementation method, the multimodal diffusion model is a diffusion language model with multimodal processing capabilities. The multimodal diffusion model receives target sensor data, extracts features to obtain corresponding visual features, and then, based on the extracted visual features, simultaneously fills in the environmental perception content, decision rationale content, and action suggestion content within the target prompt template to generate complete target prompt information. Subsequently, the target prompt information and visual features are used as the reasoning basis for the multimodal diffusion model to sequentially generate control sub-information corresponding to each time step, thus obtaining target driving control information.
[0088] It should be noted that by using both visual features and target prompts as the basis for reasoning, we can ensure the originality and accuracy of the perceived data by relying on visual features, avoiding feature loss and information deviation during template filling. At the same time, we can use the structured reasoning logic of target prompts to standardize the generation path of time-series control sub-information, so that the control sub-information at each moment not only matches the perception results of the actual driving environment, but also conforms to the causal logic of driving decisions. In addition, by using a parallel filling method to complete the filling operation of multiple types of content, we can further shorten the reasoning time, adapt to the needs of real-time driving control, and ensure the efficiency and reliability of target driving control information generation.
[0089] As an optional implementation, target prompt information is obtained by filling the content to be filled in the target prompt template based on target sensor data, including: Visual features and textual features are extracted from the target sensor data and the target prompt template, respectively; Visual and textual features are aligned and concatenated to generate fused features; The target prompt information is obtained by filling in the content to be filled in the target prompt template based on the fusion features.
[0090] It should be noted that visual features are obtained by encoding target sensor data (such as image data and radar data) using a multimodal diffusion model, and are used to reflect the perceived information of the driving environment. Text features are obtained by semantically encoding the target prompt template using a multimodal diffusion model, and are used to reflect semantic information such as the reasoning framework and filling constraints. Spatiotemporal alignment and dimensional concatenation of the two types of features can eliminate the modal differences between target sensor data and target prompt templates, couple perceptual information and semantic information, ensure that template filling not only conforms to the real driving environment, but also follows reasoning logic, avoid problems such as filling content deviation and logical disconnect, improve the reliability and executability of target prompt information, and provide a data and logical foundation for the generation of subsequent time-sequential control sub-information.
[0091] Optionally, visual features can be extracted from the target sensor data, including: preprocessing the target sensor data, performing convolutional coding and feature dimensionality reduction, removing redundant noise data, and then extracting visual features that characterize environmental targets and road condition information.
[0092] For example, when the target sensor data includes image data (multiple images of the same location), distortion correction and stitching preprocessing are performed on multiple frames of images, and the contour and position features of objects such as traffic lights, vehicles, and pedestrians are extracted through convolutional layers; when the target sensor data includes image data and radar data (such as radar point cloud data, millimeter-wave radar data, and ultrasonic radar data, or one or more of these), the initial features corresponding to the image data and radar data are extracted separately, and then stitched together to obtain visual features.
[0093] Optionally, text features can be extracted from the target prompt template, including: segmenting the structured text of the target prompt template into words, semantic embedding encoding, and extracting text features representing the reasoning framework and the filled region.
[0094] For example, for the target prompt template "traffic light status is __, __ vehicles ahead, __ task to be performed", the text is first standardized preprocessed to retain the core semantic sentence structure and placeholder identifiers; then word segmentation is performed, splitting it into semantic units as follows: traffic light, status, is, __, ahead, vehicles, __, vehicles, need, perform, __, task; then the above word segmentation units are encoded by a semantic encoder to transform the textual filling constraints and reasoning logic into numerical feature vectors, resulting in text features.
[0095] Optionally, the visual features and text features are aligned and concatenated to generate fused features, including: mapping the visual features and text features to the same dimension and performing feature alignment processing; and fusing the aligned visual features and text features to obtain fused features.
[0096] It should be noted that linear / nonlinear mapping layers such as mapping networks and multilayer perceptrons can be used to achieve dimensional unification and alignment calibration of features from different modalities. For example, high-dimensional visual features can be mapped to the semantic feature dimension of text features to eliminate dimensional differences and distribution biases between modalities. Then, visual features and text features in the same dimension can be spliced together to finally generate regular fused features.
[0097] Optionally, the target prompt information is obtained by filling the content to be filled in the target prompt template based on the fusion features, including: inputting the fusion features into the diffusion backbone network of the multimodal diffusion model, matching the semantic requirements of the template filling area based on the fusion features, and generating corresponding semantic content to complete the filling.
[0098] It should be noted that the fused features contain both driving environment perception data and template inference constraints. The pre-trained multimodal diffusion model learns the mapping inference ability between fused features and template filling, and can match the filling area with perception information to output compliant filling content.
[0099] Optionally, target prompt information is obtained by filling in the content to be filled in the target prompt template based on the fusion features, including: The target driving scenario corresponding to the mobile device is determined based on the pre-trained multimodal diffusion model and fusion features; Based on the target driving scenario and fusion features, the target prompt information is obtained by filling in the content to be filled in the target prompt template.
[0100] It should be noted that the target driving scenario indicates the characteristics of the driving environment, road attributes, and scenario category in which the mobile device is currently and during its planned future driving period. It is a comprehensive scenario representation result obtained after macro-semantic modeling of the driving environment. The target driving scenario can include, but is not limited to, off-road, highway, urban road, rural road, construction section, wet and slippery road section, tunnel driving, and intersection passage. The target driving scenario provides a scenario-based macro-constraint and reasoning guidance for the filling process of the content to be filled in the target prompt template, ensuring that the filled target prompt information fits the actual driving scenario requirements of the mobile device, and enabling the control sub-information generated based on the target prompt information to have scenario adaptability and driving rationality.
[0101] For example, the fusion feature includes the perception information that "the road is a two-way six-lane road, there are no intersections or pedestrians, the speed limit is 120km / h, and the road surface is dry". The pre-trained multimodal diffusion model determines the target driving scenario as a high-speed driving scenario based on this fusion feature. Then, by combining the scenario feature and the fusion feature, the content to be filled in the target prompt template, such as "driving rule constraints__, vehicle speed range constraints__, trajectory planning requirements__", is filled in to obtain the target prompt information that "driving rule constraints keep the lane straight, vehicle speed range constraints are 80-120km / h, and trajectory planning requirements extend at a constant speed along the current lane", thus achieving accurate matching between the template filling content and the high-speed driving scenario.
[0102] As an optional implementation, the multimodal diffusion model includes a visual encoder, a text encoder, a connector, and a diffusion backbone network; the visual encoder is used to extract visual features from target sensor data; the text encoder is used to extract text features from the target prompt template; the connector is used to align and splice the visual features and text features to generate fused features; the diffusion backbone network is used to fill the content to be filled in the target prompt template according to the fused features to obtain target prompt information, and to generate multiple control sub-information according to the target prompt information in a chronological order of multiple time points.
[0103] For example, taking the multimodal diffusion model as an example of a diffusion-based multimodal large language model, the following will combine... Figure 2 and Figure 3 The diagram shown illustrates the inference process of a pre-trained multimodal diffusion model, providing a schematic explanation of the inference process.
[0104] Optional, such as Figure 2As shown, the input to the multimodal diffusion model includes target sensor data and a target cue template. The target cue template guides the inference and generation process of the multimodal diffusion model. After the target sensor data is input, a visual encoder extracts high-dimensional visual features, which represent the perceived information such as the target location and road condition in the driving environment. After the target cue template is input, a text encoder extracts text features, which represent the inference framework, filling constraints, and task requirements of the template. The visual and text features are input to a connector, and after aligning the temporal and spatial dimensions, they are spliced to generate a fusion feature. This fusion feature simultaneously carries environmental perception information and template semantic constraints. The fusion feature is input to the diffusion backbone network, which, based on a progressive denoising generation paradigm, fills the content to be filled in the target cue template according to the fusion feature to obtain the target cue information. It also generates multiple control sub-information based on the target cue information in a chronological order at multiple time points, outputting the target driving control information.
[0105] Optional, such as Figure 3 As shown, the input of the multimodal diffusion model includes target sensor data and target cue template. After the target sensor data is input, the visual encoder extracts visual features; after the target cue template is input, the text encoder extracts text features; the visual features and text features are input into the connector, aligned and spliced to generate fused features; the fused features and the target cue template can be input together into the diffusion backbone network, and the diffusion backbone network outputs target cue information.
[0106] It should be noted that the fused features contain environmental perception information from the target sensor data and semantic constraint information from the target cue template. Whether the fused features are input alone into the diffusion backbone network, or both the fused features and the target cue template are input together, target cue information and target driving control information can be generated. Inputting the target cue template into the diffusion backbone network strengthens the transmission of the template inference framework and filling constraints, reduces semantic bias or logical disconnects that may occur when relying solely on fused features for decoding, and improves the accuracy of the output content. This application does not impose limitations on this aspect.
[0107] Optionally, the connector can be implemented using a multilayer perceptron (MLP), which can map visual features into an embedding sequence compatible with text feature tokens. Then, the text feature tokens and the visually projected embedding sequence are concatenated in a preset order to form a unified multimodal fusion feature.
[0108] As an optional implementation, the target prompt information includes a first reasoning step and a second reasoning step, generating multiple control sub-information based on the target prompt information in a chronological order of multiple moments, including: The first inference step is performed based on the diffusion backbone network, target cue information, and fusion features to generate a time step sequence; wherein, the time step sequence includes multiple mask markers corresponding to multiple time points; The second inference step is performed based on the diffusion backbone network, target cue information, time step sequence and fusion features. The control sub-information corresponding to each mask mark is obtained in the order of multiple time steps to generate multiple control sub-information.
[0109] It should be noted that the inference step is used to break down the generation process of target driving control information into stages, making it conform to the progressive thinking logic of human driving decision-making, thereby improving the rigor, coherence, and reliability of the generation of control sub-information. The mask markers are temporal placeholders under the diffusion generation paradigm, used to represent the decoding positions of control sub-information at each moment; the model can fill the mask markers through stepwise denoising or parallel prediction to obtain specific driving control parameters. The number of mask markers is consistent with the preset number of moments in the target driving task, and the distribution of the mask markers can be limited by causal constraints to ensure the logical coherence of temporal generation.
[0110] Optionally, the reasoning steps can be obtained by filling in the action suggestions to be filled in the target prompt template, or by a fixed execution flow predefined in the target prompt template. This application does not limit this.
[0111] For example, the first inference step can be "the diffusion backbone network generates a time step sequence containing three moments based on the sequence format constraints of the fusion features and target prompt information, and represents the temporal position in the form of mask markers"; the second inference step can be "the diffusion backbone network decodes and fills each mask marker according to the order of moments 1-3 determined by the time step sequence, and obtains continuous target driving trajectory information following the temporal order, and the filling result at each moment is the corresponding control sub-information".
[0112] The diffusion backbone network performs a first inference step based on target cues and fused features. This involves jointly encoding the fused features and target cues, initializing them with diffusion denoising, and combining this with the temporal causal logic of the driving task to complete the temporal arrangement of mask markers, generating a time-step sequence with mask markers. The second inference step, based on the target cues, time-step sequence, and fused features, uses the time-step sequence as a fixed temporal framework. Combining the inference constraints of the environmental perception information and target cues carried by the fused features, it decodes the mask markers bit-by-bit in chronological order through a progressively denoised diffusion generation method, obtaining the control sub-information corresponding to each mask marker.
[0113] Optionally, to ensure that the lower-level information (such as driving trajectory information and driving control information) can perceive the higher-level information (such as the logical intent and decision rules contained in the target prompt information), but the higher-level information is not affected by the noise of the lower-level information, a hierarchical masking mechanism can be introduced during the process of the diffusion backbone network executing the above inference steps. By constructing a causal attention mask matrix to limit the attention perception range, the decoding process of the lower-level mask tags can perceive the logical intent, temporal constraints, action planning and other information of the higher-level information, while the inference process of the higher-level information is completely unaffected by the noise of the lower-level mask tags. This ensures the stability of the higher-level inference logic and makes the generation of lower-level control parameters conform to the decision requirements of the higher level.
[0114] Optionally, to further ensure the temporal causal consistency during the inference process, a re-masking strategy can be adopted during the inference stage where the diffusion backbone network progressively denoises the masking labels of the time step sequence. This involves sampling and denoising only the masking labels at the current inference temporal position, and re-masking the masking labels that have not reached the preset generation time sequence. This forces the maintenance of the Markov chain property and mathematically prevents information from future time sequences from participating in the inference process at the current moment. This ensures that the generation of the time step sequence and the time-by-time decoding of the masking labels strictly follow the temporal causal logic of the driving task, and guarantees the coherence and rationality of the generation of multiple control sub-information in the order of time.
[0115] For example, the mathematical expression of the remasking strategy is: , in, Let represent the mask marker corresponding to the t-th temporal position when the diffusion time is i, and M represent the mask identifier. This means that a mask is re-added to the mask marker at this position, keeping it in a mask state to be decoded, and preventing it from participating in the inference and generation process of the current level. t represents the marker for all time positions t, i.e., covering all times in the time step sequence; (t): represents the generation level to which the t-th temporal position belongs, and l represents the target level being generated in the current inference step.
[0116] This can be understood as follows: during the process of the diffusion backbone network performing inference steps and generating the target level l, for all temporal positions, the level to which the target level is higher than the current target level ( For all masked tags (t)>l), their state at diffusion time i is forcibly set to the masked state until the preset generation step corresponding to the higher level is entered. This is to prevent higher-level tags that have not reached the generation time from being decoded in advance and participating in the current level reasoning, thus ensuring the causal consistency and temporal rigor of the reasoning process.
[0117] As an optional implementation, the target prompt information may include a target inference step, which generates multiple control sub-information based on the target prompt information in a sequential order of multiple time points, including: performing the target inference step based on the diffusion backbone network, the target prompt information, and the fusion features to generate control sub-information corresponding to multiple time points.
[0118] It should be noted that the purpose of the target inference step is to break the step-by-step generation mode of the time dimension and the driving control parameter dimension. By generating control sub-information with timestamps through a multi-dimensional diffusion tensor, the multimodal diffusion model can simultaneously complete the temporal localization and corresponding driving control parameter generation at each moment in a single inference process. This effectively avoids the hierarchical transmission loss problems such as information loss and parameter and time misalignment that are prone to occur in step-by-step inference. The multi-dimensional diffusion tensor refers to a fused tensor constructed by integrating the time dimension and the driving control parameter dimension. It integrates the time dimension representing the driving time sequence and the control parameter dimension representing the vehicle driving state, providing a data carrier for the model to simultaneously complete the localization at each moment and the generation of corresponding control parameters.
[0119] For example, the target inference step included in the target prompt information is "based on fused features and driving decision constraints, generate continuous control sub-information with timestamps within the next 3 seconds". The diffusion backbone network generates multiple timestamped control sub-information at various times through a progressively denoised diffusion generation method, which can be "t=1s: x-coordinate 10.2m, y-coordinate 2.5m; t=2s: x-coordinate 20.4m, y-coordinate 2.5m; t=3s: x-coordinate 30.6m, y-coordinate 1.5m". Each timestamped driving trajectory is the control sub-information at the corresponding time. By completing the collaborative generation of timing and parameters in a single inference step, the hierarchical transmission loss caused by step-by-step inference is reduced, and the inference efficiency of control sub-information is improved.
[0120] As an optional implementation, the control sub-information includes at least one control parameter, each control parameter consisting of at least two parameter parts with different priorities. The diffusion backbone network generates the control sub-information in the following manner: The diffusion backbone network generates parameter parts of each priority according to the target prompt information, time step sequence and fusion features, in a preset priority order; among them, the parameter parts of all control parameters of the same priority are generated in parallel.
[0121] It should be noted that the preset priority order can be set based on the parameter logic hierarchy of autonomous driving control and the actual needs of driving decisions, following the principle of first determining the range and then fine-tuning, to ensure that the generation process of control parameters conforms to the actual decision-making logic of vehicle driving control. Among them, the control sub-information includes control parameters that may include, but are not limited to, trajectory parameters (such as X-axis coordinates and Y-axis coordinates) and command parameters (such as acceleration parameters and steering angle parameters).
[0122] For example, when the target driving control information is the target driving trajectory information, multiple control sub-information sets are sets of driving trajectory parameters corresponding to the mobile device at multiple times. The control sub-information can include two control parameters: X-axis coordinates and Y-axis coordinates. These coordinates can be a vehicle coordinate system constructed with the mobile device itself as the origin or a global geographic coordinate system in an electronic map. Each control parameter consists of at least two parameter parts with different priorities. For example, the X-axis coordinate can be divided into two parameter parts with different priorities: an integer X-axis coordinate and a decimal X-axis coordinate, with the integer X-axis coordinate having higher priority than the decimal X-axis coordinate. Similarly, the Y-axis coordinate can be divided into two parameter parts with different priorities: an integer Y-axis coordinate and a decimal Y-axis coordinate, with the integer Y-axis coordinate having higher priority than the decimal Y-axis coordinate. The diffusion backbone network can generate the two parameter parts with the same priority (integer X-axis coordinates and integer Y-axis coordinates) in parallel, completing the macroscopic spatial positioning of the vehicle at that moment. Then, based on the already generated integer coordinate part, it generates the parameter parts with the same priority (decimal X-axis coordinates and decimal Y-axis coordinates) in parallel, achieving fine-tuning of the vehicle's spatial position.
[0123] Optionally, the control parameters may have at least two different priority parameter parts, which may be divided into an integer part and a decimal part, or into multiple parts according to the number of digits, and the priority of each part may be set sequentially from high to low according to the numerical magnitude of the parameter part.
[0124] For example, when the target driving control information is target driving control command information, the control sub-information refers to the set of driving control command parameters corresponding to the mobile device at a certain moment. This may include control parameters such as steering angle parameters, acceleration parameters, and vehicle speed adjustment parameters. Each control parameter can be divided into at least two parameter parts of different priorities according to the numerical accuracy requirements. For example, the steering angle parameter can be divided into three parameter parts of different priorities: tens digit, units digit, and decimal digit, in the order of tens digit, units digit, and decimal digit. The diffusion backbone network will first generate the highest priority parameter part of all control parameters in parallel, such as simultaneously generating the tens digit part of the steering angle and the tens digit part of the acceleration. Then, according to the order of priority from high to low, it will sequentially generate the parameter parts of each priority level in parallel until the generation of all control parameter parts is completed.
[0125] Please see Figure 4 , Figure 4 This is a schematic diagram illustrating the generation of multiple control sub-information for the multimodal diffusion model disclosed in the embodiments of this application. For example... Figure 4As shown, the target driving control information is the target driving trajectory information. Multiple control sub-information corresponds to the set of driving trajectory parameters of the mobile device at multiple consecutive time points such as time 1, time 2, and time 3. Each control sub-information consists of two types of core control parameters: X-axis coordinate and Y-axis coordinate, which are used to accurately characterize the spatial position of the mobile device at the corresponding time.
[0126] It should be noted that for control sub-information at multiple moments, the diffusion backbone network can generate it moment-by-moment using a remasking strategy, following temporal causal logic. For example, it first generates the complete control sub-information of the previous moment (e.g., moment 1), and then, based on the generated preceding moment information, it progressively derives and generates the control sub-information of the next moment (e.g., moment 2, moment 3). The remasking mechanism forcibly shields information interference from future moments, ensuring the causal consistency and logical coherence of the temporal generation. For control sub-information at a specific moment, each parameter part is generated layer by layer according to a preset priority order. High-priority parameter parts (e.g., X-axis integer coordinates, Y-axis integer coordinates) are generated first to define the macroscopic spatial range of the mobile device at that moment; then low-priority parameter parts (e.g., X-axis decimal coordinates, Y-axis decimal coordinates) are generated to fine-tune the determined integer coordinates and ensure driving accuracy.
[0127] It is understandable that when the control sub-information contains multiple control parameters (such as...) Figure 4 When generating X-axis and Y-axis coordinates (as shown), the same priority parameter portion of different control parameters can be generated in parallel using a diffusion backbone network, synchronously completing the generation operation of that priority portion of all control parameters. For example, when generating control sub-information at time 1, the model will first generate the integer X-axis coordinates and integer Y-axis coordinates in parallel, and then generate the decimal X-axis coordinates and decimal Y-axis coordinates in parallel. This ensures the hierarchical order of parameter generation, avoids logical confusion caused by mixed generation of multiple parameters, and improves the overall generation efficiency of control sub-information, reducing decision inference latency.
[0128] For example, such as Figure 4 As shown, the control sub-information at time 2 starts from the initial state (+0?.??,+0?.??). First, the integer part is filled to obtain (+00.??,+05.??), then the decimal part is filled to finally generate the complete control sub-information (+00.82,+05.34). Among them, the character "?" represents the mask mark to be decoded, indicating that the parameter part at this position has not yet been generated and needs to be gradually filled with noise reduction.
[0129] Please see Figure 5 , Figure 5 This is a schematic diagram illustrating the inference process of a pre-trained multimodal diffusion model disclosed in an embodiment of this application. Figure 5As shown, the pre-trained multimodal diffusion model includes a text encoder, a visual encoder, and a diffusion backbone network, which are used to generate control sub-information under multimodal inputs. Its inference process follows a causal hierarchy.
[0130] In the multimodal input encoding stage, the visual encoder receives target sensor data, and the text encoder receives target prompt templates generated based on target driving task and target driving state data. The two encoders encode the input information of visual modality and text modality respectively, and input the fused multimodal features into the diffusion backbone network to provide a unified semantic basis for the generation of subsequent control sub-information.
[0131] During the inference phase of the diffusion backbone network, within the causal hierarchy, the target prompt template is semantically filled to obtain target prompt information, which serves as the highest-level task constraint, providing macro-level guidance for subsequent parameter generation. Integer X-axis and integer Y-axis coordinates are generated sequentially, followed by decimal X-axis and decimal Y-axis coordinates, achieving progressive generation of parameter parts with different priorities. This gradually transforms high-level semantics into low-level executable control parameters. Simultaneously, a hierarchical masking mechanism ensures that upper-level information guides lower-level parameter generation, while noise in lower-level parameters does not negatively interfere with high-level semantic inference.
[0132] Within the causal hierarchy, parallel generation follows the paradigm of the diffusion language model within the same level. For example, parameters of the same priority (such as integer coordinates on the X-axis and integer coordinates on the Y-axis) are generated in parallel, and the decoding of the corresponding priority part of all control parameters is completed synchronously, improving generation efficiency. In the case of multiple control sub-information, the temporal causal order is strictly followed in the temporal dimension, and the interference of parameter information in future moments is forcibly shielded through a reusable masking strategy to ensure the reliability of the generated multiple control sub-information.
[0133] Through this causal hierarchical constraint reasoning architecture, the model not only ensures the logical rigor and temporal coherence of the generation of control sub-information, but also improves reasoning efficiency through the parallel generation of parameters with the same priority, which can better adapt to the real-time driving control requirements in autonomous driving scenarios.
[0134] In this embodiment, the specific architecture and number of modal branches of the multimodal diffusion model are not limited. The multimodal diffusion model is pre-trained using a supervised training paradigm. During training, a training dataset containing sample sensor data, sample prompt templates, and sample driving control information can be constructed. The sample sensor data and sample prompt templates are input into the initial multimodal diffusion model, which performs multimodal feature parsing, template inference and filling, and outputs predicted driving control information. Subsequently, the loss value between the predicted driving control information and the sample driving control information is calculated. Based on this loss value, the model parameters are iteratively optimized using a backpropagation algorithm until the loss value meets the preset convergence condition, thus obtaining the pre-trained multimodal diffusion model.
[0135] Optionally, temporal consistency constraints and multimodal alignment constraints can be introduced during training to avoid conflicts and gaps in the control sub-information output by the model at different times, while improving the model's generalization and adaptation capabilities to complex driving scenarios and ensuring the stability and reliability of the target driving control information.
[0136] S104 controls the movement of the mobile device based on multiple control sub-information.
[0137] In this embodiment of the application, after obtaining the target driving control information of the mobile device, the mobile device can be controlled to drive based on the multiple control sub-information included in the target driving control information, so as to ensure that the mobile device can follow the target driving task requirements, complete the driving operation in the driving environment, and achieve precise control of autonomous driving.
[0138] Optionally, when the target driving control information is a target driving control command, multiple control sub-information can be directly sent to the driving actuators of the mobile device (such as the power system, steering system, braking system, etc.). The driving actuators can then directly execute the corresponding driving operations according to the commands, such as adjusting the driving speed of the mobile device according to the driving speed command, and adjusting the driving direction of the mobile device according to the steering angle command, thereby realizing the driving control of the mobile device.
[0139] Optionally, when the target driving control information is the target driving trajectory information, multiple control sub-information can be converted into driving control commands that can be recognized by the driving actuators of the mobile device. The converted driving control commands are then sent to each actuator to guide the mobile device to drive gradually along the preset driving trajectory, ensuring that the driving path of the mobile device is consistent with the target driving trajectory information, thereby realizing the driving control of the mobile device.
[0140] It should be noted that the type of target driving control information can be flexibly set according to the driving scenario of the mobile device, and this application does not limit it here.
[0141] By implementing the above technical solution, compared with the significant defects of related technologies, such as the high inference latency and lack of global consistency of the autoregressive model-based paradigm, and the standard diffusion language model-based paradigm ignoring hierarchical causal structure, lacking spatiotemporal constraints, and having an uncontrollable generation process, the driving control method provided in this application provides a comprehensive environmental perception and logical constraint basis for driving decisions through the fusion encoding of multimodal sensor data and target prompt templates, avoiding future information leakage and the generation of physically unreasonable trajectories; through the parameter generation logic of first determining the range, then fine-tuning, and parallel generation of parameters with the same priority, the inference latency is reduced while ensuring the global consistency of control parameters, thereby improving the timeliness and reliability of driving control.
[0142] As an optional implementation, the driving control method provided in this application can obtain a multimodal diffusion model by training an initial model using a training dataset. Please refer to... Figure 6 , Figure 6 This is a flowchart illustrating a driving control method disclosed in an embodiment of this application. Figure 2 The details are explained below.
[0143] S601, in the first training phase, trains the initial model based on reference sensor data, reference prompt template, and reference prompt information.
[0144] As an optional implementation, the multimodal diffusion model is obtained by training an initial model based on a training dataset; the training dataset includes reference sensor data, reference prompt templates, reference prompt information, and reference driving control information; the training phase of the multimodal diffusion model includes a first training phase and a second training phase. Before inputting the target sensor data and target prompt templates into the pre-trained multimodal diffusion model to obtain the target driving control information output by the multimodal diffusion model, the method further includes: In the first training phase, the initial model is trained based on reference sensor data, reference prompt templates, and reference prompt information. In the second training phase, the initial model is trained based on reference sensor data, reference prompt templates, reference prompt information, and reference driving control information.
[0145] It should be noted that the training dataset is used for supervised training, parameter optimization, and convergence iteration of the initial model. By providing the model with sample inputs and expected outputs, it improves the model's ability to understand the driving environment, target prompts, and the accuracy and robustness of driving control information inference.
[0146] The reference sensor data refers to the data obtained by the mobile device in real-world driving scenarios, perceiving the surrounding driving environment. It serves as sample input in the training dataset. Its composition, acquisition format, and data type can refer to the aforementioned example of target sensor data to ensure that the training data maintains the same origin as the input data used in actual model inference, thereby improving the effectiveness and generalization ability of model training. The reference sensor data may include image data collected by multiple cameras mounted on the mobile device. Furthermore, the reference sensor data may also include one or more of the following: reference radar point cloud data, reference millimeter-wave radar data, and reference ultrasonic radar data, used to provide realistic environmental perception input for the initial model.
[0147] The reference prompt template is a structured reasoning framework built upon reference driving task and reference driving state data. It contains content to be filled in, maintaining consistency with the structure, type, and generation logic of the aforementioned target prompt template. It serves as an input template sample in the training dataset, standardizing the initial model's reasoning path and filling constraints, enabling the model to learn the mapping rules from reference sensor data to reference prompt information. The reference prompt template can take the form of a thought chain text template or a directed acyclic graph structure template, etc. Its content to be filled can include environmental perception content, decision rationale content, action suggestion content, etc., maintaining format consistency with the target prompt template used in actual reasoning to ensure consistency between the model training scenario and the reasoning scenario.
[0148] The reference prompt information is a structured reasoning result obtained by contextually adapting and filling the content to be filled in the reference prompt template. It is consistent with the composition, semantic logic and function of the aforementioned target prompt information. It is used as an intermediate label sample in the training dataset to guide the initial model to learn the ability to complete the template filling based on the reference sensor data. This enables the predicted prompt information output by the model to fit the environmental perception results and driving decision logic of the real driving scenario, providing a reliable reasoning basis for the generation of subsequent driving control information. It is consistent with the function and format of the target prompt information during actual reasoning.
[0149] Reference driving control information refers to the standard driving control output of the mobile device in a real driving scenario. It maintains consistency with the aforementioned target driving control information in terms of type, composition, and temporal logic. It serves as the final label sample in the training dataset, guiding the initial model to learn the ability to generate temporally continuous and causally rigorous control sub-information based on reference prompts. Reference driving control information includes multiple reference control sub-information corresponding to multiple time points. Each reference control sub-information corresponds to driving control parameters (such as trajectory coordinates, steering angle, vehicle speed, etc.) at a specific time point. Its generation logic is consistent with the target driving control information during actual inference, ensuring that the model learns the temporally sequential control output capability that meets the requirements of the driving task.
[0150] By adopting a phased training approach, the first phase focuses on the model's template filling and semantic reasoning capabilities, enabling the model to learn to understand the driving environment and generate compliant reasoning prompts. The second phase focuses on the model's temporal control generation capabilities, enabling the model to generate accurate control sub-information based on existing reasoning prompts. This approach can effectively reduce task confusion and gradient optimization difficulty in end-to-end training, improve the model's convergence speed and generalization ability, and ensure that the pre-trained multimodal diffusion model has stable performance in actual reasoning scenarios.
[0151] As an optional implementation, in the first training phase, the initial model is trained based on reference sensor data, reference cue templates, and reference cue information, including: In the first training phase, reference sensor data and reference prompt templates are input into the initial model to obtain the predicted prompt information output by the initial model; the parameters of the initial model are updated based on the difference between the predicted prompt information and the reference prompt information.
[0152] It should be noted that the difference between the predicted prompts and the reference prompts refers to the deviation between the two in the structured content (such as environmental perception, decision rationale, action suggestions, etc.). This deviation can be quantified by a preset first loss function to obtain the target loss, so as to guide the iterative update of the model parameters based on this difference, thereby improving the template filling accuracy and logical rigor of the model.
[0153] Optionally, the first loss function can be the cross-entropy loss function (CEL) or the Euclidean distance loss function (EDL). The target loss determined by the above loss function is a comprehensive quantification of the deviation between the predicted prompt information and the reference prompt information in terms of semantic logic and numerical accuracy, so as to guide the parameter iteration update of the initial model based on the target loss.
[0154] It should be noted that Cross-Entropy Loss (CEL) is a loss calculation method used to quantify the difference between two probability distributions. Its core function is to measure the degree of matching between predicted and true semantics. In this application, it is mainly adapted to quantify the deviation of semantic slots such as environmental perception and decision rationale in the reference prompt template. Euclidean Distance Loss (EDL) is a loss calculation method used to quantify the spatial distance between two consecutive values. Its core function is to measure the absolute deviation between predicted and true values.
[0155] In actual training, a single loss function or a combination of loss functions can be flexibly selected according to the type of slot to be filled in the reference prompt template: for purely semantic content, the cross-entropy loss function (CEL) is used alone; for purely numerical content, the Euclidean distance loss function (EDL) is used alone; for mixed slots that contain both semantic description and numerical information (such as "there is a construction section 50m ahead"), the two loss functions can be weighted and fused by preset weights to obtain the final target loss to fully cover the bias dimension.
[0156] The target loss is obtained by quantizing using the first loss function mentioned above. This allows for targeted adaptation to different slot types in template filling, enabling accurate representation of semantic logic deviation and numerical precision deviation. This target loss guides the iterative update of the initial model's parameters, gradually bringing the model's output prediction prompts closer to the reference prompts, thereby improving the model's template filling accuracy and scene generalization ability.
[0157] It should be noted that the initial model training process involves adjusting the model parameters to reduce the value of the first loss function below a preset threshold, or to make the target loss value meet a preset target requirement, or to satisfy other convergence conditions. The parameters of the initial model may include one or more of the following: structural parameters of the visual encoder, text encoder, connector, and diffusion backbone network (such as the number of model layers, neuron weights, layer normalization parameters, attention mask matrix parameters, etc.); if the initial model is a neural network model, its structural parameters may include at least one of the following: the number of neural network layers, width, neuron weights, activation function parameters (such as the threshold of ReLU, the slope of Sigmoid).
[0158] After determining the target loss, optimizers such as Adam and Stochastic Gradient Descent (SGD) can be used to calculate the gradient of the target loss with respect to each parameter of the model based on the backpropagation algorithm, and update the parameters along the gradient descent direction to complete a single parameter iteration. It should be noted that the Adam optimizer is an adaptive optimization algorithm that combines momentum gradient descent and adaptive learning rate. It dynamically calculates the first moment (mean) and second moment (variance) estimates of each parameter and assigns differentiated learning rates to different parameters. It has the advantages of fast convergence speed and strong robustness to initial learning rate hyperparameters, and is suitable for the rapid optimization needs of high-dimensional and multi-module parameters in the initial model. On the other hand, the Stochastic Gradient Descent (SGD) optimizer calculates the gradient and updates the parameters by randomly selecting sample batches in a single iteration. Although the convergence speed is relatively slower, the optimization process is more stable and less prone to getting trapped in local optima. It is often used as a supplementary strategy to the Adam optimizer for parameter fine-tuning in the later stages of model training, improving the stability of parameter convergence and the model's generalization ability. This application does not restrict the choice of optimizer.
[0159] Optionally, the convergence condition for the initial model in the first training phase may include, but is not limited to, any of the following: the number of parameter iterations reaches a preset threshold; the calculated target loss is less than a preset loss threshold; the target loss no longer decreases within a preset number of consecutive iterations, or the decrease is less than a preset magnitude threshold. It should be noted that by setting convergence conditions, underfitting due to insufficient training iterations can be avoided, as can overfitting due to overtraining, ultimately ensuring that the model possesses stable template filling and semantic inference capabilities after the first training phase. If the initial model after any parameter iteration update satisfies any convergence condition, the iteration process of the first training phase is terminated, and the second training phase begins.
[0160] S602, in the second training phase, the initial model is trained based on reference sensor data, reference prompt templates, reference prompt information, and reference driving control information.
[0161] As an optional implementation, in the second training phase, the initial model is trained based on reference sensor data, reference prompt templates, reference prompt information, and reference driving control information, including: The reference sensor data, reference prompt template, and reference prompt information are input into the initial model to obtain the predicted driving control information output by the initial model; the parameters of the initial model are updated based on the difference between the predicted driving control information and the reference driving control information.
[0162] It should be noted that the difference between the predicted driving control information and the reference driving control information refers to the deviation between the driving trajectory information (such as X-axis coordinates and Y-axis coordinates) or driving control commands (such as steering angle commands, vehicle speed commands, and acceleration commands) corresponding to the two. The target loss can be obtained by quantifying the difference through a preset second loss function, so as to guide the iterative update of the model parameters based on the difference, thereby improving the accuracy, temporal coherence, and physical rationality of the control sub-information generated by the model.
[0163] Optionally, when the predicted driving control information and the reference driving control information are of the type of driving trajectory information, the second loss function can be the SmoothL1 loss function, which calculates the trajectory displacement deviation of the two at each time point to achieve accurate quantification of the trajectory prediction deviation.
[0164] Optionally, when the predicted driving control information and the reference driving control information are driving control command types, the deviation of the control command prediction can be quantified by calculating the deviation between control components such as steering angle, vehicle speed, and acceleration.
[0165] Furthermore, for driving control commands of different dimensions such as steering angle, vehicle speed, and acceleration, since their dynamic range, physical constraints, and impact on driving safety vary, independent loss calculations can be used to quantify the deviations of different control command dimensions. Then, the target loss can be obtained by weighted summation through preset weights. This can avoid the single-dimensional deviation from dominating the training process and enhance the stability and convergence efficiency of model training.
[0166] It should be noted that the purpose of the second training phase is to optimize the model's temporal control generation capability. The model parameters remain consistent with those of the first training phase (including the structural parameters of the visual encoder, text encoder, connector, and diffusion backbone network). The focus is on fine-tuning the parameters related to temporal arrangement and parameter decoding in the diffusion backbone network to ensure that the generated control sub-information conforms to the inference constraints of the reference prompt information.
[0167] After determining the target loss, the Adam and SGD optimizers used in the first training stage are used to calculate the gradient and update the parameters based on the backpropagation algorithm. The Adam optimizer is used for fast convergence in the early stage, and the SGD optimizer is used for fine-tuning the parameters in the later stage, which can meet the needs of efficient optimization of parameters of multiple modules.
[0168] Optionally, the convergence conditions for the second training phase can be consistent with those for the first training phase (e.g., the number of parameter iterations reaches a preset threshold, the target loss is less than a preset threshold, and the loss decreases by a certain amount over multiple consecutive cycles), or can be flexibly adjusted according to training needs to avoid underfitting or overfitting the model. If any convergence condition is met, the iteration process terminates, and the pre-trained multimodal diffusion model is obtained.
[0169] As an optional implementation, the initial model includes a visual encoder, a text encoder, a connector, and an initial diffusion backbone network; updating the parameters of the initial model includes updating the parameters of the initial diffusion backbone network.
[0170] It should be noted that the visual encoder, text encoder, and connector are functionally universal and stable: the core function of the visual encoder is to extract visual features such as environmental targets and road conditions from sensor data (such as images and radar data); the text encoder is used to parse the textual semantic constraints of the reference prompt template; and the connector is responsible for spatiotemporal alignment and splicing of the two types of features to generate fused features. As long as the data types and structures of the training data and the actual inference data remain homogeneous, they can be directly adapted without adjustment, thereby improving the model training efficiency.
[0171] By implementing the above technical solution, a phased training mode of "first filling semantic reasoning with templates, then generating temporal control" is adopted. Combined with a targeted loss function design, this not only reduces the difficulty of task confusion and gradient optimization in end-to-end training, but also ensures that the model meets the adaptation requirements in terms of semantic logic, numerical accuracy, and temporal coherence. At the same time, with the help of a mature optimizer and convergence condition settings, the convergence speed and generalization ability of the model are guaranteed. Ultimately, the pre-trained multimodal diffusion model can generate driving control information that conforms to causal logic in actual reasoning, adapting to the real-time and safety requirements of autonomous driving.
[0172] It should be understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps. In addition, the above embodiments can be implemented independently or in combination with each other, and different implementation methods in the above embodiments can also be combined with each other, without limitation.
[0173] Based on the foregoing embodiments, this application provides a driving control device, which includes various modules and units included in each module, and can be implemented by a processor; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP) or field programmable gate array (FPGA), etc.
[0174] Please see Figure 7 , Figure 7 This is a structural block diagram of a driving control device disclosed in an embodiment of this application, such as... Figure 7 The driving control device shown includes a data acquisition module 710, a template generation module 720, a driving planning model 730, and a driving control module 740.
[0175] The data acquisition module 710 is used to acquire the target driving task, target driving status data and target sensor data of the mobile device; wherein, the target sensor data includes multiple images of the same location, each image being acquired from a different angle; The template generation module 720 is used to generate a target prompt template based on the target driving task and target driving status data; wherein, the target prompt template includes content to be filled in; The driving planning model 730 is used to input target sensor data and target prompt template into a pre-trained multimodal diffusion model to obtain target driving control information output by the multimodal diffusion model. The target driving control information includes multiple control sub-information corresponding to multiple time points. The multimodal diffusion model is used to fill the content to be filled in the target prompt template according to the target sensor data to obtain target prompt information, and to generate multiple control sub-information according to the target prompt information in the order of multiple time points. The driving control module 740 is used to control the driving of the mobile device based on multiple control sub-information.
[0176] In some embodiments, the driving planning model 730 is further configured to extract visual features and text features from the target sensor data and the target prompt template, respectively; align and splice the visual features and text features to generate fused features; and fill the content to be filled in the target prompt template according to the fused features to obtain target prompt information.
[0177] In some embodiments, the multimodal diffusion model includes a visual encoder, a text encoder, a connector, and a diffusion backbone network; the visual encoder is used to extract visual features from target sensor data; the text encoder is used to extract text features from a target prompt template; the connector is used to align and splice the visual features and text features to generate fused features; the diffusion backbone network is used to fill the content to be filled in the target prompt template according to the fused features to obtain target prompt information, and to generate multiple control sub-information according to the target prompt information in a sequential order of multiple time points.
[0178] In some embodiments, the target prompting information includes a first inference step and a second inference step. The driving control module 740 is further configured to perform the first inference step based on the diffusion backbone network, the target prompting information, and the fusion features to generate a time step sequence; wherein the time step sequence includes multiple mask markers corresponding to multiple time points; and to perform the second inference step based on the diffusion backbone network, the target prompting information, the time step sequence, and the fusion features to obtain control sub-information corresponding to each mask marker in the order of the multiple time points, so as to generate multiple control sub-information.
[0179] In some embodiments, the control sub-information includes at least one control parameter, each control parameter consisting of at least two parameter parts with different priorities. The diffusion backbone network generates the control sub-information in the following manner: the diffusion backbone network generates parameter parts of each priority according to a preset priority order based on the target prompt information, time step sequence, and fusion features; wherein, parameter parts of the same priority for all control parameters are generated in parallel.
[0180] In some embodiments, the content to be filled in the target prompt template includes environmental perception content, decision rationale content, and action suggestion content; the environmental perception content, decision rationale content, and action suggestion content are filled in parallel based on target sensor data; the environmental perception content is used to characterize the environmental perception results of the environment in which the mobile device is located; the decision rationale content is used to characterize the reasoning basis for determining the target driving control information based on the environmental perception results; and the action suggestion content is used to indicate the target driving control information generated based on the reasoning basis.
[0181] In some embodiments, the multimodal diffusion model is obtained by training an initial model based on a training dataset; the training dataset includes reference sensor data, reference prompt templates, reference prompt information, and reference driving control information; the training phase of the multimodal diffusion model includes a first training phase and a second training phase, and the driving planning model 730 is further used to train the initial model based on the reference sensor data, reference prompt templates, and reference prompt information in the first training phase; and to train the initial model based on the reference sensor data, reference prompt templates, reference prompt information, and reference driving control information in the second training phase.
[0182] In some embodiments, the driving planning model 730 is further configured to, during the first training phase, input reference sensor data and a reference prompt template into an initial model to obtain predicted prompt information output by the initial model; update the parameters of the initial model based on the difference between the predicted prompt information and the reference prompt information; and input reference sensor data, a reference prompt template, and reference prompt information into the initial model to obtain predicted driving control information output by the initial model; update the parameters of the initial model based on the difference between the predicted driving control information and the reference driving control information.
[0183] It should be noted that, in the embodiments of this application... Figure 7 The module division of the driving control device shown is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0184] Please refer to Figure 8 , Figure 8 This is a structural block diagram of a mobile device disclosed in an embodiment of this application. Figure 8 As shown, the mobile device includes a memory 810 and a processor 820, with the memory 810 storing executable program code.
[0185] The memory 810 is coupled to the processor 820; The processor 820 calls the executable program code stored in the memory 810 to execute any one of the mobile device control methods in the above method embodiments.
[0186] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in any of the mobile device control methods provided in the above embodiments.
[0187] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps in any of the mobile device control methods provided in the above embodiments.
[0188] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the mobile device to which the present application is applied. A specific mobile device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0189] It should be noted that the descriptions of the above embodiments of the apparatus, computer-readable storage medium, and computer program product are similar to the descriptions of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the embodiments of the apparatus, computer-readable storage medium, and computer program product of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0190] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be repeated here.
[0191] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.
[0192] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0193] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules above is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, and can be electrical, mechanical, or other forms.
[0194] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0195] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0196] The features disclosed in the several device embodiments provided in this application can be arbitrarily combined without conflict to obtain new device embodiments.
[0197] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A driving control method, characterized in that, include: The system acquires the target driving task, target driving status data, and target sensor data of the mobile device; wherein the target sensor data includes multiple images of the same location, each image being acquired from a different angle. A target prompt template is generated based on the target driving task and the target driving status data; wherein, the target prompt template includes content to be filled in; The target sensor data and the target prompt template are input into a pre-trained multimodal diffusion model to obtain target driving control information output by the multimodal diffusion model; wherein, the target driving control information includes multiple control sub-information corresponding to multiple time points; the multimodal diffusion model is used to fill the content to be filled in the target prompt template according to the target sensor data to obtain target prompt information, and to generate the multiple control sub-information according to the target prompt information in the chronological order of the multiple time points; The mobile device is controlled to move according to the multiple control sub-information.
2. The method according to claim 1, characterized in that, The step of filling the target prompt template with the content to be filled based on the target sensor data to obtain the target prompt information includes: Visual features and text features are extracted from the target sensor data and the target prompt template, respectively; The visual features and text features are aligned and concatenated to generate a fused feature; The target prompt information is obtained by filling the content to be filled in the target prompt template according to the fusion features.
3. The method according to claim 2, characterized in that, The multimodal diffusion model includes a visual encoder, a text encoder, a connector, and a diffusion backbone network; the visual encoder is used to extract the visual features from the target sensor data; The text encoder is used to extract the text features from the target prompt template; the connector is used to align and splice the visual features and the text features to generate the fused features; The diffusion backbone network is used to fill the content to be filled in the target prompt template according to the fusion features to obtain the target prompt information, and to generate the multiple control sub-information according to the target prompt information in the chronological order of the multiple time points.
4. The method according to claim 3, characterized in that, The target prompt information includes a first reasoning step and a second reasoning step. Generating the multiple control sub-information based on the target prompt information in the chronological order of the multiple time points includes: The first inference step is executed based on the diffusion backbone network, the target prompt information, and the fusion features to generate a time step sequence; wherein, the time step sequence includes multiple mask markers corresponding to the multiple time points; The second inference step is executed based on the diffusion backbone network, the target prompt information, the time step sequence, and the fusion features. The control sub-information corresponding to each mask mark is obtained in the order of the multiple time points to generate the multiple control sub-information.
5. The method according to claim 4, characterized in that, The control sub-information includes at least one control parameter, each control parameter consisting of at least two parameter parts with different priorities. The diffusion backbone network generates the control sub-information in the following manner: The diffusion backbone network generates parameter parts of each priority according to the target prompt information, the time step sequence and the fusion features, in a preset priority order; wherein, parameter parts of all control parameters of the same priority are generated in parallel.
6. The method according to any one of claims 1-5, characterized in that, The content to be filled in the target prompt template includes environmental perception content, decision rationale content, and action suggestion content; the environmental perception content, decision rationale content, and action suggestion content are filled in parallel based on the target sensor data; the environmental perception content is used to characterize the environmental perception results of the environment in which the mobile device is located; The decision rationale is used to characterize the reasoning basis for determining the target driving control information based on the environmental perception results; the action suggestion is used to indicate the target driving control information generated based on the reasoning basis.
7. The method according to any one of claims 1-5, characterized in that, The multimodal diffusion model is obtained by training an initial model based on a training dataset; the training dataset includes reference sensor data, reference prompt templates, reference prompt information, and reference driving control information. The training phase of the multimodal diffusion model includes a first training phase and a second training phase. Before inputting the target sensor data and the target cue template into the pre-trained multimodal diffusion model to obtain the target driving control information output by the multimodal diffusion model, the method further includes: In the first training phase, the initial model is trained based on the reference sensor data, the reference prompt template, and the reference prompt information; In the second training phase, the initial model is trained based on the reference sensor data, the reference prompt template, the reference prompt information, and the reference driving control information.
8. The method according to claim 7, characterized in that, The first training phase, which trains the initial model based on the reference sensor data, the reference cue template, and the reference cue information, includes: In the first training phase, the reference sensor data and the reference prompt template are input into the initial model to obtain the predicted prompt information output by the initial model; the parameters of the initial model are updated according to the difference between the predicted prompt information and the reference prompt information. In the second training phase, the initial model is trained based on the reference sensor data, the reference prompt template, the reference prompt information, and the reference driving control information, including: The reference sensor data, the reference prompt template, and the reference prompt information are input into the initial model to obtain the predicted driving control information output by the initial model; the parameters of the initial model are updated according to the difference between the predicted driving control information and the reference driving control information.
9. The method according to claim 8, characterized in that, The initial model includes a visual encoder, a text encoder, a connector, and an initial diffusion backbone network; updating the parameters of the initial model includes: Update the parameters of the initial diffusion backbone network.
10. A mobile device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 9.