Driving regulation control method, device and equipment and computer readable storage medium
By integrating multimodal data and world models to predict future states, and combining them with a pre-set driving control model to determine the optimal control strategy, the problem of existing technologies being unable to identify long-tail complex scenarios has been solved, thereby improving the accuracy and safety of driving control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG ZEEKR INTELLIGENT TECH CO LTD
- Filing Date
- 2026-03-02
- Publication Date
- 2026-04-24
AI Technical Summary
Existing driving control systems rely on visual language models or large single-stage/segmented models, which cannot accurately identify long-tailed complex scenarios, resulting in an inability to perform precise driving control.
By integrating multimodal data (such as images, sound, and sensor signals) to simulate environmental dynamics, using a world model to predict future states, and combining it with a pre-set driving control model to determine the optimal control strategy, the system avoids direct reliance on driver experience.
It enables accurate identification and driving control in long-tail complex scenarios, improving the accuracy and safety of driving control and reducing data acquisition costs.
Smart Images

Figure CN121912980A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of driver assistance technology, and in particular to a driving control method, device, equipment and computer-readable storage medium. Background Technology
[0002] Assisted driving uses sensors, cameras, and / or radar to acquire data about the vehicle itself or the environment around the vehicle, and then uses this data to control the driver's driving.
[0003] Currently, driving control systems typically rely on visual language models or large, segmented / partial models to transform and extract information about the objective world environment. This allows them to obtain environmental information about the vehicle's movement and then determine control rules based on that information and the driver's experience. However, using visual language models or large, segmented / partial models fails to enable machines to better understand the objective world, making it difficult to accurately identify complex, long-tailed scenarios. Therefore, a new driving control approach is urgently needed.
[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main objective of this application is to provide a driving control method, device, equipment, and computer-readable storage medium, aiming to provide a new driving control approach for driving control.
[0006] To achieve the above objectives, this application provides a driving regulation control method, the driving regulation control method comprising: Acquire vehicle environmental data, and perform multimodal data processing based on the acquired environmental data to obtain the predicted control state; The optimal control strategy is determined based on the predicted control state and the preset driving control model, so as to carry out driving control of the vehicle based on the optimal control strategy.
[0007] In one embodiment, the step of obtaining the predicted control state by performing multimodal data processing based on the collected environmental data includes: Temporal modeling data is obtained by performing time modeling based on the collected environmental data, and spatial modeling data is obtained by performing spatial modeling based on the collected environmental data. Environmental modeling data is obtained by performing environmental modeling based on the collected environmental data, and the predictive control status is determined based on the environmental modeling data, the time modeling data, and the spatial modeling data.
[0008] In one embodiment, the step of determining the predictive control state based on the environmental modeling data, the temporal modeling data, and the spatial modeling data includes: Based on the time modeling data and the spatial modeling data, the proposed target control state is determined in the preset prediction state table, and based on the environmental modeling data, the influencing control state is determined in the preset prediction state table. Based on the influence control state, the proposed target control state is modified to obtain the predicted control state.
[0009] In one embodiment, the driving control method includes: The current state input during the regulation control training is obtained, and the distribution of regulation control actions corresponding to the current state is determined based on the initial driving regulation control model. An experience sample set is determined according to the predetermined probability parameters and the distribution of regulation control actions. Based on the initial driving control model, the training state value corresponding to the experience sample in the experience sample set is determined, and the temporal difference error is determined based on the training state value. The preset driving control model is then trained based on the temporal difference error and the experience sample set.
[0010] In one embodiment, the step of determining the experience sample set based on the predetermined probability parameter and the control action distribution includes: If the random number generated within the preset numerical range is less than the predetermined probability parameter, then a control action is randomly selected from the control action distribution as the target action. If the random number generated within the preset numerical range is greater than or equal to the predetermined probability parameter, then the control action with the highest environmental reward probability in the control action distribution will be taken as the target action. The next predicted state and environmental reward are obtained based on the target action prediction, and the current state, the target action, the next predicted state, and the environmental reward are used as an experience sample set.
[0011] In one embodiment, the step of training a preset driving control model based on the time-series difference error and the experience sample set includes: The network parameters and value parameters in the initial driving control model are updated based on the time-series difference error, and the next prediction state in the empirical sample set is determined. The next predicted state is used as the current state input during the regulation control training, and the step of determining the regulation control action distribution corresponding to the current state based on the updated initial driving regulation control model is executed until a predetermined termination condition is reached, and the updated initial driving regulation control model is used as the preset driving regulation control model.
[0012] In one embodiment, the step of determining the optimal control strategy based on the predicted control state and the preset driving control model includes: Determine the predicted state value and predicted state output in the preset driving control model; The optimal control strategy is determined based on the predicted state value and the predicted state.
[0013] Furthermore, to achieve the above objectives, this application also provides a driving control device, the driving control device comprising: The data acquisition module is used to acquire environmental data of the vehicle and perform multimodal data processing based on the acquired environmental data to obtain the predicted control state. The driving regulation control module is used to determine the optimal regulation control strategy based on the predicted regulation control state and the preset driving regulation control model, so as to perform driving regulation control on the vehicle based on the optimal regulation control strategy.
[0014] In addition, to achieve the above objectives, this application also provides a driving control device, including a processor, a memory, and a driving control method program stored in the memory that can be executed by the processor, wherein when the driving control method program is executed by the processor, it implements the steps of the driving control method as described above.
[0015] This application also provides a computer-readable storage medium storing a driving regulation method program, wherein when the driving regulation method program is executed by a processor, it implements the steps of the driving regulation method as described above.
[0016] This application provides a driving control method that acquires vehicle environmental data, performs multimodal data processing based on the acquired environmental data to obtain a predicted control state, and determines an optimal control strategy based on the predicted control state and a preset driving control model. This method integrates multimodal data (such as images, sounds, and sensor signals from the acquired environmental data) to simulate environmental dynamics and predict future states. It is similar to the process by which humans learn about the world's operating mechanisms through observation, emphasizing state representation and transition models. This avoids the limitations of visual language models or large, segmented / partial models that prevent machines from better understanding the objective world, making it difficult to accurately identify complex, long-tailed scenarios and resulting in control rules that do not fully conform to the actual environment. Furthermore, it determines the optimal control strategy based on the predicted state and a preset driving control model, thus completing the driving control process, rather than directly relying on the driver's driving experience to determine control rules. Therefore, it achieves driving control based on a novel driving control approach. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the first embodiment of the driving regulation control method of this application; Figure 2This is a schematic flowchart of the first embodiment of the driving regulation control method of this application; Figure 3 This is a schematic diagram of a training process for the preset driving control model in the driving control method of this application; Figure 4 This is a schematic diagram of the driving control device module of this application; Figure 5 This is a schematic diagram of the hardware operating environment involved in the device in this application.
[0018] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings.
[0019] Explanation of icon numbers: 1001 Processing device; 1002 Read-only memory; 1003 Storage device; 1004 Random access memory; 1005 Bus; 1006 Input / output interface; 1007 Input device; 1008 Output device; 1009 Communication device; 10 Policy network module; 20 Environment model module; 30 Value network module; 40 Model update module; 50 Model iteration module. Detailed Implementation
[0020] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0021] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0022] Driving control systems typically transform and extract information about the objective world environment based on visual language models or large, segmented / segmented models to obtain environmental information about vehicle operation. These environmental information, combined with the driver's experience, are then used to determine control rules. However, using visual language models or large, segmented / segmented models fails to enable machines to better understand the objective world, making it difficult to accurately identify complex, long-tail scenarios (those special situations that occur infrequently but involve many factors and are difficult to process). This results in an inability to accurately implement driving control based on the actual environment.
[0023] Therefore, based on the shortcomings of the above-mentioned driving control methods, the driving control method of this application is proposed. The solution of this application embodiment is: to simulate environmental dynamics and predict future states by integrating multimodal data (such as images, sounds, and sensor signals from collected environmental data). It is similar to the process by which humans learn the working mechanism of the world through observation, emphasizing state representation and transition models, thereby avoiding the phenomenon that visual language models or large one-piece / segmented models cannot enable machines to better understand our objective world, making it impossible to accurately identify long-tailed complex scenes, resulting in the phenomenon that the obtained control rules cannot fully conform to the actual environment. At the same time, it can determine the optimal control strategy based on the predicted state and the preset driving control model, thereby completing the driving control process, rather than directly determining the control rules based on the driver's driving experience. Thus, driving control can be achieved based on a new driving control method.
[0024] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a device capable of performing the above functions, such as a vehicle control terminal. The following description uses a vehicle control terminal as an example to illustrate this embodiment and the subsequent embodiments.
[0025] Based on this, the embodiments of this application provide a driving control method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the driving regulation control method of this application.
[0026] Reference Figure 1 This application provides a driving regulation method. In a first embodiment of the driving regulation method, the driving regulation method includes: Step S10: Obtain the vehicle's environmental data and perform multimodal data processing based on the environmental data to obtain the predicted control state. For example, the entire driving control system can acquire information from external data acquisition devices such as cameras, radar, and map terminals as environmental data. This environmental data primarily refers to vehicle-related data, such as images, point clouds, map data, external factors, status information (e.g., speed, direction, acceleration), and task objectives (e.g., navigation path, destination). The map terminal is the terminal for acquiring map data, such as a mobile phone or a locator installed in the vehicle. Further details can be found in... Figure 2 , Figure 2This is a flowchart illustrating the first embodiment of the driving control method of this application. The camera can extract visual features as image data using C3D (3nets) + iDT + linear SVM vector data. C3D is a convolutional neural network specifically designed for video analysis, capturing both spatial (intra-frame) and temporal (inter-frame) features in the video simultaneously through 3D convolutional kernels. iDT is a traditional method based on handcrafted features, extracting motion features by tracking keypoint trajectories with dense optical flow. Linear SVM fuses the features extracted by C3D and iDT and then uses a linear support vector machine for classification. Data augmentation techniques, such as random cropping, rotation, and brightness adjustment, can also be added to improve the model's robustness to different lighting and viewing angle conditions. Furthermore, more advanced convolutional neural networks, such as EfficientNet (Efficient Network) and RegNet (Regular Network), are considered for extracting image features to obtain more discriminative visual representations. Other methods for feature extraction can also be used, but these will not be detailed here. For example, radar extracts 3D spatial features using a PointNet network. Further preprocessing of the radar-acquired point cloud data, such as noise reduction and downsampling, reduces noise interference and computational load. It's also possible to explore combining other point cloud features, such as intensity information and normal direction, to enrich the feature representation of the point cloud. More efficient point cloud processing networks, such as PointTransformer, can be used to improve feature extraction. Other methods for feature extraction can also be used, but these will not be detailed here. For example, in addition to GNN (Graph Neural Network) coordinates and road data, map terminals can incorporate more map information, such as traffic rules and lane attributes (e.g., turning restrictions, dedicated lanes). Simultaneously, variations of graph neural networks, such as GAT (Graph Attention Networks), can be considered to better capture complex relationships in map data. Other methods for feature extraction can also be used, but these will not be detailed here. It is worth noting that in order to convert data of different modalities (such as images, point clouds, and radar echoes) into a unified feature representation, the data can be processed in cameras, radar, and map terminals to use the collected data with a unified feature representation as the collected environmental data.
[0027] It's worth noting that converting data into a unified feature representation is a common processing method for cameras, radar, and map terminals. This involves vectorizing and encrypting the data, and storing the necessary data processing programs within these devices. This allows for the collection and conversion of the required data into unified features. Furthermore, the output data from cameras, radar, and map terminals requires feature alignment and fusion. For example, the attention mechanism of Transformers can be used to align and fuse features from different modalities in both temporal and spatial dimensions, forming a unified environmental representation. This ensures that the data from different modalities are precisely synchronized in time, avoiding inaccurate feature alignment due to time discrepancies. Alternatively, hardware synchronization (e.g., using the same clock source) or software synchronization (e.g., timestamp-based interpolation algorithms) can be employed to improve the temporal consistency of multimodal data. For example, when using a clock source controller for synchronization, the synchronously acquired data is marked, and the marked data is then used as the time synchronization data. Of course, a related controller can also be used to mark spatial information to ensure the accuracy of spatial information. Both time synchronization and spatial synchronization methods can use existing time synchronization and spatial synchronization programs and store them in the clock source controller. Then, when performing data acquisition, the clock source controller executes the time synchronization and spatial synchronization programs to ensure the synchronization of data in time and space, thereby ensuring the accuracy of driving control.
[0028] In this embodiment, after determining the collected environmental data, multimodal data processing is performed based on the collected environmental data to obtain the predicted control state. The predicted control state refers to the predicted information about the future environment, such as predicted vehicle speed, direction, and path. It is worth noting that the multimodal data processing of the collected environmental data can be performed using a conventional world model, because the world model integrates multimodal data (such as images, sound, and sensor signals) to simulate environmental dynamics and predict future states. It is similar to the process by which humans learn the workings of the world through observation, emphasizing state representation and transition models. Of course, other models integrating multimodal data can also be used; this application uses a world model as an example for explanation. By introducing a world model into the driving control process to integrate multimodal data (such as images, sound, and sensor signals) to simulate environmental dynamics and predict future states, and because the world model itself is a process of learning the workings of the world through observation, emphasizing state representation and transition models, driving control can be accurately performed based on the actual environment.
[0029] For example, taking the world model as an example, the world model generates high-fidelity simulation scene data based on its own properties, reducing the dependence of the end-to-end autonomous driving model on real vehicle data by more than 80%, significantly reducing data acquisition costs, i.e., generating simulation data in subsequent data processing. On the other hand, compared to traditional autonomous driving which relies on massive amounts of human driving data, but high-difficulty scenarios (such as extreme weather and sudden accidents) account for less than 1%, the world model can autonomously generate a scene library covering more than 90% of extreme cases, breaking through the data quality bottleneck. Furthermore, the world model can also combine reinforcement learning (i.e., reinforcement learning design in policy evaluation learning module 30) to emerge with long thought chain capabilities that surpass human logic, enabling the autonomous driving system to handle multi-step decision-making problems in complex traffic scenarios. Of course, the world model can also embed physical rules, on the one hand, the generated videos strictly follow traffic rules and physical laws, such as automatically adjusting the distance when braking and ensuring that the nighttime light projection conforms to the real optical effect, reducing safety hazards caused by data deviation. On the other hand, adversarial testing capabilities: the world model can generate adversarial scenarios (such as pedestrians suddenly running into the road), and reinforcement learning optimizes the strategy through repeated trial and error, improving the system's ability to cope with extreme situations. Therefore, the use of the world model can, on the one hand, accurately control driving based on the actual environment, and on the other hand, ensure accurate data processing to achieve the accuracy of the entire control process.
[0030] Step S20: Determine the optimal control strategy based on the predicted control status and the preset driving control model, so as to carry out driving control of the vehicle based on the optimal control strategy.
[0031] In this embodiment, after determining the predicted control state based on the world model, the optimal control strategy is determined based on the predicted control state and the preset driving control model. The vehicle is then controlled according to this optimal control strategy. The optimal control strategy refers to the optimal control path, and may also include planned speed, direction, etc. Vehicle actuators can control direction (e.g., the steering wheel output controller) or speed (e.g., the accelerator and brake output controller). It is worth noting that the preset driving control model can be a model trained on training data, rather than a model directly defined based on user experience, thus ensuring the accuracy of the entire driving control process.
[0032] This embodiment provides a driving control method that acquires vehicle environmental data, performs multimodal data processing based on the acquired environmental data to obtain a predicted control state, and determines an optimal control strategy based on the predicted control state and a preset driving control model. This method integrates multimodal data (such as images, sounds, and sensor signals from the acquired environmental data) to simulate environmental dynamics and predict future states. It is similar to the process by which humans learn about the world's operating mechanisms through observation, emphasizing state representation and transition models. This avoids the limitations of visual language models or large, segmented / partial models that prevent machines from better understanding our objective world, making it difficult to accurately identify complex, long-tailed scenarios and resulting in control rules that do not fully conform to the actual environment. Furthermore, it determines the optimal control strategy based on the predicted state and a preset driving control model, thus completing the driving control process, rather than directly relying on the driver's driving experience to determine control rules. Therefore, it achieves driving control based on a novel approach.
[0033] Furthermore, based on the first embodiment of this application described above, a second embodiment of the driving regulation control method of this application is proposed. In this embodiment, step S10, the step of obtaining the predicted regulation control state by performing multimodal data processing based on the collected environmental data, includes: Step S11: Perform time modeling based on the collected environmental data to obtain time modeling data, and perform spatial modeling based on the collected environmental data to obtain spatial modeling data; Step S12: Based on the collected environmental data, environmental modeling is performed to obtain environmental modeling data. Based on the environmental modeling data, time modeling data, and spatial modeling data, the predicted control status is determined.
[0034] In this embodiment, taking the world model using Transformer for data processing as an example (other data processing methods are also possible, but Transformer is generally used in world models), the self-attention mechanism captures the long-term dependencies of the collected environmental data in the time series, and then uses these relationships as time modeling data. Of course, in addition to single-time-step encoding, multi-scale time representation methods can be used to divide the time series into different time scales (such as short-term, medium-term, and long-term), and design corresponding feature representations and attention mechanisms for each time scale. This can better capture dynamic information at different time scales and improve prediction accuracy. That is, time modeling data is obtained by performing time modeling based on a time modeling program. Time modeling data refers to the data obtained after performing time modeling on the collected data, and the time modeling program refers to the program that models the data in the time series. For example, in autonomous driving scenarios, continuous frames of radar point clouds and image data can be modeled to capture the motion trajectory of objects. Simultaneously, the global attention mechanism of Transformer is used to capture the spatial interaction relationships between targets such as vehicles, pedestrians, and traffic signs. The target location is encoded as a vector and input into the Transformer along with features. Then, through attention weights, the relationships between targets are learned to predict potential collisions or interactions. Spatial modeling data refers to data obtained by spatially modeling the collected data. For example, the spatial alignment between pixels in an image and points in a radar point cloud.
[0035] Furthermore, while modeling based solely on time and space can cover most prediction scenarios, to improve accuracy, environmental modeling data can be added. This involves incorporating external factors such as weather conditions (sunny, rainy, snowy) and time (daytime, nighttime), which significantly influence object trajectories. These external factors can be encoded as vectors and input into the Transformer along with multimodal features, enabling the model to learn the impact of these factors on object motion. Environmental modeling data refers to the data obtained after modeling the collected data with environmental factors.
[0036] In one embodiment, reference may be made to Figure 2The world model architecture processes multimodal inputs (collected environmental data) at the current moment to generate high-dimensional feature representations of the environment (i.e., various modeling data). For example, the world model can preprocess input image, radar point cloud, and map data, converting them into serialized representations suitable for Transformer processing. Then, based on the preprocessing, it converts the image, radar point cloud, and map data into serialized data suitable for Transformer processing (and models various modeling data from this serialized data). Linear embedding and positional encoding are then used: a linear layer maps the preprocessed features to a high-dimensional space, and positional encoding (such as spatial location and timestamps) is added to preserve the temporal and spatial information of the data. For example, when fusing image, radar point cloud, and map data, the Transformer's self-attention mechanism can be used: the encoder uses self-attention to interactively model features from different modalities. Self-attention allows the model to dynamically focus on correlations between different modalities, such as the association between objects in an image and corresponding points in a radar point cloud. Cross-modal attention can also be used: by modifying the attention mechanism (such as introducing cross-attention), the encoder can directly model interactions between different modalities. For example, radar point cloud features can guide the attention allocation of image features, thereby enhancing the detection of occluded objects. Different encoders abstract and fuse input features layer by layer through multi-layer self-attention and feedforward neural networks, generating feature representations containing multimodal information and spatiotemporal context. The output is a high-dimensional feature sequence that can be used as input for downstream tasks such as object detection, trajectory prediction, and scene understanding, providing a basis for subsequent predictions. This involves using environmental modeling data, temporal modeling data, and spatial modeling data to determine the predictive control state, because these data can effectively reconstruct the real world, thus ensuring the accuracy of subsequent predictive control state determination.
[0037] Furthermore, the steps for determining the predicted control state based on environmental modeling data, temporal modeling data, and spatial modeling data include: Step S121: Determine the proposed target control state in the preset prediction state table based on time modeling data and spatial modeling data, and determine the impact control state in the preset prediction state table based on environmental modeling data. Step S122: Based on the influence control state, the proposed target control state is corrected to obtain the predicted control state.
[0038] In this embodiment, the current vehicle state information (such as speed, direction, and acceleration) and the task objective (such as navigation path and destination) are encoded as environmental data to obtain environmental modeling data, temporal modeling data, and spatial modeling data. Then, the predicted control state is determined based on these three data sets. Specifically, the target control state can be directly determined from a preset predicted state table based on the temporal and spatial modeling data, and the influencing control state can be determined from the preset predicted state table based on the environmental modeling data. The preset predicted state table is a mapping table of different control states corresponding to temporal and spatial modeling data, and can be defined in advance. When the modeling data are A and B, the corresponding proposed target control state is C. The proposed target control state represents a possible control state. Because environmental influences need to be considered, the influence control state corresponding to the environmental modeling data in the preset prediction state table can be further determined. The influence control state refers to the state that affects the proposed target control state. For example, if the proposed target control state is straight-line acceleration, but the environmental modeling data indicates a rainy environment and signal interruption (due to acquisition errors), then the corrected predicted control state is straight-line driving to ensure the accuracy of the predicted control state. Alternatively, environmental modeling data can be used for precise correction, such as determining that a pedestrian crossing is being passed based on a nighttime environment. Furthermore, the predicted control state can be determined by combining environmental modeling data, temporal modeling data, and spatial modeling data to ensure its accuracy.
[0039] In one embodiment, the design of the entire world model may include fully connected layers. These layers can perform non-linear transformations on the output of the attention mechanism through feature transformation, thereby extracting higher-level features and helping the model learn more complex patterns and relationships. Furthermore, by introducing fully connected layers, the Transformer model can have more parameters, increasing its capacity and expressive power, thus enabling it to handle more complex tasks and data. Fully connected layers can also alleviate the vanishing gradient problem to some extent by introducing non-linear activation functions (such as ReLU), making the model easier to train. Further, in the output layer of the Transformer, fully connected layers are typically used to map the model's output to specific categories or values, thereby achieving classification or regression tasks. For example, in machine translation tasks, fully connected layers can map the decoder's output to the vocabulary of the target language to generate translation results. Exemplarily, the world model also includes linear layers. In the Transformer's encoder and decoder, linear layers are typically used in conjunction with activation functions (such as ReLU) to form part of an FFN (Feed-Forward Network). Linear layers first perform a linear transformation on the input, providing a basis for subsequent non-linear activations. Linear layers are used to adjust the dimensionality of features between different sub-layers of the Transformer (such as self-attention layers and feedforward neural network layers), ensuring smooth data transfer between layers. For example, after a self-attention mechanism, a linear layer can map attention-weighted features to the same dimension as the input for subsequent processing. In multi-head attention mechanisms, each head generates a set of attention-weighted features. Linear layers can be used to integrate these features from different heads, extracting more comprehensive information. Linear layers can also be used to transfer information between different layers, helping the model learn more complex feature representations. World models can also perform layer normalization. Layer normalization stabilizes the data distribution by standardizing the hidden size dimension, thereby accelerating the training process and ensuring smooth model convergence. In deep neural networks, the data distribution of the output of each layer may differ, leading to internal covariate shift problems. Layer normalization helps solve this problem by reducing the differences in data distribution between layers. Transformers typically handle variable-length sequences. Their layer normalization is independent of batch size, normalizing all features within each sample instead, making them more suitable for processing variable-length sequence data. Further details can be found in [reference needed]. Figure 2By using linear layers and a classifier, target prediction, scene understanding, and trajectory prediction are performed based on high-dimensional feature sequences to generate dynamic scenes, thereby determining the prediction and control strategy. The above describes the design of a commonly used world model. This leads to a comparison with the shortcomings of visual language models or segmented / piecewise large models. Based on this design, the world model can learn from the observation of the world's operating mechanisms, emphasizing state representation and transition models, thus enabling precise driving control based on the actual environment.
[0040] Furthermore, based on the first and / or second embodiments of this application described above, a third embodiment of the driving regulation control method of this application is proposed. In this embodiment, the driving regulation control method includes: Step S30: Obtain the current state input during the regulation control training, determine the regulation control action distribution corresponding to the current state based on the initial driving regulation control model, and determine the experience sample set according to the predetermined probability parameters and the regulation control action distribution; Step S40: Based on the initial driving control model, determine the training state value corresponding to the experience sample in the experience sample set, and determine the time difference error based on the training state value. Train the preset driving control model according to the time difference error and the experience sample set.
[0041] In this embodiment, before executing the entire driving control procedure, a preset driving control model needs to be trained. This is achieved by acquiring the current state input during training and then determining the distribution of control actions corresponding to the current state based on the initial driving control model. For example, a policy network, such as an Actor network, can be used to process the current state to obtain the control action distribution. Then, an experience sample set is determined based on predetermined probability parameters and the control action distribution. Here, the control action distribution refers to the distribution of the next planned control action in the current state. The predetermined probability parameter is a pre-set probability parameter used to select the target action. A higher predetermined probability parameter indicates a greater bias towards exploration, while a higher predetermined probability parameter indicates a greater bias towards exploitation. For example, the state and state reward are determined based on the predetermined probability parameter, and then the state and state reward are used as the experience sample set. Subsequently, the training state value corresponding to the experience sample in the experience sample set is determined based on the initial driving control model. This generally refers to determining the current value of the current state and the next value of the next predicted state, and these two values are used as the training state value. In this process, experience samples from the experience sample set can be input into the value network to evaluate the current value of the current state and the next predicted value of the next state. The value network is a Critic network, meaning the values of the two states are determined based on the value network. Then, the temporal difference error is calculated based on the values of the two states. For example, the formula for calculating the temporal difference error is: δi = ri + γVφ(si') - Vφ(si); where δi is the temporal difference error at the i-th time step; ri is the environmental reward obtained by performing the target action in the current state at the i-th time step; γ is a discount factor, typically a number between 0 and 1, used to discount future rewards, reflecting the importance placed on future rewards; Vφ(si') is the estimated value function of the next predicted state si', and Vφ(si) is the estimated value function of the current state si. The environmental reward obtained by performing the target action in the current state can be stored in the experience sample set, which can be determined based on predetermined probability parameters and the distribution of control actions. Therefore, the temporal difference error measures the accuracy of prediction by comparing the sum of the immediate reward of the current state (i.e., environmental reward) and the discounted estimate of future rewards with the estimate of the current state's value function. This difference guides the updating of model parameters, thereby optimizing the performance of the policy.Then, a preset driving control model is trained based on the temporal difference error and the empirical sample set. The initial driving control model is updated based on the temporal difference error. Based on the updated initial driving control model, the next state in the empirical sample set is used to execute the steps of determining the distribution of control actions corresponding to the current state based on the initial driving control model, until a predetermined termination condition is reached. The updated initial driving control model is then used as the preset driving control model to complete the training of the preset driving control model. The predetermined termination condition is a pre-set condition for terminating the iteration, such as the current iteration number being greater than or equal to the maximum iteration number, the loss function being minimized, or the initial driving control model being updated at least a certain number of times.
[0042] For example, refer to Figure 3 , Figure 3This diagram illustrates a training process for a pre-set driving control model in the driving control method of this application. The driving control system includes: a strategy network module 10, an environment model module 20, a value network module 30, a model update module 40, and a model iteration module 50. The strategy network module 10 is connected to the environment model module 20, the value network module 30, and the model iteration module 50. The model iteration module 50 is connected to the environment model module 20, the environment model module 20 is connected to the value network module 30, and the model update module 40 is connected to the value network module 30. The strategy network module 10 is used to obtain the current state of the driving environment model and input the current state into the strategy network to generate a control action distribution. The driving control strategy, i.e., the control action distribution, is then generated through the strategy network in the strategy network module 10. The environment model module 20 is used to select a target action from the control action distribution with predetermined probability parameters, simulate the target action in the driving environment model, predict the next predicted state and environmental reward, and use the current state, target action, next predicted state, and environmental reward as an experience sample set. The environment model module 20 selects target actions from the control action distribution using predetermined probability parameters. Compared to directly selecting the control action with the highest reward probability or randomly selecting control actions, this approach balances the exploration and utilization of control strategies according to needs, avoiding the predicament of getting stuck in local optima. The value network module 30 inputs experience samples from the experience sample set into the value network to evaluate the current value of the current state and the next value of the next predicted state. The model update module 40 calculates the temporal difference error based on the current value and the next value, and updates the policy network and value network based on the temporal difference error. The model iteration module 50 sends the next predicted state as the new current state to the policy network module 10, decays the predetermined probability parameters to obtain new predetermined probability parameters, and then sends them to the environment model module 20, until a predetermined termination condition is reached, resulting in a trained driving control system for driving control. Therefore, by attenuating the predetermined probability parameters to obtain new predetermined probability parameters and then iterating, the control strategy gradually shifts from exploration to utilization. As a result, the trained driving control system can better balance exploration and utilization and gradually converge to the optimal control strategy.
[0043] In one embodiment, the step of determining the empirical sample set based on predetermined probability parameters and the distribution of regulatory actions includes: Step S31: If the random number generated within the preset numerical range is less than the predetermined probability parameter, then a control action is randomly selected from the control action distribution as the target action. Step S32: If the random number generated within the preset numerical range is greater than or equal to the predetermined probability parameter, then the control action with the highest environmental reward probability in the control action distribution is taken as the target action. Step S33: Based on the target action prediction, the next predicted state and environmental reward are obtained, and the current state, target action, next predicted state and environmental reward are used as an experience sample set.
[0044] In this embodiment, a target action needs to be randomly selected. This can be done by random selection or by generating a random number within a preset value range. The target action is then selected based on the random number. Specifically, if the random number generated within the preset value range is less than a predetermined probability parameter, a control action is randomly selected from the control action distribution as the target action. If the random number generated within the preset value range is greater than or equal to the predetermined probability parameter, the control action with the highest environmental reward probability in the control action distribution is selected as the target action. The preset value range can be a randomly defined range. For example, the control action distribution includes actions S1-SN after the current state. Therefore, an action SM can be randomly selected as the target action, or the control action SM+1 with the highest environmental reward probability can be selected as the target action. Here, SM and SM+1 belong to S1-SN. The highest environmental reward probability refers to a predefined environmental reward probability. For example, in an overspeeding state, the deceleration action definitely has the highest environmental reward probability. Alternatively, the environmental reward probability of each action can be predefined according to actual rules and design. For instance, the predetermined probability parameter can be multiplied by a decay coefficient to obtain a new predetermined probability parameter. The decay coefficient is a value greater than zero and less than 1, thus making the new predetermined probability parameter less than the predetermined probability parameter before decay. For example, in environments with unclear rewards, such as sparse or no rewards, the agent may "go in circles" and fail to explore the optimal strategy. Therefore, in this embodiment, an intrinsic reward can be calculated by obtaining the next true state and the next predicted state, and calculating the intrinsic reward based on the state error between the two states. The intrinsic reward is positively correlated with the state error. Furthermore, a new environmental reward is generated based on the intrinsic reward and the environmental reward; that is, an environmental reward is defined for each state, and a comprehensive reward can be generated by combining the intrinsic reward. The state error can be obtained by calculating the mean squared error, cross-entropy, and other loss function values between the next true state and the next predicted state. Thus, the intrinsic reward can encourage the driving control system to explore states that are novel or uncertain to it, which is beneficial for discovering potential rewards and accelerating the learning speed. Furthermore, to prevent the driving control system from maintaining a high level of curiosity about unknown exploration directions, which could lead to difficulty in convergence, the intrinsic reward can be subject to various decay methods, such as linear decay (i.e., the intrinsic reward gradually decreases linearly with the number of explorations), exponential decay (i.e., the intrinsic reward gradually decreases exponentially with the number of explorations), and custom decay (i.e., a decay function is set according to the specific task requirements to adjust the intrinsic reward). Therefore, this embodiment allows the driving control system to maintain its curiosity about the unknown environment while gradually shifting its attention to tasks that are more likely to bring practical benefits.Ultimately, the next predicted state and environmental reward can be determined based on the above methods. For example, using a conventional prediction model, such as the Bellman expectation equation, the next predicted state and environmental reward will be predicted. The current state, target action, next predicted state, and environmental reward will be used as an empirical sample set to provide a basis for subsequent processing.
[0045] In one embodiment, the step of training a preset driving control model based on time-series difference errors and an empirical sample set includes: Step S41: Update the network parameters and value parameters in the initial driving control model based on the time series difference error, and determine the next prediction state in the empirical sample set. Step S42: Take the next predicted state as the current state input during the regulation control training, and execute the step of determining the regulation control action distribution corresponding to the current state based on the updated initial driving regulation control model, until the predetermined termination condition is reached, and take the updated initial driving regulation control model as the preset driving regulation control model.
[0046] In this embodiment, the network parameters and value parameters in the initial driving control model are updated based on the temporal difference error, and the next prediction state in the empirical sample set is determined. The training and learning process described in the previous embodiment continues until a predetermined termination condition is met, at which point the updated initial driving control model is used as the preset driving control model. For example, updating the network parameters in the initial driving control model based on the temporal difference error can be achieved by using a policy gradient with entropy regularization to update the network parameters of the policy network in the initial driving control model. The formula for calculating the policy gradient with entropy regularization is: θJ(θ) = θlogπθ(ai|si)δi+λ θH(πθ(si)); in, θJ(θ) represents the gradient of the objective function J(θ) with respect to the policy parameters θ (i.e., network parameters); θlogπθ(ai|si) represents the gradient of the logarithmic policy function πθ(ai|si) with respect to the policy parameter θ, where πθ(ai|si) represents the probability of choosing action ai in the current state si; δi is the temporal difference error at the i-th time step; λ is the weight coefficient of the entropy regularization term, which characterizes the degree of influence of the entropy regularization term on the total gradient. The larger the value of λ, the more obvious the effect of the entropy regularization term, and the stronger the exploratory nature of the policy. θH(πθ(si)) represents the gradient of the entropy of policy πθ (i.e., the distribution of regulatory actions) in the current state si with respect to the policy parameter θ, and πθ(si) represents the distribution of regulatory actions in the current state si. Network parameters are then updated based on this. For example, the value parameter update can be a linear decay of the intrinsic reward. Furthermore, by introducing entropy regularization to increase the linear decay of both the policy entropy and the intrinsic reward, the uncertainty of the policy is maintained, preventing premature convergence to a deterministic policy and thus encouraging exploration.
[0047] Furthermore, based on the first, second, and / or third embodiments of this application described above, a fourth embodiment of the driving regulation control method of this application is proposed. In this embodiment, the step of determining the optimal regulation control strategy based on the predicted regulation control state and the preset driving regulation control model includes: Step S21: Determine the predicted state value and predicted state output in the preset driving control model; Step S22: Determine the optimal control strategy based on the predicted state value and the predicted state.
[0048] In this embodiment, the predicted control state can be directly input into the preset driving control model to determine the output predicted state value and predicted state. For example, taking the predicted control state T, the preset driving control model after training can output the predicted state after the next predicted action of the predicted control state T and the value of the predicted state. The predicted state refers to the state after the predicted control state T performs an action A1, and the predicted state value refers to the reward of the predicted state. Then, with each output of the predicted state value and predicted state, the subsequent predicted state value and predicted state are output step by step, and the optimal control strategy is determined by combining the overall results. In another embodiment, the optimal control policy can also be generated based on the predicted control state using reinforcement learning. The reinforcement learning program can be one of the commonly used Q-Learning (Q-Value Learning), DQN (Deep Q-Network), DDPG (Deep Deterministic Policy Gradient), or PPO (Proximal Policy Optimization). These algorithms can be directly stored in the reinforcement learning controller to generate the optimal control policy from the predicted control policy. For example, see [reference needed]. Figure 2 The predictive control policy is used as input for environmental information, and then the value function Vπ(s) of a given policy π is calculated based on the state and action. That is, the cumulative reward expectation of the agent starting from state s when following policy π, such as using the Bellman expectation equation (1): (1) Where Rt+1 is the immediate reward, i.e., the reward value of the strategy, γ is the discount factor, between 0 and 1; Vπ(St+1) is the value of the next state. At this time, the value of each state and the immediate reward can also be input into different memories, and then connected to the adder based on memory selectivity to output the expected cumulative reward. Furthermore, iterative updates will be performed based on the expected cumulative reward, such as the iterative update formula (2): (2) Where π(a|s): the probability that the policy chooses action a in state s; p(st+1,r|s,a): the state transition probability, and Vk+1(S) represents the average value of the entire policy. Then, initialize V0(s)=0. For all states, update Vk+1(s) repeatedly until convergence to determine the optimal policy. At this time, the next optimal policy will be generated based on the optimal policy. That is, the policy improvement subunit 32 can generate a better policy π′ based on the current value function Vπ. For each state s, the action that maximizes the sum of the immediate reward and the future value is selected as shown in the following formula (3): (3) Simplifying formula (3) yields formula (4): (4) Where Qπ(s,a) is the action value function.
[0049] That is, if π′≠π, then the value function Vπ′ of the new policy π′ is strictly better than Vπ (or the two are equal). Finally, the policy iteration subunit 33 alternates between policy evaluation and policy improvement until convergence to the optimal policy. In other words, the entire process is as follows: policy evaluation subunit 31 calculates the value function Vπ by fixing the current policy π; policy improvement subunit 32 generates a new policy π′ based on Vπ; and policy iteration subunit 33 finds the optimal policy π if π′=π. and corresponding value function V Otherwise, let π←π′, return and continue execution until convergence. It's worth noting that at this point, an optimal control strategy can be determined based on the reinforcement learning approach described above. This optimal strategy can then be used to refine the strategy determined based on the predicted state value and predicted state, ensuring the accuracy of the final strategy. For example, the refinement method could primarily rely on determining the optimal control strategy based on the predicted state value and predicted state, making minor adjustments in control directions where there is a significant difference between the two, or adjusting within the safe driving range, to guarantee the accuracy of the final driving control and the driving effect.
[0050] In one embodiment, an action space can be designed, i.e., a discrete action space can be defined, which is a set of discrete actions, such as acceleration, deceleration, left turn, right turn, lane keeping, lane changing, etc. Each action corresponds to a discrete label, and the reinforcement learning model outputs the probability distribution of the action. A continuous action space can be defined, i.e., a continuous action space, such as acceleration, steering angle, etc. The reinforcement learning model outputs specific numerical values for the actions, suitable for scenarios requiring more precise control. Reward functions can be designed, such as safety rewards, which encourage the model to avoid collisions. For example, when the distance between the vehicle and an obstacle is less than a safety threshold, a negative reward is given; when the vehicle maintains a safe distance, a positive reward is given. Efficiency rewards encourage efficient driving. For example, rewards are given based on the vehicle's speed, acceleration, and deviation from the target path. Task objective rewards encourage the model to complete navigation tasks. For example, a positive reward is given when the vehicle approaches the destination; a negative reward is given when the vehicle deviates from the navigation path. Comfort rewards encourage smooth driving. For example, rewards are given based on the rate of change of acceleration and steering angle. Of course, adaptive settings for self-learning rewards can also be implemented, but these will not be detailed here. For an example, please refer to... Figure 2 The world model acts as an environment simulator, predicting future environmental states using a Transformer decoder, which then serves as input to the reinforcement learning agent. The world model simulates the impact of the agent's actions on the environment, generating new states and rewards. For example, when the agent chooses to accelerate, the world model predicts changes in vehicle speed and the reactions of surrounding targets. Reinforcement learning is then used to optimize the control strategy, allowing the reinforcement learning agent to learn the optimal control strategy based on the world model's predicted states and rewards. For instance, the PPO algorithm is used to optimize the policy function, generating optimal acceleration and steering actions. Simultaneously, in actual driving, the world model's predictions are combined with real-time sensor data to dynamically adjust the control strategy. For example, when the world model predicts an obstacle ahead, the agent decelerates or changes lanes in advance, thus ensuring the accuracy of driving control through the world model and reinforcement learning.
[0051] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the driving control method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0052] This application also provides a driving control device, referring to... Figure 4 Driving control devices: The data acquisition module A10 is used to acquire environmental data of the vehicle and perform multimodal data processing based on the acquired environmental data to obtain the predicted control status. The driving control module A20 is used to determine the optimal control strategy based on the predicted control status and the preset driving control model, so as to carry out driving control of the vehicle based on the optimal control strategy.
[0053] The driving regulation control system provided in this application, employing the driving regulation control method in the above embodiments, can provide a new driving regulation control approach. Compared with the prior art, the beneficial effects of the driving regulation control system provided in this application are the same as those of the driving regulation control method provided in the above embodiments, and other technical features of the driving regulation control system are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0054] This application provides a driving control device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the driving control method in the above embodiment 1.
[0055] The following is for reference. Figure 5 The diagram illustrates a structural schematic of a driving control device suitable for implementing embodiments of this application. The driving control device in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The driving control device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0056] like Figure 5As shown, the driving control device may include a processing system 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage system 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the driving control device. The processing system 1001, the ROM 1002, and the RAM 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input system 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output system 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage system 1003 including, for example, magnetic tape, hard disk, etc.; and a communication system 1009. Communication system 1009 allows the driving control equipment to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows driving control equipment with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0057] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication system, or installed from storage system 1003, or installed from read-only memory 1002. When the computer program is executed by processing system 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0058] The driving control device provided in this application, employing the driving control method in the above embodiments, can provide a new driving control approach. Compared with the prior art, the beneficial effects of the driving control device provided in this application are the same as those of the driving control method provided in the above embodiments, and other technical features of the driving control device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0059] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0060] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0061] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the driving control method in the above embodiments.
[0062] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0063] The aforementioned computer-readable storage medium may be included in the driving control device; or it may exist independently and not be installed in the driving control device.
[0064] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the driving control device, cause the driving control device to: Acquire environmental data of the vehicle, and perform multimodal data processing based on the acquired environmental data to obtain the predicted control state; The optimal control strategy is determined based on the predicted control status and the preset driving control model, so as to carry out driving control of the vehicle based on the optimal control strategy.
[0065] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0066] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0067] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0068] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the above-described driving control method, thus providing a new driving control approach. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the driving control method provided in the above embodiments, and will not be repeated here.
[0069] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the driving control method described above.
[0070] The computer program product provided in this application can provide a new driving control method. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the driving control method provided in the above embodiments, and will not be repeated here.
[0071] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A driving regulation control method, characterized in that, The driving control method includes: Acquire vehicle environmental data, and perform multimodal data processing based on the acquired environmental data to obtain the predicted control state; The optimal control strategy is determined based on the predicted control state and the preset driving control model, so as to carry out driving control of the vehicle based on the optimal control strategy.
2. The driving regulation control method as described in claim 1, characterized in that, The step of obtaining the predicted control state by performing multimodal data processing based on the collected environmental data includes: Temporal modeling data is obtained by performing time modeling based on the collected environmental data, and spatial modeling data is obtained by performing spatial modeling based on the collected environmental data. Environmental modeling data is obtained by performing environmental modeling based on the collected environmental data, and the predictive control status is determined based on the environmental modeling data, the time modeling data, and the spatial modeling data.
3. The driving regulation control method as described in claim 2, characterized in that, The step of determining the predictive control state based on the environmental modeling data, the temporal modeling data, and the spatial modeling data includes: Based on the time modeling data and the spatial modeling data, the proposed target control state is determined in the preset prediction state table, and the influence control state is determined in the preset prediction state table based on the environmental modeling data. Based on the influence control state, the proposed target control state is modified to obtain the predicted control state.
4. The driving regulation control method according to any one of claims 1 to 3, characterized in that, The driving control method includes: The current state input during the regulation control training is obtained, and the distribution of regulation control actions corresponding to the current state is determined based on the initial driving regulation control model. An experience sample set is determined according to a predetermined probability parameter and the distribution of regulation control actions. Based on the initial driving control model, the training state value corresponding to the experience sample in the experience sample set is determined, and the temporal difference error is determined based on the training state value. The preset driving control model is then trained based on the temporal difference error and the experience sample set.
5. The driving regulation control method as described in claim 4, characterized in that, The step of determining the empirical sample set based on the predetermined probability parameters and the control action distribution includes: If the random number generated within the preset numerical range is less than the predetermined probability parameter, then a control action is randomly selected from the control action distribution as the target action. If the random number generated within the preset numerical range is greater than or equal to the predetermined probability parameter, then the control action with the highest environmental reward probability in the control action distribution will be taken as the target action. The next predicted state and environmental reward are obtained based on the target action prediction, and the current state, the target action, the next predicted state, and the environmental reward are used as an experience sample set.
6. The driving regulation control method as described in claim 4, characterized in that, The step of training a preset driving control model based on the time-series difference error and the experience sample set includes: The network parameters and value parameters in the initial driving control model are updated based on the time-series difference error, and the next prediction state in the empirical sample set is determined. The next predicted state is used as the current state input during the regulation control training, and the step of determining the regulation control action distribution corresponding to the current state based on the updated initial driving regulation control model is executed until a predetermined termination condition is reached, and the updated initial driving regulation control model is used as the preset driving regulation control model.
7. The driving regulation control method according to any one of claims 1 to 3, characterized in that, The step of determining the optimal control strategy based on the predicted control state and the preset driving control model includes: Determine the predicted state value and predicted state output in the preset driving control model; The optimal control strategy is determined based on the predicted state value and the predicted state.
8. A driving control device, characterized in that, The driving control device includes: The data acquisition module is used to acquire environmental data of the vehicle and perform multimodal data processing based on the acquired environmental data to obtain the predicted control state. The driving regulation control module is used to determine the optimal regulation control strategy based on the predicted regulation control state and the preset driving regulation control model, so as to perform driving regulation control on the vehicle based on the optimal regulation control strategy.
9. A driving control device, characterized in that, The driving control device includes a processor and a memory. The memory stores a driving control method program that can run on the processor. When the driving control method program is executed by the processor, it implements the steps of the driving control method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a driving regulation control method program, wherein when the driving regulation control method program is executed by a processor, it implements the steps of the driving regulation control method as described in any one of claims 1 to 7.