Multi-modal fusion automatic driving decision arbitration method and system based on VLA model
Patent Information
- Application Number
- CN202610747110.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-05-28
AI Technical Summary
然而,这些方法存在明显的局限性:静态优先级表不能应对动态变化的复杂场景,容易导致系统行为生硬且缺乏灵活性;加权融合由于权重固定,不能根据场景动态调整,且对硬约束(如AEB)处理不当;简单切换则因场景分类边界模糊容易误切换,且切换过程不平滑,不能处理混合场景
[0008]The beneficial effects of this application are as follows: This application obtains the temporal state information, current environmental information, and multi-source decision information of an unmanned logistics vehicle. Based on a pre-built arbitration network model, it arbitrates the temporal state information, current environmental information, and multi-source decision information to obtain the autonomous driving arbitration result of the unmanned logistics vehicle. The arbitration process for the temporal state information, current environmental information, and multi-source decision information includes: feature encoding the temporal state information to obtain state feature codes; feature encoding the current environmental information to obtain environmental feature codes; and feature encoding the multi-source decision information to obtain decision feature codes. Based on a self-attention mechanism module, feature fusion is performed on the decision fusion features, state feature codes, and environmental feature codes to obtain the multi-source decision fusion result. The system uses multi-source fusion features to classify and make decisions, resulting in an autonomous driving arbitration result. This process, through a self-attention mechanism module, automatically learns and calculates the attention weight distribution of different decision sources in the current scenario based on temporal state information, current environmental information, and multi-source decision information. Feature fusion is then performed based on this attention weight distribution, and the autonomous driving arbitration result is obtained from the fusion result. This allows the arbitration result to smoothly and flexibly adapt to various complex and changing traffic scenarios, improving the flexibility and robustness of decision-making. It solves the problems of traditional "static priority tables" being unable to cope with dynamically changing scenarios and exhibiting rigid behavior, as well as the poor adaptability caused by fixed weights in "simple weighted fusion." Furthermore, through attention... The force weight distribution provides a clear view of whether the pre-built arbitration network model relies more on "rule constraints" or "human instructions" at a specific moment. This provides a traceable basis for system fault analysis, logic verification, and safety authentication, improving the system's interpretability and debuggability. By incorporating current traffic constraints into the pre-built arbitration network model and fusing them with the current planned trajectory and current executed actions, the pre-built arbitration network model can explicitly "perceive" safety boundaries and rule restrictions when making classification decisions. This retains the intelligence of the VLA model while forcibly introducing the constraint of safety rules, effectively avoiding potential violations or dangerous behaviors in the end-to-end model and meeting the requirements of Level 4 autonomous driving. The high requirements for functional safety significantly improve the safety of autonomous driving, solving the problem that the existing VLA end-to-end model, as a "black box," lacks safety constraint checks and cannot guarantee safety in extreme scenarios. By fusing features of the current human control command, the current planned trajectory, and the current executed action, the pre-built arbitration network model can understand the difference between the intention of the human command and the current planned trajectory, and make the optimal compromise or switching decision based on the current environmental state. This not only solves the conflict of multi-source heterogeneous information, but also ensures a smooth transition of human-machine co-driving, avoids abrupt changes in vehicle actions, and solves the problem of unsmooth switching and easy oscillation in traditional methods when dealing with mixed scenarios (such as conflicts between human commands and automatic planning).
Smart Images

Figure CN122275953B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, and in particular to a multimodal fusion autonomous driving decision arbitration method and system based on VLA model. Background Technology
[0002] Level 4 autonomous driving systems typically employ a hierarchical decision-making architecture, comprising multiple layers of decision-making modules, such as a rule / safety module, a motion planning module, a VLA (Vision-Language-Action) master model, and a remote / human-machine interaction module. The rule / safety module outputs safety-related hard constraint commands based on preset rules, primarily used in scenarios such as AEB (Autonomous Emergency Braking), FCW (Forward Collision Warning), electronic fences, and solid line constraints. The motion planning module calculates the optimal trajectory based on optimization algorithms (e.g., Model Predictive Control (MPC), convex optimization, etc.). The VLA master model outputs driving decisions based on end-to-end deep learning models (e.g., Transformer-based models, hybrid convolutional neural network-recurrent neural network models, etc.). The remote / human-machine interaction module receives human commands from a remote cockpit in the cloud (e.g., "pull over," "brake immediately," control the steering wheel, control the accelerator, etc.). However, in complex scenarios, the outputs of different decision-making modules may conflict, so an arbitration mechanism is needed to coordinate and resolve the conflicts.
[0003] Traditional arbitration methods mainly include static priority tables, weighted fusion, and simple switching. However, these methods have significant limitations: static priority tables cannot handle complex, dynamically changing scenarios, easily leading to rigid and inflexible system behavior; weighted fusion, due to fixed weights, cannot be dynamically adjusted according to the scenario and improperly handles hard constraints (such as AEB); simple switching is prone to erroneous switching due to ambiguous scenario classification boundaries, and the switching process is not smooth, making it unable to handle mixed scenarios. Existing VLA models typically function as a single decision source, lacking effective fusion capabilities with other decision sources (such as rule modules or manual instructions); furthermore, their end-to-end black-box nature results in poor interpretability, hindering security authentication and problem debugging; more importantly, VLA models lack security constraint checks on the output, making it difficult to guarantee security in extreme scenarios.
[0004] Therefore, there is an urgent need for a decision arbitration method that can dynamically adapt to complex scenarios while ensuring security and interpretability. Summary of the Invention
[0005] In view of the shortcomings of the prior art described above, this application provides a multimodal fusion autonomous driving decision arbitration method and system based on VLA model to solve the above technical problems.
[0006] According to one aspect of the embodiments of this application, a multimodal fusion autonomous driving decision arbitration method based on a VLA model is provided. This method acquires temporal state information, current environmental information, and multi-source decision information of an unmanned logistics vehicle. The multi-source decision information includes: current traffic constraints, current planned trajectory, current executed action, and current manual control command. The current traffic constraints are triggered and generated based on the temporal state information and the current environmental information. The current executed action is generated by the VLA model with the temporal state information, the current environmental information, and the moving target state information as input. The current planned trajectory is generated by a motion planning module with the temporal state information, the current environmental information, and the moving target state information as input. Based on a pre-built arbitration network model, the temporal state information, the current environmental information, and the multi-source decision information are arbitrated to obtain the autonomous driving decision of the unmanned logistics vehicle. The autonomous driving arbitration result includes: autonomous driving trajectory and autonomous driving actions; wherein, the arbitration process for the temporal state information, the current environment information, and the multi-source decision information includes: feature encoding of the temporal state information to obtain state feature encoding; feature encoding of the current environment information to obtain environment feature encoding; and feature encoding of the multi-source decision information to obtain decision feature encoding; the decision feature encoding includes: traffic constraint encoding, planned trajectory encoding, executed action encoding, and control command encoding; based on the self-attention mechanism module, feature fusion is performed on the decision fusion feature, the state feature encoding, and the environment feature encoding to obtain multi-source fusion feature; the decision fusion feature is obtained by fusing based on the attention weight distribution between the decision feature encodings; classification decision is performed on the multi-source fusion feature to obtain the autonomous driving arbitration result.
[0007] According to another aspect of the embodiments of this application, a multimodal fusion autonomous driving decision arbitration system based on a VLA model is also provided, comprising: an information acquisition module, used to acquire temporal state information, current environmental information, and multi-source decision information of an unmanned logistics vehicle, wherein the multi-source decision information includes: current traffic constraints, current planned trajectory, current executed action, and current manual control command; the current traffic constraints are triggered and generated based on the temporal state information and the current environmental information; the current executed action is generated by the VLA model under the input conditions of the temporal state information, the current environmental information, and the moving target state information; the current planned trajectory is generated by the motion planning module under the input conditions of the temporal state information, the current environmental information, and the moving target state information; and an information arbitration module, used to arbitrate the temporal state information, the current environmental information, and the multi-source decision information based on a pre-built arbitration network model, to obtain... The autonomous driving arbitration result of the unmanned logistics vehicle includes: autonomous driving trajectory and autonomous driving execution actions; wherein, the arbitration process of the temporal state information, the current environment information, and the multi-source decision information includes: feature encoding of the temporal state information to obtain state feature encoding; feature encoding of the current environment information to obtain environment feature encoding; and feature encoding of the multi-source decision information to obtain decision feature encoding; the decision feature encoding includes: traffic constraint encoding, planned trajectory encoding, execution action encoding, and control command encoding; based on the self-attention mechanism module, feature fusion is performed on the decision fusion feature, the state feature encoding, and the environment feature encoding to obtain multi-source fusion feature; the decision fusion feature is obtained by fusing based on the attention weight distribution between the decision feature encodings; classification decision is performed on the multi-source fusion feature to obtain the autonomous driving arbitration result.
[0008] The beneficial effects of this application are as follows: This application obtains the temporal state information, current environmental information, and multi-source decision information of an unmanned logistics vehicle. Based on a pre-built arbitration network model, it arbitrates the temporal state information, current environmental information, and multi-source decision information to obtain the autonomous driving arbitration result of the unmanned logistics vehicle. The arbitration process for the temporal state information, current environmental information, and multi-source decision information includes: feature encoding the temporal state information to obtain state feature codes; feature encoding the current environmental information to obtain environmental feature codes; and feature encoding the multi-source decision information to obtain decision feature codes. Based on a self-attention mechanism module, feature fusion is performed on the decision fusion features, state feature codes, and environmental feature codes to obtain the multi-source decision fusion result. The system uses multi-source fusion features to classify and make decisions, resulting in an autonomous driving arbitration result. This process, through a self-attention mechanism module, automatically learns and calculates the attention weight distribution of different decision sources in the current scenario based on temporal state information, current environmental information, and multi-source decision information. Feature fusion is then performed based on this attention weight distribution, and the autonomous driving arbitration result is obtained from the fusion result. This allows the arbitration result to smoothly and flexibly adapt to various complex and changing traffic scenarios, improving the flexibility and robustness of decision-making. It solves the problems of traditional "static priority tables" being unable to cope with dynamically changing scenarios and exhibiting rigid behavior, as well as the poor adaptability caused by fixed weights in "simple weighted fusion." Furthermore, through attention... The force weight distribution provides a clear view of whether the pre-built arbitration network model relies more on "rule constraints" or "human instructions" at a specific moment. This provides a traceable basis for system fault analysis, logic verification, and safety authentication, improving the system's interpretability and debuggability. By incorporating current traffic constraints into the pre-built arbitration network model and fusing them with the current planned trajectory and current executed actions, the pre-built arbitration network model can explicitly "perceive" safety boundaries and rule restrictions when making classification decisions. This retains the intelligence of the VLA model while forcibly introducing the constraint of safety rules, effectively avoiding potential violations or dangerous behaviors in the end-to-end model and meeting the requirements of Level 4 autonomous driving. The high requirements for functional safety significantly improve the safety of autonomous driving, solving the problem that the existing VLA end-to-end model, as a "black box," lacks safety constraint checks and cannot guarantee safety in extreme scenarios. By fusing features of the current human control command, the current planned trajectory, and the current executed action, the pre-built arbitration network model can understand the difference between the intention of the human command and the current planned trajectory, and make the optimal compromise or switching decision based on the current environmental state. This not only solves the conflict of multi-source heterogeneous information, but also ensures a smooth transition of human-machine co-driving, avoids abrupt changes in vehicle actions, and solves the problem of unsmooth switching and easy oscillation in traditional methods when dealing with mixed scenarios (such as conflicts between human commands and automatic planning).
[0009] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0010] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0011] Figure 1 This is a schematic diagram illustrating an exemplary system architecture as shown in an exemplary embodiment of this application;
[0012] Figure 2 This is a flowchart illustrating an exemplary embodiment of the multimodal fusion autonomous driving decision arbitration method based on the VLA model, as shown in this application.
[0013] Figure 3 This is a schematic diagram illustrating the structure of a pre-built arbitration network model, as shown in an exemplary embodiment of this application.
[0014] Figure 4 This is a flowchart illustrating a pre-built arbitration network model, as shown in an exemplary embodiment of this application;
[0015] Figure 5 This is a schematic diagram illustrating the attention weight distribution of multi-source decision information, as shown in an exemplary embodiment of this application.
[0016] Figure 6 This is a block diagram illustrating a multimodal fusion autonomous driving decision arbitration system based on a VLA model, as shown in an exemplary embodiment of this application. Detailed Implementation
[0017] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0018] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. The drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0019] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the present application. However, it will be apparent to those skilled in the art that embodiments of the present application may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the present application.
[0020] Figure 1 This is a schematic diagram illustrating an exemplary system architecture as shown in an exemplary embodiment of this application.
[0021] Reference Figure 1 As shown, the system architecture may include a storage device 101 and a processing device 102. The processing device 102 may be at least one of a desktop graphics processing unit (GPU) computer, a GPU computing cluster, or a neural network computer. Those skilled in the art can use the processing device 102 to acquire the temporal state information, current environmental information, and multi-source decision information of the unmanned logistics vehicle. Based on a pre-built arbitration network model, the temporal state information, current environmental information, and multi-source decision information are arbitrated to obtain the autonomous driving arbitration result of the unmanned logistics vehicle. The arbitration process for the temporal state information, current environmental information, and multi-source decision information includes: feature encoding of the temporal state information to obtain state feature encoding; feature encoding of the current environmental information to obtain environmental feature encoding; feature encoding of the multi-source decision information to obtain decision feature encoding; feature fusion of the decision fusion feature, state feature encoding, and environmental feature encoding based on a self-attention mechanism module to obtain multi-source fusion features; and classification decision based on the multi-source fusion features to obtain the autonomous driving arbitration result. The storage device 101 is used to store the time-series status information, current environmental information and multi-source decision information of the unmanned logistics vehicle, and to provide the time-series status information, current environmental information and multi-source decision information of the unmanned logistics vehicle to the processing device 102 for processing.
[0022] The implementation details of the technical solutions in the embodiments of this application are described in detail below:
[0023] Figure 2 This is a flowchart illustrating an exemplary embodiment of the multimodal fusion autonomous driving decision arbitration method based on a VLA model, as shown in this application. (Refer to...) Figure 2As shown, the multimodal fusion autonomous driving decision arbitration method based on the VLA model includes at least steps S210 to S220, which are described in detail below:
[0024] In step S210, the temporal state information, current environmental information, and multi-source decision information of the unmanned logistics vehicle are acquired. In one embodiment of this application, the multi-source decision information includes: current traffic constraints, current planned trajectory, current executed action, and current manual control command. The current traffic constraints are generated based on the temporal state information and current environmental information, and are implemented through a rule / safety module. The rule / safety module can be a pre-trained deep learning network model with functions such as traffic light constraint detection, lane keeping constraint detection, speed limit and right-of-way constraint detection, collision detection, dynamic feasibility module, and right-of-way game. The current executed action is generated by the VLA model under the input conditions of temporal state information, current environmental information, and moving target state information. The current planned trajectory is generated by the motion planning module under the input conditions of temporal state information, current environmental information, and moving target state information. The motion planning module uses MPC (Model Predictive Control) algorithm, convex optimization algorithm, etc., to generate the current planned trajectory under the input conditions of temporal state information, current environmental information, and moving target state information. Currently, manual control commands are obtained from the cloud-based remote cockpit and human-machine interface devices of the unmanned logistics vehicle. The temporal state information includes: the current state information and the state information within a preset time period prior to the current moment. This state information includes: driving position, driving speed, driving acceleration, and driving angular velocity. The driving position is determined by a global navigation satellite system; the driving speed is collected by speed sensors, driving acceleration by acceleration sensors, and driving angular velocity by angular velocity sensors. Current environmental information includes: road signage information (e.g., lane boundary lines, lane center lines, lane driving direction) and static obstacle information (e.g., fence location information, sign location information). This environmental information is perceived through LiDAR, cameras, millimeter-wave radar, etc.
[0025] In step S220, based on a pre-built arbitration network model, the time-series state information, current environmental information, and multi-source decision information are arbitrated to obtain the autonomous driving arbitration result of the unmanned logistics vehicle. In one embodiment of this application, the autonomous driving arbitration result includes: autonomous driving trajectory, autonomous driving execution actions, etc.; the pre-built arbitration network model includes an input layer, a feature fusion layer, and an output layer. The input layer includes: a state feature encoder, an environmental feature encoder, a traffic constraint encoder, a planned trajectory encoder, an execution action encoder, and a control command encoder. The state feature encoder is used to perform feature encoding on the time-series state information to obtain state feature encoding; the environmental feature encoder is used to perform feature encoding on the current environmental information to obtain environmental feature encoding; the traffic constraint encoder is used to encode the current traffic constraints to obtain traffic constraint encoding; the planned trajectory encoder is used to encode the current planned trajectory to obtain planned trajectory encoding; the execution action encoder is used to encode the current execution action to obtain execution action encoding; and the control command encoder is used to encode the current manual control command to obtain control command encoding; the feature fusion layer uses a 4-layer Transformer. The encoder distinguishes decision information from different sources through positional encoding, calculates the attention weight distribution among decision feature encodings using a self-attention mechanism, and fuses the decision feature encodings based on this distribution. It also calculates the attention distribution among decision fusion features, state feature encodings, and environment feature encodings using the same self-attention mechanism, and performs feature fusion on these three types of encodings. The output layer includes a decision output head and a weight output head. When the network structure of the decision output head consists of a 512-dimensional linear layer, a 256-dimensional linear layer, and a 100-dimensional linear layer, it outputs the autonomous driving trajectory. When the network structure of the decision output head consists of a 512-dimensional linear layer, a 256-dimensional linear layer, and a Tanh activation function, it outputs the autonomous driving actions. When the network structure of the weight output head consists of a 512-dimensional linear layer, a 256-dimensional linear layer, a 100-dimensional linear layer, and a Softmax activation function, it outputs a visualized attention weight distribution of multi-source decision information.
[0026] In one embodiment of this application, the process of arbitrating temporal state information, current environmental information, and multi-source decision information includes:
[0027] The temporal state information is feature-encoded to obtain state feature codes; the current environmental information is feature-encoded to obtain environmental feature codes; and multi-source decision information is feature-encoded to obtain decision feature codes. In one embodiment of this application, the decision feature codes include: traffic constraint codes, planned trajectory codes, execution action codes, and control command codes.
[0028] Based on a self-attention mechanism module, feature fusion is performed on decision fusion features, state feature codes, and environment feature codes to obtain multi-source fusion features. In one embodiment of this application, the decision fusion features are obtained by fusing based on the attention weight distribution between decision feature codes. The process of fusing decision feature codes based on the attention weight distribution between decision feature codes includes: calculating the correlation between decision feature codes based on a self-attention mechanism, using the correlation between decision feature codes as the attention weight distribution between decision feature codes, updating each decision feature code according to the attention weight distribution between decision feature codes to obtain each updated decision feature code, and concatenating each updated decision feature code to obtain the decision fusion features. The formula for calculating the correlation between decision feature codes is shown below:
[0029] Equation (1)
[0030] in, Indicates correlation. Represents the query vector. Represents the key vector. This represents the dimension of the key vector.
[0031] In one embodiment of this application, when there is correlation between traffic constraint code and planned trajectory code, the query vector is obtained by the dot product of traffic constraint code and learnable query projection matrix, and the key vector is obtained by the dot product of planned trajectory code and learnable key projection matrix. The correlation between traffic constraint codes, the correlation between traffic constraint code and execution action code, and the correlation between traffic constraint code and control command code are all calculated according to formula (1).
[0032] Taking the calculation of the updated decision feature code as an example, the calculation formula for the updated decision feature code is as follows:
[0033] Equation (2)
[0034] in, This represents the updated decision feature encoding. This indicates the correlation between decision feature codes and decision feature codes. It is represented by the dot product of the decision feature encoding and the projected matrix of the learnable values. This indicates the correlation between decision feature encoding and planning trajectory encoding. This is represented by the dot product of the planned trajectory code and the projected matrix of the learnable values. This indicates the correlation between traffic constraint codes and action codes. This is represented by the dot product of the action encoding and the projected matrix of the learnable values. This indicates the correlation between traffic constraint codes and control instruction codes. This is represented by the dot product of the control instruction encoding and the projected matrix of the learnable value.
[0035] In one embodiment of this application, the calculation process of the updated planning trajectory code, the calculation process of the updated execution action code, and the calculation process of the updated control command code are all the same as the calculation process of the updated decision feature code.
[0036] The classification decision is performed on the multi-source fusion features to obtain the autonomous driving arbitration result. In one embodiment of this application, the process of classifying the multi-source fusion features is implemented through a decision output head and a weight output head. The multi-source fusion features are input into the decision output head to obtain the autonomous driving execution action or autonomous driving execution trajectory, and the multi-source fusion features are input into the weight output head to obtain the visualized attention weight distribution of the multi-source decision information.
[0037] In one embodiment of this application, the process of pre-constructing an arbitration network model includes:
[0038] The system acquires historical operational scenario data and conflict scenario simulation data for unmanned logistics vehicles. In one embodiment of this application, both the historical operational scenario data and the conflict scenario simulation data cover scenarios such as normal driving, rule conflicts, and remote intervention, which is beneficial for scenario diversification.
[0039] Historical operational scenario data and conflict scenario simulation data are used as sample data. The expected weight distribution of decision information from different input sources in the sample data is labeled, and the expected execution result of autonomous driving for each frame of data in each scenario is labeled, resulting in sample data with labeled information. In one embodiment of this application, the number of sample data reaches 1 million. The process of labeling the expected execution result of autonomous driving for each frame of data in each scenario is achieved through manual labeling or simulation optimization. The method of labeling the expected weight distribution of decision information from different input sources in the sample data is based on expert experience or through simulation analysis.
[0040] The process involves perturbing the labeled sample data to obtain enhanced sample data, and then diversifying the decisions based on the enhanced sample data to obtain sample data with multimodal decision-making annotations. In one embodiment of this application, the process of perturbing the labeled sample data is implemented based on simulation perturbation (changing illumination parameters, weather parameters, traffic density parameters, etc.) or causal logic perturbation, and the process of diversifying the decisions based on the enhanced sample data is implemented through a pre-trained diffusion model or a pre-trained exploration-utilization balancer.
[0041] A pre-trained arbitration network model is obtained by supervising the training of a pre-set arbitration network model using sample data with multimodal decision annotation information. In one embodiment of this application, the process of supervising the training of the pre-set arbitration network model using sample data with multimodal decision annotation information uses an optimizer (AdamW) to minimize the difference between the autonomous driving arbitration prediction result and the autonomous driving expected execution result, as well as the difference between the expected weight distribution and the predicted weight distribution. A cosine decomposition algorithm is used to achieve smooth convergence of the parameters in the pre-set arbitration network model. The batch size during training is set to 64, and the total number of training rounds is set to 100 to 200 cycles until the loss function converges, thereby outputting the pre-trained arbitration network model.
[0042] Conflict scenario data from sample data with multimodal decision annotation information is selected, and a pre-trained arbitration network model is iteratively trained to obtain a pre-constructed arbitration network model. In one embodiment of this application, strengthening the pre-trained arbitration network model with conflict scenario data is beneficial to improving the pre-constructed arbitration network model's ability to handle multi-source decision conflict scenarios, ensuring that the pre-constructed arbitration network model can still output stable, safe, and traffic-compliant autonomous driving arbitration results under complex conflict scenarios. During the iterative training process, the PPO (Proximal Policy Optimization) algorithm or the SAC (Soft Actor-Critic) algorithm is used to construct a composite reward function that includes safety constraints, traffic efficiency rewards, and ride comfort penalties. This function optimizes and adjusts the network parameters in the pre-trained arbitration network model, enabling the pre-trained arbitration network model to output stable, safe, and traffic-compliant autonomous driving arbitration results while maximizing cumulative rewards, thereby significantly improving the overall decision-making performance and robustness of the pre-trained arbitration network model.
[0043] In one embodiment of this application, the process of supervised training of a pre-defined arbitration network model using sample data with multimodal decision annotation information includes:
[0044] The process extracts first historical temporal state information, first historical environment information, and first multi-source historical decision information from sample data with multimodal decision annotation information. In one embodiment of this application, the information type in the first historical temporal state information is the same as the information type of the temporal state information, and the latest timestamp of the first historical temporal state information is located before the latest timestamp of the temporal state information. The information type of the first historical environment information is the same as the information type of the current environment information, and the latest timestamp of the first historical environment information is located before the timestamp of the current environment information. The information type of the multi-source decision information is the same as the information type of the first multi-source historical decision information, and the latest timestamp of the first multi-source historical decision information is located before the timestamp of the multi-source decision information. The process of extracting the first historical temporal state information is implemented through a temporal-aware word segmenter, the process of extracting the first historical environment information is implemented through a hierarchical visual backbone network model, and the process of extracting the first multi-source historical decision information is implemented through an action / trajectory encoder.
[0045] Based on a pre-defined arbitration network model, arbitration is performed on the first historical time-series state information, the first historical environment information, and the first multi-source historical decision information to obtain the autonomous driving arbitration prediction result and the prediction weight distribution of different input decision information in the sample data with multi-modal decision annotation information. In one embodiment of this application, the process of arbitrating the first historical time-series state information, the first historical environment information, and the first multi-source historical decision information using the pre-defined arbitration network model includes: the pre-defined arbitration network model performs feature encoding on the first historical time-series state information, the first historical environment information, and the first multi-source historical decision information respectively; then, the feature-encoded information is fused based on a self-attention mechanism; and finally, the fused information is input into the decision output head and the weight output head respectively to obtain the corresponding output result.
[0046] With the objective of minimizing the difference between the autonomous driving arbitration prediction result and the expected autonomous driving execution result, as well as the sum of the differences between the predicted weight distribution and the expected weight distribution, the parameters in the preset arbitration network model are adjusted to obtain a pre-trained arbitration network model. In one embodiment of this application, if the difference between the autonomous driving arbitration prediction result and the expected autonomous driving execution result, as well as the sum of the differences between the predicted weight distribution and the expected weight distribution, are represented by a loss function, the expression of the loss function is as follows:
[0047] Equation (3)
[0048] in, Represents the loss function. Indicates the loss weight of the decision item. Indicates the loss of the decision item. This indicates the predicted results of the autonomous driving arbitration. This indicates the expected execution result of autonomous driving. Represents the loss weight of the distribution term. Represents the loss of the distribution term. Indicates the predicted weight distribution. This represents the expected weight distribution, where the sum of the loss weight of the decision term and the loss weight of the distribution term equals 1.
[0049] The expression for the decision term loss is as follows:
[0050] Equation (4)
[0051] in, Indicates the loss of the decision item. This indicates the predicted results of the autonomous driving arbitration. This indicates the expected execution result of autonomous driving. Indicates the number of samples. The first part of the expected execution result of autonomous driving One expected result Indicating the first in the autonomous driving arbitration prediction results One prediction result, This represents the smoothed L1 loss function;
[0052] The formula for calculating the loss of the distribution term is as follows:
[0053] Equation (5)
[0054] in, Represents the loss of the distribution term. Indicates the predicted weight distribution. Indicates the expected weight distribution. This indicates the number of categories of input decision information. Indicates the first The expected weights of each category of input decision information. Indicates the first The prediction weights for each category of input decision information. This represents the smoothing term, with a value of 0.0001.
[0055] In one embodiment of this application, the process of arbitrating first historical time-series state information, first historical environmental information, and first multi-source historical decision information based on a preset arbitration network model includes:
[0056] The first historical time-series state information is feature-encoded to obtain a first state feature code; the first historical environmental information is feature-encoded to obtain a first environmental feature code; and the first multi-source historical decision information is feature-encoded to obtain a first decision feature code. In one embodiment of this application, the first decision feature code includes: a first historical traffic constraint code, a first historical planned trajectory code, a first historical executed action code, and a first historical control command code. The state feature encoder performs feature encoding on the first historical time-series state information to obtain the first state feature code; the environmental feature encoder performs feature encoding on the first historical environmental information to obtain the first environmental feature code; the first multi-source historical decision information includes: a first historical traffic constraint, a first historical planned trajectory, a first historical executed action, and a first historical manual control command. The traffic constraint encoder encodes the first historical traffic constraint to obtain the first historical traffic constraint code; the planned trajectory encoder encodes the first historical planned trajectory to obtain the first historical planned trajectory code; the executed action encoder encodes the first historical executed action to obtain the first historical executed action code; and the control command encoder encodes the first historical manual control command to obtain the first historical control command code.
[0057] A self-attention mechanism module based on preset parameters performs feature fusion on a first decision fusion feature, a first state feature encoding, and a first environment feature encoding to obtain a first multi-source fusion feature. In one embodiment of this application, the first decision fusion feature is obtained by fusing based on the attention weight distribution among the first decision feature encodings; the self-attention mechanism module with preset parameters is obtained after parameter adjustment. The process of fusing the first decision fusion feature, the first state feature encoding, and the first environment feature encoding based on the self-attention mechanism module with preset parameters to obtain the first multi-source fusion feature is the same as the process of fusing the decision fusion feature, the state feature encoding, and the environment feature encoding based on the self-attention mechanism module to obtain the multi-source fusion feature. The process of fusing based on the attention weight distribution among the first decision feature encodings to obtain the first decision fusion feature is the same as the process of fusing based on the attention weight distribution among the decision feature encodings to obtain the decision fusion feature.
[0058] A classification decision is made on the first multi-source fusion feature to obtain the autonomous driving arbitration prediction result and the prediction weight distribution. In one embodiment of this application, the process of making a classification decision on the first multi-source fusion feature to obtain the autonomous driving arbitration prediction result and the prediction weight distribution is the same as the process of making a classification decision on the multi-source fusion feature to obtain the autonomous driving arbitration result.
[0059] In one embodiment of this application, the process of iteratively training a pre-trained arbitration network model by selecting conflict scenario data from sample data with multimodal decision annotation information includes:
[0060] The second historical temporal state information, second historical environment information, and second multi-source historical decision information are extracted from the conflict scenario data. In one embodiment of this application, the information type in the second historical temporal state information is the same as that in the first historical temporal state information, and the latest timestamp of the second historical temporal state information is located before the latest timestamp of the temporal state information. The information type in the second historical environment information is the same as that in the first historical environment information, and the latest timestamp of the second historical environment information is located before the timestamp of the current environment information. The information type in the second multi-source historical decision information is the same as that in the first multi-source historical decision information, and the latest timestamp of the second multi-source historical decision information is located before the timestamp of the multi-source decision information. The process of extracting the second historical temporal state information is implemented through a temporal-aware word segmenter, the process of extracting the second historical environment information is implemented through a hierarchical visual backbone network model, and the process of extracting the second multi-source historical decision information is implemented through an action / trajectory encoder.
[0061] Based on a pre-trained arbitration network model, the second historical time-series state information, the second historical environment information, and the second multi-source historical decision information are arbitrated to obtain the predicted trajectory and predicted state of the unmanned logistics vehicle, as well as the predicted trajectory and predicted state of the moving target. In one embodiment of this application, the process of arbitrating the second historical time-series state information, the second historical environment information, and the second multi-source historical decision information based on the pre-trained arbitration network model includes: the pre-trained arbitration network model performs feature encoding on the second historical time-series state information, the second historical environment information, and the second multi-source historical decision information respectively to obtain encoded information; the encoded information is fused based on a self-attention mechanism to obtain fused information; the fused information is input into a trajectory generation head to obtain the predicted trajectory of the unmanned logistics vehicle and the predicted trajectory of the moving target; the fused information is input into a state regression head to obtain the predicted state of the unmanned logistics vehicle and the predicted state of the moving target.
[0062] Based on the predicted trajectories of the unmanned logistics vehicle and the moving target, determine the driving distance and collision situation between the unmanned logistics vehicle and the moving target, the human intervention status and violation status of the unmanned logistics vehicle, and the distance between the unmanned logistics vehicle and the destination; and determine the predicted speed, predicted acceleration, and predicted angular velocity of the unmanned logistics vehicle from the predicted status of the unmanned logistics vehicle. In one embodiment of this application, the predicted trajectory of the unmanned logistics vehicle includes the predicted position of the unmanned logistics vehicle at different time steps, the predicted trajectory of the moving target includes the predicted position of the moving target at different time steps, and the travel distance is the straight-line distance between the predicted position of the moving target and the predicted position of the unmanned logistics vehicle at the same time step. After obtaining the straight-line distance between the predicted position of the moving target and the predicted position of the unmanned logistics vehicle at the same time step, the straight-line distance between the predicted position of the moving target and the predicted position of the unmanned logistics vehicle at the same time step is compared with a preset distance threshold. If the straight-line distance between the predicted position of the moving target and the predicted position of the unmanned logistics vehicle at the same time step is less than or equal to the preset distance threshold, it is determined that the moving target and the unmanned logistics vehicle have collided at that moment. If the straight-line distance between the predicted position of the moving target and the predicted position of the unmanned logistics vehicle at the same time step is greater than the preset distance threshold, it is determined that the moving target and the unmanned logistics vehicle have not collided at that moment. The predicted state of the unmanned logistics vehicle includes: the predicted speed, predicted acceleration, heading angle, and steering angle of the unmanned logistics vehicle at different time steps. The predicted angular velocity of the unmanned logistics vehicle is the ratio of the difference in heading angle between adjacent time steps to that per unit time step. Since the unmanned logistics vehicle will deviate laterally from the reference path when manually taken over in extreme road conditions or in the event of a malfunction, the manual takeover state of the unmanned logistics vehicle is determined by comparing the lateral deviation of the unmanned logistics vehicle from the reference path (e.g., the center line of the lane) at different time steps with a preset deviation threshold. When the duration of the lateral deviation being greater than the preset deviation threshold is greater than a preset duration threshold, the unmanned logistics vehicle is determined to be in a manual takeover state at that time step. The process of determining violations by unmanned logistics vehicles includes: mapping the predicted trajectory of the unmanned logistics vehicle onto an electronic map containing lane attributes and traffic facility information; detecting whether the predicted trajectory overlaps with prohibited areas (such as solid lines and pedestrian crossings) on the map, and recording a violation if such overlap exists; comparing the predicted speed values of the trajectory at different time steps with the speed limit range of the road segment, and recording a violation if the predicted speed value is outside the speed limit range; and counting all violations. The distance between the unmanned logistics vehicle and its destination is determined based on the distance between the predicted position of the unmanned logistics vehicle at different time steps and the destination position.
[0063] The safety bonus for the unmanned logistics vehicle is determined based on the driving distance and collision situation. In one embodiment of this application, the formula for calculating the safety bonus is as follows:
[0064] Equation (6)
[0065] in, Indicates a security reward. Indicates the weight of the distance penalty term. This represents the total number of time steps within the prediction time domain. express Distance penalty function for time steps, Indicates the weight of the collision penalty term. express The collision indicator function at each time step is set to 1 when a collision occurs and 0 when no collision occurs. express The driving distance at each time step, the weight of the distance penalty term, and the weight of the collision penalty term can be set according to the actual situation. For example, the weight of the distance penalty term can be set to a small value, and the weight of the collision penalty term can be set to a large value. The weight of the collision penalty term is a multiple of the weight of the distance penalty term. In equation (6), when no collision occurs, the collision indicator function term is 0, and the safety reward is determined according to the distance penalty function. That is, the smaller the driving distance, the greater the safety reward. When a collision occurs, the collision indicator function term is 1, and the product of the collision indicator function and the weight of the collision penalty term is the main contributor to the safety reward.
[0066] The expression for the distance penalty function at each time step is as follows:
[0067] Equation (7)
[0068] in, express Distance penalty function for time steps, express The distance traveled in time step, Indicates the physical collision threshold. Indicates the radius of safety impact. This represents the collision penalty coefficient.
[0069] The driving efficiency bonus of the unmanned logistics vehicle is determined based on its predicted speed and acceleration, as well as the distance between the vehicle and its destination. In one embodiment of this application, the formula for calculating the driving efficiency bonus includes:
[0070] Equation (8)
[0071] in, This indicates a reward for improved driving efficiency. This represents the total number of time steps within the prediction time domain. Indicates the reward weight of the speed item. This indicates that unmanned logistics vehicles are in Prediction speed of time step This indicates the speed limit for unmanned logistics vehicles. Indicates the reward weight based on the distance to the finish line. Indicates in The time step measures the distance between the unmanned logistics vehicle and its destination. Indicates in The time step measures the distance between the unmanned logistics vehicle and its destination. Indicates a unit of time. This indicates the penalty weight for the acceleration term. This indicates that unmanned logistics vehicles are in Predicted acceleration at time step This represents the maximum acceleration of the unmanned logistics vehicle. The reward weights for speed, destination distance, and acceleration can all be set according to the actual situation. Equation (8) is used to characterize the fact that the greater the predicted speed, the greater the difference between the distance between the unmanned logistics vehicle and the destination in the previous time step and the distance between the unmanned logistics vehicle and the destination in the current time step, and the smaller the predicted acceleration, the greater the reward obtained. This encourages rapid and effective displacement during the driving process and suppresses rapid acceleration and deceleration, which is conducive to ensuring the smoothness of the driving trajectory while driving quickly and effectively.
[0072] The trajectory comfort reward of the unmanned logistics vehicle is determined based on its predicted acceleration and predicted angular velocity. In one embodiment of this application, the formula for calculating the trajectory comfort reward is as follows:
[0073] Equation (9)
[0074] in, Indicates a reward for comfort during the trajectory. This represents the total number of time steps within the prediction time domain. Indicates the weight of the accelerometer term. This indicates that unmanned logistics vehicles are in Predicted acceleration at time step This indicates that unmanned logistics vehicles are in Predicted acceleration at time step This indicates the maximum acceleration of the unmanned logistics vehicle. Indicates a unit of time. Indicates the weight of the angular velocity term. This indicates that unmanned logistics vehicles are in Predicted angular velocity at the time step This represents the maximum angular velocity of the unmanned logistics vehicle. The weights of the jerk and angular velocity terms are set according to the actual situation. Formula (9) is used to characterize that when the predicted jerk and angular velocity are large, the penalty is greater, which is beneficial to suppressing large predicted jerk and increasing angular velocity, and ensuring the comfort of the trajectory.
[0075] The remote control reward for the unmanned logistics vehicle is determined based on the status of manual takeover. In one embodiment of this application, the calculation formula for the remote control reward is as follows:
[0076] Equation (10)
[0077] in, Indicates a reward for remote control. This represents the total number of time steps within the prediction time domain. Indicates the weight of the manually managed items. This represents the manual takeover indicator function. When the unmanned logistics vehicle is in a manually controlled state, the manual takeover indicator function is 1; when the unmanned logistics vehicle is not in a manually controlled state, the manual takeover indicator function is 0. Indicates the weight of the switching delay time item. express The switching delay time when the time step is in manual takeover mode is collected by the arbitrator of the underlying control authority. This represents the maximum switching delay time. The weights of the manual takeover item and the switching delay time item are set according to the actual situation. The arbitrator of the underlying control is a core software module running on the chassis domain controller or central computing platform, used to determine and decide in real time who has the final control of the vehicle among the instructions of the autonomous driving system, remote manual driving, and the in-vehicle safety driver. Equation (10) is used to characterize that when the current state is manual takeover and the switching delay time is larger, the penalty received is greater, thereby achieving the effect of encouraging autonomous driving.
[0078] Based on the nature of the violation, a rule compliance reward for the unmanned logistics vehicle is determined. In one embodiment of this application, the formula for calculating the rule compliance reward is as follows:
[0079] Equation (11)
[0080] in, This indicates that following the rules will result in a reward. This represents the total number of time steps within the prediction time domain. Indicates the weight of the violation item. express The violation indication function for the time step, when the unmanned logistics vehicle is in If a time step violation occurs, then The violation indication function for the time step is 1, when the unmanned logistics vehicle is in If no violation occurs during the time step, then The violation indication function for each time step is 0. The weight of each violation is set according to the actual situation. Equation (11) is used to characterize that the more violations occur, the heavier the penalty, thereby encouraging driving in accordance with the rules.
[0081] The safety reward, driving efficiency reward, trajectory comfort reward, remote control reward, and rule compliance reward are weighted and summed to form a multi-objective reward. The parameters of the pre-trained arbitration network model are adjusted to maximize this multi-objective reward, resulting in a pre-constructed arbitration network model. In one embodiment of this application, the calculation formula for the multi-objective reward is as follows:
[0082] Equation (12)
[0083] in, Indicates multi-objective rewards. Indicates the security reward weight. Indicates a security reward. Indicates the weight of driving efficiency rewards. This indicates a reward for improved driving efficiency. Indicates the weight of trajectory comfort reward. Indicates a reward for comfort during the trajectory. This indicates the reward weight for following the rules. This indicates that following the rules will be rewarded. Indicates the weight of remote control rewards. The weights for remote control rewards, safety rewards, driving efficiency rewards, trajectory comfort rewards, rule compliance rewards, and remote control rewards are all set according to actual conditions.
[0084] In one embodiment of this application, the process of arbitrating the second historical time-series state information, the second historical environment information, and the second multi-source historical decision information based on a pre-trained arbitration network model includes:
[0085] The second historical time-series state information is feature-encoded to obtain a second state feature code; the second historical environmental information is feature-encoded to obtain a second environmental feature code; and the second multi-source historical decision information is feature-encoded to obtain a second decision feature code. In one embodiment of this application, the second decision feature code includes: a second historical traffic constraint code, a second historical planned trajectory code, a second historical executed action code, and a second historical control command code; the state feature encoder performs feature encoding on the second historical time-series state information to obtain the second state feature code; the environmental feature encoder performs feature encoding on the second historical environmental information to obtain the second environmental feature code; the first multi-source historical decision information includes: second historical traffic constraints, second historical planned trajectories, second historical executed actions, and second historical manual control commands; the traffic constraint encoder encodes the second historical traffic constraints to obtain the second historical traffic constraint code; the planned trajectory encoder encodes the second historical planned trajectory to obtain the second historical planned trajectory code; the executed action encoder encodes the second historical executed actions to obtain the second historical executed action code; and the control command encoder encodes the second historical manual control commands to obtain the second historical control command code.
[0086] Based on the self-attention mechanism module, feature fusion is performed on the second decision fusion feature, the second state feature encoding, and the second environment feature encoding to obtain the second multi-source fusion feature. In one embodiment of this application, the second decision fusion feature is obtained by fusing based on the attention weight distribution among the second decision feature encodings; the process of obtaining the second decision fusion feature based on the attention weight distribution among the second decision feature encodings is the same as the process of obtaining the decision fusion feature based on the attention weight distribution among the decision feature encodings. The process of feature fusion based on the self-attention mechanism module on the second decision fusion feature, the second state feature encoding, and the second environment feature encoding is the same as the process of feature fusion based on the self-attention mechanism module on the decision fusion feature, the state feature encoding, and the environment feature encoding to obtain the multi-source fusion feature.
[0087] The second multi-source fusion feature is used to make classification decisions to obtain the predicted trajectory and predicted state of the unmanned logistics vehicle, as well as the predicted trajectory and predicted state of the moving target. In one embodiment of this application, the process of making classification decisions on the second multi-source fusion feature includes: inputting the second multi-source fusion feature into a trajectory generation head to obtain the predicted trajectory of the unmanned logistics vehicle and the predicted trajectory of the moving target; and inputting the second multi-source fusion feature into a state regression head to obtain the predicted state of the unmanned logistics vehicle and the predicted state of the moving target.
[0088] In one embodiment of this application, the process of fusing decision fusion features, state feature encoding, and environment feature encoding based on a self-attention mechanism module to obtain multi-source fusion features includes:
[0089] Based on the self-attention mechanism module, the attention weight distribution among decision fusion features, state feature codes, and environment feature codes is calculated. In one embodiment of this application, the attention weight distribution among decision fusion features, state feature codes, and environment feature codes includes: the correlation between decision fusion features, the correlation between decision fusion features and state feature codes, the correlation between decision fusion features and environment feature codes, the correlation between state feature codes and decision fusion features, the correlation between state feature codes and environment feature codes, the correlation between environment feature codes and decision fusion features, and the correlation between environment feature codes and state feature codes. The correlation calculation formula is shown in Equation (1).
[0090] Based on the attention weight distribution, the decision fusion features are updated to obtain decision update features; based on the attention weight distribution, the state feature encoding is updated to obtain state update encoding; based on the attention weight distribution, the environment feature encoding is updated to obtain environment update encoding. In one embodiment of this application, the calculation formula for the decision update features is as follows:
[0091] Equation (13)
[0092] in, This indicates the decision update feature. This indicates the correlation between decision fusion features and decision fusion features. Indicates the characteristics of decision fusion. This indicates the correlation between decision fusion features and state feature encoding. Represents state feature encoding, This indicates the correlation between decision fusion features and environmental feature encoding. This represents the encoding of environmental characteristics.
[0093] In one embodiment of this application, both the state update code and the environment update code can be calculated according to the form of formula (13).
[0094] A residual connection is performed on the decision update feature and the decision fusion feature to obtain the first decision connection feature; a residual connection is performed on the state update code and the state feature code to obtain the first state connection feature; and a residual connection is performed on the environment update code and the environment feature code to obtain the first environment connection feature. In one embodiment of this application, the residual connection helps to retain key information in the decision fusion feature, the state feature code, and the environment feature code, avoiding the problem of gradient vanishing or loss of effective information in the multi-layer propagation of the decision fusion feature, the state feature code, and the environment feature code, thus ensuring the stability and integrity of feature transmission.
[0095] The first decision connection feature is input into the feedforward network module to obtain the decision feedforward output feature; the first state connection feature is input into the feedforward network module to obtain the state feedforward output feature; and the first environment connection feature is input into the feedforward network module to obtain the environment feedforward output feature. In one embodiment of this application, the feedforward network module is composed of two linear transformation layers and one nonlinear activation layer stacked together, which can further perform spatial transformation and feature purification on the first decision connection feature, the first state connection feature, and the first environment connection feature, thereby enhancing the expressive power of different types of features.
[0096] A second decision connection feature is obtained by performing a residual connection on the first decision connection feature and the decision feedforward output feature; a second state connection feature is obtained by performing a residual connection on the first state connection feature and the state feedforward output feature; and a second environment connection feature is obtained by performing a residual connection on the first environment connection feature and the environment feedforward output feature. In one embodiment of this application, residual connections help to further retain the effective information in the first decision connection feature, the first state connection feature, and the first environment connection feature, reducing information loss during feature transformation, thereby reducing the arbitration capability of the pre-built arbitration network model.
[0097] Layered normalization is applied to the second decision connection features, the second state connection features, and the second environment connection features. These layered normalized decision connection features, state connection features, and environment connection features are then concatenated to obtain multi-source fusion features. In one embodiment of this application, layered normalization helps reduce the impact of differences in feature distributions on the pre-built arbitration network model, improving the stability of the inference inference of the pre-built arbitration network model. By concatenating the normalized features, the differential information of different types of features is preserved, while integrating multiple types of features into a unified feature output, providing complete feature input for subsequent classification decisions. This avoids the loss of decision information due to the dispersion of output features from different modules, ensuring the accuracy of the final arbitration result. The self-attention mechanism captures the potential correlation between different types of features, fully mining hidden features useful for the current decision from historical information. Compared to traditional simple concatenation and fusion, this method can more accurately achieve multimodal information fusion, improving the adaptability of the pre-built arbitration network model to complex scenarios.
[0098] Figure 3 This is a schematic diagram illustrating the structure of a pre-built arbitration network model, as shown in an exemplary embodiment of this application. Figure 3 In the pre-built arbitration network model, there are three layers: an input layer, a feature fusion layer, and an output layer. The input layer includes a state feature encoder, an environmental feature encoder, a traffic constraint encoder, a planned trajectory encoder, an executed action encoder, and a control command encoder. The output layer includes a decision output head and a weight output head. The decision output head includes a trajectory generation head, a behavior classification head, and a state regression head.
[0099] In one embodiment of this application, a state feature encoder is used to encode the temporal state information to obtain a state feature code; an environmental feature encoder is used to encode the current environmental information to obtain an environmental feature code; a traffic constraint encoder is used to encode the current traffic constraints to obtain a traffic constraint code; a planned trajectory encoder is used to encode the current planned trajectory to obtain a planned trajectory code; an execution action encoder is used to encode the current execution action to obtain an execution action code; and a control command encoder is used to encode the current manual control command to obtain a control command code. The feature fusion layer employs a 4-layer Transformer Encoder. It distinguishes decision information from different sources through position encoding, calculates the attention weight distribution between decision feature codes through a self-attention mechanism, and fuses the decision feature codes according to the attention weight distribution. Furthermore, it calculates the attention distribution between the decision fusion features, state feature codes, and environmental feature codes through a self-attention mechanism, and fuses the decision fusion features, state feature codes, and environmental feature codes according to the attention distribution between them. The weight output head is used to output a visualized attention weight distribution of multi-source decision information; the trajectory generation head is used to output the autonomous driving trajectory; the behavior classification head is used to output the autonomous driving execution actions; and the state regression head is used to output the driving state (e.g., speed, acceleration, heading angle, etc.).
[0100] Figure 4 This is a flowchart illustrating a pre-built arbitration network model, as shown in an exemplary embodiment of this application. Figure 4 In the process of pre-constructing the arbitration network model, the following steps are taken: (1) Obtaining historical operation scenario data and conflict scenario simulation data of unmanned logistics vehicles; (2) Using historical operation scenario data and conflict scenario simulation data as sample data; labeling the expected weight distribution of decision information from different input sources in the sample data, labeling the expected execution result of autonomous driving in each frame of data in each scenario in the sample data, and obtaining sample data with labeled information; (3) Perturbing the sample data with labeled information to obtain enhanced sample data; and diversifying the decision on the enhanced sample data to obtain sample data with multimodal decision labeling information; (4) Supervising the training of the pre-constructed arbitration network model using the sample data with multimodal decision labeling information to obtain the pre-trained arbitration network model; (5) Selecting conflict scenario data from the sample data with multimodal decision labeling information, iteratively training the pre-trained arbitration network model to obtain the pre-constructed arbitration network model.
[0101] Figure 5 This is a schematic diagram illustrating the attention weight distribution of multi-source decision information in an exemplary embodiment of this application, as shown below. Figure 5As shown, the multi-source decision information includes: current traffic constraints, current planned trajectory, current action to be executed, and current manual control command. The weights of the current traffic constraints, current planned trajectory, current action to be executed, and current manual control command are 0.0. Figure 5 The attention weight distribution of multi-source decision information can be clearly seen, which increases the interpretability of the output results of the pre-built arbitration network model. This is conducive to researchers intuitively assessing the dependence of the pre-built arbitration network model on decision information from different sources. When decision arbitration errors occur, the causes of model decision bias can be quickly located, and it can be quickly identified which source of decision information has caused excessive interference to the model. This facilitates targeted adjustments to the model structure or training data, further improving the reliability of the arbitration results.
[0102] The arbitration method used in this application has the following advantages compared to the traditional static priority arbitration method:
[0103] Table 1
[0104] flexibility Difference good Significant improvement Scene adaptability limited powerful Significant improvement Decision quality generally high Increase by 40% Explainability generally powerful Significant improvement
[0105] As shown in Table 1, the arbitration method proposed in this application significantly improves upon the traditional static priority arbitration method in four dimensions: flexibility, scenario adaptability, decision quality, and interpretability. This application achieves dynamic fusion of multimodal information through a pre-constructed arbitration network model, which can automatically adjust the attention weights of each decision information according to different scenarios. It is more flexible and can adapt to various complex urban scenarios, mountain road scenarios, and special operation scenarios, greatly enhancing its adaptability to complex scenarios. The decision quality is improved by 40% compared to traditional methods. At the same time, it can also output a visualized attention weight distribution, significantly enhancing the interpretability of the arbitration results. It can better meet the decision arbitration needs of autonomous driving of unmanned logistics vehicles in multiple scenarios, improve the safety and stability of autonomous driving, and improve the problems of poor flexibility, poor scenario adaptability, and low decision quality and interpretability of traditional static priority methods.
[0106] The attention-based fusion method in this application has the following advantages compared to the simple weighted fusion method:
[0107] Table 2
[0108] Weight dynamism fixed dynamic Significant improvement Hard constraint handling improper priority Significant improvement Fusion quality generally high Increase by 35% Smoothness Difference good Increase by 50%
[0109] As shown in Table 2, the attention-based fusion method in this application significantly improves upon the simple weighted fusion method in four dimensions: weight dynamism, hard constraint handling, fusion quality, and smoothness. This application, through its self-attention mechanism, can automatically allocate dynamic weights based on the correlation of different input features, rather than using fixed weights in the simple weighted method. This results in stronger dynamic weight adaptation capabilities. When faced with scenarios involving hard constraints such as traffic rules and operational requirements, it can automatically increase the attention weight of constraint information, ensuring that hard constraints are prioritized. This solves the problem that the simple weighted method cannot effectively handle hard constraint requirements. The fusion quality is improved by 35% compared to simple weighted fusion, and the smoothness of the decision output is improved by 50%, enabling the output of more stable and continuous arbitration decisions. This avoids the problem of large fluctuations in the output of simple weighted fusion, further improving the comfort and safety during autonomous driving and better adapting to the decision arbitration needs of multimodal fusion.
[0110] The advantages of the decision arbitration method compared to the end-to-end learning method in this application are as follows:
[0111] Table 3
[0112] Explainability Difference good Significant improvement Security Guarantee weak powerful Significant improvement Human-machine collaboration difficulty smooth Significant improvement Debugging difficulty high Low Significantly reduced
[0113] As shown in Table 3, the decision arbitration method in this application has significant advantages over the end-to-end learning method in four dimensions: interpretability, security assurance, human-machine collaboration, and debugging difficulty. The modular hierarchical decision architecture proposed in this application has clear functional boundaries for each module and can also output a visualized attention weight distribution, which is more interpretable than the black-box decision process of the end-to-end method. This application relies on multi-source decision fusion to achieve final arbitration, and can ensure that autonomous driving decisions meet safety requirements from the mechanism level through rule pre-screening, hard constraint weight reinforcement, etc., with stronger safety fallback capability. When facing special scenarios that require human intervention, this application can naturally integrate human control commands as part of the multi-source decision into the arbitration process, realizing a smooth human-machine collaboration switch and solving the problem that the end-to-end learning method is difficult to adapt to human-machine collaboration needs. At the same time, due to the clear modular structure, the problematic module can be quickly located after an error, and the debugging difficulty is significantly reduced compared with the end-to-end method. It is more suitable for the development of autonomous driving systems for mass production and can better meet the decision-making needs of specific commercial autonomous driving scenarios such as unmanned logistics vehicles, further improving the safety and reliability of autonomous driving.
[0114] This application introduces a lightweight arbitration network to dynamically integrate information from different decision sources, outputting a unified, safe, smooth, and interpretable final decision. This effectively solves the multi-source decision conflict problem in L4 autonomous driving systems, improves the system's intelligence and safety, reduces traffic accident rates, alleviates public concerns about autonomous driving, and promotes technology popularization. It is particularly suitable for the "human-machine collaboration" operation mode of unmanned logistics vehicles and is applicable to various operation scenarios.
[0115] The following describes an embodiment of the apparatus described in this application, which can be used to execute the VLA-based multimodal fusion autonomous driving decision arbitration system described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the VLA-based multimodal fusion autonomous driving decision arbitration method described above in this application.
[0116] Figure 6 This is a block diagram illustrating a multimodal fusion autonomous driving decision arbitration system based on a VLA model, as shown in an exemplary embodiment of this application.
[0117] like Figure 6 As shown, the exemplary VLA-based multimodal fusion autonomous driving decision arbitration system 600 includes:
[0118] The information acquisition module 601 is used to acquire the time-series status information, current environmental information and multi-source decision information of the unmanned logistics vehicle.
[0119] The information arbitration module 602 is used to arbitrate time-series state information, current environment information and multi-source decision information based on a pre-built arbitration network model to obtain the autonomous driving arbitration result of the unmanned logistics vehicle.
[0120] The process of arbitrating temporal state information, current environmental information, and multi-source decision information includes:
[0121] The temporal state information is feature-encoded to obtain the state feature code; the current environment information is feature-encoded to obtain the environment feature code; and the multi-source decision information is feature-encoded to obtain the decision feature code. The decision feature code includes: traffic constraint code, planning trajectory code, execution action code, and control command code.
[0122] Based on the self-attention mechanism module, feature fusion is performed on decision fusion features, state feature encoding, and environment feature encoding to obtain multi-source fusion features; the decision fusion features are obtained by fusing the attention weight distribution among the decision feature encodings.
[0123] The multi-source fusion features are classified and used to make decisions, resulting in the arbitration result for autonomous driving.
[0124] In one embodiment of this application, the multi-source decision information includes: current traffic constraints, current planned trajectory, current executed action, and current manual control command. The current traffic constraints are generated based on temporal state information and current environmental information, and are implemented through a rule / safety module. This rule / safety module can be a pre-trained deep learning network model with functions such as traffic light constraint detection, lane keeping constraint detection, speed limit and right-of-way constraint detection, collision detection, dynamic feasibility module, and right-of-way game theory. The current executed action is generated by the VLA model with temporal state information, current environmental information, and moving target state information as input. The current planned trajectory is generated by the motion planning module with temporal state information, current environmental information, and moving target state information as input. The motion planning module uses MPC (Model Predictive Control) algorithms, convex optimization algorithms, etc., to generate the current planned trajectory with temporal state information, current environmental information, and moving target state information as input. The current manual control command is obtained from the cloud-based remote cockpit and human-machine interface of the unmanned logistics vehicle. The temporal status information includes: the status information at the current moment and the status information within a preset time period prior to the current moment. The status information includes: driving position, driving speed, driving acceleration, driving angular velocity, etc. The driving position is obtained through global navigation satellite system positioning, driving speed is collected through speed sensors, driving acceleration is collected through acceleration sensors, etc., and driving angular velocity is collected through angular velocity sensors. The current environmental information includes: road indication information (e.g., lane boundary lines, lane center lines, lane driving direction) and static obstacle information (e.g., fence location information, sign location information). The current environmental information is obtained through LiDAR, cameras, millimeter-wave radar, etc.
[0125] In one embodiment of this application, the autonomous driving arbitration result includes: autonomous driving trajectory, autonomous driving execution actions, etc.; the pre-built arbitration network model includes an input layer, a feature fusion layer, and an output layer. The input layer includes: a state feature encoder, an environmental feature encoder, a traffic constraint encoder, a planned trajectory encoder, an execution action encoder, and a control command encoder. The state feature encoder is used to encode the temporal state information to obtain a state feature code; the environmental feature encoder is used to encode the current environmental information to obtain an environmental feature code; the traffic constraint encoder is used to encode the current traffic constraints to obtain a traffic constraint code; the planned trajectory encoder is used to encode the current planned trajectory to obtain a planned trajectory code; the execution action encoder is used to encode the current execution action to obtain an execution action code; and the control command encoder is used to encode the current manual control command to obtain a control command code. The feature fusion layer uses a 4-layer Transformer architecture. The encoder distinguishes decision information from different sources through positional encoding, calculates the attention weight distribution among decision feature encodings using a self-attention mechanism, and fuses the decision feature encodings based on this distribution. It also calculates the attention distribution among decision fusion features, state feature encodings, and environment feature encodings using the same self-attention mechanism, and performs feature fusion on these three types of encodings. The output layer includes a decision output head and a weight output head. When the network structure of the decision output head consists of a 512-dimensional linear layer, a 256-dimensional linear layer, and a 100-dimensional linear layer, it outputs the autonomous driving trajectory. When the network structure of the decision output head consists of a 512-dimensional linear layer, a 256-dimensional linear layer, and a Tanh activation function, it outputs the autonomous driving actions. When the network structure of the weight output head consists of a 512-dimensional linear layer, a 256-dimensional linear layer, a 100-dimensional linear layer, and a Softmax activation function, it outputs a visualized attention weight distribution of multi-source decision information.
[0126] In one embodiment of this application, the decision fusion feature is obtained by fusing the attention weight distribution between decision feature codes. The process of fusing decision feature codes based on the attention weight distribution between decision feature codes includes: calculating the correlation between decision feature codes based on a self-attention mechanism, using the correlation between decision feature codes as the attention weight distribution between decision feature codes, updating each decision feature code according to the attention weight distribution between decision feature codes to obtain each updated decision feature code, and concatenating each updated decision feature code to obtain the decision fusion feature. The formula for calculating the correlation between decision feature codes is shown in formula (1). Taking the calculation of the updated decision feature code as an example, the formula for calculating the updated decision feature code is shown in formula (2).
[0127] In one embodiment of this application, the calculation process of the updated planning trajectory code, the calculation process of the updated execution action code, and the calculation process of the updated control command code are all the same as the calculation process of the updated decision feature code.
[0128] In one embodiment of this application, the process of classifying and making decisions based on multi-source fusion features is implemented through a decision output head and a weight output head. The multi-source fusion features are input into the decision output head to obtain the autonomous driving execution action or autonomous driving execution trajectory, and the multi-source fusion features are input into the weight output head to obtain the visualized attention weight distribution of multi-source decision information.
[0129] It should be noted that the VLA-based multimodal fusion autonomous driving decision arbitration system and the VLA-based multimodal fusion autonomous driving decision arbitration method provided in the above embodiments belong to the same concept. The specific methods by which each module and unit performs its operations have been described in detail in the method embodiments and will not be repeated here. In practical applications, the VLA-based multimodal fusion autonomous driving decision arbitration system provided in the above embodiments can be configured to have different functional modules perform the functions as needed, i.e., the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. This is not a limitation here.
[0130] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A multimodal fusion decision arbitration method for autonomous driving based on a VLA model, characterized in that, The method includes: The system acquires the temporal state information, current environmental information, and multi-source decision information of the unmanned logistics vehicle. The multi-source decision information includes: current traffic constraints, current planned trajectory, current executed action, and current manual control command. The current traffic constraints are generated based on the temporal state information and the current environmental information. The current executed action is generated by a VLA model with the temporal state information, the current environmental information, and the moving target state information as input. The current planned trajectory is generated by a motion planning module with the temporal state information, the current environmental information, and the moving target state information as input. Based on a pre-built arbitration network model, the time-series state information, the current environment information, and the multi-source decision information are arbitrated to obtain the autonomous driving arbitration result of the unmanned logistics vehicle; the autonomous driving arbitration result includes: autonomous driving trajectory and autonomous driving execution actions; The process of arbitrating the temporal state information, the current environment information, and the multi-source decision information includes: The temporal state information is feature-encoded to obtain a state feature code; the current environment information is feature-encoded to obtain an environment feature code; and the multi-source decision information is feature-encoded to obtain a decision feature code; the decision feature code includes: traffic constraint code, planning trajectory code, execution action code, and control command code; Based on the self-attention mechanism module, feature fusion is performed on the decision fusion feature, the state feature encoding, and the environment feature encoding to obtain multi-source fusion features; the decision fusion feature is obtained by fusing the attention weight distribution among the decision feature encodings. The multi-source fusion features are classified and a decision is made to obtain the autonomous driving arbitration result; The process of pre-constructing the arbitration network model includes: Acquire the historical operation scenario data and conflict scenario simulation data of the unmanned logistics vehicle; The historical operation scenario data and the conflict scenario simulation data are used as sample data; the expected weight distribution of decision information from different input sources in the sample data is labeled, and the expected execution result of autonomous driving in each frame of data in each scenario in the sample data is labeled to obtain sample data with labeled information. The labeled sample data is perturbed to obtain enhanced sample data; and the enhanced sample data is then subjected to decision diversification to obtain sample data with multimodal decision labeling information. The pre-trained arbitration network model is obtained by supervising the training of the pre-set arbitration network model using sample data with multimodal decision annotation information. Conflict scenario data from the sample data with multimodal decision annotation information is selected, and the pre-trained arbitration network model is iteratively trained to obtain the pre-constructed arbitration network model. The process of selecting conflict scenario data from the sample data with multimodal decision annotation information and iteratively training the pre-trained arbitration network model includes: Extract the second historical temporal state information, the second historical environmental information, and the second multi-source historical decision information from the conflict scenario data; Based on the pre-trained arbitration network model, the second historical time-series state information, the second historical environment information, and the second multi-source historical decision information are arbitrated to obtain the predicted trajectory and predicted state of the unmanned logistics vehicle, as well as the predicted trajectory and predicted state of the moving target. Based on the predicted trajectory of the unmanned logistics vehicle and the predicted trajectory of the moving target, determine the driving distance and collision situation between the unmanned logistics vehicle and the moving target, the manual takeover status and violation status of the unmanned logistics vehicle, and the distance between the unmanned logistics vehicle and the destination; and determine the predicted speed, predicted acceleration, and predicted angular velocity of the unmanned logistics vehicle from the predicted status of the unmanned logistics vehicle. The safety bonus for the unmanned logistics vehicle is determined based on the travel distance and the collision situation. The driving efficiency bonus of the unmanned logistics vehicle is determined based on its predicted speed and predicted acceleration, as well as the distance between the unmanned logistics vehicle and the destination. The trajectory comfort reward of the unmanned logistics vehicle is determined based on the predicted acceleration and predicted angular velocity of the unmanned logistics vehicle. The remote control reward for the unmanned logistics vehicle is determined based on the state of manual takeover. Based on the aforementioned violations, a rule compliance reward will be determined for the unmanned logistics vehicle. The safety reward, driving efficiency reward, trajectory comfort reward, remote control reward, and rule compliance reward are weighted and summed to form a multi-objective reward. The parameters in the pre-trained arbitration network model are adjusted with the goal of maximizing the multi-objective reward to obtain the pre-constructed arbitration network model.
2. The multimodal fusion autonomous driving decision arbitration method based on the VLA model according to claim 1, characterized in that, The process of supervised training of the preset arbitration network model using sample data with multimodal decision annotation information includes: Extract the first historical temporal state information, the first historical environmental information, and the first multi-source historical decision information from the sample data with multimodal decision annotation information; Based on the preset arbitration network model, the first historical time-series state information, the first historical environment information, and the first multi-source historical decision information are arbitrated to obtain the autonomous driving arbitration prediction result and the prediction weight distribution of different input decision information in the sample data with multimodal decision annotation information. With the goal of minimizing the difference between the autonomous driving arbitration prediction result and the autonomous driving expected execution result, as well as the sum of the differences between the predicted weight distribution and the expected weight distribution, the parameters in the preset arbitration network model are adjusted to obtain the pre-trained arbitration network model.
3. The multimodal fusion autonomous driving decision arbitration method based on the VLA model according to claim 2, characterized in that, Based on the preset arbitration network model, the process of arbitrating the first historical time-series state information, the first historical environment information, and the first multi-source historical decision information includes: The first historical time-series state information is feature-encoded to obtain a first state feature code; the first historical environment information is feature-encoded to obtain a first environment feature code; and the first multi-source historical decision information is feature-encoded to obtain a first decision feature code; the first decision feature code includes: a first historical traffic constraint code, a first historical planning trajectory code, a first historical execution action code, and a first historical control instruction code; The self-attention mechanism module based on preset parameters performs feature fusion on the first decision fusion feature, the first state feature encoding, and the first environment feature encoding to obtain the first multi-source fusion feature; the first decision fusion feature is obtained by fusing based on the attention weight distribution among the first decision feature encodings; the self-attention mechanism module based on preset parameters is obtained after parameter adjustment. The first multi-source fusion feature is used to make a classification decision to obtain the autonomous driving arbitration prediction result and the prediction weight distribution.
4. The multimodal fusion autonomous driving decision arbitration method based on the VLA model according to claim 2, characterized in that, If the sum of the differences between the autonomous driving arbitration prediction result and the autonomous driving expected execution result, and the differences between the predicted weight distribution and the expected weight distribution, is represented by a loss function, then the expression of the loss function includes: , in, Represents the loss function. Indicates the loss weight of the decision item. Indicates the loss of the decision item. This indicates the predicted results of the autonomous driving arbitration. This indicates the expected execution result of autonomous driving. Represents the loss weight of the distribution term. Represents the loss of the distribution term. Indicates the predicted weight distribution. Indicates the expected weight distribution; The expression for the loss of the decision item is as follows: , in, Indicates the number of samples. The first part of the expected execution result of autonomous driving The expected result The first in the autonomous driving arbitration prediction results One prediction result, This represents the smoothed L1 loss function; The formula for calculating the loss of the distribution term includes: , in, This indicates the number of categories of input decision information. Indicates the first The expected weights of each category of input decision information. Indicates the first The prediction weights for each category of input decision information. This indicates the smoothing term.
5. The multimodal fusion autonomous driving decision arbitration method based on the VLA model according to claim 1, characterized in that, Based on the pre-trained arbitration network model, the process of arbitrating the second historical time-series state information, the second historical environment information, and the second multi-source historical decision information includes: The second historical time-series state information is feature-encoded to obtain a second state feature code; the second historical environment information is feature-encoded to obtain a second environment feature code; and the second multi-source historical decision information is feature-encoded to obtain a second decision feature code; the second decision feature code includes: a second historical traffic constraint code, a second historical planning trajectory code, a second historical execution action code, and a second historical control instruction code; Based on the self-attention mechanism module, the second decision fusion feature, the second state feature encoding, and the second environment feature encoding are fused to obtain the second multi-source fusion feature; the second decision fusion feature is obtained by fusing the attention weight distribution among the second decision feature encodings. The second multi-source fusion feature is used to make a classification decision to obtain the predicted trajectory and predicted state of the unmanned logistics vehicle, as well as the predicted trajectory and predicted state of the moving target.
6. The multimodal fusion autonomous driving decision arbitration method based on the VLA model according to claim 1, characterized in that, The formula for calculating the multi-objective reward includes: , in, Indicates multi-objective rewards. Indicates the security reward weight. Indicates a security reward. Indicates the weight of driving efficiency rewards. This indicates a reward for improved driving efficiency. Indicates the weight of trajectory comfort reward. Indicates a reward for comfort during the trajectory. This indicates the reward weight for following the rules. This indicates that following the rules will be rewarded. Indicates the weight of remote control rewards. Indicates a reward for remote control; The formula for calculating the security reward includes: , in, Indicates the weight of the distance penalty term. This represents the total number of time steps within the prediction time domain. express Distance penalty function for time steps, Indicates the weight of the collision penalty term. express The collision indicator function at each time step is set to 1 when a collision occurs and 0 when no collision occurs. express The distance traveled in a time step; The The expression for the distance penalty function at each time step is as follows: , in, Indicates the physical collision threshold. Indicates the radius of safety impact. Indicates the collision penalty coefficient; The formula for calculating the driving efficiency bonus includes: , in, Indicates the reward weight of the speed item. This indicates that unmanned logistics vehicles are in Prediction speed of time step This indicates the speed limit for unmanned logistics vehicles. Indicates the reward weight based on the distance to the finish line. Indicates in The time step measures the distance between the unmanned logistics vehicle and its destination. Indicates in The time step measures the distance between the unmanned logistics vehicle and its destination. Indicates a unit of time. This indicates the penalty weight for the acceleration term. This indicates that unmanned logistics vehicles are in Predicted acceleration at time step This indicates the maximum acceleration of the unmanned logistics vehicle; The formula for calculating the trajectory comfort reward includes: , in, Indicates the weight of the accelerometer term. This indicates that unmanned logistics vehicles are in Predicted acceleration at time step This indicates the maximum acceleration of the unmanned logistics vehicle. Indicates the weight of the angular velocity term. This indicates that unmanned logistics vehicles are in Predicted angular velocity at the time step This represents the maximum angular velocity of the unmanned logistics vehicle; The formula for calculating the remote control reward includes: , in, Indicates the weight of the manually managed items. This represents the manual takeover indicator function. When the unmanned logistics vehicle is in a manually controlled state, the manual takeover indicator function is 1; when the unmanned logistics vehicle is not in a manually controlled state, the manual takeover indicator function is 0. Indicates the weight of the switching delay time item. express Switching delay time when the time step is in manual takeover mode. Indicates the maximum handover delay time; The rules follow the reward calculation formula, including: , in, Indicates the weight of the violation item. express The violation indication function for the time step, when the unmanned logistics vehicle is in If a time step violation occurs, then The violation indication function for the time step is 1, when the unmanned logistics vehicle is in If no violation occurs during the time step, then The violation indicator function for the time step is 0.
7. The multimodal fusion autonomous driving decision arbitration method based on the VLA model according to any one of claims 1-3, characterized in that, The process of fusing decision fusion features, state feature encoding, and environment feature encoding based on the self-attention mechanism module to obtain multi-source fusion features includes: Based on the self-attention mechanism module, the attention weight distribution among the decision fusion feature, the state feature encoding, and the environment feature encoding is calculated; The decision fusion features are updated according to the attention weight distribution to obtain decision update features; the state feature encoding is updated according to the attention weight distribution to obtain state update encoding; and the environment feature encoding is updated according to the attention weight distribution to obtain environment update encoding. A residual connection is performed on the decision update feature and the decision fusion feature to obtain a first decision connection feature; a residual connection is performed on the state update code and the state feature code to obtain a first state connection feature; a residual connection is performed on the environment update code and the environment feature code to obtain a first environment connection feature; The first decision connection feature is input into the feedforward network module to obtain the decision feedforward output feature; the first state connection feature is input into the feedforward network module to obtain the state feedforward output feature; the first environment connection feature is input into the feedforward network module to obtain the environment feedforward output feature. A second decision connection feature is obtained by performing a residual connection on the first decision connection feature and the decision feedforward output feature; a second state connection feature is obtained by performing a residual connection on the first state connection feature and the state feedforward output feature; and a second environment connection feature is obtained by performing a residual connection on the first environment connection feature and the environment feedforward output feature. The second decision connection feature is layer-normalized, the second state connection feature is layer-normalized, and the second environment connection feature is layer-normalized. The layer-normalized decision connection feature, the layer-normalized state connection feature, and the layer-normalized environment connection feature are then concatenated to obtain the multi-source fusion feature.
8. A multimodal fusion autonomous driving decision arbitration system based on a VLA model, characterized in that, The system is used to implement the multimodal fusion autonomous driving decision arbitration method based on the VLA model as described in claim 1, and the system includes: The information acquisition module is used to acquire the temporal state information, current environmental information, and multi-source decision information of the unmanned logistics vehicle. The multi-source decision information includes: current traffic constraints, current planned trajectory, current executed action, and current manual control command. The current traffic constraints are generated based on the temporal state information and the current environmental information. The current executed action is generated by the VLA model with the temporal state information, the current environmental information, and the moving target state information as input. The current planned trajectory is generated by the motion planning module with the temporal state information, the current environmental information, and the moving target state information as input. The information arbitration module is used to arbitrate the time-series state information, the current environment information, and the multi-source decision information based on a pre-built arbitration network model to obtain the autonomous driving arbitration result of the unmanned logistics vehicle; the autonomous driving arbitration result includes: autonomous driving trajectory and autonomous driving execution actions; The process of arbitrating the temporal state information, the current environment information, and the multi-source decision information includes: The temporal state information is feature-encoded to obtain a state feature code; the current environment information is feature-encoded to obtain an environment feature code; and the multi-source decision information is feature-encoded to obtain a decision feature code; the decision feature code includes: traffic constraint code, planning trajectory code, execution action code, and control command code; Based on the self-attention mechanism module, feature fusion is performed on the decision fusion feature, the state feature encoding, and the environment feature encoding to obtain multi-source fusion features; the decision fusion feature is obtained by fusing the attention weight distribution among the decision feature encodings. The multi-source fusion features are classified and a decision is made to obtain the autonomous driving arbitration result.
Citation Information
Patent Citations
Hybrid decision arbitration method and system based on automatic driving
CN121246855A
Multi-modal large model training and deployment method for automatic driving
CN121457291A