Multi-unmanned vehicle coordination control method based on environment alignment large language model

By adopting a large language model based on environmental alignment in multi-unit vehicle systems, the problems of limited information sharing and global state construction in collaborative control of multi-unit vehicle are solved, efficient global state completion and dynamic environmental adaptation are achieved, and collaborative control efficiency and decision-making quality are improved.

CN120178879APending Publication Date: 2025-06-20TONGJI UNIV

Patent Information

Application Number
CN202510317496.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing multi-unit vehicle collaborative control method has problems such as limited information sharing, slow training convergence speed, and low coordination efficiency in high-dimensional decision-making issues, and has failed to fully consider the information sharing, global state construction and collaborative decision-making issues in multi-unit vehicle systems.

Method used

A large language model based on environmental alignment is adopted, and local observations are obtained from the environment through each unmanned vehicle, and policies and communication prompts are obtained to the prompt word system. In the centralized training stage, the global state is obtained through the communication network, the policy network parameters are optimized, and the distributed execution stage is obtained through the optimized policy network, and the next action is sampled to obtain.

Benefits of technology

The completion of global state is achieved, the system's dependence on accessing global state is reduced, the collaborative control capability of multiple unmanned vehicle systems is improved, the adaptability to complex dynamic environments is enhanced, and the decision quality and system reliability are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120178879A_ABST
    Figure CN120178879A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-unmanned vehicle coordination control method based on an environment alignment big language model, and the method comprises the steps: enabling a cue word system to convert the local observation information of each unmanned vehicle into a structured text, generating a strategy prompt and a communication prompt, and generating an executable action sequence and communication information through a pre-trained big language model. And each unmanned vehicle selects the optimal action through random sampling according to the generated action execution probability, calculates the value of a global state through a commentator network, and optimizes strategy network parameters. All unmanned vehicles share a group of strategy network parameters, and collective learning and optimization are carried out by using collective trajectory data, so that the cooperative capability and decision accuracy of the system are improved. Through the method, multiple unmanned vehicles can efficiently cooperate in a dynamic and uncertain environment to jointly complete complex tasks. The task execution efficiency, the collaborative decision-making capability and the robustness of the multi-unmanned vehicle system can be remarkably improved, and the method has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of coordinated control of unmanned vehicles, and particularly relates to a multi-unmanned vehicle coordinated control method based on an environment-aligned large language model. Background Art

[0002] In recent years, multi-unmanned vehicle systems have been widely used in fields such as intelligent transportation, military reconnaissance, logistics distribution, and disaster relief. The core goal of multi-unmanned vehicle collaborative control is to achieve efficient perception, decision-making, and execution in a dynamic and uncertain environment to complete complex collaborative tasks.

[0003] Currently, multi-unmanned vehicle collaborative control mainly adopts reinforcement learning (RL), multi-agent reinforcement learning (MARL), game theory methods, and rule-based decision-making systems. Among them, the reinforcement learning-based method shows strong adaptability in high-dimensional decision-making problems, but there are still problems such as limited information sharing among multi-unmanned vehicles, slow training convergence speed, and low collaborative efficiency.

[0004] In the prior art, there has been research on using large language models to train unmanned vehicles. For example, Chinese Patent Application CN118761453A discloses a safety reinforcement learning method for an autonomous driving system based on LLM and KG constraints, including: collecting environmental information at the initial moment; generating questions through a constraint reward question-answerer constructed by a large language model plus a knowledge graph, and then generating real-time constraint text; combining the constraint text with the state information and transmitting it to the unmanned vehicle to make an action; obtaining the new state, prompt questions, and constraint text at time t + 1, and using the constraint reward question-answerer to evaluate the action at the previous moment and give a reward; inputting the constraint text, state information, and comprehensive reward into the unmanned vehicle together, and the unmanned vehicle makes an action at the moment; performing iterative safety learning and converging the strategy. The limitation of this method is that it only constraints the decision-making process of a single unmanned vehicle, ignoring the generation of the global state and the coordination of actions in multi-unmanned vehicle collaborative work, and fails to fully consider information sharing, global state construction, and collaborative decision-making problems in multi-unmanned vehicle systems. Therefore, there are still the following deficiencies in the multi-unmanned vehicle environment. This method mainly uses a constraint reward question-answerer to perform reward constraints on a single unmanned vehicle. In a multi-unmanned vehicle system, the optimization of individual rewards does not necessarily lead to the overall system optimum, and the overall collaborative performance may decline due to inconsistent individual goals. Summary of the Invention

[0005] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a multi-unmanned vehicle coordinated control method based on an environment-aligned large language model.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] On the one hand, the present invention provides a multi-unmanned vehicle coordinated control method based on an environment-aligned large language model, including the following steps:

[0008] Each unmanned vehicle obtains local observations from the environment;

[0009] Each unmanned vehicle sends the local observations to the prompt system; based on the local observations, the prompt system obtains the policy prompts and communication prompts for each unmanned vehicle;

[0010] In the centralized training stage, based on the communication prompts, the global state is obtained through the communication network. Based on the policy prompts, each unmanned vehicle generates an executable action sequence and its action execution probability through the policy network. Each unmanned vehicle obtains the next action by sampling based on the action execution probability, executes the action and obtains the reward value, and inputs the global state and the reward value into the critic network to optimize the parameters of the policy network of each unmanned vehicle through the critic network;

[0011] In the distributed execution stage, each unmanned vehicle obtains the action probability distribution function through the optimized policy network based on the policy prompts; each unmanned vehicle obtains the next action by sampling based on the action probability distribution function.

[0012] Further, the environment-aligned large language model includes a policy network based on the environment-aligned large language model, a communication network based on the environment-aligned large language model, a prompt system, and a critic network based on a multi-layer perceptron. Both the policy network and the communication network are pre-trained large language models.

[0013] Further, the local observation is a set of partial information sensed by a single unmanned vehicle from the environment at the current moment, including position information, speed information, path planning information, sensor information, and environmental state information.

[0014] Further, the process of obtaining the policy prompts includes:

[0015] The prompt system receives the local observations sent by each unmanned vehicle, formats the local observation information of each unmanned vehicle, converts the local observation information into structured text, and combines it with the preset task instructions of each unmanned vehicle to form policy prompts. The policy prompts include a policy instruction part, a policy input part, and a policy response part, where:

[0016] The policy instruction part is used to describe the task objective and specify the decision-making purpose of each unmanned vehicle, including navigation, obstacle avoidance, and cooperative task execution. The text description corresponding to the current task instruction of each unmanned vehicle is extracted from the preset task library according to the preset task instructions;

[0017] The policy input part is the textualized local observation information of all unmanned vehicles, including the states of all unmanned vehicles, the perceived environmental information, and the key information that may affect the decisions of other unmanned vehicles, which is described in the form of structured text;

[0018] The policy response part is a set of formatted executable actions. All legal actions of each unmanned vehicle under the current local observation conditions are extracted from a predefined action library and converted into text form. The set of executable actions includes moving forward, moving backward, turning left, turning right, stopping, accelerating, and decelerating.

[0019] Furthermore, the communication prompt acquisition process includes:

[0020] The prompt word system receives the local observation information sent by each unmanned vehicle, formats the local observation information, converts the local observation information into a textual description, and combines it with a preset communication instruction to generate a communication prompt. The communication prompt includes a communication instruction part, a communication input part, and a communication response part, where:

[0021] The communication instruction part is used to describe the communication purpose and specify the information interaction objectives between unmanned vehicles, including sharing environmental information, exchanging status data, and collaborative decision-making support. The corresponding textual description is extracted from a preset communication instruction library according to the communication instruction;

[0022] The communication input part is the textualized local observation information of all unmanned vehicles, including the states of all unmanned vehicles, the perceived environmental information, and the key information that may affect the decisions of other unmanned vehicles, which is described in the form of structured text;

[0023] The communication response part is the textualized global state information at the previous moment. Based on the local observation information provided by all unmanned vehicles at the previous moment, the global state is inferred through a communication network based on the environment-aligned large language model, and the global state information is converted into a text representation.

[0024] Furthermore, based on the communication prompt, obtaining the global state through the communication network specifically includes:

[0025] The communication prompt is processed by a tokenizer to obtain the corresponding token sequence, and the token sequence is input into the communication network. The communication network generates the textualized global state at the current moment based on the communication purpose in the communication prompt, the local observation information of all unmanned vehicles, and the textualized global state information at the previous moment.

[0026] Furthermore, based on the policy prompt, each unmanned vehicle generates an executable action sequence and its action execution probability through a policy network, specifically including:

[0027] The policy prompt is processed by a tokenizer to generate a sequence of policy prompt tokens, which are provided as input to the policy network of the large language model aligned with the environment.

[0028] Based on the sequence of policy prompt tokens, the policy network generates an executable sequence of action tokens and their corresponding generation probabilities. The formula is as follows:

[0029]

[0030] where P(a i |o) represents the generation probability of generating the action sequence a i under the condition of given all local observations o of the unmanned vehicles, is the number of tokens of the action sequence a i , T o is the number of tokens of all local observation sequences o of the unmanned vehicles, and P(x i |x1,x2,…,x i-1 ) represents the probability of generating the current token x i under the condition of given the previous i - 1 tokens. The executable actions include one or more tokens.

[0031] For the generated sequence of executable action tokens, softmax normalization is performed, and finally the action execution probability of each executable action is obtained. The formula is as follows:

[0032]

[0033] where P sofrmax (a i |o) represents the action execution probability of action a i under the condition of given all local observations o of the unmanned vehicles, |a i | is the number of tokens in the executable action a i , and A represents the set of all executable actions.

[0034] Furthermore, each unmanned vehicle obtains the next action by sampling based on the action execution probability, which specifically includes:

[0035] According to the action execution probability generated by each unmanned vehicle, a temperature sampling method is used to select an action from the set of executable actions.

[0036] The temperature sampling method includes: adjusting the smoothness of the action execution probability distribution of each action through a temperature parameter T, using the adjusted action execution probability of each action as the weight for the action to be selected, and randomly selecting an action from the set of actions. The formula is as follows:

[0037]

[0038] Among them, P i is the action execution probability of the i-th action after adjustment;

[0039] By adjusting the temperature T, a higher temperature is used at the initial stage of training to encourage the driverless vehicle to try different actions; the temperature is gradually reduced at the later stage of training.

[0040] Furthermore, inputting the global state and the reward value into the critic network to optimize the parameters of each driverless vehicle policy network specifically includes:

[0041] Input the global state S t at the current moment and the corresponding reward value r t+1 into the critic network, and the critic network calculates the state value function V t according to the current global state S π (S t ). The calculation formula is:

[0042]

[0043] Among them, γ is the discount factor, π represents the joint policy π i is the policy of the i-th driverless vehicle, s0 is the global state at the initial moment, and V π (S t ) represents the expected value of the cumulative reward that can be obtained in the future starting from the global state S t under the policy π obtained by the current policy network;

[0044] The critic network calculates the Q-value function Q t according to the current global state and the action a π (S t ,a t ). The calculation formula is:

[0045]

[0046] Among them, a0 is the action taken at the initial moment, and Q π (S t ,a t ) represents the expected value of the cumulative reward that can be obtained in the future after taking the action a t in the global state S t under the policy π obtained by the current policy network;

[0047] Based on the state value function V π (S t ) and the Q-value function Q π (S t ,at ), calculate the advantage function $A$ π (S t , $A$ t ), and the calculation formula is:

[0048] $A$ π (S t , $A$ t ) = $Q$ π (S t , $A$ t ) - $V$ π (S t )

[0049] where, $A$ π (S t , $A$ t ) represents the degree of superiority or inferiority of the average performance of executing a specific action $A$ t under the global state $S$ t relative to the policy $\pi$;

[0050] Based on the advantage function $A$ π (S t , $A$ t ), calculate the policy optimization objective function $J(\theta)$, and the calculation formula is:

[0051]

[0052] where, $\pi$ θ (A i |o i ) represents the action probability of the current policy network, represents the policy of the previous iteration step, and the ratio of the two is the importance sampling ratio, which measures the change in action selection between the old and new policies, is the advantage function of the previous round of policy, and the clip operation limits the sampling ratio within the range of $[1 - \epsilon, 1 + \epsilon]$, where $\epsilon$ is a hyperparameter that defines the maximum change range of policy update;

[0053] Optimize $J(\theta)$ through gradient descent to adjust the policy network parameters.

[0054] Furthermore, all unmanned vehicles share a set of policy network parameters, and the shared policy network parameters are updated using collective trajectory data, where the collective trajectory data includes: the global state and local observation of each unmanned vehicle at each moment, the executed action at each moment, the reward value at each moment, the state transition information at each moment, and the action execution probability generated by the policy network;

[0055] Compared with the prior art, the present invention has the following advantages:

[0056] (1) The present invention realizes the completion of the global state through a large language model, reducing the system's dependence on accessing the global state: For multi-robot reinforcement learning, a large language model based on environment alignment is used to aggregate the local observations of each robot into the global state of the system, allowing the robots to obtain global information in an environment where the global state cannot be accessed, and using the global state dependence to assist in evaluating the policy network, thereby reducing the system's dependence on accessing the global state.

[0057] (2) The present invention realizes the complementary advantages of multi-robot reinforcement learning and large language models, solving the problems of the lack of alignment between the large language model and the environment and the dependence of multi-robot reinforcement learning based on the CTDE paradigm on the global state. The large language model based on environment alignment can be pre-trained using an offline dataset, and the communication model can be fine-tuned on the offline dataset through instruction prompts. When deployed during the online training process, it can still maintain sufficient generalization ability.

[0058] (3) The present invention guides the large language model to generate the global state and actions through a prompt system, realizing information sharing among individual robots, enabling each robot to make optimal decisions based on global information, avoiding problems such as decision conflicts and path overlaps caused by limited local observations in traditional methods, and improving the cooperative control ability of multi-robots.

[0059] (4) The present invention infers the global state through a communication network, uses the large language model aligned with the environment combined with the local observation data of the robot to infer the global state information at the current moment, and provides it as input to each robot, enabling it to optimize the policy based on the global information. Compared with the constraint text that only relies on the knowledge graph in the prior art, this method improves the adaptability of the robot group to complex dynamic environments.

[0060] (5) The present invention optimizes policy generation through the large language model and policy prompts to improve decision-making quality. The present invention uses the large language model to analyze task requirements and generates actions through policy prompts, enabling the robot to not only rely on the action probability calculated by traditional reinforcement learning methods when making decisions, but also be able to optimize in combination with text-based task instructions. Compared with the method of only relying on reinforcement learning for action decision-making in the comparative document, the method of the present invention can generate control strategies that better meet the task objectives in complex environments.

[0061] (6) In the present invention, a group of policy network parameters are shared among all unmanned vehicles, and the policy is updated using collective trajectory data. During the training process, the present invention collects the trajectory data of all unmanned vehicles, including global states, local observations, executed actions, reward values, etc., and uses this data to optimize the shared policy network, enabling each unmanned vehicle to maintain consistent decision-making when facing different environments. Compared with the method of only optimizing the policy of a single unmanned vehicle in the comparative document, the present invention significantly improves the collaborative consistency and decision-making generalization ability of the unmanned vehicle group. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is a flowchart of the multi-unmanned vehicle coordinated control method of the present invention;

[0063] Figure 2 is a schematic diagram of the multi-unmanned vehicle coordinated control method of the present invention;

[0064] Figure 3 is a system relationship diagram of the environment alignment large language model module and method of an embodiment of the present invention;

[0065] Figure 4 is a schematic flowchart of an example of the multi-unmanned vehicle coordinated control method based on the environment alignment large language model of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0067] Embodiment 1:

[0068] This embodiment provides a multi-unmanned vehicle coordinated control method based on an environment alignment large language model, as Figure 1 、 Figure 2 shown, including the following steps:

[0069] Each unmanned vehicle obtains local observations from the environment;

[0070] Each unmanned vehicle sends the local observations to the prompt system; based on the local observations, policy prompts and communication prompts for each unmanned vehicle are obtained through the prompt system;

[0071] In the centralized training phase, based on the communication prompt, the global state is obtained through the communication network. Based on the policy prompt, each unmanned vehicle generates an executable action sequence and its action execution probability through the policy network. Each unmanned vehicle samples to obtain the next action based on the action execution probability, executes the action and obtains the reward value, inputs the global state and the reward value into the critic network, and optimizes the parameters of the policy network of each unmanned vehicle through the critic network.

[0072] In the distributed execution phase, each unmanned vehicle obtains the action probability distribution function through the optimized policy network based on the policy prompt; each unmanned vehicle samples to obtain the next action based on the action probability distribution function.

[0073] Furthermore, the large language model based on environment alignment includes a policy network based on the large language model of environment alignment, a communication network based on the large language model of environment alignment, a prompt system, and a critic network based on a multi-layer perceptron. Both the policy network and the communication network are pre-trained large language models.

[0074] The present invention constructs a policy network and a communication network using a large language model based on environment alignment, enabling unmanned vehicles to efficiently analyze task requirements and environmental information, thereby generating more reasonable control strategies. Compared with traditional reinforcement learning methods, the present invention utilizes the natural language understanding ability of the large language model, making policy decisions no longer rely solely on numerical optimization, but reasoning in combination with high-level task descriptions, thus enhancing the flexibility, generalization ability, and environmental adaptability of the policy. At the same time, the communication network generates unified global state information by reasoning about the local observation data of each unmanned vehicle, making individual decisions more accurate, avoiding the uncertainty brought by limited information sharing, and improving the cooperation efficiency and task completion rate of the multi-unmanned vehicle system.

[0075] The prompt system of the present invention constructs policy prompts and communication prompts, enabling the policy network and the communication network to more effectively analyze the task objectives and environmental information of unmanned vehicles, ensuring the rationality of policy generation. Combining with the critic network based on a multi-layer perceptron for policy optimization, the present invention reduces the dependence on a large amount of online interaction data, speeds up the policy convergence speed, and improves the stability of the policy. Overall, this method improves the decision-making quality of the multi-unmanned vehicle system in complex dynamic environments, enables unmanned vehicles to execute cooperative tasks more efficiently, and improves the reliability and adaptability of the system.

[0076] Furthermore, the local observation is a set of partial information sensed by a single unmanned vehicle from the environment at the current moment, including position information, speed information, path planning information, sensor information, and environmental state information.

[0077] Furthermore, the process of obtaining the policy prompt includes:

[0078] The prompt system receives the local observations sent by each unmanned vehicle, formats the local observation information of each unmanned vehicle, converts the local observation information into structured text, and combines it with the preset task instructions of each unmanned vehicle to form a policy prompt. The policy prompt includes a policy instruction part, a policy input part, and a policy response part, where:

[0079] The policy instruction part is used to describe the task objectives, specify the decision-making purposes of each unmanned vehicle, including navigation, obstacle avoidance, and collaborative task execution, and extract the text-based descriptions corresponding to the current task instructions of each unmanned vehicle from the preset task library according to the preset task instructions;

[0080] The policy input part is the text-based local observation information of all unmanned vehicles, including the states of all unmanned vehicles, the perceived environmental information, and the key information that may affect the decisions of other unmanned vehicles, and is described in the form of structured text;

[0081] The policy response part is a formatted set of executable actions, extracts all legal actions under the current local observation conditions of each unmanned vehicle from the predefined action library, and converts them into text form. The set of executable actions includes forward, backward, left turn, right turn, stop, acceleration, and deceleration.

[0082] Through this structured policy prompt, the system can generate personalized control strategies for each unmanned vehicle more efficiently, avoiding the over-reliance or neglect of complex environmental information and task requirements in traditional methods. The introduction of structured text makes the policy network more flexible and has strong reasoning ability when processing complex information, thus improving the quality and consistency of policy generation. In addition, based on the combination of the predefined action set and the task library, the decision-making process is further simplified, and the decision-making speed and accuracy are improved. Finally, the construction of this policy prompt not only improves the collaborative ability of the multi-unmanned vehicle system in a dynamic environment, but also enhances the efficiency and reliability of task execution.

[0083] Furthermore, the process of obtaining the communication prompt includes:

[0084] The prompt system receives the local observation information sent by each unmanned vehicle, formats the local observation information, converts the local observation information into text-based descriptions, and combines it with the preset communication instructions to generate a communication prompt. The communication prompt includes a communication instruction part, a communication input part, and a communication response part, where:

[0085] The communication instruction part is used to describe the communication purpose, specify the information interaction objectives between each unmanned vehicle, including sharing environmental information, exchanging status data, and collaborative decision-making support, and extract the corresponding text-based descriptions from the preset communication instruction library according to the communication instructions;

[0086] The communication input part is the partial observation information of all unmanned vehicles in text form, including the states of all unmanned vehicles, the perceived environmental information, and the key information that may affect the decisions of other unmanned vehicles, which is described in the form of structured text;

[0087] The communication response part is the globally textualized state information at the previous moment. Based on the partial observation information provided by all unmanned vehicles at the previous moment, the global state is inferred through a communication network based on an environment-aligned large language model, and the global state information is converted into a text representation.

[0088] By converting the partial observation information of each unmanned vehicle into unified structured text and combining preset communication instructions to generate communication prompts, the present invention realizes efficient information sharing and collaborative decision-making in a multi-unmanned vehicle system. This process ensures that each unmanned vehicle can communicate environmental information and status data in a timely and accurate manner, avoiding the problem of information silos. At the same time, the global state is generated through the inference ability of the large language model, improving the accuracy of decision-making and collaborative efficiency, thus significantly enhancing the collaborative working ability and task execution efficiency of multi-unmanned vehicles in complex dynamic environments.

[0089] Further, obtaining the global state through the communication network based on the communication prompt specifically includes:

[0090] The communication prompt is processed by a tokenizer to obtain the corresponding token sequence, and the token sequence is input into the communication network. The communication network generates the globally textualized state at the current moment according to the communication purpose in the communication prompt, the partial observation information of all unmanned vehicles, and the globally textualized state information at the previous moment.

[0091] Further, based on the policy prompt, each unmanned vehicle generates an executable action sequence and its action execution probability through a policy network, specifically including:

[0092] The policy prompt is processed by a tokenizer to generate a policy prompt token sequence, which is provided as input to the policy network based on the environment-aligned large language model;

[0093] The policy network generates an executable action token sequence and its corresponding generation probability according to the policy prompt token sequence. The formula is:

[0094]

[0095] where P(a i |o) represents the generation probability of generating the action sequence a i under the condition of the partial observation o of all unmanned vehicles, is the number of tokens of the action sequence a i , and T ois the number of tokens for all local observation sequences o of the unmanned vehicles, P(x i |x1,x2,…,x i-1 ) represents the probability of generating the current token x i given the previous i - 1 tokens. The executable actions include one or more tokens;

[0096] For the generated sequence of executable action tokens, perform softmax normalization to finally obtain the action execution probability of each executable action. The formula is:

[0097]

[0098] where P softmax (a i |o) represents the action execution probability of action a i given all local observations o of the unmanned vehicles. |a i | is the number of tokens in the executable action a i , and A represents the set of all executable actions.

[0099] The token probabilities generated by the policy network are then used to calculate the probabilities of the corresponding actions. Note that each action may consist of a different number of tokens (e.g., the tokens for "stop" and "move north one step" are 1 and 4 respectively). When the probability is calculated as the product of token probabilities, actions with more tokens will result in smaller cumulative values. To address this issue, dividing the probability by the token length ensures that the probabilities of actions with different token lengths are on the same scale. Additionally, the calculated probability P(a i |o) is essentially the logits of each token and is not normalized. Therefore, an additional softmax operation is applied to the logits to obtain the execution probability of each available action.

[0100] Furthermore, each unmanned vehicle obtains the next action through sampling based on the action execution probability, specifically including:

[0101] According to the action execution probability generated by each unmanned vehicle, select an action from the set of executable actions using the temperature sampling method;

[0102] The temperature sampling method includes: adjusting the smoothness of the action execution probability distribution of each action through the temperature parameter T, using the adjusted action execution probability of each action as the weight for that action to be selected, and randomly selecting an action from the set of actions. The formula is:

[0103]

[0104] Among them, P i is the action execution probability of the i-th adjusted action;

[0105] A higher temperature is used in the initial stage of training to encourage the unmanned vehicle to try different actions and discover potential high-reward strategies. The temperature is gradually reduced in the later stage of training to converge the strategy to the optimal action and reduce randomness.

[0106] In this step, each unmanned vehicle uses the temperature sampling method to select the next action according to the action execution probability it generates. The core idea of temperature sampling is to balance the relationship between exploration and exploitation by adjusting the temperature parameter TT. In the context of multi-unmanned vehicle cooperative control, this process helps to effectively improve the learning efficiency of the agent and the final decision-making quality.

[0107] First of all, the temperature sampling method adjusts the smoothness of the action execution probability by introducing the temperature parameter TT. A high temperature value will make the probability distribution smoother, so that actions with lower probabilities also have the opportunity to be selected, encouraging the unmanned vehicle to try different actions in the initial stage of training, thereby exploring more policy spaces. This process helps to discover potential high-reward strategies, avoid the agent falling into a local optimal solution in the initial stage, and enhance the exploration ability of the system.

[0108] In the later stage of training, gradually reducing the temperature can reduce the randomness of sampling, making the agent more inclined to select actions with high probabilities, thereby converging to the optimal strategy. The gradual reduction of temperature plays a role similar to "annealing" in reinforcement learning, making the strategy converge by reducing randomness and improving the certainty of actions. In this way, the agent can ensure extensive exploration in the initial stage while finally gradually locking in to the optimal decision-making strategy to ensure the efficient execution of the cooperative task.

[0109] The technical effect and advantage of this method are that it enables the unmanned vehicle system to balance exploration and exploitation in the learning process. By flexibly adjusting the temperature parameter, the learning ability of the system and the stability of the strategy can be effectively improved, and it can adapt to the needs of different stages, ensuring full exploration in the early stage and rapid convergence to the optimal strategy in the later stage, ultimately improving the success rate and execution efficiency of the multi-unmanned vehicle cooperative control task.

[0110] Furthermore, inputting the global state and the reward value into the critic network to optimize the parameters of each unmanned vehicle's policy network specifically includes:

[0111] Inputting the global state S t at the current moment t+1Input the critic network, which calculates the state-value function V t based on the current global state S π (S t ), and the calculation formula is:

[0112]

[0113] where γ is the discount factor and π represents the joint policy π i is the policy of the autonomous vehicle i, s0 is the global state at the initial time, and V π (S t ) represents the expected value of the cumulative reward that can be obtained in the future starting from the global state S t under the policy π obtained by the current policy network;

[0114] The critic network calculates the Q-value function Q t based on the current global state and the action a executed at the current time π (S t , a t ), and the calculation formula is:

[0115]

[0116] where a0 is the action taken at the initial time, and Q π (S t , a t ) represents the expected value of the cumulative reward that can be obtained in the future after taking the action a t in the global state S t under the policy π obtained by the current policy network;

[0117] Based on the state-value function V π (S t ) and the Q-value function Q π (S t , a t ), the advantage function A π (S t , a t ) is calculated, and the calculation formula is:

[0118] A π (S t , a t ) = Q π (S t , a t ) - V π (S t )

[0119] where A π (S t , a t) represents a measure of the global state S t under which a specific action a t is executed, and it indicates the degree of superiority or inferiority of the average performance relative to the policy π;

[0120] Based on the advantage function A π (S t , a t ), the policy optimization objective function J(θ) is calculated, and the calculation formula is:

[0121]

[0122] where π θ (a i | o i ) represents the action probability of the current policy network, represents the policy of the previous iteration step, and the ratio of the two is the importance sampling ratio, which measures the change in action selection between the old and new policies, is the advantage function of the previous round of policy, and the clip operation limits the sampling ratio within the range of [1 - ∈, 1 + ∈], where ∈ is a hyperparameter that defines the maximum change range of policy update;

[0123] The policy network parameters are adjusted by optimizing J(θ) through gradient descent.

[0124] By calculating the state value function, Q-value function, and advantage function through the critic network, the collaborative decision-making process of multiple unmanned vehicles can be effectively evaluated and optimized. By introducing the advantage function to measure the pros and cons of each action and further optimizing the policy, it can ensure adaptive adjustment according to the global state and real-time rewards in a dynamic environment, thereby improving the decision-making efficiency and accuracy of the multiple unmanned vehicle system. At the same time, using the gradient descent method of the policy optimization objective function J(θ) for policy update can steadily improve the performance of the entire multiple unmanned vehicle cooperation system and ensure its stronger adaptability and robustness when facing complex tasks and changing environments.

[0125] Furthermore, all unmanned vehicles share a set of policy network parameters, and the shared policy network parameters are updated using collective trajectory data, where the collective trajectory data includes: the global state and local observations of each unmanned vehicle at each moment, the executed actions at each moment, the reward values at each moment, the state transition information at each moment, and the action execution probability generated by the policy network;

[0126] Through the sharing of collective trajectory data and collective learning, the collaborative development among unmanned vehicles can be promoted, avoiding the inefficiencies and coordination problems that may arise from the independent training of individual unmanned vehicles. Sharing the parameters of the policy network helps reduce redundant calculations and avoid policy inconsistencies, improving the overall performance of the multi-unmanned vehicle system. Through the update of collective trajectory data, the experience of the entire system can be fully utilized to ensure that each unmanned vehicle can learn from the collective actions and continuously optimize its own policy during the global collaboration process. This method significantly improves the execution efficiency and decision-making accuracy of multi-unmanned vehicles in complex tasks, thereby enhancing the robustness and adaptability of the system.

[0127] Embodiment 2:

[0128] As Figure 2 、 Figure 3 shown, the dual-motor vision intelligent vehicle is a four-wheel industrial intelligent vehicle with strong environmental adaptability that uses front and rear dual motors and is equipped with vision sensors. In this embodiment, this dual-motor vision intelligent vehicle will be used as the unmanned vehicle, environmental exploration will be used as the task, and the technical solution of the present invention will be implemented on this premise.

[0129] The multi-unmanned vehicle reinforcement learning module based on the environment-aligned large language model will distinguish two task execution situations for the dual-motor vision intelligent vehicle cluster according to different execution policies, providing support for the full-process logical analysis and diagnosis. The specific situations are as follows:

[0130] Situation 1: When the dual-motor vision intelligent vehicle cluster needs to execute an environmental exploration task that can access the global state;

[0131] Situation 2: When the dual-motor vision intelligent vehicle cluster needs to execute an environmental exploration task that cannot access the global state.

[0132] As Figure 4 shown in Fig. a, it is the flowchart of the dual-motor vision intelligent vehicle cluster needing to execute an environmental exploration task that can access the global state. As Figure 4 shown in Fig. b, it is the flowchart of the dual-motor vision intelligent vehicle cluster needing to execute an environmental exploration task that cannot access the global state. When the dual-motor vision intelligent vehicle faces Situation 1, the specific steps for executing relevant tasks through the environment-aligned large language model module and the multi-unmanned vehicle reinforcement learning method include:

[0133] Step S101: Build an integrated equipment platform including the control system of the intelligent vehicle cluster, calibrate the sensors, and arrange the multi-unmanned vehicle reinforcement learning network module based on the environment-aligned large language model;

[0134] Step S102: The intelligent vehicle cluster control system issues a collaborative exploration task instruction for the dual-motor vision intelligent vehicle cluster to the multi-unmanned vehicle system;

[0135] Step S103: Input the task instructions of the intelligent vehicle cluster control system and the data captured by the vision sensor into the environment-aligned large language model through the wireless receiving device, and analyze the data by the environment-aligned large language model;

[0136] Step S104: Process all the input information through the prompt system; Take the processed policy prompt as input and pass it into the trained policy network; Determine the exploration actions of the current dual-motor vision intelligent vehicle cluster according to the policy function, and send the action policy to the intelligent vehicle cluster control system;

[0137] Step S105: The intelligent vehicle cluster control system receives the action policy sent by the policy network and conducts an evaluation and analysis. If the actions executed by the current dual-motor vision intelligent vehicle cluster meet the requirements of the task instructions issued by the control system, it is default to allow the dual-motor vision intelligent vehicle cluster to continue the current exploration; If the actions executed by the current dual-motor vision intelligent vehicle cluster do not meet the requirements of the task instructions issued by the intelligent vehicle cluster control system, the intelligent vehicle cluster control system will send an instruction to stop the exploration actions of the current dual-motor vision intelligent vehicle cluster, and require the system module to collect all data again and repeat steps S3 to S5 until the exploration results of the dual-motor vision intelligent vehicle cluster meet the requirements of the task instructions issued by the intelligent vehicle cluster control system, then stop the process.

[0138] When the dual-motor vision intelligent vehicle cluster faces situation 2, the specific steps for performing related tasks through the environment-aligned large language model module and the multi-unmanned vehicle reinforcement learning method include:

[0139] Step S201: Build an integrated equipment platform including the intelligent vehicle cluster control system, calibrate the sensors and arrange the multi-unmanned vehicle reinforcement learning network module based on the environment-aligned large language model;

[0140] Step S202: The intelligent vehicle cluster control system issues a collaborative exploration task instruction for the dual-motor vision intelligent vehicle cluster to the multi-unmanned vehicle system;

[0141] Step S203: Input the task instructions of the intelligent vehicle cluster control system and the data captured by the vision sensor into the environment-aligned large language model through the wireless receiving device, and analyze the current situation by the environment-aligned large language model;

[0142] Step S204: Process all the input information through the prompt system, take the processed communication prompt as input and pass it into the communication network; Complete the information through the communication network and generate the global state of the current dual-motor vision intelligent vehicle cluster, take the policy prompt as input and pass it into the trained policy network, determine the actions of the current dual-motor vision intelligent vehicle cluster according to the policy function, and send the action policy to the intelligent vehicle cluster control system;

[0143] Step S205: The intelligent vehicle cluster control system receives the action policies issued by the policy network for evaluation and analysis. If the actions performed by the current dual-motor vision intelligent vehicle cluster meet the requirements of the task instructions issued by the intelligent vehicle cluster control system, it is default to allow the dual-motor vision intelligent vehicle cluster to continue the current exploration. If the actions performed by the current multi-dual-motor vision intelligent vehicle cluster do not meet the requirements of the task instructions issued by the intelligent vehicle cluster control system, the intelligent vehicle cluster control system will send an instruction to stop the exploration actions of the current dual-motor vision intelligent vehicle cluster, and require the system module to re-collect all data and repeat steps S3 to S5 until the exploration results of the dual-motor vision intelligent vehicle cluster meet the requirements of the task instructions issued by the intelligent vehicle cluster control system, then the process is stopped.

[0144] The multi-unmanned vehicle reinforcement learning method based on environment-aligned large language model of the present invention integrates the advantages of large language models and multi-unmanned vehicle reinforcement learning. With the powerful prior knowledge and semantic understanding ability of the large language model, it can provide policy initialization and a more effective exploration process for multi-unmanned vehicle reinforcement learning. This combination reduces the ineffectiveness of initial random exploration, improves sample efficiency, and accelerates training convergence.

[0145] The multi-unmanned vehicle reinforcement learning method based on environment-aligned large language model of the present invention uses the large language model as a communication network to infer the global state through causal reasoning, reducing the dependence on complex communication mechanisms and global states. This design improves the scalability of the multi-unmanned vehicle system, enabling it to make efficient decisions in partially observable environments.

[0146] The multi-unmanned vehicle reinforcement learning method based on environment-aligned large language model of the present invention designs compact and efficient prompt words for the policy network and communication network, reducing redundant representations of the input and optimizing memory usage. It effectively supports the expansion of the multi-unmanned vehicle system, ensuring the efficiency and adaptability of the framework.

[0147] The principle of the present invention is as follows:

[0148] The core principle of the large language model LLM is to predict the probability of the next token in the sequence based on the previous context of the given sequence. Formally, the joint probability of a token sequence (x1, x2, …, x k ) can be decomposed as:

[0149]

[0150] where P LLMs (x i |x1, x2, …, x i-1 ) represents the probability of x i given all the prior tokens in the sequence for the LLMThe conditional probability. In MARL, the policy network is defined as a mapping that generates a probability distribution over actions \(P(a i |o i ), and the communication network is defined as a mapping that generates a probability distribution over the global state through communication \(P(s|o 1 ,…,o n ). To integrate these into the framework, the corresponding actions, local observations, and global states are transformed into text representations through prompt design. Then the tokenizer processes these text prompts to generate the corresponding token sequences and \(s = \{x_1,\ldots,x |s| \}\), where \(|\cdot|\) represents the length of the token sequence. The policy network processed by the LLM can be written as:

[0151]

[0152] For the communication network \(P LLMs (s|o 1 ,…,o n ) computed by the LLMs, it can be expressed as:

[0153]

[0154] The left side of the formula represents the probability of generating the state sequence \(s\) given all observations. The right side is in the form of a product of probabilities, starting from the index and ending at . Each product term \(P(x i |x_1,x_2,\ldots,x i-1 ) represents the probability of generating the \(i\)-th element, conditionally dependent on all previous elements. The observation sequence \(o 1 ,…,o n is concatenated into a fixed context, occupying the first positions. The state sequence \(s\) is generated starting from the position after the observations end, i.e., from to At each step of generating the state element \(x i , the conditional history includes all observations and the generated state elements. For example, if the total length of the observations is 7 and the state length is 5, the state sequence generation range is \(i = 8\) to \(i = 12\), and the conditional history for each \(x i is \(x_1,x_2,\ldots,x i-1 . The formula generates the state sequence step by step in an autoregressive manner, with the probability at each step based on the complete historical information, and the final probability is the product of the joint probabilities of each generation step.

[0155] Based on the above analysis, a construction method that defines LLMs as a policy and communication network is established. The initial state of the policy network is not a random policy, but the prior knowledge embedded in LLMs can be utilized to reduce ineffective exploration. Similarly, for the communication network, the causal reasoning ability of LLMs can be used to extract the global state from local observations. This fundamental insight constitutes the motivation for constructing the EALLMs framework.

[0156] Specifically, we first texturize the current observations and available actions of the unmanned vehicle to form a policy prompt. The local observations are formatted as inputs, and each valid action is formatted as a response. These components are then combined with task-specific instructions (e.g., collaborating with the team to complete a task) to form the policy prompt. The prompt is tokenized by a tokenizer and provided as input to EALLMs-P.

[0157] The token probabilities generated by the policy network EALLMs-P are then used to calculate the probabilities of the corresponding actions. It should be noted that each action may consist of a different number of tokens (e.g., the tokens for "stop" and "move north one step" are 1 and 4 respectively). When the probability is calculated by the product of token probabilities, actions with more tokens will result in smaller cumulative values. To address this issue, dividing the probability by the token length ensures that the probabilities of actions with different token lengths are on the same scale. Additionally, the calculated probability P LLMs (a i |o i ) is essentially the logits of each token and is not normalized. Therefore, an additional softmax operation is applied to the logits to obtain the execution probability of each available action:

[0158]

[0159] The response part of the communication prompt contains the texturized global state. For example, "Enemy A is located at (3,5), and the rest of the area is safe", the response part provides a clear text description of the global state, which can guide LLMs to learn how to correctly infer global information from local observations and ensure the format consistency of generating the global state.

[0160] In the centralized training phase, the action policies of the unmanned vehicle swarm are obtained through the EALLMs-P policy network. Based on the communication prompt, the global state is obtained through the EALLMs-C network. Based on the global state, the state value function is obtained through the critic network. Based on the state value function, the decision-making quality of the EALLMs-P network is evaluated; based on the evaluation results, the parameters of the EALLMs-P network are optimized;

[0161]

[0162] Represents the expected value of the cumulative reward obtained under the condition that the global state is s. Where γ is the discount factor and π represents the joint policy s0 is the global state at time t = 0.

[0163]

[0164] Represents the expected value of the cumulative reward obtained under the condition that the global state is s and the action taken is a. a0 is the action taken at time t = 0. The difference from the state value function is that Q π (s,a) adds the initial action, and the initial action of V π (s) is determined by the policy.

[0165] A π (s,a) = Q π (s,a) - V π (s)

[0166] Represents the difference between the expected reward obtained by executing action a and the average performance of the policy under the same policy π and global state a. The advantage function measures the quality of executing action a in state s relative to the average performance of policy π. If A π (s,a) > 0, the return of action a is higher than the average return of policy π in state s, indicating that this action is better.

[0167]

[0168] Where π θ (a i |o i ) represents the current policy, represents the policy of the previous iteration step. The ratio of the two is the importance sampling ratio, indicating the new policy π θ and the old policy in the ratio of action selection probabilities. The advantage function measures the quality of executing action a in state s relative to the old policy 's average performance. The clip operation limits the importance sampling ratio within the range of [1 - ∈, 1 + ∈] to prevent the policy update from being too large. ∈ is a hyperparameter that defines the clipping range of the sampling ratio. The min function takes the smaller value of the two to ensure that the policy update neither deviates too far from the old policy nor fails to utilize the positive signal of the advantage function.

[0169] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0170] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A multi-unmanned vehicle coordinated control method based on an environment-aligned large language model, characterized in that: The following steps are involved: Each unmanned vehicle obtains local observations from the environment; Each unmanned vehicle sends the local observation to a prompt word system; based on the local observation, the prompt word system obtains strategy prompts and communication prompts of each unmanned vehicle; In the centralized training phase, based on the communication prompt, the global state is obtained through the communication network. Based on the policy prompt, each unmanned vehicle generates an executable action sequence and its action execution probability through the policy network. Based on the action execution probability, each unmanned vehicle obtains the next action through sampling, executes the action and obtains the reward value. The global state and reward value are input into the critic network, and the parameters of each unmanned vehicle policy network are optimized through the critic network. In the distributed execution stage, each unmanned vehicle obtains the action probability distribution function through the optimized strategy network based on the strategy prompt; each unmanned vehicle obtains the next action through sampling based on the action probability distribution function.

2. According to claim 1, a method for coordinated control of multiple unmanned vehicles based on an environment-aligned large language model is characterized in that: The environment-aligned large language model includes a strategy network based on the environment-aligned large language model, a communication network based on the environment-aligned large language model, a prompt word system, and a critic network based on a multi-layer perceptron. Both the strategy network and the communication network are pre-trained large language models.

3. The method for coordinated control of multiple unmanned vehicles based on an environment-aligned large language model according to claim 1 is characterized in that: The local observation is a collection of partial information perceived by a single unmanned vehicle from the environment at the current moment, including position information, speed information, path planning information, sensor information, and environmental status information.

4. The method for coordinated control of multiple unmanned vehicles based on an environment-aligned large language model according to claim 1 is characterized in that: The strategy prompt acquisition process includes: The prompt word system receives the local observations sent by each unmanned vehicle, formats the local observation information of each unmanned vehicle, converts the local observation information into structured text, and combines it with the preset task instructions of each unmanned vehicle to form a strategy prompt. The strategy prompt includes a strategy instruction part, a strategy input part and a strategy response part, wherein: The strategy instruction part is used to describe the mission objectives and stipulate the decision-making purpose of each unmanned vehicle, including navigation, obstacle avoidance, and collaborative task execution. According to the preset mission instructions, the text description corresponding to the current mission instructions of each unmanned vehicle is extracted from the preset mission library; The policy input part is the textual local observation information of all unmanned vehicles, including the status of all unmanned vehicles, the perceived environmental information, and key information that may affect the decision-making of other unmanned vehicles, described in the form of structured text; The strategy response part is a formatted set of executable actions. All legal actions of each unmanned vehicle under the current local observation conditions are extracted from the predefined action library and converted into text form. The executable action set includes forward, backward, left turn, right turn, stop, accelerate, and decelerate.

5. The method for coordinated control of multiple unmanned vehicles based on an environment-aligned large language model according to claim 1, characterized in that: The communication prompt acquisition process includes: The prompt word system receives the local observation information sent by each unmanned vehicle, formats the local observation information, converts the local observation information into a text description, and generates a communication prompt in combination with a preset communication instruction. The communication prompt includes a communication instruction part, a communication input part and a communication response part, wherein: The communication instruction part is used to describe the purpose of communication and stipulate the information interaction goals between each unmanned vehicle, including sharing environmental information, exchanging status data, and collaborative decision support. According to the communication instruction, the corresponding text description is extracted from the preset communication instruction library; The communication input part is the textual local observation information of all unmanned vehicles, including the status of all unmanned vehicles, the perceived environmental information, and key information that may affect the decision-making of other unmanned vehicles, described in the form of structured text; The communication response part is the textualized global state information of the previous moment. Based on the local observation information provided by all unmanned vehicles at the previous moment, the global state is inferred through a communication network based on the environment-aligned large language model, and the global state information is converted into a text representation.

6. The method for coordinated control of multiple unmanned vehicles based on an environment-aligned large language model according to claim 1, characterized in that: The acquiring of the global state through the communication network based on the communication prompt specifically includes: The communication prompt is processed through a word segmenter to obtain a corresponding word sequence, and the word sequence is input into the communication network. The communication network generates a textualized global state at the current moment based on the communication purpose in the communication prompt, the local observation information of all unmanned vehicles, and the textualized global state information at the previous moment.

7. The method for coordinated control of multiple unmanned vehicles based on an environment-aligned large language model according to claim 1 is characterized in that: Based on the strategy prompts, each unmanned vehicle generates executable action sequences and their action execution probabilities through the strategy network, including: The strategy prompt is processed by the tokenizer to generate a strategy prompt token sequence, which is then provided as input to the strategy network based on the environment-aligned large language model. The policy network generates executable action token sequences and their corresponding generation probabilities based on the policy prompt token sequence. The formula is: Among them, P(a i |o) means generating action sequence a given the local observation o of all unmanned vehicles i The probability of generating For action sequence a i The number of tokens, T o is the number of tokens of all unmanned vehicle local observation sequences o, P(x i |x1,x2,…,x i-1 ) indicates that the current token x is generated given the previous i-1 tokens. i The executable action includes one or more tokens; For the generated executable action token sequence, softmax normalization processing is performed to finally obtain the action execution probability of each executable action, and the formula is: Among them, P softmax (a i |o) means that given the local observation o of all unmanned vehicles, action a i The probability of executing an action, |a i | is an executable action a i The number of tokens in A represents the set of all executable actions.

8. The method for coordinated control of multiple unmanned vehicles based on an environment-aligned large language model according to claim 1 or 7, characterized in that: Each unmanned vehicle obtains the next action through sampling based on the action execution probability, specifically including: According to the action execution probability generated by each unmanned vehicle, an action is selected from the executable action set using the temperature sampling method; The temperature sampling method includes: adjusting the smoothness of the action execution probability distribution of each action by the temperature parameter T, taking the adjusted action execution probability of each action as the weight of the action to be selected, and randomly selecting an action from the action set, the formula is: Among them, P i is the adjusted action execution probability of the i-th action; By adjusting the temperature T, a higher temperature is used in the early stages of training to encourage the driverless car to try different actions; the temperature is gradually lowered in the later stages of training.

9. The method for coordinated control of multiple unmanned vehicles based on an environment-aligned large language model according to claim 1, characterized in that: The global state and reward value are input into the critic network, and the parameters of each unmanned vehicle strategy network are optimized through the critic network, specifically including: The global state S at the current moment t And the corresponding reward value r t+1 Input critic network, critic network is based on the current global state S t Calculate the state value function V π (S t ), the calculation formula is: Among them, γ is the discount factor, and π represents the joint strategy π i is the strategy of the unmanned vehicle i, s0 is the global state at the initial moment, V π (S t ) means that under the strategy π obtained by the current policy network, from the global state S t Initially, the expectation of cumulative rewards to be obtained in the future; The critic network is based on the current global state and the action a performed at the current moment. t Calculate the Q value function Q π (S t ,a t ), the calculation formula is: Among them, a0 is the action taken at the initial moment, Q π (S t ,a t ) Under the strategy π obtained by the current policy network, in the global state S t Take action a t After that, the expectation of the cumulative rewards that can be obtained in the future; Based on the state value function V π (S t ) and the Q-value function Q π (S t ,a t ), calculate the advantage function A π (S t ,a t ), the calculation formula is: A π (S t ,a t )=Q π (S t ,a t )-V π (S t ) Among them, A π (S t ,a t ) represents the measurement in the global state S t Next, perform a specific action a t How good or bad is the average performance of strategy π? Based on the advantage function A π (S t ,a t ) Calculate the strategy optimization objective function J(θ), the calculation formula is: Among them, π θ (a i |o i ) represents the action probability of the current policy network, represents the strategy of the previous iteration step. The ratio of the two is the importance sampling ratio, which measures the change in action selection between the new and old strategies. is the advantage function of the previous round of strategy. The clip operation limits the sampling ratio to the range of [1-∈, 1+∈]. ∈ is a hyperparameter that defines the maximum range of strategy update. The policy network parameters are adjusted by optimizing J(θ) through gradient descent.

10. The method for coordinated control of multiple unmanned vehicles based on an environment-aligned large language model according to claim 1, characterized in that: All unmanned vehicles share a set of policy network parameters and use collective trajectory data to update the shared policy network parameters. The collective trajectory data includes: the global state and local observation of each unmanned vehicle at each moment, the execution action at each moment, the reward value at each moment, the state transition information at each moment, and the action execution probability generated by the policy network.

Citation Information

Patent Citations

  • Automatic driving system safety reinforcement learning method based on LLM and KG constraints

    CN118761453A

Cited By

  • Unmanned aerial vehicle air-ground collaborative way-finding method and system based on federal large model

    CN122488803A