Multi-vehicle cooperative decision method and device, equipment, storage medium and program product
By constructing a multi-vehicle collaborative decision-making method and optimizing the policy network using individual, neighborhood, and global advantage functions, the problem of insufficient multi-vehicle collaborative decision-making capability in complex traffic scenarios is solved, achieving more efficient collaborative decision-making and better environmental understanding.
Patent Information
- Application Number
- CN202511484951.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-10-17
AI Technical Summary
Existing technologies have poor multi-vehicle collaborative decision-making capabilities in complex traffic scenarios and lack effective handling of chaotic vehicle interactions, resulting in insufficient interpretability and reliability of decision-making models.
By constructing a multi-vehicle collaborative decision-making method, we quantify vehicle interactions using individual observations of each vehicle, and construct individual, neighborhood, and global advantage functions by combining a pre-trained policy network. We then optimize the policy network to achieve local and global collaborative decision-making.
It improves the interpretability and reliability of collaborative decision-making in complex traffic scenarios, reduces model error, and enhances the algorithm's generalization ability and training stability.
Smart Images

Figure CN120952585B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, specifically to multi-vehicle collaborative decision-making methods, devices, equipment, storage media, and program products. Background Technology
[0002] With the development of artificial intelligence and in-vehicle hardware, autonomous driving technology is receiving increasing attention and application. Decision-making is a crucial part of autonomous driving tasks, especially when road conditions are complex and surrounding traffic is dense; inappropriate decisions can put the vehicle in a dangerous situation.
[0003] Currently, multi-agent reinforcement learning methods have been widely applied to the collaborative decision-making problem in autonomous driving. Existing technical solutions avoid the adverse effects of a single vehicle's optimal decision on other vehicles by constructing a collaborative decision-making model for multiple autonomous vehicles, thus achieving multi-vehicle collaborative decision-making. However, they do not fully consider the interaction intentions between vehicles and lack effective mechanisms to handle issues such as dense vehicle traffic and chaotic interactions in complex traffic scenarios, resulting in insufficient interpretability and reliability of the decision-making model.
[0004] Other solutions also improve traffic efficiency and safety by using self-attention networks and priority lists in specific scenarios to assess and optimize vehicle actions. However, their applicability in other complex traffic scenarios is also insufficient, and it is difficult to quantify the degree of collaboration. Summary of the Invention
[0005] In view of this, the present invention provides a multi-vehicle collaborative decision-making method, apparatus, device, storage medium and program product to solve the problem of poor collaborative decision-making ability of the prior art in complex traffic scenarios.
[0006] In a first aspect, the present invention provides a multi-vehicle collaborative decision-making method, wherein the time span of the environment in which the multiple vehicles are located includes multiple time steps, and the method includes: performing the following steps at each time step:
[0007] Based on the initial environmental state, each vehicle's individual observations of other vehicles are determined. These individual observations are then input into a pre-trained initial policy network to obtain the target action for each vehicle.
[0008] Determine the individual reward for each vehicle after performing the corresponding target action in the initial environmental state. Based on the individual rewards, construct the individual advantage function for each vehicle, the neighborhood advantage function for each vehicle's neighborhood, and the global advantage function for the environment in which multiple vehicles are located.
[0009] Based on the initial collaborative model, the individual advantage function and the neighborhood advantage function are coordinated. The initial policy network is then optimized using the obtained collaborative advantage function and the individual advantage function to obtain the target policy network.
[0010] Based on the cooperative advantage function, the global advantage function, the initial policy network, and the target policy network, the initial cooperative model is optimized to obtain the target cooperative model;
[0011] Based on the target policy network and target cooperation model, the initial policy network and initial cooperation model for the next time step are determined so that each vehicle can determine the target action for the next time step.
[0012] This application utilizes each vehicle's individual observations of other vehicles to quantify vehicle interactions in complex traffic scenarios and combines this with a pre-trained initial policy network to obtain the target actions performed by each vehicle under individual observations. By determining the individual reward for each vehicle after performing the corresponding target action in the initial environmental state, an individual advantage function for each vehicle, a neighborhood advantage function for the local neighborhood, and a global advantage function for the global environment are constructed to measure the advantages that a vehicle gains from performing the target action at the individual, neighborhood, and global levels. Individual decision optimization is performed using the individual advantage function, and the initial collaborative model is used to coordinate the individual advantage function and the neighborhood advantage function for local neighborhood decision optimization, resulting in the target policy network, which can make better decisions in complex traffic scenarios. Finally, the initial collaborative model is optimized at the global level, thereby continuously improving vehicle decisions in the local neighborhood.
[0013] In some optional implementations, the individual advantage function and the neighborhood advantage function are coordinated based on the initial coordination model, including:
[0014] Determine the initial collaboration factors for the initial collaboration model;
[0015] The individual advantage function and the neighborhood advantage function are weighted and superimposed using the initial synergy factor to obtain the synergy advantage function.
[0016] This application utilizes an initial synergy factor to coordinate the individual advantage function and the neighborhood advantage function, enabling individual vehicles to allocate reward preferences among themselves and other individuals. This avoids excessive individual gains affecting neighborhood rewards or vice versa, thereby achieving reward equilibrium.
[0017] In some optional implementations, the initial policy network is optimized using the obtained cooperative advantage function and individual advantage function to obtain the target policy network, including:
[0018] An individual objective function is constructed based on the individual advantage function. The initial policy network is then preliminarily optimized using the individual objective function to obtain the intermediate policy network.
[0019] A cooperative objective function is constructed based on the cooperative advantage function; wherein, the optimization parameters of the cooperative objective function include the intermediate policy parameters of the intermediate policy network and the initial cooperative factor.
[0020] By fixing the initial cooperative factor in the cooperative objective function, the intermediate policy parameters of the intermediate policy network are optimized to obtain the target policy network.
[0021] This application first uses an individual objective function to perform preliminary optimization of the policy network at the individual level, and then uses a cooperative objective function to further optimize the policy network at the neighborhood level, resulting in a target policy network that allows both individual and neighborhood vehicles to achieve higher returns.
[0022] In some optional implementations, the initial cooperative model is optimized based on the cooperative advantage function, the global advantage function, the initial policy network, and the target policy network to obtain the target cooperative model, including:
[0023] Determine the initial policy parameters and the target policy parameters;
[0024] Based on the cooperative advantage function, the global advantage function, the initial policy parameters, and the target policy parameters, a global objective function is constructed; wherein, the optimization parameters of the global objective function include the initial cooperative factor.
[0025] The initial collaborative model is optimized using a global objective function to obtain the target collaborative model.
[0026] This application constructs a global objective function with optimization parameters as initial collaborative factors, and optimizes the collaborative factors at the global level in order to optimize the decision-making strategies of neighboring vehicles, thereby achieving the goal of local collaborative optimization of each vehicle strategy and global collaborative optimization of the collaborative factors in each local collaboration.
[0027] In some optional implementations, after each vehicle performs the corresponding target action in the initial environmental state, the initial environmental state is updated to the target environmental state; based on individual rewards, an individual advantage function is constructed for each vehicle, including:
[0028] Based on the individual reward of each vehicle, we obtain the first individual value function of each vehicle in the target environment state and the second individual value function in the initial environment state;
[0029] For each vehicle, an individual advantage function is obtained based on the vehicle's individual reward, the first volume value function, and the second volume value function; whereby the individual advantage function is used to characterize the advantage that the vehicle gains from performing the corresponding target action in the initial environmental state.
[0030] This application constructs a first individual value function under the target environmental state and a second individual value function under the initial environmental state, thereby combining the individual reward, the first individual value function and the second individual value function to obtain an individual advantage function. This function evaluates the individual advantage brought about by the vehicle performing the corresponding target action to change the initial environmental state to the target environmental state, so as to improve individual returns.
[0031] In some optional implementations, the vehicle lifecycle includes a start time step upon entering the environment and an end time step upon leaving the environment; based on the individual reward of each vehicle, a first individual value function for each vehicle in the target environment state and a second individual value function in the initial environment state are obtained, including:
[0032] For each vehicle, based on its individual reward, calculate the average discounted reward for that vehicle from the next time step to the end time step, and obtain the first volume value function;
[0033] For each vehicle, based on its individual reward, calculate the average discounted reward for that vehicle from the current time step to the end time step, thus obtaining the second body value function.
[0034] This application calculates the average discounted reward from the next time step to the end of the vehicle's lifecycle to obtain the first volume value function corresponding to the vehicle performing the corresponding target action in the initial environmental state. By calculating the average discounted reward from the current time step to the end of the vehicle's lifecycle, the second volume value function corresponding to the vehicle not performing the corresponding target action in the initial environmental state is evaluated, so as to calculate the advantage of performing the action compared to not performing the action.
[0035] In some optional implementations, a neighborhood advantage function is constructed for each vehicle's neighborhood based on individual rewards, including:
[0036] Identify the neighboring vehicles within the neighborhood of the target vehicle among multiple vehicles. Based on the individual rewards of the target vehicle and its neighboring vehicles, obtain the neighborhood reward of the neighborhood of the target vehicle, the first neighborhood value function under the target environment state, and the second neighborhood value function under the initial environment state.
[0037] Based on the neighborhood reward, the first neighborhood value function, and the second neighborhood value function, the neighborhood advantage function of the neighborhood where the target vehicle is located is obtained; where the neighborhood advantage function is used to characterize the advantage brought by the target vehicle and neighboring vehicles to perform the corresponding target action in the initial environmental state.
[0038] This embodiment constructs a first neighborhood value function under the target environment state and a second neighborhood value function under the initial environment state, thereby combining the neighborhood reward, the first individual neighborhood function and the second individual neighborhood function to obtain a neighborhood advantage function. This function evaluates the individual advantage brought about by local vehicles performing corresponding target actions to change the initial environment state to the target environment state, so as to improve local neighborhood benefits.
[0039] In some alternative implementations, based on the initial environmental states, each vehicle's individual observations of other vehicles are determined, including:
[0040] Based on the initial environmental state, the state characteristics of each vehicle are obtained;
[0041] For each vehicle, a query vector is generated based on the vehicle's state characteristics, and a key vector and a value vector are generated based on the state characteristics of other vehicles.
[0042] For each vehicle, an attention weight vector is calculated based on the similarity between the vehicle's query vector and key vector. An attention matrix is then calculated based on the attention weight vector and the value vector. The attention matrix includes the vehicle's individual observations of other vehicles.
[0043] In this application, since collaborative decision-making involves the interaction between each vehicle and other vehicles, the vehicle's state features are transformed into a query vector. By calculating the similarity between the query vector and a key vector containing descriptive features of all other vehicles, the vehicle's attention to all other vehicles is evaluated, thus obtaining an attention weight vector. This attention weight vector is then used to weight the value vector containing the value features of all other vehicles to obtain the vehicle's individual observations of other vehicles. This filters out redundant information in complex traffic scenarios, allowing the vehicle to consider only the key information relevant to the decision through the attention mechanism.
[0044] Secondly, the present invention provides a multi-vehicle collaborative decision-making device, wherein the time span of the environment in which the multiple vehicles are located includes multiple time steps, and at each time step the device includes:
[0045] The first processing module is used to determine each vehicle's individual observations of other vehicles based on the initial environmental state, input the individual observations into the pre-trained initial policy network, and obtain the target action of each vehicle.
[0046] The second processing module is used to determine the individual reward of each vehicle after performing the corresponding target action in the initial environmental state. Based on the individual reward, the individual advantage function of each vehicle, the neighborhood advantage function of each vehicle's neighborhood, and the global advantage function of the environment in which multiple vehicles are located are constructed.
[0047] The third processing module is used to coordinate the individual advantage function and the neighborhood advantage function based on the initial cooperative model, and to optimize the initial policy network using the obtained cooperative advantage function and the individual advantage function to obtain the target policy network.
[0048] The fourth processing module is used to optimize the initial cooperative model based on the cooperative advantage function, the global advantage function, the initial policy network, and the target policy network to obtain the target cooperative model.
[0049] The fifth processing module is used to determine the initial policy network and initial cooperation model for the next time step based on the target policy network and the target cooperation model, so that each vehicle can determine the target action for the next time step.
[0050] Thirdly, the present invention provides a computer device, comprising:
[0051] The memory and processor are interconnected and communicate with each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the multi-vehicle cooperative decision-making method described in the first aspect or any of its corresponding embodiments.
[0052] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the multi-vehicle cooperative decision-making method described in the first aspect or any corresponding embodiment thereof.
[0053] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the multi-vehicle cooperative decision-making method described in the first aspect or any corresponding embodiment thereof.
[0054] The beneficial effects of this invention are as follows:
[0055] This application utilizes each vehicle's individual observations of other vehicles to quantify vehicle interactions in complex traffic scenarios and combines this with a pre-trained initial policy network to obtain the target actions performed by each vehicle under individual observations. By determining the individual reward for each vehicle after performing the corresponding target action in the initial environmental state, an individual advantage function for each vehicle, a neighborhood advantage function for the local neighborhood, and a global advantage function for the global environment are constructed to measure the advantages that a vehicle gains from performing the target action at the individual, neighborhood, and global levels. Individual decision optimization is performed using the individual advantage function, and the initial collaborative model is used to coordinate the individual advantage function and the neighborhood advantage function for local neighborhood decision optimization, resulting in the target policy network, which can make better decisions in complex traffic scenarios. Finally, the initial collaborative model is optimized at the global level, thereby continuously improving vehicle decisions in the local neighborhood.
[0056] This application naturally explains the interaction between the vehicle and other vehicles by outputting the attention matrix around the vehicle, identifying the neighboring vehicles with the highest probability of interaction, and greatly improving the interpretability and reliability of collaborative decision-making. Through local collaboration, the agent vehicle can choose the appropriate role positioning according to different driving conditions, better understand the environment and learn strategies, thereby reducing model errors and improving the level of collaborative decision-making. Through global collaborative optimization, the process of locally collaboratively optimizing the strategy of each agent vehicle and globally collaboratively optimizing the collaborative factors in local collaboration is realized, continuously improving the agent strategy and enhancing the generalization ability and training stability of the algorithm. Attached Figure Description
[0057] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0058] Figure 1 This is a flowchart illustrating a multi-vehicle collaborative decision-making method according to an embodiment of the present invention;
[0059] Figure 2 This is a flowchart illustrating another multi-vehicle collaborative decision-making method according to an embodiment of the present invention;
[0060] Figure 3 This is a schematic diagram of a model architecture based on interactive attention according to an embodiment of the present invention;
[0061] Figure 4 This is a schematic diagram of the architecture of a vehicle attention head according to an embodiment of the present invention;
[0062] Figure 5 This is a schematic diagram of a local collaborative framework according to an embodiment of the present invention;
[0063] Figure 6 This is a schematic diagram of a global collaboration framework according to an embodiment of the present invention;
[0064] Figure 7 This is a schematic diagram of a multi-vehicle collaborative decision-making framework according to an embodiment of the present invention;
[0065] Figure 8 This is a structural block diagram of a multi-vehicle collaborative decision-making device according to an embodiment of the present invention;
[0066] Figure 9 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation
[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0068] According to an embodiment of the present invention, a multi-vehicle cooperative decision-making method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0069] This embodiment provides a multi-vehicle collaborative decision-making method, wherein the time span of the environment in which the multiple vehicles are located includes multiple time steps. Figure 1 This is a flowchart of the multi-vehicle cooperative decision-making method according to an embodiment of the present invention at each time step, such as... Figure 1 As shown, the process includes the following steps:
[0070] Step S101: Based on the initial environment state, determine the individual observations of each vehicle regarding other vehicles, input the individual observations into the pre-trained initial policy network, and obtain the target action of each vehicle.
[0071] Specifically, the multi-vehicle decision-making problem in an environment can be abstracted as a set of distributed Markov decision processes, using Markov tuples. The time span of this environment Divided into multiple time steps , This is the set of indices for all vehicles in the environment, representing the set of vehicles in the environment at each time step. ,in, Indicates the current time step The number of existing vehicles in China Represents the environmental state space. Represents the environment action space. Represents the state transition distribution. For reward space, It is the initial state distribution (used for environment state initialization). It is an observation space that includes the observation of all individual vehicles, and , Let this be the observation function space containing the local observation functions of all vehicles. It is a discount factor (which can be set in advance).
[0072] It should be noted that the discount factor In the Markov decision process, this is a core parameter reflecting the vehicle's preference for current and future rewards, used to balance them. When the value approaches 1, the vehicle tends to execute actions that result in future consequences of the target action; when... When the odds approach zero, vehicles focus more on immediate rewards and may engage in more short-sighted behaviors. In this embodiment, a discount factor is used when calculating the individual advantage function of each vehicle, the neighborhood advantage function of each vehicle's neighborhood, and the global advantage function of the environment in which multiple vehicles are located. Perform reward balancing. For example, when calculating the individual advantage function and the neighborhood advantage function, a discount factor is used. A discount factor of 0.99 can be used when calculating the global advantage function. You can choose 1, but the specific value can be set based on your preference for current and future rewards.
[0073] In step S101, the current time step initial environmental state Initial environment state A collection of vehicles The state characteristics of each vehicle in the system consist of coordinates, speed, acceleration, driving intention (e.g., light signals), driving direction, and driving actions. These state characteristics can be acquired through various sensors such as LiDAR, speed sensors, acceleration sensors, and cameras. Through its own local observation function and initial environment state Get the current time step Individual observations of other vehicles .
[0074] It should be noted that individual vehicle observation includes observation data of the vehicle's own state characteristics and those of other vehicles in the surrounding environment. The vehicle's own observation data, i.e., its own state characteristics, can be acquired using various sensors. As for the vehicle's observation data of other vehicles, since the state characteristics of each vehicle constitute the current time step... initial environmental state For a vehicle, the state characteristics of other vehicles in the real environment can also be observed by the vehicle using various sensors. Individual observation is used to characterize the vehicle's observation of its own state characteristics and those of other vehicles. How to obtain individual observations using the vehicle's own local observation function is described in the following embodiments and will not be repeated here.
[0075] In step S101, each vehicle A pre-trained initial policy network was deployed. The initial policy network is based on individual observations. Select the current time step target action ,Right now Among them, the target action From the vehicle Action space To choose, that is Current time step Joint action space A collection of vehicles The motion space of each vehicle The union of sets, i.e. .
[0076] It should be further noted that the policy network is based on policy parameter parameterization. The policy network can adopt a neural network model, and the policy parameters of the policy network are all the model parameters in the neural network model. Taking the neural network model as an example, by constructing a training dataset (including a candidate observation set and the real action label corresponding to each candidate observation), the candidate observation set is used as the input of the neural network model to obtain the action output by the neural network model. The loss value is calculated by the output action of the neural network model and the real action label. The model parameters of the neural network model are continuously optimized using the loss value, so that the policy network can output the optimal action policy (i.e., the target action) based on vehicle observations in deep learning, thus obtaining the pre-trained initial policy network.
[0077] To facilitate understanding, this embodiment distinguishes between the policy networks before and after optimization at each time step. For the current time step... In other words, it will be at the current time step The unoptimized policy network in the current time step is called the initial policy network, and it will be used at the current time step. The policy network trained in subsequent processes is called the target policy network. Compared to the initial policy network, the policy parameters of the target policy network have been optimized. Similarly, in this embodiment, the current time step... The unoptimized policy parameters are called the initial policy parameters, and will be used at the current time step. The policy parameters trained in subsequent processes are called target policy parameters.
[0078] In other words, for the current time step In other words, the current time step The initial policy network used is the one from the previous time step. A trained target policy network.
[0079] Step S102: Determine the individual reward for each vehicle after performing the corresponding target action in the initial environmental state. Based on the individual rewards, construct the individual advantage function of each vehicle, the neighborhood advantage function of each vehicle's neighborhood, and the global advantage function of the environment in which multiple vehicles are located.
[0080] It's important to understand that reward, value function, and advantage function are fundamental concepts in deep learning. Reward refers to the real-time feedback after performing the target action; if the target action is good, a positive reward is given, otherwise a negative reward is given. The value function is used to characterize the expected long-term cumulative reward that performing the target action can bring, which is beneficial for the vehicle to make long-term decisions. The advantage function is used to characterize the advantage that performing the target action brings compared to not performing the target action.
[0081] The multi-vehicle collaborative decision-making system in this embodiment is divided into three levels: individual level (referring to a single vehicle), local level (referring to neighboring vehicles that are close to the vehicle itself), and global level (referring to all vehicles). This application evaluates the system based on the reward, value function, and advantage function corresponding to the above three levels, so that each vehicle agent can make a decision that is better than itself without affecting the benefits brought to the local and global levels.
[0082] In this embodiment, the reward for the vehicle performing the target action is used as the individual reward, the reward for all vehicles in the vehicle's neighborhood performing the corresponding target action is used as the neighborhood reward, and the reward for all vehicles globally performing the corresponding target action is used as the global reward. Furthermore, the value function of the vehicle performing the target action is used as the individual value function, the value function of all vehicles in the neighborhood performing the corresponding target action is used as the neighborhood value function, and the value function of all vehicles globally performing the corresponding target action is used as the global value function. Additionally, the advantage function of the vehicle performing the target action is used as the individual advantage function, the advantage function of all vehicles in the neighborhood performing the corresponding target action is used as the neighborhood advantage function, and the advantage function of all vehicles globally performing the corresponding target action is used as the global advantage function.
[0083] Specifically, in determining the vehicle set Target actions of all vehicles These actions will then form a combined action. , This joint action includes the target action for each vehicle. Vehicle set Each vehicle in the initial environmental state Execute the corresponding target action. Afterwards, the initial environmental state Change to the target environment state (As the initial environment state for the next time step). That is, in the initial environment state The following joint action was carried out. The transition and update of environmental states follow a state transition distribution. .
[0084] In step S102, the reward space Including vehicles reward function Through the reward function Determine the vehicle At the current time step Perform the target action Subsequent individual rewards ,Right now .
[0085] In some embodiments, a reward function can be defined for each vehicle. Driven by rewards Speed Reward Destination rewards and penalties It consists of four parts, namely ,in, Indicates the driving reward weight, Indicates the speed reward weight. Indicates the destination reward weight. This indicates the weight of the penalty item.
[0086] For example, driving rewards can be determined by judging the vehicle Perform the target action The subsequent driving parameters are obtained by checking whether they meet the set driving objectives. For example, it can be determined whether the vehicle is driving along the current driving route, which includes a series of consecutive driving waypoints. If the vehicle is currently driving to the corresponding driving waypoint, it means that the vehicle is driving along the current driving route, and a driving reward of 1 is given; otherwise, a driving reward of 0 is given.
[0087] For example, speed rewards can be determined by judging the vehicle Perform the target action Do the subsequent velocity parameters match the target action? The corresponding instructions are used to obtain the data. For example, the target action. To accelerate to a certain target speed, determine the vehicle's speed. Perform the target action If the subsequent speed parameter is the target speed, a speed bonus of 1 is given; otherwise, a speed bonus of 0 is given.
[0088] For example, destination rewards By judging the vehicle Perform the target action Whether or not the vehicle reaches its destination will be determined later. For example, if the vehicle... Perform the target action If the current location is the location of the destination, then the destination will be rewarded with 1; otherwise, the destination will be rewarded with 0.
[0089] For example, the weight of the penalty item can be determined by judging the vehicle. Perform the target action Whether a collision occurred afterward is determined by this. For example, if the vehicle... Perform the target action If a collision is detected, a positive penalty of 1 is applied (i.e., a penalty is imposed); otherwise, a negative penalty of -1 is applied (i.e., no penalty is imposed).
[0090] It should be noted that the driving reward weight, speed reward weight, destination reward weight, and penalty weight are used to measure the degree of importance attached to driving rewards, speed rewards, destination rewards, and penalties, respectively. For example, if driving rewards are valued more, the driving reward weight can be higher than the other reward weights. The specific weight can be set according to the actual preference for different rewards and penalties, as long as the penalty weight is negative. For example, the driving reward weight... The speed reward weight is 0.1. The speed reward weight is 0.1. The penalty term weight is 0.1. It is -0.1.
[0091] Furthermore, through each vehicle At the current time step Perform the target action Subsequent individual rewards The individual advantage function, the neighborhood advantage function, and the global advantage function of each vehicle are calculated. The individual advantage function, the neighborhood advantage function, and the global advantage function are used to characterize the initial environmental state of the individual vehicle, the local neighboring vehicles, and all vehicles in the global environment, respectively. Execute the corresponding target action. Compared to the advantages of not performing the target action.
[0092] Step S103: Based on the initial cooperative model, the individual advantage function and the neighborhood advantage function are coordinated, and the initial policy network is optimized using the obtained cooperative advantage function and the individual advantage function to obtain the target policy network.
[0093] Specifically, the initial collaborative model is parameterized by the initial collaborative factor, which is used to collaborate the individual advantage function and the neighborhood advantage function to obtain the collaborative advantage function.
[0094] It should be noted that the collaborative model is based on collaborative factor parameterization. The collaborative model is used to weight the individual advantage function of a single vehicle and the neighborhood advantage function of its local area to obtain the collaborative advantage function. This collaborative advantage function then finds a balance between individual and local neighborhood gains, facilitating the vehicle to make decisions superior to both individual and local gains. For ease of understanding, this embodiment distinguishes between the collaborative model before and after optimization at each time step. For the current time step... In other words, it will be at the current time step The unoptimized collaborative model in the current time step is called the initial collaborative model, and it will be used in the current time step. The trained collaborative model in subsequent processes is called the target collaborative model. Compared to the collaborative model, the collaborative factor of the target collaborative model has been optimized. Similarly, in this embodiment, the current time step... The unoptimized collaborative factor is called the initial collaborative factor, and it will be used at the current time step. The co-factors trained in subsequent processes are called target co-factors.
[0095] In other words, for the current time step In other words, the current time step The initial collaborative model used is the one from the previous time step. A well-trained target collaboration model.
[0096] This application optimizes the policy network at the individual vehicle level using an individual advantage function to improve individual gains, and uses a collaborative advantage function to perform collaborative optimization at the local neighborhood level, thereby learning a policy that can balance individual gains and local neighborhood gains. This enables the target policy network to better understand the environment in complex traffic scenarios, thereby reducing decision-making errors and improving decision-making level.
[0097] Step S104: Based on the cooperative advantage function, the global advantage function, the initial policy network, and the target policy network, the initial cooperative model is optimized to obtain the target cooperative model.
[0098] Specifically, based on step S103, the decision-making scope is expanded from the local neighborhood to the global environment. The initial cooperative model is globally optimized by integrating the unoptimized initial policy network, the optimized target policy network, the cooperative advantage function of the local neighborhood, and the global advantage function of the global environment. This optimizes the cooperative factors in each local cooperative at the global level, enabling the target cooperative model to achieve better local cooperation in each local neighborhood.
[0099] Step S105: Based on the target policy network and the target cooperation model, determine the initial policy network and the initial cooperation model for the next time step, so that each vehicle can determine the target action for the next time step.
[0100] Specifically, the target policy network is used as the initial policy network for the next time step, and the target cooperative model is used as the initial cooperative model for the next time step. This allows each vehicle in the environment to make better decision-making actions in future time steps by utilizing the policy network and local cooperation, thereby ensuring driving safety.
[0101] The multi-vehicle cooperative decision-making method provided in this embodiment quantifies vehicle interactions in complex traffic scenarios by utilizing each vehicle's individual observations of other vehicles. It then combines this with a pre-trained initial policy network to obtain the target action performed by each vehicle under its individual observations. By determining the individual reward for each vehicle after performing the corresponding target action in the initial environmental state, it constructs an individual advantage function for each vehicle, a neighborhood advantage function for its local neighborhood, and a global advantage function for the global environment. This measures the advantages that a vehicle gains after performing the target action at the individual, neighborhood, and global levels. Individual decision optimization is performed using the individual advantage function, and the initial cooperative model is used to coordinate the individual advantage function and the neighborhood advantage function for local neighborhood decision optimization, resulting in a target policy network that can make better decisions in complex traffic scenarios. Finally, the initial cooperative model is optimized at the global level, thereby continuously improving vehicle decisions in the local neighborhood.
[0102] It should be further noted that the "multi-vehicle" in this application refers to a multi-vehicle intelligent agent, which is capable of autonomously perceiving the environment, making dynamic decisions, and executing corresponding actions to achieve its goals. For ease of understanding, this embodiment is illustrated using a collaborative decision-making scenario of multi-vehicle intelligent agents as an example. However, in practical applications, this application is not limited to this. In addition to multi-vehicle intelligent agents, the multi-agent collaborative system can also be other types of multi-agent collaborative systems. The collaborative decision-making method of this application can also be applied to other multi-agent collaborative systems. Any variations that can be conceived by those skilled in the art should fall within the protection scope of this application.
[0103] For example, in logistics and delivery scenarios, multiple delivery robots need to collaborate to plan routes. The collaborative decision-making method of this application can also be used to make collaborative decisions through local and global collaboration, determine the target action of each delivery robot agent, improve the overall delivery speed, and reduce path conflicts.
[0104] For example, in the simulation system of autonomous vehicles, it is necessary to simulate and test the driving behavior of multiple vehicles. During the test, each virtual vehicle intelligent agent is controlled to cooperate according to the collaborative decision-making method of this application, so as to determine the target action to be executed at each time step and improve the accuracy of the simulation test.
[0105] For example, in a game scenario, a group of multiple non-player virtual characters can perform corresponding actions by adopting the collaborative decision-making method of this application, thereby simulating realistic social behavior and action strategies, making the game experience and immersion of real players more intense.
[0106] This embodiment provides a multi-vehicle collaborative decision-making method, wherein the time span of the environment in which the multiple vehicles are located includes multiple time steps. Figure 2 This is a flowchart of the multi-vehicle cooperative decision-making method according to an embodiment of the present invention at each time step, such as... Figure 2 As shown, the process includes the following steps:
[0107] Step S201: Based on the initial environment state, determine the individual observations of each vehicle regarding other vehicles, input the individual observations into the pre-trained initial policy network, and obtain the target action of each vehicle.
[0108] In some embodiments, individual observations of each vehicle relative to other vehicles are determined through the following steps:
[0109] Step a1: Based on the initial environmental state, obtain the state characteristics of each vehicle.
[0110] Specifically, collaborative decision-making among vehicles involves the interaction between each vehicle and other vehicles, thus requiring an environmental state space. It should at least include a description of the vehicle's information about every other vehicle in the vicinity. This should be achieved through the environmental state space. Get the current time step initial environmental state Thus, the vehicle set is obtained. The state characteristics of each vehicle, such as the state parameters acquired by various sensors.
[0111] Step a2: For each vehicle, generate a query vector based on the vehicle's state features, and generate key vectors and value vectors based on the state features of other vehicles.
[0112] like Figure 3 As shown, this embodiment provides a model architecture based on interactive attention. This model architecture consists of several linearly identical encoders, a vehicle attention head, and a linear decoder. Each vehicle attention head architecture is as follows: Figure 4 As shown, For linear projection of the query vector, This is a linear projection of the key vector (primarily for extracting the index of state features). This is a linear projection of the value vector (primarily extracting the values of state features). , , They are all linear layers. For column vector dimensions, The dimension is the row vector, and the key vector is the key vector. Sum value vector It is composed of the state features of all vehicles, while the query vector is... Provided solely by the vehicle itself.
[0113] Step a3: For each vehicle, calculate the attention weight vector based on the similarity between the query vector and the key vector of that vehicle, and calculate the attention matrix based on the attention weight vector and the value vector; wherein, the attention matrix includes the individual observations of that vehicle on other vehicles.
[0114] Specifically, the vehicle's state features are first obtained through its embedded linear projection. Calculate and issue a single query vector Then this query vector With a set of key vectors containing descriptive features (i.e., indices) for each vehicle. Comparison, key vectors It is through shared linear projection It was calculated. The query was performed using the calculation. With any key dot product between Evaluation Inquiry With any key The similarity between them. These similarities Then, from the inverse square root of the row vector dimension Scaling is performed, and finally the softmax function is applied. Normalization is performed across vehicles to obtain the attention weight vector. The attention matrix includes the attention values of the vehicle to other vehicles in the surrounding area.
[0115] Furthermore, by calculating the attention weight vector AND value vector The product of the two is the attention matrix. , where the value vector Each value Both are linear projections of a shared value vector. This was calculated. It should be noted that the attention matrix includes each vehicle's individual observations of other vehicles. That is, the calculation process of the attention matrix mentioned above is the observation process of the vehicle's local observation function.
[0116] In this embodiment, since collaborative decision-making involves the interaction between each vehicle and other vehicles, the vehicle's state features are transformed into a query vector. The similarity between the query vector and a key vector containing descriptive features of all other vehicles is calculated to evaluate the vehicle's attention to all other vehicles, thus obtaining an attention weight vector. This attention weight vector is then used to weight the value vector containing the value features of all other vehicles, yielding the vehicle's individual observations of other vehicles. This filters out redundant information in complex traffic scenarios, allowing the vehicle to consider only key information relevant to the decision through the attention mechanism.
[0117] For a detailed explanation of how to input individual observations into a pre-trained initial policy network to obtain the target action for each vehicle, please refer to [link to documentation]. Figure 1 The detailed description of step S101 in the illustrated embodiment will not be repeated here.
[0118] Step S202: Determine the individual reward for each vehicle after performing the corresponding target action in the initial environmental state. Based on the individual rewards, construct the individual advantage function for each vehicle, the neighborhood advantage function for each vehicle's neighborhood, and the global advantage function for the environment in which multiple vehicles are located.
[0119] In some embodiments, the individual advantage function for each vehicle is constructed through the following steps:
[0120] Step b1: Based on the individual reward of each vehicle, obtain the first individual value function of each vehicle in the target environment state and the second individual value function in the initial environment state.
[0121] This application defines the environment round as... ,in It is the time span of the environment, vehicles The lifecycle includes the start time step when entering the environment. and the end time step when leaving the environment Therefore, an environment round contains a set of vehicle rounds. .vehicle Initial policy network From policy parameters Parameterization, then the vehicle In an environment round, an individual objective is defined as the sum of discounted individual rewards, i.e., the vehicle objective. Individual returns during their life cycle ,in, It expresses expectation.
[0122] It should be noted that the initial policy network For a pre-trained neural network model, the initial policy network can be used. Before deployment to vehicles, the neural network model can be trained to optimize its parameters; the trained model parameters constitute the initial policy network. strategy parameters In the initial policy network After deployment on the vehicle, the initial policy network The system can select the target action for a vehicle based on individual observations and continuously adjust the initial policy network in subsequent processes. strategy parameters Optimize.
[0123] In this embodiment, the Independent Proximal Policy Optimization (IPPO) algorithm is used to maximize the individual objective of each vehicle. For each vehicle, based on that vehicle... Individual rewards Calculate the vehicle From the next time step End time step The average discount reward will yield the first volume value function. For each vehicle, based on that vehicle... Individual rewards Calculate the vehicle From the current time step End time step The average discount reward will yield the second volume value function. .
[0124] This embodiment calculates the average discounted reward from the next time step to the end of the vehicle's lifecycle to obtain the first volume value function corresponding to the vehicle performing the corresponding target action in the initial environmental state. By calculating the average discounted reward from the current time step to the end of the vehicle's lifecycle, the second volume value function corresponding to the vehicle not performing the corresponding target action in the initial environmental state is evaluated, so as to calculate the advantage of performing the action compared to not performing the action.
[0125] Step b2: For each vehicle, obtain the vehicle's individual advantage function based on the vehicle's individual reward, the first body value function, and the second body value function.
[0126] Specifically, calculate vehicles At the current time step Individual advantage function ,in, Indicates that the target action will not be performed. Thus, the individual advantage function is utilized. Characterizing vehicles In the initial environmental state Execute the corresponding target action. The advantages it brings. Among them, At this point, 0.99 can be used.
[0127] This application embodiment constructs a first individual value function under the target environment state and a second individual value function under the initial environment state, thereby combining the individual reward, the first individual value function and the second individual value function to obtain an individual advantage function, and evaluates the individual advantage brought by the vehicle performing the corresponding target action to change the initial environment state to the target environment state, so as to improve the individual gain.
[0128] In some embodiments, the neighborhood advantage function for each vehicle's neighborhood is constructed through the following steps:
[0129] Step c1: Determine the neighboring vehicles in the neighborhood of the target vehicle among the multiple vehicles. Based on the individual rewards of the target vehicle and the neighboring vehicles, obtain the neighborhood reward of the neighborhood of the target vehicle, the first neighborhood value function in the target environment state, and the second neighborhood value function in the initial environment state.
[0130] Specifically, the target vehicle For any one of the multiple vehicles, pass through the target vehicle The corresponding attention weight vector is used to filter target vehicles. Neighboring vehicles within the target vehicle's neighborhood, for example, other vehicles whose attention weight exceeds the attention weight threshold are considered neighboring vehicles. This is achieved by calculating the target vehicle's... Target vehicles within the surrounding area And the average of the individual rewards of neighboring vehicles, to obtain the target vehicle Neighborhood rewards ,in, N The number of vehicles in the neighboring area.
[0131] Furthermore, such as Figure 5 As shown, in order to incorporate collaborative factors into the training process to improve the overall collaborative decision-making performance, a neighborhood value function is used to approximate the neighborhood reward. Total discounts According to the target vehicle Neighborhood rewards Calculate the target vehicle The surrounding area is in the target environment state First neighborhood value function and in the initial environmental state The second neighborhood value function .
[0132] Step c2: Based on the neighborhood reward, the first neighborhood value function, and the second neighborhood value function, obtain the neighborhood advantage function of the neighborhood where the target vehicle is located.
[0133] Specifically, see again Figure 5 Calculate the target vehicle Neighborhood dominance function Thus, the neighborhood advantage function is utilized. Characterizing the target vehicle and neighboring vehicles in the initial environmental state The advantages of executing the corresponding target action. Among them, At this point, 0.99 can be used.
[0134] This application constructs a first neighborhood value function under the target environment state and a second neighborhood value function under the initial environment state, thereby combining the neighborhood reward, the first individual neighborhood function and the second individual neighborhood function to obtain a neighborhood advantage function. This function evaluates the individual advantage brought about by local vehicles performing corresponding target actions to change the initial environment state to the target environment state, so as to improve local neighborhood benefits.
[0135] In some embodiments, an additional global value function is used. To estimate the global reward The value of is then used to calculate the global advantage function. Compared to neighborhood, global expands the vehicle scope from the neighborhood to all vehicles in the environment. At this point, we can choose 1, therefore the global reward is... Global Advantage Function For the specific calculation process, please refer to the detailed steps of neighborhood reward and neighborhood advantage function, which will not be repeated here.
[0136] Step S203: Based on the initial cooperative model, the individual advantage function and the neighborhood advantage function are coordinated, and the initial policy network is optimized using the obtained cooperative advantage function and the individual advantage function to obtain the target policy network.
[0137] Specifically, step S203 includes:
[0138] Step S2031: Determine the initial collaboration factors of the initial collaboration model.
[0139] Specifically, the initial collaborative model consists of initial collaborative factors. Parameterization, current time step The initial collaborative model, i.e., the previous time step The trained target collaborative model.
[0140] Step S2032: The individual advantage function and the neighborhood advantage function are weighted and superimposed using the initial synergy factor to obtain the synergy advantage function.
[0141] Specifically, see again Figure 5 The individual advantage function and the neighborhood advantage function are weighted and superimposed using the initial synergy factor, resulting in the synergy advantage function. The calculation formula is: .
[0142] This embodiment utilizes an initial synergy factor to coordinate the individual advantage function and the neighborhood advantage function, enabling individual vehicles to allocate reward preferences among themselves and other individuals. This avoids excessive individual gains affecting neighborhood rewards or vice versa, thereby achieving reward equilibrium.
[0143] Step S2033: Construct an individual objective function based on the individual advantage function, and use the individual objective function to perform preliminary optimization on the initial policy network to obtain the intermediate policy network.
[0144] Specifically, for each vehicle, the gradient function of the individual target is calculated using the policy gradient method. The pruning importance factor in the Proximal Policy Optimization (PPO) algorithm is used to mitigate the distribution shift that occurs after policy updates, and the final individual objective function is constructed as follows: ,in, These are hyperparameters (which can be set in advance). The gain of the new strategy compared to the old strategy. , Indicates the old strategy, For the new strategy, This is a truncation operation.
[0145] Furthermore, by defining the individual objective function as... The policy parameters are subjected to gradient ascent to increase the expected return for each individual, thereby initially optimizing the initial policy network and obtaining an intermediate policy network. The specific process of gradient optimization and updating can be found in the detailed descriptions of relevant techniques, and will not be elaborated upon here.
[0146] Step S2034: Construct a collaborative objective function based on the collaborative advantage function.
[0147] Specifically, see again Figure 5 Utilizing the collaborative advantage function Construct a collaborative objective function Among them, the collaborative objective function The optimization parameters include the intermediate policy parameters of the intermediate policy network. and initial cooperability factor .
[0148] Step S2035: By fixing the initial cooperative factor in the cooperative objective function, the intermediate policy parameters of the intermediate policy network are optimized to obtain the target policy network.
[0149] Specifically, the intermediate policy parameters of the intermediate policy network can be determined through stochastic gradient ascent. Further updates are made to optimize local collaborative decision-making. The specific process of stochastic gradient ascent can be found in the detailed descriptions of relevant techniques, and will not be elaborated upon here.
[0150] This application embodiment first uses an individual objective function to perform preliminary optimization of the policy network at the individual level, and then uses a cooperative objective function to further optimize the policy network at the neighborhood level, resulting in a target policy network that allows both individual and neighborhood vehicles to obtain higher returns.
[0151] Step S204: Based on the cooperative advantage function, the global advantage function, the initial policy network, and the target policy network, the initial cooperative model is optimized to obtain the target cooperative model.
[0152] Specifically, step S204 includes:
[0153] Step S2041: Determine the initial policy parameters of the initial policy parameters and the target policy parameters of the target policy parameters. Based on the cooperative advantage function, the global advantage function, the initial policy parameters, and the target policy parameters, construct the global objective function.
[0154] Specifically, the initial policy parameters are expressed as follows: The target policy parameter is expressed as: The global objective of the overall environment is defined as the expected sum of individual rewards for all vehicles within the current round, from the start to the end of their lifecycle. for:
[0155]
[0156] in, Indicates the environment round Current time step All vehicles that exist.
[0157] Specifically, the global objective Decompose into individual global goals Among them, individual global goals for:
[0158]
[0159] What needs to be understood is the sum of individual global goals. , ,in, For vehicles The strategy parameters for maximizing the individual global objective can be used to determine the strategy parameters for each vehicle. It can maximize its individual global goals To achieve the overall goal Maximize, so that each vehicle maximizes The individual's global goal is equivalent to maximizing the global goal. Global reward. As the current time step The average reward for all vehicles, then Vehicles Cumulative global rewards over the lifecycle.
[0160] Furthermore, the gradient of the global objective is calculated using the chain rule. ,Right now , and This can be viewed as two parts that are calculated and processed separately. The first part... That is, policy gradient calculation, where the sample From the initial policy parameters before optimization Generation. Through a first-order Taylor expansion, the second part... ,in, Let be the initial policy gradient, where The learning rate is the learning rate. The learning rate is used to control the convergence speed of the second part of the global objective. If the learning rate is too large, the convergence speed is too fast, and training may become unstable; if the learning rate is too small, the convergence speed is too slow, and it may get stuck in local optima. Setting a reasonable learning rate is crucial. Balancing convergence speed and training stability, for example, the learning rate. Take 0.0003.
[0161] In this way, the calculation of any vehicle can be derived from the two parts of the gradient calculation. The global objective function is The global objective function The optimization parameters include the initial cooperability factor. This global objective function is used to achieve global collaborative optimization.
[0162] Step S2042: Optimize the initial collaborative model using the global objective function to obtain the target collaborative model.
[0163] Specifically, such as Figure 6 As shown, by analyzing about Perform stochastic gradient ascent to update and optimize the initial cooperability factor. The goal of the target collaboration model is to obtain the target collaboration factor, which is to achieve the goal of optimizing the collaboration factor in each local collaboration of each vehicle strategy and optimizing the collaboration factor in each local collaboration of the global collaboration, thereby achieving the purpose of optimizing the decision of the entire environmental system.
[0164] This application embodiment constructs a global objective function with optimization parameters as initial collaborative factors, and optimizes the collaborative factors at the global level in order to optimize the decision-making strategies of neighboring vehicles, thereby achieving the goal of local collaborative optimization of each vehicle strategy and global collaborative optimization of the collaborative factors in each local collaboration.
[0165] Step S205: Based on the target policy network and the target cooperation model, determine the initial policy network and the initial cooperation model for the next time step, so that each vehicle can determine the target action for the next time step. See details for further information. Figure 2 The detailed description of step S105 in the illustrated embodiment will not be repeated here.
[0166] Figure 7 For the multi-vehicle collaborative decision-making framework of this application, see [link to relevant documentation]. Figure 7 This application first encodes the state features of the vehicle and other vehicles, and after linear projection processing, obtains a query vector, a key vector, and a value vector. By calculating the similarity between the query vector and the key vector, an attention weight vector is obtained. The value vector is then weighted using the attention weight vector to obtain an attention weight matrix. By outputting the attention matrix around the vehicle, the interactions between the vehicle and other vehicles are naturally explained, identifying the neighborhood vehicles with the highest probability of interaction, greatly improving the interpretability and reliability of collaborative decision-making.
[0167] At the local collaboration level, a neighborhood advantage function is calculated based on neighborhood rewards within the local neighborhood. An initial collaboration factor is used to collaborate the individual advantage function and the neighborhood advantage function to construct a neighborhood objective function, thereby optimizing and updating the policy network to obtain the target policy network. This enables the intelligent agent vehicle to select appropriate role positioning according to different driving conditions, better understand the environment and learn policies, thus reducing model errors and improving collaborative decision-making.
[0168] At the global cooperation level, a global advantage function is calculated based on the global rewards of all vehicles within the environment. This global advantage function is then used to construct a global objective function, thereby updating the initial cooperation factors. By continuously improving the agent's strategy through local cooperation optimization of each agent's vehicle strategy and global cooperation optimization of the cooperation factors within local cooperation, the algorithm's generalization ability and training stability are enhanced.
[0169] This application addresses the technical problems of insufficient collaborative decision-making ability, poor algorithm generalization ability, and unstable training in existing technologies in complex traffic scenarios by using interactive attention mechanism, reward collaboration based on collaborative factors, and two-layer collaborative decision optimization. It significantly improves the collaborative decision-making level of autonomous vehicles in complex traffic scenarios and enhances driving safety.
[0170] This embodiment also provides a multi-vehicle collaborative decision-making device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0171] This embodiment provides a multi-vehicle collaborative decision-making device, such as... Figure 8 As shown, it includes:
[0172] The first processing module 801 is used to determine each vehicle's individual observations of other vehicles based on the initial environment state, input the individual observations into the pre-trained initial policy network, and obtain the target action of each vehicle.
[0173] The second processing module 802 is used to determine the individual reward of each vehicle after performing the corresponding target action in the initial environmental state, and based on the individual reward, construct the individual advantage function of each vehicle, the neighborhood advantage function of the neighborhood where each vehicle is located, and the global advantage function of the environment where multiple vehicles are located.
[0174] The third processing module 803 is used to coordinate the individual advantage function and the neighborhood advantage function based on the initial cooperative model, and to optimize the initial policy network using the obtained cooperative advantage function and the individual advantage function to obtain the target policy network.
[0175] The fourth processing module 804 is used to optimize the initial cooperative model based on the cooperative advantage function, the global advantage function, the initial policy network, and the target policy network to obtain the target cooperative model;
[0176] The fifth processing module 805 is used to determine the initial policy network and initial cooperation model for the next time step based on the target policy network and the target cooperation model, so that each vehicle can determine the target action for the next time step.
[0177] In some optional implementations, the first processing module 801 is further configured to:
[0178] Based on the initial environmental state, the state characteristics of each vehicle are obtained;
[0179] For each vehicle, a query vector is generated based on the vehicle's state characteristics, and a key vector and a value vector are generated based on the state characteristics of other vehicles.
[0180] For each vehicle, an attention weight vector is calculated based on the similarity between the vehicle's query vector and key vector. An attention matrix is then calculated based on the attention weight vector and the value vector. The attention matrix includes the vehicle's individual observations of other vehicles.
[0181] In some optional implementations, after each vehicle performs the corresponding target action in the initial environmental state, the initial environmental state is updated to the target environmental state; the second processing module 802 is further configured to:
[0182] Based on the individual reward of each vehicle, we obtain the first individual value function of each vehicle in the target environment state and the second individual value function in the initial environment state;
[0183] For each vehicle, an individual advantage function is obtained based on the vehicle's individual reward, the first volume value function, and the second volume value function; whereby the individual advantage function is used to characterize the advantage that the vehicle gains from performing the corresponding target action in the initial environmental state.
[0184] In some optional implementations, the vehicle's lifecycle includes a start time step when entering the environment and an end time step when leaving the environment; the second processing module 802 is further configured to:
[0185] For each vehicle, based on its individual reward, calculate the average discounted reward for that vehicle from the next time step to the end time step, and obtain the first volume value function;
[0186] For each vehicle, based on its individual reward, calculate the average discounted reward for that vehicle from the current time step to the end time step, thus obtaining the second body value function.
[0187] In some optional implementations, the second processing module 802 is further configured to:
[0188] Identify the neighboring vehicles within the neighborhood of the target vehicle among multiple vehicles. Based on the individual rewards of the target vehicle and its neighboring vehicles, obtain the neighborhood reward of the neighborhood of the target vehicle, the first neighborhood value function under the target environment state, and the second neighborhood value function under the initial environment state.
[0189] Based on the neighborhood reward, the first neighborhood value function, and the second neighborhood value function, the neighborhood advantage function of the neighborhood where the target vehicle is located is obtained; where the neighborhood advantage function is used to characterize the advantage brought by the target vehicle and neighboring vehicles to perform the corresponding target action in the initial environmental state.
[0190] In some optional implementations, the third processing module 803 is further configured to:
[0191] Determine the initial collaboration factors for the initial collaboration model;
[0192] The individual advantage function and the neighborhood advantage function are weighted and superimposed using the initial synergy factor to obtain the synergy advantage function.
[0193] In some optional implementations, the third processing module 803 is further configured to:
[0194] An individual objective function is constructed based on the individual advantage function. The initial policy network is then preliminarily optimized using the individual objective function to obtain the intermediate policy network.
[0195] A cooperative objective function is constructed based on the cooperative advantage function; wherein, the optimization parameters of the cooperative objective function include the intermediate policy parameters of the intermediate policy network and the initial cooperative factor.
[0196] By fixing the initial cooperative factor in the cooperative objective function, the intermediate policy parameters of the intermediate policy network are optimized to obtain the target policy network.
[0197] In some alternative implementations, the fourth processing module 804 is further configured to:
[0198] Determine the initial policy parameters and the target policy parameters;
[0199] Based on the cooperative advantage function, the global advantage function, the initial policy parameters, and the target policy parameters, a global objective function is constructed; wherein, the optimization parameters of the global objective function include the initial cooperative factor.
[0200] The initial collaborative model is optimized using a global objective function to obtain the target collaborative model.
[0201] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0202] In this embodiment, the multi-vehicle collaborative decision-making device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0203] This invention also provides a computer device having the above-described features. Figure 8 The multi-vehicle collaborative decision-making device shown.
[0204] Please see Figure 9 , Figure 9 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 9 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 9 Take a processor 10 as an example.
[0205] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPA), or any combination thereof.
[0206] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.
[0207] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0208] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0209] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 9 Taking the example of a connection between China and Israel via a bus.
[0210] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touchscreen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touchscreen.
[0211] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0212] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0213] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A multi-vehicle collaborative decision-making method, characterized in that, The time span of the environment in which the multiple vehicles are located includes multiple time steps, and the method includes: performing the following steps at each time step: Based on the initial environmental state, each vehicle's individual observations of other vehicles are determined, and these individual observations are input into a pre-trained initial policy network to obtain the target action of each vehicle. Determine the individual reward for each vehicle after performing the corresponding target action in the initial environmental state. Based on the individual reward, construct the individual advantage function for each vehicle, the neighborhood advantage function for each vehicle's neighborhood, and the global advantage function for the environment in which multiple vehicles are located. Based on the initial collaborative model, the individual advantage function and the neighborhood advantage function are coordinated, and the initial policy network is optimized using the obtained collaborative advantage function and the individual advantage function to obtain the target policy network. Based on the cooperative advantage function, the global advantage function, the initial policy network, and the target policy network, the initial cooperative model is optimized to obtain the target cooperative model. Based on the target policy network and the target cooperation model, the initial policy network and initial cooperation model for the next time step are determined so that each vehicle can determine the target action for the next time step. The process of determining each vehicle's individual observations of other vehicles based on the initial environmental state includes: Based on the initial environmental state, the state characteristics of each vehicle are obtained; For each vehicle, a query vector is generated based on the vehicle's state characteristics, and a key vector and a value vector are generated based on the state characteristics of other vehicles; For each vehicle, an attention weight vector is calculated based on the similarity between the vehicle's query vector and the key vector. An attention matrix is then calculated based on the attention weight vector and the value vector. The attention matrix includes the vehicle's individual observations of other vehicles. By using the attention weight vector corresponding to the target vehicle, neighboring vehicles within the target vehicle's neighborhood are selected, and other vehicles whose attention weight exceeds the attention weight threshold are considered as neighboring vehicles.
2. The multi-vehicle collaborative decision-making method according to claim 1, characterized in that, The coordination of individual advantage functions and neighborhood advantage functions based on the initial coordination model includes: Determine the initial collaboration factors of the initial collaboration model; The individual advantage function and the neighborhood advantage function are weighted and superimposed using the initial synergy factor to obtain the synergy advantage function.
3. The multi-vehicle collaborative decision-making method according to claim 2, characterized in that, The step of optimizing the initial policy network using the obtained collaborative advantage function and individual advantage function to obtain the target policy network includes: Based on the individual advantage function, an individual objective function is constructed, and the initial policy network is initially optimized using the individual objective function to obtain an intermediate policy network. A cooperative objective function is constructed based on the cooperative advantage function; wherein the optimization parameters of the cooperative objective function include the intermediate policy parameters of the intermediate policy network and the initial cooperative factor. By fixing the initial cooperative factor in the cooperative objective function, the intermediate policy parameters of the intermediate policy network are optimized to obtain the target policy network.
4. The multi-vehicle collaborative decision-making method according to claim 2, characterized in that, The optimization of the initial cooperative model based on the cooperative advantage function, the global advantage function, the initial policy network, and the target policy network to obtain the target cooperative model includes: Determine the initial policy parameters and the target policy parameters; Based on the synergistic advantage function, the global advantage function, the initial policy parameters, and the target policy parameters, a global objective function is constructed; wherein, the optimization parameters of the global objective function include the initial synergistic factor; The initial collaborative model is optimized using the global objective function to obtain the target collaborative model.
5. The multi-vehicle collaborative decision-making method according to claim 1, characterized in that, After each vehicle performs the corresponding target action in the initial environmental state, the initial environmental state is updated to the target environmental state. Based on the individual rewards, an individual advantage function is constructed for each vehicle, including: Based on the individual reward of each vehicle, we obtain the first individual value function of each vehicle in the target environment state and the second individual value function in the initial environment state; For each vehicle, an individual advantage function is obtained based on the vehicle's individual reward, the first individual value function, and the second individual value function; wherein, the individual advantage function is used to characterize the advantage that the vehicle gains from performing the corresponding target action in the initial environmental state.
6. The multi-vehicle collaborative decision-making method according to claim 5, characterized in that, The vehicle lifecycle includes a start time step when entering the environment and an end time step when leaving the environment; the process of obtaining the first individual value function of each vehicle in the target environment state and the second individual value function in the initial environment state based on the individual reward of each vehicle includes: For each vehicle, based on the vehicle's individual reward, calculate the average discounted reward of the vehicle from the next time step to the end time step to obtain the first volume value function; For each vehicle, based on the vehicle's individual reward, the average discounted reward of the vehicle from the current time step to the end time step is calculated to obtain the second body value function.
7. The multi-vehicle collaborative decision-making method according to claim 5, characterized in that, Based on the individual rewards, a neighborhood advantage function is constructed for each vehicle's neighborhood, including: Identify the neighboring vehicles within the neighborhood of the target vehicle among multiple vehicles. Based on the individual rewards of the target vehicle and its neighboring vehicles, obtain the neighborhood reward of the neighborhood of the target vehicle, the first neighborhood value function under the target environment state, and the second neighborhood value function under the initial environment state. Based on the neighborhood reward, the first neighborhood value function, and the second neighborhood value function, a neighborhood advantage function for the neighborhood where the target vehicle is located is obtained; wherein, the neighborhood advantage function is used to characterize the advantage brought by the target vehicle and neighboring vehicles performing the corresponding target action in the initial environmental state.
8. A multi-vehicle collaborative decision-making device, characterized in that, The time span of the environment in which the multiple vehicles are located includes multiple time steps, and in each time step the device includes: The first processing module is used to determine each vehicle's individual observations of other vehicles based on the initial environmental state, and input the individual observations into a pre-trained initial policy network to obtain the target action of each vehicle. The second processing module is used to determine the individual reward of each vehicle after performing the corresponding target action in the initial environmental state, and based on the individual reward, construct the individual advantage function of each vehicle, the neighborhood advantage function of the neighborhood where each vehicle is located, and the global advantage function of the environment in which multiple vehicles are located. The third processing module is used to coordinate the individual advantage function and the neighborhood advantage function based on the initial cooperative model, and optimize the initial policy network using the obtained cooperative advantage function and the individual advantage function to obtain the target policy network. The fourth processing module is used to optimize the initial cooperative model based on the cooperative advantage function, the global advantage function, the initial policy network, and the target policy network to obtain the target cooperative model. The fifth processing module is used to determine the initial strategy network and initial cooperation model for the next time step based on the target strategy network and the target cooperation model, so that each vehicle can determine the target action for the next time step. The first processing module is also used for: Based on the initial environmental state, the state characteristics of each vehicle are obtained; For each vehicle, a query vector is generated based on the vehicle's state characteristics, and a key vector and a value vector are generated based on the state characteristics of other vehicles; For each vehicle, an attention weight vector is calculated based on the similarity between the vehicle's query vector and the key vector. An attention matrix is then calculated based on the attention weight vector and the value vector. The attention matrix includes the vehicle's individual observations of other vehicles. By using the attention weight vector corresponding to the target vehicle, neighboring vehicles within the target vehicle's neighborhood are selected, and other vehicles whose attention weight exceeds the attention weight threshold are considered as neighboring vehicles.
9. A computer device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the multi-vehicle cooperative decision-making method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the multi-vehicle cooperative decision-making method according to any one of claims 1 to 7.
11. A computer program product, characterized in that, It includes computer instructions for causing a computer to execute the multi-vehicle cooperative decision-making method according to any one of claims 1 to 7.