Queue cruise control method based on subsequent features and migration reinforcement learning
Through the method based on successor features and transfer reinforcement learning, the problems of high training cost and difficult deployment of multi-agent reinforcement learning algorithms caused by agent heterogeneity are solved, and the real-time rapid deployment of heterogeneous ICV queue CACC tasks are achieved and the performance optimization efficiency is improved.
Patent Information
- Application Number
- CN202510213583.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-06
AI Technical Summary
Agent heterogeneity leads to high training costs for multi-agent reinforcement learning algorithms, making it difficult to quickly deploy CACC tasks in heterogeneous ICV queues in real time, and has low performance optimization efficiency.
The queue cruise control method based on successor features and transfer reinforcement learning is adopted, and by extracting knowledge and building an efficient transfer reinforcement learning mechanism between multi-agent systems, the learning efficiency and control performance of the target domain multi-agent system is improved.
It effectively reduces the training cost of multi-agent reinforcement learning algorithms, realizes the real-time rapid deployment of heterogeneous ICV queue CACC tasks, and improves performance optimization efficiency.
Smart Images

Figure CN120096564A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent connected vehicles, and in particular to a queue cruise control method based on successor features and transfer reinforcement learning. Background Art
[0002] Coordinated Adaptive Cruise Control (CACC) technology is designed to improve the safety and efficiency of Intelligent Connected Vehicle (ICV) in a connected vehicle environment. Through vehicle-to-everything (V2X) communication, CACC technology will guide the ICVs in the queue to adjust their control variables and maintain ideal vehicle distance, relative speed, etc., which can effectively alleviate traffic congestion and improve the safety and economy of the intelligent driving system.
[0003] The main schemes adopted by CACC can be divided into two categories, model-based and model-free, depending on whether the system dynamics are modeled. Among the model-based methods, Model Predictive Control (MPC) is the most widely developed and applied. It builds a prediction model of vehicle dynamics to achieve the optimization and control rate solution of the system in the future prediction time domain. However, the performance of MPC is highly dependent on the accurate model established for the actual complex system, which is very difficult in the actual road traffic environment.
[0004] In recent years, the rapid development of artificial intelligence has provided new opportunities to enhance the capabilities of CACC systems. In particular, a platoon equipped with CACC can be viewed as a Multi-Agent System (MAS), in which individual vehicles act as autonomous agents to make decisions based on their states while considering the global objectives of the platoon. A single MAS needs to solve a Decentralized Markov Decision Process (Dec-MDP) problem.
[0005] Multi-Agent Reinforcement Learning (MARL) technology can provide a solution to this problem. Since MARL is a reinforcement learning technology that can achieve model-free decision-making, its biggest advantage over the MPC method is that MARL technology only requires reward feedback from the environment, without the need for a system model.
[0006] However, MARL often initializes the control strategy randomly and gradually learns the optimal state-action value function (Q value) and optimal strategy in real-time interaction with the environment. However, the control performance is often not ideal before the learning is completed, and a long period of recursive training is required, which is not conducive to the actual deployment of CACC.
[0007] In recent years, transfer reinforcement learning based on successor feature (SF) has been proposed. In this method, the Q value is expressed as the product of the successor feature and the task weight, and it is believed that different agents share the successor feature, which reveals the global feature of the long-term reward mechanism of the task, and only the task weights are different from each other; further, the knowledge transfer mechanism between agents can be constructed based on SF, so that reinforcement learning training does not need to start from scratch, but training the initial agent can show higher performance, thereby improving the actual deployment performance of CACC based on MARL.
[0008] However, the existing SF-based transfer reinforcement learning work still has some shortcomings: 1. Assuming that the Q value can be represented by SF through a linear combination of task weights, it will limit the application of SF in some complex nonlinear Q value tasks; 2. The existing SF-based transfer mechanism has failed to study the knowledge transfer between heterogeneous multi-agent systems under continuous state and action space.
[0009] For the CACC task of ICV, the ICV queue can be abstracted as a multi-agent system (MAS), and the ICV can be abstracted as an agent in the MAS. The reward distribution of each agent is different. In this task, the reward mechanism of the CACC task between the MAS and a single agent is complex, the parameters between agents are heterogeneous, and the value function topology between different MASs is different. This makes the distribution of the agent value function not only variable, but also the decision-making of the agent is coupled with the adjacent agents. This will make the actual road deployment and the training cost of the intelligent connected queue with multi-agent reinforcement learning as the control algorithm high, and the CACC task is difficult to deploy quickly and in real time. Summary of the invention
[0010] The purpose of this application is to provide a queue cruise control method based on successor features and transfer reinforcement learning, in order to solve the problem that the heterogeneity of agents leads to high training costs of multi-agent reinforcement learning algorithms and it is difficult to quickly deploy CACC tasks for heterogeneous ICV queues in real time, and to achieve the improvement of learning efficiency and control performance of the target domain multi-agent system; The solution of the present invention adopts the subsequent features of extracted knowledge, the heterogeneity of long-term cumulative discounted rewards of intelligent agents, and the efficient transfer reinforcement learning mechanism between multi-agent systems constructed based on the features, as a technical route for real-time queue deployment transfer reinforcement learning.
[0011] The present invention provides a platoon cruise control method based on successor features and transfer reinforcement learning. The method is applied to the coordinated adaptive cruise control (CACC) task deployment of heterogeneous intelligent connected vehicle (ICV) platoons. The specific steps of the method are as follows: Step 1: During the source domain learning process, complete the global high-level state transfer features of the MAS of the M source domains , subsequent features Learning and joint strategies training; Step 2: In the process of target domain knowledge transfer, knowledge transfer is performed based on the subsequent features of the source domain and the joint strategy; Step 3 completes the training of the joint strategy for the target domain. Will continuously update and iterate to the optimal strategy , so that the CACC task of the ICV queue in the target domain achieves optimal performance.
[0012] Preferably, the global high-level state transition feature of the heterogeneous multi-agent in step 1 is used for a single MAS in the continuous state space and action space and strategies Under the action of , the Q value function of this MAS is obtained, and the Q value represents the expected value of the long-term reward of the single MAS under the guidance of this strategy. When the Q value distribution of MAS will show heterogeneity, the reward distribution heterogeneity is constructed by high-level features The Q-value function of the MAS collaborative learning is used as a feature extractor to complete the state transfer feature migration across heterogeneous MAS. , and heterogeneous reward functions for different MAS Extraction.
[0013] Furthermore, step 2 of knowledge transfer based on source domain subsequent features and joint strategies refers to MAS is used as the source domain MAS, MAS as the target domain MAS accepts knowledge transfer from all source domains MAS, and each MAS has Agent, then The status of the MAS ,action ,in , They are Moment The state and action variables of each agent; For MAS, considering the continuous state space and action space , for strategy The Q value function of the following MAS is: (1) in, and Respectively MAS in The state and action variables at each moment, and Represent the specific values of state and action respectively. It is MAS in The reward value at that moment, is the specific value of the reward value. For Starting time, Take Action, and The selection of the time start strategy follows The corresponding Q value in the case of It is a joint strategy of MAS. is the reward distribution probability; is the state transition probability; Represents the state variable of MAS Belonging space, Represents the action variable of MAS Belonging space, is the discount factor, is the differential symbol, and The integral symbol represents the integration of state and action variables. It can be seen from formula (1) that when the reward distribution probability presents heterogeneity, and the state transition probability When the two are the same, the Q value distribution will be heterogeneous, and the high-level features that characterize the reward distribution heterogeneity are: (2) in, is an advanced state transition feature, The conditional probability of the high-level state transition feature to the state variable and the action variable; The Heterogeneous reward distribution of MAS, Different distribution; Conditional probabilities of high-level state transition features on state variables and action variables The homogeneous state transition probability distribution of the intelligent agent is concentrated. After the source domain MAS is constructed, it is directly transferred to the target domain MAS for the construction of Q value. In the source domain, Joint state-action transition data of MAS For input, For continuous feature space, design state transition features The extraction network consists of two parts: a heterogeneous network branch and a shared backbone network, wherein the shared backbone network completes arrive The mapping relationship of heterogeneous network branches represents the characteristics To Reward Mapping relationship: Each MAS inputs data into the shared backbone network, and the output Then input the corresponding heterogeneous network branch to complete arrive The mapping of The heterogeneous network branch model corresponding to each MAS is: ; Given features As Input, for the extracted features have: (3) in, Indicates MAS in The reward value at the moment; Then design the training loss as follows: (4) in, are the parameters of the heterogeneous network branch model, Representative The mapping relationship between the heterogeneous network branches corresponding to each MAS; Specific to the CACC task, MAS in Status of the moment Defined as ; in, For the first ICV (intelligent agent) in The state at the moment, including ; in, for The distance between the car and the front car, for Relative speed between the vehicle and the vehicle in front; Action is the vector of accelerations of all ICVs in the queue; reward , that is, MAS in The reward at each moment takes into account the distance error between all ICVs and the preceding vehicle, and the relative speed between all ICVs and the preceding vehicle. in, is the reference target spacing, For the Vehicle energy consumption, Customize weights to adjust the weights of different goals in rewards; Similarly, the shared backbone network model approximates The distribution of , then the training loss is designed as follows: (5) in, To share the backbone network model parameters, Represents the mapping relationship of the shared backbone network.
[0014] Preferably, the state transfer feature of the migration across heterogeneous MAS is completed , and heterogeneous reward functions for different MAS The extraction is to decouple and extract the common state transition features and heterogeneous reward features of different MAS, further learn the successor features that characterize the distribution of the CACC task value function, and transfer the knowledge of the successor features to the target domain; Formula (2) is given a more compact form: (6) in, , representing the state transition probability The cumulative discount sum is called For strategy The subsequent characteristics under influence; Since in classical Q-learning, there is ,in Get a value for the reward The space where you are located; Therefore, for a single MAS, The learning of is equivalent to the learning of Q value, and the corresponding state transition characteristics The role is equivalent to the immediate reward in Q-value learning , used to form The iteration goal of Iterative method: (7) in, , represent In the The iterative value of the moment, represents the learning rate; k is the number of steps projected from the current time to the future time; When the transfer data is obtained When , the new transfer feature data pair is directly calculated ; For transfer features The sampling value of ,have , and for the values of other transfer features ,have ,but The iterative method is: (8) in, Indicated in The characteristic mapping value of the state transition data pair at the moment is equal to ; It means any; because The learning of is equivalent to the learning of Q value, so, Iterative reflection The actual probability distribution of The long-term cumulative discount and.
[0015] Furthermore, the source domain successor feature (SF) in step 1 is used as the global shared successor feature. Reflecting the overall Q-value distribution of MAS with arbitrary heterogeneous reward distribution, the source domain MAS completes the After learning, the target domain MAS is Migrate, and construct and iterate the Q value distribution of the target domain MAS; According to formula (6) Construct the target domain Q value and combine it with the real-time state transfer data of the target domain To learn, the steps are as follows: Step A reconstructs the state transition feature extraction network for the target domain MAS; Step B transfers data pairs in real-time state in the target domain are input and output samples; Step C takes the loss function shown in formula (4) of the target domain MAS as the target, and optimizes the parameters of the shared backbone network part of the network according to the classical learning method of neural network, namely the backpropagation (BP) algorithm and the gradient descent method; Step D replaces Step C The parameter optimization results are applied to the CACC task of intelligent connected vehicles to optimize the ICV platoon following performance.
[0016] Preferably, the transfer reinforcement learning of the successor features of the source domain in step 2 adopts the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm and the CACC method of the Successor Feature-based MADDPG (SF-MADDPG). The SF-MADDPG iterative algorithm follows the actor-critic architecture of MADDPG, in which the actor completes the policy mapping from state to action, the critic completes the value mapping of the state-action data pair to the value function, and the critic judges the long-term return of the actor's strategy.
[0017] Furthermore, in step 1, the MAS global high-level state transfer features of M source domains are completed during the source domain learning process , subsequent features Learning and joint strategies The training refers to the learning of the global high-level state transition features and subsequent features of heterogeneous multi-agents in the source domain of knowledge. The number of source domains and training cycles are used as units to complete the cycle. State transition characteristics of the source domain , subsequent features and joint strategies The specific steps of training are as follows: Step 1.1 The source domain In the source domain, At the beginning of a training cycle, reset the environment, that is, initialize ;in, For the first ICV (intelligent agent) in =The state at time 0; Step 1.2 At this moment, according to the joint strategy given by the actor and status , we get the joint action of MAS ,in, ; For the first ICV (intelligent agent) in The state of the moment, For the ICV (Intelligent Agent) = Action at time 0 .
[0018] Step 1.3 Execute the MAS action , and get and MAS Total Rewards ; Step 1.4: Calculate state transition features ; Step 1.5: Store state transition data pairs To the experience replay pool , the experience replay pool is maintained and updated by the Internet of Vehicles data platform; Step 1.6: From the Experience Pool Randomly sample data Group data; Step 1.7: Design training loss functions based on formulas (4) and (5) and , and used for state transition characteristics Training of the extraction network: (4) Similarly, the shared backbone network model approximates The distribution of , then the training loss is designed as follows: (5) The following updates are performed on the feature extraction network: (9) (10) in, is the parameter of the shared backbone network The learning rate, It is for Heterogeneous network branch parameters corresponding to source domain MAS The learning rate; Step 1.8: Based on state transition characteristics , build a fully connected neural network to fit the subsequent features The distribution of the fully connected neural network is ,right Perform the following updates: (11) in, It is in strategy Next The subsequent characteristics of the step iteration; ; With the updated fully connected neural network parameters Calculate successor features ; is the parameter of the successor feature The learning rate; Step 1.9: According to formula (6), the Q value of MAS is obtained: (6) Step 1.10: Based on the actor-critic structure, the Q value is used as the critic to judge the decision performance, and the joint strategy of MAS based on Q value To update: (12) in, , for and The KL divergence between .
[0019] Preferably, the strategy in step 3 Will continuously update and iterate to the optimal strategy , so that the CACC task of the ICV queue in the target domain can achieve the optimal performance, which means that the transfer reinforcement learning based on the successor features in the knowledge transfer process is cycled in the target domain knowledge transfer process in units of training cycles to complete the joint strategy of the target domain The training finally optimizes the CACC strategy based on the car-following performance of the intelligent connected vehicle. The specific steps are as follows: Step 2.1: In the target domain, At the beginning of a training cycle, reset the environment, that is, initialize ; Step 2.2: At this moment, according to the joint strategy given by the actors and status , get the action ,in ; Step 2.3: Execute the target domain MAS action , and get and MAS Total Rewards ; Step 2.4: Utilize state feature network parameters transferred from the source domain , forward calculation to obtain the feature ; Step 2.5: Store state transition data pairs To the experience replay pool , Experience Replay Pool Maintained and updated by the Internet of Vehicles data platform; Step 2.6: From the experience pool Randomly sample data Group data; Step 2.7: Design training loss function based on formula (4) , and used for state transition feature extraction network Training: (4) in, is the heterogeneous network branch model parameter; For feature extraction network Perform the following updates: (10) Step 2.8: The Q value constructed by the reviewer contains the future discounted cumulative reward information. For each target domain MAS, the source domain number to be migrated is selected according to the maximum Q value principle. : (13) Among them, Q value The calculation is based on the state transition characteristics , subsequent features and source domain and joint strategies According to formula (6), MAS is obtained based on the The Q value under the knowledge of the source domain as a reference : (6) The first The state transition features, successor features, and joint strategies of the source domain will be transferred to the target domain; Step 2.9: Based on the actor-critic structure, the Q value is used as the critic to judge the decision performance, and the joint strategy of MAS based on the Q value To update: (14) in, is the target domain strategy learning rate, , for and The KL divergence between .
[0020] Furthermore, the final strategy in step 3 Will be continuously updated and iterated , so that the CACC task of the ICV platoon in the target domain achieves the optimal performance, which means that the CACC strategy is finally achieved based on the following performance of the intelligent connected vehicle Optimization. Preferably, the ICV fleet is viewed as a multi-agent system (MAS), where each ICV is an agent, and its model is expressed as: (15) in, and The vehicles status and control inputs, , and The vehicles The position, velocity and acceleration of Usually represents the desired acceleration. , is a given constant matrix that satisfies: (16) in, is the inertial time lag parameter of the vehicle longitudinal dynamics; In a connected environment, the CACC task of an intelligent vehicle described by equations (15) and (16) is modeled as a Markov decision process (MDP) of heterogeneous multi-agents. MDP is mathematically described by a multi-tuple: ; in, and Represent the state variables of the agent Space and action variables Belonging space; represents the state transition probability, Reward function , is the discount factor, which is the discount term for the future reward value when the agent calculates the cumulative expected reward; the agent's strategy is generally used Represented by, and divided into random strategies: ~ and a deterministic strategy: ;by Indicates status Next select action The probability of maximizing the cumulative reward of the entire process is called return, express The reward value at the moment is The reward function at the moment The definition is as follows: (17) The discount factor Existence and rewards The boundedness of a single agent Status is a vector, ,action For scalars, .
[0021] The present invention provides a platoon cruise control method based on successor features and transfer reinforcement learning. Due to the adoption of a double-layer network structure of heterogeneous network branches and a shared backbone network, the state transition features and heterogeneous reward functions of heterogeneous multi-agent state transition data are extracted, and successor features reflecting the Q-value distribution in the entire area are constructed based on the state transition features. A transfer reinforcement learning method is further proposed, which effectively solves the problems in the prior art of high training cost of multi-agent reinforcement learning algorithms, difficulty in real-time deployment of CACC tasks, and low performance optimization efficiency caused by the heterogeneity of agents, thereby achieving the effect of rapid startup of the reinforcement learning performance of CACC of heterogeneous intelligent connected vehicles in the target domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a queue cruise control method based on successor features and transfer reinforcement learning. Extraction network diagram of Figure 2(a) is a simulation result diagram of a platoon cruise control method based on successor features and transfer reinforcement learning and a CACC method based on SF-MADDPG trained in the target domain; Figure 2(b) is a simulation result diagram of the training process of the CACC method based on MADDPG in the target domain for a platoon cruise control method based on successor features and transfer reinforcement learning; Figure 3 It is a schematic diagram of the source domain learning process and knowledge transfer process of a platoon cruise control method based on successor features and transfer reinforcement learning; Figure 4 It is a flow chart of a platoon cruise control method based on successor features and transfer reinforcement learning; DETAILED DESCRIPTION Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims; The main goal of CACC in the existing technology is to make the vehicles in the ICV platoon travel at the same speed and predetermined distance, while reducing the energy consumption of the platoon. CACC is usually achieved by adopting a collaborative decision-making mechanism between ICVs in a V2X communication environment. Such a platoon can be regarded as a multi-agent system (MAS), in which each ICV is an agent, and its model is expressed as: (15) in, and The vehicles status and control inputs, , and The vehicles The position, velocity and acceleration of Usually represents the desired acceleration. Get a value for the reward The space where you are located; , is a given constant matrix that satisfies: (16) in, is the inertial time lag parameter of the vehicle longitudinal dynamics. The present invention considers the inertial time lag of all intelligent bodies. The above dynamic characteristics and parameters are unknown to all agents in the multi-agent system under consideration, which is closer to the real-world scenario.
[0023] In a connected environment, the CACC task of an intelligent vehicle described by equations (15) and (16) can be modeled as a Markov decision process (MDP) of heterogeneous multi-agents. MDP can be mathematically described by a multi-tuple: ,in and Represent the state variables of the agent Space and action variables Belonging space. represents the state transition probability. Reward function , it can return the reward that the agent obtains from the environment at each step, and the reward value has an upper bound; when the agent interacts with the environment, the reward value is the feedback of the environment to the current state and action of the agent, guiding the agent to make correct decisions. is the discount factor, which is the discount term for the future moment reward value when the agent calculates the cumulative expected reward. Assume that the marginal probability of the state at the next moment is Just and It is related to the state and action at that moment, but has nothing to do with the historical state and action of the agent, that is, it satisfies the Markov Property.
[0024] The agent's strategy is usually Represented by, and divided into random strategies: ~ and a deterministic strategy: ; The random strategy generates a probability distribution in the action space according to the current state, and the agent randomly selects an action in the distribution. The deterministic strategy directly generates an action value according to the current state of the agent; both have their own advantages in different scenarios, but neither affects the construction of the basic model. Indicates status Next select action The optimization problem of an MDP is to maximize the cumulative reward of the entire process, called the return. express The reward value at the moment, then the return function at that moment is defined as follows: (17) In formula (17), It is a temporary symbol that summarizes the previous polynomial using the summation symbol, representing The number of steps projected from moment to future moment; discount factor Existence and rewards The boundedness of is bounded. The closer it is to 0, the smaller the reward value takes into account the rewards at future moments. Size and strategy Therefore, the goal of the agent is to find an optimal strategy , so that The expected value of is maximized at any moment.
[0025] Specifically for the CACC task, define a single agent Status is a vector, ,action For scalars, ; In actual situations, the reward index of vehicle cruise control not only includes the following performance, but also takes into account the vehicle energy consumption index and constructs a control strategy with certain energy consumption economy. Therefore, the reward function There are different distributions of , which can also be called reward distribution heterogeneity.
[0026] In the embodiment of the present invention, in order to solve the problem that the heterogeneity of agents leads to high training cost of multi-agent reinforcement learning algorithms and difficulty in real-time and rapid deployment of CACC tasks for heterogeneous ICV queues, it is necessary to form a solution for constructing value functions based on CACC task knowledge.
[0027] The embodiment of the present invention provides a queue cruise control method based on successor features and transfer reinforcement learning, such as Figure 4 As shown: This method is applied to the deployment of Coordinated Adaptive Cruise Control (CACC) tasks for heterogeneous intelligent connected vehicle (ICV) platoons. The specific steps of this method are as follows: Step 1: During the source domain learning process, complete the global high-level state transfer features of the MAS (Multi-Agent System) of M source domains , subsequent features Learning and joint strategies training; Step 2: In the process of target domain knowledge transfer, knowledge transfer is performed based on the subsequent features of the source domain and the joint strategy; Step 3 completes the training of the joint strategy for the target domain. Will continuously update and iterate to the optimal strategy , so that the CACC task of the ICV queue in the target domain achieves optimal performance.
[0028] The present invention provides a queue cruise control method based on successor features and transfer reinforcement learning. Due to the adoption of a double-layer network structure of heterogeneous network branches and a shared backbone network, the state transition features and heterogeneous reward functions of heterogeneous multi-agent state transition data are extracted, and successor features reflecting the Q value distribution in the entire area are constructed on the basis of the state transition features. A transfer reinforcement learning method is further proposed, which effectively solves the problems in the prior art of high training cost of multi-agent reinforcement learning algorithms caused by the heterogeneity of agents, difficulty in real-time and rapid deployment of CACC tasks for heterogeneous ICV queues, and low performance optimization efficiency, thereby achieving the effect of rapid startup of the reinforcement learning performance of heterogeneous intelligent connected vehicles CACC in the target domain, and finally enables the CACC task of the real-time ICV queue to achieve rapid startup performance by relying on the knowledge migration of SF.
[0029] Embodiment 1 The embodiment of the present invention provides a queue cruise control method based on successor features and transfer reinforcement learning. The technical solution first combines the ideas of federated learning and multi-task learning to construct a global high-level state transition feature of heterogeneous multi-agents to decompose the reward distribution heterogeneity and homogeneous state transition probability distribution between MASs; then, based on the relationship between the accumulated discounted reward of the agent and the state transition feature, a successor feature (SuccessorFeature, SF) is proposed based on the state transition feature, and an iterative algorithm of the global SF is constructed based on the TD (0) learning idea; Furthermore, a method for quickly constructing arbitrary value functions is proposed based on the global SF, and a transfer reinforcement learning mechanism between MASs is studied with SF as the core; ultimately, the CACC task of the real-time ICV queue can achieve fast startup performance by relying on the knowledge transfer of SF. The specific technical solution is as follows: 1. Learning mechanism of global high-level state transition features and subsequent features of heterogeneous multi-agents For knowledge transfer, consider the source domain that generates knowledge and the target domain that receives the knowledge transfer. The purpose of the transfer is to make the learning of the target domain agent no longer start from scratch, but to be based on a certain understanding of the task and environment, which helps to accelerate the evolution of the agent.
[0030] for MAS is used as the source domain MAS, MAS as the target domain MAS accepts knowledge transfer from all source domains. Each MAS has Agent, then The status of the MAS ,action ,in , They are Moment The state and action variables of each agent, T represents transposition.
[0031] For a single agent, consider a continuous state space and action space , for strategy The Q value function of a single agent is: (18) The Q value is the core indicator to measure the long-term value of the local decision of the agent. Meaning Time is the starting point and state Take Action, and The selection of the time start strategy follows The corresponding Q value in the case of .
[0032] From (18), we can see that when the reward distribution probability presents heterogeneity, and the state transition probability The same will make the Q-value distribution heterogeneous, increasing the difficulty of constructing the value function of multi-agent collaborative learning. Considering the use of high-level features to characterize the heterogeneity of reward distribution, we have: (19) in, is a high-level state transition feature, which is the conditional probability of the state variable and the action variable It embodies the homogeneous state transition probability distribution of the object agent, and can be directly transferred to the target domain MAS after the source domain MAS is constructed for the construction of Q value. This reflects the heterogeneous reward distribution, The distribution is different.
[0033] The knowledge transfer of the present invention occurs between MASs. Joint state-action transition data of MAS For input, For continuous feature space, design state transition features The extraction network consists of two parts: a heterogeneous network branch and a shared backbone network, wherein the shared backbone network completes arrive The mapping relationship of heterogeneous network branches represents the characteristics To Reward Mapping relationship: Each MAS inputs data into the shared backbone network, and the output Then input the corresponding heterogeneous network branch to complete arrive The mapping of The heterogeneous network branch model corresponding to each MAS is: ; Given features As Input, for the extracted features have: (3) Then design the training loss as follows: (4) in, are the parameters of the heterogeneous network branch model, Representative The mapping relationship between the heterogeneous network branches corresponding to each MAS; Similarly, the shared backbone network model approximates The distribution of , then the training loss is designed as follows: (5) in, To share the backbone network model parameters, Represents the mapping relationship of the shared backbone network; This forms a two-layer neural network structure consisting of heterogeneous network branches and a shared backbone network, which acts as a feature extractor to complete the state transfer features that can migrate across heterogeneous MAS. , and heterogeneous reward functions for different MAS Extraction, such as Figure 1 As shown, the state transfer data pair For input, , construct a state transition feature extraction network.
[0034] Therefore, the common state transition features and heterogeneous reward features of different MAS can be extracted, and the state transition features can be used as the basis for the next step of knowledge migration to the target domain.
[0035] Furthermore, formula (2) can be written in a more compact form: (6) in, , representing the state transition probability The cumulative discount sum is called For strategy The subsequent characteristics under influence; In classical Q-learning, we have , we can see that for a single MAS, The learning and iteration of is equivalent to the learning and iteration of Q value, and the corresponding state transition characteristics The role is equivalent to the immediate reward in Q learning , used to form The iteration goal of Iterative method: (7) in, , represent In the The iterative value of the moment, represents the learning rate; When the transfer data is obtained When , the new transfer feature data pair is directly calculated ; For the observed transfer characteristics ,have , and for the values of other transfer features ,have ,but The iterative method is: (8) in, Indicated in The characteristic mapping value of the state transition data pair at the moment is equal to .because The iteration of is the same as the iteration of Q value in Q learning, so, Iterative reflection The actual probability distribution of The long-term cumulative discount and.
[0036] 2. Transfer reinforcement learning based on successor features As a globally shared successor feature, Reflecting the overall Q-value distribution of MAS with arbitrary heterogeneous reward distribution, the source domain MAS completes the After learning, the target domain MAS is The Q value distribution of the target domain MAS is further constructed and iterated.
[0037] The construction of the target domain Q value also requires , because previously The distribution of is known, and we only need to combine the real-time state transfer data of the target domain The learning can be done, including: reconstructing the target domain MAS Figure 1 The state transition feature extraction network shown in the figure is based on the real-time state transition data of the target domain. As input and output samples, the loss function shown in formula (4) of the target domain MAS is taken as the target, and the parameters of the shared backbone network part of the network are optimized according to the classical learning method of neural network, namely backpropagation (BP) and gradient descent method.
[0038] Since the mapping relationship of the Q value of the target domain MAS does not need to be generated from scratch and iteratively based on a large amount of training data, but is quickly started by relying on the knowledge transfer of the source domain MAS, better performance can be quickly achieved in actual deployment. Specifically, in the CACC task of intelligent connected vehicles, the optimization of the ICV platoon following performance can be quickly achieved.
[0039] In one embodiment, the CACC scenario of the following ICV queue: the energy consumption of a single ICV is , the speed is , the acceleration is Generally, the rolling resistance, air resistance, slope resistance, acceleration resistance and efficiency of the power system of the vehicle are taken into consideration. The relationship between energy consumption, speed and acceleration is: (20) in, arrive Represents various coefficients, and different combinations of coefficients represent the energy consumption distribution of different vehicles. Consider that each ICV fleet contains 3 ICVs with different configurations. The vehicle energy consumption parameter combination of the first MAS in the source domain is: (21.a) (21.b) (21.c) The vehicle energy consumption parameter combination of the second MAS in the source domain is: (21.d) (21.e) (21.f) The vehicle energy consumption parameter combination of the target domain MAS is: (21.g) (21.h) (21.i) Among them, the target domain vehicles are large vehicles with generally high energy consumption, while the source domain vehicles are small vehicles with generally low energy consumption. For a single MAS, the reward function of the CACC task is set as: (twenty two) in, , , They are A car in Time and distance to the vehicle ahead, Vehicle speed, Vehicle energy consumption, As the reference distance is set to 30m, the task objective of CACC is to ensure the stability of the distance between vehicles and relative speed within the queue and to minimize energy consumption.
[0040] In an embodiment of the present invention, in this given CACC task of intelligent connected vehicles, a CACC method of Successor Feature-based MADDPG (SF-MADDPG) is proposed in combination with the existing multi-agent reinforcement learning algorithm MADDPG (Multi-Agent Deep Deterministic Policy Gradient). The SF-MADDPG algorithm follows the actor-critic architecture of MADDPG.
[0041] The comparison of two algorithms is set: MADDPG directly acts on the target domain ICV queue, and the SF-MADDPG proposed in this invention is compared with the total reward in a single training cycle. The training setting is: one training cycle contains an Urban Dynamometer Driving Schedule (UDDS) working condition. SF-MADDPG is first trained for 100 cycles under the source domain MAS to fully extract the subsequent features, and then trained for 300 cycles in the target domain context with MADDPG. The discount factor of the Q value is 0.99, the experience pool size is 20000, the experience replay size is 1024, and the learning rate of each module is 0.0025. Then Figure 2 (a) and Figure 2 (b) show that CACC based on SF-MADDPG and MADDPG is trained in the target domain, and the total reward of the two algorithms in a single cycle under 300 training cycles is compared. Each algorithm has 10 experiments to ensure full verification. The solid line is the average value and the gray area is the distribution of 10 experimental data.
[0042] It can be found that the performance of SF-MADDPG and MADDPG in the target domain has been greatly improved. Since SF-MADDPG continues to train on the basis of knowledge transfer from the source domain, it has a higher starting point and eventually converges to better performance (the average reward is -45952 at the 300th cycle), while the average reward of MADDPG at the 300th cycle is -68627. The present invention has achieved a 49.34% performance improvement. If the 100 cycles of the source domain are also included in the SF-MADDPG training cost, the performance / training cost is still improved by 12.01% relative to the MADDPG algorithm. It proves that the algorithm SF-MADDPG proposed in the present invention can help the target domain algorithm obtain a better algorithm performance iteration starting point (i.e., a fast startup effect), and achieve comprehensive performance improvement in the following performance and energy consumption optimization of the CACC task.
[0043] The embodiment of the present application provides a platoon cruise control method based on successor features and transfer reinforcement learning. Due to the adoption of a two-layer network structure of heterogeneous network branches and a shared backbone network, the state transition features and heterogeneous reward functions of heterogeneous multi-agent state transition data are extracted, and successor features reflecting the Q value distribution in the entire area are constructed based on the state transition features. A transfer reinforcement learning method is further proposed, which effectively solves the problems in the prior art that the heterogeneity of agents leads to high training costs of multi-agent reinforcement learning algorithms, difficulty in real-time deployment, and low performance optimization efficiency, thereby achieving the effect of rapid startup of the reinforcement learning performance of heterogeneous intelligent connected vehicles CACC in the target domain.
[0044] Embodiment 2 In an embodiment of the present invention, in a given CACC task of an intelligent connected vehicle, a CACC method of a successor feature-based MADDPG (Successor State Transition Feature-based MADDPG, SF-MADDPG) is proposed in combination with the existing multi-agent deep deterministic policy gradient algorithm MADDPG (Multi-Agent Deep Deterministic Policy Gradient). The SF-MADDPG algorithm follows the actor-critic architecture of MADDPG.
[0045] In one embodiment, the present invention provides a complete system flow of an intelligent network-connected platoon cruise control method based on transfer reinforcement learning, as shown in the following figure: Figure 3 As shown in the figure, the overall process can be divided into the source domain learning process and the knowledge transfer process. In the source domain learning process, the soft update is based on the online actors and online critics to complete the update of the target actors and target critics respectively; the TD learning is based on the target actors and target critics to update the online actors and online critics in real time.
[0046] The technical solution of the present invention mainly includes the following steps in the embodiment: Step 1 Source domain learning process - learning of global high-level state transition features and subsequent features of heterogeneous multi-agents: In one embodiment, an agent that defines a CACC task Status Defined as ,in for The distance between the car and the front car, for Relative speed between the vehicle and the vehicle in front; Action For acceleration; reward , is the reference target spacing, Customize the weights for each goal to adjust the weights of different goals in the reward. , , in order to highlight the emphasis on the following distance target.
[0047] At the same time, the following parameters are determined: in M source domains, each source domain initializes the actor and critic modules for N agents; initializes the state transition feature extraction network; Source domain MAS initialization parameters and its learning rate , initialize the global reward feature parameters and its learning rate ; Number of training cycles ; Training duration in a single training cycle ;Experience pool size ; Experience replay dataset size ;Discount factor ; Parameters of subsequent features and its learning rate ;The set of actor parameters for all agents in the MAS and their common learning rate ;Soft update factor ; Step 1.1: The source domain In the source domain, At the beginning of a training cycle, reset the environment, that is, initialize ; Step 1.2: At this moment, according to the joint strategy given by the actor and status , and the joint action of the MAS is obtained ,in ; Step 1.3: Execute the MAS action , and get and MAS Total Rewards ; Step 1.4: Calculate state transition features
[0048] Step 1.5: Store state transition data pairs To the experience replay pool ,This experience pool is maintained and updated by the Internet of Vehicles data platform; Step 1.6: From the Experience Pool Randomly sample data Group data; Step 1.7: Design training loss functions based on formulas (4) and (5) and , and used for state transition characteristics Training of the extraction network: (4) in are the neural network model parameters.
[0049] Similarly, another neural network model is used to represent , then the design training loss is as follows: (5) in, To share the backbone network model parameters, Represents the mapping relationship of the shared backbone network. And further, the following updates are performed on the feature extraction network: (9) (10) in, is the parameter of the shared backbone network The learning rate, It is for Heterogeneous network branch parameters corresponding to source domain MAS The learning rate; Step 1.8: Based on state transition characteristics , build a fully connected neural network to fit the subsequent features The distribution of the fully connected neural network is ,right Perform the following updates: (11) in, , and further, with the updated neural network parameters Calculate the subsequent characteristics of the MAS .
[0050] Step 1.9: According to formula (6), the Q value of the MAS is obtained: (6) Step 1.10: Based on the actor-critic structure, the Q value is used as the critic to judge the decision performance, and the joint strategy of MAS based on Q value To update: (12) in , for and The KL divergence between .
[0051] The above steps are repeated in units of source domain number and training cycle, and the State transition characteristics of the source domain , subsequent features and joint strategies training.
[0052] Step 2: Target domain knowledge transfer process - Transfer reinforcement learning based on successor features: Given M state transition features trained in source domains , the subsequent features and the optimal joint strategy , ; Initialize the state transfer feature extraction network for the target domain MAS, including initialization parameters and its learning rate , initialize the global reward feature parameters and its learning rate ; Number of training cycles ; Training duration in a single training cycle ;Experience pool size ; Experience replay dataset size ;Discount factor ; Parameters of subsequent features and its learning rate ;The set of actor parameters for all agents in the MAS and their common learning rate ;Soft update factor
[0053] Step 2.1: In the target domain, At the beginning of a training cycle, reset the environment, that is, initialize ; Step 2.2: At this moment, according to the joint strategy given by the actors and status , get the action ,in ; Step 2.3: Execute the target domain MAS action , and get and MAS Total Rewards ; Step 2.4: State feature network parameters transferred from the source domain , the feature is obtained by forward calculation of the neural network ; Step 2.5: Store state transition data pairs To the experience replay pool ,This experience pool is maintained and updated by the Internet of Vehicles data platform; Step 2.6: From the experience pool Randomly sample data Group data; Step 2.7: Design training loss function based on formula (4) , and used for state transition feature extraction network Training: (4) in is the neural network model parameter. And further, the feature extraction network Perform the following updates: (10) Step 2.8: The Q value constructed by the reviewer contains the future discounted cumulative reward information. The larger the Q value, the better the optimization performance of the decision. Therefore, for each target domain agent, the source domain number to be migrated is selected according to the maximum Q value principle. : (13) Among them, Q value The calculation is based on State transition characteristics of the source domain , subsequent features and joint strategies According to formula (6), the MAS is The Q value under the knowledge of the source domain as a reference : (6) No. State transition characteristics of the source domain , subsequent features and joint strategies will be migrated to the target domain.
[0054] Step 2.9: Based on the actor-critic structure, the Q value is used as the critic to judge the decision performance, and the joint strategy of MAS based on the Q value To update: (14) in, , for and The KL divergence between .
[0055] The above steps are repeated in training cycles to complete the joint strategy of the target domain. The training is then conducted and the CACC strategy is finally optimized based on the car-following performance of the intelligent connected vehicle. Will be continuously updated and iterated , so that the CACC of the ICV queue in the target domain can achieve the optimal result.
[0056] The embodiment of the present invention provides a queue cruise control method based on successor features and transfer reinforcement learning, which relates to the field of intelligent connected vehicles. Due to the adoption of a double-layer network structure of heterogeneous network branches and a shared backbone network, the state transition features and heterogeneous reward functions of heterogeneous multi-agent state transition data are extracted, and successor features reflecting the Q value distribution in the entire area are constructed on the basis of the state transition features. A transfer reinforcement learning method is further proposed, which effectively solves the problems in the prior art of high training cost of multi-agent reinforcement learning algorithms caused by the heterogeneity of agents, difficulty in real-time and rapid deployment of CACC tasks of heterogeneous ICV queues, and low performance optimization efficiency, thereby achieving the effect of rapid startup of the reinforcement learning performance of heterogeneous intelligent connected vehicles CACC in the target domain, and finally enables the CACC task of the real-time ICV queue to achieve rapid startup performance by virtue of the knowledge migration of successor features.
Claims
1. A queue cruise control method based on successor features and transfer reinforcement learning, the method is applied to the CACC task deployment of heterogeneous ICV queues, characterized in that: The specific steps of the method are as follows: Step 1: During the source domain learning process, complete the global high-level state transfer features of the MAS of the M source domains , subsequent features Learning and joint strategies training; Step 2: In the process of target domain knowledge transfer, knowledge transfer is performed based on the subsequent features of the source domain and the joint strategy; Step 3 completes the training of the joint strategy for the target domain. Will continuously update and iterate to the optimal strategy , so that the CACC task of the ICV queue in the target domain achieves optimal performance.
2. A queue cruise control method based on successor features and transfer reinforcement learning according to claim 1, characterized in that: The global high-level state transition feature of the heterogeneous multi-agent in step 1 is used for a single MAS in the continuous state space and action space and strategies Under the action of , the Q value function of the single MAS is obtained, and the Q value represents the expected value of the long-term reward of the single MAS under the guidance of the strategy. When the Q value distribution of MAS will show heterogeneity, the reward distribution heterogeneity is constructed by high-level features The Q-value function of the MAS collaborative learning is used as a feature extractor to complete the state transfer feature migration across heterogeneous MAS. , and heterogeneous reward functions for different MAS Extraction.
3. A queue cruise control method based on successor features and transfer reinforcement learning according to claim 1, characterized in that: The step 2 of transferring knowledge based on the subsequent features of the source domain and the joint strategy refers to MAS is used as the source domain MAS, MAS as the target domain MAS accepts knowledge transfer from all source domains MAS, and each MAS has Agent, then The status of the MAS ,action ,in , They are Moment The state and action variables of each agent; For MAS, considering the continuous state space and action space , for strategy The Q value function of the following MAS is: (1) in, and Respectively MAS in The state and action variables at each moment, and Represent the specific values of state and action respectively. It is MAS in The reward value at that moment, is the specific value of the reward value. For Starting time, Take Action, and The selection of the time start strategy follows The corresponding Q value in the case of It is a joint strategy of MAS. is the reward distribution probability; is the state transition probability; Represents the state variable of MAS Belonging space, Represents the action variable of MAS Belonging space, is the discount factor, is the differential symbol, and The integration symbol represents the integration of state and action variables; From formula (1), we can conclude that when the reward distribution probability Presenting heterogeneity, state transition probability When the Q value distribution is the same, it shows heterogeneity. The high-level features characterize the heterogeneity of reward distribution as follows: (2) in, is an advanced state transition feature, The conditional probability of the high-level state transition feature to the state variable and the action variable; The Heterogeneous reward distribution of MAS; Conditional probabilities of high-level state transition features on state variables and action variables The homogeneous state transition probability distribution of the intelligent agent is concentrated. After the source domain MAS is constructed, it is directly transferred to the target domain MAS for the construction of Q value. In the source domain, Joint state-action transition data of MAS For input, For continuous feature space, design state transition features The extraction network consists of two parts: a heterogeneous network branch and a shared backbone network. The shared backbone network completes arrive The mapping relationship of heterogeneous network branches represents the characteristics To Reward Mapping relationship: Each MAS inputs data into the shared backbone network, and the output Then input the corresponding heterogeneous network branch to complete arrive The mapping of The heterogeneous network branch model corresponding to each MAS is: ; Given features As Input, extract the completed features for: (3) in, Indicates MAS in The reward value at the moment; Then design the training loss as follows: (4) in, are the parameters of the heterogeneous network branch model, Representative The mapping relationship between the heterogeneous network branches corresponding to the MAS; MAS in Status of the moment Defined as ; For the first ICV (intelligent agent) in The state at the moment, including ; for The distance between the car and the front car, for Relative speed between the vehicle and the vehicle in front; Action is a vector of accelerations of all ICVs in the queue; reward , that is, MAS in The reward at each moment is the error between the distance between all ICVs and the vehicle in front, and the relative speed between all ICVs and the vehicle in front; is the reference target spacing, For the Vehicle energy consumption, They are custom weights respectively; Similarly, the shared backbone network model approximates The distribution of , then the training loss is designed as follows: (5) in, To share the backbone network model parameters, Represents the mapping relationship of the shared backbone network.
4. A queue cruise control method based on successor features and transfer reinforcement learning according to claim 2, characterized in that: The state transfer characteristics of completing the migration across heterogeneous MAS , and heterogeneous reward functions for different MAS The extraction is to decouple and extract the common state transition features and heterogeneous reward features of different MAS, further learn to obtain the successor features of the CACC task value function distribution, and transfer the knowledge of the successor features to the target domain; Formula (2) is given a more compact form: (6) in, , representing the state transition probability The cumulative discount sum is called For strategy The subsequent characteristics under influence; In classical Q-learning, we have ,in, Get a value for the reward The space where you are located; For a single MAS, The learning of is equivalent to the learning of Q value, and the state transition characteristics Immediate rewards in Q-value learning The same function, composition The iterative target method is: (7) in, , represent In the The iterative value of the moment, represents the learning rate; k is the number of steps projected from the current time to the future time; When the transfer data is obtained When , the new transfer feature data pair is directly calculated ; Transfer Features The sampling value of , with , the values of other transfer features , with ,but The iterative method is: (8) in, Indicated in The characteristic mapping value of the state transition data pair at the moment is equal to ; Indicates any.
5. The queue cruise control method based on successor features and transfer reinforcement learning according to claim 1 is characterized in that: The SF of the source domain in step 1 is used as the global shared successor feature. The overall Q value distribution of MAS reflecting any heterogeneous reward distribution is completed by the source domain MAS After learning, the target domain MAS is Migrate, and construct and iterate the Q value distribution of the target domain MAS; According to formula (6) Construct the target domain Q value and combine it with the real-time state transfer data of the target domain To learn, the steps are as follows: Step A reconstructs the state transition feature extraction network for the target domain MAS; Step B transfers data pairs in real-time state in the target domain are input and output samples; Step C takes the loss function shown in formula (4) of the target domain MAS as the target, and uses the classical neural network learning method, namely the BP algorithm and the gradient descent method, to learn the heterogeneous network branches of the network. Optimize the parameters of Step D replaces Step C The parameter optimization results are applied to the CACC task of intelligent connected vehicles to optimize the ICV platoon following performance.
6. A queue cruise control method based on successor features and transfer reinforcement learning according to claim 1, characterized in that: The transfer reinforcement learning of the successor features of the source domain in step 2 adopts the MADDPG algorithm and the CACC method of SF-MADDPG based on the successor features. The SF-MADDPG iterative algorithm adopts the actor-critic architecture of MADDPG. The actor completes the strategy mapping from state to action, the critic completes the value mapping of state-action data pairs to value functions, and the critic judges the long-term return of the actor's strategy.
7. A queue cruise control method based on successor features and transfer reinforcement learning according to claim 1, characterized in that: The step 1 completes the MAS global high-level state transfer features of M source domains during the source domain learning process. , subsequent features Learning and joint strategies The training refers to the learning of the global high-level state transition features and subsequent features of heterogeneous multi-agents in the source domain of knowledge. The number of source domains and training cycles are used as units to complete the cycle. State transition characteristics of the source domain , subsequent features and joint strategies The specific steps of training are as follows: Step 1.1 The source domain In the source domain, At the beginning of a training cycle, reset the environment, that is, initialize ;in, For the first ICV (intelligent agent) in =The state at time 0; Step 1.2 At this moment, according to the joint strategy given by the actor and status , we get the joint action of MAS ,in, ; For the first ICV (intelligent agent) in The state of the moment, For the ICV (Intelligent Agent) = Action at time 0 ; Step 1.3 Execute the MAS action , and get and MAS Total Rewards ; Step 1.4: Calculate state transition features ; Step 1.5: Store state transition data pairs To the experience replay pool , the experience replay pool is maintained and updated by the Internet of Vehicles data platform; Step 1.6: From the Experience Pool Randomly sample data Group data; Step 1.7: Design training loss functions based on formulas (4) and (5) and , and used for state transition characteristics Training of the extraction network: (4) Similarly, the shared backbone network model approximates The distribution of , then the training loss is designed as follows: (5) The following updates are performed on the feature extraction network: (9) (10) in, is the parameter of the shared backbone network The learning rate, It is for Heterogeneous network branch parameters corresponding to source domain MAS The learning rate; Step 1.8: Based on state transition characteristics , build a fully connected neural network to fit the subsequent features The distribution of the fully connected neural network is ,right Perform the following updates: (11) in, It is in strategy Next The subsequent characteristics of the step iteration; ; With the updated fully connected neural network parameters Calculate successor features ; is the parameter of the successor feature The learning rate; Step 1.9: According to formula (6), the Q value of MAS is obtained: (6) Step 1.10: Based on the actor-critic structure, the Q value is used as the critic to judge the decision performance, and the joint strategy of MAS based on the Q value To update: (12) in, , for and The KL divergence between .
8. The queue cruise control method based on successor features and transfer reinforcement learning according to claim 1 is characterized in that: The strategy in step 3 Will continuously update and iterate to the optimal strategy , so that the CACC task of the ICV queue in the target domain can achieve the optimal performance, which means that the transfer reinforcement learning based on the successor features in the knowledge transfer process is cycled in the target domain knowledge transfer process in units of training cycles to complete the joint strategy of the target domain The training finally optimizes the CACC strategy based on the car-following performance of the intelligent connected vehicle. The specific steps are as follows: Step 3.1: In the target domain, At the beginning of a training cycle, reset the environment, that is, initialize ; Step 3.2: At this moment, according to the joint strategy given by the actor and status , get the action ,in ; Step 3.3: Execute the target domain MAS action , and get and MAS Total Rewards ; Step 3.4: Utilize state feature network parameters transferred from the source domain , forward calculation to obtain the feature ; Step 3.5: Store state transition data pairs To the experience replay pool , Experience Replay Pool Maintained and updated by the Internet of Vehicles data platform; Step 3.6: From the Experience Pool Randomly sample data Group data; Step 3.7: Design training loss function based on formula (4) , and used for state transition feature extraction network Training: (4) in, is the heterogeneous network branch model parameter; For feature extraction network Perform the following updates: (10) Step 3.8: The Q value constructed by the reviewer contains the future discounted cumulative reward information. For each target domain MAS, the source domain number to be migrated is selected according to the maximum Q value principle. : (13) Among them, Q value The calculation is based on the state transition characteristics , subsequent features and source domain and joint strategies According to formula (6), MAS is obtained based on the The Q value under the knowledge of the source domain as a reference : (6) The first The state transition features, successor features, and joint strategies of the source domain will be transferred to the target domain; Step 3.9: Based on the actor-critic structure, the Q value is used as the critic to judge the decision performance, and the joint strategy of MAS based on the Q value To update: (14) in, is the target domain strategy learning rate, , for and The KL divergence between .
9. The queue cruise control method based on successor features and transfer reinforcement learning according to claim 1 is characterized in that: The final strategy in step 3 Will be continuously updated and iterated , so that the CACC task of the ICV platoon in the target domain achieves the optimal performance, which means that the CACC strategy is finally achieved based on the following performance of the intelligent connected vehicle Optimization.
10. A queue cruise control method based on successor features and transfer reinforcement learning according to any one of claims 1 to 9, characterized in that: The ICV fleet is considered as a multi-agent system (MAS), where each ICV is an agent and its model is expressed as: (15) in, and The vehicles status and control inputs, , and The vehicles The position, velocity and acceleration of Usually represents the desired acceleration. , is a given constant matrix that satisfies: (16) in, is the inertial time lag parameter of the vehicle longitudinal dynamics; In a connected environment, the CACC task of an intelligent vehicle described by equations (15) and (16) is modeled as a heterogeneous multi-agent MDP, which is mathematically described by a multi-tuple ; in, and Represent the state variables of the agent Space and action variables Belonging space; represents the state transition probability, Represents the reward function , is the discount factor, which is the discount term for the future reward value when the agent calculates the cumulative expected reward; the agent's strategy is generally used Represented by, and divided into random strategies: ~ and a deterministic strategy: ;by Indicates status Next select action The probability of maximizing the cumulative reward of the entire process is called the return, express The reward value at the moment is The reward function at the moment The definition is as follows: (17) The discount factor Existence and rewards The boundedness of a single agent Status is a vector, ,action For scalars, .
Citation Information
Cited By
Pilotless automobile queue control method based on physical information reinforcement learning
CN121232881A