Multi-Agent Cooperative Control Method for Traffic Scenarios
By applying multi-agent collaborative control methods in traffic scenarios, and using deep reinforcement learning, perceptual supervision and fusion feedback technology, the shortcomings of existing methods in communication resource consumption and detection results are solved, more efficient information utilization and model generalization are achieved, and the operation efficiency of the transportation system is improved.
Patent Information
- Application Number
- CN202411263550.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-09-10
AI Technical Summary
The existing multi-agent collaborative control method is difficult to meet the requirements of low communication resource consumption and high detection result accuracy in large-scale networked systems, and the controller performance is limited by the sensor's perception accuracy.
A multi-agent collaborative control method for traffic scenarios is proposed. Through deep reinforcement learning, perceptual supervision and fusion feedback technology, sensor control data information sharing and online evolutionary learning among multiple traffic agents is realized, forming a two-way optimization framework for the agent.
It improves information utilization and model generalization capabilities, enhances the self-evolution capabilities of the model, reduces communication overhead, and improves the operation efficiency of the transportation system.
Smart Images

Figure CN119129641B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent control, and particularly to a multi-agent collaborative control method for traffic scenarios. Background Art
[0002] With the wide application of Artificial Intelligence (AI) technology in modern traffic, various transportation tools are given the concept of agents, which are important components of the traffic system. As one of the common research objects of machine learning, an agent is usually defined as an entity in the environment that can obtain and interpret data reflecting events occurring in the environment, select actions based on the interpretation, and finally execute actions that have an impact on the environment, that is, it has the ability of perception and decision-making. An agent usually consists of a sensor, a controller, and an actuator, where the actuator that only executes actions is mainly related to the physical state of the agent itself. In traffic scenarios, the cooperation of multiple traffic agents is often required to cope with the complex and changing road environment and limited computing and communication resources. To meet different application requirements, many researchers focus on the field of multi-agent collaborative perception and control for research.
[0003] In terms of multi-agent collaborative perception in large-scale networked systems, different agents compensate each other by sharing their sensing information, thereby improving the effectiveness of subsequent local decisions. Multi-agent perception can be divided into three basic modes according to the different shared information: pre-fusion, mid-fusion, and post-fusion. The information shared in pre-fusion is the original data, and the fusion of each original data can be achieved without loss. However, transmitting the original data in the pre-fusion method will consume a large amount of communication resources and it is difficult to meet the scenarios with high requirements for perception real-time. On the contrary, post-fusion shares the perception results independently generated by each agent. The post-fusion consumes less communication bandwidth, but since the results are generated based on their respective incomplete observations, there are noises or errors in themselves, and the accuracy of the fusion results is low. In order to meet the requirements of both low communication resource consumption and high detection result accuracy, the mid-fusion method is proposed and has received more attention from researchers. Mid-fusion means that each agent processes the original data into intermediate layer features and then shares them, and the final perception result is obtained based on the fused features.
[0004] Meanwhile, the multi-agent cooperative control of large-scale networked systems optimizes the behavior of multiple agents by coordinating their actions. Different from traditional control methods for single agents, cooperative control technology relies on communication between agents to exchange information about their states and goals, must consider the complex interactions and dependencies between agents, and can adapt to changing conditions in real time. For example, the ramp merging strategy based on game theory is used to optimize the merging coordination of vehicles in mixed traffic, and this strategy can determine the dynamic merging order and corresponding longitudinal / lateral control. In addition to the multi-agent cooperative control method based on game theory, some researchers formulate the cooperative control problem as an optimization problem based on optimization methods, and the goal is to find the optimal actions for each agent to maximize the overall performance of the vehicle queue in mixed traffic. For example, the mixed-integer linear programming (MILP) model cooperatively optimizes the driving trajectories of autonomous vehicles from the perspective of the total vehicle delay. In addition, the method based on reinforcement learning trains the control algorithm through the trial-and-error method, that is, the algorithm learns from its own experience in the environment and adapts to the changing environment. For example, the architecture based on blockchain integrated multi-agent deep reinforcement learning (Block-MADRL).
[0005] Most of the existing methods are studied based on the framework where the sensor unidirectionally provides environmental information to the controller, and the multi-agent cooperative algorithms for the multi-agent cooperative perception and cooperative control in large-scale networked systems are executed separately. In this unidirectional framework, the upper limit of the controller performance is often limited by the sensing accuracy of the sensor. Therefore, there is an urgent need to introduce a sensing and control collaborative decision-making mechanism based on multi-agent online evolutionary learning, and through the reverse optimization channel from the controller to the sensor, form a two-way optimization framework for agents, so that the sensing and control units can be optimized collaboratively to adapt to the unknown changes in the dynamic environment and improve the operation efficiency of the traffic system. Summary of the Invention
[0006] To solve the above technical problems, the present invention proposes a multi-agent cooperative control method for traffic scenarios. In the method, various sensing and control task models are made to continuously improve their generalization in the environment with changing data distributions without artificial supervision information, and at the same time, decision sharing and decision improvement among traffic agents are realized. By applying local cooperative decision-making technology, the local decision-making subsystem of the large-scale networked system can quickly adapt to non-stationary environmental changes and improve the decision execution efficiency and accuracy.
[0007] To achieve the above object, the technical solution of the present invention is as follows:
[0008] A multi-agent cooperative control method for traffic scenarios includes the following steps:
[0009] Receive the sensing data information of multiple traffic agents in an intelligent transportation system based on deep reinforcement learning; the multiple agents include traffic agents with interaction functions and traffic agents without interaction functions, and the sensing data information includes multi-source perception information of traffic agents with and without interaction functions and feedback information of traffic agents with interaction functions; a perception end model is set on each traffic agent side;
[0010] Extract the multi-source perception features and feedback features of the multi-source perception information and feedback information, and perform multi-source heterogeneous information fusion on the multi-source perception features and feedback features to obtain the sensing fusion features;
[0011] Transmit the sensing fusion features to the multiple traffic agents in real time, adopt sensing supervision and fusion feedback technology to perform real-time information interaction with the multiple traffic agents to realize new category data detection, and perform incremental learning based on the new category data to realize online update and iteration of the model to assist the multiple traffic agents to update strategies and actions.
[0012] Preferably, the multi-source heterogeneous information fusion of the multi-source perception features and feedback features to obtain the sensing fusion feature V s , adopts self-supervised transfer learning, and the formula of its loss function is as follows:
[0013] E(θ n , θ o ) = ∑L(G n (V s , θ n ), y s ) - α∑L(G o (V S , θ o ), y s )
[0014] Among them, the sensing fusion feature V s represents the target domain; the source domain y s is the deterministic physical world feature in the control feedback; G n and G o represent the prediction results of the target classifier and the source domain classifier, L represents the classification loss, θ n and θ o are the internal parameters corresponding to the target classifier and the source domain classifier respectively, and α is a balance hyperparameter.
[0015] Preferably, adopt sensing supervision and fusion feedback technology to perform real-time information interaction with multiple traffic agents to realize new category data detection, and perform online training and learning based on the new category data to realize the update and iteration of the model to assist the multiple traffic agents to update strategies and actions, which specifically includes the following steps:
[0016] During the evolution of the data stream, the model jointly uses entropy and probability to obtain the score of new category data, and the formula is as follows:
[0017] Sco(V s ) = min(λc, 1)Prob(V s ) + max(1 - λc, 0)∑-V s logV s
[0018] In the formula, Sco is the score of the new data obtained according to V s . If it is higher than the set threshold, it is determined as a new category; c is the number of blocks starting from 0 in each new category cycle, and the max / min operator keeps the weight between (0, 1); λ is a trade-off parameter used to dynamically balance the two criteria; Prob is the degree of similarity between the new data and the newly discovered new categories;
[0019] When the perception-end model is updated according to the new category data, the change of the model parameters further updates the state information of the traffic agent. The traffic agent updates its strategy and actions according to the instant state information and the surrounding environment information, and the formula is as follows:
[0020]
[0021] In the formula, is the reward of action a in state s, is the state transition matrix, E π is the expectation function, R t+1 is the reward obtained by the agent at time t + 1, Q π is the action value function of the current policy function π, t is the current time, S t is the state of the agent at time t, and the specific value is represented by s, A t is the action to be executed by the agent at time t, and the specific value is represented by a, and γ is the discount factor with a value range of 0 - 1.
[0022] Preferably, it further includes: based on the iterated model, realizing the autonomous lane change and autonomous collision avoidance of the traffic agents with interactive functions in the mixed traffic section, and the specific steps are as follows:
[0023] During the movement of the traffic agent with interactive functions, the driving efficiency is composed of the path plan completion rate and the moving deceleration ratio, and its formula can be constructed as follows:
[0024] μ = δ·σ(r, a) + (1 - δ)·ρ(a)
[0025] Among them, δ is the balance coefficient, a represents the motion decision, r is the path decision, σ represents the path plan completion rate, which is the ratio of the successful completion of the path plan during the entire driving process of the vehicle, and ρ represents the moving deceleration ratio, which is the proportion of decelerated motion in all motion decisions taken by the vehicle during the entire driving process;
[0026] In the construction of the cost function of path and motion decision, based on the surrounding environment of traffic agents with interactive functions, the driving state of traditional vehicles, and the road environment, the total decision cost function θ can be constructed as:
[0027] θ = w s Q s (r, a) + w c Q c (r) + w d Q d (r, a) + Z
[0028] Among them, Q s 、Q c and Q d are the cost functions of static safety, comfort, and dynamic safety respectively, normalized to values between 0 and 1, w s 、w c and w d are the weights of the three cost functions, and Z measures the cost function of violating traffic rules during the driving process of the vehicle:
[0029]
[0030] Among them, Q g is the cost caused by road geometry, indicating that the vehicle is repelled by a force pointing to the passable lane when it exceeds the lane boundary, is the traffic rule constraint of the road section, Q a is the acceleration cost, w g 、 and w a are the weights of the three cost functions.
[0031] Based on the above technical solutions, the beneficial effects of the present invention are:
[0032] 1) The present invention improves the utilization rate of integrated information: The integration of multi-source information of multiple traffic agents realizes the full mining of the value of multi-source and multi-modal data information, and further expands the general ability of the model. In the multi-traffic-agent autonomous collision avoidance simulation experiment, compared with the model that does not adopt the multi-source information integration technology, the data utilization rate is increased by 5.4%;
[0033] 2) The present invention enhances the model generalization ability: Online evolutionary learning enables the model to be independent of artificial supervision information during the training process, thereby improving the model's generalization ability for novel categories. In the experiment, when facing new categories, the recognition accuracy of the model using online evolution is about 2.7% higher than that of the model without online evolution;
[0034] 3) The present invention increases the model self-evolution ability: The introduction of online evolutionary learning and reinforcement learning enables the model to collect real-time information on changes in the real environment for self-iteration and evolution, and can make self-adjustments according to the real-time changing environment to cope with various scenarios that may be encountered in the future;
[0035] 4) The present invention reduces communication overhead: Through the sensing and control cooperation technology, effective information can be transmitted in real time among multiple traffic agents and applied to the construction of local system knowledge sharing and regulation mechanisms. In the experiment, compared with the traditional QMIX algorithm, the training efficiency is significantly improved, and the communication overhead is reduced by 10%. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of a multi-agent cooperative control method for traffic scenarios in an embodiment;
[0037] Figure 2 is a schematic diagram of multi-source heterogeneous information fusion in a multi-agent cooperative control method for traffic scenarios in an embodiment;
[0038] Figure 3 is a schematic diagram of cooperative online evolutionary learning in a multi-agent cooperative control method for traffic scenarios in an embodiment;
[0039] Figure 4 is a schematic diagram of the sensing and control cooperative online evolution framework in a multi-agent cooperative control method for traffic scenarios in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0041] As Figures 1 to 4 shown, in this embodiment, aiming at the local autonomous decision-making problem in the intelligent transportation system for traffic scenarios, taking self-supervised learning and multi-agent reinforcement learning theories as the main theoretical basis, aiming at the problems of association integration of agents and adaptive learning of perception and control tasks, a multi-agent cooperative control method for traffic scenarios is proposed through the association information integration technology, discriminant information fusion technology, and cooperative perception and control data adaptive learning technology among multiple traffic agents, including the following steps:
[0042] Step 1: Receive the sensing data information of multiple traffic intelligent agents in the intelligent transportation system based on deep reinforcement learning; the multiple agents include traffic intelligent agents with interaction functions and those without interaction functions, and the sensing data information includes multi-source perception information of traffic intelligent agents with and without interaction functions and feedback information of traffic intelligent agents with interaction functions; a perception end model is set on each traffic intelligent agent side.
[0043] In this embodiment, the objects are traffic intelligent agents with interaction functions and those without interaction functions in a mixed traffic scenario. The traffic intelligent agents with interaction functions include buses, autonomous vehicles, drone equipment, etc., and the traffic intelligent agents without interaction functions include traditional vehicles, roadside devices, etc. Basic preparation stage: 1) Data preparation: The sensors of buses, autonomous vehicles, traditional vehicles, roadside devices, and drone equipment capture the surrounding environment data and perform necessary preprocessing; 2) Basic model loading: Buses, autonomous vehicles, traditional vehicles, roadside devices, and drones load the perception end models suitable for themselves to ensure that the traffic intelligent agents have basic capabilities.
[0044] Step 2: Extract the multi-source perception features and feedback features of the multi-source perception information and feedback information, and perform multi-source heterogeneous information fusion on the multi-source perception features and feedback features to obtain the sensing and control fusion features.
[0045] In this embodiment, multi-source heterogeneous information fusion is adopted: Each traffic intelligent agent transmits the data after feature extraction to the multi-source information fusion technology module, and transmits the fused data to the information fusion center for aggregation processing to obtain the sensing and control fusion feature V s . The specific description is as follows:
[0046] Based on the theory of hierarchical network system modeling and characterization, each intelligent agent in the hierarchical community structure model has independent dynamic behaviors. Traffic intelligent agents can independently adjust their states such as speed and steering. Through the analysis and judgment of the paths and motion planning of multiple traffic intelligent agents, the deep associations between the sensing data of multiple traffic intelligent agents and their paths and motion tasks can be effectively summarized. First, as Figure 2 shown, perform joint modeling on multi-source heterogeneous information (images, point clouds, ultrasonic waves, controller information, etc.):
[0047] V S =φ(P1, P2, …, P n , F1, F2, …, F m )
[0048] where φ is the feature fusion function, and V s represents the multi-source sensing and control data features after fusion in different scenarios, and P nDenote the multi-source perception features of the same dimension obtained by using different feature extraction networks for different traffic agents as F m Denote the feedback features extracted from the feedback information of the controller of the traffic agent with interaction function. Input the sensing and control fusion feature V s into the subsequent task network to realize the online evolution of multiple traffic agents.
[0049] Based on the self-classifiers of multiple traffic agents, in order to better utilize the subsequent real-time obtained sensing and control fusion information for online evolutionary learning, a self-supervised domain adaptation method is adopted, that is, the source domain and the target domain have the same task, the data in the two domains are similar but different, there are a large number of labels in the source domain, and there are no labels in the target domain. The sensing and control fusion feature V s represents the target domain, while the source domain y s is the deterministic physical world feature in the control feedback. Therefore, it is hoped to train a mapping from y s to V s on the source domain through domain adaptation, and use this mapping to estimate V s constrained by y s . At this time, the loss for self-supervised transfer learning can be obtained:
[0050] E(θ n , θ o ) = ∑L(G n (V s , θ n ), y s ) - α∑L(G o (V S , θ o ), y s )
[0051] where Q n and G o represent the prediction results of the target classifier and the source domain classifier, L represents the classification loss, θ n and θ o are the internal parameters corresponding to the target classifier and the source domain classifier respectively, and α is the balancing hyperparameter. For the multi-source information perceived by the traffic agent and the control feedback information of the interactive traffic agent, an adaptive method is used to realize the self-supervised constraint of the deterministic physical world feature in the control feedback on the sensing and control information fusion, and obtain a more accurate overall feature estimate V s than any single traffic agent participating in the fusion estimation. By studying the real-time optimization and representation learning of multiple proxy tasks, a pre-trained model suitable for the overall goal is obtained to realize the accommodation of complex knowledge of multi-source data.
[0052] Step 3, input the sensing and control fusion feature V sTransmitted to multiple traffic agents in real time, and adopt perception supervision and fusion feedback technology to interact with multiple traffic agents in real time for information to achieve new category data detection, and perform incremental learning based on the new category data to realize the online update and iteration of the model to assist the multiple traffic agents to update strategies and actions.
[0053] In this embodiment, the perception supervision and fusion feedback technology: according to the real-time states of each traffic agent and the surrounding environment information, and according to the different capabilities of each traffic agent, adopt the perception supervision and fusion feedback technology to interact with the traffic agent in real time, so that the traffic agent can fully master the surrounding information for the next decision-making. Collaborative online evolutionary learning: The perception-side models of each traffic agent achieve new category data detection under the multi-traffic agent perception-control collaborative framework, and perform training and learning based on the new category data to realize the update and iteration of the model. Specifically as follows:
[0054] Design a multi-traffic agent collaborative online evolutionary learning method to achieve the adaptive learning of multiple traffic agents for tasks and reduce human intervention. Through the perception-control fusion of the perception information of multiple traffic agents and the control feedback information of the interactive traffic agents, and at the same time based on the information sharing mechanism of each traffic agent among the multi-traffic agent systems, use more accurate perception-control fusion feature V s Give the corresponding supervision information to each perception traffic agent to realize the online model optimization of the perception traffic agent; at the same time, give more comprehensive feedback information to each control traffic agent to realize the online model optimization of the control traffic agent. The above process is as Figure 3 Example.
[0055] Based on the model deployed at the perception end of each traffic agent, it is initially trained for several known categories, and the features of the new category are unknown. During the evolution of the data stream, the model jointly obtains the new data score using entropy and probability:
[0056] Sco(V s ) = min(λc, 1)Prob(V s ) + max(1 - λc, 0)∑-V s logV s
[0057] In the formula, Sco is based on V sThe newly obtained data score, if higher than the set threshold, is determined as a new class; c is the number of blocks starting from 0 in each new class cycle, and the max / min operator keeps the weight between (0,1); λ is a trade-off parameter used to dynamically balance the two criteria; Prob is the degree of similarity between the new data and the newly discovered new classes. If an instance is similar to the new class prototype, it is very likely to be a new class. Here, it is the normalized similarity of the novel class. Different new classes emerging at different stages are defined according to the new data score. After the new class data is updated, the model conducts incremental learning based on the new class data. As the data stream evolves, the model accumulates prior information and establishes appropriate feature prototypes for new instances for subsequent discrimination of new classes.
[0058] After the perception-end model is updated according to the new class data, the changes in the model parameters further update the state information of the traffic agent. The traffic agent updates its strategy and actions based on the immediate state information and the surrounding environment information:
[0059]
[0060] In the formula, is the reward for action a in state s, is the state transition matrix, E π is the expectation function, R t+1 is the reward obtained by the agent at time t + 1, Q π is the action value function of the current policy function π, t is the current time, S t is the state where the agent is at time t, specifically represented by s, A t is the action that the agent needs to execute at time t, specifically represented by a, and γ is the discount factor with a value range of 0 - 1. The traffic agent continuously takes different actions according to the surrounding environment information, and at the same time its state is updated in real time and dynamically. In this process of dynamic change and adjustment, the perception information and the control information achieve real-time interaction and coordination.
[0061] The advantages of the multi-traffic-agent perception and control coordination are reflected in: 1) The controlled traffic agent can continuously obtain more accurate fusion perception information and more comprehensive control feedback information from V s and can continuously optimize the online adaptation of the controller to non-stationary environmental changes to achieve better control effects and response speeds; 2) The perceptual traffic agent can continuously obtain more accurate feature supervision information from V s than its own perceptual feature estimation, and at the same time can continuously optimize the online adaptation of the perceptron to non-stationary environmental changes to achieve better perception capabilities; 3) Under the framework of multi-traffic agents, the supervised learning and reinforcement learning based on deep networks can be unified, and an online evolutionary learning independent of manually labeled information can be formed.
[0062] Study the impact on online evolutionary learning by exploring the correlation between the real-time rewards obtained by the controller and the perception information in multi-traffic-agent reinforcement learning; for path and motion planning tasks, study and establish a corresponding online evolutionary learning simulation environment, conduct a flexible search for the number of traffic agents, and establish a reliable traffic agent large-scale deployment plan for sensor-control collaboration through multiple rounds of judging the impact relationship of the number of traffic agents on the overall, individual performance, and generalization ability.
[0063] In a multi-agent collaborative control method for traffic scenarios in an embodiment, it further includes providing a model based on iteration to implement the autonomous lane-changing and autonomous collision avoidance processes of traffic agents with interactive functions in a mixed traffic section. The specific steps of this process are as follows:
[0064] In a mixed traffic scenario, based on the multi-source heterogeneous information fusion mechanism of multi-traffic agents and the multi-traffic-agent sensor-control collaborative online evolutionary method, autonomous lane-changing and autonomous collision avoidance of vehicles in a mixed traffic section can be realized, improving the driving efficiency, and thus enhancing the traffic efficiency of the section. In the processing of collaborative path and motion decision tasks of multiple physical control traffic agents, realizing efficient autonomous reasoning and decision-making of multi-traffic agents is an urgent problem to be solved in this task. In a mixed traffic scenario, improving the driving efficiency of vehicles is the main goal, and at the same time, it is also crucial to define and measure the costs generated by vehicles taking different paths and motion decisions during driving. Therefore, the goal of multi-traffic-agent collaborative decision-making is transformed into maximizing the driving efficiency μ of vehicles while minimizing the decision cost θ generated by vehicle collisions and traffic rule violations.
[0065] During the movement of physical traffic agents, the driving efficiency is mainly composed of the path plan completion rate and the moving deceleration ratio, and its formula can be constructed as follows:
[0066] μ = δ·σ(r,a) + (1 - δ)·ρ(a)
[0067] Where δ is the balance coefficient, a represents motion decisions such as acceleration, deceleration, and parking, r is the path decision such as lane-changing behaviors like changing to the left lane or the right lane. σ represents the path plan completion rate, which is the ratio of the successful completion of the path plan by the vehicle during the entire driving process, and ρ represents the moving deceleration ratio, which is the proportion of decelerated motion in all motion decisions taken by the vehicle during the entire driving process.
[0068] In the construction of the path and motion decision cost function, based on the surrounding environment of physical control traffic agents, the driving state of traditional vehicles, and the road environment, the total decision cost function θ can be constructed as:
[0069] θ = w s Q s (r,a) + w c Q c(r)+w d Q d (r,a)+Z
[0070] where Q s 、Q c and Q d are the cost functions of static safety (the cost of avoiding static obstacles on the road), comfort, and dynamic safety (the cost of avoiding obstacles for traditional vehicles on the road), respectively, and are normalized to values between 0 and 1. w s 、w c and w d are the weights of the three cost functions. A higher w s may result in a safer but slower lane-changing path, while a higher w c may result in a shorter path, and a higher w d may cause the vehicle to maintain its original speed and overtake moving vehicles. Z measures the cost function of violating traffic rules during vehicle driving:
[0071]
[0072] where, Q g is the cost generated by road geometry, indicating that when the vehicle exceeds the lane boundary, it receives a repulsive force (cost) pointing to the passable lane. is the traffic rule constraint of the road section. Q a is the acceleration cost, mainly due to tire friction limitation and vehicle throttle limitation. w g 、 and w a are the weights of the three cost functions.
[0073] At the road section layer of mixed traffic, multiple physical traffic intelligent agents (vehicles, roadside units, drones, etc.) follow traffic rules, traffic signal instructions, and the regulation instructions of the upper-layer road section-level traffic intelligent agents. Through the autonomous decision-making of sensing and cooperation, the physical control traffic intelligent agents maximize the driving efficiency μ and minimize the decision cost θ to obtain their respective real-time paths and autonomous motion decisions.
[0074] It should be understood that although the steps in the above flowchart are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear description in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed and completed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0075] The above is only the preferred implementation mode of the multi-agent collaborative control method for traffic scenarios disclosed in the present invention, and is not used to limit the protection scope of the embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of this specification shall be included within the protection scope of the embodiments of this specification.
Claims
1. A multi-agent collaborative control method for traffic scenarios, characterized in that: The steps include: Receiving sensory control data information of multiple traffic agents in an intelligent traffic system based on deep reinforcement learning; the multiple agents include traffic agents with interactive functions and traffic agents without interactive functions, and the sensory control data information includes multi-source perception information of traffic agents with interactive functions and traffic agents without interactive functions and feedback information of traffic agents with interactive functions; each traffic agent is provided with a perception end model; Extracting multi-source sensing features and feedback features of the multi-source sensing information and feedback information, and performing multi-source heterogeneous information fusion on the multi-source sensing features and feedback features to obtain sensing and control fusion features; The sensing and control fusion features are transmitted to multiple traffic agents in real time. The sensing supervision and fusion feedback technology is used to interact with multiple traffic agents in real time to detect new category data. Incremental learning is performed based on new category data to achieve online update iteration of the model to assist multiple traffic agents in updating strategies and actions. The specific steps include the following: During the evolution of the data stream, the model uses entropy and probability to jointly derive the score of the new category data. The formula is as follows: Sco(V s )=min(λc,1)Prob(V s )+max(1-λc,0)Σ-V s logV s In the formula, Sco is the sensor-control fusion feature V s If the score of the new data is higher than the set threshold, it is judged as a new class; c is the number of blocks starting from 0 in each new class cycle, and the max / min operator keeps the weight between (0,1); λ is a trade-off parameter used to dynamically balance the two criteria; Prob is the degree of similarity between the new data and the discovered new class; When the perception model is updated according to the new category data, the change of model parameters further updates the state information of the traffic agent. The traffic agent updates its strategy and action according to the real-time state information and surrounding environment information. The formula is as follows: In the formula, is the reward for action a in state s, is the state transfer matrix, E π is the expected function, R t+1 is the reward obtained by the agent at time t+1, Q π is the action value function of the current policy function π, t is the current moment, S t is the state of the agent at time t, and its specific value is represented by s. t is the action that the agent is to perform at time t. Its specific value is represented by a. γ is the discount factor, and its value range is 0-1.
2. The multi-agent collaborative control method for traffic scenarios according to claim 1 is characterized in that: The multi-source heterogeneous information fusion of multi-source perception features and feedback features is performed to obtain the sensing and control fusion feature V s , self-supervised transfer learning is used, and the formula of its loss function is as follows: E(θ n ,i o )=ΣL(G n (V s ,i n ),y s )-αΣL(G o (V S ,i o ),y s ) Among them, the sensing and control fusion feature V s represents the target domain; the source domain y s To control the deterministic physical world characteristics in feedback; G n and G o represents the prediction results of the target classifier and the source domain classifier, L represents the classification loss, and θ n and θ o They correspond to the internal parameters of the target classifier and the source domain classifier respectively, and α is a balancing hyperparameter.
3. The multi-agent collaborative control method for traffic scenarios according to claim 1 is characterized in that: Also includes: Based on the iterative model, autonomous lane changing and autonomous collision avoidance of traffic agents with interactive functions in mixed traffic sections are realized. The specific steps are as follows: During the movement of the interactive traffic agent, the driving efficiency is composed of the path plan completion rate and the moving deceleration ratio, and its formula can be constructed as follows: μ=δ·σ(r,a)+(1-δ)·ρ(a) Among them, δ is the balance coefficient, a represents the motion decision, r is the path decision, σ represents the path plan completion rate, which is the ratio of the vehicle's path plan successfully completed during the entire driving process, and ρ represents the mobile deceleration ratio, which is the proportion of the vehicle's deceleration movement in all the motion decisions taken during the entire driving process; In the construction of the path and motion decision cost function, based on the surrounding environment of the interactive traffic agent, the driving status of traditional vehicles and the road environment, the total decision cost function Can be constructed as: Where Q s , Q c and Q d are the cost functions for static safety, comfort, and dynamic safety, normalized to values between 0 and 1, and w s 、w c and w d It is the weight of the three cost functions. Z measures the cost function of violating traffic rules during vehicle driving: Among them, Q g The cost generated by the road geometry indicates that the vehicle exceeds the lane boundary and is repelled toward the passable lane. is the traffic rule constraint of the road section, Q a is the acceleration cost, w g , and w a are the weights of the three cost functions.
Citation Information
Patent Citations
Signal lamp cooperative control method based on multi-agent reinforcement learning and multi-mode signal perception
CN116612636A