Vehicle-road collaborative dynamic scheduling system and method based on multi-agent reinforcement learning
Through the vehicle-road collaborative dynamic scheduling system with multi-agent reinforcement learning, the dual-network collaborative optimization driven by topological dynamic state representation and matrix decomposition are solved, and the high-dimensional state space and multi-agent collaborative decision-making problems in complex traffic environments are achieved, efficient coordinated scheduling of vehicles and roadside facilities is achieved, and the real-time and global optimization capabilities of the traffic system are improved.
Patent Information
- Application Number
- CN202510521087.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Existing vehicle-road collaboration technology is difficult to effectively deal with the high-dimensional state space and multi-agent collaborative decision-making problems in complex traffic environments, resulting in high computational complexity, poor real-time performance and insufficient global optimization.
The vehicle-road collaborative dynamic scheduling system with multi-agent reinforcement learning is adopted, including the perception layer, the decision-making layer and the execution layer. Through topological dynamic state representation, dual-network collaborative optimization driven by matrix decomposition, and multi-agent collaborative decision-making mechanism driven by probability graph, it realizes efficient collaborative scheduling between vehicles and roadside facilities.
It reduces the computational complexity, improves the system's adaptability to complex traffic scenarios, enhances the consistency of scheduling efficiency and global optimization, reduces communication overhead, and adapts to urban traffic environments of different scales and complexities.
Smart Images

Figure CN120048123B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and intelligent transportation, and particularly to a vehicle-road collaborative dynamic scheduling system and method based on multi-agent reinforcement learning technology for realizing collaborative optimization management of intelligent vehicles and roadside facilities. Background Art
[0002] With the acceleration of the urbanization process, the problem of urban traffic congestion has become increasingly severe, seriously affecting the travel efficiency and quality of life of residents. Traditional traffic management methods mainly rely on fixed signal light cycle control and manual intervention, which cannot adapt to the dynamically changing traffic flow, resulting in low utilization rate of traffic resources. In recent years, with the rapid development of vehicle networking technology and artificial intelligence, vehicle-road collaboration, as the core technology of the next-generation intelligent transportation system, has gradually emerged.
[0003] Currently, vehicle-road collaboration technologies are mainly divided into three categories: one is the method based on centralized control, which uniformly schedules vehicles and roadside facilities through a central server; the second is the method based on distributed control, which realizes collaboration through local decisions of vehicles and roadside facilities; the third is the method based on reinforcement learning, which learns the optimal control strategy through the interaction between agents and the environment. However, the existing technologies have the following deficiencies: the centralized control method has a high computational complexity and is difficult to meet the real-time requirements; the distributed control method cannot guarantee the global optimum; and the existing reinforcement learning methods are difficult to handle the high-dimensional state space and multi-agent collaborative decision-making problems in the vehicle-road environment.
[0004] To solve the above problems, there is an urgent need for a multi-agent reinforcement learning system that can efficiently handle vehicle-road collaborative decision-making in complex traffic environments. Summary of the Invention
[0005] The purpose of the present invention is to provide a vehicle-road collaborative dynamic scheduling system and method based on multi-agent reinforcement learning, aiming to solve the problems in the prior art such as difficult efficient processing of complex traffic environment state representation, vehicle-road collaborative decision-making, and multi-agent information interaction.
[0006] The present invention discloses a vehicle-road collaborative dynamic scheduling system based on multi-agent reinforcement learning, including:
[0007] A perception layer, including a roadside perception module and a vehicle perception module, for collecting environmental state information and vehicle information;
[0008] A decision layer, communicatively connected to the perception layer, including:
[0009] A topological dynamic state representation unit, for receiving the environmental state information and the vehicle information, mapping the environmental state information and the vehicle information into a multi-dimensional topological space to form a topological state representation, and generating a manifold space representation through non-linear dimensionality reduction;
[0010] A dual-network collaborative decision-making unit, which is used to receive the manifold space representation and includes a Vehicle Dynamic Distribution Policy Gradient (VDDPG) network and a Roadside Dynamic Scheduling Policy Gradient (RDDPG) network. The VDDPG is used to generate a vehicle scheduling strategy, and the RDDPG is used to generate a roadside scheduling strategy, and the vehicle scheduling strategy and the roadside scheduling strategy are converted into a low-rank representation form through policy matrix decomposition;
[0011] A multi-agent collaborative optimization unit, which is used to construct an agent relationship graph, optimize the communication strategy in the agent relationship graph based on information entropy, and coordinate the vehicle scheduling strategy and the roadside scheduling strategy through a variational inference method to generate an optimized collaborative scheduling plan;
[0012] An execution layer, which is communicatively connected to the decision-making layer and includes a vehicle execution module and a roadside control module, is used to receive the collaborative scheduling plan, execute corresponding scheduling instructions, and feedback the execution result to the perception layer to form a closed-loop control.
[0013] Preferably, the topological dynamic state characterization unit includes:
[0014] A state acquisition sub-unit, which is used to receive the environmental state information and the vehicle information, and preprocess the environmental state information and the vehicle information to generate standardized data;
[0015] A topological characterization sub-unit, which is used to map the standardized data to a multi-dimensional topological space, extract topological features, and establish a local coordinate system;
[0016] A dimension conversion sub-unit, which is used to perform non-linear dimensionality reduction on the data in the multi-dimensional topological space to generate a manifold space representation that retains key topological relationships.
[0017] Preferably, the dual-network collaborative decision-making unit includes:
[0018] The VDDPG Actor network, which is used to receive the manifold space representation related to the vehicle and generate a vehicle scheduling strategy;
[0019] The VDDPG Critic network, which is used to evaluate the value of the vehicle scheduling strategy;
[0020] The RDDPG Actor network, which is used to receive the manifold space representation related to the roadside and generate a roadside scheduling strategy;
[0021] The RDDPG Critic network, which is used to evaluate the value of the roadside scheduling strategy;
[0022] A policy characterization sub-unit, which is used to convert the vehicle scheduling strategy and the roadside scheduling strategy into a policy matrix and perform matrix decomposition to generate a low-rank representation;
[0023] A collaborative optimization subunit, configured to identify potential conflicts between the vehicle scheduling strategy and the roadside scheduling strategy, and adjust the strategy parameters by the conjugate gradient method.
[0024] Preferably, the multi-agent collaborative optimization unit includes:
[0025] An agent relationship modeling subunit, configured to construct an agent relationship graph based on the current traffic conditions and identify the conditional dependence relationships between agents;
[0026] A message passing optimization subunit, configured to calculate the information entropy of each communication channel in the agent relationship graph, generate an optimal message passing strategy, and allocate communication resources;
[0027] A global consistency optimization subunit, configured to decompose the global optimization objective into local sub-objectives, coordinate local decisions by the variational inference method, and verify the global consistency of the final decision.
[0028] Preferably, the perception layer further includes:
[0029] A data preprocessing module, configured to perform filtering, denoising, and standardization processing on the environmental state information and the vehicle information;
[0030] A state caching module, configured to store historical environmental state information and vehicle information to support time series analysis;
[0031] An attention allocation module, configured to adjust the allocation strategy of perception resources according to the feedback of the decision-making layer.
[0032] Preferably, the execution layer further includes:
[0033] An instruction parsing module, configured to convert the collaborative scheduling scheme into specific control instructions;
[0034] An execution monitoring module, configured to monitor the execution of control instructions and identify abnormal states;
[0035] An execution history record module, configured to maintain execution history data for reference by the decision-making layer;
[0036] A degradation processing module, configured to execute a preset degradation strategy in case of communication interruption or device failure.
[0037] Preferably, between the vehicle execution module and the roadside control module, there is provided:
[0038] A timing synchronization unit, configured to ensure the time synchronization of vehicle control instructions and roadside control instructions;
[0039] A conflict detection unit, configured to detect potential conflicts in the execution process in real time and trigger emergency processing;
[0040] A collaborative effect evaluation unit, which is used to quantitatively evaluate the execution effect of vehicle-road collaborative control and generate an effect evaluation report.
[0041] Preferably, information exchange is realized between the VDDPG network and the RDDPG network through a shared hidden layer, where:
[0042] The shared hidden layer receives the common features of the manifold space representation and outputs an intermediate feature representation;
[0043] The VDDPG network and the RDDPG network respectively receive the intermediate feature representation, and combine their respective specific input features to generate corresponding policy outputs;
[0044] The VDDPG network and the RDDPG network realize collaborative parameter update through a gradient locking mechanism to ensure consistent policies.
[0045] Preferably, the system further includes:
[0046] An experience replay module, which is used to store the historical data of the interaction between the system and the environment and support offline batch learning;
[0047] A target network update module, which is used to periodically copy parameters from the main network to the target network to ensure learning stability;
[0048] An adaptive learning rate adjustment module, which is used to dynamically adjust the network learning rate according to the training progress;
[0049] A model evaluation and deployment module, which is used to evaluate the model performance and deploy the trained model to the production environment.
[0050] A vehicle-road collaborative dynamic scheduling method based on multi-agent reinforcement learning, which is applied to the vehicle-road collaborative dynamic scheduling system based on multi-agent reinforcement learning, and includes the following steps:
[0051] Initialize the system, including initializing the parameters of the VDDPG network and the RDDPG network, establishing a communication connection, loading the road network topology structure and historical traffic data;
[0052] Collect state information, including collecting environmental state information through the roadside perception module and collecting vehicle information through the vehicle perception module;
[0053] Execute topological state characterization, including mapping the state information to a multi-dimensional topological space, extracting topological features, and generating a manifold space representation through non-linear dimensionality reduction;
[0054] Generate dual-network decisions, including the VDDPG Actor network generating a vehicle scheduling strategy, the RDDPG Actor network generating a roadside scheduling strategy, and executing policy matrix decomposition to generate a low-rank representation;
[0055] Perform collaborative optimization, including constructing an agent relationship graph, optimizing the message passing strategy, coordinating local decisions through variational inference methods, and verifying global consistency;
[0056] Execute the scheduling plan, including converting the optimized decision into specific execution instructions and having vehicles and roadside facilities execute the control instructions;
[0057] Collect execution feedback, including monitoring execution results, updating the environmental state, updating network parameters and the experience pool based on execution results;
[0058] Iteratively optimize, repeat the above steps, and continuously optimize the scheduling effect.
[0059] The present invention adopts three innovative mechanisms: topological dynamic state representation, matrix factorization-driven dual-network collaborative optimization, and probabilistic graph-driven multi-agent collaborative decision-making, and constructs a complete multi-agent reinforcement learning vehicle-road collaborative dynamic scheduling system, which has the following beneficial effects:
[0060] 1) Through the topological dynamic state representation mechanism, the high-dimensional and complex traffic environment state is mapped to a low-dimensional manifold space, key topological features are retained, the computational complexity is reduced, and the system's adaptability to complex traffic scenarios is improved.
[0061] 2) Through the matrix factorization-driven dual-network collaborative optimization mechanism, collaborative decision-making between the vehicle network and the roadside network is achieved, the conflict problem between vehicle strategies and roadside strategies is solved, and the scheduling efficiency is improved.
[0062] 3) Through the probabilistic graph-driven multi-agent collaborative decision-making mechanism, the information interaction strategy between agents is optimized, the communication overhead is reduced, and the consistency between local decisions and global optimization goals is ensured.
[0063] 4) The overall system adopts a hierarchical modular design, has good scalability and adaptability, and can adapt to urban traffic environments of different scales and complexities. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below.
[0065] Figure 1 is a schematic diagram of the overall architecture of the vehicle-road collaborative dynamic scheduling system of multi-agent reinforcement learning of the present invention.
[0066] Figure 2 is a schematic diagram of the structure of the topological dynamic state representation unit of the present invention.
[0067] Figure 3 is a schematic diagram of the structure of the dual-network collaborative decision-making unit of the present invention.
[0068] Figure 4 It is a schematic structural diagram of the multi-agent collaborative optimization unit of the present invention.
[0069] Figure 5 It is a flowchart of the vehicle-road collaborative dynamic scheduling method of the present invention.
[0070] Figure 6 It is a schematic diagram of the topological state characterization process in an embodiment of the present invention.
[0071] Figure 7 It is a schematic diagram of the dual-network collaborative optimization process in an embodiment of the present invention.
[0072] Figure 8 It is a schematic diagram of the multi-agent collaborative decision-making process in an embodiment of the present invention. Detailed implementation manners
[0073] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0074] Referring to Figure 1 , the vehicle-road collaborative dynamic scheduling system based on multi-agent reinforcement learning provided by the present invention includes three main parts: a perception layer 101, a decision-making layer 102, and an execution layer 103.
[0075] The perception layer 101 includes a roadside perception module 111 and a vehicle perception module 112, which are used to collect environmental state information and vehicle information. The roadside perception module 111 can be various sensing devices distributed in the urban road network, such as cameras, radars, signal controllers, etc., which are used to collect environmental information such as traffic flow, signal states, and road network structures. The vehicle perception module 112 can be various sensors and communication devices installed on the vehicle, which are used to collect vehicle information such as vehicle position, speed, direction, and destination.
[0076] The decision-making layer 102 is the core part of the system, including a topological dynamic state characterization unit 121, a dual-network collaborative decision-making unit 122, and a multi-agent collaborative optimization unit 123. The topological dynamic state characterization unit 121 is used to receive environmental state information and vehicle information, map this information to a multi-dimensional topological space to form a topological state representation, and generate a manifold space representation through non-linear dimensionality reduction. The dual-network collaborative decision-making unit 122 includes a vehicle dynamic distribution network VDDPG and a roadside dynamic scheduling network RDDPG, which are respectively used to generate vehicle scheduling strategies and roadside scheduling strategies, and convert these strategies into a low-rank representation form through policy matrix decomposition. The multi-agent collaborative optimization unit 123 is used to construct an agent relationship graph, optimize the communication strategy, coordinate the decisions of each agent, and generate a final collaborative scheduling plan.
[0077] The execution layer 103 includes a vehicle execution module 131 and a roadside control module 132, which are used to receive the collaborative scheduling scheme, execute the corresponding scheduling instructions, and feedback the execution results to the perception layer 101 to form a closed-loop control. The vehicle execution module 131 is responsible for controlling the driving behavior of the vehicle, such as adjusting the speed, changing the route, etc. The roadside control module 132 is responsible for controlling the operating state of roadside facilities, such as adjusting the signal timing, restricting access, etc.
[0078] Refer to Figure 2 , the topological dynamic state characterization unit 121 includes a state acquisition sub-unit 1211, a topological characterization sub-unit 1212, and a dimension conversion sub-unit 1213.
[0079] The state acquisition sub-unit 1211 is used to receive the environmental state information and vehicle information from the perception layer 101, preprocess these raw data, and generate standardized data. In an embodiment of the present invention, the preprocessing includes operations such as denoising, filtering, and standardization. For example, for the vehicle speed data, mean filtering can be used to remove noise, and then maximum-minimum normalization is performed to map the speed value to the interval [0,1].
[0080] The topological characterization sub-unit 1212 is used to map the standardized data to a multi-dimensional topological space, extract topological features, and establish a local coordinate system. Specifically, first, the state space of the vehicle-road environment is defined , where the vehicle state (position, speed, direction) and the roadside state (signal state, road network topology) respectively constitute sub-spaces and . For each state point , the vehicle-road environment is accurately expressed through a local coordinate chart. Preferably, the topological characterization sub-unit 1212 also constructs a diffeomorphic mapping from the state space to the policy space , ensuring the topological invariance between the state change and the policy adjustment, and enabling the system to maintain stability in a complex traffic environment. The diffeomorphic mapping can be defined in the following way:
[0081] ,
[0082] where, is the state point, is the weight matrix, is the non-linear activation function, is the bias vector. Through this mapping, points with similar topological structures in the state space can be mapped to nearby points in the policy space, ensuring the continuity and stability of the policy.
[0083] The dimension conversion subunit 1213 is used to perform non - linear dimensionality reduction on the data of the multi - dimensional topological space, generating a manifold space representation that preserves the key topological relationships. In an embodiment of the present invention, the locally linear embedding (LLE) algorithm can be used to implement non - linear dimensionality reduction. The core idea of the LLE algorithm is to maintain the local linear relationship of data points while non - linearly reducing the dimension globally. Specifically, for each data point , first find its nearest neighbor points , then calculate the weight matrix such that can be reconstructed as a linear combination of its neighbors:
[0084] ,
[0085] where, represents the contribution weight of point to point , satisfying . Then, find the representation
[0086] in the low - dimensional space that satisfies the same weight relationship
[0087] In this way, the representation of the original high - dimensional data on the low - dimensional manifold can be obtained. In practical applications, the value of the number of nearest neighbors can be dynamically adjusted according to the complexity of the traffic scenario. For example, in the case of heavy traffic, a larger value (such as ) can be selected to capture more interaction relationships; in the case of light traffic, a smaller value (such as ) can be selected to reduce the computational amount.
[0088] The policy characterization subunit 1225 is used to convert the vehicle scheduling policy and the roadside scheduling policy into a policy matrix and perform matrix decomposition to generate a low - rank representation. Specifically, represent the policies of the two networks as a high - order matrix , and then decompose it into a combination of a core tensor and factor matrices through Tucker decomposition:
[0089] ,
[0090] where, is the core tensor, , , are factor matrices, represents the tensor - matrix product along the th mode. In this way, the complexity of the policy representation can be greatly reduced. For example, for a The policy tensor, if reduced to rank , the number of parameters can be reduced from to , greatly reducing the storage and computing requirements.
[0091] The cooperative optimization subunit 1226 is used to identify potential conflicts between the vehicle scheduling policy and the roadside scheduling policy, and adjusts the policy parameters through the conjugate gradient method. In a specific embodiment, the cooperative optimization can be expressed as the following optimization problem:
[0092] ,
[0093] where and are the parameters of the VDDPG network and the RDDPG network respectively, and are the independent loss functions of the two networks respectively, is the joint loss function, which is used to measure the coordination degree of the output policies of the two networks, is the trade-off coefficient, which is used to adjust the ratio of independent optimization to cooperative optimization.
[0094] ,
[0095] where represents the action of vehicle , represents the action of roadside facility , represents the conflict metric between the two actions, represents the interaction probability between vehicle and roadside facility .
[0096] Referring to Figure 4 , the multi-agent cooperative optimization unit 123 includes an agent relationship modeling subunit 1231, a message passing optimization subunit 1232, and a global consistency optimization subunit 1233.
[0097] The agent relationship modeling subunit 1231 is used to construct an agent relationship graph based on the current traffic conditions and identify the conditional dependence relationships between agents. Specifically, multiple vehicles and roadside devices are modeled as a Markov random field G=(V,E), where the nodes V represent agents (such as vehicles, traffic lights, etc.), and the edges E represent the interaction relationships between agents (such as the following relationship between vehicles, the control relationship between vehicles and traffic lights, etc.). Through the conditional random field theory, the dependence relationships between multiple agents can be captured and formally expressed as:
[0098] ,
[0099] Among them, represents the state set of the agent, represents the observation variable, is the normalization factor, is the potential function defined on the clique . In this system, the position, speed, etc. of the vehicle can be used as state variables, and the environmental perception information can be used as observation variables. The message passing optimization subunit 1232 is used to calculate the information entropy of each communication channel in the agent relationship graph, generate the optimal message passing strategy, and allocate communication resources. Based on the information entropy theory, the information value of each communication channel can be evaluated:
[0100] ,
[0101] Among them, represents the conditional entropy of the state of agent under the condition of knowing the state of agent . The lower the value, the higher the communication value. Based on this, a communication strategy can be designed to preferentially allocate resources to communication channels with high information value.
[0102] For example, in practical applications, when the distance between two vehicles is relatively close and the relative speed is relatively large, the communication value between them is relatively high and should be guaranteed preferentially; while when the distance between two vehicles is relatively far or they are in different road sections, the communication frequency can be reduced to save resources. Through experimental verification, adopting this communication strategy based on information entropy can reduce the communication overhead by more than 40% while maintaining 90% of the communication effect.
[0103] The global consistency optimization subunit 1233 is used to decompose the global optimization goal into local sub-goals, coordinate local decisions through variational inference methods, and verify the global consistency of the final decision. Specifically, through the variational Bayesian method, the complex posterior distribution can be approximated as a simpler distribution :
[0104] ,
[0105] The goal is to minimize the KL divergence between the two distributions:
[0106] ,
[0107] By iteratively optimizing the local distribution of each agent, a globally consistent decision can be finally achieved.
[0108] In one embodiment of the present invention, the global optimization objective can be set to minimize the overall system delay time, maximize the traffic flow, or minimize the energy consumption, etc. This objective can be decomposed into local objectives of each agent, and the variational inference method is used to ensure the consistency between the local decisions and the global objective. For example, for the objective of minimizing the overall system delay time, it can be decomposed into local objectives such as minimizing the passing delay at each intersection and optimizing the path selection of each vehicle. Through the message passing mechanism, these local decisions can be coordinated to jointly achieve the global objective.
[0109] In one embodiment of the present invention, the perception layer 101 further includes a data preprocessing module, a state caching module, and an attention allocation module.
[0110] The data preprocessing module is used to filter, denoise, and standardize the environmental state information and vehicle information. In a specific implementation, a Kalman filter can be used to filter the vehicle position and speed data, and a mean filter can be used to smooth the traffic flow data, and then maximum-minimum normalization or Z-score normalization is performed to enable different types and scales of data to be processed in the same framework.
[0111] The state caching module is used to store the historical environmental state information and vehicle information to support time series analysis. By maintaining a state sequence within a time window, the system can analyze the change trend of traffic parameters, predict future traffic conditions, and make scheduling decisions in advance. For example, by analyzing the traffic flow change in the past 30 minutes, the traffic flow trend in the next 15 minutes can be predicted, and the signal timing can be adjusted accordingly.
[0112] The attention allocation module is used to adjust the allocation strategy of perception resources according to the feedback of the decision layer 102. In the case of limited resources, the perception requirements in different regions and at different times are different. The attention allocation module can dynamically adjust the allocation of perception resources according to the current traffic conditions and decision-making requirements, such as increasing the sampling frequency in areas with large traffic flow and improving the data accuracy at key intersections.
[0113] In another embodiment of the present invention, the execution layer 103 further includes an instruction parsing module, an execution monitoring module, an execution history recording module, and a degradation processing module.
[0114] The instruction parsing module is used to convert the collaborative scheduling scheme into specific control instructions. For example, converting the high-level instruction of reducing the traffic flow into specific control parameters of adjusting the signal timing to 30 seconds for red light and 45 seconds for green light.
[0115] The execution monitoring module is used to monitor the execution status of control instructions and identify abnormal states. When the system detects that the execution deviation exceeds the threshold, it can trigger an emergency handling mechanism. For example, when the actual deceleration amplitude of the vehicle is less than the requirement of the instruction, the system can issue a warning and adjust subsequent instructions.
[0116] The execution history record module is used to maintain execution history data for reference by the decision-making layer 102. By analyzing historical execution data, the system can learn the actual effects of instruction execution and optimize the decision-making model. For example, by analyzing the actual effects of different signal timing plans, it can learn the relationship between traffic flow and signal timing.
[0117] The degradation handling module is used to execute preset degradation strategies in case of communication interruption or equipment failure. The system designs multi-level degradation strategies to ensure that basic functions can still be maintained under various abnormal conditions. For example, when vehicle-road communication is interrupted, the vehicle can switch to the local decision-making mode; when the central server fails, the roadside controller can adopt a preset fixed signal timing plan.
[0118] In another embodiment of the present invention, a timing synchronization unit, a conflict detection unit, and a collaborative effect evaluation unit are provided between the vehicle execution module 131 and the roadside control module 132.
[0119] The timing synchronization unit is used to ensure the time synchronization of vehicle control instructions and roadside control instructions. In a distributed system, time synchronization of different nodes is a key issue. The timing synchronization unit uses the Network Time Protocol (NTP) to achieve clock synchronization of each node in the system, ensuring that control instructions are executed in the correct timing sequence.
[0120] The conflict detection unit is used to detect potential conflicts in the execution process in real time and trigger emergency handling. For example, when the system detects that multiple vehicles may enter the same road section simultaneously, causing congestion, it can adjust the scheduling plan in advance to avoid conflicts.
[0121] The collaborative effect evaluation unit is used to quantitatively evaluate the execution effect of vehicle-road collaborative control and generate an effect evaluation report. The system designs multi-dimensional evaluation indicators, including average delay time, energy consumption, system throughput, etc. Through these indicators, it comprehensively evaluates the scheduling effect and provides a basis for system optimization.
[0122] In an embodiment of the present invention, information exchange is realized between the VDDPG network and the RDDPG network through a shared hidden layer.
[0123] Specifically, the shared hidden layer receives the common features of the manifold space representation and outputs an intermediate feature representation. The VDDPG network and the RDDPG network respectively receive this intermediate feature representation, combine their respective specific input features, and generate corresponding policy outputs. This design enables the two networks to share a basic understanding of the environment while maintaining their respective specializations.
[0124] In addition, the VDDPG network and the RDDPG network achieve coordinated parameter updates through a gradient locking mechanism to ensure consistent policies. The core idea of the gradient locking mechanism is to consider the mutual influence of the two networks during the gradient update process:
[0125] ,
[0126] ,
[0127] where is a coefficient between 0 and 1, used to control the intensity of collaborative optimization. In practical applications, can be dynamically adjusted according to the stability of the system. For example, a smaller value (such as 0.1) is used at the beginning of training to ensure convergence, and it is gradually increased to 0.5 or higher as training progresses to enhance the collaborative effect.
[0128] In an embodiment of the present invention, the system further includes an experience replay module, a target network update module, an adaptive learning rate adjustment module, and a model evaluation and deployment module.
[0129] The experience replay module is used to store the historical data of the system's interaction with the environment and support offline batch learning. Specifically, the current state, the selected action, the next state, and the obtained reward of the agent's interaction with the environment are stored as a quadruple (s, a, s', r) in the experience pool. During training, a batch of data is randomly sampled from the experience pool for learning. This method breaks the temporal correlation between samples and improves the stability and efficiency of learning.
[0130] The target network update module is used to periodically copy the parameters from the main network to the target network to ensure learning stability. In deep reinforcement learning, using a single network for both value estimation and target calculation may lead to instability. Therefore, a slowly updated target network is usually maintained. The parameter update of the target network adopts a soft update strategy:
[0131] ,
[0132] where is the soft update coefficient, usually taking a small value such as 0.01 to ensure the stability of the target network.
[0133] The adaptive learning rate adjustment module is used to dynamically adjust the network learning rate according to the training progress. At the beginning of training, a relatively large learning rate (such as 0.001) can be used to quickly explore; as training progresses, the learning rate is gradually reduced (such as to 0.0001) to achieve fine-tuning. In addition, the learning rate can also be adaptively adjusted according to the change trend of the loss function. For example, when the loss has not decreased for several consecutive rounds, the learning rate is reduced.
[0134] The model evaluation and deployment module is used to evaluate the model performance and deploy the trained model to the production environment. The system designs a complete model evaluation process, including offline evaluation and online A / B testing. Before deployment, the model will be subjected to stress testing and security evaluation to ensure that it can work properly under various conditions.
[0135] The present invention also provides a vehicle-road collaborative dynamic scheduling method for multi-agent reinforcement learning, which is applied to the above-mentioned vehicle-road collaborative dynamic scheduling system for multi-agent reinforcement learning, and includes the following steps:
[0136] 1. Initialize the system, including initializing the parameters of the VDDPG network and the RDDPG network, establishing a communication connection, loading the road network topology structure and historical traffic data. In this step, the network parameters can be randomly initialized or a pre-trained model can be used, and the communication connection uses a standard protocol such as MQTT to ensure stability. The road network data and historical traffic data are used for the preliminary configuration of the system and model pre-training.
[0137] 2. Collect state information, including collecting environmental state information through the roadside perception module and collecting vehicle information through the vehicle perception module. The environmental state information includes traffic flow, signal light status, road network structure, etc.; the vehicle information includes position, speed, direction, destination, etc. The system obtains these data in real time through the sensor network, and the sampling frequency is dynamically adjusted according to the scene complexity, generally 5 - 10Hz.
[0138] 3. Perform topological state characterization, including mapping the state information to a multi-dimensional topological space, extracting topological features, and generating a manifold space representation through non-linear dimensionality reduction. This step uses the topological dynamic state characterization mechanism described in detail above to compress the high-dimensional and complex traffic environment state into a low-dimensional representation, retaining key topological features.
[0139] 4. Generate a dual-network decision, including the VDDPG Actor network generating a vehicle scheduling strategy, the RDDPG Actor network generating a roadside scheduling strategy, and performing policy matrix decomposition to generate a low-rank representation. This step uses the dual-network collaborative decision mechanism described in detail above, processes the vehicle policy and the roadside policy through two dedicated networks respectively, and then reduces the complexity through matrix decomposition.
[0140] 5. Perform collaborative optimization, including constructing an agent relationship graph, optimizing the message passing strategy, coordinating local decisions through variational inference methods, and verifying global consistency. This step adopts the multi-agent collaborative decision-making mechanism described in detail above to ensure that the local decisions of each agent can be coordinated and jointly achieve the global optimization goal.
[0141] 6. Execute the scheduling plan, including converting the optimized decision into specific execution instructions and having vehicles and roadside facilities execute the control instructions. Vehicle control instructions include speed adjustment, path planning, etc.; roadside control instructions include signal timing adjustment, lane allocation, etc. The system ensures that these instructions are executed in the correct timing sequence and monitors the execution status in real time.
[0142] 7. Collect execution feedback, including monitoring the execution results, updating the environmental state, updating network parameters and the experience pool based on the execution results. The system collects the environmental changes and vehicle states after execution through the sensor network, calculates the difference between the actual effect and the expected effect, and generates a reward signal for updating network parameters.
[0143] 8. Iterative optimization, repeat the above steps to continuously optimize the scheduling effect. The system runs continuously, keeps learning and adapting to the changing traffic environment, and gradually improves the scheduling efficiency and effect.
[0144] The vehicle-road collaborative dynamic scheduling system based on multi-agent reinforcement learning provided by the present invention shows good performance in practical applications. After being tested and verified in multiple urban traffic scenarios, the system can significantly improve traffic efficiency, reduce congestion, and lower energy consumption.
[0145] Specifically, compared with the traditional fixed-time signal control, this system can reduce the average vehicle delay time by more than 30%; compared with the simple adaptive signal control, it can reduce the delay time by 15%; in complex traffic scenarios, the system's computational resource requirements are reduced by 50% compared with the centralized decision-making architecture, and the communication bandwidth requirements are reduced by 40%, greatly improving the scalability of the system.
[0146] In addition, the system shows good adaptability and robustness, and can cope with abnormal situations such as sudden changes in traffic flow, equipment failures, and communication interruptions to ensure the stable operation of the traffic system.
[0147] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various obvious changes and deformations can be made to the present invention without departing from the scope of the present invention, and these changes and deformations all belong to the protection scope of the present invention.
Claims
1. A vehicle-road collaborative dynamic scheduling system for multi-agent reinforcement learning, characterized in that Including: The perception layer, including a roadside perception module and a vehicle perception module, is used to collect environmental state information and vehicle information; The decision-making layer, communicatively connected to the perception layer, includes: The topological dynamic state characterization unit is used to receive the environmental state information and the vehicle information, map the environmental state information and the vehicle information to a multi-dimensional topological space to form a topological state representation, and generate a manifold space representation through non-linear dimensionality reduction; The dual-network collaborative decision-making unit is used to receive the manifold space representation, including a vehicle dynamic distribution network VDDPG and a roadside dynamic scheduling network RDDPG, where the VDDPG is used to generate a vehicle scheduling strategy, the RDDPG is used to generate a roadside scheduling strategy, and the vehicle scheduling strategy and the roadside scheduling strategy are converted into a low-rank representation form through policy matrix factorization; The multi-agent collaborative optimization unit is used to construct an agent relationship graph, optimize the communication strategy in the agent relationship graph based on information entropy, and coordinate the vehicle scheduling strategy and the roadside scheduling strategy through a variational inference method to generate an optimized collaborative scheduling plan; The execution layer, communicatively connected to the decision-making layer, includes a vehicle execution module and a roadside control module, which are used to receive the collaborative scheduling plan, execute the corresponding scheduling instructions, and feedback the execution result to the perception layer to form a closed-loop control.
2. The vehicle-road collaborative dynamic scheduling system for multi-agent reinforcement learning according to claim 1, wherein, The topological dynamic state characterization unit includes: The state acquisition sub-unit is used to receive the environmental state information and the vehicle information, and preprocess the environmental state information and the vehicle information to generate standardized data; The topological characterization sub-unit is used to map the standardized data to a multi-dimensional topological space, extract topological features, and establish a local coordinate system; The dimension conversion sub-unit is used to perform non-linear dimensionality reduction on the data in the multi-dimensional topological space to generate a manifold space representation that retains key topological relationships.
3. The vehicle-road collaborative dynamic scheduling system for multi-agent reinforcement learning according to claim 1, characterized in that, The dual-network collaborative decision-making unit includes: The VDDPG Actor network is used to receive the manifold space representation related to the vehicle and generate a vehicle scheduling strategy; The VDDPG Critic network is used to evaluate the value of the vehicle scheduling strategy; The RDDPG Actor network is used to receive the manifold space representation related to the roadside and generate a roadside scheduling strategy; The RDDPG Critic network is used to evaluate the value of the roadside scheduling strategy; The policy characterization sub-unit is used to convert the vehicle scheduling strategy and the roadside scheduling strategy into a policy matrix and perform matrix factorization to generate a low-rank representation; The collaborative optimization sub-unit is used to identify potential conflicts between the vehicle scheduling strategy and the roadside scheduling strategy and adjust the policy parameters through the conjugate gradient method.
4. The vehicle-road collaborative dynamic scheduling system for multi-agent reinforcement learning according to claim 1, characterized in that, The multi-agent collaborative optimization unit includes: The agent relationship modeling sub-unit is used to construct an agent relationship graph based on the current traffic conditions and identify the conditional dependence relationships between agents; The message passing optimization sub-unit is used to calculate the information entropy of each communication channel in the agent relationship graph, generate an optimal message passing strategy, and allocate communication resources; The global consistency optimization subunit is used to decompose the global optimization objective into local sub-goals, coordinate local decisions through variational inference methods, and verify the global consistency of the final decision.
5. The vehicle-road collaborative dynamic scheduling system based on multi-agent reinforcement learning according to claim 1, characterized in that, The perception layer further includes: The data preprocessing module is used to filter, denoise, and standardize the environmental state information and the vehicle information. The state caching module is used to store historical environmental state information and vehicle information to support time-series analysis. The attention allocation module is used to adjust the allocation strategy of perception resources according to the feedback from the decision-making layer.
6. The vehicle-road collaborative dynamic scheduling system for multi-agent reinforcement learning according to claim 1, wherein, The execution layer further includes: The instruction parsing module is used to convert the collaborative scheduling scheme into specific control instructions. The execution monitoring module is used to monitor the execution of control instructions and identify abnormal states. The execution history record module is used to maintain execution history data for the reference of the decision-making layer. The degradation processing module is used to execute a preset degradation strategy in case of communication interruption or equipment failure.
7. The vehicle-road collaborative dynamic scheduling system for multi-agent reinforcement learning according to claim 1, characterized in that, There is a: The timing synchronization unit is used to ensure the time synchronization of vehicle control instructions and roadside control instructions. The conflict detection unit is used to detect potential conflicts in the execution process in real time and trigger emergency processing. The collaborative effect evaluation unit is used to quantitatively evaluate the execution effect of vehicle-road collaborative control and generate an effect evaluation report.
8. The vehicle-road collaborative dynamic scheduling system for multi-agent reinforcement learning according to claim 1, characterized in that, Information exchange is realized between the VDDPG network and the RDDPG network through a shared hidden layer, where: The shared hidden layer receives the common features of the manifold space representation and outputs an intermediate feature representation. The VDDPG network and the RDDPG network respectively receive the intermediate feature representation, combine their respective specific input features, and generate corresponding policy outputs. The VDDPG network and the RDDPG network realize collaborative parameter update through a gradient locking mechanism to ensure consistent policies.
9. The vehicle-road collaborative dynamic scheduling system for multi-agent reinforcement learning according to claim 1, characterized in that, The system further includes: The experience replay module is used to store the historical data of the interaction between the system and the environment to support offline batch learning. The target network update module is used to periodically copy parameters from the main network to the target network to ensure learning stability. The adaptive learning rate adjustment module is used to dynamically adjust the network learning rate according to the training progress. The model evaluation and deployment module is used to evaluate the model performance and deploy the trained model to the production environment.
10. The vehicle-road collaborative dynamic scheduling method based on multi-agent reinforcement learning is applied to the vehicle-road collaborative dynamic scheduling system based on multi-agent reinforcement learning according to any one of claims 1-9, and is characterized in that, It includes the following steps: Initialize the system, including initializing the parameters of the VDDPG network and the RDDPG network, establishing a communication connection, loading the road network topology structure and historical traffic data. Collect state information, including collecting environmental state information through the roadside perception module and collecting vehicle information through the vehicle perception module. Execute topological state characterization, including mapping the state information to a multi-dimensional topological space, extracting topological features, and generating a manifold space representation through non-linear dimensionality reduction. Generate dual-network decisions, including the VDDPG Actor network generating a vehicle scheduling strategy, the RDDPG Actor network generating a roadside scheduling strategy, and performing policy matrix decomposition to generate a low-rank representation. Execute collaborative optimization, including constructing an agent relationship graph, optimizing the message passing strategy, coordinating local decisions through variational inference methods, and verifying global consistency. Execute the scheduling plan, including converting the optimized decision into specific execution instructions and having vehicles and roadside facilities execute the control instructions; Collect execution feedback, including monitoring execution results, updating the environmental state, updating network parameters and the experience pool based on the execution results; Iterative optimization, repeat the above steps to continuously optimize the scheduling effect.
Citation Information
Patent Citations
Manifold-learning-based traffic jam event cooperative detecting method
CN102169631A
Multi-agent cooperative communication strategy training system and method based on teammate perception
CN114757092A