Multi-agent reinforcement learning-based vehicle-road collaborative dynamic scheduling system and method thereof

Through the vehicle-road collaborative dynamic scheduling system with multi-agent reinforcement learning, the topological dynamic state representation and dual-network collaborative decision-making mechanism are used to solve the problem of vehicle-road collaborative decision-making in complex traffic environments, and efficient traffic management and optimization scheduling effects are achieved.

CN120048123AActive Publication Date: 2025-05-27CHONGQING DEXIN ROBOT TESTING CENT CO LTD

Patent Information

Application Number
CN202510521087.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently handle vehicle-road collaborative decision-making in complex traffic environments, especially in multi-agent collaborative decision-making and high-dimensional state space processing.

Method used

A vehicle-road collaborative dynamic scheduling system using multi-agent reinforcement learning, including the perception layer, the decision-making layer and the execution layer. The perception layer collects information through roadside and vehicle perception modules, the decision-making layer uses topological dynamic state representation, dual-network collaborative decision-making and multi-agent collaborative optimization mechanisms to make decisions, and the execution layer executes scheduling instructions and feedbacks the results.

Benefits of technology

It realizes efficient state representation and collaborative decision-making for complex traffic environments, improves the system's adaptability to traffic scenarios, reduces computational complexity and communication overhead, and ensures the consistency between local decisions and global optimization goals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048123A_ABST
    Figure CN120048123A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence and intelligent traffic, in particular to a multi-agent reinforcement learning-based vehicle-road collaborative dynamic scheduling system and method, which comprises a sensing layer, a decision-making layer and an execution layer. The sensing layer is provided with a roadside and vehicle sensing module and collects environment and vehicle information, the decision-making layer receives the information, the information is mapped to a multi-dimensional topological space through the topological dynamic state representation unit and is subjected to nonlinear dimensionality reduction to form manifold space representation, and the dual-network collaborative decision-making unit comprises a vehicle dynamic distribution network VDDPG and a roadside dynamic scheduling network RDDPG. Respectively generating a vehicle scheduling strategy and a roadside scheduling strategy, and decomposing a strategy matrix into low-rank representation; the multi-agent collaborative optimization unit constructs an agent relation graph, optimizes a communication strategy based on information entropy, deduces a coordinated scheduling strategy by using variational, and generates an optimized coordinated scheduling scheme; the execution layer receives the scheme, the vehicle execution module and the roadside control module execute instructions and feed back results to the sensing layer, and closed-loop control is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and intelligent transportation, and particularly to a vehicle-road collaborative dynamic scheduling system and method based on multi-agent reinforcement learning technology for realizing the collaborative optimization management of intelligent vehicles and roadside facilities. Background Art

[0002] With the acceleration of the urbanization process, the problem of urban traffic congestion has become increasingly severe, seriously affecting the travel efficiency and quality of life of residents. Traditional traffic management methods mainly rely on fixed signal light cycle control and manual intervention, which cannot adapt to the dynamically changing traffic flow, resulting in low utilization rate of traffic resources. In recent years, with the rapid development of vehicle networking technology and artificial intelligence, vehicle-road collaboration, as the core technology of the next-generation intelligent transportation system, has gradually emerged.

[0003] Currently, vehicle-road collaboration technologies are mainly divided into three categories: one is the method based on centralized control, which uniformly schedules vehicles and roadside facilities through a central server; the second is the method based on distributed control, which realizes collaboration through the local decisions of vehicles and roadside facilities; the third is the method based on reinforcement learning, which learns the optimal control strategy through the interaction between agents and the environment. However, the existing technologies have the following deficiencies: the centralized control method has a high computational complexity and is difficult to meet the real-time requirements; the distributed control method cannot guarantee the global optimum; and the existing reinforcement learning methods are difficult to handle the high-dimensional state space and multi-agent collaborative decision-making problems in the vehicle-road environment.

[0004] To solve the above problems, there is an urgent need for a multi-agent reinforcement learning system that can efficiently handle vehicle-road collaborative decision-making in complex traffic environments. Summary of the Invention

[0005] The purpose of the present invention is to provide a vehicle-road collaborative dynamic scheduling system and method based on multi-agent reinforcement learning, aiming to solve the problems in the prior art such as the difficulty in efficiently handling the state representation of complex traffic environments, vehicle-road collaborative decision-making, and multi-agent information interaction.

[0006] The present invention discloses a vehicle-road collaborative dynamic scheduling system based on multi-agent reinforcement learning, including: A perception layer, including a roadside perception module and a vehicle perception module, for collecting environmental state information and vehicle information; A decision layer, communicatively connected to the perception layer, including: A topological dynamic state representation unit, for receiving the environmental state information and the vehicle information, mapping the environmental state information and the vehicle information into a multi-dimensional topological space to form a topological state representation, and generating a manifold space representation through non-linear dimensionality reduction; The dual-network collaborative decision-making unit is used to receive the manifold space representation, including the Vehicle Dynamic Distribution Policy Gradient (VDDPG) and the Roadside Dynamic Scheduling Policy Gradient (RDDPG). The VDDPG is used to generate a vehicle scheduling strategy, and the RDDPG is used to generate a roadside scheduling strategy, and the vehicle scheduling strategy and the roadside scheduling strategy are converted into a low-rank representation form through policy matrix decomposition; The multi-agent collaborative optimization unit is used to construct an agent relationship graph, optimize the communication strategy in the agent relationship graph based on information entropy, and coordinate the vehicle scheduling strategy and the roadside scheduling strategy through variational inference to generate an optimized collaborative scheduling plan; The execution layer is communicatively connected to the decision-making layer and includes a vehicle execution module and a roadside control module, which are used to receive the collaborative scheduling plan, execute corresponding scheduling instructions, and feedback the execution results to the perception layer to form a closed-loop control.

[0007] Preferably, the topological dynamic state characterization unit includes: The state acquisition sub-unit is used to receive the environmental state information and the vehicle information, and preprocess the environmental state information and the vehicle information to generate standardized data; The topological characterization sub-unit is used to map the standardized data to a multi-dimensional topological space, extract topological features, and establish a local coordinate system; The dimension conversion sub-unit is used to perform non-linear dimensionality reduction on the data in the multi-dimensional topological space to generate a manifold space representation that retains key topological relationships.

[0008] Preferably, the dual-network collaborative decision-making unit includes: The VDDPG Actor network is used to receive the manifold space representation related to the vehicle and generate a vehicle scheduling strategy; The VDDPG Critic network is used to evaluate the value of the vehicle scheduling strategy; The RDDPG Actor network is used to receive the manifold space representation related to the roadside and generate a roadside scheduling strategy; The RDDPG Critic network is used to evaluate the value of the roadside scheduling strategy; The policy characterization sub-unit is used to convert the vehicle scheduling strategy and the roadside scheduling strategy into a policy matrix and perform matrix decomposition to generate a low-rank representation; The collaborative optimization sub-unit is used to identify potential conflicts between the vehicle scheduling strategy and the roadside scheduling strategy and adjust the policy parameters through the conjugate gradient method.

[0009] Preferably, the multi-agent collaborative optimization unit includes: An agent relationship modeling subunit, configured to construct an agent relationship graph based on the current traffic conditions and identify conditional dependence relationships between agents; A message passing optimization subunit, configured to calculate the information entropy of each communication channel in the agent relationship graph, generate an optimal message passing strategy, and allocate communication resources; A global consistency optimization subunit, configured to decompose the global optimization objective into local sub-goals, coordinate local decisions through variational inference methods, and verify the global consistency of the final decision.

[0010] Preferably, the perception layer further includes: A data preprocessing module, configured to perform filtering, denoising, and normalization processing on the environmental state information and the vehicle information; A state caching module, configured to store historical environmental state information and vehicle information to support time series analysis; An attention allocation module, configured to adjust the allocation strategy of perception resources according to the feedback of the decision-making layer.

[0011] Preferably, the execution layer further includes: An instruction parsing module, configured to convert the collaborative scheduling scheme into specific control instructions; An execution monitoring module, configured to monitor the execution of control instructions and identify abnormal states; An execution history record module, configured to maintain execution history data for reference by the decision-making layer; A degradation processing module, configured to execute a preset degradation strategy in case of communication interruption or equipment failure.

[0012] Preferably, between the vehicle execution module and the roadside control module, there is provided: A timing synchronization unit, configured to ensure the time synchronization of vehicle control instructions and roadside control instructions; A conflict detection unit, configured to detect potential conflicts during the execution process in real time and trigger emergency handling; A collaborative effect evaluation unit, configured to quantitatively evaluate the execution effect of vehicle-road collaborative control and generate an effect evaluation report.

[0013] Preferably, information exchange is realized between the VDDPG network and the RDDPG network through a shared hidden layer, where: The shared hidden layer receives the common features of the manifold space representation and outputs an intermediate feature representation; The VDDPG network and the RDDPG network respectively receive the intermediate feature representation, combine their respective specific input features, and generate corresponding policy outputs; The VDDPG network and the RDDPG network realize collaborative parameter update through a gradient locking mechanism to ensure that the policies are coordinated.

[0014] Preferably, the system further includes: An experience replay module for storing historical data of the interaction between the system and the environment, supporting offline batch learning; A target network update module for periodically copying parameters from the main network to the target network to ensure learning stability; An adaptive learning rate adjustment module for dynamically adjusting the network learning rate according to the training progress; A model evaluation and deployment module for evaluating the model performance and deploying the trained model to the production environment.

[0015] A vehicle-road collaborative dynamic scheduling method based on multi-agent reinforcement learning, applied to the vehicle-road collaborative dynamic scheduling system based on multi-agent reinforcement learning, includes the following steps: Initialize the system, including initializing the parameters of the VDDPG network and the RDDPG network, establishing a communication connection, loading the road network topology structure and historical traffic data; Collect state information, including collecting environmental state information through the roadside perception module and collecting vehicle information through the vehicle perception module; Execute topological state representation, including mapping the state information to a multi-dimensional topological space, extracting topological features, and generating a manifold space representation through non-linear dimensionality reduction; Generate dual-network decisions, including the VDDPG Actor network generating a vehicle scheduling strategy, the RDDPG Actor network generating a roadside scheduling strategy, and performing policy matrix decomposition to generate a low-rank representation; Execute collaborative optimization, including constructing an agent relationship graph, optimizing the message passing strategy, coordinating local decisions through variational inference methods, and verifying global consistency; Execute the scheduling plan, including converting the optimized decision into specific execution instructions, and the vehicle and roadside facilities executing the control instructions; Collect execution feedback, including monitoring the execution result, updating the environmental state, and updating the network parameters and experience pool based on the execution result; Iterative optimization, repeating the above steps to continuously optimize the scheduling effect.

[0016] The present invention adopts three innovative mechanisms: topological dynamic state representation, dual-network collaborative optimization driven by matrix decomposition, and multi-agent collaborative decision-making driven by probability graphs, to construct a complete vehicle-road collaborative dynamic scheduling system based on multi-agent reinforcement learning, having the following beneficial effects: 1) Through the topological dynamic state representation mechanism, the high-dimensional and complex traffic environment state is mapped to a low-dimensional manifold space, retaining key topological features, reducing the computational complexity, and improving the system's adaptability to complex traffic scenarios.

[0017] 2) Through the dual-network collaborative optimization mechanism driven by matrix factorization, the collaborative decision-making between the vehicle network and the roadside network is realized, the conflict problem between vehicle strategies and roadside strategies is solved, and the scheduling efficiency is improved.

[0018] 3) Through the multi-agent collaborative decision-making mechanism driven by probability graphs, the information interaction strategy between agents is optimized, the communication overhead is reduced, and the consistency between local decisions and global optimization goals is ensured.

[0019] 4) The overall system adopts a hierarchical modular design, has good scalability and adaptability, and can adapt to urban traffic environments of different scales and complexities. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below.

[0021] Figure 1 It is a schematic diagram of the overall architecture of the vehicle-road collaborative dynamic scheduling system for multi-agent reinforcement learning of the present invention.

[0022] Figure 2 It is a schematic diagram of the structure of the topological dynamic state characterization unit of the present invention.

[0023] Figure 3 It is a schematic diagram of the structure of the dual-network collaborative decision-making unit of the present invention.

[0024] Figure 4 It is a schematic diagram of the structure of the multi-agent collaborative optimization unit of the present invention.

[0025] Figure 5 It is a flowchart of the vehicle-road collaborative dynamic scheduling method of the present invention.

[0026] Figure 6 It is a schematic diagram of the topological state characterization process in the embodiment of the present invention.

[0027] Figure 7 It is a schematic diagram of the dual-network collaborative optimization process in the embodiment of the present invention.

[0028] Figure 8 It is a schematic diagram of the multi-agent collaborative decision-making process in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not used to limit the present invention.

[0030] Refer to Figure 1, the vehicle-road collaborative dynamic scheduling system based on multi-agent reinforcement learning provided by the present invention includes three main parts: a perception layer 101, a decision-making layer 102, and an execution layer 103.

[0031] The perception layer 101 includes a roadside perception module 111 and a vehicle perception module 112, which are used to collect environmental state information and vehicle information. The roadside perception module 111 can be various sensing devices distributed in the urban road network, such as cameras, radars, signal controllers, etc., which are used to collect environmental information such as traffic flow, signal status, and road network structure. The vehicle perception module 112 can be various sensors and communication devices installed on the vehicle, which are used to collect vehicle information such as vehicle position, speed, direction, and destination.

[0032] The decision-making layer 102 is the core part of the system, including a topological dynamic state representation unit 121, a dual-network collaborative decision-making unit 122, and a multi-agent collaborative optimization unit 123. The topological dynamic state representation unit 121 is used to receive environmental state information and vehicle information, map this information to a multi-dimensional topological space to form a topological state representation, and generate a manifold space representation through non-linear dimensionality reduction. The dual-network collaborative decision-making unit 122 includes a vehicle dynamic distribution network VDDPG and a roadside dynamic scheduling network RDDPG, which are respectively used to generate vehicle scheduling strategies and roadside scheduling strategies, and convert these strategies into a low-rank representation form through policy matrix decomposition. The multi-agent collaborative optimization unit 123 is used to construct an agent relationship graph, optimize the communication strategy, coordinate the decisions of each agent, and generate a final collaborative scheduling plan.

[0033] The execution layer 103 includes a vehicle execution module 131 and a roadside control module 132, which are used to receive the collaborative scheduling plan, execute the corresponding scheduling instructions, and feedback the execution results to the perception layer 101 to form a closed-loop control. The vehicle execution module 131 is responsible for controlling the driving behavior of the vehicle, such as adjusting the speed and changing the route. The roadside control module 132 is responsible for controlling the operating state of roadside facilities, such as adjusting signal timing and restricting access.

[0034] Referring to Figure 2 , the topological dynamic state representation unit 121 includes a state acquisition subunit 1211, a topological representation subunit 1212, and a dimension conversion subunit 1213.

[0035] The state acquisition subunit 1211 is used to receive environmental state information and vehicle information from the perception layer 101, preprocess this raw data, and generate standardized data. In an embodiment of the present invention, the preprocessing includes operations such as denoising, filtering, and standardization. For example, for vehicle speed data, mean filtering can be used to remove noise, and then maximum-minimum standardization can be performed to map the speed value to the [0,1] interval.

[0036] The topological characterization subunit 1212 is used to map the standardized data into a multi-dimensional topological space, extract topological features, and establish a local coordinate system. Specifically, the state space of the vehicle-road environment is first defined , where the vehicle state (position, speed, direction) and the roadside state (signal light state, road network topology) respectively constitute subspaces and . For each state point , the vehicle-road environment is accurately expressed through a local coordinate chart. Preferably, the topological characterization subunit 1212 also constructs a diffeomorphic mapping from the state space to the policy space , ensuring topological invariance between state changes and policy adjustments, and enabling the system to maintain stability in complex traffic environments. The diffeomorphic mapping can be defined in the following way: , where is the state point, is the weight matrix, is the non-linear activation function, is the bias vector. Through this mapping, points with similar topological structures in the state space can be mapped to nearby points in the policy space, ensuring the continuity and stability of the policy.

[0037] The dimension conversion subunit 1213 is used to perform non-linear dimensionality reduction on the data in the multi-dimensional topological space, generating a manifold space representation that retains key topological relationships. In one embodiment of the present invention, the locally linear embedding (LLE) algorithm can be used to achieve non-linear dimensionality reduction. The core idea of the LLE algorithm is to maintain the local linear relationship of data points while non-linearly reducing the dimension globally. Specifically, for each data point , first find its nearest neighbor points , then calculate the weight matrix such that can be reconstructed by a linear combination of its neighbors: , where represents the contribution weight of point to point , satisfying . Then, find the representation in the low-dimensional space that satisfies the same weight relationship In this way, the representation of the original high-dimensional data on the low-dimensional manifold can be obtained. In practical applications, the number of nearest neighbors can be dynamically adjusted according to the complexity of the traffic scenario.Value. For example, in the case of heavy traffic, a larger value (such as ) can be selected to capture more interaction relationships; in the case of light traffic, a smaller value (such as ) can be selected to reduce the computational amount.

[0038] The policy representation subunit 1225 is used to convert the vehicle scheduling policy and the roadside scheduling policy into a policy matrix and perform matrix factorization to generate a low-rank representation. Specifically, the policies of the two networks are represented as a high-order matrix , and then it is decomposed into a combination of a core tensor and factor matrices through Tucker factorization: , where is the core tensor, , , are factor matrices, and represents the tensor-matrix product along the th mode. In this way, the complexity of the policy representation can be greatly reduced. For example, for a policy tensor, if the rank is reduced to , the number of parameters can be reduced from to , greatly reducing the storage and computational requirements.

[0039] The collaborative optimization subunit 1226 is used to identify potential conflicts between the vehicle scheduling policy and the roadside scheduling policy and adjust the policy parameters through the conjugate gradient method. In a specific embodiment, the collaborative optimization can be expressed as the following optimization problem: , where and are the parameters of the VDDPG network and the RDDPG network respectively, and are the independent loss functions of the two networks respectively, is the joint loss function, which is used to measure the coordination degree of the output policies of the two networks, is the trade-off coefficient, which is used to adjust the ratio of independent optimization to collaborative optimization.

[0040] , where represents the action of vehicle , represents the action of roadside facility , represents the conflict measure between the two actions, Represents a vehicle and roadside facilities between the interaction probability.

[0041] Refer to Figure 4 , the multi-agent collaborative optimization unit 123 includes an agent relationship modeling subunit 1231, a message passing optimization subunit 1232, and a global consistency optimization subunit 1233.

[0042] The agent relationship modeling subunit 1231 is used to construct an agent relationship graph based on the current traffic conditions and identify the conditional dependence relationships between agents. Specifically, multiple vehicles and roadside devices are modeled as a Markov random field G=(V,E), where the nodes V represent agents (such as vehicles, traffic lights, etc.), and the edges E represent the interaction relationships between agents (such as the following relationship between vehicles, the control relationship between vehicles and traffic lights, etc.). Through the conditional random field theory, the dependence relationships between multiple agents can be captured and formally represented as: , where, represents the state set of the agent, represents the observed variable, is the normalization factor, is the potential function defined on the clique . In this system, the position, speed, etc. of the vehicle can be used as state variables, and the environmental perception information can be used as the observed variable. The message passing optimization subunit 1232 is used to calculate the information entropy of each communication channel in the agent relationship graph, generate an optimal message passing strategy, and allocate communication resources. Based on the information entropy theory, the information value of each communication channel can be evaluated: , where, represents the conditional entropy of the state of agent given the state of agent , and the lower the value, the higher the communication value. Based on this, a communication strategy can be designed to preferentially allocate resources to the communication channels with high information value.

[0043] For example, in practical applications, when the distance between two vehicles is relatively close and the relative speed is relatively large, the communication value between them is relatively high and should be guaranteed preferentially; while when the distance between two vehicles is relatively far or they are in different sections of the road, the communication frequency can be reduced to save resources. Through experimental verification, adopting this communication strategy based on information entropy can reduce the communication overhead by more than 40% while maintaining 90% of the communication effect.

[0044] The global consistency optimization subunit 1233 is used to decompose the global optimization objective into local sub-objectives, coordinate local decisions through variational inference methods, and verify the global consistency of the final decision. Specifically, through the variational Bayesian method, a complex posterior distribution can be approximated as a simpler distribution : , The goal is to minimize the KL divergence between the two distributions: , By iteratively optimizing the local distributions of each agent , a globally consistent decision can be finally achieved.

[0045] In an embodiment of the present invention, the global optimization objective can be set to minimize the overall system delay time, maximize the traffic flow, or minimize the energy consumption, etc. This objective can be decomposed into local objectives of each agent, and the variational inference method is used to ensure the consistency between the local decision and the global objective. For example, for the objective of minimizing the overall system delay time, it can be decomposed into local objectives such as minimizing the passing delay at each intersection and optimizing the path selection of each vehicle. Through the message passing mechanism, these local decisions can be coordinated to jointly achieve the global objective.

[0046] In an embodiment of the present invention, the perception layer 101 further includes a data preprocessing module, a state caching module, and an attention allocation module.

[0047] The data preprocessing module is used to filter, denoise, and standardize the environmental state information and vehicle information. In specific implementation, a Kalman filter can be used to filter the vehicle position and speed data, and mean filtering can be used to smooth the traffic flow data, and then maximum-minimum normalization or Z-score normalization is performed to enable different types and scales of data to be processed in the same framework.

[0048] The state caching module is used to store the historical environmental state information and vehicle information to support time-series analysis. By maintaining a state sequence within a time window, the system can analyze the change trend of traffic parameters, predict future traffic conditions, and make scheduling decisions in advance. For example, by analyzing the traffic flow changes in the past 30 minutes, the flow trend in the next 15 minutes can be predicted, and the signal timing can be adjusted accordingly.

[0049] The attention allocation module is used to adjust the allocation strategy of sensing resources according to the feedback from the decision-making layer 102. In the case of limited resources, the sensing requirements in different regions and at different times are different. The attention allocation module can dynamically adjust the allocation of sensing resources according to the current traffic conditions and decision-making requirements. For example, it can increase the sampling frequency in areas with heavy traffic and improve the data accuracy at key intersections.

[0050] In another embodiment of the present invention, the execution layer 103 further includes an instruction parsing module, an execution monitoring module, an execution history record module, and a degradation processing module.

[0051] The instruction parsing module is used to convert the collaborative scheduling scheme into specific control instructions. For example, it converts the high-level instruction of reducing vehicle flow into specific control parameters such as adjusting the signal timing to 30 seconds for red light and 45 seconds for green light.

[0052] The execution monitoring module is used to monitor the execution of control instructions and identify abnormal states. When the system detects that the execution deviation exceeds the threshold, it can trigger an emergency handling mechanism. For example, when the actual deceleration amplitude of the vehicle is less than the requirement of the instruction, the system can issue a warning and adjust the subsequent instructions.

[0053] The execution history record module is used to maintain the execution history data for the reference of the decision-making layer 102. By analyzing the historical execution data, the system can learn the actual effects of instruction execution and optimize the decision-making model. For example, by analyzing the actual effects of different timing schemes, it can learn the relationship between traffic flow and signal timing.

[0054] The degradation processing module is used to execute a preset degradation strategy in case of communication interruption or device failure. The system designs a multi-level degradation strategy to ensure that basic functions can still be maintained under various abnormal conditions. For example, when the vehicle-road communication is interrupted, the vehicle can switch to the local decision-making mode; when the central server fails, the roadside controller can adopt a preset fixed timing scheme.

[0055] In yet another embodiment of the present invention, a timing synchronization unit, a conflict detection unit, and a collaborative effect evaluation unit are provided between the vehicle execution module 131 and the roadside control module 132.

[0056] The timing synchronization unit is used to ensure the time synchronization of vehicle control instructions and roadside control instructions. In a distributed system, the time synchronization of different nodes is a key issue. The timing synchronization unit uses the Network Time Protocol (NTP) to achieve clock synchronization of each node in the system, ensuring that control instructions are executed in the correct timing.

[0057] The conflict detection unit is used to detect potential conflicts during the execution process in real time and trigger emergency handling. For example, when the system detects that multiple vehicles may enter the same road section simultaneously, causing congestion, it can adjust the scheduling plan in advance to avoid conflicts.

[0058] The collaborative effect evaluation unit is used to quantitatively evaluate the execution effect of vehicle-road collaborative control and generate an effect evaluation report. The system designs multi-dimensional evaluation indicators, including average delay time, energy consumption, system throughput, etc. Through these indicators, the scheduling effect is comprehensively evaluated to provide a basis for system optimization.

[0059] In an embodiment of the present invention, information exchange is achieved between the VDDPG network and the RDDPG network through a shared hidden layer.

[0060] Specifically, the shared hidden layer receives the common features represented in the manifold space and outputs an intermediate feature representation. The VDDPG network and the RDDPG network respectively receive this intermediate feature representation, and combine their respective specific input features to generate corresponding policy outputs. This design enables the two networks to share a basic understanding of the environment while maintaining their respective professionalism.

[0061] In addition, the VDDPG network and the RDDPG network achieve collaborative parameter update through a gradient locking mechanism to ensure that the policies are coordinated. The core idea of the gradient locking mechanism is to consider the mutual influence of the two networks during the gradient update process: , , where is a coefficient between 0 and 1, which is used to control the intensity of collaborative optimization. In practical applications, can be dynamically adjusted according to the stability of the system. For example, a smaller value (such as 0.1) is used at the initial stage of training to ensure convergence, and it is gradually increased to 0.5 or higher as the training progresses to enhance the collaborative effect.

[0062] In an embodiment of the present invention, the system further includes an experience replay module, a target network update module, an adaptive learning rate adjustment module, and a model evaluation and deployment module.

[0063] The experience replay module is used to store the historical data of the interaction between the system and the environment and support offline batch learning. Specifically, the current state, the selected action, the next state, and the obtained reward of the interaction between the agent and the environment are stored in the experience pool as a quadruple (s, a, s', r). During the training process, batch data is randomly sampled from the experience pool for learning. This method breaks the temporal correlation between samples and improves the stability and efficiency of learning.

[0064] The target network update module is used to periodically copy parameters from the main network to the target network to ensure learning stability. In deep reinforcement learning, using a single network for both value estimation and target calculation may lead to instability. Therefore, a slowly updated target network is usually maintained. The parameter update of the target network adopts a soft update strategy: , where, is the soft update coefficient, usually taking a small value such as 0.01 to ensure the stability of the target network.

[0065] The adaptive learning rate adjustment module is used to dynamically adjust the network learning rate according to the training progress. At the beginning of training, a larger learning rate (such as 0.001) can be used for rapid exploration; as training progresses, the learning rate is gradually decreased (such as to 0.0001) for fine-tuning. In addition, the learning rate can be adaptively adjusted according to the change trend of the loss function. For example, when the loss has not decreased for multiple consecutive rounds, the learning rate is decreased.

[0066] The model evaluation and deployment module is used to evaluate the model performance and deploy the trained model to the production environment. The system designs a complete model evaluation process, including offline evaluation and online A / B testing. Before deployment, the model will be subjected to stress testing and security evaluation to ensure that it can work properly under various conditions.

[0067] The present invention also provides a vehicle-road collaborative dynamic scheduling method for multi-agent reinforcement learning, which is applied to the above-mentioned vehicle-road collaborative dynamic scheduling system for multi-agent reinforcement learning, and includes the following steps: 1. Initialize the system, including initializing the parameters of the VDDPG network and the RDDPG network, establishing a communication connection, loading the road network topology structure and historical traffic data. In this step, the network parameters can be randomly initialized or a pre-trained model can be used, and the communication connection adopts a standard protocol such as MQTT to ensure stability. The road network data and historical traffic data are used for the preliminary configuration of the system and model pre-training.

[0068] 2. Collect state information, including collecting environmental state information through the roadside perception module and collecting vehicle information through the vehicle perception module. The environmental state information includes traffic flow, signal light status, road network structure, etc.; the vehicle information includes position, speed, direction, destination, etc. The system obtains these data in real time through the sensor network, and the sampling frequency is dynamically adjusted according to the scene complexity, generally 5 - 10Hz.

[0069] 3. Perform topological state characterization, including mapping the state information to a multi-dimensional topological space, extracting topological features, and generating a manifold space representation through non-linear dimensionality reduction. This step adopts the topological dynamic state characterization mechanism described in detail above to compress the high-dimensional and complex traffic environment state into a low-dimensional representation and retain the key topological features.

[0070] 4. Generate a dual-network decision, including the VDDPG Actor network generating a vehicle scheduling strategy, the RDDPG Actor network generating a roadside scheduling strategy, and performing policy matrix decomposition to generate a low-rank representation. This step adopts the dual-network collaborative decision-making mechanism described in detail above, processes the vehicle policy and the roadside policy through two dedicated networks respectively, and then reduces the complexity through matrix decomposition.

[0071] 5. Execute collaborative optimization, including constructing an agent relationship graph, optimizing the message passing strategy, coordinating local decisions through variational inference methods, and verifying global consistency. This step adopts the multi-agent collaborative decision-making mechanism described in detail above to ensure that the local decisions of each agent can be coordinated and jointly achieve the global optimization goal.

[0072] 6. Execute the scheduling plan, including converting the optimized decision into specific execution instructions, and the vehicles and roadside facilities execute the control instructions. The vehicle control instructions include speed adjustment, path planning, etc.; the roadside control instructions include signal timing adjustment, lane allocation, etc. The system ensures that these instructions are executed in the correct time sequence and monitors the execution status in real time.

[0073] 7. Collect execution feedback, including monitoring the execution results, updating the environmental state, updating the network parameters and the experience pool based on the execution results. The system collects the environmental changes and vehicle states after execution through the sensor network, calculates the difference between the actual effect and the expected effect, and generates a reward signal for updating the network parameters.

[0074] 8. Iterative optimization, repeat the above steps, and continuously optimize the scheduling effect. The system runs continuously, learns and adapts to the changing traffic environment, and gradually improves the scheduling efficiency and effect.

[0075] The vehicle-road collaborative dynamic scheduling system based on multi-agent reinforcement learning provided by the present invention shows good performance in practical applications. After being tested and verified in multiple urban traffic scenarios, the system can significantly improve traffic efficiency, reduce congestion, and reduce energy consumption.

[0076] Specifically, compared with the traditional fixed-time signal control, this system can reduce the average vehicle delay time by more than 30%; compared with the simple adaptive signal control, it can reduce the delay time by 15%; in complex traffic scenarios, the system's computing resource requirements are reduced by 50% compared with the centralized decision-making architecture, and the communication bandwidth requirements are reduced by 40%, greatly improving the scalability of the system.

[0077] In addition, the system shows good adaptability and robustness, and can cope with abnormal situations such as sudden changes in traffic flow, equipment failures, and communication interruptions to ensure the stable operation of the traffic system.

[0078] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various obvious changes and modifications can be made to the present invention without departing from the scope of the present invention, and these changes and modifications all fall within the protection scope of the present invention.

Claims

1. A multi-agent reinforcement learning vehicle-road cooperative dynamic scheduling system, characterized by: include: The perception layer includes a roadside perception module and a vehicle perception module, which are used to collect environmental status information and vehicle information; The decision layer is connected to the perception layer in communication, and includes: A topological dynamic state representation unit, configured to receive the environmental state information and the vehicle information, map the environmental state information and the vehicle information to a multidimensional topological space to form a topological state representation, and generate a manifold space representation by nonlinear dimensionality reduction; A dual-network collaborative decision-making unit, used to receive the manifold space representation, including a vehicle dynamic allocation network VDDPG and a roadside dynamic dispatch network RDDPG, wherein the VDDPG is used to generate a vehicle dispatch strategy, and the RDDPG is used to generate a roadside dispatch strategy, and convert the vehicle dispatch strategy and the roadside dispatch strategy into a low-rank representation through strategy matrix decomposition; A multi-agent collaborative optimization unit, used to construct an agent relationship graph, optimize the communication strategy in the agent relationship graph based on information entropy, coordinate the vehicle scheduling strategy and the roadside scheduling strategy through a variational inference method, and generate an optimized collaborative scheduling solution; The execution layer is in communication with the decision layer, and includes a vehicle execution module and a roadside control module, which are used to receive the collaborative scheduling scheme, execute corresponding scheduling instructions, and feed back the execution results to the perception layer to form a closed-loop control.

2. The multi-agent reinforcement learning vehicle-road cooperative dynamic scheduling system according to claim 1 is characterized in that: The topology dynamic state characterization unit comprises: A state acquisition subunit, used for receiving the environmental state information and the vehicle information, and preprocessing the environmental state information and the vehicle information to generate standardized data; A topological characterization subunit, used to map the standardized data into a multidimensional topological space, extract topological features, and establish a local coordinate system; The dimension conversion subunit is used to perform nonlinear dimensionality reduction on the data in the multidimensional topological space to generate a manifold space representation that retains key topological relationships.

3. The vehicle-road cooperative dynamic scheduling system based on multi-agent reinforcement learning according to claim 1 is characterized in that: The dual network collaborative decision unit includes: VDDPGActor network, used to receive the manifold space representation related to vehicles and generate vehicle scheduling strategies; VDDPGCritic network, used to evaluate the value of the vehicle dispatching strategy; RDDPGActor network, used to receive the manifold space representation related to the roadside and generate the roadside dispatch strategy; RDDPGCritic network, used to evaluate the value of the roadside dispatch strategy; A strategy representation subunit, used for converting the vehicle dispatch strategy and the roadside dispatch strategy into a strategy matrix, and performing matrix decomposition to generate a low-rank representation; The collaborative optimization subunit is used to identify potential conflicts between the vehicle scheduling strategy and the roadside scheduling strategy, and adjust strategy parameters through the conjugate gradient method.

4. The multi-agent reinforcement learning vehicle-road cooperative dynamic scheduling system according to claim 1 is characterized in that: The multi-agent collaborative optimization unit comprises: The agent relationship modeling subunit is used to build an agent relationship graph based on the current traffic conditions and identify the conditional dependencies between agents; The message transmission optimization subunit is used to calculate the information entropy of each communication channel in the agent relationship graph, generate the optimal message transmission strategy, and allocate communication resources; The global consistency optimization subunit is used to decompose the global optimization objective into local sub-objectives, coordinate local decisions through variational inference methods, and verify the global consistency of the final decision.

5. The multi-agent reinforcement learning vehicle-road cooperative dynamic scheduling system according to claim 1 is characterized in that: The perception layer also includes: A data preprocessing module, used for filtering, denoising and standardizing the environmental status information and the vehicle information; The state cache module is used to store historical environment state information and vehicle information and support time series analysis; The attention allocation module is used to adjust the allocation strategy of perception resources according to the feedback from the decision layer.

6. The multi-agent reinforcement learning vehicle-road cooperative dynamic scheduling system according to claim 1 is characterized in that: The execution layer also includes: An instruction parsing module, used to convert the collaborative scheduling scheme into specific control instructions; An execution monitoring module is used to monitor the execution of control instructions and identify abnormal conditions; Execution history module, used to maintain execution history data for reference by decision-makers; The degradation processing module is used to execute the preset degradation strategy in the event of communication interruption or equipment failure.

7. The multi-agent reinforcement learning vehicle-road cooperative dynamic scheduling system according to claim 1 is characterized in that: Between the vehicle execution module and the roadside control module is provided: A timing synchronization unit, used to ensure the time synchronization between vehicle control instructions and roadside control instructions; Conflict detection unit, used to detect potential conflicts in the execution process in real time and trigger emergency processing; The collaborative effect evaluation unit is used to quantitatively evaluate the execution effect of vehicle-road collaborative control and generate an effect evaluation report.

8. The multi-agent reinforcement learning vehicle-road cooperative dynamic scheduling system according to claim 1 is characterized in that: The VDDPG network and the RDDPG network exchange information via a shared hidden layer, wherein: The shared hidden layer receives the common features represented in the manifold space and outputs an intermediate feature representation; The VDDPG network and the RDDPG network respectively receive the intermediate feature representation, and generate corresponding strategy outputs in combination with their respective specific input features; The VDDPG network and the RDDPG network implement collaborative parameter updates through a gradient locking mechanism to ensure strategy coordination and consistency.

9. The vehicle-road cooperative dynamic scheduling system based on multi-agent reinforcement learning according to claim 1 is characterized in that: The system further comprises: The experience replay module is used to store historical data of the interaction between the system and the environment and supports offline batch learning; The target network update module is used to periodically copy parameters from the main network to the target network to ensure learning stability; Adaptive learning rate adjustment module, used to dynamically adjust the network learning rate according to the progress of training; The model evaluation and deployment module is used to evaluate model performance and deploy the trained model to the production environment.

10. A vehicle-road cooperative dynamic scheduling method based on multi-agent reinforcement learning, applied to a vehicle-road cooperative dynamic scheduling system based on multi-agent reinforcement learning as claimed in any one of claims 1 to 9, characterized in that: The following steps are involved: Initialize the system, including initializing VDDPG network and RDDPG network parameters, establishing communication connections, and loading road network topology and historical traffic data; Collecting status information, including collecting environmental status information through a roadside perception module and collecting vehicle information through a vehicle perception module; Perform topological state representation, including mapping state information into a multidimensional topological space, extracting topological features, and generating a manifold space representation through nonlinear dimensionality reduction; Generate dual network decisions, including VDDPGActor network to generate vehicle dispatch strategy, RDDPGActor network to generate roadside dispatch strategy, and execute strategy matrix decomposition to generate low-rank representation; Perform collaborative optimization, including building agent relationship graphs, optimizing message passing strategies, coordinating local decisions through variational inference methods, and verifying global consistency; Execute the dispatch plan, including converting the optimized decision into specific execution instructions, and executing control instructions for vehicles and roadside facilities; Collect execution feedback, including monitoring execution results, updating environment status, and updating network parameters and experience pool based on execution results; Iterate optimization and repeat the above steps to continuously optimize the scheduling effect.

Citation Information

Patent Citations

  • Manifold-learning-based traffic jam event cooperative detecting method

    CN102169631A

  • Multi-agent cooperative communication strategy training system and method based on teammate perception

    CN114757092A

  • Traffic light control method and system based on deep learning

    CN117935562A

  • Internet of vehicles multi-parameter monitoring method and system based on 5G multi-access edge computing

    CN119233227A

  • Intelligent networked automobile collaborative decision-making method based on double interactive perception

    CN119428755A

Cited By

  • Construction method and system of multi-agent collaborative decision graph for traffic control

    CN121505868A

  • Multi-agent-based OHT dynamic scheduling and traffic control optimization method and system

    CN121921977A