Traffic control method and device in man-machine mixed driving environment and roadside equipment
By building multiple agents at the intersection and using deep reinforcement learning algorithms to determine the control parameters of the vehicle queue, the inefficiency problem of traffic management in the mixed driving environment of man-machine is solved, efficient traffic control without signal light adjustment is achieved, adapting to various vehicle queues, and improving the applicability and reliability of traffic management.
Patent Information
- Application Number
- CN202510492757.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art cannot effectively control artificially driven vehicles in a mixed-drive environment, resulting in low traffic management performance, or frequent adjustment of signal timing to adapt to environmental changes, resulting in complex management and inability to apply to signal light-free intersections.
By constructing multiple agents corresponding to the intersection control line, using deep reinforcement learning algorithms to determine the target lane, speed and acceleration of the vehicle queue, differentiated control of manual and autonomous vehicles is achieved, and a proximity strategy optimization model is built to improve control stability and flexibility.
It realizes efficient traffic management in the human-machine hybrid driving environment without adjusting the signal timing, and is suitable for signal light intersections, improving the applicability, management ability, flexibility and reliability of traffic control.
Smart Images

Figure CN120388470A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic control, and in particular, to a traffic control method, device, and roadside equipment in a human-machine co-driving environment. Background Art
[0002] With the acceleration of the urbanization process and the improvement of residents' living standards, the urban road traffic network is facing unprecedented pressures and challenges. The surging traffic flow, limited road resources, and increasingly diverse traffic patterns have made it an important issue that urgently needs to be solved globally on how to achieve the intelligentization, high efficiency, and greening of the urban traffic system. In recent years, the rapid development of intelligent connected vehicles and intelligent transportation systems has provided a new technical direction for solving the above problems. Through communications between vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I), etc., intelligent connected technology realizes the real-time sharing and collaborative control of traffic information. This technology provides an important foundation for building a more efficient and safe traffic management system.
[0003] In the prior art, there are mainly two ways to control traffic in a human-machine co-driving environment. One is to not consider manually driven vehicles and only control autonomous vehicles in a human-machine co-driving environment, which has the following technical problems: it is unable to effectively control manually driven vehicles in a human-machine co-driving environment, resulting in the technical problem of low traffic management. The second is to achieve traffic control by adjusting the signal timing at intersections. The specific means is: to count information such as the flow of autonomous vehicles and manually driven vehicles, and adjust the signal timing while keeping this information unchanged. It has the following technical problems: as the human-machine co-driving environment is constantly changing, it is necessary to continuously adjust the signal timing. The adjustment method is complex and not applicable to intersections where the signal timing cannot be adjusted, thus resulting in low traffic management performance.
[0004] Therefore, there is an urgent need to provide a traffic control method, device, and roadside equipment in a human-machine co-driving environment to achieve efficient traffic management in a human-machine co-driving environment without adjusting the signal timing. Summary of the Invention
[0005] In view of this, it is necessary to provide a traffic control method, device, and roadside equipment in a human-machine co-driving environment to solve the technical problem that the traffic management performance in a human-machine co-driving environment is low due to the existing solutions of only controlling autonomous vehicles or dynamically adjusting signal timing.
[0006] In a first aspect, to solve the above technical problem, the present invention provides a traffic control method in a human-machine co-driving environment, including: Obtain multiple control lines at the intersection and the vehicle queues to pass through the control lines, and construct multiple agents corresponding one-to-one to the multiple control lines. The vehicle queues are a mixed queue of human-driven vehicles and autonomous vehicles; Train the proximal policy optimization models of the constructed multiple agents to obtain a target policy network; Determine the control parameters of the vehicle queues based on the target policy network. The control parameters include the target lanes, target speeds, and target accelerations for the vehicle queues to pass through the control lines; Control the human-driven vehicles to drive according to the target lanes, and control the autonomous vehicles to drive according to the target lanes, the target speeds, and the target accelerations.
[0007] In one possible implementation, the training of the proximal policy optimization models of the constructed multiple agents includes: Generate initial actions based on an initial policy network and an environmental state. The initial actions include the lanes, speeds, and accelerations for the vehicle queues to pass through the control lines; Determine the reward value and the new environmental state for executing the initial actions, and respectively determine the estimated environmental state value and the estimated new environmental state value based on the environmental state and the new environmental state; Determine the advantage function value based on the reward value, the estimated environmental state value, and the estimated new environmental state value; Update the initial policy network based on the advantage function value and the proximal optimization policy to obtain an updated policy network; Judge whether the convergence condition is satisfied. When it is satisfied, the updated policy network is the target policy network.
[0008] In one possible implementation, the convergence condition is: the number of iterations reaches the maximum number of iterations or the change value of the updated policy network is less than the change threshold.
[0009] In one possible implementation, the determination of the reward value for executing the initial actions includes: Determine the delay index and the safety index for executing the initial actions; Take the weighted sum of the delay index and the safety index as the reward value.
[0010] In one possible implementation, the updating of the initial policy network based on the advantage function value and the proximal optimization policy to obtain an updated policy network includes: Determine the optimization direction of the initial policy network based on the advantage function value; Determine the optimization amplitude of the initial policy network based on the proximal optimization policy; Update the initial policy network based on the optimization direction and the optimization amplitude to obtain the updated policy network.
[0011] In a possible implementation, the advantage function value is:
[0012] In the formula, is the advantage function value; is the reward value; is the discount factor; is the estimated value of the new environmental state; is the estimated value of the environmental state.
[0013] In a possible implementation, the proximal optimization policy is:
[0014]
[0015] In the formula, is the optimization amplitude; is the expected value; is the minimum value function; is the probability ratio of the actions selected by the new and old policies; is the advantage function value at time t; is the truncation function; is the hyperparameter; are the parameters of the initial policy network.
[0016] In a possible implementation, the control line is the actual stop line of the intersection.
[0017] In a second aspect, the present invention also provides a traffic control device in a human-machine mixed driving environment, including: A multi-agent construction unit, configured to obtain a plurality of control lines at an intersection and vehicle queues to pass through the control lines, and construct a plurality of agents corresponding to the plurality of control lines one by one, where the vehicle queues are a mixed queue of human-driven vehicles and autonomous vehicles; A target policy network determination unit, configured to train the proximal policy optimization models of the constructed plurality of agents to obtain a target policy network; A control parameter determination unit, configured to determine control parameters of the vehicle queue based on the target policy network, where the control parameters include a target lane, a target speed, and a target acceleration for the vehicle queue to pass through the control line; A traffic control unit is configured to control the human-driven vehicle to travel along the target lane, and control the autonomous vehicle to travel along the target lane, at the target speed, and with the target acceleration.
[0018] In a third aspect, the present invention further provides a roadside device, including a memory and a processor, wherein, The memory is used for storing programs; The processor is coupled to the memory and is configured to execute the programs stored in the memory to implement the steps in the traffic control method in the human-machine mixed driving environment in any of the above possible implementation manners.
[0019] The beneficial effects of the present invention are as follows: The traffic control method in the human-machine mixed driving environment provided by the present invention constructs multiple agents corresponding one-to-one to multiple control lines at intersections, and uses the deep reinforcement learning algorithm to determine the control parameters of the vehicle queue in the human-machine mixed driving environment, that is, the target lane, the target speed, and the target acceleration, realizing the direct control of the vehicle queue. Compared with the existing solution of dynamically adjusting signal timing, there is no need for the traffic lights at intersections to have the ability of dynamic adjustment, and it can even be applied to intersections without traffic lights, realizing traffic control in the human-machine mixed driving environment without dynamic adjustment of signal timing, improving the applicability of the traffic control method. Moreover, it can adapt to a wider range of vehicle queues, improving the management ability of the traffic control method.
[0020] Furthermore, the present invention constructs a proximal policy optimization model. Compared with the traditional deep reinforcement learning model, by introducing the proximal policy, the learning process is made more stable, improving the stability and reliability of the control of the vehicle queue.
[0021] Even further, when controlling the vehicle queue, the present invention differentially controls the human-driven vehicle and the autonomous vehicle. On the premise of realizing the simultaneous control of the human-driven vehicle and the autonomous vehicle, it takes into account the human behavior of the human-driven vehicle, which is more in line with the actual situation, that is: further improving the flexibility, rationality, and reliability of the traffic control method. Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0023] Figure 1 It is a schematic flowchart of an embodiment of the traffic control method in the human-machine mixed driving environment provided by the present invention; Figure 2 A schematic flowchart of an embodiment of step S102 provided by the present invention; Figure 3 A schematic flowchart of an embodiment of step S204 provided by the present invention; Figure 4 A schematic structural diagram of an embodiment of a traffic control device in a human-machine co-driving environment provided by the present invention; Figure 5 A schematic structural diagram of an embodiment of a roadside device provided by the present invention. Detailed implementation manners
[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.
[0025] It should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present invention illustrate the operations implemented according to some embodiments of the present invention. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present invention. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor systems and / or microcontroller systems.
[0026] Referring to "embodiment" in this article means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0027] The present invention provides a traffic control method, device and roadside device in a human-machine co-driving environment, which will be described separately below.
[0028] Figure 1 A schematic flowchart of an embodiment of the traffic control method in a human-machine co-driving environment provided by the present invention, as Figure 1As shown in the figure, the traffic control method in the human-machine co-driving environment includes: S101. Obtain a plurality of control lines at an intersection and vehicle queues that need to pass through the control lines, and construct a plurality of agents corresponding one-to-one to the plurality of control lines. The vehicle queues are mixed queues of manually driven vehicles and self-driving vehicles.
[0029] Among them, the control line refers to a line used to determine the vehicle queue that needs to pass through this line.
[0030] Specifically, the control line is perpendicular to the center line of the lane. The control line can be a virtual line or the actual stop line of the intersection.
[0031] In a preferred embodiment of the present invention, the control line is the actual stop line of the intersection.
[0032] Specifically, the method for obtaining the plurality of control lines is as follows: Store the spatial positions of the control lines in the storage medium of the roadside unit (RSU) in advance, and when performing step S101, call and obtain them from the storage medium.
[0033] It should be noted that: The acquisition of the control line includes not only the acquisition of the spatial position but also the acquisition of the status of the control line. Its status includes: passable and non-passable.
[0034] Among them, there are two ways to determine passable and non-passable. One is determined by the signal timing of the intersection. When the signal timing corresponding to the control line is green, the status is passable, otherwise it is non-passable. The other is to judge whether there is a conflict between the control line and other control lines. When there is a conflict, the status is non-passable, and when there is no conflict, the status is passable. Among them, the judgment of conflict is: When two control lines are both passable, whether there is a traffic accident.
[0035] Specifically, the method for obtaining the vehicle queue is as follows: The roadside unit is communicatively connected to each vehicle in the vehicle queue to obtain the vehicle queue in real time.
[0036] S102. Train the proximal policy optimization model of the constructed multiple agents to obtain a target policy network.
[0037] Among them, the construction of the proximal policy optimization model includes the construction of the action space, state space, policy network, value network, reward function, etc. of each agent.
[0038] S103. Determine the control parameters of the vehicle queue based on the target policy network. The control parameters include the target lane, target speed, and target acceleration for the vehicle queue to pass through the control line.
[0039] It should be understood that: The actions determined by the target policy network are the control parameters.
[0040] S104. Control the manually driven vehicle to travel in the target lane, and control the autonomous vehicle to travel in the target lane, at the target speed, and with the target acceleration.
[0041] Specifically, the manually driven vehicle is locked in the target lane for travel, and its traveling speed and acceleration can be controlled by the driver himself, providing an operating space for the driver of the manually driven vehicle. For the autonomous vehicle, it needs to strictly travel in the target lane, at the target speed, and with the target acceleration.
[0042] It should be understood that: The traffic control method in the human-machine mixed driving environment in the embodiments of the present invention can be implemented in any device based on the traffic control method in the human-machine mixed driving environment, such as: roadside devices, etc. Specifically, the traffic control method in the human-machine mixed driving environment is stored in the above-mentioned device in the form of a compiled program. When the device is started, the program is called, and the traffic control method in the human-machine mixed driving environment is implemented.
[0043] Compared with the prior art, the traffic control method in the human-machine mixed driving environment provided by the embodiments of the present invention constructs multiple intelligent agents corresponding one-to-one to multiple control lines at the intersection, and uses the deep reinforcement learning algorithm to determine the control parameters of the vehicle queue in the human-machine mixed driving environment, that is: the target lane, the target speed, and the target acceleration, realizing the direct control of the vehicle queue. Compared with the scheme of dynamically adjusting signal timing in the prior art, there is no need for the traffic lights at the intersection to have the ability of dynamic adjustment, and it can even be applied to intersections without traffic lights, realizing the traffic control in the human-machine mixed driving environment without dynamically adjusting signal timing, improving the applicability of the traffic control method. And it can adapt to a wider range of vehicle queues, improving the management ability of the traffic control method.
[0044] Furthermore, the embodiments of the present invention construct a proximal policy optimization model. Compared with the traditional deep reinforcement learning model, by introducing the proximal policy, the learning process is made more stable, improving the stability and reliability of the control of the vehicle queue.
[0045] Even further, when the embodiments of the present invention control the vehicle queue, they perform differential control on the manually driven vehicle and the autonomous vehicle. On the premise of realizing the simultaneous control of the manually driven vehicle and the autonomous vehicle, the manual behavior of the manually driven vehicle is taken into account, which is more in line with the actual situation, that is: further improving the flexibility, rationality, and reliability of the traffic control method.
[0046] In some embodiments of the present invention, as Figure 2 shown, step S102 includes: S201. Generate an initial action based on the initial policy network and the environmental state, and the initial action includes the lane, speed, and acceleration of the vehicle queue passing through the control line.
[0047] Among them, the initial policy network can be obtained by either random setting or pre-training.
[0048] It should be understood that action selection is shared among multiple agents. The agents share the environmental state and action selection through V2X communication for global coordination.
[0049] It should be noted that the initial action refers to the initial action of the leading vehicle in the vehicle queue, and then the actions of other vehicles are determined based on the action topology relationship between the leading vehicle and all the vehicles behind it. In other words, the traffic control method in the embodiments of the present invention only controls the leading vehicle, and after the control parameters of the leading vehicle are determined, all the vehicles behind it can be adaptively determined according to the action topology relationship.
[0050] S202. Determine the reward value and the new environmental state for executing the initial action, and respectively determine the environmental state value estimate and the new environmental state value estimate based on the environmental state and the new environmental state.
[0051] Among them, the value estimate is determined based on the value function, and the value function can also be obtained by either random setting or pre-training.
[0052] S203. Determine the advantage function value based on the reward value, the environmental state value estimate, and the new environmental state value estimate; S204. Update the initial policy network based on the advantage function value and the proximal optimization policy to obtain an updated policy network; S205. Determine whether the convergence condition is satisfied. When it is satisfied, update the policy network to the target policy network.
[0053] It should be noted that when the convergence condition is not satisfied, use the updated policy network as the initial policy network, use the updated values such as the new environmental state value estimate as the environmental state value estimate, etc. as the initial values, then return to step S201, and repeat steps S201 - S205 until the convergence condition is satisfied to obtain the target policy network.
[0054] In the specific embodiments of the present invention, the convergence condition is: the number of iterations reaches the maximum number of iterations or the change value of the updated policy network is less than the change threshold.
[0055] Among them, the change value of the updated policy network refers to the change value compared with the policy network before the current update.
[0056] To take into account the traffic flow passing efficiency and passing safety under the traffic control method, in some embodiments of the present invention, determining the reward value for executing the initial action in step S202 includes: Determine the delay index and the safety index for executing the initial action; The weighted sum of the delay index and the safety index is used as the reward value.
[0057] In the embodiment of the present invention, the weighted sum of the delay index and the safety index is used as the reward value, so that the parameters in the two dimensions of delay and safety are both considered in the reward value, ensuring that the control parameters determined based on the reward value take into account both the traffic flow passing efficiency and safety.
[0058] Among them, the delay index can be determined based on the overall delay time of the intersection, and the safety index can be determined based on the collision probability, which will not be elaborated here one by one.
[0059] Since the behavior of the driver has a great influence on the manually driven vehicle, when both the delay index and the safety index are not high, sometimes it is not because the target policy network is unreasonable, but may be caused by the behavior of the driver, such as answering the phone, chatting, etc.
[0060] In some embodiments of the present invention, in order to eliminate the interference of the driver's behavior on the determination of the target policy network to a certain extent, the reward value is the weighted sum of the first reward value of the manually driven vehicle and the second reward value of the autonomous vehicle. Among them, both the first reward value and the second reward value are determined by the delay index and the safety index. The difference is that the first reward value is the product of the weighted sum of the delay knowledge and the safety index and the influence coefficient, while the second reward value is the weighted sum of the delay knowledge and the safety index.
[0061] In other words, the first reward value is set as the product of the second reward value and the influence coefficient, and the influence of the driver's individual behavior is characterized by the influence coefficient, thereby ensuring the accuracy of the reward value. Therefore, the accuracy of the determined target policy network can be further improved to improve the control accuracy and reliability of the vehicle queue.
[0062] In some embodiments of the present invention, as Figure 3 shown, step S204 includes: S301. Determine the optimization direction of the initial policy network based on the advantage function value; S302. Determine the optimization amplitude of the initial policy network based on the proximal optimization policy; S303. Update the initial policy network based on the optimization direction and the optimization amplitude to obtain an updated policy network.
[0063] Specifically, the advantage function provides the direction of action improvement, and the proximal optimization policy restricts the update amplitude to ensure the training smoothness, that is, realizes efficient and stable policy optimization in complex tasks.
[0064] In the specific embodiment of the present invention, the advantage function value is:
[0065] In the formula, is the advantage function value; is the reward value; is the discount factor; is the estimated value of the new environmental state; is the estimated value of the environmental state.
[0066] The proximal optimization strategy is:
[0067]
[0068] In the formula, is the optimization amplitude; is the expected value; is the minimum value function; is the probability ratio of the actions selected by the old and new strategies; is the advantage function value at time t; is the truncation function; is the hyperparameter; is the parameter of the initial policy network.
[0069] It should be noted that: The goal of training the proximal policy optimization models of multiple constructed agents is to achieve the best balance in the long-term cumulative rewards of each agent, ensuring both the optimization of the local goals of each agent and the optimization of the global traffic flow. Through multiple training iterations, it converges to an optimal policy, enabling the agents to cooperate in a complex traffic environment, maximizing traffic efficiency, and ensuring driving safety.
[0070] In summary, the traffic control method in the human-machine mixed driving environment proposed in the embodiments of the present invention combines multi-agent collaborative strategies and deep reinforcement learning algorithms to achieve dynamic regulation and optimization of vehicle flow, can flexibly adapt to different traffic scenarios, and improve the adaptability and compatibility of traffic control.
[0071] To better implement the traffic control method in the human-machine mixed driving environment in the embodiments of the present invention, correspondingly, the embodiments of the present invention also provide a traffic control device in the human-machine mixed driving environment, as Figure 4 shown. The traffic control device 400 in the human-machine mixed driving environment includes: A multi-agent construction unit 401, configured to obtain a plurality of control lines at an intersection and vehicle queues passing through the control lines, and construct a plurality of agents corresponding one-to-one to the plurality of control lines, where the vehicle queues are mixed queues of manually driven vehicles and self-driving vehicles; The target policy network determination unit 402 is configured to train the proximal policy optimization models of multiple agents constructed to obtain a target policy network; The control parameter determination unit 403 is configured to determine the control parameters of the vehicle queue based on the target policy network, where the control parameters include the target lane, target speed, and target acceleration of the vehicle queue passing through the control line; The traffic control unit 404 is configured to control the manually-driven vehicle to travel in the target lane, and control the autonomous vehicle to travel in accordance with the target lane, target speed, and target acceleration.
[0072] The traffic control device 400 in the human-machine mixed driving environment provided in the above embodiment can implement the technical solutions described in the above embodiment of the traffic control method in the human-machine mixed driving environment. The specific implementation principles of the above modules or units can be referred to the corresponding content in the above embodiment of the traffic control method in the human-machine mixed driving environment, which will not be elaborated here.
[0073] As Figure 5 shown, the present invention also correspondingly provides a roadside device 500. The roadside device 500 includes a processor 501, a memory 502, and a display 503. Figure 5 Only some components of the roadside device 500 are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0074] In some embodiments, the processor 501 may be a central processing unit (CPU), a microprocessor, or other data processing chips, and is configured to run the program code stored in the memory 502 or process data, such as the traffic control method in the human-machine mixed driving environment of the present invention.
[0075] In some embodiments, the memory 502 may be an internal storage unit of the roadside device 500, such as the hard disk or memory of the roadside device 500. In other embodiments, the memory 502 may also be an external storage device of the roadside device 500, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the roadside device 500.
[0076] Furthermore, the memory 502 may also include both the internal storage unit and the external storage device of the roadside device 500. The memory 502 is used to store the application software installed on the roadside device 500 and various types of data.
[0077] The display 503 may be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. in some embodiments. The display 503 is used to display information of the roadside device 500 and to display a visual user interface. The components 501-503 of the roadside device 500 communicate with each other via a system bus.
[0078] In some embodiments of the present invention, when the processor 501 executes the traffic control program in the human-machine co-driving environment in the memory 502, the following steps can be achieved: Obtain a plurality of control lines at an intersection and vehicle queues that need to pass through the control lines, and construct a plurality of agents corresponding one-to-one to the plurality of control lines. The vehicle queues are a mixed queue of human-driven vehicles and autonomous vehicles; Train the proximal policy optimization model of the constructed plurality of agents to obtain a target policy network; Determine the control parameters of the vehicle queue based on the target policy network. The control parameters include the target lane, target speed, and target acceleration for the vehicle queue to pass through the control line; Control the human-driven vehicles to drive in the target lane, and control the autonomous vehicles to drive in accordance with the target lane, target speed, and target acceleration.
[0079] It should be understood that when the processor 501 executes the traffic control program in the human-machine co-driving environment in the memory 502, in addition to the above functions, other functions can also be achieved. For details, reference can be made to the description of the corresponding method embodiments above.
[0080] Those skilled in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The computer program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.
[0081] The above has introduced in detail a traffic control method, device, and roadside device in a human-machine co-driving environment provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A traffic control method in a human-machine co-driving environment, characterized in that, including: Obtain multiple control lines at an intersection and vehicle queues passing through the control lines, and construct multiple agents corresponding one-to-one to the multiple control lines, where the vehicle queues are a mixed queue of human-driven vehicles and autonomous vehicles; Train the proximal policy optimization model of the constructed multiple agents to obtain a target policy network; Determine control parameters of the vehicle queue based on the target policy network, where the control parameters include the target lane, target speed, and target acceleration for the vehicle queue to pass through the control line; Control the human-driven vehicles to drive in the target lane, and control the autonomous vehicles to drive in the target lane, at the target speed, and with the target acceleration; 2. The traffic control method in a human-machine co-driving environment according to claim 1, wherein The training of the proximal policy optimization model of the constructed multiple agents includes: Generate initial actions based on an initial policy network and an environmental state, where the initial actions include the lane, speed, and acceleration for the vehicle queue to pass through the control line; Determine a reward value and a new environmental state for executing the initial actions, and respectively determine an environmental state value estimate and a new environmental state value estimate based on the environmental state and the new environmental state; Determine an advantage function value based on the reward value, the environmental state value estimate, and the new environmental state value estimate; Update the initial policy network based on the advantage function value and the proximal optimization policy to obtain an updated policy network; Judge whether a convergence condition is satisfied. When it is satisfied, the updated policy network is the target policy network.
3. The traffic control method in a human-machine co-driving environment according to claim 2, characterized in that, The convergence condition is: the number of iterations reaches the maximum number of iterations or the change value of the updated policy network is less than a change threshold.
4. The traffic control method in the human-machine co-driving environment according to claim 2, wherein, The determination of the reward value for executing the initial actions includes: Determine a delay index and a safety index for executing the initial actions; Use the weighted sum of the delay index and the safety index as the reward value.
5. The traffic control method in a human-machine co-driving environment according to claim 2, wherein, The updating of the initial policy network based on the advantage function value and the proximal optimization policy to obtain an updated policy network includes: Determine an optimization direction of the initial policy network based on the advantage function value; Determine an optimization amplitude of the initial policy network based on the proximal optimization policy; Update the initial policy network based on the optimization direction and the optimization amplitude to obtain the updated policy network.
6. The traffic control method in a human-machine co-driving environment according to claim 5, wherein The advantage function value is: Wherein, is the advantage function value; is the reward value; is the discount factor; is the estimated value of the new environmental state; is the estimated value of the environmental state.
7. The traffic control method in the human-machine co-driving environment according to claim 5, characterized in that The proximal optimization policy is: In the formula, is the optimization amplitude; is the expected value; is the minimum value function; is the probability ratio of the new and old policy selection actions; is the advantage function value at time t; is the truncation function; is the hyperparameter; are the parameters of the initial policy network.
8. The traffic control method in a human-machine co-driving environment according to claim 1, wherein The control line is the actual stop line of the intersection.
9. A traffic control device in a human-machine co-driving environment, characterized in that, including: A multi-agent construction unit for obtaining multiple control lines at an intersection and vehicle queues passing through the control lines, and constructing multiple agents corresponding one-to-one to the multiple control lines, where the vehicle queues are a mixed queue of human-driven vehicles and autonomous vehicles; A target policy network determination unit for training the proximal policy optimization model of the constructed multiple agents to obtain a target policy network; A control parameter determination unit for determining control parameters of the vehicle queue based on the target policy network, where the control parameters include the target lane, target speed, and target acceleration for the vehicle queue to pass through the control line; A traffic control unit is configured to control the manually-driven vehicle to travel along the target lane, and control the autonomous vehicle to travel along the target lane, at the target speed and with the target acceleration.
10. A roadside device, characterized in that, It includes a memory and a processor, wherein, the memory is used for storing programs; the processor is coupled to the memory and is configured to execute the programs stored in the memory to implement the steps in the traffic control method in the human-machine co-driving environment according to any one of claims 1 to 8 above.