A large model-based multi-task concurrency and line planning speed control method

By adopting a multi-task concurrency and route planning speed control method based on a large model, the problem of insufficient prediction accuracy caused by static speed assumptions in convoy control is solved. This enables dynamic optimization and accurate control of convoys and traffic lights, improving the prediction accuracy and execution consistency of convoy arrival time.

CN121393152BActive Publication Date: 2026-05-15QINGDAO HISENSE TRANS TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QINGDAO HISENSE TRANS TECH
Filing Date
2025-12-22
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing fleet control methods rely on static speed assumptions and lack the ability to adapt and fuse constraints, resulting in insufficient accuracy in predicted arrival times and inconsistencies between predictions and actual execution.

Method used

A multi-task concurrent and route planning speed control method based on a large model is adopted. The speed sequence and time sequence of the convoy are obtained by the prediction model, and the traffic control strategy, including the adjustment strategy of convoy and traffic lights, is generated by the strategy generation model. The optimization model is used for verification and optimization.

Benefits of technology

It improves the accuracy and effectiveness of fleet control, and realizes global coordination and dynamic optimization under conditions of multiple fleets, multiple routes and complex disturbances, ensuring the on-time arrival of fleets and efficient coordination of traffic signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121393152B_ABST
    Figure CN121393152B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent transportation, and in particular relates to a multi-task concurrency and route planning speed control method based on a large model. In the embodiment of the application, the speed sequence and the time sequence of a plurality of vehicle platoons are predicted through a prediction model, so that the control of the plurality of vehicle platoons is no longer based on a single static speed. Then, a strategy generation model is used to generate a vehicle platoon control strategy based on the speed sequence and the time sequence, so that the vehicle platoon control strategy is used for control, the existing vehicle platoon control method depends on a static speed assumption, lacks adaptive and constraint fusion capability, and the prediction accuracy of the predicted arrival time of the vehicle platoon is insufficient and the prediction and actual execution are inconsistent, so that the accuracy and effect of the vehicle platoon control are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent transportation technology, and in particular to a method for multi-task concurrency and route planning speed control based on a large model. Background Technology

[0002] In related technologies, when there are large-scale events, traffic control plans are planned in advance, especially the routes and arrival times of each convoy.

[0003] However, existing technologies mainly rely on static speed assumptions, that is, assuming the speed of each convoy and that the convoys travel at a constant speed, in order to assess the conflict situation of the convoys and make corresponding adjustments based on the conflict situation. This convoy control method lacks the ability to adapt and integrate constraints. Under conditions of multiple convoys, multiple routes, and complex disturbances, it is difficult to achieve global coordination and dynamic optimization, and it often leads to problems such as insufficient accuracy in predicting the estimated arrival time of the convoys and inconsistency between prediction and actual execution.

[0004] Therefore, there is an urgent need for a fleet control method that can accurately generate fleet control strategies. Summary of the Invention

[0005] This application provides a multi-task concurrent and route planning speed control method based on a large model to solve the problems of existing fleet control methods that rely on static speed assumptions, lack adaptive and constraint fusion capabilities, resulting in insufficient accuracy in predicting the estimated arrival time of the fleet and inconsistency between prediction and actual execution.

[0006] In a first aspect, embodiments of this application provide a multi-task concurrency and route planning speed control method based on a large model, the method comprising:

[0007] Obtain road network information and driving routes of multiple vehicle fleets within a designated area;

[0008] A prediction model is used to determine the speed sequence and time sequence of the multiple convoys based on the road network information and the driving routes of the multiple convoys; the time sequence is used to indicate the predicted time for the corresponding convoy to arrive at at least one preset node;

[0009] A strategy generation model is used to generate traffic control strategies based on the speed sequences of the multiple vehicle fleets, the time sequences, the road network information, and preset constraints. The traffic control strategies include vehicle fleet control strategies that adjust the speed sequences of the vehicle fleets, and signal control strategies that adjust traffic lights.

[0010] Secondly, embodiments of this application also provide a multi-task concurrency and route planning speed control device based on a large model, the device comprising:

[0011] The prediction module is used to acquire road network information of a set area and the driving routes of multiple convoys; using a prediction model, based on the road network information and the driving routes of the multiple convoys, it determines the speed sequence and time sequence of the multiple convoys; the time sequence is used to indicate the predicted time for the corresponding convoy to arrive at at least one preset node;

[0012] The control module is used to generate traffic control strategies based on the speed sequences of the multiple vehicle fleets, the time sequences, the road network information, and preset constraints using a strategy generation model; wherein the traffic control strategies include vehicle fleet control strategies that adjust the speed sequences of the vehicle fleets, and signal control strategies that adjust traffic lights.

[0013] Thirdly, embodiments of this application also provide an electronic device, which includes at least a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the steps of the multi-task concurrency and route planning speed control method based on a large model as described above.

[0014] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the multi-task concurrency and route planning speed control method based on a large model as described above.

[0015] In this embodiment, the electronic device acquires road network information of a designated area and the driving routes of each convoy; using a prediction model, based on the road network information and the driving routes of multiple convoys, it determines the speed sequences and time sequences of multiple convoys; the time sequences are used to indicate the predicted time for the corresponding convoy to arrive at at least one preset node; using a strategy generation model, it generates a traffic control strategy based on the speed sequences, time sequences, road network information, and preset constraints of multiple convoys; wherein, the traffic control strategy includes a convoy control strategy that adjusts the speed sequences of convoys, and a signal control strategy that adjusts traffic lights. In this embodiment, by predicting the speed sequences and time sequences of convoys using a prediction model, convoy control is no longer based on a single static speed. Then, by using a strategy generation model, a convoy control strategy is generated based on the speed sequences and time sequences, thereby enabling control based on the convoy control strategy. This solves the problem that existing convoy control methods rely on static speed assumptions, lack adaptive and constraint fusion capabilities, resulting in insufficient prediction accuracy of convoy arrival times and inconsistencies between predictions and actual execution, thus improving the accuracy and effectiveness of convoy control. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A schematic diagram of a fleet control process provided in an embodiment of this application;

[0018] Figure 2 A schematic diagram illustrating the training process of the strategy generation model provided in this application embodiment;

[0019] Figure 3 A flowchart of the overall fleet control provided in this application embodiment;

[0020] Figure 4 A schematic diagram of the fleet control system provided in the embodiments of this application;

[0021] Figure 5 This application provides a schematic diagram of a fleet control device structure.

[0022] Figure 6 This is a schematic diagram of an electronic device structure provided in an embodiment of this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0024] The terms "first" and "second" in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising" and any variations thereof are intended to cover non-exclusive protection. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. The term "multiple" in this application can mean at least two, for example, two, three, or more, and is not limited by the embodiments of this application.

[0025] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These embodiments should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that in the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solutions of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0026] Past traffic control schemes relied heavily on rule engines, threshold determination, and offline simulation. They were characterized by static speed assumptions (i.e., vehicle speed was treated as constant) and post-event evaluation, lacking adaptive and constraint fusion capabilities. Under conditions of multiple fleets, multiple routes, and complex disturbances, it was difficult to achieve global coordination and dynamic optimization, often resulting in problems such as insufficient ETA prediction accuracy, low utilization of green window resources, and inconsistencies between simulation and actual execution. There was a lack of a reliable and verifiable closed-loop control system.

[0027] In recent years, the integration of Large Language Models (LLMs) and Reinforcement Learning (RL) has driven the development of intelligent decision-making technologies. Inference models such as DeepSeek-R1, Qwen3, and QwQ-32B have demonstrated strong task understanding and policy generation capabilities, but they still have limitations in the field of traffic control optimization: general-purpose LLMs lack structured representations of road topology and signal constraints, making it difficult to output executable control policies within a framework of continuous spatiotemporal variables and safety constraints.

[0028] In summary, the relevant technologies generally rely on fixed thresholds or static speed assumptions for conflict warning and speed adjustment, making it difficult to balance safety, punctuality, and coordination in scenarios with multiple fleets operating concurrently.

[0029] Based on this, this application proposes a multi-task concurrent and route planning speed control method, device, equipment, and medium based on a large model. The method acquires road network information of a designated area and the driving routes of multiple vehicle fleets. A prediction model is used to determine the speed sequences and time sequences of multiple vehicle fleets based on the road network information and the driving routes of the multiple vehicle fleets. The time sequences are used to indicate the predicted arrival times of the corresponding vehicle fleets at multiple preset nodes. A strategy generation model is used to generate traffic control strategies based on the speed sequences, time sequences, road network information, and preset constraints of the multiple vehicle fleets. The traffic control strategies include vehicle fleet control strategies that adjust the speed sequences of the vehicle fleets, and signal control strategies that adjust traffic lights.

[0030] Example 1:

[0031] Figure 1 A schematic diagram of a multi-task concurrency and route planning speed control process based on a large model is provided for embodiments of this application. The process includes:

[0032] S101: Obtain road network information and driving routes of multiple convoys in a designated area; use a prediction model to determine the speed sequence and time sequence of the multiple convoys based on the road network information and the driving routes of the multiple convoys; the time sequence is used to indicate the predicted time for the corresponding convoy to arrive at at least one preset node.

[0033] This application provides a method for multi-task concurrency and route planning speed control based on a large model, which can be applied to electronic devices such as PCs, servers, and traffic control systems.

[0034] In this embodiment, when there is a need for fleet management, a fleet management command can be sent to the electronic device through a client, other devices, or the display interface of the electronic device. This fleet management command carries the fleet identifier of each fleet to be controlled and the area identifier of the designated area. The designated area includes the event venue; for example, the designated area can be a 5km radius around the event venue.

[0035] After receiving the fleet management request, the electronic device obtains the driving route of each fleet based on the identifier of each fleet among the multiple fleets carried in the fleet management request; the electronic device also obtains the road network information of the designated area based on the identifier of the designated area carried in the fleet management request. This road network information includes, but is not limited to, road network topology, signal timing, trunk line coordination parameters, priority and meeting time windows, closure and speed limit information, floating car and detector traffic data, and historical operation records.

[0036] Among them, electronic devices can obtain road network information of the aforementioned designated area from the intelligent transportation platform, and can also obtain road network information of the aforementioned designated area from other devices or servers.

[0037] For example, electronic devices access multiple interfaces of the intelligent transportation platform in real time, including but not limited to: high-precision map services providing centimeter-level accuracy in road geometry; traffic signal control systems providing real-time phase status and timing schemes for intersections across the entire road network; an IoT sensing layer providing real-time traffic flow and speed data; and an urban event management center pushing temporary control information. The electronic devices obtain road network topology from the high-precision map service, signal timing and arterial coordination parameters from the traffic signal control system, traffic flow data and historical operation records from the IoT sensing layer, and priority, venue time windows, lockdown, and speed limit information from the urban event management center.

[0038] Additionally, the electronic device can obtain the fleet task package corresponding to each fleet during task scheduling. This fleet task package can be a structured JavaScript Object Notation (JSON) object, containing the unique identifier of the corresponding fleet, fleet size, priority level, origin and destination geofences, core time windows, and fleet route. After obtaining the fleet task package corresponding to each fleet, the electronic device parses the fleet task package corresponding to each fleet to determine the corresponding driving route for each fleet.

[0039] In addition, in this embodiment of the application, the above-mentioned fleet management requirements may also include the driving routes of each fleet and the road network information of the designated area.

[0040] After acquiring the road network information of the designated area and the driving route of the convoy, the electronic device can perform unified spatiotemporal synchronization, anomaly removal, and interpolation repair on the road network information and driving route, thereby forming a standardized data stream and providing basic input for modeling.

[0041] For example, electronic devices can use Global Positioning System (GPS) timestamps as a reference and interpolate and synchronize road network information and driving routes using Kalman filtering; electronic devices can also use sliding window statistical methods or deep learning anomaly detection models to remove abnormal information in road network information and driving routes; electronic devices can also use shortest path backtracking based on road network topology to fill in missing points in road network information and driving routes; electronic devices can also convert raw data into model-friendly feature vectors, such as converting social traffic flow into saturation indicators.

[0042] Based on this, electronic devices can invoke predictive models to predict the driving conditions of each convoy, providing forward-looking information for subsequent convoy control processes. This predictive model can be a Spatio-Temporal Foundation Model (STFM). In this embodiment, STFM employs a two-stage learning process: large-scale pre-training and supervised fine-tuning (SFT) in the transportation domain. Specifically, STFM first learns spatio-temporal dependencies based on multi-city trajectories, and then performs supervised fine-tuning in a specific escort scenario, outputting speed sequences and beat distributions.

[0043] In this embodiment, STFM outputs two key results: first, a continuous velocity field v(x,t) for each convoy on a future spatiotemporal grid, i.e., the velocity sequence of each convoy, thus providing the convoy with macroscopic perception of the traffic environment; second, a time series for each convoy, consisting of predicted arrival times for each node on its route. Additionally, STFM can output confidence intervals to represent the confidence level of each convoy's velocity and time series. The velocity sequence of each convoy can be output as a velocity curve, which is not limited here.

[0044] Specifically, the electronic device inputs the road network information of the designated area and the driving routes of each convoy into a trained prediction model. Based on the road network information of the designated area and the driving routes of each convoy, the prediction model determines and outputs the speed sequence and time sequence for each convoy. The time sequence indicates the predicted time for the corresponding convoy to arrive at at least one preset node. This preset node can be any intersection or other location within a road segment, and includes the destination, i.e., the location of the meeting place.

[0045] In one possible implementation, the inputs to the STFM can be the convoy's driving path, the signal sequence determined based on signal timing, trunk line coordination parameters, closure status, and historical disturbance annotations carried in historical operation records; the output of the STFM can be the speed field. Node beat The STFM output provides initial values ​​and uncertainty boundaries for subsequent fleet management processes, supporting early conflict avoidance. The velocity field is the velocity sequence in the above embodiments, and the node beats are the time sequences in the above embodiments.

[0046] S102: Using a strategy generation model, a traffic control strategy is generated based on the speed sequences of the multiple vehicle fleets, the time sequences, the road network information, and preset constraints; wherein, the traffic control strategy includes a vehicle fleet control strategy that adjusts the speed sequences of the vehicle fleets, and a signal control strategy that adjusts the traffic lights.

[0047] In this embodiment, after determining the speed and time sequences of each vehicle fleet, the electronic device employs a policy generation model to generate corresponding traffic control policies. This policy generation model can be a large language model with a small number of parameters, such as a small language model with approximately 3 bytes of parameters. This policy generation model possesses contextual understanding and multi-constraint reasoning capabilities, enabling it to express the policy space in a linguistic manner and achieve a more flexible "knowledge-policy-constraint" mapping. For example, this policy generation model can be a large spatiotemporal model.

[0048] Specifically, the electronic device inputs all the outputs of the prediction model, road network information, and a set of formalized pre-defined constraints into the policy generation model. These pre-defined constraints can be compiled into a constraint graph, a structured knowledge base that transforms rules into computable logic, explicitly including requirements such as node time isolation, minimum following distance, green window compatibility, trunk line coordination continuity, and opportunity constraints.

[0049] For example, the preset constraints include, but are not limited to, at least one of the following:

[0050] Constraint 1: Node time isolation constraint, used to ensure that the arrival time difference between different convoys at the same node is not less than the safe time interval. To prevent encounters and conflicts.

[0051] For example, the node's time isolation constraint satisfies the following formula:

[0052]

[0053] in, The first time that convoy f arrives at node n. For the second time when convoy g arrives at node n, For safety time intervals.

[0054] Constraint 2: Following distance constraint, used to ensure that the distance between the two convoys at any given time is not less than the safe following distance. Avoid following too closely.

[0055] For example, the following formula is satisfied by the following following distance constraint:

[0056]

[0057] in, Let be the distance between convoy f and convoy g at time t. For safe following distance.

[0058] Constraint 3: Green window compatibility constraint, which limits the arrival time of the convoy to fall within the green light period to ensure consistent signal control.

[0059] For example, the green window compatibility constraint satisfies the following formula:

[0060]

[0061] in, The first time that convoy f arrives at node n. Let be the green window time window for node n.

[0062] Constraint 4: Coordination and continuity constraint condition. When making green extension, red cut-off, or offset adjustments, ensure the continuity of trunk bandwidth and phase offset, and do not disrupt the overall coordination.

[0063] Constraint 5: Opportunity Constraint, used to express the reliability level of task punctuality, requires that within a pre-set probability threshold... The convoy successfully passed through the venue within the scheduled time window.

[0064] For example, the chance constraint satisfies the following formula:

[0065]

[0066] in, Let f be the arrival time of convoy f. For the time window of the meeting, This is the start time of the venue's time window. This is the end time of the time window for the event. Let f be the confidence probability that vehicle f will successfully complete the passage within the scheduled time window at the venue. A preset probability threshold is set.

[0067] In this embodiment, after receiving the speed sequence, time sequence, road network information, and preset constraints for each fleet, the policy generation model performs an inference process. The internal mechanism of this policy generation model combines the contextual understanding capabilities of LLM with policy priors learned through reinforcement learning. The output of the policy generation model is a highly structured JSON object containing traffic control policies. These traffic control policies include fleet management policies that adjust the speed sequences of the fleets, and signal management policies that adjust traffic lights.

[0068] For example, a fleet management strategy could be a speed adjustment percentage specified for each fleet, and a signal control strategy could be a green extension for each node. The fleet management strategy could also include decision rationale in natural language.

[0069] For example, if STFM predicts a 5-second conflict between convoy A and convoy B at node N3, the strategy generation model might output: increase the speed of the high-priority convoy A by 5%, allowing it to pass earlier; simultaneously, apply a -3% deceleration to the low-priority convoy B, and, in conjunction with the N3 redline cutoff of -2 seconds, delay its arrival time. This expands the arrival time difference Δt between convoy A and convoy B to 10 seconds, satisfying the safety isolation requirement, while ensuring that convoy A remains on time.

[0070] In this embodiment, the electronic device predicts the speed and time series of the convoy through a prediction model, so that convoy control is no longer based on a single static speed. Then, a strategy generation model generates a convoy control strategy based on the speed and time series, and control is performed based on the convoy control strategy. This solves the problems of existing convoy control methods relying on static speed assumptions and lacking adaptive and constraint fusion capabilities, which leads to insufficient accuracy in predicting the expected arrival time of the convoy and inconsistency between prediction and actual execution, thereby improving the accuracy and effectiveness of convoy control.

[0071] To improve the accuracy of fleet control, based on the above embodiments, the method in this application embodiment further includes:

[0072] Based on the aforementioned fleet management strategy, the speed sequences and time sequences corresponding to the multiple fleets are adjusted;

[0073] According to the signal control strategy, the signal sequence is adjusted, wherein the signal sequence includes the green window duration corresponding to each preset node;

[0074] An optimization model is used to verify and optimize the speed sequence and the signal sequence based on the updated speed sequence, time sequence, road network information, updated signal sequence, and preset constraints, thereby generating optimized speed sequence and signal sequence.

[0075] In this embodiment of the application, after the electronic device obtains the traffic control strategy output by the strategy generation model, the electronic device can further optimize and refine the traffic control strategy through the optimization model, and perform real-time re-optimization with constraints to ensure the absolute safety and global optimality of the final execution plan.

[0076] Specifically, the electronic equipment adjusts the speed and time sequences for each fleet according to the fleet management strategy. It also adjusts the signal sequences according to the signal management strategy, including the green window duration for each preset node. The electronic equipment inputs the adjusted speed, time, and signal sequences, road network information, and preset constraints into the optimization model. This optimization model then verifies and optimizes the speed and signal sequences based on the updated speed, time, road network information, updated signal sequences, and preset constraints for each fleet, generating optimized speed and signal sequences.

[0077] The optimization model can be a Model Predictive Control (MPC) architecture, which dynamically and jointly optimizes the updated speed sequence and signal sequence to achieve punctual, smooth and efficient passage for multiple vehicle fleets and routes while ensuring safety and coordination.

[0078] The following describes a possible optimization method:

[0079] The electronic equipment performs precise mathematical transformations on the speed profile of each platoon based on the platoon control strategy within the traffic management policy. For example, if the platoon control strategy recommends an acceleration of +5%, the new speed sequence is the original speed multiplied by 1.05. The electronic equipment can then recalculate the position and arrival time using numerical integration to obtain the updated time series.

[0080] The electronic equipment modifies the signal timing files of each node according to the signal control strategy in the traffic control strategy, and generates an updated signal sequence. The updated signal sequence clearly records the start and end times of the green light for each node in the future time period.

[0081] The electronic equipment inputs the updated speed sequence, time sequence, road network information, updated signal sequence, and preset constraints of each vehicle fleet into the MPC framework. This MPC framework constructs a complex optimization problem within a finite time domain (e.g., 60 seconds). Its optimization variables include the discretized speed profiles of each vehicle fleet and the signal parameter sets of each intersection (e.g., green extension, red cutoff, phase offset). The objective function within this MPC framework is a multi-dimensional weighted sum, comprehensively considering Estimated Time of Arrival Deviation (ETA) deviation, social delay costs, conflict risk costs, speed smoothness, and signal switching penalties. The weights of each factor can be dynamically adjusted according to task priority.

[0082] The objective functions include ETA deviation to minimize the difference between actual and planned arrival times, social delay costs to reduce the impact on other vehicles, conflict risk costs to penalize the probability of node encounters, speed smoothness to limit acceleration abrupt changes to ensure ride comfort, and signal switching penalties to avoid frequent adjustments that could disrupt signal stability.

[0083] For example, the objective function satisfies the following formula:

[0084]

[0085] in, , , , and To preset weights, Indicates the estimated arrival time. Indicates the actual arrival time. This represents the absolute value of the ETA deviation between the estimated and actual arrival times. Indicates the social cost of delay. Indicates the risks and costs of conflict. Indicates speed smoothness, This indicates a penalty for signal switching.

[0086] In this embodiment, the constraints of the MPC framework include, but are not limited to, hard constraints such as speed range, acceleration range, minimum vehicle spacing, and node time isolation, as well as opportunity constraints constructed from the uncertainty of the STFM output, namely, the confidence probability that the convoy will successfully complete the passage within the predetermined meeting time window. , To pre-set the probability threshold and ensure that the time window requirements are met even in the worst case, there are also coordination constraints such as trunk offset change to ensure that the green wave band is not disrupted.

[0087] In this embodiment, the MPC framework employs an efficient nonlinear programming solver to solve the problem in each control cycle, strictly adhering to the rolling time-domain control principle. This means that only the first control step of the solution is executed, with the remainder discarded. The next cycle re-optimizes based on the latest vehicle and road network conditions, forming a robust dynamic closed loop. This allows for continuous adaptation to minor environmental disturbances, maintaining the timeliness of the strategy. In this embodiment, when environmental disturbances or task conditions change, the MPC framework can be automatically triggered for re-optimization, achieving dynamic closed-loop adjustment.

[0088] Furthermore, after each optimization, the MPC framework not only outputs executable speed and signal sequences but also generates a verifiable control certificate. This control certificate includes, but is not limited to, a Δt non-meeting certificate and a probabilistic ETA report. The Δt non-meeting certificate is a formal proof listing the Δt values ​​for all convoy pairs at all nodes, confirming that they are all greater than or equal to the safety threshold. The probabilistic ETA report, considering the uncertainty of STFM, provides the confidence probability of each convoy arriving on time. This control certificate not only serves as the basis for execution but also provides tamper-proof data credentials for mission auditing and post-mission review, ensuring the reliability, traceability, and auditability of the entire control process, thus achieving a seamless transition from policy generation to safe execution.

[0089] To improve the accuracy of fleet control, based on the above embodiments, in the embodiments of this application,

[0090] The strategy generation model is a spatiotemporal large model;

[0091] The training process of the policy generation model includes:

[0092] Acquire training samples, which include sample road network information, sample speed sequences of multiple vehicle fleets, and sample time sequences;

[0093] The model is generated using the strategy to be trained. Based on the preset constraints, sample road network information, sample speed sequences and sample time sequences of the multiple vehicle fleets, a sample traffic control strategy is generated.

[0094] Based on the traffic control strategy described above, the sample speed sequences and sample time sequences of the multiple vehicle fleets are adjusted.

[0095] A simulation model is used to perform simulation evaluation based on the sample road network information, the adjusted sample speed sequence and sample time sequence of the multiple vehicle fleets, and to determine the sample evaluation results.

[0096] The GRPO algorithm is optimized using a group-relative strategy. Based on the sample evaluation results, the preset reward function, the adjusted sample speed sequences of the multiple vehicle teams, and the sample time sequences, the reward function value is determined, and the parameters of the strategy generation model to be trained are adjusted according to the reward function value.

[0097] In this embodiment, the strategy generation model can be trained by deeply interacting with a high-fidelity differentiable simulation environment and using an advanced Group Relative Policy Optimization (GRPO) algorithm, so that the strategy generation model can learn and master complex multi-vehicle cooperative control strategies from scratch.

[0098] Among them, the policy generation model can be a language model with a small number of parameters. This policy generation model can achieve semantic understanding and action generation through instruction tuning and GRPO reinforcement learning. The trained policy generation model can maximize the on-time rate of tasks, minimize delays and conflicts, ensure green window coordination and comfort, and at the same time have interpretable output (policy rationale and confidence distribution).

[0099] Figure 2 This is a schematic diagram of the training process of the policy generation model provided in the embodiments of this application, as shown below. Figure 2 As shown, the training process of this strategy-generated model includes:

[0100] S201: Obtain training samples.

[0101] In this embodiment of the application, the electronic device can obtain training samples from a sample database, which includes sample road network information, sample speed sequences for each vehicle fleet, and sample time sequences.

[0102] In addition, the sample database can include countless complete simulation scenarios automatically generated by the scene engine. Each scenario contains randomly selected or combined road network segments, 2 to 10 convoys with different routes, priorities and departure times, as well as different levels of social traffic flow and randomly injected background disturbances (such as accidents and road closures) to ensure that the model can be exposed to various extreme and normal scenarios.

[0103] S202: The policy generation model to be trained makes predictions.

[0104] In this embodiment of the application, the electronic device can input training samples into the strategy generation model to be trained. The strategy generation model generates sample traffic control strategies based on the preset constraints, sample road network information, sample speed sequence and sample time sequence carried in the training samples.

[0105] S203: Simulation evaluation based on sample traffic control strategies.

[0106] In this embodiment of the application, the electronic device can adjust the sample speed sequence and sample time sequence of each vehicle according to the sample traffic control strategy, and use a simulation model to perform simulation based on the adjusted sample speed sequence and sample time sequence to obtain the sample evaluation result corresponding to each vehicle.

[0107] The simulation model can be a differentiable simulation model, which can approximate traditional discrete and non-differentiable traffic processes (such as vehicle queuing and green window occupancy) with smooth mathematical functions, such as using the difference of the sigmoid function to construct a smooth green window occupancy indicator function. This differentiable design allows the differentiable simulation model to output differentiable evaluation metrics, such as total delay J, conflict risk R, and coordination disruption degree. These metrics can be directly used for gradient calculation, greatly improving training efficiency.

[0108] S204: Use the GRPO algorithm for policy updates.

[0109] In this embodiment, after the electronic device obtains the sample evaluation results output by the simulation model, it can use the GRPO algorithm for policy update. GRPO is a policy gradient algorithm that uses relative ranking within a group. For multiple policy variants generated in the same batch, GRPO calculates the reward function value corresponding to their relative advantage, rather than the absolute reward, thereby significantly reducing the variance of the policy gradient and significantly improving the stability and convergence speed of training.

[0110] Subsequently, based on the reward function value, the electronic devices iteratively update the model parameters by maximizing the expected cumulative reward, enabling the policy generation model to gradually evolve into an expert proficient in multi-vehicle coordination and capable of making optimal decisions under complex constraints. Furthermore, because the policy generation model architecture itself possesses language generation capabilities, it also learns to generate policies with natural language justifications during training. This not only enhances the interpretability of the policies but also facilitates human review and intervention.

[0111] To improve the accuracy of fleet control, based on the above embodiments, in this embodiment, before determining the reward function value using the group relative strategy optimization GRPO algorithm according to the sample evaluation results, the preset reward function, the adjusted sample speed sequence of the multiple fleets, and the sample time sequence, the method further includes:

[0112] Based on the information in the simulation results regarding whether the convoy arrived on time, the on-time bonus value is determined.

[0113] The arrival time deviation is determined based on the simulated arrival time of the convoy carried in the simulation results and the predicted arrival time carried in the sample time series.

[0114] The signal switching penalty value is determined based on the number of signal light changes carried in the sample traffic control strategy.

[0115] Based on the information about whether the convoys meet, carried in the simulation results, determine the conflict risk penalty value;

[0116] The reward function value is determined based on the on-time reward value, the arrival time deviation, the signal switching penalty value, the conflict risk penalty value, and the preset weight.

[0117] In this embodiment of the application, the reward function can be a multi-dimensional, configurable weighted sum function, each term of which corresponds to a key business objective, including but not limited to at least one of the following: on-time reward value, arrival time deviation, signal switching penalty value, and conflict risk penalty value.

[0118] The on-time reward value is awarded to convoys for arriving on schedule, incentivizing them to arrive on time. This on-time reward value is composed of a sub-value corresponding to each convoy. For example, for any convoy, if its simulated arrival time in the simulation results precedes the sample arrival time carried in the sample time series (i.e., the simulation results contain information about the convoy's on-time arrival), then the convoy's on-time reward value is determined as the base reward (e.g., +10); otherwise, the convoy's on-time reward sub-value is determined as zero. The electronic equipment determines the on-time reward value as the sum of the on-time reward sub-values ​​corresponding to each convoy.

[0119] In addition, to encourage the model to get as close to the target as possible, a soft boundary mechanism can be introduced, that is, when outside the time window but within a tolerable range (such as ±30 seconds), a partial reward is given, thereby preventing the policy from going to extremes.

[0120] The arrival time deviation is the difference between the sample arrival time carried in the sample time series and the simulated arrival time carried in the simulation results, used to constrain time errors. This arrival time deviation can also be composed of arrival time sub-deviations corresponding to each convoy. For example, for any convoy, the arrival time sub-deviation is determined based on the simulated arrival time of the convoy carried in the simulation results and the predicted arrival time carried in the corresponding sample time series. The electronic equipment determines the arrival time deviation by summing the arrival time sub-deviations corresponding to each convoy.

[0121] The signal switching penalty value is determined based on the number of signal light changes carried in the sample traffic control strategy to maintain control smoothness. The electronic device counts the number of signal light changes at all intersections; in the example, one valid green light transition or red light cutoff is counted as one change. The electronic device stores a correspondence between the number of signal light changes and the signal switching penalty value. Based on this correspondence, the electronic device can determine the signal switching penalty value corresponding to the number of signal light changes. Alternatively, the electronic device can directly determine the number of signal light changes as the signal switching penalty value.

[0122] A conflict risk penalty value is used to prevent convoys from meeting at nodes. The simulation model continuously monitors the entire simulation process, and once it detects any event that violates safety constraints, such as the arrival time difference between two convoys at the same node being less than a safety threshold, or the distance between vehicles on the same road segment being less than the minimum safe distance, the simulation model will record it in the simulation results. Based on this, electronic devices can determine the conflict risk penalty value according to the information on whether the convoys have met carried in the simulation results. For example, if the simulation results carry information about convoys meeting, a base conflict risk penalty value is assigned.

[0123] In addition, in this embodiment of the application, the conflict risk penalty value can also be superimposed. For each encounter, a basic conflict risk penalty value is superimposed on the conflict risk penalty value.

[0124] The electronic device determines the reward function value by weighting the punctuality reward value, arrival time deviation, signal switching penalty value, and conflict risk penalty value. The weight of each parameter can be flexibly configured.

[0125] For example, the reward function value satisfies the following formula:

[0126]

[0127] in, This represents the reward function value. , , and To preset weights, As a reward value for punctuality, For arrival time deviation, This is the penalty value for signal switching. This represents the penalty value for conflict risk.

[0128] In addition, in the embodiments of this application, when determining the reward function value, the electronic device may also add a social delay penalty value to reflect the overall traffic efficiency.

[0129] For example, the reward function value can also satisfy the following formula:

[0130]

[0131] in, This represents the reward function value. , , , and To preset weights, As a reward value for punctuality, For arrival time deviation, The social delay penalty value, This is the penalty value for signal switching. This represents the penalty value for conflict risk.

[0132] To improve the accuracy of fleet control, based on the above embodiments, the method in this application embodiment further includes:

[0133] The preset constraints, the sample evaluation results, the sample speed sequences and sample time sequences adjusted by the multiple vehicle teams are used as inputs to the next iteration of the policy generation model to be trained.

[0134] In this embodiment, the training process of the policy generation model is not a one-way, one-time process, but a closed-loop process that is self-reinforcing and continuously evolving. By introducing a context-enhanced feedback mechanism during this training process, the model's learning efficiency and generalization ability can be significantly improved.

[0135] Specifically, after each round of GRPO training, the electronic device will not discard all the information from this interaction. Instead, it will use the preset constraints, sample evaluation results, and the sample speed and time sequences adjusted by each team as inputs for the next round of the policy generation model to be trained.

[0136] The following specific example illustrates the input and output of the policy generation model during training: the input includes instructions, data input, constraint construction, prediction layer output, and simulation feedback.

[0137] The instructions include: You are a traffic control agent. Please generate a globally coordinated speed and signal adjustment scheme based on the traffic conditions, signal configurations, constraints, and simulation feedback of multiple vehicle fleets and routes.

[0138] Data input includes the following:

[0139] Road network and signaling: Nodes N1–N4 (main lines), N3 intersection point; current timing: N1 green extension 2s, N3 red extension 3s; coordination offset remains unchanged.

[0140] Team status:

[0141] Team A: 3 vehicles, main line L1, ETA deviation 5s;

[0142] Team B: 2 vehicles, branch line L2, ETA deviation +3s.

[0143] Environmental data: Social traffic flow 1200veh / h, moderate disturbance, construction zone present.

[0144] Constraint construction includes the following:

[0145] Time and space constraints: node time isolation Δt_min = 8s; following distance d_min = 25m; upper speed limit v_max = 60km / h, minimum acceleration a_min = 3m / s²;

[0146] Signal coordination constraints: Maintain trunk bandwidth continuity, limit green extension / red cutoff amplitude to no more than ±5s, and offset continuity error ≤2s;

[0147] Safety and social constraints: Maintain a risk probability P_risk ≤ 0.1 for interfering vehicles and social traffic flow to ensure that the main bandwidth of the trunk line is not damaged;

[0148] Task and priority constraints: High-priority convoys can dynamically expand the green window within the coordination zone. The passage priority of high-priority convoys is > medium-priority > low-priority, to maintain the fairness of global scheduling.

[0149] The prediction layer output includes the following:

[0150] Team A's beat τ_A=[N1:15:02:05, N2:15:03:10, N3:15:04:00, N4:15:05:10];

[0151] Team B's beat τ_B = [N3:15:03:55, N4:15:04:45]; potential conflict N3 (Δt = 5s < Δt_min);

[0152] Velocity field prediction v(x,t): A is slightly slower, B is slightly faster, and the confidence level of risk node N3 is 0.85.

[0153] Simulation feedback includes the following:

[0154] Delay J = 12.3s; Risk R = 0.27; Cost of disrupted coordination = 1.8;

[0155] Simulation results show that A queues in front of N3 with a 5-second delay, while B enters the green window area too early.

[0156] Based on the input of the above strategy generation model, the output of the strategy generation model includes, but is not limited to: a set of control strategies for each fleet, which includes fleet identification, driving route, speed adjustment percentage (+ for acceleration, - for deceleration), node green window duration adjustment, and predicted arrival time status; and a multi-fleet global coordination strategy, which includes conflict node location, optimization adjustment strategy, and the logical basis for strategy generation (integrating STFM prediction, constraints, and simulation feedback).

[0157] In the example above, the inputs to the strategy generation model include STFM prediction results (prediction layer output), cost signals from the simulation model (simulation feedback), preset constraints (constraint construction), road network information (road network and information, environmental data), the driving routes of each convoy (convoy status), and prompts (instructions). The outputs of the strategy generation model are candidate speed templates, signal adjustments, priority scheduling suggestions, and explanations of decision rationales, generated in either linguistic or structured parameter form.

[0158] In addition, the output of the model generated by this strategy can be converted by the parsing module into executable initial values ​​or constraint weights for the MPC framework, which are then input into the MPC framework to achieve seamless integration of language decision-making and control optimization.

[0159] To improve the accuracy of fleet control, based on the above embodiments, the method in this application embodiment further includes:

[0160] Implement the traffic control strategy and monitor the operation information of the multiple fleets in real time;

[0161] If a conflict risk is determined based on the operational information of the multiple fleets, a rollback command is sent to the fleet with the conflict risk, so that the speed of the fleet with the conflict risk is restored to the default value.

[0162] In this embodiment, although the traffic control strategy is already very sophisticated, the real world is still full of unpredictable uncertainties, such as sudden vehicle malfunctions, driver errors, or extreme weather. Therefore, while executing the traffic control strategy, the electronic equipment also monitors the operational information of each vehicle fleet in real time.

[0163] The electronic devices can continuously acquire operational information for each convoy, including real-time location, speed, and heading angle, through Vehicle-to-Everything (V2X) communication or high-frequency GPS data reporting. Furthermore, the electronic devices can employ algorithms such as Extended Kalman Filtering (EDF) to fuse the operational information of each convoy, resulting in smoother and more accurate state estimates and providing a reliable basis for risk assessment.

[0164] The electronic equipment determines whether there is a risk of conflict based on the operational information of each fleet. If the electronic equipment determines that there is a risk of conflict, it sends a rollback command to the fleet with the risk of conflict, restoring the speed of the fleet with the risk of conflict to the default value.

[0165] To improve the accuracy of fleet control, based on the above embodiments, in this embodiment, determining the risk of conflict based on the operational information of each fleet includes:

[0166] For any given fleet, if the distance between the location of the fleet and the location of a candidate fleet is determined to be less than a preset distance threshold based on the location information carried in the operation information of the multiple fleets, then it is determined that there is a risk of conflict between the candidate fleet and the candidate fleet.

[0167] In this embodiment, the electronic device calculates the real-time distance between the preceding and following vehicles on the same road segment and compares it with a preset distance threshold used to represent a safe distance. The electronic device also predicts the arrival time of all intersecting convoys at key nodes and checks whether the time difference is less than a preset isolation time threshold. In addition, the electronic device monitors whether the convoys strictly comply with traffic light instructions to prevent violations such as running red lights.

[0168] Specifically, for any given convoy, the electronic equipment utilizes the road network topology for rapid filtering. For example, it identifies all other convoys traveling on the same road segment as the given convoy by matching road segment IDs. The equipment can also calculate the intersection of their routes to identify other convoys sharing one or more downstream intersections with the given convoy, thus narrowing down the large convoy group into a smaller candidate set with potential interactions. Subsequently, the electronic equipment performs precise processing of the real-time location information of these candidate convoys to determine whether there is a risk of conflict between each candidate convoy and the given convoy.

[0169] Specifically, if the electronic device determines that the distance between the location of the current convoy and the location of the candidate convoy is less than a preset distance threshold, then it determines that there is a risk of conflict between the candidate convoy and the candidate convoy.

[0170] This distance threshold can be derived from a combination of factors, such as: the basic legal minimum following distance, the reaction distance calculated based on the current relative speed and the preset driver reaction time, and the braking distance calculated based on the current speed and the comfort deceleration. The electronic device determines the distance threshold as the sum of the legal minimum following distance, the reaction distance, and the braking distance.

[0171] In addition to calculating the distance between the convoy and candidate convoys, electronic devices can also determine the vehicles' direction of travel and intentions to better mitigate conflict risks. By calculating the difference in the current heading angles of the convoy and candidate convoys, it can be determined whether they are traveling in the same direction, towards each other, or turning to merge. Only when the heading angle difference is less than a preset threshold is it considered that the two vehicles intend to approach each other.

[0172] To further improve accuracy, the electronic equipment will also make simple predictions about the future trajectories of the two convoys based on the current speed and heading, such as predicting the positions of the two convoys within the next 5 seconds, and determining whether these predicted trajectories have spatial overlap.

[0173] The electronic device will only determine that there is a risk of conflict if all three conditions are met simultaneously, or at least one of them is met: the real-time distance between candidate convoys is less than the distance threshold, the heading angle difference is less than the threshold, and the predicted trajectories intersect within a certain period of time.

[0174] In this embodiment, the electronic device performs the fleet control schemes described in the above embodiments at regular intervals and records the traffic control strategy and the corresponding fleet operation information each time, so as to facilitate subsequent auditing and task playback.

[0175] Example, Figure 3 The overall flowchart of fleet control provided in the embodiments of this application is as follows. Figure 3 As shown, the process includes:

[0176] S301: The cycle begins.

[0177] S302: Data loading and constraint generation.

[0178] In this embodiment, the electronic device acquires road network information for a designated area and the driving routes of each vehicle fleet. The electronic device also acquires preset constraints.

[0179] S303: State prediction based on prediction model.

[0180] The electronic equipment uses STFM to determine the speed sequence and time sequence of each convoy based on road network information and the driving route of each convoy.

[0181] S304: Generate traffic control strategies based on strategy generation models.

[0182] The electronic equipment uses a strategy generation model to generate traffic control strategies based on the speed sequence, time sequence, road network information, and preset constraints of each fleet. The traffic control strategies include fleet control strategies that adjust the speed sequence of the fleets and signal control strategies that adjust traffic lights.

[0183] S305: Optimize traffic control strategies based on optimization models.

[0184] The electronic equipment adjusts the speed sequence and time sequence for each fleet according to the fleet management strategy; it also adjusts the signal sequence according to the signal management strategy, whereby the signal sequence includes the green window duration for each preset node; and it uses an optimization model to verify and optimize the speed sequence and signal sequence based on the updated speed sequence, time sequence, road network information, updated signal sequence, and preset constraints for each fleet, generating optimized speed sequence and signal sequence.

[0185] S306: Implement traffic control strategies and monitor them.

[0186] The electronic equipment executes traffic control strategies and monitors the operation information of each fleet in real time. If a conflict risk is identified based on the operation information of each fleet, a rollback command is sent to the fleet with the conflict risk, so that the speed of the fleet with the conflict risk is restored to the default value.

[0187] S307: Enter The cycle is repeated, and S302 is re-executed.

[0188] Based on the above embodiments, the multi-task concurrency and route planning speed control method based on a large model provided in this application can also be applied to a fleet control system, which includes, but is not limited to, a data and constraint input layer, a strategy generation module, a feedback evaluation module, and an online control layer.

[0189] Example, Figure 4 The diagram illustrates the structure of a fleet control system provided in this embodiment. As shown, the fleet control system includes, but is not limited to, a data and constraint input layer 401, a strategy generation module 402, a feedback evaluation module 403, and an online control layer 404. Specifically, the data and constraint input layer 401 receives input road network information, fleet routes, and preset constraints; the strategy generation module 402 can be a reinforcement learning spatiotemporal large-scale model used to learn the speed-signal coupling patterns of multiple fleets and routes, and generate strategies; the feedback evaluation module 403 can be an in-loop simulation agent used to output delay, risk, and ETA confidence intervals for the online control layer to reference; and the online control layer 404 can be an MPC rolling optimization execution model used to jointly optimize speed and signal sequences, generating a Δt non-meeting certificate and a probabilistic ETA.

[0190] This application provides a multi-vehicle, multi-route speed control and simulation method based on a reinforcement learning spatiotemporal large model and MPC rolling optimization. First, a reinforcement learning spatiotemporal large model is used to learn the spatiotemporal coupling relationship between speed, signal, and constraints in a multi-vehicle, multi-route simulation environment, constructing a strategy model with global coordination capabilities. Second, a differentiable simulation agent is used to evaluate the strategy results in real time, providing feedback on performance indicators such as queuing, delay, and green window occupancy. Then, the MPC module performs constrained optimization and dynamic correction of the strategy output in the rolling time domain, achieving adaptive adjustment to complex disturbances. Finally, the safety and punctuality of the control effect are quantified through Δt non-meeting certificates and probabilistic ETA, constructing a reliable, verifiable, and reproducible spatiotemporal coordinated control system for multi-vehicle, multi-route operations.

[0191] This application's embodiments construct a large-scale traffic model with spatiotemporal understanding and policy generation capabilities through reinforcement learning, learning the speed and signal coordination patterns among multiple vehicle fleets and routes. During the online operation phase, the MPC module performs constrained rolling optimization and correction, realizing a reliable closed-loop control system from policy generation to simulation evaluation to optimized execution. This scheme effectively combines the scenario recognition capabilities of the large-scale model with the constraint solving capabilities of the MPC, significantly improving the interpretability and feasibility of multi-vehicle fleet and multi-route collaborative control.

[0192] Compared with existing technologies, the multi-task concurrency and route planning speed control method based on a large model in this application has the following advantages:

[0193] 1. This application constructs a spatiotemporal policy model through reinforcement learning, which can adaptively generate joint speed and signal control strategies in complex traffic networks, and achieve real-time constrained optimization with the help of MPC, significantly improving the safety, punctuality and robustness of multi-vehicle collaborative operation.

[0194] 2. By introducing multi-dimensional weighted rewards (punctuality, delay, risk, handover smoothness, etc.), this application enables the reinforcement learning model to autonomously learn the spatiotemporal dependency between speed and signal during training, thereby enabling the strategy to dynamically balance safety and traffic efficiency and achieve optimal collaborative control of multiple vehicle fleets and nodes.

[0195] 3. This application provides quantitative safety and credibility indicators for each round of control output through Δt non-meeting certificate, probabilistic ETA and risk budget mechanism, realizes an interpretable, traceable and verifiable control closed loop, and ensures the reliability and auditability of traffic operation plan.

[0196] Based on the above embodiments, this application also provides a multi-vehicle fleet control device based on a spatiotemporal large model. Figure 5 A schematic diagram of a multi-vehicle fleet control device based on a spatiotemporal large model is provided for embodiments of this application. The device includes:

[0197] The prediction module 501 is used to acquire road network information of a set area and the driving routes of multiple convoys; using a prediction model, based on the road network information and the driving routes of the multiple convoys, it determines the speed sequence and time sequence of the multiple convoys; the time sequence is used to indicate the predicted time for the corresponding convoy to arrive at at least one preset node;

[0198] The control module 502 is used to generate a traffic control strategy based on the speed sequence of the multiple vehicle fleets, the time sequence, the road network information, and preset constraints using a strategy generation model; wherein the traffic control strategy includes a vehicle fleet control strategy that adjusts the speed sequence of the vehicle fleets, and a signal control strategy that adjusts the traffic lights.

[0199] In one possible implementation, the control module 502 is further configured to adjust the speed sequences and time sequences corresponding to the multiple fleets according to the fleet management strategy; adjust the signal sequences according to the signal management strategy, wherein the signal sequences include the green window duration corresponding to each preset node; and use an optimization model to verify and optimize the speed sequences and signal sequences based on the updated speed sequences, time sequences, road network information, updated signal sequences, and preset constraints, thereby generating optimized speed sequences and signal sequences.

[0200] In one possible implementation, the strategy generation model is a spatiotemporal large model;

[0201] The device further includes:

[0202] Training module 503 is used to acquire training samples, which include sample road network information, sample speed sequences and sample time sequences of multiple vehicle fleets; using the strategy generation model to be trained, a sample traffic control strategy is generated based on the preset constraints, sample road network information, and sample speed sequences and sample time sequences of the multiple vehicle fleets; the sample speed sequences and sample time sequences of the multiple vehicle fleets are adjusted according to the sample traffic control strategy; a simulation model is used to perform simulation evaluation based on the sample road network information and the adjusted sample speed sequences and sample time sequences of the multiple vehicle fleets to determine the sample evaluation results; a group relative strategy optimization (GRPO) algorithm is used to determine the reward function value based on the sample evaluation results, a preset reward function, the adjusted sample speed sequences and sample time sequences of the multiple vehicle fleets, and the parameters of the strategy generation model to be trained are adjusted according to the reward function value.

[0203] In one possible implementation, the training module 503 is further configured to: determine an on-time reward value based on the information regarding whether the convoy arrives on time carried in the simulation results; determine an arrival time deviation based on the simulated arrival time of the convoy carried in the simulation results and the predicted arrival time carried in the sample time series; determine a signal switching penalty value based on the number of signal light changes carried in the sample traffic control strategy; determine a conflict risk penalty value based on the information regarding whether the convoys meet carried in the simulation results; and determine the reward function value based on the on-time reward value, the arrival time deviation, the signal switching penalty value, the conflict risk penalty value, and a preset weight.

[0204] In one possible implementation, the training module 503 is further configured to use the preset constraints, the sample evaluation results, the sample speed sequences and sample time sequences adjusted by the multiple vehicle fleets as inputs to the next iteration of the strategy generation model to be trained.

[0205] In one possible implementation, the control module 502 is further configured to execute the traffic control strategy and monitor the operation information in real time; if a conflict risk is determined based on the operation information, a rollback command is sent to the convoy with the conflict risk, so that the speed of the convoy with the conflict risk is restored to the default value.

[0206] In one possible implementation, the control module 502 is specifically used to determine, for any convoy, if the distance between the location of the convoy and the location of the candidate convoy is less than a preset distance threshold based on the location information carried in the operation information, then determine that there is a conflict risk between the candidate convoy and the candidate convoy.

[0207] Based on the above embodiments, this application also provides an electronic device. Figure 6 This application provides a schematic diagram of an electronic device structure, such as... Figure 6 As shown, it includes: processor 601, communication interface 602, memory 603 and communication bus 604, wherein processor 601, communication interface 602 and memory 603 communicate with each other through communication bus 604.

[0208] The memory 603 stores a computer program that, when executed by the processor 601, causes the processor 601 to perform the steps of any of the large-model-based multi-task concurrency and route planning speed control methods provided in the above embodiments.

[0209] Since the principle of the above-mentioned electronic device in solving the problem is similar to the multi-task concurrency and route planning speed control method based on a large model, the implementation of the above-mentioned electronic device can be found in the implementation examples of the method, and repeated details will not be repeated.

[0210] The communication bus mentioned in the aforementioned electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not indicate that there is only one bus or one type of bus. Communication interface 602 is used for communication between the aforementioned electronic device and other devices. The memory can include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor.

[0211] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0212] Based on the above embodiments, this invention also provides a computer-readable storage medium storing a computer program executable by a processor. When the program runs on the processor, it causes the processor to execute any of the steps of multi-task concurrency and route planning speed control based on a large model as provided in the above embodiments.

[0213] Since the principle of the computer-readable storage medium in solving the problem is similar to the multi-task concurrency and route planning speed control method based on a large model, the implementation of the computer-readable storage medium can be found in the embodiments of the method, and repeated details will not be repeated.

[0214] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0215] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0216] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0217] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0218] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A multi-task concurrency and route planning speed control method based on a large model, characterized in that, The method includes: Obtain road network information and driving routes of multiple vehicle fleets within a designated area; A prediction model is used to determine the speed sequence and time sequence of the multiple convoys based on the road network information and the driving routes of the multiple convoys; the time sequence is used to indicate the predicted time for the corresponding convoy to arrive at at least one preset node; A strategy generation model is used to generate traffic control strategies based on the speed sequences of the multiple vehicle fleets, the time sequences, the road network information, and preset constraints. The traffic control strategies include vehicle fleet control strategies that adjust the speed sequences of the vehicle fleets, and signal control strategies that adjust traffic lights. The method further includes: Based on the fleet management strategy, the speed sequences and time sequences corresponding to the multiple fleets are adjusted; According to the signal control strategy, the signal sequence is adjusted, wherein the signal sequence includes the green window duration corresponding to each preset node; An optimization model is used to verify and optimize the speed sequence and the signal sequence based on the updated speed sequence, time sequence, road network information, updated signal sequence, and preset constraints, thereby generating optimized speed sequence and signal sequence. The strategy generation model is a spatiotemporal large model; The training process of the policy generation model includes: Acquire training samples, which include sample road network information, sample speed sequences of multiple vehicle fleets, and sample time sequences; A model is generated using a strategy to be trained. Based on the preset constraints, the sample road network information, and the sample speed sequences and sample time sequences of the multiple vehicle fleets, a sample traffic control strategy is generated. Based on the traffic control strategy described above, the sample speed sequences and sample time sequences of the multiple vehicle fleets are adjusted. A simulation model is used to perform simulation evaluation based on the sample road network information, the sample speed sequence adjusted by multiple vehicle fleets, and the sample time sequence, and to determine the sample evaluation result. The on-time bonus value is determined based on the information regarding whether the fleet arrived on time, carried in the sample evaluation results. The arrival time deviation is determined based on the simulated arrival time of the convoy carried in the sample evaluation results and the predicted arrival time carried in the sample time series. The signal switching penalty value is determined based on the number of signal light changes carried in the sample traffic control strategy. Based on the information regarding whether the convoys encountered each other, carried in the sample evaluation results, a conflict risk penalty value is determined; The reward function value is determined based on the on-time reward value, the arrival time deviation, the signal switching penalty value, the conflict risk penalty value, and the preset weights. The parameters of the policy generation model to be trained are adjusted based on the reward function value.

2. The method according to claim 1, characterized in that, The method further includes: The preset constraints, the sample evaluation results, the sample speed sequences and sample time sequences adjusted by the multiple vehicle teams are used as inputs to the next iteration of the policy generation model to be trained.

3. The method according to claim 1, characterized in that, The method further includes: Implement the traffic control strategy and monitor the operation information of the multiple vehicle fleets in real time; If a conflict risk is determined based on the operational information, a rollback command is sent to the convoy with the conflict risk, causing the speed of the convoy with the conflict risk to be restored to the default value.

4. The method according to claim 3, characterized in that, The determination of conflict risk based on the operational information includes: For any given fleet, if the distance between the location of the fleet and the location of the candidate fleet is determined to be less than a preset distance threshold based on the location information carried in the operation information, then it is determined that there is a risk of conflict between the candidate fleet and the candidate fleet.

5. A multi-task concurrent and route planning speed control device based on a large model, characterized in that, The device includes: The prediction module is used to acquire road network information of a set area and the driving routes of multiple convoys; using a prediction model, based on the road network information and the driving routes of the multiple convoys, it determines the speed sequence and time sequence of the multiple convoys; the time sequence is used to indicate the predicted time for the corresponding convoy to arrive at at least one preset node; The control module is used to generate traffic control strategies based on the speed sequences of the multiple vehicle fleets, the time sequences, the road network information, and preset constraints using a strategy generation model; wherein the traffic control strategies include vehicle fleet control strategies that adjust the speed sequences of the vehicle fleets, and signal control strategies that adjust traffic lights; The control module is further configured to adjust the speed sequences and time sequences corresponding to the multiple fleets according to the fleet management strategy; adjust the signal sequences according to the signal management strategy, wherein the signal sequences include the green window duration corresponding to each preset node; and use an optimization model to verify and optimize the speed sequences and signal sequences based on the updated speed sequences, time sequences, road network information, updated signal sequences, and preset constraints, thereby generating optimized speed sequences and signal sequences. The strategy generation model is a spatiotemporal large model; The device further includes: The training module is used to acquire training samples, which include sample road network information, sample speed sequences and sample time sequences of multiple vehicle fleets; using a strategy generation model to be trained, a sample traffic control strategy is generated based on the preset constraints, the sample road network information, and the sample speed sequences and sample time sequences of the multiple vehicle fleets; the sample speed sequences and sample time sequences of the multiple vehicle fleets are adjusted according to the sample traffic control strategy; a simulation model is used to perform simulation evaluation based on the sample road network information, the adjusted sample speed sequences of the multiple vehicle fleets, and the sample time sequences, and the sample evaluation result is determined; based on the information carried in the sample evaluation result... The system calculates the on-time reward value based on whether the convoy arrives on time; the arrival time deviation is determined based on the simulated arrival time of the convoy carried in the sample evaluation results and the predicted arrival time carried in the sample time series; the signal switching penalty value is determined based on the number of signal light changes carried in the sample traffic control strategy; the conflict risk penalty value is determined based on the information on whether the convoys meet carried in the sample evaluation results; and the reward function value is determined based on the on-time reward value, the arrival time deviation, the signal switching penalty value, the conflict risk penalty value, and the preset weights. The parameters of the strategy generation model to be trained are then adjusted according to the reward function value.

6. An electronic device, characterized in that, The electronic device includes at least a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the steps of the multi-task concurrency and route planning speed control method based on a large model as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the steps of the multi-task concurrency and route planning speed control method based on a large model as described in any of claims 1-4.