Method for coordinating multi-mobile bodies by reinforcing learning self-adapting regulation ant colony

CN122617084BActive Publication Date: 2026-09-18CIVIL AVIATION UNIV OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202611116927.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-27
Publication Date
2026-09-18
Estimated Expiration
2046-07-27

AI Technical Summary

Technical Problem

自适应因子仅是基于迭代次数的开环时间表(余弦衰减函数),无法实时感知蚁群搜索过程中的种群聚集程度与进化停滞状态

Benefits of technology

[0019] (1) By collecting population aggregation degree and evolutionary stagnation markers in real time through reinforcement learning agents, the pheromone heuristic factor and expectation heuristic factor of ant colony algorithm can be dynamically adjusted. This enables the algorithm to adaptively adjust the balance between global exploration and local development at different search stages, reducing the risk of premature convergence or search stagnation caused by parameter solidification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122617084B_ABST
    Figure CN122617084B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of intelligent optimization and cooperative control, in particular to a multi-mobile body cooperative scheduling method of reinforcement learning adaptive regulation ant colony, comprising: modeling a scheduling scene physical network as a directed graph and obtaining mobile body task data; initializing ant colony algorithm parameters and reinforcement learning intelligent agent; iteratively executing a search process, and outputting an optimal non-interference scheduling scheme after iteration ends; through deep integration of reinforcement learning and ant colony algorithm, the present application realizes adaptive closed-loop regulation and control of parameters, effectively improving efficiency and robustness of multi-mobile body cooperative scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent optimization and cooperative control technology, and in particular to a multi-moving body cooperative scheduling method for ant colony adaptive adjustment using reinforcement learning. Background Technology

[0002] The Multi-Agent Cooperative Scheduling Problem (MACSP) aims to plan interference-free spatiotemporal paths from a starting point to a destination for multiple mobile objects in a shared networked physical environment. It is a typical nondeterministic, polynomial-time difficult (NP-hard) combinatorial optimization problem. Such problems are widely found in scenarios such as airport taxiing, automated terminal vehicle scheduling, mine locomotive transportation, and warehouse robot path planning. With the continuous increase in traffic density, competition for shared resources (nodes, edges) is becoming increasingly fierce, and trajectory interactions and resource interference between mobile objects have become key bottlenecks restricting system efficiency and economy.

[0003] Existing solution methods are mainly divided into two categories: exact operations research methods and metaheuristic algorithms. Exact operations research methods (such as mixed-integer linear programming) can theoretically find the global optimum, but they suffer from state space explosion in large-scale, highly dynamic scenarios, making them difficult to meet real-time scheduling requirements. Metaheuristic algorithms have gained widespread attention due to their efficient search capabilities in discrete spaces. Among them, the Ant Colony Optimization (ACO) algorithm, with its positive feedback mechanism and parallel search characteristics, demonstrates unique advantages in path planning problems.

[0004] For example, patent document CN117852841B discloses a joint airport scheduling method integrating bidirectional particle swarm optimization and multi-strategy ant colony optimization. This method uses a bidirectional particle swarm optimization algorithm to optimize the initial pheromone distribution of the ant colony algorithm and introduces an adaptive factor that decays with the number of iterations to adjust the pheromone concentration, aiming to avoid premature convergence. It also employs a two-stage queue method to handle the actual runway scheduling time. However, this technical solution still has the following shortcomings:

[0005] First, its core control parameters, the pheromone heuristic factor and the expected heuristic factor, are still set to fixed values ​​based on offline experience. The adaptive factor is only based on an open-loop timetable (cosine decay function) of the number of iterations, and cannot perceive the degree of population aggregation and evolutionary stagnation during the ant colony search process in real time. In high-density, highly dynamic scheduling scenarios, fixed parameter configuration can easily lead the algorithm into local optima, and may even induce systemic deadlock by over-enhancing certain congested paths.

[0006] Second, its heuristic functions are mainly based on physical distance and preset time windows, lacking the ability to perceive and intrinsically avoid real-time queuing pressure at nodes (intersections, resource points). When multiple mobile entities gather near the same node, this type of method struggles to proactively avoid high-risk areas at the path generation level, typically requiring passive intervention and relief through external queue sorting or waiting mechanisms, thus reducing the overall smoothness and robustness of the scheduling scheme.

[0007] Third, its objective function mainly focuses on the total time to complete the task, without fully considering the energy consumption differences of the mobile vehicle under different operating states (normal driving, idling and waiting), and it does not provide sufficient support for environmental and economic efficiency in the context of green and low-carbon development.

[0008] In summary, overcoming the limitations of traditional ant colony optimization (ACO) algorithms, such as fixed parameters and lack of real-time state awareness, and deeply integrating adaptive decision-making mechanisms with the discrete graph search advantages of ACO algorithms to design a collaborative scheduling method capable of dynamically sensing the population's evolutionary state and proactively avoiding congestion and deadlock, is a pressing technical challenge in this field. This invention is particularly applicable to networked scheduling scenarios with high density and multiple interference characteristics, such as airports, automated container terminals, and rail transit hubs. Summary of the Invention

[0009] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows:

[0010] This invention provides a reinforcement learning-based adaptive adjustment method for multi-moving body cooperative scheduling of ant colonies, the method comprising the following steps:

[0011] S100, acquire physical access network data of the scheduling scenario and task data of multiple mobile bodies to be scheduled; wherein, the scheduling scenario refers to an engineering scenario composed of a physical access network and multiple mobile bodies moving within the physical access network; the physical access network is modeled as a directed graph, wherein the nodes of the directed graph represent location points and the edges represent passable road segments, and the task data includes at least the starting node, ending node and earliest allowed departure time of each mobile body.

[0012] S200, initialize the pheromone heuristic factor and expectation heuristic factor of the ant colony algorithm, and initialize the reinforcement learning agent, which is associated with a pre-constructed state-action value function table.

[0013] S300, iteratively execute the preset search process until the termination condition is met, and generate the optimal scheduling scheme, which includes the non-interference spatiotemporal path of each moving body from the starting point to the ending point.

[0014] The preset search process includes the following sub-steps:

[0015] S301, the reinforcement learning agent collects the aggregation degree, iteration progress and evolutionary stagnation marker of the population scheduling cost during the current ant colony search process, selects an action according to the state-action value function table, and generates dynamic adjustment instructions for the pheromone heuristic factor and the desired heuristic factor.

[0016] S302, the ant colony algorithm uses the adjusted pheromone heuristic factor and the expected heuristic factor as running parameters to perform path search on the directed graph. Each ant iteratively selects the next node according to the transfer probability that combines the static path distance and the real-time congestion potential of the node, and constructs candidate paths for all the mobile bodies to be scheduled in turn.

[0017] S303, perform spatiotemporal interference constraint verification and comprehensive scheduling cost calculation on the candidate path, generate a reward signal based on the verification result and comprehensive scheduling cost, and update the state-action value function table of the reinforcement learning agent with the reward signal.

[0018] The present invention has at least the following beneficial effects:

[0019] (1) By collecting population aggregation degree and evolutionary stagnation markers in real time through reinforcement learning agents, the pheromone heuristic factor and expectation heuristic factor of ant colony algorithm can be dynamically adjusted. This enables the algorithm to adaptively adjust the balance between global exploration and local development at different search stages, reducing the risk of premature convergence or search stagnation caused by parameter solidification.

[0020] (2) Couple the static path distance and the real-time congestion potential of the node to the transition probability, so that the ants can avoid congested nodes first during the path construction process, reduce path interference and waiting time, thereby improving the feasibility of the scheduling scheme.

[0021] (3) By adopting spatiotemporal interference constraint verification and comprehensive scheduling cost calculation, and combining the reward signal to update the state-action value function table, the agent can be guided to gradually optimize the parameter adjustment strategy and generate a scheduling scheme that meets the safety interval and has low total running time and total energy consumption.

[0022] (4) The optimal scheduling scheme output after the iteration terminates includes the non-interference spatiotemporal path of each mobile body from the starting point to the end point, which is applicable to engineering scenarios such as airports, automated docks, and warehouse robots that require multi-mobile body collaborative scheduling.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 A flowchart of a multi-moving body cooperative scheduling method for reinforcement learning adaptive adjustment ant colonies provided in an embodiment of the present invention;

[0026] Figure 2 This is a schematic diagram of the physical network of the airport surface. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0029] It should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of these steps can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the steps can be rearranged. A process can be terminated when its operation is complete, but it may also have additional steps not included in the figures. A process can correspond to a method, function, procedure, subroutine, subroutine, etc.

[0030] This invention provides a reinforcement learning-based adaptive adjustment method for multi-movement cooperative scheduling of ant colonies, such as... Figure 1 As shown, the method may include the following steps:

[0031] S100: Obtain physical access network data for the scheduling scenario and task data for multiple mobile entities to be scheduled.

[0032] The nodes of the directed graph represent location points, and the edges represent passable road segments. The task data includes at least the start node, end node, and earliest allowed departure time of each mobile entity. The physical network of the scheduling scenario is modeled as a directed graph, and task data of multiple mobile entities to be scheduled are obtained.

[0033] The scheduling scenario refers to an engineering scenario consisting of a physical traffic network (including nodes and road segments) and multiple mobile bodies moving within that physical traffic network. The nodes and edges of the physical traffic network correspond to location points and passable road segments in the actual engineering environment, respectively. Illustrative scenarios include, but are not limited to: airport taxiing scenarios, automated container terminal guided vehicle scenarios, rail transit hub train route scenarios, and warehouse robot scenarios.

[0034] A mobile entity refers to an entity that moves autonomously or under controlled conditions within the physical network of access, such as aircraft, vehicles, trains, and robots. Each mobile entity has an independent starting point, ending point, and motion constraints.

[0035] The physical access network is modeled as a directed graph, where nodes represent locations within the network, including but not limited to intersections, parking spaces, work sites, path endpoints, or hub nodes. Edges represent passable road segments (such as taxiways, lanes, or tracks) between nodes. Each edge is associated with a physical length used to calculate the travel time of a moving vehicle across that edge. The directed graph supports directional constraints on road segments, such as one-way traffic segments or taxiways that allow passage in a specific direction.

[0036] During the initialization phase, based on the directed graph and the physical length of each edge, shortest path algorithms such as Dijkstra's algorithm or Floyd-Warshall's algorithm are used to pre-calculate and cache the shortest path distances between all node pairs. These shortest path distances are then used to calculate the static spatial guidance term in subsequent heuristic functions.

[0037] Task data refers to information describing the scheduling operations required for each mobile entity, including at least the entity's start node, destination node, and earliest allowed departure time. In scenarios requiring time window constraints, the task data may also include a latest arrival time constraint at the destination. Task data is obtained by reading from the information management platform of the scheduling system, such as airport collaborative decision-making systems, port operation control systems, train scheduling systems, or pre-set flight / task schedules.

[0038] In one illustrative embodiment, the scheduling scenario is an airport surface taxiing scheduling system, and the moving body is an aircraft. Task data is obtained from the flight planning management system and includes information such as flight number, aircraft type, takeoff and landing times, parking positions, and runway allocation.

[0039] The airport surface taxiing scheduling problem aims to plan a conflict-free spatiotemporal trajectory for a group of departing and arriving flights from the starting point (parking stand or runway exit) to the destination (runway entrance or parking stand), and to achieve comprehensive optimization of total taxiing time and fuel consumption while satisfying constraints on surface topology, time window and safety interval.

[0040] (1) Scene network model

[0041] An illustrative airport surface physical network can be as follows Figure 2 As shown. The airport surface physical network is abstracted as a directed graph G=(V,E), where:

[0042] The node set V represents key physical locations on the airport surface, including but not limited to: aircraft standpoints, taxiway intersections, holding points, runway entrances, and runway exits. Each node corresponds to a location point in the diagram.

[0043] The edge set E represents the gliding segment connecting the nodes. Each edge in E has a physical length (in meters) and supports bidirectional traffic.

[0044] During taxiing, the aircraft moves at a constant speed, and turning losses are converted into an increase in edge weights. The aircraft can wait at any node, and fuel consumption at idle is calculated as a constant value per unit time.

[0045] (2) Task data

[0046] Let F be the set of aircraft to be scheduled. For each aircraft i∈F, its mission data includes at least:

[0047] Starting node o i ∈V, where the starting node is either a parking position or a runway exit;

[0048] End node d i ∈V, where the endpoint node is either the runway entrance or the parking position;

[0049] Earliest permitted departure time T i early .

[0050] For departing aircraft, an additional constraint is imposed: T i arr ≤D i T i arr D is the actual time when aircraft i arrives at the destination node. i The preset latest arrival time;

[0051] Aircraft model information is used to determine taxiing speed and fuel consumption rate;

[0052] All flight schedules within the planning period are known in advance, including launch time, aircraft type, origin and destination.

[0053] (3) Path representation and temporal evolution

[0054] Let the taxiing path of aircraft i be represented as a sequence of nodes (v i,0 ,v i,1 ,…,v i,k ,…,v i,N(i) ), where the starting point v i,0 =o i , destination v i,N(i) =d i And there are directed edges between adjacent nodes, where N(i) is the number of edges traversed by aircraft i, which is the node sequence length minus 1. i,k Let i be the (k+1)th node in the node sequence, where k ranges from 1 to N(i). Let e ​​be the kth edge of aircraft i. i,k =(v i,k-1 ,v i,k The physical length of this edge is L. i,k .

[0055] The taxiing speed of aircraft i is a constant value V i (Unit: meters per second), then through edge e i,k The required travel time is:

[0056] t i,k travel =L i,k / V i .

[0057] Define aircraft i at node v i,k-1 The waiting time at point v is w. i,k (Unit: seconds), where w i,1 This indicates that at the starting point v i,0 Waiting time at the location.

[0058] Let T be the earliest permitted departure time for aircraft i. i early Then, the time T when entering the first edge... i,1 in It is given by the following formula:

[0059] T i,1 in =T i early +w i,1 ;

[0060] For the k-th edge (k≥1), aircraft i leaves edge e. i,k The moment T i,k out(i.e., reaching node v) i,k (moment) and entering the edge e i,k The moment T i,k in (i.e., reaching node v) i,k-1 The relationship (and the time when the waiting ends and the process begins to enter that side) is as follows:

[0061] T i,k out =T i,k in +t i,k travel ;

[0062] For the k-th edge, the time T when aircraft i enters the next edge. i,k+1 in The departure time of the current edge and the waiting time w of the next node are determined by the current edge's departure time. i,k+1 Decide:

[0063] T i,k+1 in =T i,k out +w i,k+1 .

[0064] Finally, the moment T is when aircraft i reaches the destination, that is, leaves the last edge. i arr For T i,N(i) out .

[0065] Specifically, when k=1, T i,1 in The earliest departure time and the starting point waiting time have been determined.

[0066] (4) Constraints

[0067] The scheduling scheme must meet the following constraints:

[0068] Path connectivity: The path of each aircraft must be a continuous path from the origin to the destination on graph G.

[0069] Node capacity constraint: At any given time, each node can be occupied by at most one aircraft.

[0070] Safety separation constraints: A minimum safe time interval must be met between any two aircraft, specifically including the following three types of interference constraints:

[0071] Head-to-head interference: When two aircraft are traveling towards each other on the same road segment, their time windows of occupation do not overlap.

[0072] Tail-end interference: When two aircraft are traveling in the same direction on the same road segment, the time when the later aircraft enters the road segment is no earlier than (i.e., greater than or equal to) the sum of the time when the earlier aircraft leaves the road segment and the minimum safe time interval.

[0073] Cross-interference: When two aircraft pass through the same cross node one after the other, the time difference between the two aircraft passing through the cross node is not less than the minimum safe time interval.

[0074] The runway area is managed separately by air traffic control (ATC) and is not within the scope of taxiway scheduling in this embodiment.

[0075] (5) Optimization Objective

[0076] To balance operational efficiency and environmental economy, the overall scheduling cost J is the weighted sum of the total taxiing time and total fuel consumption of all aircraft, i.e.: .in, and These are weighting coefficients, satisfying the condition that their sum equals 1 and all are greater than or equal to 0. In one illustrative embodiment, = =0.5. Fuel i This represents the total fuel consumption of aircraft i.

[0077] The fuel consumption of aircraft i is differentiated based on its operating state: taxiing (running) and idling (standing still). Total fuel consumption is calculated using the following formula:

[0078] .

[0079] Where, μ taxi The fuel consumption rate per unit time (unit: kg / s) of aircraft i in taxiing state is pre-calibrated using aircraft model parameters;

[0080] μ wait The fuel consumption rate per unit time of aircraft i in idling is also pre-calibrated through aircraft type parameters;

[0081] t i taxi The total travel time of aircraft i across all road segments;

[0082] t i wait This represents the total waiting time for aircraft i across all road segments.

[0083] Where, μ taxi >μ wait , t i taxi +t i wait =Ti arr -T i early .

[0084] The goal of scheduling optimization is to minimize J, that is, to find the scheduling scheme that minimizes the weighted total cost.

[0085] S200, initialize the pheromone heuristic factor and expectation heuristic factor of the ant colony algorithm, and initialize the reinforcement learning agent, which is associated with a pre-constructed state-action value function table.

[0086] In one illustrative embodiment, the initialization process is as follows:

[0087] (1) Ant colony algorithm parameter initialization

[0088] Set the initial values ​​for the pheromone heuristic factor α and the expected heuristic factor β in the ant colony algorithm. To ensure sufficient exploration capability in the early stages of the search, the initial value of α is α0. init With the initial value of β init It can be preset to be equal and moderate, for example, α init =1.0, β init =1.0. At the same time, hyperparameters such as pheromone evaporation factor ρ, pheromone intensity constant G, ant number M, and maximum number of iterations Gmax are set. All of these parameters are pre-configured according to the scale of the scheduling scenario.

[0089] In the ant colony algorithm used in this invention, each ant corresponds to an independent scheduling scheme search agent. The complete walking process of an ant on the directed graph is the process of sequentially searching and constructing a set of candidate paths for all scheduled moving objects. Specifically:

[0090] Starting from the origin of the first moving object, the ant selects the next node according to the transition probability and gradually constructs a path from the origin to the destination. After completing the path of the current moving object, the ant moves to the next moving object and repeats the above process until a path is generated for all moving objects.

[0091] A complete set of candidate paths (covering all moving vehicles) constitutes a scheduling scheme.

[0092] Multiple ants search in parallel, generating multiple different scheduling schemes. Through the deposition and volatilization mechanism of pheromones, subsequent searches are guided to gradually converge toward the best scheme.

[0093] The pheromone concentration update employs a max-min ant system (MMAS) strategy: pheromone deposition is only allowed on the optimal candidate path in the current iteration, and the pheromone concentration on each path segment is limited to a preset lower threshold τ. min With upper limit threshold τ maxBetween. The pheromone concentration is updated according to the following formula:

[0094] τ uv (t+1) = (1-ρ)×τ uv (t) + Δτ ij best ;

[0095] Where, τ uv (t) represents the pheromone concentration on edge (u, v) after the t-th iteration, where u and v are node indices in the directed graph, and edge (u, v) represents the directed path from node u to node v. τ ij (t+1) represents the updated pheromone concentration, Δτ ij best Let Δτ be the pheromone increment deposited on edge (i,j) by the optimal scheduling scheme in the current iteration, defined as: if edge (i,j) belongs to the optimal candidate path in the current iteration, Δτ ij best Equal to G / J best Otherwise, △τ ij best The value is equal to 0. J best The weighted total cost is the comprehensive optimization objective value corresponding to the optimal scheduling scheme in the current iteration.

[0096] After the update is completed, the pheromone concentration on each road segment is limited: the updated pheromone concentration is compared with the upper limit threshold τ. max and lower limit threshold τ min Compare, if the concentration is greater than τ max Then set it to τ max If the concentration is less than τ min Then set it to τ min If the concentration is at τ min With τ max Between these, it remains unchanged.

[0097] Upper limit threshold τ max It is usually set as the initial pheromone value multiplied by a large coefficient, such as 10 times the initial pheromone concentration, or calculated based on the expected cost of the optimal solution, such as τ. max =1 / (ρ·J opt ), where J opt This represents the weighted total cost of the currently known optimal scheduling scheme. The lower bound threshold τ... min Let it be τ max / a, where a is a constant greater than 1, typically a = 2·|V|, where |V| is the number of nodes, to ensure that each path retains a certain exploration probability. In an illustrative embodiment, this can be pre-tuned offline according to the scale of the scheduling scenario.

[0098] The above formula for updating pheromone concentration reflects the core mechanism of the Max-Min Ant System (MMAS): the volatile term (1-ρ) gradually decays the pheromone on non-optimal paths, weakening their attractiveness; simultaneously, it only applies the path quality (comprehensive scheduling cost J) to the optimal path in the current iteration. best The pheromone increment Δτ is inversely proportional to the pheromone level. ij best This fosters an elite-oriented approach, strengthening the guiding role of high-quality solutions and thus significantly improving convergence efficiency while maintaining a certain level of exploratory capability. Furthermore, limiting the pheromone concentration to [τ] min ,τ max Within the interval: lower limit τ min Ensure that each path retains a probability of being explored to prevent the search from completely stalling; upper limit τ max To prevent premature convergence due to excessive pheromone accumulation on a particular path, this strategy only allows the optimal ant to release pheromones, enhancing the algorithm's ability to converge towards the global optimum while maintaining continuous exploration potential through concentration limiting.

[0099] (2) Initialization of reinforcement learning agent

[0100] The reinforcement learning agent employs the Q-Learning algorithm, the core of which is a state-action value function table (Q-table) used to store the estimated cumulative reward for taking different actions in different states, i.e., the Q-value. Q(s,a) represents the expected value of the discounted cumulative reward that can be obtained after performing action a in state s according to the optimal policy.

[0101] The agent receives an immediate reward, i.e., a reward signal r, after the current iteration (step t). t+1 The reward signal is a composite reward function, comprising at least: a performance reward term reflecting the improvement in scheduling cost of the current scheduling scheme compared to the historical best scheme; a feasibility reward term for incentivizing the generation of interference-free scheduling schemes; and a penalty term for suppressing global path interlocking; wherein the weight of the penalty term is greater than the weights of the performance reward term and the feasibility reward term. The formula is expressed as:

[0102] r t+1 =λ1·(J t best -J t+1 best ) / J t best +λ2·I feasible -λ3·I deadlock .

[0103] Among them, J t best and J t+1 bestThese are the optimal scheduling costs (overall optimization objective values) for generation t and generation (t+1), respectively; the optimal scheduling cost is the overall scheduling cost J (the weighted sum of the total running time and total energy consumption of all mobile entities), and the smaller the value, the better the scheduling scheme. feasible This is a feasibility indicator function, which takes a value of 1 when the current scheduling scheme satisfies all spatiotemporal interference constraints, and a value of 0 otherwise, constituting a feasibility reward term; I deadlock This is a deadlock indicator function. It takes a value of 1 when a global path interlock (scheduling deadlock) is detected, and a value of 0 otherwise, which constitutes a penalty.

[0104] Global path interlocking refers to a directed graph containing a non-empty set of mobile entities. For each mobile entity in this set, there exists another mobile entity in the set that is waiting for that other mobile entity to release the node or road segment it currently occupies or is about to occupy, thus forming a circular dependency. This prevents all mobile entities in the set from continuing to move towards their respective destinations; in other words, a circular waiting chain exists. Interference-free scheduling schemes refer to scheduling schemes where the planned paths of all mobile entities satisfy node capacity constraints and safety interval constraints.

[0105] The weighting coefficients λ1, λ2, and λ3 are positive numbers, and are assigned according to a hierarchical principle of prioritizing safety, followed by feasibility, and finally performance optimization.

[0106] The feasibility reward weight is greater than the performance reward weight: take λ2>λ1, so that the positive reward of obtaining a conflict-free feasible scheduling scheme is higher than the positive reward of a single performance optimization, thereby driving the agent to prioritize searching for feasible schemes that meet the constraints.

[0107] The deadlock penalty weight is much greater than the feasibility reward weight: λ3 is much greater than λ2. The penalty imposed by a single global path interlock (scheduling deadlock) is sufficient to offset the positive benefits of multiple feasible solutions, effectively avoiding scheduling deadlock from the perspective of reward and punishment mechanism.

[0108] In an illustrative embodiment, the initial settings are λ1=1, λ2=10, and λ3=50. The deadlock penalty coefficient λ3 is determined through offline pre-experiments: the algorithm is run within the parameter set {20, 50, 100}, using minimizing the total aircraft waiting time at the airport as the selection criterion, and the optimal value is ultimately determined to be λ3=50. The above values ​​are only examples and can be adjusted according to the scale of the scenario in actual applications.

[0109] In this invention, the Q-value (state-action value function) is iteratively updated by the agent through interaction with the environment (receiving immediate rewards) and using the Bellman equation to gradually approach the optimal value. Specifically, the update rule for the Q-value is as follows:

[0110] .

[0111] Among them, Qt (s) t a t ) represents the state s at step t. t Select action a below t The current Q value; Q t+1 (s) t a t ) represents the updated Q value; η is the learning rate, controlling the step size of each update, which can be set to 0.1 in an illustrative embodiment; for example, 0.1; γ is the discount factor, weighing the importance of current and future rewards, with a value range of [0,1); in an illustrative embodiment, it can be set to 0.9; r t+1 The immediate reward obtained after performing an action; For the next state s t+1 The maximum Q value among all possible actions represents an estimate of the optimal future reward.

[0112] Through repeated iterations, the Q-table gradually converges, enabling the intelligent agent to select the action that maximizes the expected cumulative reward based on the current state.

[0113] The rows of the Q-table correspond to the discretized environment states, and the columns correspond to the actions that the agent can choose. During initialization, all entries in the Q-table are set to zero, indicating that the agent has no preference for any state-action pairs at the initial moment, that is, it has no prior knowledge and learns and updates itself entirely through reward signals obtained from subsequent interactions with the environment.

[0114] In each iteration step, the agent's input is the currently collected state information, which is discretized and mapped to a state index. The agent queries the Q-table and selects an action using an ε-greedy strategy, and the output is the parameter adjustment instruction corresponding to the selected action.

[0115] The improvement of the reward signal in this invention lies in the construction of a composite reward function that includes a performance reward, a feasibility reward, and a deadlock penalty, and the adoption of a safety-first hierarchical weight configuration (the penalty weight is greater than the sum of the feasibility reward and the performance reward). Compared with existing technologies that treat safety conditions as hard constraints or linear weights of single scalar rewards, this application actively guides the agent to avoid global path interlocks through a dedicated penalty term and alleviates the sparse reward problem through a feasibility reward term. This enables reinforcement learning to simultaneously optimize operational efficiency and energy consumption while ensuring scheduling safety, achieving an organic unity of multi-objective collaboration and safety priority.

[0116] (3) Definition of state-action space

[0117] The state space is composed of a three-dimensional discretization of the population scheduling cost's clustering degree, iteration progress, and evolutionary stagnation flag. The state information collected by the agent includes the specific values ​​of the population scheduling cost's clustering degree, iteration progress, and evolutionary stagnation flag at the current moment.

[0118] The clustering degree of population scheduling cost is quantified by the coefficient of variation of the comprehensive scheduling cost of all candidate paths constructed by all ants in the current generation. The comprehensive scheduling cost (i.e., the path cost of each ant) ​​refers to the weighted total cost corresponding to the complete scheduling scheme constructed by that ant, specifically including the weighted sum of total gliding time and total fuel consumption. The calculation formula is as follows:

[0119] I div (t) = σ(J) t ) / [μ(J t )+ε].

[0120] Among them, I div (t) represents the clustering degree of the population scheduling cost in generation t, i.e., the current generation, σ(J t ) represents the standard deviation of the path costs of all ants in the current generation, μ(J) t The mean of the path costs of all ants in the current generation is denoted as ε, which is a preset positive minimum value to avoid the denominator being zero. When the path costs of all ants are exactly the same, the mean is zero. In an illustrative embodiment, it can be 10. -5 .

[0121] A higher clustering value for the population scheduling cost indicates a more dispersed distribution of solutions among the ants in the population, suggesting the search is in a fully explored state; a lower value indicates highly homogeneous solutions, with the population tending to cluster. When the clustering value of the population scheduling cost is less than a preset clustering threshold, the population is determined to have entered a clustered state, warning of premature convergence risk. At this point, a parameter adjustment mechanism, such as a random exploration mode or a heuristic enhancement mode, needs to be triggered to restore population diversity. The preset clustering threshold can be set based on the needs of the actual scheduling scenario. In an illustrative embodiment, this threshold can be 0.05. Iteration Progress I stage This is the ratio of the current iteration count to the preset maximum iteration count, used to distinguish between the early exploration phase and the later development phase of the search. When the iteration progress is less than the preset phase boundary threshold, it is determined to be in the exploration phase, where the ant colony algorithm focuses on global exploration; when the iteration progress is greater than or equal to the preset phase boundary threshold, it is determined to be in the development phase, where the ant colony algorithm focuses on local development. The preset phase boundary threshold can be set based on the needs of the actual scheduling scenario. In an illustrative embodiment, this threshold can be 0.5.

[0122] Evolutionary Stagnation Marker I sym This flag indicates whether the global optimal solution has improved within a preset number of consecutive algebras. If the global optimal solution has not improved within a preset number of algebras, such as three consecutive algebras, then the flag is activated, i.e., I. sym The first preset value is 1; otherwise, I symThis is a second preset value, such as 0. When this flag is active, it triggers parameter adjustment mechanisms, such as enabling random exploration mode or heuristic enhancement mode, to help the algorithm escape local optima.

[0123] The three dimensions (population aggregation, iteration progress, and evolutionary stagnation indicator) are each discretized into a finite number of intervals by their respective preset thresholds. Then, the discrete intervals of each dimension are combined to form a finite number of discrete states, which constitute the state space of the reinforcement learning agent.

[0124] In one illustrative embodiment, each dimension can be discretized into two intervals, which are combined to obtain a discrete state set corresponding to the total number of states; the number of rows in the Q table is equal to the total number of states, and the number of columns is equal to the number of actions in the action space, so as to ensure that the agent can complete policy learning within a finite number of iterations.

[0125] Furthermore, unlike the traditional MMAS strategy where the volatile factor ρ is set to a fixed constant, this invention proposes a mechanism for dynamically adjusting the pheromone volatile factor. Specifically, the volatile factor ρ(t) at the t-th iteration is adjusted according to the population aggregation degree I of the current iteration. div (t) Dynamic adjustment. The specific mapping rule is as follows:

[0126] When I div When (t) is less than the preset aggregation threshold, the current population solution distribution is considered to have good diversity and no search stagnation has occurred. At this time, a smaller pheromone evaporation factor, such as 0.05, is used to slow down the decay rate of historical pheromones, maintain the guiding stability of the current high-quality solution direction, and promote smooth population convergence. When I div When (t) exceeds a preset aggregation threshold, the population solution is deemed highly homogeneous, posing a risk of premature convergence and getting trapped in a local optimum. At this point, a stalemate-breaking mechanism is triggered, employing a larger pheromone evaporation factor, such as 0.3, to accelerate the decay of historical pheromones. Combined with the pheromone lower bound constraint mechanism of the Max-Min Ant System (MMAS), the larger evaporation factor can quickly smooth out the pheromone concentration difference between the current locally optimal path and unexplored paths, thereby significantly weakening the guiding effect of the current suboptimal path and forcing the ant colony to jump out of the local optimum region and re-explore other feasible regions in the solution space.

[0127] This single-threshold-based state-triggered mechanism eliminates complex continuous parameter mapping, enabling the pheromone update process to adaptively adjust according to the discrete search state of the population. It maintains the memory effect of historical pheromones during normal population evolution and quickly breaks pheromone monopolies when premature convergence risk is detected. This achieves a balance between global exploration and local exploitation with lower computational overhead, thereby improving the algorithm's robustness.

[0128] Action Space: Contains multiple preset search patterns, which serve as actions that the agent can choose. Each preset search pattern corresponds to a set of preset pheromone heuristic factor increment values ​​and desired heuristic factor increment values. These increment values ​​are the dynamic adjustment instructions output by the agent.

[0129] Specifically, the preset search modes include:

[0130] Random exploration mode: This mode is used to simultaneously reduce the values ​​of the pheromone heuristic factor and the expected heuristic factor, so that the transition probability tends to be evenly distributed, thereby enhancing the global random exploration capability and helping the algorithm escape local optima.

[0131] Heuristic Enhancement Mode: This mode reduces the value of the pheromone heuristic factor and increases the value of the expected heuristic factor to weaken the dependence on historical pheromones, enhance the impact of real-time congestion potential on path selection, and guide the vehicle to avoid highly congested nodes. Path Enhancement Mode: This mode simultaneously increases the values ​​of both the pheromone heuristic factor and the expected heuristic factor to strengthen the guiding role of historically optimal paths in subsequent searches and accelerate algorithm convergence.

[0132] Parameter reset mode: used to restore the pheromone heuristic factor and the expected heuristic factor to their initial values, in order to eliminate parameter drift in long-term iterations and maintain the balance between exploration and development.

[0133] The specific values ​​of the pheromone heuristic factor increment and the expected heuristic factor increment for each mode can be determined by conducting grid pre-experiments on typical scenarios, such as performing 10 to 50 grid traversals, and offline calibration. The calibration criterion is to maximize the average improvement rate of the objective function.

[0134] In an illustrative embodiment, the pheromone heuristic factor increment and expected heuristic factor increment for the random exploration mode are -0.5 and -1.0, respectively; the pheromone heuristic factor increment and expected heuristic factor increment for the heuristic enhancement mode are -0.2 and +0.8, respectively; and the pheromone heuristic factor increment and expected heuristic factor increment for the path enhancement mode are +0.5 and -0.5, respectively. The parameter reset mode restores the parameters to their initial baseline values, independent of incremental accumulation.

[0135] In one illustrative embodiment, the initial reference value of α can be set to 1.0, and the initial reference value of β can be set to 2.0.

[0136] The upper and lower limits of the parameter value range are determined as follows: based on the sensitivity analysis of the parameters to the algorithm's convergence and exploration capabilities, combined with offline optimization experiments in typical scheduling scenarios, boundary values ​​that can ensure stable operation of the algorithm without parameter drift are selected. Empirical rules indicate that: if the value of α is too small, it will weaken the guiding role of historical pheromones; if the value is too large, it will cause the algorithm to converge prematurely. If the value of β is too large, it will make path selection too sensitive to instantaneous congestion changes; if the value is too small, it will ignore real-time heuristic information. Accordingly, in this embodiment, the value range of α is [0.1, 5.0], and the value range of β is [0.0, 8.0].

[0137] The agent uses an ε-greedy strategy to select actions from the Q-table. The selection is based on the evolutionary stagnation flag I in the current state. sym Iteration progress I stage Clustering degree I with population scheduling cost div In one illustrative embodiment, the following rules can be set:

[0138] If I sym It is the first preset value and I div If the population density is less than the preset aggregation threshold, the population is determined to be highly homogeneous, and a random exploration mode is triggered.

[0139] If I sym The first preset value, but I div If the aggregation degree exceeds the preset threshold, the population is determined to maintain a certain level of diversity, triggering a heuristic enhancement mode.

[0140] If I stage Greater than the preset stage boundary threshold and I sym If the second preset value is used, it is determined that the population solution distribution is sufficiently dispersed, triggering the path reinforcement mode;

[0141] If I sym It is the second preset value and I div If the population density exceeds the preset aggregation threshold, it is determined that the population is within the normal distribution range, triggering the parameter reset mode.

[0142] Parameter updates follow a saturation truncation rule: the updated parameter value equals the original parameter value plus the increment corresponding to the selected action, and then a truncation function is used to constrain the result between a preset lower and upper limit of the parameter. That is:

[0143] θ t+1 =Clip(θ t +Δ θ (a k ),θ min θ max ),θ∈{α,β}.

[0144] Where, θ tLet θ be the parameter value (α or β) for the t-th iteration (or t-th generation). t+1 This is the updated parameter value. Δ θ (a k Select action a for the agent k Then, the increment value corresponding to the parameter, a positive number indicates an increase, and a negative number indicates a decrease. k This represents the k-th preset search mode in the action space, such as random exploration, heuristic enhancement, path reinforcement, or parameter reset. Clip(·) is the truncation function, θ min and θ max These are the lower and upper limits of the parameters, which can be calibrated offline based on the scene scale or set based on empirical rules. This mechanism prevents parameters from drifting indefinitely and ensures algorithm stability.

[0145] (4) The relationship between intelligent agents and ant colony algorithms

[0146] The reinforcement learning agent and the ant colony algorithm are linked in the following way: After each generation of ant colony search, the agent queries the Q-table based on the current population state, selects an action, and outputs the corresponding parameter adjustment instructions; the ant colony algorithm receives the adjusted α and β values ​​for the next generation of path search. The agent's reward signal is generated based on the quality of the scheduling scheme (a weighted sum of total coasting time and fuel consumption) and whether the constraints are met, and the Q-table is updated accordingly.

[0147] Under this association mechanism, the reinforcement learning agent dynamically adjusts the core parameters of the ant colony algorithm by sensing the population state in real time, enabling the algorithm to adaptively balance global exploration and local exploitation at different search stages. The technical effects of this association mechanism include:

[0148] This enables the ant colony algorithm to adjust its search behavior in real time according to the population aggregation degree and evolutionary stagnation state, thus alleviating the search stagnation or premature convergence problem caused by fixed parameters.

[0149] By guiding the agent to gradually optimize the parameter adjustment strategy through reward signals, the probability of generating a scheduling scheme that is interference-free and has a low overall cost is increased.

[0150] In various multi-mobile-body cooperative scheduling scenarios (including high-density network environments), it can stably output feasible solutions that meet safety constraints and achieve a dynamic balance between global exploration and local development.

[0151] S300: Iteratively execute the preset search process until the termination condition is met to generate the optimal scheduling scheme; and at the end of each iteration, update the pheromone concentration on the directed graph based on the optimal scheduling scheme of the current generation.

[0152] The optimal scheduling scheme includes a non-interference spatiotemporal path for each mobile body from the starting point to the ending point. The non-interference spatiotemporal path means that when a mobile body runs along the planned path, it meets the preset safety interval constraint with other mobile bodies. The safety interval constraint includes node capacity constraint and safety interval constraint. The node capacity constraint means that each node can be occupied by at most one mobile body at the same time. The safety interval constraint means that the time or space interval of interference types such as head-on collision, tail-on collision, and intersection is not less than a threshold.

[0153] At the end of each iteration, the pheromone concentration on the directed graph can be updated based on the aforementioned pheromone update rule, namely the MMAS strategy and its amplitude limiting process.

[0154] Furthermore, the preset search process includes the following sub-steps:

[0155] S301, the reinforcement learning agent collects the aggregation degree, iteration progress and evolutionary stagnation marker of the population scheduling cost during the current ant colony search process, selects an action according to the state-action value function table, and generates dynamic adjustment instructions for the pheromone heuristic factor and the desired heuristic factor.

[0156] The agent, based on the collected state information (clustering degree of population scheduling cost, iteration progress, and evolutionary stagnation indicator of the global optimal solution), queries a pre-constructed state-action value function table (Q-table) and selects an action using an ε-greedy strategy: randomly selecting an action with probability ε, and selecting the action with the largest Q-value in the current state with probability 1-ε. The selected action corresponds to a set of preset pheromone heuristic factor increment values ​​Δα and expected heuristic factor increment values ​​Δβ. The agent generates dynamic adjustment instructions based on the selected action, instructing the current pheromone heuristic factor α and expected heuristic factor β to be updated to α+Δα and β+Δβ, respectively.

[0157] S302, the ant colony algorithm uses the adjusted pheromone heuristic factor and the expected heuristic factor as operating parameters to perform path search on the directed graph. Each ant iteratively selects the next node based on a transition probability that integrates the static path distance and the node's real-time congestion potential, constructing candidate paths for all scheduled moving bodies in sequence.

[0158] Specifically, the transition probability P of ant x (x takes values ​​from 1 to M) choosing the next node q at node z. zq (t) is:

[0159] , q∈N(z,x).

[0160] Where, τ zq (t) represents the pheromone concentration on edge (z,q) at the t-th iteration, τ zs(t) represents the pheromone concentration on edge (z, s) at the t-th iteration, and N(z, x) is the set of reachable nodes that ant x can choose from at node z. η zq (t) represents the heuristic function value on edge (z,q) at the t-th iteration, and η zs Let η(t) be the heuristic function value on edge (z,s) at iteration t. The heuristic function includes a static spatial guidance term and a dynamic congestion penalty term; the static spatial guidance term is calculated based on the physical length of the road segment and the shortest distance from the node to the target point, while the dynamic congestion penalty term is calculated based on the real-time congestion potential energy of the node. The formula for calculating the heuristic function is: η zq (t) = [1 / (d)] zq +D short (q, d) p ))]·E q (t).

[0161] Where, d zq Let D be the physical length of the road segment between node z and node q. short (q, d) x Let d be the target node from node q to ant x. x The shortest glide distance (based on physical length) is pre-calculated and cached during initialization using Dijkstra's algorithm, and is directly looked up in the table at runtime; E q (t) represents the real-time congestion potential energy of node q in the t-th iteration, with a range of (0,1], reflecting the degree of congestion of the node.

[0162] In this embodiment, the real-time congestion potential energy of a node is constructed as follows: Based on the number of moving objects waiting at each node in the current iteration or the cumulative waiting time step, the cumulative waiting potential energy of each node is calculated; after normalization and exponential transformation, the cumulative waiting potential energy is used as the real-time congestion potential energy of that node. Specifically, the formula for calculating the real-time congestion potential energy of a node is: E q (t) = exp(-h·W q (t) / W max ).

[0163] Among them, W q (t) represents the accumulated waiting potential energy of node q in the t-th iteration, defined as the cumulative waiting time steps or the number of waiting moving bodies for all aircraft at node q in the current iteration. exp() represents the natural exponential function. h is the congestion sensitivity factor, a preset positive number used to control the decay rate of the exponential penalty. W max As a normalization benchmark, the maximum accumulated waiting potential energy in the current iteration or a preset constant can be used. When the node is open (W) q (t)≈0), penalty term E q(t) approaches 1, and the heuristic function degenerates into a purely static path distance-oriented function; when nodes are congested (W q (t) is much greater than 0), penalty term E q When (t) approaches 0, ants tend to choose an alternate route.

[0164] In one illustrative embodiment, h=3, where the heuristic attraction of the most congested node decreases to approximately 4.98%, providing sufficient repulsion while preventing the heuristic value from dropping to zero.

[0165] Each ant constructs a path for all the moving objects to be scheduled in sequence according to the following steps:

[0166] Step 1: Initialization. Arrange the mobile entities in the set to be scheduled in any order to form a queue to be processed.

[0167] Step 2: Path Construction. For the current moving entity in the queue to be processed, starting from the starting node of the moving entity, repeatedly perform the node selection operation: calculate the transition probabilities of all reachable next nodes based on the current node, and select the next node according to the roulette wheel strategy, until the ending node of the moving entity is reached; record the sequence of nodes traversed and the corresponding directed edges.

[0168] Step 3: Congestion Information Update. During the current path construction process of the moving vehicle, if the moving vehicle is delayed at a node due to spatiotemporal interference, the waiting time steps of that node are accumulated and used to calculate the accumulated waiting potential energy of that node.

[0169] Step 4: Iterative processing. Set the next moving object in the queue to be processed as the current moving object, and repeat steps 2 to 3 until the paths of all moving objects in the queue to be processed have been constructed.

[0170] Step 5: Solution Output. Combine the paths of all moving entities into a candidate scheduling solution.

[0171] In the above process, the real-time congestion potential energy in the heuristic function used by the ant to calculate the transfer probability is taken from the latest value of the current iteration and is dynamically updated as the iteration process progresses.

[0172] S303, perform spatiotemporal interference constraint verification and comprehensive scheduling cost calculation on the candidate path, generate a reward signal based on the verification result and comprehensive scheduling cost, and update the state-action value function table of the reinforcement learning agent with the reward signal.

[0173] For interference detected during the verification, the following strategy is adopted to resolve it:

[0174] Time-based waiting strategy: make one of the two moving bodies involved in the interference wait in front of the interference node until the safety interval is met;

[0175] Local path replanning strategy: Replan local detour routes for moving objects that are interfering, avoiding interference nodes or road segments.

[0176] The calculation method for the integrated scheduling cost and reward signal is as described above. After receiving the reward signal, the agent updates the value of the corresponding entry in the Q table according to the Bellman equation, and gradually learns the optimal parameter adjustment strategy.

[0177] The iteration terminates when any of the following conditions are met:

[0178] Reach the preset maximum number of iterations;

[0179] The dynamic adjustment instructions output by the reinforcement learning agent remain unchanged for a preset number of generations (e.g., 5 generations) and the convergence value of the comprehensive scheduling cost is stable within a preset threshold range (i.e., the relative fluctuation of the optimal cost in adjacent generations is less than the given tolerance).

[0180] The iteration terminates when any of the above conditions are met, and the current optimal scheduling scheme is output.

[0181] Furthermore, the method also includes:

[0182] S400: The optimal scheduling scheme generated in S300 is converted into motion control commands for each mobile body and sent to the actuator of the mobile body to drive the mobile body to move along the interference-free spatiotemporal path.

[0183] Specifically, the optimal scheduling scheme includes at least the following information: the complete spatiotemporal path of each mobile entity, i.e., the spatial location sequence of the mobile entity at each time node; the planned time for each mobile entity to pass through each road segment and node; and the spatiotemporal interval arrangement between each mobile entity to ensure that the non-interference constraint remains effective during the physical execution phase.

[0184] The above scheduling scheme is converted into control instructions executable by each mobile unit, including:

[0185] Path tracking instructions: include the sequence of nodes the mobile vehicle needs to pass through along the planned path, the speed of each road segment, and the passage order of each road segment;

[0186] Time constraint instructions: These include the planned arrival times of mobile vehicles at key nodes such as intersections, parking spaces, and work sites, as well as the allowable time deviation ranges.

[0187] Priority instructions: When multiple mobile entities compete for the same passage resources, such as narrow road sections or intersections, the passage rights and yielding order of each mobile entity are clearly defined.

[0188] Control commands are issued through a communication network between the dispatch center and the mobile vehicle. The dispatch center encapsulates these commands in a standardized data format and sends them to the vehicle controller or airborne controller of each mobile vehicle via wired or wireless communication links. Upon receiving the commands, the controller parses them into low-level motion control signals such as speed setting, steering angle, and braking commands, driving the mobile vehicle along the planned path.

[0189] During execution, the dispatch center continuously receives real-time location, speed, and status information from each mobile entity. If a mobile entity is detected to have deviated from the planned path or time due to external disturbances such as temporary obstacles, communication delays, or execution errors, and the deviation exceeds a preset tolerance threshold, a local rescheduling process is triggered: the actual state of all mobile entities is used as the new initial conditions, steps S100 to S300 are re-executed to generate an updated scheduling scheme, and the updated control commands are reissued to the affected mobile entities, achieving closed-loop dynamic adjustment and ensuring that the overall system operation continuously meets the interference-free constraint.

[0190] Through the aforementioned execution and feedback mechanisms, the scheduling scheme generated by S300 smoothly transitions from the virtual planning stage to the physical execution stage, and has the ability to correct deviations online, thereby ensuring the executability and system stability of the scheduling scheme in the actual engineering environment.

[0191] In a high-saturation scenario of 100 flights at a large hub airport, the proposed method (RLACO) achieved a total taxiing time of 734.30 minutes, a 16.1% reduction compared to the suboptimal algorithm I-QL (875.75 minutes) and a 61.2% reduction compared to the genetic algorithm (GA). The total waiting time was only 49 seconds, while I-QL and GA reached 2602 seconds and 2891 seconds respectively, effectively eliminating systemic deadlock. Ablation experiments showed that RLACO reduced fuel consumption by 56.2% and total waiting time by 98.1% compared to the basic ACO. In a generalization experiment of 80 flights at another compact hub airport, RLACO achieved an average taxiing time of 6.71 minutes, a 13.6% reduction compared to I-QL (7.77 minutes), with a total waiting time of only 42 seconds (compared to 1282 seconds for I-QL). These experiments demonstrate that the proposed method exhibits significant advantages and strong robustness under high-density, cross-scenario conditions.

[0192] This invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being configured to perform the method described in this invention.

[0193] This invention also provides a computer-readable storage medium storing computer-executable instructions for performing the methods described in this invention.

[0194] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.

[0195] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A reinforcement learning-based adaptive adjustment method for multi-moving body cooperative scheduling of ant colonies, characterized in that, The method includes the following steps: S100, acquire physical access network data of the scheduling scenario and task data of multiple mobile bodies to be scheduled; wherein, the scheduling scenario refers to an engineering scenario composed of a physical access network and multiple mobile bodies moving within the physical access network; the physical access network is modeled as a directed graph, wherein the nodes of the directed graph represent location points and the edges represent passable road segments, and the task data includes at least the starting node, ending node and earliest allowed departure time of each mobile body; S200, initialize the pheromone heuristic factor and the expectation heuristic factor of the ant colony algorithm, and initialize the reinforcement learning agent, which is associated with a pre-constructed state-action value function table; S300, iteratively execute the preset search process until the termination condition is met, and generate the optimal scheduling scheme, which includes the non-interference spatiotemporal path of each moving body from the starting point to the ending point. The preset search process includes the following sub-steps: S301, the reinforcement learning agent collects the aggregation degree, iteration progress and evolutionary stagnation flag of the population scheduling cost during the current ant colony search process, selects an action according to the state-action value function table, and generates dynamic adjustment instructions for the pheromone heuristic factor and the desired heuristic factor. S302, the ant colony algorithm uses the adjusted pheromone heuristic factor and the expected heuristic factor as running parameters to perform path search on the directed graph. Each ant iteratively selects the next node according to the transfer probability that combines the static path distance and the real-time congestion potential of the node, and constructs candidate paths for all the mobile bodies to be scheduled in turn. S303, perform spatiotemporal interference constraint verification and comprehensive scheduling cost calculation on the candidate path, generate a reward signal based on the verification result and comprehensive scheduling cost, and update the state-action value function table of the reinforcement learning agent with the reward signal; The heuristic function in the transition probability includes a static spatial guidance term and a dynamic congestion penalty term; the static spatial guidance term is calculated based on the physical length of the road segment and the shortest distance from the node to the target point; the dynamic congestion penalty term is calculated based on the real-time congestion potential energy of the node; the formula for calculating the real-time congestion potential energy of the node is: E q (t) = exp(-h·W q (t) / W max ); Among them, E q (t) represents the real-time congestion potential energy of node q in the t-th iteration, W q (t) represents the accumulated waiting potential energy of node q in the t-th iteration, exp() represents the natural exponential function, h is the congestion sensitivity factor, and W max This serves as the normalized benchmark.

2. The method according to claim 1, characterized in that, The clustering degree of the population scheduling cost is quantified by the coefficient of variation of the scheduling costs of candidate paths constructed by all ants in the current generation; the evolutionary stagnation indicator is determined by judging whether the global optimal solution has been improved over several consecutive generations; the comprehensive scheduling cost includes the weighted sum of the total running time and total energy consumption of the moving body.

3. The method according to claim 1, characterized in that, The state space of the state-action value function table is composed of three-dimensional discretization of the population scheduling cost, iteration progress, and evolutionary stagnation marker; the action space of the state-action value function table contains multiple preset search patterns, and each preset search pattern corresponds to a set of preset pheromone heuristic factor increment values ​​and expected heuristic factor increment values.

4. The method according to claim 3, characterized in that, The preset search modes include: random exploration mode, heuristic enhancement mode, path reinforcement mode, and parameter reset mode; the random exploration mode is used to simultaneously reduce the values ​​of the pheromone heuristic factor and the desired heuristic factor; the heuristic enhancement mode is used to reduce the value of the pheromone heuristic factor and increase the value of the desired heuristic factor; the path reinforcement mode is used to simultaneously increase the values ​​of the pheromone heuristic factor and the desired heuristic factor; and the parameter reset mode is used to restore the pheromone heuristic factor and the desired heuristic factor to their initial values.

5. The method according to claim 1, characterized in that, The real-time congestion potential of a node is constructed in the following way: the cumulative waiting potential of each node is calculated based on the number of moving bodies waiting at each node in the current iteration or the cumulative waiting time step; the cumulative waiting potential is normalized and exponentially transformed to serve as the real-time congestion potential of that node.

6. The method according to claim 1, characterized in that, The ant colony algorithm only allows the best candidate path in the current iteration to deposit pheromones, and limits the pheromone concentration on each road segment to between a preset lower threshold and an upper threshold.

7. The method according to claim 1, characterized in that, The reward signal is a composite reward function, which includes at least: a performance reward term reflecting the improvement in scheduling cost of the current scheduling scheme compared to the historical best scheme, a feasibility reward term for incentivizing the generation of interference-free scheduling schemes, and a penalty term for suppressing global path interlocking; wherein the weight of the penalty term is greater than the weights of the performance reward term and the feasibility reward term.

8. The method according to claim 1, characterized in that, The spatiotemporal interference constraint verification involves interference types including head-on interference, tail-on interference, and cross-interference; for interference detected during verification, a time-based waiting strategy or a local path replanning strategy is adopted to resolve it.

9. The method according to claim 1, characterized in that, The termination condition is reaching the preset maximum number of iterations, or the dynamic adjustment instructions output by the reinforcement learning agent remain unchanged for multiple consecutive generations and the convergence value of the scheduling cost is stable within a preset threshold range.

Citation Information

Patent Citations

  • An airport joint scheduling method integrating bidirectional particle swarm and multi-strategy ant colony

    CN117852841B

  • Self-adaptive ant colony optimization based flexible workshop dispatching technology

    CN103246938A

  • Improved ant colony algorithm based on road congestion problem

    CN114418056A