Cooperative scheduling optimization method for rail and trackless combined transportation of underground mine
By constructing a multi-objective optimization model and a dynamic time window mutual exclusion constraint mechanism, combined with multi-agent Markov decision process and deep reinforcement learning algorithm, the problem of coordinated scheduling of tracked and trackless equipment in underground mine transportation system is solved, realizing global optimization and intelligent management, and improving transportation efficiency and system adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-03-27
AI Technical Summary
In the existing underground mine transportation system, the lack of a unified scheduling and coordination mechanism for both rail-mounted electric locomotives and trackless rubber-tired vehicles leads to problems such as resource conflicts and path blockages, resulting in a decline in transportation efficiency.
A multi-objective optimization model is constructed, which combines a dynamic time window mutual exclusion constraint mechanism and a multi-agent Markov decision process. The collaborative scheduling of rail and trackless transportation systems is optimized through deep reinforcement learning algorithms to achieve global optimization and intelligent management.
It achieves efficient collaboration between rail and trackless transportation systems, significantly reduces transportation costs and risks, improves the level of intelligent scheduling, and enhances the system's adaptability and engineering feasibility.
Smart Images

Figure CN121745537A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent transportation scheduling technology in underground mines, and particularly relates to a collaborative scheduling optimization method for combined rail and trackless transportation in underground mines. Background Technology
[0002] Underground mine transportation systems are a crucial link in mining operations, responsible for efficiently transporting ore from underground to the surface. Their operational efficiency directly impacts mine production capacity and cost control. With increasing mining depth and expanding mining areas, transportation routes become more complex, significantly increasing scheduling difficulties. Especially in the context of current smart mine construction, achieving intelligent management and efficient collaboration in underground transportation processes has become a core issue of concern for the industry.
[0003] Currently, underground mine transportation mainly relies on two types of equipment: rail-mounted electric locomotives and trackless rubber-tired vehicles. Rail transportation is suitable for long-distance, high-volume scenarios with low operating costs but lacks scheduling flexibility. Trackless transportation, on the other hand, is highly mobile and suitable for complex tunnels and localized transportation, but its unit cost is higher and it is easily affected by traffic congestion and road conditions. These two modes of transportation typically operate independently, lacking a unified scheduling and coordination mechanism. Especially at key nodes such as ore chute loading points and intersecting tunnels, resource conflicts and path blockages easily occur, leading to decreased transportation efficiency. Therefore, there is an urgent need to construct a collaborative scheduling method that integrates optimization modeling and intelligent algorithms to achieve efficient collaboration and global optimization between rail-mounted and trackless transportation systems. Summary of the Invention
[0004] This invention proposes a collaborative scheduling optimization method for combined rail and trackless transportation in underground mines, which solves the problems of insufficient coordination of transportation equipment and frequent path conflicts in existing transportation schemes.
[0005] To address the aforementioned technical problems, this invention provides a collaborative scheduling optimization method for combined rail and trackless transportation in underground mines, comprising the following steps:
[0006] S1: With minimizing transportation costs and minimizing safety penalty costs as optimization objectives, a multi-objective optimization model is constructed. The production capacity and grade information of underground mine ore passes, the path data of transportation roadways, and the operating parameters of rail and trackless transportation equipment are collected, and the constraints of the multi-objective optimization model are set.
[0007] S2: Construct a dynamic time window mutual exclusion constraint mechanism for rail-mounted electric locomotives and trackless rubber-tired vehicles at loading points and path intersections. Based on the dynamic time window mutual exclusion constraint mechanism, identify scheduling time window overlaps and running path conflicts in the scheduling scheme, and write the dynamic time window mutual exclusion constraint mechanism into the constraint conditions of the multi-objective optimization model.
[0008] S3: Solve the constructed multi-objective optimization model to generate an initial scheduling scheme.
[0009] Preferably, the expression for the multi-objective optimization model in S1 is:
[0010] ;
[0011] In the formula, The overall goal is to integrate transportation costs with safety penalties; The number of ore passes; This refers to the number of electric trams. This refers to the number of trackless rubber-wheeled vehicles. The unit transportation cost of the tram; For the first A tram from the first The first chute loaded The quantity of graded ore; Unit transportation cost for trackless rubber-tired vehicles; For the first A trackless rubber-wheeled vehicle from the first The first chute loaded The quantity of graded ore; This is a weighting coefficient for security penalties; Unit conflict cost; The total number of equipment intersection conflicts within a shift; The weighting factor is the duration of the conflict. The total duration of equipment intersection conflicts within a shift, expressed in minutes.
[0012] Preferably, the constraints in S1 include:
[0013] (1) Ore balance constraint:
[0014] ;
[0015] In the formula, Target ore quantity for each shift;
[0016] (2) Well ore production constraints:
[0017] ;
[0018] In the formula, For the first The maximum output of a single well;
[0019] (3) Ore grade constraints:
[0020] ;
[0021] In the formula, Indicates the first Taiwan tram from the first The first chute was transported via a rail transport route. The actual average grade of graded ore; Indicates the first Taiwan trackless rubber-wheeled vehicle from the first The first chute was transported via a trackless transport route. The actual average grade of graded ore; To meet the target quality requirements;
[0022] (4) Constraints to avoid ore loading conflicts:
[0023] ;
[0024] ;
[0025] In the formula, Indicates the first Taiwan tram in the first The well chute was the first The start time of loading ore into the ores; Indicates the relationship with the first The first incident of the Taiwan tram conflict Taiwan tram in the first The well chute was the first The end time of loading for similar ores; Indicates the first Taiwan trackless rubber-wheeled vehicle in the first The well chute was the first The start time of loading ore into the ores; Indicates the relationship with the first The first conflict between Taiwan and the trackless rubber-wheeled vehicle Taiwan trackless rubber-wheeled vehicle in the first The well chute was the first The end time of loading for similar ores; Let 0-1 be the decision variable, indicating that in the 0th... In the chute, the operational sequence of the railcars is as follows: When the first... The ore loading operation of the tram was in the first When the tram was in progress, =1, otherwise =0; Let 0-1 be the decision variable, indicating that in the 0th... In the chute, the working sequence of the trackless rubber-wheeled vehicles is as follows: When the first... The loading operation of the trackless rubber-wheeled vehicle was in the first After the trackless rubber-wheeled vehicle proceeded, =1, otherwise =0; It is a sufficiently large positive number;
[0026] (5) Path intersection conflict avoidance constraint:
[0027] For the intersection of track and trackless roads Any rail-mounted electric locomotive With trackless rubber-wheeled vehicles The passage time must meet the non-overlapping constraint:
[0028] ;
[0029] ;
[0030] ;
[0031] In the formula, Number the intersections of the paths; The time when the vehicle enters the intersection; For the first Taiwan's tram enters the first The moment of each intersection; For the first Taiwan tram leaves The moment of each intersection; For the first A trackless rubber-wheeled vehicle entered the first The moment of each intersection; For the first A trackless rubber-wheeled vehicle left the first The moment of each intersection; In order to be with the first The trams involved in the conflict in Taiwan Entering the The moment of each intersection; In order to be with the first The trams involved in the conflict in Taiwan Leaving the The moment of each intersection; In order to be with the first Trackless rubber-tired vehicles involved in the conflict with trams in Taiwan Entering the The moment of each intersection; In order to be with the first Trackless rubber-tired vehicles involved in the conflict with trams in Taiwan Leaving the The moment of each intersection;
[0032] (6) Vehicle capacity constraints:
[0033] ;
[0034] ;
[0035] In the formula, Rated loading capacity of the railcar Rated loading capacity for trackless rubber-tired vehicles; For the first Taiwan tram in the first The chute was loaded with the first The amount of ore loaded in a single batch; For the first Taiwan trackless rubber-wheeled vehicle in the first The chute was loaded with the first The amount of ore loaded in a single batch;
[0036] (7) Time forecasting constraints: - ;
[0037] - ;
[0038] In the formula, For the first Taiwan tram in the first The chute was loaded with the first Predicted loading time for similar ores; For the first Taiwan trackless rubber-wheeled vehicle in the first The chute was loaded with the first Predicted loading time for similar ores; For the first The first tram The well chute was the first The end time of loading for similar ores; For the first The first tram The well chute was the first The start time for loading ore types; For the first The first trackless rubber-wheeled vehicle The well chute was the first End time of loading for similar ores; For the first The first trackless rubber-wheeled vehicle The well chute was the first The start time for loading ore types;
[0039] (8) Schedule constraints:
[0040] ; .
[0041] Preferably, the dynamic time window mutual exclusion constraint mechanism is constructed in S2, including the following steps:
[0042] S21: Based on vehicle arrival time, loading / passage time, and task urgency, define an operation time window for each transportation device. The expression for the operation time window is:
[0043] ;
[0044] In the formula, This refers to the vehicle's arrival time. Loading / passage time; The initial time window for each vehicle; The start time of operation for each piece of transport equipment; This refers to the end time of operation for each piece of transport equipment;
[0045] S22: Calculate the priority of transportation equipment by combining task urgency, unit transportation cost and equipment utilization rate;
[0046] S23: When overlapping time windows of different devices are detected, it is determined to be a scheduling conflict. The time windows of the devices with lower priority among the conflicting devices are extended as a whole. The specific adjustment method is as follows:
[0047] ;
[0048] In the formula, Adjusted operating time window for low-priority equipment; Adjusted start time for low-priority equipment; The end time of the operation after adjustment for low-priority equipment; The end time for high-priority vehicles; The loading / passage time for low-priority vehicles.
[0049] Preferably, the expression for calculating the priority of the transport equipment in S22 is:
[0050] ;
[0051] In the formula, Priority indicators; To determine the urgency of the task; Unit transportation cost; To improve equipment utilization; , , These are the weighting coefficients.
[0052] Preferably, when solving the constructed multi-objective optimization model in S3, the multi-objective optimization model is transformed into a multi-agent Markov decision process model, including the following steps:
[0053] S31: Constructing the state space:
[0054] Transportation task status matrix: 0 = corresponding chute task not assigned to a vehicle, 1 = corresponding chute task assigned to a vehicle;
[0055] Equipment operation capability matrix: 1 = vehicle can operate in the current chute-path combination, 0 = vehicle cannot operate in the current chute-path combination;
[0056] Process duration matrix: The values represent the estimated time for a vehicle to complete the corresponding process;
[0057] Equipment transfer time matrix: The values represent the one-way travel time of the corresponding equipment from the current work point to the target work point;
[0058] Energy consumption matrix: The values represent the unit energy consumption of the corresponding equipment under the combination of the current work point and the target work point;
[0059] S32: Constructing the Action Space: Scheduling actions include task allocation for rail-mounted electric locomotives and trackless rubber-tired vehicles, setting scheduling priorities, and calling up backup equipment. Action selection rules are set based on transportation costs, safety risks, remaining vehicle loading capacity, chute usage status, and route safety risks.
[0060] Rule 1: Prioritize the task combination that best balances transportation costs and safety risks;
[0061] Rule 2: Prioritize task allocation schemes with short idle mileage and high proportion of safe paths;
[0062] Rule 3: High-value mineral transport tasks should be prioritized for equipment with low unit cost;
[0063] Rule 4: When the path safety risk exceeds the threshold, schedule equipment with a low-cost detour plan;
[0064] Rule 5: When the ore supply in the ore pass is insufficient, dynamically transfer the tasks to be assigned to the ore pass with sufficient ore supply;
[0065] Rule 6: When there is sufficient ore in the ore pass, priority should be given to scheduling low-unit-cost equipment at full load;
[0066] Rule 7: In the event of equipment failure, priority should be given to dispatching low-cost backup equipment with safety compatibility;
[0067] S33: After the network outputs an action, the agent determines the scheduled process according to the corresponding scheduling rules, and the environment provides an immediate reward. This reflects the impact of the scheduling scheme.
[0068] Preferably, the expression for the instant reward in S33 is:
[0069] ;
[0070] In the formula, For instant rewards; To integrate transportation costs with safety penalties to achieve the overall target cost; Total vehicle waiting time; Negative reward weighting for security; Unit conflict cost; This represents the total number of conflicts. The weighting factor is the duration of the conflict.
[0071] The total duration of equipment intersection conflicts within a shift, in minutes; , , These are the weighting coefficients.
[0072] Preferably, in step S3, after transforming the multi-objective optimization model into a multi-agent Markov decision process model and defining immediate rewards, the optimal scheduling strategy is trained and optimized using the deep reinforcement learning algorithm PPO, including the following steps:
[0073] Step (1): Initialize the environment and agent parameters, and initialize the matrix and agent in the state space;
[0074] Step (2): Within the set scheduling period, construct the current scheduling state vector based on task allocation, equipment status and resource parameters, and select the transportation resources that can participate in the scheduling of this period by combining the constraints in the multi-objective optimization model.
[0075] Step (3): The agent evaluates the expected benefits of candidate actions based on the rules in the action space, assigns priorities based on the estimated values, and selects the optimal scheduling action;
[0076] Step (4): After executing the selected action, update the parameters of the policy network and value network based on the feedback state transition results and immediate reward values;
[0077] Step (5): If the current shift reaches the set shift time constraint and the cumulative transported ore volume reaches the target value, output the optimal scheduling scheme; otherwise, repeat steps (2) to (4) based on the updated status information until the task is completed.
[0078] Preferably, the training process of the PPO algorithm includes the following steps:
[0079] Step (1): Based on the current state vector, the agent generates the probability distribution of each scheduled action through the policy execution network Actor, selects actions according to the probability distribution, interacts with the environment, collects interaction data, and uses the value network Critic to calculate the value of the final state.
[0080] Step (2): Starting with the value of the final state, calculate the cumulative reward of each historical state using a backward method to form a reward set. Input the historical states into the Critic network in sequence to obtain the corresponding state value estimation set. Calculate the advantage function based on the reward set and the state value set, and construct the loss function of the Critic network.
[0081] Step (3): Input the state value set into the Actor network before and after the update respectively, obtain the corresponding action probability distribution, and calculate the ratio of the action probability of the old and new policies under the same state-action pair;
[0082] Step (4): Construct the total loss function by combining the loss function of the Critic network and the loss function of the Actor network;
[0083] Step (5): Perform backpropagation based on the total loss function, and update the parameters of the Critic network and the Actor network simultaneously;
[0084] Step (6): Repeat steps (3) to (5), adjust the weights of the Actor network multiple times based on real-time transportation data, and perform local iterative updates until the strategy is stable in the current round;
[0085] Step (7): Repeat steps (1) to (6) until multiple training rounds are completed. Then, the iteration terminates and the trained model is output.
[0086] Preferably, the expression for the total loss function constructed in step (4) is:
[0087] ;
[0088] In the formula, Total loss; This is an approximate proportional shear loss; For the value function loss; For policy entropy loss; , , These are the weight parameters.
[0089] The beneficial effects of the present invention include at least the following:
[0090] 1. Achieve coordinated scheduling and global optimization: Break through the limitations of independent operation of rail and trackless transportation systems, and achieve global optimal resource allocation through unified management of scheduling models and constraint mechanisms;
[0091] 2. Significantly reduce transportation costs and risks: Under multiple constraints such as ore balance, grade, and safety, optimize equipment scheduling paths and time arrangements to reduce empty load rate, congestion rate, and number of conflicts;
[0092] 3. Enhance the intelligence level of scheduling: Model the scheduling problem as a multi-agent decision-making process, introduce reinforcement learning strategy, support dynamic and real-time scheduling adjustment, and significantly improve the system's adaptive ability and scheduling efficiency;
[0093] 4. Enhance the adaptability of the scheduling system: By adopting a priority allocation mechanism and a backup equipment call strategy, other task paths or equipment can be automatically selected in the event of equipment failure or unavailable path, ensuring the continuous operation of transportation tasks;
[0094] 5. Possesses broad promotional value: This invention is applicable to underground mines of different scales and with different transportation structures, and has good engineering feasibility and application prospects. Attached Figure Description
[0095] Figure 1 This is a flowchart of the coordinated scheduling optimization method for rail-rail and trackless combined transportation according to an embodiment of the present invention;
[0096] Figure 2 This is a flowchart illustrating the operation of the time window mutual exclusion rule in an embodiment of the present invention.
[0097] Figure 3 This is a schematic diagram illustrating the principle of the PPO algorithm in an embodiment of the present invention;
[0098] Figure 4 This is a schematic diagram of the training process of the PPO algorithm in an embodiment of the present invention. Detailed Implementation
[0099] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0100] Before describing specific embodiments of the present invention, the following explanations are provided:
[0101] Given a pre-determined transportation plan, suppose at a certain moment, multiple ore passes within a mining area require transportation operations. The ore volume of each pass is available according to the plan. Multiple transportation equipment of various types are involved in the operation, with each piece of equipment having potentially different production capacities and energy consumption. Each transportation task consists of multiple processes, including loading, transporting, and unloading ore. The time required for each process is available based on on-site data. Each transportation route follows the same production process and is repeated multiple times, but the sequence of processes must be followed. Any process in any transportation operation must be performed using one or more designated pieces of equipment. Specifically, the following definitions apply:
[0102] (1) The same equipment can only perform one transportation-related process at a time, and parallel operations are not allowed;
[0103] (2) Ignore auxiliary time such as equipment start-up and shutdown, and only calculate the time of core processes such as loading, transportation and unloading;
[0104] (3) Only one piece of equipment can be used for loading ore at any given time in the same pass to avoid spatial conflicts;
[0105] (4) The equipment operation process must not be interrupted; once started, the current task must be completed.
[0106] (5) There are no sudden events to interfere with the information such as the amount of ore in the ore pass, the status of equipment, and the passage of transportation routes. The information is known and stable.
[0107] (6) All equipment and ore passes were in a usable state at shift 0, with no initial idleness or malfunction;
[0108] (7) The transfer time of the equipment between different chutes is known;
[0109] (8) The number of operations at each ore pass must be controlled according to the production capacity limit, and the amount of ore loaded in a single operation must not exceed the rated load capacity of the equipment. For example... Figure 1 As shown, this invention provides a collaborative scheduling optimization method for combined rail and trackless transportation in underground mines. Using four key parameters—working time, transportation cost, conflict situation, and work plan—as input, the method employs three stages: multi-objective modeling, dynamic time window conflict detection, and deep reinforcement learning strategy optimization. The final output is a collaborative scheduling scheme for rail and trackless transportation equipment in underground mines, including the following steps:
[0110] S1: With the optimization objectives of minimizing transportation costs and minimizing safety penalty costs, a multi-objective optimization model is constructed. The production capacity and grade information of underground mine ore passes, the path data of transportation roadways, and the operating parameters of rail and trackless transportation equipment are collected, and the constraints of the multi-objective optimization model are set.
[0111] Specifically, based on the underground mining operation plan, it is assumed that within the underground mining area developed through the integrated use of vertical shafts and inclined ramps, there are... The ore volume from each pass needs to be transferred to the standard unloading point via the transportation system. The amount of ore to be transported from each pass within an 8-hour shift on a given day is set according to the mining operation plan. Configuration Taiwanese transportation equipment participated in ore transfer operations, including Taiwan tram and Each trackless rubber-tired vehicle has a production capacity of [number] units. All operations must ensure that the single loading volume does not exceed the rated value. The amount of ore to be transported in each pass must be transferred through the core transportation process of passing loading → route transportation → unloading point unloading. The transportation of ore in each stope follows the same process, and is adjusted according to the amount of ore to be transported and the equipment capacity. Each round trip transportation operation must adhere to the following sequence: transportation can only begin after loading is complete, unloading can only begin after transportation is complete, and the next loading operation can only proceed after unloading is complete. Every transportation step within any single operation must be performed using a designated transportation machine or one of several similar machines. The operation time for the same type of step in different transportation operations within the same stope may differ on the same transportation machine; the operation time for the same transportation step may also differ on different transportation machines. Furthermore, the transfer time of transportation equipment between different stope chutes must be taken into consideration during scheduling to ensure continuous and efficient transportation operations.
[0112] After completing data collection and constraint setting, a multi-objective optimization model is constructed based on transportation cost and safety indicators, with the goal of minimizing transportation cost and safety penalty cost. This provides a foundation for generating the initial scheduling scheme. Through variable modeling and objective function design, global optimization of transportation task allocation and route selection is achieved.
[0113] The expression for the multi-objective optimization model is:
[0114] ;
[0115] In the formula, The overall goal is to integrate transportation costs with safety penalties; The number of ore passes; This refers to the number of electric trams. This refers to the number of trackless rubber-wheeled vehicles. The unit transportation cost of the tram; For the first A tram from the first The first chute loaded The quantity of graded ore; Unit transportation cost for trackless rubber-tired vehicles; For the first A trackless rubber-wheeled vehicle from the first The first chute loaded The quantity of graded ore; This is a weighting coefficient for security penalties; Unit conflict cost; The total number of equipment intersection conflicts within a shift; The weighting factor is the duration of the conflict. The total duration of equipment intersection conflicts within a shift, expressed in minutes.
[0116] After establishing the multi-objective optimization model, several scheduling constraints need to be introduced to ensure the feasibility and safety of the scheduling scheme, including mineral quantity balance, resource capacity limit, conflict avoidance mechanism, etc. The optimization model is solved through constraint modeling to obtain an initial scheduling scheme that meets the requirements, specifically including the following constraints:
[0117] (1) Ore balance constraint, used to ensure that the ore transported within a shift meets the target requirements:
[0118] ;
[0119] In the formula, Target ore quantity for each shift; The number of ore passes; This refers to the number of electric trams. This refers to the number of trackless rubber-wheeled vehicles.
[0120] (2) Well ore production constraints, used to prevent over-exploitation of well ore:
[0121] ;
[0122] In the formula, Let be the upper limit of the production capacity of the i-th well.
[0123] (3) Ore grade constraints, used to ensure that the overall grade of transported ore meets the standards:
[0124] ;
[0125] In the formula, Indicates the first Taiwan tram from the first The first chute was transported via a rail transport route. The actual average grade of graded ore; Indicates the first Taiwan trackless rubber-wheeled vehicle from the first The first chute was transported via a trackless transport route. The actual average grade of graded ore; The target grade requirement is that the weighted average grade of the ore transported along each transport route must not be lower than this value, in order to ensure that the overall quality of the ore transported in the end meets the set standard.
[0126] (4) Loading conflict avoidance constraint, used to ensure that no loading conflict occurs in the same pass at the same time:
[0127] ;
[0128] ;
[0129] In the formula, Indicates the first Taiwan tram in the first The well chute was the first The start time of loading ore into the ores; Indicates the relationship with the first The first incident of the Taiwan tram conflict Taiwan tram in the first The well chute was the first The end time of loading for similar ores; Indicates the first Taiwan trackless rubber-wheeled vehicle in the first The well chute was the first The start time of loading ore into the ores; Indicates the relationship with the first The first conflict between Taiwan and the trackless rubber-wheeled vehicle Taiwan trackless rubber-wheeled vehicle in the first The well chute was the first The end time of loading for similar ores; Let 0-1 be the decision variable, indicating that in the 0th... In the chute, the operational sequence of the railcars is as follows: When the first... The ore loading operation of the tram was in the first When the tram was in progress, =1, otherwise =0; Let 0-1 be the decision variable, indicating that in the 0th... In the chute, the working sequence of the trackless rubber-wheeled vehicles is as follows: When the first... The loading operation of the trackless rubber-wheeled vehicle was in the first After the trackless rubber-wheeled vehicle proceeded, =1, otherwise =0; It is a sufficiently large positive number.
[0130] (5) Path intersection conflict avoidance constraint:
[0131] For the intersection of track and trackless roads For example, at alleyway bends, at level crossings, any rail-mounted electric locomotive With trackless rubber-wheeled vehicles The passage time must meet the non-overlapping constraint:
[0132] ;
[0133] ;
[0134] ;
[0135] In the formula, Number the intersections of the paths; The time when the vehicle enters the intersection; For the first Taiwan's tram enters the first The moment of each intersection; For the first Taiwan tram leaves The moment of each intersection; For the first A trackless rubber-wheeled vehicle entered the first The moment of each intersection; For the first A trackless rubber-wheeled vehicle left the first The moment of each intersection; In order to be with the first The trams involved in the conflict in Taiwan Entering the The moment of each intersection; In order to be with the first The trams involved in the conflict in Taiwan Leaving the The moment of each intersection; In order to be with the first Trackless rubber-tired vehicles involved in the conflict with trams in Taiwan Entering the The moment of each intersection; In order to be with the first Trackless rubber-tired vehicles involved in the conflict with trams in Taiwan Leaving the The moment of each intersection;
[0136] If the above conditions are not met, it is determined to be a path conflict and is included in the total number of conflicts. The duration of the conflict is:
[0137] ;
[0138] And Cumulative to the total duration of the conflict .
[0139] (6) Vehicle capacity constraints, used to ensure that the vehicle's load does not exceed its rated load capacity:
[0140] ;
[0141] ;
[0142] In the formula, Rated loading capacity of the railcar Rated loading capacity for trackless rubber-tired vehicles; For the first Taiwan tram in the first The chute was loaded with the first The amount of ore loaded in a single batch; For the first Taiwan trackless rubber-wheeled vehicle in the first The chute was loaded with the first The amount of ore loaded in a single batch.
[0143] (7) Time forecasting constraint, used to constrain whether the actual loading time conforms to the forecast:
[0144] - ;
[0145] - ;
[0146] In the formula, For the first Taiwan tram in the first The chute was loaded with the first Predicted loading time for similar ores; For the first Taiwan trackless rubber-wheeled vehicle in the first The chute was loaded with the first Predicted loading time for similar ores; For the first The first tram The well chute was the first The end time of loading for similar ores; For the first The first tram The well chute was the first The start time for loading ore types; For the first The first trackless rubber-wheeled vehicle The well chute was the first End time of loading for similar ores; For the first The first trackless rubber-wheeled vehicle The well chute was the first The start time for loading ore into the ores.
[0147] (8) Shift time constraints, used to ensure that transportation tasks are completed within the specified 8-hour shifts:
[0148] ; .
[0149] S2: Construct a dynamic time window mutual exclusion constraint mechanism for rail-mounted electric locomotives and trackless rubber-tired vehicles at ore loading points and path intersections. Based on the dynamic time window mutual exclusion constraint mechanism, identify scheduling time window overlaps and running path conflicts in the scheduling scheme, and write the dynamic time window mutual exclusion constraint mechanism into the constraint conditions of the multi-objective optimization model.
[0150] Figure 2 This is a flowchart illustrating the time window conflict detection and adjustment process in an embodiment of the present invention. The diagram details the operation of the dynamic time window mutual exclusion constraint mechanism, including starting with real-time collection of transportation equipment status data, sequentially calculating the initial operation time window, performing scheduling conflict detection, determining whether there is time overlap or path resource conflict, and iterating continuously through dynamic priority adjustment and scheduling sequence updates until a scheduling scheme satisfying the cooperative constraints is generated. This mechanism can effectively coordinate the operation sequence of rail-mounted electric locomotives and trackless rubber-tired vehicles, reduce conflict risks, and improve scheduling execution efficiency.
[0151] Specifically, the following steps are included:
[0152] S21: Based on vehicle arrival time, loading / passage time, and task urgency, define an operation time window for each transportation device. The expression for the operation time window is:
[0153] ;
[0154] In the formula, This refers to the vehicle's arrival time. Loading / passage time; The initial time window for each vehicle; The start time of operation for each piece of transport equipment; This refers to the end time of operation for each transport device.
[0155] S22: After identifying the scheduling conflict between rail-mounted electric locomotives and trackless rubber-tired vehicles, in order to reasonably adjust the scheduling order, it is necessary to calculate the priority index of each transportation equipment, and prioritize the operation of critical tasks, high-utilization and low-cost equipment. The following priority evaluation function is used for calculation:
[0156] ;
[0157] In the formula, Priority indicators; To determine the urgency of the task; Unit transportation cost; To improve equipment utilization; , , These are the weighting coefficients;
[0158] S23: When overlapping time windows of different devices are detected, i.e., the start loading time of vehicle A is less than the end loading time of vehicle B and the start loading time of vehicle B is less than the end loading time of vehicle A. If a scheduling conflict occurs, the time window of the lower-priority device among the conflicting devices will be extended as a whole. The specific adjustment method is as follows:
[0159] ;
[0160] In the formula, Adjusted operating time window for low-priority equipment; Adjusted start time for low-priority equipment; The end time of the operation after adjustment for low-priority equipment; The end time for high-priority vehicles; The loading / passage time for low-priority vehicles.
[0161] S24: Priority is recalculated every 10 minutes to adapt to real-time changes in operating conditions, such as the insertion of temporary high-urgency tasks or changes in utilization after equipment failure repair.
[0162] S3: Solve the constructed multi-objective optimization model to generate an initial scheduling scheme.
[0163] To achieve intelligent scheduling and optimization of rail-based and trackless transportation equipment in underground mines, the aforementioned multi-objective optimization model is transformed into a multi-agent Markov decision process model. By systematically modeling the state space, action space, and reward mechanism, each transportation device acts as an agent, continuously interacting and learning in the scheduling environment to optimize strategies and achieve efficient task allocation and path selection.
[0164] like Figure 3 As shown, to more clearly illustrate the interaction mechanism between the agent and the scheduling environment, this diagram illustrates how the agent, based on its current state, interacts with the scheduling environment during each scheduling cycle. Select scheduling action The environment returns to a new state based on the selected action. and instant rewards This constitutes a learning loop for continuous strategy optimization. The above-described interaction process forms the basic framework of the multi-agent reinforcement learning model in this embodiment of the invention, specifically including the following steps:
[0165] S31: Constructing the state space:
[0166] Transportation task status matrix: 0 = corresponding chute task not assigned to a vehicle, 1 = corresponding chute task assigned to a vehicle;
[0167] Equipment operation capability matrix: 1 = vehicle can operate in the current chute-path combination, 0 = vehicle cannot operate in the current chute-path combination;
[0168] Process duration matrix: The values represent the estimated time for a vehicle to complete the corresponding process;
[0169] Equipment transfer time matrix: The values represent the one-way travel time of the corresponding equipment from the current work point to the target work point;
[0170] Energy Consumption Matrix: The values represent the unit energy consumption of the corresponding equipment under the combination of the current work point and the target work point.
[0171] These five matrices represent five different states and serve as the data input for the model. This data needs to be determined based on the on-site environment and production information. The system comprehensively covers information such as the chute usage status, real-time vehicle location, remaining loading capacity, elapsed time, and predicted work hours based on advanced algorithms, providing rich data for the agent's decision-making.
[0172] S32: Constructing the action space: This includes scheduling actions such as task allocation for rail-mounted electric locomotives and trackless rubber-tired vehicles, setting scheduling priorities, and calling up backup equipment. Action selection rules are set based on transportation costs, safety risks, remaining vehicle loading capacity, chute usage status, and route safety risks.
[0173] Rule 1: Prioritize the task combination that best balances transportation costs and safety risks;
[0174] Rule 2: Prioritize task allocation schemes with short idle mileage and high proportion of safe paths;
[0175] Rule 3: High-value mineral transport tasks should be prioritized for equipment with low unit cost;
[0176] Rule 4: When the path safety risk exceeds the threshold, schedule equipment with a low-cost detour plan;
[0177] Rule 5: When the ore supply in the ore pass is insufficient, dynamically transfer the tasks to be assigned to the ore pass with sufficient ore supply;
[0178] Rule 6: When there is sufficient ore in the ore pass, priority should be given to scheduling low-unit-cost equipment at full load;
[0179] Rule 7: In the event of equipment failure, priority should be given to dispatching low-cost backup equipment with safety compatibility;
[0180] S33: After the network outputs an action, the AI determines the scheduled process according to the corresponding scheduling rules, and then the environment provides an immediate reward. This reflects the impact of the scheduling scheme.
[0181] Specifically, to guide the multi-agent scheduling system towards optimal transportation efficiency and minimal safety conflicts, an immediate reward function needs to be set for each scheduling decision action. This positive and negative incentive mechanism improves overall scheduling performance. The immediate reward function comprehensively considers multiple factors such as transportation costs, waiting times, and safety conflicts, and its expression is:
[0182] ;
[0183] In the formula, For instant rewards; To integrate transportation costs with safety penalties to achieve the overall target cost; Total vehicle waiting time; Negative reward weighting for security; Unit conflict cost; This represents the total number of conflicts. The weighting factor is the duration of the conflict. The total duration of equipment intersection conflicts within a shift, in minutes; , , These are the weighting coefficients.
[0184] After transforming the multi-objective optimization model into a multi-agent Markov decision process model and setting an immediate reward function, the Proximal Policy Optimization (PPO) algorithm from deep reinforcement learning is used to train the agents' policies in order to obtain a stable and effective cooperative scheduling strategy. The training process is based on multi-round state-action interactions, continuously updating the parameters of the policy network and value network to improve the performance of the scheduling strategy. Specifically, it includes the following steps:
[0185] Step (1): Initialize the environment and agent parameters, and initialize the matrix and agent in the state space;
[0186] Step (2): Within the set scheduling period, construct the current scheduling state vector based on task allocation, equipment status and resource parameters, and select the transportation resources that can participate in the scheduling of this period by combining the constraints in the multi-objective optimization model.
[0187] Step (3): The agent evaluates the expected benefits of candidate actions based on the rules in the action space, assigns priorities based on the estimated values, and selects the optimal scheduling action;
[0188] Step (4): After executing the selected action, update the parameters of the policy network and value network based on the feedback state transition results and immediate reward values;
[0189] Step (5): If the current shift reaches the set shift time constraint and the cumulative transported ore volume reaches the target value, output the optimal scheduling scheme; otherwise, repeat steps (2) to (4) based on the updated status information until the task is completed.
[0190] Figure 4 This demonstrates the complete process from data sampling, value assessment, advantage function calculation, policy improvement, to loss function construction and network parameter updating, and clearly describes the collaborative scheduling policy training mechanism based on the Actor-Critic architecture, specifically including the following steps:
[0191] Step (1): Based on the current state vector, the agent generates the probability distribution of each scheduled action through the policy execution network Actor, selects actions according to the probability distribution, interacts with the environment, collects interaction data, and uses the value network Critic to calculate the value of the final state.
[0192] Step (2): Starting with the value of the final state, calculate the cumulative reward of each historical state using a backward method to form a reward set. Input these historical states into the Critic network in sequence to obtain the corresponding state value estimation set. Calculate the advantage function based on the reward set and the state value set, and construct the loss function of the Critic network.
[0193] Step (3): Input the state value set into the Actor network before and after the update respectively, obtain the corresponding action probability distribution, and calculate the ratio of the action probability of the old and new policies under the same state-action pair. The old probability distribution is the action probability distribution output by the Actor network before the update, and the new probability distribution is the action probability distribution output by the Actor network after the update. The ratio is used to quantify the improvement of the policy before and after the update: if the ratio is close to 1, it means that the difference between the new policy and the old policy is small; if the ratio deviates from 1 by a lot, it means that the adjustment of the old policy by the new policy is large.
[0194] Step (4): Construct the total loss function by combining the loss function of the Critic network and the loss function of the Actor network;
[0195] Step (5): Perform backpropagation based on the total loss function, and update the parameters of the Critic network and the Actor network simultaneously;
[0196] Step (6): Repeat steps (3) to (5), adjust the weights of the Actor network multiple times based on real-time transportation data, and perform local iterative updates until the strategy is stable in the current round;
[0197] Step (7): Repeat steps (1) to (6) until multiple training rounds are completed. Then, the iteration terminates and the trained model is output.
[0198] In the training process described above, to achieve joint optimization of the policy network (Actor) and the value network (Critic), a total loss function is constructed within the Actor-Critic framework to guide the network parameter updates. Its function expression is as follows:
[0199] ;
[0200] In the formula, Total loss; This is an approximate proportional shear loss; For the value function loss; For policy entropy loss; , , These are the weight parameters.
[0201] This invention analyzes the scheduling process and shift completion process of combined rail-mounted electric locomotives and trackless rubber-tired vehicles in the joint development of underground mine shafts and ramps. Addressing the spatial conflicts between the two types of vehicles at loading points and route intersections, as well as industry pain points such as resource allocation imbalances and difficulty in controlling transportation costs, a multi-objective mixed-integer programming model is established with the dual optimization objectives of minimizing total shift transportation costs and minimizing safety penalty costs. The rail-mounted electric locomotives and trackless rubber-tired vehicles are treated as independent intelligent agents, and their real-time scheduling process based on the transportation environment is modeled as a multi-agent Markov decision process, with targeted design of the state space, action space, and reward function. Based on collected vehicle transportation parameters and ore pass production data, a deep reinforcement learning (PPO) algorithm is used to train the intelligent agent model, generating optimal solutions for ore allocation, ore pass allocation, and scheduling timing.
[0202] Specifically, the embodiments of the present invention effectively solve the spatial conflict problem at loading points and path intersections by designing action rules with dynamic time window mutual exclusion constraints and multi-dimensional priority scheduling, combined with the prediction of conflict risks in the state space. Non-productive time such as vehicle waiting time is greatly reduced, ensuring the efficient and orderly progress of rail and trackless joint transportation operations and avoiding efficiency losses caused by scheduling chaos.
[0203] By using a multi-objective mixed integer programming model to scientifically allocate ore reserves and ore passes, and by using reinforcement learning to guide low-cost and low-risk actions, resource waste by rail-mounted electric locomotives and trackless rubber-tired vehicles has been reduced, and equipment utilization has been significantly improved. At the same time, by integrating comprehensive cost control of energy consumption, equipment wear and tear, and conflict losses, the total cost of shift transportation has been effectively reduced, thus significantly improving the economic efficiency of the mine transportation process.
[0204] The deep reinforcement learning algorithm used in this embodiment of the invention can quickly respond to changes in dynamic working conditions underground. When there are fluctuations in the amount of ore in the ore pass or temporary vehicle failures, the intelligent agent can autonomously adjust the scheduling strategy based on the real-time status, maintaining the optimality of the scheduling scheme without human intervention, and fully meeting the flexible scheduling requirements of underground mine joint transportation.
[0205] This invention combines multi-objective mixed integer programming with multi-agent reinforcement learning technology to the scenario of tracked and trackless joint transportation scheduling in underground mines, breaking through the limitations of traditional independent scheduling modes. Furthermore, through modular state and action design, it can adapt to the joint transportation needs of underground mines of different scales and under different development conditions, and has broad application prospects in the field of intelligent transportation scheduling in underground mines with integrated development of vertical shafts and inclined ramps.
[0206] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described; only preferred embodiments of the present invention are illustrated. The descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. As long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0207] It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept, and these all fall within the scope of protection of this invention. Therefore, the scope of protection of this invention should be determined by the appended claims.
Claims
1. A collaborative scheduling optimization method for combined rail and trackless transportation in underground mines, characterized in that, Includes the following steps: S1: With minimizing transportation costs and minimizing safety penalty costs as optimization objectives, a multi-objective optimization model is constructed. The production capacity and grade information of underground mine ore passes, the path data of transportation roadways, and the operating parameters of rail and trackless transportation equipment are collected, and the constraints of the multi-objective optimization model are set. S2: Construct a dynamic time window mutual exclusion constraint mechanism for rail-mounted electric locomotives and trackless rubber-tired vehicles at loading points and path intersections. Based on the dynamic time window mutual exclusion constraint mechanism, identify scheduling time window overlaps and running path conflicts in the scheduling scheme, and write the dynamic time window mutual exclusion constraint mechanism into the constraint conditions of the multi-objective optimization model. S3: Solve the constructed multi-objective optimization model to generate an initial scheduling scheme.
2. The method for coordinated scheduling optimization of combined rail and trackless transportation in underground mines according to claim 1, characterized in that: The expression for the multi-objective optimization model described in S1 is: ; In the formula, The overall goal is to integrate transportation costs with safety penalties; The number of ore passes; This refers to the number of electric trams. This refers to the number of trackless rubber-wheeled vehicles. The unit transportation cost of the tram; For the first A tram from the first The first chute loaded The quantity of graded ore; Unit transportation cost for trackless rubber-tired vehicles; For the first A trackless rubber-wheeled vehicle from the first The first chute loaded The quantity of graded ore; This is a weighting coefficient for security penalties; Unit conflict cost; The total number of equipment intersection conflicts within a shift; The weighting factor is the duration of the conflict. The total duration of equipment intersection conflicts within a shift, expressed in minutes.
3. The method for coordinated scheduling optimization of combined rail and trackless transportation in underground mines according to claim 2, characterized in that: The constraints described in S1 include: (1) Ore balance constraint: ; In the formula, Target ore quantity for each shift; (2) Well ore production constraints: ; In the formula, For the first The maximum output of a single well; (3) Ore grade constraints: ; In the formula, Indicates the first Taiwan tram from the first The first chute was transported via a rail transport route. The actual average grade of graded ore; Indicates the first Taiwan trackless rubber-wheeled vehicle from the first The first chute was transported via a trackless transport route. The actual average grade of graded ore; To meet the target quality requirements; (4) Constraints to avoid ore loading conflicts: ; ; In the formula, Indicates the first Taiwan tram in the first The well chute was the first The start time of loading ore into the ores; Indicates the relationship with the first The first incident of the Taiwan tram conflict Taiwan tram in the first The well chute was the first The end time of loading for similar ores; Indicates the first Taiwan trackless rubber-wheeled vehicle in the first The well chute was the first The start time of loading ore into the ores; Indicates the relationship with the first The first conflict between Taiwan and the trackless rubber-wheeled vehicle Taiwan trackless rubber-wheeled vehicle in the first The well chute was the first The end time of loading for similar ores; Let 0-1 be the decision variable, indicating that in the 0th... In the chute, the operational sequence of the railcars is as follows: When the first... The ore loading operation of the tram was in the first When the tram was in progress, =1, otherwise =0; Let 0-1 be the decision variable, indicating that in the 0th... In the chute, the working sequence of the trackless rubber-wheeled vehicles is as follows: When the first... The loading operation of the trackless rubber-wheeled vehicle was in the first After the trackless rubber-wheeled vehicle proceeded, =1, otherwise =0; It is a sufficiently large positive number; (5) Path intersection conflict avoidance constraint: For the intersection of track and trackless roads Any rail-mounted electric locomotive With trackless rubber-wheeled vehicles The passage time must meet the non-overlapping constraint: ; ; ; In the formula, Number the intersections of the paths; The time when the vehicle enters the intersection; For the first Taiwan's tram enters the first The moment of each intersection; For the first Taiwan tram leaves The moment of each intersection; For the first A trackless rubber-wheeled vehicle entered the first The moment of each intersection; For the first A trackless rubber-wheeled vehicle left the first The moment of each intersection; In order to be with the first The trams involved in the conflict in Taiwan Entering the The moment of each intersection; In order to be with the first The trams involved in the conflict in Taiwan Leaving the The moment of each intersection; In order to be with the first Trackless rubber-tired vehicles involved in the conflict with trams in Taiwan Entering the The moment of each intersection; In order to be with the first Trackless rubber-tired vehicles involved in the conflict with trams in Taiwan Leaving the The moment of each intersection; (6) Vehicle capacity constraints: ; ; In the formula, Rated loading capacity of the railcar Rated loading capacity for trackless rubber-tired vehicles; For the first Taiwan tram in the first The chute was loaded with the first The amount of ore loaded in a single batch; For the first Taiwan trackless rubber-wheeled vehicle in the first The chute was loaded with the first The amount of ore loaded in a single batch; (7) Time forecasting constraints: - ; - ; In the formula, For the first Taiwan tram in the first The chute was loaded with the first Predicted loading time for similar ores; For the first Taiwan trackless rubber-wheeled vehicle in the first The chute was loaded with the first Predicted loading time for similar ores; For the first The first tram The well chute was the first The end time of loading for similar ores; For the first The first tram The well chute was the first The start time for loading ore types; For the first The first trackless rubber-wheeled vehicle The well chute was the first End time of loading for similar ores; For the first The first trackless rubber-wheeled vehicle The well chute was the first The start time for loading ore types; (8) Schedule constraints: ; 。 4. The collaborative scheduling optimization method for combined rail and trackless transportation in underground mines according to claim 1, characterized in that: The dynamic time window mutual exclusion constraint mechanism is constructed in S2, including the following steps: S21: Based on vehicle arrival time, loading / passage time, and task urgency, define an operation time window for each transportation device. The expression for the operation time window is: ; In the formula, This refers to the vehicle's arrival time. Loading / passage time; The initial time window for each vehicle; The start time of operation for each piece of transport equipment; This refers to the end time of operation for each piece of transport equipment; S22: Calculate the priority of transportation equipment by combining task urgency, unit transportation cost and equipment utilization rate; S23: When overlapping time windows of different devices are detected, it is determined to be a scheduling conflict. The time windows of the devices with lower priority among the conflicting devices are extended as a whole. The specific adjustment method is as follows: ; In the formula, Adjusted operating time window for low-priority equipment; Adjusted start time for low-priority equipment; The end time of the operation after adjustment for low-priority equipment; The end time for high-priority vehicles; The loading / passage time for low-priority vehicles.
5. The collaborative scheduling optimization method for combined rail and trackless transportation in underground mines according to claim 4, characterized in that: The expression for calculating the priority of transport equipment in S22 is: ; In the formula, Priority indicators; To determine the urgency of the task; Unit transportation cost; To improve equipment utilization; , , These are the weighting coefficients.
6. The collaborative scheduling optimization method for combined rail and trackless transportation in underground mines according to claim 1, characterized in that: When solving the constructed multi-objective optimization model in S3, the multi-objective optimization model is transformed into a multi-agent Markov decision process model, including the following steps: S31: Constructing the state space: Transportation task status matrix: 0 = corresponding chute task not assigned to a vehicle, 1 = corresponding chute task assigned to a vehicle; Equipment operation capability matrix: 1 = vehicle can operate in the current chute-path combination, 0 = vehicle cannot operate in the current chute-path combination; Process duration matrix: The values represent the estimated time for a vehicle to complete the corresponding process; Equipment transfer time matrix: The values represent the one-way travel time of the corresponding equipment from the current work point to the target work point; Energy consumption matrix: The values represent the unit energy consumption of the corresponding equipment under the combination of the current work point and the target work point; S32: Constructing the Action Space: Scheduling actions include task allocation for rail-mounted electric locomotives and trackless rubber-tired vehicles, setting scheduling priorities, and calling up backup equipment. Action selection rules are set based on transportation costs, safety risks, remaining vehicle loading capacity, chute usage status, and route safety risks. Rule 1: Prioritize the task combination that best balances transportation costs and safety risks; Rule 2: Prioritize task allocation schemes with short idle mileage and high proportion of safe paths; Rule 3: High-value mineral transport tasks should be prioritized for equipment with low unit cost; Rule 4: When the path safety risk exceeds the threshold, schedule equipment with a low-cost detour plan; Rule 5: When the ore supply in the ore pass is insufficient, dynamically transfer the tasks to be assigned to the ore pass with sufficient ore supply; Rule 6: When there is sufficient ore in the ore pass, priority should be given to scheduling low-unit-cost equipment at full load; Rule 7: In the event of equipment failure, priority should be given to dispatching low-cost backup equipment with safety compatibility; S33: After the network outputs an action, the agent determines the scheduled process according to the corresponding scheduling rules, and the environment provides an immediate reward. This reflects the impact of the scheduling scheme.
7. The collaborative scheduling optimization method for combined rail and trackless transportation in underground mines according to claim 6, characterized in that: The expression for the instant reward described in S33 is: ; In the formula, For instant rewards; To integrate transportation costs with safety penalties to achieve the overall target cost; Total vehicle waiting time; Negative reward weighting for security; Unit conflict cost; This represents the total number of conflicts. The weighting factor is the duration of the conflict. The total duration of equipment intersection conflicts within a shift, in minutes; , , These are the weighting coefficients.
8. The collaborative scheduling optimization method for combined rail and trackless transportation in underground mines according to claim 7, characterized in that: In S3, after transforming the multi-objective optimization model into a multi-agent Markov decision process model and defining immediate rewards, the optimal scheduling strategy is trained and optimized using the deep reinforcement learning algorithm PPO, including the following steps: Step (1): Initialize the environment and agent parameters, and initialize the matrix and agent in the state space; Step (2): Within the set scheduling period, construct the current scheduling state vector based on task allocation, equipment status and resource parameters, and select the transportation resources that can participate in the scheduling of this period by combining the constraints in the multi-objective optimization model. Step (3): The agent evaluates the expected benefits of candidate actions based on the rules in the action space, assigns priorities based on the estimated values, and selects the optimal scheduling action; Step (4): After executing the selected action, update the parameters of the policy network and value network based on the feedback state transition results and immediate reward values; Step (5): If the current shift reaches the set shift time constraint and the cumulative transported ore volume reaches the target value, output the optimal scheduling scheme; otherwise, repeat steps (2) to (4) based on the updated status information until the task is completed.
9. The method for coordinated scheduling optimization of combined rail and trackless transportation in underground mines according to claim 8, characterized in that: The training process of the PPO algorithm includes the following steps: Step (1): Based on the current state vector, the agent generates the probability distribution of each scheduled action through the policy execution network Actor, selects actions according to the probability distribution, interacts with the environment, collects interaction data, and uses the value network Critic to calculate the value of the final state. Step (2): Starting with the value of the final state, calculate the cumulative reward of each historical state using a backward method to form a reward set. Input the historical states into the Critic network in sequence to obtain the corresponding state value estimation set. Calculate the advantage function based on the reward set and the state value set, and construct the loss function of the Critic network. Step (3): Input the state value set into the Actor network before and after the update respectively, obtain the corresponding action probability distribution, and calculate the ratio of the action probability of the old and new policies under the same state-action pair; Step (4): Construct the total loss function by combining the loss function of the Critic network and the loss function of the Actor network; Step (5): Perform backpropagation based on the total loss function, and update the parameters of the Critic network and the Actor network simultaneously; Step (6): Repeat steps (3) to (5), adjust the weights of the Actor network multiple times based on real-time transportation data, and perform local iterative updates until the strategy is stable in the current round; Step (7): Repeat steps (1) to (6) until multiple training rounds are completed. Then, the iteration terminates and the trained model is output.
10. The method for coordinated scheduling optimization of combined rail and trackless transportation in underground mines according to claim 9, characterized in that: The expression for the total loss function constructed in step (4) is: ; In the formula, Total loss; This is an approximate proportional shear loss; For the value function loss; For policy entropy loss; , , These are the weight parameters.