Vehicle scheduling and cooperative control system and method for surface mine

Through digital twin technology and multi-agent reinforcement learning, a digital twin environment for open-pit mines is built, which solves the problem of mutual influence of path selection in vehicle scheduling in open-pit mines, realizes efficient path planning and construction cost optimization, and improves the overall operation effect of the mine.

CN120258428APending Publication Date: 2025-07-04JIANGSU XCMG STATE KEY LAB TECH CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510348719.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing open-pit mine vehicle scheduling and path planning methods are difficult to adapt to frequently changing operating conditions, resulting in the mutual influence of vehicle path selection, which is inefficient in overall efficiency, and the existing technology has failed to effectively model the mining resources with digital intelligent construction equipment such as unmanned mine vehicles and unmanned loaders, resulting in insufficient model accuracy and weak adaptability.

Method used

Digital twin technology and multi-agent reinforcement learning methods are adopted, combined with multi-objective optimization and heuristic algorithms, a digital twin environment for open-pit mines is built to realize vehicle collaborative scheduling and path planning, and real-time data processing and security authentication are performed through the hierarchical architecture of the equipment management layer, edge layer, cloud and secure access control layer to optimize vehicle scheduling and path selection.

Benefits of technology

The intelligent level of vehicle scheduling in open-pit mines has been improved, the overall nature and coordinated control capabilities of path planning have been optimized, construction efficiency and cost have been balanced, and the economy and efficiency of mine production have been ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258428A_ABST
    Figure CN120258428A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle scheduling and cooperative control system and method in the technical field of surface mine construction vehicle scheduling, and the system comprises an equipment management layer which is used for obtaining surface mine environment and equipment state data and controlling equipment, and an edge layer which is used for carrying out the preprocessing and preliminary analysis of the data uploaded by the equipment management layer. The cloud end is used for constructing a digital twin environment of the surface mine, calculating a vehicle collaborative scheduling scheme and a vehicle optimal path of the surface mine by utilizing a multi-objective optimization and heuristic algorithm and a multi-agent reinforcement learning algorithm, and performing vehicle trajectory tracking control; the cloud end is used for forwarding a control instruction issued by the cloud end or the edge layer to the equipment management layer; and the security access control layer is used for performing dynamic authority security authentication and access control. According to the invention, dynamic path planning and multi-task cooperative scheduling in a surface mine scene can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a vehicle scheduling and collaborative control system and method for open-pit mines, belonging to the technical field of open-pit mine construction vehicle scheduling. Background Art

[0002] With the continuous expansion of the scale and increasing complexity of mine exploitation, the production scheduling and path planning of open-pit mines face multiple complex challenges. The exploitation process of open-pit mines involves a large number of equipment and processes, such as drilling, blasting, loading, and transportation. Moreover, the mine site environment is complex and dynamic, and the scheduling of multiple vehicles, equipment, and tasks needs to be processed. Traditional scheduling and path planning methods are difficult to adapt to the frequently changing operating conditions. In addition, the path selection between vehicles will affect each other, and how to improve the overall operation efficiency through path collaborative optimization has become a difficult problem. At the same time, the optimization of construction efficiency and cost in mine production is a key requirement, and the existing centralized scheduling scheme is difficult to support efficient real-time calculation due to high latency and bandwidth bottlenecks.

[0003] In recent years, digital twin technology has been widely applied in the industrial field. Digital twin realizes the synchronous operation of the physical world and the digital world by creating a virtual model of a physical entity and establishing real-time data interaction between the real environment and the virtual model. As a distributed artificial intelligence system, the multi-agent system (MAS) can perform collaborative work and decision optimization through multiple independent individuals (agents) in complex scenarios, and is suitable for multi-equipment collaboration and task allocation in mines. Each agent can represent a type of equipment (such as a forklift, a transport vehicle, etc.), and make an overall optimal decision considering factors such as operation safety, its own state, task requirements, and environmental changes.

[0004] If digital twin technology and multi-agent systems are applied to open-pit mines, it will be beneficial to solve the problems encountered in open-pit mine collaborative scheduling. However, the current mine modeling methods under digital twin technology usually only focus on the mine resources themselves, and do not model the mine resources and digital intelligent construction equipment such as driverless trucks and driverless loaders as a whole, resulting in problems such as insufficient model accuracy and delayed data update. And the reinforcement learning method based on multi-agents has not been fully studied and applied in the field of mine construction. Currently, mine scheduling usually adopts single-agent reinforcement learning, which only focuses on the optimal decision of its own vehicles in the mine operation area, and often faces problems such as insufficient coordination, lack of a global perspective, and weak self-adaptability, and it is difficult to achieve the optimization of overall efficiency. Summary of the Invention

[0005] The objective of the present invention is to provide a vehicle scheduling and collaborative control system and method for open-pit mines. Through digital twin, multi-agent reinforcement learning, and optimization scheduling algorithms, dynamic path planning and multi-task collaborative scheduling in the open-pit mine scenario are realized, and efficient path selection and construction cost optimization are achieved in the dynamic mine environment.

[0006] To achieve the above objective / To solve the above technical problems, the present invention is implemented by adopting the following technical solutions:

[0007] In the first aspect, the present invention provides a vehicle scheduling and collaborative control system for open-pit mines, including:

[0008] The device management layer is used to obtain the environmental data and device status data of the open-pit mine, and is also used to schedule and control the vehicle equipment of the open-pit mine according to the control instructions issued by the cloud or the edge layer;

[0009] The edge layer is used to preprocess and preliminarily analyze the data uploaded by the device management layer, upload the preprocessed data and the preliminary analysis results to the cloud, and generate real-time control instructions according to the preliminary analysis results of the open-pit mine;

[0010] The cloud is used to construct a digital twin environment of the open-pit mine containing mine resources and mine vehicle information according to the data uploaded by the edge layer, and use multi-objective optimization, heuristic algorithms, and multi-agent reinforcement learning algorithms in the digital twin environment of the open-pit mine to calculate the vehicle collaborative scheduling scheme and the optimal vehicle path of the open-pit mine, and perform vehicle trajectory tracking control;

[0011] The security access control layer is used to forward the control instructions issued by the cloud or the edge layer to the device management layer, and perform dynamic security authentication and access control on the edge layer and the cloud during the forwarding process.

[0012] In combination with the first aspect, further, the digital twin environment of the open-pit mine includes mine resource information, production plans, and mine vehicle information. Among them, the mine vehicle information includes the type of vehicle, the affiliated formation, fuel / olectricity quantity, motion state, working state, loading state, and driver information.

[0013] In combination with the first aspect, further, the edge layer performs preliminary analysis on the data uploaded by the device management layer, including vehicle collision avoidance warning, vehicle emergency braking, equipment health grading, and basic energy efficiency calculation;

[0014] The specific operation of the vehicle collision avoidance warning is: calculating the distance between different construction vehicles based on the real-time positions of the vehicles uploaded by the device management layer, and generating a vehicle collision avoidance warning instruction according to the distance;

[0015] The specific operation of the vehicle emergency braking is as follows: according to the preset safety threshold and the data uploaded by the device management layer, when a certain value breaks through the preset safety threshold, an emergency braking instruction is generated.

[0016] The specific operation of the device health grading is as follows: based on the data uploaded by the device management layer, judge the device health status according to the preset rules, and divide the device health status into normal, warning, and failure;

[0017] The specific operation of the basic energy efficiency calculation is as follows: based on the data uploaded by the device management layer and the current energy consumption rate, calculate and display the remaining battery life of the device in real time.

[0018] Combined with the first aspect, further, the security access control layer adopts zero-trust security technology. When the cloud or edge layer controls the device, a control access request is sent to the security access control layer. The trust evaluation engine conducts a trust evaluation on the access request, establishes a one-time secure access connection according to the trust evaluation result, and uses the access control engine to perform zero-credit policy decision calculation to achieve access permission control.

[0019] In the second aspect, based on the vehicle scheduling and coordination control system for open-pit mines provided in the first aspect, the present invention also provides a vehicle scheduling and coordination control method for open-pit mines, including the following steps:

[0020] Build a digital twin environment for the open-pit mine using digital twin technology according to the open-pit mine data;

[0021] Calculate the vehicle coordination scheduling plan using a vehicle scheduling model based on multi-objective optimization in the digital twin environment of the open-pit mine;

[0022] According to the vehicle coordination scheduling plan, use a heuristic algorithm to perform vehicle path planning to obtain the optimal vehicle path;

[0023] Control the operation of the construction vehicle according to the optimal vehicle path, and use a vehicle trajectory tracking control model based on multi-agent reinforcement learning to perform vehicle trajectory tracking control.

[0024] Combined with the second aspect, further, the calculation of the vehicle coordination scheduling plan using a vehicle scheduling model based on multi-objective optimization in the digital twin environment of the open-pit mine includes:

[0025] Take each possible vehicle coordination scheduling plan as an individual to generate an initial population;

[0026] Calculate the fitness value of each individual using a vehicle scheduling model based on multi-objective optimization;

[0027] Update the individual according to the fitness value of each individual;

[0028] Merge the updated individuals with the individuals before the update, perform non-dominated sorting, and select the next generation population according to the crowding distance;

[0029] In each iteration process, select the optimal solution from the Pareto front solutions by the entropy weight TOPSIS method;

[0030] Repeat the iteration until the iteration termination condition is satisfied, and obtain the vehicle collaborative scheduling scheme according to the optimal solution.

[0031] Combined with the second aspect, further, the updating of the individual according to the fitness value of each individual includes:

[0032] Given a random number r in the range of [0, 1];

[0033] When the random number r ≤ 0.2, update the individual by using the crossover operation based on BWO, and the expression is as follows:

[0034] ;

[0035] Where represents the individual i after the (t + 1)-th update, and are two randomly selected individuals respectively, and val is a random value of 0 or 1;

[0036] When the random number 0.2 < r ≤ 0.8, update the individual by using the global search based on SO, and the expression is as follows:

[0037] ;

[0038] Where is a random function, and upper and lower are the upper and lower boundaries of the decision variable respectively;

[0039] When the random number r > 0.8, update the individual by using the levy flight, and the expression is as follows:

[0040] ;

[0041] Where is the best individual of the current population, and LevyFlight is a random step size used to enhance the global search ability.

[0042] Combined with the second aspect, further, in each iteration process, selecting the optimal solution from the Pareto front solutions by the entropy weight TOPSIS method includes:

[0043] Perform normalization processing according to the objective types of each objective function in the vehicle scheduling model to obtain the decision matrix;

[0044] Calculate the objective weights of each objective by the entropy weight TOPSIS method;

[0045] Obtain the weighted decision matrix based on the objective weights of each objective and the decision matrix;

[0046] In each iteration process, calculate the Euclidean distances from each individual in the population to the positive ideal solution and the negative ideal solution, and obtain the ideal solution closeness of each individual;

[0047] Select the individual with the maximum closeness as the local optimal solution;

[0048] Update the global optimal solution according to the local optimal solution.

[0049] Combined with the second aspect, further, regard the vehicles in the open-pit mine as agents, define the state information of the vehicle itself and the state information of all observable vehicles within the vehicle's perception range as the state space of the agent corresponding to the vehicle, define the vehicle steering angle and acceleration as the action space of the agent corresponding to the vehicle, and construct a vehicle trajectory tracking control model based on multi-agent reinforcement learning by designing a reward function, building an action network and a critic network.

[0050] Combined with the second aspect, further, the agent The reward function at time step t Is defined as:

[0051] ;

[0052] Among them, Represents the collision reward, Represents the arrival reward, Represents the tracking accuracy reward, Represents the efficiency reward, Represents the comfort reward, Are the weight coefficients of the tracking accuracy reward 、efficiency reward And comfort reward Respectively;

[0053] The calculation formula of the tracking accuracy reward Is:

[0054] ;

[0055] Among them, Is the distance from the vehicle center point to the planned path;

[0056] The calculation formula of the efficiency reward Is:

[0057] ;

[0058] Among them, is the speed of the vehicle itself, , are respectively the maximum speed limit and the minimum speed limit in the construction scenario;

[0059] Comfort reward The calculation formula is:

[0060] ;

[0061] Among them, is the acceleration at the current moment, is the acceleration at the previous moment, are respectively the maximum and minimum accelerations.

[0062] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0063] The present invention proposes a vehicle scheduling and collaborative control system and method for open-pit mines, models and simulates mine resources and equipment such as mine vehicles, realizes low-latency real-time operation and feedback through algorithms deployed in the edge layer, realizes vehicle scheduling, path optimization and collaborative control operations with a certain latency through various algorithms deployed in the cloud, and finally realizes the safety control of mine vehicles through a security access control layer with a security authentication function, improving the overall operation effect of open-pit mines.

[0064] The present invention deploys a vehicle scheduling model based on multi-objective optimization, a heuristic algorithm for vehicle path planning, and a multi-agent reinforcement learning algorithm for vehicle trajectory tracking control in the cloud. Through the algorithm based on multi-objective optimization, a vehicle coordination and scheduling scheme for different tasks is obtained. Through the heuristic algorithm, the optimal path is planned for each vehicle. Path optimization enables each vehicle to not only optimize its own path when selecting a path, but also avoid resource conflicts and congestion problems through mutual cooperation, thereby achieving the global optimum of path planning. Finally, by training each agent in the digital twin environment, the present invention deploys a high-precision and low-risk path planning strategy to the mine site, greatly improving the intelligent level of vehicle scheduling and path planning in open-pit mines, enhancing the overall collaborative control ability of the mine, further balancing the construction efficiency and cost of the mine, and ensuring the economy and efficiency of production in the mine scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 FIG. shows a schematic structural diagram of a vehicle scheduling and collaborative control system for an open-pit mine provided by an embodiment of the present invention;

[0066] Figure 2 FIG. shows a schematic diagram of the steps of a vehicle scheduling and collaborative control method for an open-pit mine provided by an embodiment of the present invention;

[0067] Figure 3 The figure shows a schematic diagram of the network structure of the multi-agent reinforcement learning model in an embodiment of the present invention;

[0068] Figure 4 The figure shows a schematic diagram of using the multi-agent reinforcement learning model for vehicle trajectory tracking control in an embodiment of the present invention. Specific embodiments

[0069] The technical solution of the present invention will be described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.

[0070] Embodiment 1

[0071] This embodiment introduces a vehicle scheduling and collaborative control system for open-pit mines, as Figure 1 shown, the main architecture of the system includes a device management layer, an edge layer, a cloud, and a secure access control layer.

[0072] The device management layer is the connection bridge between most of the production equipment (mainly engineering vehicles) in the open-pit mine, other working systems and the system of the present invention. An equipment Internet of Things network is constructed through an industrial controller (PLC) and intelligent sensing terminals to realize environmental perception, vehicle monitoring, and instruction execution. In terms of environmental perception, the present invention deploys a temperature and humidity, vibration, and displacement sensor group to collect mine geological environment parameters in real time. In terms of vehicle monitoring, the present invention integrates an on-vehicle controller to obtain device status data such as positioning, load, and energy consumption.

[0073] The device management layer collects environmental and device status data in open-pit mine operations, monitors information such as the positions and working states of different vehicles in the mining area in real time, and uploads it to the edge layer to form business data, providing data support for vehicle path planning and scheduling optimization, making the device information comprehensively and real-time available. The device management layer can also control the mine vehicles according to the control instructions issued by the cloud or the edge layer to realize various functions such as task planning, task allocation, and vehicle scheduling. The device management layer establishes a closed-loop control interface for vehicle steering, throttle, and braking in order to execute instructions.

[0074] The device management layer encapsulates the original sensor data using an industrial Internet of Things protocol (such as OPC UA) to achieve uplink transmission; it receives the scheduling instructions from the cloud / edge layer through an encrypted instruction channel to achieve downlink control.

[0075] The edge layer is the first layer of data processing in the system of the present invention and is generally located at the mine construction site, where it can provide real-time data analysis functions. The edge layer is mainly used to receive the data uploaded by the device management layer, preprocess and preliminarily analyze and feedback the data. By preprocessing the collected sensor data, the preprocessed data is obtained, and then the preprocessed data is uploaded to the cloud to reduce the transmission load. The edge layer can also perform preliminary task calculations based on the preprocessed data and store the preprocessed data and the preliminary analysis results in the edge-side database, maintaining a connection with the cloud database.

[0076] The operations for the edge layer to perform data preprocessing include: performing data cleaning to eliminate outliers and compensate for missing data; implementing spatio-temporal alignment to unify multi-source data to the same time reference and coordinate system; and performing data compression to reduce the amount of transmitted data using differential coding technology.

[0077] When performing certain real-time tasks, the edge layer calculates based on the preprocessed data and task requirements, generates key control instructions according to the calculation results, and quickly feeds back the key control instructions to the device management layer via the secure access control layer to control the vehicle actions to meet the real-time requirements of the tasks and achieve rapid response.

[0078] In the embodiment of the present invention, the content of the preliminary analysis performed by the edge layer mainly includes millisecond-level safety response and status simplification evaluation. The millisecond-level safety response includes vehicle collision prevention warning and vehicle emergency braking, etc., and the status simplification evaluation includes equipment health grading and basic energy efficiency calculation, etc.

[0079] The specific operation of the vehicle collision prevention warning is: calculating the distance between different construction vehicles based on the real-time positions of the vehicles uploaded by the device management layer, and generating a vehicle collision prevention warning instruction according to the distance.

[0080] The specific operation of the vehicle emergency braking is: according to the preset safety threshold and the data uploaded by the device management layer, when a certain value breaks through the preset safety threshold, an emergency braking instruction is generated.

[0081] The specific operation of the equipment health grading is: based on the data uploaded by the device management layer, judging the equipment health status according to the preset rules, and classifying the equipment health status into three levels: normal, warning, and failure.

[0082] The specific operation of the basic energy efficiency calculation is: calculating and displaying the remaining endurance time of the equipment in real time based on the data uploaded by the device management layer and the current energy consumption rate.

[0083] The present invention sets up edge-cloud collaboration and establishes a hierarchical processing mechanism. The edge layer can periodically upload regular data to the cloud in batches and push emergency event data to the cloud in real time to trigger emergency response. In addition, the edge side model of the edge layer is connected to the cloud model in the cloud, supporting synchronous updates with the cloud to ensure the real-time execution of the latest path planning and scheduling strategies.

[0084] The cloud is mainly used to conduct in-depth analysis and data processing of the collaborative control tasks of open-pit mines based on the pre-processed data and preliminary analysis results uploaded by the edge layer. The cloud uses digital twin technology to build an open-pit mine digital twin environment containing mine vehicle information, and uses multi-objective optimization, heuristic algorithms and multi-agent reinforcement learning algorithms in the open-pit mine digital twin environment to calculate the optimal paths of different vehicles in the open-pit mine and the collaborative scheduling schemes of different tasks, and perform vehicle trajectory tracking control. The calculation process and results of relevant business data will be stored in the cloud database for query, display, and statistical analysis. The various computing models encapsulated in the cloud can be processed in large quantities based on the preliminary calculation results of the edge layer and the pre-processed data, and the required cloud non-real-time high-precision processing results are sent to the device management layer via the security access control layer.

[0085] In an embodiment of the present invention, the operations of digital twin modeling in the cloud include: 1. Environmental model: integrating geological scanning data and real-time monitoring information to build a three-dimensional mine model; 2. Vehicle model: establishing a parameterized digital mirror including mechanical characteristics and energy consumption patterns; 3. Dynamic simulation system: simulating vehicle movement and operation processes based on a physical engine, and building virtual sensors to achieve virtual-real environment state mapping.

[0086] Unlike the mine modeling methods under the digital twin technology in the prior art, which usually only focus on the mine resources themselves, the present invention can model and simulate mine resources and mine engineering machinery. The open-pit mine digital twin environment includes mine resource information, production plans and mine engineering vehicle information, among which the mine engineering vehicle information includes the type of vehicle, its formation, fuel / power level, movement status, working status, cargo status and driver information, etc.

[0087] The cloud combines transportation tasks, equipment status, and path topology for multi-objective optimization, and uses a hybrid planning method to solve vehicle scheduling and path planning problems. Based on the results of digital twin simulation, equipment anomalies and efficiency bottlenecks are predicted, and production plans and equipment operating parameters are dynamically adjusted.

[0088] Construction workers or smart devices in open-pit mines can combine real-time feedback data from the edge layer and high-precision feedback data from the cloud with a certain delay to make corresponding decisions, thereby better realizing vehicle scheduling and collaborative control.

[0089] The security access control layer is connected between the cloud and the device management layer, and between the edge layer and the device management layer. It is used to forward the control instructions sent from the cloud or the edge layer to the device management layer, and perform dynamic security authentication and access control on the edge layer and the cloud during the forwarding process to ensure the security of mine data transmission.

[0090] The security authentication system of the system of the present invention mainly includes dynamic identity verification, two-way certificate authentication during device access, regular rotation and update of session keys, hierarchical permission management, establishment of a multi-dimensional permission matrix of device-operation-personnel, and multiple approvals and confirmations required for the execution of key instructions, etc.

[0091] The system of the present invention introduces zero-trust security technology. When a user controls a device through the cloud or the edge layer, a control access request needs to be sent to the security access control layer. The trust evaluation engine and the access control engine are used to implement dynamic identity authentication and authorization. Once the access request is permitted, the system dynamically configures the data plane, and the access proxy accepts the traffic data of the access subject and establishes a one-time secure access connection. During the secure access process, the trust evaluation engine continuously conducts trust evaluation and provides the evaluation data to the access control engine for zero-trust policy decision calculation to determine whether the access control policy needs to be changed. When necessary, the connection is disconnected in a timely manner through the access proxy to quickly achieve resource protection. Based on dynamic trust evaluation and permission control, the present invention can effectively prevent unauthorized access and ensure the secure transmission of data in a complex mine environment.

[0092] In the digital twin environment of the cloud, the present invention constructs a virtual intelligent body for each mine vehicle and uses the reinforcement learning method to achieve the collaborative optimization of the paths of multiple intelligent bodies. Each intelligent body will fully consider the strategies of other intelligent bodies when choosing a path to avoid path overlap and resource conflicts. Based on the Markov decision process (MDP) of reinforcement learning, each intelligent body conducts collaborative training with the goal of improving transportation efficiency, avoiding congestion, and reducing energy consumption, and finally obtains the optimal strategy for the vehicle path.

[0093] During the reinforcement learning training process, the system of the present invention designs a multi-objective reward mechanism for the intelligent body, such as reducing the distance between vehicles, optimizing the path length, and reducing fuel consumption. Through the shared reward mechanism and the distributed value function, each vehicle intelligent body will influence each other when choosing a path to maximize the overall scheduling efficiency.

[0094] The present invention deploys an optimization scheduling system based on heuristic algorithms (such as genetic algorithms, particle swarm optimization algorithms, simulated annealing, etc.) in the cloud. This system combines on-site real-time data and production plans to obtain the Pareto optimal solution of construction efficiency and construction cost, and finally outputs a scheduling plan with global optimality. Task instructions are sent to the mine site in real time through the V2X network (vehicle-to-everything network) to achieve closed-loop management of scheduling.

[0095] The system of the present invention establishes a hierarchical encryption strategy, adopts a lightweight encryption algorithm in the device management layer to ensure the security of terminal communication, and uses a national cryptographic standard algorithm for data transmission encryption in the edge-cloud channel. A traffic monitoring system is deployed to identify abnormal access patterns and establish an instruction whitelist to filter illegal control requests.

[0096] Embodiment 2

[0097] Based on the system introduced in Embodiment 1, this embodiment introduces a vehicle scheduling and cooperative control method for open-pit mines, as Figure 2 shown, including the following steps:

[0098] Step A: Utilize digital twin technology to build a digital twin environment for open-pit mines based on elements such as mine resources, types of engineering vehicles, motion states, working states, task plans, load capacities, driver information, etc.

[0099] Compared with the prior art, when building the virtual open-pit mine environment, the present invention fully considers the information of each intelligent device in the mine, improves the authenticity and integrity of the virtual environment, and provides stronger support for vehicle scheduling and cooperative control.

[0100] Step B: According to the data collected by the device management layer, calculate the vehicle cooperative scheduling plan in the digital twin environment of the open-pit mine using a pre-constructed vehicle scheduling model based on multi-objective optimization. The vehicle cooperative scheduling plan includes the scheduling time, scheduling order, target point coordinates, etc. of each vehicle in the open-pit mine.

[0101] The present invention determines the constraint conditions with the transportation task, equipment status, and path topology as the objectives according to the requirements of open-pit mine equipment scheduling and control, and establishes a vehicle scheduling model based on multi-objective optimization.

[0102] In the embodiment of the present invention, taking the charging scheduling of electric construction vehicles as an example, a vehicle scheduling model related to charging is built. According to actual requirements, the charging vehicle scheduling model mainly includes a vehicle charging and discharging model and a charging station model. Among them, the vehicle charging and discharging model mainly considers the battery capacity, charging status, charging speed, arrival time, etc. of each vehicle to reasonably arrange charging resources and control the charging process. The battery capacity of each vehicle is usually set to a fixed value, with the unit of kWh.

[0103] In the embodiment of the present invention, the charging status represents the proportion of the current battery power to the total battery capacity, and is defined as:

[0104] (1)

[0105] Where represents the charging status of vehicle at moment, Represents the vehicle At The current battery capacity at the moment, Represents the vehicle The total battery capacity of.

[0106] For an uncharged vehicle, the power consumption per time step is set to (Unit: kWh). Therefore, the update formula for the vehicle's battery level when not charging is:

[0107] (2)

[0108] Among them, Represents the battery level of vehicle i at At the moment. When charging, the charging amount of the vehicle per time step is r (unit: kWh). The update formula for the battery level after charging is:

[0109] (3)

[0110] In the charging station model, it is assumed that the charging station purchases electricity from the regional power system and combines a photovoltaic system and an energy storage system to supply power to construction vehicles. The energy storage system can achieve the purpose of cost savings by charging during off-peak electricity prices and discharging during peak electricity prices. The charging station transfers energy to the construction vehicle through a charging pile. The calculation formula for the vehicle's charging demand is as follows:

[0111] (4)

[0112] Among them, Represents the vehicle The charging demand of, Represents the vehicle Battery capacity, Is the target charging state (e.g., 90%), For the vehicle The current battery level.

[0113] Within each time step, the update formula for the battery level of the charging vehicle is as follows:

[0114] (5)

[0115] Among them, For the vehicle At The current charging power at the moment, Is the time step size.

[0116] Set the charging optimization objective function and introduce constraint conditions based on the above formulas.

[0117] The charging vehicle scheduling model mainly includes two objectives. Objective 1: Maximize the total driving time. Objective 2: Minimize the total charging cost. Among them, the expression of the total driving time is as follows:

[0118] (6)

[0119] Among them, represents the battery level of vehicle at the next time step, and N is the total number of vehicles.

[0120] Formula (6) represents the sum of the available driving times of all vehicles (the negative sign is taken because this objective needs to be maximized in the optimization algorithm).

[0121] The expression of the total charging cost is as follows:

[0122] (7)

[0123] Among them, is the electricity price (per kilowatt-hour), is whether vehicle i is charging at the current time step, = 0 indicates that vehicle i is not charging, = 1 indicates that vehicle i is charging.

[0124] Formula (7) calculates the total charging cost of all vehicles at the current time step.

[0125] The constraint conditions of the charging vehicle scheduling model are as follows:

[0126] Constraint 1: Limit on the number of charging vehicles:

[0127] (8)

[0128] Among them, represents Constraint 1, and U represents the total number of charging piles.

[0129] Constraint 2: The battery level of the vehicle is limited between 30% and 90%:

[0130] (9)

[0131] Among them, represents Constraint 2.

[0132] Based on Formulas (6) - (9), use the multi-objective stochastic competition optimization algorithm to solve the optimal allocation plan for the charging of construction vehicles in open-pit mines, that is, the vehicle collaborative scheduling plan.

[0133] Using the multi-objective stochastic competition optimization algorithm to solve the optimal vehicle collaborative scheduling plan, the main steps include:

[0134] Step B01: Generate an initial population, where each individual represents a possible vehicle collaborative scheduling scheme.

[0135] Step B02: Calculate the fitness of each individual. Here, the fitness function uses a charging scheduling model (a vehicle scheduling model based on multi-objective optimization). The optimization objectives of the present invention include maximizing the driving time and minimizing the charging cost.

[0136] Step B03: Update the individuals according to the fitness value of each individual, generate offspring individuals, and prevent them from exceeding the boundaries.

[0137] In the embodiment of the present invention, the core idea of population update lies in combining the global search ability of BWO and the dynamic exploration strategy of snake optimization (SO), and introducing the Levy flight strategy to avoid local optima.

[0138] The present invention gives a random number r within the range of [0, 1]. When the random number r ≤ 0.2, the individuals are randomly updated using the crossover operation based on BWO, and the expression is as follows:

[0139] (10)

[0140] Where, represents the individual i after the (t + 1)-th update, and are two randomly selected individuals respectively, and val is a random value of 0 or 1.

[0141] When the random number 0.2 < r ≤ 0.8, the individuals are updated using the global search based on SO. Specifically, the individuals are updated based on the boundaries, and the expression is as follows:

[0142] (11)

[0143] Where, is a random function, = r, and upper and lower are the upper and lower boundaries of the decision variable (such as the charging instruction of each vehicle) respectively.

[0144] When the random number r > 0.8, the individuals are updated using the Levy flight, and the expression is as follows:

[0145] (12)

[0146] Where, is the best individual of the current population, and LevyFlight is a random step size used to enhance the global search ability.

[0147] Step B04: Merge the individuals before update and after update, perform non-dominated sorting, and select the next-generation population according to the crowding distance.

[0148] Step B05: Select the optimal solution from the Pareto front solutions by the entropy weight TOPSIS method.

[0149] (1) Normalize the decision matrix according to the objective types (maximization / minimization) of each objective function in the vehicle scheduling model.

[0150] (2) Calculate the objective weights of each objective by the entropy weight TOPSIS method.

[0151] (3) Obtain the weighted decision matrix based on the objective weights of each objective and the decision matrix.

[0152] (4) In each iteration process, calculate the Euclidean distances from each individual in the population to the positive ideal solution (PIS) and the negative ideal solution (NIS), and obtain the ideal solution closeness degree of each individual.

[0153] (5) Select the individual with the maximum closeness degree as the local optimal solution.

[0154] (6) Update the global optimal solution according to the local optimal solution.

[0155] Step B06: Repeat Step B02 - Step B05 until the iteration termination condition is met, end the iteration, and obtain the optimal vehicle collaborative scheduling plan according to the global optimal solution. Through the above steps, the optimal vehicle collaborative scheduling plan is obtained. The optimal plan of the charging vehicle scheduling model includes the start charging time, end charging time, target charging pile coordinates (i.e., target points) of each vehicle in the open-pit mine, etc. According to the optimal charging allocation plan, the cloud can issue instructions for charging and completion of charging to each vehicle, and at the same time send the coordinates of the target point to each vehicle.

[0156] Step C: According to the target point coordinates in the vehicle collaborative scheduling plan, use the heuristic algorithm to perform vehicle path planning to obtain the optimal vehicle path.

[0157] The present invention can use variants of the A* algorithm or Dijkstra algorithm for vehicle path planning. In the embodiment of the present invention, taking the A* algorithm as an example, the specific operations of Step C are as follows:

[0158] Step C01: Obtain the vehicle scheduling target point coordinates according to the vehicle collaborative scheduling plan, set the starting position of the vehicle to be path-planned as the starting point, and set the target point coordinates as the end point. Generate a topological map according to the GNSS coordinates of the actual road. On the topological map, each node represents the intersection or key position of the road, and the edge represents the road segment and its connectivity.

[0159] Step C02: Add the starting point to the open list and initialize the heuristic value and cumulative cost of all nodes in the topological map.

[0160] Common heuristic estimation functions are Euclidean distance or Manhattan distance (more suitable for urban road networks). Taking Euclidean distance as an example, the formula is as follows:

[0161] (13)

[0162] Among them, represents the estimated cost from the current node to the target node, and represent the coordinates of the current node and the target node respectively.

[0163] Cumulative cost represents the actual cost from the starting point to the current node 𝑛. A common calculation method is to accumulate the weights of each edge (such as distance, time, or consumption). The formula is as follows:

[0164] (14)

[0165] Among them, is the parent node of the current node, that is, the previous node in the search path, represents the actual cumulative cost from the starting point to the parent node, and t is time.

[0166] The total cost of each node is composed of the sum of the heuristic value and the cumulative cost, and is used to sort the node priorities in the open list. The formula for the total cost is as follows:

[0167] (15)

[0168] Step C03: Select the node with the smallest value in the open list for expansion. If the node has not been visited or the cumulative cost of the new path is smaller, update the path and cost information of the node and add its adjacent nodes to the open list.

[0169] Step C04: Repeat Step C03 until the path termination condition is met, and output the optimal path of the vehicle.

[0170] In the embodiments of the present invention, the path termination conditions include: when the end point is expanded or the open list is empty. If the end point is expanded, trace the path to determine the final path.

[0171] Step D: According to the optimal vehicle path obtained in Step C, control instructions are sent from the cloud to control the operation of the construction vehicle. At the same time, the vehicle trajectory tracking and control model based on multi-agent reinforcement learning pre-constructed is used to perform vehicle trajectory tracking and control on the operating construction vehicle.

[0172] In the present invention, the intelligent vehicle in the open-pit mine is regarded as an agent, the state and action space of the agent are defined, the reward function is designed, the action network and the critic network are built, and the vehicle trajectory tracking control model (MADDPG model) based on multi-agent reinforcement learning is constructed.

[0173] The structures of the action network and the critic network in the vehicle trajectory tracking control model are as Figure 3 shown. Based on the Actor-Critic framework in the present invention, the execution actions and state spaces of all agents are aggregated into a shared Critic network for training. During the training process, the Critic network gives guidance to the Actor network of each agent at the same time. When executing, the Actor network of each agent outputs the execution action completely independently.

[0174] In the embodiment of the present invention, the state space of the agent is defined as the state information of the vehicle itself and the state information of all observable vehicles within the vehicle's perception range, which is a matrix with a dimension of M×F. Among them, M is the number of all observable vehicles (including the vehicle itself) of the agent and F is the number of feature dimensions of the vehicle state. The feature vector of the vehicle is expressed as:

[0175] Vehicle is expressed as:

[0176] (16)

[0177] Among them, indicates whether vehicle k is observable by the vehicle itself, is the lateral position of vehicle k, is the longitudinal position of vehicle k, is the speed of vehicle k, is the heading angle of vehicle k, is the lane id where vehicle k is located, is the lateral distance between the center point of the vehicle and the target path.

[0178] The global state space is defined as the joint state space of all agents within a certain control range, which is expressed as:

[0179] (17)

[0180] Among them, represents the state space of vehicle v, represents the number of all controlled vehicles.

[0181] The action space of the agent is defined as a set of high-level behavior decisions, including vehicle steering angle and acceleration. The global action space is defined as the action combination of all agents within a certain control range, expressed as:

[0182] (18)

[0183] Among them, represents the vehicle's action space.

[0184] The reward function of the agent at time step t is defined as:

[0185] (19)

[0186] Among them, represents the collision reward, represents the arrival reward, represents the tracking accuracy reward, represents the efficiency reward, represents the comfort reward,

[0187] are the weight coefficients of the tracking accuracy reward

[0188]

[0189]

[0190]

[0191]

[0192]

[0193] ​​​​​​​​​​​​​​​(22)

[0194] Efficiency reward The calculation formula is:

[0195] (23)

[0196] Wherein, is the vehicle speed of the own vehicle, , are respectively the maximum speed limit and the minimum speed limit in the construction scenario. In the present invention, the minimum speed limit is defined as 0.

[0197] Comfort reward The calculation formula is:

[0198] (24)

[0199] Wherein, is the acceleration at the current moment, is the acceleration at the previous moment, are respectively the maximum and minimum accelerations.

[0200] During training, the parameters of each Actor network are , and are updated through the policy gradient of the expected return of each agent, specifically expressed as:

[0201] (25)

[0202] Wherein, D is the experience replay pool for storing data, and each stored data form is .

[0203] The parameters of the centralized training Critic network are , and are updated through the loss function, expressed as:

[0204] (26)

[0205] The present invention trains the vehicle trajectory tracking control model based on the above loss function, updates the model parameters, and obtains the trained vehicle trajectory tracking control model. Multi-agent reinforcement learning can more comprehensively consider the cooperation between vehicles and reduce the potential collision risk between vehicles. Effective information exchange can be carried out between agents to coordinate their respective actions, thereby improving the overall traffic efficiency and safety.

[0206] In the embodiments of the present invention, each agent obtains its own vehicle state information and the state information of other vehicles within the sensing range through the constructed digital twin environment, so as to obtain the state space of each agent. The vehicle state information includes the lateral position, longitudinal position, speed, heading angle of the vehicle, and the identifier of the lane where it is located.

[0207] Based on the state space of each agent, the trained vehicle trajectory tracking control model is used to perform trajectory tracking control on the mine construction vehicle. In the embodiments of the present invention, the execution actions output by the multi-agent deep reinforcement learning MADDPG model are executed by the simulation environment. When the execution in the simulation environment is correct, they are then sent by the cloud to the construction vehicle. The trajectory tracking process is as Figure 4 shown.

[0208] In summary, the embodiments of the present invention aim to provide a new digital twin-based intelligent system framework for mine construction, which can model and simulate mine resources and mine construction machinery, and can dispatch vehicles through algorithms deployed in the cloud to realize the vehicle dispatching and cooperative control functions of open-pit mines.

[0209] The present invention realizes the modeling of the open-pit mine construction scenario based on the digital twin technology. When modeling, it fully considers the mine resources and various intelligent mine construction vehicles, and establishes a cloud-edge collaborative system based on the digital twin, making full use of the computing power of the edge layer and the cloud to realize shallow-layer real-time operation and deep-layer high-precision operation, improving the utilization rate of system resources and the vehicle dispatching control performance. The system of the present invention also introduces a zero-trust architecture to enhance the security of system data transmission.

[0210] At the method level, the present invention enables each vehicle to not only optimize its own path when selecting a path through a heuristic algorithm, but also avoid resource conflicts and congestion problems through mutual cooperation, improving the rationality and reliability of path planning. The present invention uses a vehicle trajectory tracking control model based on multi-agent reinforcement learning to real-time track the running trajectories of each construction vehicle in the open-pit mine, considering factors such as overall traffic efficiency, safety, and path tracking accuracy, which is beneficial to improving vehicle dispatching performance.

[0211] The architecture of the present invention can realize various functions such as virtual model of open-pit mines, optimization of vehicle cooperative dispatching schemes, optimization of vehicle paths, and vehicle trajectory tracking control. Through vehicle dispatching and cooperative control from a global perspective, it improves the operational coordination and self-adaptability of open-pit mines, and realizes safe, real-time, and efficient intelligent dispatching management of mines.

[0212] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0213] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0214] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0215] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0216] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present invention and without departing from the spirit and scope protected by the present invention's claims, can still make many forms, and all of these fall within the protection scope of the present invention.

Claims

1. A vehicle scheduling and collaborative control system for open-pit mines, characterized in that, include: The equipment management layer is used to obtain the environmental data and equipment status data of the open-pit mine, and is also used to dispatch and control the vehicles and equipment of the open-pit mine according to the control instructions issued by the cloud or edge layer; The edge layer is used to pre-process and preliminarily analyze the data uploaded by the equipment management layer, upload the pre-processed data and preliminary analysis results to the cloud, and generate real-time control instructions based on the preliminary analysis results of the open-pit mine; The cloud is used to build an open-pit mine digital twin environment containing mine resources and mine vehicle information based on the data uploaded by the edge layer. In the open-pit mine digital twin environment, multi-objective optimization, heuristic algorithms, and multi-agent reinforcement learning algorithms are used to calculate the open-pit mine vehicle collaborative scheduling plan and the optimal vehicle path, and to perform vehicle trajectory tracking control; The security access control layer is used to forward the control instructions issued by the cloud or edge layer to the device management layer, and perform dynamic security authentication and access control on the edge layer and the cloud during the forwarding process.

2. The vehicle scheduling and coordination control system according to claim 1, characterized in that The open-pit mine digital twin environment includes mine resource information, production plan and mine vehicle information, wherein the mine vehicle information includes vehicle type, formation, fuel / power level, movement status, working status, cargo status and driver information.

3. The vehicle scheduling and collaboration control system according to claim 1, wherein The edge layer performs preliminary analysis on the data uploaded by the equipment management layer, including vehicle anti-collision warning, vehicle emergency braking, equipment health classification and basic energy efficiency calculation; The specific operation of the vehicle anti-collision warning is: calculating the distance between different construction vehicles based on the real-time position of the vehicle uploaded by the equipment management layer, and generating a vehicle anti-collision warning instruction according to the distance; The specific operation of the vehicle emergency braking is: according to the preset safety threshold and the data uploaded by the device management layer, when a certain value exceeds the preset safety threshold, an emergency braking instruction is generated; The specific operation of the equipment health classification is: based on the data uploaded by the equipment management layer, the equipment health status is judged according to the preset rules, and the equipment health status is classified into normal, warning, and fault; The specific operation of the basic energy efficiency calculation is: based on the data uploaded by the device management layer and the current energy consumption rate, calculate and display the remaining battery life of the device in real time.

4. The vehicle scheduling and collaboration control system according to claim 1, characterized in that The security access control layer adopts zero-trust security technology. When the cloud or edge layer controls the device, a control access request is initiated to the security access control layer, and the access request is trusted by the trust assessment engine. A one-time secure access connection is established based on the trust assessment result, and the access control engine is used to perform zero-trust policy decision calculations to achieve access permission control.

5. A method for a vehicle scheduling and collaborative control system of the open-pit mine according to claim 1, characterized in that, The steps include: Based on the open-pit mine data, use digital twin technology to build an open-pit mine digital twin environment; Calculate the vehicle collaborative scheduling scheme using a vehicle scheduling model based on multi-objective optimization in the open-pit mine digital twin environment; According to the vehicle collaborative dispatching scheme, the heuristic algorithm is used to plan the vehicle path and obtain the optimal vehicle path; The operation of construction vehicles is controlled according to the optimal vehicle path, and the vehicle trajectory tracking control is performed using a vehicle trajectory tracking control model based on multi-agent reinforcement learning.

6. The vehicle scheduling and cooperative control method for open-pit mines according to claim 5, characterized in that, Calculating the vehicle collaborative scheduling plan using the vehicle scheduling model based on multi-objective optimization in the digital twin environment of open-pit mines, including: Regarding each possible vehicle collaborative scheduling plan as an individual to generate an initial population; Calculating the fitness value of each individual using the vehicle scheduling model based on multi-objective optimization; Updating the individual according to the fitness value of each individual; Performing non-dominated sorting after merging the updated individuals with the pre-updated individuals, and selecting the next-generation population according to the crowding distance; In each iteration process, selecting the optimal solution from the Pareto front solutions through the entropy weight TOPSIS method; Repeating the iteration until the iteration termination condition is met, and obtaining the vehicle collaborative scheduling plan according to the optimal solution.

7. The vehicle scheduling and cooperative control method for open-pit mines according to claim 6, characterized in that, The updating the individual according to the fitness value of each individual includes: Giving a random number r within the range of [0, 1]; When the random number r ≤ 0.2, updating the individual using the crossover operation based on BWO, and the expression is as follows: ; Among them, represents individual i after the (t + 1)-th update, and are two randomly selected individuals respectively, and val is a random value of 0 or 1; When the random number 0.2 < r ≤ 0.8, updating the individual using the global search based on SO, and the expression is as follows: ; Among them, is a random function, and upper and lower are the upper and lower boundaries of the decision variable, respectively; When the random number r > 0.8, updating the individual using the levy flight, and the expression is as follows: ; Among them, is the best individual of the current population, and LevyFlight is a random step size used to enhance the global search ability.

8. The vehicle scheduling and collaborative control method for open-pit mines according to claim 6, characterized in that In each iteration process, selecting the optimal solution from the Pareto front solutions through the entropy weight TOPSIS method, including: Performing normalization processing according to the objective types of each objective function in the vehicle scheduling model to obtain the decision matrix; Calculating the objective weights of each objective through the entropy weight TOPSIS method; Obtaining the weighted decision matrix based on the objective weights of each objective and the decision matrix; In each iteration process, calculating the Euclidean distances from each individual in the population to the positive ideal solution and the negative ideal solution to obtain the ideal solution closeness of each individual; Selecting the individual with the largest closeness as the local optimal solution; Updating the global optimal solution according to the local optimal solution.

9. The vehicle scheduling and collaborative control method for open-pit mines according to claim 5, characterized in that, Regarding the vehicles in the open-pit mine as agents, defining the state information of the vehicle itself and the state information of all observable vehicles within the vehicle's sensing range as the state space of the agent corresponding to the vehicle, defining the vehicle steering angle and acceleration as the action space of the agent corresponding to the vehicle, and constructing a vehicle trajectory tracking control model based on multi-agent reinforcement learning by designing a reward function, building an action network and a critic network.

10. The vehicle scheduling and collaborative control method for open-pit mines according to claim 8, characterized in that, Agent Reward function at time step t is defined as: ; Among them, represents the collision reward, represents the arrival reward, represents the tracking accuracy reward, represents the efficiency reward, represents the comfort reward, are respectively the weight coefficients of the tracking accuracy reward , the efficiency reward and the comfort reward ; Tracking accuracy reward The calculation formula is as follows: ; Among them, is the distance from the vehicle center point to the planned path; Efficiency Reward The calculation formula is as follows: ; Among them, is the speed of the host vehicle, , are respectively the maximum speed limit and the minimum speed limit in the construction scenario; Comfort Reward The calculation formula is as follows: ; wherein, is the acceleration at the current moment, is the acceleration at the previous moment, are the maximum and minimum accelerations respectively.

Citation Information

Cited By

  • Multi-machine collaborative scheduling mine intelligent transportation system and method

    CN121146458A

  • A mine intelligent transportation system and method based on multi-machine cooperative scheduling

    CN121146458B

  • Intelligent dispatching method and system for underground mine ore removal shoveling and transporting equipment

    CN121526257A

  • An intelligent scheduling method and system for underground mine ore extraction shovel-truck equipment

    CN121526257B

  • Multi-agent model parameter initialization method and device based on imitation learning and storage medium

    CN121683850A