Coal mine underground vehicle dispatching method, device and electronic equipment

CN115860599BActive Publication Date: 2026-09-11YULIN SHENHUA ENERGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211446653.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-18
Publication Date
2026-09-11
Estimated Expiration
2042-11-18

AI Technical Summary

Technical Problem

[0003]本申请的主要目的在于提供一种煤矿井下车辆的派单方法、装置、计算机可读存储介质和电子设备,以解决现有技术中由于物料设备与车型的匹配难度较高以及多种类型车辆分布的不合理性性,会导致井下煤矿运输的效率较低的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115860599B_ABST
    Figure CN115860599B_ABST
Patent Text Reader

Abstract

The application provides a coal mine underground vehicle dispatching method and device and electronic equipment. In the scheme, the optimal vehicle, that is, the target vehicle, can be selected for the target object to be transported according to the related information of the target object to be transported in the coal mine and the state information of the plurality of vehicles. The matching degree of the selected target vehicle and the target object to be transported is the highest. The target vehicle is dispatched, the target object is transported, the matching efficiency of the target object and the vehicle can be improved, the distribution of the plurality of vehicles is updated, the optimal vehicle can be determined more accurately when dispatching in the subsequent process, the distribution of the plurality of vehicles is more reasonable, and the problem of high empty rate of the vehicle generally does not occur, thereby improving the efficiency of the coal mine transportation in the coal mine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a method, apparatus, computer-readable storage medium, and electronic device for dispatching vehicles in underground coal mines. Background Technology

[0002] The transportation scenario in underground coal mines is extremely complex. At fixed shift change times, the dispatch platform (server) receives a large number of travel requests every minute. Vehicles of various types are constantly in motion, and the real-time transportation status of vehicles in underground coal mines changes rapidly over time. Due to the unique nature of underground coal mine transportation, different equipment needs to be transported to designated locations at specified times each day to meet production demands. Furthermore, different materials and equipment are compatible with different vehicle types, making equipment and vehicle matching highly challenging. Each intelligent decision made by the dispatch server affects future vehicle and order distribution. Therefore, the current solution suffers from low efficiency in underground coal mine transportation due to the high difficulty in matching materials and equipment with vehicle types and the irrational distribution of various vehicle types. Summary of the Invention

[0003] The main objective of this application is to provide a method, apparatus, computer-readable storage medium, and electronic device for dispatching vehicles in underground coal mines, in order to solve the problem that the efficiency of underground coal mine transportation is low due to the difficulty in matching materials and equipment with vehicle models and the unreasonable distribution of various types of vehicles in the prior art.

[0004] According to one aspect of the present invention, a method for dispatching vehicles in a coal mine is provided, comprising: acquiring relevant information about a target object to be transported in the coal mine, the relevant information including at least one of the following: the location of the target object, the type of the target object, and the weight of the target object, wherein the target object includes goods or people; acquiring status information of multiple vehicles in the coal mine, the status information including at least one of the following: the type of the vehicle, the load capacity of the vehicle, the location of the vehicle, and whether the vehicle is currently transporting goods; selecting one vehicle from the multiple vehicles as a target vehicle based on the relevant information of the target object and the status information of the multiple vehicles, and controlling the target vehicle to transport the target object; updating the distribution of the multiple vehicles in the coal mine based on the status information of the target vehicle and the status information of non-target vehicles, and controlling the movement of the target vehicle and the non-target vehicles according to the distribution, wherein the distribution refers to the location distribution of the vehicles in the coal mine.

[0005] Optionally, selecting one vehicle as the target vehicle from the plurality of vehicles based on the relevant information of the target object and the status information of the plurality of vehicles includes: selecting a plurality of initial target vehicles that meet a first preset condition from the plurality of vehicles based on the relevant information of the target object and the status information of the plurality of vehicles, wherein the first preset condition includes at least one of the following: the vehicle is not transporting goods, the type of the vehicle is related to the type of the target object, the load capacity of the vehicle is greater than or equal to the weight of the target object, and the distance between the location of the vehicle and the location of the target object is less than a first distance; selecting a target vehicle that meets a second preset condition from the plurality of initial target vehicles based on the relevant information of the target object and the status information of the plurality of initial target vehicles, wherein the second preset condition includes at least one of the following: the initial target vehicle is not transporting goods, the type of the initial target vehicle is related to the type of the target object, the load capacity of the initial target vehicle is greater than or equal to the weight of the target object, and the distance between the location of the initial target vehicle and the location of the target object is less than a second distance, wherein the target vehicle is the optimal vehicle for transporting the target object among the plurality of initial target vehicles, and the first distance is greater than the second distance.

[0006] Optionally, updating the distribution of multiple vehicles underground in the coal mine based on the status information of the target vehicle and the status information of non-target vehicles includes: acquiring environmental information underground in the coal mine, the environmental information including at least one of the following: relevant information of multiple target objects, road condition information underground in the coal mine, and location information of the working area underground in the coal mine; when the target vehicle is in one of the first, second, third, or fourth states, updating the distribution of multiple vehicles underground in the coal mine based on the environmental information, the status information of the target vehicle, and the status information of the non-target vehicles, wherein the first state refers to the state where the target vehicle is heading to the location of the target object, the second state refers to the state where the target vehicle is loading the target object, the third state refers to the state where the target vehicle is transporting the target object to the target location, and the fourth state refers to the state where the target vehicle is idle.

[0007] Optionally, after controlling the target vehicle to transport the target object, the method further includes: constructing a first model, the output of which is the reward result of the actual target vehicle selected based on the relevant information of the target object and the status information of multiple vehicles; constructing a second model, the output of which is the reward result of the predicted target vehicle selected based on historical data, wherein the historical data refers to the historical relevant information of the historical target object and the historical status information of multiple vehicles obtained within a historical time period; and determining the dispatch accuracy based on the output of the first model and the output of the second model.

[0008] Optionally, constructing a first model includes: taking the timely completion of orders by the actual target vehicle as a positive reward value, taking the relationship between the transportation cost and transportation mileage of the actual target vehicle completing the order as a first negative reward value, taking the downtime of the actual target vehicle completing the order as a second negative reward value, and obtaining the sum of the positive reward value, the first negative reward value, and the second negative reward value to obtain an initial first model; summing the output results of multiple initial first models of the actual target vehicle within the target time period to obtain the first model.

[0009] Optionally, constructing a second model includes: constructing the second model, wherein the second model is trained using multiple sets of training data, each set of training data including historical information related to the historical target object acquired within a historical time period, historical state information of multiple vehicles, and historical reward results.

[0010] Optionally, determining the order dispatch accuracy based on the output results of the first model and the second model includes: determining the order dispatch accuracy as a first accuracy when the similarity between the output results of the first model and the output results of the second model is greater than or equal to a similarity threshold; and determining the order dispatch accuracy as a second accuracy when the similarity between the output results of the first model and the output results of the second model is less than the similarity threshold, wherein the first accuracy is higher than the second accuracy.

[0011] According to another aspect of the present invention, a dispatching device for vehicles in a coal mine is also provided, comprising: a first acquisition unit, configured to acquire relevant information of a target object to be transported in the coal mine, the relevant information including at least one of the following: the location of the target object, the type of the target object, and the weight of the target object, wherein the target object includes goods or people; a second acquisition unit, configured to acquire status information of multiple vehicles in the coal mine, the status information including at least one of the following: the type of the vehicle, the load capacity of the vehicle, the location of the vehicle, and whether the vehicle is currently transporting goods; a dispatching unit, configured to select one vehicle from the multiple vehicles as a target vehicle based on the relevant information of the target object and the status information of the multiple vehicles, and control the target vehicle to transport the target object; and a processing unit, configured to update the distribution of the multiple vehicles in the coal mine based on the status information of the target vehicle and the status information of non-target vehicles, and control the movement of the target vehicle and the non-target vehicles according to the distribution, wherein the distribution refers to the location distribution of the vehicles in the coal mine.

[0012] According to another aspect of the present invention, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored program, wherein the program executes any one of the methods described.

[0013] According to another aspect of the present invention, an electronic device is also provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include methods for performing any one of the methods described.

[0014] In this embodiment of the invention, firstly, relevant information about the target object to be transported underground in the coal mine is obtained. Then, the status information of multiple vehicles underground in the coal mine is obtained. Next, based on the relevant information of the target object and the status information of the multiple vehicles, one vehicle is selected as the target vehicle, and the target vehicle is controlled to transport the target object. Finally, based on the status information of the target vehicle and the status information of non-target vehicles, the distribution of the multiple vehicles underground in the coal mine is updated, and the target vehicle and non-target vehicles are controlled to move according to the distribution. In this scheme, the optimal vehicle, i.e., the target vehicle, is selected based on the relevant information of the target object to be transported underground in the coal mine and the status information of multiple vehicles. The selected target vehicle has the highest matching degree with the target object to be transported. Dispatching the target vehicle to transport the target object improves the matching efficiency between the target object and the vehicle. Furthermore, this scheme updates the distribution of multiple vehicles, so that when dispatching orders later, the optimal vehicle can be determined more efficiently and accurately because the distribution of multiple vehicles has been updated. This results in a more reasonable distribution of multiple vehicles and generally avoids the problem of high empty load rates, thereby improving the efficiency of coal transportation underground in the coal mine. Attached Figure Description

[0015] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0016] Figure 1 A flowchart illustrating a method for dispatching underground vehicles in a coal mine according to an embodiment of this application is shown.

[0017] Figure 2 A flowchart illustrating another method for dispatching vehicles underground in coal mines is shown.

[0018] Figure 3 A schematic diagram showing the selection of the target vehicle is shown;

[0019] Figure 4 A schematic diagram of the model construction is shown;

[0020] Figure 5 A schematic diagram of the structure of a dispatching device for underground coal mine vehicles according to an embodiment of this application is shown. Detailed Implementation

[0021] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0022] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0024] It should be understood that when an element (such as a layer, film, region, or substrate) is described as being "on" another element, the element may be directly on the other element, or there may be an intermediate element present. Furthermore, in the specification and claims, when an element is described as being "connected" to another element, the element may be "directly connected" to the other element, or "connected" to the other element via a third element.

[0025] This solution falls at the intersection of the coal mining industry and artificial intelligence technology, primarily focusing on intelligent modeling of underground auxiliary transportation systems in coal mines. These systems are crucial for mine production and construction, responsible for transporting personnel, equipment, and materials. Mine safety is heavily reliant on underground auxiliary transportation systems and various auxiliary transportation equipment. For many coal mining enterprises, continuously improving the efficiency, safety, and convenience of underground traffic management, as well as reducing overall costs, is of paramount importance.

[0026] The transportation scenario in underground coal mines is extremely complex. Firstly, at fixed shift change times, the dispatch platform receives a large number of travel requests every minute. Secondly, various vehicle types are constantly in motion, and the real-time transportation status in underground coal mines changes rapidly over time. Furthermore, due to the unique nature of underground coal mine transportation, different equipment needs to transport goods to specific locations at specific times each day to meet production demands, and different materials and equipment have different compatible vehicle types, further increasing the difficulty of matching equipment with vehicle models. Finally, each dispatch affects future vehicle and order distribution. These industry challenges place higher demands on the algorithm: it not only needs to be highly efficient, capable of quickly and dynamically matching drivers and equipment in real time, making decisions within seconds, but also needs to predict future situations and consider the long-term benefits of the matching algorithm. In addition, it must ensure user experience while considering driver income, and globally optimize overall transportation efficiency.

[0027] A crucial aspect of underground auxiliary transportation systems in coal mines is matching specific vehicle types with the required personnel, equipment, and goods. The core of this is the intelligent agent dispatching model (where the intelligent agent is the vehicle). Furthermore, the model needs to meet constraints such as time period constraints, vehicle mileage constraints, and limited number of personnel.

[0028] Many domestic experts and scholars have conducted research on the mechanisms and algorithms of intelligent order dispatch models. They have delved into the implementation mechanisms of intelligent order dispatch on food delivery platforms, using wave frequency rules, execution rules, calculation rules, and planning period rules to classify and optimize order dispatch scenarios, and employing machine learning, operations research, and simulation analysis for optimization decisions. Experts have proposed a gridded Manhattan path algorithm for finding empty taxis and passengers, which grids the area between taxis and passenger-carrying core points to find the Manhattan path with the highest probability of picking up passengers and recommend it to empty taxi drivers. Experts have also proposed an intelligent taxi dispatch method based on the artificial fish swarm algorithm to achieve global dispatch and rational allocation of taxi resources. The paper proposes improved schemes for the foraging function, clustering function, and tail-chasing function in the standard artificial fish swarm algorithm, as well as an optimization strategy that limits the current optimal state threshold, giving the improved algorithm a good search capability for finding the global optimal solution. Furthermore, experts have proposed a gridded real-time dynamic dispatch reinforcement learning control method for taxis, using dynamic adjustment of empty routes as a control means to establish a dispatch reinforcement learning model to further improve taxi service efficiency and increase taxi revenue.

[0029] Current research on intelligent dispatching algorithms largely focuses on intelligent scheduling for taxis, with fewer algorithms specifically designed for underground coal mine scenarios, and even fewer research on traffic scheduling algorithms within specific topological relationships. Research on intelligent dispatching can be broadly categorized as follows: First, there's the optimization of traditional path planning algorithms for specific dispatching scenarios and road conditions. This method heavily relies on the fundamental performance of traditional path planning algorithms, and gridded or cellular map information requires high connectivity of topological information, making it unsuitable for direct application in underground coal mine dispatching tasks. Second, there's the study of the internal mechanisms of intelligent dispatching models and the use of operations research to optimize decision-making. This type of method requires experts to manually design various constraints and objective functions. For dispatching tasks in complex scenarios, the solver may not be able to provide a satisfactory solution quickly. Underground coal mine scenarios are highly sensitive to time consumption, especially during peak dispatching periods. Underground transportation systems cannot tolerate long waiting times, and in scenarios with frequently changing traffic conditions, the optimal solution returned by the solver may lose real-time reliability, further increasing the risk of transportation plan failure. Then, heuristic algorithms, such as genetic algorithms and ant colony algorithms, are used to model and optimize the dispatch scenario. Taking genetic algorithms as an example, the local search capability of this method is poor, which makes the simple genetic algorithm relatively time-consuming and prone to low search efficiency in the later stage of algorithm training.

[0030] Therefore, some solutions suffer from problems such as difficulty in matching materials and equipment with vehicles, unreasonable distribution of various types of vehicles, high vehicle vacancy rates, and low efficiency in underground coal mine transportation.

[0031] As mentioned in the background section, the existing technology suffers from low efficiency in underground coal mine transportation due to the difficulty in matching materials and equipment with vehicle models and the unreasonable distribution of various types of vehicles. In order to solve the above problems, a typical embodiment of this application provides a method, device, computer-readable storage medium, and electronic device for dispatching underground coal mine vehicles.

[0032] According to an embodiment of this application, a method for dispatching vehicles in underground coal mines is provided.

[0033] Figure 1 This is a flowchart of a method for dispatching orders for underground vehicles in coal mines according to an embodiment of this application. Figure 1 As shown, the method includes the following steps:

[0034] Step S101: Obtain relevant information about the target object to be transported in the coal mine. The relevant information includes at least one of the following: the location of the target object, the type of the target object, and the weight of the target object. The target object may include goods or people.

[0035] In step S101 above, relevant information about the target object to be transported in the coal mine can be obtained. For example, when the target object is cargo, the location of the target object is a specific work area, as well as the type of the target object, such as steel, coal, or wood, and the weight of the target object. Only in this way can the optimal vehicle (target vehicle) be determined and dispatched based on the relevant information about the target object to be transported in the coal mine.

[0036] Step S102: Obtain the status information of multiple vehicles in the coal mine. The status information includes at least one of the following: the type of the vehicle, the load capacity of the vehicle, the location of the vehicle, and whether the vehicle is transporting goods.

[0037] In step S102 above, the status information of multiple vehicles in the coal mine can be obtained, such as whether the vehicle type can transport goods or people, whether the vehicle type is for transporting goods (coal or timber), the load capacity of the vehicle, etc. In this way, the optimal vehicle (target vehicle) can be determined and dispatched based on the status information of multiple vehicles in the coal mine.

[0038] Step S103: Based on the above-mentioned relevant information of the target object and the above-mentioned status information of the multiple vehicles, select one of the multiple vehicles as the target vehicle, and control the target vehicle to transport the target object.

[0039] In step S103 above, after obtaining relevant information about the target object to be transported in the coal mine and the status information of multiple vehicles, the optimal vehicle, i.e. the target vehicle, can be determined. The selected target vehicle has the highest matching degree with the target object to be transported. In this way, dispatching orders to the target vehicle to transport the target object can improve the matching efficiency between the target object and the vehicle.

[0040] For example, there are three vehicles in a coal mine: vehicle A, vehicle B, and vehicle C. Vehicle A is located at position A and is used to transport steel with a load capacity of 100 kg. Vehicle B is located at position B and is used to transport timber with a load capacity of 50 kg. Vehicle C is located at position C and is used to transport coal with a load capacity of 200 kg. If an order is received that 150 kg of coal needs to be transported from position D, then according to the selection algorithm in this scheme, vehicle C at position C will be selected as the target vehicle.

[0041] Step S104: Based on the status information of the target vehicle and the status information of the non-target vehicle, update the distribution of the multiple vehicles in the coal mine, and control the movement of the target vehicle and the non-target vehicle according to the distribution, wherein the distribution refers to the location distribution of the vehicles in the coal mine.

[0042] In step S104 above, the distribution of multiple vehicles in the coal mine can be updated when the target vehicle is performing a task or is being dispatched. This is beneficial for subsequent dispatching because if multiple vehicles in a region are accepting orders or transporting goods, there are no vehicles in that region that can accept subsequent orders. So, by directly updating the distribution of multiple vehicles, the nearest or most suitable target vehicle can still be found when dispatching orders. At the same time, because the distribution of multiple vehicles has been updated, the target vehicle can also reach the target location more quickly to transport goods.

[0043] The above method first obtains relevant information about the target object to be transported underground in the coal mine. Then, it obtains the status information of multiple vehicles underground. Next, based on the target object's information and the vehicle status information, it selects one vehicle as the target vehicle and controls it to transport the target object. Finally, based on the target vehicle's status information and the status information of non-target vehicles, it updates the distribution of multiple vehicles underground and controls the movement of both the target and non-target vehicles according to this distribution. This scheme selects the optimal vehicle (target vehicle) based on the target object's information and the vehicle status information. The selected target vehicle has the highest matching degree with the target object, thus improving the matching efficiency between the target object and the vehicle. Furthermore, this scheme updates the distribution of multiple vehicles, allowing for more efficient and accurate allocation of the optimal vehicle during subsequent dispatching. This results in a more reasonable vehicle distribution and generally avoids high empty-load rates, thereby improving the efficiency of coal transportation underground.

[0044] Specifically, in this solution, a Markov decision model can be used to model the dispatching task to solve the sequential decision problem. A deep neural network model can be used to extract features from the vehicle's state information. An intelligent dispatching algorithm based on deep reinforcement learning can be used to comprehensively evaluate the traffic congestion at intersections and take into account the state information of multiple vehicles, so as to match vehicle orders to the most suitable vehicle and driver.

[0045] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0046] To select the target vehicle more efficiently and accurately, a two-stage selection process can be employed. In one embodiment of this application, based on the aforementioned relevant information of the target object and the aforementioned status information of multiple vehicles, one vehicle is selected as the target vehicle from among the multiple vehicles. This specifically includes the following steps:

[0047] Step S301: Based on the above-mentioned relevant information of the target object and the above-mentioned status information of the multiple vehicles, select multiple initial target vehicles that meet the first preset conditions from the multiple vehicles. The first preset conditions include at least one of the following: the vehicle is not transporting goods, the type of the vehicle is related to the type of the target object, the load capacity of the vehicle is greater than or equal to the weight of the target object, and the distance between the position of the vehicle and the position of the target object is less than a first distance.

[0048] Step S302: Based on the aforementioned relevant information of the target object and the aforementioned status information of the multiple initial target vehicles, select the target vehicle that meets the second preset condition from the multiple initial target vehicles. The second preset condition includes at least one of the following: the initial target vehicle is not transporting goods; the type of the initial target vehicle is related to the type of the target object; the load capacity of the initial target vehicle is greater than or equal to the weight of the target object; the distance between the location of the initial target vehicle and the location of the target object is less than a second distance. The target vehicle is the optimal vehicle for transporting the target object among the multiple initial target vehicles, and the first distance is greater than the second distance.

[0049] For example, there are five vehicles in a coal mine: vehicle A, vehicle B, vehicle C, vehicle D, and vehicle E. After the first screening based on the first preset condition, vehicles A, B, and C all meet the first preset condition. However, to avoid all three vehicles from loading and unloading goods at the same time, a second screening is performed based on the second preset condition, and vehicle C is selected as the target vehicle. The specific selection algorithm is not limited; it can be selected based on status information or based on a weighted algorithm.

[0050] In the above steps S301 to S302, a first screening is performed based on the first preset condition to select multiple initial target vehicles. Then, a second screening is performed based on the second preset condition to select the final target vehicle from the multiple initial target vehicles. This is to avoid multiple vehicles meeting the first preset condition at the same time, which would result in multiple vehicles receiving orders simultaneously. This further ensures that the selected target vehicle is the best vehicle.

[0051] The training of deep reinforcement learning agents relies heavily on interaction data between the agent and its environment. A broader range of data from these interactions generally leads to better handling of state information in the training results. A key challenge in the training process is balancing exploration and utilization, while ensuring safe interaction between the agent and its environment as the breadth of exploration increases. Therefore, introducing safety constraints specific to underground mining scenarios becomes a significant challenge in the modeling process of reinforcement learning agents. Under the definition of a Markov decision process, the order dispatching process involves the platform obtaining the state of each vehicle to be assigned in each round of order dispatch and setting all pending orders as one of the actions the driver can perform. The optimization goal is to reduce downtime and transportation costs while ensuring a good user experience, and further optimize actions using action masks. After the agent outputs vehicle-related actions, constraints are applied through vehicle-task matching to remove vehicles that do not meet the requirements (this removal is not an actual deletion from the mine, but rather a cessation of further computation within the algorithm), and then further selection is performed from the vehicles that meet the basic requirements.

[0052] In reality, the distribution of multiple vehicles underground in the coal mine has already changed while the target vehicle is moving. Therefore, the distribution of vehicles underground in the coal mine can be updated based on the status of the target vehicle. In another embodiment of this application, the distribution of multiple vehicles underground in the coal mine is updated based on the status information of the target vehicle and the status information of the non-target vehicles, specifically including the following steps:

[0053] Step S401: Obtain environmental information from underground coal mines. The environmental information includes at least one of the following: relevant information of multiple target objects, road condition information from underground coal mines, and location information of the working area underground coal mines.

[0054] Step S402: When the target vehicle is in one of the first, second, third, or fourth states, the distribution of the multiple vehicles in the coal mine is updated based on the environmental information, the status information of the target vehicle, and the status information of the non-target vehicles. The first state refers to the state where the target vehicle is heading to the location of the target object; the second state refers to the state where the target vehicle is loading the target object; the third state refers to the state where the target vehicle is transporting the target object to the target location; and the fourth state refers to the state where the target vehicle is idle.

[0055] In steps S401 to S402 above, the distribution of multiple vehicles underground in the coal mine can be updated, which can ensure that the distribution of multiple vehicles underground in the coal mine is relatively reasonable.

[0056] After the target vehicle has transported the target object, the accuracy rate of the dispatch can be determined. This allows us to assess the effectiveness of the dispatch and optimize the algorithm based on the dispatch accuracy rate. In another embodiment of this application, after controlling the target vehicle to transport the target object, the method further includes the following steps:

[0057] Step S105: Construct a first model. The output of the first model is the reward result of the actual target vehicle selected based on the relevant information of the target object and the state information of the multiple vehicles.

[0058] In underground transportation scenarios, different equipment needs to be transported to designated locations at different times. Therefore, the response time for issuing dispatch instructions can be within seconds. Furthermore, each decision can focus on improving long-term benefits. In one specific embodiment of this application, a first model is constructed, including: using the timely completion of orders by the actual target vehicle as a positive reward value; using the relationship between the transportation cost and transportation mileage of the actual target vehicle completing the order as a first negative reward value; and using the downtime of the actual target vehicle completing the order as a second negative reward value. The sum of the positive reward value, the first negative reward value, and the second negative reward value is obtained to form an initial first model. The output results of multiple initial first models for the actual target vehicles within a target time period are summed to obtain the first model. In this embodiment, reinforcement learning and combinatorial optimization fusion can be performed using the constructed first model to solve for the globally optimal solution.

[0059] Specifically, the first model can be a Markov decision model. The main idea of ​​the algorithm is that the problem first conforms to the definition of sequential decision making, and the algorithm uses a Markov Decision Process (MDP) for modeling and solves it using a reinforcement learning model. Secondly, for the many-to-many matching between vehicles and equipment, the problem is modeled as a combinatorial optimization problem to obtain the global optimum. After reinforcement learning obtains the available vehicles, it performs type matching between equipment and vehicles, and then performs mask optimization on the output actions of reinforcement learning.

[0060] The Markov decision process consists of the following modules:

[0061] Agent: Each vehicle is defined as an agent. Using a single vehicle as an agent significantly reduces the state and action space, resulting in higher accuracy and effectiveness of the solution.

[0062] State information: State 's' defines vehicle type, load capacity, location, and whether it is currently hauling goods. For simplicity, time is quantified into 15-minute intervals, but any other feasible time intervals are also possible. Therefore, a complete episode consists of 96 time slices.

[0063] Action: This section designs four actions for the intelligent agent: vehicle heading to load equipment, waiting for equipment to be loaded onto the vehicle, delivering equipment to the destination, and idle operation.

[0064] State transition and reward function: Completing an order automatically causes a transition in the vehicle's spatiotemporal state, which simultaneously generates a reward. We define the reward function... For reasons of downtime and transportation costs.

[0065] The specific formula for positive reward value is as follows: ,in, Indicates a positive reward value. It is the positive reward value at each moment, if the order can be completed on time. =1, otherwise =0, This indicates the time period for completing the order task, while the positive reward value is a statistical value indicating whether the order was completed on time.

[0066] The first negative reward value is a statistical value of transportation costs, calculated using the following formula: ,in This represents the first negative reward value. Indicates the first The total mileage of a single order, and the first negative reward value for a task is the cost of completing that task across all orders. This indicates the relationship between transportation costs and transportation distance. This indicates the number of orders.

[0067] The formula relating transportation costs to transportation mileage is: ,in, Indicates the distance traveled.

[0068] The specific formula for the second negative reward value is as follows: ,in, This indicates the second negative reward value. This indicates the number of people the order carries. Indicates the first The first order The time lost by each miner.

[0069] The total reward is then calculated using the following formula: The calculation of the total reward value is the process of constructing the initial first model. After defining the basic elements of the MDP, the next step is to train an optimal policy based on the deep reinforcement learning framework to maximize the cumulative expected reward.

[0070] Based on the above scenario abstraction, we can design a set of discrete action reinforcement learning algorithms. The algorithm needs to provide the optimal action that maximizes the reward under the given state.

[0071] First, define the reward function (i.e., the first model) as a Markov reward chain from... The sum of all decaying reward benefits from the start of the timeline: ,in, This represents the first model, where each term in the function is the output of an initial first model.

[0072] In practical algorithm design, two neural networks can be designed: a critic network, primarily used to fit the second model (input is the state, output is the value of the current state); and an actor network, used to fit the policy function and output the action (input is the state, output is the probability of the action to be executed in the current state). During training, the agent continuously interacts with the environment through the output of the neural network actor, obtaining training samples. When enough training samples are collected, the parameters of both neural networks are updated.

[0073] The actor network contains four outputs, each corresponding to a probability value of a different output action (the first state, the second state, the third state, and the fourth state). In the last layer of the network, actions need to be sampled in the discrete action space according to probability.

[0074] During training, an experience replay mechanism (experience pool) is used to store the sampled data. This approach enriches the distribution of data samples and reduces training failures caused by uneven data distribution. However, because the parameters of the actor network are constantly updated, the data in the experience pool may come from different actor network parameters, leading to training instability. To address this, KL divergence is used to measure the difference between old and new parameters. A threshold is set to truncate the difference, ensuring that the updates of old and new parameters remain within a certain range and guaranteeing training stability.

[0075] Step S106: Construct a second model. The output of the second model is the reward result of the target vehicle selected based on historical data. The historical data refers to the historical relevant information of the target object and the historical status information of multiple vehicles obtained within the historical time period.

[0076] In another specific embodiment of this application, constructing a second model includes: constructing the aforementioned second model, wherein the second model is trained using multiple sets of training data, each set of training data including historical information related to the target object acquired within a historical time period, historical state information of multiple vehicles, and historical reward results. In this embodiment, through the constructed second model, a predicted value for the reward result of the target vehicle can be obtained, thus obtaining a prediction result, which can subsequently be used to determine the accuracy of the result calculated by the first model.

[0077] Specifically, how to assess the value of a vehicle appearing in a specific time / space is also a core issue in intelligent dispatching systems. The second model can quantify this in both time and space dimensions. Based on the platform's massive historical data resources, a deep neural network is used to fit the state value function (second model). This method solves for the expected revenue of the vehicle in each spatiotemporal state, i.e., the value function.

[0078] In an alternative embodiment, the second model can also be represented by the Bellman expectation equation:

[0079] In this formula, This represents the output of the second model. Indicates the expected value. Indicates the first The reward value at any moment, As a discount factor, Indicates the first The value function at time t, This indicates the vehicle's status information. Indicates the first The vehicle's state information at any given time is expressed by the formula as follows: the current state value function equals the current instantaneous reward plus the discounted expected value of the future value function. This scheme uses deep reinforcement learning technology for modeling, specifically the PPO algorithm. A deep neural network is used to fit the first and second models. The neural network result is a multi-layer fully connected network with the number of input layer neurons equal to the state space dimension and the number of output layer neurons equal to 1. This output value represents the accuracy.

[0080] Specifically, the action strategy optimization formula in this scheme is as follows:

[0081] In this formula, This indicates the result of the output action. This represents the total length with time as the step size. This represents the new policy function. This represents the old policy function. Indicates the first The vehicle's actions at each time point are defined as the policy function. Indicates the first Vehicle status information at any given time. Represents the dominance function. These are hyperparameters associated with the dominance function, used to determine the minimum and maximum values. Since is a constant, the overall formula expresses the following meaning: Based on the ratio of the action probabilities of the old and new policy functions, determine the specific value of the advantage function. First, find the minimum value between the policy ratio and the fixed value. Then, find the overall formula that maximizes this minimum value. parameter, The calculation formula is:

[0082] ,

[0083] The optimization formula for the second model is: In this formula, Indicates the first The value function parameters are updated next time. This represents the total length with time as the step size. For value function, The actual reward value function, Indicates the first The vehicle's status information at time t, the overall formula means the t... The network parameters of the secondary value function are determined by the minimum mean square error between the value function and the actual reward value. Among them, For the above text .

[0084] Step S107: Determine the order dispatch accuracy based on the output results of the first model and the second model.

[0085] Combining the results calculated by the first model and the second model can form a complete evaluation method. In order to further determine the accuracy of the matching algorithm of this scheme, in another specific embodiment of this application, the accuracy of dispatching is determined based on the output results of the first model and the second model, including: when the similarity between the output results of the first model and the output results of the second model is greater than or equal to a similarity threshold, the accuracy of dispatching is determined to be a first accuracy; when the similarity between the output results of the first model and the output results of the second model is less than the similarity threshold, the accuracy of dispatching is determined to be a second accuracy, wherein the first accuracy is higher than the second accuracy.

[0086] Specifically, combining the outputs of the first and second models forms a complete dispatch evaluation method. The algorithm includes offline and online components. The offline component essentially discretizes any spatial location at any given time into a spatiotemporal grid, calculating the expected revenue for that grid up to the end of the day based on dispatch records (including drivers participating in dispatch but without orders). The online component evaluates the matching degree based on four criteria: First, if the dispatched vehicle can arrive on time, the action safety constraint needs to comprehensively consider both current immediate rewards and future value. Within the action safety constraint, state data is extrapolated based on the actions performed by each agent, and the new state is used as input to a deep neural network. Then, each agent compares the value function of the new state and selects the action with the higher value function. Finally, the higher the transportation cost, the lower the matching degree.

[0087] For example, suppose there are currently 5 vehicles. After the first model, the state values ​​of 3 vehicles all choose the action of hauling goods, i.e., π_1 = hauling goods, π_2 = hauling goods, and π_3 = hauling goods. The value functions of the three agents are v_1(st), v_2(st), and v_3(st), respectively. At this point, a safety constraint layer is needed for further selection. The specific selection process is as follows: First, assuming the action of hauling goods is chosen, we will get a reward value r_1, and the state is changed to s_t+1'. s_t+1' is input into the value function neural network to get v_1(s_t+1'). Similarly, we get v_2(s_t+1') and v_3(s_t+1'). The safety constraint layer needs to compare r_1, r_2, r_3, and the three value functions mentioned above. It needs to compare not only the size of r, but also the size of v(s_t+1'). After comprehensive comparison, the better specific action is selected.

[0088] In steps S105 to S107 above, the accuracy of the matching algorithm of this scheme can be determined by constructing the first model and the second model. This allows for further optimization of the algorithm to improve the matching accuracy of this scheme.

[0089] This scheme uses deep reinforcement learning algorithms to model the dispatch system. Compared with traditional heuristic algorithms and operations research algorithms, the intelligent agent modeling method studied has a faster solution speed and higher solution accuracy. While ensuring the timeliness of the solution, it can also meet the solution accuracy requirements of the underground mining scenario.

[0090] For the technical aspects used in implementing this solution, heuristic algorithms such as evolutionary algorithms, ant colony algorithms, and fish colony algorithms can also be selected for modeling and solving. The aforementioned modeling algorithms, such as ant colony algorithms and evolutionary algorithms, are all heuristic algorithms. While these algorithms can yield relatively good feasible solutions, they are not necessarily the optimal solutions.

[0091] For the overall solution, evolutionary algorithms or fish swarm algorithms can be used for agent modeling. Compared to deep reinforcement learning algorithms, these methods rely heavily on computational resources during solution development and do not depend on gradient information when updating the model. The advantages of this approach are its simple and intuitive algorithmic approach, ease of modification by users, and the ability to obtain a very good feasible solution when computational resources are sufficient. However, the disadvantage is that it cannot guarantee a globally optimal solution, and the specific performance depends on the problem and the designer's expert experience.

[0092] In another optional embodiment, the dispatching method for underground coal mine vehicles follows the procedure as follows: Figure 2As shown, firstly, a massive amount of historical data is acquired to simulate the environment and obtain the simulation environment. Then, it is determined whether the environmental information and vehicle status information can interact, i.e., whether the environmental information and vehicle status information are acquired. If they are acquired, the agent is initialized and controlled to interact with the environmental information, i.e., the process of selecting the target vehicle. Then, it is determined whether the data is sufficient. If it is insufficient, the agent is re-initialized. If it is sufficient, the agent is selected, trained, and its parameters are updated until a standard agent is obtained.

[0093] The process of selecting the target vehicle is as follows Figure 3 As shown, the vehicle type, vehicle location information, cargo location information, cargo type, etc. are first obtained and input into the model as vector inputs or matrices. Multiple initial target vehicles that meet the first preset conditions are selected. Then, according to the second preset conditions, the target vehicles are selected based on the legality of the time information and input into the simulation environment.

[0094] The process of building a model is as follows Figure 4 As shown, first, the agent's actions are obtained and then input into the actor variable as the env parameter. The agent's state and reward value are also input into the actor variable. Finally, historical data is obtained as a historical sequence. The target agent is obtained based on multiple batches of data, and parameters can be updated during the process.

[0095] This application also provides a dispatching device for underground coal mine vehicles. It should be noted that the dispatching device for underground coal mine vehicles in this application can be used to execute the dispatching method for underground coal mine vehicles provided in this application. The dispatching device for underground coal mine vehicles provided in this application will be described below.

[0096] Figure 5 This is a schematic diagram of the dispatching device for underground vehicles in a coal mine according to an embodiment of this application. Figure 5 As shown, the dispatching device includes:

[0097] The first acquisition unit 10 is used to acquire relevant information about the target object to be transported in the coal mine. The relevant information includes at least one of the following: the location of the target object, the type of the target object, and the weight of the target object. The target object includes goods or people.

[0098] The first acquisition unit mentioned above can acquire relevant information about the target object to be transported in the coal mine. For example, when the target object is cargo, the location of the target object is a specific work area, as well as the type of the target object, such as steel, coal, or wood, and the weight of the target object. Only in this way can the optimal vehicle (target vehicle) be determined and dispatched based on the relevant information about the target object to be transported in the coal mine.

[0099] The second acquisition unit 20 is used to acquire the status information of multiple vehicles in the coal mine. The status information includes at least one of the following: the type of the vehicle, the load capacity of the vehicle, the location of the vehicle, and whether the vehicle is transporting goods.

[0100] The second acquisition unit mentioned above can acquire the status information of multiple vehicles in the coal mine, such as whether the vehicle type can transport goods or people, whether the vehicle type is for transporting goods (coal or timber), the load capacity of the vehicle, etc. In this way, the optimal vehicle (target vehicle) can be determined and dispatched based on the acquired status information of multiple vehicles in the coal mine.

[0101] The dispatching unit 30 is used to select one of the vehicles as the target vehicle from the multiple vehicles based on the relevant information of the target object and the status information of the multiple vehicles, and to control the target vehicle to transport the target object.

[0102] The dispatching unit described above, having already obtained relevant information about the target object to be transported in the coal mine and the status information of multiple vehicles, can determine the optimal vehicle, i.e. the target vehicle. The selected target vehicle has the highest matching degree with the target object to be transported. By dispatching the target vehicle to transport the target object, the matching efficiency between the target object and the vehicle can be improved.

[0103] The processing unit 40 is used to update the distribution of multiple vehicles in the coal mine according to the status information of the target vehicle and the status information of the non-target vehicle, and to control the movement of the target vehicle and the non-target vehicle according to the distribution, wherein the distribution refers to the positional distribution of the vehicles in the coal mine.

[0104] The aforementioned processing unit can update the distribution of multiple vehicles underground in the coal mine while the target vehicle is performing a task or being dispatched. This is beneficial for subsequent order dispatch because if multiple vehicles in a region are all accepting orders or transporting goods, there will be no vehicles available to accept subsequent orders in that region. By directly updating the distribution of multiple vehicles, the nearest or most suitable target vehicle can still be found when dispatching orders. At the same time, because the distribution of multiple vehicles has been updated, the target vehicle can also reach the target location more quickly to transport goods.

[0105] In the aforementioned device, the first acquisition unit acquires relevant information about the target object to be transported underground in the coal mine, and the second acquisition unit acquires the status information of multiple vehicles underground in the coal mine. The dispatching unit selects one vehicle from the multiple vehicles as the target vehicle based on the target object's information and the status information of the multiple vehicles, and controls the target vehicle to transport the target object. The processing unit updates the distribution of the multiple vehicles underground in the coal mine based on the status information of the target vehicle and the non-target vehicles, and controls the movement of the target vehicle and non-target vehicles according to the distribution. This scheme can select the optimal vehicle (target vehicle) for the target object to be transported based on the relevant information of the target object underground in the coal mine and the status information of the multiple vehicles. The selected target vehicle has the highest matching degree with the target object to be transported. Dispatching the target vehicle to transport the target object improves the matching efficiency between the target object and the vehicle. Furthermore, this scheme updates the distribution of multiple vehicles, so that when dispatching orders later, the optimal vehicle can be determined more efficiently and accurately because the distribution of multiple vehicles has been updated. This results in a more reasonable distribution of multiple vehicles and generally avoids the problem of high empty vehicle rates, thereby improving the efficiency of coal transportation underground in the coal mine.

[0106] To select target vehicles more efficiently and accurately, a two-stage selection process can be employed. In one embodiment of this application, the dispatch unit includes a first selection module and a second selection module, with the functions of each module as follows:

[0107] The first selection module is used to select multiple initial target vehicles that meet the first preset conditions from the multiple vehicles based on the above-mentioned relevant information of the target object and the above-mentioned status information of the multiple vehicles. The first preset conditions include at least one of the following: the vehicle is not transporting goods, the type of the vehicle is related to the type of the target object, the load capacity of the vehicle is greater than or equal to the weight of the target object, and the distance between the position of the vehicle and the position of the target object is less than a first distance.

[0108] The second selection module is used to select a target vehicle that meets a second preset condition from the multiple initial target vehicles based on the relevant information of the target object and the status information of the multiple initial target vehicles. The second preset condition includes at least one of the following: the initial target vehicle is not transporting goods; the type of the initial target vehicle is related to the type of the target object; the load capacity of the initial target vehicle is greater than or equal to the weight of the target object; the distance between the position of the initial target vehicle and the position of the target object is less than a second distance. The target vehicle is the optimal vehicle for transporting the target object from the multiple initial target vehicles, and the first distance is greater than the second distance.

[0109] The first selection module and the second selection module described above first perform a first screening based on the first preset condition to select multiple initial target vehicles, and then perform a second screening based on the second preset condition to select the final target vehicle from the multiple initial target vehicles. This is to avoid multiple vehicles meeting the first preset condition at the same time, which would result in multiple vehicles receiving orders at the same time, and thus further ensure that the selected target vehicle is the most suitable vehicle.

[0110] In reality, the distribution of multiple vehicles underground in the coal mine has already changed while the target vehicle is moving. Therefore, the distribution of vehicles underground in the coal mine can be updated based on the status of the target vehicle. In another embodiment of this application, the processing unit includes an acquisition module and an update module, and the functions of each module are as follows:

[0111] The acquisition module is used to acquire environmental information in the coal mine, which includes at least one of the following: relevant information of multiple target objects, road condition information in the coal mine, and location information of the working area in the coal mine.

[0112] The update module is used to update the distribution of multiple vehicles in the coal mine based on the environmental information, the status information of the target vehicle, and the status information of the non-target vehicles when the target vehicle is in one of the first, second, third, or fourth states. The first state refers to the state where the target vehicle is heading to the location of the target object, the second state refers to the state where the target vehicle is loading the target object, the third state refers to the state where the target vehicle is transporting the target object to the target location, and the fourth state refers to the state where the target vehicle is idle.

[0113] The aforementioned acquisition and update modules can update the distribution of multiple vehicles underground in a coal mine, thus ensuring a more reasonable distribution of these vehicles.

[0114] After the target vehicle has transported the target object, the accuracy rate of the dispatch can be determined. This determines the effectiveness of the dispatch and allows for subsequent algorithm optimization based on the dispatch accuracy rate. In another embodiment of this application, the above-mentioned device further includes a first construction unit, a second construction unit, and a determination unit. The functions of each unit are as follows:

[0115] The first construction unit is used to construct a first model after controlling the target vehicle to transport the target object. The output of the first model is the reward result of the actual target vehicle selected based on the relevant information of the target object and the status information of the multiple vehicles.

[0116] In underground transportation scenarios, different equipment needs to be transported to designated locations at different times. Therefore, the dispatch instruction can be responded to within seconds, and each decision can focus on improving long-term benefits. In a specific embodiment of this application, the first construction unit includes a first construction module and a second construction module. The first construction module is used to take the timely completion of the order by the actual target vehicle as a positive reward value, the relationship between the transportation cost and transportation mileage of the actual target vehicle completing the order as a first negative reward value, and the downtime of the actual target vehicle completing the order as a second negative reward value, and obtain the sum of the positive reward value, the first negative reward value, and the second negative reward value to obtain an initial first model. The second construction module is used to sum the output results of multiple initial first models of the actual target vehicles within the target time period to obtain the first model. In this embodiment, reinforcement learning and combinatorial optimization fusion can be performed through the constructed first model, and the global optimal solution can be obtained by solving the problem using the first model.

[0117] The second building unit is used to build the second model. The output of the second model is the reward result of the target vehicle selected based on historical data. The historical data refers to the historical relevant information of the target object and the historical status information of multiple vehicles obtained within the historical time period.

[0118] In another specific embodiment of this application, the second building unit includes a third building module, which is used to build the aforementioned second model. The second model is trained using multiple sets of training data. Each set of training data includes historical information related to the target object acquired within a historical time period, historical state information of multiple vehicles, and historical reward results. In this embodiment, the constructed second model can obtain a predicted value for the reward result of the target vehicle, thus providing a prediction result. This prediction result can then be used to determine the accuracy of the result calculated by the first model.

[0119] The determination unit is used to determine the accuracy of order dispatch based on the output results of the first model and the second model.

[0120] Combining the results calculated by the first model and the second model can form a complete evaluation method. In order to further determine the accuracy of the matching algorithm of this scheme, in another specific embodiment of this application, the determining unit includes a first determining module and a second determining module. The first determining module is used to determine the order dispatch accuracy as a first accuracy when the similarity between the output result of the first model and the output result of the second model is greater than or equal to a similarity threshold. The second determining module is used to determine the order dispatch accuracy as a second accuracy when the similarity between the output result of the first model and the output result of the second model is less than the similarity threshold. The first accuracy is higher than the second accuracy.

[0121] The first construction unit, the second construction unit, and the determining unit described above can determine the accuracy of the matching algorithm of this scheme through the constructed first model and the second model. This allows for further optimization of the algorithm to improve the matching accuracy of this scheme.

[0122] The dispatching device for underground coal mine vehicles includes a processor and a memory. The first acquisition unit, the second acquisition unit, the dispatching unit, and the processing unit are all stored in the memory as program units. The processor executes the program units stored in the memory to achieve the corresponding functions.

[0123] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured, and adjusting kernel parameters can address the problem of low efficiency in underground coal mine transportation caused by the difficulty in matching material handling equipment with vehicle models and the unreasonable distribution of various vehicle types.

[0124] The memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0125] This invention provides a computer-readable storage medium storing a program that, when executed by a processor, implements the above-described method for dispatching vehicles in underground coal mines.

[0126] This invention provides a processor for running a program, wherein the program executes the dispatching method for underground coal mine vehicles.

[0127] This application also provides an electronic device including one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include methods for performing any of the above-described methods.

[0128] This invention provides a device including a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs at least the following steps:

[0129] Step S101: Obtain relevant information about the target object to be transported in the coal mine. The relevant information includes at least one of the following: the location of the target object, the type of the target object, and the weight of the target object. The target object may include goods or people.

[0130] Step S102: Obtain the status information of multiple vehicles in the coal mine. The status information includes at least one of the following: the type of the vehicle, the load capacity of the vehicle, the location of the vehicle, and whether the vehicle is transporting goods.

[0131] Step S103: Based on the above-mentioned relevant information of the target object and the above-mentioned status information of the multiple vehicles, select one of the multiple vehicles as the target vehicle, and control the target vehicle to transport the target object.

[0132] Step S104: Based on the status information of the target vehicle and the status information of the non-target vehicle, update the distribution of the multiple vehicles in the coal mine, and control the movement of the target vehicle and the non-target vehicle according to the distribution, wherein the distribution refers to the location distribution of the vehicles in the coal mine.

[0133] The devices mentioned in this article can be servers, PCs, tablets, mobile phones, etc.

[0134] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialization program having at least the following method steps:

[0135] Step S101: Obtain relevant information about the target object to be transported in the coal mine. The relevant information includes at least one of the following: the location of the target object, the type of the target object, and the weight of the target object. The target object may include goods or people.

[0136] Step S102: Obtain the status information of multiple vehicles in the coal mine. The status information includes at least one of the following: the type of the vehicle, the load capacity of the vehicle, the location of the vehicle, and whether the vehicle is transporting goods.

[0137] Step S103: Based on the above-mentioned relevant information of the target object and the above-mentioned status information of the multiple vehicles, select one of the multiple vehicles as the target vehicle, and control the target vehicle to transport the target object.

[0138] Step S104: Based on the status information of the target vehicle and the status information of the non-target vehicle, update the distribution of the multiple vehicles in the coal mine, and control the movement of the target vehicle and the non-target vehicle according to the distribution, wherein the distribution refers to the location distribution of the vehicles in the coal mine.

[0139] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0140] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units described above can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0141] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0142] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0143] If the aforementioned integrated units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0144] As can be seen from the above description, the embodiments of this application achieve the following technical effects:

[0145] 1) The dispatching method for underground vehicles in this application first obtains relevant information about the target object to be transported in the underground coal mine, then obtains the status information of multiple vehicles in the underground coal mine, then selects one vehicle as the target vehicle based on the relevant information of the target object and the status information of multiple vehicles, and controls the target vehicle to transport the target object, and finally updates the distribution of multiple vehicles in the underground coal mine based on the status information of the target vehicle and the status information of non-target vehicles, and controls the movement of the target vehicle and non-target vehicles according to the distribution. This scheme selects the optimal vehicle (target vehicle) for transporting coal from underground mines based on information about the target objects and the status information of multiple vehicles. The selected target vehicle has the highest matching degree with the target objects. Dispatching the target vehicle to transport the target objects improves the matching efficiency between target objects and vehicles. Furthermore, this scheme updates the distribution of multiple vehicles, allowing for more efficient and accurate assignment of the optimal vehicle during subsequent dispatching. This results in a more reasonable distribution of vehicles and generally avoids high empty-load rates, thereby improving the efficiency of coal transportation in underground mines.

[0146] 2) The dispatching device for underground coal mine vehicles of this application comprises a first acquisition unit acquiring relevant information of the target object to be transported underground, a second acquisition unit acquiring the status information of multiple vehicles underground, a dispatching unit selecting one vehicle as the target vehicle based on the relevant information of the target object and the status information of multiple vehicles, and controlling the target vehicle to transport the target object, and a processing unit updating the distribution of multiple vehicles underground based on the status information of the target vehicle and the status information of non-target vehicles, and controlling the movement of the target vehicle and non-target vehicles according to the distribution. This scheme selects the optimal vehicle (target vehicle) for transporting coal from underground mines based on information about the target objects and the status information of multiple vehicles. The selected target vehicle has the highest matching degree with the target objects. Dispatching the target vehicle to transport the target objects improves the matching efficiency between target objects and vehicles. Furthermore, this scheme updates the distribution of multiple vehicles, allowing for more efficient and accurate assignment of the optimal vehicle during subsequent dispatching. This results in a more reasonable distribution of vehicles and generally avoids high empty-load rates, thereby improving the efficiency of coal transportation in underground mines.

[0147] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method of dispatching a coal mine underground vehicle, characterised by, include: Obtain relevant information about a target object to be transported underground in a coal mine, the relevant information including at least one of the following: the location of the target object, the type of the target object, and the weight of the target object, the target object including goods or people; Obtain status information of multiple vehicles underground in a coal mine, the status information including at least one of the following: vehicle type, vehicle load capacity, vehicle location, and whether the vehicle is transporting goods; Based on the relevant information of the target object and the status information of the multiple vehicles, one vehicle is selected from the multiple vehicles as the target vehicle, and the target vehicle is controlled to transport the target object. Based on the status information of the target vehicle and the status information of the non-target vehicle, the distribution of multiple vehicles in the coal mine is updated, and the movement of the target vehicle and the non-target vehicle is controlled according to the distribution. The distribution refers to the positional distribution of the vehicles in the coal mine. After controlling the target vehicle to transport the target object, the method further includes: constructing a first model, the output of which is the reward result of the actual target vehicle selected based on the relevant information of the target object and the status information of multiple vehicles; constructing a second model, the output of which is the reward result of the predicted target vehicle selected based on historical data, wherein the historical data refers to the historical relevant information of the target object and the historical status information of multiple vehicles obtained within a historical time period; and determining the order dispatch accuracy based on the output results of the first model and the second model. Constructing a first model includes: using the timely completion of orders by the actual target vehicle as a positive reward value, the relationship between the transportation cost and transportation mileage for the actual target vehicle to complete the order as a first negative reward value, and the downtime for the actual target vehicle to complete the order as a second negative reward value; and obtaining the sum of the positive reward value, the first negative reward value, and the second negative reward value to obtain an initial first model; summing the output results of multiple initial first models for the actual target vehicle within the target time period to obtain the first model. Constructing a second model includes: constructing the second model, wherein the second model is trained using multiple sets of training data, each set of training data including historical information related to the target object acquired within a historical time period, historical state information of multiple vehicles, and historical reward results; The second model is embodied by the Bellman expectation equation: , This means that the output of the second model equals the immediate reward at the current moment plus the discounted expected value of the future value function. Indicates the expected value. Let γ represent the reward value at time t+1, and γ be the discount factor. Let S represent the value function at time t+1, and let S represent the vehicle's state information. This represents the vehicle's state information at time t.

2. The method according to claim 1, characterized in that, Based on the relevant information of the target object and the status information of the multiple vehicles, selecting one vehicle from the multiple vehicles as the target vehicle includes: Based on the relevant information of the target object and the status information of the multiple vehicles, a plurality of initial target vehicles that meet the first preset conditions are selected from the multiple vehicles. The first preset conditions include at least one of the following: the vehicle is not transporting goods, the type of the vehicle is related to the type of the target object, the load capacity of the vehicle is greater than or equal to the weight of the target object, and the distance between the location of the vehicle and the location of the target object is less than a first distance. Based on the relevant information of the target object and the status information of the multiple initial target vehicles, a target vehicle that meets the second preset condition is selected from the multiple initial target vehicles. The second preset condition includes at least one of the following: the initial target vehicle is not transporting goods; the type of the initial target vehicle is related to the type of the target object; the load capacity of the initial target vehicle is greater than or equal to the weight of the target object; the distance between the location of the initial target vehicle and the location of the target object is less than a second distance. The target vehicle is the optimal vehicle for transporting the target object among the multiple initial target vehicles, and the first distance is greater than the second distance.

3. The method according to claim 1, characterized in that, Based on the status information of the target vehicle and the status information of the non-target vehicles, the distribution of multiple vehicles underground in the coal mine is updated, including: Obtain environmental information from underground coal mines, wherein the environmental information includes at least one of the following: relevant information of multiple target objects, road condition information underground coal mines, and location information of the working area underground coal mines; When the target vehicle is in one of the first, second, third, or fourth states, the distribution of multiple vehicles in the coal mine is updated based on the environmental information, the state information of the target vehicle, and the state information of the non-target vehicles. The first state refers to the state where the target vehicle is heading to the location of the target object; the second state refers to the state where the target vehicle is loading the target object; the third state refers to the state where the target vehicle is transporting the target object to the target location; and the fourth state refers to the state where the target vehicle is idle.

4. The method according to claim 1, characterized in that, Based on the output results of the first model and the second model, the accuracy of order dispatch is determined, including: If the similarity between the output of the first model and the output of the second model is greater than or equal to the similarity threshold, the accuracy rate of dispatching is determined to be the first accuracy rate. If the similarity between the output of the first model and the output of the second model is less than the similarity threshold, the accuracy rate of dispatching is determined to be the second accuracy rate, wherein the first accuracy rate is higher than the second accuracy rate.

5. A dispatching device for underground vehicles in a coal mine, characterized in that, include: The first acquisition unit is used to acquire relevant information about the target object to be transported in the coal mine. The relevant information includes at least one of the following: the location of the target object, the type of the target object, and the weight of the target object. The target object includes goods or people. The second acquisition unit is used to acquire the status information of multiple vehicles in the coal mine, the status information including at least one of the following: the type of the vehicle, the load capacity of the vehicle, the location of the vehicle, and whether the vehicle is transporting goods. The dispatching unit is used to select one of the vehicles as the target vehicle from the multiple vehicles based on the relevant information of the target object and the status information of the multiple vehicles, and to control the target vehicle to transport the target object. The processing unit is configured to update the distribution of multiple vehicles in the coal mine according to the status information of the target vehicle and the status information of the non-target vehicle, and control the movement of the target vehicle and the non-target vehicle according to the distribution, wherein the distribution refers to the positional distribution of the vehicles in the coal mine. The device further includes a first construction unit, a second construction unit, and a determination unit. The first construction unit is used to construct a first model after controlling the target vehicle to transport the target object. The output of the first model is the reward result of the actual target vehicle selected based on the relevant information of the target object and the status information of multiple vehicles. The second construction unit is used to construct a second model. The output of the second model is the reward result of the predicted target vehicle selected based on historical data, where the historical data refers to the historical relevant information of the target object and the historical status information of multiple vehicles obtained within a historical time period. The determination unit is used to determine the accuracy rate of order dispatch based on the output results of the first model and the second model. The first construction unit includes a first construction module and a second construction module. The first construction module is used to take the timely completion of orders by the actual target vehicle as a positive reward value, the relationship between the transportation cost and transportation mileage of the actual target vehicle in completing the order as a first negative reward value, and the downtime of the actual target vehicle in completing the order as a second negative reward value, and obtain the sum of the positive reward value, the first negative reward value, and the second negative reward value to obtain an initial first model. The second construction module is used to sum the output results of multiple initial first models of the actual target vehicle within a target time period to obtain the first model. The second building unit includes a third building module, which is used to build the second model. The second model is trained using multiple sets of training data. Each set of training data includes historical information related to the target object acquired within a historical time period, historical state information of multiple vehicles, and historical reward results. The second model is embodied by the Bellman expectation equation: , This means that the output of the second model equals the immediate reward at the current moment plus the discounted expected value of the future value function. Indicates the expected value. Let γ represent the reward value at time t+1, and γ be the discount factor. Let S represent the value function at time t+1, and let S represent the vehicle's state information. This represents the vehicle's state information at time t.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program performs the method according to any one of claims 1 to 4.

7. An electronic device, characterized in that, include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising methods for performing any one of claims 1 to 4.

Citation Information

Patent Citations

  • Freight adjustment and control method for combined spelling

    CN112801336A

  • Mining service transportation trackless rubber-tyred vehicle scheduling method and device

    CN114626717A