Mine sump coal slime cleaning and digging path planning method
By applying reinforcement learning and neural networks in the siltation environment of the underground water tank, dynamically planning the clearing path, the problem of silting path planning in the existing technology relying on manual experience, and a safe and efficient silting process is achieved.
Patent Information
- Application Number
- CN202411844593.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, the lack of automatic path planning during the silt process of underground water tanks has resulted in the reliance on manual experience in the clearing and excavation path planning, which can easily lead to the risk of coal slime collapse and robots being buried.
Adopting a reinforcement learning framework, a rasterized model of the water tank clearing environment is established, and a Q network is built through a fully connected neural network to realize dynamic optimal decision-making of the clearing path.
Dynamic path planning of underground silt robots is realized, ensuring the minimum cost of silt, the fastest silt and the minimum risk, and improving the safety and efficiency of the silt process.
Smart Images

Figure CN119940668A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of mine water bunker excavation, and in particular to a mine water bunker coal slime excavation path planning method. Background Art
[0002] The mine water produced during the coal mine production process contains a large amount of impurities, mainly coal slime and sand, which settle at the bottom of the water tank, reducing the effective volume of the water tank. If the siltation is not cleared in time, the water tank will not only fail to play a buffering and storage role, but may also cause a flooding accident; the impurities in the gushing water will also enter the suction well of the main pump room, causing the wear of the flow parts of the water pump to increase and accelerate the damage of the water pump; when there is too much silt in the suction well, it will also block the suction tap, causing drainage difficulties or even impossible drainage, which seriously threatens the safety of the mine. Therefore, timely siltation of mine water tanks is one of the effective ways to ensure mine drainage safety, extend the life of water pumps, and improve the efficiency of drainage systems.
[0003] At present, there are few studies on automatic path planning in the process of underground water tank desilting. In the project, it is determined manually on site by workers, and the path and method of desilting are all determined by the experience of the operators. During the desilting process, if the path planning of the desilting robot is incorrect, it may cause the thick layer of coal slime to collapse and the robot to be buried. Therefore, it is necessary to perform dynamic path planning for the desilting process of underground water tanks to ensure safe and efficient operation of desilting.
[0004] The patent of this invention proposes to use reinforcement learning as a framework, take the minimum excavation cost as the goal, establish a rasterized model of the silo excavation environment, and on this basis use reinforcement learning to achieve dynamic optimal decision-making of the excavation path. Summary of the invention
[0005] In order to solve the problems existing in the background technology, the present invention proposes a method for planning the path of coal slime excavation in a water silo in an underground mine.
[0006] A method for planning a path for clearing coal slime in a water bunker in a mine, comprising the steps of:
[0007] S100, obtaining image information of the silt removal area at the bottom of the water tank;
[0008] S200, performing rasterization processing on the desilting area in the image;
[0009] S300, identifying and obtaining siltation degree information of the desilting area, and quantifying it into each grid respectively;
[0010] S400, constructing a dredging path planning model;
[0011] S500 , planning a dredging path through a dredging path planning model according to the siltation degree information of each grid and the position information of the dredging robot.
[0012] Based on the above, in step S100, image information of the desilting area at the bottom of the water tank is obtained through machine vision and / or laser radar.
[0013] Based on the above, in step S200, the grid width is the cleaning width of the dredging equipment each time, and the number of rows and columns of the dredging area grid is:
[0014] List OK
[0015] Among them, the width of the bottom of the water tank is w, the height is h, and the cleaning width is d; if a square grid is used, the width of the last column of grids is w-(n-1)d; the height of the last row of grids is h-(m-1)d.
[0016] Based on the above, in step S300, a siltation degree classification and recognition neural network is constructed; the acquired real-time water tank image is sent to the siltation degree classification and recognition neural network for recognition and output of the siltation degree information of the water tank.
[0017] Based on the above, in step S300, the siltation degree of the grid in the i-th row and j-th column is:
[0018] yij∈[0,1]
[0019] Among them, 0 means no siltation and 1 means severe siltation.
[0020] Based on the above, in step S400, a fully connected neural network is used to establish a Q network, represented as Q(s, a; w), and the input is the state variable s k , the input dimension is mn+1; the output is Q(s k ,a k ), the output dimension is four directional dimensions; action a k for:
[0021]
[0022] ε∈(0,1) is the generated random number, ε0 is a constant, and A is the action space.
[0023] Based on the above, the decision-making process of dredging route planning includes the following elements:
[0024] (1) State space S: the space spanned by all siltation levels of the dredging environment and the current position;
[0025] (2) State variable s k :The siltation degree of the dredging environment at the k-th decision step is Current position is 5 k ,but
[0026] (3) Action space A: The space composed of the directions of the dredging robot’s movement. Let a1, a2, a3, and a4 represent the four movements of up, down, left, and right respectively. The action space is expressed as
[0027]
[0028] (4) Execute action
[0029] (5) Single-step return value r k+1|k :Taking the minimum dredging cost as the goal, the dredging speed v is introduced k+1|k and risk factor They are defined as follows
[0030]
[0031] The reward value function is
[0032]
[0033] Among them, λ∈(0,1) represents the proportion of dredging speed in the return value, which is a constant.
[0034] Compared with the prior art, the present invention has outstanding substantial features and significant progress. Specifically, the present invention establishes a dredging path planning model of a neural network, takes the minimum dredging cost as the goal, establishes a rasterized model of the sump dredging environment, and adopts a reinforcement learning method to realize dynamic optimal decision-making of the dredging path, thereby realizing dynamic path planning of the underground dredging robot, taking the minimum dredging cost as the goal, and ensuring the fastest dredging and minimum risk at the same time. It has the advantages of being simple, convenient, efficient and practical. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a schematic diagram of the actual water tank of the present invention and its rasterization.
[0036] Figure 2 It is a schematic diagram of the visualization result after the siltation degree modeling after the dredging area is rasterized in the present invention.
[0037] Figure 3 It is a schematic diagram of modeling the decision-making process of dredging path planning of the present invention.
[0038] Figure 4 It is a schematic diagram of an underground water tank environment and equipment installation of the present invention.
[0039] Figure 5 It is a schematic diagram of the corresponding relationship between the position label and the calibration disk of the present invention.
[0040] Figure 6 It is a schematic diagram of the image coordinate system and the recognition area of the present invention.
[0041] Figure 7 It is a schematic diagram of the neural network structure for classification and identification of siltation degree of the present invention. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] A method for planning a path for coal slime excavation in a water tank in a mine comprises the following steps: S100, obtaining image information of a desilting area at the bottom of the water tank; S200, performing rasterization processing on the desilting area in the image; S300, identifying and obtaining siltation degree information of the desilting area, and quantifying it to each grid respectively; S400, constructing a desilting path planning model; S500, planning a desilting path through the desilting path planning model according to the siltation degree information of each grid and the position information of the desilting robot.
[0044] Specifically, Figure 1 As shown, taking a water tank in a well as an example, Figure 1 (a) is an image of the actual siltation of the underground water tank. The machine vision (existing machine vision equipment such as high-definition cameras, etc.) installed on the top of the dredging robot and the top laser radar are used to obtain the siltation image of the bottom area of the water tank. On this basis, the rasterization abstract processing of the dredging area is completed, as shown in Figure 2. Figure 1 (b) as shown.
[0045] Assume that the width of the bottom of the water tank is w and the height is h. A square grid is used, and the grid width is the width of the dredging equipment at one time, usually the width of the dredging drum, set to d. Then the number of dredging areas to be gridded is: OK in Indicates the rounding down of a floating point number. After rasterization abstraction, the following is formed Figure 1 (c) shows the grid environment. The width of the last column grid is w-(n-1)d; the height of the last row grid is h-(m-1)d.
[0046] The siltation degree of the bottom area is obtained by machine vision or laser radar. In this embodiment, a water tank in a well is taken as an example. Figure 4As shown in the figure, it is a schematic diagram of the environment and equipment installation of a water tank in a mine. A camera and a light source are set on the top of the water tank and face the direction of the water tank respectively. The camera is the core equipment for environmental perception. A camera that meets the safety standards of coal mine underground is selected. In this embodiment, the resolution is 1080P, that is, the single-frame image pixel is 1920×1080. The light source is a device that provides lighting for the water tank. In this embodiment, a light with a focusing unidirectional irradiation function that meets the safety of coal mine underground is selected, with a power of about 150W. In other embodiments, the power or number of light sources can be increased or decreased according to actual conditions. The position sign is used to refer to the position information. The position sign is installed on the top of the water tank. There are at least two position signs, and the one farthest from the camera must be installed on the rear retaining wall of the water tank. Increasing the number of position signs can increase the recognition accuracy of the position. In this embodiment, four are taken as an example; in other embodiments, a reasonable number can be set according to actual conditions. In addition, during the installation of the position sign, the following two requirements must be met: 1) A certain fixed position of all signboards, such as the corner point of a corner, must be on a line; 2) The rear signboard cannot be blocked by the front signboard in the imaging. The calibration plate is composed of a two-color checkerboard and is used to calibrate the vertical position of the position sign. The calibration plates are located in the same horizontal plane and each calibration plate corresponds to a position sign and is vertically arranged below the position sign. A reference angle of the calibration plate is on a vertical line with the reference angle of the corresponding position sign. Figure 5 The camera captured the image of the siltation in the water tank, as shown in Figure 6 As shown in the figure, assume that the image width is D pixels and the height is H pixels, and establish the coordinate systems x and y along the width and height respectively. A1, A2, A3, and A4 represent the image points of the reference corner points of the position signs No. 1 to No. 4 in the image; B1, B2, and B3 represent the image points of the reference corner points corresponding to the calibration disk in the image. Since the water tank, location sign, calibration plate and camera are all fixed, after acquiring an image and establishing the image coordinate system, the image coordinate system will be applicable to each image. By manually analyzing and measuring pixels, the coordinate information of A1-A4 and B1-B3 and the linear function relationship of the two sides of the bottom of the water tank in the image coordinate system can be obtained. Then, the coordinate point information of C1-C3 is obtained according to the coordinate information of B1-B3. Then, the coordinate range corresponding to the recognition area in the figure is divided according to the coordinate points B1-B3 and C1-C3, that is, the area enclosed by B1B2C2C1 (S1 area), the area enclosed by C2B2B3C3 (S2 area) and the area enclosed by C3B3 and the endpoints of the two sides ( Figure 6 The area enclosed by the image coordinate system (not shown in the figure) (S3 area) is divided into three recognition areas. Since the recognition area is also fixed, it can be applied to other subsequent monitoring images after one division and calibration. After the image coordinate system and the recognition area are calibrated, the calibration disk can also be removed. Construct a neural network for classification and recognition of siltation degree, and select ResNet50 as the main body of the classification network. The structure is as follows Figure 7As shown, 1) first use the gradient operator module to extract the gradient information of the image; 2) in the splicing module, the original image and the gradient information are spliced into new feature information, and the dimension of the information is Ih×Id×4, where Ih represents the height of the image, Id represents the width of the image, and 4 represents the number of feature information channels, which are GRB three channels and gradient channels respectively; 3) the Resize module resets the width and height of the image to a 224X224 image to form 224×224×4 feature information; 4) ResNet50 is the backbone network, where the number of channels of the input layer is set to 4, and the output of the final output layer is 1000 categories; the rest of the settings remain the original settings; 5) Linear is a fully connected layer, whose input is 1000 and output is 3, forming a discriminant classification of the siltation degree of the underground water tank; in this embodiment, the three categories of category 1 represent no siltation, category 2 represent a small amount of siltation, and category 3 represent a large amount of siltation. A large number of data samples are collected, and the samples are divided into three categories according to the workers' experience: no siltation, a small amount of siltation, and a large amount of siltation. They are stored in three folders for use in training neural networks. The number of samples in each category is kept balanced, and the number of samples in a single category is at least 100. When training the neural network, the cross entropy loss function is selected as the loss function, the Adam is selected as the optimizer, the learning rate is 0.01, the location of the data sample folder is input, and training and classification are performed according to the number of folders. After the acquired real-time water tank image is divided into regions, the image of each recognition area is sent to the trained siltation degree classification recognition neural network for recognition, and the siltation degree information of each recognition area is obtained respectively. In this embodiment, only three recognition areas are taken as an example to illustrate the recognition principle and process of the siltation degree, and in practice, more recognition areas can be divided according to the demand for the number of grids, that is, each grid is a recognition area; since the recognition of the siltation degree is not the invention point of this application, it will not be repeated.
[0047] After obtaining the identification and classification results of each grid identification area, both classification 1 (indicating no siltation) and classification 2 (indicating a small amount of siltation) are marked as 0, and classification 3 (indicating a large amount of siltation) is marked as 1. The siltation degree of each place is quantified to Figure 1 (c) In the corresponding grids. Let y ij represents the siltation degree of the grid in row i and column j, y ij∈[0,1] , where 0 represents no siltation and 1 represents the most severe siltation. Figure 2 Visualization of the modeled siltation level in the rasterized desilted area.
[0048] Modeling the decision-making process of dredging path planning:
[0049] Assume that the current time is k, and the current position of the dredging robot is e in the environment k =E ij ; The current location of the siltation level is The key elements of dredging process decision modeling are as follows: Figure 3 As shown:
[0050] (1) State space S: represents the space spanned by all siltation levels of the dredging environment and the current position;
[0051] (2) State variable s k :The siltation degree of the dredging environment at the k-th decision step is Current position is 5 k , then $ k 67,
[0052] (3) Action space A: The space composed of the directions of the dredging robot's movement. Let a1, a2, a3, and a4 represent the four movements of up, down, left, and right respectively. The action space is expressed by the following formula:
[0053]
[0054] (4) Execute action
[0055] (5) Single-step return value r k+1|k :Taking the minimum dredging cost as the goal, the dredging speed v is introduced k+1|k and risk factor They are defined as follows
[0056]
[0057] Formula (2) is the dredging speed, which can be adjusted according to the needs. If the next moment enters the environment e k+1 The larger the value of, the more sludge can be cleaned at that location. Formula (3) is a risk system, which is adjusted according to needs. If the next moment enters the environment e k+1 The higher the degree of surrounding siltation, the more likely it is that the current silt removal will cause danger. The exponential function is used in both formulas to consider the silt removal speed and risk factor at the same order of magnitude from 0 to 1. Therefore, the return value function is defined as
[0058]
[0059] In the formula, λ∈(0,1) represents the proportion of dredging speed in the return value, which is usually set as a constant. k+1|k The higher the value, the lower the excavation cost.
[0060] If the overall siltation level is low, indicating low risk, speed is prioritized; if the overall siltation level is high, indicating high risk, risk control is prioritized. Therefore, λ is selected as follows:
[0061]
[0062] In the formula, the thresholds 0.2, 0.8 and 0.4 are all set according to actual engineering conditions.
[0063] Modeling of dredging path planning model:
[0064] According to the aforementioned dredging (or warehouse clearing) path planning decision process modeling, the Deep Q-learning architecture is used to dynamically plan the path of the dredging robot.
[0065] (1) Use neural network to establish Q network.
[0066] The Q network uses a fully connected neural network with an input of s k , so the input dimension is mn+1; the output is Q(s k ,a k ), that is, evaluate in state s k Under this condition, action a k Therefore, the output dimension is 4 (four directions: up, down, left, and right).
[0067]
[0068]
[0069] Let the weight of Q network be w, and Q network is represented by Q(s,a;w);
[0070] (2) Establish Target Q network
[0071] According to the Deep Q learning architecture, a neural network with the same structure as the Q network is established as TargetQnetwork, which is represented by Q T (s,a;w T ). T After Q(s,a;w) is trained to a certain round, w is directly assigned to w T .
[0072] (3) Establish the loss function for Q(s,a;w) neural network training
[0073]
[0074] (4) Action a k Selection Mechanism
[0075] Action a k The e-greedy algorithm is used for selection.
[0076]
[0077] ε∈(0,1) is a generated random number, and ε0 is a constant. When the Q network is in the training phase, ε0 is set to a larger value for easy exploration; in the application phase, the value of ε0 is set to a smaller value to facilitate the neural network to take the optimal value.
[0078] Neural network training process:
[0079] (1) Initialize Q(s, a; w) and pass the value of w directly to Q T (s,a;w T ), complete Q T (s,a;w T ) initialization;
[0080] (2) Initialize the starting state s 0 , set ε0, specify w T Updated training round N;
[0081] (3) According to a k Selection mechanism, according to s k Select action a k ;
[0082] (4) Execute action a k , and get the new state s k+1 ; According to formula (4), the return value r is obtained k+1|k , forming an experience value {s k ,s k+1 ,a k ,r k+1|k}, store the experience in the experience pool;
[0083] (5) Obtain experience from the experience pool, calculate the loss function according to formula (6), return the loss value, and update the weight w;
[0084] (6) Repeat steps (3)-(5) of the training process. When the number of training times is equal to an integer multiple of N, directly assign w to w T . until the neural network converges.
[0085] Use of dredging route planning:
[0086] (1) Input the current environment and the robot's position into Q(s, a; w) to obtain the value of Q(s, a; w);
[0087] (2) Select the execution action based on Q(s,a;w) and equation (7).
[0088] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
Claims
1. A method for planning a path for clearing coal slime in a water silo in a mine, characterized in that: Includes steps: S100, obtaining image information of the silt removal area at the bottom of the water tank; S200, performing rasterization processing on the desilting area in the image; S300, identifying and obtaining siltation degree information of the desilting area, and quantifying it into each grid respectively; S400, constructing a dredging path planning model; S500 , planning a dredging path through a dredging path planning model according to the siltation degree information of each grid and the position information of the dredging robot.
2. The method for planning the path for clearing coal slime in a mine water silo according to claim 1, characterized in that: In step S100, image information of the desilting area at the bottom of the water tank is obtained by machine vision and / or laser radar.
3. The method for planning the path for clearing coal slime in a mine water silo according to claim 1, characterized in that: In step S200, the grid width is the cleaning width of the dredging equipment each time, and the number of rows and columns of the dredging area grid is: List OK Among them, the width of the bottom of the water tank is w , height is h, cleaning width is d ; If a square grid is used, the width of the last column of the grid is w-(n-1)d; the height of the last row of the grid is h-(m-1)d.
4. The method for planning the path for clearing coal slime in a mine water silo according to claim 1, characterized in that: In step S300, a siltation degree classification and recognition neural network is constructed; the acquired real-time water tank image is sent to the siltation degree classification and recognition neural network for recognition and output of siltation degree information of the water tank.
5. The method for planning the path for clearing coal slime in a water silo in a mine according to claim 4, characterized in that: In step S300, the siltation degree of the grid in row i and column j is: y ij ∈[0,1] Among them, 0 means no siltation and 1 means severe siltation.
6. The method for planning the path for clearing coal slime in a mine water silo according to claim 1, characterized in that: In step S400, a fully connected neural network is used to establish a Q network, represented as Q(s, a; w), and the input is the state variable s k , the input dimension is mn+1; the output is Q(s k ,a k ), the output dimension is four directional dimensions; action a k for: ε∈(0,1) is the generated random number, ε0 is a constant, and A is the action space.
7. The method for planning the path for clearing coal slime in a mine water silo according to claim 1, characterized in that: The decision-making process for desilting route planning includes the following elements: (1) State space S: the space spanned by all siltation levels of the dredging environment and the current position; (2) State variable s k :The siltation degree of the dredging environment at the k-th decision step is Current location is e k , then s k ∈S, (3) Action space A: The space composed of the directions of the dredging robot’s movement. Let a1, a2, a3, and a4 represent the four movements of up, down, left, and right respectively. The action space is expressed as (4) Execute action (5) Single-step return value r k+1|k :Taking the minimum dredging cost as the goal, the dredging speed v is introduced k+1|k and the risk factor ζ k+1|k , respectively defined as follows The reward value function is Among them, λ∈(0,1) represents the proportion of dredging speed in the return value, which is a constant.