Real-time deviation correction method, device and equipment for end sill mining shaft track, medium and product

By using multi-source data fusion and reinforcement learning algorithms for trajectory control of end-face coal mining machines, the problems of large trajectory control errors and low efficiency in existing technologies have been solved. Real-time and accurate adaptive correction has been achieved, improving safety and efficiency.

CN121596754BActive Publication Date: 2026-05-15TAIYUAN INST OF CHINA COAL TECH & ENG GROUP +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610122858.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-29
Publication Date
2026-05-15
Estimated Expiration
2046-01-29

AI Technical Summary

Technical Problem

Existing trajectory control technology for end-face coal mining machines is difficult to achieve real-time and accurate adaptive correction in complex underground environments, resulting in large errors, low efficiency, and potential safety hazards.

Method used

By employing multi-source data fusion and reinforcement learning algorithms, and through UWB positioning, operating condition and vibration data preprocessing, real-time error correction control is achieved using policy networks and dual-value networks, combined with reward functions and PID controllers to realize automatic closed-loop control.

Benefits of technology

It improves the accuracy and efficiency of trajectory correction, reduces human intervention, enhances safety and system stability, and adapts to complex downhole environment changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121596754B_ABST
    Figure CN121596754B_ABST
Patent Text Reader

Abstract

The application discloses a real-time deviation rectification method and device for end-slope mining track, equipment, medium and product, relates to the field of intelligent track control, and comprises the following steps: acquiring original data of an end-slope coal mining machine; determining a current state according to the original data; determining a continuous action by using a strategy network according to the current state; the continuous action is used for rectifying and controlling the end-slope mining track; the continuous action comprises a differential speed of a caterpillar belt, a pushing speed and a roller pitch angle; the strategy network is optimized based on an evaluation result of a double-value network in a training stage; a cumulative return value is determined by using the double-value network according to the current state and the continuous action; the cumulative return value is used for evaluating the continuous action and optimizing the strategy network. The application can realize real-time deviation rectification of the end-slope coal mining machine track.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of trajectory control, and in particular to a method, apparatus, equipment, medium and product for real-time trajectory correction in end-to-end sampling tunnels. Background Technology

[0002] End-face mining is a crucial technology for increasing production and efficiency in open-pit coal mines. It involves mechanized coal mining using parallel mining chambers arranged along the end-face coal pillars. When the coal mining machine operates within the narrow, enclosed chambers, precise trajectory control is paramount. If the trajectory deviates, it can lead to irregular chamber cross-sections and uneven coal pillar widths, resulting in reduced recovery rates. In severe cases, it may cause coal pillar instability, chamber collapse, and other safety accidents, threatening personnel and equipment safety. Therefore, achieving high-precision, adaptive movement of the coal mining machine along a preset trajectory is a core technical challenge for ensuring the safe and efficient operation of end-face mining.

[0003] Currently, the industry mainly relies on two types of technologies for navigation and trajectory control of end-face coal mining machines: one is the offline calibration and dead reckoning method based on laser gyroscopes, fiber optic inertial navigation, etc.; the other is a simple feedback control combining odometers and periodic manual correction. However, these existing technologies are susceptible to magnetic interference and rock powder displacement in practical applications, have limited sensing data, resulting in an error of 3°–5°, and have low working efficiency, making it difficult to meet the real-time and accurate correction requirements under complex underground working conditions.

[0004] Existing trajectory control technology for end-face coal mining machines is limited by the limitations of perception and the lack of self-learning optimization capabilities, making it difficult to achieve real-time and accurate adaptive correction in dynamic, unstructured underground environments. Therefore, there is an urgent need for an optimal control strategy that can integrate multi-source information, evaluate the status in real time, and dynamically generate an optimal control strategy through reinforcement learning algorithms to achieve automated closed-loop control of end-face mining, thereby simultaneously improving operational efficiency and safety. Summary of the Invention

[0005] The purpose of this application is to provide a method, device, equipment, medium, and product for real-time correction of the trajectory of end-slope mining tunnels, so as to improve the efficiency and accuracy of real-time correction of the trajectory of end-slope mining tunnels.

[0006] To achieve the above objectives, this application provides the following solution:

[0007] Firstly, this application provides a method for real-time correction of the trajectory of an end-to-end sampling tunnel, including:

[0008] Acquire multi-source raw data from the end-face coal mining machine, wherein the multi-source raw data includes at least UWB positioning data, operating condition data, vibration data, and environmental data;

[0009] The multi-source raw data is preprocessed, including time synchronization, iterative extended Kalman filter fusion, feature extraction and normalization, to obtain a current state vector that represents the current comprehensive state of the end-side coal mining machine. The current state vector includes at least positioning error, attitude information, working condition characteristics and vibration characteristics.

[0010] The current state vector is input to the strategy network, and the strategy network outputs continuous actions for real-time correction control of the end-side mining tunnel trajectory. The continuous actions include track differential speed, propulsion speed and drum pitch angle.

[0011] During the training phase of the policy network, the current state vector and the continuous actions are input into the dual-value network, which outputs the corresponding cumulative reward value. The continuous actions are evaluated based on the cumulative reward value and fed back to the policy network to optimize the parameters of the policy network. The dual-value network uses two parallel and independent value networks, and takes the minimum value of their outputs as the cumulative reward value. During the inference phase of the policy network, the optimized policy network directly outputs the continuous actions used for real-time correction control of the end-face mining tunnel trajectory.

[0012] In one embodiment, in the dual-value network, the formula for calculating the cumulative return value, which takes the minimum value of the two outputs, is as follows:

[0013]

[0014] in, Indicates the target of the current update. Value, i.e., cumulative return value; This represents the reward value obtained after performing an action; This represents the weighting coefficient used to balance immediate rewards and future rewards; This indicates that the first value function estimate is in The following measures of Value, of which, This represents the state vector at the next moment. For policy networks in The following actions are given; For the second value function estimation in The following measures of value.

[0015] In one embodiment, the method further includes: determining a reward function based on the current state vector and continuous actions, specifically including:

[0016] Based on the next state vector and the current state vector obtained after the execution of the continuous actions, the lateral estimation deviation and heading angle deviation are determined.

[0017] The reward function is determined based on the lateral estimation deviation and the heading angle deviation, and on the preset effective propulsion speed and the corresponding preset propulsion acceleration.

[0018] In one embodiment, the reward function is:

[0019]

[0020] in, In time step The reward value obtained after performing an action; This represents the lateral estimation bias; This refers to the deviation in heading angle; To effectively accelerate progress; To accelerate; The weighting coefficients representing the lateral estimation bias; The weighting coefficient representing the heading angle deviation is used to balance the attitude consistency of the end-face coal mining machine; A weighting coefficient representing the effective advance speed, used to balance the operating efficiency of the end-face coal mining machine; The weighting coefficient representing the propulsion acceleration is used to balance the operational stability of the end-face coal mining machine.

[0021] In one embodiment, the method further includes: monitoring the signal-to-noise ratio of the UWB anchor points that collect the UWB positioning data and the triaxial acceleration of the end-side coal mining machine; when the signal-to-noise ratio is lower than a preset threshold and / or the triaxial acceleration is greater than a set acceleration, the continuous action output by the strategy network is amplitude-clamped and switched to a trajectory control strategy based on a PID controller.

[0022] Secondly, this application provides a real-time trajectory correction device for end-slope mining tunnels, comprising:

[0023] The acquisition module is used to acquire multi-source raw data of the end-side coal mining machine. The multi-source raw data includes at least UWB positioning data, working condition data, vibration data, and environmental data.

[0024] The current state determination module is used to preprocess the multi-source raw data. The preprocessing includes time synchronization, iterative extended Kalman filter fusion, feature extraction and normalization to obtain a current state vector that represents the current comprehensive state of the end-side coal mining machine. The current state vector includes at least positioning error, attitude information, working condition characteristics and vibration characteristics.

[0025] The continuous action determination module is used to input the current state vector into the strategy network, and the strategy network outputs continuous actions for real-time correction control of the end-side mining tunnel trajectory. The continuous actions include track differential speed, propulsion speed and drum pitch angle.

[0026] An evaluation module is used during the training phase of the policy network to input the current state vector and the continuous actions into a dual-value network, whereby the dual-value network outputs a corresponding cumulative reward value. The continuous actions are evaluated based on the cumulative reward value and fed back to the policy network to optimize its parameters. The dual-value network employs two parallel and independent value networks, and the minimum of their outputs is taken as the cumulative reward value. During the inference phase of the policy network, the optimized policy network directly outputs continuous actions for real-time correction control of the end-face mining tunnel trajectory.

[0027] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the computer program to implement the aforementioned real-time trajectory correction method for end-to-end sampling.

[0028] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned real-time trajectory correction method for end-to-end sampling.

[0029] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned real-time trajectory correction method for end-to-end sampling.

[0030] According to the specific embodiments provided in this application, the following technical effects are disclosed:

[0031] This application provides a method, device, equipment, medium, and product for real-time trajectory correction in an end-side mining tunnel, comprising: acquiring multi-source raw data of an end-side coal mining machine; determining a current state vector based on the multi-source raw data; determining continuous actions using a strategy network based on the current state vector; the continuous actions being used for real-time trajectory correction control of the end-side mining tunnel; the continuous actions including track differential speed, propulsion speed, and drum pitch angle; the strategy network optimizing parameters based on the evaluation results of a dual-value network during the training phase; determining a cumulative reward value using the dual-value network based on the current state and continuous actions; and the cumulative reward value being used to evaluate the continuous actions and optimize the strategy network. Automatic closed-loop control without manual intervention is achieved through the strategy network and dual-value network, improving work efficiency. The evaluation based on the dual-value network avoids overfitting and improves error correction accuracy. Attached Figure Description

[0032] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0033] Figure 1 Flowchart of the method for real-time correction of the trajectory of the end-to-end sampling tunnel;

[0034] Figure 2 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0036] The present application will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0037] In one exemplary embodiment, such as Figure 1 As shown, a method for real-time trajectory correction of end-to-end sampling tunnels is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method includes the following steps:

[0038] Step 101: Obtain multi-source raw data from the end-side coal mining machine. The multi-source raw data includes at least UWB positioning data, operating condition data, vibration data, and environmental data.

[0039] Step 102: Preprocess the multi-source raw data. The preprocessing includes time synchronization, iterative extended Kalman filter fusion, feature extraction and normalization to obtain a current state vector that represents the current comprehensive state of the end-side coal mining machine. The current state vector includes at least positioning error, attitude information, working condition characteristics and vibration characteristics.

[0040] Step 103: Input the current state vector into the strategy network, and the strategy network outputs continuous actions for real-time correction control of the end-side mining tunnel trajectory. The continuous actions include track differential speed, propulsion speed and drum pitch angle.

[0041] In this embodiment, the dual-value network evaluates the quality of consecutive actions in parallel during the training phase and feeds the evaluation results back to the policy network to optimize the parameters of the policy network.

[0042] During the training phase: For the same state-action, the two value networks output their respective values ​​in parallel. The smaller of the two values ​​is taken. Value as target The value is then used to guide the policy network in reverse.

[0043] During the inference phase: the policy network directly outputs continuous actions based on the current state vector.

[0044] The role of the dual value network is to "evaluate the quality of actions" without directly outputting actions; the role of the policy network is to "generate actions".

[0045] Step 104: During the training phase of the policy network, the current state vector and the continuous action are input into the dual-value network, and the dual-value network outputs the corresponding cumulative reward value. The continuous action is evaluated based on the cumulative reward value and fed back to the policy network to optimize the parameters of the policy network. The dual-value network uses two parallel and independent value networks, and takes the minimum value of their outputs as the cumulative reward value. During the inference phase of the policy network, the continuous action used for real-time correction control of the end-face mining tunnel trajectory is directly output based on the optimized policy network.

[0046] In this embodiment, the cooperation between the policy network and the dual-value network enables automatic closed-loop control without human intervention, thereby improving work efficiency. The conservative valuation mechanism of the dual-value network can suppress value overestimation and improve trajectory correction accuracy and system stability.

[0047] In an exemplary embodiment, the multi-source raw data is preprocessed to explicitly incorporate uncontrollable disturbances from the end-face mining scenario. The multi-source raw data includes at least UWB positioning data, operating condition data, vibration data, and environmental data. These are then fused through time synchronization and iterative extended Kalman filtering, followed by feature extraction and normalization, to form a current state vector containing positioning error, attitude, operating condition characteristics, and vibration characteristics. This allows subsequent real-time correction control to not only close the loop on geometric errors but also on a "error + load + vibration + environmental quality" loop, better reflecting the actual operating conditions of the end-face mining machine (rock abrupt changes, obstruction, multi-source disturbances).

[0048] In practical applications, the formula for calculating the training objective by taking the minimum value of the two outputs of the dual-value network is as follows:

[0049]

[0050] in, Indicates the target of the current update. Value, i.e., cumulative return value; This represents the reward value obtained after performing an action; This represents the weighting coefficient used to balance immediate rewards and future rewards; This indicates that the first value function estimate is in The following measures of Value, of which, This represents the state vector at the next moment. For policy networks in The following actions are given; For the second value function estimation in The following measures of value.

[0051] In an exemplary embodiment, the method further includes: determining a reward function based on the current state vector and the continuous actions, specifically including: determining the lateral estimation deviation and the heading angle deviation based on the next state vector obtained after the continuous actions are executed and the current state vector; and determining the reward function based on the lateral estimation deviation and the heading angle deviation, and based on a preset effective propulsion speed and a corresponding preset propulsion acceleration.

[0052] In this embodiment, a closed-loop reinforcement learning framework is employed, in which a policy network outputs continuous actions and a dual-value network evaluates the quality of these continuous actions. First, multi-source raw data is collected from multiple sensors. After preprocessing operations such as time synchronization and fusion, the current state vector is obtained. The policy network then outputs continuous actions (track differential speed, propulsion speed, drum pitch angle, etc.) for real-time correction control. Within this closed-loop reinforcement learning framework, a reward function is constructed to quantify the engineering objective of end-slope mining tunnel trajectory correction into an instantaneous scalar feedback, enabling the dual-value network to execute actions better and driving the policy network to iteratively optimize towards higher rewards.

[0053] In practical applications, the reward function is:

[0054]

[0055] in, In time step The reward value obtained after performing an action; This represents the lateral estimation bias; This refers to the deviation in heading angle; To effectively accelerate progress; To accelerate; The weighting coefficients representing the lateral estimation bias; The weighting coefficient representing the heading angle deviation is used to balance the attitude consistency of the end-face coal mining machine; A weighting coefficient representing the effective advance speed, used to balance the operating efficiency of the end-face coal mining machine; The weighting coefficient representing the propulsion acceleration is used to balance the operational stability of the end-face coal mining machine.

[0056] In this embodiment, the reward function is designed based on four types of hard indicators in the actual application scenario of end-to-end sampling correction, and these four types of hard indicators are mapped to four observable and computable quantities:

[0057] 1. Trajectory accuracy (lateral deviation of the end-face coal mining machine): Lateral estimation deviation It directly reflects the lateral deviation between the centerline of the end-side coal mining machine and the target trajectory.

[0058] 2. Attitude / Direction Consistency (Heading Deviation of End-Side Coal Mining Machine): Heading Angle Deviation This reflects the deviation of the end-face coal mining machine from the direction relative to the target.

[0059] 3. Operational efficiency (contribution of the end-face coal mining machine to propulsion): effective propulsion speed This reflects the speed of movement along the target direction.

[0060] 4. Operational stability (impact / violent movement suppression of end-face coal mining machine): propulsion acceleration It is a quantity that characterizes the smoothness of the executed action and is used to suppress the mechanical shock, vibration and control divergence risks caused by sudden acceleration, sudden stop and sudden turn.

[0061] Therefore, the reward function is structurally composed of various reward and penalty terms: deviation term (the larger the deviation, the greater the penalty) + efficiency term (the larger the efficiency, the greater the reward) + stability term (the more severe the stability, the greater the penalty), and weight coefficients are set to balance the importance of each reward and penalty term.

[0062] In this embodiment, the core purpose of using absolute values ​​for lateral estimation deviation and heading angle deviation is to ensure that the corresponding reward / penalty items only consider the magnitude of the deviation, without considering the direction of the deviation. In practical applications, whether the real-time trajectory deviates to the left or right, or whether a left or right yaw occurs, there is a safety risk. Therefore, absolute value processing is used to achieve symmetrical penalties for both sides.

[0063] An increase in the absolute value of lateral estimation deviation and heading angle deviation indicates a worsening of correction and an increased risk of narrowing coal pillars. Therefore, using a negative sign to increase the corresponding values ​​will reduce the reward value. Effective propulsion speed reflects operational efficiency and propulsion contribution. Therefore, a positive sign is used to encourage forward movement rather than idling or stopping conservatively. The greater the propulsion acceleration, the stronger the impact, vibration, and control irregularities, and the easier it is to cause equipment failure or model divergence. Therefore, a negative sign is used for calculation.

[0064] In practical applications, the variables selected for constructing the reward function must meet three criteria:

[0065] 1. Observable: All data can be obtained directly from state construction or from the fusion results of multi-source sensors (observable data includes at least UWB positioning data, operating condition data, vibration data, and environmental data), and can be continuously obtained in a closed mining environment.

[0066] 2. Controllable: The continuous actions output by the strategy network (track differential speed, propulsion speed, drum pitch angle) have a direct impact on the above observable data, thus forming a closed-loop link of "action → status → reward → update".

[0067] 3. Safe to implement: When the environmental quality deteriorates or the impact is too great, this application has a safety protection with action amplitude clamping and switching control strategy, which makes the reward function driven learning and control have engineering usability.

[0068] In practical applications, UWB positioning data collected by ultra-wideband (UWB) technology is obtained through UWB anchor points deployed on the top of the mining tunnel. Each UWB anchor point is installed on the top of the mining tunnel via a quick-release bracket at the top of the retractable carbon fiber skeleton support rod on the end-face coal mining machine. The spacing between adjacent UWB anchor points is 25m. Each UWB anchor point is installed on the quick-release bracket at the top of the retractable carbon fiber skeleton support rod on the end-face coal mining machine. The upper and lower ends of the retractable carbon fiber skeleton support rod are respectively equipped with anti-slip rubber pads and a spiral tightening mechanism. Each UWB anchor point is equipped with a microelectromechanical system (MEMS) steering antenna, which automatically optimizes the angle to improve the signal-to-noise ratio of the UWB anchor point.

[0069] The real-time trajectory correction method for end-side mining tunnels also includes monitoring the signal-to-noise ratio of the UWB anchor points that collect the UWB positioning data and the triaxial acceleration of the end-side coal mining machine; when the signal-to-noise ratio is lower than a preset threshold and / or the triaxial acceleration is greater than a set acceleration, the continuous action output by the strategy network is amplitude-clamped and switched to a trajectory control strategy based on a PID controller.

[0070] This application utilizes ultra-wideband (UWB) positioning-inertial measurement unit (IMU) fusion combined with reinforcement learning (RL) to correct tunnel trajectory deviations in real time during end-face excavation, while maintaining parallel / dip control. This application can maintain UWB signal coverage and suppress multipath effects within a 300m-long closed mining tunnel, converting IMU-UWB fusion positioning results into real-time attitude fine-tuning commands for the end-face coal mining machine, and adaptively adjusting the correction strategy under complex coal and rock variations.

[0071] In this embodiment, when the signal-to-noise ratio of the UWB anchor point is lower than a preset threshold or the triaxial acceleration of the end-face coal mining machine is greater than a set acceleration, the continuous action output by the strategy network is amplitude-clamped, and the system switches to a PID controller for control. This provides engineering safety protection for the solution in this application. This application constructs a hierarchical control strategy of reinforcement learning + PID. Specifically, it monitors the signal-to-noise ratio of the UWB anchor point and the triaxial acceleration of the end-face coal mining machine in real time. Once the signal quality is poor or the vibration is too large (hard rock impact is detected), the system automatically clamps the amplitude of the strategy network output and forces a switch back to the classic PID controller for trajectory control.

[0072] In this embodiment, the input to the policy network is the current state vector (such as positioning error, attitude, working condition characteristics, and vibration characteristics), and the output is continuous actions (such as track differential speed, propulsion speed, and drum pitch angle). The goal is to update the output continuous actions in a direction with high value.

[0073] The input to the dual-value network is the current state vector and the continuous actions, and the output is the estimated cumulative reward value. The objective is to minimize the Bellman error, which is a key indicator for evaluating the accuracy of the value function or policy network optimization in reinforcement learning. Its mathematical essence is the residual of the Bellman equation, which reflects the deviation between the current value network estimate and the theoretical optimal value.

[0074] In this embodiment, an experience replay pool is also constructed to improve sample efficiency and training stability in reinforcement learning.

[0075] A slow-updating target network is maintained for both the policy network and the dual-value network, and updated according to a predetermined iteration path each time. This is used to stabilize the training process and avoid valuation oscillations.

[0076] This application also applies a deep reinforcement learning algorithm (Deep Deterministic Policy Gradient, DDPG) for continuous action space to trajectory correction in end-face mining tunnels, constructing a scenario of "end-face mining UWB-IMU fusion + DDPG correction control". The specific implementation method is as follows:

[0077] State Vector Setting: In the end-to-end sampling trajectory correction scenario of this application, the state vector is composed of multi-source sensor data after fusion and feature extraction, mainly including: positioning error (lateral deviation and attitude deviation calculated by UWB-IMU fusion), attitude information (nose pitch angle, roll angle, and heading angle), operating condition characteristics (indicators reflecting load status such as propulsion current, cutting load, and oil pressure), and vibration characteristics (root mean square value and main frequency amplitude obtained from triaxial acceleration analysis).

[0078] Set continuous motion: propulsion speed, drum pitch angle, track differential speed.

[0079] The reward value obtained after performing an action is set: it can be incorporated into the instantaneous performance function based on the negative value of the target trajectory deviation penalty, the positive reward of cutting efficiency, and the operational stability (avoiding excessive vibration). It represents the instantaneous performance function value obtained by performing continuous actions under the state vector, and is used to comprehensively evaluate the action effect.

[0080] The training process of DDPG: The agent built based on DDPG continuously interacts with the simulation environment (or historical working condition data) to gradually learn real-time correction strategies for different coal seam hardness and resistance conditions, and finally achieves the control objective of stabilizing the trajectory deviation of the coal mining machine within the allowable range.

[0081] Through the structure of a policy network-dual value network and offline experience playback, the intelligent agent can efficiently learn the optimal control strategy in the continuous action space. When applied to the trajectory correction of a coal mining machine, it can achieve closed-loop intelligent control with "unmanned intervention and real-time correction".

[0082] This application presents a real-time deviation correction method for DDPG based on a dual-value network structure, designed to address the problem of trajectory deviation estimation distortion caused by factors such as UWB signal obstruction, multi-source disturbances, and abrupt lithological changes during underground endwall coal mining. The method introduces a second value network into the traditional DDPG algorithm, using the smaller of the two value networks as the target value for training. This effectively alleviates the problems of "over-correction" and "estimation oscillation" in trajectory control of the policy network.

[0083] The multi-source raw data collected in this application are shown in Table 1.

[0084] Table 1

[0085]

[0086] The preprocessing process of multi-source raw data is as follows: hard timestamp synchronization (unified sampling frequency) is performed on the multi-source raw data, and then iterative extended Kalman filter fusion (UWB+IMU), FFT (Fast Fourier Transform) / sliding window processing (working condition + disturbance) and normalization processing are performed in sequence to finally obtain the current state vector representing the current comprehensive state of the end-side coal mining machine.

[0087] After time synchronization, fusion, and feature extraction of the multi-source raw data, a current state vector of approximately 30-40 dimensions is formed, containing positioning error, attitude information, operating condition characteristics, and vibration characteristics. It also includes additional information such as lateral estimation deviation, heading angle deviation, historical actions, and environmental perception. This current state vector is used as input to the policy network to generate continuous actions (track differential speed, propulsion speed, and drum pitch angle) to achieve trajectory correction control. The composition of the current state vector is shown in Table 2.

[0088] Table 2

[0089]

[0090] Based on the contents of Table 2, a multi-dimensional current state vector is formed, which is then input into the policy network and the dual-value network and stored in the experience replay pool.

[0091] In this application, the current state vector is also input into the architecture of "end-side mining UWB-IMU fusion + DDPG correction control", and the output of this architecture is shown in Table 3 below.

[0092] Table 3

[0093]

[0094] In practical applications, the experience replay pool stores quadruples. ,in, This represents the current state vector, which includes positioning error, attitude information, operating condition characteristics, and vibration characteristics. This represents the state vector at the next moment; Indicates continuous motion (track differential speed, propulsion speed, roller pitch angle); In time step The reward value obtained after performing an action.

[0095] All network training (loss calculation for dual-value networks, policy updates for policy networks) samples training data from this experience replay pool.

[0096] The significance of the experience replay pool constructed in this application is to ensure that the end-face coal mining machine, during the actual mining tunnel trajectory correction process, does not need to rely on the latest data, but can "learn repeatedly from past experience," thereby training a more stable and reliable correction strategy. A sample group of a preset batch size is selected from the experience replay pool to calculate the minimum value of the dual-value network and synchronously update the strategy network.

[0097] Based on the same inventive concept, this application also provides a real-time trajectory correction device for end-slope mining tunnels to implement the aforementioned real-time trajectory correction method. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the real-time trajectory correction device for end-slope mining tunnels provided below can be found in the limitations of the real-time trajectory correction method for end-slope mining tunnels described above, and will not be repeated here.

[0098] In one exemplary embodiment, a real-time trajectory correction device for end-slope sampling tunnels is provided, comprising:

[0099] The acquisition module is used to acquire multi-source raw data of the end-side coal mining machine. The multi-source raw data includes at least UWB positioning data, working condition data, vibration data, and environmental data.

[0100] The current state determination module is used to preprocess the multi-source raw data. The preprocessing includes time synchronization, iterative extended Kalman filter fusion, feature extraction and normalization to obtain a current state vector that represents the current comprehensive state of the end-side coal mining machine. The current state vector includes at least positioning error, attitude information, working condition characteristics and vibration characteristics.

[0101] The continuous action determination module is used to input the current state vector into the strategy network, and the strategy network outputs continuous actions for real-time correction control of the end-side mining tunnel trajectory. The continuous actions include track differential speed, propulsion speed and drum pitch angle.

[0102] An evaluation module is used during the training phase of the policy network to input the current state vector and the continuous actions into a dual-value network, whereby the dual-value network outputs a corresponding cumulative reward value. The continuous actions are evaluated based on the cumulative reward value and fed back to the policy network to optimize its parameters. The dual-value network employs two parallel and independent value networks, and the minimum of their outputs is taken as the cumulative reward value. During the inference phase of the policy network, the optimized policy network directly outputs continuous actions for real-time correction control of the end-face mining tunnel trajectory.

[0103] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 2As shown, the computer device includes a processor, memory, input / output interfaces, and a communication interface. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores real-time trajectory correction data for end-to-end sampling tunnels. The input / output interfaces allow the processor to exchange information with external devices. The communication interface allows communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a real-time trajectory correction method for end-to-end sampling tunnels.

[0104] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the above-described method embodiments.

[0105] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described method embodiments.

[0106] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0107] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for real-time trajectory correction of a tunnel end-slope sampling method, characterized in that, include: Acquire multi-source raw data from the end-face coal mining machine, wherein the multi-source raw data includes at least UWB positioning data, operating condition data, vibration data, and environmental data; The multi-source raw data is preprocessed, including time synchronization, iterative extended Kalman filter fusion, feature extraction and normalization, to obtain a current state vector that represents the current comprehensive state of the end-side coal mining machine. The current state vector includes at least positioning error, attitude information, working condition characteristics and vibration characteristics. The current state vector is input to the strategy network, and the strategy network outputs continuous actions for real-time correction control of the end-side mining tunnel trajectory. The continuous actions include track differential speed, propulsion speed and drum pitch angle. During the training phase of the policy network, the current state vector and the continuous actions are input into a dual-value network, which outputs a corresponding cumulative reward value. The continuous actions are evaluated based on the cumulative reward value and fed back to the policy network to optimize its parameters. The dual-value network uses two parallel and independent value networks, and the minimum value of their outputs is taken as the cumulative reward value. During the inference phase of the policy network, the optimized policy network directly outputs continuous actions for real-time correction control of the end-face mining tunnel trajectory. The signal-to-noise ratio of the UWB anchor points that collect the UWB positioning data and the triaxial acceleration of the end-side coal mining machine are monitored. When the signal-to-noise ratio is lower than a preset threshold and / or the triaxial acceleration is greater than a set acceleration, the amplitude of the continuous action output by the strategy network is clamped and switched to a trajectory control strategy based on a PID controller.

2. The method for real-time correction of the trajectory of the end-face mining tunnel according to claim 1, characterized in that, In the dual-value network, the formula for calculating the cumulative return value, taking the minimum of the two outputs, is as follows: in, Indicates the target of the current update. Value, i.e., cumulative return value; This represents the reward value obtained after performing an action; This represents the weighting coefficient used to balance immediate rewards and future rewards; This indicates that the first value function estimate is in The following measures action Value, of which, This represents the state vector at the next moment. For policy networks in The following actions are given; For the second value function estimation in The following measures action value.

3. The method for real-time trajectory correction of end-face mining tunnels according to claim 1, characterized in that, Also includes: Based on the current state vector and continuous actions, the reward function is determined, specifically including: Based on the next state vector and the current state vector obtained after the execution of the continuous actions, the lateral estimation deviation and heading angle deviation are determined. The reward function is determined based on the lateral estimation deviation and the heading angle deviation, and on the preset effective propulsion speed and the corresponding preset propulsion acceleration.

4. The method for real-time trajectory correction of end-face mining tunnels according to claim 3, characterized in that, The reward function is: in, In time step The reward value obtained after performing an action; This represents the lateral estimation bias; This refers to the deviation in heading angle; To effectively accelerate progress; To accelerate; The weighting coefficients representing the lateral estimation bias; The weighting coefficient representing the heading angle deviation is used to balance the attitude consistency of the end-face coal mining machine; A weighting coefficient representing the effective advance speed, used to balance the operating efficiency of the end-face coal mining machine; The weighting coefficient representing the propulsion acceleration is used to balance the operational stability of the end-face coal mining machine.

5. A real-time trajectory correction device for end-face mining tunnels, characterized in that, For performing the real-time trajectory correction method for end-slope mining tunnels as described in any one of claims 1-4, the real-time trajectory correction device for end-slope mining tunnels comprises: The acquisition module is used to acquire multi-source raw data of the end-side coal mining machine. The multi-source raw data includes at least UWB positioning data, working condition data, vibration data, and environmental data. The current state determination module is used to preprocess the multi-source raw data. The preprocessing includes time synchronization, iterative extended Kalman filter fusion, feature extraction and normalization to obtain a current state vector that represents the current comprehensive state of the end-side coal mining machine. The current state vector includes at least positioning error, attitude information, working condition characteristics and vibration characteristics. The continuous action determination module is used to input the current state vector into the strategy network, and the strategy network outputs continuous actions for real-time correction control of the end-side mining tunnel trajectory. The continuous actions include track differential speed, propulsion speed and drum pitch angle. An evaluation module is used during the training phase of the policy network to input the current state vector and the continuous actions into a dual-value network, whereby the dual-value network outputs a corresponding cumulative reward value. The continuous actions are evaluated based on the cumulative reward value and fed back to the policy network to optimize its parameters. The dual-value network employs two parallel and independent value networks, and the minimum of their outputs is taken as the cumulative reward value. During the inference phase of the policy network, the optimized policy network directly outputs continuous actions for real-time correction control of the end-face mining tunnel trajectory.

6. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that the processor executes the computer program to implement the real-time trajectory correction method for end-to-end sampling tunnels as described in any one of claims 1-4.

7. A computer-readable storage medium storing a computer program, characterized in that, When executed by a processor, the computer program implements the real-time trajectory correction method for end-to-end sampling tunnels as described in any one of claims 1-4.

8. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the real-time trajectory correction method for end-to-end sampling tunnels as described in any one of claims 1-4.