Autonomous walking track planning method and device for heading machine based on reinforcement learning algorithm

Through autonomous walking trajectory planning using reinforcement learning algorithms, the problems of collision and path deviation of traditional tunnel boring machines in complex tunnel environments have been solved, achieving safe and efficient autonomous operation.

CN120630977APending Publication Date: 2025-09-12SHANXI TIANDI COAL MINING MACHINERY +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510645220.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The traditional way of traveling of tunnel boring machines relies on manual control, which makes them prone to collisions and path deviations in complex tunnel environments, affecting efficiency and safety.

Method used

An autonomous walking trajectory planning method based on reinforcement learning algorithm is adopted. By collecting the body posture and lane data in real time, a trajectory planning model is constructed, a reward function is designed, and the reinforcement learning algorithm is applied to adjust the trajectory to achieve autonomous decision-making.

Benefits of technology

It improves the safety and efficiency of roadheader operations in complex tunnel environments, reduces manual intervention, and can flexibly respond to different geological conditions, avoid collisions, and optimize paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120630977A_ABST
    Figure CN120630977A_ABST
Patent Text Reader

Abstract

The invention discloses an autonomous walking track planning method and equipment of a heading machine based on a reinforcement learning algorithm, and the method achieves the optimal track planning of the heading machine from the current position to a roadway head-on by comprehensively considering the key parameters of the real-time body pose (including the position, the pitch angle, the yaw angle and the roll angle) of the heading machine, the roadway size and the like. Through a reinforcement learning algorithm, the heading machine can autonomously decide a walking path, the dependence on manual intervention is reduced, and the autonomy of a walking track is improved; the walking track is dynamically adjusted according to the position and posture of the machine body and the size information of the roadway collected in real time, and unnecessary adjustment time and energy consumption are reduced; through an intelligent decision and obstacle avoidance strategy, collision between the heading machine and a roadway wall is avoided, and operation safety is ensured; by adjusting model parameters or retraining, the heading machine can quickly adapt to roadway environments with different sizes and geological conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous walking trajectory planning for roadheaders, and in particular relates to a method and device for autonomous walking trajectory planning for roadheaders based on a reinforcement learning algorithm. Background Art

[0002] As a vital piece of equipment for coal mine tunneling, roadheaders (TBMs) operate in a complex and dynamic environment, characterized by narrow tunnel spaces, variable slopes, and irregular tunneling directions. Traditional manual control relies on the driver's experience and judgment. However, in practice, factors such as visual blind spots and misjudgment often lead to collisions between the TBM and tunnel walls, or deviations from the intended trajectory, impacting tunneling efficiency and safety. Summary of the Invention

[0003] The present invention aims to improve the operating capability of a roadheader in a complex tunnel environment through intelligent decision-making and autonomous learning.

[0004] The first object of the present invention is to provide the following technical solution: a method for autonomous walking trajectory planning of a roadheader based on a reinforcement learning algorithm, comprising:

[0005] Collect the autonomous travel trajectory planning data of the roadheader and build the autonomous travel trajectory planning model of the roadheader;

[0006] Based on the autonomous trajectory planning model of the roadheader, a reward function for autonomous trajectory planning of the roadheader is designed;

[0007] According to the reward function, the reinforcement learning algorithm is applied to adjust the walking trajectory of the roadheader and realize autonomous walking trajectory planning of the roadheader.

[0008] Furthermore, the autonomous walking trajectory planning data of the roadheader is collected, including:

[0009] Through sensors and measuring equipment, the machine's body posture data and tunnel dimension data are collected in real time as data for the machine's autonomous walking trajectory planning.

[0010] The fuselage posture data includes fuselage position coordinate data, pitch angle data, yaw angle data and roll angle data; the tunnel size data includes tunnel width data W and tunnel height data H.

[0011] Furthermore, the autonomous walking trajectory planning model of the tunnel boring machine describes the environment through a four-tuple (S, A, R, T);

[0012] The state space S represents the set of all states of the roadheader; the autonomous walking trajectory planning data of the roadheader is used as the state variable in the state space;

[0013] The action space A represents the set of all actions taken by the roadheader in each state. Action types include forward, backward, left turn, right turn, and stop. By parameterizing each action, the action is converted into a combination of velocity and angular velocity.

[0014] The reward function R represents the immediate reward obtained when the first action is performed in the first state and the transition is made to the second state;

[0015] The state transition function T represents the probability of executing the first action in the first state and transitioning to the second state.

[0016] Furthermore, the types of reward functions designed include:

[0017] Safety reward: When the distance between the tunnel boring machine and the tunnel wall exceeds the safety threshold, a positive reward is given; otherwise, a negative reward is given to encourage the tunnel boring machine to maintain a safe distance;

[0018] Efficiency rewards: Based on the traveling speed and path length of the TBM, certain efficiency rewards are given to encourage the TBM to complete tasks quickly and efficiently.

[0019] Smoothness bonus: Penalizes the acceleration and angular acceleration of the roadheader to reduce bumps and vibrations during travel and improve work quality.

[0020] Furthermore, the autonomous walking trajectory planning model of the tunnel boring machine also includes a strategy and a value function; wherein,

[0021] A strategy is a mapping from state to action; the choice of strategy directly affects the trajectory and operating efficiency of the tunnel boring machine.

[0022] The value function is used to evaluate the expected long-term cumulative reward obtained by following the first strategy starting from the first state;

[0023] The action-value function is used to evaluate the expected long-term cumulative reward obtained by starting from the first state, performing the first action according to the first policy, and acting according to the first policy.

[0024] Furthermore, the goal of reinforcement learning is to find an optimal policy π such that for all states s, there is a maximum value function V(s) or action value function Q*(s,a); this is achieved by maximizing the expected cumulative reward, which is expressed as the sum of discounted rewards, as follows:

[0025]

[0026] Where γ is a discount factor (0≤γ<1) that is used to balance the importance of immediate rewards and future rewards;

[0027] At this point, the reinforcement learning problem of autonomous walking trajectory planning for the tunnel boring machine is simplified to finding the optimal strategy, and the optimization strategy maximizes the cumulative discounted reward from the initial state to the terminal state.

[0028] Furthermore, based on the reward function, a reinforcement learning algorithm is applied to adjust the trajectory of the roadheader, including:

[0029] Initialization, including initialization of the strategy and value function; at the same time, setting up an experience replay pool to store the new four-tuple of historical state-action-reward-new state;

[0030] Execute in a loop, observe the current state, select actions according to the strategy, execute the actions selected according to the strategy, and observe the new state and reward;

[0031] Store the new quadruple into the experience replay pool; randomly extract samples from the experience replay pool to train the value function;

[0032] Update the weights of the value function and optimize the loss function using gradient descent;

[0033] The iteration continues until the maximum number of iterations is reached, and the output result is a state-action pair, which serves as the optimal walking trajectory of the tunnel boring machine.

[0034] The second object of the present invention is to provide a device for autonomous walking trajectory planning of a roadheader based on a reinforcement learning algorithm, comprising:

[0035] The model building module is used to collect the autonomous walking trajectory planning data of the roadheader and build the autonomous walking trajectory planning model of the roadheader;

[0036] The model optimization module is used to design the reward function for the autonomous trajectory planning of the roadheader based on the autonomous trajectory planning model of the roadheader;

[0037] The trajectory planning module is used to adjust the trajectory of the roadheader based on the reward function and apply the reinforcement learning algorithm to realize autonomous trajectory planning of the roadheader.

[0038] The third object of the present invention is to provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute each step in the method of the aforementioned technical solution.

[0039] A fourth object of the present invention is to provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute each step in the method according to the aforementioned technical solution.

[0040] Compared with the prior art, the advantages of the present invention are:

[0041] The present invention's autonomous trajectory planning method for a roadheader (TBM) based on a reinforcement learning algorithm comprehensively considers the TBM's real-time body posture (including position, pitch, yaw, and roll angles) and key parameters such as tunnel dimensions to achieve optimal trajectory planning from the TBM's current position to the tunnel's head. The reinforcement learning algorithm enables the TBM to autonomously determine its path, reducing reliance on human intervention and improving its trajectory autonomy. The TBM's trajectory is dynamically adjusted based on real-time information collected about the body's posture and tunnel dimensions, reducing unnecessary adjustment time and energy consumption. Intelligent decision-making and obstacle avoidance strategies prevent collisions between the TBM and tunnel walls, ensuring operational safety. By adjusting model parameters or retraining, the TBM can quickly adapt to tunnel environments of varying sizes and geological conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A schematic flow chart of a method for autonomous walking trajectory planning of a roadheader based on a reinforcement learning algorithm provided by the present invention;

[0043] Figure 2 A schematic structural diagram of a device for autonomous walking trajectory planning of a roadheader based on a reinforcement learning algorithm provided by the present invention;

[0044] Figure 3 A schematic structural diagram of a non-transitory computer-readable storage medium storing computer instructions provided by the present invention. DETAILED DESCRIPTION

[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0046] like Figure 1 Shown: A method for autonomous walking trajectory planning of a roadheader based on a reinforcement learning algorithm, comprising:

[0047] S110: Collecting autonomous travel trajectory planning data of the roadheader and building an autonomous travel trajectory planning model for the roadheader.

[0048] Data collection is the basis for autonomous trajectory planning of the roadheader. The autonomous trajectory planning data of the roadheader includes:

[0049] Through sensors and measuring equipment, the machine's body posture data (x, y, z, θ, φ, ψ) and tunnel dimension data (width W, height H) are collected in real time as data for the machine's autonomous walking trajectory planning.

[0050] The fuselage posture data includes the fuselage position spatial coordinate data, pitch angle data, yaw angle data and roll angle data; the tunnel size data includes tunnel width data W and tunnel height data H.

[0051] When constructing the autonomous walking trajectory planning model for a roadheader, the environment is described by a four-tuple (S, A, R, T);

[0052] The state space S represents the set of all states of the roadheader; the autonomous walking trajectory planning data of the roadheader is used as the state variable in the state space; the fuselage yaw angle, fuselage posture information, tunnel dimensions, etc. are used as state variables in the state space. t is the description of the environment of the tunnel boring machine at time t, s t ={x t, y t ,z t ,θ t ,φ t ,ψ t ,W t ,H t}.

[0053] The action space A represents the set of all actions taken by the tunnel boring machine in each state; action a t is the operation that the tunnel boring machine can perform at time t, including five actions: forward, backward, left turn, right turn and stop. These actions can be parameterized as a combination of velocity v and angular velocity ω. t ={v t ,ω t}.

[0054] The reward function R represents the immediate reward obtained when executing the first action in the first state and transitioning to the second state. The design of the reward function should be able to reflect the quality of the roadheader's walking trajectory, such as reducing collisions with the tunnel wall and optimizing the walking path, r = R(s, a, s′).

[0055] The state transition function T represents the probability of performing the first action in the first state and transitioning to the second state. Since the movement of the tunnel boring machine is continuous, the state transition function can be approximated by a physical model or simulation environment, T(s,a,s′)=P(s′|s,a).

[0056] The autonomous walking trajectory planning model of the tunnel boring machine also includes strategy and value function; among them,

[0057] Strategy π: A mapping from state to action, i.e., π(s) = a, representing the action to be taken in state s. The choice of strategy directly affects the trajectory and operating efficiency of the roadheader.

[0058] Value function V(s): This function evaluates the expected long-term cumulative reward obtained by following policy π starting from state s. The value function is used to assess the potential value of the tunnel boring machine in different states.

[0059] Action-value function Q(s,a): Similarly, it evaluates the expected long-term cumulative reward obtained by taking action a in state s and acting according to policy π. The action-value function is used to guide the tunnel boring machine to choose the optimal action in a specific state.

[0060] S120: Based on the autonomous walking trajectory planning model of the roadheader, design a reward function for the autonomous walking trajectory planning of the roadheader.

[0061] The design of the reward function is a key step in reinforcement learning, which directly determines the optimization direction of the algorithm. For the problem of autonomous walking trajectory planning of a roadheader, the reward function can be designed as:

[0062] Safety reward: When the distance between the tunnel boring machine and the tunnel wall exceeds the safety threshold, a positive reward is given; otherwise, a negative reward is given to encourage the tunnel boring machine to maintain a safe distance.

[0063] Efficiency rewards: Based on the traveling speed and path length of the tunnel boring machine, certain efficiency rewards are given to encourage the tunnel boring machine to complete the task quickly and efficiently.

[0064] Smoothness bonus: Penalizes the acceleration and angular acceleration of the roadheader to reduce bumps and vibrations during travel and improve work quality.

[0065] S130: According to the reward function, a reinforcement learning algorithm is applied to adjust the walking trajectory of the roadheader to achieve autonomous walking trajectory planning of the roadheader.

[0066] The goal of reinforcement learning is to find an optimal policy π such that for all states s, there is a maximum value function V(s) or action value function Q*(s,a); this is achieved by maximizing the expected cumulative reward, which is expressed as the sum of discounted rewards, as follows:

[0067]

[0068] Where γ is a discount factor (0≤γ<1) that is used to balance the importance of immediate rewards and future rewards;

[0069] At this point, the reinforcement learning problem of autonomous walking trajectory planning for the tunnel boring machine is simplified to finding the optimal strategy, and the optimization strategy maximizes the cumulative discounted reward from the initial state to the terminal state.

[0070] The reinforcement learning problem of autonomous trajectory planning for a roadheader can be simplified to finding an optimal policy π* that maximizes the cumulative discounted reward from the initial state to the final state (or in an infinite timeframe). This is typically achieved by solving the following optimization problem:

[0071] optimization

[0072] where s t+1 ~T(s t ,π(s t )); T represents the total number of time steps, but in many practical problems, one may consider an infinite time horizon (i.e., T→∞) and use a discount factor to ensure convergence of the reward.

[0073] Based on the reward function, a reinforcement learning algorithm is applied to adjust the trajectory of the roadheader, including:

[0074] Initialization, including the initialization of the strategy and value function; at the same time, set up the experience replay pool to store the new four-tuple of historical state-action-reward-new state; specifically,

[0075] Initialize the agent’s strategy π(a t ∣s t )(such as random strategy);

[0076] Initialize the value function Q(s t ,a t );

[0077] Set up an experience replay pool D to store the historical state-action-reward-new state quadruple.

[0078] Execute in a loop, observe the current state, select actions according to the strategy, execute the actions selected according to the strategy, and observe the new state and reward; specifically,

[0079] Observe the current state s t ;

[0080] According to the strategy π(a t |s t )(or ε-greedy strategy) select action a t ;

[0081] Execute action a t , observe the new state s t+1 and reward r t+1 ;

[0082] Store the new quadruple into the experience replay pool; randomly extract samples from the experience replay pool to train the value function;

[0083] Update the weights of the value function and optimize the loss function using gradient descent;

[0084] The quadruple (s t ,a t ,r t+1 ,s t+1 ) is stored in the experience replay pool D;

[0085] Randomly extract a batch of samples from D to train the value function Q(s t ,a t );

[0086] Update the value function Q(s t ,a t ), the gradient descent method is usually used to optimize the loss function;

[0087] Update Status t ←s t+1 ;

[0088] Iterate until the maximum number of iterations, and the output result is a state-action pair (s t ,a t ), as the optimal trajectory of the tunnel boring machine. These state-action pairs represent the optimal actions to be taken under different environmental conditions.

[0089] Through the above steps, the present invention constructs an autonomous walking trajectory planning model for a tunnel boring machine based on reinforcement learning, which can automatically adjust the walking trajectory of the tunnel boring machine according to real-time collected information such as the machine body posture and tunnel dimensions, thereby ensuring the safety and efficiency of the tunnel boring operation.

[0090] In the method of the present invention, through autonomous learning and decision-making, the tunnel boring machine can cope with complex tunnel environments more flexibly. This method can intelligently plan the optimal walking path without human intervention by autonomously sensing the body posture and tunnel environment, and flexibly cope with complex and changing work requirements. Its high efficiency is reflected in path optimization and speed improvement, which significantly reduces unnecessary actions and ensures both work efficiency and quality. At the same time, the system has a high degree of safety, monitors and avoids collisions in real time, and protects the safety of equipment and tunnels by optimizing walking smoothness. Its adaptability allows the system to easily cope with tunnels of different shapes, sizes and complexities. The system can also adjust according to the different initial positions and postures of the tunnel boring machine to ensure autonomous walking under any circumstances.

[0091] like Figure 2 As shown, the present invention provides a device 200 for planning an autonomous walking trajectory of a roadheader based on a reinforcement learning algorithm, comprising:

[0092] The model building module 210 is used to collect the autonomous walking trajectory planning data of the roadheader and build an autonomous walking trajectory planning model for the roadheader;

[0093] A model optimization module 220 is used to design a reward function for autonomous travel trajectory planning of the roadheader based on the autonomous travel trajectory planning model of the roadheader;

[0094] The trajectory planning module 230 is used to apply the reinforcement learning algorithm according to the reward function to adjust the walking trajectory of the roadheader and realize autonomous walking trajectory planning of the roadheader.

[0095] In order to implement the embodiment, the present invention also proposes an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute each step in the method of the aforementioned technical solution.

[0096] like Figure 3 As shown, the non-transitory computer-readable storage medium 900 includes a memory 910 of instructions and an interface 930, and the instructions can be executed by a processor 920 to complete the method. Alternatively, the storage medium can be a non-transitory computer-readable storage medium, for example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0097] In order to implement the embodiments, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method according to the embodiments of the present invention is implemented.

[0098] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0099] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0100] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0101] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.

[0102] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the embodiments described, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0103] Those skilled in the art will understand that all or part of the steps of the method for implementing the embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0104] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0105] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present invention have been shown and described above, it is understood that the embodiments are exemplary and are not to be construed as limiting the present invention. Those skilled in the art may make changes, modifications, substitutions, and variations to the embodiments within the scope of the present invention.

Claims

1. A method for autonomous walking trajectory planning of a roadheader based on a reinforcement learning algorithm, characterized in that: include: Collect the autonomous travel trajectory planning data of the roadheader and build the autonomous travel trajectory planning model of the roadheader; Based on the autonomous walking trajectory planning model of the roadheader, a reward function for autonomous walking trajectory planning of the roadheader is designed; According to the reward function, a reinforcement learning algorithm is applied to adjust the walking trajectory of the roadheader to achieve autonomous walking trajectory planning of the roadheader.

2. The method for autonomous walking trajectory planning of a roadheader based on a reinforcement learning algorithm according to claim 1, characterized in that: Collect autonomous travel trajectory planning data for the roadheader, including: The body posture data and tunnel dimension data of the tunnel boring machine are collected in real time through sensors and measuring equipment as the autonomous walking trajectory planning data of the tunnel boring machine; wherein, The fuselage posture data includes fuselage position coordinate data, pitch angle data, yaw angle data and roll angle data; the lane size data includes lane width data W and lane height data H.

3. The method for autonomous walking trajectory planning of a roadheader based on a reinforcement learning algorithm according to claim 2, characterized in that: The autonomous walking trajectory planning model of the roadheader describes the environment through a four-tuple (S, A, R, T); wherein, The state space S represents the set of all states of the roadheader; the autonomous walking trajectory planning data of the roadheader is used as the state variable in the state space; The action space A represents the set of all actions taken by the roadheader in each state. Action types include forward, backward, left turn, right turn, and stop. By parameterizing each action, the action is converted into a combination of velocity and angular velocity. The reward function R represents the immediate reward obtained when the first action is performed in the first state and the transition is made to the second state; The state transition function T represents the probability of executing the first action in the first state and transitioning to the second state.

4. The method for autonomous walking trajectory planning of a roadheader based on a reinforcement learning algorithm according to claim 3 is characterized in that: Types of reward functions include: Safety reward: When the distance between the tunnel boring machine and the tunnel wall exceeds the safety threshold, a positive reward is given; otherwise, a negative reward is given to encourage the tunnel boring machine to maintain a safe distance; Efficiency rewards: Based on the traveling speed and path length of the TBM, certain efficiency rewards are given to encourage the TBM to complete tasks quickly and efficiently. Smoothness bonus: Penalizes the acceleration and angular acceleration of the roadheader to reduce bumps and vibrations during travel and improve work quality.

5. The method for autonomous walking trajectory planning of a roadheader based on a reinforcement learning algorithm according to claim 3 is characterized in that: The autonomous walking trajectory planning model of the roadheader also includes a strategy and a value function; wherein, The strategy is a mapping from state to action; the choice of the strategy directly affects the travel trajectory and operating efficiency of the roadheader; The value function is used to evaluate the expected long-term cumulative reward obtained by following the first strategy starting from the first state; The action-value function is used to evaluate the expected long-term cumulative reward obtained by starting from the first state, performing the first action according to the first policy, and acting according to the first policy.

6. The method for autonomous walking trajectory planning of a roadheader based on a reinforcement learning algorithm according to claim 5, characterized in that: The goal of reinforcement learning is to find an optimal policy π such that for all states s, there is a maximum value function V(s) or action value function Q*(s,a); this is achieved by maximizing the expected cumulative reward, which is expressed as the sum of discounted rewards, as follows: Where γ is a discount factor (0≤γ<1) that is used to balance the importance of immediate rewards and future rewards; At this point, the reinforcement learning problem of autonomous walking trajectory planning for the tunnel boring machine is simplified to finding the optimal strategy, and the optimization strategy maximizes the cumulative discounted reward from the initial state to the terminal state.

7. The method for autonomous walking trajectory planning of a roadheader based on a reinforcement learning algorithm according to claim 1, characterized in that: According to the reward function, a reinforcement learning algorithm is applied to adjust the walking trajectory of the roadheader, including: Initialization, including initialization of the strategy and value function; at the same time, setting up an experience replay pool to store the new four-tuple of historical state-action-reward-new state; Execute in a loop, observe the current state, select actions according to the strategy, execute the actions selected according to the strategy, and observe the new state and reward; Storing the new quadruple in the experience replay pool; and randomly extracting samples from the experience replay pool to train the value function; Update the weights of the value function and optimize the loss function using gradient descent; The iteration continues until the maximum number of iterations is reached, and the output result is a state-action pair, which serves as the optimal walking trajectory of the tunnel boring machine.

8. A roadheader autonomous walking trajectory planning device based on reinforcement learning algorithm, characterized in that: include: The model building module is used to collect the autonomous walking trajectory planning data of the roadheader and build the autonomous walking trajectory planning model of the roadheader; A model optimization module is used to design a reward function for autonomous walking trajectory planning of the roadheader based on the autonomous walking trajectory planning model of the roadheader; The trajectory planning module is used to apply a reinforcement learning algorithm according to the reward function to adjust the walking trajectory of the roadheader and realize autonomous walking trajectory planning of the roadheader.

9. An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform each step in the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute each step of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Real-time deviation rectifying method, device, equipment, medium and product for end slope mining cave track

    CN121596754A

  • VLA model autonomous generalization method, system, device and medium

    CN121638318A