Automatic loading control method, training method, device, equipment and medium
By applying reinforcement learning technology on the shovel machine and training the intelligent body to optimize shovel control, the problem of difficulty in adapting to the dynamic changes of the ore pile to be shoveled is solved, and efficient and excellent quality shoveling effect is achieved.
Patent Information
- Application Number
- CN202510307103.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-16
AI Technical Summary
The existing automatic shoveling control method of shoveling machines is difficult to adapt to the dynamic changes of the ore pile to be shoveled, resulting in low shoveling efficiency and quality.
The automatic shovel installation control method based on reinforcement learning is adopted, and the dynamic changes in the shovel installation process are optimized by training the intelligent body to use the state data, control instructions and reward values of the shovel installation.
It realizes high-quality and efficient shoveling of the dynamically changing ore piles to be shoveled, which improves the success rate and loading capacity of shoveling.
Smart Images

Figure CN119830019B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of mining, and in particular, to an automatic loading control method, a training method, a device, equipment and a medium. Background Art
[0002] A scraper is a type of multi-functional construction machinery that integrates loading, transportation and unloading functions, and plays an important role in the ore mining and tunneling operations of underground mines. As the mining activities continue to extend deeper underground, the operating conditions of the scraper become more severe, facing challenges such as high temperature, dust and noise. These adverse factors not only pose threats to the performance and lifespan of the equipment, but also seriously endanger the physical and mental health of the operating workers. In view of this, the automatic control technology for scraper operation has become a hot research field. In this field, the automatic loading control technology of the scraper is particularly crucial.
[0003] In the existing automatic loading control methods for scrapers, the force feedback-based method relies on on-site parameter adjustment and is difficult to adapt to the dynamic changes of the ore pile to be loaded. Summary of the Invention
[0004] The present application provides an automatic loading control method, a training method, a device, equipment and a medium, which can solve one of the problems existing in the background art.
[0005] To achieve the above object, the present application adopts the following technical solutions:
[0006] In a first aspect, a training method for an automatic loading intelligent body is provided. The intelligent body acts on a scraper, and the method includes:
[0007] Initializing the intelligent body; and
[0008] Training the intelligent body by using the collected data,
[0009] The collected data includes: the state data of the scraper, the control instructions of the intelligent body for the scraper, and the reward value, where
[0010] The state data includes: the telescopic amount of the oil cylinder, the steering articulation angle, the coordinate position of the scraper, the acceleration and angular velocity along the front vehicle body,
[0011] The control instructions include: the telescopic speed of the oil cylinder, the angular velocity of the steering hinge, and the angular velocity of the wheel.
[0012] Based on the above technical solution, the intelligent agent is trained using the collected data composed of the state data, control commands, and reward values of the scraper to obtain an intelligent agent for realizing automatic loading. In particular, the state data is selected as the telescopic amount of the oil cylinder, the steering articulation angle, the coordinate position of the scraper, the acceleration and angular velocity along the front vehicle body, and the control commands are selected as the telescopic speed of the oil cylinder, the angular velocity of the steering hinge, and the angular velocity of the wheels. It is found in the simulation experiment and practical application that for the intelligent agent obtained by training, the output control commands can control the scraper and effectively perform high-quality and efficient loading on the ore heap to be shoveled with dynamic changes during the loading process.
[0013] In a possible design of the first aspect, the reward function r in the intelligent agent is defined as: where, is the coordinate of the bucket in the advancing direction of the scraper towards the ore heap; a is the row vector composed of the current action control amounts corresponding to the control commands; F is the amount of ore in the bucket; is the elongation of the tilting cylinder; both k and f are coefficients; f1 is the amount of ore when the bucket is full, and f1 is a decimal less than f2 and greater than 0.
[0014] In a possible design of the first aspect, the intelligent agent includes: a policy network and a value network. The input of the policy network is the current state, and the output is the probability of performing an action in the current state; the input of the value network is the current state, and the output is the value of the current state.
[0015] In a possible design of the first aspect, the objective function of the intelligent agent includes: the policy loss and value loss defined by the reward function, and the entropy regularization term.
[0016] In a possible design of the first aspect, the objective function is optimized by mini-batch gradient descent to update the parameters of the policy network and the value network.
[0017] In the second aspect, an automatic loading control method is provided. The control method is based on the intelligent agent obtained by training according to any one of the first aspect, and the intelligent agent acts on the scraper to complete the automatic loading control.
[0018] In the third aspect, a training device for an automatic loading intelligent agent is provided. The intelligent agent acts on the scraper, and the training device includes:
[0019] An initialization unit for initializing the intelligent agent; and
[0020] A training unit for training the intelligent agent using the collected data,
[0021] The collected data includes: the status data of the scraper, the control instructions of the intelligent agent for the scraper, and the reward value, where
[0022] The status data includes: the telescopic amount of the oil cylinder, the steering articulation angle, the coordinate position of the scraper, the acceleration and angular velocity along the front vehicle body.
[0023] The control instructions include: the telescopic speed of the oil cylinder, the angular velocity of the steering hinge, and the angular velocity of the wheels.
[0024] Fourthly, a scraper is provided, and the scraper is provided with the intelligent agent obtained by training according to any one of the first aspects, and the intelligent agent acts on the scraper to complete automatic loading control.
[0025] Fifthly, an electronic device is provided, and the electronic device includes: a processor, and a memory coupled to the processor, where the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the training method according to any one of the possible implementation manners of the first aspect, or executes the automatic loading control method according to the second aspect.
[0026] Sixthly, a computer-readable storage medium is provided, including a computer program or instruction, and when the computer program or instruction runs on a computer, the computer is enabled to execute the training method according to any one of the possible implementation manners of the first aspect, or execute the automatic loading control method according to the second aspect.
[0027] Seventhly, a computer program product is provided, including: a computer program or instruction, and when the computer program or instruction runs on a computer, the computer is enabled to execute the training method according to any one of the possible implementation manners of the first aspect, or execute the control method according to the second aspect. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0029] Figure 1 It is the implementation flowchart of the automatic loading method of the scraper based on reinforcement learning provided by the embodiment of the present application;
[0030] Figure 2 It is the structural schematic diagram of the scraper provided by the embodiment of the present application;
[0031] Figure 3 It is a comparison chart of the ore loading amount provided by the embodiment of the present application;
[0032] Figure 4 It is a comparison chart of the loading time provided by the embodiment of the present application. Detailed implementation manners
[0033] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application.
[0034] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division in the device or a different order in the flowchart. The terms "first", "second", etc. in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.
[0036] As Figure 1 shown, the embodiment of the present application proposes a method for automatic loading of a scraper based on reinforcement learning, and this method is carried out based on the system structure as Figure 2 shown.
[0037] First, in the controller module, training is carried out according to the following steps to obtain a reinforcement learning agent.
[0038] Step 1: Initialize two neural networks, which are respectively used to represent the policy and the value. Among them, the input of the policy network is the current state s, and the output is the probability of executing the action a in the current state s, denoted as ; the input of the value network is the current state s, and the output is the value of the current state s, denoted as . Among them, θ and are respectively the parameters of the two neural networks.
[0039] Among them, the state s is composed of the observation data obtained by the sensor and is defined as: In the formula, W b , W c are respectively the telescopic amount of the boom cylinder and the telescopic amount of the bucket cylinder of the scraper; j z is the steering articulation angle of the scraper; is the wheel speed; [dx , d y , d z are the position coordinates of the scraper in the local coordinate system with the center of the ore pile to be shoveled as the origin; [[a x , a y , a z , [g x , g y , g z are the linear acceleration and angular acceleration of the front body of the scraper along the x, y, and z coordinate axes respectively.
[0040] Action a is the control instruction output by the reinforcement learning agent and is defined as: . Where v b , v c are the telescopic speeds of the boom cylinder and the bucket cylinder respectively, with elongation being positive; is the angular velocity of the steering hinge of the scraper; is the angular velocity of the wheel.
[0041] Step 2: Collect data. Interact in the environment based on the current policy and collect and save the data generated during the interaction: In the formula, s t is the current state, a t is the current action, r t is the current reward, s t+1 is the next state, and done t indicates whether the current sequence has ended, 1 for ended and 0 for not ended. The criterion for determining the end of the interaction episode is: the bucket reaches the set rotation limit position or the length of the interaction episode reaches the limit value. Continuously interact with the environment and save the data until the set maximum buffer length B m .
[0042] Among them, the reward r t is defined as: In the formula, is the coordinate of the bucket in the forward direction of the scraper towards the ore pile; a is the row vector composed of the current action control amount; F is the amount of ore in the bucket; is the elongation of the dump cylinder; k and f are both coefficients. Among them, f1 is the amount of ore when the bucket is full, and f1 is a decimal less than f2 and greater than 0.
[0043] Step 3: Calculate the objective function. The objective function used in this method consists of three parts: policy loss, value loss, and entropy regularization term: . In the formula, c1 and c2 are hyperparameters used to control the weights of the value loss and entropy regularization.
[0044] Among them, the policy loss: , in the formula, Among them, is the discount factor, is a parameter that controls the smoothness of the advantage, represents the old policy, represents the new policy, is the clipping threshold parameter. clip() is the clipping function, and min() is the function to take the minimum value. Their mathematical definitions are as follows: where a and b are preset values.
[0045] Value loss: , where, .
[0046] Entropy regularization term: , where, is the policy at state s t and β is the regularization coefficient.
[0047] Step 4: Optimize the objective function. Optimize the objective function through mini-batch gradient descent to update the parameters of the policy and value networks: where α is the learning rate in the formula.
[0048] Step 5: Iterative optimization. Repeat Step 2 to Step 4 until the model converges or reaches the set maximum number of training steps.
[0049] After the training is completed, reload the agent in the control module. Encode the data collected by the sensor according to and input it into the policy network of the agent. Then transfer the action output by the policy network to the actuator module, and the scraper performs the corresponding action instructions until the loading action is completed.
[0050] Next, through a specific application example, the above embodiments will be exemplarily described.
[0051] The purpose of this example is to provide a method for automatic loading of a scraper based on deep reinforcement learning, which can adapt to dynamically changing ore piles; and provides a modeling method for the action space, state space, and reward function of reinforcement learning for automatic loading of a scraper, greatly improving the convergence success rate and application effect of training.
[0052] To achieve the above purpose, the specific implementation is as follows:
[0053] Step 1: Construct an agent. Initialize two neural networks, which are respectively used to represent the policy and value. Among them, the input of the policy network is the current state s, and the output is the probability of executing action a in the current state s, denoted as ; the input of the value network is the current state s, and the output is the value of the current state s, denoted as Among them, θ and are the parameters of two neural networks respectively. Both neural networks have two hidden layers with 64 nodes, and the activation function is tanh.
[0054] Among them, the state s is composed of the observed data obtained by the sensor and is defined as: In the formula, W b , W c are the telescopic amounts of the boom cylinder and the bucket cylinder of the scraper respectively; j z is the steering articulation angle of the scraper; is the wheel speed; [d x , d y , d z is the position coordinate of the scraper in the local coordinate system with the center of the ore pile to be shoveled as the origin; [a x , a y , a z , [g x , g y , g z are the linear accelerations and angular accelerations of the front body of the scraper along the x, y, and z coordinate axes respectively.
[0055] The action a is the control instruction output by the reinforcement learning agent and is defined as: . Among them, v b , v c are the telescopic speeds of the boom cylinder and the bucket cylinder respectively, with elongation being positive; is the angular velocity of the steering hinge of the scraper; is the angular velocity of the wheel.
[0056] Step 2: Agent training.
[0057] Step 2.1: Data collection. Based on the current policy interact in the environment, collect and save the data generated during the interaction: In the formula, s t is the current state, a t is the current action, r t is the current reward, s t+1 is the next state, done t is whether the current sequence ends, 1 means end, 0 means not end. The criterion for determining the end of the interaction episode is that the bucket reaches the set rotation limit position or the length of the interaction episode reaches the limit value. Continuously interact with the environment and save the data until the set maximum buffer length of 2048 is reached.
[0058] Among them, the reward r t is defined as: In the formula, is the coordinate of the bucket in the direction of the loader moving toward the ore pile; a is the row vector composed of the current action control quantity; F is the amount of ore in the bucket; is the extension of the bucket cylinder.
[0059] Step 2.2: Calculate the objective function. The objective function used in this method consists of three parts: policy loss, value loss, and entropy regularization term: .
[0060] The strategy loss is: In the formula, in, represents the old policy (the old policy given state s t When selecting action a t probability), represents a new strategy (the current strategy given state s t When selecting action a t The probability of ). clip() is the clipping function, min() is the minimum value, and their mathematical definitions are as follows: Where a and b are preset values.
[0061] Value loss: ,in, .
[0062] Entropy regularization term: ,in, It’s a strategy In status t The entropy under .
[0063] Step 2.3: Optimize the objective function. Optimize the objective function through mini-batch gradient descent and update the parameters of the policy and value networks as shown in the following formula: Step 2.4: Iterative optimization. Repeat steps 2.1 to 2.3 until the model reaches the set maximum number of training steps of 20,000,000.
[0064] Step 3: Get the status information of the scraper and convert the status data into Encode and pass it into the agent's policy network.
[0065] Step 5: The agent outputs action instructions based on state input.
[0066] Step 6: The scraper performs corresponding actions according to the action instructions.
[0067] Step 7: Repeat steps 3 to 6 until the shovel installation is complete.
[0068] In order to show the superiority of the method proposed in this aspect, Figure 3 As shown in the literature, this method is compared with [1] To the literature[4] Forty groups of loading experiments were conducted to compare with the existing methods and human drivers. Compared with the existing methods, the proposed method shows the best performance in terms of loading success rate and loading capacity, approaching that of human drivers; and it outperforms human drivers in terms of loading time and loading consistency.
[0069] Literature [1] to Literature [4] List:
[0070] Literature [1] Eriksson D, Ghabcheloo R, Geimer M. Automatic Loading ofUnknown Material with a Wheel Loader Using Reinforcement Learning[C] / / 2024IEEE International Conference on Robotics and Automation (ICRA). 2024: 3646-3652.
[0071] Literature [2] Eriksson D, Ghabcheloo R, Geimer M. Optimizing a BucketFilling Strategy for Wheel Loaders Inside a Dream Environment[A]. 2024.
[0072] Literature [3] Shen C, Sloth C. Generalized Framework for Wheel LoaderAutomatic Shoveling Task with Expert Initialized Reinforcement Learning[C] / / 2024 IEEE / SICE International Symposium on System Integration (SII). 2024:382-389.
[0073] Literature [4]Dadhich S, Sandin F, Bodin U, et al. Adaptation of a wheelloader automatic bucket filling neural network using reinforcement learning[C] / / 2020 International Joint Conference on Neural Networks (ijcnn). NewYork: Ieee, 2020.
[0074] The above-mentioned method for automatic loading of scrapers based on deep reinforcement learning can also generate a model of the ore-rock pile to be shoveled underground in a physical simulator for simulation training and simulation experiments. By generating the model of the ore-rock pile to be shoveled, a batch of ore-rock block models with different shapes and a specified distribution of block sizes can be automatically generated and form an ore-rock pile. This model can be used for research in fields such as automatic loading control and the law of bulk movement.
[0075] The embodiment of the present application also provides a training device for an automatic loading intelligent body, where the intelligent body acts on a scraper, and the training device includes:
[0076] An initialization unit for initializing the intelligent body; and
[0077] A training unit for training the intelligent body using the collected data,
[0078] The collected data includes: the state data of the scraper, the control instructions of the intelligent body for the scraper, and the reward value, where
[0079] The state data includes: the telescopic amount of the oil cylinder, the steering hinge angle, the coordinate position of the scraper, the acceleration and angular velocity along the front vehicle body,
[0080] The control instructions include: the telescopic speed of the oil cylinder, the angular velocity of the steering hinge, and the angular velocity of the wheels.
[0081] The embodiment of the present application also provides a scraper, which is provided with the intelligent body trained as described in any one of the first aspects, and the intelligent body acts on the scraper to complete automatic loading control.
[0082] The embodiment of the present application also provides an electronic device, including: a processor, and a memory coupled to the processor, where the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the method described in any one of the above embodiments.
[0083] The electronic device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device may include, but is not limited to, a processor and a memory.
[0084] The so-called processor can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The processor is the control center of the electronic device, connecting various parts of the entire device through various interfaces and circuits.
[0085] The memory can be used to store the computer program. The processor realizes various functions of the electronic device by running or executing the computer program stored in the memory and calling the data stored in the memory.
[0086] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0087] The embodiments of the present application also provide a storage medium, which is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0088] The embodiments of the present application also provide a computer program product, including: a computer program or instruction. When the computer program or instruction runs on a computer, the computer is enabled to execute the method of any one of the above possible implementation manners.
[0089] The above is the preferred implementation manner of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present application.
Claims
1. A training method for an automatic shovel loading agent, characterized in that: The intelligent agent acts on a scraper, and the method comprises: Initializing the agent; and Using the collected data, the intelligent agent is trained, The collected data includes: the state data of the scraper, the control instructions of the intelligent agent to the scraper, and the reward value, wherein: The state data includes: cylinder extension amount, steering articulation angle, scraper coordinate position, acceleration and angular velocity of the front vehicle body, The control instructions include: cylinder extension speed, steering hinge angular velocity and wheel angular velocity, The reward function r in the agent is defined as: in, is the coordinate of the bucket in the direction of the loader moving toward the ore pile; a is the row vector composed of the current action control quantity corresponding to the control instruction; F is the amount of ore in the bucket; is the extension of the bucket cylinder; k and f are coefficients; f1 is the amount of ore when the bucket is full, and f1 is a decimal less than f2 and greater than 0.
2. The training method according to claim 1, characterized in that: The intelligent agent includes: a strategy network and a value network. The input of the strategy network is the current state, and the output is the probability of executing an action in the current state; the input of the value network is the current state, and the output is the value of the current state.
3. The training method according to claim 2, characterized in that: The objective function of the agent includes: strategy loss and value loss defined by a reward function, and an entropy regularization term.
4. The training method according to claim 3, characterized in that: The objective function is optimized by mini-batch gradient descent to update the parameters of the policy network and the value network.
5. An automatic shovel loading control method, characterized in that: The control method is based on the intelligent agent trained by the training method as described in any one of claims 1 to 4, and the intelligent agent acts on the shovel loader to complete automatic shoveling control.
6. A training device for an automatic shovel loading agent, characterized in that: The intelligent agent acts on a scraper, and the training device comprises: an initialization unit, used to initialize the agent; and A training unit, used to train the intelligent agent using the collected data, The collected data includes: the state data of the scraper, the control instructions of the intelligent agent to the scraper, and the reward value, wherein: The state data includes: cylinder extension amount, steering articulation angle, scraper coordinate position, acceleration and angular velocity of the front vehicle body, The control instructions include: oil cylinder extension speed, steering hinge angular velocity and wheel angular velocity.
7. A scraper, characterized in that: The scraper loader is provided with the intelligent agent trained by the training method as described in any one of claims 1 to 4, and the intelligent agent acts on the scraper loader to complete automatic loading control.
8. An electronic device, characterized in that: The electronic device comprises: a processor, and a memory coupled to the processor, The memory is used to store a computer program; and The processor is used to execute the computer program stored in the memory so that the electronic device executes the training method as described in any one of claims 1 to 4, or executes the automatic shoveling control method as described in claim 5.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a computer program or instructions. When the computer program or instructions are executed on a computer, the computer executes the training method as described in any one of claims 1 to 4, or executes the automatic shoveling control method as described in claim 5.
Citation Information
Patent Citations
Scraper autonomous unloading control system based on reinforcement learning
CN119571880A