Multi-agent autonomous learning method for hydraulic support

By employing a multi-agent autonomous learning method for hydraulic supports, and combining real and simulated data, displacement/pressure-time action sequences are generated and corrected. This solves the problems of equipment control and data fusion in coal mining, and realizes an efficient and safe intelligent decision-making system.

CN121543624AActive Publication Date: 2026-02-17TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511703880.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-02-17
Estimated Expiration
2045-11-19

AI Technical Summary

Technical Problem

Existing coal mining technologies face challenges such as complex geological environments, diverse equipment control strategies, high levels of personnel involvement, low production efficiency, and a lack of closed-loop data fusion analysis combining real and simulated data.

Method used

A multi-agent autonomous learning method for hydraulic supports is adopted. By pre-deploying learning-type real-loop agents, learning-type simulation loop agents, and learning-type dual-loop agents, and combining scene features and preset strategies, displacement/pressure-time action sequences are generated, detected, and corrected, achieving deep integration of real data and simulation data.

Benefits of technology

A robust, highly efficient, safe, redundant, and knowledge-accumulating intelligent control system for hydraulic supports has been constructed, achieving a leap from single-device control to intelligent decision-making across all scenarios, thereby enhancing the autonomy, collaboration, and safety of coal mining operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121543624A_ABST
    Figure CN121543624A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-agent autonomous learning method for a hydraulic support, and belongs to the technical field of coal mine intellectualization. Comprising the steps that scene features input by a learning type reality ring Agent and a displacement / pressure-time action sequence output by the learning type reality ring Agent serve as input of a learning type double-ring Agent, the learning type double-ring Agent corrects the displacement / pressure-time action sequence, and the corrected displacement / pressure-time action sequence is obtained; the corrected displacement / pressure-time action sequence is returned to the learning type simulation ring Agent again to verify whether conflicts exist or not, if conflicts exist after correction, correction is continued, loop iteration is conducted till the corrected displacement / pressure-time action sequence does not have conflicts, and a final actual action strategy is expanded to a preset strategy set. The method is used for training the learning type reality ring Agent with higher quality. According to the invention, closed-loop data fusion analysis combining real data and simulation data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent coal mine technology, and particularly relates to a hydraulic support multi-agent autonomous learning method. BACKGROUND

[0002] Coal mining, as an important part of energy supply, has important economic and social significance. However, the traditional coal mining process faces many challenges such as complex geological environment, diversified equipment control strategy, high degree of personnel participation, and low production and mining efficiency. In the current era of artificial intelligence, promoting the intelligent transformation of the coal mining field and improving the autonomy, collaboration, adaptability and safety of coal mining operations have become the key to solving the current difficulties.

[0003] The current intelligent system of the coal mining face faces limitations such as difficulty in multi-device collaboration, poor adaptability to dynamic scenarios, low knowledge reuse, and high personnel workload. Existing technologies often only rely on real historical data of coal production to develop Agent models or develop simulation environments according to real-world rules and coal mining processes to generate simulation data, lacking closed-loop data fusion analysis that combines real data with simulation data. SUMMARY

[0004] To solve the above technical problems, the present application provides a hydraulic support multi-agent autonomous learning method. The technical solution of the present application is as follows: A hydraulic support multi-agent autonomous learning method is implemented based on pre-deployed learning-type real loop Agents, learning-type simulation loop Agents, and learning-type double-loop Agents, comprising: S1, acquiring scene features and matching the scene features with a preset strategy set; S2, if there is a scene in the preset strategy set that matches the scene features, outputting a displacement / pressure-time action sequence corresponding to the scene features in the preset strategy set; S3, if there is no scene in the preset strategy set that matches the scene features, inputting the scene features into the learning-type real loop Agent, the learning-type real loop Agent outputting a displacement / pressure-time action sequence under the scene features, and the displacement / pressure-time action sequence output by the learning-type real loop Agent serving as input for the learning-type simulation loop Agent; S4, the learning simulation ring agent detects the displacement / pressure-time action sequence, and feeds back a plurality of evaluation indexes based on a preset evaluation index set, the evaluation indexes are used to reflect whether the displacement / pressure-time action sequence output by the learning reality ring agent is in conflict; if the displacement / pressure-time action sequence is not in conflict, the actual action strategy is directly given; if the displacement / pressure-time action sequence is in conflict, the evaluation indexes will be part of the reward function of the learning double-ring agent to evaluate the value of the state under the scene characteristics; S5, the input of the learning reality ring agent, i.e. the scene characteristics and the output displacement / pressure-time action sequence, is used as the input of the learning double-ring agent, the learning double-ring agent corrects the displacement / pressure-time action sequence under the scene characteristics to obtain a corrected displacement / pressure-time action sequence with the best value under the scene characteristics, and returns the corrected displacement / pressure-time action sequence to the learning simulation ring agent again to verify whether there is a conflict, if the corrected displacement / pressure-time action sequence is not in conflict, the actual action strategy is directly given, if the corrected displacement / pressure-time action sequence is in conflict, the corrected displacement / pressure-time is fed back to the learning double-ring agent again for correction, and the cycle is iterated until the corrected displacement / pressure-time action sequence is not in conflict, the final actual action strategy is expanded to the preset strategy set, which is used to train a learning reality ring agent with higher quality.

[0005] Preferably, the matching of the scene characteristics with the preset strategy set comprises: The preset strategy set is searched by using a search algorithm.

[0006] Preferably, the learning reality ring agent comprises a learning reality ring agent of a hydraulic support, and the S3 further comprises: constructing a learning reality ring agent of a hydraulic support, specifically comprising: First, according to the action characteristics of the four actions of column lowering, support moving, column lifting and pushing of the hydraulic support in the well, the output of the learning reality ring agent of the hydraulic support is constructed as a displacement-time action sequence and a pressure-time action sequence; Second, the input of the learning reality ring agent of the hydraulic support is constructed according to the environmental physical constraints, the action target constraints, the device coordination constraints and the action constraints of the current support; Then, the loss function of the learning reality ring agent of the hydraulic support is constructed according to the hardware limit constraints, the motion target constraints and the action time sequence constraints; Finally, the learning reality ring agent of the hydraulic support is trained according to the output, the input and the loss function of the learning reality ring agent of the hydraulic support.

[0007] Preferably, in the input of the learning reality loop Agent of the hydraulic support, the following is included: In terms of environmental physical constraints, the angle of the hydraulic support base, the specific quantity of the top and bottom plate material is quantified as a mathematical index, and the angle of the hydraulic support base is set to The friction coefficient of the bottom plate is taken as an index reflecting the material of the bottom plate The compressive strength of the top plate is taken as an index reflecting the material of the top plate ; In terms of action target constraints, the target displacement values of the moving frame and the pushing are set to and The target minimum pressure of the descending column and the target initial support force of the ascending column are set to and ; In terms of equipment coordination constraints, the traversal direction of the coal mining machine is set to forward or reverse, and is quantified as a mathematical index, represented as 0 and 1 respectively; the action state of the adjacent frame is divided into 0 and 1 whether it reaches the state of the action of the frame; specifically, when the coal mining machine passes, the hydraulic support performs four actions in order: descending column, moving frame, ascending column, and pushing, and the adjacent frame performs three actions: descending column, moving frame, and ascending column, and then the frame performs the action, so the adjacent frame completes the three actions: descending column, moving frame, and ascending column, and the state is divided into 0, and vice versa, which is divided into 1; In terms of frame action constraints, the first two frames of pressure and displacement are taken as input, denoted as and .

[0008] Preferably, in the construction of the loss function of the learning reality loop Agent of the hydraulic support, the following is included: Set the loss function of the learning reality loop Agent of the hydraulic support as m, then: In terms of hardware limit constraints, set the real-time speed of descending column, moving frame, ascending column, and pushing to , and the corresponding single-action limit speed to If the hardware limit constraint aspect violates the real world rules, give the corresponding loss, denoted as: ; ; ; ; In terms of motion target constraints, set the moving frame displacement value to The pushing displacement value is , the descending column minimum pressure is , the ascending column initial support force is , and the target displacement values of the moving frame and the pushing are and The target minimum pressure of the descending column is The target initial support force of the ascending column is The loss is determined according to the difference ratio of real-time value and target value, which is expressed as: ; ; ; ; In terms of action timing constraints, if the following error traverses the direction, then If the following error traverses the correct direction but conflicts with the adjacent frame, i.e., the adjacent frame action state is 1, the frame starts to act, and sets .

[0009] Preferably, the S4 further comprises constructing a learning type simulation ring Agent, specifically comprising constructing a hydraulic system mathematical model, a hydraulic support mathematical model, and a support and surrounding rock coupling model. The hydraulic system mathematical model is expressed as: ; ; ; Wherein, is the total supply flow of the emulsion pump, is the actual consumption flow of the executive element, is the system leakage flow; is the action speed of the executive element, is the inlet cavity area of the executive element; is the supply pressure of the hydraulic system, is the load of the executive element, is the back pressure of the outlet cavity of the executive element, is the outlet cavity area of the executive element; The hydraulic support mathematical model comprises an ascending column model, a descending column model, a frame moving model, and a pushing model: The ascending column model is expressed as: ; Wherein, is the ascending column thrust, is the supply pressure of the hydraulic system, is the rodless cavity area of the column, is the rod cavity area of the column, is the outlet cavity back pressure; The descending column model is expressed as: ; Wherein, To reduce column force; The moving frame model is represented as: ; in, For the shifting thrust, This indicates the area of ​​the rodless cavity of the jack. This indicates the area of ​​the rod cavity in the jack; The push-pull model is represented as: ; in, For pushing and pulling forces; The coupling model between the support and the surrounding rock includes a stiffness coupling model, a strength coupling model, and a stability coupling model: The stiffness coupling model is characterized by the compaction stiffness of the top plate and is expressed as follows: ; in, This refers to the compaction strength of the top slab on the support after the frame is moved. This is the initial support force for the lifting column after the frame is moved. This represents the total displacement of the moving frame; The strength coupling model is characterized by the dynamic resistance of the surrounding rock and is expressed as follows: ; in, The dynamic resistance of the surrounding rock during frame shifting / pushing. The drag-rate coefficient, For displacement rate; The stability coupling model, characterized by the risk of roof collapse, is expressed as follows: ; in, This represents the displacement increment during frame relocation. This is the time increment during the transfer of the rack. This represents the maximum moving rate.

[0010] Preferably, after constructing the mathematical model of the hydraulic system, the mathematical model of the hydraulic support, and the coupling model between the support and the surrounding rock, the method further includes: Step 1: Verify the action speed based on the mathematical model of the hydraulic system, specifically including: Action speed verification: If the displacement-time action sequence generated by the learning simulation loop agent is calculated... This indicates that the action speed is too fast, exceeding the maximum total liquid supply flow rate of the emulsion pump. The speed needs to be reduced; Step 2: Verification based on the mathematical model of the hydraulic support, specifically including: Rising column pressure verification: if the rising column pressure calculated based on the pressure-time action sequence generated by the learning simulation loop Agent exceeds the rated working resistance of the column calculated ; , it is determined that the pressure is too high, wherein the rising column pressure represents the initial support force size during the rising of the column; Displacement verification: if the speed calculated based on the displacement-time action sequence generated by the learning simulation loop Agent is not less than the safe speed of the moving frame and the safe speed of the pusher , it is determined that the displacement is too large, wherein: ; ; wherein, and are the maximum flow allowed for the moving frame and the pusher, respectively; Third step: verification based on the support and surrounding rock coupling model, specifically including: (1) Stiffness coupling verification, specifically including: if the <experience value> calculated based on the pressure-time action sequence generated by the learning simulation loop Agent is less than the rated working resistance of the column (2) Strength coupling verification, specifically including: if the calculated based on the pressure-time action sequence generated by the learning simulation loop Agent is greater than the rated working resistance of the column , it is determined that the dynamic resistance of the surrounding rock exceeds the bearing capacity of the support, and the action speed needs to be reduced; wherein, represents the upper limit of the dynamic resistance of the surrounding rock that the hydraulic support can withstand under the rated working condition, is the rated working pressure of the column, and k is the mapping coefficient of the dynamic resistance of the surrounding rock and the column pressure; (3) Roof caving risk verification, specifically including: if the calculated based on the displacement-time action sequence generated by the learning simulation loop Agent is greater than the rated working resistance of the column ; The learning simulation loop Agent determines whether the action speed of the displacement / pressure-time action sequence is too fast and whether the initial support force of the rising column needs to be adjusted based on the verification results of the first to third steps to form an evaluation index set [Y, Z]. If the action speed is too fast, Y=1, otherwise 0. If the initial support force of the rising column needs to be adjusted, Z=1, otherwise 0.

[0011] Preferably, S5 further comprises constructing a learning-type double-loop Agent, specifically comprising: Firstly, constructing the input and output of the learning-type real loop Agent as the state space of the learning-type double-loop Agent; Secondly, setting the displacement / pressure-time action sequence at the designated time of correction as the action space of the learning-type double-loop Agent; Thirdly, determining the reward function of the learning-type double-loop Agent according to the hardware limit constraint, the motion target constraint, the action timing constraint and the simulation feedback constraint; Finally, training the learning-type double-loop Agent according to the state space, the action space and the reward function of the learning-type double-loop Agent.

[0012] Preferably, the reward function of the learning-type double-loop Agent is set as r , Then: The real-time speed of the descending column, the moving frame, the ascending column and the pushing trolley is set as , and the corresponding single-action limit speed is If the hardware limit constraint violates the real-world rules, the corresponding penalty is given, which is represented as: ; ; ; ; In terms of the motion target constraint, the displacement value of the moving frame is set as , the displacement value of the pushing trolley is set as , the minimum pressure of the descending column is set as , the initial support force of the ascending column is set as , the target displacement values of the moving frame and the pushing trolley are set as and respectively, the target minimum pressure of the descending column is set as , and the target initial support force of the ascending column is set as The penalty is determined according to the difference between the real-time value and the target value, which is represented as: ; ; ; ; In terms of the action timing constraint, if the following error traverses the direction, then ; if the following error traverses the correct direction but conflicts with the adjacent frame, i.e. the adjacent frame action state is 1, and the current frame starts to act, then ; Regarding the simulation feedback constraints, the set of evaluation metrics given by the learning simulation loop agent is set as follows: ,if ,but ;if =1, then ;if ,but ;if ,but .

[0013] Preferably, the learning dual-loop agent is trained using the DDPG algorithm based on the state space, action space, and reward function of the learning dual-loop agent.

[0014] All of the above-mentioned optional technical solutions can be combined arbitrarily, and the present invention will not provide a detailed description of the structure after each combination.

[0015] By means of the above solution, the beneficial effects of the present invention are as follows: By pre-deploying a learning-based reality loop agent, a learning-based simulation loop agent, and a learning-based dual-loop agent, the learning-based reality loop agent can generate displacement / pressure-time action sequences based on real-world patterns according to the current scene characteristics of the coal mining face. When the learning-based simulation loop agent detects conflicts in the displacement / pressure-time action sequences under the scene characteristics, the learning-based dual-loop agent corrects the displacement / pressure-time action sequences under the scene characteristics until there are no conflicts. The final actual action strategy is then expanded to the preset strategy set to train a higher-quality learning-based reality loop agent. This achieves closed-loop data fusion analysis that combines real data with simulated data, thereby ensuring the construction of a more accurate learning-based reality loop agent.

[0016] The autonomous learning method of the multi-agent closed-loop architecture is not only limited to generating more accurate action sequences, but also more conducive to realizing the leap from single device control to full-scenario intelligent decision-making system. Through the deep integration of real and simulated data, it not only retains the physical authenticity of the real scene, but also breaks through the sample limitations of real data through simulation. Finally, it builds a new generation of intelligent control system for hydraulic supports with strong robustness, high iteration efficiency, safety redundancy, knowledge accumulation and autonomous decision-making, thus providing core technical support for the intelligentization of coal mines.

[0017] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, the preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings. Attached Figure Description

[0018] Figure 1is a flowchart of a hydraulic support multi-agent autonomous learning method provided by an embodiment of the present application.

[0019] Figure 2 is a flowchart of a learning-type real loop Agent for constructing a hydraulic support in an embodiment of the present application.

[0020] Figure 3 is a flowchart of a learning-type simulation loop Agent for constructing a hydraulic support in an embodiment of the present application.

[0021] Figure 4 is a flowchart of a learning-type double-loop Agent for constructing a hydraulic support in an embodiment of the present application. DETAILED DESCRIPTION

[0022] The specific embodiments of the present application will be further described in detail below with reference to the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.

[0023] As shown in Figure 1 The hydraulic support multi-agent autonomous learning method provided by the embodiment of the present application is realized based on a pre-deployed learning-type real loop Agent, a learning-type simulation loop Agent, and a learning-type double-loop Agent. The learning-type real loop Agent is a successful strategy automatic integration-analysis-modeling algorithm for operation records and working condition data of an actual coal mining face production process, which is used to realize development of experience data of a real coal mining process, so as to generate a real rule-based action sequence based on coal mining site data. The learning-type simulation loop Agent is a feasible strategy automatic generation-simulation-optimization-induction algorithm for production and operation of a coal mining face in a simulation system, which realizes exploration of simulation data of a virtual coal mining process. The learning-type double-loop Agent is a data and control fusion analysis algorithm and a virtual-real strategy credible evaluation algorithm between real and simulation systems, which is used to realize continuous correction and improvement of each Agent model. On this basis, the hydraulic support multi-agent autonomous learning method includes the following steps S1 to S5: S1, acquire scene features and match the scene features with a preset strategy set.

[0024] Specifically, when matching the scene features with the preset strategy set, S1 can search the preset strategy set by using a search algorithm such as an exhaustive method. The preset strategy set includes a plurality of corresponding relationships between scene features and displacement / pressure-time action sequences. The displacement / pressure-time action sequence includes a displacement-time action sequence and a pressure-time action sequence.

[0025] S2, if there is a scene in the preset strategy set that matches the scene features, output the displacement / pressure-time action sequence corresponding to the scene features in the preset strategy set.

[0026] S3. If there is no scene matching the scene features in the preset strategy set, the scene features are input into the learning reality loop agent. The learning reality loop agent outputs the displacement / pressure-time action sequence under the scene features, and the displacement / pressure-time action sequence output by the learning reality loop agent is used as the input of the learning simulation loop agent.

[0027] It should be noted that before inputting scene features into the learning reality loop agent, the learning reality loop agent needs to be constructed first. The learning reality loop agent is essentially a generator based on a diffusion model. The idea is to construct an action generator from existing reality data using the diffusion model approach, which can generate action sequences based on reality patterns. In this embodiment of the invention, an action generator for hydraulic supports is constructed based on the coal mining face, i.e., a learning reality loop agent for hydraulic supports. For example... Figure 2 As shown, building a learning reality loop agent for hydraulic supports requires three steps: setting the input, setting the loss function, and setting the output.

[0028] First, the output settings for the learning-type reality loop agent of the hydraulic support are defined. The output of the learning-type reality loop agent of the hydraulic support is based on the actions of the hydraulic support, which mainly perform four actions downhole: lowering the support column, moving the support, raising the support column, and pushing the conveyor. These four actions are reflected on the sensors as a sequence of stroke changes over time and a curve of support column pressure changes over time. Therefore, the output is set as a displacement-time action sequence and a pressure-time action sequence. In this embodiment, the time steps of the displacement-time action sequence and the pressure-time action sequence are consistent and the number of steps is the same, i.e., displacement and pressure action sequences generated at the same time. Therefore, the output settings for the learning-type reality loop agent of the hydraulic support can be summarized as follows: based on the action characteristics of the four actions performed by the hydraulic support downhole—lowering the support column, moving the support, raising the support column, and pushing the conveyor—the output of the learning-type reality loop agent of the hydraulic support is constructed as a displacement-time action sequence and a pressure-time action sequence. Secondly, the inputs to the learning reality loop agent of the hydraulic support are set. The input of the learning reality loop agent needs to consider the preceding states of the output, that is, it needs to consider why such a sequence of actions is generated. Therefore, in this embodiment of the invention, the input of the learning reality loop agent of the hydraulic support is constructed based on environmental physical constraints, action target constraints, equipment coordination constraints, and the action constraints of the hydraulic support itself.

[0029] Regarding environmental physical constraints, considering the physical factors of the environment, since the base angle of the hydraulic support determines its travel angle, and the material of the top and bottom plates determines the material of the environment in which the hydraulic support is located, this embodiment of the invention treats the base angle and the material of the top and bottom plates as variables. In conjunction with the above, this embodiment of the invention, in terms of environmental physical constraints, quantifies the base angle and the material of the top and bottom plates of the hydraulic support into mathematical indicators, setting the base angle of the hydraulic support as... The coefficient of friction of the base plate is used as an indicator of the base plate material. The compressive strength of the roof slab is used as an indicator of the roof slab material. .

[0030] Regarding action target constraints, based on the actions themselves, hydraulic supports require action target values ​​to perform actions. The moving and pushing actions require corresponding target displacement values. For example, the maximum value for moving and pushing is the cutting depth of the coal mining machine (865mm). However, the actual situation needs to be dynamically adjusted based on reality. For instance, if the value of the stroke sensor before moving is 800mm, then the target displacement value for a single moving action can only be 800mm, and needs to be dynamically changed based on the actual value. The target displacement value for pushing needs to be dynamically changed based on the stroke value after the moving action ends. For example, if the stroke value after the moving action ends is 50mm, then the target displacement value for pushing is the cutting depth of the coal mining machine minus the stroke value after the moving action ends, which is 815mm. Lowering the column is the process of the column detaching from the roof, with pressure decreasing from high to low. The target minimum pressure is when it drops to near 0. Raising the column needs to reach the range of the rated initial support force required by the hydraulic support, such as 80% of 31.5Mpa to the maximum, i.e., 24-31.5Mpa. In conjunction with the above, in terms of action target constraints, this embodiment of the invention sets the target displacement values ​​for frame shifting and pushing as follows: and The target minimum pressure for lowering the column and the target initial support force for raising the column are set as follows: and .

[0031] Regarding equipment coordination constraints, since hydraulic supports do not operate independently in the coal mining face, but rather form a cluster of hydraulic supports that coordinate their actions according to certain rules, it is necessary to consider the action status of adjacent supports to determine whether the current support can take action. Whether to refer to the status of the support on the left or right depends on the traversal direction of the coal mining machine. Based on the above, this embodiment of the invention sets the traversal direction of the coal mining machine as either forward or reverse, quantified as mathematical indices represented by 0 and 1 respectively. It also classifies whether the action status of an adjacent support reaches the state where the current support can take action as 0 or 1. Specifically, as the coal mining machine passes by, the hydraulic supports sequentially perform four actions: lowering the support, moving the support, raising the support, and pushing the conveyor. After an adjacent support completes the three actions of lowering, moving, and raising the support, the current support can take action. Therefore, if an adjacent support completes the three actions of lowering, moving, and raising the support, the state is classified as 0; otherwise, the state is classified as 1.

[0032] Regarding the motion constraints of this frame, since the stroke and pressure values ​​are not constant, the first two frames of pressure and displacement need to be used as input. Based on the above, in this embodiment of the invention, the first two frames of pressure and displacement are used as input for the motion constraints of this frame, denoted as... and .

[0033] Next, a loss function is set for the learning-based reality loop agent of the hydraulic support. Specifically, when setting the loss function, relevant reality constraints need to be incorporated to guide the output action sequence to not violate reality rules. Considering hardware limits, such as the rate of change of column lowering / raising pressure not exceeding the cylinder limit, and the rate of support shifting / pushing speed not exceeding the cylinder limit, penalties are incurred if these limits are exceeded. Regarding the target displacement of support shifting / pushing, support shifting is determined based on the previous two frames of motion. For example, if the stroke value before support shifting is 800, then the target displacement value is 800. When the action is completed, the further the shift value deviates from 800, the higher the penalty. For pushing, the target displacement value is the difference between the coal mining machine's cutting depth and the actual value reached by support shifting; the greater the deviation from the target displacement value, the higher the penalty. Regarding the target for lowering the support column, if the pressure is within a reasonable range at the end of the lowering process, there is no penalty. If the pressure does not reach the reasonable range, a penalty is imposed based on the deviation between the actual pressure value and the target pressure value. For raising the support column, the focus is on whether the initial support force reaches a reasonable range. If it does not, a penalty is imposed based on the deviation from the target initial support force range. Regarding the timing constraints, if the coal mining machine traversal sequence is 0, the left adjacent support takes priority, and vice versa. If the sequence is correct and the adjacent support action is 0, the current support can begin action; otherwise, a severe penalty is imposed if the timing constraints are violated. In conjunction with this part, this embodiment of the invention constructs a loss function for the learning-type reality loop agent of the hydraulic support based on hardware limit constraints, motion target constraints, and action timing constraints.

[0034] In one specific embodiment, the loss function for constructing the learning reality loop agent of the hydraulic support includes: Let m be the loss function of the learning-based reality loop agent for the hydraulic support. Then: Regarding hardware limit constraints, the real-time speeds for lowering the column, moving the frame, raising the column, and pushing the conveyor are set as follows: Its corresponding single-action limit speed is If the hardware limits violate real-world rules, a corresponding loss is given, expressed as: ; ; ; ; Regarding the constraints on the motion target, the displacement value of the moving frame is set to... The displacement value of the pusher is The minimum pressure of the column is The initial support force of the lifting column is The target displacement values ​​for the moving frame and the pushing slide are respectively and The target minimum pressure for the column is The initial support force of the rising column is According to real-time values With target value The loss is determined by the proportion of the difference, expressed as: ; ; ; ; Regarding action timing constraints, if the traversal direction is followed by an error, then... If it follows the correct direction but conflicts with the adjacent frame (i.e., the adjacent frame's action status is 1), this frame begins to move and is set. .

[0035] Finally, the learning reality loop agent for the hydraulic support is trained based on the output, input, and loss function of the learning reality loop agent for the hydraulic support.

[0036] Thus, the learning-based reality loop agent is essentially set up based on a diffusion model generator. It can then extract and process real-world data to train itself. This extraction and processing of real-world data involves transforming the real-time collected data into a format that matches the input format of the learning-based reality loop agent. For example, if the input format of the learning-based reality loop agent is scene features, then the scene features will be extracted during the data processing.

[0037] S4, the learning simulation loop agent detects the displacement / pressure-time action sequence and feeds back several evaluation indicators based on a preset evaluation indicator set. The evaluation indicators are used to reflect whether there is a conflict in the displacement / pressure-time action sequence output by the learning reality loop agent. If there is no conflict in the displacement / pressure-time action sequence, the actual action strategy is given directly. If there is a conflict in the displacement / pressure-time action sequence, the evaluation indicators will be used as part of the reward function of the learning dual-loop agent to evaluate the value of the state under the scene features.

[0038] It should be noted that a learning simulation loop agent needs to be constructed before step S4. Essentially, the learning simulation loop agent simplifies the complex and abstract real world into a mathematical model, sets relevant constraints, and simulates the action sequences generated by the learning reality loop agent. For example... Figure 3 As shown, this embodiment of the invention constructs a hydraulic system mathematical model, a hydraulic support mathematical model, and a support-surrounding rock coupling model based on the working principle of hydraulic supports, the characteristics of actuators, and the support-surrounding rock coupling theory. The models are logically coherent and their parameters are interconnected, as detailed below: 1. Mathematical model of hydraulic system: The core formulas of a hydraulic system include fluid supply flow balance, flow-speed relationship, and pressure-load balance. Therefore, the mathematical model of the hydraulic system is expressed as: ; ; ; in, This represents the total flow rate (L / min) of the emulsion pump. The actual flow rate consumed by the actuator (column / jack). System leakage flow rate (L / min); For the speed of the actuator's operation, The area of ​​the liquid inlet chamber of the actuator ; The hydraulic system is supplied with hydraulic pressure (MPa). For the load of the actuator (KN). The back pressure (MPa) of the liquid outlet chamber of the actuator. The area of ​​the liquid outlet chamber of the actuator.

[0039] Specifically, regarding Since the hydraulic support has four actions—lowering the column, raising the column, moving the support, and pushing the slide—let's assume the generator output sequence is a curve of displacement versus pressure over time. Moving the support and pushing the slide are performed by jacks as actuators. In this case... Regarding the lowering and raising of the column, the actuator is the column itself, which is reflected in the data as pressure changes over time. However, because this embodiment of the invention uses a simplified mathematical model, it is assumed that the height of each lowering and raising operation is the same. Then at this time At this time It is derived from the pressure-time change curve, specifically the duration of the lifting column's movement.

[0040] 2. Mathematical model of hydraulic support: The column and jack are the "core actuators" of the hydraulic support's operation. Based on the steady-state force principle of a double-acting hydraulic cylinder, this invention establishes simplified mathematical models for the column to characterize the mechanical relationship between column raising / lowering and the jack to characterize the mechanical relationship between jack moving / pushing, for four actions: column raising, column lowering, support moving, and jack pushing. Specifically, by establishing mechanical formulas for the column raising, column lowering, support moving, and jack pushing models, the conversion relationship between "pressure → force" is quantified, as follows: The rising column model is represented as: ; in, For the lifting column thrust, To supply fluid pressure to the hydraulic system, The area of ​​the rodless cavity of the column is given. Let be the area of ​​the cavity of the column with rods. This is the back pressure of the liquid outlet chamber; The descending column model is represented as: ; in, To reduce the column force and unload the hydraulic support, the column lowering and raising chambers are switched.

[0041] The moving frame model is represented as: ; in, For the shifting thrust, This indicates the area of ​​the rodless cavity of the jack. This indicates the area of ​​the rod cavity in the jack; The push-pull model is represented as: ; in, To achieve the pushing and pulling force, the resistance of the scraper conveyor must be overcome.

[0042] 3. Coupled model of support and surrounding rock: The support-surrounding rock coupling model is the "core of system stability." This invention, based on the three theories of stiffness coupling, strength coupling, and stability coupling, combines displacement and pressure data output by the generator to quantify the adaptation relationship between "support action → surrounding rock response." The support-surrounding rock coupling model includes a stiffness coupling model, a strength coupling model, and a stability coupling model, as detailed below: The stiffness coupling model is characterized by the compaction stiffness of the top plate and is expressed as follows: ; in, This refers to the compaction strength of the top slab on the support after the frame is moved. This is the initial support force for the lifting column after the frame is moved. This is the total displacement of the moving frame; if If the value is less than the engineering experience value, it indicates that the top slab is soft and the initial support force needs to be increased.

[0043] The strength coupling model is characterized by the dynamic resistance of the surrounding rock and is expressed as follows: ; in, The dynamic resistance of the surrounding rock during frame shifting / pushing. The drag-rate coefficient, For displacement rate; The stability coupling model, characterized by the risk of roof collapse, is expressed as follows: ; in, This represents the displacement increment during frame relocation. This is the time increment during the transfer of the rack. This is the maximum moving speed. If Greater than If this is confirmed, there is a risk of roof collapse, and the scaffolding movement must be stopped immediately, the initial support force increased, and the roof reinforced.

[0044] At this point, a simplified mathematical model of the learning simulation loop agent has been constructed. The following details how to verify the action sequence based on this mathematical model: Step 1: Verify the action speed based on the mathematical model of the hydraulic system, specifically including: Action speed verification: If the displacement-time action sequence generated by the learning simulation loop agent is calculated... This indicates that the action speed is too fast, exceeding the maximum total liquid supply flow rate of the emulsion pump. The speed needs to be reduced.

[0045] Step 2: Verification based on the mathematical model of the hydraulic support, specifically including: Columnar pressure verification: If the columnar pressure is extracted based on the pressure-time action sequence generated by the learning simulation loop agent... Calculated Exceeding the rated working resistance of the column If so, the pressure is determined to be too high, including the lifting column pressure. This represents the initial support force during the lifting of the column; Displacement verification: If the velocity calculated from the displacement-time action sequence generated by the learning simulation loop agent is not less than the safe rate of the frame movement. and the safe speed of pushing If the displacement is too large, then it is determined that: ; ; in, and These are the maximum allowable flow rates for moving the frame and pushing the conveyor, respectively. This allows us to calculate the maximum allowable speeds for moving the frame and pushing the conveyor.

[0046] Step 3: Validation based on the coupling model between the support and the surrounding rock, specifically including: (1) Stiffness coupling verification, specifically including: If the stress-time action sequence generated by the learning simulation loop agent is calculated... If the value is less than the empirical value, it indicates that the top plate compaction stiffness is insufficient, and the initial support force of the lifting column needs to be increased. (2) Strength coupling verification, specifically including: If the stress-time action sequence generated by the learning simulation loop agent is calculated... If the dynamic resistance of the surrounding rock exceeds the bearing capacity of the support, the operating rate needs to be reduced; among which, This indicates the upper limit of the dynamic resistance of the surrounding rock that the hydraulic support can withstand under rated operating conditions. The rated working pressure of the column is given by k, which is the mapping coefficient between the dynamic resistance of the surrounding rock and the pressure of the column. (3) Verification of roof collapse risk, specifically including: If the displacement-time action sequence generated by the learning simulation loop agent is calculated... Therefore, it is determined that the mutation rate needs to be reduced and the mutation range expanded. ; The learning simulation loop agent judges whether the action rate of the displacement / pressure-time action sequence is too fast and whether the initial support force of the lifting column needs to be adjusted based on the verification results of the first to third steps, forming an evaluation index set [Y,Z]. If the action rate is too fast, Y=1, otherwise it is 0; if the initial support force of the lifting column needs to be adjusted, Z=1, otherwise it is 0.

[0047] Based on the learning-based simulation loop agent constructed above, step S4, when detecting the displacement / pressure-time action sequence, checks evaluation indicators such as whether the action rate of the displacement / pressure-time action sequence is too fast and whether the initial support force of the lifting column has been adjusted. The actual action strategy includes when to perform actions such as lowering the column, moving the frame, raising the column, and pushing the slide, as well as the specific values ​​of pressure / displacement during these actions. The specific content of the reward function for the learning-based dual-loop agent will be explained in the following steps.

[0048] S5, the input to the learning-type reality loop agent, namely the scene features and the output displacement / stress-time action sequence, is used as the input to the learning-type dual-loop agent. The learning-type dual-loop agent corrects the displacement / stress-time action sequence under the scene features to obtain the corrected displacement / stress-time action sequence with the best value under the scene features. The corrected displacement / stress-time action sequence is then returned to the learning-type simulation loop agent to verify whether there is a conflict. If the corrected displacement / stress-time action sequence has no conflict, the actual action strategy is directly given. If the corrected displacement / stress-time action sequence has a conflict, the corrected displacement / stress-time action sequence is fed back to the learning-type dual-loop agent for correction. This process is repeated until the corrected displacement / stress-time action sequence has no conflict. Finally, the actual action strategy is expanded to a preset strategy set to train a higher-quality learning-type reality loop agent.

[0049] It should be noted that the S5 section also includes the construction of a learning-based dual-loop agent, such as... Figure 4 As shown, constructing a learning-based dual-loop agent specifically includes: First, the inputs and outputs of the learning reality loop agent are constructed as the state space of the learning dual-loop agent. That is, the state space of the learning dual-loop agent includes environmental physical constraints, action target constraints, device coordination constraints, local action constraints, and the output of the learning reality loop agent. Details regarding the environmental physical constraints, action target constraints, device coordination constraints, local action constraints, and the output of the learning reality loop agent are provided in the above embodiments and will not be repeated here.

[0050] Secondly, the action space of the learning dual-loop agent is defined as the action space for modifying the displacement / pressure-time action sequence at a specified moment. Specifically, the action space of the learning dual-loop agent is set to modify the displacement / pressure value at a certain moment, thereby achieving the goal of modifying the action sequence.

[0051] Next, the reward function of the learning dual-loop agent is determined based on hardware limit constraints, motion target constraints, action timing constraints, and simulation feedback constraints.

[0052] In one specific embodiment, the reward function of the learning dual-loop agent is set to r. , but: Set the real-time speeds for lowering the column, moving the support frame, raising the column, and pushing the conveyor belt as follows: Its corresponding single-action limit speed is If hardware limitations violate real-world rules, a corresponding penalty will be imposed, as follows: ; ; ; ; Regarding the constraints on the motion target, the displacement value of the moving frame is set to... The displacement value of the pusher is The minimum pressure of the column is The initial support force of the lifting column is The target displacement values ​​for the moving frame and the pushing slide are respectively and The target minimum pressure for column reduction is The initial support force of the rising column is The penalty is determined based on the ratio of the difference between the real-time value and the target value, expressed as: ; ; ; ; Regarding action timing constraints, if the traversal direction is followed by an error, then... If it follows the correct direction but conflicts with the adjacent frame (i.e., the adjacent frame's action status is 1), this frame begins to move and is set. ; Regarding the simulation feedback constraints, the set of evaluation metrics given by the learning simulation loop agent is set as follows: ,if ,but ;if =1, then ;if ,but ;if ,but .

[0053] Finally, the learning double-loop agent is trained based on the state space, action space, and reward function of the learning double-loop agent.

[0054] Specifically, when training a learning-type two-loop agent, the DDPG (Deep Deterministic Policy Gradient) algorithm can be used to train the learning-type two-loop agent based on its state space, action space, and reward function.

[0055] It should be noted that the specific type of the learning reality loop agent is not specifically limited in this embodiment of the invention. For example, the learning reality loop agent can be a manual control decision-making model after the automation of the central hydraulic support cluster, an intelligent prediction and control model for the straightness of the working face, a dynamic prediction model for spatiotemporal pressure events in the intelligent fully mechanized mining face, or an instant prediction model for the pressure-bearing effect after the initial support of the hydraulic support, and so on. The learning simulation loop agent and the learning dual-loop agent correspond to the learning reality loop agent. For example, the learning simulation loop agent corresponding to the intelligent prediction and control model for the straightness of the working face is a simulation model that controls the straightness of the working face, and the corresponding learning dual-loop agent is a combination of a reality model and a simulation model that controls the straightness of the working face.

[0056] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multi-agent autonomous learning method for hydraulic supports, characterized in that, Implementation based on pre-deployed learning-based real-world loop agents, learning-based simulated loop agents, and learning-based dual-loop agents, including: S1, Obtain scene features and match scene features with a preset strategy set; S2, If there is a scene in the preset strategy set that matches the scene features, then output the displacement / pressure-time action sequence corresponding to the scene features in the preset strategy set; S3, if there is no scene matching the scene features in the preset strategy set, then the scene features are input into the learning reality loop agent, the learning reality loop agent outputs the displacement / pressure-time action sequence under the scene features, and the displacement / pressure-time action sequence output by the learning reality loop agent is used as the input of the learning simulation loop agent; S4, the learning simulation loop agent detects the displacement / pressure-time action sequence and feeds back several evaluation indicators based on a preset evaluation indicator set. The evaluation indicators are used to reflect whether there is a conflict in the displacement / pressure-time action sequence output by the learning reality loop agent. If there is no conflict in the displacement / pressure-time action sequence, the actual action strategy is given directly. If there is a conflict in the displacement / pressure-time action sequence, the evaluation indicators will be used as part of the reward function of the learning dual-loop agent to evaluate the value of the state under the scene features. S5, the learning-type reality loop agent's input, namely the scene features and the output displacement / stress-time action sequence, serves as the input to the learning-type dual-loop agent. The learning-type dual-loop agent corrects the displacement / stress-time action sequence under the scene features to obtain the corrected displacement / stress-time action sequence with the best value under the scene features. The corrected displacement / stress-time action sequence is then returned to the learning-type simulation loop agent to verify whether there are any conflicts. If the corrected displacement / stress-time action sequence has no conflicts, the actual action strategy is directly given. If the corrected displacement / stress-time action sequence has conflicts, the corrected displacement / stress-time is fed back to the learning loop dual-loop agent for correction. This process is iterated until the corrected displacement / stress-time action sequence has no conflicts. Finally, the actual action strategy is expanded to a preset strategy set to train a higher-quality learning-type reality loop agent.

2. The multi-agent autonomous learning method for hydraulic supports according to claim 1, characterized in that, The matching of scene features with a preset strategy set includes: The search algorithm is used to search the preset strategy set.

3. The multi-agent autonomous learning method for hydraulic supports according to claim 1, characterized in that, The learning-based reality loop agent includes a learning-based reality loop agent for hydraulic supports. Before step S3, the process further includes: constructing a learning-based reality loop agent for the hydraulic supports, specifically including: First, based on the action characteristics of the four actions of hydraulic supports in the well, namely lowering the support column, moving the support, raising the support column, and pushing the slide, a learning-type reality loop of hydraulic supports is constructed. The output of the agent is the displacement-time action sequence and the pressure-time action sequence. Secondly, the input of the learning-type reality environment agent for the hydraulic support is constructed based on environmental physical constraints, action target constraints, equipment coordination constraints, and the frame's action constraints. Next, a loss function for the learning-type reality loop agent of the hydraulic support is constructed based on hardware limit constraints, motion target constraints, and action timing constraints. Finally, the learning reality loop agent for the hydraulic support is trained based on the output, input, and loss function of the learning reality loop agent for the hydraulic support.

4. The multi-agent autonomous learning method for hydraulic supports according to claim 3, characterized in that, The inputs to the learning reality loop agent for constructing hydraulic supports include: Regarding environmental physical constraints, the angle of the hydraulic support base and the materials of the top and bottom plates are specifically quantified into mathematical indicators, and the angle of the hydraulic support base is set as follows: The coefficient of friction of the base plate is used as an indicator of the base plate material. The compressive strength of the roof slab is used as an indicator of the roof slab material. ; Regarding the action target constraints, the target displacement values ​​for shifting the frame and pushing the slide are set as follows: and The target minimum pressure for lowering the column and the target initial support force for raising the column are set as follows: and ; Regarding equipment coordination constraints, the traversal direction of the coal mining machine is set as either forward or reverse, quantified into mathematical indicators represented as 0 and 1 respectively; whether the action state of the adjacent support reaches the state where this support is to take action is divided into 0 and 1; specifically, because when the coal mining machine passes by, the hydraulic support performs four actions in sequence: lowering the column, moving the support, raising the column, and pushing the conveyor. After the adjacent support completes the three actions of lowering the column, moving the support, and raising the column, this support takes action. Therefore, when the adjacent support completes the three actions of lowering the column, moving the support, and raising the column, the state is divided into 0, and otherwise the state is divided into 1. Regarding the motion constraints of this frame, the first two frames of pressure and displacement are used as input, denoted as... and .

5. The multi-agent autonomous learning method for hydraulic supports according to claim 3 or 4, characterized in that, When constructing the loss function for the learning-based reality loop agent of the hydraulic support, the following are included: Let m be the loss function of the learning-based reality loop agent for the hydraulic support. Then: Regarding hardware limit constraints, the real-time speeds for lowering the column, moving the frame, raising the column, and pushing the conveyor are set as follows: Its corresponding single-action limit speed is If the hardware limits violate real-world rules, a corresponding loss is given, expressed as: ; ; ; ; Regarding the constraints on the motion target, the displacement value of the moving frame is set to... The displacement value of the pusher is The minimum pressure of the column is The initial support force of the lifting column is The target displacement values ​​for the moving frame and the pushing slide are respectively and The target minimum pressure for the column is The initial support force of the rising column is The loss is determined by the ratio of the difference between the real-time value and the target value, expressed as: ; ; ; ; Regarding action timing constraints, if the traversal direction is followed by an error, then... If it follows the correct direction but conflicts with the adjacent frame (i.e., the adjacent frame's action status is 1), this frame begins to move and is set. .

6. The multi-agent autonomous learning method for hydraulic supports according to claim 1, characterized in that, S4 is preceded by the construction of a learning simulation loop Agent, which specifically includes the construction of a hydraulic system mathematical model, a hydraulic support mathematical model, and a support-surrounding rock coupling model. The mathematical model of the hydraulic system is expressed as follows: ; ; ; in, This refers to the total liquid supply flow rate of the emulsion pump. The actual flow rate consumed by the actuator. This refers to the system's leaked traffic. For the speed of the actuator's operation, The area of ​​the liquid inlet chamber of the actuator; To supply fluid pressure to the hydraulic system, For the load of the actuator, For the back pressure of the liquid outlet chamber of the actuator, The area of ​​the liquid outlet chamber of the actuator; The mathematical model of the hydraulic support includes a column raising model, a column lowering model, a support shifting model, and a sliding model: The rising column model is represented as: ; in, For the lifting column thrust, To supply fluid pressure to the hydraulic system, The area of ​​the rodless cavity of the column is given. Let be the area of ​​the cavity of the column with rods. This is the back pressure of the liquid outlet chamber; The descending column model is represented as: ; in, To reduce column force; The moving frame model is represented as: ; in, For the shifting thrust, This indicates the area of ​​the rodless cavity of the jack. This indicates the area of ​​the rod cavity in the jack; The push-pull model is represented as: ; in, For pushing and pulling forces; The coupling model between the support and the surrounding rock includes a stiffness coupling model, a strength coupling model, and a stability coupling model: The stiffness coupling model is characterized by the compaction stiffness of the top plate and is expressed as follows: ; in, This refers to the compaction strength of the top slab on the support after the frame is moved. This is the initial support force for the lifting column after the frame is moved. This represents the total displacement of the moving frame; The strength coupling model is characterized by the dynamic resistance of the surrounding rock and is expressed as follows: ; in, The dynamic resistance of the surrounding rock during frame shifting / pushing. The drag-rate coefficient, For displacement rate; The stability coupling model, characterized by the risk of roof collapse, is expressed as follows: ; in, This represents the displacement increment during frame relocation. This is the time increment during the transfer of the rack. This represents the maximum moving rate.

7. The multi-agent autonomous learning method for hydraulic supports according to claim 6, characterized in that, After constructing the mathematical model of the hydraulic system, the mathematical model of the hydraulic support, and the coupling model between the support and the surrounding rock, the following is also included: Step 1: Verify the action speed based on the mathematical model of the hydraulic system, specifically including: Action speed verification: If the displacement-time action sequence generated by the learning simulation loop agent is calculated... This indicates that the action speed is too fast, exceeding the maximum total liquid supply flow rate of the emulsion pump. The speed needs to be reduced; Step 2: Verification based on the mathematical model of the hydraulic support, specifically including: Columnar pressure verification: If the columnar pressure is extracted based on the pressure-time action sequence generated by the learning simulation loop agent... Calculated Exceeding the rated working resistance of the column If so, the pressure is determined to be too high, including the lifting column pressure. This represents the initial support force during the lifting of the column; Displacement verification: If the velocity calculated from the displacement-time action sequence generated by the learning simulation loop agent is not less than the safe rate of the frame movement. and the safe speed of pushing and sliding If the displacement is too large, then it is determined that: ; ; in, and These are the maximum allowable flow rates for moving the frame and pushing the conveyor, respectively. Step 3: Validation based on the coupling model between the support and the surrounding rock, specifically including: (1) Stiffness coupling verification, specifically including: If the stress-time action sequence generated by the learning simulation loop agent is calculated... If the value is less than the empirical value, it indicates that the top plate compaction stiffness is insufficient, and the initial support force of the lifting column needs to be increased. (2) Strength coupling verification, specifically including: If the stress-time action sequence generated by the learning simulation loop agent is calculated... If the dynamic resistance of the surrounding rock exceeds the bearing capacity of the support, the operating rate needs to be reduced; among which, This indicates the upper limit of the dynamic resistance of the surrounding rock that the hydraulic support can withstand under rated operating conditions. The rated working pressure of the column is given by k, which is the mapping coefficient between the dynamic resistance of the surrounding rock and the pressure of the column. (3) Verification of roof collapse risk, specifically including: If the displacement-time action sequence generated by the learning simulation loop agent is calculated... Therefore, it is determined that the mutation rate needs to be reduced and the mutation range expanded. ; The learning simulation loop agent judges whether the action rate of the displacement / pressure-time action sequence is too fast and whether the initial support force of the lifting column needs to be adjusted based on the verification results of the first to third steps, forming an evaluation index set [Y,Z]. If the action rate is too fast, Y=1, otherwise it is 0; if the initial support force of the lifting column needs to be adjusted, Z=1, otherwise it is 0.

8. The multi-agent autonomous learning method for hydraulic supports according to claim 1, characterized in that, The S5 preceding S5 also includes constructing a learning-based dual-loop agent, specifically including: First, the input and output of the learning reality loop agent are constructed as the state space of the learning dual-loop agent; Secondly, the displacement / pressure-time action sequence at a specified time is set as the action space of the learning dual-loop agent; Next, the reward function of the learning dual-loop agent is determined based on hardware limit constraints, motion target constraints, action timing constraints, and simulation feedback constraints. Finally, the learning double-loop agent is trained based on the state space, action space, and reward function of the learning double-loop agent.

9. The multi-agent autonomous learning method for hydraulic supports according to claim 8, characterized in that, Let the reward function of the learning-type two-loop agent be r. , but: Set the real-time speeds for lowering the column, moving the support frame, raising the column, and pushing the conveyor to be as follows: Its corresponding single-action limit speed is If hardware limitations violate real-world rules, a corresponding penalty will be imposed, expressed as follows: ; ; ; ; Regarding the constraints on the motion target, the displacement value of the moving frame is set to... The displacement value of the pusher is The minimum pressure of the column is The initial support force of the lifting column is The target displacement values ​​for the moving frame and the pushing slide are respectively and The target minimum pressure for column reduction is The initial support force of the rising column is The penalty is determined based on the ratio of the difference between the real-time value and the target value, expressed as: ; ; ; ; Regarding action timing constraints, if the traversal direction is followed by an error, then... ; If it follows the correct direction but conflicts with the adjacent frame (i.e., the adjacent frame's action state is 1), this frame begins to move and is set. ; Regarding the simulation feedback constraints, the set of evaluation metrics given by the learning simulation loop agent is set as follows: ,if ,but ;if =1, then ;if ,but ;if ,but .

10. The multi-agent autonomous learning method for hydraulic supports according to claim 8 or 9, characterized in that, The learning dual-loop agent is trained using the DDPG algorithm based on the state space, action space, and reward function of the learning dual-loop agent.

Citation Information

Patent Citations

  • Virtual fully mechanized mining production system deduction method based on multi-agent deep reinforcement learning

    CN114329936A

  • Neural task planner for autonomous vehicles

    US20210223774A1