Hydraulic support multi-agent autonomous learning method

By employing a multi-agent autonomous learning method for hydraulic supports, and combining real and simulated data, the problems of complex geological environments and diverse equipment control in coal mining have been solved, enabling efficient and safe intelligent decision-making across all scenarios and improving the level of intelligence in coal mines.

CN121543624BActive Publication Date: 2026-05-26TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TAIYUAN UNIVERSITY OF TECHNOLOGY
Filing Date
2025-11-19
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing coal mining technologies face challenges such as complex geological environments, diverse equipment control strategies, high levels of personnel involvement, low production efficiency, and a lack of closed-loop data fusion analysis combining real and simulated data.

Method used

A multi-agent autonomous learning method for hydraulic supports is adopted. By pre-deploying learning-type reality loop agents, learning-type simulation loop agents, and learning-type dual-loop agents, scene feature matching, action sequence generation, conflict detection and correction are achieved, thus constructing a closed-loop data fusion analysis system.

Benefits of technology

It has achieved more precise action sequence generation, improved the leap from single device control to full-scenario intelligent decision-making system, and built a robust, highly efficient, safe, redundant, and knowledge-accumulating intelligent control system for hydraulic supports, supporting intelligent coal mines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121543624B_ABST
    Figure CN121543624B_ABST
Patent Text Reader

Abstract

This invention relates to a multi-agent autonomous learning method for hydraulic supports, belonging to the field of intelligent coal mine technology. It includes: using the scene features input by a learning-type real-loop agent and the output displacement / pressure-time action sequence as input to a learning-type dual-loop agent; the learning-type dual-loop agent corrects the displacement / pressure-time action sequence to obtain a corrected displacement / pressure-time action sequence; the corrected displacement / pressure-time action sequence is then returned to the learning-type simulation loop agent to verify for conflicts; if conflicts are found after correction, the correction continues, iterating until the corrected displacement / pressure-time action sequence is conflict-free; finally, the final actual action strategy is expanded to a preset strategy set for training a higher-quality learning-type real-loop agent. This invention achieves closed-loop data fusion analysis that combines real data and simulated data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent coal mining technology, and in particular to a multi-agent autonomous learning method for hydraulic supports. Background Technology

[0002] Coal mining, as a crucial component of energy supply, holds significant economic and social importance. However, traditional coal mining processes face numerous challenges, including complex geological environments, diverse equipment control strategies, high levels of human involvement, and low production efficiency. In the current era of artificial intelligence, promoting the intelligent transformation of coal mining and enhancing the autonomy, collaboration, adaptability, and safety of mining operations has become key to overcoming these current difficulties.

[0003] Current intelligent systems for coal mining faces limitations such as difficulty in multi-equipment coordination, poor adaptability to dynamic scenarios, low knowledge reusability, and high workload for personnel. Existing technologies often only address one aspect, relying on historical data from coal production to develop agent models or on simulation environments developed based on real-world rules and coal mining processes to generate simulated data. They lack closed-loop data fusion analysis that combines real and simulated data. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a multi-agent autonomous learning method for hydraulic supports. The technical solution of this invention is as follows:

[0005] A multi-agent autonomous learning method for hydraulic supports, based on pre-deployed learning-type reality loop agents, learning-type simulation loop agents, and learning-type dual-loop agents, includes:

[0006] S1, Obtain scene features and match scene features with a preset strategy set;

[0007] S2, If there is a scene in the preset strategy set that matches the scene features, then output the displacement / pressure-time action sequence corresponding to the scene features in the preset strategy set;

[0008] S3, if there is no scene matching the scene features in the preset strategy set, then the scene features are input into the learning reality loop agent, the learning reality loop agent outputs the displacement / pressure-time action sequence under the scene features, and the displacement / pressure-time action sequence output by the learning reality loop agent is used as the input of the learning simulation loop agent;

[0009] S4, the learning simulation loop agent detects the displacement / pressure-time action sequence and feeds back several evaluation indicators based on a preset evaluation indicator set. The evaluation indicators are used to reflect whether there is a conflict in the displacement / pressure-time action sequence output by the learning reality loop agent. If there is no conflict in the displacement / pressure-time action sequence, the actual action strategy is given directly. If there is a conflict in the displacement / pressure-time action sequence, the evaluation indicators will be used as part of the reward function of the learning dual-loop agent to evaluate the value of the state under the scene features.

[0010] S5, the input to the learning-type reality loop agent, namely the scene features and the output displacement / stress-time action sequence, is used as the input to the learning-type dual-loop agent. The learning-type dual-loop agent corrects the displacement / stress-time action sequence under the scene features to obtain the corrected displacement / stress-time action sequence with the best value under the scene features. The corrected displacement / stress-time action sequence is then returned to the learning-type simulation loop agent to verify whether there is a conflict. If the corrected displacement / stress-time action sequence has no conflict, the actual action strategy is directly given. If the corrected displacement / stress-time action sequence has a conflict, the corrected displacement / stress-time action sequence is fed back to the learning-type dual-loop agent for correction. This process is repeated until the corrected displacement / stress-time action sequence has no conflict. Finally, the actual action strategy is expanded to a preset strategy set to train a higher-quality learning-type reality loop agent.

[0011] Preferably, the matching of scene features with a preset strategy set includes:

[0012] The search algorithm is used to search the preset strategy set.

[0013] Preferably, the learning reality loop agent includes a learning reality loop agent for a hydraulic support, and before step S3, it further includes: constructing a learning reality loop agent for the hydraulic support, specifically including:

[0014] First, based on the action characteristics of the four actions of hydraulic supports in the well, namely lowering the support column, moving the support, raising the support column, and pushing the slide, a learning-type reality loop of hydraulic supports is constructed. The output of the agent is the displacement-time action sequence and the pressure-time action sequence.

[0015] Secondly, the input of the learning-type reality environment agent for the hydraulic support is constructed based on environmental physical constraints, action target constraints, equipment coordination constraints, and the frame's action constraints.

[0016] Next, a loss function for the learning-type reality loop agent of the hydraulic support is constructed based on hardware limit constraints, motion target constraints, and action timing constraints.

[0017] Finally, the learning reality loop agent for the hydraulic support is trained based on the output, input, and loss function of the learning reality loop agent for the hydraulic support.

[0018] Preferably, the input to the learning reality loop agent for constructing the hydraulic support includes:

[0019] Regarding environmental physical constraints, the angle of the hydraulic support base and the materials of the top and bottom plates are specifically quantified into mathematical indicators, and the angle of the hydraulic support base is set as follows: The coefficient of friction of the base plate is used as an indicator of the base plate material. The compressive strength of the roof slab is used as an indicator of the roof slab material. ;

[0020] Regarding the action target constraints, the target displacement values ​​for shifting the frame and pushing the slide are set as follows: and The target minimum pressure for lowering the column and the target initial support force for raising the column are set as follows: and ;

[0021] Regarding equipment coordination constraints, the traversal direction of the coal mining machine is set as either forward or reverse, quantified into mathematical indicators represented as 0 and 1 respectively; whether the action state of the adjacent support reaches the state where this support is to take action is divided into 0 and 1; specifically, because when the coal mining machine passes by, the hydraulic support performs four actions in sequence: lowering the column, moving the support, raising the column, and pushing the conveyor. After the adjacent support completes the three actions of lowering the column, moving the support, and raising the column, this support takes action. Therefore, when the adjacent support completes the three actions of lowering the column, moving the support, and raising the column, the state is divided into 0, and otherwise the state is divided into 1.

[0022] Regarding the motion constraints of this frame, the first two frames of pressure and displacement are used as input, denoted as... and .

[0023] Preferably, the loss function for constructing the learning-based reality loop agent of the hydraulic support includes:

[0024] Let m be the loss function of the learning-based reality loop agent for the hydraulic support. Then:

[0025] Regarding hardware limit constraints, the real-time speeds for lowering the column, moving the frame, raising the column, and pushing the conveyor are set as follows: Its corresponding single-action limit speed is If the hardware limits violate real-world rules, a corresponding loss is given, expressed as:

[0026] ;

[0027] ;

[0028] ;

[0029] ;

[0030] Regarding the constraints on the motion target, the displacement value of the moving frame is set to... The displacement value of the pusher is The minimum pressure of the column is The initial support force of the lifting column is The target displacement values ​​for the moving frame and the pushing slide are respectively and The target minimum pressure for the column is The initial support force of the rising column is The loss is determined by the ratio of the difference between the real-time value and the target value, expressed as:

[0031] ;

[0032] ;

[0033] ;

[0034] ;

[0035] Regarding action timing constraints, if the traversal direction is followed by an error, then... If it follows the correct direction but conflicts with the adjacent frame (i.e., the adjacent frame's action status is 1), this frame begins to move and is set. .

[0036] Preferably, before step S4, the method further includes constructing a learning simulation loop agent, specifically including constructing a hydraulic system mathematical model, a hydraulic support mathematical model, and a support-surrounding rock coupling model;

[0037] The mathematical model of the hydraulic system is expressed as follows:

[0038] ;

[0039] ;

[0040] ;

[0041] in, This refers to the total fluid supply flow rate of the emulsion pump. The actual flow rate consumed by the actuator. This refers to the system's leaked traffic. For the speed of the actuator's operation, The area of ​​the liquid inlet chamber of the actuator; To supply fluid pressure to the hydraulic system, For the load of the actuator, For the back pressure of the liquid outlet chamber of the actuator, The area of ​​the liquid outlet chamber of the actuator;

[0042] The mathematical model of the hydraulic support includes a column raising model, a column lowering model, a support shifting model, and a sliding model:

[0043] The rising column model is represented as:

[0044] ;

[0045] in, For the lifting column thrust, To supply fluid pressure to the hydraulic system, The area of ​​the rodless cavity of the column is given. Let be the area of ​​the cavity of the column with rods. This is the back pressure of the liquid outlet chamber;

[0046] The descending column model is represented as:

[0047] ;

[0048] in, To reduce column force;

[0049] The moving frame model is represented as:

[0050] ;

[0051] in, For the shifting thrust, This indicates the area of ​​the rodless cavity of the jack. This indicates the area of ​​the rod cavity in the jack;

[0052] The push-pull model is represented as:

[0053] ;

[0054] in, For pushing and pulling forces;

[0055] The coupling model between the support and the surrounding rock includes a stiffness coupling model, a strength coupling model, and a stability coupling model:

[0056] The stiffness coupling model is characterized by the compaction stiffness of the top plate and is expressed as follows:

[0057] ;

[0058] in, This refers to the compaction strength of the top slab on the support after the frame is moved. This is the initial support force for the lifting column after the frame is moved. This represents the total displacement of the moving frame;

[0059] The strength coupling model is characterized by the dynamic resistance of the surrounding rock and is expressed as follows:

[0060] ;

[0061] in, The dynamic resistance of the surrounding rock during frame shifting / pushing. The drag-rate coefficient, For displacement rate;

[0062] The stability coupling model, characterized by the risk of roof collapse, is expressed as follows:

[0063] ;

[0064] in, This represents the displacement increment during frame relocation. This is the time increment during the transfer of the rack. This represents the maximum moving rate.

[0065] Preferably, after constructing the mathematical model of the hydraulic system, the mathematical model of the hydraulic support, and the coupling model between the support and the surrounding rock, the method further includes:

[0066] Step 1: Verify the action speed based on the mathematical model of the hydraulic system, specifically including:

[0067] Action speed verification: If the displacement-time action sequence generated by the learning simulation loop agent is calculated... This indicates that the action speed is too fast, exceeding the maximum total liquid supply flow rate of the emulsion pump. The speed needs to be reduced;

[0068] Step 2: Verification based on the mathematical model of the hydraulic support, specifically including:

[0069] Columnar pressure verification: If the columnar pressure is extracted based on the pressure-time action sequence generated by the learning simulation loop agent... Calculated Exceeding the rated working resistance of the column If so, the pressure is determined to be too high, including the lifting column pressure. This represents the initial support force during the lifting of the column;

[0070] Displacement verification: If the velocity calculated from the displacement-time action sequence generated by the learning simulation loop agent is not less than the safe rate of the frame movement. and the safe speed of pushing If the displacement is too large, then it is determined that:

[0071] ;

[0072] ;

[0073] in, and These are the maximum allowable flow rates for moving the frame and pushing the conveyor, respectively.

[0074] Step 3: Validation based on the coupling model between the support and the surrounding rock, specifically including:

[0075] (1) Stiffness coupling verification, specifically including:

[0076] If the stress-time action sequence generated by the learning simulation loop agent is calculated... If the value is less than the empirical value, it indicates that the top plate compaction stiffness is insufficient, and the initial support force of the lifting column needs to be increased.

[0077] (2) Strength coupling verification, specifically including:

[0078] If the stress-time action sequence generated by the learning simulation loop agent is calculated... If the dynamic resistance of the surrounding rock exceeds the bearing capacity of the support, the operating rate needs to be reduced; among which, This indicates the upper limit of the dynamic resistance of the surrounding rock that the hydraulic support can withstand under rated operating conditions. The rated working pressure of the column is given by k, which is the mapping coefficient between the dynamic resistance of the surrounding rock and the pressure of the column.

[0079] (3) Verification of roof collapse risk, specifically including:

[0080] If the displacement-time action sequence generated by the learning simulation loop agent is calculated... Therefore, it is determined that the mutation rate needs to be reduced and the mutation range expanded. ;

[0081] The learning simulation loop agent judges whether the action rate of the displacement / pressure-time action sequence is too fast and whether the initial support force of the lifting column needs to be adjusted based on the verification results of the first to third steps, forming an evaluation index set [Y,Z]. If the action rate is too fast, Y=1, otherwise it is 0; if the initial support force of the lifting column needs to be adjusted, Z=1, otherwise it is 0.

[0082] Preferably, before step S5, the method further includes constructing a learning-based dual-loop agent, specifically including:

[0083] First, the input and output of the learning reality loop agent are constructed as the state space of the learning dual-loop agent;

[0084] Secondly, the displacement / pressure-time action sequence at a specified time is set as the action space of the learning dual-loop agent;

[0085] Next, the reward function of the learning dual-loop agent is determined based on hardware limit constraints, motion target constraints, action timing constraints, and simulation feedback constraints.

[0086] Finally, the learning double-loop agent is trained based on the state space, action space, and reward function of the learning double-loop agent.

[0087] Preferably, the reward function of the learning-type dual-loop agent is set to r. , but:

[0088] Set the real-time speeds for lowering the column, moving the support frame, raising the column, and pushing the conveyor to be as follows: Its corresponding single-action limit speed is If hardware limitations violate real-world rules, a corresponding penalty will be imposed, expressed as follows:

[0089] ;

[0090] ;

[0091] ;

[0092] ;

[0093] Regarding the constraints on the motion target, the displacement value of the moving frame is set to... The displacement value of the pusher is The minimum pressure of the column is The initial support force of the lifting column is The target displacement values ​​for the moving frame and the pushing slide are respectively and The target minimum pressure for column reduction is The initial support force of the rising column is The penalty is determined based on the ratio of the difference between the real-time value and the target value, expressed as:

[0094] ;

[0095] ;

[0096] ;

[0097] ;

[0098] Regarding action timing constraints, if the traversal direction is followed by an error, then... If it follows the correct direction but conflicts with the adjacent frame (i.e., the adjacent frame's action status is 1), this frame begins to move and is set. ;

[0099] Regarding the simulation feedback constraints, the set of evaluation metrics given by the learning simulation loop agent is set as follows: ,if ,but ;if =1, then ;if ,but ;if ,but .

[0100] Preferably, the learning dual-loop agent is trained using the DDPG algorithm based on the state space, action space, and reward function of the learning dual-loop agent.

[0101] All of the above-mentioned optional technical solutions can be combined arbitrarily, and the present invention will not provide a detailed description of the structure after each combination.

[0102] By means of the above solution, the beneficial effects of the present invention are as follows:

[0103] By pre-deploying a learning-based reality loop agent, a learning-based simulation loop agent, and a learning-based dual-loop agent, the learning-based reality loop agent can generate displacement / pressure-time action sequences based on real-world patterns according to the current scene characteristics of the coal mining face. When the learning-based simulation loop agent detects conflicts in the displacement / pressure-time action sequences under the scene characteristics, the learning-based dual-loop agent corrects the displacement / pressure-time action sequences under the scene characteristics until there are no conflicts. The final actual action strategy is then expanded to the preset strategy set to train a higher-quality learning-based reality loop agent. This achieves closed-loop data fusion analysis that combines real data with simulated data, thereby ensuring the construction of a more accurate learning-based reality loop agent.

[0104] The autonomous learning method of the multi-agent closed-loop architecture is not only limited to generating more accurate action sequences, but also more conducive to realizing the leap from single device control to full-scenario intelligent decision-making system. Through the deep integration of real and simulated data, it not only retains the physical authenticity of the real scene, but also breaks through the sample limitations of real data through simulation. Finally, it builds a new generation of intelligent control system for hydraulic supports with strong robustness, high iteration efficiency, safety redundancy, knowledge accumulation and autonomous decision-making, thus providing core technical support for the intelligentization of coal mines.

[0105] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, the preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings. Attached Figure Description

[0106] Figure 1 This is a flowchart of the multi-agent autonomous learning method for hydraulic supports provided in an embodiment of the present invention.

[0107] Figure 2 This is a flowchart of the learning reality loop agent for constructing the hydraulic support in an embodiment of the present invention.

[0108] Figure 3 This is a flowchart of constructing a learning simulation loop agent in an embodiment of the present invention.

[0109] Figure 4 This is a flowchart of constructing a learning-type dual-loop Agent for hydraulic supports in an embodiment of the present invention. Detailed Implementation

[0110] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0111] like Figure 1 As shown, the multi-agent autonomous learning method for hydraulic supports provided in this embodiment of the invention is based on pre-deployed learning-type reality loop agents, learning-type simulation loop agents, and learning-type dual-loop agents. The learning-type reality loop agent is an automatic integration-analysis-modeling algorithm for successful strategies of operation records and working condition data from the actual coal mining face production process. It is used to develop experience data from the real coal mining process, thereby generating action sequences based on real-world patterns from coal mining site data. The learning-type simulation loop agent is an automatic generation-simulation-optimization-induction algorithm for feasible strategies of coal mining face production operation in the simulation system, realizing the exploration of simulated data for the virtual coal mining process. The learning-type dual-loop agent is a data and control fusion analysis algorithm between the real and simulated systems, as well as a virtual-real strategy credibility evaluation algorithm, used to continuously correct and improve the models of each agent. Based on this, the multi-agent autonomous learning method for hydraulic supports includes the following steps S1 to S5:

[0112] S1, acquire scene features and match scene features with a preset strategy set.

[0113] Specifically, when matching scene features with a preset strategy set, S1 can use a search algorithm such as exhaustive search to search the preset strategy set. The preset strategy set includes the correspondence between several scene features and displacement / pressure-time action sequences. The displacement / pressure-time action sequences include displacement-time action sequences and pressure-time action sequences.

[0114] S2, If there is a scene in the preset strategy set that matches the scene features, then output the displacement / pressure-time action sequence corresponding to the scene features in the preset strategy set.

[0115] S3. If there is no scene matching the scene features in the preset strategy set, the scene features are input into the learning reality loop agent. The learning reality loop agent outputs the displacement / pressure-time action sequence under the scene features, and the displacement / pressure-time action sequence output by the learning reality loop agent is used as the input of the learning simulation loop agent.

[0116] It should be noted that before inputting scene features into the learning reality loop agent, the learning reality loop agent needs to be constructed first. The learning reality loop agent is essentially a generator based on a diffusion model. The idea is to construct an action generator from existing reality data using the diffusion model approach, which can generate action sequences based on reality patterns. In this embodiment of the invention, an action generator for hydraulic supports is constructed based on the coal mining face, i.e., a learning reality loop agent for hydraulic supports. For example... Figure 2 As shown, building a learning reality loop agent for hydraulic supports requires three steps: setting the input, setting the loss function, and setting the output.

[0117] First, the output settings for the learning-type reality loop agent of the hydraulic support are defined. The output of the learning-type reality loop agent of the hydraulic support is based on the actions of the hydraulic support, which mainly perform four actions downhole: lowering the support column, moving the support, raising the support column, and pushing the conveyor. These four actions are reflected on the sensors as a sequence of stroke changes over time and a curve of support column pressure changes over time. Therefore, the output is set as a displacement-time action sequence and a pressure-time action sequence. In this embodiment, the time steps of the displacement-time action sequence and the pressure-time action sequence are consistent and the number of steps is the same, i.e., displacement and pressure action sequences generated at the same time. Therefore, the output settings for the learning-type reality loop agent of the hydraulic support can be summarized as follows: based on the action characteristics of the four actions performed by the hydraulic support downhole—lowering the support column, moving the support, raising the support column, and pushing the conveyor—the output of the learning-type reality loop agent of the hydraulic support is constructed as a displacement-time action sequence and a pressure-time action sequence.

[0118] Secondly, the inputs to the learning reality loop agent of the hydraulic support are set. The input of the learning reality loop agent needs to consider the preceding states of the output, that is, it needs to consider why such a sequence of actions is generated. Therefore, in this embodiment of the invention, the input of the learning reality loop agent of the hydraulic support is constructed based on environmental physical constraints, action target constraints, equipment coordination constraints, and the action constraints of the hydraulic support itself.

[0119] Regarding environmental physical constraints, considering the physical factors of the environment, since the base angle of the hydraulic support determines its travel angle, and the material of the top and bottom plates determines the material of the environment in which the hydraulic support is located, this embodiment of the invention treats the base angle and the material of the top and bottom plates as variables. In conjunction with the above, this embodiment of the invention, in terms of environmental physical constraints, quantifies the base angle and the material of the top and bottom plates of the hydraulic support into mathematical indicators, setting the base angle of the hydraulic support as... The coefficient of friction of the base plate is used as an indicator of the base plate material. The compressive strength of the roof slab is used as an indicator of the roof slab material. .

[0120] Regarding action target constraints, based on the actions themselves, hydraulic supports require action target values ​​to perform actions. The moving and pushing actions require corresponding target displacement values. For example, the maximum value for moving and pushing is the cutting depth of the coal mining machine (865mm). However, the actual situation needs to be dynamically adjusted based on reality. For instance, if the value of the stroke sensor before moving is 800mm, then the target displacement value for a single moving action can only be 800mm, and needs to be dynamically changed based on the actual value. The target displacement value for pushing needs to be dynamically changed based on the stroke value after the moving action ends. For example, if the stroke value after the moving action ends is 50mm, then the target displacement value for pushing is the cutting depth of the coal mining machine minus the stroke value after the moving action ends, which is 815mm. Lowering the column is the process of the column detaching from the roof, with pressure decreasing from high to low. The target minimum pressure is when it drops to near 0. Raising the column needs to reach the range of the rated initial support force required by the hydraulic support, such as 80% of 31.5Mpa to the maximum, i.e., 24-31.5Mpa. In conjunction with the above, in terms of action target constraints, this embodiment of the invention sets the target displacement values ​​for frame shifting and pushing as follows: and The target minimum pressure for lowering the column and the target initial support force for raising the column are set as follows: and .

[0121] Regarding equipment coordination constraints, since hydraulic supports do not operate independently in the coal mining face, but rather form a cluster of hydraulic supports that coordinate their actions according to certain rules, it is necessary to consider the action status of adjacent supports to determine whether the current support can take action. Whether to refer to the status of the support on the left or right depends on the traversal direction of the coal mining machine. Based on the above, this embodiment of the invention sets the traversal direction of the coal mining machine as either forward or reverse, quantified as mathematical indices represented by 0 and 1 respectively. It also classifies whether the action status of an adjacent support reaches the state where the current support can take action as 0 or 1. Specifically, as the coal mining machine passes by, the hydraulic supports sequentially perform four actions: lowering the support, moving the support, raising the support, and pushing the conveyor. After an adjacent support completes the three actions of lowering, moving, and raising the support, the current support can take action. Therefore, if an adjacent support completes the three actions of lowering, moving, and raising the support, the state is classified as 0; otherwise, the state is classified as 1.

[0122] Regarding the motion constraints of this frame, since the stroke and pressure values ​​are not constant, the first two frames of pressure and displacement need to be used as input. Based on the above, in this embodiment of the invention, the first two frames of pressure and displacement are used as input for the motion constraints of this frame, denoted as... and .

[0123] Next, a loss function is set for the learning-based reality loop agent of the hydraulic support. Specifically, when setting the loss function, relevant reality constraints need to be incorporated to guide the output action sequence to not violate reality rules. Considering hardware limits, such as the rate of change of column lowering / raising pressure not exceeding the cylinder limit, and the rate of support shifting / pushing speed not exceeding the cylinder limit, penalties are incurred if these limits are exceeded. Regarding the target displacement of support shifting / pushing, support shifting is determined based on the previous two frames of motion. For example, if the stroke value before support shifting is 800, then the target displacement value is 800. When the action is completed, the further the shift value deviates from 800, the higher the penalty. For pushing, the target displacement value is the difference between the coal mining machine's cutting depth and the actual value reached by support shifting; the greater the deviation from the target displacement value, the higher the penalty. Regarding the target for lowering the support column, if the pressure is within a reasonable range at the end of the lowering process, there is no penalty. If the pressure does not reach the reasonable range, a penalty is imposed based on the deviation between the actual pressure value and the target pressure value. For raising the support column, the focus is on whether the initial support force reaches a reasonable range. If it does not, a penalty is imposed based on the deviation from the target initial support force range. Regarding the timing constraints, if the coal mining machine traversal sequence is 0, the left adjacent support takes priority, and vice versa. If the sequence is correct and the adjacent support action is 0, the current support can begin action; otherwise, a severe penalty is imposed if the timing constraints are violated. In conjunction with this part, this embodiment of the invention constructs a loss function for the learning-type reality loop agent of the hydraulic support based on hardware limit constraints, motion target constraints, and action timing constraints.

[0124] In one specific embodiment, the loss function for constructing the learning reality loop agent of the hydraulic support includes:

[0125] Let m be the loss function of the learning-based reality loop agent for the hydraulic support. Then:

[0126] Regarding hardware limit constraints, the real-time speeds for lowering the column, moving the frame, raising the column, and pushing the conveyor are set as follows: Its corresponding single-action limit speed is If the hardware limits violate real-world rules, a corresponding loss is given, expressed as:

[0127] ;

[0128] ;

[0129] ;

[0130] ;

[0131] Regarding the constraints on the motion target, the displacement value of the moving frame is set to... The displacement value of the pusher is The minimum pressure of the column is The initial support force of the lifting column is The target displacement values ​​for the moving frame and the pushing slide are respectively and The target minimum pressure for the column is The initial support force of the rising column is According to real-time values With target value The loss is determined by the proportion of the difference, expressed as:

[0132] ;

[0133] ;

[0134] ;

[0135] ;

[0136] Regarding action timing constraints, if the traversal direction is followed by an error, then... If it follows the correct direction but conflicts with the adjacent frame (i.e., the adjacent frame's action status is 1), this frame begins to move and is set. .

[0137] Finally, the learning reality loop agent for the hydraulic support is trained based on the output, input, and loss function of the learning reality loop agent for the hydraulic support.

[0138] Thus, the learning-based reality loop agent is essentially set up based on a diffusion model generator. It can then extract and process real-world data to train itself. This extraction and processing of real-world data involves transforming the real-time collected data into a format that matches the input format of the learning-based reality loop agent. For example, if the input format of the learning-based reality loop agent is scene features, then the scene features will be extracted during the data processing.

[0139] S4, the learning simulation loop agent detects the displacement / pressure-time action sequence and feeds back several evaluation indicators based on a preset evaluation indicator set. The evaluation indicators are used to reflect whether there is a conflict in the displacement / pressure-time action sequence output by the learning reality loop agent. If there is no conflict in the displacement / pressure-time action sequence, the actual action strategy is given directly. If there is a conflict in the displacement / pressure-time action sequence, the evaluation indicators will be used as part of the reward function of the learning dual-loop agent to evaluate the value of the state under the scene features.

[0140] It should be noted that a learning simulation loop agent needs to be constructed before S4. Essentially, the learning simulation loop agent simplifies the complex and abstract real world into a mathematical model, sets relevant constraints, and simulates the action sequences generated by the learning reality loop agent. For example... Figure 3 As shown, this embodiment of the invention constructs a hydraulic system mathematical model, a hydraulic support mathematical model, and a support-surrounding rock coupling model based on the working principle of hydraulic supports, the characteristics of actuators, and the support-surrounding rock coupling theory. The models are logically coherent and their parameters are interconnected, as detailed below:

[0141] 1. Mathematical model of hydraulic system:

[0142] The core formulas of a hydraulic system include fluid supply flow balance, flow-speed relationship, and pressure-load balance. Therefore, the mathematical model of the hydraulic system is expressed as:

[0143] ;

[0144] ;

[0145] ;

[0146] in, This represents the total flow rate (L / min) of the emulsion pump. The actual flow rate consumed by the actuator (column / jack). System leakage flow rate (L / min); For the speed of the actuator's operation, The area of ​​the liquid inlet chamber of the actuator ; The hydraulic system is supplied with hydraulic pressure (MPa). For the load of the actuator (KN). The back pressure (MPa) of the liquid outlet chamber of the actuator. The area of ​​the liquid outlet chamber of the actuator.

[0147] Specifically, regarding Since the hydraulic support has four actions—lowering the column, raising the column, moving the support, and pushing the slide—let's assume the generator output sequence is a curve of displacement versus pressure over time. Moving the support and pushing the slide are performed by jacks as actuators. In this case... Regarding the lowering and raising of the column, the actuator is the column itself, which is reflected in the data as pressure changes over time. However, because this embodiment of the invention uses a simplified mathematical model, it is assumed that the height of each lowering and raising operation is the same. Then at this time At this time It is derived from the pressure-time change curve, specifically the duration of the lifting column's movement.

[0148] 2. Mathematical model of hydraulic support:

[0149] The column and jack are the "core actuators" of the hydraulic support's operation. Based on the steady-state force principle of a double-acting hydraulic cylinder, this invention establishes simplified mathematical models for the column to characterize the mechanical relationship between column raising / lowering and the jack to characterize the mechanical relationship between jack moving / pushing, for four actions: column raising, column lowering, support moving, and jack pushing. Specifically, by establishing mechanical formulas for the column raising, column lowering, support moving, and jack pushing models, the conversion relationship between "pressure → force" is quantified, as follows:

[0150] The rising column model is represented as:

[0151] ;

[0152] in, For the lifting column thrust, To supply fluid pressure to the hydraulic system, The area of ​​the rodless cavity of the column is given. Let be the area of ​​the cavity of the column with rods. This is the back pressure of the liquid outlet chamber;

[0153] The descending column model is represented as:

[0154] ;

[0155] in, To reduce the column force and unload the hydraulic support, the column lowering and raising chambers are switched.

[0156] The moving frame model is represented as:

[0157] ;

[0158] in, For the shifting thrust, This indicates the area of ​​the rodless cavity of the jack. This indicates the area of ​​the rod cavity in the jack;

[0159] The push-pull model is represented as:

[0160] ;

[0161] in, To achieve the pushing and pulling force, the resistance of the scraper conveyor must be overcome.

[0162] 3. Coupled model of support and surrounding rock:

[0163] The support-surrounding rock coupling model is the "core of system stability." This invention, based on the three theories of stiffness coupling, strength coupling, and stability coupling, combines displacement and pressure data output by the generator to quantify the adaptation relationship between "support action → surrounding rock response." The support-surrounding rock coupling model includes a stiffness coupling model, a strength coupling model, and a stability coupling model, as detailed below:

[0164] The stiffness coupling model is characterized by the compaction stiffness of the top plate and is expressed as follows:

[0165] ;

[0166] in, This refers to the compaction strength of the top slab on the support after the frame is moved. This is the initial support force for the lifting column after the frame is moved. This is the total displacement of the moving frame; if If the value is less than the engineering experience value, it indicates that the top slab is soft and the initial support force needs to be increased.

[0167] The strength coupling model is characterized by the dynamic resistance of the surrounding rock and is expressed as follows:

[0168] ;

[0169] in, The dynamic resistance of the surrounding rock during frame shifting / pushing. The drag-rate coefficient, For displacement rate;

[0170] The stability coupling model, characterized by the risk of roof collapse, is expressed as follows:

[0171] ;

[0172] in, This represents the displacement increment during frame relocation. This is the time increment during the transfer of the rack. This is the maximum moving speed. If Greater than If this is confirmed, there is a risk of roof collapse, and the scaffolding movement must be stopped immediately, the initial support force increased, and the roof reinforced.

[0173] At this point, a simplified mathematical model of the learning simulation loop agent has been constructed. The following details how to verify the action sequence based on this mathematical model:

[0174] Step 1: Verify the action speed based on the mathematical model of the hydraulic system, specifically including:

[0175] Action speed verification: If the displacement-time action sequence generated by the learning simulation loop agent is calculated... This indicates that the action speed is too fast, exceeding the maximum total liquid supply flow rate of the emulsion pump. The speed needs to be reduced.

[0176] Step 2: Verification based on the mathematical model of the hydraulic support, specifically including:

[0177] Columnar pressure verification: If the columnar pressure is extracted based on the pressure-time action sequence generated by the learning simulation loop agent... Calculated Exceeding the rated working resistance of the column If so, the pressure is determined to be too high, including the lifting column pressure. This represents the initial support force during the lifting of the column;

[0178] Displacement verification: If the velocity calculated from the displacement-time action sequence generated by the learning simulation loop agent is not less than the safe rate of the frame movement. and the safe speed of pushing If the displacement is too large, then it is determined that:

[0179] ;

[0180] ;

[0181] in, and These are the maximum allowable flow rates for moving the frame and pushing the conveyor, respectively. From this, the maximum allowable speeds for moving the frame and pushing the conveyor can be calculated.

[0182] Step 3: Validation based on the coupling model between the support and the surrounding rock, specifically including:

[0183] (1) Stiffness coupling verification, specifically including:

[0184] If the stress-time action sequence generated by the learning simulation loop agent is calculated... If the value is less than the empirical value, it indicates that the top plate compaction stiffness is insufficient, and the initial support force of the lifting column needs to be increased.

[0185] (2) Strength coupling verification, specifically including:

[0186] If the stress-time action sequence generated by the learning simulation loop agent is calculated... If the dynamic resistance of the surrounding rock exceeds the bearing capacity of the support, the operating rate needs to be reduced; among which, This indicates the upper limit of the dynamic resistance of the surrounding rock that the hydraulic support can withstand under rated operating conditions. The rated working pressure of the column is given by k, which is the mapping coefficient between the dynamic resistance of the surrounding rock and the pressure of the column.

[0187] (3) Verification of roof collapse risk, specifically including:

[0188] If the displacement-time action sequence generated by the learning simulation loop agent is calculated... Therefore, it is determined that the mutation rate needs to be reduced and the mutation range expanded. ;

[0189] The learning simulation loop agent judges whether the action rate of the displacement / pressure-time action sequence is too fast and whether the initial support force of the lifting column needs to be adjusted based on the verification results of the first to third steps, forming an evaluation index set [Y,Z]. If the action rate is too fast, Y=1, otherwise it is 0; if the initial support force of the lifting column needs to be adjusted, Z=1, otherwise it is 0.

[0190] Based on the learning-based simulation loop agent constructed above, step S4, when detecting the displacement / pressure-time action sequence, checks evaluation indicators such as whether the action rate of the displacement / pressure-time action sequence is too fast and whether the initial support force of the lifting column has been adjusted. The actual action strategy includes when to perform actions such as lowering the column, moving the frame, raising the column, and pushing the slide, as well as the specific values ​​of pressure / displacement during these actions. The specific content of the reward function for the learning-based dual-loop agent will be explained in the following steps.

[0191] S5, the input to the learning-type reality loop agent, namely the scene features and the output displacement / stress-time action sequence, is used as the input to the learning-type dual-loop agent. The learning-type dual-loop agent corrects the displacement / stress-time action sequence under the scene features to obtain the corrected displacement / stress-time action sequence with the best value under the scene features. The corrected displacement / stress-time action sequence is then returned to the learning-type simulation loop agent to verify whether there is a conflict. If the corrected displacement / stress-time action sequence has no conflict, the actual action strategy is directly given. If the corrected displacement / stress-time action sequence has a conflict, the corrected displacement / stress-time action sequence is fed back to the learning-type dual-loop agent for correction. This process is repeated until the corrected displacement / stress-time action sequence has no conflict. Finally, the actual action strategy is expanded to a preset strategy set to train a higher-quality learning-type reality loop agent.

[0192] It should be noted that the S5 section also includes the construction of a learning-based dual-loop agent, such as... Figure 4 As shown, constructing a learning-based dual-loop agent specifically includes:

[0193] First, the inputs and outputs of the learning reality loop agent are constructed as the state space of the learning dual-loop agent. That is, the state space of the learning dual-loop agent includes environmental physical constraints, action target constraints, device coordination constraints, local action constraints, and the output of the learning reality loop agent. Details regarding the environmental physical constraints, action target constraints, device coordination constraints, local action constraints, and the output of the learning reality loop agent are provided in the above embodiments and will not be repeated here.

[0194] Secondly, the action space of the learning dual-loop agent is defined as the action space for modifying the displacement / pressure-time action sequence at a specified moment. Specifically, the action space of the learning dual-loop agent is set to modify the displacement / pressure value at a certain moment, thereby achieving the goal of modifying the action sequence.

[0195] Next, the reward function of the learning dual-loop agent is determined based on hardware limit constraints, motion target constraints, action timing constraints, and simulation feedback constraints.

[0196] In one specific embodiment, the reward function of the learning-based dual-loop agent is set to r. , but:

[0197] Set the real-time speeds for lowering the column, moving the support frame, raising the column, and pushing the conveyor to be as follows: Its corresponding single-action limit speed is If hardware limitations violate real-world rules, a corresponding penalty will be imposed, expressed as follows:

[0198] ;

[0199] ;

[0200] ;

[0201] ;

[0202] Regarding the constraints on the motion target, the displacement value of the moving frame is set to... The displacement value of the pusher is The minimum pressure of the column is The initial support force of the lifting column is The target displacement values ​​for the moving frame and the pushing slide are respectively and The target minimum pressure for column reduction is The initial support force of the rising column is The penalty is determined based on the ratio of the difference between the real-time value and the target value, expressed as:

[0203] ;

[0204] ;

[0205] ;

[0206] ;

[0207] Regarding action timing constraints, if the traversal direction is followed by an error, then... If it follows the correct direction but conflicts with the adjacent frame (i.e., the adjacent frame's action status is 1), this frame begins to move and is set. ;

[0208] Regarding the simulation feedback constraints, the set of evaluation metrics given by the learning simulation loop agent is set as follows: ,if ,but ;if ,but ;if ,but ;if ,but .

[0209] Finally, the learning double-loop agent is trained based on the state space, action space, and reward function of the learning double-loop agent.

[0210] Specifically, when training a learning-type two-loop agent, the DDPG (Deep Deterministic Policy Gradient) algorithm can be used to train the learning-type two-loop agent based on its state space, action space, and reward function.

[0211] It should be noted that the specific type of the learning reality loop agent is not specifically limited in this embodiment of the invention. For example, the learning reality loop agent can be a manual control decision model after the automation of the central hydraulic support cluster, an intelligent prediction and control model for the straightness of the working face, a dynamic prediction model for spatiotemporal pressure events in the intelligent fully mechanized mining face, or an instant prediction model for the pressure-bearing effect after the initial support of the hydraulic support, and so on. The learning simulation loop agent and the learning dual-loop agent correspond to the learning reality loop agent. For example, the learning simulation loop agent corresponding to the intelligent prediction and control model for the straightness of the working face is a simulation model that controls the straightness of the working face, and the corresponding learning dual-loop agent is a combination of a reality model and a simulation model that controls the straightness of the working face.

[0212] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A hydraulic support multi-agent autonomous learning method, characterized in that, Implementation based on pre-deployed learning-based real-world loop agents, learning-based simulated loop agents, and learning-based dual-loop agents, including: S1, Obtain scene features and match scene features with a preset strategy set; S2, If there is a scene in the preset strategy set that matches the scene features, then output the displacement / pressure-time action sequence corresponding to the scene features in the preset strategy set; S3, if there is no scene matching the scene features in the preset strategy set, then the scene features are input into the learning reality loop agent, the learning reality loop agent outputs the displacement / pressure-time action sequence under the scene features, and the displacement / pressure-time action sequence output by the learning reality loop agent is used as the input of the learning simulation loop agent; S4, the learning simulation loop agent detects the displacement / pressure-time action sequence and feeds back several evaluation indicators based on a preset evaluation indicator set. The evaluation indicators are used to reflect whether there is a conflict in the displacement / pressure-time action sequence output by the learning reality loop agent. If there is no conflict in the displacement / pressure-time action sequence, the actual action strategy is given directly. If there is a conflict in the displacement / pressure-time action sequence, the evaluation indicators will be used as part of the reward function of the learning dual-loop agent to evaluate the value of the state under the scene features. S5, the input to the learning-type reality loop agent, namely the scene features and the output displacement / stress-time action sequence, is used as the input to the learning-type dual-loop agent. The learning-type dual-loop agent corrects the displacement / stress-time action sequence under the scene features to obtain the corrected displacement / stress-time action sequence with the best value under the scene features. The corrected displacement / stress-time action sequence is then returned to the learning-type simulation loop agent to verify whether there is a conflict. If the corrected displacement / stress-time action sequence has no conflict, the actual action strategy is directly given. If the corrected displacement / stress-time action sequence has a conflict, the corrected displacement / stress-time action sequence is fed back to the learning-type dual-loop agent for correction. This process is repeated until the corrected displacement / stress-time action sequence has no conflict. Finally, the actual action strategy is expanded to a preset strategy set to train a higher-quality learning-type reality loop agent.

2. The hydraulic support multi-agent autonomous learning method according to claim 1, characterized in that, The matching of scene features with a preset strategy set includes: The search algorithm is used to search the preset strategy set.

3. The hydraulic support multi-agent autonomous learning method according to claim 1, characterized in that, The learning-based reality loop agent includes a learning-based reality loop agent for hydraulic supports. Before step S3, the process further includes: constructing a learning-based reality loop agent for the hydraulic supports, specifically including: First, based on the action characteristics of the four actions of hydraulic supports in the well, namely lowering the support column, moving the support, raising the support column, and pushing the slide, a learning-type reality loop of hydraulic supports is constructed. The output of the agent is the displacement-time action sequence and the pressure-time action sequence. Secondly, the input of the learning-type reality environment agent for the hydraulic support is constructed based on environmental physical constraints, action target constraints, equipment coordination constraints, and the frame's action constraints. Next, a loss function for the learning-type reality loop agent of the hydraulic support is constructed based on hardware limit constraints, motion target constraints, and action timing constraints. Finally, the learning reality loop agent for the hydraulic support is trained based on the output, input, and loss function of the learning reality loop agent for the hydraulic support.

4. The hydraulic support multi-agent autonomous learning method according to claim 3, characterized in that, The inputs to the learning reality loop agent for constructing hydraulic supports include: In terms of environmental physical constraints, the angle of the hydraulic support base, the top and bottom plate material are quantified as mathematical indexes, and the angle of the hydraulic support base is set as The friction coefficient of the bottom plate is set as an index reflecting the material of the bottom plate The compressive strength of the top plate is set as an index reflecting the material of the top plate ; In the action target constraint aspect, the target displacement values of the moving frame and the pushing frame are respectively set as and The target minimum pressure of the lowering column and the target initial support force of the lifting column are respectively set as and ; Regarding equipment coordination constraints, the traversal direction of the coal mining machine is set as either forward or reverse, quantified into mathematical indicators represented as 0 and 1 respectively; whether the action state of the adjacent support reaches the state where this support is to take action is divided into 0 and 1; because when the coal mining machine passes by, the hydraulic support performs four actions in sequence: lowering the column, moving the support, raising the column, and pushing the conveyor. After the adjacent support completes the three actions of lowering the column, moving the support, and raising the column, this support takes action. Therefore, if the adjacent support completes the three actions of lowering the column, moving the support, and raising the column, the state is divided into 0, and otherwise the state is divided into 1. In the aspect of the frame motion constraint, the first two frames of pressure and displacement are taken as the input, denoted as and .

5. The hydraulic support multi-agent autonomous learning method according to claim 3 or 4, characterized in that, When constructing the loss function for the learning-based reality loop agent of the hydraulic support, the following are included: Let m be the loss function of the learning-based reality loop agent for the hydraulic support. Then: In terms of hardware limit constraints, the real-time speed of setting down column, moving frame, lifting column and pushing trolley is set as , and the corresponding single-action limit speed is If the hardware limit constraint aspect violates the real-world rules, the corresponding loss is given, which is represented as: ; ; ; ; In terms of the motion target constraint, the setting of the displacement value of the moving frame is , the displacement value of the pushing and pulling is , the minimum pressure of the descending column is , the initial support force of the ascending column is , the target displacement values of the moving frame and the pushing and pulling are and , the target minimum pressure of the descending column is , and the target initial support force of the ascending column is . The loss is determined according to the difference between the real-time value and the target value, which is represented as: ; ; ; ; Regarding action timing constraints, if the traversal direction is followed by an error, then... If it follows the correct direction but conflicts with the adjacent frame (i.e., the adjacent frame's action status is 1), this frame begins to move and is set. .

6. The multi-agent autonomous learning method for hydraulic supports according to claim 1, characterized in that, S4 is preceded by the construction of a learning simulation loop Agent, which specifically includes the construction of a hydraulic system mathematical model, a hydraulic support mathematical model, and a support-surrounding rock coupling model. The mathematical model of the hydraulic system is expressed as follows: ; ; ; in, This refers to the total fluid supply flow rate of the emulsion pump. The actual flow rate consumed by the actuator. This refers to the system's leaked traffic. For the speed of the actuator's operation, The area of ​​the liquid inlet chamber of the actuator; To supply fluid pressure to the hydraulic system, For the load of the actuator, For the back pressure of the liquid outlet chamber of the actuator, The area of ​​the liquid outlet chamber of the actuator; The mathematical model of the hydraulic support includes a column raising model, a column lowering model, a support shifting model, and a sliding model: The rising column model is represented as: ; in, For the lifting column thrust, To supply fluid pressure to the hydraulic system, The area of ​​the rodless cavity of the column is given. The area of ​​the column with a rod cavity. This refers to the back pressure of the liquid outlet chamber; The descending column model is represented as: ; in, To reduce column force; The moving frame model is represented as: ; in, For the shifting thrust, This indicates the area of ​​the rodless cavity of the jack. This indicates the area of ​​the rod cavity in the jack; The push-pull model is represented as: ; in, For pushing and pulling forces; The coupling model between the support and the surrounding rock includes a stiffness coupling model, a strength coupling model, and a stability coupling model: The stiffness coupling model is characterized by the compaction stiffness of the top plate and is expressed as follows: ; in, This refers to the compaction strength of the top slab on the support after the frame is moved. This is the initial support force for the lifting column after the frame is moved. This represents the total displacement of the moving frame; The strength coupling model is characterized by the dynamic resistance of the surrounding rock and is expressed as follows: ; in, The dynamic resistance of the surrounding rock during frame shifting / pushing. The drag-rate coefficient, For displacement rate; The stability coupling model, characterized by the risk of roof collapse, is expressed as follows: ; in, This represents the displacement increment during frame relocation. This is the time increment during the transfer of the rack. This represents the maximum moving rate.

7. The multi-agent autonomous learning method for hydraulic supports according to claim 6, characterized in that, After constructing the mathematical model of the hydraulic system, the mathematical model of the hydraulic support, and the coupling model between the support and the surrounding rock, the following is also included: Step 1: Verify the action speed based on the mathematical model of the hydraulic system, specifically including: Action speed verification: If the displacement-time action sequence generated by the learning simulation loop agent is calculated... This indicates that the action speed is too fast, exceeding the maximum total liquid supply flow rate of the emulsion pump. The speed needs to be reduced; Step 2: Verification based on the mathematical model of the hydraulic support, specifically including: Columnar pressure verification: If the columnar pressure is extracted based on the pressure-time action sequence generated by the learning simulation loop agent... Calculated Exceeding the rated working resistance of the column If so, the pressure is determined to be too high, including the lifting column pressure. This represents the initial support force during the lifting of the column; Displacement verification: If the velocity calculated from the displacement-time action sequence generated by the learning simulation loop agent is not less than the safe rate of the frame movement. and the safe speed of pushing If the displacement is too large, then it is determined that: ; ; in, and These are the maximum allowable flow rates for moving the frame and pushing the conveyor, respectively. Step 3: Validation based on the coupling model between the support and the surrounding rock, specifically including: (1) Stiffness coupling verification, specifically including: If the stress-time action sequence generated by the learning simulation loop agent is calculated... If the value is less than the empirical value, it indicates that the top plate compaction stiffness is insufficient, and the initial support force of the lifting column needs to be increased. (2) Strength coupling verification, specifically including: If the stress-time action sequence generated by the learning simulation loop agent is calculated... If the dynamic resistance of the surrounding rock exceeds the bearing capacity of the support, the operating rate needs to be reduced; among which, This indicates the upper limit of the dynamic resistance of the surrounding rock that the hydraulic support can withstand under rated operating conditions. The rated working pressure of the column is given by k, which is the mapping coefficient between the dynamic resistance of the surrounding rock and the pressure of the column. (3) Verification of roof collapse risk, specifically including: If the displacement-time action sequence generated by the learning simulation loop agent is calculated... Therefore, it is determined that the mutation rate needs to be reduced and the mutation range expanded. ; The learning simulation loop agent judges whether the action rate of the displacement / pressure-time action sequence is too fast and whether the initial support force of the lifting column needs to be adjusted based on the verification results of the first to third steps, forming an evaluation index set [Y,Z]. If the action rate is too fast, Y=1, otherwise it is 0; if the initial support force of the lifting column needs to be adjusted, Z=1, otherwise it is 0.

8. The multi-agent autonomous learning method for hydraulic supports according to claim 1, characterized in that, The S5 preceding S5 also includes constructing a learning-based dual-loop agent, specifically including: First, the input and output of the learning reality loop agent are constructed as the state space of the learning dual-loop agent; Secondly, the displacement / pressure-time action sequence at a specified time is set as the action space of the learning dual-loop agent; Next, the reward function of the learning dual-loop agent is determined based on hardware limit constraints, motion target constraints, action timing constraints, and simulation feedback constraints. Finally, the learning double-loop agent is trained based on the state space, action space, and reward function of the learning double-loop agent.

9. The multi-agent autonomous learning method for hydraulic supports according to claim 8, characterized in that, Let the reward function of the learning-type two-loop agent be r. , but: Set the real-time speeds for lowering the column, moving the support frame, raising the column, and pushing the conveyor to be as follows: Its corresponding single-action limit speed is If hardware limitations violate real-world rules, a corresponding penalty will be imposed, expressed as follows: ; ; ; ; Regarding the constraints on the motion target, the displacement value of the moving frame is set to... The displacement value of the pusher is The minimum pressure of the column is The initial support force of the lifting column is The target displacement values ​​for the moving frame and the pushing slide are respectively and The target minimum pressure for the column is The initial support force of the rising column is The penalty is determined based on the ratio of the difference between the real-time value and the target value, expressed as: ; ; ; ; Regarding action timing constraints, if the traversal direction is followed by an error, then... ; If it follows the correct direction but conflicts with the adjacent frame (i.e., the adjacent frame's action status is 1), this frame begins to move and is set. ; Regarding the simulation feedback constraints, the set of evaluation metrics given by the learning simulation loop agent is set as follows: ,if ,but ;if ,but ;if ,but ;if ,but .

10. The multi-agent autonomous learning method for hydraulic supports according to claim 8 or 9, characterized in that, The DDPG algorithm is used to train a learning dual-loop agent based on the state space, action space, and reward function of the learning dual-loop agent.