Multi-agent cooperative control decision-making method for hydraulic support

By building a strategy value network and strategy estimation model, combined with deep learning and reinforcement learning technology, the problem of adaptive and coordinated control of hydraulic support systems in complex coal mining scenarios is solved, and the adaptive coordinated control of hydraulic support clusters is realized, which improves the production efficiency and safety of intelligent mining work surfaces.

CN119957279AActive Publication Date: 2025-05-09TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510445682.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

In the complex and changeable coal mining face production scenarios, the single-cured closed-loop control logic is difficult to achieve normal application, making it difficult to adapt and coordinate the various control modes of the hydraulic support.

Method used

By building a strategy value network and strategy estimation model of the hydraulic support centralized control center, as well as the action estimation model of multiple hydraulic support electro-hydraulic controllers, deep learning and reinforcement learning technology are used to form an intelligent decision-making mechanism to realize adaptive collaborative control of the hydraulic support cluster.

Benefits of technology

It realizes dynamic decision-making of the coordinated distributed control rules of the hydraulic support electro-hydraulic control system, improves the adaptive control capabilities of the hydraulic support cluster, and supports the safe and efficient production of hydraulic support in the intelligent production working surface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119957279A_ABST
    Figure CN119957279A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-agent cooperative control decision-making method for a hydraulic support, and belongs to the technical field of coal mine intellectualization. Comprising the steps of obtaining real-time data of operation of the hydraulic support multi-agent under a current task, taking the real-time data as a current state set of the hydraulic support multi-agent, and randomly generating an initial linkage action sequence; based on the current state set and the initial connection action sequence, determining a current optimal connection action sequence of the hydraulic support multi-agent through a strategy value network and a strategy estimation model corresponding to the current task; decomposing the optimal connection action sequence into action sequences of the hydraulic supports through an action estimation model; and each hydraulic support is controlled to execute the next action sequence through the action sequence of each hydraulic support. The invention provides a hydraulic support multi-agent cooperative control decision-making method based on experience development and simulation exploration dual drive, and self-adaptive cooperative control of a hydraulic support cluster can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent coal mine technology, and in particular to a multi-agent collaborative control decision-making method for a hydraulic support. Background Art

[0002] The hydraulic support system is a combined equipment system with the largest number of coal mining faces and the most complex state changes. There are hard constraints between each other based on the coal mining process. Through the orderly, timely, accurate and coordinated control and execution of thousands of actions of hundreds of hydraulic supports in the working face, multiple processes such as following the machine to move the frame, roof support, and scraper conveyor are completed. At present, the construction of intelligent coal mine working faces is in its early stages, and the electro-hydraulic control automation technology of hydraulic supports is relatively mature. Through the sensor collection of working face environment and equipment status information, the electro-hydraulic control program of various control modes of hydraulic supports is set according to the requirements of coal mining process, and the lifting, lowering, moving, pushing, side guards and telescopic beams of hydraulic supports are realized. The automatic control of actions such as lifting, lowering, moving, pushing, side guards and telescopic beams of hydraulic supports is realized, which is the key technology for the safe and efficient production of intelligent mining working faces with few people.

[0003] The working face production scene is a complex scene composed of mining environment, comprehensive mining equipment and people. The single solidified closed-loop control logic cannot adapt to the complex, changeable and dynamic working face production scene. In actual production, it is difficult to implement various control modes of hydraulic supports in a normalized manner. Therefore, the adaptive control of hydraulic supports based on intelligent decision-making is one of the technical bottlenecks that urgently needs to be broken through in the current intelligent mining research, that is, based on big data analysis and artificial intelligence, an intelligent decision-making mechanism is formed through autonomous learning and data modeling to realize the adaptive collaborative control of hydraulic support clusters.

[0004] However, the hydraulic support system belongs to a large-scale key equipment group in coal mines, which has the characteristics of spatiotemporal constraints of coal mining technology, multi-source sensor information, complex cluster action control, and strong correlation with scene status. Its control decision is an optimal decision problem of control sequence with multi-level groups and individuals and multi-spatiotemporal scales of states and actions. From the perspective of artificial intelligence theory, the control of hydraulic support clusters can be regarded as a typical multi-agent autonomous learning and collaborative control problem. With the help of new generation artificial intelligence theories and methods such as deep learning, reinforcement learning, game and collaborative control, combined with professional knowledge in coal mining technology, coal mining machinery, hydraulic transmission, etc., the control decision mechanism of hydraulic support multi-agent system is studied, which is an important theoretical basis for realizing the adaptive control technology of hydraulic support clusters. Summary of the invention

[0005] In order to solve the above technical problems, the present invention provides a multi-agent collaborative control decision method for a hydraulic support. The technical solution of the present invention is as follows: A multi-agent collaborative control decision-making method for a hydraulic support, comprising: S1, taking a single hydraulic support as a single agent and the hydraulic support causal decision model as the agent environment model, constructs the strategy value network and strategy estimation model of the hydraulic support centralized control center, as well as the action estimation model of multiple hydraulic support electro-hydraulic controllers; S2, obtain the real-time data of the hydraulic support multi-agent under the current task, use it as the current state set of the hydraulic support multi-agent, and randomly generate the initial connection action sequence; S3, based on the current state set and initial connection action sequence of the hydraulic support multi-agent, the optimal connection action sequence of the hydraulic support multi-agent is determined through the strategy value network and strategy estimation model corresponding to the current task; S4, decomposing the current optimal connection action sequence of the hydraulic support multi-agent into the action sequence of each hydraulic support through the action estimation model; S5, controlling each hydraulic support to execute the next action sequence through the action sequence of each hydraulic support.

[0006] Optionally, when S3 is implemented, it includes: S31, inputting the current state set and the initial connection action sequence of the hydraulic support multi-agent into the strategy value network corresponding to the current task, and the strategy value network corresponding to the current task outputs the value of the hydraulic support multi-agent executing the initial connection action sequence under the current state set; S32, inputting the value output by the policy value network into the policy estimation model, and the policy estimation model outputs a new connection action sequence; S33, inputting the new connection action sequence into the strategy value network corresponding to the current task, and the strategy value network outputs the value of the hydraulic support multi-agent executing the new connection action sequence under the current state set; S34, repeatedly executing S32 and S33 until the refresh time of the collaborative control decision model is reached; S35, recording the value of each connection action sequence, and determining the current optimal connection action sequence of the hydraulic support multi-agent according to the value of each connection action sequence.

[0007] Optionally, when S5 is implemented, it includes: S51, controlling each hydraulic support to sequentially execute each action indicated in the next action sequence through the action sequence of each hydraulic support; S52, when the time for each hydraulic support to execute an action reaches the refresh time of the collaborative control decision model, a new optimal connection action sequence of the hydraulic support multi-agent is obtained, and each hydraulic support is controlled to execute the next action sequence through the new optimal connection action sequence of the hydraulic support multi-agent.

[0008] Optionally, before S2, the step further includes: establishing an experience pool, specifically including: S21, obtain the global sensor data of the intelligent mining working face with a duration of more than six months and all the action data of the hydraulic support multi-agent; S22, extract the correspondence between the global sensor data and all the action data, obtain the result state set obtained after the hydraulic support multi-agent executes the connection action sequence under the initial state set, and establish the correspondence between the initial state set, connection action sequence and result state set of the hydraulic support multi-agent to form an experience pool.

[0009] Optionally, the hydraulic support causal decision model includes a correspondence between an initial state set, a connection action sequence and a result state set, and the connection action sequence in the hydraulic support causal decision model is randomly generated.

[0010] Optionally, the strategy value network is a value function that maps the state set and connection action sequence of the hydraulic support multi-agent to value. The strategy value network of each task is trained with the connection action sequence and state set in the experience pool and the hydraulic support causal decision model as the sample set, and the value set obtained by the value function of each task as the output set.

[0011] Optionally, before S31, the method further includes: training a strategy value network corresponding to the current task, specifically including: S311, determine the value function of the hydraulic support multi-agent under the current task based on the experience pool and the hydraulic support causal decision structure model; S312, substituting multiple initial state sets and random connection action sequences in the experience pool and the hydraulic support causal decision structure model into the value function to obtain a value set; S313, using the experience pool and multiple initial state sets and connection action sequences in the hydraulic support causal decision structure model as sample sets, and using the value set as the output set to train the strategy value network corresponding to the current task.

[0012] Optionally, after determining the current optimal connection action sequence of the hydraulic support multi-agent, S3 further includes: The experience pool and / or the hydraulic support causal decision model are updated according to the current state set of the hydraulic support multi-agent and the current optimal connection action sequence of the hydraulic support multi-agent.

[0013] All the above optional technical solutions can be combined arbitrarily, and the present invention does not provide detailed descriptions of the structures after the combinations.

[0014] By means of the above scheme, the beneficial effects of the present invention are as follows: By constructing an experience pool and a hydraulic support causal decision model, and taking a single hydraulic support as a single agent and the hydraulic support causal decision model as the agent environment model, a strategy value network and strategy estimation model of the hydraulic support centralized control center, as well as an action estimation model of multiple hydraulic support electro-hydraulic controllers are constructed. After obtaining the real-time data of the hydraulic support multi-agent running under the current task, the current optimal connection action sequence of the hydraulic support multi-agent can be determined based on the strategy value network and the strategy estimation model, and the current optimal connection action sequence of the hydraulic support multi-agent can be decomposed into the action sequence of each hydraulic support through the action estimation model, so that each hydraulic support can be controlled to execute the next action sequence through the action sequence of each hydraulic support. This provides a hydraulic support multi-agent collaborative control decision method based on dual-drive of experience development and simulation exploration, realizes dynamic decision-making of the collaborative distributed control law of the hydraulic support electro-hydraulic control system, and can realize adaptive collaborative control of hydraulic support clusters, providing a practical theoretical method and feasible technical path for the adaptive control of hydraulic supports in intelligent mining working faces.

[0015] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flow chart of the multi-agent collaborative control decision-making method for a hydraulic support provided in an embodiment of the present invention.

[0017] Figure 2 Schematic diagram of a training strategy value network in an embodiment of the present invention.

[0018] Figure 3 It is a schematic diagram of a causal decision model of a hydraulic support under a hydraulic support following machine task in an embodiment of the present invention.

[0019] Figure 4 It is an overall flow chart of the multi-agent collaborative control decision-making method for a hydraulic support provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0021] The multi-agent collaborative control decision method for hydraulic supports provided in the embodiment of the present invention can be implemented by any electronic device with computing functions, such as a PC, mobile terminal or server. In view of the technical difficulties faced by the adaptive control of hydraulic support systems under complex dynamic coal mining scenarios and conditions with limited power resources, the embodiment of the present invention adopts multidisciplinary cross-disciplinary methods such as mining professional knowledge reasoning, on-site data mining and utilization, hydraulic support causal decision model modeling, deep reinforcement learning, etc., to study and construct a hydraulic support multi-agent collaborative control decision method based on dual-drive of experience development and simulation exploration, to realize dynamic decision-making of the coordinated distributed control law of the hydraulic support electro-hydraulic control system, and to provide a practical theoretical method and feasible technical path for the adaptive control of hydraulic supports on intelligent mining working faces. Figure 1 As shown, the hydraulic support multi-agent collaborative control decision method provided by the embodiment of the present invention includes the following steps S1 to S5: S1, taking a single hydraulic support as a single intelligent agent and the hydraulic support causal decision model as the intelligent agent environment model, constructs the strategy value network and strategy estimation model of the hydraulic support control center, as well as the action estimation model of multiple hydraulic support electro-hydraulic controllers.

[0022] Specifically, under the theoretical framework of multi-agent reinforcement learning, the embodiment of the present invention takes a single hydraulic support as a single agent and a hydraulic support causal decision model as an agent environment model, designs a hydraulic support multi-agent A (Actor) -C (Critic) framework, takes a single hydraulic support electro-hydraulic controller as an Actor, and takes a hydraulic support centralized control center as a Critic, to form a hydraulic support multi-agent collaborative control decision method based on the AC framework. The method adopts a deep network modeling method to respectively construct a strategy estimation model and a strategy value network of the hydraulic support centralized control center (Critic), as well as an action estimation model of multiple hydraulic support electro-hydraulic controllers (Actor). The input and output are shown in Table 1.

[0023]

[0024] The hydraulic support causal decision model includes the correspondence between multiple hydraulic support multi-agent connection action sequences and state sets, and the state set includes an initial state set and a result state set connected by the connection action sequence. The hydraulic support causal decision model can generate a corresponding state set according to the connection action sequence.

[0025] In addition, the method provided by the embodiment of the present invention also needs to be implemented based on a pre-established experience pool. Therefore, it is necessary to establish an experience pool first. The steps of establishing an experience pool specifically include: S21, obtaining the global sensor data of the intelligent mining working face with a duration greater than six months and all the action data of the hydraulic support multi-agent; S22, extracting the correspondence between the global sensor data and all the action data, obtaining the result state set obtained after the hydraulic support multi-agent executes the connection action sequence under the initial state set, and establishing the correspondence between the initial state set, the connection action sequence and the result state set of the hydraulic support multi-agent to form an experience pool.

[0026] Specifically, relying on the developed intelligent mining face big data platform, global sensor data of working faces with a duration of more than six months are collected. Based on the observable variables defined in the established hydraulic support causal decision model, the global sensor data types collected include but are not limited to: sensor data such as coal mining machine inertial navigation trajectory, position, direction, speed, mining height, etc., sensor data such as hydraulic support angle, pressure, stroke, height, etc., hydraulic support multi-agent automatic / manual mode and all action data such as lowering column, raising column, pushing and sliding, balancing, side protection, and side protection. The data is used at an interval of 0.5s, and each data contains a timestamp. Based on the data pool composed of the above data, the connection action sequence of the hydraulic support multi-agent and its corresponding state set can be extracted from the data pool, such as what action the hydraulic support No. performed at what time and in what state, and what state it reached, to form an experience pool.

[0027] Based on the hydraulic support causal decision model, the result state set obtained by the hydraulic support multi-agent executing the connection action sequence under the initial state set can be continuously generated. Therefore, the hydraulic support causal decision model can continuously supplement data for the experience pool. In the embodiment of the present invention, the method of learning through the experience pool is called experience learning, and the method of learning through the hydraulic support causal decision model is called simulation exploration.

[0028] The strategy value network is a value function that maps the state set and connection action sequence of the hydraulic support multi-agent to the value. Specifically, for a certain task, the value function is a function that maps the state set and connection action sequence under the task to the value. It generates the connection action sequence and state set through the experience pool composed of real data and the hydraulic support causal decision model. According to experience, the most suitable function is actually constructed to fit the value of the initial state set executing this connection action sequence under the task. By continuously modifying the fitted value function, a very accurate value function for the task can be obtained. For value functions under different tasks, the variables are different. For example, if the task is to follow the machine, the value function is a variable related to the speed of following the machine; if the task is to support, the value function is a variable related to the support state, support time, etc.

[0029] The strategy value network of each task is trained by using multiple connected action sequences and state sets in the experience pool and the hydraulic support causal decision model as sample sets, and the value set obtained by the value function of each task as the output set.

[0030] The input of the policy estimation model is the value of the connection action sequence, and the output is the connection action sequence. Specifically, during training, the value obtained by the policy value network is used as input to the policy estimation model, and the output of the policy estimation model is a new connection action sequence. In a cycle, the policy estimation model will record the value of each connection action sequence, and at this time, the optimal connection action sequence under the initial state set is continuously obtained through the Nash equilibrium algorithm.

[0031] When actually selecting the policy value network and policy estimation model for modeling, you can choose deep learning network architectures with different parameter magnitudes according to the amount of data, such as the Transformer architecture and new architectures such as large models. After training on a large amount of data, they can be well suited for the current task.

[0032] S2, obtains the real-time data of the hydraulic support multi-agent running under the current task, takes it as the current state set of the hydraulic support multi-agent, and randomly generates the initial connection action sequence.

[0033] Among them, the real-time data of the hydraulic support multi-agent running under the current task is based on different task types, and the types of real-time data obtained are also different. For example, if the current task is to follow the machine, it is related to data such as pressure and displacement. Therefore, the real-time data obtained includes pressure and displacement, etc. Specifically, when acquiring real-time data, it is implemented based on the sensors that have been installed on the intelligent mining working face. The current state set includes multiple state parameters related to the current task of the hydraulic support. The connection action sequence includes multiple actions that the hydraulic support multi-agent needs to perform in sequence when performing the current task under the current state set. The hydraulic support multi-agent in an embodiment of the present invention refers to a hydraulic support related to the current task. For example, for a following machine control task, when following the machine, it may only be related to several hydraulic supports near the coal mining machine. Then the hydraulic support multi-agent under this task is several hydraulic supports near the coal mining machine.

[0034] When determining the initial connection action sequence, a connection action sequence is randomly selected from the experience pool or the hydraulic support causal decision model as the initial connection action sequence.

[0035] S3, based on the current state set and initial connection action sequence of the hydraulic support multi-agent, the current optimal connection action sequence of the hydraulic support multi-agent is determined through the strategy value network and strategy estimation model corresponding to the current task.

[0036] In a specific embodiment, S3 can be implemented by the Nash equilibrium algorithm, which is a calculation method for finding the Nash equilibrium point in game theory. The strategy selected by the Nash equilibrium algorithm is indeed the best choice for each agent in the current state. At the Nash equilibrium point, each agent chooses the most advantageous strategy for itself after considering the strategies of other agents, so no participant can gain more benefits by unilaterally changing the strategy. In the embodiment of the present invention, S3 specifically includes S31 to S35: S31, input the current state set and initial connection action sequence of the hydraulic support multi-agent into the strategy value network corresponding to the current task, and the strategy value network corresponding to the current task outputs the value of the hydraulic support multi-agent executing the initial connection action sequence under the current state set.

[0037] Among them, before the specific implementation of S31, it is necessary to first train the strategy value network corresponding to the current task, specifically including: S311, determining the value function of the hydraulic support multi-agent under the current task based on the experience pool and the hydraulic support causal decision structure model; S312, substituting multiple initial state sets and random connection action sequences in the experience pool and the hydraulic support causal decision structure model into the value function to obtain the value set; S313, using the multiple initial state sets and connection action sequences in the experience pool and the hydraulic support causal decision structure model as the sample set, and the value set as the output set to train the strategy value network corresponding to the current task.

[0038] Specifically, when S311 is implemented, according to the correspondence between the initial state set, the connection action sequence and the result state set in the experience pool and the hydraulic support causal decision structure model, the state parameter fitting value function related to the current task is selected. The fitting process can be performed by an expert or implemented by a corresponding function fitting algorithm.

[0039] Regarding the process of training the strategy value network corresponding to the current task, reference may be made to the existing artificial intelligence neural network training process, which will not be elaborated in detail in the present invention.

[0040] S32, input the value output by the policy value network into the policy estimation model, and the policy estimation model outputs a new connection action sequence.

[0041] S33, input the new connection action sequence into the strategy value network, and the strategy value network outputs the value of the hydraulic support multi-agent executing the new connection action sequence under the current state set.

[0042] S34, repeatedly executing S32 and S33 until the refresh time of the collaborative control decision model is reached.

[0043] S35, recording the value of each connection action sequence, and determining the current optimal connection action sequence of the hydraulic support multi-agent according to the value of each connection action sequence.

[0044] For the refresh time of each collaborative control decision model, the above steps S31 to S35 are executed to obtain the optimal connection action sequence within the refresh time of each collaborative control decision model.

[0045] It should be noted that the refresh time of the collaborative control decision model is set as needed, and its time interval is less than or equal to the time it takes for the hydraulic support multi-agent to complete the execution of the current optimal connection action sequence. Preferably, the refresh time of the collaborative control decision model is less than the time it takes for the hydraulic support multi-agent to complete the execution of the current optimal connection action sequence, so that the next optimal connection action sequence has been obtained before the hydraulic support multi-agent completes the execution of the current optimal connection action sequence.

[0046] S4, through the action estimation model, the current optimal connection action sequence of the hydraulic support multi-agent is decomposed into the action sequence of each hydraulic support.

[0047] Specifically, when the connection action sequence is input into the action estimation model, the action estimation model will decompose the connection action sequence to obtain the next action sequence of each hydraulic support single agent under the current task. Among them, the action sequence of each hydraulic support may include one or more actions that need to be performed when the hydraulic support moves next time, and the time interval between each action is t.

[0048] S5, controlling each hydraulic support to execute the next action sequence through the action sequence of each hydraulic support.

[0049] In a specific embodiment, when S5 is implemented, it includes: S51, controlling each hydraulic support to execute each action indicated in the next action sequence in sequence through the action sequence of each hydraulic support; S52, when the time for each hydraulic support to execute the action reaches the refresh time of the collaborative control decision model, obtaining the new optimal connection action sequence of the hydraulic support multi-agent, and controlling each hydraulic support to execute the next action sequence through the new optimal connection action sequence of the hydraulic support multi-agent.

[0050] Specifically, when the action sequence of each hydraulic support controls each hydraulic support to sequentially execute each action indicated in the next action sequence, the method provided by the embodiment of the present invention will further determine the current new optimal connection action sequence of the hydraulic support multi-agent by continuing to go through steps S2 and S3. When the time for each hydraulic support to execute the action reaches the refresh time T of the collaborative control decision model, m (T m >t), the number of actions performed by each hydraulic support is Tm / t, at this time, each hydraulic support no longer continues to execute other subsequent actions, but starts to execute the actions indicated in the new optimal connection action sequence of the hydraulic support multi-agent. The method of obtaining the new optimal connection action sequence of the hydraulic support multi-agent can be achieved through steps S2 and S3, which will not be repeated here.

[0051] Furthermore, after determining the current optimal connection action sequence of the hydraulic support multi-agent, S3 can also: update the experience pool and / or the hydraulic support causal decision model according to the current state set of the hydraulic support multi-agent and the current optimal connection action sequence of the hydraulic support multi-agent, so as to continuously expand the data in the experience pool and / or the hydraulic support causal decision model, and update the strategy value network and strategy estimation model based on the updated data to make the strategy value network and strategy estimation model more accurate.

[0052] In order to facilitate understanding of the method provided by the embodiment of the present invention, the method provided by the embodiment of the present invention is described below by taking a specific example. Specifically, the example takes the current task of two hydraulic supports as a follow-up task as an example.

[0053] According to the cyclic process characteristics of the hydraulic support following task, it is defined that two hydraulic supports complete a following and moving frame control task as a separate scene (the task from the beginning to the completion is a separate scene). In each control cycle, the working surface inclination angle, fluid supply flow rate, and pressure limit are the initial state conditions of the scene, that is, the working surface inclination angle, fluid supply flow rate, and pressure limit are used as state concentration parameters. Each hydraulic support is an intelligent agent, and the "result" variables such as following speed, pressure, and displacement in the hydraulic support causal structure model are used as state indicators after the hydraulic support intelligent agent acts. Each hydraulic support intelligent agent can define seven actions such as starting to lower the column, stopping to lower the column, starting to pull the frame, stopping to pull the frame, starting to raise the column, stopping to raise the column, and no action, that is, the action of the i-th hydraulic support intelligent agent , then the connection action sequence of the two hydraulic support agents is In the process of the hydraulic support multi-agent completing a machine following task, a hydraulic support multi-agent "state set-action" connection action sequence set will be formed: "initial state set → connection action sequence → result state set 1 → connection action sequence → result state set 2 → ... → end of the scene".

[0054] Furthermore, the connection action sequence of the hydraulic support multi-agent is defined as a series of directed sequence actions at discrete moments (the step size can be fixed, greater than the planning calculation time and ensuring the control accuracy), that is, , then the connection action of the two hydraulic support agents at time t is , its connection action sequence is:

[0055] Assume that the shearer speed is The target values ​​of the pulling distances of No. 1 and No. 2 hydraulic supports are , The target values ​​of the initial support force are , The initial state set of working surface inclination, hydraulic upper limit, and fluid supply flow rate is , the pulling values ​​of No. 1 hydraulic support and No. 2 hydraulic support are l 1 , l 2 The initial support forces are P 1 , P 2 , the overall following speed of the two hydraulic supports is v Then, the value function of the j-th cycle following control target of the two hydraulic supports can be expressed as .

[0056] Based on the above content, the steps for implementing the scenario provided in this example are as follows: In the first step, the state set and connection action sequence generated by the experience pool and the hydraulic support causal decision model are used to fit the The actual function of the task is obtained by using the value function of the machine-following task.

[0057] In the second step, in order to avoid the "dimensionality disaster" problem of the table storage method caused by massive data samples, Deep QNetwork (DQN) is used to build a strategy value network, which includes: using the initial state set and connection action sequence generated by the experience pool and the hydraulic support causal structure model, and using the value function The obtained value set is used to train the strategy value network of the following task, such as Figure 2 shown.

[0058] The third step is to use the value set and experience pool obtained by the value function and the set of connection action sequences in the hydraulic support causal decision model to train the strategy estimation model. The training process uses the random game Nash Q-Learning algorithm to solve the Nash equilibrium of the multi-agent random game so that the connection strategy of the training process satisfies the following formula:

[0059] In addition to Another strategy is that at the Nash equilibrium, all agents cannot obtain greater rewards by simply changing their own strategies. By designing a Nash equilibrium convergence algorithm, we can obtain To achieve the optimal connection action sequence, multi-agent behavior game learning of hydraulic support is realized.

[0060] The fourth step is to obtain the initial state set of the two hydraulic supports currently following the machine, input it into the strategy value network of the following task, and obtain the optimal connection action sequence.

[0061] In the fifth step, the optimal connection action sequence is sent as input to the action estimation model, and the action estimation model will decompose the optimal connection action sequence into the action sequence that the first hydraulic support and the second hydraulic support should do.

[0062] At this point, a multi-agent collaborative control decision-making model for the two hydraulic support following tasks is designed.

[0063] Furthermore, in order to facilitate the understanding of the hydraulic support causal decision model, the hydraulic support causal decision model is illustrated below with an example.

[0064] Based on the knowledge of relevant fields such as coal mining process theory, coal mining machinery linkage principle, support and surrounding rock coupling theory, hydraulic transmission principle, control engineering principle, etc., the field experience is summarized through on-site observation of coal mining working face and communication with front-line experts. The observable scene state information of coal mining working face and the action type and timing output by the electro-hydraulic controller of hydraulic support are used as the "cause" node variables, and the observable information such as hydraulic support posture, support pressure, spatial displacement, etc. are used as the "result" node variables. The cause-effect diagram representation method is used to design the cause-effect path diagram of various control modes of hydraulic support such as following, straightening, and support, so as to characterize the spatio-temporal relationship between the dynamic scene state of the working face and the controllable action of the hydraulic support on the result state.

[0065] Taking the task of controlling the timing of the middle follower action of the hydraulic support as an example, the schematic diagram of the causal structure model of the hydraulic support is initially drawn as follows: Figure 3 shown. Figure 3 Notes on some symbols in the figure: J represents the action of lowering the column; S represents the action of raising the column; L represents the action of pulling the frame; -s represents the starting time of the action; -e represents the ending time of the action; -t represents the time width; -l represents the spatial displacement distance; -p represents the pressure value of the column (initial support force).

[0066] Specifically: 1J-s is expressed as the starting time of the first hydraulic support column lowering action that needs to follow the machine outside the safety distance based on the position of the coal mining machine, 1J-t is expressed as the execution time of the first hydraulic support column lowering action, 1J2J-t is expressed as the interval time between the start of the first hydraulic support column lowering and the start of the second hydraulic support column lowering, 1J-l is expressed as the descending distance of the first hydraulic support column lowering, 1J-p is expressed as the column pressure when the first hydraulic support column is lowered, and the others are similar.

[0067] Figure 3The figure shows the causal relationship path diagram of the timing control mode of the lowering-shifting-lifting action of two hydraulic support followers drawn initially. The square is the observed variable, the circle is the latent variable, the starting point of the arrow represents the "dependent" variable, and the end point of the arrow represents the dynamic support target that the hydraulic support follower should complete, and the key indicators can be measured by sensors, etc. The hydraulic support follower speed, initial support force, and frame moving distance are designed as the main "result" variables; considering the key time parameters of the hydraulic support electro-hydraulic control timing control program, different action time intervals are designed as intermediate latent variables. Due to the complexity of the hydraulic support cluster action, the above path diagram example can only qualitatively represent the partial causal relationship of the hydraulic support follower support process "action (scenario condition)-state", and the actual causal structure model of the entire hydraulic support cluster action is more complex.

[0068] In summary, the present invention pre-constructs an experience pool and a causal decision model for a hydraulic support, and uses a single hydraulic support as a single agent and the causal decision model for the hydraulic support as an agent environment model to construct a strategy value network and a strategy estimation model for the hydraulic support control center, as well as an action estimation model for multiple hydraulic support electro-hydraulic controllers. This allows the present optimal connection action sequence of the hydraulic support multi-agent to be determined based on the strategy value network and the strategy estimation model after obtaining the real-time data of the hydraulic support multi-agent running under the current task, and the action estimation model is used to decompose the present optimal connection action sequence of the hydraulic support multi-agent into the action sequence of each hydraulic support, thereby controlling each hydraulic support to execute the next action sequence through the action sequence of each hydraulic support. This provides a multi-agent collaborative control decision method for hydraulic supports based on dual drive of experience development and simulation exploration, which can realize adaptive collaborative control of hydraulic support clusters. Figure 4 As shown, it is an overall flow chart of the method provided by an embodiment of the present invention.

[0069] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the technical principles of the present invention, and these improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A multi-agent collaborative control decision-making method for a hydraulic support, characterized in that: include: S1, taking a single hydraulic support as a single agent and the hydraulic support causal decision model as the agent environment model, constructs the strategy value network and strategy estimation model of the hydraulic support centralized control center, as well as the action estimation model of multiple hydraulic support electro-hydraulic controllers; S2, obtain the real-time data of the hydraulic support multi-agent under the current task, use it as the current state set of the hydraulic support multi-agent, and randomly generate the initial connection action sequence; S3, based on the current state set and initial connection action sequence of the hydraulic support multi-agent, the optimal connection action sequence of the hydraulic support multi-agent is determined through the strategy value network and strategy estimation model corresponding to the current task; S4, decomposing the current optimal connection action sequence of the hydraulic support multi-agent into the action sequence of each hydraulic support through the action estimation model; S5, controlling each hydraulic support to execute the next action sequence through the action sequence of each hydraulic support.

2. The multi-agent collaborative control decision-making method for hydraulic support according to claim 1 is characterized in that: When S3 is specifically implemented, it includes: S31, inputting the current state set and the initial connection action sequence of the hydraulic support multi-agent into the strategy value network corresponding to the current task, and the strategy value network corresponding to the current task outputs the value of the hydraulic support multi-agent executing the initial connection action sequence under the current state set; S32, inputting the value output by the policy value network into the policy estimation model, and the policy estimation model outputs a new connection action sequence; S33, inputting the new connection action sequence into the strategy value network corresponding to the current task, and the strategy value network outputs the value of the hydraulic support multi-agent executing the new connection action sequence under the current state set; S34, repeatedly executing S32 and S33 until the refresh time of the collaborative control decision model is reached; S35, recording the value of each connection action sequence, and determining the current optimal connection action sequence of the hydraulic support multi-agent according to the value of each connection action sequence.

3. The multi-agent collaborative control decision-making method for hydraulic support according to claim 2 is characterized in that: When S5 is specifically implemented, it includes: S51, controlling each hydraulic support to sequentially execute each action indicated in the next action sequence through the action sequence of each hydraulic support; S52, when the time for each hydraulic support to execute an action reaches the refresh time of the collaborative control decision model, a new optimal connection action sequence of the hydraulic support multi-agent is obtained, and each hydraulic support is controlled to execute the next action sequence through the new optimal connection action sequence of the hydraulic support multi-agent.

4. The multi-agent collaborative control decision-making method for hydraulic support according to claim 1 is characterized in that: The step S2 also includes: establishing an experience pool, which specifically includes: S21, obtain the global sensor data of the intelligent mining working face with a duration of more than six months and all the action data of the hydraulic support multi-agent; S22, extract the correspondence between the global sensor data and all the action data, obtain the result state set obtained after the hydraulic support multi-agent executes the connection action sequence under the initial state set, and establish the correspondence between the initial state set, connection action sequence and result state set of the hydraulic support multi-agent to form an experience pool.

5. The multi-agent collaborative control decision-making method for hydraulic support according to claim 1 is characterized in that: The hydraulic support causal decision model includes the corresponding relationship between the initial state set, the connection action sequence and the result state set, and the connection action sequence in the hydraulic support causal decision model is randomly generated.

6. The multi-agent collaborative control decision-making method for hydraulic support according to claim 1 or 2, characterized in that: The strategy value network is a value function that maps the state set and connection action sequence of the hydraulic support multi-agent to value. The strategy value network of each task is trained with the connection action sequence and state set in the experience pool and the hydraulic support causal decision model as the sample set, and the value set obtained by the value function of each task as the output set.

7. The multi-agent collaborative control decision-making method for hydraulic support according to claim 2 is characterized in that: Before S31, the method further includes: training a strategy value network corresponding to the current task, specifically including: S311, determine the value function of the hydraulic support multi-agent under the current task based on the experience pool and the hydraulic support causal decision structure model; S312, substituting multiple initial state sets and random connection action sequences in the experience pool and the hydraulic support causal decision structure model into the value function to obtain a value set; S313, using the experience pool and multiple initial state sets and connection action sequences in the hydraulic support causal decision structure model as sample sets, and using the value set as the output set to train the strategy value network corresponding to the current task.

8. The multi-agent collaborative control decision-making method for hydraulic support according to claim 1 is characterized in that: After determining the optimal connection action sequence of the hydraulic support multi-agent, S3 further includes: The experience pool and / or the hydraulic support causal decision model are updated according to the current state set of the hydraulic support multi-agent and the current optimal connection action sequence of the hydraulic support multi-agent.

Citation Information

Patent Citations

  • Artificial intelligent type hydraulic support electric-hydraulic control system

    CN103352713A

  • Multi-mobile emergency power supply toughness optimization scheduling method based on data driving

    CN118449131A

  • Coal mine intelligent fully-mechanized coal mining method and system with coal mining being data mining

    CN118704951A

  • Man-machine hybrid formation intelligent decision generation method based on diffusion model and feedback learning

    CN118838164A

  • Cluster-oriented multi-agent cooperative task planning method

    CN119536258A