Multi-agent Cooperative Control Decision-making Method for Hydraulic Powered Support
By building a strategic value network and action estimation model of hydraulic support, using deep learning and reinforcement learning technology, the control problem of hydraulic support system in complex coal mining scenarios is solved, and the adaptive collaborative control of hydraulic support clusters is realized, and the control efficiency and production safety are improved.
Patent Information
- Application Number
- CN202510445682.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-10
AI Technical Summary
In the complex and changeable coal mining face production scenarios, the single-cured closed-loop control logic is difficult to achieve normal application, resulting in the difficulty of effective adaptability and coordinated control of multiple control modes of hydraulic support.
By building a strategy value network and strategy estimation model of the hydraulic support centralized control center, as well as the action estimation model of multiple hydraulic support electro-hydraulic controllers, deep learning and reinforcement learning technology are used to form an intelligent decision-making mechanism to realize adaptive collaborative control of the hydraulic support cluster.
It realizes dynamic decision-making of the coordinated distributed control rules of the hydraulic support electro-hydraulic control system, and can adaptively adjust the control strategy to improve the control efficiency and production safety of the hydraulic support cluster.
Smart Images

Figure CN119957279B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of coal mine intelligentization, and particularly to a multi-agent collaborative control decision-making method for hydraulic supports. Background Art
[0002] The hydraulic support system is a combined equipment system with the largest number and the most complex state changes in the coal mining face, and there are hard constraints among them based on the coal mining process. Through the orderly, timely, accurate, and coordinated control and execution of thousands of actions of hundreds of hydraulic supports in the working face, multiple processes such as following the shearer to move the supports, roof support, and scraper conveyor pushing are completed. At present, the construction of intelligent coal mining faces is in the initial stage, and the electro-hydraulic control automation technology of hydraulic supports has been relatively mature. By sensing and collecting the working face environment and equipment status information, and setting the electro-hydraulic control program of various control modes of hydraulic supports according to the requirements of the coal mining process, the automatic control of actions such as raising, lowering, moving, pushing the scraper conveyor, protecting the rib plate, and telescopic beam of hydraulic supports is realized, which is the key technology for the intelligent coal mining face to achieve less manpower, safety, and high efficiency in production.
[0003] The working face production scenario is a complex scenario jointly composed of the mining environment, fully mechanized mining equipment, and people. A single and fixed closed-loop control logic cannot adapt to the complex, changeable, and dynamic working face production scenario, and it is difficult to apply and implement various control modes of hydraulic supports normally in actual production. Therefore, the adaptive control of hydraulic supports based on intelligent decision-making is one of the technical bottlenecks that need to be urgently broken through in the current intelligent mining research, that is, based on big data analysis and artificial intelligence, an intelligent decision-making mechanism is formed through autonomous learning and data modeling to achieve the adaptive collaborative control of hydraulic support clusters.
[0004] However, the hydraulic support system belongs to a large-scale key equipment group in coal mines, with characteristics such as spatio-temporal constraints of coal mining processes, multi-source sensing information, complex cluster action control, and strong correlation with scene states. Its control decision-making is an optimal decision-making problem for control sequences at multiple levels of group-individual and multiple spatio-temporal scales of state-action. From the perspective of artificial intelligence theory, the cluster control of hydraulic supports can be regarded as a typical multi-agent autonomous learning and collaborative control problem. By using new-generation artificial intelligence theories and methods such as deep learning, reinforcement learning, game theory, and collaborative control, and combining professional field knowledge such as coal mining processes, coal mining machinery, and hydraulic transmission, studying the control decision-making mechanism of the multi-agent system of hydraulic supports is an important theoretical basis for realizing the adaptive control technology of hydraulic support clusters. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a multi-agent collaborative control decision-making method for hydraulic supports. The technical solution of the present invention is as follows:
[0006] A multi-agent collaborative control decision-making method for hydraulic supports, which includes:
[0007] S1. Take a single hydraulic support as a single intelligent agent, and use the causal decision-making model of the hydraulic support as the intelligent agent environment model to construct the policy value network and policy estimation model of the hydraulic support centralized control center, as well as the action estimation model of the electro-hydraulic controllers of multiple hydraulic supports;
[0008] S2. Obtain the real-time data of the operation of the multi-intelligent agents of the hydraulic support under the current task, use it as the current state set of the multi-intelligent agents of the hydraulic support, and randomly generate an initial connection action sequence;
[0009] S3. Based on the current state set and the initial connection action sequence of the multi-intelligent agents of the hydraulic support, determine the current optimal connection action sequence of the multi-intelligent agents of the hydraulic support through the policy value network and policy estimation model corresponding to the current task;
[0010] S4. Decompose the current optimal connection action sequence of the multi-intelligent agents of the hydraulic support into the action sequences of each hydraulic support through the action estimation model;
[0011] S5. Control each hydraulic support to execute the next action sequence through the action sequences of each hydraulic support.
[0012] Optionally, when specifically implemented, S3 includes:
[0013] S31. Input the current state set and the initial connection action sequence of the multi-intelligent agents of the hydraulic support into the policy value network corresponding to the current task, and the policy value network corresponding to the current task outputs the value of the multi-intelligent agents of the hydraulic support executing the initial connection action sequence under the current state set;
[0014] S32. Input the value output by the policy value network into the policy estimation model, and the policy estimation model outputs a new connection action sequence;
[0015] S33. Input the new connection action sequence into the policy value network corresponding to the current task, and the policy value network corresponding to the current task outputs the value of the multi-intelligent agents of the hydraulic support executing the new connection action sequence under the current state set;
[0016] S34. Repeat S32 and S33 until it stops when reaching the refresh time of the cooperative control decision model;
[0017] S35. Record the value of each connection action sequence, and determine the current optimal connection action sequence of the multi-intelligent agents of the hydraulic support according to the value of each connection action sequence.
[0018] Optionally, when specifically implemented, S5 includes:
[0019] S51. Control each hydraulic support to sequentially execute each action indicated in the next action sequence through the action sequences of each hydraulic support;
[0020] S52. When the execution time of each hydraulic support reaches the refresh time of the cooperative control decision model, obtain a new optimal connection action sequence of the multi-agent of the hydraulic support, and control each hydraulic support to execute the next action sequence through the new optimal connection action sequence of the multi-agent of the hydraulic support.
[0021] Optionally, before S2, it further includes: establishing an experience pool, specifically including:
[0022] S21. Obtain the global sensing data with a duration of more than six months in the intelligent coal mining face and all the action data of the multi-agent of the hydraulic support;
[0023] S22. Extract the corresponding relationship between the global sensing data and all the action data, obtain the result state set obtained after the multi-agent of the hydraulic support executes the connection action sequence under the initial state set, and establish the corresponding relationship between the initial state set, the connection action sequence and the result state set of the multi-agent of the hydraulic support to form an experience pool.
[0024] Optionally, the causal decision model of the hydraulic support includes the corresponding relationship between the initial state set, the connection action sequence and the result state set, and the connection action sequence in the causal decision model of the hydraulic support is randomly generated.
[0025] Optionally, the policy value network is a value function that maps the state set and the connection action sequence of the multi-agent of the hydraulic support to values. The policy value network of each task is trained with the experience pool, the connection action sequence in the causal decision model of the hydraulic support, and the state set as the sample set, and the value set obtained by the value function of each task as the output set.
[0026] Optionally, before S31, it further includes: training the policy value network corresponding to the current task, specifically including:
[0027] S311. Determine the value function of the multi-agent of the hydraulic support under the current task according to the experience pool and the causal decision structure model of the hydraulic support;
[0028] S312. Substitute multiple initial state sets and random connection action sequences in the experience pool and the causal decision structure model of the hydraulic support into the value function to obtain a value set;
[0029] S313. Use multiple initial state sets and connection action sequences in the experience pool and the causal decision structure model of the hydraulic support as the sample set, and the value set as the output set to train the policy value network corresponding to the current task.
[0030] Optionally, after S3 determines the current optimal connection action sequence of the multi-agent of the hydraulic support, it further includes:
[0031] Update the experience pool and / or the hydraulic support causal decision-making model according to the current state set of the multi-agent hydraulic supports and the current optimal connection action sequence of the multi-agent hydraulic supports.
[0032] All of the above optional technical solutions can be combined arbitrarily, and the present invention does not elaborate on the structures after combination one by one.
[0033] With the above solutions, the beneficial effects of the present invention are as follows:
[0034] By constructing an experience pool and a hydraulic support causal decision-making model, using a single hydraulic support as a single agent and the hydraulic support causal decision-making model as an agent environment model, constructing a policy value network and a policy estimation model for the hydraulic support centralized control center, and an action estimation model for the electro-hydraulic controllers of multiple hydraulic supports, after obtaining the real-time data of the multi-agent hydraulic supports running under the current task, the current optimal connection action sequence of the multi-agent hydraulic supports can be determined based on the policy value network and the policy estimation model, and the current optimal connection action sequence of the multi-agent hydraulic supports can be decomposed into the action sequences of each hydraulic support through the action estimation model, so as to control each hydraulic support to execute the next action sequence through the action sequences of each hydraulic support, thereby providing a method for collaborative control decision-making of multi-agent hydraulic supports driven by both empirical development and simulation exploration, realizing the dynamic decision-making of the collaborative distributed control law of the electro-hydraulic control system of hydraulic supports, achieving the adaptive collaborative control of the hydraulic support cluster, and providing a practical theoretical method and a feasible technical path for the adaptive control of hydraulic supports in intelligent mining faces.
[0035] The above description is only an overview of the technical solution of the present invention. In order to understand the technical means of the present invention more clearly and implement it according to the content of the specification, the following takes the preferred embodiments of the present invention and combines with the drawings to elaborate in detail as follows. Description of the Drawings
[0036] Figure 1 is the flowchart of the collaborative control decision-making method for multi-agent hydraulic supports provided by the embodiment of the present invention.
[0037] Figure 2 is the schematic diagram of training the policy value network in the embodiment of the present invention.
[0038] Figure 3 is the schematic diagram of the hydraulic support causal decision-making model under the hydraulic support following task in the embodiment of the present invention.
[0039] Figure 4 is the overall flowchart of the collaborative control decision-making method for multi-agent hydraulic supports provided by the embodiment of the present invention. Detailed Embodiments
[0040] The following further describes in detail the specific implementation manners of the present invention in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0041] The multi-agent collaborative control decision-making method for hydraulic supports provided by the embodiments of the present invention can be implemented by any electronic device with computing functions, such as a PC, a mobile terminal, or a server. In view of the technical problems faced by the adaptive control of the hydraulic support system under complex dynamic coal mining scenarios and limited power resources, the embodiments of the present invention adopt interdisciplinary methods such as mining professional domain knowledge reasoning, on-site data mining and utilization, hydraulic support causal decision model modeling, and deep reinforcement learning to research and construct a multi-agent collaborative control decision-making method for hydraulic supports driven by both empirical development and simulation exploration, realizing the dynamic decision-making of the collaborative distributed control law of the electro-hydraulic control system of hydraulic supports, and providing a practical theoretical method and a feasible technical path for the adaptive control of hydraulic supports in intelligent coal mining faces. As Figure 1 shown, the multi-agent collaborative control decision-making method for hydraulic supports provided by the embodiments of the present invention includes the following steps S1 to S5:
[0042] S1. Taking a single hydraulic support as a single agent and the hydraulic support causal decision model as the agent environment model, construct the policy value network and policy estimation model of the hydraulic support centralized control center, and the action estimation model of multiple hydraulic support electro-hydraulic controllers.
[0043] Specifically, in the framework of multi-agent reinforcement learning theory, the embodiments of the present invention take a single hydraulic support as a single agent and the hydraulic support causal decision model as the agent environment model, design a multi-agent A (Actor)-C (Critic) framework for hydraulic supports, use a single hydraulic support electro-hydraulic controller as the Actor and the hydraulic support centralized control center as the Critic to form a multi-agent collaborative control decision-making method for hydraulic supports based on the A-C framework. It adopts a deep network modeling method to respectively construct the policy estimation model and policy value network of the hydraulic support centralized control center (Critic), and the action estimation model of multiple hydraulic support electro-hydraulic controllers (Actor), and their input and output are shown in Table 1.
[0044]
[0045] Among them, the hydraulic support causal decision model includes the corresponding relationships between multiple connection action sequences and state sets of multi-agents of hydraulic supports. The state set includes an initial state set and a result state set connected by the connection action sequence. The hydraulic support causal decision model can generate the corresponding state set according to the connection action sequence.
[0046] In addition, the method provided by the embodiments of the present invention also needs to be implemented based on a pre-established experience pool. Therefore, it is necessary to establish an experience pool first. The steps of establishing the experience pool specifically include: S21, obtaining the global sensing data of the intelligent mining face with a duration of more than six months and all the action data of the multi-agent of the hydraulic support; S22, extracting the corresponding relationship between the global sensing data and all the action data, obtaining the result state set obtained after the multi-agent of the hydraulic support executes the connection action sequence under the initial state set, and establishing the corresponding relationship between the initial state set, the connection action sequence and the result state set of the multi-agent of the hydraulic support to form an experience pool.
[0047] Specifically, relying on the developed big data platform of the intelligent mining face, the global sensing data of the working face with a duration of more than six months is collected. According to the observable variables defined in the established causal decision model of the hydraulic support, the types of the collected global sensing data include but are not limited to: the inertial navigation trajectory, position, direction, speed, mining height and other sensing data of the shearer, the angle, pressure, stroke, height and other sensing data of the hydraulic support, all the action data of the multi-agent of the hydraulic support in the automatic / manual mode, such as lowering the support, raising the support, pushing the conveyor, balancing, side protection, rib protection, etc. The data is taken at intervals of 0.5 s, and each piece of data contains a timestamp. According to the data pool composed of the above data, the connection action sequence of the multi-agent of the hydraulic support and its corresponding state set can be extracted from the data pool, such as at what time the No. hydraulic support did what action in what state and what state was generated, to form an experience pool.
[0048] Based on the causal decision model of the hydraulic support, the result state set obtained by the multi-agent of the hydraulic support executing the connection action sequence under the initial state set can be continuously generated. Therefore, the causal decision model of the hydraulic support can continuously supplement data to the experience pool. The method of learning through the experience pool in the embodiments of the present invention is called experience learning, and the method of learning through the causal decision model of the hydraulic support is called simulation exploration.
[0049] The policy value network is a value function that maps the state set and connection action sequence of the multi-agent of the hydraulic support to a value. Specifically, for a certain task, the value function is a function that maps the state set and connection action sequence under this task to a value. It generates the connection action sequence and state set through the experience pool composed of real data and the causal decision model of the hydraulic support. According to experience, the most suitable function is actually constructed to fit the value of the initial state set executing this connection action sequence under this task. By continuously modifying the fitting value function, the very accurate value function of this task can be finally obtained. For the value functions under different tasks, their variables are different. For example, if the task is following the shearer, the value function is a variable related to the following speed; if the task is support, the value function is a variable related to the support state, support time, etc.
[0050] The policy value network for each task is trained with the experience pool, multiple connected action sequences and state sets in the hydraulic support causal decision-making model as the sample set, and the value set obtained by the value function of each task as the output set.
[0051] The input of the policy estimation model is the value of the connected action sequence, and the output is the connected action sequence. Specifically, during training, the value obtained from the policy value network is used as the input and fed into the policy estimation model. The output of the policy estimation model is a new connected action sequence. Repeating this process, the policy estimation model will record the value of each connected action sequence. At this time, the optimal connected action sequence under the initial state set is continuously obtained through the Nash equilibrium algorithm.
[0052] When actually choosing the policy value network and the policy estimation model for modeling, different deep learning network architectures with different parameter scales can be selected according to the size of the data volume, such as new architectures like the Transformer architecture and large models. After being trained with a large amount of data, they can be well applied to the current task.
[0053] S2. Obtain the real-time data of the hydraulic support multi-agent running under the current task, use it as the current state set of the hydraulic support multi-agent, and randomly generate an initial connected action sequence.
[0054] Among them, the real-time data of the hydraulic support multi-agent running under the current task is based on different task types, and the types of real-time data obtained are also different. For example, if the current task is following the shearer, it is related to data such as pressure and displacement. Therefore, the obtained real-time data includes pressure and displacement, etc. Specifically, when obtaining the real-time data, it is realized based on the sensors already installed in the intelligent mining face. The current state set includes multiple state parameters related to the current task of the hydraulic support. The connected action sequence includes multiple actions that the hydraulic support multi-agent needs to perform in sequence when executing the current task under the current state set. The hydraulic support multi-agent in the embodiments of the present invention refers to the hydraulic supports related to the current task. For example, for the following-the-shearers control task, when specifically following the shearer, it may only be related to several hydraulic supports near the shearer. Then, the hydraulic support multi-agent under this task is several hydraulic supports near the shearer.
[0055] When determining the initial connected action sequence, randomly select a connected action sequence from the experience pool or the hydraulic support causal decision-making model as the initial connected action sequence.
[0056] S3. Based on the current state set of the hydraulic support multi-agent and the initial connected action sequence, determine the current optimal connected action sequence of the hydraulic support multi-agent through the policy value network and the policy estimation model corresponding to the current task.
[0057] In a specific embodiment, when specifically implemented, S3 can be achieved through the Nash equilibrium algorithm. The Nash equilibrium algorithm is a calculation method used to find the Nash equilibrium point in game theory. The strategy selected through the Nash equilibrium algorithm is indeed the best choice for each agent in the current state. At the Nash equilibrium point, each agent, after considering the strategies of other agents, selects the strategy that is most beneficial to itself. Therefore, no participant can obtain more benefits by unilaterally changing the strategy. In the embodiment of the present invention, S3 specifically includes S31 to S35:
[0058] S31, input the current state set and the initial connection action sequence of the hydraulic support multi-agent into the policy value network corresponding to the current task, and the policy value network corresponding to the current task outputs the value of the hydraulic support multi-agent executing the initial connection action sequence under the current state set.
[0059] Among them, before specifically implementing S31, it is necessary to first train the policy value network corresponding to the current task, specifically including: S311, determine the value function of the hydraulic support multi-agent under the current task according to the experience pool and the hydraulic support causal decision structure model; S312, substitute multiple initial state sets and random connection action sequences in the experience pool and the hydraulic support causal decision structure model into the value function to obtain a value set; S313, use multiple initial state sets and connection action sequences in the experience pool and the hydraulic support causal decision structure model as the sample set, and use the value set as the output set to train the policy value network corresponding to the current task.
[0060] Specifically, when specifically implementing S311, according to the corresponding relationship between the initial state set, the connection action sequence, and the result state set in the experience pool and the hydraulic support causal decision structure model, select the state parameters related to the current task to fit the value function. The fitting process can be carried out by an expert or implemented through a corresponding function fitting algorithm.
[0061] Regarding the process of training the policy value network corresponding to the current task, reference can be made to the existing artificial intelligence neural network training process, and the present invention will not elaborate on this in detail.
[0062] S32, input the value output by the policy value network into the policy estimation model, and the policy estimation model outputs a new connection action sequence.
[0063] S33, input the new connection action sequence into the policy value network, and the policy value network outputs the value of the hydraulic support multi-agent executing the new connection action sequence under the current state set.
[0064] S34, repeatedly execute S32 and S33 until it stops when reaching the refresh time of the cooperative control decision model.
[0065] S35. Record the value of each connection action sequence, and determine the current optimal connection action sequence of the hydraulic support multi-agent according to the value of each connection action sequence.
[0066] For the refresh time of each cooperative control decision model, perform according to the above steps S31 to S35, and the optimal connection action sequence within the refresh time of each cooperative control decision model can be obtained.
[0067] It should be noted that the refresh time of the cooperative control decision model is set as needed, and its time interval is less than or equal to the time for the hydraulic support multi-agent to execute the current optimal connection action sequence. Preferably, the refresh time of the cooperative control decision model is less than the time for the hydraulic support multi-agent to execute the current optimal connection action sequence, so that before the hydraulic support multi-agent executes the current optimal connection action sequence, the next optimal connection action sequence has been obtained.
[0068] S4. Decompose the current optimal connection action sequence of the hydraulic support multi-agent into the action sequences of each hydraulic support through the action estimation model.
[0069] Specifically, when the connection action sequence is input into the action estimation model, the action estimation model will decompose the connection action sequence to obtain what the action sequence of each single hydraulic support agent will be in the next step under the current task. Among them, the action sequence of each hydraulic support may include one or more actions that the hydraulic support needs to execute during the next action, and the time interval between each action is t.
[0070] S5. Control each hydraulic support to execute the next action sequence through the action sequences of each hydraulic support.
[0071] In a specific embodiment, when implementing S5, it includes: S51. Control each hydraulic support to sequentially execute each action indicated in the next action sequence through the action sequences of each hydraulic support; S52. When the time for each hydraulic support to execute the action reaches the refresh time of the cooperative control decision model, obtain the new optimal connection action sequence of the hydraulic support multi-agent, and control each hydraulic support to execute the next action sequence through the new optimal connection action sequence of the hydraulic support multi-agent.
[0072] Specifically, when the action sequences of each hydraulic support control each hydraulic support to sequentially execute each action indicated in the next action sequence, the method provided by the embodiment of the present invention will also continue to determine the current new optimal connection action sequence of the hydraulic support multi-agent through steps S2 and S3. When the time for each hydraulic support to execute the action reaches the refresh time T m (T m >t), the number of actions executed by each hydraulic support is Tm / t. At this time, each hydraulic support no longer performs subsequent other actions, but starts to perform the actions indicated in the new optimal connection action sequence of the multi-agent of the hydraulic support. Regarding the method for obtaining the new optimal connection action sequence of the multi-agent of the hydraulic support, it can be implemented through steps S2 and S3, which will not be elaborated here.
[0073] Further, after determining the current optimal connection action sequence of the multi-agent of the hydraulic support in S3, it can also: update the experience pool and / or the causal decision model of the hydraulic support according to the current state set of the multi-agent of the hydraulic support and the current optimal connection action sequence of the multi-agent of the hydraulic support, so as to continuously expand the data in the experience pool and / or the causal decision model of the hydraulic support, and update the policy value network and the policy estimation model based on the updated data, making the policy value network and the policy estimation model more accurate.
[0074] To facilitate the understanding of the method provided by the embodiments of the present invention, the method provided by the embodiments of the present invention will be illustrated by a specific example below. Specifically, this example takes the current task of two hydraulic supports as the following-the-shearer task.
[0075] According to the cyclic process characteristics of the following-the-shearer task of the hydraulic support, defining the completion of a following-the-shearer support moving control task by two hydraulic supports as a single episode (a task from start to completion is a single episode), the working face inclination angle, the liquid supply flow rate, and the pressure limit at each control cycle are the initial state conditions of this episode, that is, the working face inclination angle, the liquid supply flow rate, and the pressure limit are used as parameters in the state set. Each hydraulic support is an agent, and the "result" variables such as the following-the-shearer speed, pressure, and displacement in the causal structure model of the hydraulic support are used as the state indicators after the action of the hydraulic support agent. Each hydraulic support agent can define seven actions including the start of lowering the support, the stop of lowering the support, the start of pulling the support, the stop of pulling the support, the start of raising the support, the stop of raising the support, and no action, that is, the action of the i-th hydraulic support agent , then the connection action sequence of the two hydraulic support agents is . During the process of the multi-agent of the hydraulic support completing a following-the-shearer task, a connection action sequence set of "state set - action" of the multi-agent of the hydraulic support will be formed as "initial state set → connection action sequence → result state set 1 → connection action sequence → result state set 2 →... → episode end".
[0076] Further, the connection action sequence of the multi-agent of the hydraulic support is defined as a directed sequence of actions at a series of discrete time moments (the step size can be fixed, greater than the planning calculation time and ensuring the control accuracy), that is , then the connection action of the two hydraulic support agents at time t is , and its connection action sequence is:
[0077]
[0078] Assume the shearer speed is , and the target values of the advancing distances of the No. 1 and No. 2 hydraulic supports are respectively 、 , and the target values of the initial support forces are respectively 、 . The initial state set such as the face dip angle, the hydraulic upper limit, and the supply flow rate is . The advancing values of the No. 1 hydraulic support and the No. 2 hydraulic support are respectively l 1 、 l 2 , and the initial support forces are respectively P 1 、 P 2 . The following-the-shearer speed of the two hydraulic supports as a whole is v . Then, the value function of the following-the-shearer control target for the j-th cycle of the two hydraulic supports can be expressed as .
[0079] Based on the above content, when the scenario provided in this example is specifically implemented, the steps are as follows:
[0080] First step, using the state set and the connected action sequence generated by the experience pool and the hydraulic support causal decision model, according to experience, fit the actual function to obtain the value function of the following-the-shearer task.
[0081] Second step, to avoid the "curse of dimensionality" problem of the tabular storage method caused by a large number of data samples, use Deep Q Network (DQN) to construct a policy value network, specifically including: using the initial state set and the connected action sequence generated by the experience pool and the hydraulic support causal structure model, and the value set obtained by using the value function to train the policy value network of the following-the-shearer task, as shown in Figure 2 .
[0082] Third step, use the value set obtained by the value function, the experience pool, and the set of connected action sequences in the hydraulic support causal decision model to train the policy estimation model. The random game Nash Q-Learning algorithm is used in the training process to solve the Nash equilibrium of the multi-agent random game, so that the connected strategy in the training process satisfies the following formula:
[0083]
[0084] represents except Another strategy is that at the Nash equilibrium, for all agents, they cannot obtain greater rewards by only changing their own strategies. By designing a Nash equilibrium convergence algorithm, the optimal connection action sequence at the current state is obtained to realize the multi-agent behavior game learning of hydraulic supports.
[0085] Step 4: Obtain the initial state set of the current following-the-shearer of two hydraulic supports and input it into the policy-value network of the following-the-shearer task to obtain the optimal connection action sequence.
[0086] Step 5: Take the optimal connection action sequence as the input and send it to the action estimation model. The action estimation model will decompose the optimal connection action sequence into the action sequences that the No. 1 hydraulic support and the No. 2 hydraulic support should perform.
[0087] So far, a multi-agent collaborative control decision model for the following-the-shearer task of two hydraulic supports is designed and formed.
[0088] Furthermore, to facilitate the understanding of the causal decision model of hydraulic supports, the causal decision model of hydraulic supports is illustrated by examples below.
[0089] Based on the relevant field knowledge such as coal mining technology theory, coal mining machinery linkage principle, support and surrounding rock coupling theory, hydraulic transmission principle, and control engineering principle, through on-site observation in the coal mining face and communication with front-line expert experience, etc., the on-site experience is summarized. Taking the observable scene state information of the coal mining face and the action types, time sequences, etc. that the electro-hydraulic controller of the hydraulic support can output as the "cause" node variables, and taking the observable information such as the pose, support pressure, and spatial displacement of the hydraulic support as the "effect" node variables, using the causal graph representation method, the causal relationship path diagrams of various control modes such as following-the-shearer, straightening, and support of the hydraulic support are designed to characterize the spatio-temporal action relationship between the dynamic scene state of the working face and the controllable actions of the hydraulic support on the result state.
[0090] Taking the time sequence control task of the middle following-the-shearer action of the hydraulic support as an example, the schematic diagram of the causal structure model of the hydraulic support drawn initially is as Figure 3 shown. Figure 3 Partial symbol annotations in the figure: J represents the action of lowering the support; S represents the action of raising the support; L represents the action of pulling the support; -s represents the start time of the action; -e represents the end time of the action; -t represents the time width; -l represents the spatial displacement distance; -p represents the value of the column pressure (initial support force).
[0091] Specifically, 1J-s represents the start time of the lowering action of the first hydraulic support that needs to follow the shearer outside the safe distance based on the position of the shearer, 1J-t represents the execution time of the lowering action of the first hydraulic support, 1J2J-t represents the time interval between the start of the lowering of the first hydraulic support and the start of the lowering of the second hydraulic support, 1J-l represents the lowering distance of the first hydraulic support during the lowering action, 1J-p represents the column pressure of the first hydraulic support during the lowering action, and so on for others.
[0092] Figure 3 The figure shows the causal relationship path diagram of the sequential control mode of the lowering-shifting-raising actions of two hydraulic supports following the shearer. The observed variables are within the squares, and the latent variables are within the circles. The starting point of the arrow represents the "cause" variable, and the ending point of the arrow represents that the hydraulic support following the shearer should achieve the dynamic support target, and key indicators can be measured by sensors, etc. The speed of the hydraulic support following the shearer, the initial support force, the shifting distance, etc. are designed as the main "effect" variables; considering the key time parameters of the electro-hydraulic sequential control program of the hydraulic support, different action time intervals are designed as intermediate latent variables. Due to the complexity of the collective actions of the hydraulic supports, the above path diagram example can only qualitatively represent part of the causal relationship of the "action (scenario conditions)-state" in the process of the hydraulic support following the shearer, and the actual causal structure model of all the collective actions of the hydraulic supports is more complex.
[0093] In summary, the present invention pre-constructs an experience pool and a causal decision model for hydraulic supports, uses a single hydraulic support as a single intelligent agent, and uses the causal decision model for hydraulic supports as the intelligent agent environment model to construct a policy value network and a policy estimation model for the centralized control center of hydraulic supports, and an action estimation model for the electro-hydraulic controllers of multiple hydraulic supports. After obtaining the real-time data of the operation of the multiple intelligent agents of the hydraulic supports under the current task, it is possible to determine the current optimal connection action sequence of the multiple intelligent agents of the hydraulic supports based on the policy value network and the policy estimation model, and decompose the current optimal connection action sequence of the multiple intelligent agents of the hydraulic supports into the action sequences of each hydraulic support through the action estimation model, so as to control each hydraulic support to execute the next action sequence through the action sequences of each hydraulic support, thereby providing a method for collaborative control decision-making of multiple intelligent agents of hydraulic supports driven by both empirical development and simulation exploration, which can achieve the adaptive collaborative control of the hydraulic support cluster. As Figure 4 shown, it is the overall flowchart of the method provided by the embodiment of the present invention.
[0094] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A multi-agent collaborative control decision-making method for a hydraulic support, characterized in that: include: S1, taking a single hydraulic support as a single agent and the hydraulic support causal decision model as the agent environment model, constructs the strategy value network and strategy estimation model of the hydraulic support centralized control center, as well as the action estimation model of multiple hydraulic support electro-hydraulic controllers; S2, obtain the real-time data of the hydraulic support multi-agent under the current task, use it as the current state set of the hydraulic support multi-agent, and randomly generate the initial connection action sequence; S3, based on the current state set and initial connection action sequence of the hydraulic support multi-agent, the optimal connection action sequence of the hydraulic support multi-agent is determined through the strategy value network and strategy estimation model corresponding to the current task; S4, decomposing the current optimal connection action sequence of the hydraulic support multi-agent into the action sequence of each hydraulic support through the action estimation model; S5, controlling each hydraulic support to execute the next action sequence through the action sequence of each hydraulic support; When S3 is specifically implemented, it includes: S31, inputting the current state set and the initial connection action sequence of the hydraulic support multi-agent into the strategy value network corresponding to the current task, and the strategy value network corresponding to the current task outputs the value of the hydraulic support multi-agent executing the initial connection action sequence under the current state set; S32, inputting the value output by the policy value network into the policy estimation model, and the policy estimation model outputs a new connection action sequence; S33, inputting the new connection action sequence into the strategy value network corresponding to the current task, and the strategy value network outputs the value of the hydraulic support multi-agent executing the new connection action sequence under the current state set; S34, repeatedly executing S32 and S33 until the refresh time of the collaborative control decision model is reached; S35, recording the value of each connection action sequence, and determining the current optimal connection action sequence of the hydraulic support multi-agent according to the value of each connection action sequence.
2. The multi-agent collaborative control decision-making method for hydraulic support according to claim 1 is characterized in that: When S5 is specifically implemented, it includes: S51, controlling each hydraulic support to sequentially execute each action indicated in the next action sequence through the action sequence of each hydraulic support; S52, when the time for each hydraulic support to execute an action reaches the refresh time of the collaborative control decision model, a new optimal connection action sequence of the hydraulic support multi-agent is obtained, and each hydraulic support is controlled to execute the next action sequence through the new optimal connection action sequence of the hydraulic support multi-agent.
3. The multi-agent collaborative control decision-making method for hydraulic support according to claim 1 is characterized in that: The step S2 also includes: establishing an experience pool, which specifically includes: S21, obtain the global sensor data of the intelligent mining working face with a duration of more than six months and all the action data of the hydraulic support multi-agent; S22, extract the correspondence between the global sensor data and all the action data, obtain the result state set obtained after the hydraulic support multi-agent executes the connection action sequence under the initial state set, and establish the correspondence between the initial state set, connection action sequence and result state set of the hydraulic support multi-agent to form an experience pool.
4. The multi-agent collaborative control decision-making method for hydraulic support according to claim 1 is characterized in that: The hydraulic support causal decision model includes the corresponding relationship between the initial state set, the connection action sequence and the result state set, and the connection action sequence in the hydraulic support causal decision model is randomly generated.
5. The multi-agent collaborative control decision-making method for hydraulic support according to claim 1 is characterized in that: The strategy value network is a value function that maps the state set and connection action sequence of the hydraulic support multi-agent to value. The strategy value network of each task is trained with the connection action sequence and state set in the experience pool and the hydraulic support causal decision model as the sample set, and the value set obtained by the value function of each task as the output set.
6. The multi-agent collaborative control decision-making method for hydraulic support according to claim 1 is characterized in that: Before S31, the method further includes: training a strategy value network corresponding to the current task, specifically including: S311, determine the value function of the hydraulic support multi-agent under the current task based on the experience pool and the hydraulic support causal decision structure model; S312, substituting multiple initial state sets and random connection action sequences in the experience pool and the hydraulic support causal decision structure model into the value function to obtain a value set; S313, using the experience pool and multiple initial state sets and connection action sequences in the hydraulic support causal decision structure model as sample sets, and using the value set as the output set to train the strategy value network corresponding to the current task.
7. The multi-agent collaborative control decision-making method for hydraulic support according to claim 1 is characterized in that: After determining the optimal connection action sequence of the hydraulic support multi-agent, S3 further includes: The experience pool and / or the hydraulic support causal decision model are updated according to the current state set of the hydraulic support multi-agent and the current optimal connection action sequence of the hydraulic support multi-agent.
Citation Information
Patent Citations
Artificial intelligent type hydraulic support electric-hydraulic control system
CN103352713A
Multi-mobile emergency power supply toughness optimization scheduling method based on data driving
CN118449131A