Equipment cluster maintenance decision-making method based on multi-agent deep reinforcement learning

The equipment cluster maintenance decision-making method based on multi-agent deep reinforcement learning solves the problem of maintenance decision-making for equipment clusters under complex conditions, realizes adaptive optimization and collaborative maintenance of equipment clusters, improves mission performance and reduces maintenance costs.

CN121998620APending Publication Date: 2026-05-08UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF ELECTRONICS SCI & TECH OF CHINA
Filing Date
2026-01-27
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing equipment cluster maintenance decision-making methods are difficult to achieve efficient, collaborative, and intelligent maintenance decisions under complex task requirements and uncertain degradation conditions. In particular, when there are many pieces of equipment, high state dimensions, and dynamically changing task requirements, the computational complexity is high and the adaptive learning ability is lacking.

Method used

A maintenance decision model for equipment clusters is constructed using a multi-agent deep reinforcement learning approach. The equipment cluster is represented as multiple independent agents at geographically deployed locations. The optimal maintenance strategy is solved through a multi-agent deep reinforcement learning algorithm, and collaborative optimization is achieved using a centralized value assessment network. A local policy network is constructed and trained in a simulation environment, and finally, adaptive decision-making is carried out in actual operation.

Benefits of technology

It has achieved adaptive optimization of equipment cluster maintenance strategy, improved the overall mission performance of equipment cluster, reduced maintenance costs, and realized real-time decision-making and autonomous learning under complex conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998620A_ABST
    Figure CN121998620A_ABST
Patent Text Reader

Abstract

The invention discloses an equipment cluster maintenance decision-making method based on multi-agent deep reinforcement learning, and aims to solve the problem of collaborative maintenance of equipment clusters deployed in a multi-site distributed manner under the condition of dynamic change of task requirements. According to the method, on the basis of equipment health status, performance level and task requirements, an equipment cluster maintenance decision process is modeled as a Markov decision process, a centralized training-distributed execution multi-agent deep reinforcement learning framework is adopted, a local strategy network is constructed for each geographical deployment site, and the maintenance decision process is optimized. And meanwhile, a centralized state value evaluation network is utilized to evaluate the overall operation state of the equipment cluster, so that collaborative optimization of equipment cluster maintenance decisions deployed across sites is realized. By training a multi-agent strategy in a simulation environment and performing online execution in actual operation, each agent can autonomously generate a maintenance strategy based on a local equipment state, and comprehensive optimal control of overall task capability and maintenance cost is realized on a system level. According to the method provided by the invention, the self-adaptive optimization of the equipment cluster maintenance strategy can be realized under the complex conditions of a large number of equipment, high state dimension and dynamic change of task requirements, so that the overall task performance of the equipment cluster is improved, and the maintenance cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of equipment operation and maintenance support technology, specifically involving an equipment cluster maintenance decision-making method based on multi-agent deep reinforcement learning. Background Technology

[0002] With the systematization, networking, and clustering of high-end equipment, modern missions are typically accomplished collaboratively by multiple pieces of equipment distributed across different geographical locations, forming equipment clusters with spatial distribution characteristics and mission coupling relationships, such as aviation equipment clusters, ship formations, and distributed manufacturing systems. During mission execution, each equipment unit continuously experiences performance degradation and random failures, and its operational status directly affects the overall mission capability of the cluster. Against this backdrop, how to implement efficient, collaborative, and intelligent maintenance decisions for distributed equipment clusters under complex mission requirements and uncertain degradation conditions has become a key technical problem urgently needing to be solved in the field of intelligent equipment operation and maintenance.

[0003] Existing equipment cluster maintenance decisions typically rely on rule-based maintenance strategies, threshold-based state-based maintenance methods, or centralized optimization scheduling methods. These methods often focus on individual equipment units, failing to adequately consider the coupling relationship between equipment status and mission capabilities across different locations and clusters. Furthermore, when facing cluster scenarios with numerous pieces of equipment, vast state spaces, and dynamically changing mission requirements, centralized modeling and optimization methods are prone to the "curse of dimensionality," with computational complexity increasing exponentially with equipment size, making real-time decision-making difficult. In addition, most existing methods rely on manually set rules or static model parameters, lacking the ability to adaptively learn from operational data and continuously optimize cluster maintenance strategies, making it difficult to meet long-term mission support requirements in complex and uncertain environments. Therefore, there is an urgent need for a method that can adaptively optimize equipment cluster maintenance decisions through agent collaboration and autonomous learning in distributed equipment cluster environments, thereby improving the mission efficiency of equipment clusters. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a multi-agent deep reinforcement learning-based equipment cluster maintenance decision-making method. This method can adaptively optimize equipment cluster maintenance strategies under complex conditions such as a large number of equipment, high state dimensions, and dynamically changing task requirements, thereby improving the overall task performance of the equipment cluster and reducing maintenance costs.

[0005] The technical solution adopted in this invention is: a method for equipment cluster maintenance decision-making based on multi-agent deep reinforcement learning, the specific steps of which are as follows:

[0006] Step 1: Collect the spatial distribution structure of the equipment cluster, the number of equipment, and the task coordination relationship between equipment; obtain the performance indicators and status monitoring data of each piece of equipment; and evaluate the health status and corresponding performance level of each piece of equipment based on the status monitoring data.

[0007] Step 2: Based on the health status and task requirements of each piece of equipment in the cluster, construct an equipment cluster maintenance decision model and model it as a Markov decision process. Define the system state space, maintenance action space, state transition probability function, reward function and Bellman equation to characterize the dynamic relationship between equipment status, maintenance behavior and system performance.

[0008] Step 3: Use a multi-agent deep reinforcement learning algorithm to solve the optimal maintenance strategy for the equipment cluster. Each geographical deployment location in the equipment cluster is treated as an independent agent. A policy network based on a deep neural network is constructed for each agent to output maintenance actions according to the local equipment status. At the same time, a centralized value evaluation network is constructed to evaluate the long-term benefits of joint maintenance decisions based on the global equipment status, thereby achieving collaborative optimization among multiple agents.

[0009] Step 4: Construct a simulation environment for the operation and degradation of equipment clusters. During the offline training phase, each agent interacts with the simulation environment to collect sample data such as equipment status, maintenance actions, system rewards, and the status at the next moment. The advantage function of the joint strategy is calculated based on the centralized value assessment network, and the strategy network parameters of each agent are updated using the strategy gradient and shearing optimization method, so that the multiple agents gradually learn to form the optimal collaborative maintenance strategy.

[0010] Step 5: During the actual operation of the equipment cluster, the status information of each piece of equipment is collected in real time. Each agent independently outputs local maintenance decisions based on the trained policy network, and continuously and adaptively updates the maintenance strategy as the equipment status and task requirements change dynamically.

[0011] Furthermore, step 2 is specifically as follows:

[0012] The equipment cluster is represented as a system consisting of multiple geographically deployed locations, where each location Includes several equipment units Each piece of equipment is a multi-state unit, and its set of health states is as follows: The higher the health status value, the higher the equipment's health level and the better its usability. Each piece of equipment... The state at time t is defined as follows The state transitions of the equipment follow a discrete-time Markov process, determined by the state transitions. Transition to state The probability is:

[0013] (1)

[0014] Furthermore, the one-step transition probability matrix of the equipment state is defined as:

[0015] (2)

[0016] For every health status Define the corresponding performance function This is used to characterize the contribution of equipment to the system's mission capabilities under this health condition, then the location... Equipment clusters in The overall performance at any given time is defined as follows:

[0017] (3)

[0018] The entire equipment cluster system in The overall performance at time step is defined as follows:

[0019] (4)

[0020] Equipment clusters The task requirements that need to be met at all times are random variables. It follows a normal distribution. and mean It varies with time and follows a discrete-time Markov process, with its state space being: ,in This represents the number of discrete levels of the mean task requirement. The state transitions at the mean task requirement level are defined by the one-step transition probability matrix:

[0021] (5)

[0022] Equipment in The maintenance decision variable at any given time is defined as follows: , Indicate location Medium equipment exist Maintenance should be carried out at all times. This indicates that the equipment continues to operate. When equipment is selected for maintenance, it will be restored to its optimal health state in the next moment. The goal of equipment cluster maintenance decisions is to achieve this within a finite mission cycle. The system formulates maintenance strategies based on equipment status to maximize the mission benefits of the equipment cluster. This is modeled as a Markov decision process, with the specific definitions of the state space, action space, state transition probability function, reward function, and Bellman equation as follows:

[0023] (a) State space: Equipment cluster in The state at any given moment is defined as the set of health states of all equipment:

[0024] (6)

[0025] Correspondingly, the state space is defined as:

[0026] (7)

[0027] (b) Operational Space: Equipment clusters in A moment's action is defined as the collection of all equipment maintenance actions:

[0028] (8)

[0029] Correspondingly, the action space is defined as:

[0030] (9)

[0031] (c) State transition probability function: The state transition of the equipment cluster system is determined by the maintenance strategy of each piece of equipment. If in If equipment is maintained regularly, it will be restored to its optimal health state in the next moment. Otherwise based on formula The state transition probability matrix is ​​used to perform state degradation. Location Medium equipment The state transition probability function can be expressed as:

[0032] (10)

[0033] If the degradation of each piece of equipment within the cluster is independent, then the state transition probability function of the cluster system is defined as:

[0034] (11)

[0035] (d) Reward function: In Task requirements at all times Let be a random variable, and let be the overall performance of the equipment cluster system. Therefore, the task completion volume and the task demand gap are defined as follows:

[0036] (12)

[0037] and

[0038] (13)

[0039] Define the unit task completion benefit coefficient as The unit task demand gap penalty coefficient is ,but The rewards for completing tasks and the penalties for missing task requirements for the constantly equipped cluster are as follows:

[0040] (14)

[0041] and

[0042] (15)

[0043] The maintenance cost of an equipment cluster includes the maintenance cost of individual equipment and fixed costs. The cost of performing maintenance activities on a single piece of equipment is... ; local location exist A fixed cost will be triggered when at least one piece of equipment is under maintenance at any given time. .make

[0044] (16)

[0045] The total maintenance cost is:

[0046] (17)

[0047] Based on the definitions of task rewards, gap penalties, and maintenance costs for equipment cluster systems, the system reward function is defined as follows:

[0048] (18)

[0049] (e) Bellman equation: in the time domain of finite programming Inside, defined from The optimal state value function from time t is:

[0050] (19)

[0051] Then it satisfies the Bellman optimality equation in the finite-time domain:

[0052] (20)

[0053] And at the end of the planning period The system state-value function is determined by the system reward at that moment:

[0054] (twenty one)

[0055] The corresponding optimal maintenance strategy is:

[0056] (twenty two)

[0057] Furthermore, step 3 is specifically as follows:

[0058] Geographical deployment locations within the equipment cluster Each is modeled as an independent intelligent agent, and each agent can only observe the operational status of equipment within its local point. At that moment, the Local observations of an agent are defined as those deployed at a location. Equipment status set:

[0059] (twenty three)

[0060] Construct a local policy network for each agent It is based on local observation and time index As input, output the maintenance action vectors for each piece of equipment at that location:

[0061] (twenty four)

[0062] The equipment cluster consists of local maintenance actions of each intelligent agent. Joint maintenance decision-making at all times:

[0063] (25)

[0064] Further build a centralized value assessment network Its input is the global state of the equipment cluster. With time index , is used to output the value function of the system.

[0065] The multi-agent reinforcement learning algorithm is executed under a centralized training-distributed execution framework. All agents share the system-level value signal given by the centralized value evaluation network, which is used to characterize the contribution of maintenance decisions at various locations to the overall performance of the equipment cluster, thereby achieving collaborative consistency in equipment maintenance decisions at multiple locations.

[0066] Furthermore, step 4 is specifically as follows:

[0067] In an equipment cluster operation simulation environment, a multi-agent proximal policy optimization algorithm with centralized training and distributed execution is used to jointly train the policy networks of each agent and the centralized value evaluation network. Based on the interaction between the current policy networks of each agent and the simulation environment, trajectory sample sequences are collected within the task cycle. The advantage function is calculated based on a centralized state-value function network to measure the improvement in benefits brought about by joint maintenance actions under the system state. The advantage function is expressed in the form of generalized advantage estimation as follows:

[0068] (26)

[0069] in,

[0070] (27)

[0071] in, As a discount factor, These are the parameters for estimating the advantage.

[0072] For each intelligent agent The probability ratio is constructed based on the action probabilities output by its local policy network:

[0073] (28)

[0074] in, This is the historical policy network for sampling trajectories.

[0075] The policy network of each agent is updated using a shearing policy objective function, which is defined as follows:

[0076] (29)

[0077] in, This is the shearing threshold.

[0078] The parameters of the centralized state-value function network are updated using a value regression loss function. The value regression loss function is defined as:

[0079] (30)

[0080] in, The objective is a value regression target constructed based on the trajectory sample returns.

[0081] The collected trajectory samples are divided into multiple mini-batches and iteratively optimized in multiple rounds, with the policy network parameters of each agent being updated alternately. With centralized state value function network parameters This leads to a collaborative maintenance strategy for equipment clusters.

[0082] Furthermore, step 5 is specifically as follows:

[0083] During actual equipment operation, the trained network outputs the optimal maintenance strategy for online decision-making, as detailed below:

[0084] Step 51: During the actual operation of the equipment cluster, collect real-time health status data of equipment at each geographical deployment location, and construct local observation data for each location's intelligent agent. ;

[0085] Step 52: Local observation Enter the policy network for the corresponding location. Each intelligent agent independently outputs maintenance actions for the equipment within that location. ;

[0086] Step 53: Combine the maintenance actions output by each intelligent agent to form a joint maintenance decision for the equipment cluster. ;

[0087] Step 54: Based on the joint maintenance decision Perform maintenance on the corresponding equipment in the equipment cluster and continue operating the equipment cluster until the next decision-making opportunity.

[0088] Step 55: Update the equipment status based on the equipment cluster operation data, and repeatedly execute steps 51 to 54 to enable the equipment cluster to continuously perform adaptive maintenance scheduling according to the trained collaborative maintenance strategy throughout the entire mission cycle.

[0089] The beneficial effects of this invention are as follows: Based on equipment health status, performance level, and mission requirements, the method of this invention models the equipment cluster maintenance decision-making process as a Markov decision process. It employs a multi-agent deep reinforcement learning framework with centralized training and distributed execution to construct local policy networks for each geographically deployed location. Simultaneously, a centralized state value evaluation network is used to uniformly evaluate the overall operational status of the equipment cluster, thereby achieving collaborative optimization of equipment maintenance decisions across locations. By training the multi-agent policies in a simulation environment and executing them online in actual operation, each location can autonomously generate maintenance decisions based on its local equipment status, achieving comprehensive optimal control of overall mission capability and maintenance cost at the system level. This invention's method can achieve adaptive optimization of equipment cluster maintenance strategies under complex conditions of numerous equipment, high state dimensions, and dynamically changing mission requirements, thereby improving the overall mission performance of the equipment cluster and reducing maintenance costs. Attached Figure Description

[0090] Figure 1 This is a flowchart of an equipment cluster maintenance decision-making method based on multi-agent deep reinforcement learning according to the present invention.

[0091] Figure 2 This is a schematic diagram of the multi-agent deep reinforcement learning algorithm framework in an embodiment of the present invention.

[0092] Figure 3 This is a schematic diagram of the algorithm training curve in an embodiment of the present invention. Detailed Implementation

[0093] The technical solution of the invention will be further described below with reference to the accompanying drawings and embodiments.

[0094] like Figure 1 The flowchart of an equipment cluster maintenance decision-making method based on multi-agent deep reinforcement learning according to the present invention is shown below. The specific steps are as follows:

[0095] Step 1: Collect the spatial distribution structure of the equipment cluster, the number of equipment, and the task coordination relationship between equipment; obtain the performance indicators and status monitoring data of each piece of equipment; and evaluate the health status and corresponding performance level of each piece of equipment based on the status monitoring data.

[0096] Step 2: Based on the health status and task requirements of each piece of equipment in the cluster, construct an equipment cluster maintenance decision model and model it as a Markov decision process. Define the system state space, maintenance action space, state transition probability function, reward function and Bellman equation to characterize the dynamic relationship between equipment status, maintenance behavior and system performance.

[0097] Step 3: Use a multi-agent deep reinforcement learning algorithm to solve the optimal maintenance strategy for the equipment cluster. Each geographical deployment location in the equipment cluster is treated as an independent agent. A policy network based on a deep neural network is constructed for each agent to output maintenance actions according to the local equipment status. At the same time, a centralized value evaluation network is constructed to evaluate the long-term benefits of joint maintenance decisions based on the global equipment status, thereby achieving collaborative optimization among multiple agents.

[0098] Step 4: Construct a simulation environment for the operation and degradation of equipment clusters. During the offline training phase, each agent interacts with the simulation environment to collect sample data such as equipment status, maintenance actions, system rewards, and the status at the next moment. The advantage function of the joint strategy is calculated based on the centralized value assessment network, and the strategy network parameters of each agent are updated using the strategy gradient and shearing optimization method, so that the multiple agents gradually learn to form the optimal collaborative maintenance strategy.

[0099] Step 5: During the actual operation of the equipment cluster, the status information of each piece of equipment is collected in real time. Each agent independently outputs local maintenance decisions based on the trained policy network, and continuously and adaptively updates the maintenance strategy as the equipment status and task requirements change dynamically.

[0100] In this embodiment, step 1 is specifically as follows:

[0101] This embodiment uses a manufacturing equipment cluster system as an example for illustration. This embodiment is only used to illustrate the implementation process of the method of the present invention and does not constitute a limitation on the type of equipment or the application field. The manufacturing equipment cluster includes 8 manufacturing equipment units in two factories. Each factory deploys 4 manufacturing equipment units with the same function to jointly complete production tasks. Each manufacturing equipment unit has 4 health states. State 1 represents a complete failure state, state 4 represents the optimal operating state, and the remaining states represent different degrees of performance degradation. The production capacity of equipment varies under different health states, defined as... ,in Indicates the state The daily production batch of the equipment is determined. The manufacturing equipment gradually degrades during continuous operation, and its health state evolves over time following a discrete-time Markov process. With a time step of one day, the equipment progresses from state to condition... Transition to state The probability is:

[0102] (1)

[0103] Furthermore, the state transition probability matrix of the equipment is as follows:

[0104] (2)

[0105] The manufacturing equipment cluster needs to complete a certain scale of production tasks every day. The daily production demand is It follows a normal distribution. The average demand level A time-varying Markov process that follows a finite state has the following state space: The state transition probability matrix for each day is:

[0106] (3)

[0107] Revenue per unit batch of products is Yuan, the penalty for the unit not completing the task is Yuan. This refers to the cost of performing maintenance on a single piece of equipment to restore it to its optimal condition, and incurring maintenance costs. Yuan. Additionally, if at least one piece of equipment is repaired on the same day at the same factory, a fixed factory-level maintenance start-up cost must be paid. Yuan is used for personnel scheduling, logistics, and ensuring resource consumption.

[0108] In this embodiment, step 2 is specifically as follows:

[0109] The production capacity of each factory is the sum of the production capacities of all equipment within the factory, that is:

[0110] (4)

[0111] The entire equipment cluster system in The total production capacity at any given time is:

[0112] (5)

[0113] Equipment in The maintenance decision variable at any given time is defined as follows: , Indicates factory Medium equipment exist Maintenance should be carried out at all times. This indicates that the equipment continues to operate. When equipment is selected for maintenance, it will be restored to its optimal health state in the next moment. The goal of equipment cluster maintenance decisions is to achieve this within a finite mission cycle. Within a given day, maintenance strategies are formulated based on equipment status to maximize the mission benefits of the equipment cluster. This is modeled as a Markov decision process, with the specific definitions of the state space, action space, state transition probability function, reward function, and Bellman equation as follows:

[0114] (a) State space: Equipment cluster in The state at any given moment is defined as the set of health states of all equipment:

[0115] (6)

[0116] Correspondingly, the state space is defined as:

[0117] (7)

[0118] (b) Operational Space: Equipment clusters in A moment's action is defined as the collection of all equipment maintenance actions:

[0119] (8)

[0120] Correspondingly, the action space is defined as:

[0121] (9)

[0122] (c) State transition probability function: The state transition of the equipment cluster system is determined by the maintenance strategy of each piece of equipment. If in If equipment is maintained regularly, it will be restored to its optimal health state in the next moment. Otherwise based on formula The state transition probability matrix is ​​used to perform state degradation. Location Medium equipment The state transition probability function can be expressed as:

[0123] (10)

[0124] If the degradation of each piece of equipment within the cluster is independent, then the state transition probability function of the cluster system is defined as:

[0125] (11)

[0126] (d) Reward function: In Task requirements at all times Let be a random variable, and let be the overall performance of the equipment cluster system. Therefore, the task completion volume and the task demand gap are defined as follows:

[0127] (12)

[0128] and

[0129] (13)

[0130] Define the unit task completion benefit coefficient as The unit task demand gap penalty coefficient is ,but The rewards for completing tasks and the penalties for missing task requirements for the constantly equipped cluster are as follows:

[0131] (14)

[0132] and

[0133] (15)

[0134] The maintenance cost of an equipment cluster includes the maintenance cost of individual equipment and fixed costs. The cost of performing maintenance activities on a single piece of equipment is... ; local location exist A fixed cost will be triggered when at least one piece of equipment is under maintenance at any given time. .make

[0135] (16)

[0136] The total maintenance cost is:

[0137] (17)

[0138] Based on the definitions of task rewards, gap penalties, and maintenance costs for equipment cluster systems, the system reward function is defined as follows:

[0139] (18)

[0140] (e) Bellman equation: in the time domain of finite programming Inside, defined from The optimal state value function from time t is:

[0141] (19)

[0142] Then it satisfies the Bellman optimality equation in the finite-time domain:

[0143] (20)

[0144] And at the end of the planning period The system state value function is determined by the system reward at that moment:

[0145] (twenty one)

[0146] The corresponding optimal maintenance strategy is:

[0147] (twenty two)

[0148] In this embodiment, step 3 is specifically as follows:

[0149] The equipment cluster includes various factories Each is modeled as an independent intelligent agent, and each agent can only observe the operating status of the equipment within its own factory. At that moment, the Local observations by an agent are defined as those deployed in the factory. Equipment status set:

[0150] (twenty three)

[0151] Construct a local policy network for each agent It is based on local observation and time index As input, output the maintenance action vectors for each piece of equipment at that location:

[0152] (twenty four)

[0153] Equipment clusters are composed of local maintenance actions of each intelligent agent. Joint maintenance decision-making at all times:

[0154] (25)

[0155] Further build a centralized value assessment network Its input is the global state of the equipment cluster. With time index This is used to output the system's value function. The multi-agent deep reinforcement learning framework used in this embodiment is as follows: Figure 2 As shown, the policy network adopts a multilayer perceptron structure, including two fully connected hidden layers, each containing 128 neurons; the value network also adopts a multilayer perceptron structure, including two fully connected hidden layers, each containing 256 neurons.

[0156] The multi-agent reinforcement learning algorithm is executed under a centralized training-distributed execution framework. All agents share the system-level value signal given by the centralized value evaluation network, which is used to characterize the contribution of maintenance decisions at various locations to the overall performance of the equipment cluster, thereby achieving collaborative consistency in equipment maintenance decisions across multiple plants.

[0157] In this embodiment, step 4 is specifically as follows:

[0158] In an equipment cluster operation simulation environment, a multi-agent proximal policy optimization algorithm with centralized training and distributed execution is used to jointly train the policy networks of each agent and the centralized value evaluation network. Based on the interaction between the current policy networks of each agent and the simulation environment, trajectory sample sequences are collected within the task cycle. The advantage function is calculated based on a centralized state-value function network to measure the improvement in benefits brought about by joint maintenance actions under the system state. The advantage function is expressed in the form of generalized advantage estimation as follows:

[0159] (26)

[0160] in,

[0161] (27)

[0162] in, As a discount factor, These are the parameters for estimating the advantage.

[0163] For each intelligent agent The probability ratio is constructed based on the action probabilities output by its local policy network:

[0164] (28)

[0165] in, This is the historical policy network for sampling trajectories.

[0166] The policy network of each agent is updated using a shearing policy objective function, which is defined as follows:

[0167] (29)

[0168] in, This is the shearing threshold.

[0169] The parameters of the centralized state-value function network are updated using a value regression loss function. The value regression loss function is defined as:

[0170] (30)

[0171] in, The objective is a value regression target constructed based on the trajectory sample returns.

[0172] The collected trajectory samples are divided into multiple mini-batches and iteratively optimized in multiple rounds, with the policy network parameters of each agent being updated alternately. With centralized state value function network parameters This leads to a collaborative maintenance strategy for equipment clusters.

[0173] In this embodiment, step 5 is specifically as follows:

[0174] During actual equipment operation, the trained network outputs the optimal maintenance strategy for online decision-making, as detailed below:

[0175] Step 51: During the actual operation of the equipment cluster, collect real-time health status data of equipment in each factory and construct local observations of local intelligent agents. ;

[0176] Step 52: Local observation Enter the policy network for the corresponding location. Each intelligent agent independently outputs maintenance actions for the equipment within that location. ;

[0177] Step 53: Combine the maintenance actions output by each intelligent agent to form a joint maintenance decision for the equipment cluster. ;

[0178] Step 54: Based on the joint maintenance decision Perform maintenance on the corresponding equipment in the equipment cluster and continue operating the equipment cluster until the next decision-making opportunity.

[0179] Step 55: Update the equipment status based on the equipment cluster operation data, and continuously execute steps 51 to 54 to enable the equipment cluster to continuously perform adaptive maintenance scheduling according to the trained collaborative maintenance strategy throughout the entire mission cycle.

[0180] The training curve of the multi-agent deep reinforcement learning algorithm in this embodiment is as follows: Figure 3 As shown, after approximately 200 iterations, the algorithm's curve tends to stabilize and maintains a high return level, indicating that the multi-agent strategy converges to a stable collaborative maintenance decision-making scheme. Furthermore, the method of this invention is compared and evaluated with a traditional heuristic strategy based on maintenance thresholds, and the results are shown in Table 1 (unit: 10). 6Table 1 shows that the total task benefit obtained by the multi-agent deep reinforcement learning strategy in the equipment cluster task execution is significantly higher than that of the heuristic strategy, indicating that the method of the present invention can adaptively learn the optimal trade-off between production benefits, task assurance and maintenance costs, thereby achieving a higher level of system-level collaborative fault recovery and task assurance capabilities.

[0181] Table 1

[0182] method Total task rewards Production revenue Out-of-stock costs Repair costs Deep reinforcement learning 7.13 9.12 4.63 1.53 Heuristic strategies 7.21 9.12 3.46 1.56

[0183] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. A method for equipment cluster maintenance decision-making based on multi-agent deep reinforcement learning, characterized in that, Includes the following steps: Step 1: Collect the spatial distribution structure of the equipment cluster, the number of equipment, and the task coordination relationship between equipment; obtain the performance indicators and status monitoring data of each piece of equipment; and evaluate the health status and corresponding performance level of each piece of equipment based on the status monitoring data. Step 2: Based on the health status and task requirements of each piece of equipment in the cluster, construct an equipment cluster maintenance decision model and model it as a Markov decision process. Define the system state space, maintenance action space, state transition probability function, reward function and Bellman equation to characterize the dynamic relationship between equipment status, maintenance behavior and system performance. Step 3: Use a multi-agent deep reinforcement learning algorithm to solve the optimal maintenance strategy for the equipment cluster. Each geographical deployment location in the equipment cluster is treated as an independent agent. A policy network based on a deep neural network is constructed for each agent to output maintenance actions according to the local equipment status. At the same time, a centralized value evaluation network is constructed to evaluate the long-term benefits of joint maintenance decisions based on the global equipment status, thereby achieving collaborative optimization among multiple agents. Step 4: Construct a simulation environment for the operation and degradation of equipment clusters. During the offline training phase, each agent interacts with the simulation environment to collect sample data such as equipment status, maintenance actions, system rewards, and the status at the next moment. The advantage function of the joint strategy is calculated based on the centralized value assessment network, and the strategy network parameters of each agent are updated using the strategy gradient and shearing optimization method, so that the multiple agents gradually learn to form the optimal collaborative maintenance strategy. Step 5: During the actual operation of the equipment cluster, the status information of each piece of equipment is collected in real time. Each agent independently outputs local maintenance decisions based on the trained policy network, and continuously and adaptively updates the maintenance strategy as the equipment status and task requirements change dynamically.

2. The equipment cluster maintenance decision-making method based on multi-agent deep reinforcement learning according to claim 1, characterized in that, Step 2 is described in detail below: The equipment cluster is represented as a system consisting of multiple geographically deployed locations, where each location Includes several equipment units Each piece of equipment is a multi-state unit, and its set of health states is as follows: The higher the health status value, the higher the equipment's health level and the better its usability. Each piece of equipment... The state at time t is defined as follows The state transitions of the equipment follow a discrete-time Markov process, determined by the state transitions. Transition to state The probability is: (1) Furthermore, the one-step transition probability matrix of the equipment state is defined as: (2) For every health status Define the corresponding performance function This is used to characterize the contribution of equipment to the system's mission capabilities under this health condition, then the location... Equipment clusters in The overall performance at any given time is defined as follows: (3) The entire equipment cluster system in The overall performance at time step is defined as follows: (4) Equipment clusters in The task requirements that need to be met at all times are random variables. It follows a normal distribution. and mean It varies with time and follows a discrete-time Markov process, with its state space being: ,in This represents the number of discrete levels of the mean task requirement. The state transitions at the mean task requirement level are defined by the one-step transition probability matrix: (5) Equipment in The maintenance decision variable at any given time is defined as follows: , Indicate location Medium equipment exist Maintenance should be carried out at all times. This indicates that the equipment continues to operate. When equipment is selected for maintenance, it will be restored to its optimal health state in the next moment. The goal of equipment cluster maintenance decisions is to achieve this within a finite mission cycle. The system formulates maintenance strategies based on equipment status to maximize the mission benefits of the equipment cluster. This is modeled as a Markov decision process, with the specific definitions of the state space, action space, state transition probability function, reward function, and Bellman equation as follows: (a) State space: Equipment cluster in The state at any given moment is defined as the set of health states of all equipment: (6) Correspondingly, the state space is defined as: (7) (b) Operational Space: Equipment clusters in A moment's action is defined as the collection of all equipment maintenance actions: (8) Correspondingly, the action space is defined as: (9) (c) State transition probability function: The state transition of the equipment cluster system is determined by the maintenance strategy of each piece of equipment. If in If equipment is maintained regularly, it will be restored to its optimal health state in the next moment. Otherwise based on formula The state transition probability matrix is ​​used to perform state degradation. Location Medium equipment The state transition probability function can be expressed as: (10) If the degradation of each piece of equipment within the cluster is independent, then the state transition probability function of the cluster system is defined as: (11) (d) Reward function: In Task requirements at all times Let be a random variable, and let be the overall performance of the equipment cluster system. Therefore, the task completion volume and the task demand gap are defined as follows: (12) and (13) Define the unit task completion benefit coefficient as The unit task demand gap penalty coefficient is ,but The rewards for completing tasks and the penalties for missing task requirements for the constantly equipped cluster are as follows: (14) and (15) The maintenance cost of an equipment cluster includes the maintenance cost of individual equipment and fixed costs. The cost of performing maintenance activities on a single piece of equipment is... ; local location exist A fixed cost will be triggered when at least one piece of equipment is under maintenance at any given time. .make (16) The total maintenance cost is: (17) Based on the definitions of task rewards, gap penalties, and maintenance costs for equipment cluster systems, the system reward function is defined as follows: (18) (e) Bellman equation: in the time domain of finite programming Inside, defined from The optimal state value function from time t is: (19) Then it satisfies the Bellman optimality equation in the finite-time domain: (20) And at the end of the planning period The system state value function is determined by the system reward at that moment: (21) The corresponding optimal maintenance strategy is: (22)。 3. The equipment cluster maintenance decision-making method based on multi-agent deep reinforcement learning according to claim 1, characterized in that, Step 3 is described in detail below: Geographical deployment locations within the equipment cluster Each is modeled as an independent intelligent agent, and each agent can only observe the operational status of equipment within its local point. At that moment, the Local observations of an agent are defined as those deployed at a location. Equipment status set: (23) Construct a local policy network for each agent It is based on local observation and time index As input, output the maintenance action vectors for each piece of equipment at that location: (24) Equipment clusters are composed of local maintenance actions of each intelligent agent. Joint maintenance decision-making at all times: (25) Further build a centralized value assessment network Its input is the global state of the equipment cluster. With time index , is used to output the value function of the system.

4. The multi-agent reinforcement learning algorithm is executed under a centralized training-distributed execution framework. All agents share the system-level value signal given by the centralized value evaluation network, which is used to characterize the contribution of maintenance decisions at various locations to the overall performance of the equipment cluster, thereby achieving collaborative consistency of equipment maintenance decisions at multiple locations.

5. The equipment cluster maintenance decision-making method based on multi-agent deep reinforcement learning according to claim 1, characterized in that, Step 4 is as follows: In an equipment cluster operation simulation environment, a multi-agent proximal policy optimization algorithm with centralized training and distributed execution is used to jointly train the policy networks of each agent and the centralized value evaluation network. Based on the interaction between the current policy networks of each agent and the simulation environment, trajectory sample sequences are collected within the task cycle. The advantage function is calculated based on a centralized state-value function network to measure the improvement in benefits brought about by joint maintenance actions under the system state. The advantage function is expressed in the form of generalized advantage estimation as follows: (26) in, (27) in, As a discount factor, These are the parameters for estimating the advantage.

6. For each agent The probability ratio is constructed based on the action probabilities output by its local policy network: (28) in, This is the historical policy network for sampling trajectories.

7. The policy network of each agent is updated using a shearing policy objective function, which is defined as follows: (29) in, This is the shearing threshold.

8. Update the parameters of the centralized state-value function network using the value regression loss function. The value regression loss function is defined as: (30) in, The objective is a value regression target constructed based on the trajectory sample returns.

9. Divide the collected trajectory samples into multiple mini-batches and perform multiple rounds of iterative optimization, alternately updating the policy network parameters of each agent. With centralized state value function network parameters This leads to a collaborative maintenance strategy for equipment clusters.

10. The equipment cluster maintenance decision-making method based on multi-agent deep reinforcement learning according to claim 1, characterized in that, Step 5 is described in detail below: During actual equipment operation, the trained network outputs the optimal maintenance strategy for online decision-making, as detailed below: Step 51: During the actual operation of the equipment cluster, collect real-time health status data of equipment at each geographical deployment location, and construct local observation data for each location's intelligent agent. ; Step 52: Local observation Enter the policy network for the corresponding location. Each intelligent agent independently outputs maintenance actions for the equipment within that location. ; Step 53: Combine the maintenance actions output by each intelligent agent to form a joint maintenance decision for the equipment cluster. ; Step 54: Based on the joint maintenance decision Perform maintenance on the corresponding equipment in the equipment cluster and continue operating the equipment cluster until the next decision-making opportunity. Step 55: Update the equipment status based on the equipment cluster operation data, and continuously execute steps 51 to 54 to enable the equipment cluster to continuously perform adaptive maintenance scheduling according to the trained collaborative maintenance strategy throughout the entire mission cycle.