Method for joint optimization of maintenance and inspection in manufacturing network based on deep reinforcement learning
By constructing machine reliability and quality models using deep reinforcement learning, this approach addresses the challenges of complex structures and reliability-quality interaction in large-scale manufacturing networks. It achieves a balance between economic benefits and operational risks, and provides optimal control strategies for manufacturing systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHENGZHOU UNIV
- Filing Date
- 2023-03-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing manufacturing system optimization methods are inadequate when dealing with large-scale manufacturing networks with complex structures and reliability-quality interactions. Traditional methods struggle to effectively balance economic benefits and operational risks and do not fully consider the continuous nature of quality inspection.
A joint optimization method for maintenance and inspection in manufacturing networks based on deep reinforcement learning is adopted to construct a machine reliability and quality model. The optimal strategy for quality inspection and maintenance is learned through a deep deterministic policy gradient algorithm, thereby achieving joint optimization of reliability-quality interaction behavior.
It achieves a better balance between economic benefits and operational risks in complex manufacturing networks, adapts to dynamic and diverse scenarios, and provides optimal manufacturing system control strategies.
Smart Images

Figure CN116384969B_ABST
Abstract
Description
A Joint Optimization Method for Manufacturing Network Repair and Inspection Based on Deep Reinforcement Learning Technical Field
[0001] This invention relates to the field of quality and reliability technology, and in particular to a joint optimization method for maintenance and inspection of manufacturing networks based on deep reinforcement learning. Background Technology
[0002] Mass customization brings greater production flexibility but increases the complexity of manufacturing systems. Against this backdrop, more high-tech manufacturing equipment with flexible capabilities has emerged, greatly enriching product process routes and causing manufacturing systems to exhibit complex network characteristics, with machines as nodes and work-in-process flow as edges. The increased machine flexibility and structural complexity both amplify the nonlinear characteristics of manufacturing systems, which will increase the operational and management difficulties of network-structured manufacturing systems and diminish the profit growth brought by flexible machines.
[0003] Operational control of manufacturing systems refers to optimizing system performance through production management methods, with machine maintenance and work-in-process quality inspection being two crucial management measures. Integrated optimization through joint control of production scheduling, work-in-process quality, and machine reliability has become the preferred approach for improving manufacturing system performance. Existing research has explored different integration forms, such as single preventive maintenance (PM), integration of production planning and maintenance, or integration of production planning, maintenance, and quality inspection. Currently, research on joint optimization of manufacturing systems mainly focuses on simulation-based methods. In addition, traditional dynamic programming, integer programming, and heuristic algorithms are also important methods for solving this problem. These traditional optimization methods have great potential for optimizing small-scale systems, such as single-machine manufacturing systems, simple serial manufacturing systems, or unitized manufacturing systems. However, these methods remain insufficient for optimizing large-scale manufacturing systems with complex system structures. Although heuristic algorithms such as genetic algorithms can effectively optimize multi-level serial or parallel manufacturing systems, their effectiveness in large-scale manufacturing systems with complex system structures has not yet been verified.
[0004] In many manufacturing systems, the diversity of process routes gives them the characteristics of complex networks. Besides the large system size and structural complexity, interactive behavior is another key factor contributing to the difficulty of joint optimization of manufacturing networks, especially the interaction between product quality and machine reliability. Currently, the mutual influence between reliability and quality in manufacturing systems has attracted much interest from researchers. For example, in the joint optimization of production and maintenance, the literature [Hajej Z, Rezg N, Gharbi A. Quality issue in forecasting problem of production and maintenance policy for production unit. International Journal of Production Research. 2018; 56: 6147-63.] proposes a cumulative defect rate function affected by machine failure rate; in the maintenance optimization of serial multi-station manufacturing systems, the literature [Zhou X, Lu B. Preventive maintenance scheduling for serial multi-station manufacturing systems with interaction between station reliability and product quality. Computers & Industrial Engineering. 2018; 122: 283-91] establishes a model of the mutual influence between machine reliability and product quality; in the performance evaluation of automated production lines and serial-parallel manufacturing systems, the literature [Ye Z, Cai Z, Si S, Zhang S, Yang H. Competing Failure Modeling for Performance Analysis of Automated Manufacturing Systems with Serial Structures and Imperfect Quality Inspection. IEEE Transactions on [Industrial Informatics. 2020; 16: 6476-86] A decision graph model and a stochastic model were constructed to characterize the interaction between reliability and quality.Furthermore, the paper [Wang L, Bai Y, Huang N, Wang Q. Fractal-based Reliability Measure for Heterogeneous Manufacturing Networks. IEEE Transactions on Industrial Informatics. 2019; 15: 6407-14] proposes a reliability metric for manufacturing networks using a fractal-based approach. These studies attempt to address the aforementioned problems of interaction behavior or large-scale operations; however, they only independently investigate one aspect of interaction behavior and structural complexity, failing to fully grasp the characteristics of manufacturing networks.
[0005] As mentioned above, joint optimization of manufacturing networks with complex structures and interactive behaviors is a challenging problem, especially when attempting to solve it using traditional methods. Traditional methods perform well in controlling or optimizing specific manufacturing scenarios, but their effectiveness is poor in dynamic and diverse manufacturing environments. However, the development of artificial intelligence (AI) has brought hope for effectively solving this real-time control and optimization problem. AI-based methods can achieve higher value-added manufacturing and create greater flexibility for smart, customized factories. Therefore, they have been widely applied in related fields such as risk assessment, intelligent maintenance, quality control, and dynamic loading strategies for repairable systems. Among the many AI-based methods, reinforcement learning has achieved remarkable success in various control tasks for nonlinear, high-dimensional, and dynamic systems. Due to its effectiveness, much reinforcement learning-based control research has been applied to the optimization of dynamic systems, such as maintenance optimization of serial production lines, multi-state engineering systems, joint production-maintenance optimization for degrading manufacturing systems, and coverage and connectivity maintenance of wireless sensor networks.
[0006] In research on manufacturing system operation control, structural complexity has not received sufficient attention. Furthermore, simplifying system states into single continuous or discrete types does not reflect actual production conditions. Meanwhile, discrete maintenance activities have become the mainstream choice for operational settings; for example, discrete maintenance activities can be combined with production scheduling to construct joint control problems. However, current research has not considered the continuous nature of quality inspection activities. The choice of reinforcement learning algorithms reveals the reason for this gap. Existing manufacturing system operation control algorithms perform well in finite-range Markov decision processes, but are still insufficient for large-scale manufacturing systems with complex structures and diverse state behaviors. In addition, research on non-manufacturing system control using more advanced reinforcement learning algorithms provides a good example of large-scale problems with state-behavior diversity. For example, the actor-critic algorithm and deep deterministic policy gradient algorithm for 13-component systems are applied to 14-component parallel-cascade systems. Furthermore, some scholars have considered inspection activities in some action settings, but this maintenance inspection differs significantly from quality inspection activities in manufacturing systems. In these studies, inspection is considered merely a discrete action. However, in actual production, quality inspection in manufacturing systems has a continuous action space, allowing for sampling inspection of work-in-process at any ratio within [0,1]. Most importantly, research on non-manufacturing systems does not consider the interaction behavior between components. Therefore, whether these advanced algorithms are applicable to large-scale manufacturing networks with reliability-quality interactions requires further investigation. Summary of the Invention
[0007] To balance the economic benefits and operational risks of manufacturing networks, this invention proposes a joint optimization method for maintenance and inspection of manufacturing networks based on deep reinforcement learning (DRL). Through machine learning-based reliability-quality joint control, it realizes the optimization of maintenance and quality inspection of large-scale manufacturing networks with reliability-quality interactive behavior.
[0008] The technical solution of this invention is implemented as follows:
[0009] A joint optimization method for maintenance and inspection in manufacturing networks based on deep reinforcement learning, the steps of which are as follows:
[0010] Step 1: At the machine level, considering the dynamic production speed caused by machine failure downtime, a machine reliability model that considers the impact of feed quality and a processing quality model that considers the impact of machine reliability were constructed.
[0011] Step 2: Conduct a systematic evaluation of the manufacturing network status and performance based on reliability and quality models; and build a joint optimization model for manufacturing network maintenance and quality inspection.
[0012] Step 3: At the system level, the economic operation of the manufacturing network is used as the standard for strategy evaluation. A deep deterministic strategy gradient algorithm is designed to learn the optimal strategy for quality inspection and maintenance under a given manufacturing network state.
[0013] At the machine level, the methods for constructing reliability and quality models are as follows:
[0014] Calculate dynamic production speed:
[0015] Each machine is treated as a node, and a directed acyclic graph G(V,E) is used to model the acyclic manufacturing network of n nodes, where V={v1,v2,…,v…} n} represents the set of nodes that create the network. Let i be the set of directed edges in the network; i and j are nodes.
[0016] When there are no machine downtimes in the manufacturing network, the machine's production speed is defined as the maximum production speed, denoted as P. rm =[P rm1 ,P rm2 ,…,P rmn The actual production speed of machines in the manufacturing network is denoted as P. ra (t)=[P ra1 (t),P ra2 (t),…,P ran [(t)], and satisfy P ra ≤P rm Among them, P rmn P represents the maximum production speed of the nth machine. ran (t) represents the actual production speed of the nth machine;
[0017] When the production speed of node i changes by ΔP rai At time (t), the change in production rate ΔP of the node immediately adjacent to it is... rak (t) and the change in production rate ΔP of the immediate downstream node raj (t) are respectively represented as:
[0018]
[0019]
[0020] in, Let i represent the set of immediately upstream neighbors connected to node i. This represents the set of immediately downstream nodes connected to node i; Let represent the probability that all work-in-process items flowing into node i originate from upstream node k. This represents the probability that the work-in-process flowing out of node i will flow into downstream node j;
[0021] Calculate dynamic maintenance costs:
[0022] When considering the dynamic nature of production speed, the failure rate when processing qualified feed materials is defined as the basic failure rate r. b (t):
[0023] r b (t)=(β / α)·(t r / α) β-1 ;
[0024] Where α is the scale parameter and β is the shape parameter; It is the relative running time of the machine calculated based on the maximum production speed; t is the actual running time of the machine.
[0025] Considering the potential impact of substandard feed, the actual failure rate r(t) is defined as:
[0026]
[0027] Where, Δr i′ Let N(t) be the cumulative failure rate increment caused by defective feed into the machine, and let N(t) be the number of defective work-in-process products processed by the machine within the time interval [0, t). The probability distribution function F(t) for machine failure is derived as follows:
[0028] F(t) = 1 - exp(-(t) r / α) β -∫Δr(t)dt);
[0029] in, Let Δr(t) be the cumulative failure rate increment caused by defective feed material during machine processing up to time t; the integral of Δr(t) is calculated using the following formula:
[0030]
[0031] Among them, t i′ This represents the actual occurrence time of the i′th failure rate increment;
[0032] Assume that the maintenance time for corrective and preventive maintenance of the machine follows the following normal distributions: and And μ cm ≥μ pm , The unit time maintenance costs for corrective maintenance and preventive maintenance are c, respectively. cm and c pmThen, the total maintenance cost of the machine during the time interval [0, t) is c. m (t) can be expressed as:
[0033]
[0034] Where, N cm (t) represents the number of corrective maintenance operations performed on the machine within the time interval [0, t), N. pm (t) represents the number of preventative maintenance operations performed on the machine within the time interval [0, t), where t cm_i1 The time spent on the i1st corrective maintenance of the machine, t pm_j1 The time spent on the machine's j1st preventative maintenance;
[0035] Construct a dynamic model for processing quality and inspection activities:
[0036] Let M(t) be defined as the number of defective products generated in the time interval [0, t). M(t) is a non-homogeneous Poisson process that satisfies the intensity function λ(t).
[0037] λ(t)=ω-ε·e -δ·r(t) ;
[0038] Where ω>0 represents the maximum nonconforming product strength, ε>0 and δ>0 are both influence coefficients of the failure rate on the strength function, and ω-ε<λ(t)<ω <P ra (t); define n d Let be the number of defective products generated by the machine within the time interval [t, t+Δt), and its probability be:
[0039]
[0040] in, Let n be the expected value of non-conforming products produced within the time period [0, t), where Δt is the quality statistical period; define n q The total number of qualified products processed by the machine within the time interval [t, t+Δt) is expressed as:
[0041]
[0042] In the inspection process, the incorrect judgment of non-conforming and conforming products is classified as: Type I error, which represents incorrect rejection, with a probability of p. I Type II error, which means incorrect acceptance, has a probability of p. II Assume the sampling ratio in the detection activity is s. a s a If ∈[0,1], then the joint probability of any qualified work-in-process occurring a Type I error is s. a ·pI The joint probability of any nonconforming work-in-process causing a Type II error is s. a ·p II Define the number of Type I errors and Type II errors as n, respectively. fr and n fa And they respectively follow a binomial distribution B(n) q ,s a ·p I ) and B(n d ,s a ·p II Therefore, the number of defective work-in-process items M'(t) leaving the machine is a non-homogeneous Poisson process M(t) and a binomial distribution B(n). d ,s a ·p II The composition of n) fa The probability that a defective work-in-process item leaves the machine within the time interval [t, t+Δt) is:
[0043]
[0044] Where m′(t) is the average number of defective work-in-process items leaving the machine during the time interval [0, t]; furthermore, the probability that qualified work-in-process items are rejected during the time interval [t, t+Δt) is:
[0045]
[0046] Where D(t) represents the number of qualified work-in-process items that are rejected at time t. Indicates from n q Randomly select n from the defective products fr The number of possible combinations for each sample;
[0047] The number of products correctly identified as defective by the machine within the time period [t, t+Δt) is n. cr The number n that was correctly judged as qualified ca They are represented as follows:
[0048]
[0049] Based on this, the proportion of non-conforming products identified by the machine through inspection activities within the time period [t, t+Δt) can be obtained as follows:
[0050]
[0051] Assume the inspection cost for a single product by machine is c. insThe total cost of the machine's inspection within the time interval [t, t+Δt) is:
[0052]
[0053] Assume the cumulative value increment of a single work-in-process item from the input node of the manufacturing network to the current processing node i is v. si Then, the average process value increment v brought about by the machined work-in-process unit corresponding to node i is... ai It can be defined as:
[0054]
[0055] Among them, v sk This represents the cumulative value increment of node i's upstream node k;
[0056] The value loss caused by nonconforming work-in-process due to Type II errors in the inspection process is defined as... Therefore, the net value increment v of all work-in-process processed by machine i during the time interval [0,t) is... net_i The sum of the value increment of all correctly accepted work-in-process, the value loss of incorrectly accepted work-in-process, and the value loss of all rejected work-in-process:
[0057]
[0058] Where, n ca_i n represents the number of work-in-process items correctly accepted by the machine corresponding to node i. fa_i n represents the number of work-in-process items accepted by the machine corresponding to node i with errors. cr_i n represents the number of work-in-process items correctly rejected by the machine corresponding to node i. fr_i These represent the number of work-in-process items that were incorrectly rejected by the machine corresponding to node i.
[0059] The method for systematically evaluating the status and performance of the manufacturing network is as follows:
[0060] For a manufacturing network with n machines, construct the state matrix S. K Represents the state at time t:
[0061] S K =[Q K ;D K H K ;O K ];
[0062] Where t = KΔt, The mass state of each machine during the time interval [t-Δt, t); DK =[t r1 ,t r2 ,…,t rn H represents the degenerate state of the machine at time t. K =[h1,h2,…,h n [] represents the machine's health status at time t, h i In the range {0,1}, 1 represents a faulty state of the machine, and 0 represents a fault-free state; O K =[o1,o2,…,o n [] represents the idle state of the machine at time t, o i In {0,1}, 1 represents the idle state and 0 represents the working state;
[0063] Define reward r K To create a network from state S in time interval [t, t+Δt). K Transition to S K+1 Net gains generated during the process:
[0064]
[0065] Among them, c I_i c is the total detection cost of the machine corresponding to node i. m_i c is the total maintenance cost of the machine corresponding to node i. D It is the decision-making cost of maintenance and inspection.
[0066] To evaluate the state S during the time interval [KΔt, K′Δt), K To S K′ The cumulative performance will benefit G K Defined as the long-term reward of the manufacturing network, it is calculated through cumulative rewards:
[0067]
[0068] Where K′>K.
[0069] The method for constructing the joint optimization model for manufacturing network maintenance and quality inspection is as follows:
[0070] Quality inspection and preventive maintenance are considered as actions, denoted as... in, For quality inspection actions, For preventative maintenance;
[0071] At t = KΔt, the actions of all machines in the manufacturing network depend on state S. K Therefore, it is denoted as A. k =π(S) K), where π(·) represents the policy function:
[0072]
[0073] Where, π * (·) represents the optimal policy function, and Q(·) represents the optimal policy function in state S. K Take A K The long-term reward function during the action;
[0074] Under the optimal policy, the value function and the Q function satisfy:
[0075]
[0076] Where V(·) represents the state S K The maximum long-term return at that time.
[0077] The method described above for a Deep Deterministic Policy Gradient Algorithm (DDPG) to learn the optimal strategy for quality inspection and maintenance under a given manufacturing network state is as follows:
[0078] Step 1: Execute the current actions to simulate the operation of the manufacturing network in the time interval [KΔt, K′Δt), where t = KΔt;
[0079] Step 1.1: Evaluate the state of the manufacturing network at time t = KΔt: In the learning environment, based on the proposed dynamic reliability and quality model, evaluate the machine state at time t = KΔt, and then evaluate the state S of the manufacturing network at time t = KΔt. K , so that state S K This is provided to the Agent as an observation of the manufacturing network;
[0080] Step 1.2: Generate an action based on the current policy function π(S): an action It is possible to input state S into the Actor network μ(S) K The DDPG algorithm obtains this by adding a normally distributed random noise N. r Try permissible actions to explore better strategies, i.e., A K =π(S) K )+N r Then, according to the preventive maintenance guideline c d The preventive maintenance action is converted into a discrete executable action {0,1}, where 0 indicates that preventive maintenance is not performed and 1 indicates that preventive maintenance is performed.
[0081] Step1.3: Execute actions to simulate the operation of the manufacturing network within the time interval \([K\Delta t, K\Delta t+\Delta t)\): After obtaining the action \(A\) at the moment \(t = K\Delta t\), the quality inspection of the corresponding machine will immediately adopt a new sampling ratio during subsequent operations. K Meanwhile, preventive maintenance actions are performed on the corresponding machines. The time \(t\) of preventive maintenance is obtained from the normal distribution \(N(\mu\) , \(\sigma\) pm pm 2 pm ). If a machine fails during the period \([K\Delta t, K\Delta t+\Delta t)\), a corrective maintenance will be immediately performed at the moment of failure. The time \(t\) of corrective maintenance is obtained from the normal distribution \(N(\mu\) cm , \(\sigma\) cm 2 cm ). cm 2 )
[0082] Step1.4: Evaluate the reward \(r\): Calculate the reward \(r\), and update \(i2 = i2 + 1\); if \(i2 < K'\), return to Step1.1; otherwise, execute Step2. K i2
[0083] Step2: Obtain the state transition record: After the operation within the time interval \([K\Delta t, K'\Delta t)\), obtain the long-term return \(G\); according to the method in Step1.1, obtain the state \(S\) at \(t = K'\Delta t\); then, obtain the state transition record \(\{S\) K , \(A\) K′ , \(G\) K , \(S\) K \} and store it in the experience buffer. K K′
[0084] Step3: Update the Actor network and the Critic network: Randomly sample a small batch of \(M\) transition records from the experience buffer to update the Actor network \(\mu(S)\) and the Critic network \(Q(S,A)\); at this time, the maximum storage capacity of the experience buffer is \(L\). When the number of transition records reaches \(L\), discard the earliest record.
[0085] Step4: Judge the end condition: If the number of training episodes Episode reaches the predetermined maximum number of training episodes or obtains a stable long-term return \(G\) K , stop the training; otherwise, update the simulation period: \([K\Delta t, K'\Delta t)\leftarrow[K'\Delta t, 2K'\Delta t - K\Delta t)\) and return to Step1.
[0086] The update method of the Actor network and the Critic network is as follows:
[0087] a) Randomly select M transfer records from the experience buffer:
[0088] S3.3.2: For the transfer records i3 = 1, 2, ..., M, calculate the future target action. and target future long-term returns And set the value function objective.
[0089] b): Update the parameters θ of the Critic network Q(S,A) by minimizing the loss function. Q :
[0090]
[0091] c): Update the parameters θ of the Actor network μ(S) by maximizing the expected cumulative long-term return. μ :
[0092]
[0093] d): Update the parameters of the target actor network and the target critic network:
[0094] θ μ′ =τθ μ +(1-τ)θ μ′ ;
[0095] θ Q′ =τθ Q +(1-τ)θ Q′ ;
[0096] Where τ is the smoothing factor.
[0097] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0098] 1) This invention proposes a mathematical model for constructing a nonlinear, high-dimensional, and dynamic environment for manufacturing networks, providing an effective state transition model for the control of manufacturing networks.
[0099] 2) Based on the consideration of reliability-quality interaction behavior, this invention constructs an effective DRL model suitable for joint reliability-quality control of manufacturing networks, which can model physical systems that simultaneously contain discrete-continuous mixed states and mixed actions; this DRL model has better adaptability to dynamic and diverse manufacturing scenarios.
[0100] 3) In order to achieve reliability-quality joint control of dynamic manufacturing network, this invention constructs a deep neural network driven machine learning model based on DDPG algorithm under the established maintenance and quality inspection hybrid action space model; according to the learning results under different manufacturing scenarios, this method can achieve the optimal manufacturing system control strategy.
[0101] 4) The model proposed in this invention can effectively balance the contradiction between the economic benefits and operational risks of manufacturing networks. Attached Figure Description
[0102] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0103] Figure 1 shows the manufacturing network formed by multiple process paths.
[0104] Figure 2 illustrates the cascading effects of machine failures in nodes of a network-structured manufacturing system.
[0105] Figure 3 is a flowchart of the present invention.
[0106] Figure 4 shows the neural network structure of the DDPG algorithm of the present invention; where (a) is the Actor network and (b) is the Critic network.
[0107] Figure 5 is a flowchart of the training process of the DRL model based on the DDPG algorithm of this invention.
[0108] Figure 6 is a flowchart of the update process of the Actor and Critic networks in the DDPG algorithm of this invention.
[0109] Figure 7 shows a directed acyclic manufacturing network according to an example of the present invention.
[0110] Figure 8 shows the training trajectory of the highest returns of the DRL model and genetic algorithm of the present invention under different manufacturing scenarios; where (a) step size Δt = 50, (b) step size Δt = 100, and (c) step size Δt = 500.
[0111] Figure 9 is a scatter plot of the unit-time rewards obtained by the three trained agents.
[0112] Figure 10 shows the connectivity curves of the manufacturing network under different training DRL Agents; where (a) step size Δt = 50, (b) step size Δt = 100, and (c) step size Δt = 500. Detailed Implementation
[0113] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0114] In flexible manufacturing environments, the diversity of process routes leads to the randomness of work-in-process (WIP) flow, giving the manufacturing system a network-like characteristic. As shown in Figure 1, each machine can be viewed as a node in the manufacturing network, and the WIP flow can be seen as the edges connecting the nodes. In this context, for a given machine, its upstream machines refer to all machines preceding it in the relevant process route, and its downstream machines refer to all machines following it in the relevant process route. In a multi-level manufacturing network, the WIP flow brings raw materials to each machine at different stages, resulting in an interaction between machine reliability and feed quality. Defective feed increases the machine's failure risk; conversely, degraded machines are more likely to produce defective products and become defective feed flowing into adjacent downstream machines (NDM). Without intervention, the interaction between machine reliability and feed quality propagates along the process route, increasing the failure risk of downstream machines.
[0115] In manufacturing systems, machines have two failure modes: hard failures, which cause the machine to stop immediately upon malfunction, and soft failures, which are typically not directly detectable and can only be identified and corrected during inspection. Hard failures cause machine downtime, reducing production speed to zero. Soft failures are caused by component degradation, triggered when the cumulative degradation first exceeds a threshold. Before a soft failure occurs, machine degradation only affects the machine's processing quality and not its production speed. In real-world industrial production, machine maintenance and quality inspection are crucial activities for improving manufacturing system performance. The former allows machines to recover from failures, while the latter reduces losses promptly by cutting off the flow of defective products between different machines. When a hard failure occurs or a soft failure is detected, corrective maintenance (CM) can be implemented to repair the downtime machine. Furthermore, when degradation does not lead to a soft failure, preventive maintenance can be implemented to restore the degraded machine. On the other hand, preventive maintenance can reduce the probability of hard failures. Moreover, preventive maintenance can also improve the processing quality of machine tools.
[0116] In the above scenario, controlling the operation of the manufacturing network becomes a complex issue. When a machine fails and is not repaired in time, the impact of the failure will propagate to its upstream and downstream machines. This will reduce the production speed of these machines. Furthermore, machines may become idle (referred to as idle machines) due to congestion or starvation. As shown in Figure 2, because M5 is the only immediate downstream machine of M3, a failure of machine M5 in node 5 will cause machine M3 to become idle. And, because machine M6 can continue operating with a reduced production speed, the production speeds of machines M4 and M7 will also decrease. Further, downstream and upstream machines further away will be impacted until the failed machine is repaired. Once the failed machine is repaired, its own production speed and that of the affected machines will recover. However, fluctuations in production speed introduce more uncertainty into machine degradation, weakening the effectiveness of pre-planned maintenance and quality inspection plans.
[0117] Against this backdrop, it is essential to optimize reliability and quality control in manufacturing networks based on dynamic machine maintenance and quality inspection. Therefore, this invention proposes a joint optimization method for maintenance-inspection in manufacturing networks based on deep reinforcement learning, considering the following assumptions: (1. Considering the discrete parts manufacturing process. 2. Each machine has quality inspection activities, sampling work-in-process (WIP) at any rate of [0%, 100%]. 3. WIP has two quality states: acceptable (no defects) and unacceptable (substandard), which can be distinguished through the quality inspection process. 4. Machines have two failure modes: hard failure and soft failure. 5. Two maintenance activities will be performed: corrective maintenance and preventive maintenance, both of which can restore the machine to a new state.) This method explores optimal reliability and quality control strategies under economical operation of manufacturing networks, as shown in Figure 3. First, at the machine level, considering the dynamic production speed caused by machine downtime due to faults, a reliability model considering the impact of feed quality (QR impact) and a quality model considering the impact of machine reliability (RQ impact) are constructed. The proposed models are then used to systematically evaluate the machine state and the manufacturing network state. Secondly, at the system level, taking the economic operation of the manufacturing network as the standard for strategy evaluation, an optimization model based on deep reinforcement learning is proposed to learn the optimal strategy for quality inspection and maintenance under a given manufacturing network state.
[0118] Dynamic production speed: Treating each machine as a node, a directed acyclic graph G(V,E) is used to model the acyclic manufacturing network of n nodes, where V={v1,v2,…,v…} n} represents the set of nodes in a manufacturing network, that is, the set of machines. Let W be the set of directed edges in the manufacturing network, representing the flow direction of work-in-process; i and j are nodes. +This represents the positive weight matrix of the network edges.
[0119]
[0120] in, The forward weights of the directed edges represent the probability that work-in-process flowing out of node i will flow into downstream node j. Additionally, the inverse weight matrix W of the directed edges... - It can be done through matrix W + The calculation yielded, where Let represent the probability that all work-in-process items flowing into node i originate from upstream node k. Let the set of immediately adjacent downstream nodes connected to node i be defined as . The set of immediately upstream neighbors (NUM) connected to node i is:
[0121] When there are no machine downtimes in the manufacturing network, the machine's production speed is defined as the maximum production speed, denoted as P. rm =[P rm1 ,P rm2 ,…,P rmn The actual production speed of machines in the manufacturing network is denoted as P. ra (t)=[P ra1 (t),P ra2 (t),…,P ran [(t)], and satisfy P ra ≤P rm Among them, P rmn P represents the maximum production speed of the nth machine. ran (t) represents the actual production speed of the nth machine; during the production process, the actual production speed of the machine is closely related to the immediate upstream and immediate downstream nodes to maintain balanced production. When the production speed of node i changes by ΔP rai At time (t), the change in production rate ΔP of the node immediately adjacent to it is... rak (t) and the change in production rate ΔP of the immediate downstream node raj (t) are respectively represented as:
[0122]
[0123]
[0124] After that, ΔP rai The effect of (t) will propagate along its process route to the source and end nodes, and cause corresponding changes in the corresponding process route on each machine.
[0125] Dynamic Machine Reliability and Maintenance Activities: In manufacturing networks, machines exhibit dynamic failure probabilities during operation. This dynamism can be attributed to three aspects: the uncertainty of machine failure modes, the dynamism of production speed, and the instability of feed quality. Fortunately, the Weibull distribution can fit different machine failure modes, such as decreasing, constant, and increasing failure rates. Therefore, failure rates following a Weibull distribution are suitable for modeling machine failure risk. The failure rate when only qualified feed is processed is defined as the base failure rate r. b When considering the dynamics of production speed, the basic failure rate (t) is expressed as:
[0126] r b (t)=(β / α)·(t r / α) β-1 (4)
[0127] Where α is the scale parameter and β is the shape parameter; t r This refers to the relative time the machine runs at maximum production speed; t is the actual time the machine runs; and the relative time t r It can also be used as an indicator to measure machine degradation.
[0128]
[0129] Considering the potential impact of substandard feed, the actual failure rate r(t) is defined as:
[0130]
[0131] Where, Δr i′ Let be the cumulative failure rate increment caused by defective feed, following a Beta distribution Beta(a,b). N(t) is the number of defective work-in-process items processed by the machine within the time interval [0,t). The probability distribution function F(t) for machine failure is derived as follows:
[0132] F(t) = 1 - exp(-(t) r / α) β -∫Δr(t)dt) (7)
[0133] in, Let Δr(t) be the cumulative failure rate increment caused by defective feed material during machine processing up to time t; the integral of Δr(t) is calculated using the following formula:
[0134]
[0135] Among them, t i′ This represents the actual time when the i′-th failure rate increment occurs.
[0136] Corrective and preventive maintenance are effective methods for handling both hard and soft machine failures. This invention assumes that both corrective and preventive maintenance will restore the machine's failure rate to the level at t=0. Furthermore, it assumes that the maintenance times for corrective and preventive maintenance follow the following normal distributions: and Corrective maintenance typically aims to restore machines that have experienced chaotic downtime due to faults, and the failures encountered are often more sudden than those faced by preventative maintenance. Therefore, it is assumed that corrective maintenance will take longer to complete, hence μ. cm ≥μ pm , Assume the unit time maintenance costs for corrective maintenance and preventive maintenance are c, respectively. cm and c pm Then, the total maintenance cost of the machine during the time interval [0, t) is c. m (t) can be expressed as:
[0137]
[0138] Where, N cm (t) represents the number of corrective maintenance operations performed on the machine within the time interval [0, t), N. pm (t) represents the number of preventative maintenance operations performed on the machine within the time interval [0, t), where t cm_i1 The time spent on the i1st corrective maintenance of the machine, t pm_j1 The time spent on the j1st preventive maintenance of the machine.
[0139] Constructing a dynamic model of processing quality and inspection activities: Processing quality is another important standard for machine reliability, which can be described by the number of nonconforming products M(t) generated within a specific time period [0,t). Work-in-process that does not meet quality specifications is called a nonconforming product. Due to the instability of machine reliability, processing quality also exhibits time-varying characteristics. Therefore, the random variable M(t) is modeled by a non-homogeneous Poisson process (NHPP) with an intensity function λ(t). The expression for λ(t) is:
[0140] λ(t)=ω-ε·e -δ·r(t) (10)
[0141] Where ω>0 represents the maximum nonconforming product intensity, ε>0 and δ>0 are the influence coefficients of the failure rate on the intensity function, and ω-ε<λ(t)<ω <P ra (t); when ω = P is defined ra When (t), it can be proven that ε = g × ω, where g ∈ [0,1] is the initial percentage of defect-free products produced by the machine. Define n dLet be the number of defective products generated by the machine within the time interval [t, t+Δt), and its probability be:
[0142]
[0143] in, Let n be the expected value of non-conforming products produced within the time interval [0, t), where Δt is the quality statistical period. q The total number of qualified products processed by the machine within the time interval [t, t+Δt) is expressed as:
[0144]
[0145] In a manufacturing network, inspection activities are performed after machining to ensure timely detection of defective work-in-process. Inspection activities typically involve Type I errors (rejection) and Type II errors (acceptance). Type I errors have a probability of p. I Type II error, with probability p II Correspondingly, the correct judgment of nonconforming and conforming products is called correct acceptance and correct rejection, respectively denoted by 1-p. I and 1-p II Let p represent the corresponding probability. If the processing machine has no detection activity, p can be used. I =0 and p II =1 indicates.
[0146] Assume the sampling ratio (random sampling) in the detection activity is s. a s a Given that ∈[0,1] and the sampling and inspection activities are independent of each other, the joint probability of any qualified work-in-process occurring a Type I error is s. a ·p I The joint probability of any nonconforming work-in-process causing a Type II error is s. a ·p II Define the number of Type I errors and Type II errors as n, respectively. fr and n fa And they respectively follow a binomial distribution B(n) q ,s a ·p I ) and B(n d ,s a ·p II Therefore, the number of non-conforming work-in-process items M'(t) leaving the machine is considered to be a non-homogeneous Poisson process M(t) and a binomial distribution B(n). d ,s a ·p II The composite of ); n faThe probability that a defective work-in-process item leaves the machine within the time interval [t, t+Δt) is:
[0147]
[0148] Where m′(t) is the average number of defective work-in-process items leaving the machine during the time interval [0, t]; furthermore, the probability that qualified work-in-process items are rejected during the time interval [t, t+Δt) is:
[0149]
[0150] Where D(t) represents the number of qualified work-in-process items that are rejected at time t. Indicates from n q Randomly select n from the defective products fr The number of possible combinations of a sample; the number of products processed by the machine within the time period [t, t+Δt) that are correctly identified as defective, n. cr The number n that was correctly judged as qualified ca They are represented as follows:
[0151]
[0152] Based on this, the proportion of non-conforming products identified by the machine through inspection activities within the time period [t, t+Δt) can be obtained as follows:
[0153]
[0154] Assume the inspection cost for a single product by machine is c. ins The total cost of the machine's inspection within the time interval [t, t+Δt) is:
[0155]
[0156] Assume the cumulative value increment of a single work-in-process item from the input node of the manufacturing network to the current processing node i is v. si The value loss caused by classifying a single work-in-process item as non-conforming at node i is also v. si Then, the average process value increment v brought about by the work-in-process inventory corresponding to node i is... ai It can be defined as:
[0157]
[0158] Among them, v skThis represents the cumulative value increment of upstream node k of node i. Furthermore, when a Type II error occurs during the inspection of non-conforming work-in-process, the non-conforming work-in-process will flow into downstream machines and consume more production resources, resulting in a greater value loss than if it were correctly identified as non-conforming in the current machine. The value loss caused by a Type II error in the inspection process is defined as... Therefore, the net value increment v of all work-in-process processed by machine i during the time interval [0,t) is... net_i The sum of the value increment of all correctly accepted work-in-process, the value loss of incorrectly accepted work-in-process, and the value loss of all rejected work-in-process:
[0159]
[0160] Where, n ca_i n represents the number of work-in-process items correctly accepted by the machine corresponding to node i. fa_i n represents the number of work-in-process items accepted by the machine corresponding to node i with errors. cr_i n represents the number of work-in-process items correctly rejected by the machine corresponding to node i. fr_i These represent the number of work-in-process items that were incorrectly rejected by the machine corresponding to node i.
[0161] Evaluating the manufacturing network state: Machine performance manifests in various ways, such as failures caused by hard or soft faults, idle time caused by starvation or congestion, variations in processing quality, and the degree of degradation; all of these affect the performance of the manufacturing network. For a manufacturing network with n machines, construct the state matrix S. K Represents the state at time t:
[0162] S K =[Q K ;D K H K ;O K (20)
[0163] Where t = KΔt, The defect ratio of each machine within the time interval [t-Δt,t) represents the quality status of the machine and is calculated using equation (16). K =[t r1 ,t r2 ,…,t rn ] represents the degradation state of the machine at time t, represented by the relative time from the last maintenance to the present for each machine, as shown in equation (5). K =[h1,h2,…,h n [] represents the machine's health status at time t, h iIn the range {0,1}, 1 represents a faulty state of the machine, and 0 represents a fault-free state; O K =[o1,o2,…,o n [] represents the idle state of the machine at time t, o i In {0,1}, 1 represents an idle state and 0 represents a working state.
[0164] Evaluating the cumulative performance of a manufacturing network: Cumulative performance evaluation is used to determine the cost-effectiveness of maintenance and quality inspection programs for all machines. From an economic perspective, the cumulative performance of a manufacturing network can be represented by net revenue. Define the reward r. K To create a network from state S in time interval [t, t+Δt). K Transition to S K+1 The net benefit generated during the process; where Δt is the period of quality statistics, which is also the step size for measuring state transition. Therefore, the reward per unit step is equal to the cumulative net value increment after deducting maintenance, inspection, and decision-making costs;
[0165]
[0166] Among them, c I_i c is the total detection cost of the machine corresponding to node i. m_i c is the total maintenance cost of the machine corresponding to node i. D This refers to the decision-making cost of maintenance and inspection actions. To evaluate the cost of changes from state S within the time interval [KΔt, K′Δt). K To S K′ The cumulative performance, where the reward GK is defined as the long-term return of the manufacturing network, is calculated through cumulative rewards:
[0167]
[0168] Where K′>K.
[0169] A joint optimization model based on Markov decision processes: The dynamic reliability and quality model constructed earlier provides a state transition model for the manufacturing network. When quality inspection and maintenance are considered as actions, a typical control model based on Markov decision processes can be constructed, where the state set, action set, reward function, and state transition model are known. Simultaneously, this model aims to find the optimal preventive maintenance and quality inspection policy functions to achieve the best long-term reward for the manufacturing network. In practice, corrective maintenance will automatically initiate when a machine fails, without policy support. Therefore, the policy only refers to the behavior of preventive maintenance and quality inspection at time t = KΔt, denoted as... in, For quality inspection actions, This is a preventative maintenance action. Furthermore, at t = KΔt, the actions of all machines in the manufacturing network depend on state S. K Therefore, it is denoted as A. K =π(S) K ), where π(·) represents the policy function:
[0170]
[0171] Where, π * (·) represents the optimal policy function, and Q(·) represents the optimal policy function in state S. K Take A K The long-term reward function during the action.
[0172] Under the optimal policy, the value function and the Q function satisfy:
[0173]
[0174] Where V(·) represents the state S K The maximum long-term return at that time.
[0175] Traditional dynamic programming or heuristic algorithms can complete optimization tasks in finite-domain Markov decision processes with enumerable state and action spaces. However, in this study, the defect ratio Q... K and degenerate state D K The state space is continuous, and the action space for quality inspection is continuous, with sampling ratios that can be any value within [0,1]. Therefore, the Markov decision process of the manufacturing network has uncountable state and action spaces. Furthermore, the state and action spaces of the constructed Markov decision process grow exponentially with the number of machines, leading to the "curse of dimensionality." Therefore, traditional dynamic programming methods cannot solve this infinite temporal sequence decision problem. Although heuristic algorithms have strong search capabilities and can provide optimal solutions, their weak transfer learning ability prevents them from consistently guaranteeing optimal manufacturing system performance as the manufacturing scenario changes.
[0176] In current algorithm research, the learning ability of the DRL algorithm has been proven effective in handling infinite time-domain Markov decision processes. Among them, the DQN, Actor-Critic, and DDPG algorithms have been shown to effectively solve different Markov decision processes in maintenance schemes. First, all three algorithms are applicable to Markov decision processes with continuous or discrete state spaces. However, the DQN algorithm is only applicable to discrete action spaces, the DDPG algorithm is only applicable to continuous action spaces, and the Actor-Critic algorithm is applicable to both continuous and discrete action spaces. Furthermore, DDPG leverages the advantages of neural network-based Q-functions and the Actor-Critic framework, exhibiting superior stability compared to the other two algorithms. Therefore, this invention adopts the DDPG algorithm.
[0177] The DDPG algorithm is built on the Actor-Critic framework, where two θ values are used to construct the algorithm. μ and θ Q A neural network with parameters is used to approximate the policy function and value function. The policy function and value function are constructed using the Actor network μ(S) and Critic network Q(S,A) from the Actor-Critic framework, as shown in Figure 4. Furthermore, the activation function is the ReLU function, and the hidden layers are fully connected layers. In the Actor network, the state matrix S... K It is the input, and the corresponding action A K The output is designed to maximize long-term returns. Since the final output layer has the same normalized neurons, the quality detection action... and preventive maintenance actions There are continuous output values in [0,1]. To address the inconsistency between continuous quality inspection actions and discrete preventive maintenance actions, this study uses a given preventive maintenance criterion c. d Discretize preventive maintenance actions using the ∈[0,1]. Specifically, preventive maintenance in It will not be executed at that time, but When executed. The state matrix S K and action vector A K As input to the Critic network, the corresponding long-term expected return Q-value Q(S) is used. K A K ) as output.
[0178] Based on the constructed neural network, the agent iteratively interacts with the manufacturing network environment. During the interaction, the state of the manufacturing network within the time interval [t, t+Δt) (where t = KΔt) changes from S... K Convert to S K+1The process is defined as a step in the DRL algorithm. An Epoch refers to the evaluation phase of long-term reward within the time interval [KΔt, K′Δt), consisting of multiple steps. An Episode refers to the process of the Agent performing the task, consisting of multiple Epochs. During training, the Agent will iteratively simulate the DDPG algorithm until the maximum number of steps is reached. At time t = 0 (K = 0), assume the initial action A... 0 No quality inspections or preventative maintenance were performed. Similarly, assume the initial state S... 0 The manufacturing network is free from defects, machine degradation, machine malfunctions, and machine idleness. The detailed training process is shown in Figure 5 below.
[0179] Step 1: Execute each action to simulate the operation of the manufacturing network in the time interval [KΔt, K′Δt), where t = KΔt;
[0180] Step 1.1: Evaluate the state of the manufacturing network at time t = KΔt: In the learning environment, based on the proposed dynamic reliability and quality model, evaluate the machine state at time t = KΔt, and then evaluate the state S of the manufacturing network at time t = KΔt. K , so that state S K This is provided to the Agent as an observation of the manufacturing network;
[0181] Step 1.2: Generate an action based on the current policy function π(S): an action It is possible to input state S into the Actor network μ(S) K The DDPG algorithm obtains this by adding a normally distributed random noise N. r Try permissible actions to explore better strategies, i.e., A K =π(S) K )+N r Then, according to the preventive maintenance guideline c d Preventive maintenance actions are converted into discrete executable actions {0,1}, where 0 indicates that preventive maintenance is not performed and 1 indicates that preventive maintenance is performed; in order to avoid criterion c d The resulting preference, hence the definition of c d =0.5, which represents the median of the Actor network's output range [0,1].
[0182] Step 1.3: Execute actions to simulate the operation of the manufacturing network within the time interval [KΔt, KΔt+Δt): Obtain action A at time t = KΔt. K Subsequently, the quality inspection of the corresponding machines will immediately adopt the new sampling ratio during subsequent operations. Simultaneously, preventive maintenance actions are performed on the corresponding machines, with the preventive maintenance time t. pmObtained from the normal distribution N(μ pm , σ 2 pm ); if the machine fails during the period [KΔt, KΔt + Δt), a corrective maintenance is immediately performed at the time of failure, and the corrective maintenance time t cm is obtained from the normal distribution N(μ cm , σ 2 cm );
[0183] Step1.4: Evaluate the reward r K : Calculate the reward r i2 , and update i2 = i2 + 1; if i2 < K′, return to Step1.1; otherwise, execute Step2;
[0184] Step2: Obtain the transition record: After running in the time period [KΔt, K′Δt), obtain the long - term return G K which can; According to the method of Step1.1, obtain the state S at t = K′Δt K′ ; Then, the state transition record {S K , A K , G K , S K′} is stored in the experience buffer;
[0185] Step3: Update the Actor network and the Critic network: Randomly sample a batch of M state transition records from the experience buffer to update the Actor network μ(S) and the Critic network Q(S, A); At this time, the experience buffer has a maximum storage capacity, that is, the buffer length L. When the number of transition records reaches L, in order to store new transition records, the earliest records will be discarded.
[0186] Step4: Judge the end condition; If the Episode reaches the预定的最大训练次数 (predetermined maximum number of training times) or obtains a stable long - term return G K , stop training; otherwise, update the simulation period: [KΔt, K′Δt) ← [K′Δt, 2K′Δt - KΔt), and return to Step1.
[0187] Before the training step, use the Actor network and the Critic neural network with the same structure to construct the target Actor network μ′(S) and the target Critic network Q′(S, A) respectively, and initialize the Actor network μ(S) and the Critic network Q(S, A) with random parameters θ μ and θ Q , and use θ μ' = θ μ and θQ' =θ Q Initialize the target actor network μ′(S) and the target critic network Q′(S,A). To improve the stability of the optimization, the target actor network and the target critic network are updated periodically based on the latest Actor and Critic parameters. The neural network is updated in each training epoch according to the state transition records in the experience buffer, as shown in Figure 6, and the update algorithm is as follows.
[0188] a) Randomly select M state transition records from the experience buffer:
[0189] b): For the transfer records i3 = 1, 2, ..., M, calculate the future target action. and target future long-term returns And set the value function objective.
[0190] c): Update the parameters θ of the Critic network Q(S,A) by minimizing the loss function. Q :
[0191]
[0192] d): Update the parameters θ of the Actor network μ(S) by maximizing the expected cumulative long-term return. μ :
[0193]
[0194] e): Update the parameters of the target actor network and the target critic network:
[0195] θ μ′ =τθ μ +(1-τ)θ μ′ (27)
[0196] θ Q′ =τθ Q +(1-τ)θ Q′ (28)
[0197] Where τ is the smoothing factor.
[0198] Case Study: Manufacturing Network and Agent Parameters
[0199] In this example, a directed acyclic network consisting of 30 nodes is used to model a manufacturing network with multiple process routes. As shown in Figure 7, machines are represented as nodes with different reliability parameters, while the flow of work-in-process between machines is represented by edges. Work-in-process flows randomly between machines, and its quantity is limited by machine capacity to ensure production balance. This manufacturing network has 4 source nodes, 4 terminal nodes, and 51 directed edges, resulting in a total of 724 different process routes.
[0200] This example uses MATLAB R2021a software for training. Based on the learning rates used in relevant research, the learning rates for the Actor network and Critic network in this training are set to 2×10⁻⁶. -3 and 1×10 -3 Since the number of hidden layer neurons is highly dependent on the dimensionality of the problem, and considering similar studies of DRL, the number of hidden layer neurons is set to L. s1 =128, L s2 =256, L s3 =128, L a1 =64, L a2 =128, L c1 =256, L c2 =128, L c3 =64, L1=128, L2=256, L3=128. In addition, based on existing research, other parameters are: empirical buffer length L=1×10 6 The sampling batch size M = 1280, and the smoothing factor τ = 1 × 10⁻⁶. -3 Decision cost c D =100.
[0201] Training the DRL Agent: The DRL Agent will be trained to obtain the maximum gain over 5000 time units, i.e., (K′-K)×Δt=5000. Considering the sensitivity of Agent performance to step size, three different manufacturing scenarios are constructed using different step sizes Δt and decision frequencies K′-K: {Δt=50,K′-K=100}, {Δt=100,K′-K=50}, and {Δt=500,K′-K=10}. This means that the Agent will generate action A in each Epoch. K The system interacted with the manufacturing network 100, 50, and 10 times, respectively. Based on the constructed manufacturing network environment and DRL Agent, three training episodes were conducted for the three manufacturing scenarios, and the results are shown in Table 1.
[0202] Furthermore, an optimization model based on a genetic algorithm was used as a benchmark for this method. This algorithm constructed a population of 70 individuals, each with a 1×60 matrix as its chromosome (solution), representing the preventative maintenance and quality inspection actions of 30 nodes in the constructed manufacturing network. The reward G is calculated over 5000 time units. K The fitness function is used to evaluate individuals. For the three manufacturing scenarios, the maximum number of generations (MEG) is set to 7 × 10 for the manufacturing network to run. 5 The time taken is equal to the number of training steps under the DRL algorithm, i.e., {Δt=50,MEG=100}, {Δt=100,MEG=200} and {Δt=500,MEG=1000}. The resulting benefits are shown in Table 1.
[0203] Table 1. Profits from training with DRL and genetic algorithms
[0204]
[0205] When the returns tend to stabilize, their mean and standard deviation (SD) are calculated in Table 1. Taking into account the differences in scale, the coefficient of variation (CV) in Equation (29) is used to analyze the relative dispersion of the returns.
[0206] CV = SD / mean (29)
[0207] The mean, standard deviation, and coefficient of variation for the DRL algorithm are calculated based on the returns of the most recent 100 epochs when the returns tend to stabilize. The mean, standard deviation, and coefficient of variation for the genetic algorithm are calculated based on the returns of the most recent 50 epochs when the returns tend to stabilize. For the DRL algorithm, the highest returns are 4.83 × 10⁻⁶ when Δt = 50, 100, and 500. 4 3.63×10 4 6.47×10 4 For the genetic algorithm, the highest returns are 2.63 × 10⁻⁶ when Δt = 50, 100, and 500. 4 4.25×10 4 6.39×10 4 Only when Δt = 100 can the genetic algorithm help the manufacturing network achieve a better average return than the DRL algorithm. Therefore, compared with the genetic algorithm, the DRL algorithm proposed in this invention has better adaptability to various manufacturing scenarios in complex manufacturing networks. Moreover, in most training sessions, the coefficient of variation under the genetic algorithm is greater than that under the DRL algorithm, meaning that the stability of the returns obtained by the genetic algorithm is lower than that of the DRL algorithm. The training trajectories of the highest returns of the DRL algorithm and the genetic algorithm under different manufacturing scenarios are shown in Figure 8.
[0208] The payoff trajectories under the DRL algorithm demonstrate that the constructed DRL Agent can improve the long-term returns of the manufacturing network through interaction, proving the model's effectiveness. Furthermore, the different patterns of the payoff trajectories illustrate the differences in training results under different manufacturing scenarios. First, a higher interaction frequency between the manufacturing environment and the DRL agent (when step size Δt = 50) not only leads to slower convergence in the learning process but also results in significant losses during training, i.e., negative returns as shown in Figure 8. Second, the trajectory with Δt = 100 indicates that faster convergence can be achieved at lower interaction frequencies, but the obtained returns may not be optimal. In summary, due to the nonlinearity, high dimensionality, and dynamics of the manufacturing network, the learning performance of the DRL algorithm is highly sensitive to manufacturing scenarios of varying step sizes. Moreover, the payoff trajectories suggest that DRL training may incur potential losses. Therefore, the digital manufacturing network environment must be optimized based on machine learning to avoid potential losses when the Agent interacts with the real manufacturing system.
[0209] Experiments based on trained agents: Using the optimal DRL Agent trained in the three manufacturing scenarios described above, this invention implemented an experiment to control the manufacturing network's repair and quality inspection within 50,000 time units. The coefficient of variation and cumulative reward under different manufacturing scenarios are shown in Figure 9. With the help of the DRL Agent, the manufacturing network can obtain continuously increasing cumulative rewards in all manufacturing scenarios. Similarly, when the interaction step size Δt = 500, the manufacturing network can obtain the highest cumulative reward under the control of the proposed DRL Agent. Meanwhile, the coefficient of variation represents the relationship of the dispersion of rewards under different step sizes: CV 500 <CV 100 <CV 50 Therefore, the agent's performance is consistent with the training results, indicating that the experimental process is stable and effective.
[0210] The reward per unit time for each step is calculated based on the experimental results. For the Kth step, the reward per unit time is r. u(K) It can be obtained from equation (30):
[0211] r u(K) =r K / Δt (30)
[0212] Figure 9 is a scatter plot of the unit-time rewards obtained by three trained agents. In the scatter plot, when the step size Δt = 100 or 500, the unit-time rewards are stable and concentrated, with the unit-time reward at Δt = 500 being greater than that at Δt = 100. However, when Δt = 50, the unit-time rewards become dispersed and unstable. This phenomenon indicates that due to the high nonlinearity, high dimensionality, and dynamics of the fabrication network, smaller step sizes make it difficult for the agent to consistently make optimal decisions. On the other hand, the unit-time rewards also explain why the agents trained with step sizes of 50 and 100 have lower cumulative rewards.
[0213] Finally, this example analyzes the connectivity of a manufacturing network over 50,000 time units under the intervention of three trained DRL agents. Connectivity refers to the probability that at least one process path between the source and destination nodes of the manufacturing network remains connected. The connectivity curve is shown in Figure 10. When the step size Δt = 500, the connectivity of the manufacturing network is the worst, and the fluctuation range is the largest, indicating the worst and most unstable connectivity. Conversely, when the step size Δt = 100, the connectivity is the best, and the fluctuation range is the smallest, verifying the concept that "high returns come with high risks." Specifically, when the manufacturing network maintains a high long-term return (Δt = 500), there will be a high risk of operational interruption (the worst connectivity).
[0214] This invention investigates the joint optimization problem of maintenance and quality inspection in manufacturing networks based on the DRL algorithm, under the condition of interaction between machine reliability and work-in-process quality. First, a mathematical model for constructing a nonlinear, high-dimensional, and dynamic environment of the manufacturing network is proposed, providing an effective state transition model for network control. Second, an effective DRL model suitable for joint reliability-quality control of manufacturing networks is constructed, capable of simultaneously modeling discrete-continuous mixed states and mixed actions. Furthermore, training and experimental results verify the effectiveness of the proposed DRL model. Compared with genetic algorithms, the proposed DRL algorithm exhibits better adaptability to dynamic and diverse manufacturing scenarios. Simultaneously, the model proposed in this invention can effectively balance the contradiction between the economic benefits and operational risks of the manufacturing network.
[0215] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A joint optimization method for maintenance and inspection in manufacturing networks based on deep reinforcement learning, characterized in that, The steps are as follows: Step 1: At the machine level, considering the dynamic production speed caused by machine downtime due to malfunctions, a machine reliability model considering the impact of feed quality and a processing quality model considering the impact of machine reliability are constructed; Step 2: Based on the reliability model and the quality model, a systematic evaluation of the manufacturing network state and performance is conducted; the evaluation method is as follows: for a manufacturing network with n machines, a state matrix is constructed. Represents the state at time t: ;in, , For each machine Quality status over a period of time; This represents the degenerate state of the machine at time t. It represents the machine's health status at time t. In this context, 1 indicates a faulty state of the machine, and 0 indicates a fault-free state of the machine. This represents the idle state of the machine at time t. In this context, 1 represents an idle state, and 0 represents a working state; a reward is defined. To create a network in From state in time Transition to Net gains generated during the process: ;in, It is the total detection cost of the machine corresponding to node i. It is the total maintenance cost of the machine corresponding to node i. It is the decision-making cost of maintenance and inspection actions; in order to assess the cost of... Within a time period, from state arrive The cumulative performance will benefit Defined as the long-term reward of the manufacturing network, it is calculated through cumulative rewards: ;in, Furthermore, a joint optimization model for manufacturing network maintenance and quality inspection was established; specifically, quality inspection and preventive maintenance were treated as actions, denoted as... ,in, For quality inspection actions, For preventative maintenance actions; in The actions of all machines in the time-manufacturing network depend on the state. Therefore, it is recorded as , where π(·) represents the policy function: in, Describes the optimal policy function. Indicates that in state S K Take A K The long-term reward function during an action; under the optimal policy, the value function and Q function satisfy: ;in, Indicates that in state S K The maximum long-term return; Step 3: At the system level, using the economic operation of the manufacturing network as the standard for strategy evaluation, the optimal strategy for quality inspection and maintenance under a given manufacturing network state is learned through a designed deep deterministic policy gradient algorithm; The specific method is as follows: Step 1: Execute the current actions to simulate the manufacturing network under the current state. The operation within time, of which Step 1.1: Assessment The state of the network is constantly being generated: in a learning environment, based on the proposed dynamic reliability and quality model, it is evaluated. The machine state at any given moment, then evaluated. Constantly creating the network state S K , so that state S K This is provided to the Agent as an observation of the manufacturing network; Step 1.2: Generate an action based on the current policy function π(S): an action Able to access the Actor network Input state S K The DDPG algorithm obtains this by adding a normally distributed random noise N. r Try permissible actions to explore better strategies, that is Then, according to the preventive maintenance guideline c d The preventive maintenance actions are converted into discrete executable actions {0,1}, where 0 represents not performing preventive maintenance and 1 represents performing preventive maintenance; Step 1.3: Execute the actions, simulating the manufacturing network in Operation within time: In obtaining Action A at a given moment K Subsequently, the quality inspection of the corresponding machines will immediately adopt the new sampling ratio during subsequent operations. Simultaneously, preventive maintenance actions are performed on the corresponding machines, with the preventive maintenance time being t. pm From normal distribution Get; if the machine is If a fault occurs during the specified period, a corrective maintenance procedure shall be performed immediately upon the occurrence of the fault, with a corrective maintenance time of t. cm From normal distribution Obtain; Step 1.4: Evaluate the reward r K Calculate rewards And update ;if Otherwise, return to step 1.1; otherwise, proceed to step 2; Step 2: Obtain the state transition record: in Over time, the long-term return G is obtained. K According to the method in Step 1.1, we obtain state of time Then, obtain the state transition record. Storing data in the experience buffer; Step 3: Updating the Actor and Critic networks: Randomly select a small batch of M transition records from the experience buffer to update the Actor network μ(S) and the Critic network Q(S, A); At this time, the maximum storage capacity of the experience buffer is L. When the number of transition records reaches L, the oldest record is discarded; Step 4: Determining the termination condition: If the training episode reaches the predetermined maximum number of training sessions or a stable long-term reward G is obtained. K If the simulation fails, training will stop; otherwise, the simulation period will be updated. Then return to step 1.
2. The joint optimization method for maintenance and inspection of manufacturing networks based on deep reinforcement learning according to claim 1, characterized in that, At the machine level, the reliability and quality models are constructed as follows: Calculate the dynamic production rate: treat each machine as a node and use a directed acyclic graph. right Modeling an acyclic manufacturing network with nodes, where To create a set of nodes for a network, To create a set of directed edges in the network; All are nodes; when there are no machine downtimes in the manufacturing network, the machine's production speed is defined as the maximum production speed, denoted as... The actual production speed of machines in the manufacturing network is denoted as . And satisfy ;in, This represents the maximum production speed of the nth machine. This represents the actual production speed of the nth machine; when the production speed of node i changes... At that time, the change in the production rate of its immediate upstream node Changes in production speed of adjacent downstream nodes They are represented as follows: ; ;in, Let i represent the set of immediately upstream neighbors connected to node i. This represents the set of immediately downstream nodes connected to node i; Let represent the probability that all work-in-process items flowing into node i originate from upstream node k. This represents the probability that the work-in-process flowing out of node i will flow into downstream node j; Calculate dynamic maintenance costs: when considering the dynamic nature of production speed, the failure rate when processing qualified feed materials is defined as the base failure rate. : Where α is the scale parameter and β is the shape parameter; The relative running time of the machine is calculated based on the maximum production speed; t is the actual running time of the machine; the actual failure rate is taken into account for the impact caused by defective feed. Defined as: ;in, Let N(t) be the cumulative failure rate increment caused by defective feed into the machine, and let N(t) be the number of defective work-in-process products processed by the machine within the time interval [0, t). The probability distribution function of machine failure is derived. for: ;in, This represents the cumulative failure rate increment caused by defective feed material during machine processing up to time t. The formula for calculating the integral is: ;in, For the first The actual time when the failure rate increment occurs; assume that the maintenance time for corrective and preventive maintenance of the machine follows the following normal distribution: and ;and , The unit time maintenance costs for corrective maintenance and preventive maintenance are respectively... and Then, the total maintenance cost of the machine within the time interval [0, t) is... Represented as: ;in, Let be the number of corrective maintenance operations performed on the machine within the time interval [0, t). This represents the number of preventative maintenance procedures performed on the machine within the time interval [0, t). For the machine The time spent on corrective repairs For the machine The time spent on preventative maintenance; constructing a dynamic processing quality and inspection activity model: Defined as the number of defective products generated within the time interval [0, t). Satisfy intensity function Non-homogeneous Poisson process: Where ω>0 represents the maximum nonconforming product intensity, ε>0 and δ>0 are both influence coefficients of the failure rate on the intensity function, and ;definition Let be the number of defective products generated by the machine within the time interval [t, t+∆t), and its probability be: ;in, Let be the expected value of non-conforming product output within the time period [0, t). For quality statistics period; defined The total number of qualified products processed by the machine within the time interval [t, t+∆t) is expressed as: In the testing process, the incorrect judgment of non-conforming and conforming products is classified as: Type I error, which indicates incorrect rejection, with a probability of 1 / 2. Type II error, indicating incorrect acceptance, has a probability of . Assume the sampling ratio in the testing activity is ; , The joint probability of any qualified work-in-process occurring a Type I error is: The joint probability of any nonconforming work-in-process causing a Type II error is: The number of Type I errors and Type II errors are defined as follows: and And they respectively follow a binomial distribution. and Therefore, the number of defective work-in-process items leaving the machine It is a non-homogeneous Poisson process M(t) and a binomial distribution The combination of, then The probability that a defective work-in-process item leaves the machine within the time period [t, t+∆t) is: ;in, This represents the average number of defective work-in-process items leaving the machine during the time period [0, t]. Furthermore, the probability that acceptable work-in-process items are rejected during the time period [t, t+∆t) is: Where D(t) represents the number of qualified work-in-process items rejected at time t. Indicates from Randomly selected from the defective products The number of combinations of samples; the number of products processed by the machine within the time period [t, t+∆t) that are correctly identified as defective. cr The number n that was correctly judged as qualified ca They are represented as follows: Based on this, the proportion of non-conforming products identified by the machine during the inspection activity within the time period [t, t+∆t) is obtained as follows: Assume the machine's inspection cost for a single product is... The total cost of the machine's inspection within the time interval [t, t+∆t) is: Assume that the cumulative value increment of a single work-in-process item from the input node of the manufacturing network to the current processing node i is... Then the average process value increment brought about by the work-in-process inventory of the machined unit corresponding to node i. Defined as: ;in, This represents the cumulative value increment of upstream node k of node i; the value loss caused by non-conforming work-in-process due to Type II errors during the inspection process is defined as... φ>1; therefore, the net value increment of all work-in-process processed by machine i in time [0, t) is... The sum of the value increment of all correctly accepted work-in-process, the value loss of incorrectly accepted work-in-process, and the value loss of all rejected work-in-process: ;in, This represents the number of work-in-process items correctly accepted by the machine corresponding to node i. This represents the number of work-in-process items that were incorrectly accepted by the machine corresponding to node i. This represents the number of work-in-process items correctly rejected by the machine corresponding to node i. These represent the number of work-in-process items that were incorrectly rejected by the machine corresponding to node i.
3. The joint optimization method for maintenance and inspection of manufacturing networks based on deep reinforcement learning according to claim 1, characterized in that, The update method for the Actor network and Critic network is as follows: a) Randomly select M transfer records in batches from the experience buffer: ; S3.3.2: For the transition records i3=1,2,…,M, calculate the future target action. and target future long-term returns And set the value function objective. b) Update the parameters θ of the Critic network Q(S, A) by minimizing the loss function. Q : c): Update the parameters θ of the Actor network μ(S) by maximizing the expected cumulative long-term return. μ : ; d): Update the parameters of the target actor network and the target critic network: ; ; where τ is the smoothing factor.