Multi-stage manufacturing system joint optimization method

Through deep reinforcement learning DDPG algorithm, the multi-stage manufacturing system model is constructed, and the flow direction and quality inspection ratio in the work-in-process products are optimized, which solves the multi-stage manufacturing system optimization problem that traditional methods are difficult to adapt to dynamic changes, and realizes the system's real-time response and strategy optimization.

CN120297110APending Publication Date: 2025-07-11NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510293915.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional manufacturing system optimization methods are limited to single-stage or static environments, and are difficult to adapt to dynamically changing production needs and complex system interactions, and cannot effectively optimize production planning, material flow and equipment scheduling in multi-stage manufacturing systems.

Method used

The deep reinforcement learning DDPG algorithm is used to build a machine degradation model, a buffer area inventory model and a product quality status inspection model. Through interactive training between the agent and the manufacturing system, the product flow behavior and quality inspection ratio in the manufacturing system are optimized.

Benefits of technology

Real-time response and strategy optimization of multi-stage manufacturing systems are realized, which reduces the complexity of system evaluation and increases the reward value in the production process, and is suitable for quality inspection and production scheduling in different manufacturing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297110A_ABST
    Figure CN120297110A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of manufacturing system optimization, in particular to a multi-stage manufacturing system joint optimization method, which comprises the following steps of: establishing a machine degradation model and a processing quality model; establishing a stock model of the cache region; establishing a quality state inspection model of the in-process product; according to the degradation model, the processing quality model, the stock model and the quality state inspection model of the work-in-process, constructing a state space and a reward of a manufacturing system; constructing an action space according to the optimization target; and according to the state space, the action space and the reward, based on a DDPG algorithm, interactive training is carried out through the intelligent agent and the manufacturing system until the training is stable, and the optimal reward and action are obtained. According to the method, the deep reinforcement learning DDPG algorithm is used for performing joint optimization on the manufacturing system, the corresponding strategy with the highest reward value in the production process is obtained and compared with other common strategies for analysis, and the superiority of the method is verified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of manufacturing system optimization, and particularly relates to a multi-stage manufacturing system joint optimization method. Background Art

[0002] In modern manufacturing, production efficiency and resource optimization are the keys to an enterprise's competitiveness. With the advancement of Industry 4.0, manufacturing systems are gradually transforming towards intelligence and automation. The "Implementation Opinions of the Ministry of Industry and Information Technology and Other Six Departments on Promoting the Innovative Development of Future Industries" released in January 2024 pointed out that it is necessary to grasp the global scientific and technological innovation and industrial development trends, focus on promoting the future manufacturing development direction, build a future industry observation station, and use technologies such as artificial intelligence and advanced computing to accurately identify and cultivate high-potential future industries. Play the role of an incrementalizer for frontier technologies, aim at high-end, intelligent, and green directions, accelerate the transformation and upgrading of traditional industries, and provide new impetus for building a modern industrial system. It can be seen that the development level of intelligent manufacturing directly concerns the quality level of China's manufacturing industry and plays an important role in consolidating the foundation of the real economy, building a modern industrial system, and realizing new industrialization. As the core of the production process, the optimization problem of multi-stage manufacturing systems (MMSs) has always been the focus of research in academia and industry. Traditional optimization methods are often limited to single-stage or static environments and are difficult to adapt to dynamically changing production requirements and complex system interactions.

[0003] In recent years, Deep Reinforcement Learning (DRL), a technology that combines deep learning and reinforcement learning, has received extensive attention due to its advantages in dealing with high-dimensional state spaces and complex decision-making problems. DRL learns the optimal policy to maximize the cumulative reward through the interaction between the agent and the environment, which coincides with the decision-making process in multi-stage manufacturing systems. In multi-stage manufacturing systems, joint optimization means that multiple interdependent decision-making stages such as production planning, material flow, and equipment scheduling need to be considered simultaneously. The decisions in these stages interact with each other, forming a complex dynamic system. The joint optimization method based on DRL can dynamically adjust the policy by learning historical data and real-time feedback to adapt to the uncertainties and changes in the production process. The application of DRL in the field of intelligent manufacturing is growing exponentially year by year. By combining the advantages of Deep Neural Networks (DNN) and Reinforcement Learning (RL), it can make accurate and rapid decisions in the face of dynamic and complex situations with its powerful learning ability. This ability of DRL makes it an ideal choice for solving the optimization problems of multi-stage manufacturing systems, especially in the personalized intelligent manufacturing paradigm that requires adaptive and flexible solutions. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a joint optimization method for multi-stage manufacturing systems. The method first analyzes the series-parallel manufacturing system, establishes a system machine degradation model and a buffer stock model; then depicts the production process of the manufacturing system, determines the activities of the manufacturing system that need to be optimized, and constructs an optimization objective; finally, uses the deep reinforcement learning DDPG algorithm to jointly optimize the manufacturing system, obtains the corresponding policy with the highest reward value during the production process, and conducts a comparative analysis with other common policies to verify the superiority of the method.

[0005] The first object of the present invention is to provide a joint optimization method for multi-stage manufacturing systems, which is used for optimizing the in-process flow behavior and quality inspection ratio in multi-stage manufacturing systems. The manufacturing system includes multiple series-connected machines and a buffer for storing in-process products corresponding to each machine. The optimization method includes:

[0006] Establish a machine degradation model and a processing quality model according to the mutual influence between product quality and machine reliability in the manufacturing system;

[0007] Establish a buffer stock model; and establish an in-process quality state inspection model;

[0008] Construct the state space and the reward of the manufacturing system according to the degradation model, the processing quality model, the inventory model, and the in-process quality status inspection model.

[0009] Construct the action space according to the optimization objectives, where the optimization objectives include the proportion of in-process products flowing downstream from each machine and the quality inspection ratio.

[0010] Based on the state space, the action space, and the reward, and using the DDPG algorithm, through the interaction and training between the intelligent agent and the manufacturing system, until the training is stable, obtain the optimal reward and actions.

[0011] Preferably, when establishing the machine degradation model, the machine degradation modes include basic degradation and strong degradation; the machine degradation model includes: obtaining the failure probability distribution function and the probability density function of the machine under the mixed mode of basic degradation and strong degradation; obtaining the mixed failure rate of the machine according to the failure probability distribution function and the probability density function of the machine; and obtaining the reliability of the machine according to the failure probability distribution function of the machine.

[0012] Preferably, the failure probability distribution function of the machine is obtained according to the cumulative failure probability distribution function of the machine during the basic degradation failure process and the cumulative failure probability distribution function during the strong degradation failure process.

[0013] The probability density function of the machine is obtained according to the cumulative probability density function of the machine during the basic degradation failure process and the cumulative probability density function during the strong degradation failure process.

[0014] Among them, basic degradation means that under the condition of qualified feeding quality, the working load and working environment of the machine will be in a stable state, and the machine degradation process under this condition is called the basic degradation process.

[0015] Strong degradation means that under the condition of unqualified feeding quality, the working load or working environment of the machine will have different degrees of disturbance, and the machine degradation process under this condition is called the strong degradation process.

[0016] Preferably, the processing quality model is obtained according to the mixed failure rate of the machine, and the calculation formula is as follows:

[0017]

[0018] In the formula, ρ ∈ (0, 1] is the initial quality level of the machine; θ is a constant to make U(t) ∈ [0, 1]; r m (t) is the mixed failure rate of the machine; e is the natural constant.

[0019] Preferably, establish the inventory model of the buffer area, including:

[0020]

[0021] In the formula, n i upstream is the number of upstream machines of machine M i ;

[0022] dt i is the relative running time of machine M i ;

[0023] t delta is the relative time occupied by the machine due to maintenance and starvation reasons;

[0024] is the inventory level of buffer B i at time t;

[0025] is the inventory level of buffer B i at the next inspection time point t';

[0026] is the production rate of machine M i ;

[0027] is the production rate of machine M j ;

[0028] i and j respectively represent the corresponding numbers of the machine or buffer. M i and M j represent the machines with numbers i and j, and B i and B j represent the buffers with numbers i and j.

[0029] Preferably, an in-process quality status inspection model is established, including:

[0030] p I∩II (t) = U(t)·p II +(1 - U(t))·p I

[0031] In the formula, p I∩II (t) is the probability of making a wrong judgment on the in-process products with unknown quality status during the inspection process;

[0032] U(t) is the probability that the in-process products with unknown quality status are non-conforming products;

[0033] p I is the probability of rejecting qualified products during the quality inspection process;

[0034] p II is the probability of accepting non-conforming products during the quality inspection process.

[0035] Preferably, the state space includes:

[0036] stage k =[MS k ,BS k ,Pcum k ,DT k ,Stage k

[0037] wherein, is the state of each machine in the manufacturing system at the decision-making moment, which is a binary state, i.e., ms = 1 indicates that the machine is running, and ms = 0 indicates that the machine is in a maintenance or idle state;

[0038] is the state of each buffer area in the manufacturing system at the decision-making moment. This state has two representation forms: First, bs i =0, 1, 2…, sk Bi , that is, the inventory of the buffer area is used as the state output of the buffer area; Second, bs has three possible situations, bs = 0 represents that the inventory is empty, bs = 1 represents that the inventory is normal, and bs = 2 represents that the inventory exceeds the storage upper limit;

[0039] is the cumulative nonconforming product rate of each machine in the manufacturing system from the previous decision-making moment to the current decision-making moment;

[0040] is the relative running time of each machine in the manufacturing system at the current decision-making moment from the moment when the last maintenance is completed;

[0041] Stage k =[k, k,…, k] indicates that the state obtained by the machine belongs to stage k, which is used to make corresponding decisions better in different stages.

[0042] Preferably, the action space includes:

[0043]

[0044] wherein, is the proportion of the in-process products flowing from machine 1 to machine 2 in the k-th stage;

[0045] is the proportion of the in-process products flowing from machine 4 and machine 5 to machine 6 in the k-th stage;

[0046] is the quality inspection proportion of machine M1 in the k-th stage;

[0047] is machine Mn in the k-th stage m ​The quality inspection ratio.

[0048] Preferably, the reward includes:

[0049]

[0050] In the formula, is the quality inspection cost of machine M i within the time period [0, t); c i buffer (t) is the penalty cost of buffer B i within the time period [0, t); r i (t) is the net product value increase of the products processed by machine M i within the time period [0, t), r i (t) = N WIP (t) × r single, N WIP (t) is the number of qualified products processed by machine M i within the time period [0, t), r single is the net increase value of processing a single product; n m is the total number of machines.

[0051] The second object of the present invention is to provide a multi-stage manufacturing system joint optimization system, including:

[0052] A model establishment module, configured to establish a machine degradation model and a processing quality model according to the mutual influence between product quality and machine reliability in the manufacturing system; establish an inventory model of the buffer; and establish an in-process product quality status inspection model;

[0053] A state space construction module, configured to construct a state space and the reward of the manufacturing system according to the degradation model, the processing quality model, the inventory model, and the in-process product quality status inspection model; construct an action space according to the optimization objectives, where the optimization objectives include the proportion of in-process products flowing downstream from each machine and the quality inspection ratio;

[0054] An interaction training module, configured to, based on the DDPG algorithm, through interaction training between the agent and the manufacturing system according to the state space, the action space, and the reward, obtain the optimal reward and action after the training is stable.

[0055] The present invention has at least the following beneficial effects:

[0056] The present invention provides a method for jointly optimizing a multi-stage manufacturing system. This method first analyzes the series-parallel manufacturing system, establishes a system machine degradation model and a buffer stock model; then depicts the production process of the manufacturing system, determines the activities of the manufacturing system that need to be optimized, and constructs an optimization objective; finally, uses the deep reinforcement learning DDPG algorithm to jointly optimize the manufacturing system, obtains the corresponding strategy with the highest reward value during the production process, and conducts a comparative analysis with other common strategies to verify the superiority of this method.

[0057] The method provided by the present invention realizes the determination of the quality inspection ratio of each machine in the manufacturing system and the ratio of transporting work-in-progress to downstream machines from the perspective of actual production, and infers the real-time strategy of each stage of the manufacturing system under the maximum reward.

[0058] Based on the DDPG algorithm, the present invention combines with the production process of the multi-stage manufacturing system, effectively reduces the complexity of evaluating the multi-stage manufacturing system, and the proposed method for jointly optimizing the multi-stage manufacturing system can respond to different states of the manufacturing system in real time and provide strategies in a timely manner.

[0059] The modeling process of the present invention is simple and efficient, and has good applicability to the formulation of quality inspection and production scheduling strategies in other manufacturing scenarios. Description of the Drawings

[0060] Figure 1 It is a schematic structural diagram of a multi-stage series-parallel manufacturing system with a buffer.

[0061] Figure 2 It is a flowchart of the production process of the series-parallel manufacturing system.

[0062] Figure 3 It is the basic process of the DDPG algorithm.

[0063] Figure 4 It is the average reward of the DDPG algorithm training in this case.

[0064] Figure 5 It is the quality inspection ratio and cumulative nonconforming product rate of each machine at each stage in this case.

[0065] Figure 6 It is the change of the weight of work-in-progress flowing to downstream at different stages in this example.

[0066] Figure 7 It is the change of the cumulative reward and stage reward with the stage under different methods in this example. Detailed Implementation Manner

[0067] In order to elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will be described in detail in conjunction with embodiments.

[0068] The objective of the present invention is to provide a joint optimization method for multi-stage manufacturing systems. This method first analyzes series-parallel manufacturing systems, establishes system machine degradation models and buffer stock models; then depicts the manufacturing system production process, determines the manufacturing system activity content that needs to be optimized, and constructs optimization objectives; finally, uses the deep reinforcement learning DDPG algorithm to jointly optimize the manufacturing system to obtain the corresponding strategy with the highest reward value during the production process.

[0069] To achieve the above objective, the present invention provides a joint optimization method for multi-stage manufacturing systems, which is used for optimizing the in-process flow behavior and quality inspection ratio in multi-stage manufacturing systems. Among them, the manufacturing system includes multiple machines and a buffer area corresponding to each machine for storing in-process products. The optimization method includes:

[0070] S1. Establish a machine degradation model and a processing quality model based on the mutual influence between product quality and machine reliability in the manufacturing system;

[0071] When establishing the machine degradation model, the machine degradation modes include basic degradation and strong degradation; the machine degradation model includes: obtaining the failure probability distribution function and probability density function of the machine under the mixed mode of basic degradation and strong degradation; obtaining the mixed failure rate of the machine according to the failure probability distribution function and probability density function of the machine; and obtaining the reliability of the machine according to the failure probability distribution function of the machine.

[0072] The failure probability distribution function of the machine is obtained based on the cumulative failure probability distribution function of the machine during the basic degradation failure process and the cumulative failure probability distribution function of the machine during the strong degradation failure process;

[0073] The probability density function of the machine is obtained based on the cumulative probability density function of the machine during the basic degradation failure process and the cumulative probability density function of the machine during the strong degradation failure process;

[0074] Among them, basic degradation means that under the condition of qualified incoming material quality, the working load and working environment of the machine will be in a stable state, and the machine degradation process under this condition is called the basic degradation process;

[0075] Strong degradation means that under the condition of unqualified incoming material quality, the working load or working environment of the machine will have different degrees of disturbance, and the machine degradation process under this condition is called the strong degradation process.

[0076] The processing quality model is obtained based on the mixed failure rate of the machine, and the calculation formula is as follows:

[0077]

[0078] where ρ ∈ (0, 1] is the initial quality level of the machine; θ is a constant such that U(t) ∈ [0, 1]; r m (t) is the mixed failure rate of the machine; e is the natural constant.

[0079] It should be noted that according to the mutual influence between product quality and machine reliability in the manufacturing system, it means that when the failure rate of the machine increases, the processing quality of the machine decreases, and its U(t) is related to r m (t). When r m (t) increases, U(t) increases, and the number of defective products produced increases. This is the influence of machine reliability on product quality. When the number of defective products processed by the machine increases, according to the formula of F m (t), the cumulative failure probability of the machine will increase, that is, the reliability of the machine will decrease and the failure rate will increase. This is the influence of product quality on machine reliability.

[0080] U(t) is the work-in-progress with unknown quality status, and its probability of being a defective product; r m (t) is the failure rate of the machine under mixed quality feeding conditions; F m (t) is the failure probability distribution function of the machine under mixed conditions. The specific calculation is as follows.

[0081] S2. Establish the inventory model of the buffer area; and establish the inspection model of the quality status of the work-in-progress;

[0082] Establish the inventory model of the buffer area, including:

[0083]

[0084] where n i upstream is the number of upstream machines of machine M i ; dt i is the relative running time of machine M i ; t delta is the relative time occupied by the machine due to maintenance and starvation reasons; is the inventory level of buffer B i at time t; is the inventory level of buffer B i at the next inspection time point t'; is the productivity of machine M i ; is the productivity of machine M j ; i and j respectively represent the numbers corresponding to the machine or buffer area. M i and M j represent the machines numbered i and j, and B i and B j represent the buffers numbered i and j.

[0085] Establish a work-in-progress quality status inspection model, including:

[0086] p I∩II (t) = U(t)·p II +(1 - U(t))·p I

[0087] In the formula, p I∩II (t) is the probability of making a wrong judgment on the work-in-progress with unknown quality status during the detection process; U(t) is the work-in-progress with unknown quality status, and it is the probability of being a non-conforming product; p I is the probability of rejecting a conforming product during the quality inspection process; p II is the probability of accepting a non-conforming product during the quality inspection process.

[0088] S3. According to the degradation model, the processing quality model, the inventory model, and the work-in-progress quality status inspection model, construct the state space and the reward of the manufacturing system;

[0089] The state space includes:

[0090] state k = [MS k , BS k , Pcum k , DT k , Stage k

[0091] In the formula, is the state of each machine in the manufacturing system at the decision-making moment, which is a binary state, that is, ms = 1 means the machine is running, and ms = 0 means the machine is in a maintenance or idle state;

[0092] is the state of each buffer area in the manufacturing system at the decision-making moment. This state has two representation forms: First, bs i = 0, 1, 2…, sk Bi , that is, the inventory of the buffer area is used as the state output of the buffer area; Second, bs has three possible situations, bs = 0 represents that the inventory is empty, bs = 1 represents that the inventory is normal, and bs = 2 represents that the inventory exceeds the storage upper limit;

[0093] is the cumulative non-conforming product rate of each machine in the manufacturing system from the previous decision-making moment to the current decision-making moment;

[0094] is the relative running time of each machine in the manufacturing system at the current decision-making moment from the moment when the last maintenance is completed;

[0095] Stage​k = [k, k, …, k] indicates that the state obtained by the machine belongs to stage k, which is used to better make corresponding decisions at different stages.

[0096] The rewards include:

[0097]

[0098] In the formula, is the quality inspection cost of machine M i within the time period [0, t); c i buffer (t) is the penalty cost of buffer B i within the time period [0, t); r i (t) is the net value increase of the products processed by machine M i within the time period [0, t), r i (t) = N WIP (t) × r single, N WIP (t) is the number of qualified products processed by machine M i within the time period [0, t), r single is the net value increase for processing a single product; n m is the total number of machines.

[0099] S4. Construct an action space according to the optimization objectives, where the optimization objectives include the proportion of in-process products flowing downstream from each machine and the quality inspection ratio.

[0100] The action space includes:

[0101]

[0102] In the formula, is the proportion of in-process products flowing from machine 1 to machine 2 at the k-th stage; is the proportion of in-process products flowing from machines 4 and 5 to machine 6 at the k-th stage; is the quality inspection ratio of machine M1 at the k-th stage; is the quality inspection ratio of machine Mn m at the k-th stage.

[0103] S5. Based on the state space, action space, and rewards, and using the DDPG algorithm, train through the interaction between the intelligent agent and the manufacturing system until the training is stable, and then obtain the optimal rewards and actions.

[0104] A multi-stage manufacturing system joint optimization method provided by the present invention mainly performs joint optimization on a multi-stage manufacturing system based on deep reinforcement learning, and specifically includes the following steps:

[0105] Step 1: Analyze the component composition of the manufacturing system, and establish a degradation model and a processing quality model for common components in the manufacturing system, such as machines; establish an inventory model for buffer areas; and establish an inspection process model for inspection stations.

[0106] Step 2: Analyze the production operation rules of the manufacturing system and depict the production operation process. For example, how the machines in the manufacturing system transport work-in-progress to the downstream, how quality inspections are carried out, and when to inspect the manufacturing system.

[0107] Step 3: Use the DDPG algorithm to establish an optimization process framework. Take the manufacturing system as the environment, determine the state space returned by the environment, and determine the action space of the intelligent agent. Through continuous learning of the intelligent agent, obtain the optimal control strategy under the condition of the maximum reward.

[0108] Step 4: Compare and analyze the optimal results obtained in Step 3 with other manufacturing system operation strategies to confirm the superiority of the results obtained by this algorithm.

[0109] To illustrate a multi-stage manufacturing system joint optimization method provided by the present invention, it is described in conjunction with the accompanying drawings.

[0110] First, introduce the manufacturing system, such as Figure 1 As shown, it is a multi-stage manufacturing system. The squares represent machines, which are used to process and produce work-in-progress; the circles represent buffer areas, which are used to store the work-in-progress flowing from the upstream. It can be seen that the buffer areas correspond one-to-one with the machines, so the number of machines n m in the manufacturing system is equal to the number of buffer areas n b , that is, n b =n m . Then, model each component of the manufacturing system as follows:

[0111] (1) Machines. Assume that the machines have two degradation modes: basic degradation and severe degradation. Basic degradation process: Under the condition that the incoming material quality is qualified, the working load and working environment of the machine will be in a good stable state, and the machine degradation process under this condition is called the basic degradation process. Severe degradation process: Under the condition that the incoming material quality is unqualified, the working load or working environment of the machine will be disturbed to varying degrees, and the machine degradation process under this condition is called the severe degradation process. Both degradation modes follow the Weibull distribution. Assume that the cumulative failure probability distribution functions of the basic degradation and severe degradation failure processes are F b (t) and F s (t) respectively, and there is

[0112]

[0113] Among them, α b ∈(0, +∞) and α s ∈(0, +∞) are the scale parameters of the Weibull distribution, and β b ∈(0, +∞) and β s ∈(0, +∞) are the shape parameters of the Weibull distribution. Usually, the failure rate should be monotonically non-decreasing with time when the machine processes qualified feedstock, and the failure rate should be monotonically increasing with time when the machine processes unqualified feedstock. Therefore, the shape parameter should satisfy β b ≥1, β s >1. Denote the failure rates r b (t) and r s (t) of the machine during the basic degradation process and the strong degradation process as,

[0114] and In addition, it can be known from actual analysis that the degradation rate of the machine when processing unqualified feedstock should be faster than that when processing qualified feedstock, thus increasing the risk of machine failure, that is, satisfying the inequality F b (t) < F s (t). After simplifying the inequality, it can be obtained that the parameters of the two degradation models need to satisfy the inequality

[0115]

[0116] When the machine is under the mixed condition of processing qualified and unqualified feedstock, let the formulas of the failure probability distribution function F m (t) and the probability density function f m (t) of the machine under the mixed condition be

[0117] F m (t) = τF s (t) + (1 - τ)F b (t),

[0118] f m (t) = τf s (t) + (1 - τ)f b (t).

[0119] Among them, τ is the proportion of the strong degradation process of the machine, which is obtained by calculating the quantity proportion of the qualified and unqualified in-process products processed by the machine. The calculation formula is τ = n qf (t) / (n qf (t) + n uqf (t)), where n uqf (t) is the number of unqualified in-process products received (i.e., processed) by the machine within the time interval [0, t), and n qf (t) is the number of qualified in-process products received by the machine within the time interval [0, t).

[0120] Therefore, the failure rate r of the machine under the condition of mixed quality feed m (t) is

[0121]

[0122] At this time, the reliability R(t) of the machine is

[0123] R(t) = 1 - F m (t),

[0124] After obtaining the machine reliability, analyze and establish the machine processing quality model. Taking the machine performance as the influencing variable, conduct mathematical modeling and analysis on the probability of producing non-conforming products in the processing quality. Let the probability U(t) of the machine producing non-conforming products be

[0125]

[0126] where a, b, and c are constants, and r m (t) is the mixed failure rate of the machine. Analyze the above formula. When the machine is in an ideal state (i.e., r m (t) → 0), the probability that the machine processes work-in-progress into non-conforming products approaches a + b; when the machine is in a severely degraded state (i.e., r m (t) → ∞), then approaches 0, and then the probability that the machine produces non-conforming products approaches a. Therefore, the probability that the machine produces non-conforming products is

[0127] 0 ≤ a + b < U(t) < a ≤ 1.

[0128] From this, it can be deduced that the probability parameters need to satisfy the following conditions

[0129]

[0130] Based on the above conclusion, in order to meet the conditions, assume a = 1, b < 0 and b ∈ (0, 1]. That is, when the failure rate of the machine is large enough, all products processed by the machine are non-conforming products. Then when t = 0, the proportion u of the initial non-conforming products initial is

[0131]

[0132] Denote ρ = -b, then ρ can be the initial qualified proportion of the machine, that is, the initial quality level of the machine. Therefore, the probability U(t) that the machine produces non-conforming products can be rewritten as

[0133]

[0134] Among them, ρ ∈ (0, 1] is the initial quality level of the machine, and θ is a constant such that U(t) ∈ [0, 1].

[0135] That is, the probability that the machine produces non-conforming products is used as the processing quality model.

[0136] (2) Quality inspection. Usually, two types of inspection errors may occur during the quality inspection process: rejecting conforming products (Type I error) and accepting non-conforming products (Type II error). It is assumed that the occurrence of these two types of error events is independent of each other. The probability of a Type I error is denoted as p I , and the probability of a Type II error is denoted as p II . Therefore, the probability of accepting a conforming product, that is, the probability that the conforming product flows downstream, is 1 - p I , and the probability of rejecting a non-conforming product is 1 - p II . This invention assumes that there is a sampling inspection process after all machines are processed. When p I = 0 and p II = 1, it means that no sampling inspection activity is carried out. In addition, due to the rework activity of the manufacturing system of this invention, the work-in-progress judged to be of unqualified quality by the quality inspection activity in the model of this invention will no longer flow in the manufacturing system but will be removed from the manufacturing system.

[0137] Assume that the sampling ratio of the machine is s p ∈ [0, 1], and the two events of quality inspection of work-in-progress and judgment of the quality of work-in-progress are independent of each other. Then the probability of a Type I error for any work-in-progress with qualified quality is s p · p I ; the probability of a Type II error for any work-in-progress with unqualified quality is s p · p II . Then for work-in-progress with unknown quality status, the probability that it is a non-conforming product is U(t). Then the probability of making a wrong judgment on work-in-progress with unknown quality status during the inspection process is

[0138] p I∩II (t) = U(t) · p II + (1 - U(t)) · p I

[0139] That is, the probability of making a wrong judgment on work-in-progress with unknown quality status during the inspection process is the quality status inspection model of work-in-progress.

[0140] (3) Buffer. After machine M i finishes processing, the work-in-progress will be put into the buffer downstream of M i . If the upstream buffer is empty, then machine M i will be in a starving state. If the downstream buffer reaches its upper limit, then machine M iwill be blocked. Specifically, for buffer B1, the present invention assumes that its storage upper limit is infinite, and a batch of raw materials will flow in every once in a while, so that machine M1 will not be in a starvation state; for the machine M nm , which will not be in a blocked state. Let the productivity of machine M i be V Mi . When the machine is in a starvation, blocked, repaired or failed state, its productivity V Mi = 0; when the machine is in other states, the productivity V Mi is a fixed constant. Assume that the inventory of buffer B i at time t is sk Bi (t), then its inventory sk Bi (t') at the next detection point is,

[0141]

[0142] where n i upstream is the number of upstream machines of machine M i . dt i is the relative running time of machine M i , and its calculation formula is as follows, t delta is the relative time occupied by the machine due to maintenance, starvation, etc.

[0143] dt=(t - t′)-t delta

[0144] After determining the model of the components, combined with Figure 1 , in the present invention, n m = 9. Assume that the initial raw materials of the manufacturing system, that is, the feedstock flowing into machine M1, are all qualified raw materials. The inventory of each buffer at the initial time t = 0 is also considered to be qualified. Then, for this manufacturing system in the kth stage, the evaluation process framework of the production process with a stage production duration of Δt is as Figure 2 shown, and the main steps of this evaluation method are as follows:[[]]

[0145] Step 1: Calculate the proportion τ i , i = 1, 2,..., n m of the unqualified product feedstock of all machines during this time period.[[]]

[0146] Step 2: Calculate the mixed failure probability distribution function F m (t), the mixed probability density function f m (t) and the actual failure rate r m (t) of all machines during this time period according to the formula. Among them, t = k·Δt.[[]]

[0147] Step 3: Calculate the number of non-conforming products and conforming products processed by all machines in this stage, and conduct quality inspections on the work-in-progress products of various qualities.

[0148] Step 4: Use different methods to determine which downstream machine the work-in-progress products processed by each machine flow to.

[0149] Among them, in this application, it is obtained through intelligent training, that is, the proportion given by the DDPG algorithm flows randomly. That is, the proportion given by the DDPG algorithm for machine 1 to machine 2 is 30%, then the proportion for machine 1 to machine 3 is 70%, and the work-in-progress products produced by machine 1 will flow downstream randomly according to the ratio of 3:7.

[0150] After determining the production process of the manufacturing system, based on Figure 1 the multi-stage manufacturing system shown, construct a manufacturing system environment under dynamic production flow. It is set that the intelligent agent Agent will obtain the state observation value and the cumulative revenue of the manufacturing system within Δt time from the manufacturing system every time the manufacturing system runs for Δt time, and determine the weight of the work-in-progress products produced by each machine in the manufacturing system flowing downstream. In the present invention, each moment after experiencing Δt time is called a checkpoint of the manufacturing system, and it is defined that t = kΔt, k = 0, 1, 2, 3,.... is the decision-making moment. At the kth decision-making moment, Agent can obtain the following state observation values from the manufacturing network,

[0151] state k =[MS k ,BS k ,Pcum k ,DT k ,Stage k

[0152] Among them, is the state of each machine in the manufacturing system at the decision-making moment t = kΔt, which is a binary state, that is, ms = 1 means the machine is running, and ms = 0 means the machine is in maintenance or idle state;

[0153] is the state of each buffer area in the manufacturing system at the decision-making moment t = kΔt. This state has two representation forms: First, bs i = 0, 1, 2…, sk Bi , that is, the inventory of the buffer area is used as the state output of the buffer area; Second, bs has three possible situations, bs = 0 represents that the inventory is empty, bs = 1 represents that the inventory is normal, and bs = 2 represents that the inventory exceeds the storage upper limit.

[0154] ​The cumulative nonconforming product rate of each machine in the manufacturing system within the time interval [t - Δt, t) (t = kΔt) is used to describe the quality status of the manufacturing network; the cumulative nonconforming product rate p cum (t) is defined as the ratio of the number of nonconforming work-in-processes processed by all machines to the total number of work-in-processes processed within the time interval [0, t) in the manufacturing system. Its calculation method is as follows:

[0155]

[0156] where n i qf (t) is the number of conforming work-in-processes received by machine M i within the time interval [0, t), and n i uqf (t) is the number of nonconforming work-in-processes received by machine M i within the time interval [0, t).

[0157] is the relative running time of each machine in the manufacturing system at the current decision-making moment t = kΔt from the moment when the last maintenance was completed.

[0158] Stage k = [k, k, …, k] indicates that the state obtained by the machine belongs to stage k, which is used to make corresponding decisions better at different stages. It should be noted that the cumulative nonconforming product rate Pcum k is a continuous state space, and the rest are discrete state spaces. Therefore, the constructed state space of the manufacturing system belongs to a state space that combines continuous and discrete states.

[0159] In the present invention, it is assumed that at any decision-making moment t = kΔt, the Agent needs to determine the proportion of work-in-process flowing downstream from each machine based on the observed state and reward information of the manufacturing system. Therefore, the action of the Agent to allocate the proportion of the flow direction of work-in-process for the next operating stage [kΔt, (k + 1)Δt) of the manufacturing system at each decision-making moment t = kΔt is denoted as

[0160]

[0161] where represents the proportion of work-in-process flowing from machine 1 to machine 2 in the k-th stage [kΔt, (k + 1)Δt), then the proportion of work-in-process flowing from machine 1 to machine 3 is Similarly, the proportion of work-in-process processed by machines 4 and 5 flowing to machines 6 and 7 can be deduced. represents the quality inspection proportion of machine Mn m in the k-th stage [kΔt, (k + 1)Δt).

[0162] Suppose the reward obtained by the manufacturing system in each operation stage [kΔt, (k + 1)Δt) (or the k-th stage) is

[0163]

[0164] where r i (t) is the net value increase of the products processed by machine M i within the time interval [0, t), and r i (t) = N WIP (t)

[0165] ×r single . N WIP (t) is the number of qualified products processed by machine M i within the time interval [0, t), and r single is the net value increase for processing a single product. is the quality inspection cost of machine M i within the time interval [0, t), n ins is the number of quality inspection activities carried out by this machine within the time interval [0, t); c inspect is the cost when the machine conducts 100% quality inspection; s j p is the sampling inspection ratio for the j-th quality inspection of the machine. c i buffer (t) is the penalty cost of buffer B i within the time interval [0, t), where c over is the penalty for a single stock when the buffer exceeds the capacity limit. n over,i (t) is the number of times the buffer capacity limit is exceeded within the time interval [0, t), and its calculation formula is as follows. In the following formula, N Bi is the upper limit of the inventory capacity of buffer B i .

[0166]

[0167] It can be seen that only after-sales maintenance, quality inspection, and weight allocation are considered in this production activity. Based on the constructed environment, corresponding state space, action space, and reward, the basic process of the DDPG algorithm considering the interaction between the agent and the manufacturing system is as follows, and its main steps are as follows:

[0168] Step 1 Initialization: Initialize the Actor network, Critic network, Target Actor network, Target Critic network of the DDPG algorithm, and the experience replay buffer; Initialize the relevant parameters of the machines, buffers, etc. in the manufacturing system. Set the iteration number iteration = 1, stage k = 1, and the current running time t of the manufacturing system = kΔt.

[0169] Step 2 Interaction between the Agent and the manufacturing system:

[0170] Step 2.1 If k = 1, generate a random action. Otherwise, according to the current environmental state (state), obtain the behavior action through the Actor network.

[0171] Step 2.2 The manufacturing system conducts production activities according to the action: After each machine produces work-in-progress, the work-in-progress flows downstream according to the weight. Obtain the state (state') of the manufacturing system after the action is executed.

[0172] Step 2.3 Collect data, and pack the state (state), action, reward, and next state (state') into a tuple (state, action, reward, state') and put it into the experience pool.

[0173] Step 2.4 If the number of tuples in the experience pool meets the batch requirement, then take multiple tuples as a batch and put them into the neural network for learning.

[0174] Step 2.5 Determine whether the current running time t of the manufacturing system is greater than the total running time of the manufacturing system: If it is greater, jump to Step 3. Otherwise, jump to Step 2.1, continue the production activities of the manufacturing system in this iteration, and k = k + 1.

[0175] Step 3 Record the data of this iteration of iteration, and record the total profit of the manufacturing system after completing the manufacturing system production task in this iteration.

[0176] Step 4 Determine whether the iteration number iteration is greater than the maximum iteration number max_iteration: If it is greater, complete the training task and end the algorithm. Otherwise, jump to Step 2, iteration = iteration + 1, k = 1, and continue the training iteration.

[0177] In summary, the overall framework of the multi-stage manufacturing system joint optimization method based on deep reinforcement learning proposed by the present invention is as Figure 3As shown, after initializing the neural network and various parameters of the algorithm, the agent interacts with the established manufacturing system. After the interaction, the agent conducts learning and training until the optimal reward and action are obtained after stable training.

[0178] To further illustrate a multi-stage manufacturing system joint optimization method provided by the present invention, it will be described in conjunction with the accompanying drawings and specific examples.

[0179] This embodiment is based on Figure 1 the system structure shown, and the learning rates of the Actor network and the Critic network are set to 3×10 -3 and 3×10 -4 respectively. The capacity of the experience replay pool is 1×10 4 . When the network is updated, the minimum batch of experiences extracted is batch = 200, the discount factor γ = 0.1, and the smoothing factor τ = 1×10 -1 . For the number of layers of the neural network, the input of the Critic network is the environmental state State and the action Action, and the input of the Actor network is the environmental state State. Both networks have 3 layers, with 128 nodes in each layer. The production cycle of the manufacturing system is Δt = 100, r single = 8, c over = 100, c inspect = 40, and a total of n stage = 10 stages are carried out, then t total = 1000. Except for the first buffer, for all other buffers, if the buffer stock exceeds 100, a penalty of c over will be generated for each additional in-process product in the buffer. The remaining parameters are shown in Tables 1 to 4 below.

[0180] Table 1 Parameters related to machine degradation and processing quality

[0181]

[0182] Table 2 Parameters related to machine quality inspection activities

[0183]

[0184] Table 3 Parameters related to machine maintenance activities

[0185]

[0186]

[0187] Table 4 Parameters related to buffer

[0188]

[0189] Figure 4 Shown is a line chart of the average system reward varying with the number of training times during the DDQN training process. In the initial stage of training (about the first 100 times), the average reward rises rapidly, indicating that the system is quickly learning and improving its strategy. Subsequently, the average reward fluctuates between approximately 8000 and 12000, showing that the system has reached a certain performance level but is still trying to further optimize. Although the average reward fluctuates within a certain range, there is no obvious downward trend overall, meaning that the system is exploring different strategies to find a better solution, indicating that the algorithm has converged.

[0190] Figure 5 It is the change trend of the sampling inspection ratio (blue line) and the cumulative nonconforming product rate (red line) of the intelligent agent decision-making machines 1 - 9 at different stages after the DDQN algorithm training. Among them Figure 5 (a) - (i) in it are respectively the change trends of the sampling inspection ratio (blue line) and the cumulative nonconforming product rate (red line) of machines 1 - 9 at different stages. It can be seen that the increase in the sampling inspection ratio seems to be related to the rise in the cumulative nonconforming product rate, which may indicate that more inspections are needed to control quality at these stages. It proves the effectiveness of the intelligent agent trained by the algorithm. Although some machines do not follow this relationship at different stages, generally speaking, it can show that the basis for the change of the intelligent agent's quality inspection ratio is related to the size of the cumulative nonconforming product rate.

[0191] Figure 6 Shown is the change trend of in-process products flowing to machines 2, 3 and machines 6, 7 at different stages. It can be seen that the red lines and blue lines in both subgraphs show certain periodic fluctuations, which may be related to the operating cycle or production cycle of the machines. Combining with Table 1, it can be known that the production speed ratios of machine 2 and machine 3 and machine 6 and machine 7 are the same, both being 3:4. And the change trend of the ratio of in-process products flowing to machine 3 and flowing to machine 7 is similar. Similarly, the change trend of the ratio of in-process products flowing to machine 2 and flowing to machine 6 is similar. This also proves the rationality of the learning results of the intelligent agent Agent.

[0192] In addition, in order to further illustrate the advantages of the reinforcement learning method, the present invention designs the following two methods for the in-process products processed by the upstream machine to flow to the downstream:

[0193] Method 1: The in-process products flow to the downstream equipment in the way of equal-probability random sampling. If there is only one corresponding downstream equipment for this machine, the processed in-process products directly flow to this downstream equipment; if the corresponding downstream equipment of this machine is more than 1, the processed in-process products will flow to a certain downstream equipment randomly with equal probability.

[0194] Method 2: The work-in-progress flows downstream according to the status of downstream machines. If a downstream machine is starved, the work-in-progress preferentially flows to this machine with a priority of 2; if a downstream machine is blocked, the priority of the work-in-progress flowing to this machine decreases, with a priority of 0; in addition, the downstream machines in other states have a priority of 1. The work-in-progress selects the downstream machine to flow into according to the priority of 2>1>0. If the number of machines with the same priority is greater than 1, the selection is made according to Method 1.

[0195] Simulations were carried out on these three scheduling methods. It should be noted that the quality inspection ratios of each machine refer to the DDPG algorithm, and the other parameter settings are also the same. The obtained results are as Figure 7 shown in (a) and (b) below, respectively showing the cumulative reward and stage reward varying with the stage under different methods. In Figure 7 shown in (a) below, in the cumulative reward, as the stage increases from 1 to 9, the reinforcement learning method shows the most stable growth. The final cumulative reward is close to 14000, which is significantly higher than the other two strategies. The cumulative reward growth trend of the work-in-progress flowing randomly downstream is similar to that of the reinforcement learning method, but the final cumulative reward is slightly lower than that of the reinforcement learning method, about 12000. The cumulative reward determined by the machine status grows the slowest, and the final cumulative reward is less than 10000. In Figure 7 shown in (b) below, the performance of the three strategies fluctuates greatly. However, the reinforcement learning method is more stable than the other two, indicating that in the later stage of the simulation, the processing quality of the machines decreases, and the reinforcement learning method can well avoid this problem and strive to keep the system reward stable.

[0196] The present invention provides a multi-stage manufacturing system joint optimization system, including:

[0197] A model establishment module for establishing a machine degradation model and a processing quality model according to the mutual influence between product quality and machine reliability in the manufacturing system; establishing an inventory model for the buffer area; and establishing a work-in-progress quality status inspection model;

[0198] A space construction module for constructing a state space and the reward of the manufacturing system according to the degradation model, the processing quality model, the inventory model, and the work-in-progress quality status inspection model; constructing an action space according to the optimization objectives, where the optimization objectives include the proportion of work-in-progress flowing downstream to each machine and the quality inspection ratio;

[0199] An interaction training module for, based on the state space, the action space, and the reward, and based on the DDPG algorithm, training through the interaction between the intelligent agent and the manufacturing system until the training is stable, and then obtaining the optimal reward and actions.

[0200] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A joint optimization method for a multi-stage manufacturing system, characterized in that, Optimization of the in - process flow behavior and quality inspection ratio in a multi - stage manufacturing system, where the manufacturing system includes multiple series - connected machines and a buffer area for storing in - process products corresponding to each machine. The optimization method includes: Establish a machine degradation model and a processing quality model based on the mutual influence between product quality and machine reliability in the manufacturing system; Establish an inventory model for the buffer area; and establish an in - process product quality status inspection model; Construct a state space and the reward of the manufacturing system according to the degradation model, the processing quality model, the inventory model, and the in - process product quality status inspection model; Construct an action space according to the optimization objective, where the optimization objective includes the proportion of in - process products flowing downstream from each machine and the quality inspection ratio; According to the state space, the action space, and the reward, based on the DDPG algorithm, through the interaction and training between the intelligent agent and the manufacturing system, until the training is stable, obtain the optimal reward and actions.

2. The joint optimization method for a multi-stage manufacturing system according to claim 1, wherein When establishing the machine degradation model, the machine degradation modes include basic degradation and strong degradation; the machine degradation model includes: obtaining the failure probability distribution function and probability density function of the machine under the mixed mode of basic degradation and strong degradation; obtaining the mixed failure rate of the machine according to the failure probability distribution function and probability density function of the machine; and obtaining the reliability of the machine according to the failure probability distribution function of the machine.

3. The joint optimization method for a multi-stage manufacturing system according to claim 2, characterized in that, The failure probability distribution function of the machine is obtained according to the cumulative failure probability distribution function of the machine in the basic degradation failure process and the cumulative failure probability distribution function of the machine in the strong degradation failure process; The probability density function of the machine is obtained according to the cumulative probability density function of the machine in the basic degradation failure process and the cumulative probability density function of the machine in the strong degradation failure process; Among them, basic degradation means that under the condition of qualified incoming material quality, the working load and working environment of the machine will be in a stable state, and the machine degradation process under this condition is called the basic degradation process; Strong degradation means that under the condition of unqualified incoming material quality, the working load or working environment of the machine will have different degrees of disturbances, and the machine degradation process under this condition is called the strong degradation process.

4. The joint optimization method of the multi-stage manufacturing system according to claim 2, characterized in that The processing quality model is obtained according to the mixed failure rate of the machine, and the calculation formula is as follows: where ρ ∈ (0, 1] is the initial quality level of the machine; θ is a constant such that U(t) ∈ [0, 1]; r m (t) is the mixed failure rate of the machine; e is the natural constant.

5. The joint optimization method of the multi-stage manufacturing system according to claim 1, wherein Establishing an inventory model for the buffer area includes: where n i upstream is the number of upstream machines of machine M i ; dt i For machine M i 's relative running time; t delta The relative time occupied by the machine due to maintenance and starvation reasons; Is buffer B i Inventory at time t; For buffer B i The inventory at the next detection time point t′; For machine M i production rate; · For machine M j production rate; i and j respectively represent the numbers corresponding to the machines or buffer areas, M i and M j represent the machines numbered i and j, B i and B j represent the buffer areas numbered i and j.

6. The joint optimization method for a multi-stage manufacturing system according to claim 1, characterized in that, Establishing an in - process product quality status inspection model includes: p I∩II (t) = U(t)·p II +(1 - U(t))·p I where p I∩II (t) is the probability of making a wrong judgment on the work-in-progress with unknown quality status during the detection process; U(t) is the in - process product with unknown quality status, and its probability of being a non - conforming product; p I The probability of rejecting qualified products during the quality inspection process; p II is the probability of accepting nonconforming products during the quality inspection process.

7. The joint optimization method for a multi-stage manufacturing system according to claim 1, characterized in that The state space includes: state k = [MS k , BS k , Pcum k , DT k , Stage k ​ wherein, represents the state of each machine in the manufacturing system at the decision-making moment, which is a binary state, i.e., ms = 1 indicates that the machine is running, and ms = 0 indicates that the machine is in a maintenance or idle state; For the state of each buffer in the manufacturing system at the decision-making moment, there are two representation forms for this state: First, bs i = 0, 1, 2…, sk Bi , that is, the inventory of the buffer is used as the state output of the buffer; Second, there are three possible situations for bs. bs = 0 represents that the inventory is empty, bs = 1 represents that the inventory is normal, and bs = 2 represents that the inventory exceeds the storage upper limit; is the cumulative nonconforming product rate of each machine in the manufacturing system between the previous decision-making moment and the current decision-making moment; The relative running time of each machine in the manufacturing system at the current decision-making moment from the moment when the last maintenance was completed; Stage k = [k, k, …, k] indicates that the state obtained by the machine belongs to stage k, which is used to make corresponding decisions better under different stages.

8. The joint optimization method for a multi-stage manufacturing system according to claim 1, wherein The action space includes: In the formula, is the proportion of the in-process products flowing from machine 1 to machine 2 at the k-th stage; is the proportion of the in-process products flowing from Machine 4 and Machine 5 to Machine 6 in the k-th stage; is the quality inspection ratio of machine M1 in the k-th stage; is the quality inspection ratio of machine Mn at the k-th stage m ​ 9. The joint optimization method for a multi-stage manufacturing system according to claim 1, wherein, The reward includes: Wherein, is the machine M i is the quality inspection cost within the time period [0, t); c i buffer (t) is the penalty cost of the buffer B i within the time period [0, t); r i (t) is the increased net value of the products processed by the machine M i within the time period [0, t), r i (t) = N WIP (t) × r single, N WIP (t) is the number of qualified products processed by the machine M i within the time period [0, t), r single is the increased net value of processing a single product; n m is the total number of machines.

10. A multi-stage manufacturing system joint optimization system, characterized in that, Includes: A model - building module for establishing a machine degradation model and a processing quality model based on the mutual influence between product quality and machine reliability in the manufacturing system; Establish an inventory model for the buffer area; and establish an in - process product quality status inspection model; A space - construction module for constructing a state space and the reward of the manufacturing system according to the degradation model, the processing quality model, the inventory model, and the in - process product quality status inspection model; Construct an action space according to the optimization objective, where the optimization objective includes the proportion of in - process products flowing downstream from each machine and the quality inspection ratio; An interactive training module, which is used to interact and train with a manufacturing system based on the DDPG algorithm according to the state space, action space, and reward, and obtain the optimal reward and action until the training is stable.