Local disturbance-oriented workshop dynamic scheduling method and system, medium and equipment

By quantitatively analyzing the availability of key equipment in the aerospace manufacturing workshop and establishing a local scheduling model, combining rules combinations and deep reinforcement learning algorithms, an optimized scheduling solution is generated, which solves the problem of low equipment availability affecting production continuity and achieves more efficient production scheduling.

CN120010403APending Publication Date: 2025-05-16SHANGHAI SHENJIAN PRECISION MASCH TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510022568.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The aerospace manufacturing workshop is facing a large amount of production tasks for key equipment and a rapid decline in performance of parts, resulting in low equipment availability and affecting the continuity of the production process. It is difficult for the existing technology to effectively solve this problem.

Method used

By quantitatively analyzing the availability of key devices, a local scheduling model considering the availability of equipment is established, a local scheduling algorithm based on rule combination and deep reinforcement learning is designed, and a scheduling scheme for process task reassignment and process task resorting within the scope of local limited equipment is generated.

Benefits of technology

It improves the executability of the scheduling scheme, enhances equipment availability considerations, optimizes the total weighted completion time of the workpiece, and improves the continuity of the production process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010403A_ABST
    Figure CN120010403A_ABST
Patent Text Reader

Abstract

The invention provides a local disturbance-oriented workshop dynamic scheduling method and system, a medium and equipment, and the method comprises the steps: proposing a key equipment availability analysis method, and calculating the average fault time, average repair time and other availability values of the equipment; establishing a local scheduling model considering the availability of the equipment, determining constraint conditions, and constructing an optimization target of the minimum total weighted completion time of the workpieces; and designing a local scheduling algorithm based on rule combination and deep reinforcement learning, and generating a local scheduling scheme for process task reassignment and process task reordering in a local limited equipment range. According to the method, aiming at the multi-variety and small-batch production characteristics of aerospace products, local disturbances such as equipment faults are considered, quantitative analysis of equipment availability during production task allocation is realized through key equipment availability analysis, and dynamic scheduling optimization under local disturbances is realized based on a local scheduling algorithm of rule combination and deep reinforcement learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of workshop scheduling, and in particular to a method, system, medium and equipment for dynamic workshop scheduling facing local disturbance. Background Art

[0002] The manufacturing model of aerospace products presents the characteristics of multi-variety and small-batch production, and faces the production demand of batch production and research and development model tasks on the same line. In recent years, the number of batch production and research and development model orders for aerospace manufacturing has continued to grow, and the need to ensure the continuity of the production process under product diversification is facing important challenges.

[0003] Faced with the higher requirements for continuous production capacity of aerospace manufacturing workshops, multi-task, high-load manufacturing equipment is the key manufacturing resource that needs attention. Compared with other equipment, key equipment has more production tasks, and the performance of its key components decays faster. In order to maintain production performance, key equipment requires more frequent repairs and maintenance. Repairs and maintenance require the use of equipment, and the production tasks of key equipment are compactly arranged. Production tasks and unavailable time periods are prone to conflict, causing the scheduling plan to fail and thus affecting the continuity of the production process. Therefore, how to take the availability of key equipment into account in the production scheduling link to improve the executability of the scheduling plan is the focus of research.

[0004] Patent application document CN114186791A discloses a dynamic scheduling method for assembly and adjustment production of complex equipment products with multiple models and small batches. It is aimed at dynamic assembly and adjustment production scenarios with multiple models on the same line and multiple disturbance factors coexisting. It establishes a dynamic assembly and adjustment production scheduling model with the goal of weighted minimization of processing cost, delivery error and cross-workshop transfer times, and designs a hybrid variable neighborhood solution algorithm to solve the model, wherein the particle swarm optimization algorithm is introduced to ensure the local optimum within the neighborhood structure set; the neighborhood structure set is systematically changed to expand the search range and ensure the global optimum; the latest hybrid variable neighborhood algorithm key parameters are obtained through the BP neural network, so that the entire algorithm can adapt to the changing environment. However, this patent cannot completely solve the existing technical problems, nor can it meet the needs of the present invention. Summary of the invention

[0005] In view of the defects in the prior art, the purpose of the present invention is to provide a workshop dynamic scheduling method, system, medium and equipment for local disturbance.

[0006] The workshop dynamic scheduling method for local disturbance provided by the present invention includes:

[0007] Step 1: Quantitatively analyze the availability of key equipment in the production process and calculate the equipment availability indicators, including mean time between failures and mean time to repair;

[0008] Step 2: Establish a local scheduling model that takes equipment availability into consideration, determine constraints, and construct an optimization objective that minimizes the total weighted completion time of workpieces;

[0009] Step 3: Design a local scheduling algorithm based on rule combination and deep reinforcement learning to generate a local scheduling solution for process task reassignment and process task reordering within a local limited equipment range.

[0010] Preferably, the step 1 comprises:

[0011] Use the procurement data of key functional components to establish a reliability description model and calculate the MTBF index of the equipment;

[0012] In the reliability description model, a three-parameter Weibull distribution including a scale parameter, a shape parameter, and a location parameter is used to describe the failure probability change of mechanical equipment and its components, and an equipment failure probability density function is constructed, as shown in the following formula:

[0013]

[0014] Among them, γ is the position parameter, which means that failure will not occur within the initial γ time; η is the scale parameter, which indicates the scaling degree of the curve; β is the shape parameter, which indicates the trend of the curve; t is the continuous use time for production;

[0015] The equipment failure probability density function is integrated to obtain a cumulative distribution function, as shown in the following formula:

[0016]

[0017] Using the failure sample data of functional components, the genetic algorithm is used to obtain the γ, η, and β parameter values ​​of the best fitting failure samples, and the cumulative distribution function of the failure probability is calculated to obtain the reliability function, as shown in the following formula:

[0018]

[0019] Where R(t) represents the reliability of the functional component when the continuous production time is t. For a manufacturing device containing C functional components, the reliability function R is calculated based on the failure samples and the genetic algorithm. 1 (t),R 2 (t),…,R c (t)…,R C (t), and thus the equipment reliability change trend is obtained:

[0020]

[0021] According to the equipment reliability change trend, the expected average working time between two equipment failures is calculated:

[0022]

[0023] According to the MTBF value and the MTTR statistical information in the equipment maintenance record table, a quantitative analysis result of the equipment availability is obtained.

[0024] Preferably, the optimization objective of minimizing the total weighted completion time of the artifacts in step 2 is expressed as:

[0025]

[0026] Among them, i represents the workpiece number, I represents the total number of workpieces, ω i represents the weighting coefficient of workpiece i, F i represents the completion time of job i.

[0027] Preferably, the step 3 comprises:

[0028] Construct a Markov decision process model corresponding to the local scheduling process. By analyzing the constraints and optimization goals faced by the local scheduling mathematical model, design the workshop state and reward function in the Markov decision model. By establishing a scheduling rule library for assigning process tasks to equipment and sorting process tasks on equipment, design scheduling actions with rule combinations.

[0029] By using the improved deep Q-network algorithm, through establishing a deep learning network, designing the loss function of reinforcement learning, and embedding the experience replay mechanism, the weight coefficient optimization in the rule combination process is realized, the reasonable combination of scheduling rules is obtained, and the local scheduling plan that minimizes the weighted completion time is obtained.

[0030] The workshop dynamic scheduling system oriented to local disturbance provided by the present invention comprises:

[0031] Module M1: Quantitatively analyze the availability of key equipment in the production process and calculate the availability indicators of the equipment, including mean time between failures and mean time to repair;

[0032] Module M2: Establish a local scheduling model that takes equipment availability into account, determine constraints, and construct an optimization objective that minimizes the total weighted completion time of workpieces;

[0033] Module M3: Design a local scheduling algorithm based on rule combination and deep reinforcement learning to generate a local scheduling solution for process task reassignment and process task reordering within a local limited equipment range.

[0034] Preferably, the module M1 comprises:

[0035] Use the procurement data of key functional components to establish a reliability description model and calculate the MTBF index of the equipment;

[0036] In the reliability description model, a three-parameter Weibull distribution including a scale parameter, a shape parameter, and a location parameter is used to describe the failure probability change of mechanical equipment and its components, and an equipment failure probability density function is constructed, as shown in the following formula:

[0037]

[0038] Among them, γ is the position parameter, which means that failure will not occur within the initial γ time; η is the scale parameter, which indicates the scaling degree of the curve; β is the shape parameter, which indicates the trend of the curve; t is the continuous use time for production;

[0039] The equipment failure probability density function is integrated to obtain a cumulative distribution function, as shown in the following formula:

[0040]

[0041] Using the failure sample data of functional components, the genetic algorithm is used to obtain the γ, η, and β parameter values ​​of the best fitting failure samples, and the cumulative distribution function of the failure probability is calculated to obtain the reliability function, as shown in the following formula:

[0042]

[0043] Where R(t) represents the reliability of the functional component when the continuous production time is t. For a manufacturing device containing C functional components, the reliability function R is calculated based on the failure samples and the genetic algorithm. 1 (t),R 2 (t),…,R c (t)…,R C (t), and thus the equipment reliability change trend is obtained:

[0044]

[0045] According to the equipment reliability change trend, the expected average working time between two equipment failures is calculated:

[0046]

[0047] According to the MTBF value and the MTTR statistical information in the equipment maintenance record table, a quantitative analysis result of the equipment availability is obtained.

[0048] Preferably, the optimization goal of minimizing the total weighted completion time of the artifacts in the module M2 is expressed as:

[0049]

[0050] Among them, i represents the workpiece number, I represents the total number of workpieces, ω i represents the weighting coefficient of workpiece i, F i represents the completion time of job i.

[0051] Preferably, the module M3 comprises:

[0052] Construct a Markov decision process model corresponding to the local scheduling process. By analyzing the constraints and optimization goals faced by the local scheduling mathematical model, design the workshop state and reward function in the Markov decision model. By establishing a scheduling rule library for assigning process tasks to equipment and sorting process tasks on equipment, design scheduling actions with rule combinations.

[0053] By using the improved deep Q-network algorithm, through establishing a deep learning network, designing the loss function of reinforcement learning, and embedding the experience replay mechanism, the weight coefficient optimization in the rule combination process is realized, the reasonable combination of scheduling rules is obtained, and the local scheduling plan that minimizes the weighted completion time is obtained.

[0054] According to the computer-readable storage medium storing a computer program provided by the present invention, when the computer program is executed by a processor, the steps of the local disturbance-oriented workshop dynamic scheduling method are implemented.

[0055] The electronic device provided according to the present invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the local disturbance oriented workshop dynamic scheduling method are implemented.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] (1) Aiming at the local scheduling problem of aerospace manufacturing workshops, the present invention considers dynamic disturbance events such as abnormalities of key resources such as equipment, designs a local scheduling model that considers equipment availability constraints, and proposes an optimization goal of minimizing the total weighted completion time of workpieces. This fully reflects the main contradiction of dynamic disturbances in aerospace manufacturing workshops and has practical significance for solving similar problems in enterprises.

[0058] (2) The present invention achieves quantitative analysis of the availability of key equipment by constructing a three-parameter Weibull failure probability distribution model and estimating the parameters;

[0059] (3) The present invention improves the feasibility of the scheduling scheme by establishing a scheduling rule base for assigning process tasks to equipment and sorting equipment process tasks, constructing a Markov decision process model, and constructing a scheduling agent based on a deep Q network algorithm to output process sorting and equipment assignment schemes that meet equipment availability constraints. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0061] Figure 1 It is a flow chart of the aerospace manufacturing workshop dynamic scheduling method facing local disturbance of the present invention;

[0062] Figure 2 It is a framework diagram of a local scheduling algorithm based on rule combination and deep reinforcement learning in an embodiment of the present invention;

[0063] Figure 3 Flow chart of an improved deep Q network (D3QN) algorithm in an embodiment of the present invention;

[0064] Figure 4 It is a bar chart of comparative experimental results of local scheduling methods in embodiments of the present invention. DETAILED DESCRIPTION

[0065] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0066] Example 1

[0067] The technical problem to be solved by the present invention is: to provide a dynamic scheduling method for aerospace manufacturing workshops oriented to local disturbances, which on the one hand realizes the quantitative analysis of equipment availability under local disturbances; on the other hand, based on deep reinforcement learning, it realizes the dynamic scheduling optimization of aerospace manufacturing workshops considering the availability constraints of key equipment, and outputs the process sorting and equipment assignment scheme that meets the equipment availability constraints.

[0068] The technical solution adopted by the method of the present invention is: a dynamic scheduling method for aerospace manufacturing workshop facing local disturbances, such as Figure 1 , including the following steps:

[0069] Step 1: Quantitatively analyze the availability of key equipment in the production process and calculate the availability indicators such as the average failure time and average repair time of the equipment;

[0070] The equipment availability is an important indicator that comprehensively reflects maintainability, reliability and security. Its calculation formula is as follows:

[0071]

[0072] Among them, MTBF (Mean Time Between Failures) represents the average working time between two adjacent failures of the equipment; MTTR (Mean Time To Repair) represents the average recovery time from the occurrence of a failure to the recovery of the equipment to an operational state. The MTBF and MTTR are used to quantitatively describe the availability of the equipment as a basis for local scheduling.

[0073] The main cause of equipment failure is the loss and degradation of key functional components. Therefore, a reliability description model is established based on the procurement data of key functional components, and then the MTBF index of the equipment is calculated. In the reliability description, a two-parameter Weibull distribution containing a scale parameter and a shape parameter is used to describe the change in the failure probability of mechanical equipment and its components to achieve better fitting accuracy. On this basis, the three-parameter Weibull distribution adds a position parameter to further improve the fitting ability of different failure modes. Therefore, the present invention constructs the equipment failure probability density function (ProbabilityDensity Function, PDF) as shown in the following formula:

[0074]

[0075] Among them, γ is the position parameter, which means that failure will not occur within the initial γ time; η is the scale parameter, which indicates the scaling degree of the curve; β is the shape parameter, which indicates the trend of the curve; is the continuous use time for production. The failure probability of mechanical manufacturing equipment will accumulate over time. By integrating the equipment failure probability density function, we can get the cumulative distribution function (CDF) shown in the following formula:

[0076]

[0077] According to the failure sample data of the functional components, reasonable three parameters γ, η, and β are fitted to construct the cumulative distribution function of the failure probability. The present invention proposes a Weibull distribution parameter optimization method based on a genetic algorithm. By using the genetic algorithm, the γ, η, and β parameter values ​​of the best fitting failure samples are obtained, and the cumulative distribution function of the failure probability shown in formula (3) is calculated, and then the reliability function shown in formula (4) is obtained:

[0078]

[0079] Where R(t) represents the reliability of the functional component when the continuous production time is t. For a manufacturing device containing C functional components, the reliability function R can be calculated based on the failure samples and genetic algorithm. 1 (t),R 2 (t),…,R c (t)…,R C (t), and thus the equipment reliability change trend is obtained:

[0080]

[0081] According to the changing trend of equipment reliability, the expected average working time between two equipment failures can be calculated according to formula (5):

[0082]

[0083] Based on the MTBF value and the MTTR statistical information in the equipment maintenance record table, the quantitative analysis results of equipment availability are achieved.

[0084] Step 2: Establish a local scheduling model that takes into account equipment availability constraints and construct an optimization goal to minimize the total weighted completion time of workpieces. The parameters and decision variables in the mathematical model are defined as shown in the table:

[0085] Table 1 Parameters and decision variables of the dynamic scheduling process

[0086]

[0087] According to the above parameter definitions, the local scheduling problem can be described as the following model:

[0088] Objective function:

[0089]

[0090] Constraints:

[0091]

[0092] Formula (7) represents the optimization goal of minimizing the total weighted completion time of the workpiece; Formula (8) indicates that each production process can only be produced on one device; Formula (9) describes that the process cannot be interrupted when it is produced on the device; Formula (10) ensures that one device can only produce one process of one workpiece at a time, and cannot process multiple processes of multiple workpieces at the same time; Formula (11) describes the process sequence constraint caused by the workpiece process route; Formula (12) describes that the completion time of the workpiece is determined by the completion time of the last process task; Formula (13) describes that the total number of processes to be scheduled on the equipment is determined by the number of processes to be scheduled for all workpieces; Formulas (14) and (15) ) ensures that each process of each workpiece will be given a sorting position; Formula (16) ensures that at the time of production scheduling, the start time of the first process on each device must be greater than the initial remaining occupancy time, that is, it is necessary to wait for the device to complete the existing production task; Formula (17) describes that the processes with a lower sorting on the same device need to receive production operations after the processes with a higher sorting; Formula (18) describes the average equipment failure interval obtained by equipment availability analysis; Formula (19) describes the equipment failure repair status obtained by equipment availability analysis; Formulas (20) and (21) indicate that when the equipment is in a fault repair state, the process production task cannot be performed.

[0093] Step 3: Design a local scheduling algorithm based on rule combination and deep reinforcement learning to generate a local scheduling solution for process task reassignment and process task reordering within a local limited equipment range.

[0094] The local scheduling algorithm specifically includes the following sub-steps:

[0095] Step 3.1: Construct a Markov decision process (MDP) model corresponding to the local scheduling process. By analyzing the constraints and optimization objectives faced by the local scheduling mathematical model, design the workshop status and reward function in the Markov decision model. By establishing a scheduling rule library for assigning process tasks to equipment and sorting process tasks on equipment, design scheduling actions with rule combinations.

[0096] Step 3.2: Using the improved deep Q network (D3QN) algorithm, by establishing a deep learning network, designing a loss function for reinforcement learning, and embedding an experience replay mechanism, the weight coefficients in the rule combination process are optimized, a reasonable combination of scheduling rules is obtained, and a local scheduling solution that minimizes the weighted completion time is obtained. The specific process is as follows: Figure 3 .

[0097] (1) Markov Decision Process (MDP) Model

[0098] The Markov decision process (MDP) model describes how the machine learning process selects a reasonable decision action A according to the environmental state S to achieve the decision process of maximizing the reward function R. The present invention focuses on the local scheduling requirements. First, it is necessary to establish the environmental state considering various constraints according to the model of the scheduling problem, and establish the reward function from the perspective of the optimization goal, and then consider the decision variable X mij and Y kij To simplify the action space, a scheduling rule base is established to form decision actions based on rule combination operations. According to the local scheduling constraints described by equations (2) to (15) in step 2, the scheduling state S = {S 1 ,S 2 ,…,S 7}As shown in Table 2.

[0099] Table 2 State space of local scheduling

[0100] Status Characteristics describe <![CDATA[S 1 ]]> The subsequent available time period of the device <![CDATA[S 2 ]]> The subsequent unavailability period of the device <![CDATA[S 3 ]]> Subsequent task information, including workpieces, processes, and working hours <![CDATA[S 4 ]]> Equipment utilization standard deviation <![CDATA[S 5 ]]> Average equipment utilization <![CDATA[S 6 ]]> The start and end time of the process tasks assigned to the equipment <![CDATA[S 7 ]]> The start and end time of the process tasks executed on the equipment

[0101] According to the optimization objective of the local scheduling model in step 2, the reward function shown in formula (22) is designed in the Markov decision process (MDP) model:

[0102]

[0103] Where τ represents a constant, and R represents the reward value obtained after completing a scheduling action. Considering that the local scheduling problem includes two scheduling contents: the assignment of process tasks to equipment and the sorting of process tasks on equipment, a scheduling rule base is constructed as shown in Table 3.

[0104] Table 3 Local scheduling rule base

[0105]

[0106] The scheduling action obtains the decision X by combining the rules shown in formula (23) in the scheduling rule base. mij and Y kij The composite rule of assignment calculates the corresponding reward return function.

[0107]

[0108] where ε i Represents the scheduling rule r in the rule combination process i The weight coefficient of .

[0109] (2) Improved Deep Q Network (D3QN) algorithm

[0110] According to the established Markov decision process (MDP) model, the improved deep Q network (D3QN) algorithm integrates the advantages of the deep Q-networks (DQN) algorithm, the double deep Q-network (DoubleDQN) algorithm and the Dueling DQN algorithm, evaluates the reward return of scheduling actions according to the state of the scheduling environment, and makes reasonable decisions on scheduling actions based on the reward return evaluation results.

[0111] Firstly, according to the environmental state and scheduling action in the Markov decision process (MDP) model, an improved deep Q network (D3QN) algorithm network structure is designed.

[0112] The number of nodes in the input layer is the same as the dimension of the environmental state, and the number of nodes in the output layer is the same as the weight coefficient corresponding to the scheduling action. A fully connected hidden layer is constructed between the input layer and the output layer, and feature extraction is used to achieve adaptive adjustment of the weight coefficient under different environmental states.

[0113] Then, the main network θ and the target network θ′ are constructed respectively, and the action a is scheduled according to the tth step. t Environmental conditions faced t , calculate the corresponding Q values ​​according to formula (22):

[0114] Q(s t ,a t ,θ)=R t (s t ,a t )-R t-1 (s t-1 ,a t-1 ) (twenty four)

[0115] Q(s t ,a t ,θ′)=R t (s t ,a t )-R t-c (s t-c ,a t-c ) (25)

[0116] Where c represents the update interval step size of the target network and γ represents the learning rate.

[0117] Finally, according to the target network trained by the improved deep Q network (D3QN) algorithm, the scheduling rules are dynamically combined under different environmental conditions, and scheduling decisions such as assigning processes to equipment and sorting processes on equipment are made according to the composite scheduling rules to obtain reasonable local scheduling plans and achieve the optimization goal of minimizing the weighted completion time.

[0118] Example 2

[0119] Aiming at the local disturbance of equipment failure in the production process of aerospace manufacturing workshop, the present invention constructs a local scheduling model with the optimization goal of minimizing the weighted completion time of workpieces taking into account the equipment availability constraint, designs a local scheduling algorithm combining rule combination and deep reinforcement learning, and proposes a dynamic scheduling method for aerospace manufacturing workshop facing local disturbance, which includes the following steps:

[0120] Step 1: Quantitatively analyze the equipment availability during the production process and calculate the equipment's average failure time, average repair time and other availability values;

[0121] This embodiment establishes a reliability description model based on the procurement data of key functional components, and then calculates the MTBF index of the equipment. The calculation formula is as follows:

[0122]

[0123] In reliability description, the two-parameter Weibull distribution containing scale parameters and shape parameters is used to describe the failure probability changes of mechanical equipment and its components, which can achieve better fitting accuracy. The three-parameter Weibull distribution adds a location parameter on this basis, further improving the fitting ability of different failure modes. Construct the equipment failure probability density function (PDF) as shown in the following formula:

[0124]

[0125] Among them, γ is the position parameter, which means that failure will not occur within the initial γ time; η is the scale parameter, which indicates the scaling degree of the curve; β is the shape parameter, which indicates the trend of the curve; is the continuous use time for production. The failure probability of mechanical manufacturing equipment will accumulate over time. By integrating equation (27), we can obtain the cumulative distribution function (CDF) shown in the following equation:

[0126]

[0127] This embodiment fits reasonable three parameters of γ, η, and β according to the failure sample data of the functional components to construct the cumulative distribution function of its failure probability. Based on the Weibull distribution parameter optimization method of the genetic algorithm, the cumulative distribution function of the equipment failure fitting failure samples is generated by designing parameter encoding, population initialization, fitness evaluation and operation operators:

[0128] (1) Parameter encoding

[0129] The three parameters γ, η, and β are represented as x 1 ,x 2 ,x3 , forming a chromosome including three gene locations [x 1 ,x 2 ,x 3 ], each position uses a continuous variable to represent the corresponding parameter value.

[0130] (2) Population initialization

[0131] The maximum likelihood estimation method is used to calculate the initial estimates of each parameter γ 0 ,η 0 ,β 0 , given the domain coefficient h, [1-h,1+h]·γ 0 ,[1-h,1+h]·η 0 ,[1-h,1+h]·β 0 As the search range of each parameter, Popsize chromosomes are randomly generated within the search range as the initial population.

[0132] (3) Fitness evaluation

[0133] For the specific nth individual in Popsize chromosomes, by sorting the sample data from small to large in time {t 1 ,t 2 ,…t sn ,...} are arranged, and according to the D test method, the difference between the theoretical distribution function of formula (29) and the empirical distribution function actually presented by the sample is calculated:

[0134] D n = sup|F(t i )-E(t i )|=max{d i} <D (29)

[0135] Where F(t i ) indicates that the three parameters are taken as γ n ,η n ,β n Time t i The failure probability at time, E(t i ) indicates that according to the sample distribution, the t i The failure probability at time d i According to t i The absolute difference of the failure probability is obtained by calculating D n The difference between the critical value D and the fitness function of the nth individual is obtained:

[0136] Fitness n =DD n (30)

[0137] (4) Selection, crossover, and mutation

[0138] According to the fitness distribution of individuals in the population calculated by formula (30), the roulette method is used to select individuals with higher fitness to enter the next generation population. First, the crossover gene position is determined, and then the two parameter values ​​of the two chromosomes at the gene position are exchanged. Each individual randomly selects a position and randomly updates the parameter value within the given search range.

[0139] Through the above genetic algorithm process, the γ, η, β parameter values ​​of the best fitting failure samples are obtained, and the cumulative distribution function of the failure probability shown in formula (28) is calculated, and then the reliability function shown in formula (31) is obtained:

[0140]

[0141] Where R(t) represents the reliability of the functional component when the continuous production time is t. For a manufacturing device containing C functional components, the reliability function R can be calculated based on the failure samples and genetic algorithm. 1 (t),R 2 (t),…,R c (t)…,R C (t), and thus the equipment reliability change trend is obtained:

[0142]

[0143] According to the changing trend of equipment reliability, the expected average working time between two equipment failures is calculated:

[0144]

[0145] Based on the MTBF value and the MTTR statistical information in the equipment maintenance record table, the quantitative analysis results of equipment availability are achieved.

[0146] Step 2: Establish a local scheduling model that takes into account equipment availability constraints and construct an optimization objective that minimizes the total weighted completion time of workpieces;

[0147] The parameters and decision variables in the local scheduling model of this embodiment are defined as shown in the following table

[0148] Table 4 Parameters and decision variables of the local scheduling process

[0149]

[0150]

[0151] According to the above parameter definitions, the local scheduling problem in this embodiment can be described as the following model:

[0152] Objective function:

[0153]

[0154] Constraints:

[0155]

[0156] Formula (34) represents the optimization goal of minimizing the total weighted completion time of the workpiece; Formula (35) indicates that each production process can only be produced on one device; Formula (36) describes that the process cannot be interrupted when it is produced on the device; Formula (37) ensures that one device can only produce one process of one workpiece at a time, and cannot process multiple processes of multiple workpieces at the same time; Formula (38) describes the process sequence constraint caused by the workpiece process route; Formula (39) describes that the completion time of the workpiece is determined by the completion time of the last process task; Formula (40) describes that the total number of processes to be scheduled on the equipment is determined by the number of processes to be scheduled for all workpieces; Formula (41) and ( 42) ensures that each process of each workpiece will be given a sorting position; Formula (43) ensures that at the time of production scheduling, the start time of the first process on each device must be greater than the initial remaining occupancy time, that is, it is necessary to wait for the device to complete the existing production task; Formula (44) describes that the processes with a lower sorting on the same device need to receive production operations after the processes with a higher sorting; Formula (45) describes the average equipment failure interval obtained by equipment availability analysis; Formula (46) describes the equipment fault repair status obtained by equipment availability analysis; Formulas (47) and (48) indicate that when the equipment is in a fault repair state, the process production task cannot be performed.

[0157] Step 3: Design a local scheduling algorithm based on rule combination and deep reinforcement learning to generate a local scheduling solution for process task reassignment and process task reordering within a local limited equipment range.

[0158] Please see Figure 2 The local scheduling algorithm adopted in this embodiment specifically includes the following sub-steps:

[0159] Step 3.1: Construct a Markov decision process (MDP) model corresponding to the local scheduling process. By analyzing the constraints and optimization objectives faced by the local scheduling mathematical model, design the workshop status and reward function in the Markov decision process (MDP) model. By establishing a scheduling rule library for assigning process tasks to equipment and sorting process tasks on equipment, design scheduling actions with rule combinations.

[0160] The Markov decision process (MDP) model describes how the machine learning process selects a reasonable decision action A according to the environmental state S to achieve the decision process of maximizing the reward function R. This embodiment focuses on the local scheduling requirements. First, it is necessary to establish the environmental state considering various constraints based on the mathematical model of the scheduling problem, and establish the reward function from the perspective of the optimization goal, and then consider the decision variable X mij and Y kij To simplify the action space, a scheduling rule base is established to form decision actions based on rule combination operations. According to the local scheduling constraints described by equations (35) to (48) in step 2, the constructed scheduling state S = {S 1 ,S 2 ,…,S 7}As shown in Table 2.

[0161] Table 5 State space of local scheduling

[0162] Status Characteristics describe <![CDATA[S 1 ]]> The subsequent available time period of the device <![CDATA[S 2 ]]> The subsequent unavailability period of the device <![CDATA[S 3 ]]> Subsequent task information, including workpieces, processes, and working hours <![CDATA[S 4 ]]> Equipment utilization standard deviation <![CDATA[S 5 ]]> Average equipment utilization <![CDATA[S 6 ]]> The start and end time of the process tasks assigned to the equipment <![CDATA[S 7 ]]> The start and end time of the process tasks executed on the equipment

[0163] According to the optimization objective of the local scheduling model in step 2, the reward function shown in equation (16) is designed in the Markov decision process (MDP) model:

[0164]

[0165] Where τ represents a constant, and R represents the reward value obtained after completing a scheduling action. Considering that the local scheduling problem includes two scheduling contents: the assignment of process tasks to equipment and the sorting of process tasks on equipment, a scheduling rule base is constructed as shown in Table 6.

[0166] Table 6 Local scheduling rule base

[0167]

[0168] The scheduling action obtains the decision X by combining the rules shown in formula (50) in the scheduling rule base. mij and Y kij The composite rule of assignment calculates the corresponding reward return function.

[0169]

[0170] where ε i Represents the scheduling rule r in the rule combination process i The weight coefficient of .

[0171] Step 3.2: Using the improved deep Q network (D3QN) algorithm, by establishing a deep learning network, designing the loss function of reinforcement learning, and embedding the experience replay mechanism, the weight coefficient optimization in the rule combination process is realized, a reasonable combination of scheduling rules is obtained, and a local scheduling solution that minimizes the weighted completion time is obtained.

[0172] This embodiment first designs an improved deep Q network (D3QN) algorithm network structure based on the environmental state and scheduling action in the Markov decision process (MDP) model, in which the number of input layer nodes is the same as the environmental state dimension, the number of output layer nodes is the same as the weight coefficient corresponding to the scheduling action, and a fully connected hidden layer is constructed between the input layer and the output layer. The weight coefficient is adaptively adjusted under different environmental states through feature extraction. Then, the main network θ and the target network θ′ are constructed respectively, and according to the scheduling action a in the tth step, the weight coefficient is adjusted according to the weight coefficient of the scheduling action. t Environmental conditions faced t , calculate the corresponding Q values ​​according to formula (49):

[0173] Q(s t ,a t ,θ)=R t (s t ,a t )-R t-1 (s t-1 ,a t-1 ) (51)

[0174] Q(s t ,a t ,θ′)=R t (s t ,a t )-R t-c (s t-c ,a t-c ) (52)

[0175] Where c represents the update interval step size of the target network and γ represents the learning rate.

[0176] Finally, this embodiment uses the target network obtained by training the improved deep Q network (D3QN) algorithm to dynamically combine scheduling rules under different environmental conditions, and makes scheduling decisions such as assigning processes to equipment and sorting processes on equipment based on composite scheduling rules, so as to obtain reasonable local scheduling solutions and achieve the optimization goal of minimizing weighted completion time.

[0177] The present invention conducts comparative experiments on the local scheduling method based on rule combination and improved deep Q network (D3QN) algorithm with other common combination rule methods, deep Q networks (DQN), and genetic algorithms.

[0178] This embodiment uses the production case data of the MES system and ERP system of a certain aerospace product manufacturing enterprise in Shanghai, and uses the workshop simulation software to simulate and model the workshop. By simulating different production situations in the simulation environment, the diversity of samples is increased to ensure that the samples obtained through simulation can cover various local disturbances of the material resource layer in the production environment. Finally, the 500 sample data obtained by the case simulation are used as the data source to train the improved deep Q network (D3QN) algorithm network model. The specific example scale is shown in Table 7.

[0179] Table 7 Example scale

[0180] parameter Example Number of product types 10 Work plan sheet (piece) 18 Number of workpieces 25 Number of processes (passes) [4,15] Working hours (hours) [0.25,36]

[0181] This embodiment takes minimizing the weighted completion time of workpieces as the evaluation indicator, and designs five different scheduling environments of 4×5×5, 8×5×5, 12×5×5, 16×5×5, and 20×5×5. Among them, 12×5×5 means that there are 12 workpieces, the number of optional equipment for each process is at most 5, and the number of processes to be processed for each workpiece is at most 5.

[0182] In terms of reinforcement learning algorithms, common parameter values ​​are shown in Table 8.

[0183] Table 8 Common parameter values

[0184]

[0185]

[0186] This embodiment uses orthogonal experiments. First, other important parameters of the algorithm are assumed. Then, the parameters of the network structure, i.e., the number of hidden layers and the number of nodes in the hidden layers, are determined by performing parameter experiments. Finally, based on the obtained network structure parameters, the learning rate and discount factor are further selected. Finally, this embodiment sets the network hidden layers of the improved deep Q network (D3QN) algorithm to [64, 64, 64, 64, 64], the reward discount factor to 0.95, and the learning rate to 1e-4.

[0187] In the comparison method of this embodiment, in terms of the deep Q network (DQN) algorithm based on the combined scheduling rule, the parameters are set as 1000 iterations, 1000 experience pool capacity, 0.5 greedy probability, and 0.9 discount factor; in terms of the genetic algorithm, an encoding strategy based on workpieces and processes is adopted to generate the initial population, and a new generation is generated through roulette wheel selection and partial matching crossover operations, while exchange mutation is introduced to increase genetic diversity, and the fitness function is aimed at minimizing the weighted completion time of the workpiece, and the termination condition sets the number of iterations to 1000; in terms of the combined rules, according to the local scheduling rule base in Table 6, a composite rule as shown in Table 9 is established.

[0188] Table 9 Combination rule settings

[0189] combination rule Combination 1 <![CDATA[r 1 &r 4 &r 5 ]]> Combination 2 <![CDATA[r 1 &r 5 &r 6 ]]> Combination 3 <![CDATA[r 2 &r 6 &r 7 ]]> Combination 4 <![CDATA[r 2 &r 7 &r 8 ]]> Combination 5 <![CDATA[r 3 &r 4 &r 8 ]]> Combination 6 <![CDATA[r 3 &r 4 &r 5 ]]>

[0190] In order to evaluate the scheduling optimization performance of different algorithms and rules, this embodiment designs a comparative experiment of different examples.

[0191] The weighted completion time of the workpiece is the optimization target. The experimental results are shown in Table 10, and the optimal scheduling result of each example is bold.

[0192] Table 10 Test results of different scale examples

[0193]

[0194]

[0195] Please see Figure 4 , and visualize some test case results by plotting them into bar graphs.

[0196] The test results of this embodiment show that for smaller-scale problem instances, each algorithm can explore the solution space more effectively and obtain a near-optimal solution. As the scale of the example increases, the performance advantage of the improved deep Q network (D3QN) algorithm becomes more significant. The local scheduling method based on rule combination and the improved deep Q network (D3QN) algorithm adaptively adjusts the weighting coefficients in the rule combination process on the basis of characteristic analysis of the environmental state of the scheduling problem. This mechanism significantly improves the algorithm's search capability and environmental adaptability.

[0197] In contrast, the deep Q network (DQN) algorithm lacks an effective rule combination strategy, which limits its performance gain when facing larger-scale problems. As a traditional optimization algorithm, the genetic algorithm has certain capabilities in solution space search, but in this embodiment, its performance is not as good as the method based on deep reinforcement learning. The performance of rule combination shows a certain degree of decline when processing larger-scale examples, which shows the limitations of composite rules in more complex problems.

[0198] Those skilled in the art know that, in addition to implementing the system, device and its various modules provided by the present invention in a purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Therefore, the system, device and its various modules provided by the present invention can be considered as a hardware component, and the modules included therein for implementing various programs can also be considered as structures within the hardware component; the modules for implementing various functions can also be considered as both software programs for implementing the method and structures within the hardware component.

[0199] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A workshop dynamic scheduling method for local disturbance, characterized in that: include: Step 1: Quantitatively analyze the availability of key equipment in the production process and calculate the equipment availability indicators, including mean time between failures and mean time to repair; Step 2: Establish a local scheduling model that takes equipment availability into consideration, determine constraints, and construct an optimization objective that minimizes the total weighted completion time of workpieces; Step 3: Design a local scheduling algorithm based on rule combination and deep reinforcement learning to generate a local scheduling solution for process task reassignment and process task reordering within a local limited equipment range.

2. The method for dynamic workshop scheduling facing local disturbance according to claim 1, characterized in that: The step 1 comprises: Use the procurement data of key functional components to establish a reliability description model and calculate the MTBF index of the equipment; In the reliability description model, a three-parameter Weibull distribution including a scale parameter, a shape parameter, and a location parameter is used to describe the failure probability change of mechanical equipment and its components, and an equipment failure probability density function is constructed, as shown in the following formula: Among them, γ is the position parameter, which means that failure will not occur within the initial γ time; η is the scale parameter, which indicates the scaling degree of the curve; β is the shape parameter, which indicates the trend of the curve; t is the continuous use time for production; The equipment failure probability density function is integrated to obtain a cumulative distribution function, as shown in the following formula: Using the failure sample data of functional components, the genetic algorithm is used to obtain the γ, η, and β parameter values ​​of the best fitting failure samples, and the cumulative distribution function of the failure probability is calculated to obtain the reliability function, as shown in the following formula: Wherein, R(t) represents the reliability of the functional component when the continuous production time is t. For a manufacturing device containing C functional components, the reliability functions R1(t), R2(t),…, R c (t)…,R C (t), and thus the equipment reliability change trend is obtained: According to the equipment reliability change trend, the expected average working time between two equipment failures is calculated: According to the MTBF value and the MTTR statistical information in the equipment maintenance record table, a quantitative analysis result of the equipment availability is obtained.

3. The method for dynamic workshop scheduling facing local disturbance according to claim 1, characterized in that: The optimization goal of minimizing the total weighted completion time of the artifacts in step 2 is expressed as: Among them, i represents the workpiece number, I represents the total number of workpieces, ω i represents the weighting coefficient of workpiece i, F i represents the completion time of job i.

4. The method for dynamic workshop scheduling facing local disturbance according to claim 1, characterized in that: The step 3 comprises: Construct a Markov decision process model corresponding to the local scheduling process. By analyzing the constraints and optimization goals faced by the local scheduling mathematical model, design the workshop state and reward function in the Markov decision model. By establishing a scheduling rule library for assigning process tasks to equipment and sorting process tasks on equipment, design scheduling actions with rule combinations. By using the improved deep Q-network algorithm, through establishing a deep learning network, designing the loss function of reinforcement learning, and embedding the experience replay mechanism, the weight coefficient optimization in the rule combination process is realized, the reasonable combination of scheduling rules is obtained, and the local scheduling plan that minimizes the weighted completion time is obtained.

5. A workshop dynamic scheduling system for local disturbances, characterized in that: include: Module M1: Quantitatively analyze the availability of key equipment in the production process and calculate the availability indicators of the equipment, including mean time between failures and mean time to repair; Module M2: Establish a local scheduling model that takes equipment availability into account, determine constraints, and construct an optimization objective that minimizes the total weighted completion time of workpieces; Module M3: Design a local scheduling algorithm based on rule combination and deep reinforcement learning to generate a local scheduling solution for process task reassignment and process task reordering within a local limited equipment range.

6. The local disturbance-oriented workshop dynamic scheduling system according to claim 5, characterized in that: The module M1 comprises: Use the procurement data of key functional components to establish a reliability description model and calculate the MTBF index of the equipment; In the reliability description model, a three-parameter Weibull distribution including a scale parameter, a shape parameter, and a location parameter is used to describe the failure probability change of mechanical equipment and its components, and an equipment failure probability density function is constructed, as shown in the following formula: Among them, γ is the position parameter, which means that failure will not occur within the initial γ time; η is the scale parameter, which indicates the scaling degree of the curve; β is the shape parameter, which indicates the trend of the curve; t is the continuous use time for production; The equipment failure probability density function is integrated to obtain a cumulative distribution function, as shown in the following formula: Using the failure sample data of functional components, the genetic algorithm is used to obtain the γ, η, and β parameter values ​​of the best fitting failure samples, and the cumulative distribution function of the failure probability is calculated to obtain the reliability function, as shown in the following formula: Wherein, R(t) represents the reliability of the functional component when the continuous production time is t. For a manufacturing device containing C functional components, the reliability functions R1(t), R2(t),…, R c (t)…,R C (t), and thus the equipment reliability change trend is obtained: According to the equipment reliability change trend, the expected average working time between two equipment failures is calculated: According to the MTBF value and the MTTR statistical information in the equipment maintenance record table, a quantitative analysis result of the equipment availability is obtained.

7. The workshop dynamic scheduling system for local disturbance according to claim 5, characterized in that: The optimization goal of minimizing the total weighted completion time of the artifacts built in the module M2 is expressed as: Among them, i represents the workpiece number, I represents the total number of workpieces, ω i represents the weighting coefficient of workpiece i, F i represents the completion time of job i.

8. The local disturbance-oriented workshop dynamic scheduling system according to claim 5, characterized in that: The module M3 comprises: Construct a Markov decision process model corresponding to the local scheduling process. By analyzing the constraints and optimization goals faced by the local scheduling mathematical model, design the workshop state and reward function in the Markov decision model. By establishing a scheduling rule library for assigning process tasks to equipment and sorting process tasks on equipment, design scheduling actions with rule combinations. By using the improved deep Q-network algorithm, through establishing a deep learning network, designing the loss function of reinforcement learning, and embedding the experience replay mechanism, the weight coefficient optimization in the rule combination process is realized, the reasonable combination of scheduling rules is obtained, and the local scheduling plan that minimizes the weighted completion time is obtained.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the local disturbance oriented workshop dynamic scheduling method according to any one of claims 1 to 4 are implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by a processor, the steps of the local disturbance oriented workshop dynamic scheduling method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Dynamic scheduling method for installation, adjustment and production of multi-model and small-batch complex equipment products

    CN114186791A