Semiconductor manufacturing scheduling method and system and computer readable storage medium

By combining intelligent algorithms and reinforcement learning algorithms, the semiconductor manufacturing scheduling method solves the problem of low production efficiency caused by high computational load in existing technologies, achieves rapid optimization and rational allocation of resources, and reduces production costs.

CN120996393APending Publication Date: 2025-11-21EX IND TECH (SUZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410623020.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing semiconductor manufacturing scheduling methods involve large computational loads when using genetic algorithms for querying, resulting in excessively long solution times. This makes it impossible to effectively and rationally allocate production capacity, improve resource utilization, and reduce production costs.

Method used

By combining the global search capability of intelligent algorithms with the local optimization techniques of reinforcement learning algorithms, a scheduling sequence is generated through a mathematical optimization model, including hill climbing and genetic algorithms. The appropriate mathematical optimization model is selected according to the data scale, and iterative calculations and training actions are performed to determine the optimal strategy for simulated production scheduling.

Benefits of technology

It has enabled rapid optimization of semiconductor manufacturing scheduling, effectively and rationally allocated production capacity, improved resource utilization, and reduced production costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996393A_ABST
    Figure CN120996393A_ABST
Patent Text Reader

Abstract

The invention discloses a semiconductor manufacturing scheduling method and system and a computer readable storage medium. The method comprises the following steps: determining a mathematical optimization model according to a data scale; generating a first scheduling sequence according to the mathematical optimization model, and performing iterative calculation on the first scheduling sequence through the mathematical optimization model under a constraint condition to obtain a second scheduling sequence; determining an optimal strategy based on an operation result of each training action, determining a target total strategy according to the optimal strategy corresponding to each state parameter, and performing simulation scheduling according to the target total strategy to obtain a third scheduling sequence; and performing scheduling operation according to the third scheduling sequence. According to the method, rapid target optimization of semiconductor manufacturing scheduling is realized, the production capacity is effectively and reasonably distributed, the resource utilization rate is improved, and the production cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of semiconductor manufacturing, in particular to a semiconductor manufacturing scheduling method, system and computer readable storage medium. BACKGROUND

[0002] Semiconductor production manufacturing belongs to one of the complex industrial manufacturing chains, which usually includes four production stages: wafer manufacturing, wafer sorting, packaging and product testing. Among them, wafer manufacturing is a complex work, and the manufacturing of machines and silicon wafers requires hundreds of process steps. A small part of the products exists in the hundreds of processes in the factory, so it is necessary to create a complex capacity allocation and job scheduling problem, that is, the scheduling problem is a problem that is closely concerned in semiconductor manufacturing production. The scheduling problem generally refers to the flow processing of n workpieces on m machines, the time spent by each workpiece on each machine is different, and each machine can only process one workpiece at the same time. The goal of scheduling is to determine the processing order of workpieces on each machine, the starting time of each process, so that the maximum completion time is minimized or other indicators are optimized.

[0003] At present, one of the semiconductor manufacturing scheduling methods is to use an intelligent optimization algorithm such as a genetic algorithm constructed according to the evolution law to query, but this method needs to construct a sequence population and perform optimization iteration in the population, and needs to calculate each individual in the population separately. When the population is larger, the amount of calculation is also larger, thereby causing a large amount of time to be consumed for solving once. SUMMARY

[0004] The purpose of the present application is to provide a semiconductor manufacturing scheduling method, system and computer readable storage medium, which can realize effective and reasonable allocation of production capacity, improve resource utilization, and reduce production cost.

[0005] In order to achieve the above purpose, the present application provides a semiconductor manufacturing scheduling method, which comprises the following steps:

[0006] Obtaining the data scale of semiconductor manufacturing scheduling, determining a mathematical optimization model according to the data scale;

[0007] Generating a first scheduling sequence according to the mathematical optimization model, and iteratively calculating the first scheduling sequence under the constraint condition to obtain a second scheduling sequence, the second scheduling sequence being the scheduling sequence with the maximum adaptive value under the constraint condition, and the constraint condition including the maximum number of iterations;

[0008] Determining all state parameters of the semiconductor manufacturing environment and the training action of the second scheduling sequence, traversing each state parameter, and determining all traversal initial strategies corresponding to the traversed state parameters based on each initial strategy.

[0009] running each of the training actions in the initial strategy, determining an optimal strategy based on a running result of each of the training actions, determining a target total strategy according to the optimal strategy corresponding to each of the state parameters, and performing simulation scheduling according to the target total strategy to obtain a third scheduling sequence;

[0010] performing scheduling operation according to the third scheduling sequence.

[0011] Further, in the semiconductor manufacturing scheduling method, the data scale includes a number of products to be processed.

[0012] Further, in the semiconductor manufacturing scheduling method, the step of iteratively calculating the first scheduling sequence under the constraint condition to obtain the second scheduling sequence by using the mathematical optimization model specifically includes:

[0013] When the data scale does not exceed the preset threshold, two random integers are generated, the two random integers are different integers between 1 and N, and N is the length of the first scheduling sequence;

[0014] The values of the sequence number positions corresponding to the two random integers in the first scheduling sequence are exchanged to obtain a first comparison sequence;

[0015] The fitness values of the first scheduling sequence and the first comparison sequence are calculated and compared. When the fitness value of the first comparison sequence is greater than the fitness value of the first scheduling sequence, the value exchange step is repeated for the first comparison sequence until the maximum number of iterations in the constraint condition is met, and a second scheduling sequence is obtained. The second scheduling sequence is the scheduling sequence with the maximum fitness value under the constraint condition.

[0016] Further, in the semiconductor manufacturing scheduling method, the step of iteratively calculating the first scheduling sequence under the constraint condition to obtain the second scheduling sequence by using the mathematical optimization model specifically includes:

[0017] When the data scale exceeds the preset threshold, a plurality of preferred sequences are obtained by performing roulette screening, crossover operation and mutation operation on the first scheduling sequence under the constraint condition, and the constraint condition includes the maximum number of iterations, the number of individuals in the population, the crossover probability and the mutation probability.

[0018] The second scheduling sequence is determined according to the plurality of preferred sequences, and the second scheduling sequence is the preferred sequence with the maximum fitness value in the plurality of preferred sequences.

[0019] Further, in the semiconductor manufacturing scheduling method, the step of performing roulette selection on the first scheduling sequence under the constraint condition comprises:

[0020] The fitness values of the plurality of random sequences in the first scheduling sequence are calculated respectively, a normalized interval is constructed according to the proportion of the fitness value of each random sequence in the sum of the fitness values of all random sequences, the normalized interval comprises a plurality of sub-intervals corresponding to each fitness value, and the normalized interval has a value range of [0, 1];

[0021] A plurality of random numbers between 0 and 1 are generated, and the number of the random numbers is the same as the number of the plurality of random sequences;

[0022] The sub-interval into which each random number falls in the normalized interval is determined, and the corresponding random sequence is selected according to the fitness value of the sub-interval into which the random number falls, to obtain the plurality of random sequences after selection.

[0023] Further, in the semiconductor manufacturing scheduling method, the step of performing crossover operation on the first scheduling sequence under the constraint condition comprises:

[0024] The random sequences to be crossed are determined according to the crossover probability and the plurality of random sequences after selection, and the random sequences to be crossed are paired to obtain a plurality of sequence groups, each sequence group comprising two random sequences;

[0025] A random integer between 1 and N is generated, N is the length of the first scheduling sequence, the values of the two random sequences in the sequence group are exchanged in a forward or reverse direction from the position of the random integer, to obtain the plurality of random sequences after crossover.

[0026] Further, in the semiconductor manufacturing scheduling method, the step of performing mutation operation on the first scheduling sequence under the constraint condition comprises:

[0027] The random sequences to be mutated are determined according to the mutation probability and the plurality of random sequences after crossover;

[0028] Two random integers between 1 and N are generated, N is the length of the first scheduling sequence, and the sub-sequence between the positions corresponding to the two random integers in the random sequence to be mutated is randomly sorted to obtain the plurality of random sequences after mutation;

[0029] A preferred sequence is calculated according to the plurality of random sequences after mutation, and the preferred sequence is the sequence with the largest fitness value in the plurality of random sequences after mutation.

[0030] Further, in the semiconductor manufacturing scheduling method, the training action includes due date sequence, customer level, product arrival time from buffer, and product initial sequence.

[0031] In addition, the application further provides a semiconductor manufacturing scheduling system for implementing the semiconductor manufacturing scheduling method.

[0032] A first determining unit is configured to determine a mathematical optimization model according to the data scale of the semiconductor manufacturing scheduling;

[0033] A first calculating unit is configured to generate a first scheduling sequence according to the mathematical optimization model, and to obtain a second scheduling sequence by iteratively calculating the first scheduling sequence under a constraint condition through the mathematical optimization model, wherein the second scheduling sequence is a scheduling sequence with the largest adaptive value under the constraint condition, and the constraint condition includes a maximum iteration number.

[0034] A second determining unit is configured to determine a target attribute priority of a product to be processed.

[0035] A second calculating unit is configured to determine all state parameters of a semiconductor manufacturing environment, and training actions of the second scheduling sequence, to traverse each state parameter, to determine all traversed initial strategies corresponding to the traversed state parameter based on each initial strategy, to run the training actions in each traversed initial strategy, to determine an optimal strategy based on a running result of each training action, to determine a target total strategy according to the optimal strategy corresponding to each state parameter, and to perform simulation scheduling according to the target total strategy to obtain a third scheduling sequence.

[0036] A scheduling unit is configured to perform scheduling operation according to the third scheduling sequence.

[0037] In addition, the application further provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the semiconductor manufacturing scheduling method.

[0038] The application combines the global search ability of the intelligent algorithm and the local optimization skill of the reinforcement learning algorithm, realizes the rapid target optimization of the semiconductor manufacturing scheduling, effectively and reasonably allocates the production capacity, improves the resource utilization rate, and reduces the production cost. BRIEF DESCRIPTION OF DRAWINGS

[0039] Figure 1 is a flowchart of the semiconductor manufacturing scheduling method of the embodiment of the application;

[0040] Figure 2 is a schematic diagram of adaptive value calculation and estimation of the scheduling sequence;

[0041] Figure 3is a schematic diagram of the first mathematical optimization model in the present application;

[0042] Figure 4 is a schematic diagram of the second mathematical optimization model in the present application;

[0043] Figure 5 is a scene schematic diagram of the semiconductor manufacturing scheduling method in the present application;

[0044] Figure 6 is a principle schematic diagram of the reinforcement learning training in the present application;

[0045] Figure 7 is a schematic diagram of the strategy Q value table in the reinforcement learning training in the present application;

[0046] Figure 8 is a flow schematic diagram of the deep regression model training in the reinforcement learning training in the present application;

[0047] Figure 9 is a structure schematic diagram of the semiconductor manufacturing scheduling system of the semiconductor manufacturing scheduling method in the embodiment of the present application;

[0048] Figure 10 is Figure 9 a structure schematic diagram of the first calculation unit in the present application. DETAILED DESCRIPTION

[0049] The present application will be described in detail below with reference to specific embodiments and drawings.

[0050] The semiconductor manufacturing scheduling method provided by the embodiment of the present application comprises the following steps:

[0051] Obtain the data scale of semiconductor manufacturing scheduling, determine a mathematical optimization model according to the data scale, generate a first scheduling sequence according to the mathematical optimization model, and obtain a second scheduling sequence by iterative calculation of the mathematical optimization model under a constraint condition, wherein the second scheduling sequence is a scheduling sequence with the largest adaptive value under the constraint condition, and the constraint condition includes a maximum iteration number; determine all state parameters of a semiconductor manufacturing environment and training actions of the second scheduling sequence, traverse each state parameter, determine all traversed initial strategies corresponding to the traversed state parameter based on each initial strategy, run the training actions in each traversed initial strategy, determine an optimal strategy based on a running result of each training action, determine a target total strategy according to the optimal strategy corresponding to each state parameter, and perform simulation scheduling according to the target total strategy to obtain a third scheduling sequence; and perform scheduling operation according to the third scheduling sequence.

[0052] Referring to Figures 1 to 6 The semiconductor manufacturing scheduling method provided by the embodiment of the present application specifically includes the following steps:

[0053] Step S11: Perform simulation initialization on a scene of semiconductor manufacturing.

[0054] In a specific implementation, when scheduling is performed on a complete manufacturing process of semiconductor manufacturing, such as a complete manufacturing process from a crystal bar to a silicon wafer, a real environment can be abstracted as an entrance for publishing products, an exit for counting completed products, two buffer zones for storing products, and various manufacturing machines. The entrance is an order generation module, the exit is a product convergence, and various idle machines in different factory areas obtain suitable products from the same buffer zone. It can be understood that the components and quantities of various scenes can be inconsistent.

[0055] Step S12: Obtain a data scale of semiconductor manufacturing scheduling, and determine a mathematical optimization model according to the data scale.

[0056] Referring to Figure 2In a specific implementation, the application is achieved by two-stage optimization to reach the optimal strategy of semiconductor scheduling, wherein the first-stage optimization is to attempt to find a better sequence by calculating the size of the fitness value corresponding to each sequence of semiconductor scheduling by an intelligent algorithm, and the calculation amount of the intelligent algorithm is related to the data scale of semiconductor manufacturing scheduling, therefore, the data scale for evaluating semiconductor manufacturing scheduling is first acquired, and a mathematical optimization model is determined according to the data scale. In the embodiment, the data scale is the number of products to be processed, and of course, the data scale can also be other condition data such as total processing time of products. In the semiconductor production process, the number of products to be processed is determined in advance according to the production plan, mainly according to the short-term production plan or the medium-term production plan, for example, according to the customer demand, a batch of products is scheduled and produced according to the plan, and the number of products 500 is determined in advance before a batch of 500 products is scheduled and produced according to the plan, and the calculation amount of the fitness value F1 of the scheduling calculation is different according to the different number of products, and the applied mathematical optimization model is also different, specifically, a threshold value can be set according to experience or big data, if the data scale does not exceed the threshold value, a first mathematical optimization model is selected and applied, and if the data scale exceeds the threshold value, a second mathematical optimization model is selected and applied.

[0057] That is, the step of determining the mathematical optimization model according to the data scale specifically includes:

[0058] When the data scale does not exceed the preset threshold value, a first mathematical optimization model is selected and applied, and when the data scale exceeds the preset threshold value, a second mathematical optimization model is selected and applied.

[0059] In the embodiment, the threshold value is 500, the first mathematical optimization model is a hill climbing algorithm model, and the second mathematical optimization model is a genetic algorithm model. That is, in the embodiment, when the number of products to be processed does not exceed 500, the hill climbing algorithm model is selected and applied, and when the number of products to be processed exceeds 500, the genetic algorithm model is selected and applied. In this way, a suitable mathematical optimization model can be selected according to different data scales and calculation amounts, so that the model optimization and production scheduling are the fastest and most effective.

[0060] Step S13: generating a first scheduling sequence according to the mathematical optimization model;

[0061] In a specific implementation, when the data scale does not exceed the preset threshold value, a first mathematical optimization model is selected and applied. In the embodiment, the first mathematical optimization model is a hill climbing algorithm model, and the data scale is the number of products to be processed, a first scheduling sequence is generated according to the number of products to be processed, and the length of the first scheduling sequence is the number of products to be processed. Specifically, when the hill climbing algorithm model is selected and applied, a first scheduling sequence P0 with a length of N is generated according to the number N of products to be processed, and the first scheduling sequence P0 is a random sequence.

[0062] When the data scale exceeds the preset threshold, a second mathematical optimization model is selected to be applied. In this embodiment, the second mathematical optimization model is a genetic algorithm model, and the data scale is the number of products to be processed. A first scheduling sequence is generated according to the number of products to be processed and the number of individuals in a population. The first scheduling sequence includes a plurality of random sequence sequences. The number of random sequence sequences is equal to the number of individuals in the population, and the length of the random sequence sequence is equal to the number of products to be processed. Specifically, when the genetic algorithm model is selected to be applied, m random sequence sequences L 11 , L 12 , …, L 1m are generated according to the number N of products to be processed and the number m of individuals in the population, where i = 1.

[0063] Step S14: A second scheduling sequence is obtained by iteratively calculating the first scheduling sequence under a constraint condition by using the mathematical optimization model. The second scheduling sequence is a scheduling sequence with the largest fitness value under the constraint condition. The constraint condition includes a maximum iteration number.

[0064] In the specific implementation, please refer to Figure 3 When the data scale does not exceed the preset threshold, the first mathematical optimization model is a hill climbing algorithm model. At this time, the constraint condition is a maximum iteration number n, which is a number estimated according to experience and target accuracy. After a first scheduling sequence P0 with a length of N is generated, for example, the first scheduling sequence P0 is 2-3-1-7-6-4-8-5, two random integers between 1 and N are generated, and then the values at the positions corresponding to the two random integers in P0 are exchanged to obtain a first comparison sequence P1, for example, the two random integers are 3 and 4. At this time, the values at the positions 3 and 4 in the first scheduling sequence P0 are exchanged to obtain the first comparison sequence P1, which is 2-3-7-1-6-4-8-5. Then, the fitness value F0 of the first scheduling sequence P0 and the fitness value F1 of the first comparison sequence P1 are calculated and compared.

[0065] If F0≥F1, it indicates that the first comparison sequence P1 is not better than the first scheduling sequence P0. At this time, the above-mentioned value exchange step of the two random integers is performed again until the first comparison sequence P1 better than the first scheduling sequence P0 is found.

[0066] If F0 < F1, it indicates that the first comparison sequence P1 is better. At this time, continue to generate two random integers between 1 and N, and then swap the values at the positions corresponding to the two random integers in P1, denoted as P2. For example, generate two random integers 6 and 7. At this time, swap the values at the 6th and 7th positions in the first scheduling sequence P1 to obtain the second comparison sequence P2 as 2-3-7-1-6-8-4-5. Then, calculate and compare the fitness value F1 of the first comparison sequence P1 and the fitness value F2 of the second comparison sequence P2,... and so on. Generate two random integers between 1 and N, and then swap the values at the positions corresponding to the two random integers in P i to obtain P i+1 , until the maximum iteration number n is satisfied (i > n), stop the iteration, that is, calculate and obtain the second scheduling sequence P n , and the second scheduling sequence P n is the scheduling sequence with the maximum fitness value under the maximum iteration number n.

[0067] For the first scheduling sequence L = [l0, l1, l2,..., l N-1 , its fitness function F[L] is:

[0068]

[0069]

[0070] where P(li) is the total processing time of this product batch, D(li) is the due date of this product batch, Begin is the initial time, N is the length of the first scheduling sequence, and l0, l1, l2,..., l N-1 are the products in the first scheduling sequence.

[0071] Under the same other conditions, the larger the fitness value of a product batch sequence, the higher its scheduling efficiency and the more optimized and reasonable it is.

[0072] That is, the specific steps of step S14 include:

[0073] When the data scale does not exceed the preset threshold, generate two random integers, and the two random integers are different integers between 1 and N, and N is the length of the first scheduling sequence;

[0074] Swap the values at the positions corresponding to the two random integers in the first scheduling sequence to obtain the first comparison sequence;

[0075] The adaptation value of the first scheduling sequence is calculated and compared with the adaptation value of the first comparison sequence; when the adaptation value of the first comparison sequence is greater than the adaptation value of the first scheduling sequence, the step of numerical interchanging is repeated for the first comparison sequence until a maximum iteration number in the constraint condition is met, to obtain a second scheduling sequence, which is a scheduling sequence with the maximum adaptation value under the constraint condition.

[0076] The step S14 further comprises:

[0077] When the adaptation value of the first comparison sequence is not greater than the adaptation value of the first scheduling sequence, the step of generating the first comparison sequence is repeated until the adaptation value of the obtained first comparison sequence is greater than the adaptation value of the first scheduling sequence.

[0078] Please refer to Figure 4 When the data size does not exceed a preset threshold, the second mathematical optimization model is a genetic algorithm model, and at this time, the constraint condition is a maximum iteration number n, a number m of individuals in a population, a crossover probability pc and a mutation probability pm, which are estimated according to experience and target accuracy. When m random sequence sequences L 11 , L 12 , …, L 1m of length N are generated, the individuals in the population are selected through a roulette model, specifically: the adaptation values F 11 , F 12 , …, F 1m of the m random sequence sequences L 11 , L 12 , …, L 1m are calculated respectively, a normalized roulette is constructed with the m adaptation values F 11 , F 12 , …, F 1m , that is, the sum of the m adaptation values F 11 , F 12 , …, F 1m is calculated first Then the proportions of the m adaptation values F 11 , F 12 , …, F 1m to are calculated respectively, so that an interval of [0, 1] is constructed, and the m adaptation values F 11 , F 12 , …, F 1m are selected according to the interval of [0, 1] interval; then, m random numbers between 0 and 1 are generated, and the individuals in the m sub-intervals corresponding to the m random numbers are selected to form the screened population, for example, a random number is 0.05, falling in the sub-interval of F 11 , and F 11 is selected. 11 In this way, m random order sequences can be selected, and the m random order sequences can be the same (falling in the same sub-interval).

[0079] Then, the m random order sequences of the screened population are subjected to a crossover operation, that is, individuals to be crossed are selected according to a crossover probability pc, and are paired two by two, for example, the screened population has 100 random order sequences, and the crossover probability pc is 0.8, then 0.8*100=80 random order sequences are selected, and are paired into 40 pairs of sequence groups; for each pair of sequence groups, a random integer between 1 and N is generated, and the values of the two random order sequences in the sequence group from the random integer position are exchanged, and when the numbers have repetition, that is, conflict, the exchange is performed at the end of the conflict to form the m random order sequences of the crossed population; for example, a sequence group subjected to the crossover includes two random order sequences P1=2, 3, 1, 7, 5, 6, 4, 8; P2=3, 2, 4, 8, 7, 6, 1, 5; a random integer 3 is generated, and the values of the two random order sequences P1 and P2 from the 3rd position to the left (or to the right) are exchanged in turn, and the 3rd value 4 of P1 and the 7th value 4 of P1 are repeated and conflicted, and then the 7th value 4 of P1 and the 1st value 1 of P2 are exchanged to avoid the problem of repeated values in the same sequence.

[0080] Finally, the m random order sequences of the crossed population are subjected to a mutation operation: individuals to be mutated are selected according to a mutation probability pm, and for each selected random order sequence, two random integers between 1 and N are generated, and the sub-sequence between the two random numbers of the sequence is randomly sorted to finally form the m random order sequences L i1 , L i2 , …, L imFor example, after the crossover, there are 100 random order sequences in a certain population, and the mutation probability pc is 0.05. Then, 5 random order sequences are selected for mutation operation, and for each pair of sequences, two random integers between 1 and N are generated. The values of the subsequence between the two random integers are randomly sorted. For example, if the sequence is 2, 3, 1, 7, 5, 6, 4, 8, and the two random integers are 3 and 5, then the values of the subsequence between the 3rd and 5th positions (1, 7, 5) are randomly sorted, and the mutated sequence is obtained, for example, 2, 3, 5, 1, 7, 6, 4, 8. In this way, after the roulette selection, crossover, and mutation operation are performed on the m random order sequences L 11 , L 12 , …, L 1m , the m sequences L' 11 , L' 12 , …, L' 1m are obtained. The fitness values of the m sequences L' 11 , L' 12 , …, L' 1m are calculated, and the optimal sequence F1 with the maximum fitness value is obtained. In this way, one iteration calculation is completed through the genetic algorithm model.

[0081] Next, new m random order sequences are generated, and the roulette selection, crossover, and mutation operation are continued on the new m random order sequences to obtain another m sequences. The fitness values of the other m sequences are calculated, and the optimal sequence F2 with the maximum fitness value is obtained. In this way, multiple iteration calculations are completed through the genetic algorithm model, and the iteration is stopped when the iteration number reaches the maximum iteration number n. The fitness values F1, F2, …, Fn of the n optimal sequences are obtained, and the sequence with the maximum fitness value is selected as the optimal sequence, which is the second scheduling order sequence. The second scheduling order sequence is the scheduling order sequence with the maximum fitness value under the constraint conditions, and the constraint conditions include the maximum iteration number, the number of individuals in the population, the crossover probability, and the mutation probability. n

[0082] The step S14 specifically includes:

[0083] When the data size exceeds the preset threshold, the roulette selection, crossover operation, and mutation operation are performed on the first scheduling order sequence under the constraint conditions to obtain multiple optimal sequences. The constraint conditions include the maximum iteration number, the number of individuals in the population, the crossover probability, and the mutation probability.

[0084] The second scheduling order sequence is determined according to the multiple optimal sequences. The second scheduling order sequence is the optimal sequence with the maximum fitness value among the multiple optimal sequences.

[0085] ​The step of roulette screening the first scheduling sequence under the constraint condition specifically comprises:

[0086] The fitness values of the plurality of random sequence sequences in the first scheduling sequence are respectively calculated, a normalized interval is constructed according to the proportion of the fitness value of each random sequence sequence in the sum of the fitness values of all random sequence sequences, the normalized interval comprises a plurality of sub-intervals corresponding to each fitness value, and the normalized interval has a value range of [0, 1];

[0087] A plurality of random numbers between 0 and 1 are generated, and the number of the random numbers is the same as the number of the plurality of random sequence sequences;

[0088] The sub-interval into which each random number falls in the normalized interval is determined, and the corresponding random sequence sequence is selected according to the fitness value of the sub-interval into which the random number falls, to obtain the plurality of screened random sequence sequences.

[0089] The step of performing the crossover operation on the first scheduling sequence under the constraint condition specifically comprises:

[0090] The random sequence sequences to be crossed are determined according to the crossover probability and the plurality of screened random sequence sequences, and the random sequence sequences to be crossed are paired to obtain a plurality of sequence groups, each sequence group comprising two random sequence sequences;

[0091] A random integer between 1 and N is generated, N is the length of the first scheduling sequence, the values of the two random sequence sequences in the sequence group are exchanged in a forward or reverse direction from the position of the random integer, to obtain the plurality of random sequence sequences after the crossover.

[0092] The step of performing the mutation operation on the first scheduling sequence under the constraint condition specifically comprises:

[0093] The random sequence sequences to be mutated are determined according to the mutation probability and the plurality of random sequence sequences after the crossover;

[0094] Two random integers between 1 and N are generated, N is the length of the first scheduling sequence, and the sub-sequence between the positions corresponding to the two random integers in the random sequence sequence to be mutated is randomly sorted to obtain the plurality of random sequence sequences after the mutation;

[0095] A preferred sequence is calculated according to the plurality of random sequence sequences after the mutation, and the preferred sequence is the sequence with the largest fitness value in the plurality of random sequence sequences after the mutation.

[0096] The step of performing the roulette screening, the crossover operation and the mutation operation on the first scheduling sequence under the constraint condition to obtain the plurality of preferred sequences further comprises:

[0097] The steps of generating the first scheduling sequence and the roulette screening, crossover operation and mutation operation are repeated to obtain a plurality of preferred sequences until a maximum iteration number in the constraint condition is met.

[0098] Step S15: determining all state parameters of the semiconductor manufacturing environment and the training actions of the second scheduling sequence, traversing each state parameter, and determining all traversed initial strategies corresponding to the traversed state parameter based on each initial strategy.

[0099] In the specific implementation, please refer to Figures 5 to 8 , a scheduling system based on reinforcement learning is proposed for running speed, production scale and scheduling mode, with decision-making in any state as the basis for scheduling, to form a real-time schedulable fast scheduling scheme while ensuring feasible solutions and optimization effect. In this embodiment, a simulation scheduling model (i.e., a preset simulation scheduling model) is established in advance, and the environmental parameters of the environment where the simulation scheduling model is located are collected, including device information, product information, process flow and processing time, etc. All state parameters, all action parameters and standard returns are defined according to the collected environmental parameters.

[0100] In this embodiment, when determining the target total strategy generated in the reinforcement learning process, each state parameter needs to be traversed, and all traversed initial strategies corresponding to the traversed state parameter at the current time are determined in the pre-set initial strategies. Among all the traversed initial strategies, the traversed state parameters are included, and the training actions in each traversed initial strategy are different.

[0101] It should be noted that in this embodiment, during the scheduling process, the products need to walk in each buffer and machine, and the process flow is completed to end the scheduling. For the machine in the scheduling process, when the buffer corresponding to the machine has products, one of the products should be selected for processing. The buffer is a place for temporarily placing products, located before a machine or machines of the same type, and a product processed by a machine needs to be obtained from the buffer. In the scheduling process, an idle machine can select a product in the buffer for processing, and each machine or a plurality of machines correspond to one or more buffers, such as buffer 1, buffer 2, buffer 3, etc. The screening method of the buffer can be as shown in Table 1.

[0102]

[0103]

[0104] Table 1

[0105] The screening method of the machine type can be as shown in Table 2.

[0106]

[0107] Table 2

[0108] In this embodiment, the setting of the action parameter can be the number of machine types multiplied by the number of buffer zones. As shown in Table 1, there are 3 actions in the buffer zone screening mode and 3 actions in the machine type screening mode, so the number of action parameters can be 9, i.e. 9 action parameters.

[0109] In this embodiment, the action parameter is to formulate a one-dimensional action for the machine selection lot, and the lot of the buffer zone is sorted, specifically: 1, the delivery time sequence; 2, the customer level; 3, the arrival time of the product from the buffer zone; 4, the initial sequence of the product; and the return is based on the inverse number of the stage time (unit: seconds), and when there is a lot overdue, the previous positive number (same number x 100) is subtracted as a penalty.

[0110] Step S16: running each training action in the traversal initial strategy, determining the optimal strategy based on the running result of each training action, determining the target total strategy according to the optimal strategy corresponding to each state parameter, and performing simulation scheduling according to the target total strategy to obtain a third scheduling sequence;

[0111] In the specific implementation, please refer to Figures 5-8 After determining all the traversal initial strategies, the simulation scheduling model is caused to sequentially execute the training actions in the traversal initial strategies, and when each training action is executed, the state of the simulation scheduling model is kept consistent with the traversal state parameter. After the simulation scheduling model trains each training action, the return value obtained by the simulation scheduling model determines the training action with the best return effect in each training action, and the training action with the best return effect is taken as the optimal training action. The initial strategy corresponding to the optimal training action is taken as the optimal strategy, and the optimal strategy includes the traversal state parameter and the optimal training action corresponding to the traversal state parameter. Then it is determined whether the optimal strategy corresponding to all state parameters is obtained. If the optimal strategy corresponding to all state parameters is obtained, it is determined that the reinforcement learning training is completed, and the optimal strategy corresponding to all state parameters is taken as the target total strategy. The simulation scheduling is performed according to the target total strategy to obtain a third scheduling sequence.

[0112] In this embodiment, the training of the optimal strategy needs to establish a Q value table to save the state S and all actions A that will be taken, i.e. Q(S, A). For example, as shown in Figure 8As shown, the Q value table includes actions A1, A2,..., An; states S1, S2,..., Sn; qn1,..., qnm. As q11 = Q(S1, A1), if qn2 of state Sn is the largest, it is determined that the optimal action of state Sn is A2. At this time, the optimal strategy corresponding to state Sn includes state Sn and action A2.

[0113] In the embodiment, the optimal strategy is determined by traversing each state parameter and running all training actions in the traversed initial strategy corresponding to the traversed state parameter, and the target total strategy is determined according to the optimal strategy corresponding to each state parameter, thereby ensuring the effectiveness of the obtained target total strategy.

[0114] When it is detected that the simulation scheduling model has executed the target training action, that is, the target training action has been run, a corresponding reward is generated. In the embodiment, the simulation scheduling model generates a corresponding reward each time an action is executed. Therefore, after the reward corresponding to the target training action is obtained, it can be determined whether the scheduling learning process based on the traversed state parameter is completed according to the reward. If not, a new action parameter is selected from all action parameters as a new training action for continuous execution until it is determined that the scheduling learning process based on the traversed state parameter is completed. After the target training action is run, the state of the simulation scheduling model is converted from the non-scheduling state parameter to the traversed state parameter. If it is determined that the scheduling learning process corresponding to the traversed state parameter is completed, it is necessary to determine whether the scheduling learning processes corresponding to all state parameters are completed. If not, the scheduling learning processes corresponding to the state parameters that have not been completed need to be continuously executed, and it is determined that the reinforcement learning training is not completed. If the scheduling learning processes corresponding to all state parameters are completed, it is determined that the reinforcement learning training is completed. The way to determine whether the scheduling learning process corresponding to the traversed state parameter is completed can be to determine the optimal action parameter corresponding to the traversed state parameter, that is, to determine the rewards corresponding to each action parameter, select the reward with the best effect (that is, compare the obtained reward with the standard reward set in advance to determine the reward with the best effect), and select the reward with the largest reward value. The action parameter corresponding to the reward is the optimal action parameter. After the optimal action parameter corresponding to the traversed state parameter is determined, it is determined that the scheduling learning process corresponding to the traversed state parameter is completed. At this time, the optimal action parameter corresponding to each state parameter in the reinforcement learning training process can be determined, and each group of state parameters and the optimal action parameter corresponding to the group of state parameters is taken as a group of optimal strategies. All optimal strategies corresponding to the state parameters are taken as the target total strategy. The reward can include one or more of time, delivery date, equipment utilization rate, switching time of machine equipment, and formula switching.

[0115] When the target total strategy is obtained, actual scheduling operation can be performed according to the target total strategy, and scheduling result of the scheduling operation is output after the scheduling operation is completed. The reinforcement learning training and the actual scheduling operation construct a simulation scheduling model and collect environment parameters to determine state parameters, action parameters and standard returns, and construct an initial strategy according to each state parameter and each action parameter. Then, the learning process of the reinforcement learning training is performed on the simulation scheduling model, that is, the state in the simulation scheduling model is initialized, and the state in the simulation scheduling model is converted into an unscheduled state. Then, each state parameter is traversed, and the initial strategy corresponding to the traversed state parameter is determined, and the action corresponding to the state (i.e. the traversed state parameter) is obtained according to the initial strategy, and the action is executed to determine whether the scheduling (i.e. scheduling learning) is completed, if not, a new action is obtained while keeping the current traversed state unchanged, and the execution is continued. If yes, that is, the scheduling is completed, it is necessary to determine whether to end the training (i.e. reinforcement learning training), if not, the scheduling operation needs to be performed on other state parameters, that is, the strategy is updated. If yes, that is, the training is ended, the target total strategy is output, and the reinforcement learning training process is ended. Then, the scheduling operation of the actual scheduling is performed. And in the scheduling process of starting the scheduling operation, all data are obtained first, and the action corresponding to the state is obtained according to the target total strategy, the action is executed, and after the action is executed, it is determined whether the scheduling is completed. If not (i.e. the scheduling is not completed), a new action is continued to be executed. If yes (i.e. the scheduling is completed), the scheduling result is output until the end.

[0116] For example, taking a certain semiconductor processing real-time scheduling project as an example, if only the lithography area processing needs to be determined in the scheduling project, the real environment can be abstracted as one entrance, one exit, one buffer area and various processing machines, and the state (i.e. state parameter) can be designed as the number of products waiting for processing in each machine + the number of products being processed in each machine + the number of products in the buffer area + the number of products being transported to the machine in the buffer area + the number of products in the entrance + the number of completed processing products in the exit. The state is represented as a vector or an array, and each element of the vector represents a specified number, such as the number of products in the buffer area 1, the number of products in the machine 1,..., the number of products in the buffer area n, the number of products in the machine n, the number of products in the entrance and the number of completed products. The action (i.e. action parameter) can be designed as a multi-action of "buffer area Lot selection" + "machine selection". The Lot selection can be the highest priority, first-in-first-out and the least optional machine. The machine selection can be the shortest processing time and the longest machine idle time. The return can be set as the inverse number of the stage time as the basis, and when the machine has the same formula for continuous processing, a positive number is added as a reward. And when determining the best vehicle corresponding to a certain state, a deep regression model can be used to determine the best vehicle, that is, the step of determining the optimal strategy based on the running results of each training action includes:

[0117] Constructing a multi-layer perceptron:

[0118]

[0119] wherein n is the number of layers, m k is the number of neurons of each layer, D(S, A) is a function of state S and action A as input, and the output is the Q value;

[0120] According to the above multi-layer perceptron, the optimal strategy can be obtained, that is:

[0121]

[0122] wherein D(S, a) is a function of state S and the optimal action a corresponding to state S.

[0123] It should be noted that in the reinforcement learning process in the embodiment, as shown in Figure 7 , the return of each stage is dynamically changing, and it is desired to maximize the total return and further improve other indicators by different combinations. That is, since there are multiple stages in the production scheduling process, the state and action corresponding to each stage can be determined, and the return corresponding to each stage can be determined, and then the strategy can be summarized and trained according to the state, action and return. The return can be the stage time (base) + continuous processing of the same formula (additional), and the action can be a multi-action of "buffer Lot selection" + "machine selection". The Lot selection can be the highest priority, first-in-first-out and the least optional machine. The machine selection can be the shortest processing time and the longest machine idle time. And the strategy can be constructed by a multi-layer perceptron and deep reinforcement learning using the PPO algorithm.

[0124] Step S17: performing production scheduling operation according to the third scheduling sequence.

[0125] In the specific implementation, after obtaining the optimal third scheduling sequence, the scheduling operation can be performed according to the third scheduling sequence, and the semiconductor manufacturing production scheduling is quickly optimized, so as to realize reasonable scheduling of resources such as manpower and equipment, help the factory to reasonably allocate production capacity, improve resource utilization, reduce production time, balance the production line, reduce enterprise cost, etc.

[0126] In addition, please refer to 9, the present application also provides a semiconductor manufacturing production scheduling system for implementing the above semiconductor manufacturing production scheduling method, the system comprises:

[0127] The first determination unit 10 is configured to determine a mathematical optimization model according to the data scale of the semiconductor manufacturing production scheduling.

[0128] The first calculation unit 20 is used to generate a first scheduling sequence according to the mathematical optimization model, and to perform iterative calculation on the first scheduling sequence under constraints through the mathematical optimization model to obtain a second scheduling sequence. The second scheduling sequence is the scheduling sequence with the largest fitness value under the constraints, and the constraints include the maximum number of iterations.

[0129] The second determining unit 30 is used to determine the priority of the target attributes of the product to be processed.

[0130] The second calculation unit 40 is used to determine all state parameters of the semiconductor manufacturing environment and the training actions of the second scheduling sequence, traverse each of the state parameters, determine all traversal initial strategies corresponding to the traversed state parameters based on each of the initial strategies, run the training actions in each of the traversal initial strategies, determine the optimal strategy based on the running results of each of the training actions, determine the target overall strategy according to the optimal strategy corresponding to each of the state parameters, and perform simulation scheduling according to the target overall strategy to obtain the third scheduling sequence.

[0131] Production scheduling unit 50 is used to perform scheduling operations according to the third scheduling sequence.

[0132] Please refer to Figure 10 The first computing unit 20 further includes:

[0133] The generation subunit 201 is used to generate two random integers when the data size does not exceed a preset threshold. The two random integers are different integers between 1 and N, where N is the length of the first scheduling sequence.

[0134] The interchange subunit 202 is used to interchange the numerical values ​​of the sequence positions of two corresponding random integers in the first scheduling sequence to obtain the first comparison sequence.

[0135] The iterative subunit 203 is used to calculate and compare the fitness value of the first scheduling sequence with the fitness value of the first comparison sequence; when the fitness value of the first comparison sequence is greater than the fitness value of the first scheduling sequence, the numerical interchange step is repeated for the first comparison sequence until the maximum number of iterations in the constraint condition is met, and a second scheduling sequence is obtained, which is the scheduling sequence with the largest fitness value under the constraint condition.

[0136] In addition, the present invention also provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the semiconductor manufacturing scheduling method described above.

[0137] Compared with the prior art, the present invention has the following beneficial effects:

[0138] 1. By changing the traditional manual scheduling mode, it is easy to face the complex high-mix production environment. The running time is short, and the processing cost of sudden events such as order insertion failure is small, so that the enterprise production is more flexible and efficient.

[0139] 2. The production plan based on real capacity makes the supplier delivery synchronized with the factory scheduling, thereby reducing the inventory cost and transportation cost caused by the advance ordering of raw materials.

[0140] 3. By combining the global search ability of the intelligent algorithm and the local optimization skill of the reinforcement learning algorithm, the production capacity is effectively and reasonably allocated, the resource utilization rate is improved, and the production cost is reduced.

[0141] In summary, by combining the global search ability of the intelligent algorithm and the local optimization skill of the reinforcement learning algorithm, the application realizes the rapid target optimization of semiconductor manufacturing scheduling, effectively and reasonably allocates the production capacity, improves the resource utilization rate, and reduces the production cost.

[0142] The description and application of the application herein are illustrative, and are not intended to limit the scope of the application to the above-mentioned embodiments. Variations and changes of the disclosed embodiments are possible, and the replacement and equivalent components of the embodiments are known to those skilled in the art. It should be clear to those skilled in the art that the application can be realized in other forms, structures, arrangements, proportions, and with other components, materials and parts without departing from the spirit or essential characteristics of the application. Other variations and changes of the disclosed embodiments can be made without departing from the scope and spirit of the application.

Claims

1. A semiconductor manufacturing scheduling method, characterized by, The method comprises the following steps: Obtaining the data scale of the semiconductor manufacturing scheduling, determining a mathematical optimization model according to the data scale; Generating a first scheduling sequence according to the mathematical optimization model, and iteratively calculating the first scheduling sequence under a constraint condition by using the mathematical optimization model to obtain a second scheduling sequence, the second scheduling sequence being a scheduling sequence with the largest adaptive value under the constraint condition, and the constraint condition including a maximum iteration number; Determining all state parameters of a semiconductor manufacturing environment and training actions of the second scheduling sequence, traversing each state parameter, and determining all traversed initial strategies corresponding to the traversed state parameter based on each initial strategy; Running the training actions in each traversed initial strategy, determining an optimal strategy based on a running result of each training action, determining a target total strategy according to the optimal strategy corresponding to each state parameter, and performing simulation scheduling according to the target total strategy to obtain a third scheduling sequence; Performing scheduling operation according to the third scheduling sequence.

2. The semiconductor manufacturing scheduling method of Claim 1, wherein The data scale includes a number of products to be processed.

3. The production scheduling method for semiconductor manufacturing according to Claim 2, wherein The step of iteratively calculating the first scheduling sequence under a constraint condition by using the mathematical optimization model to obtain a second scheduling sequence specifically comprises: When the data scale does not exceed a preset threshold, generating two random integers, the two random integers being different integers between 1 and N, N being the length of the first scheduling sequence; Exchanging values of sequence number positions corresponding to the two random integers in the first scheduling sequence to obtain a first comparison sequence; Calculating and comparing the adaptive value of the first scheduling sequence and the adaptive value of the first comparison sequence; when the adaptive value of the first comparison sequence is greater than the adaptive value of the first scheduling sequence, repeating the value exchanging step for the first comparison sequence until the maximum iteration number in the constraint condition is met to obtain the second scheduling sequence, the second scheduling sequence being a scheduling sequence with the largest adaptive value under the constraint condition.

4. The production scheduling method for semiconductor manufacturing according to Claim 2, wherein The step of iteratively calculating the first scheduling sequence under a constraint condition by using the mathematical optimization model to obtain a second scheduling sequence specifically comprises: When the data scale exceeds a preset threshold, performing roulette screening, crossover operation and mutation operation on the first scheduling sequence under a constraint condition to obtain a plurality of preferred sequences, the constraint condition including a maximum iteration number, a number of individuals in a population, a crossover probability and a mutation probability; Determining the second scheduling sequence according to the plurality of preferred sequences, the second scheduling sequence being a preferred sequence with the largest adaptive value among the plurality of preferred sequences.

5. The production scheduling method for semiconductor manufacturing according to Claim 4, wherein The step of performing roulette screening on the first scheduling sequence under a constraint condition specifically comprises: Calculating the adaptive value of each random sequence in the first scheduling sequence, constructing a normalized interval according to the proportion of the adaptive value of each random sequence in the sum of the adaptive values of all random sequences, the normalized interval including a plurality of sub-intervals corresponding to each adaptive value, and the normalized interval having a value range of [0, 1]; a plurality of random numbers between 0 and 1 are generated, the number of the random numbers is the same as the number of the plurality of random order sequences; each random number is determined to fall into a sub-interval in the normalized interval, and a corresponding random order sequence is selected according to the fitness value of the sub-interval into which the random number falls, thereby obtaining a plurality of screened random order sequences.

6. The production scheduling method for semiconductor manufacturing according to Claim 5, wherein The step of performing the crossover operation on the first scheduling order sequence under the constraint condition specifically includes: The random order sequences to be crossed are determined according to the crossover probability and the plurality of screened random order sequences, and the random order sequences to be crossed are paired to obtain a plurality of sequence groups, each sequence group including two random order sequences; a random integer between 1 and N is generated, N being the length of the first scheduling order sequence, and the values of the two random order sequences in the sequence group are exchanged in a forward or reverse direction from the random integer position, thereby obtaining a plurality of crossed random order sequences.

7. The production scheduling method for semiconductor manufacturing according to Claim 6, wherein The step of performing the mutation operation on the first scheduling order sequence under the constraint condition specifically includes: The random order sequences to be mutated are determined according to the mutation probability and the plurality of crossed random order sequences; two random integers between 1 and N are generated, N being the length of the first scheduling order sequence, and the sub-sequence between the positions corresponding to the two random integers in the random order sequence to be mutated is randomly sorted, thereby obtaining a plurality of mutated random order sequences; A preferred sequence is calculated according to the plurality of mutated random order sequences, the preferred sequence being the sequence with the largest fitness value in the plurality of mutated random order sequences.

8. The production scheduling method for semiconductor manufacturing according to Claim 1, wherein The training actions include lead time, customer level, product arrival time from the buffer zone, and product initial order.

9. A semiconductor manufacturing scheduling system which realizes the semiconductor manufacturing scheduling method according to any one of claims 1 to 8, characterized by The system includes: A first determination unit is configured to determine a mathematical optimization model according to the data scale of semiconductor manufacturing scheduling; A first calculation unit is configured to generate a first scheduling order sequence according to the mathematical optimization model, and to obtain a second scheduling order sequence by iteratively calculating the first scheduling order sequence under a constraint condition through the mathematical optimization model, the second scheduling order sequence being a scheduling order sequence with the largest fitness value under the constraint condition, the constraint condition including a maximum iteration number; A second determination unit is configured to determine a target attribute priority of a product to be processed; A second calculation unit is configured to determine all state parameters of a semiconductor manufacturing environment and training actions of the second scheduling order sequence, to traverse each state parameter, to determine all traversed initial strategies corresponding to the traversed state parameter based on each initial strategy, and to run the training actions in each traversed initial strategy, to determine an optimal strategy based on the running results of each training action, to determine a target total strategy according to the optimal strategy corresponding to each state parameter, and to perform simulation scheduling according to the target total strategy to obtain a third scheduling order sequence; A scheduling unit is configured to perform scheduling according to the third scheduling order sequence.

10. A computer readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the semiconductor manufacturing scheduling method according to any one of claims 1-8.