Scheduling simulation method and device of processing equipment, computer readable storage medium, terminal and computer program product
By real-time monitoring of equipment status and optimizing scheduling strategies using neural network models, the problem of equipment resource imbalance in the integrated circuit processing production line was solved, dynamic balancing of equipment load and rational planning of maintenance needs were achieved, and processing efficiency and equipment stability were improved.
Patent Information
- Application Number
- CN202510773982.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-12
AI Technical Summary
In integrated circuit processing production lines, the existing static priority scheduling mechanism leads to an imbalance between idle and overloaded equipment resources, affecting processing efficiency and equipment stability, and frequent maintenance needs affect the processing rhythm.
By real-time monitoring of equipment health status and manufacturing processes, a neural network model combined with deep Q network training is used to build an adaptive scheduling strategy, dynamically optimizing processing and maintenance strategies to balance equipment load and maintenance needs.
It achieves the coordinated optimization of processing efficiency and equipment stability, reduces idle equipment resources and overload, and improves the overall throughput efficiency and equipment utilization of the production line.
Smart Images

Figure CN120630864A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and specifically to a scheduling simulation method and device for processing equipment, a computer-readable storage medium, a terminal, and a computer program product. Background Art
[0002] Current integrated circuit processing production lines generally use a static priority scheduling mechanism, which allocates resources based solely on preset priorities when tasks conflict. While this model prioritizes the processing timeliness of high-level components, the lack of real-time assessment of dynamic parameters such as equipment load and process integration leads to long-term queueing of low-priority tasks, resulting in an imbalance between idle and overloaded equipment resources, severely restricting the overall throughput efficiency of the production line. Furthermore, to maintain the nanometer-level precision requirements of processing equipment, irregular equipment maintenance is required, which also affects the overall processing rhythm and reduces the overall processing efficiency of the production line.
[0003] To address the above-mentioned shortcomings, it is urgent to develop an algorithm that can perceive the dynamic allocation of tasks among multiple resources, and to build an adaptive control strategy by synchronously monitoring the health status of equipment and the manufacturing process in real time. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a scheduling simulation method and device for processing equipment, a computer-readable storage medium, a terminal, and a computer program product, which provide an opportunity to rationally plan processing or maintenance, thereby achieving coordinated optimization of processing efficiency and equipment stability.
[0005] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions.
[0006] An embodiment of the present invention provides a scheduling simulation method for processing equipment, which is applicable to a processing production line for integrated circuits. The method includes: determining historical state characteristic parameters and current state characteristic parameters of the processing equipment; determining a scheduling strategy for the processing equipment, wherein the scheduling strategy includes one or more processing strategies and / or one or more maintenance strategies; and inputting the current state characteristic parameters, historical state characteristic parameters, and scheduling strategy into a scheduling model to obtain an execution cost under the scheduling strategy.
[0007] Optionally, the historical state characteristic parameters of the processing equipment include one or more of the following: the completion time Cm of the last task assigned to the processing equipment as of the current moment; k (t0); the processing time Tpc of the workpiece processed by the processing equipment at the current moment j,k (t0); the number of processes Op that have been completed by the processing equipment at the current moment j (t0).
[0008] Optionally, the current state characteristic parameters of the processing equipment include: the utilization rate U of the processing equipment at the current moment k (t0).
[0009] Optionally, the processing strategy of the processing equipment includes one or more of the following: among all the workpieces to be processed in the next process of the processing equipment, the workpiece to be processed with the longest processing time is given priority; among all the workpieces to be processed in the next process of the processing equipment, the workpiece to be processed with the shortest processing time is given priority; among all the workpieces to be processed in the next process of the processing equipment, the workpiece to be processed with the largest process ratio is given priority, and the largest process ratio is used to indicate that the processing time of the next process accounts for the largest proportion relative to the total processing time; among all the workpieces to be processed in the next process of the processing equipment, the workpiece to be processed with the smallest process ratio is given priority, and the smallest process ratio is used to indicate that the processing time of the next process accounts for the smallest proportion relative to the total processing time; all the workpieces to be processed in the next process of the processing equipment Among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the largest number of remaining processes is given priority; among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the least number of remaining processes is given priority; among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the longest remaining processing time is given priority; among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the shortest remaining processing time is given priority; among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the highest penalty for not completing the remaining processes on time is given priority; among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the lowest storage cost after processing is completed is given priority; among all the workpieces to be processed in the next process of the processing equipment, the workpiece to be processed is randomly selected.
[0010] Optionally, the maintenance strategy of the processing equipment includes one or more of the following: a temporary maintenance strategy with a first maintenance urgency and a first maintenance duration; a parts replacement strategy with a second maintenance urgency and a second maintenance duration; a maintenance non-new overhaul strategy with a second maintenance urgency and a first maintenance duration; a minor repair strategy in the face of a fault with a first maintenance urgency and a second maintenance duration; wherein the first maintenance urgency is greater than the second maintenance urgency, and the first maintenance duration is greater than the second maintenance duration.
[0011] Optionally, before inputting the current state characteristic parameters, historical state characteristic parameters, and scheduling strategy into the scheduling model, the method further includes: providing a preset neural network model and providing training data, the training data including current training data based on the current state characteristic parameters, historical training data based on the historical state characteristic parameters, and the scheduling strategy; using the current training data, historical training data, and scheduling strategy to train the neural network model to obtain optimized weight parameters, and using the optimized neural network model as the scheduling model.
[0012] Optionally, the preset neural network model is a deep Q network; training the neural network model includes: initializing the experience pool and the experience pool capacity; initializing the action value function network Q and the random action weight parameter θ; initializing the target value function network And the target weight parameter θ - , where θ - =θ; initialize the state input vector φ1 of state S1, the state input vector φ1 includes the current training data and historical training data with time 1 as the current time; increase one by one from time 1 to time T, and perform the following steps at each time t: input state S t The state input vector φ t , the state input vector φ t Contains current training data and historical training data with time t as the current time; selects a single processing strategy and / or a single maintenance strategy as action a in the scheduling strategy t ; Determine the action a t The reward function value Rw t ; Input state S t+1 The state input vector φ t+1 , the state input vector φ t+1 Contains the current training data and historical training data with time t+1 as the current time; conversion experience sample (φ t ,a t ,Rw t ,φ t+1 ) and stored in the experience pool; randomly sample No. l from the experience pool (φ l ,a l ,Rw l ,φ l+1 ), use the following formula to determine the target Q value y l :
[0013]
[0014] The action weight parameter θ in the loss function is updated by the gradient descent method, and the function value of the loss function is minimized to iteratively optimize the action weight parameter θ until a preset number of iterations is reached or the function value of the loss function is less than or equal to a preset convergence threshold, thereby obtaining the optimized action weight parameter θ and the scheduling model; wherein γ represents the discount factor, 1≤t≤T, and t and T are rational numbers.
[0015] Optionally, in the process of increasing from time 1 to time T, after each preset number of moments, the target value function network Reset the action-value function network Q.
[0016] Optionally, the loss function is: Loss = (y l -Q(φ l ,a l ;θ)) 2 ; Among them, y l represents the target Q value, φ l Represents the state input vector in the l-state scenario, action a l represents a single processing strategy and / or a single maintenance strategy selected in the scheduling strategy in the l-state scenario, θ represents the action weight parameter of the action value function network Q, and Q() represents the action value function network Q.
[0017] Optionally, an ε-greedy algorithm is used to select a single processing strategy and / or a single maintenance strategy as action a in the scheduling strategy. t ;
[0018] The ε-greedy algorithm includes:
[0019]
[0020] Among them, ε is used to represent probability, and 0≤ε≤1, P(s t ,a t ) indicates state S t Next select action a t The probability of t represents the state input vector at time t, A(φ t ) represents the attenuation factor function at time t, a represents the attenuation factor, Q(s t ,a t ) indicates state S t Next action a t The action value function network, θ t Represents the action weight parameter at time t.
[0021] Optionally, a reward function is used to determine the action a t The reward function value Rw t;
[0022] The reward function includes:
[0023]
[0024] Among them, Rw t represents the reward function, t represents the sampling time, j represents the workpiece to be processed, k represents the processing equipment, Op j (t) represents the number of processes that have been completed by the workpiece j to be processed by the processing equipment k at time t, Cm k (t) represents the completion time of the last task assigned to processing equipment k at time t, R k (t) represents the reliability of processing equipment k at time t, SE j (t) represents the processing time of workpiece j to be processed, DE j (t) represents the delay penalty when the remaining steps of the workpiece j are not completed on time, PE j (t) represents the remaining processing time of workpiece j from time t, ME k (t) represents the maintenance demand of processing equipment k at time t.
[0025] An embodiment of the present invention also provides a scheduling simulation device for processing equipment, which is suitable for an integrated circuit processing production line. The device includes: a parameter determination module, used to determine the historical state characteristic parameters and current state characteristic parameters of the processing equipment; a strategy determination module, used to determine the scheduling strategy of the processing equipment, wherein the scheduling strategy includes one or more processing strategies and / or one or more maintenance strategies; a scheduling simulation module, used to input the current state characteristic parameters, historical state characteristic parameters, and scheduling strategy into a scheduling model to obtain the execution cost under the scheduling strategy.
[0026] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the above-mentioned scheduling simulation method for processing equipment is executed.
[0027] An embodiment of the present invention further provides a terminal comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor executes the steps of the above-mentioned scheduling simulation method for processing equipment when running the computer program.
[0028] An embodiment of the present invention further provides a computer program product, including a computer program / instruction, which implements the steps of the scheduling simulation method for processing equipment when executed by a processor.
[0029] Compared with the prior art, the technical solution of the embodiment of the present invention has the following beneficial effects:
[0030] The scheduling simulation method for processing equipment provided in an embodiment of the present invention is applicable to integrated circuit processing production lines. The scheduling strategy is determined based on the historical state characteristic parameters and current state characteristic parameters of the processing equipment. The current state characteristic parameters, historical state characteristic parameters, and the scheduling strategy are input into the scheduling model to obtain the execution cost under the scheduling strategy. Using the above scheme, it is possible to select the optimal scheduling strategy or the scheduling strategy that best meets actual needs based on the different execution costs, and it is possible to reasonably plan processing or maintenance, thereby achieving coordinated optimization of processing efficiency and equipment stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0032] Figure 1 1 is a flow chart of a scheduling simulation method for processing equipment according to an embodiment of the present invention;
[0033] Figure 2 This is a schematic diagram of a deep Q network training process after parameter configuration with experience playback in an embodiment of the present invention;
[0034] Figure 3 It is a structural diagram of a scheduling simulation device for processing equipment according to an embodiment of the present invention;
[0035] Figure 4 It is a hardware structure diagram of a scheduling simulation device for processing equipment in an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The technical solutions of the present invention are described in detail below in conjunction with specific embodiments and the accompanying drawings. The embodiments described herein are specific embodiments of the present invention and are used to illustrate the concept of the present invention. The described embodiments are only a part of the embodiments of this application, rather than all the embodiments. These descriptions are all explanatory and exemplary and should not be construed as limiting the embodiments of the present invention and the scope of protection of the present invention. In addition to the embodiments described herein, those skilled in the art can also adopt other obvious technical solutions based on the contents disclosed in the claims of this application and its specification, including technical solutions that adopt any obvious replacements and modifications to the embodiments described herein.
[0037] As described in the background technology, current integrated circuit processing production lines generally adopt a static priority scheduling mechanism, and when tasks conflict, resources are allocated only according to the preset level. Although this mode can give priority to ensuring the processing timeliness of high-level devices, due to the lack of real-time evaluation of dynamic parameters such as equipment load rate and process connection tightness, low-priority tasks are stranded in the queue for a long time, causing an imbalance between idle and overloaded equipment resources, which seriously restricts the overall throughput efficiency of the production line. At the same time, in order to maintain the nanometer-level precision requirements of the processing equipment, the equipment needs to be maintained from time to time, which will also affect the overall processing rhythm, thereby reducing the processing efficiency of the entire production line.
[0038] To address these shortcomings, there is an urgent need to develop an algorithm that can dynamically allocate tasks across multiple resources. This algorithm, through real-time, simultaneous monitoring of equipment health and manufacturing progress, can then be used to build an adaptive control strategy. This algorithm must quickly trigger a resilient fault-tolerant mechanism in the event of a sudden failure. By combining resource allocation with production task information, this algorithm can minimize repair response time and minimize the impact of equipment anomalies on the production line, thereby achieving the coordinated optimization of processing efficiency and equipment stability.
[0039] The scheduling simulation method for processing equipment provided in an embodiment of the present invention is applicable to integrated circuit processing production lines. It determines a scheduling strategy based on the historical and current state characteristic parameters of the processing equipment. By inputting the current and historical state characteristic parameters, as well as the scheduling strategy, into a scheduling model, the execution cost under the scheduling strategy can be determined. Based on the differences in the execution costs, the optimal scheduling strategy is selected. This method thus enables the rational planning of processing or maintenance, thereby achieving the coordinated optimization of processing efficiency and equipment stability.
[0040] In order to make the above-mentioned objects, features and beneficial effects of the present invention more obvious and easy to understand, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0041] See also Figure 1 , Figure 1 The present invention is a flowchart of a scheduling simulation method for processing equipment. The method can perform the following steps S11 to S13, each of which is described below.
[0042] In step S11, historical state characteristic parameters and current state characteristic parameters of the processing equipment are determined.
[0043] The historical state characteristic parameters of the processing equipment are descriptive parameters of various characteristics of the workpiece processed by the processing equipment before the current moment, which may include the time required to process the workpiece, the processing time of the workpiece on the processing equipment, the number of processing steps completed by the workpiece on the processing equipment, etc.
[0044] In specific implementations, parameters such as workpieces, elapsed processing time, and processing steps can refer to conventional definitions. For example, in the field of semiconductor manufacturing processes, a workpiece can be a single lot or a single wafer; elapsed processing time can be the time spent on the processing equipment during the entire stage process duration or the time spent on the processing equipment during a specific step process duration; and the number of processing steps can be the number of steps as the smallest process unit or the number of times a wafer enters a tool or process chamber.
[0045] The current state characteristic parameters are parameters describing various characteristics of the processing equipment and the workpiece processed in the processing equipment at the current moment, such as the utilization rate of the processing equipment at the current moment.
[0046] The utilization rate may be the percentage of busy chambers to the total number of chambers when the processing equipment includes multiple chambers; or it may be 100% utilization in a busy state or 0% utilization in an idle state when the processing equipment has a single chamber.
[0047] In some embodiments, the current state characteristic parameters may also include other appropriate parameters of the processing equipment, such as one or more of the following: the average utilization rate of the processing equipment within a preset historical period, the standard deviation of the utilization rate of the processing equipment, the reliability of the processing equipment, the fault mark of the processing equipment, the failure rate of the processing equipment, the slack ratio of the production line's pending processing time, and the ratio of the current cost of the production line processing to the expected cost.
[0048] The average utilization rate of processing equipment over a preset historical timeframe refers to the ratio of the equipment's actual processing time to its total available time within a historical timeframe (e.g., 24 hours, a week, or a month). High utilization (a ratio close to 1) indicates that the equipment is fully loaded and overload failure prevention may be necessary. Low utilization (a ratio close to 0) indicates idleness and waste, requiring optimized task allocation. In dynamic scheduling, tasks are prioritized for low-utilization equipment.
[0049] In some embodiments, the formula may be:
[0050]
[0051] The standard deviation of processing equipment utilization refers to the degree of fluctuation in equipment utilization within a historical window. A high standard deviation indicates large load fluctuations, which may cause fatigue wear or sudden failures. A low standard deviation indicates a balanced load, which is beneficial for maintaining equipment life.
[0052] In some embodiments, the formula may be:
[0053]
[0054] The reliability of processing equipment is the probability that the equipment will operate without failure within a certain period of time, t, and typically follows an exponential or Weibull distribution. When reliability drops below a certain level (e.g., R(t) < 0.9), maintenance needs to be scheduled in advance.
[0055] In some embodiments, the formula may be:
[0056] R(t)=e -λt , λ is the failure rate.
[0057] The fault mark of the processing equipment is used to indicate the fault status of the equipment in real time. When the fault status is reached, the task of the equipment is immediately assigned and / or the maintenance process is triggered.
[0058] The failure rate of processing equipment refers to the number of equipment failures per unit time. When the failure rate exceeds a preset value, the load must be reduced or the equipment must be replaced. In the cost model, equipment with a higher failure rate is weighted lower. For example, if λ ≤ 0.01 failures / hour, normal load operation is indicated; if λ < 0.01 ≤ 0.05 , the load must be reduced by 10% to 20% and monitored; if λ < 0.1 ≤ 0.1 , the load must be halved or planned maintenance must be performed; if λ > 0.1 , the equipment must be shut down immediately for replacement.
[0059] The slack ratio of a production line's pending processing time is the ratio of the remaining processing time of a workpiece processed on the production line to the remaining delivery time, reflecting the urgency of the task. If the ratio is greater than or equal to 1, it may lead to delayed delivery and require prioritization of processing.
[0060] The ratio of current production line processing costs to projected costs refers to the ratio of actual consumption costs (energy, labor, losses, and material costs) to budgeted costs. If the ratio is greater than 1, it indicates an overspending state and requires a cost loss check or an alarm.
[0061] Specifically, the historical state characteristic parameters of the processing equipment include one or more of the following:
[0062] The completion time Cm of the last task assigned to the processing equipment as of the current moment k (t0).
[0063] The processing time Tpc of the workpiece processed by the processing equipment at the current moment j,k (t0).
[0064] The number of processes Op that have been completed by the processing equipment at the current moment j (t0).
[0065] The current state characteristic parameters of the processing equipment include: the utilization rate U of the processing equipment at the current momentk (t0).
[0066] The current time is represented as t0.
[0067] In the subsequent training of the neural network model, the input state S t The state input vector φ t It is determined by the historical state characteristic parameters and current state characteristic parameters of the processing equipment. The formula is as follows:
[0068] φ t ={Op j (t), Cm k (t), U k (t), Tpc j,k (t)}, where 1≤t≤T, and t and T are rational numbers.
[0069] In step S12, a scheduling strategy for the processing equipment is determined, where the scheduling strategy includes one or more processing strategies and / or one or more maintenance strategies.
[0070] The scheduling strategy for a processing device can be either a machining strategy or a maintenance strategy. The machining strategy for a processing device refers to a set of intelligent scheduling methods based on dynamic priority rules. Its core is to evaluate the priority of workpieces to be processed using multi-dimensional parameters. The machining strategy for a processing device is a set of priority rules based on different optimization objectives that are used to dynamically schedule the processing order of workpieces to be processed.
[0071] Specific strategies include: Longest Processing Time First (LPT), which selects the workpiece with the longest processing time in the next process to balance equipment load; Shortest Processing Time First (SPT), which prioritizes the workpiece with the shortest processing time to quickly free up tasks; Highest / Lowest Process Proportion First, which selects the workpiece with the highest or lowest proportion of the next process time to its total processing time, respectively, to optimize the efficiency of key links or disperse bottleneck pressure; Highest / Lowest Remaining Process First, the former is targeted at complex workpieces requiring multiple processing, the latter is used to simplify processes; Longest / Shortest Remaining Time First, the former avoids backlogs of long-term tasks, the latter accelerates completion; Highest Penalty Cost First, which prioritizes tasks with high default risk to reduce losses; Lowest Storage Cost First, which focuses on reducing inventory pressure; and Random Selection, which serves as the basis for the rule-free strategy. These strategies can be flexibly combined to adapt to different production management needs, such as production efficiency, cost control, or risk avoidance.
[0072] Specifically, the processing strategy of the processing equipment includes one or more of the following:
[0073] Among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the longest processing time is given priority. Among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the shortest processing time is given priority.
[0074] The processing time here refers to the processing time of all workpieces to be processed in the next process, which is the collection of processing time of each workpiece to be processed in the next process, rather than the sum of processing time of all workpieces to be processed in the next process.
[0075] Among all the workpieces to be processed in the next process of the processing equipment, the workpiece to be processed with the largest process proportion is given priority, and the largest process proportion is used to indicate that the processing time of the next process accounts for the largest proportion relative to the total processing time.
[0076] Among all workpieces to be processed in the next process of the processing equipment, the workpiece with the smallest process ratio is prioritized. The smallest process ratio indicates that the processing time of the next process accounts for the smallest proportion of the total processing time. Among all workpieces to be processed in the next process of the processing equipment, the workpiece with the largest number of remaining processes is prioritized.
[0077] Among all the workpieces to be processed in the next process of the processing equipment, the workpiece to be processed with the least number of remaining processes is given priority.
[0078] In some embodiments, the remaining number of process steps refers to the remaining number of process steps of the workpiece to be processed in the equipment.
[0079] In some embodiments, the remaining number of processes refers to the remaining number of processes of the workpiece to be processed in its production line.
[0080] Among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the longest remaining processing time is given priority.
[0081] Among all the workpieces to be processed in the next process of the processing equipment, the workpiece to be processed with the shortest remaining processing time is given priority.
[0082] In some embodiments, the remaining processing time refers to the remaining processing time of the workpiece to be processed in the equipment.
[0083] In some embodiments, the remaining processing time refers to the remaining processing time of the workpiece to be processed in its production line.
[0084] Among all the workpieces to be processed in the next process of the processing equipment, the workpiece to be processed with the highest penalty for not completing the remaining processes on time is given priority.
[0085] In some embodiments, the remaining process refers to the remaining process of the workpiece to be processed in the equipment.
[0086] In some embodiments, the remaining process refers to the remaining process of the workpiece to be processed in its production line.
[0087] Penalty is a quantified cost of not completing the remaining steps of a workpiece on time. It typically includes: losses caused by unfinished processing, such as contract breach penalties, loss of customer trust, and subsequent production line blockage (such as the cost of idle downstream processes).
[0088] Among all the workpieces to be processed in the next process of the processing equipment, the workpiece to be processed with the lowest storage cost after processing is completed is given priority.
[0089] Storage cost refers to the comprehensive cost incurred when a workpiece enters the finished product warehouse or transit inventory stage after processing is completed. It usually includes the following components: inventory holding cost, management cost, and opportunity cost.
[0090] Inventory holding costs include space occupancy fees (such as warehouse rent and depreciation), capital occupancy costs (such as working capital interest caused by inventory backlogs), insurance costs (such as insurance for damage or loss of goods), and loss costs (such as deterioration of perishable goods and rust of metal parts).
[0091] Management costs: labor costs for inventory counting, handling, and sorting; information system maintenance costs (such as WMS warehouse management system).
[0092] Opportunity cost: Potential loss of revenue from not being able to take new orders due to inventory backlog.
[0093] A workpiece to be processed is randomly selected from all workpieces to be processed in the next process of the processing equipment.
[0094] The above 11 processing strategies are the processing strategies selected in this embodiment.
[0095] It is understandable that the above examples do not constitute a limitation on the number and types of the processing strategies, and those skilled in the relevant art can obviously add and / or delete the processing strategies based on this embodiment.
[0096] The maintenance strategy is a dynamic hierarchical management method based on equipment failure characteristics and maintenance resource constraints. Strategy types are divided into two dimensional parameters: urgency (reflecting the threat level of the failure to production continuity) and maintenance time (indicating the time required for repair operations). Specifically, they include:
[0097] Temporary maintenance strategy: Rapid response to sudden failures, prioritizing reducing downtime risks but requiring a longer repair cycle;
[0098] Parts replacement strategy: Standardized spare parts replacement addresses predictable component wear and balances timeliness and cost-effectiveness.
[0099] Maintenance non-new overhaul strategy: Applicable to systematic overhaul of non-critical equipment, allowing planned resource deployment;
[0100] Minor repair strategy for failures: immediate repair of local functional failures, emphasizing rapid resumption of production.
[0101] Urgency is determined using hierarchical thresholds (e.g., Level 1 > Level 2), reflecting the priority of the fault's impact. Maintenance duration is estimated through historical data modeling or real-time diagnostics, dynamically adapting to resource scheduling needs. These strategies can be executed independently or combined iteratively, and are not limited to specific parameter value ranges. Their selection is based on factors such as equipment criticality, spare parts availability, and production plan constraints, ultimately achieving multi-objective optimization of maintenance costs, capacity loss, and safety risks.
[0102] Specifically, the maintenance strategy for the processing equipment includes one or more of the following:
[0103] The temporary maintenance strategy has a first maintenance urgency and a first maintenance duration.
[0104] Part replacement strategy with second repair urgency and second repair duration.
[0105] Maintain a non-new overhaul strategy with the second-highest repair urgency and the first-highest repair duration.
[0106] The minor repair strategy for a fault has a first repair urgency and a second repair duration.
[0107] The first maintenance urgency is greater than the second maintenance urgency, and the first maintenance duration is greater than the second maintenance duration.
[0108] In other words, the maintenance strategy refers to the specific maintenance methods taken for different failure scenarios, including:
[0109] Temporary maintenance strategy (highest urgency, longest duration): emergency measures that require immediate execution but have complex processes (such as system restart or temporary reinforcement);
[0110] Minor repair strategy for faults (highest urgency, shortest duration): quick repair of local faults (such as repair welding or replacement of small parts);
[0111] Part replacement strategy (lowest urgency, shorter duration): medium-priority replacement operations that rely on spare parts (such as chip or capacitor replacement);
[0112] Maintenance without overhaul strategy (second highest urgency, longest duration): Systematic overhaul that does not require replacement of parts (such as cleaning and calibration or maintenance of aging parts).
[0113] The maintenance urgency indicates the priority of fault handling. The first maintenance urgency (such as downtime risk) requires immediate response, while the second maintenance urgency (such as performance degradation) allows for moderate delay.
[0114] Maintenance time refers to the time required to complete the strategy. The first maintenance time (e.g., several hours) involves complex operations, while the second maintenance time (e.g., minutes) is for rapid intervention.
[0115] The system balances risk control and resource efficiency through a combination of urgency and time consumption, ensuring that critical faults are handled first and avoiding excessive use of maintenance resources.
[0116] The above four maintenance strategies are the maintenance strategies selected in this embodiment.
[0117] It is understandable that the above examples do not limit the number and types of the maintenance strategies, and those skilled in the art can obviously add and / or delete the maintenance strategies based on this embodiment.
[0118] In step S13, the current state characteristic parameters, the historical state characteristic parameters, and the scheduling strategy are input into a scheduling model to obtain the execution cost under the scheduling strategy.
[0119] Before inputting the current state characteristic parameters, the historical state characteristic parameters, and the scheduling strategy into the scheduling model, the method further includes:
[0120] Providing a preset neural network model and providing training data, wherein the training data includes current training data based on the current state characteristic parameters, historical training data based on the historical state characteristic parameters, and the scheduling strategy;
[0121] The current training data, historical training data and scheduling strategy are used to train the neural network model to obtain optimized weight parameters, and the optimized neural network model is used as the scheduling model.
[0122] Among them, the selection of state characteristics (the current state characteristic parameters, the historical state characteristic parameters) plays an important role in the overall decision-making process. Only by constructing appropriate state characteristics for the production line production and operation environment can the current scheduling state (input state) be accurately represented. In other words, the input state S t The state input vector φ t It is determined by the historical state characteristic parameters and current state characteristic parameters of the processing equipment. The formula is as follows:
[0123] φ t ={Op j (t), Cm k (t), U k (t), Tpc j,k (t)}, where 1≤t≤T, and t and T are rational numbers.
[0124] When the time step is cycled from 1 to T, T state input vectors φ can be obtained t , thus obtaining T input states S t .
[0125] See also Figure 2 , Figure 2 This is a diagram of the deep Q network training process after parameter configuration with experience playback. The input vector φ is input from the input layer t The output layer Q value is calculated through hidden layer transformation and stored in the rule base. According to the scheduling rules in the rule base, the device is instructed to pass the reward function to obtain the reward value, and then the input state is updated.
[0126] The preset neural network model is a deep Q network.
[0127] Training the neural network model includes:
[0128] Initialize the experience pool and experience pool capacity;
[0129] Initialize the action value function network Q and random action weight parameters θ;
[0130] Initialize the target value function network And the target weight parameter θ - , where θ - =θ;
[0131] Initialize the state input vector φ1 of the state S1, wherein the state input vector φ1 includes current training data with time 1 as the current time and historical training data.
[0132] From moment 1 to moment T, the following steps are performed at each moment t:
[0133] Input status S t The state input vector φ t , the state input vector φ t Contains the current training data and historical training data with time t as the current time.
[0134] Select a single processing strategy and / or a single maintenance strategy as action a in the scheduling strategy t .
[0135] Determine the action a t The reward function value Rw t .
[0136] Input status S t+1 The state input vector φ t+1 , the state input vector φ t+1 Contains current training data and historical training data with time t+1 as the current time.
[0137] Transformation experience sample (φ t ,a t ,Rw t ,φ t+1 ) and stored in the experience pool.
[0138] Randomly sample l from the experience pool (φ l ,a l ,Rw l ,φ l+1 ), use the following formula to determine the target Q value y l :
[0139]
[0140] The action weight parameter θ in the loss function is updated using a gradient descent method to minimize the function value of the loss function. The action weight parameter θ is iteratively optimized until a preset number of iterations is reached or the function value of the loss function is less than or equal to a preset convergence threshold. The optimized action weight parameter θ and the scheduling model are then obtained. Here, γ represents a discount factor, 1≤t≤T, and t and T are rational numbers.
[0141] In the process of increasing from time 1 to time T, each time after a preset number of moments, the target value function network Reset the action-value function network Q.
[0142] Specifically, in the initialization phase, the experience pool can be used to store experience samples (φ t ,a t ,Rw t ,φ t+1 The experience pool capacity determines the experience pool can store the experience samples (φ t ,a t ,Rw t ,φ t+1The size of the experience pool is determined by the hardware performance and the needs of those skilled in the art. The larger the experience pool capacity, the more conducive it is to covering a variety of state-action pairs. The action-value function network Q, whose parameters are random weights θ, is used to estimate the Q value (expected cumulative reward) of the current state-action pair. The target value function network Its initial target weight parameter θ - =θ, the target network periodically synchronizes parameters from the Q network to provide stable target Q values to alleviate training fluctuations.
[0143] After the initialization phase is completed, the main training loop is entered, namely the state scene loop and the time step loop. The main training loop is a double loop in which the time step loop is embedded in the state scene loop.
[0144] The state scenario cycles from state scenario 1 to state scenario L, that is, 1≤l≤L, where l and L are rational numbers.
[0145] The state scene cycle starts, initializing state S1, and the state input vector φ1 of state S1 is
[0146] φ1={Op j (1), Cm k (1), U k (1), Tpc j,k (1)}.
[0147] The time step cycle is started, and the number of time steps cycles from 1 to T, 1≤t≤T, and t and T are rational numbers.
[0148] At each moment t, the state S t The state input vector φ t for:
[0149] φ t ={Op j (t), Cm k (t), U k (t), Tpc j,k (t)}.
[0150] Select and execute an action, that is, select a single processing strategy and / or a single maintenance strategy as an action in the scheduling strategy. t .
[0151] Furthermore, the ε greedy algorithm is used to select a single processing strategy and / or a single maintenance strategy as action a in the scheduling strategy. t .
[0152] The ε-greedy algorithm includes:
[0153]
[0154] Among them, ε is used to represent probability, and 0≤ε≤1, P(s t ,a t ) indicates state S t Next select action a t The probability of t represents the state input vector at time t, A(φ t ) represents the attenuation factor function at time t, a represents the attenuation factor, Q(s t ,a t ) indicates state S t Next action a t The action value function network, θ t Represents the action weight parameter at time t.
[0155] In a specific embodiment, a single processing strategy and / or a single maintenance strategy is selected as action a in the scheduling strategy by random selection. t .
[0156] Among them, the action a t It can be selected from the 11 processing strategies and 4 maintenance strategies mentioned above, or other appropriate strategies.
[0157] In another specific embodiment, a single processing strategy and / or a single maintenance strategy is selected as an action in the scheduling strategy according to a number sequence, wherein the number sequence is based on a priority number sequence or a random number sequence.
[0158] Furthermore, a reward function is used to determine the action a t The reward function value Rw t .
[0159] The reward function includes:
[0160]
[0161] Among them, Rw t represents the reward function, t represents the sampling time, j represents the workpiece to be processed, k represents the processing equipment, Op j (t) represents the number of processes that have been completed by the workpiece j to be processed by the processing equipment k at time t, Cm k (t) represents the completion time of the last task assigned to processing equipment k at time t, R k (t) represents the reliability of processing equipment k at time t, SE j (t) represents the time cost of replacing the workpiece j to be processed, DE j (t) represents the delay penalty when the remaining steps of the workpiece j are not completed on time, PEj (t) represents the energy consumption cost of workpiece j to be processed from time t, ME k (t) represents the maintenance cost of processing equipment k at time t.
[0162] Determine to perform the action a t The reward function value Rw t After that, update the input state S t+1 Status S t+1 The state input vector φ t+1 for
[0163] φ t+1 ={Op j (t+1), Cm k (t+1), U k (t+1), Tpc j,k (t+1)}.
[0164] The experience sample (φ t ,a t ,Rw t ,φ t+1 ) is stored in the experience pool.
[0165] In some embodiments, the experience samples in the experience pool may be represented by state scenarios. For example, the lth state scenario corresponds to the lth experience sample, and may also correspond to the experience sample at the lth moment.
[0166] Use the following formula to determine the target Q value y l :
[0167]
[0168] The action weight parameter θ in the loss function is updated using a gradient descent method to minimize the function value of the loss function. The action weight parameter θ is iteratively optimized until a preset number of iterations is reached or the function value of the loss function is less than or equal to a preset convergence threshold. The optimized action weight parameter θ and the scheduling model are then obtained. Here, γ represents a discount factor, 1≤t≤T, and t and T are rational numbers.
[0169] In some embodiments, the value of the discount factor γ satisfies 0≤γ≤1, but is not limited thereto.
[0170] Furthermore, the loss function is: Loss = (y l -Q(φ l ,a l ;θ)) 2 Among them, y l represents the target Q value, φ l Represents the state input vector of the l-state scene, action al represents a single processing strategy and / or a single maintenance strategy selected in the scheduling strategy in the l-state scenario, θ represents the action weight parameter of the action value function network Q, and Q() represents the action value function network Q.
[0171] The loss function can also be selected from existing conventional model training loss functions, such as L1 loss function, L2 loss function, mean square error loss function, cross entropy loss function, and logarithmic loss function.
[0172] It should be noted that in some embodiments, a loss function may be used to operate on the state input vector at time t between time 1 and time T. In this case, the loss function is correspondingly adjusted to:
[0173] Loss=(y t -Q(φ t ,a t ;θ)) 2
[0174] In a specific embodiment, in the process of increasing from time 1 to time T, after each preset number of time moments, the target value function network is reset to the action value function network Q. The preset number can range from 1,000 to 10,000.
[0175] In some specific embodiments, in the process of increasing from time 1 to time T, the steps from "obtaining action strategy" to "updating target value function network" can be repeated. " steps are looped until the time step loop and the state scene loop are completed.
[0176] In a specific embodiment, the steps of the loop processing may include the following one or more steps: selecting and executing an action; determining the reward function value of the action; updating the input state S t+1 ; The experience sample (φ t ,a t ,Rw t ,φ t+1 ) is stored in the experience pool; the target Q value is determined; the action weight parameter θ in the loss function is updated using the gradient descent method to obtain the updated target value function network
[0177] In the embodiment of the present invention, by adopting the loss function and the reward function, it is possible to optimize production decisions, improve resource utilization efficiency and enhance overall production benefits more scientifically and accurately on the basis of cost control.
[0178] See also Figure 3 , Figure 31 is a schematic diagram of a scheduling simulation device for processing equipment according to an embodiment of the present invention. The scheduling simulation device can be applied to an integrated circuit processing production line, and the device can include: a parameter determination module 31 for determining historical state characteristic parameters and current state characteristic parameters of the processing equipment;
[0179] a strategy determination module 32 for determining a scheduling strategy for the processing equipment, wherein the scheduling strategy includes one or more processing strategies and / or one or more maintenance strategies;
[0180] The scheduling simulation module 33 is used to input the current state characteristic parameters, historical state characteristic parameters, and scheduling strategy into a scheduling model to obtain the execution cost under the scheduling strategy.
[0181] about Figure 3 For more information on the working principle, working method and beneficial effects of the scheduling simulation device shown, please refer to the previous and Figures 1 to 2 The detailed description is omitted here.
[0182] In specific implementation, Figure 3 The scheduling simulation device shown may correspond to a chip having a data processing function in a data processing device; or correspond to a chip or chip module having a data processing function in a data processing device, or correspond to a data processing device.
[0183] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the above-mentioned scheduling simulation method for processing equipment is executed.
[0184] An embodiment of the present invention further provides a terminal comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor executes the steps of the above-mentioned scheduling simulation method for processing equipment when running the computer program.
[0185] An embodiment of the present invention further provides a computer program product, including a computer program / instruction, which implements the steps of the scheduling simulation method for processing equipment when executed by a processor.
[0186] Reference Figure 4 , Figure 4 It is a hardware structure diagram of a scheduling simulation device for processing equipment in an embodiment of the present application.
[0187] Figure 4The terminal shown includes a memory 41, a processor 42, and a transceiver 43. The processor 42 is coupled to the memory 41 and the transceiver 43. The memory 41 may be located inside or outside the terminal. The memory 41, processor 42, and transceiver 43 may be connected via a communication bus. The transceiver 43 is used to communicate with other devices or a communication network.
[0188] Optionally, the transceiver 43 may be a transmitter. The memory 41 stores a computer program that can be run on the processor 42. When the processor 42 runs the computer program, the transceiver 43 executes the steps of the communication method provided in the above embodiment.
[0189] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0190] It should also be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a ROM, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0191] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired or wireless means.
[0192] It should be understood that the term "and / or" as used herein simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " as used herein indicates that the related objects are in an "or" relationship.
[0193] The term "plurality" used in the embodiments of the present application refers to two or more.
[0194] In this application, "equal to" can be used in conjunction with "less than" or "greater than", but not with both "less than" and "greater than". When "equal to" is used in conjunction with "less than", the technical solution used for "less than" applies. When "equal to" is used in conjunction with "greater than", the technical solution used for "greater than" applies.
[0195] The first, second, etc. descriptions appearing in the embodiments of this application are only for illustration and distinction of the description objects. There is no order, nor does it indicate any special limitation on the number of devices in the embodiments of this application, and cannot constitute any limitation on the embodiments of this application.
[0196] Although the embodiments of the present application are disclosed above, the present application is not limited thereto. Any person skilled in the art may make various changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims.
Claims
1. A scheduling simulation method for processing equipment, characterized in that: Applicable to an integrated circuit processing production line, the method includes: Determining historical state characteristic parameters and current state characteristic parameters of the processing equipment; Determining a scheduling strategy for the processing equipment, wherein the scheduling strategy includes one or more processing strategies and / or one or more maintenance strategies; The current state characteristic parameters, the historical state characteristic parameters, and the scheduling strategy are input into a scheduling model to obtain the execution cost under the scheduling strategy.
2. The method according to claim 1, characterized in that The historical state characteristic parameters of the processing equipment include one or more of the following: The completion time Cm of the last task assigned to the processing equipment as of the current moment k (t0); The processing time Tpc of the workpiece processed by the processing equipment at the current moment j,k (t0); The number of processes Op that have been completed for the workpiece processed by the processing equipment at the current moment j (t0).
3. The method according to claim 1, characterized in that The current state characteristic parameters of the processing equipment include: The utilization rate U of the processing equipment at the current moment k (t0).
4. The method according to claim 1, wherein The processing strategy of the processing equipment includes one or more of the following: Among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the longest processing time is given priority; Among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the shortest processing time is given priority; Among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the largest process proportion is given priority, where the largest process proportion is used to indicate that the processing time of the next process accounts for the largest proportion of the total processing time; Among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the smallest process ratio is given priority, where the smallest process ratio indicates that the processing time of the next process accounts for the smallest proportion relative to the total processing time; Among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the largest number of remaining processes is given priority; Among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the least number of remaining processes is given priority; Among all the workpieces to be processed in the next process of the processing equipment, the workpiece with the longest remaining processing time is given priority; Among all the workpieces to be processed in the next step of the processing equipment, the workpiece with the shortest remaining processing time is given priority; Among all the workpieces to be processed in the next process of the processing equipment, the workpiece to be processed with the highest penalty for not completing the remaining processes on time is given priority; Among all the workpieces to be processed in the next step of the processing equipment, the workpiece with the lowest storage cost after processing is given priority; A workpiece to be processed is randomly selected from all workpieces to be processed in the next process of the processing equipment.
5. The method according to claim 1, wherein The maintenance strategy for the processing equipment includes one or more of the following: A temporary maintenance strategy, having a first maintenance urgency and a first maintenance duration; Part replacement strategy with second repair urgency and second repair duration; Maintain a non-new overhaul strategy with a second repair urgency and a first repair duration; A minor repair strategy for a fault has a first level of repair urgency and a second level of repair duration; The first maintenance urgency is greater than the second maintenance urgency, and the first maintenance duration is greater than the second maintenance duration.
6. The method according to claim 1, characterized in that Before inputting the current state characteristic parameters, the historical state characteristic parameters, and the scheduling strategy into the scheduling model, the method further includes: Providing a preset neural network model and providing training data, wherein the training data includes current training data based on the current state characteristic parameters, historical training data based on the historical state characteristic parameters, and the scheduling strategy; The current training data, historical training data and scheduling strategy are used to train the neural network model to obtain optimized weight parameters, and the optimized neural network model is used as the scheduling model.
7. The method according to claim 6, characterized in that The preset neural network model is a deep Q network; Training the neural network model includes: Initialize the experience pool and experience pool capacity; Initialize the action value function network Q and random action weight parameters θ; Initialize the target value function network And the target weight parameter θ - , where θ - =θ; Initialize a state input vector φ1 of state S1, wherein the state input vector φ1 includes current training data and historical training data with time 1 as the current time; From moment 1 to moment T, the following steps are performed at each moment t: Input status S t The state input vector φ t , the state input vector φ t Contains the current training data and historical training data with time t as the current moment; Select a single processing strategy and / or a single maintenance strategy as action a in the scheduling strategy t ; Determine the action a t The reward function value Rw t ; Input status S t+1 The state input vector φ t+1 , the state input vector φ t+1 Contains current training data and historical training data with time t+1 as the current time; Transformation experience sample (φ t ,a t ,Rw t ,φ t+1 ), and stored in the experience pool; Randomly sample l from the experience pool (φ l ,a l ,Rw l ,φ l+1 ), use the following formula to determine the target Q value y l : The action weight parameter θ in the loss function is updated using a gradient descent method to minimize the function value of the loss function, so as to iteratively optimize the action weight parameter θ until a preset number of iterations is reached or the function value of the loss function is less than or equal to a preset convergence threshold, thereby obtaining the optimized action weight parameter θ and the scheduling model; Where γ represents the discount factor, 1≤t≤T, and t and T are rational numbers.
8. The method according to claim 7, characterized in that In the process of increasing from time 1 to time T, each time after a preset number of moments, the target value function network Reset the action-value function network Q.
9. The method according to claim 7, characterized in that The loss function is: Loss=(y l -Q(φ l ,a l ;i)) 2 ; Among them, y l represents the target Q value, φ l Represents the state input vector in the l-state scenario, action a l represents a single processing strategy and / or a single maintenance strategy selected in the scheduling strategy in the l-state scenario, θ represents the action weight parameter of the action value function network Q, and Q() represents the action value function network Q.
10. The method according to claim 7, characterized in that Using the ε-greedy algorithm, a single processing strategy and / or a single maintenance strategy is selected as the action at in the scheduling strategy; The ε-greedy algorithm includes: Among them, ε is used to represent probability, and 0≤ε≤1, P(s t ,a t ) indicates state S t Next select action a t The probability of t represents the state input vector at time t, A(φ t ) represents the attenuation factor function at time t, a represents the attenuation factor, Q(s t ,a t ) indicates state S t Next action a t The action value function network, θ t Represents the action weight parameter at time t.
11. The method according to claim 7, characterized in that Using the reward function, determine the action a t The reward function value Rw t ; The reward function includes: Among them, Rw t represents the reward function, t represents the sampling time, j represents the workpiece to be processed, k represents the processing equipment, Op j (t) represents the number of processes that have been completed by the workpiece j to be processed by the processing equipment k at time t, Cm k (t) represents the completion time of the last task assigned to processing equipment k at time t, R k (t) represents the reliability of processing equipment k at time t, SE j (t) represents the processing time of workpiece j to be processed, DE j (t) represents the delay penalty when the remaining steps of the workpiece j are not completed on time, PE j (t) represents the remaining processing time of workpiece j from time t, ME k (t) represents the maintenance demand of processing equipment k at time t.
12. A scheduling simulation device for processing equipment, characterized in that: Applicable to integrated circuit processing production lines, the device includes: a parameter determination module, configured to determine characteristic parameters of historical states and current states of the processing equipment; a strategy determination module, configured to determine a scheduling strategy for the processing equipment, wherein the scheduling strategy includes one or more processing strategies and / or one or more maintenance strategies; The scheduling simulation module is used to input the current state characteristic parameters, historical state characteristic parameters, and scheduling strategy into a scheduling model to obtain the execution cost under the scheduling strategy.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the scheduling simulation method for processing equipment according to any one of claims 1 to 11 is executed.
14. A terminal comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein: When the processor runs the computer program, the processor performs the steps of the scheduling simulation method for processing equipment according to any one of claims 1 to 11.
15. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the scheduling simulation method for processing equipment according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Single-machine scheduling method considering maintenance decision based on deep reinforcement learning
CN115237076A
Flexible job shop scheduling method and device and readable storage medium
CN118195263A
Wafer manufacturing system furnace tube area scheduling method based on reinforcement learning
CN118446475A
Reinforcement learning agent interaction strategy network training method suitable for job shop scheduling, program product and system
CN118657337A