A lean operation method based on production scheduling optimization and equipment OEE monitoring

By building a reinforcement learning production scheduling optimization neural network and a real-time equipment OEE monitoring unit, the problem of disconnection between production scheduling strategy and equipment OEE status was solved, dynamic coordinated adjustment of production scheduling and equipment status was achieved, and production scheduling execution efficiency and equipment utilization were improved.

CN120525140BActive Publication Date: 2025-09-19深圳市永迦电子科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511029217.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-09-19
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

In existing technologies, production scheduling strategies are disconnected from the equipment's OEE status and lack a real-time linkage mechanism. This results in high-load equipment running continuously and inefficient equipment not being identified and adjusted in a timely manner, affecting resource allocation efficiency and making it difficult to achieve targeted optimization and closed-loop control.

Method used

Construct a reinforcement learning scheduling optimization neural network, combine the objective functions of equipment load balancing, task delay minimization, and task switching cost minimization, configure the equipment OEE real-time monitoring unit and behavior attribution model, and realize dynamic coordinated adjustment of production scheduling and equipment status through feedback adjustment strategy.

Benefits of technology

It improves production scheduling execution efficiency and comprehensive equipment utilization, enhances the production system's ability to perceive and respond to operational bottlenecks, and realizes intelligent production scheduling closed-loop management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120525140B_ABST
    Figure CN120525140B_ABST
Patent Text Reader

Abstract

The present application provides a lean operation method based on production scheduling optimization and equipment OEE monitoring, which relates to the field of lean operation technology, including: deploying a multi-source heterogeneous data acquisition interface, and performing data acquisition to obtain customer order data, material supply status data, production task priority data, process timing requirement data and equipment maintenance plan data, performing data preprocessing after data acquisition, and generating a semantic label set, obtaining a semantic label set, configuring a production demand dynamic feature set based on the obtained semantic label set, and calculating the feature influence weight. By deploying a multi-source heterogeneous data acquisition interface, customer order data, material supply status data, production task priority data, process timing requirement data and equipment maintenance plan data are obtained at the same time, and a semantic label set and a production demand dynamic feature set are constructed according to the time series structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lean operations, and more particularly to a lean operations method based on production scheduling optimization and equipment OEE monitoring. Background Art

[0002] As the level of intelligence and digitalization in the manufacturing industry continues to improve, companies are increasingly demanding on production efficiency, resource utilization, and equipment availability. Production scheduling, as a key link affecting the effectiveness of the production system, is directly related to the quality of completion of production tasks and the ability to deliver them on time. At the same time, equipment OEE, as a core indicator for measuring equipment utilization efficiency, has been widely used in intelligent manufacturing systems to evaluate equipment availability, performance efficiency, and yield rate. How to perceive and adjust equipment OEE performance in real time during the production scheduling optimization process has become one of the core issues for manufacturing companies to achieve lean operations.

[0003] However, existing technologies generally have problems such as disconnection between production scheduling strategies and equipment operating status, and lack of equipment OEE behavior attribution mechanism. Most current production scheduling methods rely on static rules or offline data for scheduling decisions, and lack a linkage mechanism with real-time equipment OEE status. As a result, some high-load equipment continues to run, and inefficient equipment is not identified and adjusted in a timely manner, affecting the overall resource allocation efficiency. When the equipment OEE decreases, the existing system finds it difficult to trace the specific task scheduling path, scheduling strategy or equipment failure behavior that leads to the efficiency decline, making it difficult to achieve targeted optimization and closed-loop control. Summary of the Invention

[0004] The purpose of the present invention is to address the shortcomings of the existing technology and provide a lean operation method based on production scheduling optimization and equipment OEE monitoring. The method aims to realize intelligent mapping between tasks and equipment by constructing a reinforcement learning production scheduling optimization neural network, combining the equipment load balancing objective function, the task delay minimization objective function and the task switching cost minimization objective function. At the same time, the equipment OEE real-time monitoring unit and the equipment OEE behavior attribution model are configured to accurately identify the scheduling root causes of reduced equipment efficiency. By constructing an equipment OEE closed-loop feedback adjustment strategy, combining the feedback adjustment target vector and the adjustment strategy recommendation vector, the dynamic coordinated adjustment of the production scheduling strategy and the equipment operating status is realized, thereby improving the production scheduling execution efficiency and the comprehensive utilization rate of equipment, and enhancing the production system's perception and response capabilities to operational bottlenecks.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A lean operation method based on production scheduling optimization and equipment OEE monitoring includes the following steps:

[0007] Step S100: deploy a multi-source heterogeneous data acquisition interface and perform data acquisition to obtain customer order data, material supply status data, production task priority data, process timing requirement data, and equipment maintenance plan data; perform data preprocessing based on data acquisition, generate a semantic tag set, obtain a semantic tag set, configure a production demand dynamic feature set based on the obtained semantic tag set, and calculate the feature impact weight.

[0008] Step S200: Configure the reinforcement learning production scheduling optimization strategy environment based on the dynamic feature set of production demand, define the reinforcement learning reward function, configure the reinforcement learning production scheduling optimization neural network and the historical scheduling feedback sample set, and train the reinforcement learning production scheduling optimization neural network. Output the optimal production scheduling plan based on the reinforcement learning production scheduling optimization neural network to obtain the optimal production scheduling plan.

[0009] Step S300: Deploy the equipment OEE monitoring unit and collect equipment operating status data. Calculate the equipment OEE comprehensive score based on the equipment operating status data. Configure the OEE anomaly identification mechanism based on the equipment OEE comprehensive score and detect abnormal OEE fluctuation events. Configure the equipment OEE behavior attribution model based on the abnormal OEE fluctuation events and identify influencing factors. Output the equipment OEE abnormal fluctuation attribution analysis report through the equipment OEE behavior attribution model and mark the operation bottleneck.

[0010] Step S400: trigger the production scheduling strategy feedback adjustment process based on the abnormal fluctuation event of equipment OEE, configure the feedback adjustment target vector, generate the task scheduling action set and strategy confidence vector based on the reinforcement learning production scheduling optimization neural network, configure the adaptive feedback adjustment rules and update the historical scheduling feedback sample set.

[0011] Step S500: Configure an operation performance indicator scoring mechanism, configure model incremental training samples based on the operation performance indicator scoring mechanism, configure an incremental training mechanism for the production scheduling optimization neural network and the equipment OEE behavior attribution model, and update the historical sample library.

[0012] As a preferred solution of the present invention, step S100 is specifically as follows:

[0013] Step S100.1: Deploy a multi-source heterogeneous data acquisition interface and perform data acquisition to obtain customer order data, material supply status data, production task priority data, process timing requirement data, and equipment maintenance plan data.

[0014] Step S100.2: Perform data preprocessing based on data collection, generate a semantic tag set, and obtain the semantic tag set.

[0015] Based on the collected customer order data, material supply status data, production task priority data, process timing requirement data and equipment maintenance plan data, each data is structured and mapped according to a unified time series structure. The time series structure includes: a data source identification field, a data collection timestamp field, a data content field and a data validity field. The data collection timestamp field is unified using the ISO8601 standard, and the data content field is standardized using the International System of Units.

[0016] Based on the customer order data, material supply status data, production task priority data, process timing requirement data and equipment maintenance plan data of the time series structure, a preset semantic rule set is called to generate a corresponding semantic label set, wherein the semantic label set includes: order urgency level label, material tension level label, task priority level label, process complexity label and equipment availability level label.

[0017] When the delivery deadline of the customer order data is less than 24 hours from the current time and the order priority level is the highest level, the generated order urgency level label is: urgent order.

[0018] When the material inventory in the material supply status data is 20% lower than the minimum safety stock threshold and the supplier performance score is lower than 80 points, the material tension level label generated is: high-risk material.

[0019] Step S100.3: Configure a production demand dynamic feature set based on the obtained semantic tag set, and calculate the feature impact weight.

[0020] Based on the customer order data, material supply status data, production task priority data, process timing requirement data, equipment maintenance plan data and semantic label set of the time series structure, a production demand dynamic feature set is configured. The production demand dynamic feature set is a multidimensional vector structure, specifically including: numerical feature dimension, semantic label dimension, time weight dimension and data credibility dimension.

[0021] Based on the dynamic feature set of production demand, a data fusion scoring mechanism is executed for customer order data feature items, material supply status data feature items, production task priority data feature items, process timing requirement data feature items and equipment maintenance plan data feature items. The data fusion scoring mechanism includes: scheduling impact factor, resource conflict factor and response time factor. The data fusion scoring mechanism is calculated by the scoring value = 0.5×scheduling impact factor + 0.3×resource conflict factor + 0.2×response time factor, wherein the scheduling impact factor is calculated based on the task priority and the order delivery deadline, the resource conflict factor is calculated based on the material inventory tightness and the equipment conflict probability, and the response time factor is determined based on the time interval between the data collection timestamp and the current system. When the interval exceeds 60 minutes, the score of the corresponding feature item is automatically reduced by 20%.

[0022] As a preferred solution of the present invention, step S200 is specifically as follows:

[0023] Step S200.1: Configure a reinforcement learning production scheduling optimization strategy environment based on a dynamic feature set of production demand, and define a reinforcement learning reward function.

[0024] Based on the dynamic feature set of production demand and the operating status of the equipment, a reinforcement learning production scheduling optimization strategy environment is configured. The reinforcement learning production scheduling optimization strategy environment includes: a state space, an action space, a reward function and a strategy function.

[0025] The state space consists of a set of dynamic features of production demand and the operating status of the equipment. The dynamic feature set of production demand includes: numerical feature dimension, semantic label dimension, time weight dimension and data credibility dimension. The operating status of the equipment includes: equipment number, current utilization rate, equipment availability level and historical task execution efficiency. The action space is defined as: the scheduling mapping relationship from task to equipment, that is, the allocation mapping set between task number and equipment number. The reward function is to evaluate the scheduling quality after the execution of the scheduling optimization strategy. The strategy function is to select the corresponding scheduling action based on the state space.

[0026] Call the device load balancing objective function, task delay minimization objective function, and task switching cost minimization objective function to configure the reinforcement learning reward function.

[0027] The reinforcement learning reward function is the weighted sum of the device load balancing objective function, the task delay minimization objective function, and the task switching cost minimization objective function.

[0028] Step S200.2: Configure the reinforcement learning scheduling optimization neural network and the historical scheduling feedback sample set, and train the reinforcement learning scheduling optimization neural network.

[0029] A deep deterministic policy gradient algorithm is called to configure a reinforcement learning scheduling optimization neural network, where the reinforcement learning scheduling optimization neural network includes a policy network and a value network.

[0030] The input of the policy network is: state space data, and the output is: the selection probability of each scheduling action in the action space. The policy network adopts a three-layer fully connected network structure, including: input layer, two hidden layers and output layer, and the hidden layer activation function adopts the ReLU function.

[0031] The input of the value network is the concatenation of state space data and action space data, and the output is the value function Q value of the corresponding state and action pair. The value network adopts the structure of a double-layer convolutional layer plus a fully connected layer.

[0032] A historical scheduling feedback sample set is configured, wherein the historical scheduling feedback sample set is composed of: a dynamic feature set of production demand in a historical period, equipment operating status, actual task completion time, task scheduling path and corresponding scheduling score data.

[0033] The reinforcement learning scheduling optimization neural network is trained through the historical scheduling feedback sample set. The experience replay mechanism is used to randomly extract data samples from the historical scheduling feedback sample set. The batch size is 64. The soft update mechanism is used to update the target policy network parameters and the target value network parameters. The policy network learning rate is set to 1×10⁻ 4 , the value network learning rate is set to 1×10⁻³.

[0034] Step S200.3: Output the optimal production scheduling plan based on the reinforcement learning production scheduling optimization neural network to obtain the optimal production scheduling plan.

[0035] After the reinforcement learning scheduling optimization neural network is trained, the dynamic feature set of production demand and the real-time equipment operation status are input into the strategy network, and the task-to-equipment scheduling action set corresponding to the current state is output. The task-to-equipment scheduling action set is a matrix structure, in which each matrix element represents a specific allocation decision from a task number to an equipment number.

[0036] Based on the output task-to-equipment production scheduling action set, the optimal production scheduling plan is formed, and the optimal production scheduling plan is submitted to the scheduling execution module to drive the equipment to execute the production task according to the generated optimal production scheduling plan, record the equipment operation data and task execution data of this production schedule, and update it to the historical scheduling feedback sample set.

[0037] As a preferred solution of the present invention, step S300 is specifically as follows:

[0038] Step S300.1: Deploy an equipment OEE monitoring unit, collect equipment operating status data, and calculate the equipment OEE comprehensive score based on the equipment operating status data.

[0039] Step S300.2: Configure an OEE anomaly identification mechanism based on the equipment OEE comprehensive score and detect abnormal OEE fluctuation events.

[0040] Step S300.3: Configure the equipment OEE behavior attribution model based on the OEE abnormal fluctuation event, identify the influencing factors, output the equipment OEE abnormal fluctuation attribution analysis report through the equipment OEE behavior attribution model, and mark the operation bottleneck.

[0041] Based on the detected abnormal OEE fluctuation events, an equipment OEE behavior attribution model is configured and used to analyze the abnormal equipment OEE fluctuation events. The equipment OEE behavior attribution model is based on the task scheduling path, task number, equipment operating status data, scheduling score data stored in the historical scheduling feedback sample set, and the equipment operating status data collected in the current cycle. The attribution steps include four steps:

[0042] The first attribution step includes: establishing a scheduling behavior event table to record the task number, scheduling time, execution device number and reinforcement learning reward function score.

[0043] The second attribution step includes: establishing a sequence of equipment status changes to record changes in equipment utilization rate, equipment performance efficiency, and equipment yield over time.

[0044] The third attribution step includes: for each abnormal OEE fluctuation event of the equipment, tracing back the scheduling behavior event table and equipment status change sequence within 30 minutes to identify whether there are task changes, equipment maintenance plan adjustments and frequent equipment task switching.

[0045] The fourth attribution step includes: if the above scheduling behavior exists, then the scheduling behavior is marked as a potential cause of the abnormal fluctuation event of equipment OEE.

[0046] The scoring mechanism of the equipment OEE behavior attribution model is used to quantitatively score potential causes.

[0047] After completing the attribution analysis based on the equipment OEE behavior attribution model, an equipment OEE abnormal fluctuation attribution analysis report is output. The equipment OEE abnormal fluctuation attribution analysis report includes: the downward trend of the equipment OEE comprehensive score, the duration of the abnormal fluctuation, the equipment number where the abnormal fluctuation occurred, the task number involved, the specific scheduling behavior type that induced the abnormal fluctuation, and the corresponding inducement score.

[0048] As a preferred solution of the present invention, step S400 is specifically as follows:

[0049] Step S400.1: Trigger the production scheduling strategy feedback adjustment process based on the equipment OEE abnormal fluctuation event, and configure the feedback adjustment target vector.

[0050] Step S400.2: Generate a task scheduling action set and a strategy confidence vector based on the reinforcement learning scheduling optimization neural network.

[0051] Step S400.3: Configure adaptive feedback adjustment rules and update historical scheduling feedback sample sets.

[0052] Based on the feedback adjustment target vector and the task scheduling strategy confidence vector, an adaptive feedback adjustment rule is configured. The adaptive feedback adjustment rule includes three rules:

[0053] The first rule in the adaptive feedback adjustment rules is: if the incentive score value in the feedback adjustment target vector is greater than 80 points, and the confidence of the corresponding task number in the task scheduling strategy confidence vector is less than 0.6, then the task transfer strategy is executed, and the task with the corresponding task number is preferentially transferred to the production equipment with the lowest load rate and the highest equipment availability level in the current cycle.

[0054] The second rule in the adaptive feedback adjustment rule is: if the scheduling behavior type that induces abnormal fluctuation events in the equipment OEE in the feedback adjustment target vector is frequent task switching, then based on the current continuous task sequence on the corresponding equipment number in the task scheduling action set, a task reordering operation is performed to reduce the number of continuous task switches on the equipment.

[0055] The third rule in the adaptive feedback adjustment rules is: If the scheduling behavior type that induces the abnormal fluctuation event of equipment OEE in the feedback adjustment target vector is planned maintenance delay, then freeze the new task assignment of the corresponding equipment number within 60 minutes of the next complete production cycle, and immediately reserve a time window for the equipment to perform maintenance operations.

[0056] The modified task-to-equipment production scheduling action set is generated through adaptive feedback adjustment rules.

[0057] Based on the generated revised task-to-equipment production scheduling action set, the revised task-to-equipment production scheduling action set is submitted to the scheduling execution module, driving the production equipment to execute the production tasks according to the revised task-to-equipment production scheduling action set, and recording the relevant data that triggers the scheduling strategy feedback adjustment process. At the same time, it is added and updated to the historical scheduling feedback sample set.

[0058] As a preferred solution of the present invention, step S500 is specifically as follows:

[0059] Step S500.1: Configure an operation performance indicator scoring mechanism, and configure model incremental training samples based on the operation performance indicator scoring mechanism.

[0060] Step S500.2: Configure the production scheduling optimization neural network and the equipment OEE behavior attribution model incremental training mechanism, and update the historical sample library.

[0061] Based on the incremental training samples of the model, incremental training mechanisms of the reinforcement learning scheduling optimization neural network and the equipment OEE behavior attribution model are configured respectively. The incremental training mechanism of the reinforcement learning scheduling optimization neural network includes: based on the dynamic feature set of production demand, equipment operation status data and task-to-equipment scheduling action set in the incremental training samples of the model configured in the current cycle, a deep deterministic policy gradient algorithm is used to update the network parameters, and the parameter update method continues to use the soft update mechanism.

[0062] The incremental training mechanism of the equipment OEE behavior attribution model includes: recalculating the weight factor of the incentive scoring mechanism in the equipment OEE behavior attribution model based on the equipment operation status data, equipment OEE abnormal fluctuation attribution analysis report and overall operation performance score in the model incremental training samples. 、 、 .

[0063] Execute the incremental training process of the reinforcement learning scheduling optimization neural network and the equipment OEE behavior attribution model. Based on the incremental model training samples of each production cycle, update the weight factors of the incentive scoring mechanism in the policy network, value network and equipment OEE behavior attribution model of the reinforcement learning scheduling optimization neural network respectively. When the overall operational performance score is lower than the preset performance benchmark threshold of 0.85, the incremental model training samples generated in the current cycle are marked as negative samples. If the overall operational performance scores of two consecutive cycles are higher than 0.90, the incremental model training samples generated in the corresponding production cycle are marked as positive samples, and the historical scheduling feedback sample library is updated. According to the experience replay mechanism, data samples with a batch size of 64 are randomly selected for subsequent training.

[0064] Compared with the prior art, the present invention has the following beneficial effects:

[0065] 1. By deploying multi-source heterogeneous data acquisition interfaces, the system can simultaneously obtain customer order data, material supply status data, production task priority data, process timing requirement data, and equipment maintenance plan data, and build a semantic label set and a dynamic feature set of production demand based on the time series structure. This allows production scheduling optimization decisions to be based on multi-dimensional, real-time, and semantically clear inputs, significantly improving the accuracy of production scheduling demand modeling and the ability to collaborate with upstream and downstream resources.

[0066] 2. By constructing a reinforcement learning scheduling optimization neural network and combining the equipment load balancing objective function, task delay minimization objective function and task switching cost minimization objective function, a reward function is formed to guide strategy learning, which effectively achieves a dynamic balance between resource utilization, delivery timeliness and production efficiency in the scheduling plan, and significantly enhances the system's flexibility and adaptability in dealing with complex scheduling scenarios.

[0067] 3. Build a real-time equipment OEE monitoring unit based on the edge computing architecture, and use the OEE behavior attribution model to associate changes in equipment utilization rate, performance efficiency, and yield rate with factors such as production scheduling and equipment operating status. This enables accurate identification and explainable analysis of the causes of equipment's overall performance decline, and enhances the manufacturing system's ability to autonomously discover and respond to operational bottlenecks.

[0068] 4. The production scheduling path, equipment operating status and task completion data are recorded through the historical scheduling feedback sample set, and the reinforcement learning production scheduling optimization neural network is continuously trained in combination with the experience replay and soft update mechanism, so that the model can continuously evolve and optimize itself in actual applications, thereby realizing an intelligent production scheduling closed-loop management path from data-driven to strategy evolution, significantly improving the long-term robustness and iterative optimization capabilities of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 A flowchart of a lean operation method based on production scheduling optimization and equipment OEE monitoring is provided in an embodiment of the present application. DETAILED DESCRIPTION

[0070] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0071] See also Figure 1 , Figure 1 A flowchart of a lean operation method based on production scheduling optimization and equipment OEE monitoring is provided for an embodiment of the present application.

[0072] In this embodiment, a lean operation method based on production scheduling optimization and equipment OEE monitoring may include step S100, step S200, step S300, step S400 and step S500;

[0073] Step S100: deploy a multi-source heterogeneous data acquisition interface and perform data acquisition to obtain customer order data, material supply status data, production task priority data, process timing requirement data, and equipment maintenance plan data; perform data preprocessing based on data acquisition, generate a semantic tag set, obtain a semantic tag set, configure a production demand dynamic feature set based on the obtained semantic tag set, and calculate the feature impact weight.

[0074] Step S200: Configure the reinforcement learning production scheduling optimization strategy environment based on the dynamic feature set of production demand, define the reinforcement learning reward function, configure the reinforcement learning production scheduling optimization neural network and the historical scheduling feedback sample set, and train the reinforcement learning production scheduling optimization neural network. Output the optimal production scheduling plan based on the reinforcement learning production scheduling optimization neural network to obtain the optimal production scheduling plan.

[0075] Step S300: Deploy the equipment OEE monitoring unit and collect equipment operating status data. Calculate the equipment OEE comprehensive score based on the equipment operating status data. Configure the OEE anomaly identification mechanism based on the equipment OEE comprehensive score and detect abnormal OEE fluctuation events. Configure the equipment OEE behavior attribution model based on the abnormal OEE fluctuation events and identify influencing factors. Output the equipment OEE abnormal fluctuation attribution analysis report through the equipment OEE behavior attribution model and mark the operation bottleneck.

[0076] Step S400: trigger the production scheduling strategy feedback adjustment process based on the abnormal fluctuation event of equipment OEE, configure the feedback adjustment target vector, generate the task scheduling action set and strategy confidence vector based on the reinforcement learning production scheduling optimization neural network, configure the adaptive feedback adjustment rules and update the historical scheduling feedback sample set.

[0077] Step S500: Configure an operation performance indicator scoring mechanism, configure model incremental training samples based on the operation performance indicator scoring mechanism, configure an incremental training mechanism for the production scheduling optimization neural network and the equipment OEE behavior attribution model, and update the historical sample library.

[0078] In some specific implementations, the step S100 is specifically:

[0079] Step S100.1: Deploy a multi-source heterogeneous data acquisition interface and perform data acquisition to obtain customer order data, material supply status data, production task priority data, process timing requirement data, and equipment maintenance plan data.

[0080] Before implementing the lean operation method, a data acquisition interface for multi-source heterogeneous data collection is first deployed. The data acquisition interface is connected to the customer order management system, material supply chain management system, production task scheduling system, process flow control system and equipment maintenance management system respectively, and collects customer order data, material supply status data, production task priority data, process timing requirement data and equipment maintenance plan data in real time.

[0081] The data acquisition interface collects and obtains customer order data, material supply status data, production task priority data, process timing requirement data and equipment maintenance plan data from the customer order management system, material supply chain management system, production task scheduling system, process flow control system and equipment maintenance management system respectively.

[0082] The data collection interface collects customer order data from the customer order management system in real time, wherein the customer order data includes: order number, order time, delivery deadline, order priority level, order product type and order quantity.

[0083] The data collection interface collects material supply status data from the material supply chain management system in real time, wherein the material supply status data includes: material number, material inventory, expected arrival time, supplier performance score and material supply stability rating.

[0084] The data collection interface collects production task priority data from the production task scheduling system in real time, wherein the production task priority data includes: task number, task priority level, corresponding process flow number and task scheduling time limit.

[0085] The data acquisition interface collects process timing requirement data from the process control system in real time, wherein the process timing requirement data includes: process number, process duration, predecessor dependent process number and equipment matching level.

[0086] The data collection interface collects equipment maintenance plan data from the equipment maintenance management system in real time, wherein the equipment maintenance plan data includes: equipment number, planned maintenance start time, planned maintenance end time, maintenance level classification and equipment utilization rate estimation.

[0087] Step S100.2: Perform data preprocessing based on data collection, generate a semantic tag set, and obtain the semantic tag set.

[0088] Based on the collected customer order data, material supply status data, production task priority data, process timing requirement data and equipment maintenance plan data, each data is structured and mapped according to a unified time series structure. The time series structure includes: a data source identification field, a data collection timestamp field, a data content field and a data validity field. The data collection timestamp field is unified using the ISO8601 standard, and the data content field is standardized using the International System of Units.

[0089] Based on the customer order data, material supply status data, production task priority data, process timing requirement data and equipment maintenance plan data of the time series structure, a preset semantic rule set is called to generate a corresponding semantic label set, wherein the semantic label set includes: order urgency level label, material tension level label, task priority level label, process complexity label and equipment availability level label.

[0090] When the delivery deadline of the customer order data is less than 24 hours from the current time and the order priority level is the highest level, the generated order urgency level label is: urgent order.

[0091] When the material inventory in the material supply status data is 20% lower than the minimum safety stock threshold and the supplier performance score is lower than 80 points, the material tension level label generated is: high-risk material.

[0092] Step S100.3: Configure a production demand dynamic feature set based on the obtained semantic tag set, and calculate the feature impact weight.

[0093] Based on the customer order data, material supply status data, production task priority data, process timing requirement data, equipment maintenance plan data and semantic label set of the time series structure, a production demand dynamic feature set is configured. The production demand dynamic feature set is a multidimensional vector structure, specifically including: numerical feature dimension, semantic label dimension, time weight dimension and data credibility dimension.

[0094] The numerical feature dimension is standardized to the interval [0,1] through minimum or maximum normalization, the semantic label dimension is converted into a numerical value using one-hot encoding, the time weight dimension is mapped to a time decay weight according to the interval between the data collection timestamp and the current system time, and the data credibility dimension is determined according to the preset stability score of the data source system, where the credibility score of the customer order management system is 0.95 and the credibility score of the material supply chain management system is 0.90.

[0095] Based on the dynamic feature set of production demand, a data fusion scoring mechanism is executed for customer order data feature items, material supply status data feature items, production task priority data feature items, process timing requirement data feature items and equipment maintenance plan data feature items. The data fusion scoring mechanism includes: scheduling impact factor, resource conflict factor and response time factor. The data fusion scoring mechanism is calculated by the scoring value = 0.5×scheduling impact factor + 0.3×resource conflict factor + 0.2×response time factor, wherein the scheduling impact factor is calculated based on the task priority and the order delivery deadline, the resource conflict factor is calculated based on the material inventory tightness and the equipment conflict probability, and the response time factor is determined based on the time interval between the data collection timestamp and the current system. When the interval exceeds 60 minutes, the score of the corresponding feature item is automatically reduced by 20%.

[0096] In some specific implementations, the step S200 is specifically:

[0097] Step S200.1: Configure a reinforcement learning production scheduling optimization strategy environment based on a dynamic feature set of production demand, and define a reinforcement learning reward function.

[0098] Based on the dynamic feature set of production demand and the operating status of the equipment, a reinforcement learning production scheduling optimization strategy environment is configured. The reinforcement learning production scheduling optimization strategy environment includes: a state space, an action space, a reward function and a strategy function.

[0099] The state space consists of a set of dynamic features of production demand and the operating status of the equipment. The dynamic feature set of production demand includes: numerical feature dimension, semantic label dimension, time weight dimension and data credibility dimension. The operating status of the equipment includes: equipment number, current utilization rate, equipment availability level and historical task execution efficiency. The action space is defined as: the scheduling mapping relationship from task to equipment, that is, the allocation mapping set between task number and equipment number. The reward function is to evaluate the scheduling quality after the execution of the scheduling optimization strategy. The strategy function is to select the corresponding scheduling action based on the state space.

[0100] Call the device load balancing objective function, task delay minimization objective function, and task switching cost minimization objective function to configure the reinforcement learning reward function.

[0101] The device load balancing objective function is as follows:

[0102]

[0103] Where: It is the output value of the device load balancing objective function, which is used to evaluate the uniformity of device load under the current task allocation. is the set of equipment load rates The standard deviation of the load is used to measure the degree of distribution between devices. is the set of equipment load rates The average value is used to indicate the average load rate of all current devices. is the set of task load rates of all devices, where It is The task load rate of each device.

[0104] The objective function for minimizing task delay is:

[0105]

[0106] Where: It is the output value of the objective function of minimizing task delay, which is used to evaluate the overall degree of task delay. It is The actual completion time of each task, It is The delivery deadline of each task, It's a task The length of delay (if , then this item is 0), is the total number of tasks in the current cycle.

[0107] The objective function for minimizing task switching cost is:

[0108]

[0109] Where: It is the output value of the objective function of minimizing the task switching cost, which is used to evaluate the switching time cost between continuous tasks of the device. It is a device In executing Task and The task switching time between tasks, is the total number of devices, It is a device The current number of assigned tasks.

[0110] The reinforcement learning reward function is the weighted sum of the device load balancing objective function, the task delay minimization objective function, and the task switching cost minimization objective function. Specifically,

[0111]

[0112] Where: is the reward score of the current production scheduling plan, which is used to feed back to the reinforcement learning strategy. is the output value of the device load balancing objective function, is the output value of the objective function for minimizing task delay, is the output value of the objective function minimizing the task switching cost, and the coefficient 、 、 They are weighted factors of each optimization objective, which are used to adjust the contribution ratio of the objectives, and the sum of the weights is 1.

[0113] Step S200.2: Configure the reinforcement learning scheduling optimization neural network and the historical scheduling feedback sample set, and train the reinforcement learning scheduling optimization neural network.

[0114] A deep deterministic policy gradient algorithm is called to configure a reinforcement learning scheduling optimization neural network, where the reinforcement learning scheduling optimization neural network includes a policy network and a value network.

[0115] The input of the policy network is: state space data, and the output is: the selection probability of each scheduling action in the action space. The policy network adopts a three-layer fully connected network structure, including: input layer, two hidden layers and output layer, and the hidden layer activation function adopts the ReLU function.

[0116] The input of the value network is the concatenation of state space data and action space data, and the output is the value function Q value of the corresponding state and action pair. The value network adopts a double-layer convolutional layer plus a fully connected layer structure. The loss function of the value network is defined as:

[0117]

[0118]

[0119] Where: It is the loss function of the reinforcement learning value network, which is used to measure the accuracy of the current Q value estimation. is the number of samples in the training batch, Is the current value network in state and actions The estimated Q value under , is the target Q value, used to train the supervision signal, is the mean square error of each sample, It is a sample The target Q value, It is a sample Instant rewards, is the discount factor, which is set to 0.99 and is used to control the degree of decay of future rewards. In the next state Next, according to the current strategy The action selected, the estimated Q value, are the parameters of the target value network.

[0120] A historical scheduling feedback sample set is configured, wherein the historical scheduling feedback sample set is composed of: a dynamic feature set of production demand in a historical period, equipment operating status, actual task completion time, task scheduling path and corresponding scheduling score data.

[0121] The reinforcement learning scheduling optimization neural network is trained through the historical scheduling feedback sample set. The experience replay mechanism is used to randomly extract data samples from the historical scheduling feedback sample set. The batch size is 64. The soft update mechanism is used to update the target policy network parameters and the target value network parameters. The policy network learning rate is set to 1×10⁻ 4 , the value network learning rate is set to 1×10⁻³, and the target network parameters are updated as follows:

[0122]

[0123] Where: is the parameter vector of the target network (after update), is the parameter vector of the current main network, is the soft update factor, which indicates the update ratio and is set to 0.005.

[0124] Step S200.3: Output the optimal production scheduling plan based on the reinforcement learning production scheduling optimization neural network to obtain the optimal production scheduling plan.

[0125] After the reinforcement learning scheduling optimization neural network is trained, the dynamic feature set of production demand and the real-time equipment operation status are input into the strategy network, and the task-to-equipment scheduling action set corresponding to the current state is output. The task-to-equipment scheduling action set is a matrix structure, in which each matrix element represents a specific allocation decision from a task number to an equipment number.

[0126] Based on the output task-to-equipment production scheduling action set, the optimal production scheduling plan is formed, and the optimal production scheduling plan is submitted to the scheduling execution module to drive the equipment to execute the production task according to the generated optimal production scheduling plan, record the equipment operation data and task execution data of this production schedule, and update it to the historical scheduling feedback sample set.

[0127] In some specific implementations, step S300 is specifically:

[0128] Step S300.1: Deploy an equipment OEE monitoring unit, collect equipment operating status data, and calculate the equipment OEE comprehensive score based on the equipment operating status data.

[0129] Deploy an equipment OEE monitoring unit, which includes: a sensor acquisition module, an edge computing processing module, and a data transmission module.

[0130] The sensor acquisition module is used to collect equipment operation status data in real time, wherein the equipment operation status data includes: equipment utilization rate, equipment performance efficiency and equipment yield rate.

[0131] The equipment utilization rate is the ratio of the actual operating time of the equipment to the planned available time. The equipment performance efficiency is the ratio of the standard output quantity to the actual output quantity during the actual operating time. The equipment yield rate is the ratio of the number of qualified products to the actual output quantity.

[0132] The equipment OEE monitoring unit collects equipment operating status data once a minute and performs local preprocessing through the edge computing processing module. The preprocessing content includes: outlier filtering, sliding average smoothing and timestamp normalization.

[0133] Based on the equipment operating status data collected and pre-processed by the equipment OEE monitoring unit, the equipment OEE comprehensive score calculation formula is used to perform a weighted calculation of the equipment utilization rate, equipment performance efficiency, and equipment yield rate. The specific calculation is:

[0134]

[0135] Where: It is the comprehensive score of equipment OEE, which is used to comprehensively evaluate the operating efficiency of the equipment in the current cycle. The equipment utilization rate is the ratio of the actual operating time of the equipment to the planned available time. is the equipment performance efficiency, which is the ratio of the standard output quantity to the output quantity during the actual operating time. It is the equipment yield rate, which is the ratio of the number of qualified products to the actual output quantity. is the weighted coefficient of equipment utilization rate, set to 0.4, is the weighted coefficient of equipment performance efficiency, set to 0.3, It is the weighted coefficient of equipment yield, which is set to 0.3.

[0136] Step S300.2: Configure an OEE anomaly identification mechanism based on the equipment OEE comprehensive score and detect abnormal OEE fluctuation events.

[0137] Based on the calculated equipment OEE comprehensive score, an OEE anomaly identification mechanism is configured. The OEE anomaly identification mechanism adopts a sliding window trend detection algorithm and a threshold judgment strategy. The detection process is as follows: the historical reference window length is set to minutes, and the real-time monitoring window length is The average comprehensive OEE scores of the equipment in the real-time monitoring window and the historical reference window are calculated as follows:

[0138]

[0139] Where: It is the change in the comprehensive score of equipment OEE, which is used to identify the OEE fluctuation trend. It is the time length of the current real-time monitoring window, in minutes, and is set to 5 minutes. The length of the historical reference window, in minutes, is set to 60 minutes. It is the first The comprehensive OEE score of the equipment in the minute, The first The comprehensive OEE score of the equipment in the minute, is the sum of the OEE values ​​for all minutes in the current window, It is the sum of the OEE values ​​for all minutes within the reference window.

[0140] Setting anomaly recognition thresholds When the equipment OEE comprehensive score changes by 5%, Less than the negative anomaly recognition threshold If the fluctuation lasts for more than three sampling periods, it is considered an abnormal OEE fluctuation event.

[0141] Step S300.3: Configure the equipment OEE behavior attribution model based on the OEE abnormal fluctuation event, identify the influencing factors, output the equipment OEE abnormal fluctuation attribution analysis report through the equipment OEE behavior attribution model, and mark the operation bottleneck.

[0142] Based on the detected abnormal OEE fluctuation events, an equipment OEE behavior attribution model is configured and used to analyze the abnormal equipment OEE fluctuation events. The equipment OEE behavior attribution model is based on the task scheduling path, task number, equipment operating status data, scheduling score data stored in the historical scheduling feedback sample set, and the equipment operating status data collected in the current cycle. The attribution steps include four steps:

[0143] The first attribution step includes: establishing a scheduling behavior event table to record the task number, scheduling time, execution device number and reinforcement learning reward function score.

[0144] The second attribution step includes: establishing a sequence of equipment status changes to record changes in equipment utilization rate, equipment performance efficiency, and equipment yield over time.

[0145] The third attribution step includes: for each abnormal OEE fluctuation event of the equipment, tracing back the scheduling behavior event table and equipment status change sequence within 30 minutes to identify whether there are task changes, equipment maintenance plan adjustments and frequent equipment task switching.

[0146] The fourth attribution step includes: if the above scheduling behavior exists, then the scheduling behavior is marked as a potential cause of the abnormal fluctuation event of equipment OEE.

[0147] The scoring mechanism of the equipment OEE behavior attribution model is used to quantify the potential causes. The score is calculated as follows:

[0148]

[0149] Where: It is the cause score value of abnormal OEE fluctuation, which is used to quantify the impact of different causes. It is the change in task load, which is the change in the total workload of the device before and after a certain scheduling behavior. is the increase in the number of device task switches, is the change in the number of device task switches within a certain time window, It is the delay time of the equipment planned maintenance, which is the time difference between the original maintenance start time and the actual maintenance start time. is the weight factor of the task load change, set to 0.5, is the weight factor for the increase in the number of device task switching times, set to 0.3. is the weight factor for equipment planned maintenance delay time, which is set to 0.2.

[0150] After completing the attribution analysis based on the equipment OEE behavior attribution model, an equipment OEE abnormal fluctuation attribution analysis report is output. The equipment OEE abnormal fluctuation attribution analysis report includes: the downward trend of the equipment OEE comprehensive score, the duration of the abnormal fluctuation, the equipment number where the abnormal fluctuation occurred, the task number involved, the specific scheduling behavior type that induced the abnormal fluctuation, and the corresponding inducement score.

[0151] When the comprehensive OEE score of a certain equipment decreases due to the same inducement in two consecutive scheduling cycles, and the inducement score exceeds 75 points, the equipment number and inducement type will be recorded in the operation bottleneck record library and marked as an operation bottleneck.

[0152] In some specific implementations, the step S400 is specifically:

[0153] Step S400.1: Trigger the production scheduling strategy feedback adjustment process based on the equipment OEE abnormal fluctuation event, and configure the feedback adjustment target vector.

[0154] Based on the equipment OEE comprehensive score output in real time by the equipment OEE monitoring unit, the changes in the equipment OEE comprehensive score are continuously monitored through the equipment OEE anomaly identification mechanism. When the change in the equipment OEE comprehensive score is less than the preset negative anomaly identification threshold of 5%, and the equipment OEE behavior attribution model determines that the OEE abnormal fluctuation event is caused by a specific scheduling behavior, and the corresponding incentive score exceeds 75 points, the production scheduling strategy feedback adjustment process is triggered.

[0155] Based on the equipment OEE abnormal fluctuation attribution analysis report output by the equipment OEE behavior attribution model, the equipment number, task number, specific scheduling behavior type that induced the abnormal fluctuation, inducement score value, and abnormal fluctuation event occurrence time are obtained, and the feedback adjustment target vector is configured. The feedback adjustment target vector is specifically defined as:

[0156]

[0157] Where: It is the feedback adjustment target vector, which is used to identify the five-element information structure set of each production scheduling behavior to be corrected in the OEE closed-loop feedback adjustment strategy. The device number where the abnormal fluctuation of the equipment OEE comprehensive score occurred. It is the task number being executed when the abnormal fluctuation event of the equipment OEE comprehensive score occurs. It is the abnormal attribution type output by the equipment OEE behavior attribution model, identifying the scheduling behavior category that causes fluctuations in the equipment OEE comprehensive score. It is the attribution score calculated by the equipment OEE behavior attribution model, which indicates the degree of influence of scheduling behavior on the decline of equipment OEE comprehensive score. The higher the value, the more significant the influence. The time when the abnormal fluctuation event of the equipment's OEE comprehensive score occurred, in seconds or timestamp format.

[0158] Step S400.2: Generate a task scheduling action set and a strategy confidence vector based on the reinforcement learning scheduling optimization neural network.

[0159] Based on the reinforcement learning scheduling optimization neural network, the dynamic feature set of production demand in the current cycle and the real-time collected equipment operation status data are input to obtain the task-to-equipment scheduling action set and output the task scheduling strategy confidence vector. The task scheduling strategy confidence vector is specifically:

[0160]

[0161] Where: It is the confidence set of the scheduling strategy for all tasks in the current task scheduling path. It is numbered The confidence score of the scheduling strategy of a task is in the range of , where 1 represents full trust in the production scheduling strategy, and 0 represents complete distrust. It is the total number of tasks in the current task scheduling path.

[0162] Step S400.3: Configure adaptive feedback adjustment rules and update historical scheduling feedback sample sets.

[0163] Based on the feedback adjustment target vector and the task scheduling strategy confidence vector, an adaptive feedback adjustment rule is configured. The adaptive feedback adjustment rule includes three rules:

[0164] The first rule in the adaptive feedback adjustment rules is: if the incentive score value in the feedback adjustment target vector is greater than 80 points, and the confidence of the corresponding task number in the task scheduling strategy confidence vector is less than 0.6, then the task transfer strategy is executed, and the task with the corresponding task number is preferentially transferred to the production equipment with the lowest load rate and the highest equipment availability level in the current cycle.

[0165] The second rule in the adaptive feedback adjustment rule is: if the scheduling behavior type that induces abnormal fluctuation events in the equipment OEE in the feedback adjustment target vector is frequent task switching, then based on the current continuous task sequence on the corresponding equipment number in the task scheduling action set, a task reordering operation is performed to reduce the number of continuous task switches on the equipment.

[0166] The third rule in the adaptive feedback adjustment rules is: If the scheduling behavior type that induces the abnormal fluctuation event of equipment OEE in the feedback adjustment target vector is planned maintenance delay, then freeze the new task assignment of the corresponding equipment number within 60 minutes of the next complete production cycle, and immediately reserve a time window for the equipment to perform maintenance operations.

[0167] The modified task-to-equipment production scheduling action set is generated through adaptive feedback adjustment rules.

[0168] Based on the generated revised task-to-equipment production scheduling action set, the revised task-to-equipment production scheduling action set is submitted to the scheduling execution module, driving the production equipment to execute the production tasks according to the revised task-to-equipment production scheduling action set, and recording the relevant data that triggers the scheduling strategy feedback adjustment process. At the same time, it is added and updated to the historical scheduling feedback sample set.

[0169] In some specific implementations, the step S500 is specifically:

[0170] Step S500.1: Configure an operation performance indicator scoring mechanism, and configure model incremental training samples based on the operation performance indicator scoring mechanism.

[0171] An operation performance indicator scoring mechanism is configured based on the equipment OEE comprehensive score, the output value of the task delay minimization objective function, the output value of the equipment load balancing objective function, and the output value of the task switching cost minimization objective function. The operation performance indicator scoring mechanism is calculated in units of production cycle, specifically:

[0172]

[0173] Where: is the overall operational performance score for the current production cycle, It is the average of the comprehensive OEE scores of all equipment in the current production cycle. is the average value of the output value of the equipment load balancing objective function in the current production cycle, is the average value of the output value of the objective function for minimizing task delay in the current production cycle, is the average value of the output value of the objective function for minimizing the task switching cost in the current production cycle, 、 、 、 are weight factors, and their values ​​are set as: , the sum of the weight factors is 1.

[0174] Based on the operational performance indicator scoring mechanism, the overall operational performance score of the current production cycle is marked, and model incremental training samples are configured. The model incremental training samples specifically include: a set of dynamic production demand features in the current production cycle, equipment operating status data, a set of task-to-equipment scheduling actions, an analysis report on the attribution of abnormal fluctuations in equipment OEE, a confidence vector of the task scheduling strategy, and an overall operational performance score.

[0175] The overall operational performance score is compared with the preset performance benchmark threshold, which is 0.85. When the overall operational performance score is lower than the performance benchmark threshold, the production cycle is marked as a low-performance cycle, and the equipment number, task number, specific scheduling behavior type that caused the performance decline, and the inducement score that were recorded in detail.

[0176] Step S500.2: Configure the production scheduling optimization neural network and the equipment OEE behavior attribution model incremental training mechanism, and update the historical sample library.

[0177] Based on the incremental training samples of the model, incremental training mechanisms for the reinforcement learning scheduling optimization neural network and the equipment OEE behavior attribution model are configured respectively. The incremental training mechanism of the reinforcement learning scheduling optimization neural network includes: based on the production demand dynamic feature set, equipment operation status data, and task-to-equipment scheduling action set in the incremental training samples of the model configured in the current cycle, a deep deterministic policy gradient algorithm is used to update the network parameters. The parameter update method continues to use the soft update mechanism, specifically:

[0178]

[0179] Where: is the updated target network parameter vector, is the parameter vector of the current network, is the soft update factor, and its value is set to 0.005.

[0180] The incremental training mechanism of the equipment OEE behavior attribution model includes: recalculating the weight factor of the incentive scoring mechanism in the equipment OEE behavior attribution model based on the equipment operation status data, equipment OEE abnormal fluctuation attribution analysis report and overall operation performance score in the model incremental training samples. 、 、 , the weight update formula is specifically defined as:

[0181]

[0182] Where: For the The updated value of the weight factor, For the The value of the weight factor before updating, is the learning rate, and its value is set to , is the prediction error of the incentive scoring mechanism for the overall operational performance score of the current production cycle, specifically the mean square error between the incentive score and the overall operational performance score.

[0183] Execute the incremental training process of the reinforcement learning scheduling optimization neural network and the equipment OEE behavior attribution model. Based on the incremental model training samples of each production cycle, update the weight factors of the incentive scoring mechanism in the policy network, value network and equipment OEE behavior attribution model of the reinforcement learning scheduling optimization neural network respectively. When the overall operational performance score is lower than the preset performance benchmark threshold of 0.85, the incremental model training samples generated in the current cycle are marked as negative samples. If the overall operational performance scores of two consecutive cycles are higher than 0.90, the incremental model training samples generated in the corresponding production cycle are marked as positive samples, and the historical scheduling feedback sample library is updated. According to the experience replay mechanism, data samples with a batch size of 64 are randomly selected for subsequent training.

[0184] Based on a continuously updated reinforcement learning scheduling optimization neural network and equipment OEE behavior attribution model, a long-term, multi-stage lean operations intelligent optimization framework is configured. The lean operations intelligent optimization framework is specifically divided into three stages:

[0185] The first stage is: operational data collection and monitoring, real-time collection of production site data and real-time monitoring of equipment OEE comprehensive scores and task execution status.

[0186] The second stage is: short-term response and feedback adjustment, which implements the feedback adjustment mechanism based on abnormal fluctuations in equipment OEE and dynamically adjusts the production scheduling plan.

[0187] The third stage is long-term optimization and model evolution. Through the global model incremental training mechanism, the parameters of the reinforcement learning scheduling optimization neural network and the equipment OEE behavior attribution model are continuously updated to form a long-term optimization path.

[0188] Through the long-cycle, multi-stage lean operation intelligent optimization framework described above, continuous improvement in production cycle performance is achieved, which is specifically reflected in the overall operation performance score gradually approaching the optimal level. When the overall operation performance score exceeds 0.95 for three consecutive production cycles, it is considered to have reached the optimization maturity stage, and the current model parameters are automatically solidified as the normalized operation model parameter settings for the production system.

[0189] In the above content, in actual application, first, deploy data collection interfaces with the customer order management system, material supply chain management system, production task scheduling system, process flow control system and equipment maintenance management system to collect customer order data, material supply status data, production task priority data, process timing requirement data and equipment maintenance plan data respectively, and uniformly map the collected data into a time series structure with data source identification, data collection timestamp, data content and data validity fields. The timestamp field adopts the ISO8601 format, and the data content field uses the international system of units for standardized conversion.

[0190] Secondly, for the unified and structured customer order data, material supply status data, production task priority data, process timing requirement data and equipment maintenance plan data, a semantic label set is generated based on preset semantic rules. Among them, when the delivery deadline of the customer order is less than 24 hours from the current time and the order priority level is level 1, an emergency order semantic label is generated. When the material inventory is lower than 20% of the minimum safety stock threshold and the supplier performance score is lower than 80 points, a "high-risk material" label is generated. Based on the above labels and data, a dynamic feature set of production demand is constructed, which includes numerical feature dimension, semantic label dimension, time weight dimension and data credibility dimension.

[0191] Next, a deep deterministic policy gradient algorithm was used to construct a reinforcement learning scheduling optimization neural network consisting of a policy network and a value network. The dynamic feature set of production demand and the equipment operating status were input into the policy network, and the scheduling mapping relationship between tasks and equipment was output. The constructed policy network consists of a three-layer fully connected structure, and the value network adopts a combination of a double-layer convolution and a fully connected structure. During the training process, the experience replay mechanism was used to extract 64 batches of samples, and the learning rate of the policy network was set to 1×10⁻ 4 , the value network learning rate is set to 1×10⁻³, and the soft update factor is set to 0.005.

[0192] Subsequently, an OEE monitoring unit based on edge computing architecture was deployed to collect data on the utilization rate, performance efficiency, and yield rate of each device, which constitute the three elements of OEE. Through anomaly detection logic, when the comprehensive OEE score of the device falls below the set threshold of 0.70, behavioral attribution analysis is initiated, and the OEE behavioral attribution model is called. Combined with the production scheduling path and equipment status changes, the main cause of the OEE fluctuation is automatically located and marked with production congestion, equipment aging, and process bottleneck behavior labels, which are used to feedback the adjustment strategy of the reinforcement learning production scheduling optimization neural network.

[0193] Finally, after OEE anomalies are identified, a feedback adjustment target vector structure is constructed, including the equipment number, task number, anomaly type, and correction priority. Based on the root cause of the anomaly output by the OEE behavior attribution model, the production scheduling feedback adjustment function is called to adjust the priority weight allocation mechanism in the production scheduling strategy network. When the equipment utilization rate is lower than 40%, the average task delay exceeds 30 minutes, or the equipment switching cost exceeds the 10-minute threshold, the scheduling strategy is adaptively updated. This closed-loop mechanism improves the robustness of the scheduling strategy and the overall operational efficiency of the equipment.

[0194] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A lean operation method based on production scheduling optimization and equipment OEE monitoring, characterized by: The steps include: S100. Deploy a multi-source heterogeneous data collection interface and perform data collection to obtain customer order data, material supply status data, production task priority data, process timing requirement data, and equipment maintenance plan data. Perform data preprocessing based on the collected data, generate a semantic tag set, obtain the semantic tag set, configure a production demand dynamic feature set based on the obtained semantic tag set, and calculate the feature impact weight. S200, configuring a reinforcement learning production scheduling optimization strategy environment based on a dynamic feature set of production demand, defining a reinforcement learning reward function, configuring a reinforcement learning production scheduling optimization neural network and a historical scheduling feedback sample set, and training the reinforcement learning production scheduling optimization neural network. Based on the reinforcement learning production scheduling optimization neural network, an optimal production scheduling plan is output, thereby obtaining an optimal production scheduling plan. S300: Deploy the equipment OEE monitoring unit and collect equipment operating status data. Calculate the equipment OEE comprehensive score based on the equipment operating status data. Configure the OEE anomaly identification mechanism based on the equipment OEE comprehensive score and detect abnormal OEE fluctuation events. Configure the equipment OEE behavior attribution model based on the abnormal OEE fluctuation events and identify influencing factors. Output the equipment OEE abnormal fluctuation attribution analysis report using the equipment OEE behavior attribution model and mark operational bottlenecks. This includes: Based on the detected abnormal OEE fluctuation events, an equipment OEE behavior attribution model is configured and used to analyze the abnormal equipment OEE fluctuation events. The equipment OEE behavior attribution model is based on the task scheduling path, task number, equipment operating status data, scheduling score data stored in the historical scheduling feedback sample set, and the equipment operating status data collected in the current cycle. The attribution steps include: establishing a scheduling behavior event table, establishing an equipment status change sequence, and synchronous retrieval and matching; Establish a scheduling behavior event table to record task number, scheduling time, execution device number and reinforcement learning reward function score; Establishing a sequence of equipment status changes: recording the changes in equipment utilization rate, equipment performance efficiency, and equipment yield rate over time; Synchronous retrieval and matching involves tracing back to the scheduling behavior event table and equipment status change sequence within 30 minutes for each equipment OEE abnormal fluctuation event to identify whether there are task changes, equipment maintenance plan adjustments, and frequent equipment task switching. If there are scheduling behaviors such as task changes, equipment maintenance plan adjustments, and frequent equipment task switching in the attribution step, then the scheduling behavior will be marked as a potential cause of abnormal equipment OEE fluctuation events; Use the scoring mechanism of the equipment OEE behavior attribution model to quantitatively score potential causes; After completing the attribution analysis based on the equipment OEE behavior attribution model, an equipment OEE abnormal fluctuation attribution analysis report is output. The equipment OEE abnormal fluctuation attribution analysis report includes: the downward trend of the equipment OEE comprehensive score, the duration of the abnormal fluctuation, the equipment number where the abnormal fluctuation occurred, the number of the task involved, the specific scheduling behavior type that induced the abnormal fluctuation, and the corresponding inducement score; S400: Trigger the production scheduling strategy feedback adjustment process based on abnormal equipment OEE fluctuation events, configure the feedback adjustment target vector, generate the task scheduling action set and strategy confidence vector based on the reinforcement learning production scheduling optimization neural network, configure adaptive feedback adjustment rules, and update the historical scheduling feedback sample set; S500. Configure the operation performance indicator scoring mechanism, configure the model incremental training samples based on the operation performance indicator scoring mechanism, configure the production scheduling optimization neural network and equipment OEE behavior attribution model incremental training mechanism, and update the historical sample library.

2. A lean operation method based on production scheduling optimization and equipment OEE monitoring as claimed in claim 1, characterized in that: The S100 is specifically: S100.

1. Deploy multi-source heterogeneous data collection interfaces and conduct data collection to obtain customer order data, material supply status data, production task priority data, process timing requirements data, and equipment maintenance plan data; Before implementing lean operations, a data collection interface for multi-source heterogeneous data collection is deployed. This interface connects to the customer order management system, material supply chain management system, production task scheduling system, process flow control system, and equipment maintenance management system, respectively, and collects customer order data, material supply status data, production task priority data, process timing requirements data, and equipment maintenance plan data in real time. The data acquisition interface collects and obtains customer order data, material supply status data, production task priority data, process timing requirement data and equipment maintenance plan data from the customer order management system, material supply chain management system, production task scheduling system, process flow control system and equipment maintenance management system respectively; S100.

2. Perform data preprocessing after data collection and generate a semantic tag set to obtain a semantic tag set; Based on the collected customer order data, material supply status data, production task priority data, process timing requirement data, and equipment maintenance plan data, each data is structured and mapped according to a unified time series structure. The time series structure includes: a data source identification field, a data collection timestamp field, a data content field, and a data validity field. The data collection timestamp field is unified using the ISO8601 standard, and the data content field is standardized using the International System of Units. Based on the customer order data, material supply status data, production task priority data, process timing requirement data, and equipment maintenance plan data in the time series structure, a preset semantic rule set is called to generate a corresponding semantic tag set, wherein the semantic tag set includes: an order urgency level tag, a material shortage level tag, a task priority level tag, a process complexity level tag, and an equipment availability level tag; When the delivery deadline of the customer order data is less than 24 hours from the current time and the order priority level is the highest level, the generated order urgency level label is: urgent order; When the material inventory in the material supply status data is less than 20% of the minimum safety stock threshold and the supplier performance score is less than 80 points, the material tension level label generated is: high-risk material; S100.

3. Configure a dynamic feature set of production requirements based on the obtained semantic tag set, and calculate feature impact weights; Based on the customer order data, material supply status data, production task priority data, process timing requirement data, equipment maintenance plan data and semantic label set of the time series structure, a production demand dynamic feature set is configured. The production demand dynamic feature set is a multidimensional vector structure, specifically including: a numerical feature dimension, a semantic label dimension, a time weight dimension and a data credibility dimension; Based on the dynamic feature set of production demand, a data fusion scoring mechanism is executed for customer order data feature items, material supply status data feature items, production task priority data feature items, process timing requirement data feature items and equipment maintenance plan data feature items. The data fusion scoring mechanism includes: scheduling impact factor, resource conflict factor and response time factor. The data fusion scoring mechanism is calculated by the scoring value = 0.5×scheduling impact factor + 0.3×resource conflict factor + 0.2×response time factor, wherein the scheduling impact factor is calculated based on the task priority and the order delivery deadline, the resource conflict factor is calculated based on the material inventory tightness and the equipment conflict probability, and the response time factor is determined based on the time interval between the data collection timestamp and the current system. When the interval exceeds 60 minutes, the score of the corresponding feature item is automatically reduced by 20%.

3. The lean operation method based on production scheduling optimization and equipment OEE monitoring according to claim 1, characterized in that: The S200 is specifically: S200.

1. Configure a reinforcement learning production scheduling optimization strategy environment based on the dynamic feature set of production demand, and define a reinforcement learning reward function. Based on the dynamic feature set of production demand and the operating status of the equipment, a reinforcement learning production scheduling optimization strategy environment is configured, wherein the reinforcement learning production scheduling optimization strategy environment includes: a state space, an action space, a reward function, and a strategy function; The state space consists of a set of dynamic production demand features and equipment operating status. The dynamic production demand feature set includes: numerical feature dimensions, semantic label dimensions, time weight dimensions, and data credibility dimensions. The equipment operating status includes: equipment number, current utilization rate, equipment availability level, and historical task execution efficiency. The action space is defined as the scheduling mapping relationship from task to equipment, that is, the set of allocation mappings between task numbers and equipment numbers. The reward function evaluates the scheduling quality after the execution of the scheduling optimization strategy, and the strategy function selects the corresponding scheduling action based on the state space. Call the device load balancing objective function, task delay minimization objective function, and task switching cost minimization objective function to configure the reinforcement learning reward function; The reinforcement learning reward function is the weighted sum of the device load balancing objective function, the task delay minimization objective function, and the task switching cost minimization objective function. Specifically, Where: is the reward score of the current production scheduling plan, which is used to feed back to the reinforcement learning strategy. is the output value of the device load balancing objective function, is the output value of the objective function for minimizing task delay, is the output value of the objective function minimizing the task switching cost, and the coefficient 、 、 They are weighted factors of each optimization objective, used to adjust the contribution ratio of the objective, and the sum of the weights is 1; S200.

2. Configure the reinforcement learning scheduling optimization neural network and the historical scheduling feedback sample set, and train the reinforcement learning scheduling optimization neural network; Calling a deep deterministic policy gradient algorithm to configure a reinforcement learning scheduling optimization neural network, wherein the reinforcement learning scheduling optimization neural network includes: a policy network and a value network; The input of the policy network is: state space data, and the output is: the selection probability of each scheduling action in the action space. The policy network adopts a three-layer fully connected network structure, including: input layer, two hidden layers and output layer. The hidden layer activation function adopts the ReLU function. The input of the value network is the concatenation of state space data and action space data, and the output is the value function Q value of the corresponding state and action pair. The value network adopts the structure of two convolutional layers and one fully connected layer. Configuring a historical scheduling feedback sample set, which consists of: a set of dynamic characteristics of production demand in a historical period, equipment operating status, actual task completion time, task scheduling path, and corresponding scheduling score data; The reinforcement learning scheduling optimization neural network is trained through the historical scheduling feedback sample set. The experience replay mechanism is used to randomly extract data samples from the historical scheduling feedback sample set. The batch size is 64. The soft update mechanism is used to update the target policy network parameters and the target value network parameters. The policy network learning rate is set to 1×10⁻ 4 , the value network learning rate is set to 1×10⁻³; S200.

3. Output the optimal production scheduling plan based on the reinforcement learning production scheduling optimization neural network, and obtain the optimal production scheduling plan; After the reinforcement learning scheduling optimization neural network is trained, the dynamic feature set of production demand and the real-time equipment operating status are input into the policy network, which then outputs a set of task-to-equipment scheduling actions corresponding to the current status. The task-to-equipment scheduling action set is a matrix structure, where each matrix element represents a specific allocation decision for a task number to an equipment number. Based on the output task-to-equipment production scheduling action set, the optimal production scheduling plan is formed, and the optimal production scheduling plan is submitted to the scheduling execution module to drive the equipment to execute the production task according to the generated optimal production scheduling plan, record the equipment operation data and task execution data of this production schedule, and update it to the historical scheduling feedback sample set.

4. The lean operation method based on production scheduling optimization and equipment OEE monitoring according to claim 1, characterized in that: The S300 is specifically as follows: S300.

1. Deploy equipment OEE monitoring units, collect equipment operating status data, and calculate the equipment OEE comprehensive score based on the equipment operating status data. Deploy an equipment OEE monitoring unit, which includes: a sensor acquisition module, an edge computing processing module, and a data transmission module; The sensor acquisition module is used to collect equipment operation status data in real time, wherein the equipment operation status data includes: equipment utilization rate, equipment performance efficiency and equipment yield rate; Equipment utilization rate is the ratio of the actual equipment operating time to the planned available time. Equipment performance efficiency is the ratio of the standard output quantity to the actual output quantity during the actual operating time. Equipment yield rate is the ratio of the number of qualified products to the actual output quantity. The equipment OEE monitoring unit collects equipment operating status data once a minute and performs local preprocessing through the edge computing processing module. The preprocessing includes: outlier filtering, sliding average smoothing, and timestamp normalization. Based on the equipment operating status data collected and pre-processed by the equipment OEE monitoring unit, the equipment OEE comprehensive score calculation formula is used to perform a weighted calculation of the equipment utilization rate, equipment performance efficiency, and equipment yield rate. The specific calculation is: Where: It is the comprehensive score of equipment OEE, which is used to comprehensively evaluate the operating efficiency of the equipment in the current cycle. The equipment utilization rate is the ratio of the actual operating time of the equipment to the planned available time. is the equipment performance efficiency, which is the ratio of the standard output quantity to the output quantity during the actual operating time. It is the equipment yield rate, which is the ratio of the number of qualified products to the actual output quantity. is the weighted coefficient of equipment utilization rate, set to 0.4, is the weighted coefficient of equipment performance efficiency, set to 0.3, is the weighted coefficient of the equipment yield rate, which is set to 0.3; S300.

2. Configure an OEE anomaly identification mechanism based on the equipment OEE comprehensive score and detect abnormal OEE fluctuation events; Based on the calculated equipment OEE comprehensive score, an OEE anomaly identification mechanism is configured. The OEE anomaly identification mechanism adopts a sliding window trend detection algorithm and a threshold judgment strategy. The detection process is as follows: the historical reference window length is set to minutes, and the real-time monitoring window length is Minutes, calculate the average comprehensive OEE score of the equipment in the real-time monitoring window and the historical reference window respectively; Setting anomaly recognition thresholds When the equipment OEE comprehensive score changes by 5%, Less than the negative anomaly recognition threshold If the fluctuation lasts for more than three sampling periods, it is considered an abnormal OEE fluctuation event.

5. The lean operation method based on production scheduling optimization and equipment OEE monitoring according to claim 1, characterized in that: The S400 is specifically: S400.

1. Trigger the production scheduling strategy feedback adjustment process based on abnormal equipment OEE fluctuation events and configure the feedback adjustment target vector. Based on the equipment OEE comprehensive score output by the equipment OEE monitoring unit in real time, the equipment OEE anomaly identification mechanism continuously monitors changes in the equipment OEE comprehensive score. When the change in the equipment OEE comprehensive score is less than the preset negative anomaly identification threshold of 5%, and the equipment OEE behavior attribution model determines that the abnormal OEE fluctuation event is caused by a specific scheduling behavior, and the corresponding incentive score exceeds 75 points, the production scheduling strategy feedback adjustment process is triggered; Based on the equipment OEE abnormal fluctuation attribution analysis report output by the equipment OEE behavior attribution model, the equipment number, task number, specific scheduling behavior type that induced the abnormal fluctuation, inducement score value, and abnormal fluctuation event occurrence time are obtained, and the feedback adjustment target vector is configured; S400.

2. Generate task scheduling action set and strategy confidence vector based on reinforcement learning scheduling optimization neural network; Based on the reinforcement learning scheduling optimization neural network, the dynamic feature set of production demand in the current cycle and the real-time collected equipment operation status data are input to obtain the task-to-equipment scheduling action set and output the confidence vector of the task scheduling strategy. S400.

3. Configure adaptive feedback adjustment rules and update historical scheduling feedback sample sets; Based on the feedback adjustment target vector and the task scheduling strategy confidence vector, an adaptive feedback adjustment rule is configured. The adaptive feedback adjustment rule includes three rules: The first rule in the adaptive feedback adjustment rule is: If the incentive score in the feedback adjustment target vector is greater than 80 points, and the confidence of the corresponding task number in the task scheduling strategy confidence vector is less than 0.6, then the task transfer strategy is executed, and the task with the corresponding task number is preferentially transferred to the production equipment with the lowest load rate and the highest equipment availability level in the current cycle; The second rule in the adaptive feedback adjustment rule is: If the scheduling behavior type that induces abnormal fluctuations in equipment OEE in the feedback adjustment target vector is frequent task switching, then based on the current continuous task sequence on the corresponding equipment number in the task scheduling action set, perform a task reordering operation to reduce the number of continuous task switches on the equipment; The third rule in the adaptive feedback adjustment rules states: If the scheduling behavior type that causes the abnormal fluctuation in equipment OEE in the feedback adjustment target vector is planned maintenance delay, then freeze new task assignments for the corresponding equipment number within the next 60 minutes of the full production cycle, and immediately reserve a time window for the equipment to perform maintenance operations; Generate a revised set of task-to-equipment production scheduling actions through adaptive feedback adjustment rules; Based on the generated revised task-to-equipment production scheduling action set, the revised task-to-equipment production scheduling action set is submitted to the scheduling execution module, driving the production equipment to execute the production tasks according to the revised task-to-equipment production scheduling action set, and recording the relevant data that triggers the scheduling strategy feedback adjustment process. At the same time, it is added and updated to the historical scheduling feedback sample set.

6. The lean operation method based on production scheduling optimization and equipment OEE monitoring according to claim 1, characterized in that: The S500 is specifically: S500.

1. Configure an operational performance indicator scoring mechanism and configure incremental model training samples based on the operational performance indicator scoring mechanism. An operation performance indicator scoring mechanism is configured based on the equipment OEE comprehensive score, the output value of the task delay minimization objective function, the output value of the equipment load balancing objective function, and the output value of the task switching cost minimization objective function. The operation performance indicator scoring mechanism is calculated in units of production cycle; Based on the operational performance indicator scoring mechanism, the overall operational performance score of the current production cycle is annotated, and model incremental training samples are configured. The model incremental training samples specifically include: a dynamic feature set of production demand in the current production cycle, equipment operating status data, a set of task-to-equipment scheduling actions, an attribution analysis report on abnormal fluctuations in equipment OEE, a confidence vector of the task scheduling strategy, and the overall operational performance score; S500.

2. Configure the scheduling optimization neural network and the incremental training mechanism for the equipment OEE behavior attribution model, and update the historical sample library. Based on the incremental training samples of the model, incremental training mechanisms are configured for the reinforcement learning scheduling optimization neural network and the equipment OEE behavior attribution model. The incremental training mechanism of the reinforcement learning scheduling optimization neural network includes: based on the production demand dynamic feature set, equipment operating status data, and task-to-equipment scheduling action set in the incremental training samples of the model configured in the current cycle, a deep deterministic policy gradient algorithm is used to update network parameters, and the parameter update method continues to use the soft update mechanism; The incremental training mechanism of the equipment OEE behavior attribution model includes: recalculating the weight factor of the incentive scoring mechanism in the equipment OEE behavior attribution model based on the equipment operation status data, equipment OEE abnormal fluctuation attribution analysis report and overall operation performance score in the model incremental training samples. 、 、 ; Execute the incremental training process of the reinforcement learning scheduling optimization neural network and the equipment OEE behavior attribution model. Based on the incremental model training samples of each production cycle, update the weight factors of the incentive scoring mechanism in the policy network, value network and equipment OEE behavior attribution model of the reinforcement learning scheduling optimization neural network respectively. When the overall operational performance score is lower than the preset performance benchmark threshold of 0.85, the incremental model training samples generated in the current cycle are marked as negative samples. If the overall operational performance scores of two consecutive cycles are higher than 0.90, the incremental model training samples generated in the corresponding production cycle are marked as positive samples, and the historical scheduling feedback sample library is updated. According to the experience replay mechanism, data samples with a batch size of 64 are randomly selected for subsequent training.

Citation Information

Patent Citations

  • Method and device for acquiring abnormal attribution path of index, electronic equipment and medium

    CN118820977A

  • Index transaction attribution analysis method and device, computer equipment and storage medium

    CN119378689A