Method and system for optimizing uncertain processing time of flexible job shop

Through the dynamic fuzzy Gaussian model and sliding window Kalman filter combined with meta-learning deep reinforcement learning, the uncertain processing time optimization system of flexible job workshops solves the problem of insufficient robustness in flexible workshop scheduling, and achieves efficient processing time prediction and improved equipment utilization.

CN120806230APending Publication Date: 2025-10-17ZHEJIANG SCI-TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510841182.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively characterize uncertainty problems under dynamic disturbances, resulting in insufficient robustness in flexible workshop scheduling, high prediction errors, low equipment utilization, and inability to meet the needs of intelligent manufacturing for real-time iterative optimization.

Method used

A dynamic fuzzy Gaussian model combined with a sliding window Kalman filter is used to update fuzzy parameters in real time, and meta-learning deep reinforcement learning is used to achieve rapid strategy migration with small samples. Through real-time feedback closed-loop optimization using digital twins, an uncertain processing time optimization system for flexible job shops is constructed.

Benefits of technology

Significantly reduce processing time prediction errors, improve equipment utilization, enhance the robustness of scheduling plans, increase equipment utilization by 20%, and reduce prediction errors to below 5%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806230A_ABST
    Figure CN120806230A_ABST
Patent Text Reader

Abstract

The invention discloses a flexible workshop uncertain processing time optimization method and system, relates to the field of intelligent manufacturing and industrial scheduling optimization, and aims to solve the problems of large flexible workshop processing time prediction error and poor scheduling scheme robustness in a dynamic disturbance scene. The processing time skewed distribution is dynamically represented through an asymmetric Gaussian membership function, fuzzy parameters are updated in real time in combination with sliding window Kalman filtering, and based on meta-learning deep reinforcement learning, small sample fast strategy migration is realized, the processing time prediction error is reduced, the equipment utilization rate is improved, and the new scene response efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent manufacturing and industrial scheduling optimization, in particular to a flexible job shop uncertain processing time optimization method and system. BACKGROUND

[0002] There are still significant limitations in the current flexible job shop scheduling field. The existing technology cannot effectively depict the uncertainty problem under dynamic disturbance. The main problems are that the static modeling technology cannot respond to device state fluctuations or process adjustments in real time, the symmetry distribution assumption is seriously deviated from the asymmetric characteristics of real processing time, and the defuzzification process ignores the influence of extreme events, resulting in insufficient scheduling robustness. These problems cause high prediction error and low device utilization, which cannot meet the core demand of real-time iterative optimization for intelligent manufacturing.

[0003] For example, the Chinese patent with publication number CN111445081A proposes a time-based S 3 The PR network model and the deep reinforcement learning scheduling method of the graph convolutional neural network realize adaptive optimization through a digital twin environment; however, this scheme relies on static process time parameters and does not construct a dynamic time parameter update mechanism, making it difficult to respond to device efficiency decay or order changes and other disturbances. The Chinese patent with publication number CN113361139B relates to a production line rolling optimization system based on digital twinning, which uses an evolutionary algorithm to periodically generate scheduling strategies; however, its fixed rolling window mode results in significant response delay, and it lacks small sample fast adaptation capability, so it needs to perform full calculation repeatedly when facing sudden disturbances. SUMMARY

[0004] To solve the problems of large processing time prediction error and poor robustness of scheduling schemes in the dynamic disturbance scenario, the present application proposes a flexible job shop uncertain processing time optimization method and system, which can dynamically represent the processing time skew distribution through an asymmetric Gaussian membership function, update fuzzy parameters in real time combined with a sliding window Kalman filter, and realize small sample fast strategy migration based on meta-learning deep reinforcement learning, thereby reducing processing time prediction error, improving device utilization, and improving response efficiency in new scenarios.

[0005] To achieve the above purpose, the present application adopts the following technical scheme: a flexible job shop uncertain processing time optimization method, comprising the following steps: S1, real-time collection of job shop device state and processing time data, synchronization to the digital twin layer to build a dynamic fuzzy Gaussian number model; S2, based on the historical processing time data set, initializing and binding the parameters of the dynamic fuzzy Gaussian number model to the digital twin body; S3, updating the parameters of the dynamic fuzzy Gaussian number model according to the device state change; S4, input the updated parameters into the deep reinforcement learning strategy engine to generate scheduling instructions, execute scheduling and trigger closed-loop optimization.

[0006] In the technical solution, the dynamic fuzzy Gaussian number (DFGN) model is constructed to describe the asymmetric processing time distribution, the OPC UA protocol is used to collect equipment data, the DFGN parameters are dynamically corrected, the meta-learning deep reinforcement learning (DRL) strategy migration technology is combined to solve the problems of large prediction error and response lag caused by static modeling, symmetric assumption and insufficient dynamic adaptability, significantly improve the robustness and resource utilization of the scheduling scheme, and achieve the technical effects of processing time prediction error not exceeding 5% and equipment utilization rate increasing by 20%.

[0007] Preferably, in step S2, the dynamic fuzzy Gaussian number model uses an asymmetric Gaussian membership function to independently represent the conservative distribution and pessimistic distribution characteristics of the processing time through left and right standard deviations, respectively.

[0008] Preferably, in step S2, the parameter initialization of the dynamic fuzzy Gaussian number model includes: determining the mean value according to the median of the historical processing time data set, calculating the left standard deviation based on the lower quantile, and calculating the right standard deviation based on the higher quantile.

[0009] Preferably, the step S3 includes: S31, calculate the prediction error in the sliding window every fixed time period; S32, when the prediction error exceeds the preset threshold, start the parameter update of the dynamic fuzzy Gaussian number model.

[0010] Preferably, in step S3, the update includes: when the prediction error exceeds the threshold, dynamically correct the model using the sliding window Kalman filter algorithm; when the device state changes are detected, immediately start the sliding window Kalman filter correction model; the corrected parameters need to satisfy the monotonicity constraint of the asymmetric Gaussian function.

[0011] Preferably, in step S4, the action space of the deep reinforcement learning strategy engine is designed to bind order combinations and transition sets, wherein the order combination supports parallel processing of multiple orders to optimize priority conflicts, and the transition set maps orders to specific device operations to avoid resource conflicts.

[0012] Preferably, in step S4, generating scheduling instructions includes: calculating the similarity between the current working condition and the historical strategy, if it exceeds the matching threshold, loading the historical strategy, otherwise generating a new strategy based on a small sample through a meta-learning mechanism and updating the strategy library.

[0013] Preferably, the step S4 includes: updating the DFGN parameters every 30 minutes to trigger strategy retraining.

[0014] The application also adopts the following technical scheme: a flexible job shop uncertain machining time optimization system, which executes the flexible job shop uncertain machining time optimization method, and comprises: A data acquisition module configured to acquire equipment data in real time through an industrial Internet of Things; A digital twin module comprising: a physical layer for synchronizing equipment state data; a virtual layer for integrating a dynamic fuzzy Gaussian number model and performing parameter updating; and a decision layer for deploying a deep reinforcement learning strategy engine to generate scheduling actions; A closed-loop optimization module configured to monitor scheduling effects and feed back to the virtual layer to trigger parameter re-updating.

[0015] Preferably, the virtual layer is further configured to: when the machining time exceeds the right skewness boundary, automatically migrate the task to other idle digital twin nodes, and through weighted integration, de-fuzz the reserved pessimistic boundary to buffer the equipment failure risk.

[0016] The application has the following beneficial effects: 1) Dynamic updating of machining time parameters improves the real-time response capability to equipment state fluctuations; 2) Accurately depicting the asymmetric distribution characteristics of machining time avoids scheduling deviations caused by symmetric assumptions; 3) Through the meta-learning mechanism, small sample rapid strategy migration is realized, and the training cost of new scenarios is significantly reduced; 4) A digital twin driven closed-loop optimization architecture is constructed to support continuous iterative evolution of virtual and real collaboration; 5) Enhancing the robustness of the scheduling scheme to extreme events reduces the risk of downtime caused by sudden failures. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 is a flowchart of the flexible job shop uncertain machining time optimization method of the application.

[0018] Figure 2 is a structural diagram of the flexible job shop uncertain machining time optimization system of the application. DETAILED DESCRIPTION

[0019] To make the purpose, technical scheme and advantages of the application clearer, further detailed description of the application will be given below in combination with the drawings and examples. It should be understood that the specific embodiments described herein are only one of the best embodiments of the application, which are used to explain the application and do not limit the protection scope of the application. All other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.

[0020] Example 1 The embodiment provides a flexible job shop uncertain processing time optimization method, referring to Figure 1 , comprising the following steps.

[0021] Step S1, data acquisition and preprocessing.

[0022] Real-time acquisition of workshop equipment data through industrial Internet of Things sensor network, including dynamic information of equipment state, material flow and processing time, and the equipment state includes but is not limited to vibration and temperature state.

[0023] Detect the connection state of the sensor network, if ready, start the OPC UA protocol data stream; otherwise, check and repair the sensor connection to ensure data synchronization between the physical equipment and the digital twin layer.

[0024] The preprocessing of the original data includes denoising, normalization operation and data quality verification.

[0025] Qualified data is directly written into the cache for subsequent synchronization, and unqualified data needs to be cleaned and repaired.

[0026] Specifically, in the embodiment, if the data integrity is greater than 95% and the noise is lower than the threshold, write into the cache for subsequent processing; if the data is missing or abnormal, i.e. unqualified, perform cleaning and repair, such as interpolation completion or rejection of abnormal values.

[0027] Finally, the preprocessed data is synchronized to the digital twin layer, and a virtual model is constructed for real-time simulation and prediction.

[0028] Step S2, dynamic fuzzy Gaussian number modeling.

[0029] Based on the historical data, the mean, left skew standard deviation and right skew standard deviation parameters of the dynamic fuzzy Gaussian number are initialized, and are bound to the digital twin body for dynamic updating of the parameters.

[0030] The construction of the digital twin body includes a three-layer architecture of the physical layer (PLC data), the virtual layer (DFGN model) and the decision layer (DRL), and the equipment state is synchronized in real time, as shown in Figure 2 .

[0031] In the embodiment, the left skew (conservative) and right skew (pessimistic) standard deviations are represented by asymmetric Gaussian membership functions respectively, and are dynamically adjusted through the historical data of the equipment (such as failure rate and wear degree).

[0032] Next, taking the M01 boring process of a CNC machine tool as an example, how to determine the initial parameters according to the historical data is described in detail.

[0033] Given a historical processing time data set X of the process capacity n, X satisfies ascending arrangement.

[0034] The median of the historical processing time dataset is taken as the mean u0, and μ0-Q 0.1 (X) is taken as the left-skewed standard deviation, and Q 0.9 (X)-μ0 is taken as the right-skewed standard deviation, where Q α (X) represents that there are α proportions of data points in the dataset that are less than or equal to the value, and in some other embodiments, the value of α can be adjusted according to actual needs.

[0035] Specifically, Q α (X) can be calculated by the following method.

[0036] The position index k corresponding to Q α (X) is calculated by (n+1)α, when k is an integer, Q α (X) directly takes the kth value x k in the dataset, otherwise, let k be decomposed into an integer part i and a decimal part f, and the linear interpolation is calculated, which can be specifically represented as represents rounding down.

[0037] Specifically, the historical processing time dataset X of this process is [45, 48, 50, 52, 55] minutes, the capacity n is 5, when α is 0.1, k can be calculated to be equal to 0.6, and then Q 0.1 (X) is equal to x1+0.6(x2-x1), that is, 46.8 minutes.

[0038] Similarly, when α is 0.9, k is 5.4, and then Q 0.9 (X) is equal to x5+0.4(x6-x5), assuming that x6 is 57 minutes, Q 0.9 (X) is 55.8 minutes.

[0039] Further, the initial parameters are the mean of 50 minutes, the left-skewed standard deviation σ l is 3.2 minutes, and the right-skewed standard deviation σ r is 5.8 minutes.

[0040] In the present application, the left-skewed standard deviation represents the conservative estimation time, such as unexpected shortening of the processing time, the right-skewed standard deviation represents the pessimistic estimation time, such as delay caused by equipment failure, reflecting the risk of long-tailed distribution, effectively solving the defects of traditional fuzzy numbers in skew distribution and dynamic adaptability, and the processing time prediction error can be reduced to ≤5%.

[0041] Subsequently, when the parameters are adjusted by Kalman filtering, the newly collected processing time data will generate a new dataset, and Q α (X) is updated to update the left-skewed standard deviation and the right-skewed standard deviation.

[0042] In some other embodiments, Q α The calculation of Q α (X) is added with a boundary constraint, which is automatically truncated to the first and last values when k is out of the data index range, to avoid distortion caused by extrapolation of interpolation.

[0043] After the completion of parameter initialization, based on real-time feedback of digital twin, the left and right standard deviations are dynamically updated by using sliding window Kalman filtering.

[0044] The prediction error ∈ k is calculated every fixed time within the window. k When the prediction error ∈ k is greater than 5%, the parameter update is started.

[0045] When the digital twin layer receives a device state change notification, such as a machining error exceeding the limit or a device failure, the DFGN parameter update process is triggered, and the left and right standard deviations of the DFGN model are corrected based on the current real-time observed machining time data using the Kalman filtering algorithm; if there is no change, the current parameters are maintained.

[0046] Short window response is fast but sensitive to noise, and long window is stable but has high delay. In this technical solution, two complementary triggering mechanisms are used, device state change solves the immediacy risk and reduces downtime loss, and prediction error exceeding threshold solves systematic deviation and improves long-term stability. Through the double triggering mechanism, real-time and stability are balanced, covering sudden disturbance and gradual disturbance.

[0047] It should be noted that in the overlapping scenario, the device state change is used as a higher priority event for real-time response to avoid further expansion of the error.

[0048] The specific process of windowed Kalman filtering parameter adjustment is described in detail below.

[0049] First, data window division is performed. Based on time series or event triggering mechanism, the data of the last N time points is selected to form a sliding window (e.g., N=80 to minimize the error).

[0050] Then, state prediction and parameter calculation are performed, including the following sub-steps.

[0051] Step A1, the current DFGN parameters are obtained, including the left and right standard deviations σ l and σ r .

[0052] Step A2, the device state data is collected, the parameter change trend is predicted by Kalman filtering, and the state equation and observation equation are constructed.

[0053] Step A3, the new σ l and σ rThe process noise covariance Q and the measurement noise covariance R are dynamically adjusted using the residual statistical characteristics within the window.

[0054] After that, the parameters are checked and corrected.

[0055] The new parameters are checked to see if they are within a reasonable range, such as σ1 and σ r The mathematical constraints of the asymmetric Gaussian function (such as positive definiteness and monotonicity) must be met. If the range is exceeded, boundary constraints or interpolation adjustments are made through the parameter correction algorithm.

[0056] Finally, the model is updated.

[0057] After the check is passed, the DFGN model parameters are updated, and new fuzzy rules are generated for scheduling decisions.

[0058] The asymmetric parameter updating method of the DFGN model is described in detail below.

[0059] For the updating of the left and right bias standard deviations, first calculate the left and right bias errors.

[0060] The new left bias standard deviation is equal to the original left bias standard deviation plus the correction term, which is obtained by multiplying the left bias error by the gain coefficient K l of the left standard deviation, which reflects the influence of the left bias error on parameter correction.

[0061] Similarly, the new right bias standard deviation is equal to the original right bias standard deviation plus the correction term, which is obtained by multiplying the right bias error by the gain coefficient K r of the right standard deviation, which reflects the influence of the right bias error on parameter correction.

[0062] The new mean value is equal to the original mean value plus the mean correction term, which is obtained by multiplying the mean error by the mean gain coefficient K μ , which determines the weight of the observation value on the mean update, and the mean error is obtained by subtracting the observation prediction value Hμ k from the actual observation value z k at time k, where the actual observation value at time k is the real-time monitoring processing time of the sensor, and the observation prediction value is the prediction result mapped to the observation space through the current state estimation μ k .

[0063] The Kalman gain matrix is obtained by , where P is the prediction error covariance matrix, H k is the observation matrix, and R k is the observation noise covariance matrix.

[0064] Step S3, DRL policy adjustment and scheduling execution process, including the following two steps.

[0065] First is policy synchronization and retraining decision.

[0066] The updated DFGN parameters are synchronized to the DRL policy engine, and the system determines whether policy retraining is needed.

[0067] If retraining is needed, start the DRL retraining process, prepare the training set containing the latest state data, optimize the policy model through iterative training until the convergence condition is met, such as the loss function is stable or the maximum iteration number is reached.

[0068] If retraining is not needed, try to match the existing policy library, calculate the similarity between the current state and the library policy, and preferentially load the matching policy.

[0069] Second is dynamic scheduling execution and feedback, including the following several sub-steps.

[0070] B1, if no matching policy is found, start partial training adjustment, update the policy library and feedback to the scheduling system.

[0071] B2, execute new scheduling instructions, such as adjusting device task allocation or material transportation path.

[0072] B3, monitor the scheduling effect, evaluate the impact of parameter adjustment on production efficiency, such as shortening the processing cycle or improving resource utilization.

[0073] B4, decide whether to further adjust the parameters according to the evaluation results, forming a closed-loop feedback mechanism.

[0074] The following describes the process of DRL training in detail.

[0075] First, input DFGN parameters to generate a policy with the goal of minimizing the maximum completion time.

[0076] To reflect the actual scheduling situation of the workshop after each decision, the workpiece position state is integrated into the state space; considering the need to cope with uncertain factors in the production environment, a fuzzy membership degree set is introduced into the state; since processing time directly affects scheduling efficiency and production cycle, it is part of the state, and the algorithm can directly learn the impact of different workpiece processing priorities and scheduling order on total processing time cost.

[0077] Therefore, the production environment of the workshop at time t can be designed as a combination of workpiece position state, fuzzy membership degree set and workpiece processing time.

[0078] The workpiece position state represents the position of the workpiece to be processed at the current time t. In this embodiment, define Loc(J i ,Mk 1 means the workpiece J is currently located at the device M i k i k 0 means the workpiece J is currently located at the device M i k The state information can be stored in a position state matrix.

[0079] The workpiece processing time reflects the processing time of the workpiece at the corresponding processable device, and the data is stored in a processing time table. Under the condition of considering processing uncertainty, the processing time is stored in the table in the form of a fuzzy number.

[0080] The action space of the MDP reflects the process of the environment changing from the current state S t to the next state S t+1 at time t, i.e., the process of the workpiece triggering a transition to perform a corresponding action.

[0081] The DRL policy generates actions based on the input of the workpiece position matrix, the fuzzy membership, and the state of the workpiece processing time table, i.e., at time t, the workpieces in the order set trigger respective transitions, and the action set composed of all matching actions at the moment.

[0082] In this embodiment, the action is the order combination Order M and the transition set T K triggered by the order at the current moment.

[0083] The order combination is the order set that is simultaneously scheduled to be executed at the current moment, and each order contains multiple workpieces.

[0084] It can be understood that the order combination is a combined decision of orders, supports parallel processing of multiple orders, and can solve the order priority conflict problem, such as urgent order insertion, and realizes dynamic load balancing.

[0085] The transition set is a set of triggered device-level operation instructions, which represents the process transition action of the workpiece in the order on the device, such as “workpiece 1 starts drilling on device M3”. The order decision is converted into executable device instructions through the transition set, and the process constraint is ensured.

[0086] It can be understood that the use of the transition set can automatically avoid device conflicts, such as device M1 not being assigned two workpieces at the same time.

[0087] ​​​​In the prior art, order scheduling and equipment allocation are decided separately, orders are selected first and then equipment is allocated, which leads to local optimization, such as equipment conflict. In the present application, order combination and equipment transition are bound as a unified action, which can achieve global optimization by selecting which orders to optimize order priority and by specifying which processes to be performed by equipment to optimize equipment utilization.

[0088] In the reinforcement learning framework, the reward function serves as the guidance signal for the optimization of the agent's strategy, and its design needs to be closely related to the core optimization goal of the scheduling system.

[0089] Traditional methods usually use subjective experience values of completion time as function thresholds, but static experience values are difficult to adapt to changes in dynamic production environments.

[0090] Therefore, in the present application, the experience threshold is replaced by the completion period prediction value, and the optimized reward function can not only accurately reflect the current scheduling efficiency, but also has the ability to prospectively evaluate future production situations.

[0091] This prediction-driven reward mechanism can effectively solve the problem of strategy overfitting caused by fixed thresholds.

[0092] The reward function is fused with the order completion period prediction. If the current round completion time is less than the result obtained by the current order completion period prediction, a positive reward is obtained. If the current round completion time is greater than the result obtained by the current order completion period prediction, a negative reward is obtained. If the current round completion time is equal to the result obtained by the current order completion period prediction, the reward is 0.

[0093] Avoiding the convergence speed limitation caused by fixed reward values, taking the difference between the result obtained by the current order completion period prediction and the current round completion time, the greater the single round iteration interpolation value, the more prominent the reward obtained.

[0094] Here, the DDQN algorithm is taken as an example. The algorithm can directly take actions based on the original state input to obtain maximum cumulative reward feedback, reducing the complexity of modeling. The experience return mechanism is introduced into DDQN, which randomly samples learning from experience samples, breaks the correlation between data, and enhances the stability and efficiency of learning. Two networks are used, one for selecting actions and the other for evaluating action values, to reduce the estimation risk.

[0095] The DFGN parameters are updated once every fixed time set by humans, triggering the retraining of the strategy.

[0096] Embodiment 2 The present embodiment provides a flexible job shop uncertain processing time optimization system, as shown in Figure 2 The system includes a data acquisition module, a digital twin module, and a closed-loop optimization module.

[0097] The data acquisition module collects real-time data of workshop equipment through industrial Internet of Things. The specific deployment includes but is not limited to sensor networks of vibration sensors and temperature sensors at the main shaft of numerical control machine tools and the joints of robots, and the processing time, equipment vibration amplitude and temperature data are transmitted at a millisecond level frequency through the OPC UA protocol.

[0098] The collected raw data is filtered by wavelet threshold to remove noise, and the missing values are repaired by linear interpolation, ensuring that the data quality is higher than 95% before being synchronized to the digital twin module.

[0099] The digital twin module includes a physical layer, a virtual layer and a decision layer.

[0100] The physical layer receives the data processed by the data acquisition module, constructs the device object model, such as binding a unique device ID and process relationship chain for each CNC machine tool.

[0101] The virtual layer integrates a dynamic fuzzy Gaussian number model. This model uses an asymmetric Gaussian membership function, and the left and right bias standard deviations are calculated independently. The initial parameters are determined based on the median and quantile of the historical processing time data set.

[0102] Taking the boring process of workshop CNC machine tool M01 as an example, the historical data set is 45 minutes, 48 minutes, 50 minutes, 52 minutes, and 55 minutes. The mean value μ is initialized to 50 minutes, the left bias standard deviation σ1 is 2.8 minutes, and the right bias standard deviation σ r is 3.5 minutes. The virtual layer dynamically updates the parameters through the sliding window Kalman filter algorithm. When the device state changes or the prediction error exceeds 5%, the update process is started. The sliding window size is set to 80 groups of data, and the process noise covariance Q and measurement noise covariance R are configured as unit matrices of 0.01 and 0.1 respectively.

[0103] The decision layer deploys a deep reinforcement learning strategy engine, uses a meta-learning initialization unit to pre-train the model, stores a history optimal strategy hash table in a dynamic strategy library, and generates a reward signal according to the difference between the predicted completion time and the actual completion time.

[0104] The closed-loop optimization module monitors the scheduling execution effect in real time, including equipment utilization and order delay rate.

[0105] When the processing time exceeds the right bias standard deviation boundary, for example, the actual time of a certain process is 92 minutes, which exceeds the μ+1.5σ r boundary value of 85 minutes, the system automatically triggers the task migration mechanism, assigns the task to an idle digital twin node such as CNC machine tool M04, and reserves a pessimistic buffer boundary through the weighted integral de-fuzzification algorithm.

[0106] The weighted integral calculation mode is to perform probability density integral on the time value exceeding the boundary, and the integral threshold is set to μ+1.5σ r .

[0107] The monitoring data is fed back to the virtual layer through the HTTP API. If the device utilization is lower than 85% or the order delay rate is higher than 10%, the secondary update of the parameters and the retraining of the strategy are triggered.

[0108] During the system operation, the data acquisition module detects that the M01 spindle vibration of the CNC machine tool is abnormal, and the amplitude reaches 5.2 mm / s2. After data cleaning, the device state change signal is generated.

[0109] The physical layer synchronizes the signal to the virtual layer. The virtual layer determines that the actual processing time 120 minutes exceeds the current right boundary value 80 minutes, and immediately starts the parameter update. The right deviation σ r is expanded from 15 minutes to 25 minutes.

[0110] The updated parameters are input into the decision layer, and the deep reinforcement learning strategy engine calculates the cosine similarity between the current working condition and the historical record of the strategy library. When the similarity is lower than 85%, the meta-learning mechanism is called, a new strategy is generated based on 20 groups of fault samples, and the scheduling instruction is output to stop the CNC machine tool M01 and distribute the pending orders to the CNC machine tool M03.

[0111] The closed-loop optimization module monitors that the device utilization rate rebounds to 90%, and feeds back the actual scheduling effect to the virtual layer, triggering the σ r is fine-tuned from 25 minutes to 22 minutes.

[0112] Through the flexible job shop uncertain processing time optimization system of the embodiment, the processing time prediction error can be reduced, the device utilization rate can be improved, the order delay rate can be reduced, and the scheduling robustness problem under dynamic disturbance can be effectively solved.

Claims

1. A method for optimizing uncertain processing time in a flexible job shop, characterized by: The following steps are involved: S1, real-time collection of workshop equipment status and processing time data, synchronized to the digital twin layer to build a dynamic fuzzy Gaussian model; S2, based on the historical processing time dataset, initializes and binds the parameters of the dynamic fuzzy Gaussian model to the digital twin; S3, updates the parameters of the dynamic fuzzy Gaussian model according to the device status change; S4, inputs the updated parameters into the deep reinforcement learning strategy engine to generate scheduling instructions, executes scheduling and triggers closed-loop optimization.

2. The method for optimizing uncertain processing time in a flexible job shop according to claim 1, characterized in that: In step S2, the dynamic fuzzy Gaussian number model adopts an asymmetric Gaussian membership function to independently characterize the conservative distribution and pessimistic distribution characteristics of the processing time through the left-skewed standard deviation and the right-skewed standard deviation respectively.

3. A flexible job shop uncertain processing time optimization method according to claim 1 or 2, characterized in that: In step S2, the parameter initialization of the dynamic fuzzy Gaussian model includes: determining the mean according to the median of the historical processing time data set, determining the left-skewed standard deviation based on the lower quantile calculation, and determining the right-skewed standard deviation based on the higher quantile calculation.

4. The method for optimizing uncertain processing time in a flexible job shop according to claim 1, characterized in that: The step S3 comprises: S31, calculating the prediction error within the sliding window at fixed time intervals; S32: When the prediction error exceeds a preset threshold, the parameter update of the dynamic fuzzy Gaussian model is started.

5. The method for optimizing uncertain processing time in a flexible job shop according to claim 1 or 4, characterized in that: In step S3, the update includes: when the prediction error exceeds the threshold, the sliding window Kalman filter algorithm is used to dynamically correct the model; when a change in the device state is detected, the sliding window Kalman filter is immediately started to correct the model; the corrected parameters must meet the monotonicity constraint of the asymmetric Gaussian function.

6. The method for optimizing uncertain processing time in a flexible job shop according to claim 1, characterized in that: In step S4, the action space of the deep reinforcement learning strategy engine is designed to bind order combinations and transition sets, wherein the order combinations support parallel processing of multiple orders, and the transition sets map orders to specific device operations.

7. The method for optimizing uncertain processing time in a flexible job shop according to claim 1, characterized in that: In step S4, generating a scheduling instruction includes: calculating the similarity between the current working condition and the historical strategy. If it exceeds the matching threshold, the historical strategy is loaded; otherwise, a new strategy is generated based on a small sample through a meta-learning mechanism and the strategy library is updated.

8. The method for optimizing uncertain processing time in a flexible job shop according to claim 1 or 7, characterized in that: The step S4 includes: updating the DFGN parameters every 30 minutes to trigger strategy retraining.

9. A flexible job shop uncertain processing time optimization system, which implements the flexible job shop uncertain processing time optimization method according to any one of claims 1 to 8, characterized in that: include: a data acquisition module configured to collect device data in real time through the Industrial Internet of Things; The digital twin module includes: a physical layer that synchronizes device status data; a virtual layer that integrates a dynamic fuzzy Gaussian model and performs parameter updates; and a decision layer that deploys a deep reinforcement learning strategy engine to generate scheduling actions. The closed-loop optimization module is configured to monitor the scheduling effect and feed back to the virtual layer to trigger parameter updates.

10. The flexible job shop uncertain processing time optimization system according to claim 9, characterized in that: The virtual layer is further configured to automatically migrate tasks to other idle digital twin nodes when the processing time exceeds the right-skewed standard deviation boundary, and reserve a pessimistic boundary through weighted integral defuzzification.

Citation Information

Patent Citations

  • Digital twin virtual-real adaptive iterative optimization method for product job dynamic scheduling

    CN111445081A

  • A rolling optimization system and method for production line simulation based on digital twins

    CN113361139B