A flexible job shop active fault-tolerant scheduling method based on multi-perspective transfer learning

Through the flexible job shop active fault-tolerant scheduling method based on multi-perspective transfer learning and deep reinforcement learning, the inaccuracy problem of the scheduling system caused by data missing and equipment failure is solved, efficient and flexible production scheduling and privacy protection are achieved, and production efficiency and equipment utilization are improved.

CN119472539BActive Publication Date: 2025-10-03BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411587458.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-10-03
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

In dynamic workshop scheduling, data missing and equipment failures cause the scheduling system to be unable to accurately perceive changes in the production environment, affecting production efficiency and equipment utilization. Existing technologies make it difficult to achieve flexible scheduling and privacy protection in complex and changing production environments.

Method used

An active fault-tolerant scheduling method for flexible job shops adopts multi-perspective transfer learning. Through multi-perspective feature extraction, transfer learning and deep reinforcement learning, combined with data interpolation and fault analysis, it optimizes the scheduling process, realizes data completion and fault identification, and introduces a privacy protection mechanism.

Benefits of technology

It improves the reliability and accuracy of the scheduling system, improves production efficiency and equipment utilization, enhances the model's ability to protect data privacy, adapts to complex dynamic environments, and reduces the risk of production interruption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119472539B_ABST
    Figure CN119472539B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for active fault-tolerant scheduling of flexible job shops based on multi-perspective transfer learning. The method collects time series data of each device in the shop in real time, performs data cleaning and standardization through pre-processing; predicts and fills missing or abnormal device status data through an industrial data interpolation model based on gradient penalty weighted conditions to achieve data completion; and uses the model to perform data prediction on the future state of the system; utilizes a task migration module based on traceless transformation combined with fault analysis technology to identify potential problems; utilizes a state calculation module to calculate the specific state of the intelligent agent in the environment state space based on the complete time series data after interpolation; establishes a dynamic scheduling model for the flexible job shop, and based on a multi-perspective enhanced deep reinforcement learning scheduling module, selects actions according to the state characteristics of the current production environment and generates a scheduling plan. The present invention can improve scheduling efficiency and fault tolerance in flexible production environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of smart factories, and in particular relates to a flexible job shop active fault-tolerant scheduling method based on multi-perspective transfer learning. Background Art

[0002] With the continuous advancement of advanced manufacturing technology, the manufacturing industry is gradually transforming into service-oriented manufacturing, and workshop production tasks are becoming smaller and more customized. In this context, traditional large-volume orders are gradually decreasing, while the arrival time and quantity of personalized orders are becoming more random and unpredictable. In this dynamic environment, where the number and timing of production orders are difficult to predict, companies' focus has also shifted, from a simple focus on overall completion time to more complex production management objectives, such as ensuring product delivery on time and improving the overall utilization efficiency of factory equipment. With the increasing demand for personalized and customized products, companies need more flexible and responsive scheduling systems. Therefore, how to comprehensively consider multiple production objectives and formulate scheduling strategies in a dynamic environment has become a key issue in the smart factory field.

[0003] Missing data is a common challenge in dynamic shop floor scheduling, especially when sensors fail to provide timely data due to malfunctions or other reasons, which can impact the scheduling system's decision-making. In smart factories, the uncertainty and dynamic nature of the production environment pose significant challenges to scheduling methods. Scheduling systems must not only collect and process a variety of production data in real time but also continuously gather critical information such as equipment status and task progress to adjust scheduling strategies in real time based on changes in the production environment. Scheduling methods must be highly flexible and responsive to changes in production conditions or task orders, particularly to avoid efficiency losses or even downtime caused by untimely adjustments. In dynamic scheduling, scheduling systems must strike an appropriate balance between responsiveness and decision accuracy. Relying solely on real-time data for decision-making can lead to inaccurate perception of equipment status in the event of sensor failures or data loss, hindering timely response to drastic changes in the production environment. Summary of the Invention

[0004] To address these issues, this paper proposes a multi-perspective transfer learning-based active fault-tolerant scheduling method for flexible job shops, aiming to improve scheduling efficiency and fault tolerance in flexible production environments. By leveraging multi-perspective feature extraction, transfer learning, and deep reinforcement learning, this method optimizes the scheduling process by addressing downtime caused by equipment failures, incomplete information caused by missing data, and dynamic changes in production demand.

[0005] To achieve the above-mentioned purpose, the technical solution adopted by the present invention is: a flexible job shop active fault-tolerant scheduling method based on multi-perspective transfer learning, comprising the following steps:

[0006] S10, collects time series data of each device in the workshop in real time, and uses data preprocessing technology to clean and standardize the data;

[0007] S20, through the industrial data interpolation model of the weighted conditional generative adversarial network based on gradient penalty, predicts and fills in missing or abnormal equipment status data to achieve collected data completion; and uses this model to predict the future state of the system;

[0008] S30 uses the task migration module based on untraceable transformation and fault analysis technology to identify potential problems. The state calculation module calculates the specific state of the agent in the environment state space based on the complete time series data after interpolation, providing a basis for scheduling decisions.

[0009] S40, establish a dynamic scheduling model for a flexible job shop, and based on this, use a multi-perspective enhanced deep reinforcement learning scheduling module to select actions according to the state characteristics of the current production environment and generate a scheduling plan.

[0010] Furthermore, the industrial data interpolation model based on the gradient-penalized weighted conditional generative adversarial network adopts a variational autoencoder structure as the generator, establishes a binary classification network based on a multi-layer perceptron as the discriminator, and adopts the gradient-penalized weighted conditional generative adversarial network as the training framework of the entire model, and optimizes the parameters of the generator and discriminator through the Wasserstein distance.

[0011] Industrial data interpolation model based on gradient-penalized weighted conditional generative adversarial network, including:

[0012] S21, encodes the production task type or equipment type data and converts it into a hint vector in binary format; these hint vectors are mapped through the pre-coding parameter matrix to obtain a parameter expression suitable for model processing, laying the foundation for subsequent feature extraction.

[0013] S22, multi-view feature extraction: Based on the encoded data, a multi-view feature extraction module is used to enhance the data representation capability;

[0014] S23, spatiotemporal feature encoding: After feature extraction, the spatiotemporal feature encoding module is used to capture the temporal dependency and spatial correlation of the data.

[0015] In S24, based on the spatiotemporally encoded features, the Transformer decoder is used to generate interpolated data and perform nonlinear transformations on the features to match the original data dimensions, completing the missing value completion. The decoding process includes predictive interpolation, which is used to complete future data, and restorative interpolation, which is used to complete historical data.

[0016] Furthermore, when using the multi-view feature extraction module to enhance data representation capabilities: a random mask layer with multiple mask patterns is constructed to simulate different data missing situations, and the masked data is subjected to feature extraction through a feature representor composed of a long short-term memory network and a linear layer stack, and the linear attention mechanism and a multi-layer perceptron are further used for feature aggregation to form a multi-view feature representation.

[0017] Furthermore, the temporal dependency and spatial correlation of the data are captured through the spatiotemporal feature encoding module: in the temporal dimension, stacked long short-term memory networks are used to process sequence data to model the temporal changes of the data; in the spatial dimension, convolution and Transformers structures are used to extract features from the data; finally, the input sequence is spliced ​​and fused with temporal features and spatial features in a given dimension to achieve joint encoding of multi-dimensional information, providing a comprehensive representation of spatiotemporal information for the interpolation model.

[0018] Furthermore, the Transformer decoder is used to generate interpolated data. The decoding process includes predictive interpolation and data recovery interpolation. The former is used to complete future data, and the latter is used to complete historical data.

[0019] Furthermore, in step S30, the core of the task migration module is to utilize the time series features learned in the data interpolation model and apply them to the fault diagnosis task; then, by introducing the traceless transformation technology, the latent features in the source task are transferred to the target task, thereby achieving adaptation to the fault analysis task;

[0020] S31, feature transfer based on untraceable transformation: The spatiotemporally encoded features are used as the source task features and a set of sigma points are generated through untraceable transformation to capture the variation of nonlinear features across different tasks. The generation of sigma points is based on the latent feature distribution of the interpolation task, including generating a central sigma point at the mean position and generating positive and negative offset sigma points based on the covariance matrix.

[0021] S32, fault identification: Using the generated sigma point set, the task migration module transfers the interpolated features to the fault diagnosis decoder to perform fault identification and output the identification results.

[0022] Furthermore, in step S40, a flexible job shop dynamic scheduling model is established, including:

[0023] Basic assumptions: Each machine can only process one operation at any time, and all workpieces must be completed in a predetermined order; each operation is performed on a specific subset of machines, and the processing time is fixed;

[0024] Constraint equations: Based on the basic assumptions, they are mathematically transformed into the model's constraint equations. First, using equipment utilization as a measure of failure probability, a constraint on the repair time after a failure is set for each device. Furthermore, the module replacement time between devices or workpieces is set as a fixed constant and incorporated into the processing time constraint. A workpiece urgency parameter is set, and the model prioritizes high-urgency workpieces during the optimization process.

[0025] Multi-objective optimization: Under the premise of satisfying the above constraint equations, multi-objective optimization is used to achieve the optimal scheduling effect; the optimization goals are: minimizing the total completion time and delay time, and maximizing equipment utilization.

[0026] Furthermore, in step S40, the multi-view enhanced deep reinforcement learning scheduling module includes:

[0027] S41, Modeling: Convert the flexible job shop dynamic scheduling model into a Markov decision process and set the state space, action space and reward function;

[0028] S42, state feature extraction: To capture the multi-dimensional state features in the production environment, the state space is constructed using a multi-view feature enhancement module. The multi-view feature enhancement module takes the information in the state space established in step S41 as input, extracts feature vectors using a linear layer, and captures the semantic dependencies between these features using a Transformer network. After feature enhancement, a multi-view state representation is generated.

[0029] S43, Adversarial Deep Q Network Model Decision-Making Process: The above multi-perspective state representation is directly used as the input of the Adversarial Deep Q Network Model. The model calculates the Q value to evaluate the pros and cons of different actions by evaluating the value of the current state and the advantages of the action.

[0030] Furthermore, the reward function includes:

[0031] Input: Average delay time at the current moment Maximum delay time Equipment utilization and total construction period With the next moment

[0032] Process: 1. Judgment and Is it true? If so, mark it as delayed optimization; 2. Is it true? If so, mark the equipment utilization rate as optimized; 3. Judge, Is it true? If so, mark the total construction period as optimized. 4. Accumulate the corresponding reward value according to the status of each optimization flag. If all optimization flags are True, add an additional reward. If all optimization flags are False, impose a penalty.

[0033] Output: The final calculated reward value.

[0034] Furthermore, the method further includes step S50: introducing a differential privacy mechanism in the online parameter optimization process of the scheduling strategy, and ensuring model performance while enhancing data privacy protection capabilities through differential privacy noise perturbation and differential privacy stochastic gradient descent strategies;

[0035] Differentially private optimization of online scheduling policies, including:

[0036] Differentially private stochastic gradient descent: During parameter optimization, the gradient is clipped and Gaussian noise is added to ensure that differential privacy requirements are met during each model update.

[0037] Privacy perturbation of input data: In the forward propagation phase, Gaussian noise is added to the input features to control the global sensitivity of the data and enhance privacy protection.

[0038] The beneficial effects of adopting this technical solution are:

[0039] The present invention proposes an active fault-tolerant scheduling method for a flexible job shop based on multi-perspective transfer learning. First, the use of data interpolation technology and fault analysis methods can effectively solve the problems of data missing and equipment failure in the production process, thereby improving the reliability and accuracy of the scheduling system. Secondly, the multi-perspective deep reinforcement learning model is used for scheduling optimization to achieve a reasonable allocation of production resources and improve production efficiency. In addition, a privacy protection mechanism is introduced. By introducing differential privacy noise perturbation and differential privacy stochastic gradient descent strategy in parameter updates, the model's ability in data privacy protection is enhanced while ensuring model performance. Finally, the method shows strong adaptability in a dynamic production environment, can reduce the risk of production interruption due to equipment failure, and improve equipment utilization. These advantages enable the present invention to effectively improve the overall efficiency and stability of the system when dealing with complex scheduling tasks in flexible job shops.

[0040] The present invention can improve the integrity and accuracy of data processing: a data interpolation model based on a gradient-penalty-based weighted conditional generative adversarial network is used to fill in the missing parts of production data. At the same time, through multi-perspective feature extraction, the high-dimensional representation capability of the data is enhanced, providing more accurate data support for subsequent scheduling optimization. The present invention proposes an industrial data interpolation model based on self-supervised learning, which adopts a gradient-penalty-based weighted conditional generative adversarial network training framework and combines a multi-perspective embedding module to enrich the feature representation of self-supervised learning. Finally, a multi-task learning mechanism with dynamic weights is introduced, thereby improving the stability of the model in processing multi-variable industrial missing data interpolation tasks. The interpolation model of the present invention has a high degree of responsiveness and can quickly generate reasonable data completion and prediction of future data when data missing occurs; in addition, it can cope with the frequent changes in task orders in the workshop, which puts higher requirements on the robustness of the interpolation model; the interpolation method can fully analyze historical data and combine information on the current production environment to accurately predict missing data, thereby adapting to frequent changes in production tasks.

[0041] The present invention designs a working condition and equipment mode coding mechanism. By pre-coding the working condition information and equipment operation mode into a preset matrix and introducing it into the model as prompt conditions, the model's adaptability to data distribution under different working conditions is enhanced.

[0042] This invention enables knowledge transfer and reuse: Utilizing a task migration module, learned feature representations are shared across different tasks, enabling knowledge transfer between tasks ranging from data interpolation to fault diagnosis, and enhancing the model's adaptability in practical industrial environments. Building on the existing model, this invention introduces a task migration module based on traceless transformation. This module achieves migration from data interpolation to time series feature analysis tasks by generating Sigma points, promoting information sharing and reuse across different tasks, improving the model's adaptability and reuse rate, and reducing the manpower and time investment required by enterprises to develop new models for different tasks, thereby lowering costs.

[0043] The present invention can optimize production scheduling strategies: Based on a multi-perspective deep reinforcement learning scheduling module, production resources and tasks are optimally allocated to ensure that production efficiency can be maximized even when the production environment changes due to equipment failures. This module combines the adversarial deep Q network and the multi-perspective Transformers feature enhancement module to improve the ability to understand the production status, thereby making better scheduling decisions. The present invention proposes a representation learning module based on multi-perspective Transformers, which improves the decision-making ability of the scheduling algorithm in complex dynamic environments by performing multi-level modeling of state space features and capturing global dependencies, and inputting them as high-level states into the DQN, thereby ensuring that it can still maintain efficient and stable operation in complex and changing production environments.

[0044] This paper introduces a privacy protection mechanism. By introducing differential privacy noise perturbation and differential privacy stochastic gradient descent strategy in parameter update, it enhances the model's ability in data privacy protection while ensuring model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 Schematic diagram of the framework principle of a flexible job shop active fault-tolerant scheduling method based on multi-perspective transfer learning of the present invention;

[0046] Figure 2 Schematic diagram of data recovery interpolation and predictive interpolation model in an embodiment of the present invention;

[0047] Figure 3 Schematic diagram of a multi-view missing data feature embedding module in an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described below with reference to the accompanying drawings.

[0049] In this embodiment, see Figure 1 As shown, the present invention proposes a flexible job shop active fault-tolerant scheduling method based on multi-perspective transfer learning, including the following steps:

[0050] S10, collects time series data of each device in the workshop in real time, and uses data preprocessing technology to clean and standardize the data;

[0051] S20, through the industrial data interpolation model of the weighted conditional generative adversarial network based on gradient penalty, predicts and fills in missing or abnormal equipment status data to achieve collected data completion; and uses this model to predict the future state of the system;

[0052] S30 uses the task migration module based on untraceable transformation and fault analysis technology to identify potential problems. The state calculation module calculates the specific state of the agent in the environment state space based on the complete time series data after interpolation, providing a basis for scheduling decisions.

[0053] S40, establish a dynamic scheduling model for a flexible job shop, and based on this, use a multi-perspective enhanced deep reinforcement learning scheduling module to select actions according to the state characteristics of the current production environment and generate a scheduling plan.

[0054] As an optimization scheme for the above embodiment, the industrial data interpolation model based on the weighted conditional generative adversarial network with gradient penalty adopts a variational autoencoder (VAE) as a generator, establishes a binary classification network based on a multi-layer perceptron (MLP) as a discriminator, and adopts the weighted conditional generative adversarial network with gradient penalty as the training framework of the entire model, and optimizes the parameters of the generator and discriminator through the Wasserstein distance.

[0055] Specifically, the industrial data interpolation model based on the weighted conditional generative adversarial network with gradient penalty is Figure 2 Shown, including:

[0056] S21, encodes the production task type or equipment type data and converts it into a hint vector in binary format; these hint vectors are mapped through the pre-coding parameter matrix to obtain a parameter expression suitable for model processing, laying the foundation for subsequent feature extraction.

[0057] First, define the same data domain in the training dataset Different production processes or equipment types, denoted as For each class i , which is converted into Binary vector Defined as:

[0058]

[0059] in, Represents a vector In this way, each category is ensured to have a unique representation in the encoding vector, avoiding the impact of potential order relationships between categories on model training.

[0060] is the prompt vector after one-hot encoding Converted into a continuous parameter representation that is more suitable for model processing, the present invention introduces the precoding parameter matrix The mapping process can be expressed as:

[0061]

[0062] in It is the parameter expression vector after mapping, which is used for subsequent model processing.

[0063] Then, the input data is normalized to form a standardized data set using the Z normalization method, as shown in the following formula:

[0064]

[0065] Among them, μimpu is the mean of the input data, σ impu is the variance of the input data. This industrial preprocessing method ensures that unknown industrial input data conforms to the normal distribution assumption, meeting the prerequisites for VAE. Precoding production tasks or equipment types and preprocessing input data help the model acquire more representative features from the sample during training, thereby improving the model's performance in interpolating data in the context of personalized customization.

[0066] S22, multi-view feature extraction: Based on the encoded data, a multi-view feature extraction module is used to enhance the data representation capability.

[0067] Specifically, when using the multi-view feature extraction module to enhance data representation capabilities: construct a random mask layer with multiple mask patterns to simulate different data missing situations, extract features from the masked data through a feature representor stacked by LSTM and linear layers, and further use the linear attention mechanism and MLP for feature aggregation to form a multi-view feature representation.

[0068] In order to enhance the generalization ability of the model under limited training samples and thus improve the accuracy of data interpolation, this paper proposes a multi-view feature extraction module to enrich the data representation ability. Several different mask patterns are constructed through the random mask layer to simulate various data missing situations. The time window size is Input sample Through random mask operation Mask impu , obtain a set of industrial mask data under different mask conditions,

[0069]

[0070] This will allow us to obtain more observation data samples that retain the original temporal structure and pattern characteristics.

[0071] Then, input the masked industrial data into Figure 3 The multi-view feature extraction module shown in the figure is composed of several feature representation heads and feature aggregation heads. This module is composed of several feature representers and feature aggregators. Each feature representer uses a long short-term memory network (LSTM) stacked with linear layers to process input data under different mask conditions and extract its feature representation. The fused feature representation vector is obtained through the cat operation. This process can be expressed as:

[0072]

[0073] in, Represents the i-th feature extractor. The feature aggregator consists of a linear attention layer and an MLP

[0074] The concatenated feature representations are weighted and aggregated through the linear attention mechanism to obtain the weighted feature vector.

[0075] Subsequently, the aggregated feature vectors are transformed nonlinearly through MLP to achieve dimensionality control. This process can be expressed as:

[0076]

[0077] The motivation for this strategy is that, in practical applications, transfer fine-tuning methods that add linear layers to the original model to reshape the feature dimension often reduce the model's expressive power and computational efficiency. By adopting a multi-view feature extraction module, these problems can be effectively avoided while enhancing the richness and diversity of feature representation.

[0078] S23, spatiotemporal feature encoding: After feature extraction, the spatiotemporal feature encoding module is used to capture the temporal dependency and spatial correlation of the data.

[0079] Specifically, the temporal dependency and spatial correlation of data are captured through the spatiotemporal feature encoding module: in the temporal dimension, stacked LSTM is used to process sequence data to model the temporal changes of the data; in the spatial dimension, convolution and Transformers structures are used to extract features from the data; finally, the input sequence is spliced ​​and fused with temporal features in a given dimension to achieve joint encoding of multi-dimensional information, providing a comprehensive representation of spatiotemporal information for the interpolation model.

[0080] In order to capture the temporal structure dependency while exploring the spatial correlation between the measurement data of the same device, this paper designs a spatiotemporal feature extraction module. This module uses a stacked LSTM to capture temporal dependencies. At the same time, a convolution-Transformers structure is used to aggregate spatial features. For the time dimension, a stacked LSTM structure is used to process the input sequence data to generate a temporal feature representation. As for the spatial dimension, the present invention first uses convolution operations to capture local dependencies in space, and then further processes spatial features through Transformers. The process can be described as follows:

[0081]

[0082] in, Represents the convolution kernel, * represents the convolution operation, represents the SiLU activation function, Represents the convolution bias. The specific form of the convolution operation can be expressed as:

[0083]

[0084] Where M, N are the height and width of the convolution kernel, i, j are the spatial position indexes of the feature map. After processing by the spatiotemporal feature encoder, we can get:

[0085]

[0086] in, Represents high-dimensional time features, Represents high-dimensional space features, Represents the high-dimensional spatiotemporal features obtained through Cat operation. Then, the multi-head attention layer and MLP are used to obtain The mean of the low-dimensional representation of and variance

[0087] In S24, based on the spatiotemporally encoded features, the Transformer decoder is used to generate interpolated data and perform nonlinear transformations on the features to match the original data dimensions, completing the missing value completion. The decoding process includes predictive interpolation, which is used to complete future data, and restorative interpolation, which is used to complete historical data.

[0088] Specifically, a Transformer decoder is used to generate interpolated data. The decoding process includes predictive interpolation and data recovery interpolation. The former is used to complete future data, and the latter is used to complete historical data.

[0089] This paper uses a Transformer-based decoder structure and combines it with an MLP to build a decoding network. Specifically, the Transformer decoder first processes low-dimensional features using a multi-head attention mechanism to capture complex spatiotemporal dependencies. The MLP then converts the processed features into interpolated values ​​that match the original data dimensions. This process can be expressed as:

[0090]

[0091] The advantage of this design is that the MLP dimensions can be dynamically adjusted to suit different downstream tasks, thereby improving the model's flexibility and generalization capabilities. Furthermore, this design enhances the model's ability to handle complex spatiotemporal dependencies, helping to improve interpolation accuracy.

[0092] In the training process of the industrial data interpolation model of the gradient-penalized weighted conditional generative adversarial network: the gradient-penalized weighted conditional generative adversarial network is used as the training framework of the generator, aiming to optimize the adversarial training of the generator and the discriminator, and solve the gradient vanishing problem that may occur in the generator during the training process. On the basis of the traditional WCGAN, the gradient penalty technology is introduced to further enhance the stability and training effect of the model. During the entire training process, the present invention generates reconstructed labels through random masks, and combines the construction mechanism of predicted labels to realize self-supervised learning based on time series data. The advantage of self-supervised learning is that no additional manual labeling is required, and training samples can be generated only through some missing information in the original data, which reduces the dependence on large-scale labeled data and can effectively improve the generalization ability of the model and the accuracy of data interpolation.

[0093] The core of the training strategy of the present invention is to use VAE as a generator while introducing a weighted conditional generative adversarial network with gradient penalty. The introduction of the gradient penalty strategy enhances the generation ability of the generator and the discrimination ability of the discriminator. Under this framework, VAE, as a generator, can more effectively learn the latent distribution of time series data and ensure that the generated data has reasonable time series consistency and accuracy. As an extended model of AE, VAE introduces a probabilistic generation mechanism and regularization constraints on latent variables on the basis of traditional AE, thereby improving the model's ability to model data distribution. The latent encoding of VAE not only contains the mean and variance, but also ensures the structured distribution of the latent space through regularization constraints, providing favorable conditions for knowledge transfer. The distribution of the latent variable z is regularized by KL divergence to make it close to the prior distribution, thereby making knowledge transfer between different tasks or equipment types possible. In this study, by sharing the structure of the latent space, knowledge transfer between industrial time series data and tasks can be achieved, improving the model knowledge reuse rate.

[0094] However, VAE faces problems such as weak temporal consistency of generated data and inaccurate modeling of temporal dynamics in industrial time series data interpolation tasks. To solve this problem, this paper introduces an adversarial generation mechanism. Through WCGAN, the generator can minimize the real data distribution p data and generate data distribution p g While ensuring the accuracy of data interpolation, WCGAN can provide more stable gradients than traditional GANs, avoiding mode collapse and gradient instability problems, thereby enabling more efficient training of the generator.

[0095] As an optimization solution for the above embodiment, in step S30, the core of the task migration module is to utilize the time series features learned in the data interpolation model and apply them to the fault diagnosis task; then, by introducing the traceless transformation technology, the latent features in the source task are transferred to the target task, thereby achieving adaptation to the fault analysis task;

[0096] S31, feature transfer based on untraceable transformation: The spatiotemporally encoded features are used as the source task features and a set of sigma points are generated through untraceable transformation to capture the variation of nonlinear features across different tasks. The generation of sigma points is based on the latent feature distribution of the interpolation task, including generating a central sigma point at the mean position and generating positive and negative offset sigma points based on the covariance matrix.

[0097] S32, fault identification: Using the generated sigma point set, the task migration module transfers the interpolated features to the fault diagnosis decoder to perform fault identification and output the identification results.

[0098] In task transfer, the present invention focuses on how to share knowledge between different but related tasks. Specifically, although the source task and the target task have different target output spaces, they share similar input feature spaces and feature representations. and target tasks in and denote the input space of the source task and the target task respectively, Respectively represent the corresponding output space. Generally speaking, in the context of the task of this invention, the input space is the same, but from the perspective of rigor, it is assumed that the two are similar, that is, So they can share the same feature encoder The source task model trained in the data imputation task can be expressed as:

[0099]

[0100] In order to transfer the knowledge learned by the original encoder to the fault diagnosis task, the present invention adds a new decoder for the target task. However, considering that the potential space distribution of different tasks may be different, it is not advisable to directly use the latent variables generated by the encoder. It is impossible to achieve the ideal effect in the fault diagnosis task. In order to solve this problem, the present invention introduces the unscented transformation as the core idea of ​​the unscented Kalman filter. The unscented transformation can more accurately estimate the statistical characteristics of the random variable after nonlinear transformation by generating a set of sigma points without calculating high-order derivatives. For n-dimensional independent and identically distributed latent variables Its overall mean is The covariance matrix is The steps to generate the sigma point are:

[0101] 1. Generate a sigma point at the mean position

[0102] 2. Generate 2n offset points based on the covariance matrix, where n are positive offset points and the other n are negative offset points. These offset points are generated by transforming the square root of the covariance matrix, as shown below:

[0103]

[0104] in represents the square root of the covariance matrix, λ it Indicates that the scaling parameter is used to control the distribution range of sigma points.

[0105] Through the above steps, the present invention obtains a richer data representation to capture the potential feature changes between the source domain and the target domain. Input to the new decoder In the task, the fault diagnosis results of the target domain are obtained. At this time, the loss function of task migration can be expressed as:

[0106]

[0107] in, Represents the loss function for the fault diagnosis task. This migration module enables the reuse of temporal feature knowledge learned in data interpolation tasks, enhancing the model's adaptability between different tasks within the same domain (e.g., from data interpolation tasks to fault diagnosis tasks), thereby increasing the model's application value in real industrial environments.

[0108] As an optimization solution of the above embodiment, in step S40, a flexible job shop dynamic scheduling model is established, including:

[0109] Basic assumptions: Each machine can only process one operation at any time, and all workpieces must be completed in a predetermined order; each operation is performed on a specific subset of machines, and the processing time is fixed;

[0110] Constraint equations: Based on basic assumptions, they are mathematically transformed into the constraint equations of the model. First, using equipment utilization as an indicator to measure the probability of failure, a constraint on the post-failure repair time is set for each device to ensure that the equipment can resume normal operation according to the scheduled repair time after a failure, thereby reducing the impact of failures on scheduling. At the same time, the module replacement time between devices or workpieces is set as a fixed constant and incorporated into the processing time constraint to improve scheduling accuracy and stability. The urgency parameter of the workpiece is set, and the model gives priority to high-urgency workpieces during the optimization process, thereby improving the scheduling response capability to emergency tasks.

[0111] Multi-objective optimization: Under the premise of satisfying the above constraint equations, multi-objective optimization is used to achieve the optimal scheduling effect; the optimization goals are: minimizing the total completion time and delay time, and maximizing equipment utilization.

[0112] In step S40, the multi-view enhanced deep reinforcement learning scheduling module includes:

[0113] S41, Modeling: Convert the flexible job shop dynamic scheduling model into a Markov decision process and set the state space, action space and reward function;

[0114] First, the state space comprehensively models the dynamic state of each device and job in the system, including multi-dimensional information such as equipment availability, job progress, and process requirements. Second, the action space defines the scheduling decision rules that the system can adopt under given states to ensure the successful completion of scheduling tasks. The reward function incorporates multi-objective optimization requirements, comprehensively considering factors such as production cycle time, equipment utilization, and job delays, ensuring the model's adaptability to multi-objective scheduling tasks.

[0115] (1) State space:

[0116] The state space is the core element for the agent to perceive the workshop production environment and is the basis for the agent to make accurate scheduling decisions. Therefore, when designing the state space, this paper adheres to the principle of combining simplicity and effectiveness, striving to use a minimum number of state features to describe the production process of a dynamic and flexible job shop. This reduces model complexity while still meeting scheduling objectives. As shown in Table 1, this paper selects 12 important features as state characteristics.

[0117] Table 2 State space description

[0118]

[0119] (2) Action Space

[0120] This paper designs seven different scheduling rules for multi-objective optimization problems, each prioritizing jobs based on different criteria. This approach incorporates prior knowledge into the action space design, simplifying the complexity of action selection. It should be noted that an overly redundant action space increases learning complexity and computational burden. The designed action spaces are shown in Table 2.

[0121] Table 2 Action space description

[0122]

[0123]

[0124] (3) Reward Function

[0125] To guide intelligent agents in making optimal scheduling decisions in a dynamic and flexible job shop environment, this paper designs a multi-objective optimization reward function based on multi-objective reinforcement learning theory. This reward function provides a structured reward mechanism by evaluating tardiness, equipment utilization, and total construction duration. Its design, inspired by the Pareto optimality theory, seeks a balance between multiple conflicting objectives, encouraging intelligent agents to explore optimal scheduling strategies.

[0126] The reward function includes:

[0127] Input: Average delay time at the current moment Maximum delay time Equipment utilization and total construction period With the next moment

[0128] Process: 1. Judgment and Is it true? If so, mark it as delayed optimization; 2. Is it true? If so, mark the equipment utilization rate as optimized; 3. Judge, Is it true? If so, mark the total construction period as optimized. 4. Accumulate the corresponding reward value according to the status of each optimization flag. If all optimization flags are True, an additional reward is added; if all optimization flags are False, a penalty is imposed.

[0129] Output: The final calculated reward value.

[0130] S42, state feature extraction: To capture the multi-dimensional state features in the production environment, the state space is constructed using a multi-view feature enhancement module. The multi-view feature enhancement module takes the information in the state space established in step S41 as input, extracts feature vectors using a linear layer, and captures the semantic dependencies between these features using a Transformer network. After feature enhancement, a multi-view state representation is generated.

[0131] S43, Adversarial Deep Q Network (Dueling DQN model) decision process: The above multi-perspective state representation is directly used as the input of the Dueling DQN model. The model calculates the Q value to evaluate the pros and cons of different actions by evaluating the value of the current state and the advantages of the action.

[0132] As an optimization solution for the above embodiment, in the parameter optimization process of deep reinforcement learning, the present invention also introduces a privacy protection mechanism, combined with differential privacy technology, to ensure the security protection of production feature data during the training process, thereby improving the security of the model in actual industrial scenarios.

[0133] The method further includes step S50: introducing a differential privacy mechanism during the online parameter optimization process of the scheduling strategy, and ensuring model performance while enhancing data privacy protection capabilities through differential privacy noise perturbation and differential privacy stochastic gradient descent strategies;

[0134] Differentially private optimization of online scheduling policies, including:

[0135] Differentially private stochastic gradient descent: During parameter optimization, the gradient is clipped and Gaussian noise is added to ensure that differential privacy requirements are met during each model update.

[0136] Privacy perturbation of input data: In the forward propagation phase, Gaussian noise is added to the input features to control the global sensitivity of the data and enhance privacy protection.

[0137] The present invention introduces a differential privacy mechanism to protect sensitive information in the scheduling process from being leaked during the training process. In a dynamic workshop scheduling environment, the scheduling tasks and environmental states are highly dynamic. Therefore, any two scheduling tasks can be defined as adjacent training sets. The differential privacy mechanism ensures that when the model faces these adjacent data sets, its output and parameter updates will not be obvious, thereby protecting the sensitive information of a single scheduling task. The differential privacy mechanism introduced in this method is implemented through two methods. First, in each iteration of model training, a differential privacy stochastic gradient descent algorithm is used to clip the gradient and add noise. For the gradient of the parameter θ Its clipped gradient can be expressed as:

[0138]

[0139] in, is the gradient clipping threshold. Gaussian noise is added to the clipped gradient to meet the differential privacy requirements. The process is expressed as:

[0140]

[0141] in, Represents the noise intensity parameter. The gradient update expression after adding noise is:

[0142]

[0143] Among them, γ fjsp Represents the learning rate. To further protect the privacy of the input data, Gaussian noise is added to the input data during the forward propagation of the model. The process is expressed as:

[0144]

[0145] Among them, δ fjsp ,∈ fjsp represents the differential privacy parameter, η fisp Indicates the global sensitivity of the input data.

[0146] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A flexible job shop active fault-tolerant scheduling method based on multi-perspective transfer learning, characterized by: Including steps: S10, collects time series data of each device in the workshop in real time, and uses data preprocessing technology to clean and standardize the data; S20, through the industrial data interpolation model of the weighted conditional generative adversarial network based on gradient penalty, predicts and fills in missing or abnormal equipment status data to achieve collected data completion; and uses this model to predict the future state of the system; S30 uses a task migration module based on traceless transformation and combines it with fault analysis technology to identify potential problems; Using the state calculation module, the specific state of the agent in the environment state space is calculated based on the complete time series data after interpolation, providing a basis for scheduling decisions; S40, establish a dynamic scheduling model for a flexible job shop, and based on this, use a multi-perspective enhanced deep reinforcement learning scheduling module to select actions according to the state characteristics of the current production environment and generate a scheduling plan.

2. The method for flexible job shop active fault-tolerant scheduling based on multi-perspective transfer learning according to claim 1, characterized in that: The industrial data interpolation model based on a gradient-penalized weighted conditional generative adversarial network uses a variational autoencoder as the generator and a binary classification network based on a multilayer perceptron as the discriminator. The gradient-penalized weighted conditional generative adversarial network is used as the training framework for the entire model, and the parameters of the generator and discriminator are optimized using the Wasserstein distance. Industrial data interpolation model based on gradient-penalized weighted conditional generative adversarial network, including: S21, encoding the production task type or equipment type data and converting it into a prompt vector in binary format; These hint vectors are mapped through the precoding parameter matrix to obtain a parameter expression suitable for model processing; S22, multi-view feature extraction: Based on the encoded data, a multi-view feature extraction module is used to enhance the data representation capability; S23, spatiotemporal feature encoding: After feature extraction, the spatiotemporal feature encoding module is used to capture the temporal dependency and spatial correlation of the data. S24, based on the features after spatiotemporal encoding, uses the Transformer decoder to generate interpolation data and performs nonlinear transformation on the features to match the original data dimensions to complete the missing value filling; the decoding process includes predictive interpolation and data recovery interpolation, the former is used to fill future data, and the latter is used to fill historical data.

3. The method for flexible job shop active fault-tolerant scheduling based on multi-perspective transfer learning according to claim 2, characterized in that: When using the multi-view feature extraction module to enhance data representation capabilities: construct a random mask layer with multiple mask patterns to simulate different data missing situations, extract features from the masked data through a feature representor consisting of a long short-term memory network and a stack of linear layers, and further use a linear attention mechanism and a multi-layer perceptron to aggregate features to form a multi-view feature representation.

4. The method for flexible job shop active fault-tolerant scheduling based on multi-perspective transfer learning according to claim 2, characterized in that: The temporal dependency and spatial correlation of data are captured through the spatiotemporal feature encoding module: in the temporal dimension, stacked long short-term memory networks are used to process sequence data to model the temporal changes of the data; in the spatial dimension, convolution and Transformers structures are used to extract features from the data; finally, the input sequence is spliced ​​and fused with temporal features in a given dimension to achieve joint encoding of multi-dimensional information, providing a comprehensive representation of spatiotemporal information for the interpolation model.

5. The method for flexible job shop active fault-tolerant scheduling based on multi-perspective transfer learning according to claim 2, characterized in that: The Transformer decoder is used to generate interpolated data. The decoding process includes predictive interpolation and data recovery interpolation. The former is used to complete future data, and the latter is used to complete historical data.

6. The method for flexible job shop active fault-tolerant scheduling based on multi-perspective transfer learning according to claim 1, characterized in that: In step S30, the core of the task migration module is to utilize the time series features learned in the data interpolation model and apply them to the fault diagnosis task; then, by introducing the traceless transformation technology, the latent features in the source task are transferred to the target task, thereby achieving adaptation to the fault analysis task; S31, feature transfer based on untraceable transformation: The spatiotemporally encoded features are used as the source task features and a set of sigma points are generated through untraceable transformation to capture the variation of nonlinear features across different tasks. The generation of sigma points is based on the latent feature distribution of the interpolation task, including generating a central sigma point at the mean position and generating positive and negative offset sigma points based on the covariance matrix. S32, fault identification: Using the generated sigma point set, the task migration module transfers the interpolated features to the fault diagnosis decoder to perform fault identification and output the identification results.

7. The method for flexible job shop active fault-tolerant scheduling based on multi-perspective transfer learning according to claim 1, characterized in that: In step S40, a flexible job shop dynamic scheduling model is established, including: Basic assumptions: Each machine can only process one operation at any time, and all workpieces must be completed in a predetermined order; each operation is performed on a specific subset of machines, and the processing time is fixed; Constraint equations: Based on the basic assumptions, they are mathematically transformed into the model's constraint equations. First, using equipment utilization as a measure of failure probability, a constraint on the repair time after a failure is set for each device. Furthermore, the module replacement time between devices or workpieces is set as a fixed constant and incorporated into the processing time constraint. A workpiece urgency parameter is set, and the model prioritizes high-urgency workpieces during the optimization process. Multi-objective optimization: Under the premise of satisfying the above constraint equations, multi-objective optimization is used to achieve the optimal scheduling effect; the optimization goals are: minimizing the total completion time and delay time, and maximizing equipment utilization.

8. The method for flexible job shop active fault-tolerant scheduling based on multi-perspective transfer learning according to claim 7, characterized in that: In step S40, the multi-view enhanced deep reinforcement learning scheduling module includes: S41, Modeling: Convert the flexible job shop dynamic scheduling model into a Markov decision process and set the state space, action space and reward function; S42, state feature extraction: To capture the multi-dimensional state features in the production environment, the state space is constructed using a multi-view feature enhancement module. The multi-view feature enhancement module takes the information in the state space established in step S41 as input, extracts feature vectors using a linear layer, and captures the semantic dependencies between these features using a Transformer network. After feature enhancement, a multi-view state representation is generated. S43, Adversarial Deep Q Network Model Decision-Making Process: The above multi-perspective state representation is directly used as the input of the Adversarial Deep Q Network Model. The model calculates the Q value to evaluate the pros and cons of different actions by evaluating the value of the current state and the advantages of the action.

9. The method for flexible job shop active fault-tolerant scheduling based on multi-perspective transfer learning according to claim 8, characterized in that: The reward function includes: Input: Average delay time at the current moment Maximum delay time Equipment utilization and total construction period With the next moment Process:

1. Judgment and Is it true? If so, mark it as delayed optimization; 2. Is it true? If so, mark the equipment utilization rate as optimized; 3. Judge, Is it true? If so, mark the total construction period as optimized.

4. Accumulate the corresponding reward value according to the status of each optimization flag. If all optimization flags are True, an additional reward is added; if all optimization flags are False, a penalty is imposed. Output: The final calculated reward value.

10. The method for flexible job shop active fault-tolerant scheduling based on multi-perspective transfer learning according to claim 1, characterized in that: The method further includes step S50: introducing a differential privacy mechanism during the online parameter optimization process of the scheduling strategy, and ensuring model performance while enhancing data privacy protection capabilities through differential privacy noise perturbation and differential privacy stochastic gradient descent strategies; Differentially private optimization of online scheduling policies, including: Differentially private stochastic gradient descent: During parameter optimization, the gradient is clipped and Gaussian noise is added to ensure that differential privacy requirements are met during each model update. Privacy perturbation of input data: In the forward propagation phase, Gaussian noise is added to the input features to control the global sensitivity of the data and enhance privacy protection.

Citation Information

Patent Citations

  • Multilayer factory workshop scheduling method based on reinforcement learning

    CN116594358A

  • Federal transfer learning enhanced multi-agent workshop dynamic regulation and control method

    CN117434896A