Method, device and equipment for constructing industrial data prediction model under privacy protection
Through the combination of adversarial training and average teacher algorithm, cross-domain migration of industrial data models is achieved without accessing monitoring data, solving the problems of poor security and low accuracy in the existing technology, and ensuring the privacy protection and migration effect of the model.
Patent Information
- Application Number
- CN202410873712.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-01
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-07-01
AI Technical Summary
In the prior art, the cross-domain migration of industrial data models is poor and has low accuracy, making it difficult to realize knowledge migration between prediction models of different devices without accessing monitoring data.
By obtaining multiple data of the source model and the target domain of the source domain, determining the intermediate model and conducting adversarial training on it, a second intermediate model is obtained. Then, the second intermediate model is initialized to obtain the teacher model and the student model, trained using the average teacher algorithm, and corrected the feature extractor parameters and predictor parameters of the student model, and finally the iterated student model is used as the target prediction model.
While protecting privacy, ensuring accurate cross-domain migration of industrial models improves the security and accuracy of the model.
Smart Images

Figure CN118551805B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of industrial production technology, and in particular, relates to a method, device, and equipment for constructing an industrial data prediction model under privacy protection. Background Art
[0002] Industrial production technology refers to the application of various engineering technologies and computer technologies to achieve automation, informatization and intelligence in the industrial production process. Industrial production technology is an important part of the industrial field and plays an important role in improving production efficiency, reducing costs, improving product quality, and achieving sustainable production. With the continuous development of science and technology, industrial production technology will continue to iterate, update, develop and improve, bringing more opportunities and challenges to enterprises.
[0003] By analyzing industrial monitoring signals, we can effectively predict the industrial production process and the operating status of industrial equipment to avoid safety hazards and economic losses.
[0004] However, the collection of high-quality monitoring data is difficult and expensive. The monitoring signals collected from many devices are unlabeled and cannot be directly used for time series prediction modeling. Due to information security and privacy protection considerations, the monitoring signals of many devices cannot be shared. Therefore, it is difficult to achieve knowledge transfer between prediction models of different devices without accessing the monitoring data. Summary of the invention
[0005] The present application provides a method, device, and equipment for constructing an industrial data prediction model under privacy protection, so as to solve the defects of poor security and low accuracy in cross-domain migration of industrial data models in the prior art.
[0006] In a first aspect, the present application provides a method for constructing an industrial data prediction model under privacy protection, the method comprising:
[0007] Obtain a source model of the source domain and multiple data of the target domain;
[0008] Determine, according to the source model, model parameters corresponding to the source model and an intermediate model, wherein the model structure of the intermediate model is the same as the model structure of the source model;
[0009] According to the multiple data in the target domain and the source model, adversarial training is performed on the intermediate model to obtain a second intermediate model, and the adversarial training is used to update the feature extractor of the intermediate model.
[0010] Optionally, determining, according to the source model, model parameters and an intermediate model corresponding to the source model includes:
[0011] Parsing the source model to obtain model parameters corresponding to the source model, the model parameters including: feature extractor parameters and predictor parameters of the source model;
[0012] According to the model structure of the source model, a standard model having the same model structure as the source model is constructed, and according to the feature extractor parameters and the predictor parameters, the standard model is initialized to obtain the intermediate model.
[0013] Optionally, performing adversarial training on the intermediate model according to the target domain data and the source model to obtain a second intermediate model includes:
[0014] Inputting the target domain data into the source model and the intermediate model respectively, obtaining a first prediction result output by the source model and a second prediction result output by the intermediate model;
[0015] According to the first prediction result and the second prediction result, updating the feature extractor parameters of the intermediate model and the predictor parameters of the source model to obtain an updated intermediate model and a source model;
[0016] The updated intermediate model and the source model are updated and iterated according to the target domain data until the iteration of the intermediate model is completed to obtain the second intermediate model.
[0017] Optionally, the method further includes:
[0018] Initializing the second intermediate model according to the second intermediate model to obtain a teacher model and a student model, wherein the model structure and model parameters of the teacher model and the student model are the same as those of the second intermediate model;
[0019] The teacher model and the student model are trained using an average teacher algorithm, and a plurality of target domain data are input into the teacher model to obtain a plurality of potential features and a plurality of first prediction results output by the teacher model;
[0020] According to the plurality of potential features and the plurality of first prediction results, correction processing is performed on the feature extractor parameters and the predictor parameters of the student model;
[0021] According to the corrected feature extractor parameters and predictor parameters, an exponential moving average algorithm is used to correct the feature extractor parameters and predictor parameters of the teacher model, and the target domain multiple data are input into the updated teacher model to adaptively correct the teacher model and the student model iteratively until the iteration is completed, thereby obtaining the iteratively completed teacher model and student model;
[0022] The student model completed by the iteration is used as the target prediction model.
[0023] Optionally, performing correction processing on feature extractor parameters and predictor parameters of the student model according to the plurality of potential features and the plurality of first prediction results comprises:
[0024] The following formula is used to calibrate the feature extractor parameters and predictor parameters of the student model:
[0025]
[0026] in, is the target domain data, is the target domain data set, is the expected distance, The Barlow Twin encourages the cross-correlation matrix between the outputs of the two networks to be as close to the identity matrix as possible. is the contrast loss function, is the mean square error, is the predictor of the teacher model, is the feature extractor of the teacher model, is the predictor of the student model, is the feature extractor of the student model, is a hyperparameter, is the cross-correlation matrix between the underlying representations of the teacher model and the student model, express and The similarity between and are the predicted values of the student model and the teacher model, respectively. is a hyperparameter.
[0027] Optionally, according to the updated feature extractor parameters and predictor parameters of the student model, an exponential moving average algorithm is used to correct the feature extractor parameters and predictor parameters of the teacher model, including:
[0028] The following formula is used to calibrate the feature extractor parameters and predictor parameters of the teacher model:
[0029]
[0030] in, is the smoothing coefficient, are the feature extractor parameters and predictor parameters of the student model, is the parameter of the teacher model at the t-1th iteration .
[0031] In a second aspect, the present application provides a device for constructing an industrial data prediction model, the device comprising:
[0032] An acquisition module is used to acquire a source model of a source domain and multiple data of a target domain;
[0033] A determination module, configured to determine, according to the source model, model parameters corresponding to the source model and an intermediate model, wherein the model structure of the intermediate model is the same as the model structure of the source model;
[0034] A processing module is used to perform adversarial training on the intermediate model according to the multiple data in the target domain and the source model to obtain a second intermediate model, and the adversarial training is used to update the feature extractor of the intermediate model.
[0035] Optionally, the processing module is further used to parse the source model to obtain model parameters corresponding to the source model, wherein the model parameters include: feature extractor parameters and predictor parameters of the source model;
[0036] The processing module is also used to construct a standard model with the same model structure as the source model according to the model structure of the source model, and initialize the standard model according to the feature extractor parameters and the predictor parameters to obtain the intermediate model.
[0037] Optionally, the processing module is further used to input the plurality of target domain data into the source model and the intermediate model respectively, to obtain a first prediction result output by the source model and a second prediction result output by the intermediate model;
[0038] The processing module is further used to update the feature extractor parameters of the intermediate model and the predictor parameters of the source model according to the first prediction result and the second prediction result to obtain an updated intermediate model and source model;
[0039] The processing module is further used to perform update iterative processing on the updated intermediate model and the source model according to the multiple data of the target domain until the iteration of the intermediate model is completed to obtain the second intermediate model.
[0040] Optionally, the processing module is further used to initialize the second intermediate model according to the second intermediate model to obtain a teacher model and a student model, wherein the model structure and model parameters of the teacher model and the student model are the same as those of the second intermediate model;
[0041] The processing module is further used to train the teacher model and the student model using an average teacher algorithm, input the target domain data into the teacher model, and obtain multiple potential features and multiple first prediction results output by the teacher model;
[0042] The processing module is further used to perform correction processing on the feature extractor parameters and the predictor parameters of the student model according to the multiple potential features and the multiple first prediction results;
[0043] The processing module is further used to perform correction processing on the feature extractor parameters and the predictor parameters of the teacher model using an exponential moving average algorithm according to the corrected feature extractor parameters and the predictor parameters, and input the target domain multiple data into the updated teacher model to perform adaptive correction iteration on the teacher model and the student model until the iteration is completed, thereby obtaining the iterated teacher model and the student model;
[0044] The processing module is also used to use the student model completed by the iteration as the target prediction model.
[0045] Optionally, the device for constructing the industrial data prediction model further includes: a calculation module;
[0046] The calculation module is used to correct the feature extractor parameters and predictor parameters of the student model using the following formula:
[0047]
[0048] in, is the target domain data, is the target domain data set, is the expected distance, The Barlow Twin encourages the cross-correlation matrix between the outputs of the two networks to be as close to the identity matrix as possible. is the contrast loss function, is the mean square error, is the predictor of the teacher model, is the feature extractor of the teacher model, is the predictor of the student model, is the feature extractor of the student model, is a hyperparameter, is the cross-correlation matrix between the underlying representations of the teacher model and the student model, express and The similarity between and are the predicted values of the student model and the teacher model, respectively. is a hyperparameter.
[0049] Optionally, the calculation module is further used to correct the feature extractor parameters and predictor parameters of the teacher model using the following formula:
[0050]
[0051] in, is the smoothing coefficient, are the feature extractor parameters and predictor parameters of the student model, is the parameter of the teacher model at the t-1th iteration .
[0052] In a third aspect, the present application provides a device for constructing an industrial data prediction model under privacy protection, including:
[0053] Memory;
[0054] processor;
[0055] Wherein, the memory stores computer-executable instructions;
[0056] The processor executes the computer-executable instructions stored in the memory to implement the method for constructing an industrial data prediction model under privacy protection as described in the above-mentioned first aspect and various possible implementation methods of the first aspect.
[0057] In a fourth aspect, the present application provides a computer storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement a method for constructing an industrial data prediction model under privacy protection as described in the first aspect and various possible implementation methods of the first aspect.
[0058] The present application provides a method, device, and equipment for constructing an industrial data prediction model under privacy protection. Applied to the field of industrial production technology, the method obtains a source model of a source domain and multiple data of a target domain, determines an intermediate model according to the source model, performs adversarial training on the intermediate model according to multiple data of the target domain and the source model, obtains a second intermediate model, initializes the second intermediate model, obtains a teacher model and a student model, uses an average teacher algorithm to train the teacher model and the student model, and corrects the feature extractor parameters and predictor parameters of the student model, inputs multiple data of the target domain into the student model, and performs adaptive correction iteration on the student model, and uses the iterated student model as the target prediction model. This ensures accurate cross-domain migration of industrial models while protecting privacy. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0060] Figure 1 Schematic diagram of the process of constructing an industrial data prediction model under privacy protection provided in this application Figure 1 ;
[0061] Figure 2 Schematic diagram of the process of constructing an industrial data prediction model under privacy protection provided in this application Figure 2 ;
[0062] Figure 3 A schematic diagram of the structure of a device for constructing an industrial data prediction model under privacy protection provided in this application;
[0063] Figure 4 A schematic diagram of the structure of the equipment for building an industrial data prediction model under privacy protection provided in this application.
[0064] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0066] The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present invention and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in sequences other than those illustrated or described herein.
[0067] In the embodiments of the present application, the words "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0068] Industrial production technology is an important part of industrial production, which is centered on industry and applies a variety of advanced engineering technologies and information computer technologies to improve production efficiency, reduce production costs, improve product quality, and achieve sustainable production through automation, informatization, and intelligent means. It runs through the entire industrial production process and plays a vital role in the development, innovation, and transformation of modern industrial production.
[0069] With the continuous development of science and technology, industrial production technology is constantly undergoing revolutionary changes, new technologies are emerging in an endless stream, and industrial production methods are also changing. Today's industrial production technology has achieved the process from instrument control to field control, and then to central control. Enterprises can achieve large-scale production through advanced technologies and tools such as computer networks. Industrial automation technology, cloud manufacturing technology, advanced intelligent manufacturing technology, etc. have gradually become new trends. Industrial production has changed from the traditional human needs-centered to the characteristics of engineering technology and computer technology.
[0070] By analyzing industrial monitoring signals, we can effectively predict abnormalities and failures in the industrial production process, thereby avoiding potential safety hazards and economic losses in the production process. Industrial monitoring signals can obtain real-time data in the production line, including parameters such as temperature, pressure, and power of production equipment. By analyzing these parameters, we can obtain status information of the equipment or production process, such as failure, load, abnormality, etc., and then make predictions and judgments. Through the prediction results, we can control and manage the equipment or production process, thereby preventing potential safety hazards and economic losses.
[0071] However, the existing industrial equipment health prediction methods have the following defects:
[0072] 1) Low security. The collection of high-quality monitoring data is difficult and expensive. The monitoring signals of many devices collected are unlabeled and cannot be directly used for time series prediction modeling. Due to information security and privacy protection considerations, the monitoring signals of many devices cannot be shared. Therefore, it is difficult to achieve knowledge transfer between prediction models of different devices without accessing the monitoring data.
[0073] 2) Poor accuracy. The distribution of data from different data sources varies greatly, and the knowledge transfer model is difficult to effectively represent data that is far away from the source support.
[0074] In response to the above problems, this application proposes a method for constructing an industrial data prediction model under privacy protection. The method obtains the source model of the source domain and multiple data of the target domain, determines the model parameters and the intermediate model corresponding to the source model according to the source model, and performs adversarial training on the intermediate model according to multiple target domain data and the source model to obtain a second intermediate model. The second intermediate model is initialized to obtain a teacher model and a student model, and the teacher model and the student model are trained using the average teacher algorithm, and the feature extractor parameters and predictor parameters of the student model are corrected, and the iterated student model is used as the target prediction model. This ensures accurate cross-domain migration of industrial models while protecting privacy.
[0075] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0076] Figure 1 Schematic diagram of the process of constructing an industrial data prediction model under privacy protection provided in the application embodiment Figure 1 .like Figure 1 As shown, the method for constructing an industrial data prediction model under privacy protection provided by this embodiment includes:
[0077] S101: Obtain a source model of a source domain and multiple data of a target domain.
[0078] Among them, obtaining the source model of the source domain and multiple data of the target domain is intended to clarify the source domain and the target domain, that is, from which domain to which domain the industrial data model will be migrated. During the process, necessary conversion and cleaning of the data will be performed so that it can be better applied in different fields or different systems.
[0079] For example, a possible implementation is given here: Acquiring a source model of a source domain A and multiple target domain B data means that in this industrial data model migration, the model is migrated from the source domain A to the target domain B.
[0080] S102: Determine model parameters corresponding to the source model and the intermediate model according to the source model.
[0081] Among them, the source model is a model trained in the source domain using source domain data. It is a model that can provide good performance. The structure and parameters of the source model are usually strictly optimized and adjusted to provide powerful and stable functions.
[0082] Specifically, the source model contains a feature extractor and a predictor , the feature extractor of the source model and predictor The two modules are connected in series. The feature extractor is designed as a 4-layer convolutional neural network, and the predictor is designed as a 2-layer perceptron. The model parameters of the source model are trained by the mean square error of the predictor:
[0083]
[0084] in, is the number of samples in the source domain, is the predicted value, is the source model feature extractor, is the source model predictor.
[0085] According to the structure and parameters of the source model, the intermediate model is determined, wherein the intermediate model has the same structure and parameters as the source model.
[0086] S103. Perform adversarial training on the intermediate model according to the target domain multiple data and the source model to obtain a second intermediate model.
[0087] Among them, according to multiple target domain data and source models, the intermediate model is adversarially trained to obtain a second intermediate model. The adversarial training is used to update and optimize the feature extractor of the intermediate model so that it can adapt to the distribution of target domain data and generate a second intermediate model with good generalization ability.
[0088] In adversarial training, the feature extractor of the intermediate model is regarded as a generator. The goal is to make the features extracted by the generator as similar as possible to the features of the source domain data in the target domain through the competition mechanism in the training process, solve the difference problem between data in different domains, and thus improve the generalization ability of the model.
[0089] This embodiment proposes a method for constructing an industrial data prediction model under privacy protection, which obtains a source model of a source domain and multiple data of a target domain, determines the model parameters and the intermediate model corresponding to the source model according to the source model, and performs adversarial training on the intermediate model according to multiple target domain data and the source model to obtain a second intermediate model, thereby realizing knowledge transfer between prediction models of different devices without accessing source domain monitoring data. This improves the security of cross-domain migration of industrial data models.
[0090] Figure 2 This is a schematic diagram of the process of constructing an industrial data prediction model under privacy protection provided by an embodiment of the present application. Figure 2 This embodiment is in Figure 1 Based on the embodiment, the method for constructing the industrial data prediction model under privacy protection is described in detail. Figure 2As shown, the method for constructing an industrial data prediction model under privacy protection provided by this embodiment includes:
[0091] S201: Obtain a source model of a source domain and multiple data of a target domain.
[0092] Step S201 is the same as step S101 and will not be described again here.
[0093] S202: Analyze the source model to obtain model parameters corresponding to the source model, where the model parameters include: feature extractor parameters and predictor parameters of the source model.
[0094] The parsing of the source model refers to extracting model parameters corresponding to the source model by analyzing the structure and weight of the source model, and these model parameters include feature extractor parameters and predictor parameters of the source model.
[0095] Feature extractor parameters refer to the weights and biases of the source model used to extract features from input data. They determine how the input data is represented in the model after feature extraction. Feature extractors are usually network structures composed of multiple layers, such as the convolutional layers and pooling layers of convolutional neural networks. The weights and biases of these layers constitute feature extractor parameters, and the values of these parameters can be extracted by parsing the source model.
[0096] Predictor parameters refer to the weights and biases of the part of the source model used to make the final prediction, which generate the output of the model based on the features extracted by the feature extractor. The predictor is usually a structure composed of fully connected layers or output layers, in which each neuron corresponds to a class or a specific prediction result output by the model. By parsing the source model, the weights and biases of each neuron in the predictor can be extracted.
[0097] S203. According to the model structure of the source model, a standard model having the same model structure as the source model is constructed, and according to the feature extractor parameters and the predictor parameters, the standard model is initialized to obtain an intermediate model.
[0098] Among them, according to the model structure of the source model, building a standard model with the same model structure as the source model means establishing a model with the same structure as the source model, including the same network layer and connection method, so that the two models have the same feature extraction and prediction capabilities when processing input data.
[0099] After building the standard model, we can use the feature extractor parameters and predictor parameters of the source model to initialize the standard model to obtain the intermediate model. The purpose is to set the weights and bias parameters in the standard model to the parameter values obtained by pre-training the source model, so that the intermediate model can have the same feature extraction and prediction capabilities as the source model.
[0100] S204: Input multiple target domain data into the source model and the intermediate model respectively to obtain a first prediction result output by the source model and a second prediction result output by the intermediate model.
[0101] Among them, inputting multiple target domain data into the source model and the intermediate model respectively can be used to compare the prediction ability and adaptability of the source model and the intermediate model for new data in the target domain. By obtaining the first prediction result output by the source model and the second prediction result output by the intermediate model, the initialization effect of the intermediate model and its performance on the new target domain can be evaluated.
[0102] For example, a possible implementation is given here. In order to enable the intermediate model to focus on samples in the target domain that are significantly different from the source domain data distribution, thereby improving the generalization of the intermediate model. Through adversarial training of the intermediate model, that is, through the adversarial relationship between the source model and the intermediate model, the generalization ability of the intermediate model to the target domain is enhanced.
[0103] Specifically, the predictor of the intermediate model is fixed, and the feature extractor of the source model that is trained against the intermediate model is fixed, and its parameters are fixed not to be updated. Multiple data of the target domain are input into the source model and the intermediate model respectively, and the first prediction result output by the source model is obtained by the following formula: , And the second prediction result output by the intermediate model , :
[0104] Extracting latent features:
[0105] Get the prediction results:
[0106] in, is the target domain data set, is the feature extractor of the source model, is the feature extractor of the intermediate model, is the predictor of the source model, is the predictor of the intermediate model.
[0107] S205 . Update the feature extractor parameters of the intermediate model and the predictor parameters of the source model according to the first prediction result and the second prediction result to obtain an updated intermediate model and source model.
[0108] Among them, according to the first prediction result and the second prediction result, the feature extractor parameters of the intermediate model and the predictor parameters of the source model are updated so that the intermediate model can better adapt to the specific target domain data, thereby improving the performance in the new domain. Such an update process can help the model better predict and classify new data, thereby improving the generalization ability of the overall model.
[0109] Specifically, after the target domain data is input into the source model and the intermediate model, the calibration fine-tuning loss of the source model and the intermediate model can be defined as follows:
[0110]
[0111] in, is the target domain data set, is the predictor of the source model, is the predictor of the intermediate model, is the feature extractor of the source model, It is the feature extractor of the intermediate model.
[0112] In this phase, the predictor of the source model is trained to maximize the diversity of the predictions of the two models. It aims to effectively detect samples that are far from the source domain support. At the same time, the feature extractor of the intermediate model is updated to minimize this prediction difference, thereby achieving the purpose of aligning the target and source domain features. The loss of this adversarial training can be expressed as:
[0113]
[0114] in, is the target domain data set, is the feature extractor of the intermediate model, is the predictor of the source model.
[0115] S206 , performing update iterative processing on the updated intermediate model and the source model according to the multiple data of the target domain until the iteration of the intermediate model is completed to obtain a second intermediate model.
[0116] The updated intermediate model and the source model are updated and iterated according to the target domain data until the intermediate model iteration is completed to obtain a second intermediate model. The iteration completion condition may be, for example, reaching a set number of iterations or the effect of the current iteration is not as good as the previous iteration.
[0117] pass renew and .
[0118] S207. Initialize the second intermediate model according to the second intermediate model to obtain a teacher model and a student model.
[0119] Among them, the second intermediate model is initialized according to the second intermediate model. Through the initialization process, the teacher model and the student model have the same model structure and model parameters, and the model structure and model parameters of the teacher model and the student model are the same as the second intermediate model.
[0120] Specifically, initialize the feature extractor of the teacher model and the feature extractor of the student model The parameters of the second intermediate model are ; Initialize the predictor of the teacher model and the predictor of the student model is the parameter of the second intermediate model .
[0121] S208. Use the average teacher algorithm to train the teacher model and the student model, input multiple target domain data into the teacher model, and obtain multiple potential features and multiple first prediction results output by the teacher model.
[0122] Among them, the average teacher algorithm is used to train the teacher model and the student model, and multiple target domain data are input into the teacher model, so that multiple potential features and multiple first prediction results output by the teacher model can be obtained. These features and prediction results help us to further analyze and understand the target domain data and provide important information for subsequent prediction tasks.
[0123] Specifically, multiple first potential features and multiple first prediction results output by the teacher model are obtained by the following formula:
[0124] Extracting latent features:
[0125] Get the prediction results:
[0126] in, For the student model domain data, For the teacher model domain data, is the student model feature extractor, is the teacher model feature extractor, is the student model predictor, is the teacher model predictor.
[0127] S209: According to the multiple potential features and the multiple first prediction results, calibrate the feature extractor parameters and the predictor parameters of the student model.
[0128] Among them, according to multiple potential features and multiple first prediction results, the feature extractor parameters and predictor parameters of the student model are corrected, which can help the student model better learn valuable information from the knowledge of the teacher model and improve its performance.
[0129] Specifically, the following formula is used to correct the feature extractor parameters and predictor parameters of the student model:
[0130]
[0131] in, is the target domain data, is the target domain data set, is the expected distance, The Barlow Twin encourages the cross-correlation matrix between the outputs of the two networks to be as close to the identity matrix as possible. is the contrast loss function, is the mean square error, is the predictor of the teacher model, is the feature extractor of the teacher model, is the predictor of the student model, is the feature extractor of the student model, is a hyperparameter, is the cross-correlation matrix between the underlying representations of the teacher model and the student model, express and The similarity between and are the predicted values of the student model and the teacher model, respectively. is a hyperparameter.
[0132] The consistency cost is defined as the student model (parameters and Enhancement ) is compared with the prediction of the teacher model (parameters and Enhancement The expected distance between the predictions of ,The teacher model is used to generate pseudo labels for self-supervised training, and the parameters of the student model are updated by the pseudo labels provided by the teacher model.
[0133] An alternative approach to adapting the feature space via contrastive learning. As mentioned before, the input data of the teacher and student models are augmented differently and can be viewed as different views of the target data. A self-supervised consistency loss is applied to both the student and teacher feature extractors. and We extract latent features from the teacher and student models The feature-level consistency between different views is exploited to achieve feature space adaptation, which is achieved through Barlow Twin, encouraging the cross-correlation matrix between the outputs of the two networks to be as close to the identity matrix as possible:
[0134]
[0135] in, is a hyperparameter, is the cross-correlation matrix between the underlying representations of the two networks:
[0136]
[0137] in, represents batch samples, , Represents the feature dimension of the latent representation. Latent features It is standardized here. The first term encourages the model to associate the same features across different samples, while the second term de-correlates different features.
[0138] In addition, a contrastive loss is defined, which expects data with similar pseudo labels (predictions of the teacher model) to have similar underlying representations (outputs of the student feature extractor). Cosine similarity is used to evaluate the similarity between latent representations, i.e. , express and The similarity between . The contrast loss is defined as:
[0139]
[0140] in, is a hyperparameter to prevent the numerator from being equal to 0, and are the predicted values of the student model and the teacher model, respectively.
[0141] S210. According to the corrected feature extractor parameters and predictor parameters, the exponential moving average algorithm is used to correct the feature extractor parameters and predictor parameters of the teacher model, and multiple target domain data are input into the updated teacher model to adaptively correct the teacher model and the student model iteratively until the iteration is completed, thereby obtaining the iterated teacher model and the student model.
[0142] The teacher model is used to generate pseudo labels for self-supervised training, and the parameters of the student model are updated by the pseudo labels provided by the teacher model. According to the corrected feature extractor parameters and predictor parameters, the feature extractor parameters and predictor parameters of the teacher model are corrected using the following formula:
[0143]
[0144] in, is the smoothing coefficient, are the feature extractor parameters and predictor parameters of the student model, is the parameter of the teacher model at the t-1th iteration .
[0145] The feature extractor parameters and predictor parameters of the teacher model are corrected by using an exponential moving average algorithm to adaptively correct the teacher model and the student model until the iteration is completed, thereby obtaining the iteratively completed teacher model and student model.
[0146] Optionally, in domain adaptation and continuous learning modeling, it is often necessary to regularize the loss function to prevent the parameters of the newly trained prediction model from deviating from the source domain model. Therefore, in our framework, regularization is applied at each stage of training to ensure that the model maintains its predictive power for downstream tasks. Since different parameters in a neural network have different importance for the prediction task, we want to selectively slow down the learning of weights that are important for downstream tasks to remember. Therefore, we use elastic weight consolidation (EWC) as a regularization term.
[0147] The second-order derivative is used to evaluate the importance of the parameters. Therefore, the regularized loss function of EWC can be expressed as:
[0148]
[0149] in, is the number of samples processed in batches, and are the current model parameters and the source model parameters respectively, is the Fisher diagnostic matrix, The importance of source domain knowledge is set.
[0150] The invention is divided into two stages: generalization of the intermediate model based on adversarial training, and unsupervised alignment of the target model based on mean-teacher, and the regularization function participates in the update of the neural network parameters in both stages.
[0151] S211. Use the student model after iteration as the target prediction model.
[0152] This embodiment proposes a method for constructing an industrial data prediction model under privacy protection, which method obtains a source model of a source domain and multiple data of a target domain; parses the source model to obtain model parameters corresponding to the source model and constructs an intermediate model according to the model parameters and structure of the source model; inputs multiple data of the target domain into the source model and the intermediate model respectively to obtain a first prediction result output by the source model and a second prediction result of the intermediate model, and updates the feature extractor parameters of the intermediate model and the predictor parameters of the source model; updates and iterates the updated intermediate model and the source model according to the multiple data of the target domain to obtain a second intermediate model; and uses the average teacher algorithm to initialize the second intermediate model to obtain Teacher model and student model; input multiple target domain data into the teacher model, obtain multiple potential features and multiple first prediction results output by the teacher model, and correct the feature extractor parameters and predictor parameters of the student model; according to the corrected feature extractor parameters and predictor parameters, use the exponential moving average algorithm to correct the feature extractor parameters and predictor parameters of the teacher model; input multiple target domain data into the updated teacher model to adaptively correct the teacher model and the student model until the iteration is completed, and obtain the iterated teacher model and student model; use the iterated student model as the target prediction model to improve the accuracy of cross-domain migration of industrial data models.
[0153] Figure 3 A schematic diagram of the structure of a device for constructing an industrial data prediction model under privacy protection provided in this application, such as Figure 3 As shown, the device 300 for constructing an industrial data prediction model provided in this embodiment includes:
[0154] An acquisition module 301 is used to acquire a source model of a source domain and multiple data of a target domain;
[0155] A determination module 302 is used to determine, according to the source model, model parameters corresponding to the source model and an intermediate model, wherein the model structure of the intermediate model is the same as the model structure of the source model;
[0156] The processing module 303 is used to perform adversarial training on the intermediate model according to the multiple data of the target domain and the source model to obtain a second intermediate model, and the adversarial training is used to update the feature extractor of the intermediate model.
[0157] Optionally, the processing module 303 is further used to parse the source model to obtain model parameters corresponding to the source model, wherein the model parameters include: feature extractor parameters and predictor parameters of the source model;
[0158] The processing module 303 is also used to construct a standard model with the same model structure as the source model according to the model structure of the source model, and initialize the standard model according to the feature extractor parameters and the predictor parameters to obtain the intermediate model.
[0159] Optionally, the processing module 303 is further used to input the plurality of target domain data into the source model and the intermediate model respectively, to obtain a first prediction result output by the source model and a second prediction result output by the intermediate model;
[0160] The processing module 303 is further used to update the feature extractor parameters of the intermediate model and the predictor parameters of the source model according to the first prediction result and the second prediction result to obtain an updated intermediate model and source model;
[0161] The processing module 303 is further used to perform update iterative processing on the updated intermediate model and the source model according to the multiple data of the target domain until the iteration of the intermediate model is completed to obtain the second intermediate model.
[0162] Optionally, the processing module 303 is further used to initialize the second intermediate model according to the second intermediate model to obtain a teacher model and a student model, wherein the model structure and model parameters of the teacher model and the student model are the same as those of the second intermediate model;
[0163] The processing module 303 is further used to train the teacher model and the student model using an average teacher algorithm, input the target domain data into the teacher model, and obtain multiple potential features and multiple first prediction results output by the teacher model;
[0164] The processing module 303 is further used to perform correction processing on the feature extractor parameters and the predictor parameters of the student model according to the multiple potential features and the multiple first prediction results;
[0165] The processing module 303 is further used to calibrate the feature extractor parameters and the predictor parameters of the teacher model using an exponential moving average algorithm according to the calibrated feature extractor parameters and the predictor parameters, and input the target domain multiple data into the updated teacher model to perform adaptive calibration iteration on the teacher model and the student model until the iteration is completed, thereby obtaining the iterated teacher model and the student model;
[0166] The processing module 303 is further configured to use the student model completed by the iteration as a target prediction model.
[0167] Optionally, the device for constructing the industrial data prediction model further includes: a calculation module 304;
[0168] The calculation module 304 is used to calibrate the feature extractor parameters and predictor parameters of the student model using the following formula:
[0169]
[0170] in, is the target domain data, is the target domain data set, is the expected distance, The Barlow Twin encourages the cross-correlation matrix between the outputs of the two networks to be as close to the identity matrix as possible. is the contrast loss function, is the mean square error, is the predictor of the teacher model, is the feature extractor of the teacher model, is the predictor of the student model, is the feature extractor of the student model, is a hyperparameter, is the cross-correlation matrix between the underlying representations of the teacher model and the student model, express and The similarity between and are the predicted values of the student model and the teacher model, respectively. is a hyperparameter.
[0171] Optionally, the calculation module 304 is further used to calibrate the feature extractor parameters and predictor parameters of the teacher model using the following formula:
[0172]
[0173] in, is the smoothing coefficient, are the feature extractor parameters and predictor parameters of the student model, is the parameter of the teacher model at the t-1th iteration .
[0174] Figure 4 This is a schematic diagram of the structure of the equipment for building the industrial data prediction model under privacy protection provided by this application. Figure 4 As shown, the present application provides a device for constructing an industrial data prediction model under privacy protection. The device 400 for constructing an industrial data prediction model under privacy protection includes: a receiver 401, a transmitter 402, a processor 403 and a memory 404.
[0175] Receiver 401, used for receiving instructions and data;
[0176] A transmitter 402, used for sending instructions and data;
[0177] Memory 404, for storing computer-executable instructions;
[0178] The processor 403 is used to execute the computer-executable instructions stored in the memory 404 to implement the various steps performed by the method for constructing an industrial data prediction model under privacy protection in the above embodiment. For details, please refer to the relevant description in the embodiment of the method for constructing an industrial data prediction model under privacy protection.
[0179] Optionally, the memory 404 may be independent or integrated with the processor 403 .
[0180] When the memory 404 is independently provided, the electronic device further includes a bus for connecting the memory 404 and the processor 403 .
[0181] The present application also provides a computer storage medium, which stores computer execution instructions. When a processor executes the computer execution instructions, a method for constructing an industrial data prediction model under privacy protection as executed by the construction device of the industrial data prediction model under privacy protection as mentioned above is implemented.
[0182] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transient medium). As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0183] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0184] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A method for constructing an industrial data prediction model under privacy protection, characterized in that: Applied to constructing equipment, the industrial data is an industrial monitoring signal, the industrial monitoring signal includes the temperature, pressure, and power of the production equipment, the industrial data prediction model predicts abnormalities and failures in the industrial production process, and the method includes: A source model of a source domain and multiple data of a target domain are obtained; the source model comprises a feature extractor and a predictor, the feature extractor and the predictor are connected in series, the feature extractor is a 4-layer convolutional neural network, and the predictor is a 2-layer perceptron; Determine, according to the source model, model parameters corresponding to the source model and an intermediate model, wherein the model structure of the intermediate model is the same as the model structure of the source model; According to the target domain data and the source model, the intermediate model is subjected to adversarial training to obtain a second intermediate model, wherein the adversarial training is used to update a feature extractor of the intermediate model; Initializing the second intermediate model according to the second intermediate model to obtain a teacher model and a student model, wherein the model structure and model parameters of the teacher model and the student model are the same as those of the second intermediate model; The teacher model and the student model are trained using an average teacher algorithm, and a plurality of target domain data are input into the teacher model to obtain a plurality of potential features and a plurality of first prediction results output by the teacher model; According to the plurality of potential features and the plurality of first prediction results, correction processing is performed on the feature extractor parameters and the predictor parameters of the student model; According to the corrected feature extractor parameters and predictor parameters, an exponential moving average algorithm is used to correct the feature extractor parameters and predictor parameters of the teacher model, and the target domain multiple data are input into the updated teacher model to adaptively correct the teacher model and the student model iteratively until the iteration is completed, thereby obtaining the iteratively completed teacher model and student model; The student model completed by the iteration is used as the target prediction model, wherein the target prediction model migration is migration from the source domain to the target domain, thereby realizing knowledge migration between prediction models of different devices without accessing source domain monitoring data.
2. The method according to claim 1, characterized in that The step of determining the model parameters and the intermediate model corresponding to the source model according to the source model includes: Parsing the source model to obtain model parameters corresponding to the source model, the model parameters including: feature extractor parameters and predictor parameters of the source model; According to the model structure of the source model, a standard model having the same model structure as the source model is constructed, and according to the feature extractor parameters and the predictor parameters, the standard model is initialized to obtain the intermediate model.
3. The method according to claim 2, characterized in that The method of performing adversarial training on the intermediate model according to the target domain data and the source model to obtain a second intermediate model includes: Inputting the target domain data into the source model and the intermediate model respectively, obtaining a first prediction result output by the source model and a second prediction result output by the intermediate model; According to the first prediction result and the second prediction result, updating the feature extractor parameters of the intermediate model and the predictor parameters of the source model to obtain an updated intermediate model and a source model; The updated intermediate model and the source model are updated and iterated according to the target domain data until the iteration of the intermediate model is completed to obtain the second intermediate model.
4. The method according to claim 1, characterized in that: The correcting the feature extractor parameters and the predictor parameters of the student model according to the plurality of potential features and the plurality of first prediction results comprises: correcting the feature extractor parameters and the predictor parameters of the student model using the following formula: in, is the target domain data, is the target domain data set, is the expected distance, The Barlow Twin encourages the cross-correlation matrix between the outputs of the two networks to be as close to the identity matrix as possible. is the contrast loss function, is the mean square error, is the predictor of the teacher model, is the feature extractor of the teacher model, is the predictor of the student model, is the feature extractor of the student model, is a hyperparameter, is the cross-correlation matrix between the underlying representations of the teacher model and the student model, express and The similarity between and are the predicted values of the student model and the teacher model respectively, is a hyperparameter.
5. The method according to claim 4, characterized in that The method of correcting the feature extractor parameters and the predictor parameters of the teacher model using an exponential moving average algorithm according to the updated feature extractor parameters and the predictor parameters of the student model comprises: The following formula is used to calibrate the feature extractor parameters and predictor parameters of the teacher model: in, is the smoothing coefficient, are the feature extractor parameters and predictor parameters of the student model, is the parameter of the teacher model at the t-1th iteration .
6. A device for constructing an industrial data prediction model under privacy protection, characterized in that: Applied to constructing equipment, the industrial data is an industrial monitoring signal, the industrial monitoring signal includes the temperature, pressure, and power of the production equipment, the industrial data prediction model predicts the abnormality and failure of the industrial production process, and the device includes: An acquisition module, used for acquiring a source model of a source domain and multiple data of a target domain; the source model comprises a feature extractor and a predictor, the feature extractor and the predictor are connected in series, the feature extractor is a 4-layer convolutional neural network, and the predictor is a 2-layer perceptron; A determination module, configured to determine, according to the source model, model parameters corresponding to the source model and an intermediate model, wherein the model structure of the intermediate model is the same as the model structure of the source model; A processing module, configured to perform adversarial training on the intermediate model according to the target domain data and the source model to obtain a second intermediate model, wherein the adversarial training is used to update a feature extractor of the intermediate model; The device also includes: The processing module is further used to initialize the second intermediate model according to the second intermediate model to obtain a teacher model and a student model, wherein the model structure and model parameters of the teacher model and the student model are the same as those of the second intermediate model; The processing module is further used to train the teacher model and the student model using an average teacher algorithm, input the target domain data into the teacher model, and obtain multiple potential features and multiple first prediction results output by the teacher model; The processing module is further used to perform correction processing on the feature extractor parameters and the predictor parameters of the student model according to the multiple potential features and the multiple first prediction results; The processing module is further used to perform correction processing on the feature extractor parameters and the predictor parameters of the teacher model using an exponential moving average algorithm according to the corrected feature extractor parameters and the predictor parameters, and input the target domain multiple data into the updated teacher model to perform adaptive correction iteration on the teacher model and the student model until the iteration is completed, thereby obtaining the iterated teacher model and the student model; The processing module is also used to use the iterated student model as the target prediction model; wherein the target prediction model migration is from the source domain to the target domain, thereby realizing knowledge migration between prediction models of different devices without accessing source domain monitoring data.
7. A device for constructing an industrial data prediction model under privacy protection, characterized in that: include: Memory; processor; Wherein, the memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method for constructing an industrial data prediction model under privacy protection as described in any one of claims 1-5.
8. A computer storage medium, characterized in that: The computer storage medium stores computer execution instructions, which, when executed by a processor, are used to implement a method for constructing an industrial data prediction model under privacy protection as described in any one of claims 1 to 5.