A Model Optimization and Update Method, Device, and Medium for Generative Artificial Intelligence

By cleaning and standardizing engineering data, combined with the generative artificial intelligence dynamic fine-tuning model, the problems of difficulty in model update and high repetitive data occupancy in the existing technology are solved, and the real-time adaptability and efficient training effect of the model are achieved.

CN119961682BActive Publication Date: 2025-06-24SHENZHEN EXTREME VISION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510436389.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-06-24
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

In the prior art, models of predictive maintenance, quality control and supply chain optimization are difficult to update according to real-time changes in production environment, equipment status or market demand after deployment, resulting in excessive space occupancy of duplicate data during model training.

Method used

By obtaining the engineering data to be trained, the initial sample data is extracted, the data is cleaned and standardized, the model parameters are determined, and the model is trained using batch training methods until all sample data is completed training, and the basic model is output. Combining the basic model with generative artificial intelligence, fine-tuning the basic model through dynamic monitoring and analysis of samples, and updating the model when the convergence state is optimized.

Benefits of technology

The space usage of repeated training data of behavioral analysis models with different behaviors to be analyzed during training is reduced, the real-time adaptability and accuracy of the model is improved, the difficulty of model maintenance is reduced, and the analysis efficiency and accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961682B_ABST
    Figure CN119961682B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device and medium for optimizing and updating a model of generative artificial intelligence, which is used to reduce the space occupancy of repeated training data when training behavior analysis models for different behaviors to be analyzed. The method of the present application includes: obtaining engineering data to be trained; extracting initial sample data from the engineering data according to the behavior to be analyzed; performing data cleaning on the initial sample data set to obtain a target sample data set; determining model parameters according to the behavior to be analyzed, and obtaining target model training data from the target sample data set; using a batch training method to train the model corresponding to the behavior to be analyzed through the target model training data, and outputting a basic model; combining the basic model with generative artificial intelligence to obtain a target artificial intelligence; monitoring the dynamics of the analysis samples through the target artificial intelligence, and fine-tuning the basic model; monitoring the convergence state after the basic model is fine-tuned; deploying the updated model as the basic model of the target artificial intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular, to a method, device, and medium for optimizing and updating a model of generative artificial intelligence. Background Art

[0002] Currently, the manufacturing industry generally uses static models trained offline based on historical data in the fields of predictive maintenance, quality control, and supply chain optimization. These models obtain a relatively high prediction accuracy through batch training in the initial stage. However, once deployed, it is difficult to update their parameters according to the real-time changes in the production environment, equipment status, or market demand.

[0003] In the prior art, predictive maintenance, quality control, and supply chain optimization respectively correspond to three different prediction models. The training data sources of these models are the same engineering data. When actually training the models, there is a certain intersection in the training data required by these models, resulting in duplicate data being obtained when acquiring the training data of different models through engineering data, and the duplicate data between different model training samples over-occupies the model training equipment. Summary of the Invention

[0004] To solve the above technical problems, this application provides a method, device, and medium for optimizing and updating a model of generative artificial intelligence, which is used to reduce the space occupation of duplicate training data when training behavior analysis models for different behaviors to be analyzed.

[0005] The technical solutions provided in this application are described below:

[0006] The first aspect of this application provides a method for optimizing and updating a model of generative artificial intelligence, including:

[0007] Obtain engineering data to be trained;

[0008] Extract initial sample data from the engineering data according to the behavior to be analyzed, where the behavior to be analyzed is predictive maintenance, quality control, or supply chain optimization;

[0009] Perform data cleaning on the initial sample data set to obtain a target sample data set;

[0010] Determine model parameters according to the behavior to be analyzed, and obtain target model training data from the target sample data set according to the model parameters;

[0011] Use the batch training method to train the model corresponding to the behavior to be analyzed through the target model training data until all the sample data in the target sample data set are completed in training, and output a basic model;

[0012] Combine the basic model with generative artificial intelligence to obtain a target artificial intelligence;

[0013] Monitor the dynamics of the sample through the target artificial intelligence, and fine-tune the basic model;

[0014] Monitor the convergence state after the fine-tuning of the basic model;

[0015] When the convergence state after the fine-tuning of the basic model is better than the current convergence state, deploy the updated model as the basic model of the target artificial intelligence.

[0016] Optionally, the data cleaning of the initial sample data set to obtain the target sample data set includes:

[0017] Detect and clean the missing values of the initial sample data according to preset rules;

[0018] Detect the outliers of the initial sample data, and remove or replace the outliers according to the requirements of the behavior to be analyzed;

[0019] Establish an association key between the production line ID and the time stamp, and delete duplicate records with the association key as the index to complete the data cleaning;

[0020] Perform data standardization on the initial sample data after completing the data cleaning to obtain the target sample data.

[0021] Optionally, the performing data standardization on the initial sample data after completing the data cleaning to obtain the target sample data includes:

[0022] Calculate the mean and standard deviation of each data feature in the initial sample data through the Z-score standardization method;

[0023] Standardize the data features corresponding to the mean and the standard deviation, and the standardization process is calculated through the following formula:

[0024] ;

[0025] where x is the current data point of the data feature, μ is the mean of the data feature corresponding to this data point, and σ is the standard deviation of the data feature corresponding to this data point.

[0026] Obtain all the data with a mean of 0 and a standard deviation of 1 to get the target sample data.

[0027] Optionally, the determining the model parameters according to the behavior to be analyzed and obtaining the target model training data from the target sample data set includes:

[0028] Determine the model architecture through the behavior to be analyzed, and determine the model parameters according to the model architecture;

[0029] Split the target sample data set to obtain target model training data, where the target sample data set is split into target model training data, target model validation data, and target model test data.

[0030] Optionally, using the batch training method to train the model corresponding to the behavior to be analyzed through the target model training data until all sample data in the target sample data set are completed in training, and output a basic model, including:

[0031] Split the target model training data into several small batches according to a preset split density;

[0032] Input the small batches of data into the model corresponding to the behavior to be analyzed one by one, so that the model corresponding to the behavior to be analyzed iterates, and the model corresponding to the behavior to be analyzed is trained multiple times to obtain multiple training results;

[0033] For each obtained training result, evaluate the model performance of the training result through the target model validation data;

[0034] Stop training when the model performance reaches the preset performance index and meets the early stopping mechanism, and output the basic model.

[0035] Optionally, combining the basic model with a generative artificial intelligence to obtain a target artificial intelligence, including:

[0036] Link the input end of the basic model with the input end of the initial generative artificial intelligence;

[0037] Generate structured prompt words according to the output format of the basic model, and perform boundary restriction on the feedback of the initial generative artificial intelligence according to the structured prompt words to obtain the target artificial intelligence.

[0038] Optionally, monitoring the dynamics of the analysis sample through the target artificial intelligence and fine-tuning the basic model, including:

[0039] Deploy the target artificial intelligence so that the target artificial intelligence dynamically monitors the database of the data provider of the engineering data to be trained;

[0040] Set the trigger condition for the target artificial intelligence to fine-tune the basic model;

[0041] When the trigger condition is met, control the target artificial intelligence to fine-tune the basic model according to the new sample data corresponding to the dynamics of the analysis sample.

[0042] The second aspect of this application provides a model optimization and update device for a generative artificial intelligence, and the device includes:

[0043] A first acquisition unit for acquiring engineering data to be trained;

[0044] An extraction unit for extracting initial sample data from the engineering data according to the behavior to be analyzed, where the behavior to be analyzed is predictive maintenance, quality control, or supply chain optimization;

[0045] A data cleaning unit for cleaning the initial sample data set to obtain a target sample data set;

[0046] A second acquisition unit for determining model parameters according to the behavior to be analyzed, and acquiring target model training data from the target sample data set according to the model parameters;

[0047] A model training unit for training the model corresponding to the behavior to be analyzed through the target model training data using a batch training method until all sample data in the target sample data set is completed, and outputting a basic model;

[0048] A combination unit for combining the basic model with generative artificial intelligence to obtain a target artificial intelligence;

[0049] A first monitoring unit for monitoring the dynamics of the analysis samples through the target artificial intelligence and fine-tuning the basic model;

[0050] A second monitoring unit for monitoring the convergence state after the basic model is fine-tuned;

[0051] A deployment unit for deploying the updated model as the basic model of the target artificial intelligence when the convergence state after the basic model is fine-tuned is better than the current convergence state.

[0052] Optionally, the data cleaning unit is specifically used for:

[0053] Detecting and cleaning the missing values of the initial sample data according to preset rules;

[0054] Detecting the outliers of the initial sample data and removing or replacing the outliers according to the requirements of the behavior to be analyzed;

[0055] Establishing an association key between the production line ID and the timestamp, and deleting duplicate records with the association key as the index to complete data cleaning;

[0056] Performing data standardization on the initial sample data after data cleaning is completed to obtain target sample data.

[0057] Optionally, the data cleaning unit is specifically further used for:

[0058] Calculate the mean and standard deviation of each data feature in the initial sample data through the Z-score normalization method;

[0059] Normalize the data features corresponding to the mean and the standard deviation. The normalization process is calculated through the following formula:

[0060] ;

[0061] where x is the current data point of the data feature, μ is the mean of the data feature corresponding to this data point, and σ is the standard deviation of the data feature corresponding to this data point.

[0062] Obtain all the data with a mean of 0 and a standard deviation of 1 to get the target sample data.

[0063] Optionally, the second acquisition unit is specifically used for:

[0064] Determine the model architecture through the behavior to be analyzed, and determine the model parameters according to the model architecture;

[0065] Split the target sample data set to obtain target model training data. The target sample data set is split into target model training data, target model validation data, and target model test data.

[0066] Optionally, the model training unit is specifically used for:

[0067] Split the target model training data into several small batches according to the preset split density;

[0068] Input the small batches of data into the model corresponding to the behavior to be analyzed one by one, so that the model corresponding to the behavior to be analyzed iterates, and perform multiple trainings on the model corresponding to the behavior to be analyzed to obtain multiple training results;

[0069] For each obtained training result, evaluate the model performance of the training result through the target model validation data;

[0070] Stop training when the model performance reaches the preset performance index and meets the early stopping mechanism, and output the basic model.

[0071] Optionally, the combining unit is specifically used for:

[0072] Link the input end of the basic model with the input end of the initial generative artificial intelligence;

[0073] Generate a structured prompt word according to the output format of the basic model, and perform boundary restriction on the feedback of the initial generative artificial intelligence according to the structured prompt word to obtain the target artificial intelligence.

[0074] Optionally, the first monitoring unit is specifically configured to:

[0075] Deploy the target artificial intelligence to dynamically monitor the database of the data provider of the engineering data to be trained;

[0076] Set the trigger condition for the target artificial intelligence to fine-tune the basic model;

[0077] When the trigger condition is met, control the target artificial intelligence to fine-tune the basic model according to the new sample data corresponding to the dynamic analysis sample.

[0078] The third aspect of the present application provides a model optimization and update device for a generative artificial intelligence, and the device includes:

[0079] A processor, a memory, an input / output unit, and a bus;

[0080] The processor is connected to the memory, the input / output unit, and the bus;

[0081] The memory stores a program, and the processor calls the program to execute the method of the first aspect and any optional method in the first aspect.

[0082] The fourth aspect of the present application provides a computer-readable storage medium, and a program is stored on the computer-readable storage medium, and when the program is executed on a computer, it executes the method of the first aspect and any optional method in the first aspect.

[0083] It can be seen from the above technical solutions that the present application has the following advantages:

[0084] After the present application obtains the sample data required for the corresponding behavior analysis from the historical analysis samples through the behavior to be analyzed, it performs data cleaning on the sample data to obtain higher-quality training data, and deletes duplicate data by comparison during data cleaning to reduce the occupancy of the training data. The behavior to be analyzed includes predictive maintenance, quality control, or supply chain optimization. Then, the learning model is determined according to the behavior to be analyzed, so that the corresponding behavior to be analyzed can obtain the corresponding sample data to train the basic model for analyzing the corresponding behavior to be analyzed. Different model frameworks are determined according to different behaviors to be analyzed, and the models are combined through multimodal logic to obtain the basic model for input into the artificial intelligence. Subsequently, the output end of the basic model is linked to the interactive artificial intelligence through an interface, so that while the artificial intelligence generates real-time feedback, it actively iterates the basic model. So as to achieve the integration of all models and model training data, reduce the occupancy of model training data, quickly output results through the artificial intelligence, and iterate the basic model according to the input data and database updated data, reduce the usage difficulty of the model, and improve the analysis efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] In order to more clearly illustrate the technical solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0086] Figure 1 FIG. is a schematic flowchart of an embodiment of the method for optimizing and updating the model of the generative artificial intelligence in the present application;

[0087] Figure 2a FIG. is a schematic flowchart of an embodiment of the first stage of the method for optimizing and updating the model of the generative artificial intelligence in the present application;

[0088] Figure 2b FIG. is a schematic flowchart of an embodiment of the second stage of the method for optimizing and updating the model of the generative artificial intelligence in the present application;

[0089] Figure 3 FIG. is a schematic structural diagram of an embodiment of the device for optimizing and updating the model of the generative artificial intelligence in the present application;

[0090] Figure 4 FIG. is a schematic structural diagram of another embodiment of the device for optimizing and updating the model of the generative artificial intelligence in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0091] It should be noted that the method for optimizing and updating the model of the generative artificial intelligence provided in the present application can be applied to a terminal, a system, or a server. For example, the terminal can be a smart phone, a computer, a tablet computer, a smart TV, a smart watch, a portable computer terminal, or a fixed terminal such as a desktop computer. For the convenience of description, the terminal is taken as the execution subject in the present application for illustration.

[0092] The following will clearly and completely describe the technical solutions in the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0093] Please refer to Figure 1 , the present application first provides an embodiment of the method for optimizing and updating the model of the generative artificial intelligence, and this embodiment includes:

[0094] S101. Obtain the engineering data to be trained;

[0095] Engineering data is obtained through channels such as the factory manufacturing execution system (MES), the data acquisition and monitoring system (SCADA), and the enterprise resource planning system (ERP). Engineering data includes equipment operating status, production parameters, and supply chain records. At the same time, the terminal obtains a model parameter table, which contains hyperparameters (such as learning rate, batch size, number of training epochs) and task-specific parameters (such as time window for predictive maintenance, threshold for quality control, etc.).

[0096] S102. Extract initial sample data from the engineering data according to the behavior to be analyzed, where the behavior to be analyzed is predictive maintenance, quality control, or supply chain optimization;

[0097] The behavior to be analyzed includes predictive maintenance, quality control, or supply chain optimization, and relevant engineering data is screened according to the behavior to be analyzed. For example, predictive maintenance requires equipment sensor data (temperature, vibration), quality control focuses on production parameters (speed, humidity), and supply chain optimization involves inventory levels, order data, etc. Feature selection techniques (such as Pearson correlation coefficient, mutual information analysis) are used for data extraction to screen out irrelevant variables, and the engineering data after screening is the initial sample data.

[0098] S103. Clean the initial sample data set to obtain a target sample data set;

[0099] The process of cleaning the initial sample data includes: missing value handling, outlier handling, data deduplication, and data standardization.

[0100] Specifically, since the data in the initial sample data set retains all the valid data for the behavior to be analyzed, and different behaviors to be analyzed correspond to different training models, the initial sample data set obtained according to the behavior to be analyzed contains the same data required for training different models, and these data are duplicate data. In addition, engineering data is generated based on engineering logic and the operation data of the corresponding production line, so there may be missing values in its parameters. Therefore, other methods are needed to fill in the missing values, and at the same time, the outlier values recorded by the equipment also need to be processed accordingly to obtain complete trainable data. Finally, these data are standardized. After the above-mentioned behavior processing, a target sample data set is obtained.

[0101] S104. Determine model parameters according to the behavior to be analyzed, and obtain target model training data from the target sample data set according to the model parameters;

[0102] The terminal selects a suitable model architecture according to different behaviors to be analyzed. For example, LSTM / GRU is selected for predictive maintenance, CNN is selected for quality control, and a regression model (XGBoost) or a time series model (ARIMA) is selected for supply chain optimization.

[0103] Subsequently, the target model training data is divided into a training set (70%), a validation set (15%), and a test set (15%) to ensure the generalization ability of the model on different data sets.

[0104] S105. Use the batch training method to train the model corresponding to the behavior to be analyzed through the target model training data until all sample data in the target sample data set is completed, and output the basic model.

[0105] The target model training data is split into several small batches according to the batch size (e.g., 32). Stochastic Gradient Descent (SGD) or Adam (Adaptive Moment Estimation) optimizer is used for iterative training. The model performance is evaluated through the validation set, and an early stopping mechanism is adopted, that is, the training is terminated when there is no obvious improvement in several consecutive rounds, and the basic model is output after the early stopping mechanism is triggered.

[0106] Specifically, for different model architectures, there are differences in their training processes, but the data used in the actual training process all come from the target model training data.

[0107] For predictive maintenance, training needs to be carried out through a model that can process time series data and can remember long-term dependencies. Therefore, the LSTM (Long Short-Term Memory) / GRU (Gated Recurrent Unit) model is selected, and a sliding window is used, and the RMSprop (Root Mean Square Propagation) or Adam (Adaptive Moment Estimation) optimizer is used for training.

[0108] For quality control, training needs to be carried out through a model that is applicable to images / high-dimensional data and can automatically extract features. Therefore, the CNN (Convolutional Neural Network) model is selected, and data augmentation, Batch Normalization, and Adam optimization are adopted.

[0109] There are two types of requirements for supply chain optimization. The first is time series requirements, and the second is non-time series requirements. In time series requirements, a model suitable for short-term time series prediction and with high computational efficiency is needed for training. Therefore, the ARIMA (AutoRegressive Integrated Moving Average) model is selected, and the parameters are determined through ACF / PACF (Autocorrelation Function / Partial Autocorrelation Function), and the AIC (Akaike Information Criterion) / BIC (Bayesian Information Criterion) is used to select the model for training; in non-time series requirements, high-dimensional data needs to be processed, and a model suitable for regression / classification tasks is used for training. Therefore, the XGBoost (eXtreme Gradient Boosting) model is selected, and training is carried out by adopting gradient boosting decision trees and hyperparameter tuning.

[0110] S106. Combine the basic model with generative artificial intelligence to obtain the target artificial intelligence;

[0111] Use the output of the basic model as the input of the generative AI, and optimize the generated content through structured prompts; Fine-tune with pre-trained language models (such as GPT, deepseek) so that it can generate an explanatory analysis report, such as: the device failure probability is 80%, and it is recommended to maintain within 24 hours.

[0112] S107. Monitor and analyze the dynamics of the sample through the target artificial intelligence, and fine-tune the basic model;

[0113] The target artificial intelligence continuously monitors new input data and detects changes in the data distribution; Set fine-tuning trigger conditions, such as triggering fine-tuning when the model performance degrades or the data pattern changes; Adopt incremental learning or transfer learning to update the model weights to make it adapt to new data.

[0114] S108. Monitor the convergence state of the basic model after fine-tuning;

[0115] Calculate key indicators such as the loss value, accuracy, and F1-score (F1 score) of the model after fine-tuning on the validation set; Set convergence judgment conditions, such as the loss value tends to be stable or there is no improvement in consecutive multiple epochs (training rounds); If the model does not converge, continue to fine-tune until the convergence conditions are met.

[0116] S109. When the convergence state of the basic model after fine-tuning is better than the current convergence state, deploy the updated model as the basic model of the target artificial intelligence.

[0117] Compare the performance metrics of the fine-tuned model with the current base model to ensure the optimization effect; if the performance of the new model is better than the current model, deploy it to the target artificial intelligence system to replace the old model; record the model version information to ensure traceability and achieve long-term stable optimization.

[0118] In this embodiment, first, data cleaning is performed on the model training data of the behavior to be analyzed to improve the accuracy of model training and reduce the occupancy of the terminal by model data training. By combining the trained model and artificial intelligence, a full-process closed-loop of efficient data processing, accurate model training, AI intelligent analysis, automated fine-tuning, and model adaptive optimization is achieved. This not only improves the long-term applicability of the model but also greatly reduces the manual maintenance cost, providing an efficient and scalable solution for the intelligent upgrade of the manufacturing industry.

[0119] Please refer to Figure 2a and Figure 2b , another embodiment of the model optimization and update method for generative artificial intelligence is provided in this embodiment of the application. This embodiment includes:

[0120] S201. Obtain the engineering data to be trained;

[0121] S202. Extract the initial sample data from the engineering data according to the behavior to be analyzed, where the behavior to be analyzed is predictive maintenance, quality control, or supply chain optimization;

[0122] Steps S201 to S202 in this embodiment are similar to steps S101 to S102 in the foregoing embodiment, and will not be elaborated here specifically.

[0123] S203. Detect and clean the missing values of the initial sample data according to the preset rules;

[0124] Specifically, the terminal detects the missing values of the initial sample data and adopts corresponding filling strategies according to different data types. The preset rules are as follows:

[0125] For continuous data (such as temperature, pressure, humidity), if the data missing ratio is low (such as <5%), linear interpolation or KNN (K-Nearest Neighbors) interpolation is used for filling; if the missing ratio is high, mean filling or regression prediction filling is used.

[0126] For discrete data (such as product category, equipment status), mode filling (select the value that appears most frequently) is used; for special cases (such as no suitable filling value), the "unknown" category is filled.

[0127] For time series data, forward filling or backward filling is used to maintain time continuity.

[0128] S204. Detect the outliers in the initial sample data, and remove or replace the outliers according to the requirements of the behavior to be analyzed;

[0129] The detection and handling methods of outliers include: based on statistical methods, using the Z-score method (if the Z value of a data point is greater than 3 or less than -3, it is regarded as an outlier); using the IQR (Interquartile Range) method (quartile range method), calculating the upper and lower quartiles (Q1 and Q3) of the data, and determining whether the points exceeding 1.5 times the IQR range;

[0130] Based on business logic, when the device operation parameters are abnormal (such as the temperature exceeding the rated range of the device), business logic rules may be used to eliminate or set thresholds for correction; for abnormal data that cannot be determined, the historical median or mean can be used for replacement.

[0131] S205. Establish an association key between the production line ID and the timestamp, and delete duplicate records with the association key as the index to complete data cleaning;

[0132] Establish a unique index through the production line ID or device ID and the timestamp to ensure the uniqueness of the data;

[0133] If it is found that the data with the same ID and timestamp appears multiple times, the following strategies are adopted:

[0134] Directly delete the exactly same duplicate data; when there are slight differences in the values (such as multiple records of the temperature sensor), the terminal adopts mean merging or selects the latest data; for different non-keyword fields (such as the change of operator ID), the terminal adopts the primary key + time-recent data strategy to retain the latest record.

[0135] S206. Calculate the mean and standard deviation of each data feature in the initial sample data through the Z-score standardization method;

[0136] Specifically, the mean of each feature and the standard deviation of the feature are calculated through the following formulas:

[0137] Mean:

[0138]

[0139] Among them, N is the total number of data features, x i is the i th value of the data feature, i is the index, indicating the serial number of the current data feature i = 1,2,3…N .

[0140] Standard deviation:

[0141]

[0142] Among them, N is the total number of data features, x i is the value of the i th data feature, i is the index, representing the serial number of the current data feature i = 1,2,3…N , μ is the mean of the data features.

[0143] S207. Standardize the data features corresponding to the mean and the standard deviation. The standardization process is calculated by the following formula:

[0144] ;

[0145] Among them, the x is the current data point of the data feature, μ is the mean of the data feature corresponding to this data point, σ is the standard deviation of the data feature corresponding to this data point.

[0146] Specifically, all features are standardized for each data included in the data features by the Z-score method according to the above calculation method, so that the mean of these data features is 0 and the standard deviation is 1. The standardization process can prevent the influence of data with different dimensions on model training and improve the model convergence speed.

[0147] S208. Obtain all the data with a mean of 0 and a standard deviation of 1 to get the target sample data.

[0148] Perform data standardization on the initial sample data after data cleaning to obtain the target sample data.

[0149] Specifically, after standardization, all data is converted to a standard normal distribution (mean 0, standard deviation 1). The standardized data will be used as the target sample dataset for subsequent model training.

[0150] S209. Determine the model architecture through the behavior to be analyzed, and determine the model parameters according to the model architecture;

[0151] Specifically, the behavior to be analyzed includes predictive maintenance, quality control, and supply chain optimization.

[0152] Among them, predictive maintenance, quality control, and supply chain optimization select different model architectures according to training parameters and purposes. According to the foregoing description, the model architectures corresponding to different behaviors to be analyzed are as follows:

[0153] Predictive maintenance: Using LSTM (Long Short-Term Memory Network) or GRU (Gated Recurrent Unit), suitable for processing time series data;

[0154] Quality control: Using CNN (Convolutional Neural Network) or Transformer, applicable to image analysis or high-dimensional production data;

[0155] Supply chain optimization: Using XGBoost (Gradient Boosting Decision Tree) or ARIMA (Autoregressive Integrated Moving Average Model).

[0156] And, when performing model training, hyperparameter presets are required. The specific parameters include: learning rate, batch size, number of training epochs, and Dropout rate (to prevent overfitting). These parameters have fixed options when performing model presets and are generally selected according to the volume and accuracy of the model training samples.

[0157] S210. Split the target sample dataset to obtain target model training data. The target sample dataset is split into target model training data, target model validation data, and target model test data.

[0158] In this embodiment, the target model training data, target model validation data, and target model test data are divided in a ratio of 7:1.5:1.5, that is:

[0159] Training set (Train Set, 70%): Used to train the model and update model parameters;

[0160] Validation set (Validation Set, 15%): Used to evaluate the model and adjust hyperparameters during training;

[0161] Test set (Test Set, 15%): Tests the model performance in the final stage to ensure generalization ability.

[0162] And according to the model architecture, random sampling or time series splitting (applicable to time series data) is selected to split the target sample dataset.

[0163] S211. Split the target model training data into several small batches according to the preset split density;

[0164] Specifically, the terminal sets the training data batch size according to the batch size selected by the user in the model training interface (such as batch size = 32 or 64);

[0165] The training dataset is split into multiple mini-batches. Only one mini-batch is input in each iteration to reduce the computational complexity and improve the training efficiency. The DataLoader is used to automatically load the batch data to increase the training speed.

[0166] S212. Input the mini-batch data into the model corresponding to the behavior to be analyzed one by one, so that the model corresponding to the behavior to be analyzed iterates, and the model corresponding to the behavior to be analyzed is trained multiple times to obtain multiple training results.

[0167] Specifically, the training process includes: data input, model iteration, and optimization algorithm.

[0168] Among them, for data input, in each iteration during the training process, one mini-batch of data is input. For example, if the Batch Size is set to 32, 32 sample data will be input in each round of training. Data augmentation (such as random rotation, cropping, noise perturbation) is used to optimize the training data to improve the generalization ability of the model (applicable to image or sensor data). For model iteration, the forward propagation is used to calculate the prediction result; the loss value is calculated, such as MSE (Mean Squared Error), Cross-Entropy, etc.; the model parameters are adjusted through backpropagation. For the optimization algorithm, first, a suitable optimizer is selected, such as Adam, SGD (Stochastic Gradient Descent), RMSprop, to dynamically adjust the learning rate; then Batch Normalization is used to improve the training stability.

[0169] S213. For each obtained training result, verify the model performance of the training result through the target model to evaluate the data.

[0170] For the training results, appropriate evaluation methods need to be selected according to different behaviors to be analyzed. The specific evaluation methods are as follows: For the evaluation metrics, predictive maintenance (time series prediction) will be evaluated through RMSE (Root Mean Squared Error) and MAE (Mean Absolute Error); quality control (classification problem) will be evaluated through accuracy, precision, recall, and F1-score; supply chain optimization (regression problem) will be evaluated through R² (Coefficient of Determination) and MAPE (Mean Absolute Percentage Error).

[0171] After the evaluation is completed, K-Fold Cross Validation is adopted to improve the robustness of the model. Among them, the terminal will monitor the loss change in real time, that is, by plotting the loss curve during the training process, and judging whether the training result converges according to the loss curve.

[0172] S214. Stop training when the model performance reaches the preset performance metrics and meets the early stopping mechanism, and output the basic model.

[0173] Specifically, the early stopping mechanism is a method to prevent model overfitting. During the model training process, it monitors the performance on the validation set. When the performance no longer improves, it stops training in advance to avoid unnecessary calculations and improve the generalization ability. First, set the patience value through the early stopping mechanism. Generally, if the performance does not improve after 10 training rounds, the training is terminated. This is to avoid overfitting, that is, the model performs well on the training set but the effect decreases on the test set.

[0174] After training is completed, save the best model (based on the best performance on the validation set). Adopt model compression and knowledge distillation to optimize the model calculation efficiency.

[0175] S215. Link the input end of the basic model with the input end of the initial generative artificial intelligence;

[0176] Use the prediction result output by the basic model as the input of the generative AI. The artificial intelligence will bind the prediction result with the dialog box and learn the prediction result output by the basic model. When the staff makes a question input. For example, the predictive maintenance model outputs "equipment failure probability = 80%", which is transmitted to the generative AI for explanation and suggestion generation. Use API or message queue (such as Kafka) to implement the communication between the basic model and the generative AI. If the generative AI adopts the Transformer architecture, use Prompt Engineering to optimize the input format.

[0177] S216. Generate structured prompt words according to the output format of the basic model, and perform boundary restriction on the feedback of the initial generative artificial intelligence according to the structured prompt words to obtain the target artificial intelligence.

[0178] Structured prompt words (Prompt Engineering): For example, when the input is: "equipment failure probability = 80%, please generate maintenance suggestions.", the generative AI outputs: "It is recommended to perform preventive maintenance within 24 hours and check the temperature sensor." And through boundary restriction, avoid the AI from outputting irrelevant or incorrect information, such as: limit the AI to only generate operational suggestions and not output irrelevant speculations. At the same time, adopt the temperature control mechanism to reduce the randomness of the generative AI and improve the controllability.

[0179] S217. Deploy the target artificial intelligence so that the target artificial intelligence dynamically monitors the database of the data provider of the engineering data to be trained;

[0180] After the deployment of the target artificial intelligence is completed according to the foregoing steps, continuously monitor the data of systems such as MES and SCADA through the target artificial intelligence. When monitoring, adopt streaming processing technologies (such as Apache Kafka and Spark Streaming) to support real-time data updates. When a change in data distribution (such as a production process adjustment) is detected, the AI records this change and prepares for model fine-tuning.

[0181] S218. Set the trigger condition for the target artificial intelligence to fine-tune the basic model;

[0182] Specifically, model fine-tuning is triggered when any of the following conditions is met:

[0183] Performance degradation: For example, the accuracy on the validation set drops by more than 5%;

[0184] Data distribution change: Statistical information such as the feature mean and standard deviation of new data changes significantly;

[0185] Appearance of a new failure mode: For example, a new type of failure appears in the equipment maintenance log.

[0186] The trigger mechanism can be set to trigger regularly or trigger by an event to ensure that the model always adapts to the latest data.

[0187] S219. When the trigger condition is met, control the target artificial intelligence to fine-tune the basic model according to the new sample data corresponding to the dynamics of the analysis sample.

[0188] Specifically, the fine-tuning strategy is to adopt incremental learning to perform a lightweight update of the model using new samples; if the data changes greatly, then adopt transfer learning to perform retraining on the basis of the pre-trained model.

[0189] The fine-tuning process is as follows: First, collect new samples and re-partition the training / validation data sets; then perform parameter adjustment, that is, retrain some layers of the model (such as adjusting the weights of the last few layers); finally, compare the performance of the new and old models to ensure that the new model after fine-tuning has better performance than the old model.

[0190] S220. Monitor the convergence state of the basic model after fine-tuning;

[0191] S221. When the convergence state of the basic model after fine-tuning is better than the current convergence state, deploy the updated model as the basic model of the target artificial intelligence.

[0192] Steps S220 to S221 in this embodiment are similar to steps S108 to S109 in the foregoing embodiment, and will not be elaborated here specifically.

[0193] This embodiment covers the entire process from batch training, model optimization, combination with generative artificial intelligence, target AI deployment, dynamic monitoring to automatic fine-tuning. These steps ensure the adaptability and long-term optimization ability of the AI model, realizing the intelligent upgrade of the manufacturing behavior analysis system.

[0194] The above has described in detail the model optimization and update method of generative artificial intelligence in the embodiments of this application. Next, the model optimization and update device of generative artificial intelligence will be described in detail.

[0195] Please refer to Figure 3 , an embodiment of the model optimization and update device of generative artificial intelligence is provided in the embodiments of this application. This embodiment includes:

[0196] The first acquisition unit 301 is used to acquire engineering data to be trained;

[0197] The extraction unit 302 is used to extract initial sample data from the engineering data according to the behavior to be analyzed, and the behavior to be analyzed is predictive maintenance, quality control or supply chain optimization;

[0198] The data cleaning unit 303 is used to clean the initial sample data set to obtain a target sample data set;

[0199] The second acquisition unit 304 is used to determine model parameters according to the behavior to be analyzed, and acquire target model training data from the target sample data set according to the model parameters;

[0200] The model training unit 305 is used to train the model corresponding to the behavior to be analyzed through the target model training data using the batch training method until all sample data in the target sample data set are completed, and output a basic model;

[0201] The combination unit 306 is used to combine the basic model with generative artificial intelligence to obtain a target artificial intelligence;

[0202] The first monitoring unit 307 is used to monitor the dynamics of the analysis samples through the target artificial intelligence and fine-tune the basic model;

[0203] The second monitoring unit 308 is used to monitor the convergence state after the basic model is fine-tuned;

[0204] The deployment unit 309 is used to deploy the updated model as the basic model of the target artificial intelligence when the convergence state after the basic model is fine-tuned is better than the current convergence state.

[0205] In this embodiment, the data cleaning unit 303 is specifically used for:

[0206] Detect and clean the missing values of the initial sample data according to preset rules;

[0207] Detect the outliers in the initial sample data, and remove or replace the outliers according to the requirements of the behavior to be analyzed;

[0208] Establish an association key between the production line ID and the timestamp, and delete duplicate records with the association key as the index to complete data cleaning;

[0209] Perform data standardization on the initial sample data after data cleaning to obtain target sample data.

[0210] In this embodiment, the data cleaning unit 303 is specifically further configured to:

[0211] Calculate the mean and standard deviation of each data feature in the initial sample data through the Z-score standardization method;

[0212] Standardize the data features corresponding to the mean and the standard deviation, and the standardization process is calculated through the following formula:

[0213] ;

[0214] Wherein, x is the current data point of the data feature, μ is the mean of the data feature corresponding to this data point, and σ is the standard deviation of the data feature corresponding to this data point.

[0215] Obtain all the data with a mean of 0 and a standard deviation of 1 to obtain target sample data.

[0216] In this embodiment, the second obtaining unit 304 is specifically configured to:

[0217] Determine the model architecture through the behavior to be analyzed, and determine the model parameters according to the model architecture;

[0218] Split the target sample data set to obtain target model training data, and the target sample data set is split into target model training data, target model validation data, and target model test data.

[0219] In this embodiment, the model training unit 305 is specifically configured to:

[0220] Split the target model training data into several small batches according to a preset split density;

[0221] Input the small batches of data into the model corresponding to the behavior to be analyzed one by one, so that the model corresponding to the behavior to be analyzed iterates, and perform multiple trainings on the model corresponding to the behavior to be analyzed to obtain multiple training results;

[0222] For each obtained training result, the model performance of the training result is evaluated through the target model verification data;

[0223] When the model performance reaches the preset performance index and meets the early stopping mechanism, stop training and output the basic model.

[0224] In this embodiment, the combining unit 306 is specifically configured to:

[0225] Link the input end of the basic model with the input end of the initial generative artificial intelligence;

[0226] Generate a structured prompt word according to the output format of the basic model, and perform boundary restriction on the feedback of the initial generative artificial intelligence according to the structured prompt word to obtain the target artificial intelligence.

[0227] In this embodiment, the first monitoring unit 307 is specifically configured to:

[0228] Deploy the target artificial intelligence so that the target artificial intelligence dynamically monitors the database of the data provider of the engineering data to be trained;

[0229] Set the trigger condition for the target artificial intelligence to fine-tune the basic model;

[0230] When the trigger condition is met, control the target artificial intelligence to fine-tune the basic model according to the new sample data dynamically corresponding to the analysis sample.

[0231] In this embodiment, the functions of each unit correspond to the steps in the foregoing Figure 1 、 Figure 2a and Figure 2b The embodiments shown will not be repeated here.

[0232] Please refer to Figure 4 , another embodiment of the model optimization and update device for the generative artificial intelligence provided by the embodiment of the present application includes:

[0233] A processor 401, a memory 402, an input / output unit 403, and a bus 404;

[0234] The processor 401 is connected to the memory 402, the input / output unit 403, and the bus 404;

[0235] The processor 401 specifically executes the operations corresponding to the steps in Figure 1 、 Figure 2a and Figure 2b The methods will not be repeated here specifically.

[0236] This application also relates to a computer-readable storage medium, on which a program is stored. When the program runs on a computer, the computer is caused to execute any of the above methods.

[0237] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0238] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0239] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0240] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0241] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.

Claims

1. A model optimization and updating method for generative artificial intelligence, characterized in that: The method comprises: Get the engineering data to be trained; Extracting initial sample data from engineering data according to the behavior to be analyzed, wherein the behavior to be analyzed is a predictive maintenance behavior, a quality control behavior, or a supply chain optimization behavior; Detecting and cleaning missing values ​​of the initial sample data according to preset rules; Detecting outliers in the initial sample data, and removing or replacing the outliers according to the needs of the behavior to be analyzed; Establish an association key between the production line ID and the timestamp, use the association key as the index to delete duplicate records, and complete data cleaning; Performing data standardization on the initial sample data after data cleaning to obtain target sample data; Determining model parameters according to the behavior to be analyzed, and acquiring target model training data from the target sample data set according to the model parameters; The model corresponding to the behavior to be analyzed is trained by the target model training data using a batch training method until all sample data in the target sample data set are trained, and a basic model is output, the predictive maintenance behavior uses a memory network model for model training, the quality control behavior uses an image convolution model for model training, and the supply chain optimization behavior selects a regression model or a time series model for model training; Combining the basic model with generative artificial intelligence to obtain target artificial intelligence; Monitor and analyze the dynamics of samples through the target artificial intelligence and fine-tune the basic model; Monitoring the convergence status of the base model after fine-tuning; When the convergence state of the basic model after fine-tuning is better than the current convergence state, the updated model is deployed as the basic model of the target artificial intelligence.

2. The model optimization and updating method according to claim 1, characterized in that: The performing data standardization on the initial sample data after data cleaning to obtain target sample data includes: Calculate the mean and standard deviation of each data feature in the initial sample data by using the Z-score standardization method; The data features corresponding to the mean and the standard deviation are standardized, and the standardization process is calculated by the following formula: ; Wherein, x is the current data point of the data feature, μ is the mean of the data feature corresponding to the data point, and σ is the standard deviation of the data feature corresponding to the data point; Obtain all data whose mean is 0 and whose standard deviation is 1 to obtain target sample data.

3. The model optimization and updating method according to claim 1, characterized in that: The target sample data set is divided into target model training data, target model verification data and target model test data.

4. The model optimization and updating method according to claim 3, characterized in that: The method of using a batch training method to train the model corresponding to the behavior to be analyzed by using the target model training data until all sample data in the target sample data set have completed training and outputting a basic model includes: Splitting the target model training data into a number of small batches of data according to a preset split density; Inputting the small batch data into the model corresponding to the behavior to be analyzed one by one, so that the model corresponding to the behavior to be analyzed is iterated, and training the model corresponding to the behavior to be analyzed multiple times to obtain multiple training results; Each time a training result is obtained, the model performance of the training result is evaluated by using the target model verification data; When the model performance reaches the preset performance index and satisfies the early stopping mechanism, the training is stopped and the basic model is output.

5. The model optimization and updating method according to any one of claims 1 to 4, characterized in that: The step of combining the basic model with generative artificial intelligence to obtain target artificial intelligence includes: Linking the input of the base model to the input of the initial generative artificial intelligence; Structured prompt words are generated according to the output format of the basic model, and boundary restrictions are imposed on the feedback of the initial generative artificial intelligence according to the structured prompt words to obtain the target artificial intelligence.

6. The model optimization and updating method according to any one of claims 1 to 4, characterized in that: The target artificial intelligence monitors and analyzes the dynamics of the sample and fine-tunes the basic model, including: Deploy the target artificial intelligence so that the target artificial intelligence dynamically monitors the database of the data provider of the engineering data to be trained; Setting a trigger condition for the target artificial intelligence to fine-tune the basic model; When the trigger condition is met, the target artificial intelligence is controlled to fine-tune the basic model according to the new sample data that dynamically corresponds to the analysis sample.

7. The model optimization and updating method according to any one of claims 1 to 4, characterized in that: The extracting of initial sample data from the engineering data according to the behavior to be analyzed includes: Respectively obtaining actual parameters required for predictive maintenance behavior, quality control behavior or supply chain optimization behavior in the behavior to be analyzed; The actual parameter data of the behavior to be analyzed are acquired one by one from the engineering data according to the actual parameters to obtain initial sample data.

8. A model optimization and updating device for generative artificial intelligence, characterized in that: The device comprises: A first acquisition unit, used to acquire engineering data to be trained; An extraction unit, configured to extract initial sample data from the engineering data according to the behavior to be analyzed, wherein the behavior to be analyzed is predictive maintenance, quality control, or supply chain optimization; A data cleaning unit, used to clean the initial sample data set to obtain a target sample data set; A second acquisition unit, used to determine model parameters according to the behavior to be analyzed, and acquire target model training data from the target sample data set according to the model parameters; A model training unit, used to train the model corresponding to the behavior to be analyzed through the target model training data using a batch training method until all sample data in the target sample data set have been trained, and output a basic model; A combining unit, used to combine the basic model with the generative artificial intelligence to obtain the target artificial intelligence; A first monitoring unit, used to monitor and analyze the dynamics of the sample through the target artificial intelligence and fine-tune the basic model; A second monitoring unit, used to monitor the convergence state of the basic model after fine-tuning; A deployment unit is used to deploy the updated model as the basic model of the target artificial intelligence when the convergence state of the basic model after fine-tuning is better than the current convergence state.

9. A model optimization and updating device for generative artificial intelligence, characterized in that: The device comprises: Processor, memory, input-output unit, and bus; The processor is connected to the memory, the input and output unit, and the bus; The memory stores a program, and the processor calls the program to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a program stored thereon, wherein the program, when executed on a computer, performs the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Detection model training method and device, computer equipment and storage medium

    CN118505230A

  • Power prediction method, system and equipment based on group algorithm

    CN118643941A