Generative artificial intelligence model optimization updating method and device and medium
By cleaning and standardizing engineering data, training and combining with generative artificial intelligence, the repetitive data occupation problem during model training in the existing technology is solved, and efficient and accurate model training and long-term optimization are achieved.
Patent Information
- Application Number
- CN202510436389.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In the prior art, when training models for predictive maintenance, quality control and supply chain optimization, there is repeated data obtained, resulting in excessive consumption of model training equipment.
By obtaining the engineering data to be trained, the initial sample data is extracted, the data is cleaned and standardized, the model parameters are determined, and the model is trained using batch training methods until all sample data is completed training, and the basic model is output. Combining the basic model with generative artificial intelligence, the model is updated to optimize its convergence state through dynamic monitoring and fine-tuning.
This reduces the repetitive data space occupancy during model training with different behaviors to be analyzed, improves the efficiency and accuracy of model training, reduces the equipment occupancy of model training, and realizes long-term optimization and adaptability of the model.
Smart Images

Figure CN119961682A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a method, device and medium for optimizing and updating a model of generative artificial intelligence. Background Art
[0002] The current manufacturing industry generally uses static models trained offline based on historical data in areas such as predictive maintenance, quality control, and supply chain optimization. These models achieve high prediction accuracy through batch training in the early stages, but once deployed, their parameters are difficult to update based on real-time changes in the production environment, equipment status, or market demand.
[0003] In the prior art, predictive maintenance, quality control and supply chain optimization correspond to three different prediction models respectively. The training data of these models come from the same engineering data. When the models are actually trained, there is a certain intersection in the training data required by these models, which results in duplicate data when the training data of different models is obtained through engineering data. This causes the duplicate data between different model training samples to excessively occupy the model training equipment. Summary of the invention
[0004] In order to solve the above technical problems, the present application provides a generative artificial intelligence model optimization and updating method, device and medium, which are used to reduce the space occupied by repeated training data when training behavior analysis models of different behaviors to be analyzed.
[0005] The technical solution provided in this application is described below: The first aspect of the present application provides a generative artificial intelligence model optimization and updating method, comprising: Get the engineering data to be trained; Extracting initial sample data from engineering data according to the behavior to be analyzed, wherein the behavior to be analyzed is predictive maintenance, quality control, or supply chain optimization; Perform data cleaning on the initial sample data set to obtain the target sample data set; Determining model parameters according to the behavior to be analyzed, and acquiring target model training data from the target sample data set according to the model parameters; Using a batch training method to train the model corresponding to the behavior to be analyzed through the target model training data until all sample data in the target sample data set have completed training, and outputting a basic model; Combining the basic model with generative artificial intelligence to obtain target artificial intelligence; Monitor and analyze the dynamics of samples through the target artificial intelligence and fine-tune the basic model; Monitoring the convergence status of the base model after fine-tuning; When the convergence state of the basic model after fine-tuning is better than the current convergence state, the updated model is deployed as the basic model of the target artificial intelligence.
[0006] Optionally, performing data cleaning on the initial sample data set to obtain the target sample data set includes: Detecting and cleaning missing values of the initial sample data according to preset rules; Detecting outliers in the initial sample data, and removing or replacing the outliers according to the needs of the behavior to be analyzed; Establish an association key between the production line ID and the timestamp, use the association key as the index to delete duplicate records, and complete data cleaning; Data standardization is performed on the initial sample data that has completed data cleaning to obtain target sample data.
[0007] Optionally, performing data standardization on the initial sample data after data cleaning to obtain target sample data includes: Calculate the mean and standard deviation of each data feature in the initial sample data by using the Z-score standardization method; The data features corresponding to the mean and the standard deviation are standardized, and the standardization process is calculated by the following formula: ; Wherein, x is the current data point of the data feature, μ is the mean of the data feature corresponding to the data point, and σ is the standard deviation of the data feature corresponding to the data point.
[0008] Obtain all data whose mean is 0 and whose standard deviation is 1 to obtain target sample data.
[0009] Optionally, determining model parameters according to the behavior to be analyzed, and acquiring target model training data from the target sample data set according to the model parameters, includes: Determine a model framework through the behavior to be analyzed, and determine model parameters according to the model framework; The target sample data set is split to obtain target model training data, wherein the target sample data set is split into target model training data, target model verification data and target model test data.
[0010] Optionally, the step of training the model corresponding to the behavior to be analyzed by using the target model training data using a batch training method until all sample data in the target sample data set have completed training and outputting a basic model includes: Splitting the target model training data into a number of small batches of data according to a preset split density; Inputting the small batch data into the model corresponding to the behavior to be analyzed one by one, so that the model corresponding to the behavior to be analyzed is iterated, and training the model corresponding to the behavior to be analyzed multiple times to obtain multiple training results; Each time a training result is obtained, the model performance of the training result is evaluated by using the target model verification data; When the model performance reaches the preset performance index and satisfies the early stopping mechanism, the training is stopped and the basic model is output.
[0011] Optionally, combining the basic model with generative artificial intelligence to obtain target artificial intelligence includes: Linking the input of the base model to the input of the initial generative artificial intelligence; Structured prompt words are generated according to the output format of the basic model, and boundary restrictions are imposed on the feedback of the initial generative artificial intelligence according to the structured prompt words to obtain the target artificial intelligence.
[0012] Optionally, the monitoring and analyzing the dynamics of the sample through the target artificial intelligence and fine-tuning the basic model include: Deploy the target artificial intelligence so that the target artificial intelligence dynamically monitors the database of the data provider of the engineering data to be trained; Setting a trigger condition for the target artificial intelligence to fine-tune the basic model; When the trigger condition is met, the target artificial intelligence is controlled to fine-tune the basic model according to the new sample data that dynamically corresponds to the analysis sample.
[0013] A second aspect of the present application provides a model optimization and updating device for generative artificial intelligence, the device comprising: A first acquisition unit, used to acquire engineering data to be trained; An extraction unit, configured to extract initial sample data from the engineering data according to the behavior to be analyzed, wherein the behavior to be analyzed is predictive maintenance, quality control, or supply chain optimization; A data cleaning unit, used to clean the initial sample data set to obtain a target sample data set; A second acquisition unit, used to determine model parameters according to the behavior to be analyzed, and acquire target model training data from the target sample data set according to the model parameters; A model training unit, used to train the model corresponding to the behavior to be analyzed through the target model training data using a batch training method until all sample data in the target sample data set have been trained, and output a basic model; A combining unit, used to combine the basic model with the generative artificial intelligence to obtain the target artificial intelligence; A first monitoring unit, used to monitor and analyze the dynamics of the sample through the target artificial intelligence and fine-tune the basic model; A second monitoring unit, used to monitor the convergence state of the basic model after fine-tuning; A deployment unit is used to deploy the updated model as the basic model of the target artificial intelligence when the convergence state of the basic model after fine-tuning is better than the current convergence state.
[0014] Optionally, the data cleaning unit is specifically used for: Detecting and cleaning missing values of the initial sample data according to preset rules; Detecting outliers in the initial sample data, and removing or replacing the outliers according to the needs of the behavior to be analyzed; Establish an association key between the production line ID and the timestamp, use the association key as the index to delete duplicate records, and complete data cleaning; Data standardization is performed on the initial sample data that has completed data cleaning to obtain target sample data.
[0015] Optionally, the data cleaning unit is further used for: Calculate the mean and standard deviation of each data feature in the initial sample data by using the Z-score standardization method; The data features corresponding to the mean and the standard deviation are standardized, and the standardization process is calculated by the following formula: ; Wherein, x is the current data point of the data feature, μ is the mean of the data feature corresponding to the data point, and σ is the standard deviation of the data feature corresponding to the data point.
[0016] Obtain all data whose mean is 0 and whose standard deviation is 1 to obtain target sample data.
[0017] Optionally, the second acquiring unit is specifically configured to: Determine a model framework through the behavior to be analyzed, and determine model parameters according to the model framework; The target sample data set is split to obtain target model training data, wherein the target sample data set is split into target model training data, target model verification data and target model test data.
[0018] Optionally, the model training unit is specifically used for: Splitting the target model training data into a number of small batches of data according to a preset split density; Inputting the small batch data into the model corresponding to the behavior to be analyzed one by one, so that the model corresponding to the behavior to be analyzed is iterated, and training the model corresponding to the behavior to be analyzed multiple times to obtain multiple training results; Each time a training result is obtained, the model performance of the training result is evaluated by using the target model verification data; When the model performance reaches the preset performance index and satisfies the early stopping mechanism, the training is stopped and the basic model is output.
[0019] Optionally, the combining unit is specifically used for: Linking the input of the base model to the input of the initial generative artificial intelligence; Structured prompt words are generated according to the output format of the basic model, and boundary restrictions are imposed on the feedback of the initial generative artificial intelligence according to the structured prompt words to obtain the target artificial intelligence.
[0020] Optionally, the first monitoring unit is specifically used for: Deploy the target artificial intelligence so that the target artificial intelligence dynamically monitors the database of the data provider of the engineering data to be trained; Setting a trigger condition for the target artificial intelligence to fine-tune the basic model; When the trigger condition is met, the target artificial intelligence is controlled to fine-tune the basic model according to the new sample data that dynamically corresponds to the analysis sample.
[0021] A third aspect of the present application provides a model optimization and updating device for generative artificial intelligence, the device comprising: Processor, memory, input-output unit, and bus; The processor is connected to the memory, the input and output unit, and the bus; The memory stores a program, and the processor calls the program to execute the first aspect and any optional method in the first aspect.
[0022] A fourth aspect of the present application provides a computer-readable storage medium, on which a program is stored. When the program is executed on a computer, the program executes the first aspect and any optional method in the first aspect.
[0023] It can be seen from the above technical solutions that this application has the following advantages: After obtaining the sample data required for the corresponding behavior analysis from the historical analysis samples through the behavior to be analyzed, this application performs data cleaning on the sample data to obtain better quality training data, and reduces the occupancy of training data by comparing and deleting duplicate data during data cleaning. The behavior to be analyzed includes predictive maintenance, quality control or supply chain optimization. After that, the learning model is determined according to the behavior to be analyzed, so that the corresponding behavior to be analyzed can obtain the corresponding sample data to train a basic model for analyzing the corresponding behavior to be analyzed, determine different model frameworks according to different behaviors to be analyzed, and combine the models through multimodal logic to obtain a basic model for inputting artificial intelligence, and then link the output end of the basic model to the interactive artificial intelligence through an interface, so that the artificial intelligence generates real-time feedback while actively iterating the basic model. In order to achieve the integration of all models and model training data to reduce the occupancy of model training data, the results are quickly output through artificial intelligence and the basic model is iterated according to the input data and database update data, reducing the difficulty of using the model and improving the analysis efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solution in the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0025] Figure 1 A schematic diagram of an embodiment of the method for optimizing and updating a model of generative artificial intelligence in this application; Figure 2a A flowchart of an embodiment of the first stage of the model optimization and updating method of generative artificial intelligence in this application; Figure 2b A flowchart of an embodiment of the second stage of the generative artificial intelligence model optimization and updating method in this application; Figure 3 A schematic diagram of the structure of an embodiment of a device for optimizing and updating a model of generative artificial intelligence in this application; Figure 4 This is a schematic diagram of the structure of another embodiment of the model optimization and updating device of generative artificial intelligence in this application. DETAILED DESCRIPTION
[0026] It should be noted that the model optimization and updating method of generative artificial intelligence provided in this application can be applied to terminals, systems, and servers. For example, the terminal can be a smart phone or computer, tablet computer, smart TV, smart watch, portable computer terminal, or a fixed terminal such as a desktop computer. For the convenience of explanation, this application uses the terminal as the execution subject for example.
[0027] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0028] See also Figure 1 The present application first provides an embodiment of a model optimization and updating method for generative artificial intelligence, which includes: S101, obtaining engineering data to be trained; Engineering data is obtained through factory manufacturing execution systems (MES), supervisory control and data acquisition systems (SCADA), and enterprise resource planning systems (ERP). Engineering data includes equipment operating status, production parameters, and supply chain records. At the same time, the terminal obtains the model parameter table, which contains hyperparameters (such as learning rate, batch size, number of training rounds) and task-specific parameters (such as the time window for predictive maintenance, the threshold for quality control, etc.).
[0029] S102, extracting initial sample data from engineering data according to the behavior to be analyzed, where the behavior to be analyzed is predictive maintenance, quality control, or supply chain optimization; The behaviors to be analyzed include predictive maintenance, quality control, or supply chain optimization. Relevant engineering data is screened based on the behaviors to be analyzed. For example, predictive maintenance requires equipment sensor data (temperature, vibration), quality control focuses on production parameters (speed, humidity), and supply chain optimization involves inventory levels, order data, etc. Data extraction uses feature selection techniques (such as Pearson correlation coefficient, mutual information analysis) to screen out irrelevant variables. The engineering data after screening is the initial sample data.
[0030] S103, performing data cleaning on the initial sample data set to obtain a target sample data set; The process of initial sample data cleaning includes: missing value processing, outlier processing, data deduplication and data standardization.
[0031] Specifically, because the data in the initial sample data set retains valid data of all behaviors to be analyzed, and different behaviors to be analyzed correspond to different training models, the initial sample data set obtained according to the behaviors to be analyzed contains the same data required to train different models. These data are duplicate data. In addition, the engineering data is generated based on the engineering logic and the corresponding production line operation data of the engineering, so its parameters may have missing values. Therefore, it is necessary to use other methods to complete the missing values. At the same time, the abnormal values recorded by the equipment also need to be processed accordingly to obtain complete trainable data. Finally, these data are standardized. After the above-mentioned behavior processing, the target sample data set is obtained.
[0032] S104, determining model parameters according to the behavior to be analyzed, and acquiring target model training data from the target sample data set according to the model parameters; The terminal selects the appropriate model architecture based on different behaviors to be analyzed, such as LSTM / GRU for predictive maintenance, CNN for quality control, and regression model (XGBoost) or time series model (ARIMA) for supply chain optimization.
[0033] Subsequently, the target model training data was divided into a training set (70%), a validation set (15%), and a test set (15%) to ensure the generalization ability of the model on different data sets.
[0034] S105, using a batch training method to train the model corresponding to the behavior to be analyzed through the target model training data, until all sample data in the target sample data set have completed training, and outputting a basic model; The target model training data is split into several small batches according to the batch size (e.g. 32). Stochastic gradient descent (SGD) or Adam (Adaptive Moment Estimation) optimizer is used for iterative training. The model performance is evaluated through the validation set, and the early stopping mechanism is used, that is, the training is terminated when there is no obvious improvement after several consecutive rounds, and the basic model is output after the early stopping mechanism is triggered.
[0035] Specifically, the training process is different for different model architectures, but the data used in the actual training process comes from the target model training data.
[0036] For predictive maintenance, it is necessary to train with a model that can process time series data and remember long-term dependencies. Therefore, the LSTM (Long Short-Term Memory) / GRU (Gated Recurrent Unit) model is selected, and a sliding window is used, and the RMSprop (Root Mean Square Propagation) or Adam (Adaptive Moment Estimation) optimizer is used for training. For quality control, a model suitable for image / high-dimensional data and automatic feature extraction is needed for training, so the CNN (Convolutional Neural Network) model is selected, and data enhancement, BatchNormalization, and Adam optimization are used; There are two kinds of demands for supply chain optimization, the first is time series demand, and the second is non-time series demand. In the time series demand, a model suitable for short-term time series prediction and with high computational efficiency is needed for training, so the ARIMA (autoregressive integrated moving average model) model is selected, and the parameters are determined by ACF / PACF (autocorrelation function / autocorrelation function), and the model is selected for training using AIC (Akaike Information Criterion) / BIC (Bayesian Information Criterion); in the non-time series demand, high-dimensional data needs to be processed, and a model suitable for regression / classification tasks is needed for training, so the XGBoost (eXtremeGradient Boosting) model is selected, and training is performed by using a gradient boosting decision tree and hyperparameter tuning.
[0037] S106, combining the basic model with generative artificial intelligence to obtain target artificial intelligence; The output of the basic model is used as the input of the generative AI, and the generated content is optimized through structured prompts. Pre-trained language models (such as GPT and deepseek) are used for fine-tuning to enable them to generate explanatory analysis reports, such as: the probability of equipment failure is 80%, and maintenance is recommended within 24 hours.
[0038] S107, monitoring and analyzing the dynamics of the sample through the target artificial intelligence, and fine-tuning the basic model; The target AI continuously monitors new input data and detects changes in data distribution; sets fine-tuning trigger conditions, such as when model performance degrades or data patterns change; and uses incremental learning or transfer learning to update model weights to adapt them to new data.
[0039] S108, monitoring the convergence state of the basic model after fine-tuning; Calculate the key indicators of the fine-tuned model on the validation set, such as loss value, accuracy, F1-score (F1 score); set convergence judgment conditions, such as loss value tending to be stable or no improvement for multiple consecutive epochs (training rounds); if the model does not converge, continue to fine-tune until the convergence conditions are met.
[0040] S109. When the convergence state of the basic model after fine-tuning is better than the current convergence state, the updated model is deployed as the basic model of the target artificial intelligence.
[0041] Compare the performance indicators of the fine-tuned model with the current basic model to ensure the optimization effect; if the performance of the new model is better than the current model, deploy it to the target artificial intelligence system to replace the old model; record model version information to ensure traceability and achieve long-term stable optimization.
[0042] This embodiment first cleans the model training data of the behavior to be analyzed to improve the accuracy of model training and reduce the occupancy of the terminal by model data training. It also combines the trained model with artificial intelligence to achieve a full-process closed loop of efficient data processing, accurate model training, AI intelligent analysis, automated fine-tuning, and model adaptive optimization. This not only improves the long-term applicability of the model, but also greatly reduces the cost of manual maintenance, providing an efficient and scalable solution for the intelligent upgrade of the manufacturing industry.
[0043] See also Figure 2a and Figure 2b , the present application embodiment provides another embodiment of the model optimization and updating method of generative artificial intelligence, which embodiment includes: S201, obtaining engineering data to be trained; S202, extracting initial sample data from engineering data according to the behavior to be analyzed, where the behavior to be analyzed is predictive maintenance, quality control, or supply chain optimization; Steps S201 to S202 in this embodiment are similar to steps S101 to S102 in the aforementioned embodiment, and will not be described in detail here.
[0044] S203, detecting and cleaning missing values of the initial sample data according to preset rules; Specifically, the terminal performs missing value detection on the initial sample data and adopts corresponding filling strategies according to different data types. The preset rules are as follows: For continuous data (such as temperature, pressure, and humidity), if the missing data ratio is low (such as <5%), linear interpolation or KNN (K-Nearest Neighbors) interpolation is used to fill in the missing data; if the missing data ratio is high, mean filling or regression prediction is used to fill in the missing data.
[0045] For discrete data (such as product category, equipment status), use mode filling (select the value with the most occurrences); for special cases (such as no suitable filling value), fill in the "unknown" category.
[0046] For time series data, forward filling or backward filling is used to maintain time continuity.
[0047] S204, detecting outliers in the initial sample data, and removing or replacing the outliers according to the requirements of the behavior to be analyzed; The detection and processing methods of outliers include: based on statistical methods, using the Z-score method (if the Z value of a data point is greater than 3 or less than -3, it is considered an outlier); using the IQR (Interquartile Range) method (interquartile range method) to calculate the upper and lower quartiles (Q1 and Q3) of the data and determine whether the point exceeds 1.5 times the IQR range; Based on business logic, when equipment operating parameters are abnormal (such as temperature exceeding the rated range of the equipment), business logic rules may be used to eliminate or set thresholds for correction; for abnormal data that cannot be determined, historical medians or means can be used to replace them.
[0048] S205, establishing an association key between the production line ID and the timestamp, deleting duplicate records using the association key as an index, and completing data cleaning; Establish a unique index through production line ID or equipment ID and timestamp to ensure the uniqueness of data; If data with the same ID and timestamp appears multiple times, use the following strategy: Directly delete identical duplicate data; when there are slight differences in values (such as multiple records from a temperature sensor), the terminal uses an average merge or selects the latest data; for differences in non-key fields (such as changes in operator ID), the terminal uses the primary key + most recent data strategy to retain the latest records.
[0049] S206, calculating the mean and standard deviation of each data feature in the initial sample data by using a Z-score standardization method; Specifically, the mean of each feature and the standard deviation of the feature are calculated using the following formula: Mean:
[0050] in, Nis the total number of data features, x i For the i The value of the data feature, i is the index, indicating the sequence number of the current data feature i =[ 1,2,3…N ].
[0051] Standard Deviation:
[0052] in, N is the total number of data features, x i For the i The value of the data feature, i is the index, indicating the sequence number of the current data feature i =[ 1,2,3…N ], μ is the mean of the data features.
[0053] S207, standardizing the data features corresponding to the mean and the standard deviation, and the standardization process is calculated by the following formula: ; Among them, the x is the current data point of the data feature, μ is the mean of the data feature corresponding to the data point, σ is the standard deviation of the data feature corresponding to this data point.
[0054] Specifically, all features are calculated in the above way and the data contained in the data features are standardized one by one by the Z-score method, so that the mean of these data features is 0 and the standard deviation is 1. The standardization process can prevent the influence of data of different dimensions on model training and improve the convergence speed of the model.
[0055] S208. Obtain all data whose mean is 0 and whose standard deviation is 1 to obtain target sample data.
[0056] Data standardization is performed on the initial sample data that has completed data cleaning to obtain target sample data.
[0057] Specifically, after standardization, all data are converted to a standard normal distribution (mean 0, standard deviation 1). The standardized data will be used as the target sample data set for subsequent model training.
[0058] S209, determining a model framework through the behavior to be analyzed, and determining model parameters according to the model framework; Specifically, the behaviors to be analyzed include predictive maintenance, quality control, and supply chain optimization.
[0059] Among them, predictive maintenance, quality control, and supply chain optimization select different model architectures according to training parameters and purposes. According to the above description, the model architectures corresponding to different behaviors to be analyzed are as follows: Predictive maintenance: LSTM (Long Short-Term Memory Network) or GRU (Gated Recurrent Unit) are used, which are suitable for processing time series data; Quality control: using CNN (convolutional neural network) or Transformer, suitable for image analysis or high-dimensional production data; Supply chain optimization: Use XGBoost (gradient boosted decision tree) or ARIMA (autoregressive moving average model).
[0060] Also, when training the model, it is necessary to preset hyperparameters, including learning rate, size, number of training rounds (epochs) and Dropout rate (to prevent overfitting). These parameters have fixed options when presetting the model, and are generally selected based on the volume and accuracy of the model training samples.
[0061] S210, splitting the target sample data set to obtain target model training data, wherein the target sample data set is split into target model training data, target model verification data, and target model test data.
[0062] In this embodiment, the target model training data, the target model verification data, and the target model test data are divided in a ratio of 7:1.5:1.5, that is: Training Set (70%): used to train the model and update model parameters; Validation Set (15%): used to evaluate the model and adjust hyperparameters during training; Test Set (15%): Test model performance in the final stage to ensure generalization ability.
[0063] And according to the model architecture, select random sampling or time series splitting (applicable to time series data) to split the target sample data set.
[0064] S211, splitting the target model training data into a plurality of small batches of data according to a preset splitting density; Specifically, the terminal sets the training data batch size according to the batch size selected by the user in the model training interface (e.g., batch size=32 or 64); The training data set is split into multiple mini-batches (small batches after splitting), and only one mini-batch is input in each iteration to reduce computational complexity and improve training efficiency. The DataLoader is used to automatically load batch data and improve training speed.
[0065] S212, inputting the small batch data into the model corresponding to the behavior to be analyzed one by one, so that the model corresponding to the behavior to be analyzed is iterated, and training the model corresponding to the behavior to be analyzed multiple times to obtain multiple training results; Specifically, the training process includes: data input, model iteration and optimization algorithm.
[0066] Among them, the data input is a small batch of data input for each iteration during the training process. For example, if the Batch Size is set to 32, 32 sample data will be input for each round of training. Data enhancement (such as random rotation, cropping, and noise perturbation) is used to optimize the training data and improve the generalization ability of the model (applicable to image or sensor data). The model iteration uses forward propagation to calculate the prediction results; calculate the loss value, such as MSE (mean square error), Cross-Entropy, etc.; adjust the model parameters through backpropagation. The optimization algorithm first selects a suitable optimizer, such as Adam, SGD (stochastic gradient descent), RMSprop, and dynamically adjusts the learning rate; then Batch Normalization is used to improve training stability.
[0067] S213, each time obtaining a training result, evaluating the model performance of the training result by using the target model verification data; The training results need to select appropriate evaluation methods according to different behaviors to be analyzed. The specific evaluation methods are as follows: For evaluation indicators, predictive maintenance (time series prediction) will be evaluated through RMSE (root mean square error) and MAE (mean absolute error); quality control (classification problem) will be evaluated through accuracy, precision, recall and F1-score; supply chain optimization (regression problem) will be evaluated through R² (coefficient of determination) and MAPE (mean absolute percentage error).
[0068] After the evaluation is completed, K-Fold Cross Validation is used to improve the robustness of the model. The terminal will monitor the loss changes in real time, that is, by drawing the loss curve during the training process, and judging whether the training results have converged based on the loss curve.
[0069] S214: When the model performance reaches a preset performance indicator and satisfies an early stopping mechanism, training is stopped and a basic model is output.
[0070] Specifically, the Early Stopping mechanism is a method to prevent model overfitting. During model training, the performance on the validation set is monitored and training is stopped in advance when the performance stops improving, avoiding unnecessary calculations and improving generalization. First, the patience value is set through the Early Stopping mechanism. Generally, if the performance does not improve after 10 training rounds, the training is terminated to avoid overfitting, that is, the model performs well on the training set but the performance decreases on the test set.
[0071] After training is complete, save the best model (based on the best performance on the validation set). Use model compression and knowledge distillation to optimize model computational efficiency.
[0072] S215, linking the input end of the basic model with the input end of the initial generative artificial intelligence; The prediction results output by the basic model are used as the input of the generative AI. The artificial intelligence will bind the prediction results to the dialog box and learn the prediction results output by the basic model. When the staff asks questions, for example, the predictive maintenance model outputs "equipment failure probability = 80%", which is transmitted to the generative AI for interpretation and suggestion generation. Use API or message queue (such as Kafka) to realize the communication between the basic model and the generative AI. If the generative AI adopts the Transformer architecture, use PromptEngineering (prompt word optimization) to optimize the input format.
[0073] S216. Generate structured prompt words according to the output format of the basic model, and perform boundary restrictions on the feedback of the initial generative artificial intelligence according to the structured prompt words to obtain the target artificial intelligence.
[0074] Structured prompts (Prompt Engineering): For example, when the input is: "Equipment failure probability = 80%, please generate maintenance suggestions." Generative AI outputs: "It is recommended to perform preventive maintenance within 24 hours and check the temperature sensor." And through boundary restrictions, AI outputs irrelevant or erroneous information, such as: limiting AI to only generate operational suggestions and not output irrelevant speculations. At the same time, a temperature control mechanism (Temperature Control) is used to reduce the randomness of generative AI and improve controllability.
[0075] S217, deploying the target artificial intelligence so that the target artificial intelligence dynamically monitors the database of the data provider of the engineering data to be trained; After the deployment of the target AI is completed according to the above steps, the target AI continuously monitors the data of MES, SCADA and other systems, and uses streaming processing technology (such as Apache Kafka and Spark Streaming) to support real-time data updates. When a change in data distribution (such as production process adjustment) is detected, AI records the change and prepares for model fine-tuning.
[0076] S218, setting a trigger condition for the target artificial intelligence to fine-tune the basic model; Specifically, model fine-tuning is triggered when any of the following conditions is met: Performance degradation: if the accuracy on the validation set drops by more than 5%; Data distribution changes: statistical information such as the feature mean and standard deviation of new data changes significantly; New failure modes emerge: For example, new types of failures appear in the equipment maintenance log.
[0077] The trigger mechanism can be set to trigger periodically or trigger an event, ensuring that the model is always adapted to the latest data.
[0078] S219. When the trigger condition is met, the target artificial intelligence is controlled to fine-tune the basic model according to the new sample data that dynamically corresponds to the analysis sample.
[0079] Specifically, the fine-tuning strategy is to adopt incremental learning and use new samples to perform lightweight updates on the model; if the data changes significantly, transfer learning is used to retrain based on the pre-trained model.
[0080] The fine-tuning process is as follows: first collect new samples and re-divide the training / validation data sets; then adjust the parameters, that is, retrain some layers of the model (such as adjusting the weights of the last few layers); finally, compare the performance of the new and old models to ensure that the new model after fine-tuning is better than the old model.
[0081] S220, monitoring the convergence state of the basic model after fine-tuning; S221. When the convergence state of the basic model after fine-tuning is better than the current convergence state, the updated model is deployed as the basic model of the target artificial intelligence.
[0082] Steps S220 to S221 in this embodiment are similar to steps S108 to S109 in the aforementioned embodiment, and will not be described in detail here.
[0083] This embodiment covers the entire process from batch training, model optimization, integration with generative artificial intelligence, target AI deployment, dynamic monitoring to automatic fine-tuning. These steps ensure the adaptability and long-term optimization capabilities of the AI model and realize the intelligent upgrade of the manufacturing behavior analysis system.
[0084] The above is a detailed description of the model optimization and updating method of generative artificial intelligence in the embodiment of the present application. The following is a detailed description of the model optimization and updating device of generative artificial intelligence.
[0085] See also Figure 3 The present application provides an embodiment of a model optimization and updating device for generative artificial intelligence, which includes: The first acquisition unit 301 is used to acquire engineering data to be trained; An extraction unit 302 is used to extract initial sample data from the engineering data according to the behavior to be analyzed, wherein the behavior to be analyzed is predictive maintenance, quality control or supply chain optimization; A data cleaning unit 303 is used to clean the initial sample data set to obtain a target sample data set; A second acquisition unit 304 is used to determine model parameters according to the behavior to be analyzed, and acquire target model training data from the target sample data set according to the model parameters; A model training unit 305 is used to train the model corresponding to the behavior to be analyzed through the target model training data using a batch training method until all sample data in the target sample data set are trained and output a basic model; A combining unit 306, used to combine the basic model with the generative artificial intelligence to obtain a target artificial intelligence; A first monitoring unit 307, used to monitor and analyze the dynamics of the sample through the target artificial intelligence and fine-tune the basic model; A second monitoring unit 308, used to monitor the convergence state of the basic model after fine-tuning; The deployment unit 309 is used to deploy the updated model as the basic model of the target artificial intelligence when the convergence state of the basic model after fine-tuning is better than the current convergence state.
[0086] In this embodiment, the data cleaning unit 303 is specifically used for: Detecting and cleaning missing values of the initial sample data according to preset rules; Detecting outliers in the initial sample data, and removing or replacing the outliers according to the needs of the behavior to be analyzed; Establish an association key between the production line ID and the timestamp, use the association key as the index to delete duplicate records, and complete data cleaning; Data standardization is performed on the initial sample data that has completed data cleaning to obtain target sample data.
[0087] In this embodiment, the data cleaning unit 303 is further configured to: Calculate the mean and standard deviation of each data feature in the initial sample data by using the Z-score standardization method; The data features corresponding to the mean and the standard deviation are standardized, and the standardization process is calculated by the following formula: ; Wherein, x is the current data point of the data feature, μ is the mean of the data feature corresponding to the data point, and σ is the standard deviation of the data feature corresponding to the data point.
[0088] Obtain all data whose mean is 0 and whose standard deviation is 1 to obtain target sample data.
[0089] In this embodiment, the second acquiring unit 304 is specifically used for: Determine a model framework through the behavior to be analyzed, and determine model parameters according to the model framework; The target sample data set is split to obtain target model training data, wherein the target sample data set is split into target model training data, target model verification data and target model test data.
[0090] In this embodiment, the model training unit 305 is specifically used for: Splitting the target model training data into a number of small batches of data according to a preset split density; Inputting the small batch data into the model corresponding to the behavior to be analyzed one by one, so that the model corresponding to the behavior to be analyzed is iterated, and training the model corresponding to the behavior to be analyzed multiple times to obtain multiple training results; Each time a training result is obtained, the model performance of the training result is evaluated by using the target model verification data; When the model performance reaches the preset performance index and satisfies the early stopping mechanism, the training is stopped and the basic model is output.
[0091] In this embodiment, the combining unit 306 is specifically used for: Linking the input of the base model to the input of the initial generative artificial intelligence; Structured prompt words are generated according to the output format of the basic model, and boundary restrictions are imposed on the feedback of the initial generative artificial intelligence according to the structured prompt words to obtain the target artificial intelligence.
[0092] In this embodiment, the first monitoring unit 307 is specifically used for: Deploy the target artificial intelligence so that the target artificial intelligence dynamically monitors the database of the data provider of the engineering data to be trained; Setting a trigger condition for the target artificial intelligence to fine-tune the basic model; When the trigger condition is met, the target artificial intelligence is controlled to fine-tune the basic model according to the new sample data that dynamically corresponds to the analysis sample.
[0093] In this embodiment, the functions of each unit are the same as those described above. Figure 1 , Figure 2a and Figure 2b The steps in the illustrated embodiments correspond to each other and will not be described again here.
[0094] See also Figure 4 The present application embodiment provides another embodiment of a model optimization and updating device for generative artificial intelligence, including: Processor 401, memory 402, input and output unit 403, bus 404; The processor 401 is connected to the memory 402, the input and output unit 403 and the bus 404; Processor 401 specifically executes Figure 1 , Figure 2a and Figure 2b The operations corresponding to the steps in the method are not described in detail here.
[0095] The present application also relates to a computer-readable storage medium, on which a program is stored. When the program is run on a computer, the computer executes any of the above methods.
[0096] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0097] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0098] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0099] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0100] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk and other media that can store program code.
Claims
1. A generative artificial intelligence model optimization and updating method, characterized in that: The method comprises: Get the engineering data to be trained; Extracting initial sample data from engineering data according to the behavior to be analyzed, wherein the behavior to be analyzed is a predictive maintenance behavior, a quality control behavior, or a supply chain optimization behavior; Detecting and cleaning missing values of the initial sample data according to preset rules; Detecting outliers in the initial sample data, and removing or replacing the outliers according to the needs of the behavior to be analyzed; Establish an association key between the production line ID and the timestamp, use the association key as the index to delete duplicate records, and complete data cleaning; Performing data standardization on the initial sample data after data cleaning to obtain target sample data; Determining model parameters according to the behavior to be analyzed, and acquiring target model training data from the target sample data set according to the model parameters; The model corresponding to the behavior to be analyzed is trained by the target model training data using a batch training method until all sample data in the target sample data set are trained, and a basic model is output, the predictive maintenance behavior uses a memory network model for model training, the quality control behavior uses an image convolution model for model training, and the supply chain optimization behavior selects a regression model or a time series model for model training; Combining the basic model with generative artificial intelligence to obtain target artificial intelligence; Monitor and analyze the dynamics of samples through the target artificial intelligence and fine-tune the basic model; Monitoring the convergence status of the base model after fine-tuning; When the convergence state of the basic model after fine-tuning is better than the current convergence state, the updated model is deployed as the basic model of the target artificial intelligence.
2. The model optimization and updating method according to claim 1, characterized in that: The performing data standardization on the initial sample data after data cleaning to obtain target sample data includes: Calculate the mean and standard deviation of each data feature in the initial sample data by using the Z-score standardization method; The data features corresponding to the mean and the standard deviation are standardized, and the standardization process is calculated by the following formula: ; Wherein, x is the current data point of the data feature, μ is the mean of the data feature corresponding to the data point, and σ is the standard deviation of the data feature corresponding to the data point; Obtain all data whose mean is 0 and whose standard deviation is 1 to obtain target sample data.
3. The model optimization and updating method according to claim 1, characterized in that: The target sample data set is divided into target model training data, target model verification data and target model test data.
4. The model optimization and updating method according to claim 3, characterized in that: The method of using a batch training method to train the model corresponding to the behavior to be analyzed by using the target model training data until all sample data in the target sample data set have completed training and outputting a basic model includes: Splitting the target model training data into a number of small batches of data according to a preset split density; Inputting the small batch data into the model corresponding to the behavior to be analyzed one by one, so that the model corresponding to the behavior to be analyzed is iterated, and training the model corresponding to the behavior to be analyzed multiple times to obtain multiple training results; Each time a training result is obtained, the model performance of the training result is evaluated by using the target model verification data; When the model performance reaches the preset performance index and satisfies the early stopping mechanism, the training is stopped and the basic model is output.
5. The model optimization and updating method according to any one of claims 1 to 4, characterized in that: The step of combining the basic model with generative artificial intelligence to obtain target artificial intelligence includes: Linking the input of the base model to the input of the initial generative artificial intelligence; Structured prompt words are generated according to the output format of the basic model, and boundary restrictions are imposed on the feedback of the initial generative artificial intelligence according to the structured prompt words to obtain the target artificial intelligence.
6. The model optimization and updating method according to any one of claims 1 to 4, characterized in that: The target artificial intelligence monitors and analyzes the dynamics of the sample and fine-tunes the basic model, including: Deploy the target artificial intelligence so that the target artificial intelligence dynamically monitors the database of the data provider of the engineering data to be trained; Setting a trigger condition for the target artificial intelligence to fine-tune the basic model; When the trigger condition is met, the target artificial intelligence is controlled to fine-tune the basic model according to the new sample data that dynamically corresponds to the analysis sample.
7. The model optimization and updating method according to any one of claims 1 to 4, characterized in that: The extracting of initial sample data from the engineering data according to the behavior to be analyzed includes: Respectively obtaining actual parameters required for predictive maintenance behavior, quality control behavior or supply chain optimization behavior in the behavior to be analyzed; The actual parameter data of the behavior to be analyzed are acquired one by one from the engineering data according to the actual parameters to obtain initial sample data.
8. A model optimization and updating device for generative artificial intelligence, characterized in that: The device comprises: A first acquisition unit, used to acquire engineering data to be trained; An extraction unit, configured to extract initial sample data from the engineering data according to the behavior to be analyzed, wherein the behavior to be analyzed is predictive maintenance, quality control, or supply chain optimization; A data cleaning unit, used to clean the initial sample data set to obtain a target sample data set; A second acquisition unit, used to determine model parameters according to the behavior to be analyzed, and acquire target model training data from the target sample data set according to the model parameters; A model training unit, used to train the model corresponding to the behavior to be analyzed through the target model training data using a batch training method until all sample data in the target sample data set have been trained, and output a basic model; A combining unit, used to combine the basic model with the generative artificial intelligence to obtain the target artificial intelligence; A first monitoring unit, used to monitor and analyze the dynamics of the sample through the target artificial intelligence and fine-tune the basic model; A second monitoring unit, used to monitor the convergence state of the basic model after fine-tuning; A deployment unit is used to deploy the updated model as the basic model of the target artificial intelligence when the convergence state of the basic model after fine-tuning is better than the current convergence state.
9. A model optimization and updating device for generative artificial intelligence, characterized in that: The device comprises: Processor, memory, input-output unit, and bus; The processor is connected to the memory, the input and output unit, and the bus; The memory stores a program, and the processor calls the program to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a program stored thereon, wherein the program, when executed on a computer, performs the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Social media negative emotion recognition method based on generative artificial intelligence
CN117493973A
Detection model training method and device, computer equipment and storage medium
CN118505230A
Power prediction method, system and equipment based on group algorithm
CN118643941A
Industrial equipment fault monitoring model construction method and device, equipment and medium
CN119622497A
Distributed industrial energy operation optimization platform automatically constructing intelligent models and algorithms
US11487273B1