Large-parameter model adaptive updating method for power equipment state prediction
Through the adaptive update method, combining new data screening and historical data processing, dynamically adjusting the learning rate and building a hybrid training strategy, the problem of improper update of new and old data in power equipment status prediction is solved, improving prediction accuracy and stability, and reducing costs.
Patent Information
- Application Number
- CN202510740616.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The prior art is difficult to weigh the update of new data and historical data in the prediction of power equipment status, resulting in the model forgetting historical information or incomplete prediction of new features, affecting the prediction accuracy and cost.
Adaptive update method of large-parameter model is adopted to filter new data by calculating the feature novelty and conditional coverage of new data, combine the task correlation and contribution decay factors of historical data, formulate a dynamic learning rate update strategy, and build a hybrid training strategy for model update.
It improves the accuracy and real-time performance of power equipment status prediction, reduces the storage cost of model updates, and realizes coordinated optimization of prediction accuracy and long-term stability.
Smart Images

Figure CN120258253A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model training, and particularly to an adaptive update method for a large-parameter model for power equipment state prediction. Background Technique
[0002] During the operation of power equipment, its state is affected by various factors, such as high-voltage electric fields, heat, mechanical forces, operating conditions, meteorological environments, etc. The laws of equipment state changes and fault evolutions are contained in numerous state information such as live detection, on-line monitoring, inspection tests, as well as operating conditions, environmental climate, and power grid operation. Accurately grasping the state of power equipment and timely discovering latent faults are crucial for ensuring the safe and stable operation of the power system; with the construction and development of smart grids, the means of equipment detection have been continuously enriched, and the amount of data generated by power grid operation and equipment detection has increased exponentially. At the same time, there are a large number of abnormal data in the collected data, which increases the difficulty of power state data prediction and poses higher requirements for data processing and analysis methods. With the continuous progress of artificial intelligence and machine learning technologies, new ideas and methods have been provided for power equipment state prediction. Algorithms such as deep learning can automatically extract features from a large amount of data, establish complex non-linear models, and accurately predict the state of power equipment.
[0003] Currently, in order to improve the generality and accuracy of the model, when real-time updating a large-parameter model based on artificial intelligence, there are deficiencies in the updating process. On the one hand, excessive attention is paid to the training of new data, resulting in the model possibly forgetting historical data; on the other hand, emphasis is placed on historical data, and the training of new data is not thorough enough, making the model unable to accurately predict new features. Therefore, how to balance new data and historical data during model updating is crucial. Summary of the Invention
[0004] The purpose of the present invention is to provide an adaptive update method for a large-parameter model for power equipment state prediction to solve the problems raised in the prior art.
[0005] To achieve the above purpose, the present invention provides the following technical solutions: An adaptive update method for a large-parameter model for power equipment state prediction, the method comprising the following steps: S100. Collect equipment data during the historical operation of power equipment, respectively analyze all equipment data when the power equipment is in a normal state and an abnormal state, extract characteristic data reflecting the state of the power equipment, and use the characteristic data for machine training to obtain a state prediction model of the power equipment; Further, the specific steps of using the characteristic data for machine training to obtain a state prediction model of the power equipment are: S101. Collect the device data during the historical operation of the power equipment. The device data includes sensor data, device logs, and the environmental data where the device is located. Classify the collected device data into two situations: normal status and abnormal status according to the device logs, and generate labels. Specifically: normal status Y = 0, abnormal status Y = 1; for each continuous time-series device data, generate samples with a window length L and a step size as: [X t (i) ,Y (i) ∈(0,1)], X t (i) represents the i-th type of continuous device data based on the time series, and Y (i) represents the labels of the two device status situations of the i-th type of device data, and standardize all the sample data; S102. Use the neural network encoder to extract the device features of all the device data, calculate the statistical features of all the collected device features. The statistical features include the mean value, variance, peak value, and waveform factor; extract the frequency-domain features of the device features through Fourier transform; calculate the covariance matrix of the device features, and extract the upper triangular elements in the covariance matrix as the correlation features; combine the statistical features, frequency-domain features, and correlation features to obtain the status features of the power equipment; calculate the mutual information between each status feature and the label y, sort the mutual information of all the status features and the label y from large to small, and select the top M status features in the sorting as the training features. M represents the screening threshold, and M is set according to the model training experience; S103. Use the training features of the device data in history to construct a training set and a validation set, use the combination of the time convolutional network and the attention mechanism to construct a status prediction model, set the loss function, use the training set to train the status prediction model, and use the validation set to verify the trained status prediction model.
[0006] S200. Use the status prediction model to monitor the status of the power equipment in real time when the power equipment is operating. During the monitoring process, collect the real-time data of the power equipment operation as new data, and calculate the feature novelty and conditional coverage of the new data; Furthermore, the specific steps to combine the two types of data to obtain the value evaluation index of the new data are as follows: S201. Use the status prediction model to monitor the status of the power equipment in real time when the power equipment is operating. During the monitoring process, collect the real-time device data of the power equipment operation as new data x new , use the neural network encoder to extract the device features in the new data, and calculate the feature novelty of the new data using the device features of the new data and the device features of the historical device data. The formula is: ; In the formula, N noveltyrepresents the feature novelty of new data, m represents the number of samples of new data, T(x u new ) represents the device feature of the u-th new data sample, p hist represents the historical device feature mean, u belongs to 1 to m, represents the square of the Euclidean distance between the new data device feature and the historical device feature mean; S202. For the device features of historical device data, extract device features in different dimensions, extract the maximum and minimum values in the device features of each dimension, use the maximum and minimum values to construct the historical interval B of the device features in each dimension, and extract the device feature data values v of different dimensions in the new data new , use the historical interval to judge the mechanical energy of the device feature data value of the new data. When v new ∉B, define the output result J = 1. When v new ∈B, define the output result J = 0; use the output result J to calculate the condition coverage of the new data. The formula is: ; In the formula, C cover represents the condition coverage of the new data, D represents the total dimension of the device features, J d represents the output result of the device feature in the d-th dimension, d belongs to 1 to D.
[0007] By calculating the feature novelty and condition coverage of the new data, ensure that the abnormal features (such as early insulation deterioration, partial discharge) hidden in the new data are quickly identified, shortening the model update cycle. S300. Combine the feature novelty and condition coverage of the new data to obtain a new data value evaluation index, and formulate a new data screening mechanism to judge and screen the new data value evaluation index during the monitoring process; Furthermore, the specific steps for formulating a new data screening mechanism to judge and screen the new data value evaluation index during the monitoring process are as follows: S301. Combine the feature novelty and condition coverage of the new data to obtain a new data value evaluation index. The formula is: ; In the formula, V new represents the value evaluation index of the new data, w1 represents the weight of the feature novelty of the new data, w2 represents the weight of the condition coverage of the new data; w1 and w2 are set manually; Set the value evaluation index threshold τ, and use the value evaluation index threshold τ to construct a new data screening mechanism as: V new >τ; when the new data screening mechanism is satisfied, start model update and retain the new data; S302. The change in the accuracy of the status prediction model after real-time update. The accuracy is the difference between the predicted value and the actual true value of the status prediction model. Let the change in accuracy of the status prediction model before and after real-time update be ΔZ. When the change in accuracy is positive, the output result J' = -1; when the change in accuracy is negative, the output result J' = 1; when the change in accuracy is 0, the output result J' = 0. Use the change in accuracy to formulate a strategy for updating the threshold of the value evaluation index, specifically: ; In the formula, τ' represents the updated threshold of the value evaluation index.
[0008] Automatically tighten or relax the data access conditions according to the change in the accuracy of the validation set to avoid model degradation caused by data quality fluctuations.
[0009] S400. During the monitoring process, analyze the historical data for training the power equipment status prediction model, and calculate the task relevance and contribution attenuation factor of the historical data; Furthermore, the specific steps for calculating the task relevance and contribution attenuation factor of the historical data are as follows: S401. Collect the model parameters θ in the status prediction model of the power equipment during the training process. For each historical data sample and new data sample, find the partial derivative of the loss function with respect to the model parameters, and use the partial derivative as the loss gradient. Then, perform a one-dimensional vector transformation on the loss gradient. Let the loss gradient of the historical data sample be transformed into a one-dimensional vector g hist , and the loss gradient of the new data sample be transformed into a one-dimensional vector g new ; S402. Use the one-dimensional vectors of the loss gradients of each historical data sample and new data sample to calculate the task relevance of the historical data. The formula is: ; In the formula, R rel represents the task relevance of the historical data, n represents the total number of historical data samples, g j hist represents the one-dimensional vector of the loss gradient of the jth historical data sample, cos represents the cosine similarity, and j belongs to 1 to n; S403. Extract the time from the collection of historical data to the real-time monitoring of the power equipment as the data age of the historical data. Collect the gradient L of the validation set for the model parameters during the training of the status prediction model. Use the data age of the historical data and the validation set gradient to calculate the contribution attenuation factor of the historical data. The formula is: ; In the formula, A decayIt represents the contribution degree attenuation factor of historical data, t represents the data age of historical data, β represents the attenuation coefficient, and α represents the attenuation rate. Both the attenuation coefficient and the attenuation rate are manually set based on model training experience.
[0010] By calculating the task relevance of historical data, irrelevant historical data interference is excluded to avoid false triggering caused by stale data. By using the contribution degree attenuation factor, low-value historical data is dynamically eliminated, significantly reducing the storage cost.
[0011] S500. Obtain the value evaluation index of historical data by combining the task relevance and contribution degree attenuation factor of historical data, formulate a historical data screening mechanism, and use the historical data screening mechanism to screen historical data during the monitoring process; Furthermore, the specific steps for screening historical data using the historical data screening mechanism during the monitoring process are as follows: S501. Obtain the value evaluation index of historical data by combining the task relevance and contribution degree attenuation factor of historical data. The formula is: ; In the formula, V hist represents the value evaluation index of historical data; The historical data screening mechanism is constructed as follows: Manually set the screening threshold E, sort all the value evaluation indexes of historical data from large to small, and select the top E historical data for retention.
[0012] S600. For the state prediction model of power equipment, collect the initial learning rate of the model, use the initial learning rate to formulate a dynamic learning rate update strategy, and dynamically adjust the learning rate of the state prediction model; Furthermore, the specific steps for dynamically adjusting the learning rate of the state prediction model are as follows: S601. For the state prediction model of power equipment, collect the initial learning rate of the model, use the initial learning rate to formulate a dynamic learning rate update strategy. The formula is: ; In the formula, Lr s represents the updated learning rate, Lt0 represents the initial learning rate, r represents the number of training times, and e -0.1r represents exponential decay during the r - time training process over time.
[0013] S700. Based on the adjusted learning rate, combine the screened new data and historical data to construct a hybrid training strategy, and use the hybrid training strategy to update the state prediction model.
[0014] Furthermore, the specific steps for updating the state prediction model using the hybrid training strategy are as follows: S701. Obtain the filtered new data set F by using the new data screening mechanism and the historical data screening mechanism new , and the filtered historical data set is F hist ; When the model update program is started after judging the new data, the learning rate of the model is updated in real time by using the dynamic learning rate update strategy. The state prediction model mixes the filtered new data set and the historical data set according to the updated learning rate to formulate a mixed training strategy. The formula is: ; In the formula, LF z represents the total loss function, LF(f) represents the loss function of a single new data sample, LF(h) represents the loss function of a single historical data sample, and b represents the dynamic weight coefficient; Calculate the average loss of the new data set, which directly reflects the fitting ability of the model to the new data; Through the mixed training strategy, the model is forced to give priority to learning the device state changes reflected in the new data. If the new data contains precursors of faults, the high-loss term drives the model to quickly adjust the parameters to capture the abnormal patterns.
[0015] S702. Use the value evaluation index of the new data to adjust the dynamic weight coefficient in real time. The formula is: ; In the formula, b represents the dynamic weight coefficient; Use the mixed training strategy to update the state prediction model in real time. Fix the term 11 to ensure that historical data is always involved in training to prevent complete forgetting. log(1 + V new ) increases with the increase of V new , but the growth rate slows down; In the mixed training strategy, the new data contains significant new features, increasing b forces the model to strengthen the learning of relevant historical patterns while adapting to the new features. Through historical data regularization, the over-sensitivity of the model to new data is suppressed. If the new data is related to some historical data patterns, the mixed training enhances the model's ability to model continuous state changes.
[0016] Compared with the prior art, the beneficial effects of the present invention are: 1. The present invention directly quantifies the contribution direction of historical data to the current parameter optimization by calculating the gradient cosine similarity, and improves the accuracy of correlation judgment in complex working conditions.
[0017] 2. The present invention combines the physical aging law of the device with the data-driven value by calculating the contribution degree decay factor, and solves the defect of insufficient modeling of the device life cycle by pure data-driven methods.
[0018] 3. The present invention solves the industry pain points of low data utilization rate, serious model forgetting, and high update cost in the state prediction of power equipment, and realizes the collaborative optimization of prediction accuracy, real-time performance, and long-term stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic diagram of the steps of a large-parameter model adaptive update method for power equipment state prediction according to the present invention; Figure 2 It is a schematic flowchart of a large-parameter model adaptive update method for power equipment state prediction according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0021] Embodiment: As Figure 1 - Figure 2 shown, the present invention provides a technical solution, A large-parameter model adaptive update method for power equipment state prediction, the method includes the following steps: S100. Collect the equipment data during the historical operation of the power equipment, analyze all the equipment data when the power equipment is in normal state and abnormal state respectively, extract the characteristic data reflecting the state of the power equipment, and use the characteristic data for machine training to obtain the state prediction model of the power equipment; The specific steps of using the characteristic data for machine training to obtain the state prediction model of the power equipment are: S101. Collect the equipment data during the historical operation of the power equipment. The equipment data includes sensor data, equipment logs, and equipment environment data. According to the equipment logs, the collected equipment data is divided into two situations: normal state and abnormal state, and labels are generated, specifically: normal state Y = 0, abnormal state Y = 1; for each continuous time-series equipment data, samples are generated with a window length L and a step size as: [X t (i) , Y (i) ∈ (0, 1)], X t (i) represents the i-th continuous equipment data based on the time series, Y (i) represents the label of the two equipment state situations of the i-th equipment data, and all the sample data is standardized; S102. Use a neural network encoder to extract the device features of all device data, calculate the statistical features of all the collected device features, where the statistical features include the mean, variance, peak value, and waveform factor; extract the frequency-domain features of the device features through Fourier transform; calculate the covariance matrix of the device features, and extract the upper triangular elements in the covariance matrix as the correlation features; combine the statistical features, frequency-domain features, and correlation features to obtain the state features of the power device; calculate the mutual information between each state feature and the label y, sort the mutual information of all state features and the label y from largest to smallest, and select the top M state features in the sorting as the training features, where M represents the screening threshold and M is set according to model training experience; S103. Use the training features of the device data in history to construct a training set and a validation set, use a combination of a temporal convolutional network and an attention mechanism to construct a state prediction model, set a loss function, use the training set to train the state prediction model, and use the validation set to validate the trained state prediction model.
[0022] S200. During the operation of the power device, use the state prediction model to monitor the state of the power device in real time. During the monitoring process, collect the real-time data of the power device operation as new data, and calculate the feature novelty and conditional coverage of the new data; The specific steps for combining the two types of data to obtain the value evaluation index of the new data are as follows: S201. During the operation of the power device, use the state prediction model to monitor the state of the power device in real time. During the monitoring process, collect the real-time device data of the power device operation as new data x new , use a neural network encoder to extract the device features in the new data, and calculate the feature novelty of the new data using the device features of the new data and the device features of the historical device data. The formula is: ; In the formula, N novelty represents the feature novelty of the new data, m represents the number of samples of the new data, T(x u new ) represents the device features of the u-th new data sample, p hist represents the mean of the historical device features, u belongs to 1 to m, represents the square of the Euclidean distance between the device features of the new data and the mean of the historical device features; S202. For the device features of the historical device data, extract the device features of different dimensions, extract the maximum value and the minimum value in the device features of each dimension, use the maximum value and the minimum value to construct the historical interval B of the device features in each dimension, extract the device feature data values v of different dimensions in the new data new , use the historical interval to judge the mechanical energy of the device feature data value of the new data. When v newWhen it is not in B, the output result J is defined as 1. When v new is in B, the output result J is defined as 0; The condition coverage of the new data is calculated using the output result J. The formula is: ; In the formula, C cover represents the condition coverage of the new data, D represents the total dimension of the device characteristics, and J d represents the output result of the device characteristics in the d-th dimension, where d ranges from 1 to D.
[0023] By calculating the feature novelty and condition coverage of the new data, it is ensured that the abnormal features (such as early insulation degradation and partial discharge) hidden in the new data are quickly identified, shortening the model update cycle. S300. Combine the feature novelty and condition coverage of the new data to obtain a new data value evaluation index, and formulate a new data screening mechanism to judge and screen the new data value evaluation index during the monitoring process; The specific steps for formulating a new data screening mechanism to judge and screen the new data value evaluation index during the monitoring process are as follows: S301. Combine the feature novelty and condition coverage of the new data to obtain a new data value evaluation index. The formula is: ; In the formula, V new represents the new data value evaluation index, w1 represents the weight of the new data feature novelty, and w2 represents the weight of the new data condition coverage; w1 and w2 are set manually; Set the value evaluation index threshold τ, and use the value evaluation index threshold τ to construct a new data screening mechanism as: V new > τ; When the new data screening mechanism is satisfied, the model update is started, and the new data is retained; S302. Collect the change amount of the accuracy of the state prediction model after real-time update. The accuracy is the difference between the predicted value and the actual true value of the state prediction model; Let the change amount of the accuracy of the state prediction model before and after real-time update be △Z; When the change amount of the accuracy is positive, the output result J' = -1, when the change amount of the accuracy is negative, the output result J' = 1, and when the change amount of the accuracy is 0, the output result J' = 0; Use the change amount of the accuracy to formulate a value evaluation index threshold update strategy, specifically: ; In the formula, τ' represents the updated value evaluation index threshold.
[0024] Automatically tighten or relax the data access conditions according to the change in the accuracy of the validation set to avoid model degradation caused by data quality fluctuations.
[0025] S400. During the monitoring process, analyze the historical data for training the power equipment status prediction model, and calculate the task relevance and contribution attenuation factor of the historical data; The specific steps for calculating the task relevance and contribution attenuation factor of the historical data are as follows: S401. Collect the model parameters θ in the status prediction model of the power equipment during the training process. For each historical data sample and new data sample, find the partial derivative of the loss function with respect to the model parameters, and use the partial derivative as the loss gradient. Then, perform a one-dimensional vector transformation on the loss gradient. Let the loss gradient of the historical data sample be transformed into a one-dimensional vector g hist , and the loss gradient of the new data sample be transformed into a one-dimensional vector g new ; S402. Calculate the task relevance of the historical data using the one-dimensional vectors of the loss gradients of each historical data sample and new data sample. The formula is: ; In the formula, R rel represents the task relevance of the historical data, n represents the total number of historical data samples, g j hist represents the one-dimensional vector of the loss gradient of the j-th historical data sample, cos represents the cosine similarity, and j belongs to 1 to n; S403. Extract the time from the collection of the historical data to the real-time monitoring of the power equipment as the data age of the historical data. Collect the gradient L of the validation set for the model parameters during the training of the status prediction model. Calculate the contribution attenuation factor of the historical data using the data age of the historical data and the validation set gradient. The formula is: ; In the formula, A decay represents the contribution attenuation factor of the historical data, t represents the data age of the historical data, β represents the attenuation coefficient, α represents the attenuation rate, and both the attenuation coefficient and the attenuation rate are manually set based on the model training experience.
[0026] By calculating the task relevance of the historical data, eliminate the interference of irrelevant historical data, and avoid false triggering caused by stale data. Dynamically eliminate low-value historical data through the contribution attenuation factor, significantly reducing the storage cost.
[0027] S500. Combine the task relevance and contribution attenuation factor of the historical data to obtain the value evaluation index of the historical data, formulate a historical data screening mechanism, and use the historical data screening mechanism to screen the historical data during the monitoring process; The specific steps for using the historical data screening mechanism to screen the historical data during the monitoring process are as follows: S501. Obtain the value evaluation index of historical data by combining the task relevance of historical data and the contribution degree attenuation factor. The formula is: ; In the formula, V hist represents the value evaluation index of historical data. The historical data screening mechanism is constructed as follows: Manually set the screening threshold E, sort all the value evaluation indexes of historical data from large to small, and select the top E historical data for retention.
[0028] S600. For the state prediction model of power equipment, collect the initial learning rate of the model, formulate a dynamic learning rate update strategy using the initial learning rate, and dynamically adjust the learning rate of the state prediction model; The specific steps for dynamically adjusting the learning rate of the state prediction model are as follows: S601. For the state prediction model of power equipment, collect the initial learning rate of the model, formulate a dynamic learning rate update strategy using the initial learning rate. The formula is: ; In the formula, Lr s represents the updated learning rate, Lt0 represents the initial learning rate, r represents the number of training times, and e -0.1r represents the exponential decay during the r - time training process over time.
[0029] S700. Based on the adjusted learning rate, combine the screened new data and historical data to construct a hybrid training strategy, and use the hybrid training strategy to update the state prediction model.
[0030] The specific steps for updating the state prediction model using the hybrid training strategy are as follows: S701. Use the new data screening mechanism and the historical data screening mechanism to obtain the screened new data set as F new , and the screened historical data set as F hist ; When the model update program is started after judging the new data, use the dynamic learning rate update strategy to update the learning rate of the model in real - time. The state prediction model mixes the screened new data set and the historical data set according to the updated learning rate and formulates a hybrid training strategy. The formula is: ; In the formula, LF z represents the total loss function, LF(f) represents the loss function of a single new data sample, LF(h) represents the loss function of a single historical data sample, and b represents the dynamic weight coefficient; Calculate the average loss of the new dataset, which directly reflects the fitting ability of the model to new data; use the mixed training strategy to force the model to preferentially learn the device state changes reflected in the new data. If the new data contains fault precursors, the high loss term drives the model to quickly adjust the parameters to capture abnormal patterns.
[0031] S702. Adjust the dynamic weight coefficient in real time using the value evaluation index of the new data. The formula is: ; In the formula, b represents the dynamic weight coefficient; use the mixed training strategy to update the state prediction model in real time. The fixed term 11 ensures that historical data always participates in training to prevent complete forgetting. log(1 + V new ) increases with the increase of V new , but the growth rate slows down; In the mixed training strategy, if the new data contains significant new features, increase b, forcing the model to strengthen the learning of relevant historical patterns while adapting to the new features. Through historical data regularization, suppress the over-sensitivity of the model to new data. If the new data is related to the patterns of some historical data, the mixed training enhances the model's ability to model continuous state changes.
[0032] Example 1: Suppose a converter station transformer monitoring system needs to evaluate whether 3 newly collected samples trigger model update. The feature dimensions of the new data are Feature 1 and Feature 2. The historical intervals of the two feature dimensions are [0, 1.2] and [0.5, 2.0] respectively; the device feature data values in the two feature dimensions of Sample 1 are 0.6 and 1.3, the device feature data values in the two feature dimensions of Sample 2 are 0.3 and 0.8 respectively, and the device feature data values in the two feature dimensions of Sample 3 are 1.2 and 2.1 respectively; assume that the average values of the historical device feature data are 0.5 and 1.2; The calculated feature novelty of the new data is (0.02 + 0.2 + 1.3) / 3 = 0.5; The output results of the three samples of the new data based on Feature 1 dimension are 0, and the output results based on Feature 2 dimension are 1; the calculated conditional coverage rate of the new data is 2; Let w1 and w2 be 0.7 and 0.3 respectively; the calculated value evaluation index of the new data is 0.95.
[0033] Embodiment 2: Assume that in a certain transformer monitoring system, the new data is that the detected winding temperature has abnormally increased, and the historical data is the precursor data of the past 3 similar temperature rise events screened out; calculate the dynamic weight coefficient, specifically: b = 1 + log(1 + 2.5) = 1 + 1.216 = 2.216; in the total loss LF, the weight of the historical data loss term is 2.216 times higher than that of the new data term; through the hybrid training strategy, while fitting the new temperature rise pattern, strengthen the learning of historical similar events and identify common features.
[0034] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. An adaptive update method for a large-parameter model used in power equipment status prediction, characterized in that: The method includes the following steps: S100. Collect the device data during the historical operation of the power device, analyze all the device data when the power device is in normal state and abnormal state respectively, extract the characteristic data reflecting the state of the power device, and use the characteristic data for machine training to obtain the state prediction model of the power device; S200. During the operation of the power device, use the state prediction model to monitor the state of the power device in real time. During the monitoring process, collect the real-time data of the power device operation as new data, and calculate the characteristic novelty and conditional coverage of the new data; S300. Combine the characteristic novelty and conditional coverage of the new data to obtain the new data value evaluation index, and formulate a new data screening mechanism to judge and screen the new data value evaluation index during the monitoring process; S400. During the monitoring process, analyze the historical data for training the state prediction model of the power device, and calculate the task relevance and contribution attenuation factor of the historical data; S500. Combine the task relevance and contribution attenuation factor of the historical data to obtain the historical data value evaluation index, formulate a historical data screening mechanism, and use the historical data screening mechanism to screen the historical data during the monitoring process; S600. For the state prediction model of the power device, collect the initial learning rate of the model, use the initial learning rate to formulate a dynamic learning rate update strategy, and dynamically adjust the learning rate of the state prediction model; S700. Based on the adjusted learning rate, combine the screened new data and historical data to construct a hybrid training strategy, and use the hybrid training strategy to update the state prediction model.
2. The adaptive update method for a large parameter model for power equipment status prediction according to claim 1, wherein: The specific steps of using the characteristic data for machine training to obtain the state prediction model of the power device in S100 are as follows: S101. Collect device data during the historical operation of power equipment. The device data includes sensor data, device logs, and environmental data of the device. Classify the collected device data into two situations: normal status and abnormal status according to the device logs, and generate labels. Specifically: normal status Y = 0, abnormal status Y = 1; for each continuous time-series device data, generate samples with a window length L and a step size as: [X t (i) ,Y (i) ∈(0, 1)], X t (i) represents the i-th continuous device data based on the time series, and Y (i) represents the labels of the two device status situations of the i-th device data, and standardize all sample data; S102. Use the neural network encoder to extract the device characteristics of all device data, calculate the statistical characteristics of all collected device characteristics. The statistical characteristics include mean value, variance, peak value and waveform factor; extract the frequency domain characteristics of the device characteristics through Fourier transform; Calculate the covariance matrix of the device characteristics, and extract the upper triangular elements in the covariance matrix as the correlation characteristics; combine the statistical characteristics, frequency domain characteristics and correlation characteristics to obtain the state characteristics of the power device; Calculate the mutual information between each state characteristic and the label y, sort the mutual information between all state characteristics and the label y from large to small, and screen out the top M state characteristics in the sorting as the training characteristics. M represents the screening threshold, and M is set according to the model training experience; S103. Use the training characteristics of the device data in history to construct a training set and a validation set, use the combination of the temporal convolutional network and the attention mechanism to construct the state prediction model, set the loss function, use the training set to train the state prediction model, and use the validation set to verify the trained state prediction model.
3. A large-parameter model adaptive update method for power equipment status prediction according to claim 2, characterized in that: The specific steps of combining the two types of data to obtain the new data value evaluation index in S200 are as follows: S201. When the power equipment is working, use the state prediction model to monitor the state of the power equipment in real time. During the monitoring process, collect the real-time device data of the power equipment working as the new data x new , use the neural network encoder to extract the device features in the new data, and calculate the feature novelty of the new data using the device features of the new data and the device features of the historical device data. The formula is: ; In the formula, N novelty represents the feature novelty of the new data, m represents the number of samples of the new data, T(x u new ) represents the device feature of the u-th new data sample, p hist represents the historical device feature mean, u belongs to 1 to m, represents the square of the Euclidean distance between the new data device feature and the historical device feature mean; S202. For the device features of historical device data, extract device features in different dimensions, extract the maximum and minimum values from the device features in each dimension, construct the historical interval B of the device features in each dimension using the maximum and minimum values, and extract the device feature data values v of different dimensions in the new data new , use the historical interval to judge the mechanical energy of the device feature data value of the new data. When v new ∉B, define the output result J = 1. When v new ∈B, define the output result J = 0; calculate the condition coverage of the new data using the output result J. The formula is: ; In the formula, C cover represents the condition coverage of the new data, D represents the total dimension of the device features, and J d represents the output result of the device features in the d-th dimension, where d ranges from 1 to D.
4. An adaptive update method for a large-parameter model for power equipment status prediction according to claim 3, characterized in that: The specific steps of formulating a new data screening mechanism to judge and screen the new data value evaluation index during the monitoring process in S300 are as follows: S301. Combine the characteristic novelty and conditional coverage of the new data to obtain the new data value evaluation index. The formula is: ; In the formula, V new represents the value evaluation index of the new data, w1 represents the weight of the novelty of the new data features, and w2 represents the weight of the condition coverage of the new data; w1 and w2 are set manually; Set the threshold τ of the value evaluation index, and use the threshold τ of the value evaluation index to construct a new data screening mechanism as: V new > τ; when the new data screening mechanism is satisfied, start model update and retain the new data; S302. The change in the accuracy of the acquisition status prediction model after real-time update, where the accuracy is the difference between the predicted value and the actual true value of the status prediction model; let the change in accuracy before and after the real-time update of the status prediction model be ΔZ; when the change in accuracy is positive, the output result J' = -1, when the change in accuracy is negative, the output result J' = 1, and when the change in accuracy is 0, the output result J' = 0; use the change in accuracy to formulate a strategy for updating the threshold of the value evaluation index, specifically: ; In the formula, τ' represents the updated threshold of the value evaluation index.
5. An adaptive update method for a large-parameter model for power equipment status prediction according to claim 2, characterized in that: The specific steps for calculating the task relevance and contribution decay factor of historical data in S400 are as follows: S401. Collect the model parameter θ in the state prediction model of the power equipment during the training process. For each historical data sample and new data sample, calculate the partial derivative of the loss function with respect to the model parameter, and use the partial derivative as the loss gradient. Then, perform a one-dimensional vector transformation on the loss gradient. Suppose the loss gradient of the historical data sample is transformed into a one-dimensional vector as g hist , and the loss gradient of the new data sample is transformed into a one-dimensional vector as g new ; S402. Calculate the task relevance of historical data using the one-dimensional vector of the loss gradient of each historical data sample and the new data sample. The formula is: ; In the formula, R rel represents the task relevance of historical data, n represents the total number of historical data samples, and g j hist represents the one-dimensional vector of the loss gradient of the j-th historical data sample, cos represents the cosine similarity, and j belongs to 1 to n; S403. Extract the time from the collection of historical data to the real-time monitoring of power equipment as the data age of the historical data. Collect the gradient L of the validation set for the model parameters during the training of the acquisition status prediction model. Calculate the contribution decay factor of the historical data using the data age of the historical data and the validation set gradient. The formula is: ; In the formula, A decay represents the contribution degree attenuation factor of historical data, t represents the data age of historical data, β represents the attenuation coefficient, and α represents the attenuation rate. Both the attenuation coefficient and the attenuation rate are manually set according to the model training experience.
6. The adaptive update method for a large-parameter model for power equipment status prediction according to claim 5, wherein: The specific steps for screening historical data during the monitoring process using the historical data screening mechanism in S500 are as follows: S501. Obtain the value evaluation index of historical data by combining the task relevance and contribution decay factor of historical data. The formula is: ; In the formula, V hist represents the value evaluation index of historical data; The historical data screening mechanism is constructed as follows: Manually set the screening threshold E. Set the value evaluation indexes of all historical data to be sorted from large to small, and select the top E historical data for retention.
7. An adaptive update method for a large-parameter model for power equipment status prediction according to claim 2, characterized in that: The specific steps for dynamically adjusting the learning rate of the status prediction model in S600 are as follows: S601. For the status prediction model of power equipment, collect the initial learning rate of the model, and formulate a dynamic learning rate update strategy using the initial learning rate. The formula is: ; In the formula, Lr s represents the updated learning rate, Lt0 represents the initial learning rate, r represents the number of training times, and e -0.1r represents the exponential decay during the r - time training process over time.
8. An adaptive update method for a large-parameter model for power equipment status prediction according to claim 7, characterized in that: The specific steps for updating the status prediction model using the hybrid training strategy in S700 are as follows: S701. Using the new data screening mechanism and the historical data screening mechanism, the screened new data set is F new , and the screened historical data set is F hist ; when the model update program is started after judging the new data, the learning rate of the model is updated in real time using the dynamic learning rate update strategy. The state prediction model mixes the screened new data set and the historical data set according to the updated learning rate and formulates a mixed training strategy. The formula is: ; In the formula, LF z represents the total loss function, LF(f) represents the loss function of a single new data sample, LF(h) represents the loss function of a single historical data sample, and b represents the dynamic weight coefficient; S702. Use the value evaluation index of new data to adjust the dynamic weight coefficient in real time. The formula is: ; In the formula, b represents the dynamic weight coefficient; use the hybrid training strategy to update the status prediction model in real time.
Citation Information
Patent Citations
Power equipment temperature prediction method based on PSO-LSSVM online learning
CN111523710A
Electric power artificial intelligence model system and working method
WO2025081995A1