Hyperparameter optimization method and system of icing galloping prediction model based on meta learning

By optimizing hyperparameters through a meta-learning-based dual-loop training method, and combining multimodal time-series data and feature engineering, an icing galloping prediction model is constructed. This solves the problems of complex hyperparameter tuning and poor real-time performance in existing technologies, achieving efficient icing galloping prediction and reducing power grid risks.

CN120893529BActive Publication Date: 2025-12-12STATE GRID JIANGSU ELECTRIC POWER CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511405174.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-12-12
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

Existing technologies for predicting icing galloping suffer from complex hyperparameter tuning processes, high computational costs, and poor real-time performance. They are also difficult to adapt quickly to different meteorological scenarios and line operating conditions, resulting in insufficient prediction accuracy and failing to meet the real-time safety monitoring needs of the power grid.

Method used

An ice-covered dancing prediction model based on meta-learning is adopted. Hyperparameters are optimized through an inner and outer double-loop training method. Combined with multimodal time series data and feature engineering, an ice-covered dancing time series prediction model is constructed. Adaptive hyperparameter search is performed using a large Transformer model to achieve fast and adaptive model optimization.

Benefits of technology

It significantly improves the accuracy and efficiency of icing and galloping prediction, enabling rapid response under different meteorological scenarios and line operating conditions, reducing the risk of power grid icing disasters, and meeting the needs of real-time safety monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120893529B_ABST
    Figure CN120893529B_ABST
Patent Text Reader

Abstract

The application discloses a meta-learning-based super parameter optimization method and system for an icing galloping prediction model, collects original data from meteorological monitoring equipment, line state sensors and historical icing galloping records; aligns the collected original data according to time granularity to obtain multi-modal time series data; and performs inner-outer double-cycle training on the super parameters of an icing galloping time series prediction model constructed by using the obtained multi-modal time series data, wherein the inner cycle of the inner-outer double-cycle training performs iterative training under given super parameters, and after the model parameters converge, the outer cycle performs meta-learning-level super parameter optimization based on the model parameters that have converged, and under the premise of fixing the inner cycle training paradigm, the super parameters are adaptively searched to obtain optimal super parameters. The method can quickly and adaptively search for key super parameters of a model under different meteorological scenes and line working conditions, improve prediction accuracy and convergence speed, and realize real-time warning of icing galloping risk of a power transmission line.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power safety monitoring technology. It relates to a hyperparameter optimization method and system for an ice galloping prediction model based on meta-learning. Background Technology

[0002] Ice-induced galloping of conductors and ground wires on high-voltage transmission lines poses a significant threat to the safe operation of the power grid. In areas with low temperatures, high humidity, or high altitudes, ice accumulation on conductors can trigger low-frequency, high-amplitude vibrations due to wind loads. The ice layer significantly increases the conductor load and alters aerodynamic characteristics, leading to a chain reaction of accidents such as conductor breakage, phase-to-phase flashover, and tower structural failure, posing a severe challenge to power grid stability and socio-economic security.

[0003] In existing technologies, icing galloping prediction mainly relies on traditional rule-based methods and deep learning methods based on time-series models. Traditional methods have significant shortcomings in reflecting dynamic changes, real-time response, and robustness. Based on limited historical data and empirical parameters, they struggle to effectively capture the complex nonlinear coupling relationships between meteorological conditions, ice accumulation, and structural response. Existing time-series model methods involve a large number of interconnected hyperparameters during training (such as network depth, number of hidden units, learning rate, window length, regularization coefficient, and number of attention heads). These hyperparameters exhibit non-convex coupling relationships, making the model highly sensitive to them. The parameter tuning process often requires large-scale grid / Bayesian search, resulting in high computational costs, long cycles, and heavy reliance on experience and manual trial and error, making rapid implementation in engineering fields difficult. Overcoming the shortcomings of existing time-series model deep learning methods—such as high dependence on hyperparameter configuration, high tuning costs, poor real-time performance, and susceptibility to overfitting—to meet the real-time safety monitoring needs of the power grid has become an urgent problem to solve. Summary of the Invention

[0004] The purpose of this invention is to provide a hyperparameter optimization method and system for an ice galloping prediction model based on meta-learning. This method can quickly and adaptively optimize key hyperparameters of the model under different meteorological scenarios and line operating conditions, thereby improving prediction accuracy and convergence speed, and providing real-time warnings of ice galloping risks on transmission lines.

[0005] The technical solution to achieve the purpose of this invention is as follows:

[0006] A hyperparameter optimization method for an ice-covering dancing prediction model based on meta-learning includes the following steps:

[0007] Collect raw data from meteorological monitoring equipment, line status sensors, and historical ice galloping records;

[0008] The collected raw data is uniformly aligned according to the time granularity to obtain multimodal time series data;

[0009] The obtained multi-modal time sequence data is used for inner-outer double cycle training of the hyperparameters of the constructed icing galloping time prediction model, the inner cycle of the inner-outer double cycle training is iterative training under a given hyperparameter, after the model parameters converge, the outer cycle optimizes the hyperparameters at the meta-learning level based on the converged model parameters, the hyperparameters are adaptively searched under the premise of fixing the inner cycle training paradigm, and the optimal hyperparameters are obtained.

[0010] In the preferred technical solution, the weather conditions collected by the weather monitoring device include air temperature, relative humidity, wind speed, wind direction, air pressure, precipitation, visibility, and air quality index.

[0011] The line parameters collected by the line state sensor include conductor type, cross section, number of branches, ground line / cable type, hanging point height, span, direction, and design tension.

[0012] The icing galloping state of the historical icing galloping record includes icing thickness, galloping amplitude, galloping frequency, and occurrence mark.

[0013] In the preferred technical solution, after collecting the original data, data feature engineering is further constructed, including:

[0014] Correlation screening: calculate the Pearson correlation coefficient p between each feature and the icing galloping state, and put the features with |p|≥δ into the candidate pool, and δ is the Pearson threshold;

[0015] Domain prior enhancement:

[0016] Mandatory features: conductor tension, icing thickness, measured wind speed and wind direction;

[0017] Conditional features: if the air pressure change rate is greater than γ, then add the air pressure and its first-order difference;

[0018] Multi-scale standardization:

[0019] Numerical quantity unified normalization;

[0020] The wind direction is cosine processed;

[0021] The time field is periodically encoded: month, day, hour, and minute are respectively mapped to the sin / cos space;

[0022] Take the logarithm of the icing thickness; apply Sigmoid conversion to the noisy index;

[0023] High-order interaction generation:

[0024] Construct a product term for wind speed and icing thickness to capture the wind-ice coupling effect; the wind speed square term reflects the wind kinetic energy.

[0025] In the preferred technical solution, the unification and alignment according to the time granularity include:

[0026] According to the time granularity epsilon, the discrete sampled galloping observation values are matched with the continuously recorded meteorological elements according to the time stamp, and the corresponding meteorological sequence is associated according to the region where the line is located; if the current time period is determined as the galloping occurrence period, the observed galloping amplitude, galloping frequency and original value of ice thickness are kept; if it is not a galloping period, the above three items are set to zero. At the same time, the line static attribute is added at each time step, and the line static attribute is the line parameter;

[0027] The multi-modal time series data is fused to generate a multi-modal time series data tensor with a size of TxD; wherein T is the time step, and the feature dimension D includes 22 fields: region, time stamp, weather condition, air temperature, precipitation, wind direction, wind force level, wind speed, air pressure, humidity, air quality, visibility, line name, voltage level, number of branches, conductor model, hanging point height, span, conductor orientation, ice thickness, galloping frequency and galloping amplitude.

[0028] In the preferred technical solution, the inner loop iteratively trains under given hyperparameters, including:

[0029] The integrated time series is divided into non-overlapping patches according to the patch length, the patch division is linearly mapped to the model dimension d through the residual block, and a binary mask is attached, the mask value 0 indicates that it participates in attention calculation, and 1 indicates that it is ignored to tolerate missing data; then the sine-cosine position vector is superimposed for explicit injection of time sequence; the random mask r is stepped during the training period, so that the model adapts to different context lengths;

[0030] The icing galloping time series prediction model is composed of multi-head causal self-attention and feedforward network in each layer, the self-attention stage relies on the key / value / query vector dimension d / h for parallel calculation, h is the number of heads, and the feedforward stage uses a double-layer linear-GELU-linear network with hidden dimension 4d to independently process each time position; residual connection and layer normalization are applied after all sub-layers to ensure the stability of gradient propagation; the regularization and weight decay coefficient lambda are used to suppress overfitting;

[0031] The network output end is connected with the classification task head and the regression task head at the same time, the classification task head increases attention pooling after activation at the last time step, and then outputs the galloping probability through full connection+Sigmoid function; the regression task head linearly maps all time step features to output the galloping amplitude; the variable-length patch prediction mechanism is used to allow the execution of input patch and output patch equal-length inference in one forward propagation;

[0032] The inner loop adopts a double-task joint loss L=αBCE( ,y)+βMSE(â,a), wherein alpha is the classification loss weight, beta is the regression loss weight, provided by the outer loop, BCE is the binary cross-entropy loss, For output probability, y is the label, MSE is the mean square error loss, â is the predicted value, a is the true value; the optimizer is AdamW optimizer; the linear learning rate is used in the training process, and the cosine annealing scheduling is used after t steps, and the best checkpoint is recorded after evaluating on the validation set at the end of each epoch;

[0033] When the total loss of the validation set decreases by less than a set threshold ζ for k consecutive epochs, or reaches the maximum iteration, the inner loop is determined to be converged; at this time, the obtained weight θ is frozen, and the average value of the recent three rounds of validation indicators is returned as the meta loss L_meta.

[0034] In the preferred technical solution, the hyperparameters include continuous and discrete variables, and the continuous variables include: air pressure change rate threshold, Pearson threshold δ, monitoring feature upper and lower limit η / μ, classification loss weight α, regression loss weight β, base learning rate ω, weight decay ψ, and warm-up step number t.

[0035] The discrete variables include: time granularity ε, icing dance timing prediction model layer number N, attention head number h, patch length p , model dimension d, feature dimension D, step w, and Dropout parameter q.

[0036] In the preferred technical solution, the outer loop performs hyperparameter optimization at the meta learning level based on the converged model parameters, including:

[0037] Uniform or logarithmic uniform random sampling is used for continuous variables, and Latin hypercube or grid hierarchical sampling is used for discrete variables to generate K sets of hyperparameter set candidates (t);

[0038] For each candidate (t), the inner loop is performed until the model parameters θ( (t)) converge; after the inner loop ends, the joint loss is calculated on the independent validation set and the meta loss L_meta(t) is output;

[0039] For continuous hyperparameters, a differentiable hypernetwork strategy is used to directly obtain L_meta / , , which represents the partial derivative; for discrete hyperparameters, a black box optimization approximation is used, taking the meta loss as the objective function, and using reinforcement learning type policy gradient or evolutionary search to estimate the direction;

[0040] Continuous variables are updated using AdamW-Meta optimizer with meta learning rate γ ← - gamma L_ meta , denotes the gradient, the discrete variable triggers global resampling every r outer loop period, and the first B potential optimal points are selected by Bayesian optimization to join the next batch of evaluation, forming an exploration-exploitation balance;

[0041] The mean meta-loss is obtained by m-fold cross-validation for each outer loop evaluation, and is used as the true feedback signal, and the variance is recorded to monitor the search stability;

[0042] When the meta-loss L_meta decreases by less than ζ for s consecutive rounds or reaches the maximum outer loop round, early stopping is triggered; if the search is found to be trapped in a local optimum, a random exploration phase is automatically restarted once to reinitialize the candidate hyperparameters according to the meta-loss distribution variance ;

[0043] After the outer loop, the hyperparameters with the smallest meta-loss and the corresponding model weights θ are selected as the final output, and the search trajectory, hyperparameter sensitivity analysis, and validation indicators are saved together.

[0044] In the preferred technical solution, the optimal icing oscillation timing prediction model is obtained based on the obtained optimal hyperparameters, a classification-regression dual-task architecture is constructed, and the oscillation occurrence probability and amplitude prediction are synchronously output through a heterogeneous decoder.

[0045] In the preferred technical solution, it also includes edge deployment and real-time early warning, specifically including:

[0046] The complete model optimized by meta-learning is pruned, and through channel importance evaluation, channels and nodes with contribution less than a threshold τ are removed in multi-head causal self-attention, feedforward network and residual block, so that the overall precision loss after pruning does not exceed the set value;

[0047] The pruned model is implemented with mixed precision quantization: the weight parameters are reduced to FP16 as a whole, the hot spot tensors are introduced into channel-wise symmetric quantization, and quantization-aware training is used to fine-tune and compensate for quantization errors;

[0048] The quantized model is converted to ONNX intermediate representation, and according to the embedded system of the edge device and its compiler, graph fusion, operator kernel replacement and tensor reordering optimization are performed, a deployment package is generated and burned to the power line field edge server;

[0049] The edge server receives real-time data streams from the meteorological monitoring device and the line state sensor at a fixed frequency, calls the inference engine to output the dancing probability and amplitude prediction respectively; when the probability is greater than a set threshold, the dancing alarm is triggered, and the ice thickness, amplitude and corresponding meteorological context are uploaded to the dispatch center and the recent 24h abnormal sequence is cached locally;

[0050] A meteorological mutation detection unit is built in the edge, which performs sliding window statistics on weather, temperature, precipitation, wind direction, wind power, wind speed and humidity, automatically switches to high-frequency sampling when any index continuously exceeds the threshold, and synchronously sends the data to the model inference and background monitoring to realize early risk warning.

[0051] The application further discloses a hyperparameter optimization system of an icing dancing prediction model based on meta-learning, comprising:

[0052] A data acquisition and processing module acquires original data from meteorological monitoring devices, line state sensors and historical icing dancing records

[0053] A time series splicing module uniformly aligns the original data according to time granularity to obtain multi-modal time series data.

[0054] A hyperparameter optimization module performs inner-outer double-loop training on the hyperparameters of the constructed icing dancing time series prediction model using the obtained multi-modal time series data, the inner loop of the inner-outer double-loop training performs iterative training under a given hyperparameter, and after the model parameters converge, the outer loop performs hyperparameter optimization at a meta-learning level based on the already converged model parameters, and the hyperparameters are adaptively searched under the premise of fixed inner-loop training paradigm to obtain optimal hyperparameters.

[0055] The application further discloses a computer storage medium having a computer program stored thereon, and the computer executes the computer program to implement the hyperparameter optimization method of the icing dancing prediction model based on meta-learning.

[0056] Compared with the prior art, the application has the following advantages:

[0057] The application is based on the existing icing dancing time series large model prediction method of a power transmission line and the corresponding feature extraction and feature engineering method, uses the meta-learning technology, and performs deep learning search and optimization on the hyperparameters such as default missing values, model depth, sequence length and model initial state, so that the icing prediction accuracy of the power transmission line is improved.

[0058] The method combines real-time multi-source monitoring and historical records, breaks through the bottleneck of traditional single task and manual parameter tuning, fully captures the complex time sequence dependence relationship of wind-ice-line coupling, significantly improves the icing dancing prediction accuracy and alarm timeliness, and can effectively reduce the risk of power grid icing disasters. BRIEF DESCRIPTION OF DRAWINGS

[0059] Figure 1 Flow chart of the hyperparameter optimization method of the icing galloping prediction model based on meta-learning;

[0060] Figure 2 Flow chart of the inner-outer circulation training;

[0061] Figure 3 Flow chart of the edge deployment and real-time early warning. DETAILED DESCRIPTION

[0062] The principle of the present application is: based on the icing galloping prediction model of the power transmission line, the hyperparameters involved in the model are subjected to inner-outer circulation meta-learning training, the optimal hyperparameters are searched and the optimal state icing galloping model is trained based on the hyperparameters; based on the model parameters and hyperparameters obtained above, the model is deployed to realize high-precision risk early warning.

[0063] Embodiment:

[0064] As shown in Figure 1 A hyperparameter optimization method of an icing galloping prediction model based on meta-learning, comprising the following steps:

[0065] Collecting original data from meteorological monitoring equipment, line state sensors and historical icing galloping records;

[0066] Aligning the collected original data according to the time granularity to obtain multi-modal time series data;

[0067] Performing inner-outer circulation training on the hyperparameters of the icing galloping time series prediction model constructed using the obtained multi-modal time series data, wherein the inner loop of the inner-outer circulation training performs iterative training under a given hyperparameter, and after the model parameters converge, the outer loop performs hyperparameter optimization at the meta-learning level based on the converged model parameters, and under the premise of fixing the inner loop training paradigm, the hyperparameters are adaptively searched to obtain the optimal hyperparameters.

[0068] Specifically, the original data from meteorological monitoring equipment, line state sensors and historical icing galloping records are collected, data preprocessing and data feature engineering are constructed, original data features are extracted, and data quality is enhanced.

[0069] Collecting original data and performing data cleaning includes:

[0070] S11 Multi-source data collection: according to meteorological detection equipment, line physical characteristics, and original icing galloping record collection including but not limited to:

[0071] Line parameters: conductor type, cross section, number of branches, ground line / cable type, hanging point height, span, direction and design tension;

[0072] Meteorological conditions: air temperature, relative humidity, wind speed, wind direction, air pressure, precipitation, visibility, air quality index;

[0073] Icing dance state: icing thickness (mm), dance amplitude (m), dance frequency (Hz), and occurrence marker.

[0074] S12 filter obvious abnormal or error data;

[0075] Filter numerical fields according to 3σ principle or IQR (Q3+1.5IQR);

[0076] Set physical upper / lower limit for physical quantities such as wind speed and icing thickness (e.g. directly remove wind speed > η m / s or < μ m / s).

[0077] S13 fill in or interpolate missing data with default value.

[0078] Use cubic spline interpolation for continuous meteorological sequence;

[0079] For discrete dance observation, use weighted average of adjacent two time points when sampling interval > 5 min.

[0080] Fill in static line index with mode.

[0081] Data feature engineering includes:

[0082] S14 correlation screening: calculate the Pearson correlation coefficient p between each feature and icing dance state, and put the features with |p| ≥ δ into the candidate pool;

[0083] S15 domain prior enhancement:

[0084] Mandatory features: conductor tension, icing thickness, measured wind speed and wind direction;

[0085] Conditional features: if the air pressure change rate > γ Pa / h, add air pressure and its first-order difference.

[0086] S16 multi-scale standardization:

[0087] Uniformly normalize numerical quantities: x = (x min) / (max min); x is a certain numerical value in the original data, min is the minimum value, and max is the maximum value.

[0088] Use cosine processing for wind direction: dx = cos θ, dy = sin θ; dx and dy are the unit vector components of wind direction, and θ represents the angle of wind direction.

[0089] The time field adopts periodic encoding: month, day, hour, and minute are respectively mapped to the sin / cos space;

[0090] Taking the logarithm of the ice thickness to compress the long tail;

[0091] Applying a Sigmoid transformation to the noise-containing index to enhance the gradient stability.

[0092] S17 high-order interaction generation:

[0093] Constructing a product term of wind speed and ice thickness to capture the wind-ice coupling effect; and a wind speed square term to reflect the wind kinetic energy.

[0094] The steps of generating multi-modal time series data through time alignment include:

[0095] S21 unified time granularity: taking emax as the basic granularity, for each power transmission line, the discretely sampled galloping observation values and the continuously recorded meteorological elements are one-to-one matched according to the time stamp, and the corresponding meteorological sequence is associated according to the region where the line is located; if the current time period is determined as the galloping occurrence period, the observed galloping amplitude, galloping frequency and ice thickness are kept as original values; if it is not the galloping period, the above three items are set to zero. At the same time, the line static attributes that do not change with time are added at each time step, a total of 7 items: line name, voltage level, number of branches, conductor type, hanging point height, span and conductor orientation.

[0096] S22 the integrated time series tensor formed after the above fusion has a size of TxD; where T is the time step, and the feature dimension D includes 22 fields: region, time stamp, weather condition, air temperature, precipitation, wind direction, wind force level, wind speed, air pressure, humidity, air quality, visibility, and the aforementioned 7 line static attributes, plus line ice thickness, galloping frequency and galloping amplitude.

[0097] Specifically, the power transmission line icing galloping time series prediction model is illustrated by taking a Transformer-based time series large model as an example. Based on the Transformer-based time series large model, the meta-learning training is performed on the hyperparameters involved in the model (including but not limited to the initial state parameters of the Transformer deep learning model, the loss function proportion, the model layer number, the data default value, etc.). Through the step-by-step two-stage inner-outer double loop, the inner loop uses the corresponding hyperparameters to train the model, and after the model parameters converge, the outer loop performs hyperparameter training and learning based on the converged model parameters. Through this process, the optimal hyperparameters and the optimal model parameters obtained based thereon are obtained which are adapted to the current data form.

[0098] As shown in Figure 2 , the inner-outer double loop model training method comprises:

[0099] The inner loop S31 iteratively trains the parameters of the large-scale time series model based on the Transformer, aiming to fully converge the model weight θ under a given hyperparameter combination, and provide reliable gradient feedback for the external meta-learning loop. The entire process of the inner loop can be divided into six parts: input patch and context encoding, multi-layer Transformer encoding, output window, loss function, and convergence determination.

[0100] The external loop S32 is used for hyperparameter optimization at the meta-learning level. Under the premise of fixed inner loop training paradigm, the hyperparameters are adaptively searched to minimize the meta-loss L_meta on the validation set and obtain the final optimal combination.

[0101] The inner loop model training method comprises:

[0102] S311 Input patch and context encoding is used to unify the different lengths of power line multi-modal sequences. First, the comprehensive time series is divided into non-overlapping patches according to the patch length p . The patch is linearly mapped to the model dimension d through a residual block (two-layer MLP + residual connection layer) and is accompanied by a binary mask. Mask value 0 indicates that it participates in attention calculation, and 1 indicates that it is ignored to tolerate missing data. Then, a sine-cosine position vector is superimposed to explicitly inject the time order. In order to make the network see the full range of situations from very short context to w-step context during the training period, an offset r is randomly selected, and the first r steps of the sequence are prefixed with a mask to ensure stable performance of the model in different observation windows.

[0103] S312 The main body of the Transformer layer model adopts a pure decoder-only structure, with N layers. Each layer is composed of multi-head causal self-attention and feedforward network: the self-attention stage relies on key / value / query vector dimension d / h for parallel computation, and the feedforward stage uses a double-layer linear-GELU-linear network with hidden dimension 4d to process each time position independently. All sub-layers are applied with residual connection and layer normalization (LayerNorm) to ensure the stability of gradient propagation; Dropout (dropout parameter q) and weight decay coefficient λ are used to suppress overfitting.

[0104] S313 Output with long window prediction: The network output is connected to the classification task head and the regression task head at the same time. The classification head adds attention pooling after the last time step of the Transformer activation to converge the sequence overall features into a learnable vector, and then passes through a full connection + Sigmoid to output the dance occurrence probability; the regression head directly linearly maps all time step features to output the dance amplitude. By using the variable-length patch prediction mechanism, it allows the input patch to perform 32 steps and the output patch to perform 128 steps in one forward propagation, thereby significantly reducing the number of autoregressive rolls and improving the efficiency of long sequence prediction.

[0105] S314 Loss function and optimization strategy. The inner loop uses a double-task joint loss L = aBCE( ,y) + βMSE(â,a), where a is the classification loss weight, β is the regression loss weight, provided by the outer loop, used to dynamically balance the classification and regression gradients, BCE is the binary cross-entropy loss, , y is the label, MSE is the mean square error loss, â is the predicted value, and a is the true value. The optimizer uses AdamW, and the base learning rate ω and weight decay ψ also belong to the hyperparameter set. During training, linear warm-up is used, and after t steps, cosine annealing scheduling is used. At the end of each epoch, AUC and RMSE are evaluated on the validation set, and the best checkpoint is recorded.

[0106] S315 Convergence criterion and output: When the total loss of the validation set decreases by less than a set threshold ζ for k consecutive epochs, or reaches the maximum iteration E_max, the inner loop is determined to have converged. At this time, the obtained weight θ is frozen, and the average value of the latest three validation indicators is returned as the meta-loss L_meta. The optimal weight θ and L_meta will be used as the basis for updating the hyperparameters of the outer meta-learning loop (outer loop). Through the above process, the inner loop fully extracts the information of the training data under fixed hyperparameter configuration, and provides stable, differentiable and high-quality performance signals for the outer loop.

[0107] The outer loop model training method comprises the following steps:

[0108] S321 Hyperparameter vector definition: The hyperparameters are composed of continuous and discrete variables. The continuous variables include: air pressure change rate threshold, Pearson threshold δ, monitoring feature upper and lower limit η / μ, classification loss weight a, regression loss weight β, base learning rate ω, weight decay ψ, and warm-up step number t.

[0109] The discrete variables include: time granularity ε, number of Transformer layers N, number of attention heads h, patch length pModel dimension d, feature dimension D, step size w, and Dropout parameter q, etc.

[0110] S322 Initialization Strategy: For continuous variables, uniform or log-uniform random sampling is used; for discrete variables, Latin hypercube or grid-based hierarchical sampling is used to generate K candidate sets. (t). Simultaneously, search boundaries are set based on domain experience, for example, N∈[6,24], h∈{8,12,16}, p ∈{16,32,64}, α∈[0.1,3.0], ω∈[1e-5,5e-4], etc.

[0111] S323 inner loop call: for each group of candidates (t), start the inner loop of S31 until the model parameter θ( (t) converges; after the inner loop ends, the joint loss is calculated on the independent validation set and the meta-loss L_meta(t) is output.

[0112] S324 Elementary Gradient Solving: For Continuous Hyperparameters By employing a differentiable supernetwork strategy, it is possible to directly obtain [the necessary information] through backpropagation. L_meta / , Represents partial derivatives; for discrete hyperparameters A black-box optimization approximation is adopted, treating L_meta as the objective function and estimating the direction using reinforcement learning-style policy gradients or evolutionary search. To reduce the overhead of second-order gradients, the continuous part is solved by default using an approximate first-order method (Reptile / FOMAML).

[0113] S325 Meta-learning Update: Continuous variables are updated with hyperparameters using the AdamW-Meta optimizer at a meta-learning rate γ. ← γ L_meta, The gradient is represented by a discrete variable that triggers a global resampling every r outer loop cycles. The top B potential optima are selected using Bayesian optimization (GP+EI) and added to the next batch of evaluations, forming an exploration-exploitation balance.

[0114] S326 Multi-fold Cross-validation: To mitigate the randomness of the validation set, m-fold cross-validation is used for each outer loop evaluation to obtain the mean loss, which is then used as the true feedback signal; at the same time, the variance is recorded to monitor the search stability.

[0115] S327 Early stopping and restart: Early stopping is triggered when L meta decreases less than ζ in consecutive s rounds or reaches the maximum outer loop round R max. If the search is found to be trapped in a local optimum, the random search phase is automatically restarted once to reinitialize 5% candidates according to the meta-loss distribution variance .

[0116] S328 Output results: After the outer loop, the group with the minimum meta-loss and its corresponding model weight θ are selected as the final output, along with the search trajectory, hyperparameter sensitivity analysis, and validation metrics, which are saved for subsequent deployment and pruning quantization in S4.

[0117] Through the above eight steps, the outer meta-learning loop can efficiently explore the high-dimensional hyperparameter space within a controllable computing budget, significantly improving the generalization performance and robustness of the icing dance prediction model under multiple regional and climatic conditions.

[0118] In the model stage, a Transformer time series model optimized by meta-learning is used to build a classification-regression dual-task architecture, which synchronously outputs the dance occurrence probability and amplitude prediction through a heterogeneous decoder.

[0119] As shown in Figure 3 , before the model is officially put into production, the following edge deployment and real-time warning processes are included:

[0120] S41 Structural pruning is performed on the complete model optimized by meta-learning. Through channel importance evaluation, channels and nodes with a contribution less than the threshold τ are removed in multi-head causal self-attention, feedforward networks, and residual blocks, ensuring that the overall precision loss after pruning does not exceed 0.5%.

[0121] S42 Mixed precision quantization is performed on the pruned model: the weight parameters are reduced to FP16 as a whole, the hot tensor is introduced into channel-by-channel INT8 symmetric quantization, and quantization-aware training is used to fine-tune and compensate for quantization errors.

[0122] S43 The quantized model is converted to ONNX intermediate representation, and according to the embedded system of the edge device and its compiler, graph fusion, operator kernel replacement, and tensor reordering optimization are performed to generate a deployment package and burn it to the power line site edge server;

[0123] S44 The edge server receives real-time data streams from meteorological monitoring devices and line state sensors at a fixed frequency, calls the inference engine to output dance probability and amplitude prediction; when the probability is greater than the set threshold, dance alarm is triggered, and the icing thickness, amplitude, and corresponding meteorological context are uploaded to the dispatch center and cached locally for the last 24 hours of abnormal sequences.

[0124] ​S45, a meteorological mutation detection unit is built in the edge end, which performs sliding window statistics on parameters such as weather, temperature, precipitation, wind direction, wind power, wind speed and humidity, and automatically switches to high-frequency sampling when any index continuously exceeds the threshold value, and synchronously sends the data to model reasoning and background monitoring to realize early warning of risks.

[0125] In another embodiment, a computer storage medium has a computer program stored thereon, and the computer program is executed by a computer to implement the hyperparameter optimization method of the ice accretion galloping prediction model based on meta-learning.

[0126] In another embodiment, a hyperparameter optimization system of an ice accretion galloping prediction model based on meta-learning includes:

[0127] A data acquisition and processing module acquires raw data from meteorological monitoring equipment, line state sensors and historical ice accretion galloping records

[0128] A time series splicing module uniformly aligns the acquired raw data according to time granularity to obtain multi-modal time series data.

[0129] A hyperparameter optimization module performs inner-outer double-loop training on the hyperparameters of the constructed ice accretion galloping time series prediction model using the obtained multi-modal time series data. The inner loop of the inner-outer double-loop training performs iterative training under a given hyperparameter. After the model parameters converge, the outer loop performs hyperparameter optimization at the meta-learning level based on the already converged model parameters. Under the premise of fixing the inner loop training paradigm, the hyperparameters are adaptively searched to obtain the optimal hyperparameters.

[0130] In a preferred embodiment, the hyperparameter optimization system can further include the following modules:

[0131] A lightweight deployment module: an inference engine integrated with model pruning, mixed precision quantization and ONNX graph optimization, deployed on edge computing nodes in the power transmission line field to realize millisecond-level online reasoning.

[0132] A real-time warning module: continuously reads the edge node reasoning result, and generates an alarm when the galloping probability exceeds the preset threshold. At the same time, the meteorological parameter mutation is monitored, and when the continuous threshold exceeding time reaches the set time length, the sampling frequency is automatically increased, the abnormal sequence is recorded and returned to the dispatch center through the communication link for subsequent safety scheduling and operation and maintenance decision-making.

[0133] Wherein each module communicates through an RPC bus with high throughput. The system training stage is completed in a data center GPU cluster, and the reasoning stage is implemented in the edge nodes in the power transmission line field.

[0134] The working process of the hyperparameter optimization system of the icing galloping prediction model based on meta-learning is described below with a specific example as an example, including the following steps:

[0135] I. Data collection and cleaning (corresponding to step S1)

[0136] Multi-source data collection: The on-site installed meteorological monitoring equipment collects temperature, humidity, wind speed, wind direction, air pressure and precipitation every 60 seconds; the line state sensor uploads the conductor tension, galloping amplitude and galloping frequency at a frequency of 10Hz; the historical icing galloping table is summarized daily.

[0137] Outlier rejection: 3σ-principle and physical threshold double filtering are used. For example, the wind speed is limited to 0-70m / s; the icing thickness is limited to 0-70mm.

[0138] Missing data repair: cubic spline interpolation is used for continuous meteorological variables; linear weighting is used for adjacent two frames when the interval between discrete galloping observation points is greater than 5 minutes. When the static line attribute is missing, the value with the highest frequency at that voltage level is filled.

[0139] II. Feature engineering (corresponding to step S1)

[0140] Correlation screening: the Pearson coefficient is calculated for all 48 original features, and |p|≥0.25 is retained.

[0141] Domain prior enhancement: mandatory features include conductor tension, icing thickness, measured wind speed and wind direction; when the 1h change rate of air pressure is >2Pa / h, air pressure and its first-order difference are introduced.

[0142] Multi-scale standardization: numerical quantities are normalized; wind direction is converted to positive cosine dual channel; time fields (month, day, hour, minute) are periodically encoded by positive cosine; icing thickness is taken as logarithm log(1+x); sigmoid mapping is applied to high-noise indicators.

[0143] High-order interaction: wind speed x icing thickness and wind speed quadratic term are generated to capture wind-ice coupling and wind kinetic energy effect.

[0144] III. Time alignment and multi-modal time series generation (corresponding to step S2)

[0145] With 1min as the basic granularity (ε=1), the discrete galloping sequence and the continuous meteorological sequence are paired one by one. If a certain time step is determined as a galloping period (galloping label=1 or amplitude>0.05m), the original value is retained; otherwise, the galloping amplitude, galloping frequency and icing thickness are set to zero. 7 static line attributes are added at each time step, and finally a multi-modal tensor input with a dimension of T×22 is obtained. As shown on the left, Figure 2 As shown on the left, the pre-processing delay of the present application at the data level is less than 0.1s.

[0146] Four, Internal Loop: Transformer Time Series Large Model Training (Corresponding to Step S31)

[0147] Model Base: TimesFM (200M Parameters), Initial Hyperparameter Settings (Unless otherwise specified, the hyperparameter settings mentioned in this part are the initial hyperparameter settings of the first internal loop): Hidden Dimension d = 256, Number of Layers N = 12, Number of Heads h = 8.

[0148] Patch and Position Encoding: The input sequence is divided by p = 32; the patch is mapped by a double-layer MLP residual block and superimposed with a sine and cosine position vector; during training, a random mask r ∈ [0, 512] is used to adapt the model to different context lengths.

[0149] Decoder Stack: Multi-Head Causal Self-Attention + FFN (4d) + GELU, Parallel Dropout (q = 0.1); all sub-layers are followed by residual and LayerNorm.

[0150] Result Output: The classification branch is processed by attention pooling → FC → Sigmoid to obtain the dance probability; the regression branch outputs the amplitude.

[0151] Joint Loss: L = α·BCE( ,y) + β·MSE(â,a), initial α = 1.0, β = 2.0; optimizer AdamW, learning rate ω = 3 × 1 , warm-up = 500 step, cosine annealing rate 5 × 1 ; gradient clipping threshold 1.0.

[0152] Convergence Criteria: Validation Set Loss Decreases by <1 × 1 ³ or training step reaches 50k, stop; average training time is about 7h.

[0153] Five, External Loop: Meta Learning Hyperparameter Search (Corresponding to Step S32)

[0154] Search Space: Continuous {α, β, ω, ψ,...}, Discrete {N, h, p, d, q,...}.

[0155] Strategy: Use AdamW-Meta based on differentiable super network to update continuous type; Bayesian optimization + reinforcement learning to search discrete type. Every 5 rounds of external loop, the continuous type step γ is decayed by 0.5, and 10% of discrete points are resampled.

[0156] Termination Condition: When the meta loss L_meta decreases by <2 × 1 for 5 consecutive rounds or the number of external loop steps is ≥ 50, exit; output the optimal hyperparameters and weights θ .

[0157] Six, model compression and edge deployment (corresponding to steps S4 / S41-S45)

[0158] After training, first, the model is structurally pruned (channel importance threshold τ = 0.03), and the model parameter quantity is reduced by 37%, and the accuracy loss is 0.34%. Then, mixed precision quantization is performed: weight FP16 + activation INT8. Quantization-aware training (QAT) adds 1000 step fine-tuning, and the accuracy is restored by 0.12%. The final model is exported in ONNX format and compiled by the edge device.

[0159] Seven, real-time early warning process

[0160] The edge node receives the latest weather and dancing data every minute, and obtains the dancing probability and amplitude prediction after inputting the model. When ≥0.50 triggers I-level alarm and records the anomaly; if the weather mutation detection unit finds that the wind speed + rainfall comprehensive index is continuously above the threshold for 2h, 10s high-frequency sampling is started and the warning level is upgraded. Abnormal data is returned to the dispatch center through 4G / 5G for remote formulation of deicing or tension adjustment scheme by operation and maintenance personnel.

[0161] Eight, experimental effect

[0162] The experimental test results of 18 high-voltage transmission lines in Shaanxi, Hunan and Guizhou provinces in winter 2023-2024 show that the model of the present application improves the AUC index by 7.3% compared with the traditional LSTM method, reduces the amplitude RMSE by 11.5%, and the average power consumption of edge reasoning is only 12W, which meets the long-term offline power supply requirements on site.

[0163] As can be seen from the above embodiments, the present application combines meta-learning driven hyperparameter optimization and Transformer time series large model to solve the problem of multi-climate zone, cross-scene transmission line icing and dancing risk prediction, effectively improving the prediction accuracy and deployment efficiency.

[0164] The above embodiments are preferred embodiments of the present application, but the embodiments of the present application are not limited by the above embodiments, and any changes, modifications, substitutions, combinations, simplifications made without departing from the spirit and principles of the present application shall be equivalent replacement methods, and shall be included in the protection scope of the present application.

Claims

1. A method for hyperparameter optimization of an icing galloping prediction model based on meta-learning, characterized in that, The method comprises the following steps: Collecting original data from meteorological monitoring equipment, line state sensors and historical icing dance records; Aligning the collected original data according to time granularity to obtain multi-modal time series data; Performing inner-outer double cycle training on the obtained multi-modal time series data to obtain the hyperparameters of the icing dance time series prediction model, wherein the inner cycle of the inner-outer double cycle training performs iterative training under given hyperparameters, and after the model parameters converge, the outer cycle performs hyperparameter optimization at the meta-learning level based on the converged model parameters, the hyperparameters are adaptively searched under the premise of fixing the inner cycle training paradigm, and optimal hyperparameters are obtained; The inner cycle training under given hyperparameters comprises: The comprehensive time series is divided into non-overlapping patches according to the patch length, the patch division is linearly mapped to the model dimension d through a residual block, and a binary mask is attached, wherein a mask value of 0 indicates that attention calculation is performed, and a mask value of 1 indicates that it is ignored to tolerate missing data; then, a sine-cosine position vector is superimposed to explicitly inject time order; a random mask r is used in the training period to make the model adapt to different context lengths; Each layer of the icing dance time series prediction model is composed of multi-head causal self-attention and a feedforward network in series, the self-attention stage depends on the key / value / query vector dimension d / h for parallel calculation, h is the number of heads, and the feedforward stage independently processes each time position by using a double-layer linear-GELU-linear network with a hidden dimension of 4d; residual connection and layer normalization are applied after all sub-layers to ensure the stability of gradient propagation; a regularization and weight decay coefficient λ is used to suppress overfitting; The network output end is connected with a classification task head and a regression task head at the same time, the classification task head increases attention pooling after being activated at the last time step, and then outputs the dance probability through full connection + Sigmoid function; the regression task head linearly maps the features of all time steps to output the dance amplitude; a variable-length patch prediction mechanism is used to allow the execution of input patch and output patch equal-length inference in one forward propagation; The inner loop uses a dual-task joint loss L=αBCE( ,y)+βMSE(â,a), where α is the classification loss weight, β is the regression loss weight, provided by the outer loop, and BCE is the binary cross-entropy loss. The output probability is y, the label is y, the mean squared error loss is MSE, â is the predicted value, and a is the true value. The AdamW optimizer is selected. During training, a linear learning rate is used for warm-up, followed by t steps and then cosine annealing scheduling. After each epoch, the results are evaluated on the validation set, and the best checkpoint is recorded. When the total loss of the validation set decreases by less than a set threshold ζ for k consecutive epochs, or the maximum iteration is reached, it is determined that the inner cycle converges; at this time, the obtained weight θ is frozen, and the average value of the meta-loss L_meta of the latest three rounds of validation indicators is returned.

2. The method of claim 1, wherein the method is based on meta-learning. The meteorological conditions collected by the meteorological monitoring equipment include air temperature, relative humidity, wind speed, wind direction, air pressure, precipitation, visibility and air quality index; The line parameters collected by the line state sensor include conductor type, cross section, number of branches, ground line / cable type, hanging point height, span, direction and design tension; The icing dance state of the historical icing dance record includes icing thickness, dance amplitude, dance frequency and occurrence mark.

3. The method of claim 1, wherein the method is based on meta-learning. After collecting the original data, data feature engineering is constructed, including: Correlation screening: calculate the Pearson correlation coefficient p between each feature and the icing dance state, and put the features with |p|≥δ into the candidate pool, δ being the Pearson threshold; Domain prior enhancement: Mandatory features: conductor tension, icing thickness, measured wind speed and wind direction; Conditional features: if the air pressure change rate is greater than γ, add the air pressure and its first-order difference; Multi-scale standardization: Numerical quantities are normalized uniformly; Wind direction is cosine-processed; Time fields are period-encoded: month, day, hour, and minute are respectively mapped to sin / cos space; Icing thickness is logarithmized; noise-containing indicators are subjected to Sigmoid conversion; High-order interaction generation: A product term is constructed for wind speed and icing thickness to capture wind-ice coupling effects; a wind kinetic energy term is constructed for wind speed squared.

4. The hyperparameter optimization method of the meta-learning-based ice accretion shedding prediction model according to claim 1, characterized in that, Uniform alignment by time granularity includes: According to the time granularity ε, the discrete sampling of the galloping observation value and the continuous recording of the meteorological elements are matched one by one by time stamp, and the corresponding meteorological sequence is associated according to the region where the line is located; if the current time period is determined to be the galloping occurrence period, the observed galloping amplitude, galloping frequency and original icing thickness value are kept; if it is not a galloping period, it is set to zero, and the line static attribute is added at each time step, and the line static attribute is the line parameter; Fusion generates multi-modal time series data, and the multi-modal time series data tensor size is T×D; T is the time step, and the feature dimension D includes 22 fields: region, time stamp, weather condition, air temperature, precipitation, wind direction, wind force level, wind speed, air pressure, humidity, air quality, visibility, line name, voltage level, number of branches, conductor type, hanging point height, span, conductor orientation, icing thickness, galloping frequency, and galloping amplitude.

5. The hyperparameter optimization method of the meta-learning based ice accretion flutter prediction model according to claim 1, characterized in that, Hyperparameters include continuous and discrete variables, and continuous variables include: air pressure change rate threshold, Pearson threshold δ, monitoring feature upper and lower limits η / μ, classification loss weight α, regression loss weight β, base learning rate ω, weight decay ψ, and Warm-up step t; The discrete variables include: time granularity ε, icing galloping time series prediction model layer number N, attention head number h, patch length p , model dimension d, feature dimension D, step size w, and Dropout parameter q.

6. The hyperparameter optimization method of the meta-learning-based ice accretion flutter prediction model according to claim 5, characterized in that, Outer loop hyperparameter optimization based on already converged model parameters includes: For continuous variables, uniform or log-uniform random sampling is used, and for discrete variables, Latin hypercube or grid stratification sampling is used to generate K sets of hyperparameter set candidates (t), combined with domain experience to set search boundaries; For each group of candidates (t), by inner loop until the model parameters θ( (t)) converge; after the inner loop, the joint loss is calculated on the independent validation set and the meta-loss L_meta(t) is output. For continuous hyperparameters, we adopt the differentiable hypernetwork strategy, which directly obtains L_meta / , denotes the partial derivative; for discrete hyperparameters, we adopt the black-box optimization approximation, taking the meta-loss as the objective function, and estimate the direction using reinforcement learning-like policy gradient or evolutionary search; Continuous variables are updated with AdamW-Meta optimizer with meta learning rate γ ← -γ L_meta , denotes the gradient, global re-sampling is triggered every r outer loop period for discrete variables, and the first B potential optimal points are selected by Bayesian optimization to join the next batch of evaluation, forming an exploration-exploitation balance; For each outer loop evaluation, m-fold cross-validation is used to obtain the mean meta-loss, which is used as the true feedback signal, and the variance is recorded to monitor the stability of the search; Early stopping is triggered when the meta-loss L meta decreases less than ζ in consecutive s rounds or reaches the maximum outer loop epoch. If the search is found to be stuck in a local optimum, a random search phase is automatically restarted once to reinitialize the candidate hyperparameters according to the meta-loss distribution variance ; After the outer loop, the set of hyperparameters with the minimum validation loss is selected and its corresponding model weights θ as the final output, along with the search trajectory, hyperparameter sensitivity analysis, and validation metrics are saved.

7. The hyperparameter optimization method of the meta-learning-based ice accretion shedding prediction model according to claim 1, characterized in that, It also includes obtaining the best icing galloping time series prediction model based on the obtained optimal hyperparameters, constructing a classification-regression dual task architecture, and synchronously outputting the galloping occurrence probability and amplitude prediction through a heterogeneous decoder.

8. The hyperparameter optimization method of the meta-learning-based ice accretion shedding prediction model according to claim 1, characterized in that, It also includes edge deployment and real-time early warning, which specifically includes: Structural pruning is performed on the complete model optimized by meta-learning, and through channel importance evaluation, channels and nodes with contribution less than a threshold τ are removed from multi-head causal self-attention, feedforward network and residual block, so that the overall precision loss after pruning does not exceed the set value; After pruning, the model is implemented with mixed precision quantization, the weight parameters are reduced as a whole, the hot tensor is introduced into the channel-by-channel symmetric quantization, and the quantization-aware training is used to fine-tune and compensate for the quantization error; The quantized model is converted into an ONNX intermediate representation, and according to the embedded system of the edge device and its compiler, graph fusion, operator kernel replacement and tensor reordering optimization are performed, a deployment package is generated and burned to the edge server on site of the power transmission line; The edge server receives real-time data streams from the meteorological monitoring device and the line state sensor at a fixed frequency, calls the inference engine to output the dancing probability and amplitude prediction respectively; when the probability is greater than the set threshold, the dancing alarm is triggered, and the ice thickness, amplitude and corresponding meteorological context are uploaded to the dispatch center and the abnormal sequence is cached locally; A meteorological mutation detection unit is built in the edge, which performs sliding window statistics on weather, temperature, precipitation, wind direction, wind force, wind speed and humidity. When any index continuously exceeds the threshold, high-frequency sampling is switched on, and the data is sent to the model inference and background monitoring simultaneously to realize early warning of risks.

9. A hyperparameter optimization system of an icing galloping prediction model based on meta-learning, configured to perform the method of hyperparameter optimization of an icing galloping prediction model based on meta-learning according to any one of claims 1-8. It comprises: A data acquisition and processing module acquires original data from meteorological monitoring devices, line state sensors and historical ice dancing records A time series splicing module aligns the collected original data according to time granularity to obtain multi-modal time series data; An hyperparameter optimization module performs internal and external double-loop training on the hyperparameters of the ice dancing time series prediction model constructed using the obtained multi-modal time series data. The internal loop of the internal and external double-loop training performs iterative training under a given hyperparameter. After the model parameters converge, the external loop performs hyperparameter optimization at the meta-learning level based on the already converged model parameters. Under the premise of fixed internal loop training paradigm, the hyperparameters are adaptively searched to obtain the optimal hyperparameters. The internal loop includes: The comprehensive time series is divided into non-overlapping patches according to the patch length, the patch division is linearly mapped to the model dimension d through a residual block, and a binary mask is attached. Mask value 0 indicates participation in attention calculation, and 1 indicates ignoring to tolerate missing data. Then, a sine-cosine position vector is superimposed for explicit injection of time sequence. Random mask r steps are used during training to make the model adapt to different context lengths. Each layer of the ice dancing time series prediction model is composed of multi-head causal self-attention and feedforward network in series. The self-attention stage relies on key / value / query vector dimension d / h for parallel computation, and h is the number of heads. The feedforward stage uses a double-layer linear-GELU-linear network with hidden dimension 4d to independently process each time position. Residual connection and layer normalization are applied after all sub-layers to ensure the stability of gradient propagation. Regularization and weight decay coefficient λ are used to suppress overfitting. The network output end is connected to the classification task head and the regression task head. The classification task head adds attention pooling after the last time step is activated, and then outputs the dancing probability through full connection + Sigmoid function. The regression task head linearly maps all time step features to output the dancing amplitude. The variable-length patch prediction mechanism allows the execution of input patch and output patch inference with equal length in one forward propagation. The inner loop uses a dual-task joint loss L=αBCE( ,y)+βMSE(â,a), where α is the classification loss weight, β is the regression loss weight, provided by the outer loop, and BCE is the binary cross-entropy loss. The output probability is y, the label is y, the mean squared error loss is MSE, â is the predicted value, and a is the true value. The AdamW optimizer is selected. During training, a linear learning rate is used for warm-up, followed by t steps and then cosine annealing scheduling. After each epoch, the results are evaluated on the validation set, and the best checkpoint is recorded. When the total loss of the validation set decreases by less than a set threshold ζ for k consecutive epochs, or the maximum iteration is reached, the internal loop is determined to be converged. At this time, the obtained weight θ is frozen, and the average of the latest three rounds of validation indicators is returned as the meta-loss L_meta.

10. A computer storage medium having stored thereon a computer program, characterized in that The computer executes the computer program to implement the hyperparameter optimization method of the meta-learning-based ice dancing prediction model according to any one of claims 1-8. The computer executes the computer program to implement the hyperparameter optimization method of the meta-learning-based ice dancing prediction model according to any one of claims 1-8.

Citation Information

Patent Citations

  • Method for predicting windage yaw flashover risk of power transmission line based on meta-learning

    CN117114161A

  • Imbalanced sample galloping prediction method and system based on meteorological data

    CN120579680A