Time-frequency combined feature extraction and clustering acceleration optimization model training method

By combining time-frequency feature extraction and clustering to accelerate model training, the redundant MIP problem was solved, the accuracy and speed of MIP solving were improved, and efficient MIP solver training was achieved.

CN121145962APending Publication Date: 2025-12-16CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511696595.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing technologies, when using deep learning models for training, neglect the impact of training data quality on model training speed, resulting in a large number of redundant MIP problems consuming computing power, and failing to effectively improve the accuracy and speed of MIP solutions.

Method used

A time-frequency combined feature extraction and clustering method is adopted to accelerate and optimize model training. MIP features are extracted through time-domain autoencoder and frequency-domain encoder. Combined with attention mechanism and loss term optimization, redundant MIP problem is eliminated, and typical instances are extracted for clustering.

Benefits of technology

It effectively reduces computational redundancy, improves the accuracy and speed of MIP solutions, saves computing power, and achieves almost no loss in solution efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121145962A_ABST
    Figure CN121145962A_ABST
Patent Text Reader

Abstract

The invention relates to the field of integrated energy system optimization, and particularly discloses a time-frequency combined feature extraction and clustering acceleration optimization model training method, which comprises the following steps: S1, acquiring mixed integer programming (MIP) equation data; s2, performing data enhancement, and extracting and preprocessing the MIP data; s3, time domain features are extracted, and low-dimensional MIP time domain features are extracted from variables and constraint node matrixes of the MIP bipartite graph through a time domain auto-encoder; s4, extracting a frequency domain feature, and calculating a frequency spectrum energy entropy of each node as the frequency domain feature; s5, inputting the time domain and frequency domain features into a fusion network to obtain final time-frequency combined MIP low-dimensional features; s6, training a feature extraction model to enable the obtained feature space to better serve a downstream clustering task; and S7, clustering the fusion features, and extracting typical examples. According to the method, the characteristics of each MIP can be identified more accurately, typical instances can be extracted for subsequent training of the MIP solver model, a large number of redundant MIP problems are eliminated, and the solving accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of integrated energy system optimization technology, specifically a method for accelerating optimization model training by combining time-frequency feature extraction and clustering. Background Technology

[0002] The global energy transition towards decarbonization and sustainable development has driven the rapid deployment of renewable energy. However, traditional energy systems, due to the isolated operation of electricity, heat, and gas networks, face severe challenges in absorbing a high proportion of intermittent renewable energy sources (such as solar and wind power) and achieving cross-sectoral energy efficiency. To address the inefficiency of the isolated operation of traditional energy systems, the integrated energy system optimization problem needs to be modeled as a mixed-integer programming (MIP) problem to achieve multi-energy complementarity and synergy. However, while MIP solution methods incorporating deep learning (such as deep reinforcement learning, DRL) can improve model accuracy and dynamic response capabilities, their training and real-time optimization rely on large-scale heterogeneous data processing and high-dimensional decision space search, leading to a significant increase in upfront computational costs. Therefore, research on significantly reducing computational requirements while minimizing accuracy loss is crucial.

[0003] Currently, research on solving integrated energy dispatching problems using MIP equations mainly focuses on improving the accuracy of the solution and reducing catastrophic forgetting caused by time. However, improving accuracy and reducing forgetting rate requires training a deep learning model that mimics the solver's solution. This model is trained by randomly generating a large number of MIP problems and using the solution data from these randomly generated MIP problems to train the solver model. Many of these generated MIP problems are redundant and similar, consuming significant computational resources for solving and training the model, but without improving the model's efficiency or accuracy. The process of solving MIP problems using deep neural networks involves combining branch and bound methods to train the deep neural network to predict the variables and values ​​of the next branch. Current research focuses on accelerating the solution by improving the efficiency of predicting branch variables to reduce the size of the branch tree and by stacking computational power for parallel processing. However, both methods neglect the impact of the quality of the training data on the model training speed, thus affecting the solution speed.

[0004] Therefore, there is an urgent need for a time-frequency combined feature extraction and clustering method to accelerate the training of the optimization model, eliminate a large number of redundant MIP problems, and improve the accuracy and speed of the solution. Summary of the Invention

[0005] To address the aforementioned problems in the prior art, this invention provides a time-frequency combined feature extraction and clustering method to accelerate and optimize model training, effectively solving the problem of eliminating a large amount of redundant MIP, reducing computational redundancy, and improving the accuracy and speed of the solution.

[0006] To achieve the above objectives, this invention proposes a method for accelerating and optimizing model training by combining time-frequency feature extraction and clustering, comprising: S1. Obtain a large amount of mixed integer programming (MIP) equation data in integrated energy systems; S2. Perform data augmentation by extracting and preprocessing the MIP data, and model the MIP problem as a bipartite graph from the perspective of variables and constraints, obtaining the variable node matrix, constraint node matrix and adjacency matrix of the MIP bipartite graph. S3. Extract time-domain features: Extract low-dimensional MIP time-domain features from the variable and constraint node matrices of the MIP bipartite graph using a time-domain autoencoder. S4. Extract frequency domain features. The adjacency matrix of the MIP bipartite graph is processed by a frequency domain encoder and frequency domain techniques to obtain an approximate feature vector matrix. Then, the spectral energy entropy of each node is calculated as the frequency domain feature. S5. Input the time-domain and frequency-domain features into the fusion network, and dynamically fuse the time-domain and frequency-domain features through the attention mechanism to obtain the final time-frequency combined MIP low-dimensional features; S6. Train the feature extraction model, introduce MIP attribute loss term and clustering-oriented loss term to optimize the model, so that the network focuses on learning the domain knowledge of MIP and optimizing the feature space so that the obtained feature space can better serve the downstream clustering task. S7. Cluster the fused features, extract typical instances, input the final low-dimensional features into the clustering layer to obtain cluster centers and clusters, calculate the local density for each data point in the cluster, and select the k instances with the highest local density in each cluster as typical instances.

[0007] Preferably, in S1, the MIP data includes information on the distribution of variable types, constraint sparsity, and objective function coefficients of the integrated energy system, which is represented as follows: ; In the formula, n is the total number of variables; P is the number of integer variables (i.e., the first p variables of x are integers, and the rest are real numbers); A∈R mxn The constraint coefficient matrix: b∈R m To constrain the right-hand coefficient vector; c∈R n Let the coefficient vector of the objective function be: 1, u∈R n These represent the lower and upper bound vectors of the variable, respectively.

[0008] Preferably, in S2, the specific steps for modeling the MIP problem as a bipartite graph after obtaining a large amount of MIP equation data are as follows: S21. Extract the type and boundary features of variables from the MIP equation and construct a variable feature matrix; S22. Extract the constraint type and right-hand constant, and construct the constraint feature matrix; S23. Construct the edge indices between variables and constraints to obtain the adjacency matrix of MIP.

[0009] Preferably, in S3, the constructed temporal encoder consists of a type-aware embedding layer, a feature fusion layer, a bipartite graph convolutional layer, and a fully connected layer: the type-aware layer maps discrete type information into continuous semantic vectors; the feature fusion layer integrates the structural features and type semantics of the MIP into a single feature vector; the bipartite graph convolutional layer is used for information transfer to extract features from the bipartite graph of the MIP; and the fully connected layer outputs the temporal features as features with the same dimensions as the frequency domain features.

[0010] Preferably, in S4, the frequency domain encoder consists of three parts: a Laplacian matrix, a low-rank approximation, and a calculated spectral energy entropy. The specific steps for constructing the frequency domain encoder to extract the frequency domain features of the MIP are as follows: S41. Input the adjacency matrix A of the MIP bipartite graph, and use A to calculate the normalized bipartite graph Laplace matrix. S42. Use the Nystrom method to extract the Top-k feature vectors; S43. Calculate the spectral energy distribution of each node as a frequency domain feature.

[0011] Preferably, in S5, a time-frequency fusion module is used to fuse the time-domain and frequency-domain features of the MIP. The time-frequency feature fusion module consists of two parts: a cross-attention layer and a gated feature fusion layer. The specific steps for using the time-frequency fusion module to fuse the time-domain and frequency-domain features of the MIP are as follows: S51. In the cross-attention layer, the attention is calculated using the time-domain feature Query, frequency-domain feature Key, and Value of MIP. For each time-domain node feature, the correlation with the observed frequency-domain feature is calculated. S52. In the gated feature fusion layer, the gate controls the weight of feature fusion for different nodes. When a node is identified as a variable node, the feature components are biased towards frequency domain features, and when a node is constrained, the features are biased towards time domain features.

[0012] Preferably, in S6, the optimization model introduces MIP attribute loss term and clustering-oriented loss term, and the specific steps for learning the domain knowledge of MIP and optimizing the feature space for downstream clustering tasks when extracting the network are as follows: S61. Calculate the MIP attribute loss, which consists of three parts: variable type prediction loss, constraint sparsity alignment loss, and target coefficient KL divergence loss. S62, Reconstruction Loss: The original input is reconstructed through an autoencoder, and then the variable type cross-entropy, target coefficient binning MSE, and constraint sparsity MSE are calculated with the original input to ensure that the features retain key information. S63. Calculate the clustering-oriented loss for downstream clustering tasks. Calculate the feature similarity of the obtained features, then use the feature similarity to calculate the predicted distribution q. Use q to calculate the desired sharper feature space, i.e., the target distribution p. Calculate the KL divergence for q and p as the clustering loss to optimize the feature space. The formulas for calculating the feature similarity d and the predicted distribution q are as follows: ; In the formula, The fused time-frequency characteristics; The formula for calculating the sharpened target distribution p is: ; ; ; The expression for the clustering loss term is: .

[0013] Preferably, in S61, the specific steps for calculating the MIP attribute loss are as follows: S611. Using variable type prediction loss, the variable type is predicted by fusing the first n dimensions of the features, and then the cross-entropy loss function is calculated with the true variable type. This preserves the variable type information of the MIP in the feature space. The formula for calculating the variable type prediction loss is as follows: ; In the formula, For real variable types, These are the three types of probability distributions for prediction; S612. Using constraint sparsity alignment loss, predict the sparsity of constraints using the (n+1)th dimension of the fused features, and then calculate the L1 loss with the true constraint sparsity, so that the feature space retains the constraint sparsity information of MIP. The sparsity calculation formula is: ; In the formula, nnz(A) is the number of non-zero elements in the constraint matrix. For the number of variables, For constraint numbers; The formula for calculating the constraint sparsity alignment loss is: ; S613. Using the target coefficient KL divergence loss, calculate the probability that the coefficients of the variables in the objective function fall within each interval, so that the learned feature target coefficient distribution approximates the true coefficient distribution. The KL divergence calculation formula is as follows: ; In the formula, For the learned feature distribution, This represents the true distribution.

[0014] Preferably, in S7, the obtained low-dimensional features are clustered to extract typical instances, including the following steps: S71. Preprocess the fusion features by standardizing them, setting the mean of each dimension to 0 and the variance to 1. S72. After determining that the number of clusters is one-fiftieth of the total number of instances and completing the clustering, calculate the local density of each instance using kernel density, and select the k instances with the highest density as typical instances within the cluster for output.

[0015] A system applying the time-frequency combined feature extraction and clustering accelerated optimization model training method described above, the system includes a dataset acquisition module, a MIP bipartite graph modeling module, a feature extraction module, a feature fusion module, and a clustering extraction of typical instances module; the system also includes a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory.

[0016] Therefore, this invention proposes a time-frequency combined feature extraction and clustering method to accelerate and optimize model training, with the following beneficial effects: (1) This invention can reflect the advantages of the structural and content characteristics of MIP, more accurately identify the characteristics of each MIP, and cluster a large number of randomly generated MIPs and extract representative and characteristic MIPs for the generation of subsequent training data. (2) The present invention can extract typical instances for training subsequent MIP solver models, eliminate a large number of redundant MIP problems, save a lot of computing power and almost do not lose the accuracy and speed of the solution.

[0017] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0018] Figure 1 This is an overall flowchart of the time-frequency combined feature extraction and clustering accelerated optimization model training method of the present invention. Detailed Implementation

[0019] To make the technical solutions, advantages, and objectives of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below. The described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention without creative effort are within the protection scope of this application.

[0020] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0021] This invention provides a method for accelerating and optimizing model training by combining time-frequency feature extraction and clustering, comprising: S1. Obtain a large amount of mixed integer programming (MIP) equation data in integrated energy systems; In S1, the MIP data includes information on the distribution of variable types, constraint sparsity, and objective function coefficients of the integrated energy system, which is represented as follows: ; In the formula, n is the total number of variables; P is the number of integer variables (i.e., the first p variables of x are integers, and the rest are real numbers); A∈R mxn The constraint coefficient matrix: b∈R m To constrain the right-hand coefficient vector; c∈R n Let the coefficient vector of the objective function be: 1, u∈R n These represent the lower and upper bound vectors of the variable, respectively.

[0022] S2. Perform data augmentation by extracting and preprocessing the MIP data, and model the MIP problem as a bipartite graph from the perspective of variables and constraints, obtaining the variable node matrix, constraint node matrix and adjacency matrix of the MIP bipartite graph. In S2, the specific steps for modeling the MIP problem as a bipartite graph after obtaining a large amount of MIP equation data are as follows: S21. Extract the type and boundary features of variables from the MIP equation and construct a variable feature matrix; S22. Extract the constraint type and right-hand constant, and construct the constraint feature matrix; S23. Construct the edge indices between variables and constraints to obtain the adjacency matrix of MIP.

[0023] S3. Extract time-domain features: Extract low-dimensional MIP time-domain features from the variable and constraint node matrices of the MIP bipartite graph using a time-domain autoencoder. In S3, the constructed temporal encoder consists of a type-aware embedding layer, a feature fusion layer, a bipartite graph convolutional layer, and a fully connected layer: the type-aware layer maps discrete type information into continuous semantic vectors; the feature fusion layer integrates the structural features and type semantics of the MIP into a single feature vector; the bipartite graph convolutional layer is used for information transfer to extract features from the bipartite graph of the MIP; and the fully connected layer outputs the temporal features as features with the same dimensions as the frequency domain features.

[0024] S4. Extract frequency domain features. The adjacency matrix of the MIP bipartite graph is processed by a frequency domain encoder and frequency domain techniques to obtain an approximate feature vector matrix. Then, the spectral energy entropy of each node is calculated as the frequency domain feature. In S4, the frequency domain encoder consists of three parts: the Laplacian matrix, the low-rank approximation, and the calculated spectral energy entropy. The specific steps for constructing the frequency domain encoder to extract the frequency domain features of the MIP are as follows: S41. Input the adjacency matrix A of the MIP bipartite graph, and use A to calculate the normalized bipartite graph Laplace matrix. S42. Use the Nystrom method to extract the Top-k feature vectors; S43. Calculate the spectral energy distribution of each node as a frequency domain feature.

[0025] S5. Input the time-domain and frequency-domain features into the fusion network, and dynamically fuse the time-domain and frequency-domain features through the attention mechanism to obtain the final time-frequency combined MIP low-dimensional features; In S5, a time-frequency fusion module is used to fuse the time-domain and frequency-domain features of the MIP. The time-frequency feature fusion module consists of two parts: a cross-attention layer and a gated feature fusion layer. The specific steps for using the time-frequency fusion module to fuse the time-domain and frequency-domain features of the MIP are as follows: S51. In the cross-attention layer, the attention is calculated using the time-domain feature Query, frequency-domain feature Key, and Value of MIP. For each time-domain node feature, the correlation with the observed frequency-domain feature is calculated. S52. In the gated feature fusion layer, the gate controls the weight of feature fusion for different nodes. When a node is identified as a variable node, the feature components are biased towards frequency domain features, and when a node is constrained, the features are biased towards time domain features.

[0026] S6. Train the feature extraction model, introduce MIP attribute loss term and clustering-oriented loss term to optimize the model, so that the network focuses on learning the domain knowledge of MIP and optimizing the feature space so that the obtained feature space can better serve the downstream clustering task. In S6, the MIP attribute loss term and clustering-oriented loss term are introduced to optimize the model. The specific steps for extracting the network focus on learning the domain knowledge of MIP and optimizing the feature space for downstream clustering tasks are as follows: S61. Calculate the MIP attribute loss, which consists of three parts: variable type prediction loss, constraint sparsity alignment loss, and target coefficient KL divergence loss. S62, Reconstruction Loss: The original input is reconstructed through an autoencoder, and then the variable type cross-entropy, target coefficient binning MSE, and constraint sparsity MSE are calculated with the original input to ensure that the features retain key information. S63. Calculate the clustering-oriented loss for downstream clustering tasks. Calculate the feature similarity of the obtained features, then use the feature similarity to calculate the predicted distribution q. Use q to calculate the desired sharper feature space, i.e., the target distribution p. Calculate the KL divergence for q and p as the clustering loss to optimize the feature space. The formulas for calculating the feature similarity d and the predicted distribution q are as follows: ; In the formula, The fused time-frequency characteristics; The formula for calculating the sharpened target distribution p is: ; ; ; The expression for the clustering loss term is: .

[0027] Preferably, in S61, the specific steps for calculating the MIP attribute loss are as follows: S611. Using variable type prediction loss, the variable type is predicted by fusing the first n dimensions of the features, and then the cross-entropy loss function is calculated with the true variable type. This preserves the variable type information of the MIP in the feature space. The formula for calculating the variable type prediction loss is as follows: ; In the formula, For real variable types, These are the three types of probability distributions for prediction; S612. Using constraint sparsity alignment loss, predict the sparsity of constraints using the (n+1)th dimension of the fused features, and then calculate the L1 loss with the true constraint sparsity, so that the feature space retains the constraint sparsity information of MIP. The sparsity calculation formula is: ; In the formula, nnz(A) is the number of non-zero elements in the constraint matrix. For the number of variables, For constraint numbers; The formula for calculating the constraint sparsity alignment loss is: ; S613. Using the target coefficient KL divergence loss, calculate the probability that the coefficients of the variables in the objective function fall within each interval, so that the learned feature target coefficient distribution approximates the true coefficient distribution. The target coefficient KL divergence loss includes: Target coefficient binning: The target coefficient values ​​of the variables are divided into 20 bins. Count the number of variables in each bin. ; Calculate the probability distribution: ; KL divergence calculation formula: ; Guarantee the learned feature distribution Approximating the true distribution .

[0028] S7. Cluster the fused features, extract typical instances, input the final low-dimensional features into the clustering layer to obtain cluster centers and clusters, calculate the local density for each data point in the cluster, and select the k instances with the highest local density in each cluster as typical instances.

[0029] In S7, the obtained low-dimensional features are clustered to extract typical instances, including the following steps: S71. Preprocess the fusion features by standardizing them, setting the mean of each dimension to 0 and the variance to 1. S72. After determining that the cluster is one-fiftieth of the total number of instances and completing the clustering, calculate the local density of each instance using kernel density, and select the k instances with the highest density as typical instances within the cluster for output.

[0030] A system employing a time-frequency combined feature extraction and clustering acceleration optimization model training method includes a dataset acquisition module, a MIP bipartite graph modeling module, a feature extraction module, a feature fusion module, and a clustering extraction of typical instances module; the system also includes a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory.

[0031] Example 1

[0032] Step 1: Optimize the integrated energy system into a Model-In-Package (MIP); The modeling process is as follows: ; ; ; ; ; ; ; Among them, G th For a collection of thermal power units, G re For a collection of renewable energy units, Su k,t G represents the startup cost of unit k during time period t. i ( ) is the piecewise linear cost function of unit i.

[0033] Step 2: Model the MIP as a bipartite graph; The characteristic matrix of the variable node matrix is Each unit has seven characteristics, including unit type, minimum output, maximum output, current output, start / stop status, ramp rate, and cost coefficient.

[0034] The constraint node matrix consists of 48 power balance constraints + 48 spare constraints + 96 ramp constraints; The characteristic matrix is ; The adjacency matrix is , A[i,j]=1, when variable i participates in constraint j, the unit output variable participates in the power balance constraint of the corresponding time period.

[0035] Step 3: Time-frequency domain feature extraction network; Table 1 Time-Domain Encoders ;

[0036] Table 2 Frequency Domain Encoders ;

[0037] Step 4: Time-frequency feature fusion; First, perform cross-attention calculation: ; ; Secondly, gating fusion: ; ; Among them, for unit variable nodes (g≈0.3, focusing on the frequency domain), and for constraint-related nodes (g≈0.7, focusing on the time domain).

[0038] Next, we design the MIP characteristic loss function: Variable type prediction (classification): ; Constraints on sparsity (regression): ; Target coefficient distribution (KL divergence): ; Finally, the clustering-guided loss is calculated: Calculate the feature similarity matrix Q: ; Sharpen target distribution P: ; Final loss: ; Step 5: Model Training; Input 100 JSON data points from the power grid, with training parameters: Adam optimizer (lr=1e-4), batch_size=32, epochs=100.

[0039] In the first 50 epochs of training, the model loss parameters do not include clustering loss; only the feature extraction model parameters are trained. After the first round of training completes clustering, clustering-guided loss is added to optimize the model.

[0040] Step 6: Deploy the system and computer equipment.

[0041] The hardware configuration includes an AMD 5700X3D processor, 16GB of RAM, a 1TB SSD, and an NVIDIA RTX 4070 Super graphics card for accelerating neural network training. The software implementation utilizes the PyTorch framework, Python 3.7, the CUDA 12.01 development environment, and VS Code for deployment.

[0042] Therefore, this invention provides a time-frequency combined feature extraction and clustering method to accelerate and optimize model training. By combining the powerful feature extraction capabilities of deep learning with the time-frequency combined method, it can reflect the structural and content characteristics of MIPs, more accurately identify the characteristics of each MIP, and cluster a large number of randomly generated MIPs to extract representative and characteristic MIPs for subsequent training data generation. This can extract typical instances for the training of subsequent MIP solver models, eliminate a large number of redundant MIP problems, save a lot of computing power, and hardly lose the accuracy and speed of the solution.

[0043] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for accelerating and optimizing model training by combining time-frequency feature extraction and clustering, characterized in that, include: S1. Obtain a large amount of mixed integer programming (MIP) equation data in integrated energy systems; S2. Perform data augmentation by extracting and preprocessing the MIP data, and model the MIP problem as a bipartite graph from the perspective of variables and constraints, obtaining the variable node matrix, constraint node matrix and adjacency matrix of the MIP bipartite graph. S3. Extract time-domain features: Extract low-dimensional MIP time-domain features from the variable and constraint node matrices of the MIP bipartite graph using a time-domain autoencoder. S4. Extract frequency domain features. The adjacency matrix of the MIP bipartite graph is processed by a frequency domain encoder and frequency domain techniques to obtain an approximate feature vector matrix. Then, the spectral energy entropy of each node is calculated as the frequency domain feature. S5. Input the time-domain and frequency-domain features into the fusion network, and dynamically fuse the time-domain and frequency-domain features through the attention mechanism to obtain the final time-frequency combined MIP low-dimensional features; S6. Train the feature extraction model, introduce MIP attribute loss term and clustering-oriented loss term to optimize the model, so that the network focuses on learning the domain knowledge of MIP and optimizing the feature space so that the obtained feature space can better serve the downstream clustering task. S7. Cluster the fused features, extract typical instances, input the final low-dimensional features into the clustering layer to obtain cluster centers and clusters, calculate the local density for each data point in the cluster, and select the k instances with the highest local density in each cluster as typical instances.

2. The method for accelerating and optimizing model training by combining time-frequency features and clustering according to claim 1, characterized in that, In S1, the MIP data includes information on the distribution of variable types, constraint sparsity, and objective function coefficients of the integrated energy system, which is represented as follows: ; In the formula, n is the total number of variables; P is the number of integer variables (i.e., the first p variables of x are integers, and the rest are real numbers); A∈R mxn The constraint coefficient matrix: b∈R m To constrain the right-hand coefficient vector; c∈R n Let the coefficient vector of the objective function be: 1, u∈R n These represent the lower and upper bound vectors of the variable, respectively.

3. The time-frequency combined feature extraction and clustering accelerated optimization model training method according to claim 1, characterized in that, In S2, the specific steps for modeling the MIP problem as a bipartite graph after obtaining a large amount of MIP equation data are as follows: S21. Extract the type and boundary features of variables from the MIP equation and construct a variable feature matrix; S22. Extract the constraint type and right-hand constant, and construct the constraint feature matrix; S23. Construct the edge indices between variables and constraints to obtain the adjacency matrix of MIP.

4. The method for accelerating and optimizing model training by combining time-frequency features and clustering according to claim 1, characterized in that, In S3, the constructed temporal encoder consists of a type-aware embedding layer, a feature fusion layer, a bipartite graph convolutional layer, and a fully connected layer: the type-aware layer maps discrete type information into continuous semantic vectors; the feature fusion layer integrates the structural features and type semantics of the MIP into a single feature vector; the bipartite graph convolutional layer is used for information transfer to extract features from the bipartite graph of the MIP; and the fully connected layer outputs the temporal features as features with the same dimensions as the frequency domain features.

5. The method for accelerating and optimizing model training by combining time-frequency features and clustering according to claim 1, characterized in that, In S4, the frequency domain encoder consists of three parts: the Laplacian matrix, the low-rank approximation, and the calculated spectral energy entropy. The specific steps for constructing the frequency domain encoder to extract the frequency domain features of the MIP are as follows: S41. Input the adjacency matrix A of the MIP bipartite graph, and use A to calculate the normalized bipartite graph Laplace matrix. S42. Use the Nystrom method to extract the Top-k feature vectors; S43. Calculate the spectral energy distribution of each node as a frequency domain feature.

6. The method for accelerating and optimizing model training by combining time-frequency features and clustering according to claim 1, characterized in that, In S5, a time-frequency fusion module is used to fuse the time-domain and frequency-domain features of the MIP. The time-frequency fusion module consists of two parts: a cross-attention layer and a gated feature fusion layer. The specific steps for using the time-frequency fusion module to fuse the time-domain and frequency-domain features of the MIP are as follows: S51. In the cross-attention layer, the attention is calculated using the time-domain feature Query, frequency-domain feature Key, and Value of MIP. For each time-domain node feature, the correlation with the observed frequency-domain feature is calculated. S52. In the gated feature fusion layer, the gate controls the weight of feature fusion for different nodes. When a node is identified as a variable node, the feature components are biased towards frequency domain features, and when a node is constrained, the features are biased towards time domain features.

7. The method for accelerating and optimizing model training by combining time-frequency features and clustering according to claim 1, characterized in that, In S6, the MIP attribute loss term and clustering-oriented loss term are introduced to optimize the model. The specific steps for extracting the network focus on learning the domain knowledge of MIP and optimizing the feature space for downstream clustering tasks are as follows: S61. Calculate the MIP attribute loss, which consists of three parts: variable type prediction loss, constraint sparsity alignment loss, and target coefficient KL divergence loss. S62, Reconstruction Loss: The original input is reconstructed through an autoencoder, and then the variable type cross-entropy, target coefficient binning MSE, and constraint sparsity MSE are calculated with the original input to ensure that the features retain key information. S63. Calculate the clustering-oriented loss for downstream clustering tasks. Calculate the feature similarity of the obtained features, then use the feature similarity to calculate the predicted distribution q. Use q to calculate the desired sharper feature space, i.e., the target distribution p. Calculate the KL divergence for q and p as the clustering loss to optimize the feature space. The formulas for calculating the feature similarity d and the predicted distribution q are as follows: ; In the formula, The fused time-frequency characteristics; The formula for calculating the sharpened target distribution p is: ; ; ; The expression for the clustering loss term is: 。 8. The method for accelerating and optimizing model training by combining time-frequency features and clustering according to claim 1, characterized in that, In S61, the specific steps for calculating the MIP attribute loss are as follows: S611. Using variable type prediction loss, the variable type is predicted by fusing the first n dimensions of the features, and then the cross-entropy loss function is calculated with the true variable type. This preserves the variable type information of the MIP in the feature space. The formula for calculating the variable type prediction loss is as follows: ; In the formula, For real variable types, These are the three types of probability distributions for prediction; S612. Using constraint sparsity alignment loss, predict the sparsity of constraints using the (n+1)th dimension of the fused features, and then calculate the L1 loss with the true constraint sparsity, so that the feature space retains the constraint sparsity information of MIP. The sparsity calculation formula is: ; In the formula, nnz(A) is the number of non-zero elements in the constraint matrix. For the number of variables, To constrain the quantity; The formula for calculating the constraint sparsity alignment loss is: ; S613. Using the target coefficient KL divergence loss, calculate the probability that the coefficients of the variables in the objective function fall within each interval, so that the learned feature target coefficient distribution approximates the true coefficient distribution. The KL divergence calculation formula is as follows: ; In the formula, For the learned feature distribution, This represents the true distribution.

9. The method for accelerating and optimizing model training by combining time-frequency features and clustering according to claim 1, characterized in that, In S7, the obtained low-dimensional features are clustered to extract typical instances, including the following steps: S71. Preprocess the fusion features by standardizing them, setting the mean of each dimension to 0 and the variance to 1. S72. After determining that the cluster is one-fiftieth of the total number of instances and completing the clustering, calculate the local density of each instance using kernel density, and select the k instances with the highest density as typical instances within the cluster for output.

10. A system applying the time-frequency combined feature extraction and clustering accelerated optimization model training method as described in claims 1-9, characterized in that, The system includes a dataset acquisition module, a MIP bipartite graph modeling module, a feature extraction module, a feature fusion module, and a clustering extraction of typical instances module; the system also includes a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory.

Citation Information

Patent Citations

  • Electric power system end-to-end unit combination method for driving space-time attention graph neural network through object number fusion

    CN120449698A

  • Catalytic cracking unit simulation and prediction method based on molecular-level mechanism model and big data technology

    WO2023040512A1