An intelligent data mining platform and method based on big data analysis
By building the meta-model and confidence analysis mechanism of the LSTM network, combined with the joint optimization framework of reinforcement learning and differential evolution, the problem of incoordination of model selection bias and hyperparameter optimization in big data mining is solved, and more efficient and accurate data mining is achieved.
Patent Information
- Application Number
- CN202510449839.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-11
AI Technical Summary
In the existing big data mining methods, the problem of model selection deviation, insufficient uncertainty evaluation, and incoordinated hyperparameter optimization.
A meta-model based on LSTM network is constructed, and the historical data and model snapshot characteristics of the power company are extracted through a dual-channel architecture, combined with the confidence analysis mechanism to judge the reliability of the prediction results, and a joint optimization framework for reinforcement learning and differential evolution is launched to optimize hyperparameters.
It improves the accuracy and efficiency of data mining, can cope with data uncertainty and complexity, provides global optimal solutions, and improves the adaptability and generalization capabilities of the model.
Smart Images

Figure CN119961493B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data mining and intelligent analysis, and specifically provides an intelligent data mining platform and method based on big data analysis. Background Art
[0002] With the rapid development of information technology, big data analysis technology has become the core support for all walks of life. Big data analysis technology mainly relies on data mining algorithms to extract valuable information from massive data. In recent years, data mining technology has been widely applied, especially in industries such as finance, healthcare, power, and retail, and its application scenarios involve multiple fields such as prediction, classification, clustering, and association rule discovery. To solve the high-dimensional complex problems in big data, more and more research has focused on the intelligent and automatic selection of algorithms. In particular, research based on meta-learning and adaptive algorithm optimization frameworks has become a popular direction. Meta-learning technology can predict the optimal algorithm based on the historical data of power companies and further optimize it by combining technologies such as deep learning and reinforcement learning, making the data mining process more efficient and accurate.
[0003] However, the existing algorithm selection methods based on meta-learning still have some obvious deficiencies. First, most meta-learning methods rely on traditional single algorithm selection models, ignoring the dynamic changes and diversity of data, resulting in the inability to comprehensively consider various potential algorithm combinations on complex high-dimensional data sets and adapt to the changes in data sources of different power companies. Second, most of the reinforcement learning and differential evolution (DE) optimization frameworks in the existing technology operate independently and fail to achieve effective synergy. Although reinforcement learning can perform global optimization and differential evolution can perform local search, the combination of the two still poses significant challenges, especially in the process of hyperparameter optimization, where the global optimum is often not achieved. Third, most of the existing meta-models lack effective evaluation of the uncertainty of prediction results and fail to monitor and adjust the optimization path in real time, making the model highly uncertain when dealing with unknown data. Summary of the Invention
[0004] In view of the above problems, the present invention is proposed.
[0005] Therefore, the technical problems solved by the present invention are the problems of model selection bias, insufficient uncertainty evaluation, and uncoordinated hyperparameter optimization in existing big data mining methods.
[0006] To solve the above technical problems, the present invention provides the following technical solution: An intelligent data mining method based on big data analysis, comprising:
[0007] Obtaining data to be processed from the data source of a power company and performing preprocessing to generate a processing data set for data mining;
[0008] Based on the historical data of the power company, a meta-model is constructed to predict the optimal mining algorithm; the historical data includes historical load data, historical customer demand data, and historical equipment failure data;
[0009] The meta-model includes constructing a meta-model using a recurrent neural network, establishing a dual-channel architecture through an LSTM network and the historical data of the power company, concatenating and merging the outputs to obtain the final meta-model prediction result; the dual-channel architecture includes: a first channel and a second channel; the first channel includes inputting the historical data of the power company into the LSTM layer to extract the feature and temporal information of the historical data of the power company; the second channel includes inputting the historical model snapshot into the LSTM layer to perform temporal modeling to capture the impact of the historical mining algorithm prediction model parameters on the task performance at different training stages, recording and updating each set of hyperparameter settings of the historical mining algorithm prediction model and its characteristics shown during the training process, and learning the long-term dependencies in the historical mining algorithm prediction task;
[0010] By calculating the confidence of the meta-model, uncertainty analysis is performed on the prediction result output by the meta-model. When the confidence judgment result indicates that the meta-model prediction result is uncertain, a joint optimization framework based on reinforcement learning and differential evolution is started for hyperparameter optimization; the joint optimization framework based on reinforcement learning and differential evolution is to establish a reinforcement learning-differential evolution collaborative controller, transmit the decision information of reinforcement learning to the differential evolution algorithm, and configure the strategy of reinforcement learning;
[0011] Feed the hyperparameters optimized by the joint optimization framework based on reinforcement learning and differential evolution back to the meta-model, select the optimal mining algorithm, and execute the data mining task.
[0012] As a preferred solution of the intelligent data mining method based on big data analysis according to the present invention, wherein: the data to be processed includes structured data, unstructured data, and temporal data;
[0013] Preprocess and extract features from the historical data; use the data after feature extraction for training and parameter setting. During the training process, record the historical mining task data and the historical model snapshot;
[0014] The historical mining task data is a detailed record of past data mining tasks, including: the mining algorithms used before, hyperparameter settings, and historical mining algorithm prediction results;
[0015] The historical model snapshot includes the parameters of the historical mining algorithm prediction model, network structure parameters, hyperparameters of the optimization algorithm, training strategy, and the selection of the loss function.
[0016] As a preferred solution of the intelligent data mining method based on big data analysis according to the present invention, wherein: the meta-model includes: an input layer, a dual-channel architecture layer, a merging layer, a fully connected layer, a migration decision module layer, and an output decision layer;
[0017] The input layer includes: taking the historical data and historical model snapshots of the power company as data inputs.
[0018] As a preferred solution of the intelligent data mining method based on big data analysis according to the present invention, wherein: the merging layer includes: concatenating and merging the output information from the dual-channel architecture to form a high-dimensional feature representation;
[0019] The fully connected layer includes taking the merged high-dimensional features as inputs, performing non-linear transformation, and generating intermediate results;
[0020] The migration decision module layer includes initializing the meta-model using the historical mining algorithm prediction model parameters and performing fine-tuning; during fine-tuning, updating is performed using the difference between the output of the meta-model and the historical mining algorithm prediction model parameters;
[0021] The output decision layer includes making a final decision based on the intermediate results output by the fully connected layer and combining confidence judgment.
[0022] As a preferred solution of the intelligent data mining method based on big data analysis according to the present invention, wherein: the confidence of the meta-model includes: calculating the entropy value of the meta-model; when the entropy value of the meta-model is less than or equal to the entropy threshold, it is considered that the prediction result of the meta-model is determined, and the prediction result is directly executed;
[0023] When the entropy value of the meta-model is greater than the entropy threshold, calculate the similarity of the meta-model; when the similarity of the meta-model is greater than the similarity threshold, it is considered that the prediction result of the meta-model is determined, and the prediction result is directly executed;
[0024] The calculation of the similarity of the meta-model includes: calculating the feature similarity and calculating the context similarity;
[0025] Calculate the feature similarity through the weighted Jaccard coefficient and the improved DTW distance; calculate the context similarity through the semantic matching degree of the business scenario label considering the Word2Vec word vector;
[0026] When the meta-model similarity is less than or equal to the similarity threshold, it is considered that the prediction result of the meta-model is uncertain, and the joint optimization framework based on reinforcement learning and differential evolution is started for hyperparameter optimization.
[0027] As a preferred solution of the intelligent data mining method based on big data analysis according to the present invention, wherein: the joint optimization framework based on reinforcement learning and differential evolution includes: loading the historical data of the power company into the experience pool of reinforcement learning and setting resource constraint conditions;
[0028] The resource constraint conditions include: Condition 1, stop iterating when the maximum number of iterations is reached;
[0029] Condition 2, stop iterating when the time budget is reached;
[0030] If any one of Condition 1 or Condition 2 is satisfied, stop iterating;
[0031] The joint optimization framework based on reinforcement learning and differential evolution further includes: extracting n-dimensional statistical features from the processed data set, obtaining the algorithm characteristics recommended by the meta-model, finding the Top-k similar scenarios in the reinforcement learning experience pool, calculating the scene similarity weights based on the Mahalanobis distance in the feature space, and deriving the recommended confidence interval of the parameter space of the processed data set according to the optimal parameter distribution of the Top-k similar scenarios;
[0032] The joint optimization framework based on reinforcement learning and differential evolution further includes: constructing a differential evolution population containing x hyperparameter individuals, The x hyperparameter individuals are generated within the range recommended by reinforcement learning, The x hyperparameter individuals are randomly sampled from the parameter space of the processed data set;
[0033] Perform hybrid coding on the hyperparameter individuals; the hybrid coding is to use real number coding for continuous parameter individuals and Gray code coding for discrete parameter individuals to reduce the mutation probability of adjacent values, and introduce the positive correlation constraint between batch_size and learning_rate;
[0034] Calculate the Euclidean distance between each individual and the current optimal solution, and calculate the difference degree between individuals;
[0035] Set a dynamic difference degree threshold. When the distance between an individual and the current optimal solution is less than or equal to the dynamic threshold, perform fine search;
[0036] The fine search is that the mutation step size range is P1; each individual performs a direction exploration and fine adjustment;
[0037] When the distance between an individual and the current optimal solution exceeds the dynamic threshold, perform global search;
[0038] The global search is that the mutation step size range is P2; retain the optimal gene segments of the previous b generations for crossover operation;
[0039] Apply reverse perturbation to the individuals that have not been improved for continuous b generations to prevent falling into local optimum;
[0040] Mirror map the individuals close to the constraint boundary and adjust their exploration range.
[0041] As a preferred solution of the intelligent data mining method based on big data analysis according to the present invention, wherein: the joint optimization framework based on reinforcement learning and differential evolution further includes: after every b generations of differential evolution, the population diversity coefficient is monitored in real time; when the population diversity coefficient is greater than the diversity threshold, trigger the reinforcement learning evaluation;
[0042] According to the optimal performance, shrink the parameter range and retain the parameter interval with the best performance ;
[0043] Extend the parameter space according to the gradient direction to expand the exploration area;
[0044] Set up a reward mechanism to quantify the benefits of parameter adjustment in each optimization and dynamically adjust the reinforcement learning strategy;
[0045] The reward mechanism is to calculate the accuracy increment of the meta-model and the time consumed after each iteration, measure the benefit ratio through the accuracy increment and the time consumed, and use the benefit ratio as a reward signal to feedback to the joint optimization framework of reinforcement learning and differential evolution to update the Q value of reinforcement learning;
[0046] Retain the top individuals of each generation and store them in the elite pool;
[0047] Perform elite gene recombination every a generations, adopt a weighted crossover strategy based on parameter sensitivity, and give priority to retaining high-sensitivity parameters;
[0048] The convergence determination conditions include: when the improvement rate of the optimal solution for a consecutive a generations is less than the improvement rate threshold, it is determined to converge;
[0049] When the population gene similarity exceeds the population gene similarity threshold, it is determined to converge;
[0050] When the remaining time budget is less than the time budget threshold of the time budget, force the output of the current optimal solution and stop the optimization process.
[0051] An intelligent data mining platform for big data analysis, wherein:
[0052] A data collection and preprocessing module, which is used to obtain the data to be processed from the data source of the power company and perform preprocessing to generate a processing data set for data mining; the data to be processed includes structured data, unstructured data and time series data; the unstructured data includes text, pictures and audio data;
[0053] A prediction module, configured to construct a meta-model based on historical load data, historical customer demand data, and historical equipment failure data of a power company. The meta-model adopts a dual-channel architecture to predict the optimal data mining algorithm.
[0054] An optimization module, configured to perform uncertainty analysis on the prediction results output by the meta-model according to the confidence level output by the meta-model. When the confidence level judgment result indicates that the prediction result of the meta-model is uncertain, it automatically calls an optimization framework combining reinforcement learning and differential evolution to dynamically adjust and optimize the hyperparameters, and feeds the optimization results back to the meta-model to enhance the model's adaptive ability.
[0055] An execution module, configured to automatically select the optimal mining algorithm and execute the data mining task according to the output of the optimized meta-model, so as to realize intelligent decision-making and business support for load prediction, customer demand analysis, and equipment failure detection in the power system.
[0056] A computer device, comprising: a memory and a processor; the memory stores a computer program, and is characterized in that: when the processor executes the computer program, the steps of the method described in any one of the present inventions are implemented.
[0057] A computer-readable storage medium, on which a computer program is stored, and is characterized in that: when the computer program is executed by a processor, the steps of the method described in any one of the present inventions are implemented.
[0058] The beneficial effects of the present invention: The intelligent data mining method based on big data analysis provided by the present invention improves the data quality in the data preprocessing stage, providing reliable input for subsequent mining. By constructing a meta-model based on the LSTM network, the optimal mining algorithm is automatically predicted, avoiding the deviation of manual algorithm selection in traditional methods. The confidence level analysis mechanism of the meta-model output result can judge the reliability of the prediction result. When the confidence level is low, a joint optimization framework of reinforcement learning and differential evolution is started to optimize the hyperparameters, thereby improving the performance and accuracy of the mining algorithm. The optimized hyperparameters are fed back to the meta-model to ensure the selection of the optimal algorithm and the execution of the data mining task, improving the mining efficiency and accuracy. It not only improves the accuracy and efficiency of data mining, but also can cope with the uncertainty and complexity of data, providing a globally optimal solution while ensuring the computing efficiency, and has important practical value. Description of the Drawings
[0059] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0060] Figure 1 This is the overall flowchart of an intelligent data mining method based on big data analysis provided for the first embodiment of the present invention. Specific embodiments
[0061] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0062] Embodiment 1, referring to Figure 1 , which is an embodiment of the present invention, provides an intelligent data mining method based on big data analysis, including:
[0063] S1: Obtain the data to be processed from the data source of the power company, and perform preprocessing to generate a processing data set for data mining.
[0064] The data to be processed includes historical load data of the power company, user electricity demand information, equipment failure operation and maintenance records, etc., with the characteristics of structured data, unstructured data, and time series data.
[0065] The historical data of the power company is historical load data, historical customer demand data, and historical equipment failure data.
[0066] Perform preprocessing and feature extraction on the historical data;
[0067] Use the data after feature extraction for training and parameter setting. During the training process, record historical mining task data and historical model snapshots.
[0068] The historical mining task data is a detailed record of past data mining tasks, including: the mining algorithms used before, hyperparameter settings, and historical mining algorithm prediction results; the historical model snapshots include the parameters of the historical mining algorithm prediction model, network structure parameters, optimization algorithm hyperparameters, training strategies, and the selection of loss functions.
[0069] Structured data preprocessing includes, for missing fields, filling them with the mean value, using statistical methods to identify and process outliers in the data. Compress the data to the [0,1] interval through normalization to make the scales of each feature consistent, facilitating subsequent algorithm processing.
[0070] Unstructured data includes text, pictures, and audio data. The preprocessing steps are as follows:
[0071] Text data preprocessing includes: breaking the text into vocabulary units, using the word segmentation tool (NLTK) to segment words, removing meaningless words such as "的" and "和", and punctuation marks in the text. Restore words to their roots. Use the TF-IDF method to convert text into numerical vectors.
[0072] Image data preprocessing includes: cropping the image as needed, keeping the image size consistent, scaling the pixel values to the [0,1] interval, or normalizing the image mean and standard deviation. Increase the diversity of training data through rotation, flipping, and scaling to avoid overfitting.
[0073] Time series data preprocessing includes: Forward filling is often used to fill missing values in time series data.
[0074] Use the sliding average method to remove noise from time series data. Divide the time series data into multiple windows according to time periods to extract trend and seasonal characteristics from the time series data.
[0075] The historical data of the power company includes historical mining task data and historical model snapshots. The preprocessing steps include:
[0076] Clean up errors or redundant information in the power company's historical data to ensure data accuracy. Integrate historical task data and historical model snapshots, and associate the power company's historical data through timestamp and task ID fields. Extract relevant features from historical task data, including mining algorithms and hyperparameter settings; extract network structure parameters, optimization algorithm hyperparameters, etc. from historical model snapshots.
[0077] Based on the preprocessed data, generate a processed data set for data mining:
[0078] The data is divided into training set, validation set and test set in a ratio of 8:1:1 to ensure the fairness of model training and evaluation.
[0079] Furthermore, by preprocessing structured data, unstructured data, and time series data, we can solve problems such as missing values, outliers, and noise in the data, and improve the consistency and reliability of the data. At the same time, we use methods such as normalization and TF-IDF to standardize and transform the data to ensure that all types of data can be processed on the same scale, providing a more effective training basis for subsequent algorithm models.
[0080] Furthermore, through the cleaning and integration of the historical data of the power company, not only can the accuracy of the data be improved, but also the algorithm can be helped to better understand past tasks and model performance by extracting key features. This process ensures a reasonable starting point for the model during training and enables fair and effective evaluation based on different data sets (such as training sets, validation sets, and test sets), thereby improving the generalization ability and prediction accuracy of the model.
[0081] S2: Based on the historical data of the power company, construct a meta-model to predict the optimal mining algorithm.
[0082] The meta-model includes constructing a meta-model using a recurrent neural network, establishing a dual-channel architecture through an LSTM network and the historical data of the power company, concatenating and merging the outputs, and obtaining the final meta-model prediction result.
[0083] The meta-model also includes: an input layer, a dual-channel architecture layer, a merging layer, a fully connected layer, a transfer decision module layer, and an output decision layer.
[0084] The input layer includes: taking the historical data of the power company and the historical model snapshot as data inputs.
[0085] The dual-channel architecture includes: a first channel and a second channel.
[0086] The first channel includes inputting the historical data of the power company into the LSTM layer to extract the historical data features of the power company and the temporal information of the historical data of the power company;
[0087] The formula is expressed as:
[0088] ;
[0089] Where, is the input historical data of the power company, is the output historical data features of the power company, is the weight of the LSTM network.
[0090] The second channel includes inputting the historical model snapshot into the LSTM layer to perform temporal modeling to capture the impact of the historical mining algorithm prediction model parameters on the task performance at different training stages, recording and updating each set of hyperparameter settings of the historical mining algorithm prediction model and their characteristics shown during the training process, and learning the long-term dependencies in the historical mining algorithm prediction tasks.
[0091] The formula is expressed as:
[0092] ;
[0093] Where, is the input historical model snapshot data, is the snapshot feature of the output historical model, is the weight of the LSTM network.
[0094] The merging layer includes: concatenating the output information from the dual-channel architecture to form a high-dimensional feature representation.
[0095] The formula is expressed as:
[0096] ;
[0097] Among them, is the merged high-dimensional feature representation, represents concatenating the output features of the two channels.
[0098] The fully connected layer includes taking the merged high-dimensional feature as input, performing a non-linear transformation, and generating an intermediate result.
[0099] The formula is expressed as:
[0100] ;
[0101] Among them, is the output intermediate result of the fully connected layer, and are the weight matrix and bias term of the fully connected layer respectively, and ReLU(·) is the activation function.
[0102] The migration decision module layer includes using the historical mining algorithm to predict the model parameters to initialize the meta-model and performing fine-tuning; during fine-tuning, it is updated using the difference between the output of the meta-model and the predicted model parameters by the historical mining algorithm;
[0103] The output decision layer includes making a final decision based on the intermediate result output by the fully connected layer and combining confidence judgment.
[0104] The formula is expressed as:
[0105] ;
[0106] Among them, represents the final decision result, represents the predicted mining algorithm selection result, represents the intermediate result from the fully connected layer, represents the weight matrix of the fully connected layer, represents the bias of the fully connected layer, represents converting the result of the final decision into a probability distribution, represents the final value obtained through confidence calculation.
[0107] The confidence level of the meta-model includes: calculating the entropy value of the meta-model; when the entropy value of the meta-model is, it is considered that the prediction result of the meta-model is determined, and the prediction result is directly executed.
[0108] The formula for calculating the entropy value of the meta-model is expressed as:
[0109] ;
[0110] where represents the probability distribution of the model output category , and represents the entropy value of the meta-model.
[0111] The entropy threshold and the similarity threshold are set according to historical experience.
[0112] When the entropy value of the meta-model is greater than the entropy threshold, calculate the similarity of the meta-model; when the similarity of the meta-model is greater than the similarity threshold, it is considered that the prediction result of the meta-model is determined, and the prediction result is directly executed.
[0113] Calculating the similarity of the meta-model includes: calculating the feature similarity and calculating the context similarity.
[0114] The feature similarity is calculated by the weighted Jaccard coefficient and the improved DTW distance, and the formula is expressed as:
[0115] ;
[0116] where represents the current input feature, represents the historical data feature of the power company, DTW represents the dynamic time warping distance, represents the feature similarity.
[0117] The context similarity is calculated by considering the semantic matching degree of the business scenario label based on the Word2Vec word vector, and the formula is expressed as:
[0118] ;
[0119] where is the business scenario label word vector of the current task, is the label word vector of the historical task, represents the context similarity.
[0120] The formula for the final similarity is expressed as:
[0121] Final similarity:
[0122] ;
[0123] where Represents the final similarity, Represents the context similarity, Represents the feature similarity.
[0124] When the similarity is such, it is considered that the prediction result of the meta-model is uncertain, and the joint optimization framework based on reinforcement learning and differential evolution is started for hyperparameter optimization.
[0125] The confidence formula is expressed as:
[0126] ;
[0127] Wherein, represents the entropy value of the meta-model, represents the final similarity, represents the confidence.
[0128] Furthermore, by combining the recurrent neural network (LSTM) with the historical data of the power company and the historical model snapshots, an efficient meta-model is constructed, which can accurately capture and predict the performance of the mining algorithm at different training stages. Through the dual-channel architecture, the system can respectively extract the historical data features and historical model snapshot features of the power company and fuse them into a high-dimensional feature representation, so as to realize precise decision support for the mining algorithm. The addition of the fully connected layer and the transfer decision module further improves the accuracy and flexibility of the meta-model, enabling it to quickly make optimization decisions based on the historical data and historical model snapshots of the power company.
[0129] Even further, through the confidence calculation mechanism, a flexible decision-making method is provided. When the prediction result of the meta-model has a high degree of certainty, the prediction result can be directly executed to improve the calculation efficiency; while when the confidence is low, the system will start the joint optimization framework of reinforcement learning and differential evolution to further optimize the hyperparameters, thereby improving the accuracy and performance of the mining algorithm. By dynamically adjusting the confidence, the system can adapt to different task scenarios, ensuring that each decision is data-driven and optimal, so as to achieve the effect of optimizing the model performance.
[0130] S3: Through calculating the confidence of the meta-model, perform uncertainty analysis on the prediction result output by the meta-model. When the confidence judgment result indicates that the prediction result of the meta-model is uncertain, start the joint optimization framework based on reinforcement learning and differential evolution for hyperparameter optimization.
[0131] The joint optimization framework based on reinforcement learning and differential evolution includes: establishing a reinforcement learning-differential evolution collaborative controller, transmitting the decision-making information of reinforcement learning to the differential evolution algorithm, and configuring the strategy of reinforcement learning to ensure that it dynamically affects the search space of differential evolution in each round of optimization process;
[0132] Load the historical data of the power company into the experience pool of reinforcement learning, and set resource constraints;
[0133] The resource constraints include: Condition 1, when the maximum number of iterations reaches 1000 times, stop the iteration;
[0134] Condition 2, when the time budget reaches 3 hours, stop the iteration;
[0135] If any one of Condition 1 or Condition 2 is satisfied, stop the iteration;
[0136] The joint optimization framework based on reinforcement learning and differential evolution further includes: extracting n-dimensional statistical features from the processed dataset, obtaining the algorithm characteristics recommended by the meta-model, searching for the Top-K similar scenarios in the reinforcement learning experience pool, calculating the scene similarity weight based on the Mahalanobis distance in the feature space, and deriving the recommended confidence interval of the parameter space of the processed dataset according to the optimal parameter distribution of the Top-K similar scenarios. The formula is expressed as:
[0137] ;
[0138] ;
[0139] ;
[0140] Among them, represents the feature vector of the current data, represents the historical data of the power company, represents the Euclidean distance, represents the th scenario that is most similar to the current scenario feature among all historical scenarios.
[0141] The 10-dimensional statistical features include mean, variance, standard deviation, skewness, kurtosis, maximum value, minimum value, median, interquartile range, and covariance.
[0142] According to historical experience .
[0143] The joint optimization framework based on reinforcement learning and differential evolution further includes: constructing a differential evolution population containing 100 hyperparameter individuals, generating 80 hyperparameter individuals within the range recommended by reinforcement learning, and randomly sampling 20 hyperparameter individuals from the parameter space of the processed dataset;
[0144] Perform hybrid coding on the hyperparameter individuals; the hybrid coding uses real number coding for continuous parameter individuals and Gray code coding for discrete parameter individuals to reduce the mutation probability of adjacent values, and introduce the positive correlation constraint between batch_size and learning_rate;
[0145] Calculate the Euclidean distance between each individual and the current optimal solution, and calculate the difference degree between individuals;
[0146] Set a dynamic difference degree threshold. When the distance between an individual and the current optimal solution is less than or equal to the dynamic threshold, perform a fine search.
[0147] The formula for the dynamic threshold is expressed as:
[0148] ;
[0149] Among them, σ represents the adjustment factor; ε represents the error tolerance, represents the Euclidean distance between the i-th individual and the current optimal solution; represents the range of mutation step sizes during fine search; represents the current mutation step size; represents the maximum value of the mutation step size.
[0150] The fine search is that the range of mutation step sizes is P1.
[0151] P1 ranges from 0.1 to 0.3.
[0152] Each individual conducts a = 5 times of direction exploration and fine adjustment. The formula is expressed as:
[0153] ;
[0154] Among them, represents the adjustment of the exploration direction each time, represents the new exploration direction, represents the original exploration direction.
[0155] When the distance between an individual and the current optimal solution exceeds the dynamic threshold, perform a global search.
[0156] The global search is that the range of mutation step sizes is P2.
[0157] P2 ranges from 0.7 to 1.2.
[0158] Retain the optimal gene fragments of the first b = 3 generations for crossover operation;
[0159] Apply a reverse perturbation to the individuals that have not been improved for b = 3 consecutive generations to prevent falling into local optima. The formula is expressed as:
[0160] ;
[0161] Among them, represents the perturbation factor, represents the i-th individual after update, represents the current value of the i-th individual, represents the perturbation vector.
[0162] Mirror mapping is performed on the individuals close to the constraint boundary to adjust their exploration range, which is expressed by the formula:
[0163] ;
[0164] where, represents the current parameter value of the th individual, represents the constraint boundary of the parameter, represents comparing with and returning the smaller value; represents comparing the previous minimum value with zero and returning the larger value; represents the parameter value of the individual after adjustment by mirror mapping.
[0165] After every b = 3 generations of differential evolution are completed, the population diversity coefficient is monitored in real time; when the population diversity coefficient is greater than the diversity threshold, reinforcement learning evaluation is triggered; the diversity threshold is determined based on experience.
[0166] According to the optimal performance, the parameter range is shrunk, and the parameter interval with the best performance is retained; ; .
[0167] Extend the parameter space of in the gradient direction to expand the exploration area; .
[0168] Set up a reward mechanism to quantify the benefits of parameter adjustment in each optimization, and dynamically adjust the reinforcement learning strategy;
[0169] The reward mechanism is to calculate the accuracy increment of the meta-model and the time consumed after each iteration, measure the benefit ratio through the accuracy increment and the time consumed, and use the benefit ratio as a reward signal to feedback to the joint optimization framework of reinforcement learning and differential evolution to update the Q value of reinforcement learning.
[0170] The benefit ratio formula is expressed as:
[0171] ;
[0172] where, represents the accuracy increment, represents the time consumed.
[0173] Update the Q value of reinforcement learning, which is expressed by the formula:
[0174] ;
[0175] where, Represents the current state Under this condition, take an action The Q value of; Represents the learning rate; Represents the immediate reward; Represents the discount factor; Represents the new state Under this condition, all possible actions The maximum Q value of; Represents taking an action The new state transferred to after; Represents the new state Possible actions under this condition.
[0176] Retain the top Individuals of each generation and store them in the elite pool; .
[0177] Execute elite gene recombination every a = 5 generations, adopt a weighted crossover strategy based on parameter sensitivity, and give priority to retaining high-sensitivity parameters. The formula is expressed as:
[0178] ;
[0179] Among them, Represents the newly generated offspring individual; Represents parent individual 1 and Represents parent individual 2, Represents the parameter sensitivity weight.
[0180] The convergence determination conditions include: when the improvement rate of the optimal solution for a continuous a = 5 generations is less than the improvement rate threshold, it is judged as convergence;
[0181] When the population gene similarity exceeds the population gene similarity threshold, it is determined as convergence;
[0182] When the remaining time budget is less than the time budget threshold of the time budget, force the output of the current optimal solution and stop the optimization process.
[0183] The improvement rate threshold, population gene similarity threshold, and time budget threshold are set according to historical experience.
[0184] Furthermore, by combining the advantages of reinforcement learning and differential evolution algorithm, making full use of the strategy of reinforcement learning to guide the search process of differential evolution, thus effectively optimizing the hyperparameter search space. In this framework, reinforcement learning can dynamically adjust the optimization strategy, provide accurate parameter recommendations based on the historical data and feature analysis of power companies, while differential evolution realizes efficient hyperparameter exploration through global and local search mechanisms. This collaborative control strategy ensures that during the optimization process, it can quickly search for the optimal solution and avoid falling into local optima, improving the overall optimization effect.
[0185] Furthermore, through the optimization process of dynamically adjusting and real-time monitoring parameters, the adaptability and efficiency of the algorithm in complex environments have been significantly improved. The set dynamic thresholds, fine-grained search, and global search mechanisms enable the algorithm to flexibly adjust the search strategy under different conditions, avoiding premature convergence. At the same time, through the calculation of the reward mechanism and benefit ratio of reinforcement learning, the optimization effect can be quantified and the strategy can be adjusted in real time, further enhancing the stability and accuracy of the optimization process.
[0186] S4: Feed the hyperparameters optimized by the joint optimization framework based on reinforcement learning and differential evolution back to the meta-model, select the optimal mining algorithm, and perform the data mining task.
[0187] Transfer the optimized hyperparameter set from the joint optimization framework to the meta-model to ensure that this hyperparameter set can effectively improve the performance of the meta-model. The meta-model adjusts its internal structure according to the hyperparameter set, thereby improving the prediction ability.
[0188] In the data mining task, the meta-model selects the most suitable mining algorithm for the current task according to the historical data of the power company and the distribution of the feature space.
[0189] Use the optimized algorithm to perform the mining task on the input data set and output the final mining result, which result.
[0190] Furthermore, by passing the hyperparameter set optimized by the joint optimization framework based on reinforcement learning and differential evolution to the meta-model, it can be ensured that the meta-model can use the best hyperparameter configuration, thereby effectively improving its performance. The optimized hyperparameters can guide the meta-model to adjust its own structure, enabling the model to flexibly adapt and improve the prediction accuracy when facing different data sets and tasks. This hyperparameter tuning process avoids the limitations of relying on fixed parameters, giving the meta-model higher self-adaptability and stronger generalization ability.
[0191] Furthermore, by using the optimized hyperparameter set, the meta-model can automatically select the most suitable mining algorithm for the current task according to the historical data of the power company and the distribution of the feature space, and perform the data mining task on this basis. The finally output mining result is more accurate and practical, ensuring that the model can make optimal predictions and decisions for specific data sets and target tasks, thereby enhancing the intelligent level and application effect of the overall system.
[0192] Embodiment 2, an embodiment of the present invention, provides an intelligent data mining platform based on big data analysis, including:
[0193] A data acquisition and preprocessing module is used to obtain data to be processed from the data sources of the power company and perform preprocessing to generate a processed data set for data mining. The data to be processed includes structured data, unstructured data, and time-series data. The unstructured data includes text, pictures, and audio data.
[0194] A prediction module is used to construct a meta-model based on the historical load data, historical customer demand data, and historical equipment failure data of the power company. The meta-model adopts a dual-channel architecture to predict the optimal data mining algorithm.
[0195] An optimization module is used to perform uncertainty analysis on the prediction results output by the meta-model according to the confidence level. When the confidence level judgment result indicates that the prediction result of the meta-model is uncertain, it automatically calls an optimization framework combining reinforcement learning and differential evolution to dynamically adjust and optimize the hyperparameters, and feeds the optimization results back to the meta-model to improve the model's adaptive ability.
[0196] An execution module is used to automatically select the optimal mining algorithm and execute the data mining task according to the output of the optimized meta-model, so as to realize intelligent decision-making and business support for load prediction, customer demand analysis, and equipment failure detection in the power system.
[0197] Example 3, an embodiment of the present invention, which is different from the previous two embodiments:
[0198] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
[0199] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0200] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0201] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0202] Example 4, an embodiment of the present invention, provides an intelligent data mining platform and method based on big data analysis. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through simulation experiments.
[0203] Historical load data, historical customer demand data, and historical equipment failure data of a certain power company were used in the experiment. These data include structured data (such as historical load data tables), unstructured data (such as equipment failure text descriptions), and time-series data (such as daily load change records). The specific experimental steps are as follows:
[0204] Data preprocessing: Extract three years of historical data of the power company from the power company database. After cleaning and preprocessing, the data contains missing values, outliers, and unstructured text data. Fill the missing values with the mean value, correct the outliers by the Z-score-based method, and perform tokenization on the unstructured text data. The time series data is normalized by the sliding window method to ensure that the data meets the requirements of model training.
[0205] Meta-model construction: A meta-model is constructed based on the historical data of the power company. The meta-model adopts a dual-channel architecture of LSTM (Long Short-Term Memory Network). The first channel processes the time series characteristics of historical load data, and the second channel processes the historical model snapshot information, extracting the hyperparameter settings of the model at different training stages and its characteristics shown during the training process. This dual-channel architecture can better capture the long-term dependencies and historical memories in the data.
[0206] Confidence analysis and optimization trigger: Based on the prediction results of the meta-model, calculate the entropy value to evaluate the confidence of the model. If the entropy value is high, it indicates that the model has a large uncertainty about the prediction results. Therefore, start a joint optimization framework based on reinforcement learning and differential evolution for hyperparameter optimization. Specifically, the reinforcement learning algorithm adjusts the search space of hyperparameters according to the experience pool, and the differential evolution algorithm optimizes these hyperparameters through the population evolution strategy.
[0207] Hyperparameter optimization and data mining task execution: During the hyperparameter optimization process, each hyperparameter is optimized through the joint optimization framework of reinforcement learning and differential evolution. The optimized hyperparameter set is fed back to the meta-model and used to select the most suitable mining algorithm for the current task. Finally, use the optimized algorithm to execute the data mining tasks, predict the load forecasting, customer demand analysis, and equipment fault detection respectively, and output the results.
[0208] During the experiment, the initial values, optimized values, prediction results, and relevant experiment time and accuracy improvement of each task were recorded. The following are the key data records of the experiment:
[0209] Load forecasting data: Under the initial conditions, the load forecasting value is 456.78 MW. After optimization, the forecasting value is 459.30 MW, and the actual forecasting result is 457.50 MW. The experiment time is 3.2 hours, and the accuracy improvement is 5.26%.
[0210] Customer demand analysis: The initial value is 2456.60, the optimized value is 2468.85, the actual forecasting result is 2472.10, the experiment time is 2.9 hours, and the accuracy improvement is 4.70%.
[0211] Equipment fault detection: The initial detection rate is 0.13, after optimization it is 0.10, the prediction result is 0.08, the experimental time is 4.1 hours, and the accuracy improvement is 10.00%.
[0212] Hyperparameter setting (learning rate): The initial learning rate is 0.01, and the optimized learning rate is 0.015. The accuracy improvement is not directly calculated.
[0213] Hyperparameter setting (batchsize): The initial value is 64, and the optimized value is 128. The accuracy improvement is not directly calculated.
[0214] From the above experimental data, the advantages and innovativeness of the method of the present invention in multiple tasks can be clearly seen:
[0215] After hyperparameter optimization, the accuracy of load prediction has increased by 5.26%. Traditional load prediction methods usually rely on fixed hyperparameters, while the present invention dynamically adjusts hyperparameters through joint optimization based on reinforcement learning and differential evolution, enabling the meta-model to better adapt to the changes and complexities in the data, thereby significantly improving the prediction accuracy.
[0216] Through the mining of customer demand data, the accuracy improvement is 4.70%. Traditional methods often rely on simple algorithm selection and static parameters, while the meta-model of the present invention can combine the historical data and model snapshots of power companies to dynamically select the most suitable mining algorithm, thereby improving the analysis accuracy.
[0217] The accuracy of equipment fault detection has increased by 10.00%. The improvement of this task is particularly significant. Traditional methods often cannot quickly adapt when data is abnormal, while the present invention effectively avoids this limitation through a dynamically adjusted hyperparameter optimization framework, providing more accurate fault detection results.
[0218] The joint optimization of reinforcement learning and differential evolution algorithms has significantly improved the performance of data mining tasks. By optimizing the learning rate (adjusted from 0.01 to 0.015) and batch size (adjusted from 64 to 128), the training time of the model has been shortened and the prediction accuracy has been improved. This process shows that the hyperparameter optimization framework of the present invention has strong self-adaptability, can dynamically adjust parameters according to the needs of different datasets, and improves the robustness of the model.
[0219] Compared with traditional static hyperparameter selection methods, the method of the present invention automatically adjusts hyperparameters in multiple rounds of training through the introduction of a joint optimization framework of reinforcement learning and differential evolution to complete data mining tasks with optimal configuration. Traditional methods usually train based on fixed parameters, while the dynamic adjustment mechanism of the present invention can automatically adapt according to actual data changes, improving the accuracy, efficiency and adaptability of data mining tasks.
[0220] Further experiments show that the joint optimization based on reinforcement learning and differential evolution can provide more accurate hyperparameter tuning in different data tasks. Especially when dealing with large-scale data, the optimized model can effectively improve the training speed and accuracy through a dynamic decision-making mechanism. Through multiple rounds of experimental verification, the inventive method has shown significant advantages in load forecasting, customer demand analysis, and equipment fault detection, especially in terms of the improvement in accuracy and efficiency, fully demonstrating the innovation and practicality of the present invention.
[0221] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. An intelligent data mining method based on big data analysis, characterized in that, Including: Obtain the data to be processed from the data source of the power company, perform preprocessing, and generate a processed data set for data mining; The data to be processed includes structured data, unstructured data, and time-series data; the unstructured data includes text, pictures, and audio data; Based on the historical data of the power company, construct a meta-model to predict the optimal mining algorithm; the historical data is historical load data, historical customer demand data, and historical equipment failure data; The meta-model includes constructing a meta-model using a recurrent neural network, establishing a dual-channel architecture through an LSTM network and the historical data of the power company, concatenating and merging the outputs to obtain the final meta-model prediction result; The dual-channel architecture includes: a first channel and a second channel; the first channel includes inputting the historical data of the power company into the LSTM layer to extract the feature of the historical data of the power company and the time-series information of the historical data of the power company; the second channel includes inputting the historical model snapshot into the LSTM layer to perform time-series modeling to capture the impact of the prediction model parameters of the historical mining algorithm on the task performance at different training stages, record and update each set of hyperparameter settings of the historical mining algorithm prediction model and its characteristics shown during the training process, and learn the long-term dependencies in the historical mining algorithm prediction task; Through calculating the confidence of the meta-model, perform uncertainty analysis on the prediction result output by the meta-model. When the confidence judgment result indicates that the meta-model prediction result is uncertain, start a joint optimization framework based on reinforcement learning and differential evolution to perform hyperparameter optimization; the joint optimization framework based on reinforcement learning and differential evolution is to establish a reinforcement learning-differential evolution collaborative controller, transfer the decision-making information of reinforcement learning to the differential evolution algorithm, and configure the strategy of reinforcement learning; Feed back the hyperparameters optimized by the joint optimization framework based on reinforcement learning and differential evolution to the meta-model, select the optimal mining algorithm, and execute the data mining task to support the intelligent decision-making applications of load prediction, customer demand analysis, and equipment failure detection in the power system.
2. The intelligent data mining method based on big data analysis according to claim 1, characterized in that: Perform preprocessing and feature extraction on the historical data; Use the data after feature extraction for training and parameter setting. During the training process, record the historical mining task data and the historical model snapshot; The historical mining task data is a detailed record of past data mining tasks, including: the mining algorithms used before, hyperparameter settings, and historical mining algorithm prediction results; The historical model snapshot includes the parameters of the historical mining algorithm prediction model, network structure parameters, hyperparameters of the optimization algorithm, training strategy, and the selection of the loss function.
3. The intelligent data mining method based on big data analysis according to claim 2, characterized in that: The meta-model includes: an input layer, a dual-channel architecture layer, a merging layer, a fully connected layer, a transfer decision module layer, and an output decision layer; The input layer includes: using the historical data of the power company as the data input.
4. The intelligent data mining method based on big data analysis according to claim 3, characterized in that: The merging layer includes: concatenating and merging the output information from the dual-channel architecture to form a high-dimensional feature representation; The fully connected layer includes using the merged high-dimensional features as the input to perform a non-linear transformation to generate an intermediate result; The migration decision module layer includes initializing the meta-model using the historical mining algorithm prediction model parameters and performing fine-tuning; during fine-tuning, it is updated using the difference between the output of the meta-model and the historical mining algorithm prediction model parameters. The output decision layer includes making a final decision based on the intermediate results output by the fully connected layer and combining confidence judgment.
5. The intelligent data mining method based on big data analysis according to claim 4, characterized in that: The confidence of the meta-model includes: calculating the entropy value of the meta-model; when the entropy value of the meta-model is less than or equal to the entropy threshold, it is considered that the prediction result of the meta-model is determined, and the prediction result is directly executed. When the entropy value of the meta-model is greater than the entropy threshold, calculate the similarity of the meta-model; when the similarity of the meta-model is greater than the similarity threshold, it is considered that the prediction result of the meta-model is determined, and the prediction result is directly executed. The calculation of the similarity of the meta-model includes: calculating the feature similarity and calculating the context similarity. The feature similarity is calculated by the weighted Jaccard coefficient and the improved DTW distance; the context similarity is calculated by considering the semantic matching degree of the business scenario label based on the Word2Vec word vector. When the meta-model similarity is less than or equal to the similarity threshold, it is considered that the prediction result of the meta-model is uncertain, and the joint optimization framework based on reinforcement learning and differential evolution is started for hyperparameter optimization.
6. The intelligent data mining method based on big data analysis according to claim 5, wherein: The joint optimization framework based on reinforcement learning and differential evolution includes: loading the historical data of the power company into the experience pool of reinforcement learning and setting resource constraint conditions. The resource constraint conditions include: Condition 1, stop iterating when the maximum number of iterations is reached. Condition 2, stop iterating when the time budget is reached. If any one of Condition 1 or Condition 2 is satisfied, stop iterating. The joint optimization framework based on reinforcement learning and differential evolution also includes: extracting n-dimensional statistical features from the processed dataset, obtaining the algorithm characteristics recommended by the meta-model, finding the Top-K similar scenarios in the reinforcement learning experience pool, calculating the scene similarity weight based on the Mahalanobis distance in the feature space, and deriving the recommended confidence interval of the parameter space of the processed dataset according to the optimal parameter distribution of the Top-k similar scenarios. The joint optimization framework based on reinforcement learning and differential evolution further includes: constructing a differential evolution population containing x hyperparameter individuals, The hyperparameter individuals are generated within the range recommended by reinforcement learning, For the hyperparameter individuals, random sampling is performed in the parameter space for processing the data set; Perform hybrid encoding on the hyperparameter individuals; the hybrid encoding is to use real number encoding for continuous parameter individuals and Gray code encoding for discrete parameter individuals to reduce the mutation probability of adjacent values, and introduce the positive correlation constraint between batch_size and learning_rate. Calculate the Euclidean distance between each individual and the current optimal solution, and calculate the difference degree between individuals. Set a dynamic difference degree threshold, and when the distance between an individual and the current optimal solution is less than or equal to the dynamic threshold, perform fine search. The fine search is that the mutation step size range is P1; each individual performs a direction exploration and fine adjustment. When the distance between an individual and the current optimal solution exceeds the dynamic threshold, perform global search. The global search is that the mutation step size range is P2; retain the optimal gene segments of the previous b generations for crossover operations. Apply reverse perturbation to the individuals that have not been improved for b consecutive generations to prevent falling into local optima. Perform mirror mapping on the individuals close to the constraint boundary to adjust their exploration range.
7. The intelligent data mining method based on big data analysis according to claim 6, wherein: The joint optimization framework based on reinforcement learning and differential evolution further includes: after every b generations of differential evolution are completed, the population diversity coefficient is monitored in real time; when the population diversity coefficient is greater than the diversity threshold, a reinforcement learning evaluation is triggered; According to the optimal performance, narrow down the parameter range and retain the parameter interval; Extend according to the gradient direction the parameter space and expand the exploration area; A reward mechanism is set to quantify the benefits of parameter adjustment in each optimization, and the reinforcement learning strategy is dynamically adjusted; The reward mechanism is to calculate the accuracy increment of the meta-model and the time consumed after each iteration, measure the benefit ratio through the accuracy increment and the time consumed, and use the benefit ratio as a reward signal to feedback to the joint optimization framework of reinforcement learning and differential evolution to update the Q value of reinforcement learning; Reserve the top individuals of each generation and store them in the elite pool; Elite gene recombination is performed every a generations, and a weighted crossover strategy based on parameter sensitivity is adopted, and high-sensitivity parameters are preferentially retained; The convergence determination conditions include: when the improvement rate of the optimal solution for a consecutive a generations is less than the improvement rate threshold, it is determined to converge; When the population gene similarity exceeds the population gene similarity threshold, it is determined to converge; When the remaining time budget is less than the time budget threshold of the time budget, the current optimal solution is forced to be output, and the optimization process is stopped.
8. An intelligent data mining platform based on big data analysis using the method according to any one of claims 1-7, characterized in that: A data collection and preprocessing module, configured to obtain data to be processed from the data sources of the power company and perform preprocessing to generate a processed data set for data mining; the data to be processed includes structured data, unstructured data, and time-series data; the unstructured data includes text, pictures, and audio data; A prediction module, configured to construct a meta-model based on the historical load data, historical customer demand data, and historical equipment failure data of the power company, and the meta-model adopts a dual-channel architecture to predict the optimal data mining algorithm; An optimization module, configured to perform uncertainty analysis on the prediction results output by the meta-model according to the confidence level output by the meta-model; when the confidence level judgment result indicates that the prediction result of the meta-model is uncertain, automatically call the optimization framework combined with reinforcement learning and differential evolution to dynamically adjust and optimize the hyperparameters, and feedback the optimization results to the meta-model to improve the model's adaptive ability; An execution module, configured to automatically select the optimal mining algorithm and execute the data mining task according to the output of the optimized meta-model, and realize intelligent decision-making and business support for load prediction, customer demand analysis, and equipment failure detection in the power system.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, the steps of the intelligent data mining method based on big data analysis according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the steps of the intelligent data mining method based on big data analysis according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Product change path multi-objective optimization method fusing reinforcement learning and differential evolution
CN116451577A
Reservoir level prediction method based on multi-source data fusion integration
CN117973600A