Intelligent data mining platform and method based on big data analysis
By building a metamodel based on LSTM network and a joint optimization framework combining reinforcement learning and differential evolution, the problem of incoordination of model selection bias and hyperparameter optimization in big data mining is solved, and more efficient and accurate data mining effects are achieved.
Patent Information
- Application Number
- CN202510449839.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-11
AI Technical Summary
There are problems in existing big data mining methods such as model selection bias, insufficient uncertainty evaluation, and incoordinated hyperparameter optimization.
A smart data mining method based on big data analysis is proposed. By building a meta-model based on LSTM network, predicting the optimal mining algorithm, and combining a joint optimization framework of reinforcement learning and differential evolution, hyperparameter optimization is carried out, and optimization paths are adjusted in real time.
It improves the accuracy and efficiency of data mining, can cope with data uncertainty and complexity, and provides global optimal solutions, which has important practical value.
Smart Images

Figure CN119961493A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data mining and intelligent analysis, and specifically to an intelligent data mining platform and method based on big data analysis. Background Art
[0002] With the rapid development of information technology, big data analysis technology has become the core support of all walks of life. Big data analysis technology mainly relies on data mining algorithms to extract valuable information from massive data. In recent years, data mining technology has been widely used, especially in the financial, medical, power, retail and other industries, and its application scenarios involve prediction, classification, clustering, association rule discovery and other fields. In order to solve the high-dimensional and complex problems in big data, more and more research has begun to focus on the intelligent and automated selection of algorithms, especially research based on meta-learning and adaptive algorithm optimization framework has become a hot direction. Meta-learning technology can predict the optimal algorithm based on historical data, and further optimize it by combining deep learning, reinforcement learning and other technologies, which makes the data mining process more efficient and accurate.
[0003] However, the existing algorithm selection methods based on meta-learning still have some obvious shortcomings. First, most meta-learning methods rely on the traditional single algorithm selection model, ignoring the dynamic changes and diversity of data, resulting in the inability to fully consider various potential algorithm combinations on complex high-dimensional data sets and the inability to adapt to changes in different data sources. Secondly, the reinforcement learning and differential evolution (DE) optimization frameworks in the existing technology mostly operate independently and fail to achieve effective synergy. Although reinforcement learning can perform global optimization and differential evolution can perform local search, the combination of the two still has considerable challenges, especially in the process of hyperparameter optimization, which often fails to achieve global optimality. Furthermore, most of the existing meta-models lack effective evaluation of the uncertainty of the prediction results, and fail to monitor and adjust the optimization path in real time, which makes the model have high uncertainty when dealing with unknown data. Summary of the invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problems solved by the present invention are: the problems of model selection bias, insufficient uncertainty assessment and incoordination of hyperparameter optimization in existing big data mining methods.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: an intelligent data mining method based on big data analysis, comprising: Obtain the data to be processed from the data source, perform preprocessing, and generate a processed data set for data mining; Based on historical data, build a meta-model to predict the optimal mining algorithm; The meta-model includes: using a recurrent neural network to construct a meta-model, establishing a dual-channel architecture through an LSTM network and historical data, and merging the outputs in series to obtain the final meta-model prediction result; the dual-channel architecture includes: a first channel and a second channel; the first channel includes: inputting historical data into the LSTM layer, extracting historical data features and time series information of historical data; the second channel includes: inputting historical model snapshots into the LSTM layer, performing time series modeling to capture the impact of historical mining algorithm prediction model parameters on task performance at different training stages, recording and updating each set of hyperparameter settings of the historical mining algorithm prediction model and its characteristics exhibited during the training process, and learning long-term dependencies in historical mining algorithm prediction tasks; By calculating the confidence of the meta-model, uncertainty analysis is performed on the prediction results output by the meta-model. When the confidence judgment result indicates that the prediction results of the meta-model are uncertain, a joint optimization framework based on reinforcement learning and differential evolution is started to optimize hyperparameters. The joint optimization framework based on reinforcement learning and differential evolution is to establish a reinforcement learning-differential evolution collaborative controller, pass the decision information of reinforcement learning to the differential evolution algorithm, and configure the strategy of reinforcement learning. The hyperparameters optimized by the joint optimization framework based on reinforcement learning and differential evolution are fed back to the meta-model, the optimal mining algorithm is selected, and the data mining task is performed.
[0007] As a preferred solution of the intelligent data mining method based on big data analysis described in the present invention, wherein: the data to be processed includes structured data, unstructured data and time series data; The historical data includes historical mining task data and historical model snapshots: The historical mining task data is a detailed record of past data mining tasks, including: previously used mining algorithms, hyperparameter settings, and historical mining algorithm prediction results; The historical model snapshot includes parameters of the historical mining algorithm prediction model, network structure parameters, optimization algorithm hyperparameters, training strategies and loss function selection.
[0008] As a preferred solution of the intelligent data mining method based on big data analysis described in the present invention, wherein: the meta-model includes: an input layer, a dual-channel architecture layer, a merging layer, a fully connected layer, a migration decision module layer, and an output decision layer; The input layer includes: taking historical data and historical model snapshots as data input.
[0009] As a preferred solution of the intelligent data mining method based on big data analysis described in the present invention, wherein: the merging layer comprises: merging the output information from the dual-channel architecture in series to form a high-dimensional feature representation; The fully connected layer includes taking the merged high-dimensional features as input, performing nonlinear transformation, and generating intermediate results; The migration decision module layer includes initializing the meta-model using the historical mining algorithm prediction model parameters and fine-tuning the meta-model; during fine-tuning, the difference between the meta-model output and the historical mining algorithm prediction model parameters is used for updating; The output decision layer includes making a final decision based on the intermediate results output by the fully connected layer and combining the confidence judgment.
[0010] As a preferred solution of the intelligent data mining method based on big data analysis described in the present invention, the confidence of the metamodel includes: calculating the entropy value of the metamodel; when the entropy value of the metamodel is less than or equal to the entropy threshold, the prediction result of the metamodel is considered to be determined, and the prediction result is directly executed; When the entropy value of the metamodel is greater than the entropy threshold, the similarity of the metamodel is calculated; when the similarity of the metamodel is greater than the similarity threshold, the prediction result of the metamodel is considered to be certain, and the prediction result is directly executed; The calculating the similarity of the meta-model includes: calculating feature similarity and calculating context similarity; The feature similarity is calculated by weighted Jaccard coefficient and improved DTW distance. The context similarity is calculated by considering the semantic matching of business scenario labels based on Word2Vec word vectors. When the meta-model similarity is less than or equal to the similarity threshold, the prediction result of the meta-model is considered uncertain, and the joint optimization framework based on reinforcement learning and differential evolution is started to optimize the hyperparameters.
[0011] As a preferred solution of the intelligent data mining method based on big data analysis described in the present invention, wherein: the joint optimization framework based on reinforcement learning and differential evolution includes: loading historical data into the experience pool of reinforcement learning, setting resource constraints; Resource constraints include: Condition 1: When the maximum number of iterations is reached, the iteration is stopped; Condition 2: When the time budget is reached, stop the iteration; If any one of condition 1 or condition 2 is met, the iteration stops; The joint optimization framework based on reinforcement learning and differential evolution also includes: extracting n-dimensional statistical features from the processing data set, obtaining algorithm characteristics recommended by the meta-model, searching for Top-k similar scenes in the reinforcement learning experience pool, calculating scene similarity weights based on the Mahalanobis distance in the feature space, and deriving a recommended confidence interval of the parameter space of the processing data set according to the optimal parameter distribution of the Top-k similar scenes; The joint optimization framework based on reinforcement learning and differential evolution also includes: constructing a differential evolution population including x hyperparameter individuals, Hyperparameter individuals are generated within the range recommended by reinforcement learning, Each hyperparameter individual uses random sampling of the parameter space of the processing dataset; Hybrid coding is performed on hyperparameter individuals; hybrid coding uses real number coding for continuous parameter individuals and Gray code coding for discrete parameter individuals to reduce the probability of mutation of adjacent values and introduce a positive correlation constraint between batch_size and learning_rate; Calculate the Euclidean distance between each individual and the current optimal solution, and calculate the difference between individuals; Set a dynamic difference threshold. When the distance between the individual and the current optimal solution is less than or equal to the dynamic threshold, perform a fine search. The fine search is to change the step length range to P1; each individual conducts a direction exploration and fine adjustment; When the distance between an individual and the current optimal solution exceeds a dynamic threshold, a global search is performed; The global search is to change the asynchronous length range to P2; retain the optimal gene fragments of the previous b generations for crossover operation; Apply reverse perturbations to individuals that have not improved for consecutive b generations to prevent them from falling into local optimality; Mirror images are performed on individuals close to the constraint boundary to adjust their exploration range.
[0012] As a preferred solution of the intelligent data mining method based on big data analysis described in the present invention, the joint optimization framework based on reinforcement learning and differential evolution also includes: after each b-generation differential evolution is completed, the population diversity coefficient is monitored in real time; when the population diversity coefficient is greater than the diversity threshold, the reinforcement learning evaluation is triggered; Based on the best performance, narrow the parameter range and keep the best performance Parameter interval; Extend according to the gradient direction parameter space, expanding the exploration area; Set up a reward mechanism, quantify the benefits of parameter adjustments in each optimization, and dynamically adjust the reinforcement learning strategy; The reward mechanism is to calculate the accuracy increment and the time consumed of the meta-model after each iteration, measure the benefit ratio by the accuracy increment and the time consumed, and feed the benefit ratio as a reward signal to the joint optimization framework of reinforcement learning and differential evolution to update the Q value of reinforcement learning; Keep the best previous individuals and deposit them into the elite pool; Elite gene recombination is performed every a generation, using a weighted crossover strategy based on parameter sensitivity, with high-sensitivity parameters being retained first; The convergence judgment conditions include: when the improvement rate of the optimal solution of consecutive a generations is less than the improvement rate threshold, it is judged to be converged; When the population gene similarity exceeds the population gene similarity threshold, it is judged as convergence; When the remaining time budget is less than the time budget threshold of the time budget, the current optimal solution is forced to be output and the optimization process is stopped.
[0013] An intelligent data mining platform for big data analysis, wherein: The data module obtains the data to be processed from the data source, performs preprocessing, and generates a processed data set for data mining; The prediction module builds a meta-model based on historical data and predicts the optimal mining algorithm; The optimization module calculates the confidence of the meta-model and performs uncertainty analysis on the prediction results output by the meta-model. When the confidence judgment result indicates that the prediction results of the meta-model are uncertain, the joint optimization framework based on reinforcement learning and differential evolution is started to optimize the hyperparameters. The execution module feeds back the hyperparameters optimized by the joint optimization framework based on reinforcement learning and differential evolution to the meta-model, selects the optimal mining algorithm, and executes the data mining task.
[0014] A computer device comprises: a memory and a processor; the memory stores a computer program, wherein the processor implements the steps of any one of the methods of the present invention when executing the computer program.
[0015] A computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of any one of the methods of the present invention.
[0016] Beneficial effects of the present invention: The intelligent data mining method based on big data analysis provided by the present invention improves data quality in the data preprocessing stage and provides reliable input for subsequent mining. By constructing a meta-model based on the LSTM network, the optimal mining algorithm is automatically predicted, avoiding the deviation of the manual selection algorithm in the traditional method. The confidence analysis mechanism of the meta-model output result can judge the reliability of the prediction result. When the confidence is low, the joint optimization framework of reinforcement learning and differential evolution is started to optimize the hyperparameters, thereby improving the performance and accuracy of the mining algorithm. The optimized hyperparameters are fed back to the meta-model to ensure the selection of the optimal algorithm and the execution of the data mining task, thereby improving the mining efficiency and accuracy. It not only improves the accuracy and efficiency of data mining, but also can cope with the uncertainty and complexity of data, while ensuring the computational efficiency, providing a global optimal solution, and has important practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0018] Figure 1 An overall flow chart of an intelligent data mining method based on big data analysis provided for the first embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.
[0020] Example 1, reference Figure 1 , as an embodiment of the present invention, provides an intelligent data mining method based on big data analysis, comprising: S1: Obtain the data to be processed from the data source, perform preprocessing, and generate a processed data set for data mining.
[0021] The data to be processed includes structured data, unstructured data, and time-series data; the historical data includes historical mining task data and historical model snapshots: the historical mining task data is a detailed record of past data mining tasks, including: the mining algorithms used previously, hyperparameter settings, and historical mining algorithm prediction results; the historical model snapshots include the parameters of the historical mining algorithm prediction model, network structure parameters, hyperparameters of the optimization algorithm, training strategies, and the selection of loss functions.
[0022] The preprocessing of structured data includes, for missing fields, filling them with the mean value, using statistical methods to identify and handle outliers in the data. Compressing the data into the interval [0,1] through normalization, making the scales of each feature consistent, and facilitating subsequent algorithm processing.
[0023] Unstructured data usually includes text, pictures, and audio data. The preprocessing steps are as follows: The preprocessing of text data includes: disassembling the text into lexical units, using a tokenization tool (NLTK) to perform word segmentation, removing meaningless words such as "de", "he", etc., and punctuation marks in the text.还原 the words to their roots. Converting the text into a numerical vector using the TF-IDF method.
[0024] The preprocessing of picture data includes: cropping the image as needed to keep the image size consistent, scaling the pixel values into the interval [0,1], or normalizing the image by its mean and standard deviation. Increasing the diversity of training data through rotation, flipping, and scaling methods to avoid overfitting.
[0025] The preprocessing of time-series data includes: for missing values in time-series data, often using forward filling to fill them.
[0026] Using the moving average method to remove noise in time-series data. Dividing the time-series data into multiple windows according to time periods, and extracting the trends and seasonal features in the time-series data.
[0027] Historical data includes historical mining task data and historical model snapshots. The preprocessing steps include: Cleaning up errors or redundant information in the historical data to ensure the accuracy of the data. Integrating the historical task data and historical model snapshots, and associating the historical data through the timestamp and task ID fields. Extracting relevant features from the historical task data, including mining algorithms and hyperparameter settings; extracting network structure parameters, hyperparameters of the optimization algorithm, etc. from the historical model snapshots.
[0028] Generating a processed data set for data mining based on the preprocessed data: Dividing the data into a training set, a validation set, and a test set according to the ratio of 8:1:1 to ensure the fairness of model training and evaluation.
[0029] Furthermore, by preprocessing structured data, unstructured data, and time series data, we can solve problems such as missing values, outliers, and noise in the data, and improve the consistency and reliability of the data. At the same time, we use methods such as normalization and TF-IDF to standardize and transform the data to ensure that all types of data can be processed on the same scale, providing a more effective training basis for subsequent algorithm models.
[0030] Furthermore, by cleaning and integrating historical data, we can not only improve the accuracy of the data, but also help the algorithm better understand past tasks and model performance by extracting key features. This process ensures that the model has a reasonable starting point when training, and can be evaluated fairly and effectively based on different data sets (such as training sets, validation sets, and test sets), thereby improving the generalization ability and prediction accuracy of the model.
[0031] S2: Based on historical data, build a meta-model and predict the optimal mining algorithm.
[0032] The meta-model includes building a meta-model using a recurrent neural network, establishing a dual-channel architecture through an LSTM network and historical data, and merging the outputs in series to obtain the final meta-model prediction result.
[0033] The meta-model also includes: an input layer, a dual-channel architecture layer, a merging layer, a fully connected layer, a migration decision module layer, and an output decision layer.
[0034] The input layer includes: taking historical data and historical model snapshots as data input.
[0035] The dual-channel architecture includes: a first channel and a second channel.
[0036] The first channel includes,inputting historical data into the LSTM layer, extracting historical data features and historical data time series information; The formula is: ; in, For the historical data input, is the output historical data feature, is the weight of the LSTM network.
[0037] The second channel includes inputting the historical model snapshot into the LSTM layer, performing time series modeling to capture the impact of the historical mining algorithm prediction model parameters on the task performance at different training stages, recording and updating each set of hyperparameter settings of the historical mining algorithm prediction model and its characteristics exhibited during the training process, and learning the long-term dependencies in the historical mining algorithm prediction tasks.
[0038] The formula is: ; in, is the input historical model snapshot data, is the output historical model snapshot feature, is the weight of the LSTM network.
[0039] The merging layer includes: merging the output information from the dual-channel architecture in series to form a high-dimensional feature representation.
[0040] The formula is: ; in, is the high-dimensional feature representation after merging, Indicates that the output features of the two channels are concatenated.
[0041] The fully connected layer includes taking the merged high-dimensional features as input, performing nonlinear transformation, and generating intermediate results.
[0042] The formula is: ; in, is the output intermediate result of the fully connected layer, and are the weight matrix and bias term of the fully connected layer, respectively, and ReLU(·) is the activation function.
[0043] The migration decision module layer includes initializing the meta-model using the historical mining algorithm prediction model parameters and fine-tuning the meta-model; during fine-tuning, the difference between the meta-model output and the historical mining algorithm prediction model parameters is used for updating; The output decision layer includes making a final decision based on the intermediate results output by the fully connected layer and combining the confidence judgment.
[0044] The formula is: ; in, Indicates the final decision result, indicating the predicted mining algorithm selection result, represents the intermediate result from the fully connected layer, represents the weight matrix of the fully connected layer, represents the bias of the fully connected layer, It is used to convert the final decision result into a probability distribution. Represents the final value obtained through the confidence calculation.
[0045] The confidence level of the metamodel includes: calculating the entropy value of the metamodel; when the entropy value of the metamodel , the prediction result of the metamodel is considered to be certain and the prediction result is directly executed.
[0046] The formula for calculating the entropy value of the meta-model is expressed as: ; in, Indicates the model output category The probability distribution of Represents the entropy value of the metamodel.
[0047] The entropy threshold and similarity threshold are set based on historical experience.
[0048] When the entropy value of the metamodel is greater than the entropy threshold, the similarity of the metamodel is calculated; when the similarity of the metamodel exceeds the similarity threshold, the prediction result of the metamodel is considered to be certain and the prediction result is directly executed.
[0049] Calculating metamodel similarity includes: calculating feature similarity and calculating context similarity.
[0050] The feature similarity is calculated by weighted Jaccard coefficient and improved DTW distance, and the formula is expressed as: ; in, Represents the characteristics of the current input, represents the historical data characteristics, DTW represents the dynamic time warping distance, Indicates feature similarity.
[0051] The context similarity is calculated by considering the semantic matching of business scenario labels based on Word2Vec word vectors. The formula is expressed as: ; in, is the business scenario label word vector of the current task, is the label word vector of the historical task, Represents context similarity.
[0052] The final similarity formula is expressed as: Final similarity: ; in, represents the final similarity, represents the context similarity, Indicates feature similarity.
[0053] When similarity When , the prediction results of the meta-model are considered uncertain, and the joint optimization framework based on reinforcement learning and differential evolution is started to optimize the hyperparameters.
[0054] The confidence formula is expressed as: ; in, represents the entropy value of the metamodel, represents the final similarity, Indicates confidence.
[0055] Furthermore, by combining the recurrent neural network (LSTM) with historical data and historical model snapshots, an efficient meta-model is constructed that can accurately capture and predict the performance of the mining algorithm at different training stages. Through the dual-channel architecture, the system can extract historical data features and historical model snapshot features respectively, and fuse them into high-dimensional feature representations, thereby achieving accurate decision support for the mining algorithm. The addition of the fully connected layer and the migration decision module further improves the accuracy and flexibility of the meta-model, enabling it to quickly make optimization decisions based on historical data and historical model snapshots.
[0056] Furthermore, through the confidence calculation mechanism, a flexible decision-making method is provided. When the meta-model prediction result has a high degree of certainty, the prediction result can be directly executed to improve the calculation efficiency; when the confidence is low, the system will start the joint optimization framework of reinforcement learning and differential evolution to further optimize the hyperparameters, thereby improving the accuracy and performance of the mining algorithm. By dynamically adjusting the confidence, the system can adapt to different task scenarios to ensure that each decision is data-driven and optimal, thereby achieving the effect of optimizing model performance.
[0057] S3: By calculating the confidence of the meta-model, uncertainty analysis is performed on the prediction results output by the meta-model. When the confidence judgment result shows that the prediction result of the meta-model is uncertain, a joint optimization framework based on reinforcement learning and differential evolution is started to perform hyperparameter optimization.
[0058] The joint optimization framework based on reinforcement learning and differential evolution includes: establishing a reinforcement learning-differential evolution collaborative controller, transmitting the decision information of reinforcement learning to the differential evolution algorithm, and configuring the strategy of reinforcement learning to ensure that it dynamically affects the search space of differential evolution in each round of optimization; Load historical data into the experience pool of reinforcement learning and set resource constraints; The resource constraints include: Condition 1: When the maximum number of iterations reaches 1000, the iteration is stopped; Condition 2: When the time budget of 3 hours is reached, the iteration is stopped; If any one of condition 1 or condition 2 is met, the iteration stops; The joint optimization framework based on reinforcement learning and differential evolution also includes: extracting n-dimensional statistical features from the processing data set, obtaining the algorithm characteristics recommended by the meta-model, searching for Top-K similar scenes in the reinforcement learning experience pool, calculating the scene similarity weight based on the feature space Mahalanobis distance, and deriving the recommended confidence interval of the parameter space of the processing data set according to the optimal parameter distribution of the Top-K similar scenes. The formula is expressed as follows: ; ; ; in, Represents the feature vector of the current data, Represents historical data, represents the Euclidean distance, Indicates that in all historical scenarios, scenes with characteristics most similar to the current scene.
[0059] The 10-dimensional statistical features include mean, variance, standard deviation, skewness, kurtosis, maximum value, minimum value, median, interquartile range, and covariance.
[0060] Based on historical experience .
[0061] The joint optimization framework based on reinforcement learning and differential evolution also includes: constructing a differential evolution population including 100 hyperparameter individuals, 80 hyperparameter individuals are generated within the range recommended by reinforcement learning, and 20 hyperparameter individuals are randomly sampled using the parameter space of the processing data set; Hybrid coding is performed on hyperparameter individuals; hybrid coding uses real number coding for continuous parameter individuals and Gray code coding for discrete parameter individuals to reduce the probability of mutation of adjacent values and introduce a positive correlation constraint between batch_size and learning_rate; Calculate the Euclidean distance between each individual and the current optimal solution, and calculate the difference between individuals; A dynamic difference threshold is set, and when the distance between the individual and the current optimal solution is less than or equal to the dynamic threshold, a fine search is performed.
[0062] The formula of dynamic threshold is expressed as: ; Among them, σ represents the adjustment factor; ε represents the error tolerance, Represents the Euclidean distance between the i-th individual and the current optimal solution; Indicates the variable step length range during fine search; Indicates the current asynchronous duration; Indicates the maximum value of the variable asynchronous length.
[0063] The fine search is performed with the step length range being P1.
[0064] P1 is 0.1 to 0.3.
[0065] Each individual conducts a=5 direction explorations and fine-tunes, and the formula is expressed as: ; in, Indicates the adjustment of each exploration direction. Indicates new directions for exploration. Indicates the original exploration direction.
[0066] When the distance between an individual and the current optimal solution exceeds a dynamic threshold, a global search is performed.
[0067] The global search is performed with the variable step length range being P2.
[0068] P2 is 0.7 to 1.2.
[0069] The best gene fragments of the previous b=3 generations are retained for crossover operation; Reverse perturbations are applied to individuals that have not been improved for consecutive b=3 generations to prevent them from falling into the local optimum. The formula is expressed as: ; in, represents the disturbance factor, represents the i-th individual after update, represents the current value of the i-th individual, represents the perturbation vector.
[0070] Mirror mapping is performed on individuals close to the constraint boundary to adjust their exploration range. The formula is expressed as: ; Among them, it means The current parameter value of each individual, represents the constraint bounds of the parameters, Indicates that and Compare and return the smaller value; It means comparing the previous minimum value with zero and returning the larger value; Indicates the adjusted Individual parameter values.
[0071] After completing b=3 generations of differential evolution, the population diversity coefficient is monitored in real time; when the population diversity coefficient is greater than the diversity threshold, reinforcement learning evaluation is triggered; the diversity threshold is determined based on experience.
[0072] Based on the best performance, narrow the parameter range and keep the best performance Parameter interval; .
[0073] Extend according to the gradient direction parameter space, expanding the exploration area; .
[0074] Set up a reward mechanism, quantify the benefits of parameter adjustments in each optimization, and dynamically adjust the reinforcement learning strategy; The reward mechanism is to calculate the accuracy increment and the time consumed of the meta-model after each iteration, measure the benefit ratio by the accuracy increment and the time consumed, and feed the benefit ratio as a reward signal to the joint optimization framework of reinforcement learning and differential evolution to update the Q value of reinforcement learning.
[0075] The benefit ratio formula is: ; in, represents the accuracy increment, Indicates the time consumed.
[0076] Update the Q value of reinforcement learning, the formula is expressed as: ;
[0077] in, Indicates the current state Next, take action Q value; represents the learning rate; Indicates immediate reward; represents the discount factor; In the new state All possible actions The maximum Q value of Indicates taking action The new state to which it is transferred later; In the new state The following possible actions can be taken.
[0078] Keep the best previous individuals and deposit them into the elite pool; .
[0079] Elite gene recombination is performed every a=5 generations, using a weighted crossover strategy based on parameter sensitivity, with high-sensitivity parameters being retained first. The formula is: ;
[0080] in, Represents the newly generated offspring individual; represents parent individual 1 and represents parent individual 2, Represents the parameter sensitivity weight.
[0081] The convergence judgment conditions include: when the improvement rate of the optimal solution of a = 5 consecutive generations reaches the improvement rate threshold, it is judged to be converged; When the population gene similarity exceeds the population gene similarity threshold, it is judged as convergence; When the remaining time budget is less than the time budget threshold of the time budget, the current optimal solution is forced to be output and the optimization process is stopped.
[0082] The improvement rate threshold, population gene similarity threshold and time budget threshold are set based on historical experience.
[0083] Furthermore, by combining the advantages of reinforcement learning and differential evolution algorithms, the strategy of reinforcement learning is fully utilized to guide the search process of differential evolution, thereby effectively optimizing the hyperparameter search space. In this framework, reinforcement learning can dynamically adjust the optimization strategy and provide accurate parameter recommendations based on historical data and feature analysis, while differential evolution achieves efficient hyperparameter exploration through global and local search mechanisms. This collaborative control strategy ensures that during the optimization process, the optimal solution can be quickly searched while avoiding falling into the local optimum, thereby improving the overall optimization effect.
[0084] Furthermore, by dynamically adjusting and monitoring the optimization process of parameters in real time, the adaptability and efficiency of the algorithm in complex environments are significantly improved. The set dynamic threshold, fine search and global search mechanism enable the algorithm to flexibly adjust the search strategy under different conditions to avoid premature convergence. At the same time, through the calculation of the reward mechanism and benefit ratio of reinforcement learning, the optimization effect can be quantified and the strategy can be adjusted in real time, further improving the stability and accuracy of the optimization process.
[0085] S4: Feedback the hyperparameters optimized by the joint optimization framework based on reinforcement learning and differential evolution to the meta-model, select the optimal mining algorithm, and perform the data mining task.
[0086] The optimized hyperparameter set is passed from the joint optimization framework to the meta-model to ensure that the hyperparameter set can effectively improve the performance of the meta-model. The meta-model adjusts its internal structure based on the hyperparameter set to improve its prediction ability.
[0087] In data mining tasks, the meta-model selects the most suitable mining algorithm for the current task based on the distribution of historical data and feature space.
[0088] Use the optimized algorithm to perform mining tasks on the input data set and output the final mining results.
[0089] Furthermore, by passing the hyperparameter set optimized by the joint optimization framework based on reinforcement learning and differential evolution to the meta-model, it can be ensured that the meta-model can use the best hyperparameter configuration, thereby effectively improving its performance. The optimized hyperparameters can guide the meta-model to adjust its own structure, so that the model can flexibly adapt and improve prediction accuracy when facing different data sets and tasks. This hyperparameter tuning process avoids the limitations of relying on fixed parameters and gives the meta-model higher adaptability and stronger generalization capabilities.
[0090] Furthermore, by using the optimized hyperparameter set, the meta-model can automatically select the most suitable mining algorithm for the current task based on the distribution of historical data and feature space, and perform data mining tasks on this basis. The final output mining results are more accurate and practical, ensuring that the model can make the best predictions and decisions for specific data sets and target tasks, thereby improving the overall system intelligence level and application effect.
[0091] Embodiment 2 is an embodiment of the present invention, which provides an intelligent data mining platform based on big data analysis, including: The data module obtains the data to be processed from the data source, performs preprocessing, and generates a processed data set for data mining.
[0092] The prediction module builds a meta-model based on historical data and predicts the optimal mining algorithm.
[0093] The optimization module calculates the confidence of the meta-model and performs uncertainty analysis on the prediction results output by the meta-model. When the confidence judgment result shows that the prediction result of the meta-model is uncertain, the joint optimization framework based on reinforcement learning and differential evolution is started to optimize the hyperparameters.
[0094] The execution module feeds back the hyperparameters optimized by the joint optimization framework based on reinforcement learning and differential evolution to the meta-model, selects the optimal mining algorithm, and executes the data mining task.
[0095] Embodiment 3, an embodiment of the present invention, is different from the first two embodiments in that: If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.
[0096] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0097] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0098] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0099] Example 4 is an embodiment of the present invention, which provides an intelligent data mining platform and method based on big data analysis. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through simulation experiments.
[0100] The experiment used historical load data, customer demand data, and equipment failure data from a power company, including structured data (such as historical load data tables), unstructured data (such as text descriptions of equipment failures), and time series data (such as daily load change records). The specific experimental steps are as follows: Data preprocessing: Three years of historical data were extracted from the power company database. After cleaning and preprocessing, the data contained missing values, outliers, and unstructured text data. Missing values were filled with mean values, outliers were corrected using a Z-score-based method, and unstructured text data was tokenized. Time series data was normalized using a sliding window method to ensure that the data met the requirements of model training.
[0101] Metamodel construction: A metamodel was constructed based on historical data. The metamodel uses a dual-channel architecture of LSTM (Long Short-Term Memory Network), where the first channel processes the time series characteristics of historical load data, and the second channel processes historical model snapshot information, extracting the hyperparameter settings of the model at different training stages and the characteristics exhibited during the training process. This dual-channel architecture can better capture the long-term dependencies in the data and the historical memory of the model.
[0102] Confidence analysis and optimization trigger: Based on the prediction results of the meta-model, the entropy value is calculated to evaluate the confidence of the model. If the entropy value is high, it means that the model has a large uncertainty in the prediction results, so the joint optimization framework based on reinforcement learning and differential evolution is initiated to optimize the hyperparameters. Specifically, the reinforcement learning algorithm adjusts the search space of hyperparameters based on the experience pool, while the differential evolution algorithm optimizes these hyperparameters through the population evolution strategy.
[0103] Hyperparameter optimization and data mining task execution: During the hyperparameter optimization process, each hyperparameter is optimized through a joint optimization framework of reinforcement learning and differential evolution. The optimized hyperparameter set is fed back to the meta-model and used to select the mining algorithm that best suits the current task. Finally, the optimized algorithm is used to perform data mining tasks, and the load forecasting, customer demand analysis, and equipment fault detection are predicted and the results are output.
[0104] During the experiment, the initial value, optimized value, prediction result, and related experimental time and accuracy improvement of each task were recorded. The following are the key data records of the experiment: Load forecast data: Under the initial conditions, the load forecast value was 456.78MW. After optimization, the forecast value was 459.30MW. The actual forecast result was 457.50MW. The experimental time was 3.2 hours, and the accuracy was improved by 5.26%.
[0105] Customer demand analysis: The initial value was 2456.60, the optimized value was 2468.85, the actual prediction result was 2472.10, the experimental time was 2.9 hours, and the accuracy was improved by 4.70%.
[0106] Equipment fault detection: The initial detection rate was 0.13, which was 0.10 after optimization, the prediction result was 0.08, the experimental time was 4.1 hours, and the accuracy was improved to 10.00%.
[0107] Hyperparameter settings (learning rate): The initial learning rate is 0.01, and the optimized learning rate is 0.015. The accuracy improvement is not directly calculated.
[0108] Hyperparameter setting (batchsize): The initial value is 64, the optimized value is 128, and the accuracy improvement is not directly calculated.
[0109] Through the above experimental data, we can clearly see the advantages and innovations of the method of the present invention in multiple tasks: After hyperparameter optimization, the accuracy of load forecasting increased by 5.26%. Traditional load forecasting methods usually rely on fixed hyperparameters, while the present invention dynamically adjusts hyperparameters through joint optimization based on reinforcement learning and differential evolution, so that the metamodel can better adapt to the changes and complexity in the data, thereby significantly improving the prediction accuracy.
[0110] By mining customer demand data, the accuracy is improved to 4.70%. Traditional methods often rely on simple algorithm selection and static parameters, while the meta-model of the present invention can combine historical data and model snapshots to dynamically select the most appropriate mining algorithm, thereby improving the accuracy of analysis.
[0111] The accuracy of equipment fault detection has been improved by 10.00%. The improvement in this task is particularly significant. Traditional methods often cannot adapt quickly when data is abnormal. However, this invention effectively avoids this limitation through a dynamically adjusted hyperparameter optimization framework and provides more accurate fault detection results.
[0112] The joint optimization of reinforcement learning and differential evolution algorithms significantly improves the performance of data mining tasks. By optimizing the learning rate (adjusted from 0.01 to 0.015) and the batch size (adjusted from 64 to 128), the training time of the model is shortened and the prediction accuracy is improved. This process shows that the hyperparameter optimization framework of the present invention has strong adaptability and can dynamically adjust parameters according to different data set requirements, thereby improving the robustness of the model.
[0113] Compared with the traditional static hyperparameter selection method, the method of the present invention automatically adjusts the hyperparameters in multiple rounds of training by introducing a joint optimization framework of reinforcement learning and differential evolution to complete the data mining task with the optimal configuration. Traditional methods are usually trained based on fixed parameters, while the dynamic adjustment mechanism of the present invention can automatically adapt according to actual data changes, improving the accuracy, efficiency and adaptability of data mining tasks.
[0114] Further experiments show that the joint optimization based on reinforcement learning and differential evolution can provide more accurate hyperparameter adjustment in different data tasks, especially when processing large-scale data. The optimized model can effectively improve the training speed and accuracy through a dynamic decision-making mechanism. Through multiple rounds of experimental verification, the inventive method has shown significant advantages in load forecasting, customer demand analysis, and equipment fault detection, especially in terms of accuracy and efficiency, which fully proves the innovation and practicality of the present invention.
[0115] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. An intelligent data mining method based on big data analysis, characterized in that: include: Obtain the data to be processed from the data source, perform preprocessing, and generate a processed data set for data mining; Based on historical data, build a meta-model to predict the optimal mining algorithm; The meta-model includes building a meta-model using a recurrent neural network, establishing a dual-channel architecture through an LSTM network and historical data, and combining and outputting in series to obtain a final meta-model prediction result; The dual-channel architecture includes: a first channel and a second channel; the first channel includes inputting historical data into the LSTM layer to extract historical data features and time series information of the historical data; the second channel includes inputting historical model snapshots into the LSTM layer to perform time series modeling to capture the impact of historical mining algorithm prediction model parameters on task performance at different training stages, record and update each set of hyperparameter settings of the historical mining algorithm prediction model and its characteristics exhibited during the training process, and learn long-term dependencies in historical mining algorithm prediction tasks; By calculating the confidence of the meta-model, uncertainty analysis is performed on the prediction results output by the meta-model. When the confidence judgment result indicates that the prediction results of the meta-model are uncertain, a joint optimization framework based on reinforcement learning and differential evolution is started to optimize hyperparameters. The joint optimization framework based on reinforcement learning and differential evolution is to establish a reinforcement learning-differential evolution collaborative controller, pass the decision information of reinforcement learning to the differential evolution algorithm, and configure the strategy of reinforcement learning. The hyperparameters optimized by the joint optimization framework based on reinforcement learning and differential evolution are fed back to the meta-model, the optimal mining algorithm is selected, and the data mining task is performed.
2. The intelligent data mining method based on big data analysis according to claim 1, characterized in that: The data to be processed includes structured data, unstructured data and time series data; The historical data includes historical mining task data and historical model snapshots: The historical mining task data is a detailed record of past data mining tasks, including: previously used mining algorithms, hyperparameter settings, and historical mining algorithm prediction results; The historical model snapshot includes parameters of the historical mining algorithm prediction model, network structure parameters, optimization algorithm hyperparameters, training strategies and loss function selection.
3. The intelligent data mining method based on big data analysis according to claim 2, characterized in that: The meta-model includes: an input layer, a dual-channel architecture layer, a merging layer, a fully connected layer, a migration decision module layer, and an output decision layer; The input layer includes: taking historical data and historical model snapshots as data input.
4. The intelligent data mining method based on big data analysis according to claim 3, characterized in that: The merging layer includes: merging the output information from the dual-channel architecture in series to form a high-dimensional feature representation; The fully connected layer includes taking the merged high-dimensional features as input, performing nonlinear transformation, and generating intermediate results; The migration decision module layer includes initializing the meta-model using the historical mining algorithm prediction model parameters and fine-tuning the meta-model; during fine-tuning, the difference between the meta-model output and the historical mining algorithm prediction model parameters is used for updating; The output decision layer includes making a final decision based on the intermediate results output by the fully connected layer and combining the confidence judgment.
5. The intelligent data mining method based on big data analysis according to claim 4, characterized in that: The confidence of the metamodel includes: calculating the entropy value of the metamodel; when the entropy value of the metamodel is less than or equal to the entropy threshold, the prediction result of the metamodel is considered to be certain, and the prediction result is directly executed; When the entropy value of the metamodel is greater than the entropy threshold, the similarity of the metamodel is calculated; when the similarity of the metamodel is greater than the similarity threshold, the prediction result of the metamodel is considered to be certain, and the prediction result is directly executed; The calculating the similarity of the meta-model includes: calculating feature similarity and calculating context similarity; The feature similarity is calculated by weighted Jaccard coefficient and improved DTW distance. The context similarity is calculated by considering the semantic matching of business scenario labels based on Word2Vec word vectors. When the meta-model similarity is less than or equal to the similarity threshold, the prediction result of the meta-model is considered uncertain, and the joint optimization framework based on reinforcement learning and differential evolution is started to optimize the hyperparameters.
6. The intelligent data mining method based on big data analysis according to claim 5, characterized in that: The joint optimization framework based on reinforcement learning and differential evolution includes: loading historical data into the experience pool of reinforcement learning and setting resource constraints; Resource constraints include: Condition 1: When the maximum number of iterations is reached, the iteration is stopped; Condition 2: When the time budget is reached, stop the iteration; If any one of condition 1 or condition 2 is met, the iteration stops; The joint optimization framework based on reinforcement learning and differential evolution also includes: extracting n-dimensional statistical features from the processing data set, obtaining algorithm characteristics recommended by the meta-model, searching for Top-K similar scenes in the reinforcement learning experience pool, calculating scene similarity weights based on the Mahalanobis distance in the feature space, and deriving a recommended confidence interval of the parameter space of the processing data set according to the optimal parameter distribution of the Top-k similar scenes; The joint optimization framework based on reinforcement learning and differential evolution also includes: constructing a differential evolution population including x hyperparameter individuals, Hyperparameter individuals are generated within the range recommended by reinforcement learning, Each hyperparameter individual uses random sampling of the parameter space of the processing dataset; Hybrid coding is performed on hyperparameter individuals; hybrid coding uses real number coding for continuous parameter individuals and Gray code coding for discrete parameter individuals to reduce the probability of mutation of adjacent values and introduce a positive correlation constraint between batch_size and learning_rate; Calculate the Euclidean distance between each individual and the current optimal solution, and calculate the difference between individuals; Set a dynamic difference threshold. When the distance between the individual and the current optimal solution is less than or equal to the dynamic threshold, perform a fine search. The fine search is to change the step length range to P1; each individual conducts a direction exploration and fine adjustment; When the distance between an individual and the current optimal solution exceeds a dynamic threshold, a global search is performed; The global search is to change the asynchronous length range to P2; retain the optimal gene fragments of the previous b generations for crossover operation; Apply reverse perturbations to individuals that have not improved for consecutive b generations to prevent them from falling into local optimality; Mirror images are performed on individuals close to the constraint boundary to adjust their exploration range.
7. The intelligent data mining method based on big data analysis according to claim 6, characterized in that: The joint optimization framework based on reinforcement learning and differential evolution also includes: after each b-generation differential evolution is completed, the population diversity coefficient is monitored in real time; when the population diversity coefficient is greater than the diversity threshold, the reinforcement learning evaluation is triggered; Based on the best performance, narrow the parameter range and keep the best performance Parameter interval; Extend according to the gradient direction parameter space, expanding the exploration area; Set up a reward mechanism, quantify the benefits of parameter adjustments in each optimization, and dynamically adjust the reinforcement learning strategy; The reward mechanism is to calculate the accuracy increment and the time consumed of the meta-model after each iteration, measure the benefit ratio by the accuracy increment and the time consumed, and feed the benefit ratio as a reward signal to the joint optimization framework of reinforcement learning and differential evolution to update the Q value of reinforcement learning; Keep the best previous individuals and deposit them into the elite pool; Elite gene recombination is performed every a generation, using a weighted crossover strategy based on parameter sensitivity, with high-sensitivity parameters being retained first; The convergence judgment conditions include: when the improvement rate of the optimal solution of consecutive a generations is less than the improvement rate threshold, it is judged to be converged; When the population gene similarity exceeds the population gene similarity threshold, it is judged as convergence; When the remaining time budget is less than the time budget threshold of the time budget, the current optimal solution is forced to be output and the optimization process is stopped.
8. An intelligent data mining platform based on big data analysis using the method according to any one of claims 1 to 7, characterized in that: The data module obtains the data to be processed from the data source, performs preprocessing, and generates a processed data set for data mining; The prediction module builds a meta-model based on historical data and predicts the optimal mining algorithm; The optimization module calculates the confidence of the meta-model and performs uncertainty analysis on the prediction results output by the meta-model. When the confidence judgment result indicates that the prediction results of the meta-model are uncertain, the joint optimization framework based on reinforcement learning and differential evolution is started to optimize the hyperparameters. The execution module feeds back the hyperparameters optimized by the joint optimization framework based on reinforcement learning and differential evolution to the meta-model, selects the optimal mining algorithm, and executes the data mining task.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the intelligent data mining method based on big data analysis described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the intelligent data mining method based on big data analysis described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Product change path multi-objective optimization method fusing reinforcement learning and differential evolution
CN116451577A
Reservoir level prediction method based on multi-source data fusion integration
CN117973600A
AI big data real-time processing and analysis method
CN119493823A
Machine language large model construction method and system
CN119721118A
Reinforcement learning using meta-learned intrinsic rewards
US20210089910A1
Cited By
Answer reasoning method and device based on large language model and medium
CN120509490A
Icing multi-dimensional prediction method and system based on multi-source data
CN120822152A
A Multi-Dimensional Prediction Method and System for Icing Based on Multi-Source Data
CN120822152B
Data mining method and device, equipment, storage medium and program product
CN121117481A
Production data real-time analysis method and system based on edge calculation
CN121119407A