A data modeling method, system and storage medium
Through the data model modeling method of automated feature engineering and model optimization, the complexity of feature selection and model training in traditional methods is solved, efficient feature selection and model optimization is achieved, and the performance and generalization capabilities of the data model are improved.
Patent Information
- Application Number
- CN202411349901.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-09-26
AI Technical Summary
In the field of industrial robot welding, traditional data modeling methods require a large amount of annotated data for training. The model training process is complex and the calculation overhead is large, so it is impossible to effectively select and filter feature, which affects the performance and generalization capabilities of the data model.
A data model modeling method is proposed, including obtaining the original large data set and performing standardization processing, automatically generating features, obtaining optimized features through regularized recursive feature screening, using decision trees, support vector machines and deep neural networks for preliminary modeling, Bayesian optimization algorithm for hyperparameter combination optimization, and adaptive update of the model through incremental learning.
Through automated feature engineering and model optimization, this method reduces manual intervention, improves the efficiency and effect of feature selection, reduces the complexity and computing overhead of the model, improves the performance and generalization capabilities of the data model, and can adapt to the dynamic changes of the data.
Smart Images

Figure CN119294241B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data processing technology, and in particular to a data modeling method, system and storage medium. Background Art
[0002] In the field of data science and artificial intelligence, data modeling is a key link in realizing data analysis and prediction. Through ensemble learning technology, multiple different base models are combined to improve prediction accuracy and stability. By introducing adaptive algorithms, model parameters can be adjusted according to the dynamic changes and characteristics of data to improve the adaptability of the model. Automated feature engineering technology is used to automatically generate, select and optimize features from raw data. This method reduces manual intervention, improves the efficiency and effect of feature selection, and is suitable for modeling high-dimensional data. At the same time, deep generative models such as generative adversarial networks (GANs) or variational autoencoders (VAEs) are used to generate high-quality data samples and features. These models can effectively model when data is scarce and improve the generalization ability of the model. However, in the field of industrial robot welding, traditional data modeling methods often require a large amount of labeled data for training, and the model training process is complex and computationally expensive, and it is impossible to effectively select and screen features, which affects the performance and generalization ability of the data model. Summary of the invention
[0003] Based on this, it is necessary for the present invention to provide a data modeling method, system and storage medium to solve at least one of the above technical problems.
[0004] To achieve the above object, a data modeling method comprises the following steps:
[0005] Step S1: obtaining an original big data set, and performing standardization processing on the original big data set to obtain a standard big data set; performing feature automatic generation processing on the standard big data set to obtain a big data high-level combination feature; performing regularized recursive feature screening processing on the big data high-level combination feature to obtain a large data volume related optimization feature;
[0006] Step S2: Preliminary modeling of optimization features related to the large amount of data is performed through decision trees, support vector machines, and deep neural network basic algorithms to obtain a preliminary basic model set for big data; the preliminary basic model set for big data is screened for the best model performance to obtain the model with the best big data performance;
[0007] Step S3: using the Bayesian optimization algorithm to optimize the hyperparameter combination of the model with the best big data performance, so as to obtain the big data model optimization hyperparameter combination; based on the big data model optimization hyperparameter combination, the model with the best big data performance is optimized and adjusted to obtain the big data optimization adjustment model;
[0008] Step S4: Real-time monitoring of feature updates of optimization features related to large amounts of data to obtain real-time update features of big data; based on the real-time update features of big data, an incremental learning method is used to adaptively update and optimize the big data optimization adjustment model to generate a big data adaptive adjustment model.
[0009] Further, step S1 includes the following steps:
[0010] Step S11: Obtaining the original large data set;
[0011] Step S12: performing noise elimination processing on the original large data set to obtain an original denoised large data set; performing histogram anomaly elimination processing on the original denoised large data set to obtain an original filtered large data set;
[0012] Step S13: performing missing interpolation compensation processing on the original filtered large data set to obtain an original interpolation compensation large data set; performing standardization processing on the original interpolation compensation large data set to obtain a standard large data set;
[0013] Step S14: automatically generate features for the standard big data set to obtain high-level combined features of the big data;
[0014] Step S15: Perform regularized recursive feature screening on the high-level combined features of the big data to obtain the optimization features related to the large data volume.
[0015] Further, step S14 includes the following steps:
[0016] Step S141: performing data feature engineering processing on the standard big data set to obtain a big data feature set;
[0017] Step S142: performing nonlinear interactive correlation analysis on each big data feature in the big data feature set to obtain the nonlinear interactive correlation relationship between the big data features;
[0018] Step S143: performing feature space topological analysis on each big data feature in the big data feature set to obtain a big data feature space distribution topological structure;
[0019] Step S144: performing feature space distribution mapping processing on each big data feature in the big data feature set based on the big data feature space distribution topological structure to obtain a big data analysis mapping feature space;
[0020] Step S145: Based on the nonlinear interactive correlation relationship between the big data features, the corresponding big data features in the big data analysis mapping feature space are subjected to feature interactive distribution combination to obtain the big data high-level combined features.
[0021] Further, step S15 includes the following steps:
[0022] Step S151: performing feature complexity evaluation and analysis on each big data feature in the big data high-level combined feature to obtain a complexity score value between each big data feature in the big data combined feature;
[0023] Step S152: constructing a feature complexity matrix for the complexity score values between each big data feature in the big data combination feature to generate a complexity score matrix between big data combination features;
[0024] Step S153: Calculate the regularization factor of each big data feature corresponding to the big data high-level combined feature based on the complexity score matrix between the big data combined features to obtain the regularization factor of each big data feature in the big data combined feature;
[0025] Step S154: recursively enhance each big data feature in the big data high-level combined feature according to the regularization factor of each big data feature in the big data combined feature, to obtain a big data recursive enhanced combined feature sequence;
[0026] Step S155: Perform quantity-related threshold screening processing on each big data feature in the big data recursive reinforcement combined feature sequence to obtain a large data quantity-related optimization feature.
[0027] Further, step S155 includes the following steps:
[0028] The potential interaction effect between each big data feature in the big data recursive reinforcement combined feature sequence is evaluated and analyzed to obtain the potential interaction effect relationship between each big data feature in the recursive reinforcement feature sequence;
[0029] Based on the potential interaction effect relationship between each big data feature in the recursive reinforcement feature sequence, the dynamic quantity correlation measurement calculation is performed on each corresponding big data feature in the big data recursive reinforcement combined feature sequence to obtain the dynamic quantity correlation measurement degree of each big data feature in the recursive reinforcement feature sequence;
[0030] Based on the dynamic quantity correlation measurement degree of each big data feature in the recursive reinforcement feature sequence, the corresponding big data features are screened by quantity correlation threshold to obtain large data quantity correlation optimization features.
[0031] Further, step S2 includes the following steps:
[0032] Step S21: Preliminary modeling and processing of optimization features related to large amounts of data are performed through decision trees, support vector machines, and deep neural network basic algorithms to obtain a preliminary basic model set for big data;
[0033] Step S22: performing model accuracy and recall evaluation and analysis on each big data basic model in the big data preliminary basic model set, and obtaining the model accuracy and model recall corresponding to each big data basic model;
[0034] Step S23: performing a model F1 score evaluation and analysis on each big data basic model in the big data preliminary basic model set based on the model accuracy and model recall rate corresponding to each big data basic model, and obtaining a model F1 score corresponding to each big data basic model;
[0035] Step S24: Calculate the model performance score of the corresponding big data basic models in the big data preliminary basic model set based on the model accuracy, model recall and model F1 score corresponding to each big data basic model, so as to obtain the model performance score value corresponding to each big data basic model;
[0036] Step S25: Based on the model performance score corresponding to each big data basic model, the corresponding big data basic models in the big data preliminary basic model set are screened for the best model performance to obtain the big data best performance model.
[0037] Further, step S3 includes the following steps:
[0038] Step S31: Perform model hyperparameter exploration and analysis on the model with the best big data performance to obtain a big data model hyperparameter exploration distribution matrix;
[0039] Step S32: performing model hyperparameter combination design on the big data model hyperparameter exploration distribution matrix to generate different big data model hyperparameter combination conditions;
[0040] Step S33: using the Bayesian optimization algorithm to perform model hyperparameter prior information modeling on the model with the best performance in big data to generate a model hyperparameter Bayesian optimization prior model;
[0041] Step S34: performing hyperparameter combination optimization selection on different big data model hyperparameter combination conditions based on the model hyperparameter Bayesian optimization prior model to obtain a big data model optimized hyperparameter combination;
[0042] Step S35: Based on the big data model optimization hyperparameter combination, the model with the best big data performance is optimized and adjusted to obtain a big data optimization and adjustment model.
[0043] Further, step S4 includes the following steps:
[0044] Step S41: Real-time monitoring of feature updates of optimization features related to large amounts of data to obtain real-time update features of large amounts of data;
[0045] Step S42: performing update feature flow analysis on the big data real-time update feature to obtain the big data real-time feature dynamic change trend data;
[0046] Step S43: performing feature incremental mapping analysis on the real-time update feature of the big data based on the dynamic change trend data of the real-time feature of the big data, and obtaining the real-time feature mapping increment of the big data;
[0047] Step S44: Adaptively update and optimize the big data optimization and adjustment model based on the big data real-time feature mapping increment using an incremental learning method to generate a big data adaptive adjustment model.
[0048] Furthermore, the present invention also provides a data modeling system for executing the data modeling method as described above, the data modeling system comprising:
[0049] The big data regularized recursive feature screening module is used to obtain the original big data set, and standardize the original big data set to obtain the standard big data set; perform feature automatic generation processing on the standard big data set to obtain the high-level combined features of big data; perform regularized recursive feature screening processing on the high-level combined features of big data, so as to obtain the optimization features related to the large data volume;
[0050] The big data modeling performance screening module is used to perform preliminary modeling processing on the optimization features related to the large amount of data through decision trees, support vector machines and deep neural network basic algorithms to obtain a preliminary basic model set for big data; the preliminary basic model set for big data is screened for the best model performance to obtain the best big data performance model;
[0051] The big data model hyperparameter optimization and adjustment module is used to optimize the hyperparameter combination of the model with the best big data performance using the Bayesian optimization algorithm to obtain the optimized hyperparameter combination of the big data model; based on the optimized hyperparameter combination of the big data model, the model with the best big data performance is optimized and adjusted to obtain the big data optimized adjustment model;
[0052] The big data model adaptive update and adjustment module is used to perform real-time monitoring of feature updates of optimization features related to large amounts of data to obtain real-time update features of big data; based on the real-time update features of big data, the big data optimization adjustment model is adaptively updated and optimized using an incremental learning method to generate a big data adaptive adjustment model.
[0053] Furthermore, the present invention also provides a storage medium on which a computer program is stored, and when the computer program is executed, the data modeling method as described above is implemented.
[0054] Beneficial effects of the present invention:
[0055] 1. The data modeling method proposed in the present invention, compared with the prior art, has the beneficial effect that obtaining the original large data set is the first step in the data analysis process, and its key role is to lay the foundation for subsequent data processing, analysis and modeling. The original large data set contains all data points in the problem domain, usually from various data sources, such as sensors, databases, social media or business records, etc. This step ensures the comprehensiveness and extensiveness of the data, so that the analysis can cover all aspects of the data problem. By collecting raw data, the data distribution, pattern and abnormal situation in the real world can be obtained, which is crucial to understanding the problem background and formulating analysis strategies. The integrity and accuracy of the original data directly affect the quality of subsequent data processing. By standardizing the original large data set, the standardization process is to normalize the large data set so that data with different features have the same scale. Standardization usually involves converting the data into a standard normal distribution with a mean of 0 and a standard deviation of 1. This process helps to eliminate the dimensional differences between features, so that the data can be compared at the same scale, thereby improving the training efficiency and accuracy of the machine learning model. After standardization, the features in the data set will have similar dimensions, so that the optimization algorithm (such as gradient descent) can converge faster, while reducing the impact caused by inconsistent feature scales. It also performs automatic feature generation on standard large data sets. Its main goal is to automatically generate new and more meaningful features from the original data. These new features can often capture the hidden patterns and relationships in the original data, thereby improving the expressiveness of the model. Feature generation can be achieved through a variety of methods, including feature engineering, feature extraction of deep learning models, or methods based on statistical learning. The generated high-level combined features can usually better represent the complex relationships in the data, thereby helping to improve the predictive ability and performance of the model. This automatic feature generation not only saves the time and cost of manual feature engineering, but also discovers those feature relationships that cannot be easily identified manually. At the same time, by performing regularized recursive feature screening on high-level combined features of big data, the aim is to select the features that contribute most to model performance from many features. By regularizing the features, the complexity of the model can be controlled and overfitting can be avoided. Regularization technology limits the influence of features by adding penalty terms, thereby improving the generalization ability of the model. By recursively training the model and evaluating the importance of features, features with less impact on the model are gradually removed, and finally the most useful feature set for prediction is retained. This method can effectively reduce feature dimensions and simplify the model structure, thereby reducing computational complexity and improving the interpretability of the model. The optimized feature set can not only reduce the training time of the model, but also improve the stability and accuracy of the model. The obtained feature set will be more relevant and representative, making data analysis and prediction more reliable and efficient, thereby better realizing the feature selection and screening process.Secondly, by using decision trees, support vector machines and deep neural network basic algorithms to conduct preliminary modeling of optimization features related to large amounts of data, the generalization ability and prediction accuracy of the model can be significantly improved. The decision tree splits the data set layer by layer, so that the model can grasp the nonlinear relationship between the features and the target variable. The support vector machine can find the optimal separation hyperplane in the high-dimensional space, so as to better handle complex classification problems. The deep neural network can learn complex patterns and relationships from the data through multi-level feature extraction capabilities. The construction of these basic models provides a variety of perspectives and methods for subsequent optimization and evaluation, and can lay a solid foundation for the accuracy of big data analysis. In addition, by screening the best performance of the model set of the preliminary basic model of big data, the final best performance model can be determined. This process can not only eliminate models with poor performance, but also highlight the best performance model in practical applications. Through this screening, the complexity and computational overhead of the model can be effectively reduced, while ensuring that the selected model has the best performance in practical applications. The ultimate goal of this stage is to find a model that can provide the best prediction performance in a big data environment, so as to provide the most efficient and reliable solution for practical applications. Then, the Bayesian optimization algorithm is used to optimize the hyperparameter combination of the best performance model for big data. In this stage, by optimizing different hyperparameter combination conditions and iteratively adjusting the Bayesian optimization model, the optimal hyperparameter configuration can be accurately located in the hyperparameter space. According to the influence of the hyperparameter combination on the model performance, the hyperparameter combination can be gradually selected and optimized to obtain the best hyperparameter configuration. This process not only improves the efficiency of hyperparameter search, but also reduces unnecessary computational overhead, ensuring the high quality of the optimization results, thereby significantly improving the performance and stability of the model. The model optimization adjustment is also carried out on the best performance model for big data by optimizing the hyperparameter combination based on the big data model to adjust the hyperparameters of the model, so that the model can achieve the best performance in the big data environment. This process includes the adjustment of hyperparameters such as model structure, learning rate, and regularization, so that the model can better adapt to the characteristics and needs of the data. The optimized model usually shows higher prediction accuracy, stronger generalization ability, and better computational efficiency, so that it can provide more stable and reliable results in practical applications. Finally, by performing real-time monitoring of feature updates for optimization features related to large amounts of data, it is possible to track and monitor changes in data features in real time, thereby ensuring that the data features used are always the latest and most relevant. Through real-time monitoring, emerging trends or anomalies in the data can be captured in a timely manner, which is crucial to maintaining the accuracy and effectiveness of the model. This dynamic monitoring helps to quickly identify changes in the data, allowing the model to better adapt to the ever-changing data environment and reduce prediction errors caused by outdated features.Real-time feature updates can also help improve the system's response speed, allowing the model to make adjustments in the shortest possible time, thereby more effectively supporting the decision-making process and business needs. In addition, the big data optimization and adjustment model is adaptively updated and optimized using incremental learning methods based on real-time update features of big data. The model can be adaptively updated using incremental learning methods to maintain the model's efficiency and accuracy. The incremental learning method allows the model to be incrementally updated when new data is received, rather than completely retrained. This method not only saves computing resources and time, but also can quickly adapt to new data features and changing trends. Through adaptive adjustment, the model can continuously optimize its performance to better meet the needs of actual applications, and can also maintain stability and reliability in a dynamic data environment, improving the overall performance and generalization ability of the data model.
[0056] 2. The data model building system proposed in the present invention is composed of a big data regularized recursive feature screening module, a big data modeling performance screening module, a big data model hyperparameter optimization and adjustment module, and a big data model adaptive update and adjustment module. It can implement any data model building method described in the present invention, and is used to combine the operations between computer programs running on each module to implement the data model building method. The internal structure of the system cooperates with each other, which can greatly reduce repetitive work and manpower investment, and can quickly and effectively provide a more accurate and efficient data model building process, thereby simplifying the operation flow of the data model building system. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments thereof made with reference to the following drawings:
[0058] Figure 1 A schematic diagram of the steps of the data modeling method of the present invention;
[0059] Figure 2 for Figure 1 Detailed step flow diagram of step S1;
[0060] Figure 3 for Figure 2 Detailed step flow chart of step S15 in FIG. DETAILED DESCRIPTION
[0061] The technical method of the present invention is described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by technicians in this field without creative work are within the scope of protection of the present invention.
[0062] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.
[0063] It should be understood that, although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are used only to distinguish one unit from another unit. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0064] To achieve this, please refer to Figures 1 to 3 The present invention provides a data modeling method, the method comprising the following steps:
[0065] Step S1: obtaining an original big data set, and performing standardization processing on the original big data set to obtain a standard big data set; performing feature automatic generation processing on the standard big data set to obtain a big data high-level combination feature; performing regularized recursive feature screening processing on the big data high-level combination feature to obtain a large data volume related optimization feature;
[0066] Step S2: Preliminary modeling of optimization features related to the large amount of data is performed through decision trees, support vector machines, and deep neural network basic algorithms to obtain a preliminary basic model set for big data; the preliminary basic model set for big data is screened for the best model performance to obtain the model with the best big data performance;
[0067] Step S3: using the Bayesian optimization algorithm to optimize the hyperparameter combination of the model with the best big data performance, so as to obtain the big data model optimization hyperparameter combination; based on the big data model optimization hyperparameter combination, the model with the best big data performance is optimized and adjusted to obtain the big data optimization adjustment model;
[0068] Step S4: Real-time monitoring of feature updates of optimization features related to large amounts of data to obtain real-time update features of big data; based on the real-time update features of big data, an incremental learning method is used to adaptively update and optimize the big data optimization adjustment model to generate a big data adaptive adjustment model.
[0069] In the embodiment of the present invention, please refer to Figure 1 FIG. 1 is a schematic diagram of the steps of the data modeling method of the present invention. In this example, the data modeling method includes the following steps:
[0070] Step S1: obtaining an original big data set, and performing standardization processing on the original big data set to obtain a standard big data set; performing feature automatic generation processing on the standard big data set to obtain a big data high-level combination feature; performing regularized recursive feature screening processing on the big data high-level combination feature to obtain a large data volume related optimization feature;
[0071] In an embodiment of the present invention, by extracting information from a variety of sensors and data sources, for example, in industrial robot welding applications, real-time data in the welding process can be obtained from welding equipment, industrial control systems, vision systems, etc. These data include temperature, pressure, welding current, welding time, and sensor image data. These data are stored in a centralized manner by using a data acquisition system (such as a SCADA system) and a database (such as an SQL database), and the integrity and accuracy of the data are ensured to form an original big data set, thereby obtaining an original big data set. The original large data set previously obtained is subjected to noise elimination by using a median filter algorithm or a Gaussian filter algorithm, wherein the median filter is suitable for removing random noise, while the Gaussian filter can process smooth noise, so that the random fluctuations and measurement errors in the original large data set can be reduced by applying these algorithms, and the original large data set previously obtained after noise elimination is subjected to abnormal elimination by using a histogram equalization method, so as to identify and remove abnormal data points appearing in the histogram, and at the same time, the original large data set previously obtained after outlier removal is subjected to interpolation compensation for missing values by using a linear interpolation method or an interpolation algorithm (such as Kriging interpolation) to fill in missing values through the linear relationship between known data points, and the Kriging interpolation is interpolated according to the spatial distribution characteristics of the data, and the corresponding interpolation compensation is completed, and the original large data set previously obtained after interpolation compensation is standardized by using a standardization method (such as Z-score standardization) to standardize the large data set to have a mean of 0 and a standard deviation of 1. This step ensures that all features are on the same scale, thereby obtaining a standard large data set. Then, by using feature engineering methods, such as principal component analysis (PCA) or autoencoder, high-level data features in standard big data sets are extracted. PCA transforms the original data into a set of new variables (principal components) through dimensionality reduction technology, retaining the main features of the data, while the autoencoder compresses the data into a low-dimensional representation through a neural network and reconstructs it to capture the potential features of the data. These combined features integrate the nonlinear relationship between the features to form a more complex and meaningful feature set, thereby obtaining big data high-level combined features. In addition, by applying regularization techniques (such as L1 regularization or L2 regularization) to regularize the big data high-level combined features previously obtained after combination, the feature weight of each big data feature in the big data high-level combined features is limited to reduce the risk of overfitting. Subsequently, by using the recursive feature elimination (RFE) algorithm, the model is trained and the features with greater correlation are gradually removed to determine the features that have the greatest impact on the model performance. For example, RFE is combined with a linear regression model to gradually eliminate features and calculate model performance indicators (such as mean square error) to obtain the most relevant optimized feature set for use in the large data volume related optimization features, and finally the large data volume related optimization features are obtained.
[0072] Step S2: Preliminary modeling of optimization features related to the large amount of data is performed through decision trees, support vector machines, and deep neural network basic algorithms to obtain a preliminary basic model set for big data; the preliminary basic model set for big data is screened for the best model performance to obtain the model with the best big data performance;
[0073] In an embodiment of the present invention, a preliminary modeling process is performed on the optimization features related to the large amount of data by using decision trees, support vector machines (SVMs) and deep neural networks (DNN) algorithms, wherein the decision tree model constructs a tree structure through criteria such as information gain, and performs preliminary processing on the optimization features related to the large amount of data by classification or regression; the SVM model maps the optimization features related to the large amount of data to a high-dimensional space by selecting a suitable kernel function, and finds the optimal hyperplane to distinguish different categories; and the deep neural network extracts hierarchical features of the optimization features related to the large amount of data through multiple layers of neurons, and optimizes the model parameters through a back propagation algorithm, each model trains the optimization features related to the large amount of data, and outputs a preliminary basic model set, thereby obtaining a preliminary basic model set for big data. By evaluating and calculating the model accuracy and recall rate of each big data basic model in the big data preliminary basic model set, each big data basic model is predicted by using the test set to obtain the prediction results, and the accuracy and recall rate of each big data basic model are calculated by comparing the prediction results with the true labels. The specific calculation method is: Accuracy (Precision) is the ratio of true positive examples (TP) to (true positive examples + false positive examples), and Recall (Recall) is the ratio of true positive examples (TP) to (true positive examples + false negative examples). The calculation formulas for accuracy and recall are: Accuracy = TP / (TP + FP), Recall = TP / (TP + FN) where TP is a true positive example, FP is a false positive example, and FN is a false negative example. By combining the model accuracy and model recall rate calculated previously, the F1 score of each corresponding big data basic model in the big data preliminary basic model set is evaluated and calculated, where the F1 score is the harmonic mean of the accuracy and recall rate, which is used to comprehensively evaluate the accuracy and comprehensiveness of the model. The specific calculation formula is: F1 score = 2 × (accuracy × recall) / (accuracy + recall) to calculate the F1 score corresponding to each big data base model.Then, the model accuracy, model recall and model F1 score corresponding to each big data basic model obtained by the previous quantitative calculation are normalized to the same dimension, for example, by the z-score normalization method, and the accuracy, recall and F1 score of each big data basic model are comprehensively scored by the weighted average method, wherein the weight can be set according to the requirements of the task, such as setting the F1 score weight higher, so as to quantitatively calculate the corresponding model performance score value, and by combining the model performance score corresponding to each big data basic model obtained by the previous quantitative calculation, the corresponding big data basic models in the big data preliminary basic model set are screened to determine the model with the best performance. The specific operation includes sorting the models according to the model performance score values calculated previously, arranging them from high to low, and selecting the model with the highest score as the model with the best performance. If multiple models have the same score, they can be further refined and screened by other indicators or cross-validation methods. This process ensures that the selected model has the best comprehensive performance and can provide the best performance in practical applications, and finally obtains the model with the best big data performance.
[0074] Step S3: using the Bayesian optimization algorithm to optimize the hyperparameter combination of the model with the best big data performance, so as to obtain the big data model optimization hyperparameter combination; based on the big data model optimization hyperparameter combination, the model with the best big data performance is optimized and adjusted to obtain the big data optimization adjustment model;
[0075] In an embodiment of the present invention, a hyperparameter exploration and analysis is performed on the best performance model of big data obtained after performance screening to collect a series of hyperparameter candidate values of the big data model. These candidate values can be based on previous experimental data, domain knowledge or literature research, and use grid search, random search and other methods to analyze the corresponding hyperparameters of the big data model, such as accuracy, F1 score, AUC, etc., which will be used to measure the performance of the model on the test set, and a new hyperparameter combination condition is designed by performing corresponding statistical analysis on the hyperparameter set of the big data model determined by the previous exploration. This process includes selecting the hyperparameter dimension of interest and setting its value range, and applying statistical methods such as variance analysis to determine the most influential hyperparameters. At the same time, a new hyperparameter combination is generated through an optimization algorithm (such as genetic algorithm, simulated annealing). Each combination takes into account the parameter settings that performed well in the previous exploration, and also introduces a new parameter range. At the same time, by using Use the Bayesian optimization algorithm to establish a prior distribution model of the hyperparameters corresponding to the best performance model of big data, so as to model the prior information of the hyperparameters of the best performance model of big data through methods such as Gaussian process to capture the relationship between hyperparameters and target performance. The input is the hyperparameter combination, and the output is the performance indicator of the model. Use these data to build a Bayesian optimization model to predict the potential performance of untested hyperparameter combinations. This prior model can guide the hyperparameter selection process, and optimize the hyperparameter combination for different big data model hyperparameter combinations by using the previously constructed Bayesian optimization prior model, so as to evaluate the potential benefits of different hyperparameter combinations through the acquisition function of Bayesian optimization (such as expected improvement or probability improvement), and use the prediction results of the Bayesian model to select the optimal hyperparameter combination for experiments to ensure that the selected hyperparameter combination can maximize the model performance in practical applications, thereby optimizing the optimized hyperparameter combination of the big data model. Then, by applying the previously optimized big data model optimization hyperparameter combination to the model with the best big data performance, the model with the best big data performance is retrained and adjusted to adapt to the new hyperparameter settings. During the training process, it is necessary to monitor the training error and verification error of the model to ensure that the model is not overfitting or underfitting during the optimization process. After the adjustment is completed, a final performance evaluation is performed to confirm whether the performance of the model on the large data set has been effectively improved, so as to obtain the final model after hyperparameter optimization, and finally the big data optimization and adjustment model.
[0076] Step S4: Real-time monitoring of feature updates of optimization features related to large amounts of data to obtain real-time update features of big data; based on the real-time update features of big data, an incremental learning method is used to adaptively update and optimize the big data optimization adjustment model to generate a big data adaptive adjustment model.
[0077] In an embodiment of the present invention, a big data processing platform, such as Apache Kafka or Apache Flink, is used for real-time data stream processing, and by using its data stream monitoring module, rules are configured to detect changes in the values of the optimization feature data related to the large amount of data, for example, by comparing the difference between the newly acquired related optimization feature data and the original feature data, to timely capture emerging trends or anomalies in the data, and to ensure that the data features used are always the latest and most relevant. Once a feature change is detected, the feature data is immediately updated to form a real-time feature data set, and it is stored in a distributed storage system (such as Hadoop HDFS), so as to monitor and obtain the real-time update features of the big data. At the same time, by using data analysis tools (such as Apache Spark performs flow analysis on the previously updated big data real-time update features, which is specifically implemented by reading the updated feature data from the distributed storage and loading it into the memory, and analyzing the change trend of the feature data stream by using the sliding window technology, and by combining the real-time feature dynamic change trend obtained by the previous analysis, using machine learning tools (such as TensorFlow or PyTorch) to perform incremental mapping analysis on the corresponding big data real-time update features, so as to input the trend data into the incremental learning algorithm (such as the online learning algorithm or the gradient boosting tree) to calculate the increment of the big data feature mapping, for example, by updating the feature weights and biases through the incremental learning method, and obtaining the real-time feature mapping increment in the data analysis process. Then, the real-time feature mapping increment of the big data obtained by the previous analysis is input into the big data optimization and adjustment model obtained after the previous optimization and adjustment, and the existing big data optimization and adjustment model is adaptively updated by using the previous incremental learning algorithm (such as an online learning algorithm or a gradient boosting tree). For example, the model parameters are gradually updated by using the incremental learning algorithm and combining the feature mapping increment obtained by the previous analysis, so that the model can adapt to the changes in the real-time feature increment. During the updating process, the model takes the real-time feature mapping increment data as input and adjusts the parameters in the model (such as weights and biases) to optimize the control strategy in the big data model adjustment process, thereby generating the latest adaptive adjustment model and finally generating a big data adaptive adjustment model.
[0078] Further, step S1 includes the following steps:
[0079] Step S11: Obtaining the original large data set;
[0080] Step S12: performing noise elimination processing on the original large data set to obtain an original denoised large data set; performing histogram anomaly elimination processing on the original denoised large data set to obtain an original filtered large data set;
[0081] Step S13: performing missing interpolation compensation processing on the original filtered large data set to obtain an original interpolation compensation large data set; performing standardization processing on the original interpolation compensation large data set to obtain a standard large data set;
[0082] Step S14: automatically generate features for the standard big data set to obtain high-level combined features of the big data;
[0083] Step S15: Perform regularized recursive feature screening on the high-level combined features of the big data to obtain the optimization features related to the large data volume.
[0084] As an embodiment of the present invention, refer to Figure 2 As shown, Figure 1 Detailed step flow diagram of step S1 in FIG. 1 , in this embodiment, step S1 includes the following steps:
[0085] Step S11: Obtaining the original large data set;
[0086] In an embodiment of the present invention, by extracting information from a variety of sensors and data sources, for example, in industrial robot welding applications, real-time data in the welding process can be obtained from welding equipment, industrial control systems, visual systems, etc. These data include temperature, pressure, welding current, welding time, and sensor image data. These data are stored in a centralized manner by using a data acquisition system (such as a SCADA system) and a database (such as an SQL database), and the integrity and accuracy of the data are ensured to form an original big data set, and finally an original big data set is obtained.
[0087] Step S12: performing noise elimination processing on the original large data set to obtain an original denoised large data set; performing histogram anomaly elimination processing on the original denoised large data set to obtain an original filtered large data set;
[0088] In the embodiment of the present invention, the original large data set previously obtained is subjected to noise elimination processing by using a median filter algorithm or a Gaussian filter algorithm, wherein the median filter is suitable for removing random noise, while the Gaussian filter can process smooth noise, so that the random fluctuations and measurement errors in the original large data set can be reduced by applying these algorithms, thereby obtaining the original denoised large data set. At the same time, the original denoised large data set previously obtained after denoising is subjected to abnormal elimination processing by using a histogram equalization method, so as to identify and remove abnormal data points appearing in the histogram, and finally obtain the original filtered large data set.
[0089] Step S13: performing missing interpolation compensation processing on the original filtered large data set to obtain an original interpolation compensation large data set; performing standardization processing on the original interpolation compensation large data set to obtain a standard large data set;
[0090] In an embodiment of the present invention, the original filtered large data set obtained after the outliers are removed is interpolated and compensated for missing values by using a linear interpolation method or an interpolation algorithm (such as Kriging interpolation) to fill in the missing values through the linear relationship between known data points, and the Kriging interpolation is interpolated according to the spatial distribution characteristics of the data, and the corresponding interpolation compensation is completed, thereby obtaining the original interpolation compensated large data set. At the same time, the original interpolation compensated large data set obtained after the interpolation compensation is standardized by using a standardization method (such as Z-score standardization), so that the mean of the standardized large data set is adjusted to 0 and the standard deviation is 1. This step ensures that all features are on the same scale, and finally a standard large data set is obtained.
[0091] Step S14: automatically generate features for the standard big data set to obtain high-level combined features of the big data;
[0092] In an embodiment of the present invention, feature engineering methods such as principal component analysis (PCA) or autoencoder are used to extract high-level data features from a standard large data set. PCA converts the original data into a set of new variables (principal components) through dimensionality reduction technology, retaining the main features of the data, while the autoencoder compresses the data into a low-dimensional representation through a neural network and reconstructs it to capture the potential features of the data. These combined features integrate the nonlinear relationship between the features to form a more complex and meaningful feature set, and finally obtain high-level combined features of big data.
[0093] Step S15: Perform regularized recursive feature screening on the high-level combined features of the big data to obtain the optimization features related to the large data volume.
[0094] In an embodiment of the present invention, first, a regularization process is performed on the high-level combined features of big data obtained after the previous combination by applying a regularization technique (such as L1 regularization or L2 regularization) to limit the feature weight of each big data feature in the high-level combined features of big data and reduce the risk of overfitting. Subsequently, a recursive feature elimination (RFE) algorithm is used to train the model and gradually remove features with greater correlation to determine the features that have the greatest impact on the model performance. For example, RFE is combined with a linear regression model to gradually eliminate features and calculate model performance indicators (such as mean square error) to obtain the most relevant set of optimized features for use in large-volume-related optimized features, and finally large-volume-related optimized features are obtained.
[0095] Further, step S14 includes the following steps:
[0096] Step S141: performing data feature engineering processing on the standard big data set to obtain a big data feature set;
[0097] Step S142: performing nonlinear interactive correlation analysis on each big data feature in the big data feature set to obtain the nonlinear interactive correlation relationship between the big data features;
[0098] Step S143: performing feature space topological analysis on each big data feature in the big data feature set to obtain a big data feature space distribution topological structure;
[0099] Step S144: performing feature space distribution mapping processing on each big data feature in the big data feature set based on the big data feature space distribution topological structure to obtain a big data analysis mapping feature space;
[0100] Step S145: Based on the nonlinear interactive correlation relationship between the big data features, the corresponding big data features in the big data analysis mapping feature space are subjected to feature interactive distribution combination to obtain the big data high-level combined features.
[0101] As an embodiment of the present invention, refer to Figure 3 As shown, Figure 2 Detailed step flow diagram of step S14 in the embodiment, step S14 includes the following steps:
[0102] Step S141: performing data feature engineering processing on the standard big data set to obtain a big data feature set;
[0103] In an embodiment of the present invention, key features are extracted from a standard large data set obtained after standardization, and data preprocessing tools, such as the Pandas library in Python, are used to clean, fill missing values, and process outliers on the standard large data set. Feature extraction methods such as principal component analysis (PCA) or linear discriminant analysis (LDA) are applied to convert the data into a low-dimensional space to enhance the expressive power of the features, thereby generating a feature set containing screened and transformed features that can effectively represent the main features of the data, and finally obtaining a big data feature set.
[0104] Step S142: performing nonlinear interactive correlation analysis on each big data feature in the big data feature set to obtain the nonlinear interactive correlation relationship between the big data features;
[0105] In an embodiment of the present invention, a nonlinear interactive correlation analysis method such as kernel principal component analysis (KPCA) or polynomial regression analysis is used to perform nonlinear correlation analysis on each corresponding big data feature in a big data feature set, so as to map the features to a high-dimensional space through a kernel function, and then perform nonlinear interactive correlation modeling to capture the complex relationship between the features. The specific operation includes using the KPCA function in the Scikit-learn library or a custom polynomial regression model to calculate the nonlinear interactive correlation index between the features, thereby obtaining a nonlinear correlation matrix, showing the complex nonlinear interactive correlation between the big data features, and finally obtaining the nonlinear interactive correlation between the big data features.
[0106] Step S143: performing feature space topological analysis on each big data feature in the big data feature set to obtain a big data feature space distribution topological structure;
[0107] In an embodiment of the present invention, a topological analysis of the feature space is performed on each big data feature in a big data feature set, so as to generate a topological structure diagram of the feature space by using a topological data analysis library such as GUDHI or Ripser. This diagram shows the distribution of features in space and their mutual relationships, thereby obtaining the topological structure of the feature space, including the identification of clusters and connected components, and finally obtaining the distribution topological structure of the big data feature space.
[0108] Step S144: performing feature space distribution mapping processing on each big data feature in the big data feature set based on the big data feature space distribution topological structure to obtain a big data analysis mapping feature space;
[0109] In an embodiment of the present invention, distribution mapping is performed on each big data feature in a big data feature set by combining the big data feature space distribution topological structure obtained by previous analysis, so as to map the feature space topological structure to a two-dimensional or three-dimensional space by using dimensionality reduction techniques such as t-SNE (t-Distributed Stochastic Neighbor Embedding) or UMAP (Uniform Manifold Approximation and Projection), and simplify the data while retaining the local structure by using the t-SNE or UMAP algorithm, and process the data using the Sklearn library or UMAP library in Python to obtain a mapped feature space, which can intuitively display the structure and distribution characteristics of the data, and finally obtain a big data analysis mapping feature space.
[0110] Step S145: Based on the nonlinear interactive correlation relationship between the big data features, the corresponding big data features in the big data analysis mapping feature space are subjected to feature interactive distribution combination to obtain the big data high-level combined features.
[0111] In an embodiment of the present invention, feature interaction modeling technology (such as random forest or gradient boosting tree) is used to perform distribution combination processing of feature interaction relationships on corresponding big data features in the feature space of big data analysis mapping by combining the nonlinear interactive correlation relationship between big data features obtained in the previous analysis, so as to identify the interaction effects between features by building a model, and use these interaction effects to generate combined features. The specific operation includes using the XGBoost library in Python or the random forest in Scikit-learn to implement feature interaction modeling, extracting and combining high-level features therefrom, and these combined features integrate the nonlinear relationships between features to form a more complex and meaningful feature set, and finally obtaining high-level combined features of big data.
[0112] Further, step S15 includes the following steps:
[0113] Step S151: performing feature complexity evaluation and analysis on each big data feature in the big data high-level combined feature to obtain a complexity score value between each big data feature in the big data combined feature;
[0114] In an embodiment of the present invention, by selecting an appropriate complexity evaluation method, such as information entropy, mutual information or correlation measurement, the feature complexity is evaluated and calculated between each big data feature in the high-level combined feature of big data obtained by the previous analysis, so as to quantify the information gain or mutual information between each pair of big data features and evaluate their complexity. The information entropy can be calculated by statistically analyzing the distribution of each feature, and the mutual information is evaluated by measuring the difference between the joint probability distribution and the marginal probability distribution between the features, so as to integrate the complexity results of each pair of features into the feature complexity score value, and generate a complexity score value for each pair of features, reflecting their correlation strength and complexity, and finally obtaining the complexity score value between each big data feature in the big data combination feature.
[0115] Step S152: constructing a feature complexity matrix for the complexity score values between each big data feature in the big data combination feature to generate a complexity score matrix between big data combination features;
[0116] In an embodiment of the present invention, a matrix is constructed by combining the complexity score values between each big data feature in the big data combination feature obtained by previous quantitative calculation. Each item of the matrix corresponds to the complexity score value between the features. This requires that the complexity score values of all features be aggregated into a square matrix, and the dimension of the matrix matches the number of features. In specific operations, the complexity score values obtained previously are filled into the corresponding positions of the matrix to form a symmetric matrix, in which the elements on the diagonal represent the complexity of the feature itself, and the elements on the off-diagonal represent the complexity between the features. After the matrix is constructed, the complexity matrix can be graphically displayed using a matrix visualization tool to finally generate a complexity score matrix between big data combination features.
[0117] Step S153: Calculate the regularization factor of each big data feature corresponding to the big data high-level combined feature based on the complexity score matrix between the big data combined features to obtain the regularization factor of each big data feature in the big data combined feature;
[0118] In an embodiment of the present invention, a regularization factor is statistically calculated for each corresponding big data feature in the big data high-level combination feature by combining the complexity scoring matrix between big data combination features constructed previously, so as to select a regularization algorithm, such as L1 or L2 regularization, and calculate the regularization factor of each feature by applying a regularization method to each element of the feature complexity scoring matrix. The calculation of the regularization factor involves standardizing the score value of each row or column of the complexity scoring matrix to balance the influence of the feature, which can eliminate noise and redundancy between features, thereby obtaining the regularization factor value of each feature, and finally obtaining the regularization factor of each big data feature in the big data combination feature.
[0119] Step S154: recursively enhance each big data feature in the big data high-level combined feature according to the regularization factor of each big data feature in the big data combined feature, to obtain a big data recursive enhanced combined feature sequence;
[0120] In an embodiment of the present invention, by combining the regularization factor of each big data feature in the big data combination feature obtained by previous analysis, recursive feature enhancement processing is performed on each corresponding big data feature in the big data high-level combination feature to improve the significance of the big data feature and the stability of the model. This processing involves using an enhancement algorithm (such as Boosting, Stacking) to iteratively weight each big data feature. The specific steps include inputting the regularization factor as a weight into the enhancement algorithm, iteratively optimizing the weight and selection of each feature, and gradually adjusting the importance of the feature, and finally obtaining a big data recursive enhanced combination feature sequence. These features, after multiple rounds of enhancement and weighting, can more effectively represent the potential information of the data.
[0121] Step S155: Perform quantity-related threshold screening processing on each big data feature in the big data recursive reinforcement combined feature sequence to obtain a large data quantity-related optimization feature.
[0122] In an embodiment of the present invention, each big data feature in the big data recursively enhanced combined feature sequence obtained after the previous enhancement is screened by a quantity-related threshold, so as to evaluate the correlation between each feature and the target variable by using a correlation coefficient calculation tool, and set a quantity-related threshold to compare and judge them. The threshold is based on the relationship between the correlation of the feature and the target variable. Features with a correlation lower than the set threshold are eliminated, and the remaining features (i.e., the correlation is higher than or equal to the set threshold) are quantity-related optimization features. These features have higher correlation and can more effectively reflect the actual laws of the data, and finally obtain large data quantity-related optimization features.
[0123] Further, step S155 includes the following steps:
[0124] The potential interaction effect between each big data feature in the big data recursive reinforcement combined feature sequence is evaluated and analyzed to obtain the potential interaction effect relationship between each big data feature in the recursive reinforcement feature sequence;
[0125] In an embodiment of the present invention, a feature interaction effect model is constructed by evaluating and analyzing the potential interaction effects between each big data feature in a big data recursively enhanced combined feature sequence obtained after previous enhancement. The model uses a feature interaction analysis tool based on a graph algorithm, such as a graph convolutional network (GCN), to represent and analyze the interaction relationship between features. Each big data feature is regarded as a node in the graph, and the edges between the nodes represent potential interaction effects. The model is used to recursively calculate each feature in the feature sequence to identify and evaluate the intensity of the interaction effect between the features, and to evaluate and analyze the potential interaction effect relationship between each pair of features, thereby clarifying the potential influence and interaction relationship between each feature, and finally obtaining the potential interaction effect relationship between each big data feature in the recursively enhanced feature sequence.
[0126] Preferably, based on the potential interaction effect relationship between each big data feature in the recursive reinforcement feature sequence, a dynamic quantity correlation measurement calculation is performed on each corresponding big data feature in the big data recursive reinforcement combined feature sequence to obtain a dynamic quantity correlation measurement degree of each big data feature in the recursive reinforcement feature sequence;
[0127] In an embodiment of the present invention, a dynamic quantity correlation measurement algorithm is used to perform quantity correlation measurement calculation on each corresponding big data feature in the big data recursive reinforcement combined feature sequence by combining the potential interaction effect relationship between each big data feature in the recursive reinforcement feature sequence obtained by previous analysis. The specific operation includes applying dynamic time warping (DTW) or mutual information measurement algorithm to perform correlation analysis on each feature and its corresponding interaction effect. The dynamic time warping algorithm can handle the dynamic changes of the time series between features, while the mutual information measurement algorithm evaluates the dependency relationship between features, and calculates the dynamic quantity correlation of each feature with other features, thereby quantitatively calculating a series of quantity correlation measurement values. These measurement values reflect the importance and correlation of the features in the recursive reinforcement combined sequence, and finally obtain the dynamic quantity correlation measurement degree of each big data feature in the recursive reinforcement feature sequence.
[0128] Preferably, based on the dynamic quantity correlation measurement degree of each big data feature in the recursive reinforcement feature sequence, the corresponding big data features are subjected to quantity correlation threshold screening processing to obtain large data quantity correlation optimization features.
[0129] In an embodiment of the present invention, a preset quantity correlation measurement threshold is used to perform threshold screening and judgment on the dynamic quantity correlation measurement degree of each big data feature in the recursively enhanced feature sequence obtained by the previous quantitative calculation, and only the big data features corresponding to the measurement value higher than the threshold are retained. This step can be completed by using a screening algorithm such as K-means clustering or principal component analysis (PCA) to ensure that the screened features have significant quantity correlation. The screened feature set represents the optimized large data volume-related features, which can effectively improve the performance and accuracy of the model, and ultimately obtain large data volume-related optimized features.
[0130] Further, step S2 includes the following steps:
[0131] Step S21: Preliminary modeling and processing of optimization features related to large amounts of data are performed through decision trees, support vector machines, and deep neural network basic algorithms to obtain a preliminary basic model set for big data;
[0132] In an embodiment of the present invention, a preliminary modeling process is performed on the optimization features related to the large amount of data by using decision trees, support vector machines (SVMs) and deep neural networks (DNN) algorithms, wherein the decision tree model constructs a tree structure by using criteria such as information gain, and performs preliminary processing on the optimization features related to the large amount of data by classification or regression; the SVM model maps the optimization features related to the large amount of data to a high-dimensional space by selecting a suitable kernel function, and finds the optimal hyperplane to distinguish different categories; and the deep neural network extracts hierarchical features of the optimization features related to the large amount of data through multiple layers of neurons, and optimizes the model parameters through a back propagation algorithm. Each model trains the optimization features related to the large amount of data and outputs a preliminary basic model set. In this process, the selection of decision trees, SVMs and DNNs is based on data characteristics and task requirements. For example, decision trees are suitable for situations where the relationship between features is relatively simple, SVMs are suitable for classification of small sample high-dimensional data, and DNNs are suitable for processing complex nonlinear relationships, and finally a preliminary basic model set of big data is obtained.
[0133] Step S22: performing model accuracy and recall evaluation and analysis on each big data basic model in the big data preliminary basic model set, and obtaining the model accuracy and model recall corresponding to each big data basic model;
[0134] In an embodiment of the present invention, the model accuracy and recall rate are evaluated and calculated for each big data basic model in the big data preliminary basic model set, so as to predict each big data basic model by using the test set to obtain the prediction result, and the accuracy and recall rate of each big data basic model are calculated by comparing the prediction result with the true label. The specific calculation method is: the accuracy (Precision) is the ratio of the true positive example (TP) to the (true positive example + false positive example), and the recall rate (Recall) is the ratio of the true positive example (TP) to the (true positive example + false negative example). The calculation formulas of the accuracy and recall rate are: Accuracy = TP / (TP + FP), Recall = TP / (TP + FN) wherein TP is the true positive example, FP is the false positive example, and FN is the false negative example. This process ensures that the performance of each big data basic model under different test data is accurately evaluated, and finally the model accuracy and model recall rate corresponding to each big data basic model are obtained.
[0135] Step S23: performing a model F1 score evaluation and analysis on each big data basic model in the big data preliminary basic model set based on the model accuracy and model recall rate corresponding to each big data basic model, and obtaining a model F1 score corresponding to each big data basic model;
[0136] In an embodiment of the present invention, an F1 score is evaluated and calculated for each big data basic model corresponding to the big data preliminary basic model set by combining the model accuracy and model recall rate corresponding to each big data basic model calculated previously, wherein the F1 score is the harmonic mean of the accuracy and the recall rate, and is used to comprehensively evaluate the accuracy and comprehensiveness of the model. The specific calculation formula is: F1 score = 2 × (accuracy × recall rate) / (accuracy + recall rate), thereby calculating the F1 score corresponding to each big data basic model. This process ensures that the evaluation index comprehensively considers the accuracy and recall ability, and can more comprehensively reflect the performance of the model, and finally obtains the model F1 score corresponding to each big data basic model.
[0137] Step S24: Calculate the model performance score of the corresponding big data basic models in the big data preliminary basic model set based on the model accuracy, model recall and model F1 score corresponding to each big data basic model, so as to obtain the model performance score value corresponding to each big data basic model;
[0138] In an embodiment of the present invention, the model accuracy, model recall and model F1 score corresponding to each big data basic model previously obtained by quantitative calculation are normalized to the same dimension, for example, by a z-score normalization method, and a weighted average method is used to perform a comprehensive score on the accuracy, recall and F1 score of each big data basic model, wherein the weight can be set according to the requirements of the task, such as setting the F1 score weight higher, and finally a model performance score corresponding to each big data basic model is obtained by quantitative calculation.
[0139] Step S25: Based on the model performance score corresponding to each big data basic model, the corresponding big data basic models in the big data preliminary basic model set are screened for the best model performance to obtain the big data best performance model.
[0140] In an embodiment of the present invention, the corresponding big data basic models in the big data preliminary basic model set are screened by combining the model performance score values corresponding to each big data basic model obtained by previous quantitative calculation to determine the model with the best performance. The specific operation includes sorting the models according to the model performance score values calculated previously, arranging them from high to low, and selecting the model with the highest score as the model with the best performance. If there are multiple models with the same score, they can be further refined and screened through other indicators or cross-validation methods. This process ensures that the selected model has the best overall performance and can provide the best performance in practical applications, and finally the big data model with the best performance is obtained.
[0141] Further, step S3 includes the following steps:
[0142] Step S31: Perform model hyperparameter exploration and analysis on the model with the best big data performance to obtain a big data model hyperparameter exploration distribution matrix;
[0143] In an embodiment of the present invention, a hyperparameter exploration analysis is performed on the best big data performance model obtained after performance screening to collect a series of hyperparameter candidate values for the big data model. These candidate values can be based on previous experimental data, domain knowledge or literature research, and use grid search, random search and other methods to analyze the corresponding hyperparameters of the big data model. For example, accuracy, F1 score, AUC, etc. will be used to measure the performance of the model on the test set, and a hyperparameter exploration distribution matrix is constructed. The matrix records the model performance value corresponding to each hyperparameter, thereby forming a detailed distribution matrix of the relationship between hyperparameters and model performance, and finally obtaining the big data model hyperparameter exploration distribution matrix.
[0144] Step S32: performing model hyperparameter combination design on the big data model hyperparameter exploration distribution matrix to generate different big data model hyperparameter combination conditions;
[0145] In an embodiment of the present invention, a new hyperparameter combination condition is designed by performing corresponding statistical analysis on the big data model hyperparameter exploration distribution matrix obtained by previous exploration. This process includes selecting the hyperparameter dimensions of interest and setting their value ranges, and applying statistical methods such as variance analysis to determine the most influential hyperparameters. At the same time, new hyperparameter combinations are generated through optimization algorithms (such as genetic algorithms and simulated annealing). Each combination takes into account parameter settings that have performed well in previous explorations, and also introduces new parameter ranges to ensure diversity. Ultimately, different big data model hyperparameter combination conditions are designed and generated.
[0146] Step S33: using the Bayesian optimization algorithm to perform model hyperparameter prior information modeling on the model with the best performance in big data to generate a model hyperparameter Bayesian optimization prior model;
[0147] In an embodiment of the present invention, a prior distribution model of the hyperparameters corresponding to the model with the best performance of big data is established by using a Bayesian optimization algorithm, and the prior information of the hyperparameters of the model with the best performance of big data is modeled by using methods such as Gaussian process to capture the relationship between the hyperparameters and the target performance. The input is the distribution matrix of the hyperparameters, and the output is the performance index of the model. A Bayesian optimization model is constructed using these data to predict the potential performance of untested hyperparameter combinations. This prior model can guide the hyperparameter selection process, so as to give priority to exploring the hyperparameter combination that is most likely to improve model performance under limited experimental resources, and finally generate a Bayesian optimization prior model of the model hyperparameters.
[0148] Step S34: performing hyperparameter combination optimization selection on different big data model hyperparameter combination conditions based on the model hyperparameter Bayesian optimization prior model to obtain a big data model optimized hyperparameter combination;
[0149] In an embodiment of the present invention, a previously constructed model hyperparameter Bayesian optimization prior model is used to optimize the hyperparameter combination for different big data model hyperparameter combination conditions, so as to evaluate the potential benefits of different hyperparameter combinations through the acquisition function of Bayesian optimization (such as expected improvement or probability improvement), and use the prediction results of the Bayesian model to select the optimal hyperparameter combination for experimentation. These selections are based on the predictive performance of the model and actual experimental feedback to ensure that the selected hyperparameter combination can maximize the model performance in actual applications, and finally optimize the big data model optimized hyperparameter combination.
[0150] Step S35: Based on the big data model optimization hyperparameter combination, the model with the best big data performance is optimized and adjusted to obtain a big data optimization and adjustment model.
[0151] In an embodiment of the present invention, by applying the big data model optimization hyperparameter combination obtained after previous optimization to the model with the best big data performance, the model with the best big data performance is retrained and adjusted to adapt to the new hyperparameter settings. During the training process, it is necessary to monitor the training error and verification error of the model to ensure that the model is not overfitting or underfitting during the optimization process. After the adjustment is completed, a final performance evaluation is performed to confirm whether the performance of the model on the large data set has been effectively improved, thereby obtaining the final model after hyperparameter optimization, and finally obtaining a big data optimization and adjustment model.
[0152] Further, step S4 includes the following steps:
[0153] Step S41: Real-time monitoring of feature updates of optimization features related to large amounts of data to obtain real-time update features of large amounts of data;
[0154] In an embodiment of the present invention, a big data processing platform, such as Apache Kafka or Apache Flink, is used for real-time data stream processing, and its data stream monitoring module is used to configure rules to detect changes in the values of large amounts of relevant optimized feature data. For example, by comparing the differences between newly acquired relevant optimized feature data and original feature data, emerging trends or anomalies in the data can be captured in a timely manner, and it is ensured that the data features used are always the latest and most relevant. Once a feature change is detected, the feature data is immediately updated to form a real-time feature data set, which is stored in a distributed storage system (such as Hadoop HDFS), and finally the real-time update features of the big data are monitored.
[0155] Step S42: performing update feature flow analysis on the big data real-time update feature to obtain the big data real-time feature dynamic change trend data;
[0156] In an embodiment of the present invention, a flow analysis is performed on the real-time updated features of big data that were previously updated in real time by using a data analysis tool (such as Apache Spark). Specifically, the updated feature data is read from the distributed storage and loaded into the memory, and the changing trend of the feature data stream is analyzed by using a sliding window technology. For example, a 10-minute sliding window is set to analyze the changes in data features within the past 10 minutes, and the dynamic change trend data of the real-time features are obtained by calculating the average value, variance and trend change of the feature data, and finally the dynamic change trend data of the real-time features of the big data are obtained.
[0157] Step S43: performing feature incremental mapping analysis on the real-time update feature of the big data based on the dynamic change trend data of the real-time feature of the big data, and obtaining the real-time feature mapping increment of the big data;
[0158] In an embodiment of the present invention, an incremental mapping analysis is performed on the corresponding big data real-time update features using a machine learning tool (such as TensorFlow or PyTorch) in combination with the dynamic change trend data of the big data real-time features obtained by previous analysis, so that the trend data can be input into an incremental learning algorithm (such as an online learning algorithm or a gradient boosting tree) to calculate the increment of the big data feature mapping. For example, the feature weights and biases are updated through the incremental learning method to obtain the real-time feature mapping increment in the data analysis process, and finally the big data real-time feature mapping increment is obtained.
[0159] Step S44: Adaptively update and optimize the big data optimization and adjustment model based on the big data real-time feature mapping increment using an incremental learning method to generate a big data adaptive adjustment model.
[0160] In an embodiment of the present invention, the real-time feature mapping increment of big data obtained by previous analysis is input into the big data optimization and adjustment model obtained after previous optimization and adjustment, and the existing big data optimization and adjustment model is adaptively updated by using the previous incremental learning algorithm (such as an online learning algorithm or a gradient boosting tree). For example, the model parameters are gradually updated by using the incremental learning algorithm and combining the feature mapping increment obtained by previous analysis, so that the model can adapt to the changes in the real-time feature increment. During the updating process, the model uses the real-time feature mapping increment data as input and adjusts the parameters in the model (such as weights and biases) to optimize the control strategy in the big data model adjustment process, thereby generating the latest adaptive adjustment model and finally generating a big data adaptive adjustment model.
[0161] Furthermore, the present invention also provides a data modeling system for executing the data modeling method as described above, the data modeling system comprising:
[0162] The big data regularized recursive feature screening module is used to obtain the original big data set, and standardize the original big data set to obtain the standard big data set; perform feature automatic generation processing on the standard big data set to obtain the high-level combined features of big data; perform regularized recursive feature screening processing on the high-level combined features of big data, so as to obtain the optimization features related to the large data volume;
[0163] The big data modeling performance screening module is used to perform preliminary modeling processing on the optimization features related to the large amount of data through decision trees, support vector machines and deep neural network basic algorithms to obtain a preliminary basic model set for big data; the preliminary basic model set for big data is screened for the best model performance to obtain the best big data performance model;
[0164] The big data model hyperparameter optimization and adjustment module is used to optimize the hyperparameter combination of the model with the best big data performance using the Bayesian optimization algorithm to obtain the optimized hyperparameter combination of the big data model; based on the optimized hyperparameter combination of the big data model, the model with the best big data performance is optimized and adjusted to obtain the big data optimized adjustment model;
[0165] The big data model adaptive update and adjustment module is used to perform real-time monitoring of feature updates of optimization features related to large amounts of data to obtain real-time update features of big data; based on the real-time update features of big data, the big data optimization adjustment model is adaptively updated and optimized using an incremental learning method to generate a big data adaptive adjustment model.
[0166] Furthermore, the present invention also provides a storage medium on which a computer program is stored, and when the computer program is executed, the data modeling method as described above is implemented.
[0167] Therefore, the embodiments should be regarded as illustrative and non-restrictive from all points, and the scope of the present invention is limited by the appended claims rather than the above description, and it is therefore intended that all changes falling within the meaning and range of equivalent elements of the application documents are included in the present invention.
[0168] The above description is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features invented herein.
Claims
1. A data modeling method, characterized in that: Applied to industrial robot welding, the data model building method comprises the following steps: Step S1: acquiring real-time data during the welding process from welding equipment, industrial control systems, and visual systems, wherein the real-time data includes temperature, pressure, welding current, welding time, and image data of sensors; storing these data in a centralized manner using a data acquisition system and a database to obtain an original large data set; and performing standardization processing on the original large data set to obtain a standard large data set; performing feature automatic generation processing on the standard large data set to obtain a large data high-level combination feature; performing regularized recursive feature screening processing on the large data high-level combination feature to obtain a large data volume related optimization feature; The automatic feature generation process of the standard big data set to obtain the high-level combined features of the big data includes: Step S141: performing data feature engineering processing on the standard big data set to obtain a big data feature set; Step S142: performing nonlinear interactive correlation analysis on each big data feature in the big data feature set to obtain the nonlinear interactive correlation relationship between the big data features; Step S143: performing feature space topological analysis on each big data feature in the big data feature set to obtain a big data feature space distribution topological structure; Step S144: performing feature space distribution mapping processing on each big data feature in the big data feature set based on the big data feature space distribution topological structure to obtain a big data analysis mapping feature space; Step S145: performing feature interactive distribution combination on the corresponding big data features in the big data analysis mapping feature space based on the nonlinear interactive correlation relationship between the big data features to obtain the big data high-level combined features; Step S2: Preliminary modeling of optimization features related to the large amount of data is performed through decision trees, support vector machines, and deep neural network basic algorithms to obtain a preliminary basic model set for big data; the preliminary basic model set for big data is screened for the best model performance to obtain the model with the best big data performance; Step S3: using the Bayesian optimization algorithm to optimize the hyperparameter combination of the model with the best big data performance, so as to obtain the big data model optimization hyperparameter combination; based on the big data model optimization hyperparameter combination, the model with the best big data performance is optimized and adjusted to obtain the big data optimization adjustment model; Step S4: Real-time monitoring of feature updates of optimization features related to large amounts of data to obtain real-time update features of big data; based on the real-time update features of big data, an incremental learning method is used to adaptively update and optimize the big data optimization adjustment model to generate a big data adaptive adjustment model.
2. The data modeling method according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: Obtaining the original large data set; Step S12: performing noise elimination processing on the original large data set to obtain an original denoised large data set; performing histogram anomaly elimination processing on the original denoised large data set to obtain an original filtered large data set; Step S13: performing missing interpolation compensation processing on the original filtered large data set to obtain an original interpolation compensation large data set; performing standardization processing on the original interpolation compensation large data set to obtain a standard large data set; Step S14: automatically generate features for the standard big data set to obtain high-level combined features of the big data; Step S15: Perform regularized recursive feature screening on the high-level combined features of the big data to obtain the optimization features related to the large data volume.
3. The data modeling method according to claim 2, characterized in that: Step S15 includes the following steps: Step S151: performing feature complexity evaluation and analysis on each big data feature in the big data high-level combined feature to obtain a complexity score value between each big data feature in the big data combined feature; Step S152: constructing a feature complexity matrix for the complexity score values between each big data feature in the big data combination feature to generate a complexity score matrix between big data combination features; Step S153: Calculate the regularization factor of each big data feature corresponding to the big data high-level combined feature based on the complexity score matrix between the big data combined features to obtain the regularization factor of each big data feature in the big data combined feature; Step S154: recursively enhance each big data feature in the big data high-level combined feature according to the regularization factor of each big data feature in the big data combined feature, to obtain a big data recursive enhanced combined feature sequence; Step S155: Perform quantity-related threshold screening processing on each big data feature in the big data recursive reinforcement combined feature sequence to obtain a large data quantity-related optimization feature.
4. The data modeling method according to claim 3, characterized in that: Step S155 includes the following steps: The potential interaction effect between each big data feature in the big data recursive reinforcement combined feature sequence is evaluated and analyzed to obtain the potential interaction effect relationship between each big data feature in the recursive reinforcement feature sequence; Based on the potential interaction effect relationship between each big data feature in the recursive reinforcement feature sequence, the dynamic quantity correlation measurement calculation is performed on each corresponding big data feature in the big data recursive reinforcement combined feature sequence to obtain the dynamic quantity correlation measurement degree of each big data feature in the recursive reinforcement feature sequence; Based on the dynamic quantity correlation measurement degree of each big data feature in the recursive reinforcement feature sequence, the corresponding big data features are screened by quantity correlation threshold to obtain large data quantity correlation optimization features.
5. The data modeling method according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: Preliminary modeling and processing of optimization features related to large amounts of data are performed through decision trees, support vector machines, and deep neural network basic algorithms to obtain a preliminary basic model set for big data; Step S22: performing model accuracy and recall evaluation and analysis on each big data basic model in the big data preliminary basic model set, and obtaining the model accuracy and model recall corresponding to each big data basic model; Step S23: performing a model F1 score evaluation and analysis on each big data basic model in the big data preliminary basic model set based on the model accuracy and model recall rate corresponding to each big data basic model, and obtaining a model F1 score corresponding to each big data basic model; Step S24: Calculate the model performance score of the corresponding big data basic models in the big data preliminary basic model set based on the model accuracy, model recall and model F1 score corresponding to each big data basic model, so as to obtain the model performance score value corresponding to each big data basic model; Step S25: Based on the model performance score corresponding to each big data basic model, the corresponding big data basic models in the big data preliminary basic model set are screened for the best model performance to obtain the big data best performance model.
6. The data modeling method according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: Perform model hyperparameter exploration and analysis on the model with the best big data performance to obtain a big data model hyperparameter exploration distribution matrix; Step S32: performing model hyperparameter combination design on the big data model hyperparameter exploration distribution matrix to generate different big data model hyperparameter combination conditions; Step S33: using the Bayesian optimization algorithm to perform model hyperparameter prior information modeling on the model with the best performance in big data to generate a model hyperparameter Bayesian optimization prior model; Step S34: performing hyperparameter combination optimization selection on different big data model hyperparameter combination conditions based on the model hyperparameter Bayesian optimization prior model to obtain a big data model optimized hyperparameter combination; Step S35: Based on the big data model optimization hyperparameter combination, the model with the best big data performance is optimized and adjusted to obtain a big data optimization and adjustment model.
7. The data modeling method according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: Real-time monitoring of feature updates of optimization features related to large amounts of data to obtain real-time update features of large amounts of data; Step S42: performing update feature flow analysis on the big data real-time update feature to obtain the big data real-time feature dynamic change trend data; Step S43: performing feature incremental mapping analysis on the real-time update feature of the big data based on the dynamic change trend data of the real-time feature of the big data, and obtaining the real-time feature mapping increment of the big data; Step S44: Adaptively update and optimize the big data optimization and adjustment model based on the big data real-time feature mapping increment using an incremental learning method to generate a big data adaptive adjustment model.
8. A data modeling system, characterized in that: Used to execute the data modeling method according to claim 1, the data modeling system comprises: The big data regularized recursive feature screening module is used to obtain the original big data set, and standardize the original big data set to obtain the standard big data set; perform feature automatic generation processing on the standard big data set to obtain the high-level combined features of big data; perform regularized recursive feature screening processing on the high-level combined features of big data, so as to obtain the optimization features related to the large data volume; The big data modeling performance screening module is used to perform preliminary modeling processing on the optimization features related to the large amount of data through decision trees, support vector machines and deep neural network basic algorithms to obtain a preliminary basic model set for big data; the preliminary basic model set for big data is screened for the best model performance to obtain the best big data performance model; The big data model hyperparameter optimization and adjustment module is used to optimize the hyperparameter combination of the model with the best big data performance using the Bayesian optimization algorithm to obtain the optimized hyperparameter combination of the big data model; based on the optimized hyperparameter combination of the big data model, the model with the best big data performance is optimized and adjusted to obtain the big data optimized adjustment model; The big data model adaptive update and adjustment module is used to perform real-time monitoring of feature updates of optimization features related to large amounts of data to obtain real-time update features of big data; based on the real-time update features of big data, the big data optimization adjustment model is adaptively updated and optimized using an incremental learning method to generate a big data adaptive adjustment model.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the data modeling method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Chip packaging design optimization method based on adaptive subproblem selection strategy
CN115062501A
Short-term power load prediction method based on non-standard Bayesian algorithm optimization
CN116316573A