Big data model intelligent optimization method based on machine learning
By quantitatively evaluating the preprocessing feature factors and training data features of big data models and dynamically adjusting optimization strategies, the problems of low efficiency and low accuracy caused by data quality issues in big data model training are solved, achieving high efficiency and robustness in model training.
Patent Information
- Application Number
- CN202511556818.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-01-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the current technology, during the training of big data models, data quality issues such as anomaly rate and missing rate lead to low training efficiency, low accuracy, and wasted resources, and there is a lack of real-time risk diagnosis and optimization strategies.
By acquiring the preprocessing feature parameters of the target platform's big data model over historical periods, analyzing the preprocessing feature factors, and combining them with the feature information of the training data, the model training risk is quantitatively assessed, the optimization strategy is dynamically adjusted, and the knowledge distillation compression model is used for optimization.
This improved the accuracy and efficiency of the model training process, reduced resource waste, enhanced the robustness and intelligence of model operation, and ensured the accuracy and efficiency of the optimization process.
Smart Images

Figure CN121436089A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence machine learning, in particular to a big data model intelligent optimization method based on machine learning. BACKGROUND
[0002] With the in-depth development of big data technology, the training process of machine learning models on target platforms is increasingly complex and uncertain when processing massive and high-dimensional data. The success of model training is highly dependent on the quality of input data and the setting of the training process. In practice, the original data often has problems such as high abnormal rate and high missing rate. If the features in the preprocessing stage are not effectively monitored and evaluated, noise and bias will be introduced, leading to deviation of model training from the expected result, and even failure, which constitutes a significant project risk.
[0003] At the same time, the structural complexity of the model itself, such as the number of network layers, and the number of training samples and the number of iterations of key parameters, jointly determine the capacity and convergence state of the model. If these training data features cannot be analyzed in conjunction with the quality of data preprocessing, it is difficult for the development team to identify whether the model optimization deviates from the right track and when intervention is needed in a timely and accurate manner within the training period. Traditional optimization methods often rely on manual experience for post-adjustment, lack of data-driven, forward-looking risk cycle judgment and automatic strategy generation mechanism, resulting in low optimization efficiency and huge resource consumption. Therefore, the industry urgently needs an intelligent optimization method that can deeply integrate risk assessment in the data preprocessing stage with feature analysis in the model training stage.
[0004] Chinese patent publication No. CN116578761A discloses a big data intelligent analysis method based on deep learning. The original data is obtained through a big data acquisition module and preprocessed. The data features are obtained using a deep learning model, and various types of feature vectors are selected and combined using a feature selection algorithm to obtain a prediction result. A deep learning network architecture based on an attention mechanism is used to train and classify data using a stacked autoencoder model, and a data compression algorithm is used for data analysis. The data is displayed in multiple visualization modes to show the analysis results. The present application provides comprehensive coverage of each link, provides a guarantee for the objectivity of the results, and uses excellent data processing tools such as data compression algorithms to provide a guarantee for the efficiency of data analysis. The final result is more clear and intuitive, and the practicality is further improved.
[0005] Chinese Patent Publication No. CN119003849A discloses a deep learning-based intelligent big data acquisition method. This method identifies and collects potential data sources using web crawler technology; classifies and identifies each data source and extracts corresponding key information; performs a depth-first search on the data sources to obtain the optimal data acquisition path; and constructs an associated knowledge graph to find the globally optimal acquisition strategy. This invention helps the system quickly locate and access potential data sources, improving data acquisition efficiency. It can dynamically adjust the search strategy based on search results, improving the performance and adaptability of data acquisition, making the acquired data more valuable and meaningful. It provides strong support for subsequent data processing and utilization, enhances the ability to understand data, helps users better understand the intrinsic meaning and value of data, and improves the efficiency and accuracy of data acquisition.
[0006] However, the following problems still exist in the existing technology: In existing technologies, during the training of big data models, the complex influence of data quality, including anomaly rate, missing rate, and training parameters, leads to real-time diagnosis of training risks, resulting in low model training efficiency, low accuracy of big data model operation, waste of resources, and substandard performance. Summary of the Invention
[0007] To address this, the present invention provides a machine learning-based intelligent optimization method for big data models, which overcomes the problems of existing technologies where data preprocessing and model training are separated, lack of quantitative assessment of data correlation risks such as abnormal and missing data rates, and inability to automatically identify model performance degradation caused by data problems during the training cycle, resulting in rigid training process, low resource efficiency, low model running efficiency, and low accuracy.
[0008] To achieve the above objectives, this invention provides an intelligent optimization method for big data models based on machine learning, comprising: Obtain preprocessing feature parameters from different time periods in the historical cycle of the target platform's big data model, and analyze preprocessing feature factors based on the preprocessing feature parameters; The training risk period for big data models is determined based on the comparison results between the preprocessed feature factors and the predetermined preprocessed feature factor thresholds within a time period. Based on a defined training risk period for a big data model, the training data feature information of the target platform's big data model is extracted, and the training data feature representation value is analyzed based on the training data feature information. The difference between the training data feature representation value and the predetermined training data feature representation threshold within the time period is used to determine whether the big data model optimization meets the standard. When it is determined that the optimization of the big data model does not meet the standard, the processing strategy for big data model optimization is determined based on the difference between the preprocessed feature factors of the target platform big data model and the predetermined preprocessed feature factor threshold: the adjustment range of the training feature representation threshold is determined to determine whether to enable the knowledge distillation compression model. The preprocessing feature parameters include anomaly rate and missing rate; the training data feature information includes the number of model network layers, the number of training samples, and the number of training iterations.
[0009] Preferably, the process of analyzing and preprocessing feature factors includes: Extract the anomaly rate and missing rate from different time periods in the historical data model of the target platform; The ratio of the anomaly rate in the historical period of the target platform's big data model to the predetermined anomaly rate threshold is determined as the first preprocessing feature factor; The ratio of the missing rate in the historical period of the target platform's big data model to the predetermined missing rate threshold is determined as the second preprocessing feature factor; The sum of the first preprocessing feature factor and the second preprocessing feature factor is determined as the preprocessing feature factor.
[0010] Furthermore, the process of determining the risk cycle of big data model training includes: The preprocessed feature factors within the time period are compared with the predetermined preprocessed feature factor thresholds; If the preprocessed feature factor is less than or equal to the predetermined preprocessed feature factor threshold, it is determined as a risk period for training big data models. If the preprocessed feature factor is greater than the predetermined preprocessed feature factor threshold, it is determined to be a period of significant risk in big data model training.
[0011] Preferably, the process of analyzing the feature representation values of the training data includes: Extract the number of network layers, number of training samples, and number of training iterations of the target platform's big data model within the time period; The ratio of the number of model network layers to a predetermined threshold for the number of model network layers is determined as the first training data feature representation factor; The ratio of the number of training samples to a predetermined threshold for the number of training samples is determined as the second training data feature representation factor. The ratio of the number of training iterations to a predetermined threshold number of training iterations is determined as the third training data feature representation factor. The sum of the first training data feature representation factor, the second training data feature representation factor, and the third training data feature representation factor is determined as the training data feature representation value.
[0012] Furthermore, the process of determining whether the big data model optimization meets the standards includes: Calculate the difference between the feature representation value of the training data and the predetermined feature representation threshold of the training data within the time period; If the difference is less than or equal to a predetermined training feature difference threshold, then the big data model optimization is determined to meet the standard. If the difference is greater than the predetermined training feature difference threshold, then the big data model optimization is determined to be non-compliant with the standard.
[0013] Furthermore, the process of determining the processing strategy for optimizing big data models includes: Calculate the difference between the preprocessed feature factors of the target platform's big data model and the predetermined preprocessed feature factor thresholds; If the difference is less than or equal to a predetermined preprocessing feature difference threshold, then the processing strategy for optimizing the big data model is determined to be the adjustment range of the training feature representation threshold. If the difference is greater than the predetermined preprocessing feature difference threshold, then the processing strategy for big data model optimization is determined to be to enable the knowledge distillation compression model.
[0014] Furthermore, the adjustment range of the training feature representation threshold is determined based on the difference between the preprocessed feature factors of the target platform big data model and the predetermined preprocessed feature factor threshold.
[0015] Furthermore, the anomaly rate is determined based on the ratio of the number of outliers to the number of samples in the target platform's big data model within a single time period.
[0016] Furthermore, the missing value rate is determined based on the ratio of the number of missing values to the number of samples in the target platform's big data model within a single time period.
[0017] Furthermore, the knowledge distillation compression model includes outlier removal optimization and missing value imputation optimization.
[0018] Compared with existing technologies, this invention can accurately analyze preprocessing feature factors by acquiring preprocessing feature parameters from different time periods within the historical cycle of the target platform's big data model. This allows for the determination of the big data model training risk period. Within the determined risk period, the invention can quickly analyze training data feature values by extracting training data feature information from the target platform's big data model. The difference between the training data feature values within the time period and a predetermined training data feature threshold determines whether the big data model optimization meets the standard. If it does not meet the standard, the invention determines the big data model optimization processing strategy based on the difference between the preprocessing feature factors of the target platform's big data model and the predetermined preprocessing feature factor threshold; it also determines the adjustment range of the training feature threshold to enable the knowledge distillation compression model. This overcomes the shortcomings of existing technologies, which separate data preprocessing from model training, resulting in inconsistent data correlation quantitative assessment of model data anomaly and missing rate risks. This leads to the inability to automatically identify problems caused by data issues within the training period, such as rigid training processes, low resource efficiency, low model operating efficiency, and low accuracy. This improves the accuracy and efficiency of big data model optimization.
[0019] In particular, this invention proactively identifies and warns of training risk cycles by tracking key indicators such as anomaly rate and missing rate in real time during the data preprocessing stage. Within the risk cycle, it can further analyze training characteristics such as the number of model network layers and the number of training samples. By quantifying the gap between its representation value and the expected standard, it can accurately diagnose the degree to which the model optimization deviates from the standard. When the optimization is determined to be non-compliant with the standard, the system dynamically determines the optimization strategy based on the difference between the data quality preprocessing feature factor and the threshold. This enables on-demand allocation and precise control of resources, reducing the need to blindly invest huge computing power in full training when the data quality is poor. While ensuring model performance, it significantly saves computing costs and time, and improves the robustness and intelligence of the model training process overall.
[0020] Furthermore, this invention achieves accurate quantification and comprehensive evaluation of data quality issues through preprocessing feature factors. By unifying two different data quality issues, anomaly rate and missing rate, onto a comparable scale, it provides a holistic data health index, enabling the system to have a clear basis for judging complex data quality conditions. Based on clear numerical judgments, it overcomes the shortcomings of traditional fuzzy evaluations that rely on experience, making risk warnings more objective and automated. The optimization strategy for model training can be directly correlated with preprocessed feature data, thereby reducing the occurrence of model risks and improving the robustness and intelligence level of the entire model training system.
[0021] Furthermore, this invention distinguishes the training cycle into a risky big data model training cycle and a significant big data model training cycle by comparing comprehensive preprocessing feature factors with predetermined preprocessing feature factor thresholds. This ensures the objectivity and consistency of decision-making and provides a key basis for subsequent differentiated response strategies. The risky big data model cycle triggers the extraction and analysis of training data features for deep learning diagnosis; while the significant big data model training cycle determines the activation of the knowledge distillation and compression model. This ensures that limited computing power and human resources can be prioritized for the most critical issues, thereby achieving precision and efficiency in the optimization process and effectively reducing resource waste and over-investment in low-risk tasks.
[0022] In particular, this invention constructs a comprehensive model training status evaluation index by quantifying key training parameters. Through a unified multi-dimensional measurement of the model training process, it transforms three core features—model complexity (number of network layers), data scale (number of samples), and learning process (number of iterations)—into comparable and superimposed factors through normalization. This invention enables parameters to be integrated into a single, comprehensive training data feature representation value, thereby providing a comprehensive and standardized quantitative basis for evaluating the training status. By integrating complex training status into a transferable numerical value, it ensures the systematicness, accuracy, and efficiency of model optimization decisions, effectively improving the utilization efficiency of training resources and the accuracy of model output. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating the steps of the intelligent optimization method for big data models based on machine learning, as described in an embodiment of the present invention. Figure 2 This is a flowchart illustrating the steps of analyzing and preprocessing feature factors in an embodiment of the present invention. Figure 3 A logical decision diagram for determining the risk cycle of big data model training in this embodiment of the invention; Figure 4 This is a flowchart illustrating the steps involved in analyzing the feature representation values of training data in an embodiment of the present invention. Detailed Implementation
[0024] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0025] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these systematic embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0026] It should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the term "connection" should be interpreted broadly. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0027] Please see Figure 1 The diagram shows the steps of an intelligent optimization method for big data models based on machine learning, according to an embodiment of the present invention. The present invention provides an intelligent optimization method for big data models based on machine learning, comprising: Step S1: Obtain the preprocessing feature parameters of the target platform's big data model in different time periods of its historical cycle, and analyze the preprocessing feature factors based on the preprocessing feature parameters; Step S2: Determine the risk period for training the big data model based on the comparison results between the preprocessed feature factors and the predetermined preprocessed feature factor thresholds within the time period; Step S3: Based on the determined training risk period of the big data model, extract the training data feature information of the target platform big data model, and analyze the training data feature representation value based on the training data feature information; Step S4: Determine whether the big data model optimization meets the standard based on the difference between the training data feature representation value and the predetermined training data feature representation threshold within the time period; When it is determined that the optimization of the big data model does not meet the standard, the processing strategy for big data model optimization is determined based on the difference between the preprocessed feature factors of the target platform big data model and the predetermined preprocessed feature factor threshold: the adjustment range of the training feature representation threshold is determined to determine whether to enable the knowledge distillation compression model. The preprocessing feature parameters include anomaly rate and missing rate; the training data feature information includes the number of model network layers, the number of training samples, and the number of training iterations.
[0028] It is understood that the historical period is not limited. In this embodiment, the historical period is preferably selected within the range of [1 month, 3 months]. In this embodiment, the historical period is preferably 1 month, which will not be elaborated further.
[0029] It is understood that the time period is not limited. In this embodiment, the time period is preferably selected within the range of [24 hours, 48 hours]. In this embodiment, the time period is preferably 24 hours, which will not be elaborated further.
[0030] It is understood that there are no restrictions on the target platform's big data model; it can be a convolutional neural network model or a recurrent neural network model. The implementation example is preferably a convolutional neural network model, which will not be elaborated further here.
[0031] The implementation example proactively identifies and warns of training risk cycles by tracking key indicators such as anomaly rate and missing rate during the data preprocessing stage in real time. Within the risk cycle, it can further analyze training characteristics such as the number of model network layers and the number of training samples. By quantifying the gap between its representation value and the expected standard, it can accurately diagnose the degree to which the model optimization deviates from the standard. When the optimization is determined to be non-compliant with the standard, the system dynamically determines the optimization strategy based on the difference between the data quality preprocessing feature factor and the threshold. This enables on-demand allocation and precise control of resources, reducing the need to blindly invest huge computing power in full training when the data quality is poor. At the same time, while ensuring model performance, it significantly saves computing costs and time, and improves the robustness and intelligence of the model training process overall.
[0032] The application scenarios for this example are not limited. This example is applied to model maintenance and A / B testing on large-scale internet platforms. It is preferably used in scenarios such as e-commerce, content recommendation, and advertising, where platforms need to continuously train and update massive amounts of models. As an automated monitoring system, this example determines in real time whether ongoing model training faces risks due to sudden changes in the quality of newly fed back data and automatically triggers intervention measures, thereby ensuring the stability and iteration efficiency of online model services.
[0033] Please see Figure 2 The diagram shown is a flowchart illustrating the steps of the preprocessing feature factor analysis process according to an embodiment of the present invention. The preprocessing feature factor analysis process in the embodiment includes: Extract the anomaly rate and missing rate from different time periods in the historical data model of the target platform; The ratio of the anomaly rate in the historical period of the target platform's big data model to the predetermined anomaly rate threshold is determined as the first preprocessing feature factor; The ratio of the missing rate in the historical period of the target platform's big data model to the predetermined missing rate threshold is determined as the second preprocessing feature factor; The sum of the first preprocessing feature factor and the second preprocessing feature factor is determined as the preprocessing feature factor.
[0034] It is understood that the predetermined anomaly rate threshold is predetermined. Specifically, the anomaly rate threshold is determined by multiplying the average anomaly rate of the target platform's big data model in each time period within the month prior to system operation by the accuracy coefficient. The accuracy coefficient is selected within the range [0.95, 0.99], and in this embodiment, the accuracy coefficient is preferably 0.98.
[0035] Similarly, the predetermined missing rate threshold is determined in advance. Specifically, the missing rate threshold is determined by multiplying the average missing rate of the target platform big data model in each time period within the month before the system runs by the precision coefficient. The precision coefficient is selected in the range [0.95, 0.99]. In this embodiment, the precision coefficient is preferably 0.98.
[0036] The implementation example achieves accurate quantification and comprehensive evaluation of data quality issues by preprocessing feature factors. By unifying two different data quality issues, anomaly rate and missing rate, onto a comparable scale, it provides a holistic data health index, enabling the system to have a clear basis for judging complex data quality conditions. Based on clear numerical judgments, it overcomes the shortcomings of traditional experience-based fuzzy evaluations, making risk warnings more objective and automated. The optimization strategy for model training can be directly correlated with preprocessed feature data, thereby reducing the occurrence of model risks and improving the robustness and intelligence level of the entire model training system.
[0037] Please see Figure 3 As shown, this is a logic decision diagram for determining the risk period of big data model training according to an embodiment of the present invention. The process of determining the risk period of big data model training in the embodiment includes: The preprocessed feature factors are compared with the predetermined preprocessed feature factor thresholds within the time period. If the preprocessed feature factor is less than or equal to the predetermined preprocessed feature factor threshold, it is determined as a risk period for training big data models. If the preprocessed feature factor is greater than the predetermined preprocessed feature factor threshold, it is determined to be a period of significant risk in big data model training.
[0038] Specifically, when an implementation is determined to be a period with significant risks in training a big data model, the knowledge distillation compression model is activated.
[0039] It is understood that the predetermined preprocessing feature factor threshold is predetermined. Specifically, the average value of the preprocessing feature factors of the target platform big data model in the historical period of the month before the system operation is extracted, and the product of the average value and the accuracy coefficient is determined as the preprocessing feature factor threshold. The accuracy coefficient is selected in the range [0.95, 0.99]. In a preferred embodiment, the preprocessing feature factor threshold is selected in the range [2.01, 2.16]. In a preferred embodiment, the preprocessing feature factor threshold is 2.10.
[0040] The implementation example compares comprehensive preprocessing feature factors with predetermined preprocessing feature factor thresholds to clearly distinguish the training cycle into a risky big data model training cycle and a significant risk big data model training cycle. This ensures the objectivity and consistency of decision-making and provides a key basis for subsequent differentiated response strategies. The risky big data model cycle triggers the extraction and analysis of training data features for deep learning diagnosis; while the significant risk big data model cycle determines the activation of the knowledge distillation and compression model. This ensures that limited computing power and human resources can be prioritized for the most critical issues, thereby achieving precision and efficiency in the optimization process and effectively reducing resource waste and over-investment in low-risk tasks.
[0041] Please see Figure 4 The diagram shows a flowchart illustrating the steps involved in analyzing the feature representation values of training data according to an embodiment of the present invention. The process of analyzing the feature representation values of training data in this embodiment includes: Extract the number of network layers, number of training samples, and number of training iterations of the target platform's big data model within the time period; The ratio of the number of model network layers to a predetermined threshold for the number of model network layers is determined as the first training data feature representation factor; The ratio of the number of training samples to a predetermined threshold for the number of training samples is determined as the second training data feature representation factor. The ratio of the number of training iterations to a predetermined threshold for the number of training iterations is determined as the third training data feature representation factor. The sum of the first training data feature representation factor, the second training data feature representation factor, and the third training data feature representation factor is determined as the training data feature representation value.
[0042] It is understood that the predetermined threshold for the number of model network layers is predetermined. Specifically, the threshold for the number of model network layers is determined by multiplying the average number of model network layers of the target platform's big data model within one month before the system runs by the offset coefficient. The offset coefficient is selected within the range [1.05, 1.15], and in this embodiment, the offset coefficient is preferably 1.10.
[0043] The predetermined threshold for the number of training samples is determined in advance. Specifically, the threshold for the number of training samples is determined by multiplying the average number of training samples of the target platform's big data model within one month before the system runs by the offset coefficient. The offset coefficient is selected within the range [1.05, 1.15], and in this embodiment, the offset coefficient is preferably 1.10.
[0044] Similarly, the predetermined threshold for the number of training iterations is determined in advance. Specifically, the threshold for the number of training iterations is determined by multiplying the average number of training iterations of the target platform's big data model within one month before the system runs with the offset coefficient. The offset coefficient is selected within the range [1.05, 1.15]. In this embodiment, the offset coefficient is preferably 1.10.
[0045] This embodiment constructs a comprehensive model training status evaluation index by quantifying key training parameters. Through a unified multi-dimensional measurement of the model training process, it transforms three core features—model complexity (number of network layers), data scale (number of samples), and learning process (number of iterations)—into comparable and superimposed factors through normalization. This invention enables parameters to be integrated into a single, comprehensive training data feature representation value, thereby providing a comprehensive and standardized quantitative basis for evaluating the training status. By integrating complex training status into a transferable numerical value, it ensures the systematicness, accuracy, and efficiency of model optimization decisions, effectively improving the utilization efficiency of training resources and the accuracy of model output.
[0046] Specifically, the process for determining whether a big data model optimization conforms to the standard includes: Calculate the difference between the feature representation value of the training data and the predetermined feature representation threshold of the training data within the time period; If the difference is less than or equal to a predetermined training feature difference threshold, then the big data model optimization is determined to meet the standard. If the difference is greater than the predetermined training feature difference threshold, then the big data model optimization is determined to be non-compliant with the standard.
[0047] It is understood that the predetermined training feature representation threshold is predetermined. In this embodiment, the training feature representation threshold is selected within the range [3.05, 3.35], and preferably the training feature representation threshold is 3.15.
[0048] It is understood that the training feature difference threshold is predetermined, and in this embodiment, the training feature difference threshold is selected within the range [0.05, 0.25].
[0049] This embodiment provides an objective and quantitative judgment mechanism for determining whether the optimization of a big data model meets the standards, thereby realizing automated decision-making and closed-loop management of the optimization process. By comparing the difference with an allowable difference threshold error range through comprehensive training data feature representation values, the system simplifies a multi-dimensional and complex evaluation problem into a clear yes-or-no question, namely, optimization that meets or does not meet the standards. The computer system automatically performs the judgment. When it is determined to meet the standards, the training process can continue normally or be approved for deployment; when it is determined to not meet the standards, knowledge distillation is immediately activated, thereby ensuring that the system can make optimization strategies in a timely manner for unsatisfactory training states. By setting clear benchmarks, the waste of computing resources is reduced, and the accuracy of the final output model is ensured, promoting the standardization and stability of the model development process.
[0050] Specifically, the process of determining the processing strategy for optimizing the big data model in the embodiment includes: Calculate the difference between the preprocessed feature factors of the target platform's big data model and the predetermined preprocessed feature factor thresholds; If the difference is less than or equal to a predetermined preprocessing feature difference threshold, then the processing strategy for optimizing the big data model is determined to be the adjustment range of the training feature representation threshold. If the difference is greater than the predetermined preprocessing feature difference threshold, then the processing strategy for big data model optimization is determined to be to enable the knowledge distillation compression model.
[0051] It is understood that the implementation example selects the preprocessing feature difference threshold within the range [0.05, 0.45], and preferably the preprocessing feature difference threshold is 0.25.
[0052] This implementation precisely links the severity of data quality issues with optimization strategies, enabling intelligent and tiered model optimization intervention, thereby improving the accuracy of strategy selection and the efficiency of resource utilization. By constructing a tiered processing strategy based on root cause analysis, the severity of data quality issues is quantified. When the preprocessed feature difference is less than or equal to a predetermined preprocessed feature difference threshold, the problem is determined to be in the model structure or training parameters; therefore, the strategy is to adjust the training feature representation threshold, i.e., to fine-tune and optimize the model's training process. Conversely, when the difference exceeds the predetermined preprocessed feature difference threshold, the root cause is determined to be in the data itself; in this case, a knowledge distillation and compression model is activated to obtain a relatively optimal lightweight model.
[0053] The implementation plan effectively avoids resource waste through a hierarchical strategy. For minor issues, low-cost parameter adjustments are used; for serious problems, more economical and robust model compression techniques are employed. The system achieves on-demand allocation and precise deployment of computing resources, thereby saving overall computing costs and time, ensuring the agility and economy of the model training process, and improving the efficiency of optimization decisions and resource utilization.
[0054] Specifically, the adjustment range of the training feature representation threshold described in the embodiment is based on the determination of the difference between the preprocessed feature factors of the target platform big data model and the predetermined preprocessed feature factor threshold.
[0055] The implementation example dynamically correlates data quality with model training thresholds, achieving precision and adaptability in optimization strategies, significantly improving the method's intelligence and final effectiveness. The system creates a precise quantitative feedback loop from data problems to solutions. Based on the difference between preprocessed feature factors and thresholds, it dynamically determines the adjustment range, ensuring the accuracy of intervention, reducing under- or over-adjustment, and perfectly matching optimization measures with the severity of the problem. Secondly, it endows the system with strong adaptability, enabling it to cope with scenarios of varying severity and data conditions. This ensures that the entire optimization process effectively corrects model training biases caused by data problems while utilizing resources in the most efficient way, thus maintaining high robustness and excellent performance in complex real-world application environments.
[0056] Specifically, the anomaly rate described in the embodiment is determined based on the ratio of the number of outliers to the number of samples in the target platform's big data model within a single time period.
[0057] Specifically, an excessively high anomaly rate in the implementation examples indicates a problem in the data acquisition or transmission process. If not properly corrected or eliminated, it will interfere with model training, causing the model to learn incorrect patterns, overfit outliers, and reduce prediction accuracy.
[0058] This implementation example establishes an objective and comparable foundation for data quality assessment throughout the entire model optimization process by precisely determining the anomaly rate. By transforming the vague concept of data anomalies into a precise and calculable mathematical indicator, it enables fair and consistent measurement and comparison of anomaly levels across datasets of different time periods and sizes, providing reliable and unified input for all subsequent automated analyses. Furthermore, even with drastic fluctuations in the total data volume across different periods, the calculated anomaly rates remain comparable. The system can therefore accurately identify risk periods where data quality truly deteriorates. Moreover, by calculating preprocessing feature factors, it influences the determination of risk periods and the generation of the final optimization strategy. This ensures that problems discovered at the data end are accurately and effectively transmitted to the decision-making stage at the model end, improving the scientific rigor and accuracy of the closed-loop optimization system.
[0059] Specifically, the missing value rate described in the embodiment is determined based on the ratio of the number of missing values to the number of samples in the target platform's big data model within a single time period.
[0060] Specifically, an excessively high missing value rate in the implementation examples can lead to a reduction in the amount of effective data. If the missing value mechanism is random, it may reduce the amount of information input to the model; if the missing value is regular, including data missing values in specific scenarios, it may introduce bias. Preprocessing typically involves handling missing values through mean imputation, model prediction imputation, or removal, and the missing value rate is the basis for evaluating the effectiveness of the processing strategy.
[0061] The precise definition of the missing rate in the implementation example, complementing the quantification of the outlier rate, forms the core pillar of assessing data quality and integrity, achieving a standardized measure of data missingness issues. By calculating the missing rate as the ratio of the number of missing values to the sample size, the assessment of data integrity is transformed from a qualitative description into an objective, calculable numerical indicator. This ensures consistency and fairness in judging the severity of data missingness across different time periods or data scales, providing clear and unambiguous judgment criteria for automated systems. The system can accurately identify the periods when data integrity truly deteriorates, thus guaranteeing the accuracy and comparability of risk assessments. The accurate quantification of data integrity issues and their effective transmission throughout the optimization chain make it an indispensable part of data-driven decision-making. Therefore, the precise missing rate indicator can serve as a crucial basis for triggering subsequent intelligent decision-making processes, improving the system's accuracy and efficiency.
[0062] Specifically, the knowledge distillation compression model described in the embodiments includes outlier removal optimization and missing value filling optimization.
[0063] Specifically, this embodiment combines a knowledge distillation compression model with specific data optimization operations—outlier removal and missing value imputation—creating a highly targeted and efficient composite optimization strategy. The system achieves a strategy upgrade from passive avoidance to proactive repair. While traditional knowledge distillation focuses solely on making smaller models mimic the output of larger models, this embodiment explicitly incorporates outlier removal and missing value imputation as part of the distillation process. Simultaneously generating a lightweight compressed model, the system proactively cleans and repairs problematic data, rather than simply ignoring or bypassing it. This fundamentally improves the data quality used to train the smaller model, laying a solid foundation for producing a more robust compressed model. The embodiment significantly improves the reliability and performance of the compressed model in real-world scenarios. By embedding data optimization steps in the distillation process, the resulting compressed model is built on a more complete data representation, reducing the inheritance of noise and bias from the original data into the new model. This ensures that the compressed model possesses stronger generalization ability and higher prediction accuracy when facing real-world application data.
[0064] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A method for intelligent optimization of big data models based on machine learning, characterized in that, The method comprises the following steps: acquiring pre-processing characteristic parameters of a target platform big data model in different time periods of a historical period, and analyzing pre-processing characteristic factors based on the pre-processing characteristic parameters; determining a big data model training risk period based on a comparison result of the pre-processing characteristic factors and a predetermined pre-processing characteristic factor threshold value in the time period; extracting training data characteristic information of the target platform big data model based on the determined big data model training risk period, and analyzing training data characteristic representation values based on the training data characteristic information; determining whether big data model optimization meets a standard based on a difference between the training data characteristic representation values and a predetermined training data characteristic representation threshold value in the time period; when it is determined that the big data model optimization does not meet the standard, determining a processing strategy for the big data model optimization based on a difference between the pre-processing characteristic factors of the target platform big data model and the predetermined pre-processing characteristic factor threshold value, determining an adjustment range of the training characteristic representation threshold value, and then determining whether to enable a knowledge distillation compression model. The pre-processing characteristic parameters include an abnormality rate and a missing rate, and the training data characteristic information includes a model network layer number, a training sample number, and a training iteration number. 2.The method of claim 1, wherein, The process of analyzing the pre-processing characteristic factors comprises the following steps: extracting the abnormality rate and the missing rate of the target platform big data model in different time periods of a historical period; calculating a ratio of the abnormality rate of the target platform big data model in the historical period to a predetermined abnormality rate threshold value to determine a first pre-processing characteristic factor; calculating a ratio of the missing rate of the target platform big data model in the historical period to a predetermined missing rate threshold value to determine a second pre-processing characteristic factor; determining the pre-processing characteristic factor as a sum of the first pre-processing characteristic factor and the second pre-processing characteristic factor. 3.The method of claim 1, wherein, The process of determining the big data model training risk period comprises the following steps: comparing the pre-processing characteristic factor with the predetermined pre-processing characteristic factor threshold value in the time period; if the pre-processing characteristic factor is less than or equal to the predetermined pre-processing characteristic factor threshold value, determining a big data model training risk period; if the pre-processing characteristic factor is greater than the predetermined pre-processing characteristic factor threshold value, determining a big data model training risk significant period. 4.The method of claim 1, wherein, The process of analyzing the training data characteristic representation values comprises the following steps: extracting a model network layer number, a training sample number, and a training iteration number of the target platform big data model in the time period; determining a first training data characteristic representation factor as a ratio of the model network layer number to a predetermined model network layer number threshold value; determining a second training data characteristic representation factor as a ratio of the training sample number to a predetermined training sample number threshold value; determining a third training data characteristic representation factor as a ratio of the training iteration number to a predetermined training iteration number threshold value; determining the training data characteristic representation value as a sum of the first training data characteristic representation factor, the second training data characteristic representation factor, and the third training data characteristic representation factor. 5.The method of claim 4, wherein, The process of determining whether the big data model optimization meets the standard comprises the following steps: calculating a difference between the training data characteristic representation value and the predetermined training data characteristic representation threshold value in the time period; If the difference value is less than or equal to a predetermined training feature difference threshold value, it is determined that the big data model optimization meets the standard; If the difference value is greater than the predetermined training feature difference threshold value, it is determined that the big data model optimization does not meet the standard. 6.The method of claim 5, wherein, The process of determining the processing strategy of big data model optimization includes: calculating the difference between the pre-processing feature factor of the target platform big data model and the predetermined pre-processing feature factor threshold value; If the difference value is less than or equal to a predetermined pre-processing feature difference threshold value, it is determined that the processing strategy of big data model optimization is to determine the adjustment range of the training feature representation threshold value; If the difference value is greater than the predetermined pre-processing feature difference threshold value, it is determined that the processing strategy of big data model optimization is to enable the knowledge distillation compression model. 7.The method of claim 6, wherein, The adjustment range of the training feature representation threshold value is determined based on the difference between the pre-processing feature factor of the target platform big data model and the predetermined pre-processing feature factor threshold value. 8.The method of claim 2, wherein, The abnormal rate is determined based on the ratio of the number of abnormal values to the number of samples of the target platform big data model in a single time period. 9.The method of claim 2, wherein, The missing rate is determined based on the ratio of the number of missing values to the number of samples of the target platform big data model in a single time period. 10.The method of claim 6, wherein, The knowledge distillation compression model includes abnormal value elimination optimization and missing value filling optimization.
Citation Information
Patent Citations
Big data intelligent analysis method based on deep learning
CN116578761A
Big data intelligent acquisition method based on deep learning
CN119003849A
Model training method, terminal equipment and server
CN114091101A
Optimization method based on industrial energy consumption big data model
CN119204357A
Automatic model adjusting and optimizing system based on AI training intelligent workbench
CN119398112A