Data processing method and system for energy and carbon prediction hybrid model of data center

By using ETL technology and model priority scheduling algorithm to process heterogeneous and dynamic data in data center carbon emission forecasting, efficient incremental calculation and multi-model collaborative fusion are achieved, solving the real-time and accuracy issues of data center carbon emission forecasting, and improving forecast accuracy and efficiency.

CN120671878AInactive Publication Date: 2025-09-19WUXI SHENGHECAI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312284.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies have problems with real-time and accuracy in data center carbon emissions prediction, especially when processing massive, heterogeneous and dynamic data, making it difficult to achieve efficient incremental computing and multi-model collaborative fusion.

Method used

ETL technology is used to convert heterogeneous data into a unified format for real-time collection and processing, extract features and perform multi-source data correlation analysis, use model priority scheduling algorithm to optimize call sequences, and perform weighted fusion and outlier filtering to ensure the reliability and interpretability of prediction results.

Benefits of technology

It improves the accuracy and efficiency of data center carbon emission prediction, ensures the reliability and interpretability of prediction results, adapts to the dynamic changes of data centers, and optimizes resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671878A_ABST
    Figure CN120671878A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and system of a data center energy and carbon prediction hybrid model. The method comprises the steps of collecting original heterogeneous data related to carbon emission; preprocessing the data so as to clean noise and irrelevant data; key features are extracted from the preprocessed data to form a feature set; adapting the extracted features to input format requirements of a plurality of models; fusing the matched feature set to form a comprehensive multi-model feature; dividing the data into a plurality of dynamic batches according to the data distribution of the fusion features; performing model training and prediction on the data of each dynamic batch, and outputting a prediction result of each model; and performing weighted fusion and abnormal value filtering on prediction results of all the models to obtain a final prediction result. The method can improve the accuracy of energy and carbon prediction of the data center.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a data processing method and system for a data center energy-carbon prediction hybrid model. Background Art

[0002] Hybrid models for data center carbon emissions prediction face the challenge of processing and analyzing massive amounts of carbon emissions-related data. This data often comes from diverse monitoring devices and systems and is heterogeneous and dynamic. To improve the real-time and accuracy of carbon emissions predictions, incremental processing and fusion calculations of this data are necessary. However, due to the massive and constantly changing data volume, traditional batch processing methods are unable to meet the demands of real-time predictions. Furthermore, the correlations and dependencies between different data sources pose challenges for incremental computing. Ensuring data consistency while achieving efficient incremental computing and dynamic resource scheduling has become a pressing technical challenge. Furthermore, since carbon emissions prediction involves expertise and models from multiple fields, achieving multi-model collaboration and fusion within the incremental computing process presents a complex technical challenge. This requires consideration of interface definition between different models, data format conversion, and call sequence optimization to ensure accurate and interpretable prediction results.

[0003] In existing technologies, there are already a variety of prediction models, such as linear regression, gray prediction, and random forest. These models can be used to predict population and GDP. Then, a multivariate linear regression model is used to describe the relationship between carbon emissions and population, GDP, and energy consumption, as well as the relationship between carbon emissions and various energy consumption sectors and energy types. Based on the prediction results, the most appropriate path and measures can be determined. However, these models have limitations when dealing with data center carbon emissions prediction. Since they cannot meet the needs of real-time prediction, the real-time and accuracy of the prediction results are insufficient.

[0004] In summary, existing technologies have shortcomings in processing data center carbon emission prediction, especially in terms of real-time performance and accuracy.

[0005] In summary, in the hybrid model of data center carbon emission prediction, the key challenges include processing massive, heterogeneous and dynamic data from different sources, and achieving real-time and accurate prediction. Summary of the Invention

[0006] The present invention provides a data processing method and system for a data center energy-carbon forecasting hybrid model, which effectively solves key problems in carbon emission data processing and forecasting, and improves forecasting accuracy and efficiency.

[0007] In a first aspect, to solve the above technical problems, the present invention provides a data processing method and system for a hybrid model for data center energy and carbon prediction, comprising:

[0008] Obtaining raw heterogeneous data related to carbon emissions;

[0009] Based on the original heterogeneous data, ETL technology is used to convert data from different sources into a unified format to obtain a standardized data set;

[0010] Collect and process new data in real time to obtain incremental data sets;

[0011] Extract features from the standardized dataset and the incremental dataset and perform multiple processing, use a feature fusion algorithm to perform correlation analysis on multi-source data to obtain a fused feature set;

[0012] Dynamically dividing the fused feature set into batches according to data feature distribution to obtain an adapted data set;

[0013] Using a model priority scheduling algorithm, a comprehensive priority score is calculated based on the model's computational complexity, prediction accuracy, and resource requirements, and an optimized call sequence of the model is obtained based on the comprehensive priority score;

[0014] The adapted data set is input into the model in the optimized call sequence, and the prediction results of the model are weighted fused and outlier filtered to obtain the final prediction results.

[0015] In an optional embodiment, based on the original heterogeneous data, ETL technology is used to convert data from different sources into a unified format to obtain a standardized data set, including:

[0016] By using ETL technology to clean, transform and load the original heterogeneous data, the original heterogeneous data of different sources and formats are converted into a unified standardized format to obtain a standardized data set;

[0017] The quality of the standardized data set is checked by setting a data quality threshold. If the data quality threshold is not reached, it is marked and corrected.

[0018] In an optional embodiment, the newly added data is collected and processed in real time to obtain an incremental data set, including:

[0019] Perform data cleaning, denoising, and normalization operations on the newly collected incremental data to obtain a standardized incremental data set;

[0020] In an optional embodiment, features are extracted from the standardized dataset and the incremental dataset and multiple processing is performed, and a feature fusion algorithm is used to perform correlation analysis on the multi-source data to obtain a fused feature set, including:

[0021] Extracting key features from the standardized dataset and the incremental dataset;

[0022] Use feature selection algorithms to screen out key feature subsets that have a significant impact on the target task, reduce feature dimensions, and obtain dimension-reduced key feature subsets;

[0023] Analyzing the correlation and redundancy between features in the dimensionality-reduced key feature subset, and eliminating highly correlated and redundant features based on the correlation coefficient matrix and mutual information index to obtain a simplified key feature subset;

[0024] According to the feature type and value range, a normalization method is used to map the features of different scales of the simplified key feature subset to the same scale to obtain a regular key feature subset;

[0025] According to the rule-based key feature subsets, a suitable feature fusion algorithm is selected to learn the association rules within and between feature subsets to obtain a fused feature set.

[0026] In an optional embodiment, the fused feature set is dynamically divided into batches according to data feature distribution to obtain an adapted data set, including:

[0027] According to the input data format requirements of each model, a data conversion method is used to convert the fused feature set into an input data format suitable for each model;

[0028] According to the data feature distribution of the fused feature set, a dynamic batch partitioning algorithm is used to divide the fused feature set into multiple dynamic batches, ensuring that the data feature distribution of each batch is similar, thereby obtaining an adapted data set;

[0029] In an optional embodiment, a model priority scheduling algorithm is used to calculate a comprehensive priority score based on the computational complexity, prediction accuracy, and resource requirements of the model, and an optimized call sequence of the model is obtained based on the comprehensive priority score, including:

[0030] Obtain a list of models to be scheduled, including each model's computational complexity, prediction accuracy, and resource requirements;

[0031] Determine the weight coefficients of model attributes based on business requirements and system resource conditions;

[0032] Calculate the comprehensive priority score for each model;

[0033] The comprehensive priority score is calculated using the following formula:

[0034] S=C×W C +P×W P +R×W R

[0035] In the formula, S is the comprehensive priority score, C is the computational complexity, and W is the C is the complexity weight, P is the prediction accuracy, W P is the precision weight, R is the resource requirement, W R is the resource weight;

[0036] Sorting the model list in descending order according to the comprehensive priority score to obtain an optimized call sequence;

[0037] In an optional embodiment, the adapted data set is input into the model in the optimized call sequence, and the prediction results of the model are weighted fused and outlier filtered to obtain the final prediction results, including:

[0038] Inputting the adapted data set into the model in the optimized call sequence, each base model outputs a preliminary prediction result;

[0039] According to the weight coefficient of each model, the preliminary prediction results are weighted averaged, and the isolation forest algorithm is used to filter outliers to obtain the final prediction results of each model;

[0040] The final prediction results of each model are compared with the actual business results, the prediction error is calculated, and the prediction error is fed back to each model. The model parameters are continuously optimized through online learning to improve the prediction accuracy of the model.

[0041] In a second aspect, the present invention provides a data processing method and system for a data center energy-carbon prediction hybrid model, comprising:

[0042] Data acquisition module, used to obtain raw heterogeneous data related to carbon emissions;

[0043] An ETL operation module is used to convert data from different sources into a unified format based on the original heterogeneous data using ETL technology to obtain a standardized data set;

[0044] The incremental data processing module is used to collect and process the new data in real time to obtain the incremental data set;

[0045] A feature extraction and fusion operation module is used to extract features from the standardized data set and the incremental data set and perform multiple processing, and use a feature fusion algorithm to perform correlation analysis on multi-source data to obtain a fused feature set;

[0046] A dynamic batch division module is used to dynamically divide the fused feature set into batches according to the data feature distribution to obtain an adapted data set;

[0047] A model priority scheduling module is used to use a model priority scheduling algorithm to calculate a comprehensive priority score based on the model's computational complexity, prediction accuracy, and resource requirements, and to obtain an optimized call sequence for the model based on the comprehensive priority score;

[0048] The result output module is used to input the adapted data set into the model in the optimized call sequence, and perform weighted fusion and outlier filtering on the prediction results of the model to obtain the final prediction result.

[0049] In a third aspect, the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned data center energy-carbon forecasting hybrid model data processing methods and systems.

[0050] Compared with the prior art, the present invention has the following beneficial effects: The present invention discloses a data processing method and system for a hybrid model of data center energy carbon prediction. In response to the challenges of data heterogeneity, dynamics and large-scale real-time processing in carbon emission monitoring, the present invention adopts a unified format conversion and time window incremental processing mechanism to achieve standardization and real-time collection of multi-source data. Through the streaming computing framework and dynamic resource allocation, the present invention efficiently processes incremental data. In the feature fusion stage, the present invention considers the correlation of multi-source data, adopts normalization and standardization processing, and selects a suitable fusion strategy. In order to meet the needs of multi-model collaboration, the present invention performs data adaptation and dynamic batch division. In the prediction stage, the present invention adopts a model priority scheduling algorithm to optimize the call order, and uses a hybrid prediction model for real-time prediction. Finally, the present invention performs uncertainty estimation and confidence assessment through a result fusion algorithm to ensure the reliability and interpretability of the prediction results. The present invention effectively solves the key problems in carbon emission data processing and prediction, and improves prediction accuracy and efficiency.

[0051] In summary, the present invention proposes a data processing method and system for a data center energy-carbon prediction hybrid model, which ensures the reliability and interpretability of the prediction results and effectively improves the accuracy and efficiency of carbon emission data processing and prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 This is a flow chart of a data processing method for a hybrid model for data center energy and carbon prediction provided by the first embodiment of the present invention;

[0053] Figure 2 This is a structural diagram of a data processing method and system for a data center energy-carbon prediction hybrid model provided by the second embodiment of the present invention. DETAILED DESCRIPTION

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0055] Reference Figure 1 The first embodiment of the present invention provides a data processing method and system for a data center energy-carbon prediction hybrid model, including the following steps:

[0056] S11, obtaining raw heterogeneous data related to carbon emissions;

[0057] S12, based on the original heterogeneous data, using ETL technology to convert data from different sources into a unified format to obtain a standardized data set;

[0058] S13, collecting and processing the newly added data in real time to obtain an incremental data set;

[0059] S14, extracting features from the standardized dataset and the incremental dataset and performing multiple processing, using a feature fusion algorithm to perform correlation analysis on the multi-source data to obtain a fused feature set;

[0060] S15, dynamically dividing the fused feature set into batches according to data feature distribution to obtain an adapted data set;

[0061] S16, using a model priority scheduling algorithm, calculating a comprehensive priority score based on the computational complexity, prediction accuracy, and resource requirements of the model, and obtaining an optimized call sequence for the model based on the comprehensive priority score;

[0062] S17, inputting the adapted data set into the model in the optimized call sequence, and performing weighted fusion and outlier filtering on the prediction results of the model to obtain the final prediction results.

[0063] Obtaining raw heterogeneous data related to carbon emissions;

[0064] Based on the original heterogeneous data, ETL technology is used to convert data from different sources into a unified format to obtain a standardized data set;

[0065] Collect and process new data in real time to obtain incremental data sets;

[0066] Extract features from the standardized dataset and the incremental dataset and perform multiple processing, use a feature fusion algorithm to perform correlation analysis on multi-source data to obtain a fused feature set;

[0067] Dynamically dividing the fused feature set into batches according to data feature distribution to obtain an adapted data set;

[0068] Using a model priority scheduling algorithm, a comprehensive priority score is calculated based on the model's computational complexity, prediction accuracy, and resource requirements, and an optimized call sequence of the model is obtained based on the comprehensive priority score;

[0069] The adapted data set is input into the model in the optimized call sequence, and the prediction results of the model are weighted fused and outlier filtered to obtain the final prediction results.

[0070] It should be noted that S11 to S17 constitute a complete real-time processing and prediction process for emission data, which obtains original heterogeneous data related to carbon emissions and provides materials for subsequent data processing and analysis. ETL technology is used to convert data from different sources into a unified format to obtain a standardized data set. This step provides a consistent data format for subsequent analysis and processing, ensuring the comparability and availability of data. New data is collected and processed in real time to obtain an incremental data set. This step ensures the real-time and dynamic nature of the data and ensures that the latest emission conditions can be reflected in a timely manner. Features are extracted from the standardized data set and the incremental data set, and feature fusion algorithms are used to analyze multiple The source data is subjected to correlation analysis to generate a fused feature set. Subsequently, the fused feature set is dynamically batched based on the data feature distribution to generate an adapted dataset. A model priority scheduling algorithm is then used to calculate a comprehensive priority score based on the model's computational complexity, prediction accuracy, and resource requirements. This score is then used to determine the optimal model call sequence. This step further improves prediction accuracy and efficiency by optimizing the model call sequence, ensuring the reliability and interpretability of the prediction results. Finally, the adapted dataset is fed into the optimized model for prediction. The prediction results are weighted and fused, and outlier filtering is performed to produce the final prediction result. This process aims to improve prediction accuracy and reliability. The entire process captures and processes emissions data in real time, providing a solid data foundation for feature extraction and analysis. The model priority scheduling algorithm optimizes the model call sequence, and a hybrid prediction model is used for real-time prediction. Finally, the result fusion algorithm performs uncertainty estimation and confidence assessment to ensure the reliability and interpretability of the prediction results, effectively addressing key issues in emissions data processing and prediction, and improving prediction accuracy and efficiency.

[0071] In step S11 , original heterogeneous data related to carbon emissions is obtained.

[0072] It should be noted that in step S11, carbon emission-related data are collected in real time through smart sensors deployed at key monitoring points in the data center. The carbon emission concentration data is obtained through real-time monitoring by sensors installed near the emission source. It reflects the emission dynamics of the emission source, including smoke emission volume and emission frequency, and is a key indicator for evaluating the environmental impact and energy efficiency of the data center. The energy consumption data is captured in real time through meters connected to energy-using equipment, recording the energy usage of the data center, including power consumption and resource utilization efficiency. These data are used to analyze the operating status of the data center and optimize energy management. In step S11, the carbon emission concentration data and energy consumption data collected in real time by smart sensors can comprehensively and accurately grasp the operating status and environmental changes of the data center, and provide reliable data support for subsequent data processing, analysis and decision-making.

[0073] In step S12, based on the original heterogeneous data, ETL technology is used to convert data from different sources into a unified format to obtain a standardized data set, including:

[0074] By using ETL technology to clean, transform and load the original heterogeneous data, the original heterogeneous data of different sources and formats are converted into a unified standardized format to obtain a standardized data set;

[0075] The quality of the standardized data set is checked by setting a data quality threshold. If the data quality threshold is not reached, it is marked and corrected.

[0076] It should be noted that in step S12, ETL technology (Extraction, Transformation, and Loading) is utilized to clean, transform, and load data. The data cleansing phase aims to remove or correct errors and ensure data accuracy and consistency. The data transformation phase unifies the data into a standard format, including standardization of data types, unification of units, and consistent coding, so that data from different sources can be processed uniformly. The data is then loaded into a storage system for easy management and access. In addition, to ensure data availability and reliability, clear data quality thresholds are set based on business needs and data analysis objectives. Through automated tools or manual inspection, the standardized data set is assessed for quality to confirm whether the data meets the preset quality standards. Data that does not meet the quality threshold is marked to identify the problem and is then corrected, deleted, or specially marked for processing. This quality control step is critical to ensure that only high-quality, consistent data is used in subsequent analysis steps, thereby improving the accuracy and reliability of data center energy and carbon forecasting.

[0077] In step S13, the newly added data is collected and processed in real time to obtain an incremental data set, including:

[0078] Perform data cleaning, denoising, and normalization operations on the newly collected incremental data to obtain a standardized incremental data set;

[0079] It should be noted that in step S13, by implementing an incremental processing mechanism based on a time window, new data is collected and processed in real time within a predetermined time interval to generate an incremental data set. This process includes determining an appropriate time interval and time window size to adapt to the dynamic changing characteristics of the data, using a sliding window mechanism to capture new data in real time, and dynamically adding these data to the incremental data set, performing data cleaning, denoising and normalization operations on the newly collected incremental data to ensure the standardization of the data set, and using an incremental learning algorithm to update and optimize the model in real time based on the characteristics of the incremental data set to improve the adaptability and accuracy of the model. In this process, the feature space is dynamically adjusted through feature extraction technology to remove redundant and irrelevant features, thereby reducing the data dimension and improving the efficiency of data processing. This series of operations ensures that the data center carbon emission prediction hybrid model can continuously adapt to new data and provide accurate and efficient data support for the model.

[0080] In step S14, features are extracted from the standardized dataset and the incremental dataset and multiple processing is performed. A feature fusion algorithm is used to perform correlation analysis on the multi-source data to obtain a fused feature set, including:

[0081] Extracting key features from the standardized dataset and the incremental dataset;

[0082] Use feature selection algorithms to screen out key feature subsets that have a significant impact on the target task, reduce feature dimensions, and obtain dimension-reduced key feature subsets;

[0083] Analyzing the correlation and redundancy between features in the dimensionality-reduced key feature subset, and eliminating highly correlated and redundant features based on the correlation coefficient matrix and mutual information index to obtain a simplified key feature subset;

[0084] According to the feature type and value range, a normalization method is used to map the features of different scales of the simplified key feature subset to the same scale to obtain a regular key feature subset;

[0085] According to the rule-based key feature subsets, a suitable feature fusion algorithm is selected to learn the association rules within and between feature subsets to obtain a fused feature set.

[0086] It should be noted that in step S14, feature extraction and fusion are performed on the standardized dataset and the incremental dataset to generate a comprehensive fused feature set to improve the accuracy and efficiency of data center energy and carbon forecasting. First, key features are extracted from the two datasets. These features are crucial for understanding the energy and carbon performance of data centers. Next, a feature selection algorithm is used to screen out feature subsets that significantly impact the prediction target task, thereby reducing feature dimensionality and simplifying model complexity. The correlation and redundancy between these features are then analyzed. The correlation coefficient matrix and mutual information metric are used to identify and remove highly correlated or redundant features, thereby avoiding model overfitting and improving model generalization. Furthermore, a normalization method is used to map feature values ​​of different feature types and value ranges to the same scale to ensure fairness and effectiveness of model training. Finally, an appropriate feature fusion algorithm is selected to learn association rules within and between feature subsets, and these features are integrated to form the final fused feature set. This fused feature set serves as the basis for model training and prediction, providing more accurate information for data center energy and carbon forecasting.

[0087] In step S15, the fused feature set is dynamically divided into batches according to the data feature distribution to obtain an adapted data set, including:

[0088] According to the input data format requirements of each model, a data conversion method is used to convert the fused feature set into an input data format suitable for each model;

[0089] According to the data feature distribution of the fused feature set, a dynamic batch partitioning algorithm is used to divide the fused feature set into multiple dynamic batches, ensuring that the data feature distribution of each batch is similar, thereby obtaining an adapted data set;

[0090] It should be noted that in step S15, the fused feature set is dynamically batched according to the data feature distribution to adapt to different model input requirements and collaboratively processed to obtain the final output result. First, according to the input data format requirements of each model, a data conversion method is used to convert the fused feature set into an input data format that adapts to each model. This conversion ensures that the data can be correctly read and processed by different models. Then, according to the data feature distribution of the fused feature set, a dynamic batch partitioning algorithm is used to divide the fused feature set into multiple dynamic batches. This division ensures that the data feature distribution of each batch is similar, thereby making the model training and prediction process more stable and reliable.

[0091] In step S16, a model priority scheduling algorithm is used to calculate a comprehensive priority score based on the computational complexity, prediction accuracy, and resource requirements of the model, and an optimized call sequence of the model is obtained based on the comprehensive priority score, including:

[0092] Obtain a list of models to be scheduled, including each model's computational complexity, prediction accuracy, and resource requirements;

[0093] Determine the weight coefficients of model attributes based on business requirements and system resource conditions;

[0094] Calculate the comprehensive priority score for each model;

[0095] The comprehensive priority score is calculated using the following formula:

[0096] S=C×W C +P×W P +R×W R

[0097] In the formula, S is the comprehensive priority score, C is the computational complexity, and W is the C is the complexity weight, P is the prediction accuracy, W P is the precision weight, R is the resource requirement, W R is the resource weight;

[0098] Sorting the model list in descending order according to the comprehensive priority score to obtain an optimized call sequence;

[0099] It should be noted that first, a list of models to be scheduled is obtained. This list contains detailed information about each model, such as computational complexity, prediction accuracy, and resource requirements. This information is the basis for evaluating model performance and resource consumption. Then, based on the current business needs and system resource conditions, weight coefficients are set for each attribute of the model. The weight coefficients reflect the importance of different attributes in model scheduling. Then, we calculate the comprehensive priority score of each model. This score is derived through a specific formula that comprehensively considers the computational complexity, prediction accuracy, and resource requirements of the model, as well as the corresponding weight coefficients. The formula is as follows:

[0100] S=C×W C +P×W P +R×W R

[0101] In the formula, S is the comprehensive priority score, C is the computational complexity, and W is the C is the complexity weight, P is the prediction accuracy, W P is the precision weight, R is the resource requirement, W RResource weights. Computational complexity quantifies the difficulty of executing a model's calculations. It is typically determined by analyzing the model's algorithmic complexity, required computing resources, and execution time. Prediction accuracy reflects the accuracy of the model's predictions and can be calculated by comparing the deviation between the model's predictions and the actual observed values. Resource requirements represent the amount of computing resources required to run a model, such as CPU time and memory usage. Resource requirements are typically assessed based on the model's specifications and the availability of data center resources. Complexity weights, accuracy weights, and resource weights reflect the relative importance of different model attributes in scheduling decisions. The weight coefficients are determined based on business requirements and system resource availability. The comprehensive priority score is a score calculated by multiplying the model's computational complexity, prediction accuracy, and resource requirements with their corresponding weight coefficients. This score is used to assess the scheduling priority of each model. After obtaining each model's comprehensive priority score, the model list is sorted in descending order based on these scores to obtain an optimized model call sequence. This sequence guides the order of model calls to achieve optimal resource allocation and optimization of prediction performance. During the model call process, resource occupancy will be monitored in real time. If the monitored resource occupancy exceeds the preset resource threshold, the resource scheduling mechanism will be triggered to dynamically adjust the model call order to avoid resource overload and maintain stable system operation. Through the implementation of step S16, the order of model calls can be ensured to be both scientific and efficient, so that the data center carbon prediction hybrid model can more accurately predict carbon emissions, while optimizing resource usage and improving overall operating efficiency.

[0102] In step S17, the adapted data set is input into the model in the optimized call sequence, and the prediction results of the model are weighted fused and outlier filtered to obtain the final prediction results, including:

[0103] Inputting the adapted data set into the model in the optimized call sequence, each base model outputs a preliminary prediction result;

[0104] According to the weight coefficient of each model, the preliminary prediction results are weighted averaged, and the isolation forest algorithm is used to filter outliers to obtain the final prediction results of each model;

[0105] The final prediction results of each model are compared with the actual business results, the prediction error is calculated, and the prediction error is fed back to each model. The model parameters are continuously optimized through online learning to improve the prediction accuracy of the model.

[0106] It should be noted that in step S17, the adapted dataset is first input to each model in the optimization call sequence. These models have been sorted according to their performance and resource requirements to ensure the most efficient use of resources. Each base model independently processes the data and outputs a preliminary prediction result. These preliminary prediction results are then weighted averaged based on the model's weight coefficient. This step takes into account the importance and accuracy of each model in the overall prediction. Simultaneously, an isolation forest algorithm is used to identify and filter out outliers that could affect the accuracy of the prediction results. In this way, the final prediction results of each model are obtained. These final prediction results are then compared with the actual results to calculate the prediction error. This process is key to evaluating model performance and helps understand the deviation between the model's prediction and the actual situation. Finally, the prediction error is fed back to each model, and the model parameters are continuously optimized through online learning. This continuous optimization process improves the model's prediction accuracy, ensuring that the model can continuously improve over time and better adapt to new data and business needs. Step S17 ensures that the data center carbon prediction model not only provides accurate prediction results but also continuously learns and adapts to the ever-changing business environment and data characteristics.

[0107] To facilitate understanding of the present invention, some preferred embodiments of the present invention are further described below.

[0108] In this embodiment, a data processing method and system for a data center energy-carbon prediction hybrid model are provided, aiming to accurately predict the carbon emissions of a data center, optimize energy consumption, and improve environmental sustainability.

[0109] The working process is as follows:

[0110] Step 1: Obtain raw heterogeneous data related to carbon emissions. Advanced smart sensors are deployed at key monitoring points in the data center. These sensors are designed to capture real-time data related to carbon emissions, including server energy consumption data, cooling system emissions data, and other emission sources.

[0111] Step 2: Based on the raw heterogeneous data, ETL technology is used to convert the data from different sources into a unified format, resulting in a standardized dataset. The collected raw heterogeneous data is then processed using ETL technology. The ETL process includes data cleansing to remove erroneous and inconsistent data points, formatting the data to a unified standard, and importing the cleaned and transformed data into the analytics platform.

[0112] Step 3: New data is collected and processed in real time to generate an incremental dataset. A time-window-based incremental processing mechanism is applied to the standardized dataset. An appropriate time window size is set to accommodate the dynamic nature of the data, and a sliding window mechanism is used to collect and process new data in real time. This step also includes preprocessing of the new data, such as data cleaning, denoising, and normalization, to generate a standardized incremental dataset.

[0113] Step 4: Extract features from the standardized dataset and the incremental dataset and perform multiple processing operations. A feature fusion algorithm is used to perform correlation analysis on the multi-source data to generate a fused feature set. The features extracted from the standardized dataset and the incremental dataset are analyzed for correlation using a feature fusion algorithm to generate a fused feature set. This step involves feature selection to identify the subset of features that have the greatest impact on the prediction task, identify and eliminate highly correlated or redundant features, and perform normalization to ensure that all features are compared on the same scale.

[0114] Step 5: Dynamically divide the fused feature set into batches based on the data feature distribution to generate an adapted dataset. This step ensures that the data feature distribution of each batch is similar, making model training more efficient and accurate.

[0115] Step 6: Using a model priority scheduling algorithm, a comprehensive priority score is calculated based on the model's computational complexity, prediction accuracy, and resource requirements. This score is then used to determine the optimal model call sequence. This step ensures that the order of model calls meets both business needs and the actual system resource availability.

[0116] Step 7: The adapted dataset is fed into the model in the optimized call sequence, and the model's prediction results are weighted and filtered for outliers to obtain the final prediction result. Finally, the adapted dataset is fed into the model in the optimized call sequence, and the model's prediction results are weighted and filtered for outliers to obtain the final prediction result. This step involves weighted averaging the preliminary prediction results and using anomaly detection algorithms such as isolation forests to filter out possible outliers, thereby improving the robustness of the prediction results.

[0117] In summary, the present invention provides a data processing method and system for a data center energy-carbon prediction hybrid model that can achieve real-time monitoring and accurate prediction of data center carbon emissions, optimize the operating efficiency and environmental impact of the data center, and thus promote the development of data centers in a greener and more sustainable direction.

[0118] Reference Figure 2 The present invention provides a data processing method and system for a data center energy-carbon prediction hybrid model, including:

[0119] Data acquisition module, used to obtain raw heterogeneous data related to carbon emissions;

[0120] An ETL operation module is used to convert data from different sources into a unified format based on the original heterogeneous data using ETL technology to obtain a standardized data set;

[0121] The incremental data processing module is used to collect and process the new data in real time to obtain the incremental data set;

[0122] A feature extraction and fusion operation module is used to extract features from the standardized data set and the incremental data set and perform multiple processing, and use a feature fusion algorithm to perform correlation analysis on multi-source data to obtain a fused feature set;

[0123] A dynamic batch division module is used to dynamically divide the fused feature set into batches according to the data feature distribution to obtain an adapted data set;

[0124] A model priority scheduling module is used to use a model priority scheduling algorithm to calculate a comprehensive priority score based on the model's computational complexity, prediction accuracy, and resource requirements, and to obtain an optimized call sequence for the model based on the comprehensive priority score;

[0125] The result output module is used to input the adapted data set into the model in the optimized call sequence, and perform weighted fusion and outlier filtering on the prediction results of the model to obtain the final prediction result.

[0126] In summary, the present invention provides a data processing method and system for a hybrid model for energy and carbon prediction in a data center. First, the data acquisition module captures raw heterogeneous data related to carbon emissions in real time. These data truly reflect the operating status of the data center. Next, the ETL operation module uses ETL technology to clean and transform the raw heterogeneous data, unify the data formats of different sources, and generate a standardized data set. Then, the incremental data processing module performs real-time collection and processing on the newly added data to generate an incremental data set. Subsequently, the feature extraction and fusion operation module extracts key features from the standardized data set and the incremental data set, and performs multiple processing. The feature fusion algorithm is used to perform correlation analysis on the multi-source data to obtain a fused feature set. The dynamic batch division module dynamically divides the data into multiple batches with similar feature distributions based on the data feature distribution of the fused feature set to form an adapted data set. The model priority scheduling module uses a model priority scheduling algorithm to calculate a comprehensive priority score based on the model's computational complexity, prediction accuracy, and resource requirements, and obtains an optimized call sequence for the model based on this score. Finally, the result output module inputs the adapted data set into the model in the optimized call sequence, and performs weighted fusion and outlier filtering on the prediction results of each model to obtain the final prediction result.

[0127] It should be noted that the data processing method and system for a data center energy-carbon prediction hybrid model provided in an embodiment of the present invention are used to execute all the process steps of the data processing method for a data center energy-carbon prediction hybrid model of the above embodiment. The working principles and beneficial effects of the two correspond one to one, and therefore will not be repeated here.

[0128] An embodiment of the present invention further provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a virtual teaching data interaction program based on digital twins. When the processor executes the computer program, the steps in each of the above-mentioned embodiments of the virtual teaching data interaction method based on digital twins are implemented, such as Figure 1 Alternatively, when the processor executes the computer program, the functions of the modules / units in the above-mentioned device embodiments are realized, for example, the module.

[0129] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0130] The electronic device may be a computing device such as a desktop computer, notebook, PDA, or smart tablet. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that the aforementioned components are merely examples of electronic devices and do not constitute a limitation of the electronic device. The electronic device may include more or fewer components than those described above, or a combination of certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, and the like.

[0131] The processor may be a central processing unit (CPU), other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the electronic device, and connects various parts of the entire electronic device using various interfaces and lines.

[0132] The memory can be used to store the computer programs and / or modules, and the processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Med ia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0133] Wherein, if the module / unit integrated in the electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0134] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.

[0135] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A data processing method and system for a data center energy-carbon prediction hybrid model, characterized in that: Executed by a computer, including: Obtaining raw heterogeneous data related to carbon emissions; Based on the original heterogeneous data, ETL technology is used to convert data from different sources into a unified format to obtain a standardized data set; Collect and process new data in real time to obtain incremental data sets; Extract features from the standardized dataset and the incremental dataset and perform multiple processing, use a feature fusion algorithm to perform correlation analysis on multi-source data to obtain a fused feature set; Dynamically dividing the fused feature set into batches according to data feature distribution to obtain an adapted data set; Using a model priority scheduling algorithm, a comprehensive priority score is calculated based on the model's computational complexity, prediction accuracy, and resource requirements, and an optimized call sequence of the model is obtained based on the comprehensive priority score; The adapted data set is input into the model in the optimized call sequence, and the prediction results of the model are weighted fused and outlier filtered to obtain the final prediction results.

2. The data processing method and system of the hybrid model for data center energy and carbon prediction according to claim 1 is characterized in that: Based on the original heterogeneous data, ETL technology is used to convert data from different sources into a unified format to obtain a standardized data set, including: By using ETL technology to clean, convert and load the original heterogeneous data, the original heterogeneous data from different sources and in different formats are converted into a unified standardized format to obtain a standardized data set.

3. The data processing method and system of the hybrid model for data center energy and carbon prediction according to claim 1 is characterized in that: Collect and process new data in real time to obtain incremental data sets, including: The newly collected incremental data is cleaned, denoised, and normalized to obtain a standardized incremental data set.

4. The data processing method and system of a data center energy-carbon prediction hybrid model according to claim 1 is characterized in that: Extract features from the standardized dataset and the incremental dataset and perform multiple processing, use a feature fusion algorithm to perform correlation analysis on the multi-source data, and obtain a fused feature set, including: Extracting key features from the standardized dataset and the incremental dataset; Use feature selection algorithms to screen out key feature subsets that have a significant impact on the target task, reduce feature dimensions, and obtain dimension-reduced key feature subsets; Analyzing the correlation and redundancy between features in the dimensionality-reduced key feature subset, and eliminating highly correlated and redundant features based on the correlation coefficient matrix and mutual information index to obtain a simplified key feature subset; According to the feature type and value range, a normalization method is used to map the features of different scales of the simplified key feature subset to the same scale to obtain a regular key feature subset; According to the rule-based key feature subsets, a suitable feature fusion algorithm is selected to learn the association rules within and between feature subsets to obtain a fused feature set.

5. The data processing method and system of a data center energy-carbon prediction hybrid model according to claim 1 is characterized in that: The fused feature set is dynamically divided into batches according to the data feature distribution to obtain an adapted data set, including: According to the input data format requirements of each model, a data conversion method is used to convert the fused feature set into an input data format suitable for each model; According to the data feature distribution of the fused feature set, a dynamic batch partitioning algorithm is used to divide the fused feature set into multiple dynamic batches, ensuring that the data feature distribution of each batch is similar, and obtaining an adapted data set.

6. The data processing method and system of a data center energy-carbon prediction hybrid model according to claim 1 is characterized in that: A model priority scheduling algorithm is used to calculate a comprehensive priority score based on the model's computational complexity, prediction accuracy, and resource requirements. The optimized call sequence of the model is then obtained based on the comprehensive priority score, including: Obtain a list of models to be scheduled, including each model's computational complexity, prediction accuracy, and resource requirements; Determine the weight coefficients of model attributes based on business requirements and system resource conditions; Calculate the comprehensive priority score for each model; The comprehensive priority score is calculated using the following formula: S=C×W C +P×W P +R×W R In the formula, S is the comprehensive priority score, C is the computational complexity, and W is the C is the complexity weight, P is the prediction accuracy, W P is the precision weight, R is the resource requirement, W R is the resource weight; According to the comprehensive priority scores, the model list is sorted in descending order to obtain an optimized calling sequence.

7. The data processing method and system of a data center energy-carbon prediction hybrid model according to claim 1 is characterized in that: The adapted dataset is input into the model in the optimized call sequence, and the prediction results of the model are weighted fused and outliers are filtered to obtain the final prediction results. , including: Inputting the adapted data set into the model in the optimized call sequence, each base model outputs a preliminary prediction result; According to the weight coefficient of each model, the preliminary prediction results are weighted averaged, and the isolation forest algorithm is used to filter outliers to obtain the final prediction results of each model; The final prediction results of each model are compared with the actual business results, the prediction error is calculated, and the prediction error is fed back to each model. The model parameters are continuously optimized through online learning to improve the prediction accuracy of the model.

8. A data processing method and system for a hybrid model of energy and carbon prediction in a data center, characterized in that: include: Data acquisition module, used to obtain raw heterogeneous data related to carbon emissions; An ETL operation module is used to convert data from different sources into a unified format based on the original heterogeneous data using ETL technology to obtain a standardized data set; The incremental data processing module is used to collect and process the new data in real time to obtain the incremental data set; A feature extraction and fusion operation module is used to extract features from the standardized data set and the incremental data set and perform multiple processing, and use a feature fusion algorithm to perform correlation analysis on multi-source data to obtain a fused feature set; A dynamic batch division module is used to dynamically divide the fused feature set into batches according to the data feature distribution to obtain an adapted data set; A model priority scheduling module is used to use a model priority scheduling algorithm to calculate a comprehensive priority score based on the model's computational complexity, prediction accuracy, and resource requirements, and to obtain an optimized call sequence for the model based on the comprehensive priority score; The result output module is used to input the adapted data set into the model in the optimized call sequence, and perform weighted fusion and outlier filtering on the prediction results of the model to obtain the final prediction result.

9. An electronic device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements a data processing method for a data center energy carbon prediction hybrid model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a data processing method and system for a data center energy-carbon prediction hybrid model as described in any one of claims 1 to 7.