A computing power thermal evaluation method and decision-making system based on multidimensional data analysis
Through multi-dimensional data analysis and dynamic optimization mechanisms, a hierarchical evaluation system is built, which solves the multi-dimensional shortcomings of existing computing power evaluation methods, realizes the systematicity and accuracy of computing power resource evaluation, and supports scientific decision-making of regional computing power resources.
Patent Information
- Application Number
- CN202510637944.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing computing power analysis and evaluation methods mainly have problems such as single-dimensional evaluation that it is difficult to fully reflect the complexity of regional computing power development, ignore the mutual influence between multi-dimensional indicators, cannot accurately characterize nonlinear influence characteristics, and lack of policy document analysis and static evaluation results to adapt to dynamic changes.
A multi-dimensional data analysis method is adopted to construct a hierarchical evaluation system by fusing multi-source heterogeneous data, introducing a nonlinear weighted model and dynamic visual display, combining hierarchical analysis method and regional clustering algorithm, differentiated initial index weights are allocated for different types of regions, optimize index weights, and generate thermal spatiotemporal distribution maps.
It realizes the systematicity and objectivity of computing power resource evaluation, improves the accuracy and visualization of evaluation results, and supports the rational allocation and efficient utilization of regional computing power resources.
Smart Images

Figure CN120181518B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computing power analysis technology, and in particular to a computing power thermal evaluation method based on multidimensional data analysis and a decision-making system thereof. Background Art
[0002] Existing computing power analysis and evaluation decisions primarily rely on statistical indicator analysis and empirical evaluation methods. Statistical indicator analysis primarily evaluates hardware metrics such as rack number, server scale, and CPU utilization, combined with economic indicators such as energy consumption and investment scale for comprehensive analysis. Empirical evaluation methods use historical data and expert experience to qualitatively assess and rank regional computing power development levels. In practice, some regions employ single-dimensional scoring models or simple weighted average methods for computing power resource evaluation and decision-making analysis.
[0003] However, the above technical solutions have obvious shortcomings: first, the single-dimensional evaluation method is difficult to fully reflect the complexity of regional computing power development and ignores the mutual influence between multi-dimensional indicators; second, the simple linear scoring model cannot accurately characterize the nonlinear impact characteristics of indicators on computing power development; third, the existing evaluation methods lack in-depth analysis of policy documents and it is difficult to quantify the actual impact of the policy environment on computing power development; finally, static evaluation results are difficult to adapt to the dynamic changes in computing power demand, and the evaluation results have a low degree of visualization, which is not conducive to intuitive understanding and quick decision-making. Summary of the Invention
[0004] In view of this, the present invention proposes a computing power thermal evaluation method and decision-making system based on multi-dimensional data analysis. By integrating multi-source heterogeneous data, constructing a hierarchical evaluation system, introducing nonlinear weighted models and dynamic visualization displays, it provides decision support for the rational allocation and efficient utilization of regional computing power resources.
[0005] The technical solution of the present invention is achieved as follows:
[0006] In one aspect, the present invention provides a computing power thermal evaluation method based on multidimensional data analysis, comprising:
[0007] S1. Obtain the original dataset containing active subject data, economic data, resource and environmental data, and policy documents, preprocess the original dataset, and generate a standardized dataset;
[0008] S2. Perform vectorization and heterogeneous fusion on the standardized dataset to generate a unified fusion dataset;
[0009] S3. Build a hierarchical evaluation system and extract corresponding evaluation indicator features from the fused data set. Perform correlation analysis based on the evaluation indicator features to form a computing power evaluation feature set.
[0010] S4. Using the analytic hierarchy process and regional clustering algorithm, configure differentiated initial indicator weights for different types of regions;
[0011] S5. Calculate the initial computing power thermal value of each region using a nonlinear weighted model, combining the computing power evaluation feature set and the initial indicator weights.
[0012] S6. Obtain the actual computing power requirements of each region, use the optimization algorithm to iteratively optimize the initial indicator weights, and generate optimized indicator weights;
[0013] S7. Recalculate the regional computing power thermal value based on the optimized indicator weights, generate a thermal spatiotemporal distribution map, and generate a computing power resource allocation plan based on the thermal gradient differences.
[0014] Preferably, preprocessing the original data set in step S1 includes:
[0015] Define data cleaning rules through the rule engine and clean the original data set according to the data cleaning rules. The data cleaning rules include outlier correction thresholds, missing value filling strategies, and data integrity scoring criteria.
[0016] The dynamic time warping algorithm is used to align the cleaned data of different acquisition periods to a unified reference time axis, and the periodic noise interference is eliminated based on Fourier transform to generate a standardized data set.
[0017] Preferably, active subject data include population data, household data, and enterprise data; economic data include GDP data, industrial data, transportation and logistics data, education data, financial data, commercial data, agricultural data, and medical data; resource and environmental data include power generation data, power structure data, temperature data, and IDC facility data; policy documents include computing hub policies, regional planning policies, industrial support policies, and energy management policies.
[0018] Preferably, step S2 includes:
[0019] S21. For the active subject data, economic data, and resource and environmental data in the standardized data set, a piecewise adaptive normalization method is used for vectorization processing to convert data of different dimensions into a unified numerical vector;
[0020] S22. Construct a data quality assessment model, calculate a quality score based on the three dimensions of data integrity, timeliness, and accuracy, determine the weight of each dimension, generate a comprehensive quality score, and use the comprehensive quality score as an adjustment coefficient to perform weighted adjustment on the initial numerical vector to generate a quality-weighted numerical vector;
[0021] S23. Conduct in-depth semantic analysis of policy documents to extract policy clauses related to computing power. Quantify the impact of policies based on the term frequency-inverse document frequency method and establish association rules between policy clauses and quality-weighted feature vectors.
[0022] S24, dynamically adjusting the policy-affected vectors in the quality-weighted numerical vectors according to the association rules;
[0023] S25. Align and integrate the adjusted numerical vectors according to the time dimension and the spatial dimension to generate a unified fusion data set.
[0024] Preferably, step S21 includes:
[0025] S211. Perform logarithmic transformation and normalization on economic data: Where, It is the normalized economic data; It is the original economic data; is the minimum value of economic data; is the maximum value of economic data; is the smoothing factor, and its value range is ;
[0026] S212. Construct a segmented mapping function to normalize the active subject data and determine the threshold set based on the inflection point analysis of the historical computing power demand curve. , n is the threshold number, when the active subject data When , the mapping function is as follows: Where, This is the normalized active subject data; is the linear transformation coefficient; is the bias constant; and Determined by calculating the segment endpoint values;
[0027] S213. Perform nonlinear mapping on resource and environmental data:
[0028] Where, It is the resource and environmental data after normalization; is the normalization coefficient; is the nonlinear adjustment factor; It is the original resource and environmental data; is the environmental suitability threshold;
[0029] S214. Organize all normalized values of the same data source at the same time point into a numerical vector : Where s is the unique identifier of the data source; t is the time point.
[0030] Preferably, the hierarchical evaluation system is a two-tier evaluation indicator system:
[0031] The first-level indicators include computing power demand level, computing power support conditions, and computing power development environment;
[0032] Secondary indicators include:
[0033] The computing power demand level is divided into the density of information technology enterprises, the proportion of Internet enterprises, the number of large and medium-sized enterprises, and the digital coverage of key industries;
[0034] Under the computing power support conditions, the IDC facility capacity, power supply guarantee capability, network bandwidth level, and computer room environment adaptability are determined;
[0035] The computing power development environment includes the proportion of information industry investment, industrial policy support, resource element guarantee, and infrastructure completeness.
[0036] Preferably, in step S3, the correlation analysis adopts a feature redundancy elimination algorithm based on local sensitive hashing, by calculating the Pearson correlation coefficient of the indicator pairs in the hash bucket, combined with the set threshold to filter and retain significant features as the computing power evaluation feature set.
[0037] Preferably, in step S5, the calculation formula of the nonlinear weighted model is: Where, is the initial computing power thermal value of the jth region; is the weight of the kth initial indicator; The normalized value of the kth feature in the computational power evaluation feature set for the jth region; is the nonlinear adjustment coefficient of the kth feature; is the policy impact intensity of the jth region; is the policy impact coefficient; The number of features in the feature set for computational power evaluation.
[0038] Preferably, the method for generating the thermal spatiotemporal distribution map includes: constructing an initial Voronoi spatial structure with a set of regional locations as a construction unit and a regional computing power thermal value as a threshold; dynamically adjusting the computing power thermal field structure of adjacent areas based on a spatiotemporal propagation function, wherein the spatiotemporal propagation function takes into account the geographical distance between regions and thermal diffusion parameters; and generating a spatiotemporal distribution map that supports real-time interaction by dynamically updating the Voronoi thermal area boundaries.
[0039] On the other hand, the present invention further provides a computing power decision-making system based on multidimensional data analysis, the system being used to execute any of the above methods, the system comprising:
[0040] The data acquisition module is used to collect active subject data, economic data, resource and environmental data, and policy documents, and perform standardized preprocessing to generate standardized data sets;
[0041] The data fusion module is used to perform vectorization processing and heterogeneous fusion on the standardized data set to generate a unified fused data set;
[0042] The feature analysis module is used to build a hierarchical evaluation system, extract evaluation indicator features from the fused data set, and perform correlation analysis and feature selection to form a computing power evaluation feature set;
[0043] Regional clustering module, used to configure differentiated initial indicator weights based on the analytic hierarchy process and regional clustering algorithm;
[0044] The thermal evaluation module is used to combine the computing power evaluation feature set and the initial indicator weights to calculate the initial regional computing power thermal value through a nonlinear weighted model, and optimize the indicator weights based on the output of the weight optimization module to update the computing power thermal value;
[0045] The weight optimization module is used to iteratively optimize the initial indicator weights based on the actual computing power requirements;
[0046] The result display module is used to generate a thermal spatiotemporal distribution map based on the computing power thermal value and output a computing power resource allocation plan.
[0047] The present invention has the following beneficial effects compared to the prior art:
[0048] By building a complete computing power thermal assessment method, we have achieved full automation of the entire process from data collection and feature analysis to assessment and decision-making. This improves the systematicness and objectivity of computing power resource assessment and provides a reliable decision-making basis for regional computing power resource allocation.
[0049] The use of vectorized processing and heterogeneous fusion methods to standardize multi-source data solves the problem of difficulty in unified analysis of different types of data, improving data utilization efficiency and analysis accuracy;
[0050] Based on the hierarchical analysis method and regional clustering algorithm, differentiated initial weights are configured, which overcomes the problem of unreasonable weight configuration in traditional evaluation methods and improves the accuracy and credibility of the evaluation results;
[0051] The evaluation indicators are optimized and screened through feature redundancy elimination algorithms, which reduces data redundancy, improves the effectiveness of feature expression, and makes the evaluation system more streamlined and efficient.
[0052] A nonlinear weighted model is introduced to calculate thermal values, effectively describing the nonlinear impact of evaluation indicators on computing power thermal values and improving the model's fit to actual conditions.
[0053] The spatiotemporal computing power thermal field drawing method based on dynamic Voronoi diagram realizes the real-time dynamic display of thermal distribution and enhances the visualization effect and interactive performance of the evaluation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0055] Figure 1 is a flow chart of the method of the present invention;
[0056] Figure 2 It is a technical implementation diagram of the present invention;
[0057] Figure 3 This is a system framework diagram of the present invention. DETAILED DESCRIPTION
[0058] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0059] like Figure 1 As shown, the present invention provides a computing power thermal evaluation method based on multidimensional data analysis, comprising:
[0060] S1. Obtain the original dataset containing active subject data, economic data, resource and environmental data, and policy documents, preprocess the original dataset, and generate a standardized dataset;
[0061] S2. Perform vectorization and heterogeneous fusion on the standardized dataset to generate a unified fusion dataset;
[0062] S3. Build a hierarchical evaluation system and extract corresponding evaluation indicator features from the fused data set. Perform correlation analysis based on the evaluation indicator features to form a computing power evaluation feature set.
[0063] S4. Using the analytic hierarchy process and regional clustering algorithm, configure differentiated initial indicator weights for different types of regions;
[0064] S5. Calculate the initial computing power thermal value of each region using a nonlinear weighted model, combining the computing power evaluation feature set and the initial indicator weights.
[0065] S6. Obtain the actual computing power requirements of each region, use the optimization algorithm to iteratively optimize the initial indicator weights, and generate optimized indicator weights;
[0066] S7. Recalculate the regional computing power thermal value based on the optimized indicator weights, generate a thermal spatiotemporal distribution map, and generate a computing power resource allocation plan based on the thermal gradient differences.
[0067] See also Figure 2 The present invention first collects multi-source heterogeneous data, including active subject data, economic data, resource and environmental data, and policy documents, and achieves data standardization and integration through vectorization processing. It then constructs a hierarchical evaluation system and extracts evaluation indicator features, optimizing the feature set using a feature redundancy elimination algorithm. Differentiated initial weights are then configured based on the analytic hierarchy process and regional clustering algorithm, and weights are optimized based on actual computing power requirements. Regional computing power thermal values are calculated using a nonlinear weighted model. Finally, a dynamic Voronoi diagram is used to visualize the spatiotemporal distribution of thermal power, forming a complete computing power resource evaluation decision-making scheme. In the specific implementation process, text vectorization technology is used to perform semantic analysis on policy documents and extract policy-oriented features. Correlation analysis and principal component analysis are used to screen key evaluation indicators and establish a multi-level indicator evaluation system. A regional clustering algorithm is used to classify evaluation areas and achieve differentiated weight configuration. A spatiotemporal propagation function is introduced to dynamically adjust the computing power thermal field, generating a real-time interactive distribution map. This technical solution achieves precise computing power resource evaluation and scientific decision-making through multidimensional data analysis and dynamic optimization mechanisms.
[0068] Specifically, in one embodiment of the present invention, active subject data includes population data, household data, and enterprise data; economic data includes GDP data, industrial data, transportation and logistics data, education data, financial data, commercial data, agricultural data, and medical data; resource and environmental data includes power generation data, power structure data, temperature data, and IDC facility data; policy documents include computing power hub policies, regional planning policies, industrial support policies, and energy management policies.
[0069] In this embodiment, active subject data is obtained through population census data, enterprise registration databases, and relevant statistical yearbooks released by national statistical departments; economic data comes from statistical bulletins, industry development reports, and professional data monitoring platforms released by statistical departments at all levels; resource and environmental data are mainly obtained from the power industry operation monitoring system, meteorological department observation databases, and data center infrastructure census data; policy documents are obtained through public channels such as policy documents, planning outlines, and industry management methods released on government portal websites at all levels.
[0070] Specifically, web crawler technology, application program interface (API) calls, database exchange, file parsing, and other technical means can be used to build an automated data collection system to acquire multi-source heterogeneous data. For example, distributed web crawler technology can be used to capture statistical data and policy documents from government portals in real time, and an intelligent parsing engine can be used to perform structured extraction of statistical reports in various formats such as PDF and Excel. A RESTful API interface can be used to connect with various professional data platforms to achieve regular synchronization of economic indicators and environmental data. For data with high real-time requirements (such as power data and meteorological data), a Kafka-based streaming data collection mechanism can be used to achieve real-time data access through message queues.
[0071] Specifically, in one embodiment of the present invention, preprocessing the original data set in step S1 includes:
[0072] Data cleaning rules are defined through the rule engine, and the original data set is cleaned according to the data cleaning rules. The data cleaning rules include outlier correction thresholds, missing value filling strategies, and data integrity scoring criteria.
[0073] Specifically, a rule engine based on the Rete algorithm can be used to define and execute data cleaning rules. This rule engine transforms data cleaning rules into executable rule chains by building a rule network, enabling dynamic configuration and automated execution of rules.
[0074] In this embodiment, the outlier correction threshold is calculated using the interquartile range (IQR) method, and data exceeding Q3+1.5IQR or below Q1-1.5IQR are identified as outliers; the missing value filling strategy adopts the multiple interpolation method, and mean filling, regression filling or nearest neighbor filling is selected according to the data characteristics; the data integrity scoring standard is comprehensively calculated based on the data's temporal coverage, spatial coverage and attribute completeness.
[0075] The dynamic time warping algorithm is used to align the cleaned data of different acquisition periods to a unified reference time axis, and the periodic noise interference is eliminated based on Fourier transform to generate a standardized data set.
[0076] Specifically, the Dynamic Time Warping (DTW) algorithm calculates the optimal warping path between different time series to achieve flexible data alignment in the temporal dimension, thereby mapping data from different acquisition periods onto a unified reference time axis. Next, a Fast Fourier Transform (FFT) is applied to the aligned data, identifying and filtering out periodic noise components through frequency domain analysis to achieve data smoothing and ultimately generate a standardized dataset.
[0077] Specifically, in one embodiment of the present invention, step S2 includes:
[0078] S21. For the active subject data, economic data, and resource and environmental data in the standardized data set, a piecewise adaptive normalization method is used for vectorization processing to convert data of different dimensions into a unified numerical vector;
[0079] In this embodiment, since the distribution characteristics and value ranges of different types of data are significantly different, the use of a segmented adaptive normalization method can better maintain the original distribution characteristics and relative relationships of the data. Step S21 includes:
[0080] S211. Perform logarithmic transformation and normalization on economic data: Where, It is the normalized economic data; It is the original economic data; is the minimum value of economic data; is the maximum value of economic data; is the smoothing factor, and its value range is ;
[0081] S212. Construct a segmented mapping function to normalize the active subject data and determine the threshold set based on the inflection point analysis of the historical computing power demand curve. , n is the threshold number, when the active subject data When , the mapping function is as follows: Where, This is the normalized active subject data; is the linear transformation coefficient; is the bias constant; and Determined by calculating the segment endpoint values;
[0082] S213. Perform nonlinear mapping on resource and environmental data:
[0083] Where, It is the resource and environmental data after normalization; is the normalization coefficient; is the nonlinear adjustment factor; It is the original resource and environmental data; is the environmental suitability threshold;
[0084] S214. Organize all normalized values of the same data source at the same time point into a numerical vector : Where s is the unique identifier of the data source; t is the time point.
[0085] In this embodiment, active subject data has distinct hierarchical characteristics, and a piecewise mapping function can better distinguish data features at different levels. Economic data generally exhibits exponential growth characteristics, and logarithmic transformation can convert it into a linear relationship, facilitating subsequent analysis. Resource and environmental data exhibits nonlinear fluctuations, and nonlinear mapping can better capture the dynamic changes in the data. All normalized values from the same data source at the same time point are organized into fixed-dimensional numerical vectors according to a preset feature dimension sequence, ensuring that vectors at different time points have the same dimension and feature correspondence.
[0086] S22. Build a data quality assessment model, calculate the quality score based on the three dimensions of data integrity, timeliness, and accuracy, determine the weight of each dimension, generate a comprehensive quality score, and use the comprehensive quality score as an adjustment coefficient to perform weighted adjustment on the initial numerical vector to generate a quality-weighted numerical vector.
[0087] In this embodiment, since the acquired data has the characteristics of multi-source heterogeneity, time series continuity, and large differences in numerical distribution, completeness, timeliness, and accuracy are selected as quality assessment dimensions.
[0088] Among them, integrity is calculated based on data item coverage, and the degree of data completeness is measured by the ratio of the actual amount of data obtained to the theoretical amount of data: Where, is the completeness score, with a value range of [0,1]; is the number of data items actually obtained; The expected number of data items to be obtained.
[0089] Timeliness is based on the exponential decay principle and quantifies the time delay impact of data by setting the time decay coefficient: Where, is the timeliness score, with a value range of [0,1]; is the time decay coefficient, which is used to control the time-dependent decay rate; The data update time; Assess the time for the baseline.
[0090] Accuracy is based on statistical principles and combines standard deviation and outlier ratio to assess the credibility of data: Where, is the accuracy score, with a value range of [0,1]; is the proportion of outliers; is the data standard deviation; is the maximum allowed standard deviation.
[0091] The weight of each dimension is determined by the entropy weight method, which reflects the discrete degree of each dimension by calculating the information entropy to achieve objective weighting: Where, is the weight of the i-th dimension, i=1,2,3; is the information entropy of the i-th dimension.
[0092] Finally, the weighted scores of the three dimensions are summed to obtain the comprehensive quality score: Where, It is a comprehensive quality score, and its value range is [0,1].
[0093] The comprehensive quality score Q is used as the adjustment coefficient, and the initial numerical vector V is adjusted by linear weighting: V'=Q×V to generate a quality-weighted numerical vector.
[0094] This objective weighting method based on data characteristics can dynamically reflect the data quality status and avoid the deviation caused by subjective judgment. At the same time, through the quality weighting mechanism, it realizes the automatic downgrading of low-quality data.
[0095] S23. Conduct in-depth semantic analysis of policy documents to extract policy clauses related to computing power, quantify the intensity of policy impact based on the word frequency-inverse document frequency method, and establish association rules between policy clauses and each quality-weighted feature vector.
[0096] In this embodiment, a pre-trained BERT model can be used to perform deep semantic analysis of policy documents. The model input is the original policy document, and the output is a semantic vector representation of the text and key entity information. Based on the model output, a topic identification algorithm is used to extract text fragments related to computing power. The term frequency-inverse document frequency (TF-IDF) method is then combined to filter out policy clauses related to computing power. Specifically, by setting a core dictionary and similarity threshold for the computing power domain, text fragments with semantic similarity above the threshold are identified as relevant policy clauses.
[0097] For each policy clause j, calculate its policy impact intensity P j : Where, Indicates the frequency of keywords in the policy terms. represents the inverse document frequency of the keyword, is the policy level weight coefficient.
[0098] Construct a correlation strength model between policy clauses and quality-weighted feature vectors. Let the quality-weighted feature vector be V', and for each feature component k in the vector, calculate its correlation coefficient R with policy clause j. jk : Where, Semantic similarity, obtained by calculating the cosine similarity between policy terms and feature component descriptions; Score the quality of the feature component; are the correlation benchmark coefficient and the quality adjustment coefficient respectively.
[0099] Establish fusion rules based on the correlation coefficient matrix R: when R jk When the value is greater than a set threshold, the association between policy clause j and feature component k is established, and the strength of the association is used as the weight for subsequent analysis. This approach, based on semantic analysis and association modeling, maintains logical consistency with the data processing in the previous step while achieving a quantitative mapping from policy documents to numerical features.
[0100] S24. Dynamically adjust the policy-affected vectors in the quality-weighted numerical vectors according to the association rules.
[0101] In this embodiment, the adjustment coefficient matrix at time t is assumed to be: ,in, is the identity matrix, is a global adjustment factor used to control the overall intensity of policy impact; is the policy time-effect attenuation matrix, whose elements , , is the difference between the policy release time and the current time, is the time attenuation coefficient; is the correlation coefficient matrix at time t.
[0102] The coefficient matrix will be adjusted Acting on the quality weighted numerical vector V', the adjusted numerical vector V'' is obtained: V'' t =A t ×V' t , where the correlation coefficient R jk The feature components that are smaller than the threshold keep their original values unchanged.
[0103] S25. Align and integrate the adjusted numerical vectors according to the time dimension and the spatial dimension to generate a unified fusion data set.
[0104] In this embodiment, the adjusted numerical vector V'' obtained in step S24 is first aligned in the time dimension. A unified time granularity Δt (such as day, week, or month) is set, and data with different timestamps are mapped to a standard time point. For time point t, the numerical calculation formula after alignment is: ,in, is the time offset, 、 、 is the time weighting coefficient, and satisfies , used to smooth time series data.
[0105] Then, spatial dimension alignment is performed. Based on geographic information coding, a spatial mapping function M(x,y) is established to unify data of different spatial scales into a standard spatial grid: Among them, M(x,y) is the spatial mapping matrix, which is used to process spatial distribution differences.
[0106] Finally, the fusion dataset D is constructed, and its structure is: , where T is the time index set, S is the space index set, and v is the aligned feature vector. This alignment and integration method based on the dual dimensions of time and space achieves a unified representation of heterogeneous data.
[0107] Specifically, in one embodiment of the present invention, step S3 includes: constructing a hierarchical evaluation system, which is a two-layer evaluation indicator system:
[0108] The first-level indicators include computing power demand level, computing power support conditions, and computing power development environment;
[0109] Secondary indicators include:
[0110] The computing power demand level is divided into the density of information technology enterprises, the proportion of Internet enterprises, the number of large and medium-sized enterprises, and the digital coverage of key industries;
[0111] Under the computing power support conditions, the IDC facility capacity, power supply guarantee capability, network bandwidth level, and computer room environment adaptability are determined;
[0112] The computing power development environment includes the proportion of information industry investment, industrial policy support, resource element guarantee, and infrastructure completeness.
[0113] In this embodiment, a hierarchical evaluation system is constructed based on the inherent logic of computing power development, evaluating from three dimensions: demand, supply, and environment. The computing power demand level reflects the actual intensity of a region's demand for computing resources; the computing power support conditions reflect the region's fundamental ability to provide computing power services; and the computing power development environment represents the degree to which the sustainable development of the region's computing power industry is guaranteed. This multi-dimensional evaluation system design not only conforms to the laws of computing power development but also comprehensively reflects the region's computing power level.
[0114] For the calculation of each secondary indicator, the corresponding features are extracted from the fused dataset:
[0115] Information Technology Enterprise Density: Calculates the number of IT enterprises per unit area based on enterprise registration data and industry classification data; Internet Enterprise Proportion: Calculates the proportion of internet-related enterprises using enterprise business scope data; IDC Facility Capacity: Calculates a comprehensive calculation based on data center rack count and design load; Power Supply Capacity: Calculates the power supply stability and load factor using weighted indicators; Industry Policy Support: Quantitatively calculated based on the policy document analysis results in steps S23-24. Other indicators are similarly calculated to form a complete indicator calculation system.
[0116] In the correlation analysis phase, the locality sensitive hashing (LSH) algorithm is used to eliminate feature redundancy. First, the n secondary indicator feature vectors are constructed into an n×m feature matrix F, where m is the number of samples. The execution process of the LSH algorithm is as follows:
[0117] Design a hash function family H: , where a is a random vector, b is a random offset, and r is the bucket width parameter.
[0118] Calculate the Pearson correlation coefficient between eigenvectors: , where x and y are feature pairs that fall into the same hash bucket.
[0119] Feature screening rules: When When θ is the correlation threshold, the features with larger information content are retained and redundant features are removed.
[0120] Through this method, the final computing power evaluation feature set not only maintains the integrity of the evaluation dimension but also ensures the independence between features.
[0121] Specifically, in one embodiment of the present invention, step S4 includes:
[0122] Based on the computing power evaluation feature set, the K-means clustering algorithm was used to classify regions. The clustering process focused on key factors such as the region's computing power demand, infrastructure conditions, and industrial development level, dividing regions into different types: computing power-leading, demand-driven, and infrastructure-supported.
[0123] Then, for different types of regions, we used the analytic hierarchy process to determine the relative importance of indicators at each level. When constructing the judgment matrix, we fully considered the characteristics of the regional type. For example, for regions with leading computing power, we increased the weight of computing power support conditions; for demand-driven regions, we appropriately increased the weight of indicators related to computing power demand levels; and for basic support regions, we placed greater emphasis on the weighting of the computing power development environment.
[0124] In this embodiment, the K-means clustering algorithm and the analytic hierarchy process used are both existing algorithms, and their specific implementation steps are not repeated here.
[0125] Specifically, in one embodiment of the present invention, in step S5, the calculation formula of the nonlinear weighted model is: Where, is the initial computing power thermal value of the jth region; is the weight of the kth initial indicator; The normalized value of the kth feature in the computational power evaluation feature set for the jth region; is the nonlinear adjustment coefficient of the kth feature; is the policy impact intensity of the jth region; is the policy impact coefficient; The number of features in the feature set for computational power evaluation.
[0126] Specifically, the model reflects the interaction between indicators in the form of products and introduces policy impact items to achieve a comprehensive quantitative assessment of the regional computing power development level. During the calculation process, the eigenvalues in the computing power evaluation feature set are first standardized to ensure the comparability of indicators of different dimensions. Feature standardization uses the range standardization method to map all eigenvalues to the [0,1] interval. Nonlinear adjustment coefficient The setting is based on the importance and change sensitivity of the features, determined by combining expert experience and data analysis, and is used to adjust the nonlinear change characteristics of the contribution of different features. Policy document analysis results from steps S23-24, policy impact coefficient The impact of policy factors on the hash rate heat value is determined through historical data regression analysis. Using this model, the initial hash rate heat value of all regions can be calculated.
[0127] Specifically, in one embodiment of the present invention, in step S6, the optimization algorithm adopts a hybrid optimization algorithm of batch normalization adjustment and momentum gradient descent, and introduces a computing power feature adaptive weight mechanism. The iterative process of the optimization algorithm is as follows:
[0128] (1) Prepare input data, including the computing power evaluation feature set of n regions , initial indicator weight , k=1,2,...,m (m is the number of features), actual computing power required ,j=1,2,...,n。 Feature importance coefficient Determined based on the results of feature correlation analysis in S3.
[0129] (2) Initialize the parameters, including the regularization coefficient , feature importance adjustment factor , initial learning rate , attenuation coefficient , Adaptive Index , Momentum Factor and the momentum vector .
[0130] (3) Execute iterative optimization process
[0131] At each iteration t, execute:
[0132] 1. Calculate the computing power thermal value under the current weight as the predicted value , the specific calculation method can still be obtained based on the nonlinear weighted model.
[0133] 2. Calculate the error evaluation loss function: in, is the importance coefficient of the kth feature, determined based on the feature correlation analysis results in step S3; is the regularization coefficient; m is the number of features; is the feature importance adjustment factor. In this loss function, the first term is the error term, the second term is the L2 regularization term, and the third term is the feature importance weighting term.
[0134] 3. For each weight Compute the gradient: in, That is, the gradient of the kth weight.
[0135] 4. Batch normalization gradient processing: in, is the gradient after batch normalization, is the mean of all feature gradients, is the variance of all feature gradients, A small constant to prevent division by zero in batch normalization.
[0136] 5. Update the momentum term: in, is the momentum term of the kth feature at t iterations, is the momentum factor.
[0137] 6. Calculate the adaptive learning rate: in, is the adaptive learning rate at t round iterations.
[0138] 7. Update all weights at the same time: The iteration stops when any of the following conditions is met: the change in Loss for two consecutive rounds is less than the threshold ; Reach the maximum number of iterations. The output result is the optimized indicator weight set .
[0139] Specifically, in one embodiment of the present invention, step S7 includes:
[0140] Based on the optimized indicator weights, the regional computing power thermal value is recalculated according to the nonlinear weighted model. The initial Voronoi spatial structure is constructed with the regional location set as the construction unit and the regional computing power thermal value as the threshold. The computing power thermal field structure of the adjacent areas is dynamically adjusted based on the space-time propagation function. The space-time propagation function takes into account the geographical distance between regions and the thermal diffusion parameters. By dynamically updating the Voronoi thermal area boundaries, a space-time distribution map that supports real-time interaction is generated, and a computing power resource allocation plan is generated based on the thermal gradient differences.
[0141] Specifically, in this embodiment, the process of generating the thermal spatiotemporal distribution map is as follows:
[0142] A1. Initial Voronoi spatial structure construction. Its input is the region position set ,in is the geographical coordinates of the i-th region; the computing power thermal value set ; For each regional location point , construct its Voronoi polygons : Where p is any point on the plane, represents the Euclidean distance.
[0143] A2. Dynamic adjustment of the space-time thermal field. The space-time propagation function is designed as: in, is the geographical distance between regions i and j; is the thermal diffusion parameter, which controls the thermal influence range; is the time fluctuation coefficient, reflecting the periodic change of heat over time; is the time period parameter, which controls the frequency of thermal changes; t is the time variable.
[0144] A3. Thermal field superposition calculation. The comprehensive thermal value of region i at time t is: in, is the spatial weight coefficient, , represents the distance decay effect; is the basic thermal value of area i.
[0145] A4. Dynamic update of Voronoi region boundaries.
[0146] Calculate the thermal gradient between adjacent regions:
[0147] Boundary adjustment function:
[0148] Where B(i,j) is the dynamic boundary point between regions i and j.
[0149] In this embodiment, the computing power resource allocation plan generation process is as follows:
[0150] (1) Calculate the thermal gradient matrix G:
[0151] (2) Identify high gradient region pairs: If (gradient threshold), then mark (i, j) as the region pair that needs to be adjusted.
[0152] (3) Generate allocation suggestions: For the marked region pair (i, j), calculate the recommended allocation amount: in: is the adjustment coefficient, The maximum single allocation capacity.
[0153] In summary, the computing power thermal assessment method proposed in the present invention establishes a complete technical system from data acquisition, preprocessing, feature extraction to thermal assessment through multi-dimensional data intelligent analysis and deep integration. This method standardizes and integrates multi-source heterogeneous data such as active subjects, economic development, resources and environment, and combines segmented adaptive normalization and policy semantic analysis to achieve accurate quantification of the level of regional computing power development. Through a hierarchical evaluation system and feature selection based on local sensitive hashing, as well as key technologies such as dynamic weight optimization and Voronoi thermal field visualization, this method not only provides objective and reliable computing power assessment results, but also realizes dynamic monitoring and intelligent allocation of regional computing power distribution, providing a powerful decision-making support tool for promoting the coordinated development of regional computing power and optimizing resource allocation.
[0154] In addition, if Figure 3 As shown, the present invention also provides a computing power decision-making system based on multidimensional data analysis, which is used to execute any of the above methods, and the system includes:
[0155] The data acquisition module is used to collect active subject data, economic data, resource and environmental data, and policy documents, and perform standardized preprocessing to generate standardized data sets; specifically, a rule engine is used to define data cleaning rules, including outlier correction thresholds, missing value filling strategies, and data integrity scoring criteria; a dynamic time warping algorithm is used to align data from different collection periods to a unified reference time axis, and Fourier transform is used to eliminate periodic noise interference.
[0156] The data fusion module is used to vectorize and fuse standardized data sets into heterogeneous data to generate a unified fused data set. Specifically, the active subject data, economic data, and resource and environmental data are vectorized using a segmented adaptive normalization method. A quality assessment model based on data integrity, timeliness, and accuracy is constructed. The word frequency-inverse document frequency method is used to quantify the intensity of policy impact and establish association rules between policy clauses and feature vectors.
[0157] The feature analysis module is used to build a hierarchical evaluation system, extract evaluation indicator features from the fused data set, and perform correlation analysis and feature selection to form a computing power evaluation feature set. Specifically, a two-layer evaluation indicator system is constructed that includes the computing power demand level, computing power support conditions, and computing power development environment. A feature redundancy elimination algorithm based on local sensitive hashing is used to screen significant features by calculating the Pearson correlation coefficient of the indicator pairs in the hash bucket.
[0158] The regional clustering module is used to configure differentiated initial indicator weights based on the hierarchical analysis method and the regional clustering algorithm. Specifically, the hierarchical analysis method is combined to determine the indicator level weights, and the regions are grouped through the clustering algorithm. Differentiated initial weight coefficients are configured for different types of regions based on regional development characteristics and computing power demand levels.
[0159] The thermal evaluation module is used to combine the computing power evaluation feature set and the initial indicator weight, calculate the initial regional computing power thermal value through a nonlinear weighted model, and optimize the indicator weight according to the output of the weight optimization module to update the computing power thermal value; specifically, based on the nonlinear weighted model Calculate thermal value; adjust coefficient by characteristic nonlinearity and policy impact coefficient Realize multi-dimensional dynamic evaluation; support real-time update of thermal calculation results based on optimized weights.
[0160] The weight optimization module is used to iteratively optimize the initial indicator weights according to the actual computing power requirements; specifically, it adopts a hybrid optimization algorithm of batch normalization adjustment and momentum gradient descent; introduces a computing power feature adaptive weight mechanism; and uses the feature importance coefficient And adaptive learning rate strategy to improve optimization efficiency.
[0161] The result display module is used to generate a thermal spatiotemporal distribution map based on the computing power thermal value and output the computing power resource allocation plan. Specifically, the thermal field is constructed based on the Voronoi spatial structure; the spatiotemporal propagation function is used to generate a thermal spatiotemporal distribution map based on the computing power thermal value and output the computing power resource allocation plan. Dynamically adjust the thermal field structure of adjacent areas; generate recommendations for inter-regional computing resource allocation through thermal gradient analysis; and support real-time interactive updates of thermal maps.
[0162] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A computing power thermal evaluation method based on multidimensional data analysis, characterized in that: include: S1. Obtain the original dataset containing active subject data, economic data, resource and environmental data, and policy documents, preprocess the original dataset, and generate a standardized dataset; S2. Perform vectorization and heterogeneous fusion on the standardized dataset to generate a unified fusion dataset; Step S2 includes: S21. For the active subject data, economic data, and resource and environmental data in the standardized data set, a piecewise adaptive normalization method is used for vectorization processing to convert data of different dimensions into a unified numerical vector; S22. Construct a data quality assessment model, calculate a quality score based on the three dimensions of data integrity, timeliness, and accuracy, determine the weight of each dimension, generate a comprehensive quality score, and use the comprehensive quality score as an adjustment coefficient to perform weighted adjustment on the initial numerical vector to generate a quality-weighted numerical vector; S23. Conduct in-depth semantic analysis of policy documents to extract policy clauses related to computing power. Quantify the impact of policies based on the term frequency-inverse document frequency method and establish association rules between policy clauses and quality-weighted numerical vectors. S24, dynamically adjusting the policy-affected vectors in the quality-weighted numerical vectors according to the association rules; S25, aligning and integrating the adjusted numerical vectors according to the time dimension and the spatial dimension to generate a unified fusion data set; S3. Build a hierarchical evaluation system and extract corresponding evaluation indicator features from the fused data set. Perform correlation analysis based on the evaluation indicator features to form a computing power evaluation feature set. The hierarchical evaluation system is a two-tier evaluation indicator system: The first-level indicators include computing power demand level, computing power support conditions, and computing power development environment; Secondary indicators include: The density of information technology companies, the proportion of internet companies, the number of large and medium-sized enterprises, and the digital coverage of key industries are set based on the level of computing power demand; Under the conditions of computing power support, the IDC facility capacity, power supply guarantee capability, network bandwidth level, and computer room environment adaptability should be set; Under the computing power development environment, the proportion of information industry investment, industrial policy support, resource element guarantee, and infrastructure completeness are set; S4. Using the analytic hierarchy process and regional clustering algorithm, configure differentiated initial indicator weights for different types of regions; S5. Calculate the initial computing power thermal value of each region using a nonlinear weighted model, combining the computing power evaluation feature set and the initial indicator weights. In step S5, the calculation formula of the nonlinear weighted model is: Where H j is the initial computing power thermal value of the jth region; w k is the initial index weight of the kth feature; x jk is the normalized value of the kth feature in the computing power evaluation feature set of the jth region; α k is the nonlinear adjustment coefficient of the kth feature; P j is the policy impact intensity of the jth region; β is the policy impact coefficient; m is the number of features in the computing power evaluation feature set; S6. Obtain the actual computing power requirements of each region, use the optimization algorithm to iteratively optimize the initial indicator weights, and generate optimized indicator weights; S7. Recalculate the regional computing power thermal value based on the optimized indicator weights, generate a thermal spatiotemporal distribution map, and generate a computing power resource allocation plan based on the thermal gradient differences.
2. The method for computing power thermal evaluation based on multidimensional data analysis according to claim 1 is characterized in that: The preprocessing of the original data set in step S1 includes: Define data cleaning rules through the rule engine and clean the original data set according to the data cleaning rules. The data cleaning rules include outlier correction thresholds, missing value filling strategies, and data integrity scoring criteria. The dynamic time warping algorithm is used to align the cleaned data of different acquisition periods to a unified reference time axis, and the periodic noise interference is eliminated based on Fourier transform to generate a standardized data set.
3. The computing power thermal evaluation method based on multidimensional data analysis according to claim 1 is characterized in that: Active subject data includes population data, household data, and enterprise data; economic data includes GDP data, industrial data, transportation and logistics data, education data, financial data, commercial data, agricultural data, and medical data; resource and environmental data includes power generation data, power structure data, temperature data, and IDC facility data; policy documents include computing hub policies, regional planning policies, industrial support policies, and energy management policies.
4. The method for computing power thermal evaluation based on multidimensional data analysis according to claim 1 is characterized in that: Step S21 includes: S211. Perform logarithmic transformation and normalization on economic data: Where, X' eco is the normalized economic data; X eco is the original economic data; min eco is the minimum value of economic data; max eco is the maximum value of economic data; δ is the smoothing factor, which ranges from 1×10 -5 ≤δ≤1×10 -3 ; S212, construct a segmented mapping function to normalize the active subject data, and determine the threshold set T = {t1, t2, ..., t n }, n is the threshold number, when the active subject data X pop ∈(t i-1 ,t i ], the mapping function is as follows: X’ pop =k i ·X pop +b i ,i=1,2,…,n Where, X' pop is the normalized active subject data; k i is the linear transformation coefficient; b i is the bias constant; k i and b i Determined by calculating the segment endpoint values; S213. Perform nonlinear mapping on resource and environmental data: Where, X' env is the resource and environmental data after normalization; σ is the normalization coefficient; γ is the nonlinear adjustment factor; X env is the original resource and environmental data; X0 is the environmental suitability threshold; S214. Organize all normalized values of the same data source at the same time point into a numerical vector V(s,t): V(s,t)=[{X′ eco (s,t)},{X′ pop (s,t)},{X′ env (s,t)}] Where s is the unique identifier of the data source; t is the time point.
5. The method for computing power thermal evaluation based on multidimensional data analysis according to claim 1 is characterized in that: In step S3, the correlation analysis adopts a feature redundancy elimination algorithm based on local sensitive hashing. By calculating the Pearson correlation coefficient of the indicator pairs in the hash bucket, combined with the set threshold, significant features are retained as the computing power evaluation feature set.
6. The method for computing power thermal evaluation based on multidimensional data analysis according to claim 1, characterized in that: The method for generating a thermal spatiotemporal distribution map includes: constructing an initial Voronoi spatial structure using a set of regional locations as a construction unit and a regional computing power thermal value as a threshold; dynamically adjusting the computing power thermal field structure of adjacent areas based on a spatiotemporal propagation function, wherein the spatiotemporal propagation function takes into account the geographical distance between regions and thermal diffusion parameters; and generating a spatiotemporal distribution map that supports real-time interaction by dynamically updating the Voronoi thermal area boundaries.
7. A computing power decision-making system based on multi-dimensional data analysis, characterized in that: The system is used to perform the method according to any one of claims 1 to 6, and the system includes: The data acquisition module is used to collect active subject data, economic data, resource and environmental data, and policy documents, and perform standardized preprocessing to generate standardized data sets; The data fusion module is used to perform vectorization processing and heterogeneous fusion on the standardized data set to generate a unified fused data set; The feature analysis module is used to build a hierarchical evaluation system, extract evaluation indicator features from the fused data set, and perform correlation analysis and feature selection to form a computing power evaluation feature set; Regional clustering module, used to configure differentiated initial indicator weights based on the analytic hierarchy process and regional clustering algorithm; The thermal evaluation module is used to combine the computing power evaluation feature set and the initial indicator weights to calculate the initial regional computing power thermal value through a nonlinear weighted model, and optimize the indicator weights based on the output of the weight optimization module to update the computing power thermal value; The weight optimization module is used to iteratively optimize the initial indicator weights based on the actual computing power requirements; The result display module is used to generate a thermal spatiotemporal distribution map based on the computing power thermal value and output a computing power resource allocation plan.
Citation Information
Patent Citations
Calculation power analysis method and system for shared server and storage medium
CN119988017A