Intelligent computing center-oriented multi-type hardware equipment energy efficiency evaluation system

By collecting data from all dimensions and extracting dynamic load features, combined with cross-device collaborative modeling, targeted optimization strategies are generated, which solves the shortcomings of energy efficiency assessment for multiple types of hardware devices in existing technologies and realizes energy efficiency optimization and resource management of intelligent computing centers.

CN121807666APending Publication Date: 2026-04-07LIAOYANG POWER SUPPLY COMPANY OF STATE GRID LIAONING ELECTRIC POWER SUPPLY +2
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing hardware energy efficiency assessment technologies cannot adapt to multiple types of hardware, dynamically respond to load changes, or provide cross-device collaborative modeling, resulting in biased assessment results that cannot guide energy efficiency optimization in intelligent computing centers.

Method used

Design an energy efficiency evaluation system for various types of hardware devices in intelligent computing centers. Through full-dimensional data collection, dynamic load feature extraction, multi-dimensional evaluation, and cross-device collaborative modeling, generate targeted optimization strategies and form a closed-loop management mechanism.

Benefits of technology

It enables accurate energy efficiency assessment and collaborative optimization of various types of hardware devices, improves the energy efficiency and resource utilization of the intelligent computing center, and supports continuous optimization management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121807666A_ABST
    Figure CN121807666A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent computing center-oriented multi-type hardware equipment energy efficiency evaluation system, which realizes dynamic and accurate evaluation of individual and overall collaborative energy efficiency by collecting full-dimensional data of multi-type hardware in an intelligent computing center, and further quantifies the influence of resource interaction on the energy efficiency by establishing a cross-equipment collaborative relation model, thereby improving the energy efficiency evaluation efficiency. A targeted optimization strategy is generated, and a closed-loop management mechanism integrating energy efficiency evaluation, intelligent decision, instruction execution and data traceability is finally formed, so that continuous optimization and improvement of the energy efficiency of the intelligent calculation center are comprehensively driven; the system comprises a multi-type hardware data acquisition module, a dynamic load feature extraction module, a multi-dimensional energy efficiency evaluation module, a cross-device collaborative modeling module, a dynamic optimization decision module, a visual interaction module and a data storage and traceability module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hardware performance evaluation technology, specifically to an energy efficiency evaluation system for various types of hardware devices in intelligent computing centers. Background Technology

[0002] As the core carrier of computing infrastructure, intelligent computing centers integrate hardware devices with various architectures and functions. These devices form complex collaborative relationships during task execution, and their energy efficiency directly affects the operating costs and sustainability of the intelligent computing center. However, existing hardware energy efficiency assessment technologies have many shortcomings and are unable to meet the complex needs of intelligent computing centers.

[0003] Existing evaluation technologies are mostly designed for single types of hardware devices, focusing only on the power consumption and utilization of CPUs or GPUs. They lack energy efficiency assessments for auxiliary hardware such as storage arrays and high-speed network devices, and fail to consider the overall energy efficiency correlation when multiple types of devices work together. This results in biased evaluation results that cannot reflect the true energy efficiency status of intelligent computing centers. Secondly, existing evaluation methods mostly use static evaluation indicators, pre-setting evaluation parameters based on fixed load scenarios. However, the task load of intelligent computing centers is dynamically fluctuating, and different tasks have significantly different hardware resource requirements. Static evaluations cannot adapt to load changes, resulting in insufficient evaluation accuracy and difficulty in guiding actual energy efficiency optimization.

[0004] Furthermore, existing evaluation systems rely on a limited set of dimensions, focusing primarily on hardware power consumption or resource utilization while neglecting crucial factors such as hardware device compatibility with task types, equipment operational stability, and long-term lifespan degradation. This results in evaluations that fail to comprehensively reflect the overall energy efficiency of the hardware. Simultaneously, existing systems lack cross-device collaborative modeling capabilities, failing to quantify the impact of resource competition and data transmission latency between different hardware devices on overall energy efficiency, making it difficult to generate targeted collaborative optimization strategies. Finally, existing evaluation results are disconnected from actual optimization actions, merely providing energy efficiency data displays without establishing a closed-loop mechanism for evaluation, decision-making, and execution, thus failing to effectively guide hardware resource scheduling and energy efficiency optimization practices in intelligent computing centers.

[0005] These issues make it difficult for intelligent computing centers to accurately grasp the energy efficiency status of various types of hardware devices, hindering optimal resource allocation, resulting in wasted computing power and excessive energy consumption, and restricting the green and efficient operation of intelligent computing centers. Therefore, there is an urgent need for an energy efficiency assessment system that can adapt to multiple types of hardware, dynamically respond to load changes, conduct multi-dimensional comprehensive evaluation, and perform cross-device collaborative modeling to address the shortcomings of existing technologies. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides an energy efficiency assessment system for multiple types of hardware devices in intelligent computing centers. This system collects comprehensive data on various types of hardware within the intelligent computing center, enabling dynamic and accurate assessment of individual and overall collaborative energy efficiency. Furthermore, by establishing a cross-device collaborative relationship model, it quantifies the impact of resource interaction on energy efficiency and generates targeted optimization strategies. Ultimately, it forms a closed-loop management mechanism integrating energy efficiency assessment, intelligent decision-making, instruction execution, and data traceability, thereby comprehensively driving the continuous optimization and improvement of energy efficiency in intelligent computing centers.

[0007] To achieve the above objectives, the present invention provides the following technical solution: an energy efficiency evaluation system for multiple types of hardware devices in intelligent computing centers, comprising:

[0008] Multiple types of hardware data acquisition modules are used to collect full-dimensional operational data from CPUs, GPUs, FPGAs, storage arrays, and high-speed network devices within the intelligent computing center;

[0009] The dynamic load feature extraction module receives the output data from the multi-type hardware data acquisition module and extracts the time-series features and correlation features of the dynamic load.

[0010] The multi-dimensional energy efficiency assessment module, based on the output of the dynamic load feature extraction module, calculates the individual energy efficiency index and the overall collaborative energy efficiency index of various types of hardware devices through a comprehensive energy efficiency assessment algorithm with dynamic weight adjustment.

[0011] The cross-device collaborative modeling module establishes a model of resource interaction and collaborative relationships between multiple types of hardware devices, and quantifies the impact of cross-device collaboration on overall energy efficiency.

[0012] The dynamic optimization decision-making module generates targeted hardware resource scheduling and energy efficiency optimization instructions based on the evaluation results of the multi-dimensional energy efficiency assessment module and the model output of the cross-device collaborative modeling module.

[0013] The visualization and interaction module is used to display energy efficiency assessment results, load characteristics, collaborative relationships and optimization suggestions, and supports user interaction.

[0014] The data storage and traceability module stores the collected raw data, feature data, evaluation results, and optimization decision records, and provides data traceability functionality.

[0015] Furthermore,

[0016] The multi-type hardware data acquisition module includes:

[0017] The CPU data acquisition unit collects real-time CPU clock speed, core utilization, cache hit rate, power consumption and temperature data at a frequency of 20Hz and communicates with the CPU hardware monitoring chip through the PCIe 4.0 interface.

[0018] The GPU data acquisition unit collects data on GPU computing power utilization, memory bandwidth, memory usage, core temperature, power consumption, and task execution latency at a frequency of 30Hz. Data acquisition is achieved through a hardware monitoring interface based on the CUDA Toolkit or ROCm architecture.

[0019] The FPGA data acquisition unit collects data on FPGA logic resource utilization, DSP unit occupancy, input / output interface bandwidth, power consumption, and operating temperature at a frequency of 15Hz. Data acquisition is achieved through the JTAG interface or the FPGA's built-in monitoring module.

[0020] The storage array data acquisition unit collects data on the storage array's IOPS, throughput, read / write latency, cache hit rate, power consumption, and hard drive health status at a frequency of 10Hz. Data acquisition is achieved based on the SCSI protocol or the management interface provided by the storage device manufacturer.

[0021] The high-speed network equipment data acquisition unit collects port bandwidth utilization, packet forwarding delay, packet loss rate, power consumption and port connection status data of switches and routers. The acquisition frequency is 15Hz, and the data acquisition is achieved through SNMP or NetFlow protocol.

[0022] The multi-type hardware data acquisition module adopts a distributed acquisition architecture. Each acquisition unit transmits data to the data aggregation node via 5G industrial Ethernet, and the transmission process is encrypted with AES-256.

[0023] Furthermore,

[0024] The dynamic load feature extraction module includes:

[0025] The time series feature extraction unit uses the sliding window method to segment the collected time series data. The window size is 10s and the step size is 5s. It extracts the mean, variance, peak value, valley value and trend slope features of each hardware operating parameter.

[0026] The correlation feature extraction unit calculates the correlation strength between the operating parameters of different hardware devices based on the Pearson correlation coefficient and mutual information entropy. The correlation feature calculation formula is as follows:

[0027] ,

[0028] in, For hardware devices and Mutual information entropy, for parameters and parameters The joint probability distribution, , They are respectively , The marginal probability distribution;

[0029] The workload type identification unit, based on extracted temporal and correlation features, uses a lightweight convolutional neural network model to identify workload types, including four core workloads: AI training, AI inference, big data analysis, and scientific computing. The specific identification steps are as follows:

[0030] Feature vector construction: The temporal features and related features output by the dynamic load feature extraction module are concatenated and standardized to form a unified, fixed-dimensional comprehensive feature vector, which serves as the input to the neural network;

[0031] Forward computation of the model: The constructed comprehensive feature vector is input into the pre-trained lightweight convolutional neural network. The neural network performs hierarchical processing and nonlinear transformation on the input features through its internal convolutional layers, pooling layers and fully connected layers. Finally, a confidence score is calculated for each of the four core loads at the output layer.

[0032] Probability distribution generation and type determination: The four confidence scores output by the model are normalized and converted into a probability distribution, where the probability value corresponding to each load type is between 0 and 1, and the sum of all probabilities is 1. The system finally selects the load type with the highest probability value as the identification result; if the highest probability value is lower than the preset confidence threshold, the current load is marked as "mixed type" and the result is indicated as uncertain.

[0033] Furthermore,

[0034] The multi-dimensional energy efficiency assessment module includes:

[0035] Individual energy efficiency assessment units, targeting single hardware devices, construct a multi-dimensional assessment index system, including resource utilization rate. Power efficiency Task adaptability and equipment health Calculate the individual energy efficiency index The formula is:

[0036] ,

[0037] in, , , , For dynamic weighting coefficients, satisfying , This represents the normalized power consumption per unit task.

[0038] The dynamic weight adjustment unit adaptively adjusts weight coefficients based on the load type identified by the dynamic load feature extraction module. The adjustment rules are as follows: The system pre-defines weight configuration strategies for different load types. Once the dominant load type is identified, the unit invokes the corresponding strategy to assign specific base weights to resource utilization U, power efficiency P, task adaptability A, and device health H. These weights are then normalized, with the sum of all weight coefficients equal to 1. For AI training loads, the weights of power efficiency P and task adaptability A are increased. For big data analysis loads, the weights of resource utilization U and device health H are increased. The weight adjustment cycle is consistent with the load feature extraction cycle.

[0039] The overall collaborative energy efficiency assessment unit, combined with the output of the cross-device collaborative modeling module, calculates the overall collaborative energy efficiency index. The formula is:

[0040] ,

[0041] in, The total number of hardware devices. For the first Individual energy efficiency index of hardware devices For the cooperative gain coefficient, For the first The compatibility and interoperability of an individual device with other devices The value range is [0,1].

[0042] Furthermore,

[0043] The cross-device collaborative modeling module includes:

[0044] The collaborative relationship identification unit identifies the types of collaborative relationships between devices based on real-time data and dynamic load characteristics of multiple types of hardware data acquisition modules. These include three types: computing power collaboration, data transmission collaboration, and storage access collaboration. The collaborative relationship is represented by a directed graph model, where nodes are hardware devices and edge weights represent the collaboration strength.

[0045] The collaboration strength calculation unit calculates collaboration strength based on data transmission latency, resource contention level, and task dependencies. The formula is:

[0046] ,

[0047] in, This represents the average data transmission latency between devices. This is the resource competition coefficient. For task dependency, , , Let be the proportionality coefficient, satisfying ;

[0048] The impact of synergistic energy efficiency is quantified by establishing a correlation model between synergistic strength and overall energy efficiency. This model quantifies the positive gain or negative loss of overall energy efficiency due to different synergistic relationships. The model is trained using a multiple linear regression algorithm, and its formula is as follows:

[0049]

[0050] in: The energy efficiency impact coefficient represents the degree of influence of the synergistic relationship on overall energy efficiency. , , These are computing power collaboration strength, data transmission collaboration strength, and storage access collaboration strength, respectively. This is the intercept term of the regression model; , , These are the regression coefficients corresponding to the synergy strength, representing the weights of the impact of various synergy relationships on energy efficiency; This is the random error term;

[0051] The model was trained using historical synergy intensity data and actual energy efficiency data, and its goodness of fit was [not specified]. The value should be no less than 0.85. After training, input the real-time collaborative strength vector. This will output a quantified energy efficiency impact coefficient. .

[0052] Furthermore,

[0053] The dynamic optimization decision module includes:

[0054] The optimization target determination unit determines optimization targets based on multi-dimensional energy efficiency assessment results. These targets include three categories: maximizing individual energy efficiency, maximizing overall collaborative energy efficiency, and prioritizing energy efficiency for key tasks. Users can manually select these targets or the system can automatically match them based on the load type.

[0055] The optimization strategy generation unit generates targeted optimization strategies for different optimization objectives. In the scenario of computing power collaboration, it adjusts the computing power allocation ratio of hardware devices; in the scenario of data transmission collaboration, it optimizes network bandwidth allocation and data transmission protocol; and in the scenario of storage access collaboration, it adjusts storage caching strategy and IO priority.

[0056] The instruction generation and distribution unit converts optimization strategies into executable hardware control instructions, supports interface with intelligent computing center resource scheduling platform and hardware management interface, with instruction distribution delay ≤1s, and execution results are fed back to multi-type hardware data acquisition modules in real time, forming a closed-loop control.

[0057] Furthermore,

[0058] The visual interaction module includes:

[0059] The energy efficiency status display unit uses dashboards, heat maps, and collaborative relationship diagrams to display the individual energy efficiency index, overall collaborative energy efficiency index, load characteristics, and equipment operating parameters of various types of hardware devices in real time, with the update frequency synchronized with the data acquisition frequency.

[0060] The optimization suggestion push unit pushes optimization suggestions in the form of a combination of text and charts based on the output of the dynamic optimization decision module, including the direction of resource adjustment, parameter configuration scheme and expected optimization effect;

[0061] The user interaction unit supports users to customize evaluation indicator weights, select optimization goals, and query historical evaluation data. It provides multi-dimensional data filtering and comparison functions, and the interaction response latency is ≤500ms.

[0062] Furthermore,

[0063] The data storage and traceability module includes:

[0064] The distributed storage unit uses a time-series database to store raw running data, feature data, evaluation results, and optimization decision records. The storage period is configurable, and it supports high-concurrency read and write operations as well as data compression storage with a compression ratio of no less than 5:1.

[0065] The data traceability unit adds a unique traceability identifier to each evaluation result and optimization decision record, associating it with the corresponding original data, load characteristics, and modeling parameters, and supports multi-dimensional traceability queries by time, hardware type, load type, etc.

[0066] The data security unit employs a hierarchical data storage strategy, encrypts sensitive data, sets access control, and supports operation log auditing to ensure data integrity and security.

[0067] Furthermore,

[0068] It also includes a load prediction submodule, which uses a long short-term memory network model to predict the load change trend in the next 10-60 minutes based on historical data from the dynamic load feature extraction module. The prediction error is no more than 10%. The prediction results are pushed to the multi-dimensional energy efficiency assessment module and the dynamic optimization decision module to adjust the assessment weights and optimization strategies in advance.

[0069] Furthermore,

[0070] It also includes an anomaly warning submodule, which sets individual energy efficiency index thresholds and collaborative energy efficiency index thresholds based on the evaluation results of the multi-dimensional energy efficiency evaluation module. When the evaluation index is lower than the threshold or the decline exceeds the preset value in a short period of time, an anomaly warning is triggered. The warning methods include audible and visual alarms, SMS notifications, and system pop-ups. At the same time, it pushes anomaly cause analysis and handling suggestions. The anomaly recognition delay is ≤3s.

[0071] This invention provides an energy efficiency assessment system that can adapt to multiple types of hardware, dynamically respond to load changes, perform multi-dimensional comprehensive evaluation, and conduct cross-device collaborative modeling, and has the following beneficial effects:

[0072] 1) Achieve comprehensive and accurate energy efficiency assessment across all dimensions, enhancing the comprehensiveness and accuracy of the assessment.

[0073] The system breaks through the traditional model of evaluating single hardware devices. It designs a dedicated data acquisition scheme and evaluation indicators for the characteristics of various types of hardware in intelligent computing centers, such as CPUs, GPUs, FPGAs, storage arrays, and high-speed network devices. This achieves full coverage evaluation of multiple types of hardware devices and adapts to the complex hardware architecture of intelligent computing centers. Through dynamic load feature extraction and load type identification, the system can adapt to load changes in different task scenarios and dynamically adjust evaluation weights and optimization strategies. This avoids the problem of static evaluation being out of touch with the actual load and significantly improves the accuracy and adaptability of the evaluation.

[0074] Secondly, the system has constructed a multi-dimensional evaluation index system that includes resource utilization, power efficiency, task adaptability, and device health. It not only focuses on the real-time operating status of hardware devices, but also takes into account the adaptability of devices to tasks and their long service life. The evaluation results are more comprehensive and more in line with the actual operating needs of the intelligent computing center.

[0075] 2) Establish a cross-device collaborative impact quantification model to support system performance optimization.

[0076] The design of the cross-device collaborative modeling module quantifies the collaborative relationship and energy efficiency impact between different hardware devices, breaking through the limitations of isolated evaluation in existing technologies. Based on this, combined with multi-dimensional evaluation results, the system can identify collaborative bottlenecks and energy efficiency shortcomings, and generate collaborative optimization strategies for overall efficiency, achieving a leap from single-device optimization to multi-device collaborative optimization, providing a scientific basis for global resource optimization in intelligent computing centers.

[0077] 3) Establish a closed-loop mechanism of "assessment-decision-implementation-source tracing" to promote continuous optimization of energy efficiency management.

[0078] Based on the dynamic optimization decision-making module, the system can automatically generate hardware resource scheduling and energy efficiency optimization instructions according to the energy efficiency assessment results and collaborative model output, realizing closed-loop linkage from assessment to execution. With the help of visual interaction and data storage traceability functions, it not only provides operation and maintenance personnel with intuitive energy efficiency status display and interactive operation support, but also completely retains data trajectory and decision records, supports effect review and strategy optimization, thereby building a sustainable and traceable energy efficiency optimization management system, and truly transforming energy efficiency assessment into actual energy efficiency improvement actions.

[0079] In summary, this invention provides intelligent computing centers with a comprehensive, precise, and dynamic solution for evaluating and optimizing the energy efficiency of various types of hardware devices. This effectively improves the energy efficiency, resource utilization, and operational stability of intelligent computing centers, helping them achieve green, efficient, and sustainable operation. Attached Figure Description

[0080] Figure 1 Here is a flowchart of the system of this invention;

[0081] Figure 2 This is a diagram of the Long Short-Term Memory (LSTM) network architecture. Detailed Implementation

[0082] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0083] like Figures 1-2 As shown, it illustrates a specific embodiment of the present invention:

[0084] This invention discloses an energy efficiency assessment system for multiple types of hardware devices in intelligent computing centers. The system includes a multi-type hardware data acquisition module, a dynamic load feature extraction module, a multi-dimensional energy efficiency assessment module, a cross-device collaborative modeling module, a dynamic optimization decision-making module, a visualization interaction module, and a data storage and traceability module. The system collects comprehensive operational data from multiple types of hardware, extracts dynamic load features, and, based on a comprehensive energy efficiency assessment algorithm with dynamic weight adjustment, combines a cross-device collaborative relationship model to achieve accurate energy efficiency assessment and collaborative optimization decisions for multiple types of hardware devices. The system displays assessment results and optimization suggestions through a visual interface and supports data traceability. This invention overcomes the limitations of single-hardware, static assessment, realizing comprehensive energy efficiency assessment of multiple types of devices collaboratively and dynamically adapting to loads. It provides a scientific basis for energy efficiency optimization in intelligent computing centers, contributing to green and efficient operation.

[0085] Specifically, it includes:

[0086] Multiple types of hardware data acquisition modules are used to collect full-dimensional operational data from CPUs, GPUs, FPGAs, storage arrays, and high-speed network devices within the intelligent computing center;

[0087] The dynamic load feature extraction module receives the output data from the multi-type hardware data acquisition module and extracts the time-series features and correlation features of the dynamic load.

[0088] The multi-dimensional energy efficiency assessment module, based on the output of the dynamic load feature extraction module, calculates the individual energy efficiency index and the overall collaborative energy efficiency index of various types of hardware devices through a comprehensive energy efficiency assessment algorithm with dynamic weight adjustment.

[0089] The cross-device collaborative modeling module establishes a model of resource interaction and collaborative relationships between multiple types of hardware devices, and quantifies the impact of cross-device collaboration on overall energy efficiency.

[0090] The dynamic optimization decision-making module generates targeted hardware resource scheduling and energy efficiency optimization instructions based on the evaluation results of the multi-dimensional energy efficiency assessment module and the model output of the cross-device collaborative modeling module.

[0091] The visualization and interaction module is used to display energy efficiency assessment results, load characteristics, collaborative relationships and optimization suggestions, and supports user interaction.

[0092] The data storage and traceability module stores the collected raw data, feature data, evaluation results, and optimization decision records, and provides data traceability functionality.

[0093] Specifically, the multi-type hardware data acquisition module includes:

[0094] The CPU data acquisition unit collects real-time CPU clock speed, core utilization, cache hit rate, power consumption and temperature data at a frequency of 20Hz and communicates with the CPU hardware monitoring chip through the PCIe 4.0 interface.

[0095] The GPU data acquisition unit collects data on GPU computing power utilization, memory bandwidth, memory usage, core temperature, power consumption, and task execution latency at a frequency of 30Hz. Data acquisition is achieved through a hardware monitoring interface based on the CUDA Toolkit or ROCm architecture.

[0096] The FPGA data acquisition unit collects data on FPGA logic resource utilization, DSP unit occupancy, input / output interface bandwidth, power consumption, and operating temperature at a frequency of 15Hz. Data acquisition is achieved through the JTAG interface or the FPGA's built-in monitoring module.

[0097] The storage array data acquisition unit collects data on the storage array's IOPS, throughput, read / write latency, cache hit rate, power consumption, and hard drive health status at a frequency of 10Hz. Data acquisition is achieved based on the SCSI protocol or the management interface provided by the storage device manufacturer.

[0098] The high-speed network equipment data acquisition unit collects port bandwidth utilization, packet forwarding delay, packet loss rate, power consumption and port connection status data of switches and routers. The acquisition frequency is 15Hz, and the data acquisition is achieved through SNMP or NetFlow protocol.

[0099] The multi-type hardware data acquisition module adopts a distributed acquisition architecture. Each acquisition unit transmits data to the data aggregation node via 5G industrial Ethernet, and the transmission process is encrypted with AES-256.

[0100] Specifically, the dynamic load feature extraction module includes:

[0101] The time series feature extraction unit uses the sliding window method to segment the collected time series data. The window size is 10s and the step size is 5s. It extracts the mean, variance, peak value, valley value and trend slope features of each hardware operating parameter.

[0102] The correlation feature extraction unit calculates the correlation strength between the operating parameters of different hardware devices based on the Pearson correlation coefficient and mutual information entropy. The correlation feature calculation formula is as follows:

[0103] ,

[0104] in, For hardware devices and Mutual information entropy, for parameters and parameters The joint probability distribution, , They are respectively , The marginal probability distribution;

[0105] The workload type identification unit, based on extracted temporal and correlation features, uses a lightweight convolutional neural network model to identify workload types, including four core workloads: AI training, AI inference, big data analysis, and scientific computing. The specific identification steps are as follows:

[0106] Feature vector construction: The temporal features and related features output by the dynamic load feature extraction module are concatenated and standardized to form a unified, fixed-dimensional comprehensive feature vector, which serves as the input to the neural network;

[0107] Forward computation of the model: The constructed comprehensive feature vector is input into the pre-trained lightweight convolutional neural network. The neural network performs hierarchical processing and nonlinear transformation on the input features through its internal convolutional layers, pooling layers and fully connected layers. Finally, a confidence score is calculated for each of the four core loads at the output layer.

[0108] Probability distribution generation and type determination: The four confidence scores output by the model are normalized and converted into a probability distribution, where the probability value corresponding to each load type is between 0 and 1, and the sum of all probabilities is 1. The system finally selects the load type with the highest probability value as the identification result; if the highest probability value is lower than the preset confidence threshold, the current load is marked as "mixed type" and the result is indicated as uncertain.

[0109] Specifically, the multi-dimensional energy efficiency assessment module includes:

[0110] Individual energy efficiency assessment units, targeting single hardware devices, construct a multi-dimensional assessment index system, including resource utilization rate (…). ), power efficiency ( Task adaptability ) and equipment health ( ), calculate individual energy efficiency index ( The formula is:

[0111] ,

[0112] in, , , , For dynamic weighting coefficients, satisfying , This represents the normalized power consumption per unit task.

[0113] The dynamic weight adjustment unit adaptively adjusts weight coefficients based on the load type identified by the dynamic load feature extraction module. The adjustment rules are as follows: The system pre-defines weight configuration strategies for different load types. Once the dominant load type is identified, the unit invokes the corresponding strategy to assign specific base weights to resource utilization U, power efficiency P, task adaptability A, and device health H. These weights are then normalized, with the sum of all weight coefficients equal to 1. For AI training loads, the weights of power efficiency P and task adaptability A are increased. For big data analysis loads, the weights of resource utilization U and device health H are increased. The weight adjustment cycle is consistent with the load feature extraction cycle.

[0114] The overall collaborative energy efficiency assessment unit, combined with the output of the cross-device collaborative modeling module, calculates the overall collaborative energy efficiency index. The formula is:

[0115] ,

[0116] in, The total number of hardware devices. For the first Individual energy efficiency index of hardware devices For the cooperative gain coefficient, For the first The compatibility and interoperability of an individual device with other devices The value range is [0,1].

[0117] Specifically, the cross-device collaborative modeling module includes:

[0118] The collaborative relationship identification unit identifies the types of collaborative relationships between devices based on real-time data and dynamic load characteristics of multiple types of hardware data acquisition modules. These include three types: computing power collaboration, data transmission collaboration, and storage access collaboration. The collaborative relationship is represented by a directed graph model, where nodes are hardware devices and edge weights represent the collaboration strength.

[0119] The collaboration strength calculation unit calculates collaboration strength based on data transmission latency, resource contention level, and task dependencies. The formula is:

[0120] ,

[0121] in, This represents the average data transmission latency between devices. This is the resource competition coefficient. For task dependency, , , Let be the proportionality coefficient, satisfying ;

[0122] The impact of synergistic energy efficiency is quantified by establishing a correlation model between synergistic strength and overall energy efficiency. This model quantifies the positive gain or negative loss of overall energy efficiency due to different synergistic relationships. The model is trained using a multiple linear regression algorithm, and its formula is as follows:

[0123]

[0124] in: The energy efficiency impact coefficient represents the degree of influence of the synergistic relationship on overall energy efficiency. , , These are computing power collaboration strength, data transmission collaboration strength, and storage access collaboration strength, respectively. This is the intercept term of the regression model; , , These are the regression coefficients corresponding to the synergy strength, representing the weights of the impact of various synergy relationships on energy efficiency; This is the random error term;

[0125] The model was trained using historical synergy intensity data and actual energy efficiency data, and its goodness of fit was [not specified]. The value should be no less than 0.85. After training, input the real-time collaborative strength vector. This will output a quantified energy efficiency impact coefficient. .

[0126] Specifically, the dynamic optimization decision module includes:

[0127] The optimization target determination unit determines optimization targets based on multi-dimensional energy efficiency assessment results. These targets include three categories: maximizing individual energy efficiency, maximizing overall collaborative energy efficiency, and prioritizing energy efficiency for key tasks. Users can manually select these targets or the system can automatically match them based on the load type.

[0128] The optimization strategy generation unit generates targeted optimization strategies for different optimization objectives. In the scenario of computing power collaboration, it adjusts the computing power allocation ratio of hardware devices; in the scenario of data transmission collaboration, it optimizes network bandwidth allocation and data transmission protocol; and in the scenario of storage access collaboration, it adjusts storage caching strategy and IO priority.

[0129] The instruction generation and distribution unit converts optimization strategies into executable hardware control instructions, supports interface with intelligent computing center resource scheduling platform and hardware management interface, with instruction distribution delay ≤1s, and execution results are fed back to multi-type hardware data acquisition modules in real time, forming a closed-loop control.

[0130] IO priority refers to dynamically assigning different processing priorities to storage access requests of different tasks or processes based on optimization goals, thereby optimizing the allocation of storage resources, reducing the waiting time of critical tasks, and improving the overall system energy efficiency.

[0131] Specifically, the visual interaction module includes:

[0132] The energy efficiency status display unit uses dashboards, heat maps, and collaborative relationship diagrams to display the individual energy efficiency index, overall collaborative energy efficiency index, load characteristics, and equipment operating parameters of various types of hardware devices in real time, with the update frequency synchronized with the data acquisition frequency.

[0133] The optimization suggestion push unit pushes optimization suggestions in the form of a combination of text and charts based on the output of the dynamic optimization decision module, including the direction of resource adjustment, parameter configuration scheme and expected optimization effect;

[0134] The user interaction unit supports users to customize evaluation indicator weights, select optimization goals, and query historical evaluation data. It provides multi-dimensional data filtering and comparison functions, and the interaction response latency is ≤500ms.

[0135] Specifically, the data storage and traceability module includes:

[0136] The distributed storage unit uses a time-series database to store raw running data, feature data, evaluation results, and optimization decision records. The storage period is configurable, and it supports high-concurrency read and write operations as well as data compression storage with a compression ratio of no less than 5:1.

[0137] The data traceability unit adds a unique traceability identifier to each evaluation result and optimization decision record, associating it with the corresponding original data, load characteristics, and modeling parameters, and supports multi-dimensional traceability queries by time, hardware type, load type, etc.

[0138] The data security unit employs a hierarchical data storage strategy, encrypts sensitive data, sets access control, and supports operation log auditing to ensure data integrity and security.

[0139] The data traceability unit adds a unique traceability identifier to each evaluation result and optimization decision record, and associates it with the corresponding original data, load characteristics, and modeling parameters. Its specific implementation steps are as follows:

[0140] (1) Generate a unique traceability identifier

[0141] When the system generates an evaluation result or optimization decision record, the traceability unit immediately generates a globally unique identifier. This identifier is generated using a composite rule of "timestamp-device group hash value-random number" to ensure its uniqueness.

[0142] (2) Establish association mapping relationship

[0143] The system automatically records metadata information from all data sources relied upon when generating this record and establishes an association mapping with the aforementioned unique identifier. This metadata includes:

[0144] 1) Raw data: Records the time range of the raw data on which the evaluation results were generated and a list of the hardware devices involved.

[0145] 2) Load characteristics: Records the snapshot ID of the load characteristic data used to generate this evaluation.

[0146] 3) Modeling parameters: Record the set of key parameters such as the version of the energy efficiency assessment model used during the assessment, weighting coefficients, and synergistic gain coefficients.

[0147] (3) Store association relationships

[0148] The "unique traceability identifier" and the aforementioned "metadata mapping relationship" are stored together as a single record in a dedicated traceability index table. This table serves as the hub for data retrieval.

[0149] (4) Implement source tracing query

[0150] When a user queries a record through a visual interface, the system completes the tracing process through the following steps:

[0151] 1) Location record: When a user selects an evaluation result record, the system obtains its "unique traceability identifier".

[0152] 2) Index Query: The system uses this identifier to query the "Source Index Table" and obtain the corresponding metadata mapping relationship.

[0153] 3) Data aggregation: Based on the information in the metadata, the system quickly retrieves and aggregates the corresponding raw data, load characteristic snapshots and modeling parameters from the time series database and parameter configuration library.

[0154] 4) Results Display: All retrieved related data will be displayed to the user in a unified manner, thereby completing the entire chain of traceability from the results to the original data.

[0155] Specifically, it also includes a load prediction submodule, which uses a long short-term memory network model to predict the load change trend in the next 10-60 minutes based on historical data from the dynamic load feature extraction module. The prediction error does not exceed 10%, and the prediction results are pushed to the multi-dimensional energy efficiency assessment module and the dynamic optimization decision module for adjusting the assessment weights and optimization strategies in advance.

[0156] Long Short-Term Memory (LSTM) network models process sequence information through cell states and three gating mechanisms. The core process is as follows:

[0157] 1) Forgetting: First, the forget gate decides which old information to discard from the cell state based on the current input and the hidden state of the previous moment.

[0158] 2) Memory: Next, the input gate decides which new information currently being input will be stored in the cell state.

[0159] 3) Update: Then, combine the results of the forget gate and the input gate to update the cell state and form a new state containing long-term memory.

[0160] 4) Output: Finally, the output gate generates the hidden state at the current moment based on the updated cell state, which is used as the output of this step.

[0161] The Long Short-Term Memory (LSTM) network architecture diagram is shown below. Figure 2 As shown.

[0162] The input sequence is a time sequence; The input gate activation vector; Generate candidate values ​​for the state of the memory cell; The forget gate activation vector; The output gate activation vector; This is the output vector; For activation functions; It is the hyperbolic tangent function.

[0163] Specifically, it also includes an anomaly warning submodule. Based on the evaluation results of the multi-dimensional energy efficiency assessment module, it sets individual energy efficiency index thresholds and collaborative energy efficiency index thresholds. When the evaluation index is lower than the threshold or the decline exceeds the preset value in a short period of time, an anomaly warning is triggered. The warning methods include audible and visual alarms, SMS notifications, and system pop-ups. At the same time, it pushes anomaly cause analysis and handling suggestions. The anomaly recognition delay is ≤3s.

[0164] In the anomaly warning submodule, the specific steps for setting the individual energy efficiency index threshold and the collaborative energy efficiency index threshold are as follows:

[0165] 1) Baseline Learning Phase: After initial deployment, the system will enter a learning period of 1-4 weeks. During this period, the system continuously records and analyzes historical energy efficiency indices under different time periods and load types to form a statistical baseline for energy efficiency performance.

[0166] 2) Dynamic Threshold Calculation: Based on the statistical baseline established during the learning period, a dynamic threshold algorithm is used to set the threshold. This algorithm calculates two types of thresholds for each hardware device type and device combination:

[0167] Static threshold: An absolute value threshold based on statistical distribution. For example, the threshold for an individual's energy efficiency index is set at the 5th percentile of historical data; values ​​below this are considered absolute anomalies.

[0168] Dynamic threshold: A relative threshold based on short-term changes. It calculates the threshold for the decline of the energy efficiency index within a specific time window. When the index drops sharply in the short term exceeding this allowable decline rate, an alert is triggered even if the absolute value is not lower than the static threshold.

[0169] 3) Threshold calibration and update: The system automatically recalculates the thresholds periodically, or the administrator can manually trigger calibration based on changes in business strategies to ensure that the thresholds can adapt to the long-term evolution of system operation.

[0170] This method combines absolute performance standards with relative trends, making the threshold setting both scientific and adaptive, avoiding the problem of fixed thresholds being too rigid.

[0171] The preferred embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention. These changes involve related technologies well known to those skilled in the art, and all of them fall within the protection scope of the present invention.

[0172] Many other changes and modifications can be made without departing from the concept and scope of this invention. It should be understood that this invention is not limited to the specific embodiments, and the scope of this invention is defined by the appended claims.

Claims

1. An energy efficiency evaluation system for various types of hardware devices in intelligent computing centers, characterized in that, include: Multiple types of hardware data acquisition modules are used to collect full-dimensional operational data from CPUs, GPUs, FPGAs, storage arrays, and high-speed network devices within the intelligent computing center; The dynamic load feature extraction module receives the output data from the multi-type hardware data acquisition module and extracts the time-series features and correlation features of the dynamic load. The multi-dimensional energy efficiency assessment module, based on the output of the dynamic load feature extraction module, calculates the individual energy efficiency index and the overall collaborative energy efficiency index of various types of hardware devices through a comprehensive energy efficiency assessment algorithm with dynamic weight adjustment. The cross-device collaborative modeling module establishes a model of resource interaction and collaborative relationships between multiple types of hardware devices, and quantifies the impact of cross-device collaboration on overall energy efficiency. The dynamic optimization decision-making module generates targeted hardware resource scheduling and energy efficiency optimization instructions based on the evaluation results of the multi-dimensional energy efficiency assessment module and the model output of the cross-device collaborative modeling module. The visualization and interaction module is used to display energy efficiency assessment results, load characteristics, collaborative relationships and optimization suggestions, and supports user interaction. The data storage and traceability module stores the collected raw data, feature data, evaluation results, and optimization decision records, and provides data traceability functionality.

2. The energy efficiency evaluation system for multiple types of hardware devices in intelligent computing centers according to claim 1, characterized in that, The multi-type hardware data acquisition module includes: The CPU data acquisition unit collects real-time CPU clock speed, core utilization, cache hit rate, power consumption and temperature data at a frequency of 20Hz and communicates with the CPU hardware monitoring chip through the PCIe 4.0 interface. The GPU data acquisition unit collects data on GPU computing power utilization, memory bandwidth, memory usage, core temperature, power consumption, and task execution latency at a frequency of 30Hz. Data acquisition is achieved through a hardware monitoring interface based on the CUDA Toolkit or ROCm architecture. The FPGA data acquisition unit collects data on FPGA logic resource utilization, DSP unit occupancy, input / output interface bandwidth, power consumption, and operating temperature at a frequency of 15Hz. Data acquisition is achieved through the JTAG interface or the FPGA's built-in monitoring module. The storage array data acquisition unit collects data on the storage array's IOPS, throughput, read / write latency, cache hit rate, power consumption, and hard drive health status at a frequency of 10Hz. Data acquisition is achieved based on the SCSI protocol or the management interface provided by the storage device manufacturer. The high-speed network equipment data acquisition unit collects port bandwidth utilization, packet forwarding delay, packet loss rate, power consumption and port connection status data of switches and routers. The acquisition frequency is 15Hz, and the data acquisition is achieved through SNMP or NetFlow protocol. The multi-type hardware data acquisition module adopts a distributed acquisition architecture. Each acquisition unit transmits data to the data aggregation node via 5G industrial Ethernet, and the transmission process is encrypted with AES-256.

3. The energy efficiency evaluation system for multiple types of hardware devices in intelligent computing centers according to claim 1, characterized in that, The dynamic load feature extraction module includes: The time series feature extraction unit uses the sliding window method to segment the collected time series data. The window size is 10s and the step size is 5s. It extracts the mean, variance, peak value, valley value and trend slope features of each hardware operating parameter. The correlation feature extraction unit calculates the correlation strength between the operating parameters of different hardware devices based on the Pearson correlation coefficient and mutual information entropy. The correlation feature calculation formula is as follows: , in, For hardware devices and Mutual information entropy, for parameters and parameters The joint probability distribution, , They are respectively , The marginal probability distribution; The workload type identification unit, based on extracted temporal and correlation features, uses a lightweight convolutional neural network model to identify workload types, including four core workloads: AI training, AI inference, big data analysis, and scientific computing. The specific identification steps are as follows: Feature vector construction: The temporal features and related features output by the dynamic load feature extraction module are concatenated and standardized to form a unified, fixed-dimensional comprehensive feature vector, which serves as the input to the neural network; Forward computation of the model: The constructed comprehensive feature vector is input into the pre-trained lightweight convolutional neural network. The neural network performs hierarchical processing and nonlinear transformation on the input features through its internal convolutional layers, pooling layers and fully connected layers. Finally, a confidence score is calculated for each of the four core loads at the output layer. Probability distribution generation and type determination: The four confidence scores output by the model are normalized and converted into a probability distribution, where the probability value corresponding to each load type is between 0 and 1, and the sum of all probabilities is 1. The system finally selects the load type with the highest probability value as the identification result; if the highest probability value is lower than the preset confidence threshold, the current load is marked as "mixed type" and the result is indicated as uncertain.

4. The energy efficiency evaluation system for multiple types of hardware devices in intelligent computing centers according to claim 1, characterized in that, The multi-dimensional energy efficiency assessment module includes: Individual energy efficiency assessment units, targeting single hardware devices, construct a multi-dimensional assessment index system, including resource utilization rate. Power efficiency Task adaptability and equipment health Calculate the individual energy efficiency index The formula is: , in, , , , For dynamic weighting coefficients, satisfying , This represents the normalized power consumption per unit task. The dynamic weight adjustment unit adaptively adjusts weight coefficients based on the load type identified by the dynamic load feature extraction module. The adjustment rules are as follows: The system pre-defines weight configuration strategies for different load types. Once the dominant load type is identified, the unit invokes the corresponding strategy to assign specific base weights to resource utilization U, power efficiency P, task adaptability A, and device health H. These weights are then normalized, with the sum of all weight coefficients equal to 1. For AI training loads, the weights of power efficiency P and task adaptability A are increased. For big data analysis loads, the weights of resource utilization U and device health H are increased. The weight adjustment cycle is consistent with the load feature extraction cycle. The overall collaborative energy efficiency assessment unit, combined with the output of the cross-device collaborative modeling module, calculates the overall collaborative energy efficiency index. The formula is: , in, The total number of hardware devices. For the first Individual energy efficiency index of hardware devices For the cooperative gain coefficient, For the first The compatibility and interoperability of an individual device with other devices The value range is [0,1].

5. The energy efficiency evaluation system for multiple types of hardware devices in intelligent computing centers according to claim 1, characterized in that, The cross-device collaborative modeling module includes: The collaborative relationship identification unit identifies the types of collaborative relationships between devices based on real-time data and dynamic load characteristics of multiple types of hardware data acquisition modules. These include three types: computing power collaboration, data transmission collaboration, and storage access collaboration. The collaborative relationship is represented by a directed graph model, where nodes are hardware devices and edge weights represent the collaboration strength. The collaboration strength calculation unit calculates collaboration strength based on data transmission latency, resource contention level, and task dependencies. The formula is: , in, This represents the average data transmission latency between devices. This is the resource competition coefficient. For task dependency, , , Let be the proportionality coefficient, satisfying ; The impact of synergistic energy efficiency is quantified by establishing a correlation model between synergistic strength and overall energy efficiency. This model quantifies the positive gain or negative loss of overall energy efficiency due to different synergistic relationships. The model is trained using a multiple linear regression algorithm, and its formula is as follows: , in: The energy efficiency impact coefficient represents the degree of influence of the synergistic relationship on overall energy efficiency. , , These are computing power collaboration strength, data transmission collaboration strength, and storage access collaboration strength, respectively. This is the intercept term of the regression model; , , These are the regression coefficients corresponding to the synergy strength, representing the weights of the impact of various synergy relationships on energy efficiency; This is the random error term; The model was trained using historical synergy intensity data and actual energy efficiency data, and its goodness of fit was [not specified]. The value should be no less than 0.

85. After training, input the real-time collaborative strength vector. This will output a quantified energy efficiency impact coefficient. .

6. The energy efficiency evaluation system for multiple types of hardware devices in intelligent computing centers according to claim 1, characterized in that, The dynamic optimization decision module includes: The optimization target determination unit determines optimization targets based on multi-dimensional energy efficiency assessment results. These targets include three categories: maximizing individual energy efficiency, maximizing overall collaborative energy efficiency, and prioritizing energy efficiency for key tasks. Users can manually select these targets or the system can automatically match them based on the load type. The optimization strategy generation unit generates targeted optimization strategies for different optimization objectives. In the scenario of computing power collaboration, it adjusts the computing power allocation ratio of hardware devices; in the scenario of data transmission collaboration, it optimizes network bandwidth allocation and data transmission protocol; and in the scenario of storage access collaboration, it adjusts storage caching strategy and IO priority. The instruction generation and distribution unit converts optimization strategies into executable hardware control instructions, supports interface with intelligent computing center resource scheduling platform and hardware management interface, with instruction distribution delay ≤1s, and execution results are fed back to multi-type hardware data acquisition modules in real time, forming a closed-loop control.

7. The energy efficiency evaluation system for multiple types of hardware devices in intelligent computing centers according to claim 1, characterized in that, The visual interaction module includes: The energy efficiency status display unit uses dashboards, heat maps, and collaborative relationship diagrams to display the individual energy efficiency index, overall collaborative energy efficiency index, load characteristics, and equipment operating parameters of various types of hardware devices in real time, with the update frequency synchronized with the data acquisition frequency. The optimization suggestion push unit pushes optimization suggestions in the form of a combination of text and charts based on the output of the dynamic optimization decision module, including the direction of resource adjustment, parameter configuration scheme and expected optimization effect; The user interaction unit supports users to customize evaluation indicator weights, select optimization goals, and query historical evaluation data. It provides multi-dimensional data filtering and comparison functions, and the interaction response latency is ≤500ms.

8. The energy efficiency evaluation system for multiple types of hardware devices in intelligent computing centers according to claim 1, characterized in that, The data storage and traceability module includes: The distributed storage unit uses a time-series database to store raw running data, feature data, evaluation results, and optimization decision records. The storage period is configurable, and it supports high-concurrency read and write operations as well as data compression storage with a compression ratio of no less than 5:

1. The data traceability unit adds a unique traceability identifier to each evaluation result and optimization decision record, associating it with the corresponding original data, load characteristics, and modeling parameters, and supports multi-dimensional traceability queries by time, hardware type, load type, etc. The data security unit employs a hierarchical data storage strategy, encrypts sensitive data, sets access control, and supports operation log auditing to ensure data integrity and security.

9. The energy efficiency evaluation system for multiple types of hardware devices in intelligent computing centers according to claim 1, characterized in that, It also includes a load prediction submodule, which uses a long short-term memory network model to predict the load change trend in the next 10-60 minutes based on historical data from the dynamic load feature extraction module. The prediction error is no more than 10%. The prediction results are pushed to the multi-dimensional energy efficiency assessment module and the dynamic optimization decision module to adjust the assessment weights and optimization strategies in advance.

10. The energy efficiency evaluation system for multiple types of hardware devices in intelligent computing centers according to claim 1, characterized in that, It also includes an anomaly warning submodule, which sets individual energy efficiency index thresholds and collaborative energy efficiency index thresholds based on the evaluation results of the multi-dimensional energy efficiency evaluation module. When the evaluation index is lower than the threshold or the decline exceeds the preset value in a short period of time, an anomaly warning is triggered. The warning methods include audible and visual alarms, SMS notifications, and system pop-ups. At the same time, it pushes anomaly cause analysis and handling suggestions. The anomaly recognition delay is ≤3s.

Citation Information

Cited By

  • A data collection method and system based on multi-dimensional intelligent evaluation and electronic equipment

    CN122332731A