Intelligent computing center resource management and control method, system and equipment based on computing and power collaboration

By collecting and integrating computing resources, power grid and energy storage data in real time in the intelligent computing center, and using large models to predict and generate collaborative scheduling strategies, the problem of the separation between computing power and power has been solved, resulting in reduced electricity costs and improved green electricity consumption efficiency, and promoting the transformation of the intelligent computing center towards an efficient and low-carbon operation model.

CN121364938APending Publication Date: 2026-01-20CHINA UNITED NETWORK COMM GRP CO LTD +1

Patent Information

Application Number
CN202511945391.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

The disconnect between the computing power of the intelligent computing center and the power system leads to a single dimension of decision-making, making it impossible to coordinate and optimize electricity costs and green energy consumption. It also lacks the collaborative perception and utilization of multi-dimensional power information such as grid time-of-use pricing, microgrid real-time power generation capacity, and energy storage system status.

Method used

By collecting real-time data on the computing power resource status of the intelligent computing center, the time-of-use electricity price of the external power grid, the status of the microgrid, and the status of the energy storage system, and using a pre-trained large model for data fusion and prediction, a computing-electricity collaborative scheduling strategy is generated to realize the allocation of computing power resources, the adjustment of microgrid output, and the charging and discharging control of the energy storage system.

Benefits of technology

Significantly reduce electricity costs, improve green electricity consumption rate and energy utilization efficiency, achieve multi-objective collaborative optimization, and promote the transformation of intelligent computing centers into a new energy-computing power integrated operation model with proactive collaboration and intelligent response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121364938A_ABST
    Figure CN121364938A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent computing center resource management and control method, system and device based on computing and power collaboration. The method comprises the steps of collecting computing power resource state data of an intelligent computing center, time-of-use electricity price data of an external power grid, state data of a micro-grid and state data of an energy storage system in real time; fusing the computing power resource state data, the time-of-use electricity price data, the microgrid state data and the energy storage system state data to obtain fused data; predicting the fusion data by using a pre-trained large model, and generating a computing power resource demand prediction curve and a microgrid power generation capacity prediction curve in a future set time window; then, by taking the minimization of the total power consumption cost as an optimization target, generating a power calculation cooperative scheduling strategy; and according to the calculation and power cooperative scheduling strategy, performing calculation power resource distribution, microgrid output adjustment and energy storage system charging and discharging control. The method can effectively break the separation barrier of computing power and electric power, and realizes deep cooperation of computing power scheduling and energy management.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computing power infrastructure, in particular to a resource management and control method, system and equipment of a smart computing center based on computing power and electricity cooperation. BACKGROUND

[0002] The rapid development of artificial intelligence, big data simulation and scientific computing drives the intelligent computing center (smart computing center) to become the core infrastructure supporting digital transformation and technological innovation. As a result, the scale of computing power and energy consumption of the smart computing center has increased dramatically, resulting in an increasing proportion of electricity expenditure in the operating cost. Under the background of energy saving and emission reduction, the smart computing center is under pressure to improve energy efficiency, reduce carbon emissions and control operating costs.

[0003] Currently, the management of the smart computing center in the industry generally has the problem of separation of "computing power" and "electricity" systems. The resource management and control system usually runs independently, and its monitoring and scheduling logic mainly focuses on the running state of the computing power resource itself (such as CPU / GPU utilization, memory, network bandwidth, etc.) and the task priority. The core goal is to ensure the performance and stability of the computing power service. The energy management system is relatively isolated, and may only focus on the purchase cost of electricity from the external power grid or the simple access to self-owned new energy (such as photovoltaic and wind power).

[0004] This "computing power-electricity" separation architecture has significant limitations: first, the decision-making dimension is single, and the existing solution lacks the coordinated perception and utilization of multi-dimensional power information such as time-of-use electricity price, real-time power generation capacity of micro-grid, and state of energy storage system. Its operation and alarm response mechanism is mainly designed around equipment stability, and cannot be linked with dynamic electricity price signals, thus missing the opportunity to significantly reduce electricity costs by intelligently scheduling computing power load. Secondly, the optimization goal is one-sided, and the resource scheduling system cannot actively delay non-urgent tasks to reduce electricity costs during peak electricity prices, nor can it actively increase computing power supply to maximize the consumption of green electricity during periods of rich new energy generation. This results in the computing power load curve being disconnected from the electricity price curve and the green electricity generation curve, and even going in the opposite direction. SUMMARY

[0005] The technical problem to be solved by the present application is to overcome the above-mentioned deficiencies of the prior art, and to provide a resource management and control method, system and equipment of a smart computing center based on computing power and electricity cooperation. This method can effectively break down the barriers between computing power and electricity, achieve deep cooperation between computing power scheduling and energy management, and thus transform and upgrade the smart computing center into a value center with fine regulation and control, so as to significantly reduce operating costs, efficiently consume green energy and comprehensively improve energy utilization efficiency.

[0006] In a first aspect, the present application provides a resource management and control method of a smart computing center based on computing power and electricity cooperation, which comprises the following steps:

[0007] collecting, in real time, algorithm resource state data of the intelligent computing center, time-of-use electricity price data of an external power grid, micro-grid state data, and energy storage system state data;

[0008] fusing the algorithm resource state data, the time-of-use electricity price data, the micro-grid state data, and the energy storage system state data to obtain fused data;

[0009] using a pre-trained large model to predict the fused data to generate an algorithm resource demand prediction curve and a micro-grid power generation capability prediction curve within a future set time window;

[0010] based on the algorithm resource demand prediction curve, the micro-grid power generation capability prediction curve, and a preset algorithm task priority, generating an algorithm-power collaborative scheduling strategy with the optimization objective of minimizing overall electricity cost;

[0011] based on the algorithm-power collaborative scheduling strategy, performing allocation of algorithm resources, adjustment of micro-grid output, and charge-discharge control of the energy storage system to achieve algorithm-power collaborative resource management and control of the intelligent computing center.

[0012] Further, the algorithm resource state data is collected through an adaptive integration module; the adaptive integration module has a plurality of brand algorithm card driver libraries and software development kits built-in, and through an interface abstraction layer, performance indicators of different algorithm cards are uniformly mapped into a standardized data model.

[0013] Further, after collecting, in real time, the algorithm resource state data of the intelligent computing center, the method further includes:

[0014] comparing the algorithm resource state data with a preset threshold;

[0015] when the algorithm resource state data exceeds the preset threshold, generating and pushing an alarm information;

[0016] based on the alarm information, automatically matching and executing a corresponding predefined operation and maintenance script to eliminate abnormal algorithm resource state data before data fusion.

[0017] Further, the time-of-use electricity price data includes electricity price data corresponding to a peak time period, a high peak time period, a flat rate time period, and a low valley time period.

[0018] generating the algorithm-power collaborative scheduling strategy, specifically including:

[0019] in the peak time period, generating a first strategy, the first strategy including: controlling the micro-grid and the energy storage system to supply power at maximum power, and suspending or delaying execution of an algorithm task with a priority lower than a first preset threshold;

[0020] In the peak period, a second strategy is generated, the second strategy including: controlling the micro-grid and the energy storage system to participate in power supply on demand, and suspending or delaying execution of the computing power task with a priority lower than a second preset threshold;

[0021] In the flat price period, a third strategy is generated, the third strategy including: controlling the micro-grid power to charge the energy storage system, so that the energy storage system participates in power supply after reaching a full charge state, while guaranteeing execution of the computing power task with a priority higher than a third preset threshold;

[0022] In the valley period, a fourth strategy is generated, the fourth strategy including: controlling the micro-grid power to charge the energy storage system, and executing all computing power tasks in the task queue.

[0023] Further, while the computing and power collaborative scheduling strategy is generated, the method further includes executing daily demand peak intervention;

[0024] The daily demand peak intervention is executed, specifically including:

[0025] A historical monthly peak power HP is obtained and a threshold X is set;

[0026] The real-time active power peak NP of the current external power grid is monitored;

[0027] When NP≥(HP-X), an immediate discharge instruction is generated to schedule the micro-grid and the energy storage system to discharge;

[0028] When NP<(HP-X), the daily demand peak intervention is controlled not to be triggered or stopped.

[0029] Further, the computing power task priority is divided into four levels, which are:

[0030] The first computing power task priority P0 representing real-time reasoning and key scientific research tasks, the second computing power task priority P1 representing model training tasks, the third computing power task priority P2 representing data preprocessing and model optimization tasks, and the fourth computing power task priority P3 representing batch testing and non-urgent tasks;

[0031] Wherein,

[0032] The order of priority is: the first computing power task priority P0> the second computing power task priority P1> the third computing power task priority P2> the fourth computing power task priority P3;

[0033] The first computing power task priority P0, the second computing power task priority P1, the third computing power task priority P2, and the fourth computing power task priority P3 are all bound to the user's service level agreement SLA as a scheduling decision basis.

[0034] Further, according to the computing and power collaborative scheduling strategy, the allocation of computing power resources is executed, specifically including:

[0035] According to the algorithm and power collaborative scheduling strategy, a set of to-be-allocated computing power equipment participating in scheduling is determined;

[0036] According to the intelligent inspection score of each device in the set of to-be-allocated computing power equipment, the computing power task is prioritized; wherein the device with a higher intelligent inspection score is preferentially allocated with the task;

[0037] An unavailable device with an intelligent inspection score lower than a preset risk threshold in the set of to-be-allocated computing power equipment is identified, and a resource isolation or task migration operation is performed on the identified unavailable device.

[0038] Further, the intelligent inspection score is obtained by the following steps:

[0039] An initial health score is set for each computing power device;

[0040] Based on the preset inspection items, the computing power device is periodically or triggeringly detected;

[0041] When the detection result of any inspection item is abnormal, the health score of the computing power device is deducted according to the preset deduction rule and weight of the inspection item;

[0042] The updated health score after deduction is taken as the intelligent inspection score of the computing power device.

[0043] In a second aspect, the present application provides a resource management and control system of a smart computing center based on algorithm and power collaboration, which comprises:

[0044] A data acquisition module is configured to acquire real-time computing power resource state data of the smart computing center, time-of-use electricity price data of an external power grid, micro-grid state data, and energy storage system state data;

[0045] A data fusion module is connected with the data acquisition module and is configured to fuse the computing power resource state data, the time-of-use electricity price data, the micro-grid state data, and the energy storage system state data to obtain fusion data;

[0046] A prediction module is connected with the data fusion module and is configured to use a pre-trained large model to predict the fusion data to generate a computing power resource demand prediction curve and a micro-grid power generation capacity prediction curve in a future set time window;

[0047] A strategy generation module is connected with the prediction module and is configured to generate an algorithm and power collaborative scheduling strategy based on the computing power resource demand prediction curve, the micro-grid power generation capacity prediction curve, and a preset computing power task priority, with the minimization of overall electricity cost as an optimization objective;

[0048] The strategy execution module is connected with the strategy generation module, and is configured to execute allocation of the computing power resource, adjustment of the micro-grid output, and charge-discharge control of the energy storage system according to the computing power-electricity collaborative scheduling strategy, so as to realize the resource management and control of the intelligent computing center based on the computing power-electricity collaboration.

[0049] In a third aspect, the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes the resource management and control method of the intelligent computing center based on the computing power-electricity collaboration according to the first aspect.

[0050] The present application constructs a collaborative management mechanism of deep integration of computing power and electricity, collects and fuses multi-source heterogeneous data such as the computing power resource state of the intelligent computing center, the time-of-use electricity price of the external power grid, the micro-grid power generation capacity, and the energy storage system state in real time, and jointly predicts the future computing power demand and green electricity supply based on a pre-trained large model, and then generates an integrated collaborative strategy covering computing power scheduling, micro-grid output adjustment, and energy storage charge-discharge control with the goal of minimizing the overall electricity cost, which can effectively break the information island and decision fragmentation problem under the traditional "separation of computing power and electricity" architecture.

[0051] Specific beneficial effects are:

[0052] (1) Significantly reduce electricity costs: by actively suspending or delaying low-priority computing power tasks during peak / peak electricity price periods, and calling on the micro-grid and energy storage system to jointly supply power, high-priced electricity purchases are avoided; at the same time, in the low valley or flat price period, low-cost electricity is fully utilized to execute tasks and charge the energy storage, realizing the global optimization of electricity expenses.

[0053] (2) Improve green electricity consumption rate and energy utilization efficiency: based on the accurate prediction of the micro-grid power generation capacity, dynamically increase the computing power load during the new energy rich period, so that the computing power operation curve actively matches the green electricity output curve, reduces the curtailment of wind and light, improves the proportion of local consumption of renewable energy, and helps to achieve the goal of energy saving and emission reduction.

[0054] (3) Realize multi-objective collaborative optimization: the present application considers economy, greenness and stability, and promotes the transformation of the intelligent computing center from "passive energy consumption" to "active collaboration, intelligent response" new energy-computing power integrated operation mode.

[0055] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0056] The accompanying drawings are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and are meant to explain the application by not limiting the application. The above and other features and advantages of the present application will become more apparent from the detailed description in conjunction with the accompanying drawings, in which:

[0057] Figure 1 A schematic diagram of the resource management and control method of the intelligent computing center based on algorithm and electricity collaboration provided for the embodiments of the present application is shown in the figure.

[0058] Figure 2 A flowchart of the resource management and control method of the intelligent computing center based on algorithm and electricity collaboration provided for the embodiments of the present application is shown in the figure.

[0059] Figure 3 A schematic diagram of the resource management and control system of the intelligent computing center based on algorithm and electricity collaboration provided for the embodiments of the present application is shown in the figure.

[0060] Figure 4 A block diagram of an electronic device provided for the embodiments of the present application is shown in the figure.

[0061] Reference signs: 10, data acquisition module, 20, data fusion module, 30, prediction module, 40, strategy generation module, 50, strategy execution module, 100, processor, 200, memory. DETAILED DESCRIPTION

[0062] It can be understood that the specific embodiments and the drawings described herein are only used to explain the present application, but not to limit the present application.

[0063] It can be understood that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0064] It can be understood that, for the convenience of description, only the parts related to the present application are shown in the drawings of the present application, and the parts unrelated to the present application are not shown in the drawings.

[0065] It can be understood that each unit and module involved in the embodiments of the present application can only correspond to one entity structure, or can be composed of multiple entity structures, or multiple units and modules can be integrated into one entity structure.

[0066] It can be understood that the functions and steps marked in the flowchart and block diagram of the present application can occur in an order different from that marked in the drawings without conflict.

[0067] It can be understood that the flowcharts and block diagrams of the present application show the possible implementation architecture, function and operation of the system, device, equipment and method according to the embodiments of the present application. Each block in the flowchart or block diagram can represent a unit, module, program segment, code, which contains executable instructions for implementing the specified function. Moreover, each block or combination of blocks in the block diagram and flowchart can be implemented by a hardware-based system for implementing the specified function, or by a combination of hardware and computer instructions.

[0068] It can be understood that the units and modules involved in the embodiments of the present application can be implemented by software or hardware, for example, the units and modules can be located in a processor.

[0069] Embodiment 1:

[0070] Under the dual driving of energy saving and emission reduction targets and the surge in computing power demand, intelligent computing centers are facing challenges such as high electricity cost, extensive energy structure, and fragmented computing and power systems: traditional resource scheduling only focuses on computing performance, ignoring price fluctuations and local green power supply, while energy management lacks awareness of computing load characteristics, resulting in high-peak high-price electricity consumption, insufficient green power consumption, and prominent demand charge risk.

[0071] Therefore, the present embodiment provides an intelligent computing center resource management and control method based on computing and power collaboration. The method is suitable for new green intelligent computing centers equipped with microgrids and energy storage systems. Through computing and power data fusion and intelligent prediction collaborative scheduling of computing tasks and energy resources, the service level agreement (SLA) of key business services is guaranteed, and the minimization of electricity cost and the efficient use of renewable energy are realized. The intelligent computing center resource management and control method based on computing and power collaboration specifically includes steps K1-K5, as shown in Figure 1

[0072] Step K1: Real-time collection of computing resource state data of the intelligent computing center, time-of-use electricity price data of the external power grid, microgrid state data and energy storage system state data.

[0073] As a specific implementation, the computing resource state data is collected by an adaptive integration module. The adaptive integration module has a built-in driver library and software development kit for multiple brands of computing cards, and through an interface abstraction layer, the performance indicators of different computing cards are uniformly mapped into a standardized data model.

[0074] As a specific implementation, after real-time collection of the computing resource state data of the intelligent computing center, the method further includes:

[0075] Comparing the computing resource state data with a preset threshold;

[0076] ​When the computing resource state data exceeds a preset threshold, an alarm information is generated and pushed;

[0077] According to the alarm information, a corresponding predefined operation and maintenance script is automatically matched and executed to eliminate abnormal computing resource state data before data fusion.

[0078] The time-of-use electricity price data includes electricity price data corresponding to peak time, high peak time, flat price time, and low valley time.

[0079] Step K2: The computing resource state data, time-of-use electricity price data, micro-grid state data, and energy storage system state data are fused to obtain fused data.

[0080] Step K3: The fused data is predicted using a pre-trained large model to generate a computing resource demand prediction curve and a micro-grid power generation capacity prediction curve within a future set time window.

[0081] Step K4: Based on the computing resource demand prediction curve, the micro-grid power generation capacity prediction curve, and a preset computing task priority, a computing and electricity collaborative scheduling strategy is generated with the optimization objective of minimizing the overall electricity cost; wherein the computing and electricity collaborative scheduling strategy includes collaborative control instructions for computing resources, micro-grids, and energy storage systems.

[0082] As a specific implementation, while generating the computing and electricity collaborative scheduling strategy, the method further includes performing daily demand peak intervention;

[0083] Performing daily demand peak intervention specifically includes:

[0084] Obtaining a historical monthly electricity peak HP and setting a critical value X;

[0085] Real-time monitoring of the active power peak NP of the current external power grid;

[0086] When NP≥(HP-X), an immediate discharge instruction is generated to schedule the micro-grid and energy storage system to discharge;

[0087] When NP<(HP-X), the daily demand peak intervention is not triggered or stopped.

[0088] Step K5: According to the computing and electricity collaborative scheduling strategy, the allocation of computing resources, the adjustment of micro-grid output, and the charge and discharge control of energy storage systems are executed to realize intelligent computing center resource management and control based on computing and electricity collaboration.

[0089] Generating the computing and electricity collaborative scheduling strategy specifically includes:

[0090] In the peak period, a first strategy is generated, the first strategy comprising: controlling the micro-grid and the energy storage system to supply power at maximum power, and suspending or delaying execution of computing power tasks with a priority lower than a first preset threshold;

[0091] In the peak period, a second strategy is generated, the second strategy comprising: controlling the micro-grid and the energy storage system to participate in power supply on demand, and suspending or delaying execution of computing power tasks with a priority lower than a second preset threshold;

[0092] In the flat price period, a third strategy is generated, the third strategy comprising: controlling the micro-grid power to charge the energy storage system, so that the energy storage system participates in power supply after reaching a full charge state, while ensuring execution of computing power tasks with a priority higher than a third preset threshold;

[0093] In the valley period, a fourth strategy is generated, the fourth strategy comprising: controlling the micro-grid power to charge the energy storage system, and executing all computing power tasks in the task queue.

[0094] As a specific implementation, the computing power task priority is divided into four levels, which are:

[0095] The first computing power task priority P0 represents real-time reasoning and key scientific research tasks, the second computing power task priority P1 represents model training tasks, the third computing power task priority P2 represents data preprocessing and model optimization tasks, and the fourth computing power task priority P3 represents batch testing and non-urgent tasks.

[0096] Wherein,

[0097] The order of priority is: the first computing power task priority P0 > the second computing power task priority P1 > the third computing power task priority P2 > the fourth computing power task priority P3.

[0098] The first computing power task priority P0, the second computing power task priority P1, the third computing power task priority P2, and the fourth computing power task priority P3 are all bound to the user's service level agreement SLA as a basis for scheduling decisions.

[0099] As a specific implementation, according to the computing and power collaborative scheduling strategy, the allocation of computing power resources is performed, specifically including:

[0100] According to the computing and power collaborative scheduling strategy, a set of to-be-allocated computing power devices participating in scheduling is determined;

[0101] According to the intelligent inspection score of each device in the set of to-be-allocated computing power devices, the computing power tasks are prioritized; wherein, the devices with higher intelligent inspection scores are given priority in task allocation;

[0102] Identify the unavailable devices in the to-be-allocated computing power device set whose intelligent inspection scores are lower than a preset risk threshold, and perform resource isolation or task migration operation on the identified unavailable devices.

[0103] As a specific implementation, the intelligent inspection score is obtained through the following steps:

[0104] Set an initial health score for each computing power device.

[0105] Periodically or triggerably detect the computing power device based on a preset inspection item;

[0106] When the detection result of any inspection item is abnormal, deduct the health score of the computing power device according to the preset deduction rule and weight of the inspection item;

[0107] Take the updated health score after deduction as the intelligent inspection score of the computing power device.

[0108] Based on the foregoing method framework, the embodiment further provides a specific implementation example. As shown in the figure, the implementation example includes: Figure 2

[0109] Step S1: Set the monitoring object attribute and classify according to the server, computing power card brand and type, network device, air conditioning device, power system, etc.

[0110] Specifically, this step can be implemented by a monitoring object attribute module to encode and classify the monitoring object, and set the attribute value as AN, wherein: when AN=01, the monitoring object attribute is a server; when AN=02, the monitoring object attribute is a computing power card brand and type; when AN=03, the monitoring object attribute is a network device; when AN=04, the monitoring object attribute is an air conditioning device; and when AN=05, the monitoring object attribute is a power system.

[0111] Step S2: Determine the registration object data items to be collected according to the monitoring object attribute, and set threshold values for different indicators in the data items.

[0112] Specifically, this step can be implemented by a monitoring object registration module to determine the registration object data items to be collected according to the AN value of the monitoring object attribute, and set threshold values for different indicators in the data items, with the threshold value being identified as M.

[0113] Step S3: Adapt the interfaces of each object, and encapsulate the interfaces into standard interfaces to realize the automatic access of newly registered monitoring objects through the registration interface information.

[0114] ​Specifically, this step can be implemented by an adaptation integration module, which establishes adaptation integration of different object interfaces according to the monitoring object attributes and registration information, including computing power adaptation, network adaptation, storage adaptation, security adaptation, etc., to realize the interface ecology of index collection of heterogeneous computing power. The module has a built-in compatible unit of heterogeneous computing power, which contains the driver library and SDK of various brand computing power cards (such as NVIDIA, AMD, Ascend, etc.), and through the interface abstraction layer, the performance indicators of different computing power cards are uniformly mapped to a standardized data model. At the same time, using the interface registration unit, the monitoring objects are registered as standard interfaces according to the interface specification, allowing operation and maintenance personnel to access new models of computing power cards through configuration, realizing plug-and-play and automatic discovery of devices.

[0115] Step S4: Automatically scheduling tasks according to task frequency and time to realize automatic collection of monitoring information.

[0116] Specifically, this step can be implemented by a monitoring information collection module, which uses the adaptation integration module to automatically schedule tasks according to the preset collection frequency and time through the internal task scheduling unit, to realize automatic collection of monitoring information of GPU computing power resources, server resources, storage resources, network resources, security resources, and dynamic environment resources.

[0117] Step S5: Comparing the collected data value with the preset threshold value, and pushing the alarm information in the comparison.

[0118] Specifically, this step can be implemented by a monitoring threshold alarm module, which uses the internal threshold comparison unit to compare the collected data value x with the preset threshold value M in real time: when X < M, it is determined not to alarm; when X ≥ M, it is determined to trigger an alarm, and the threshold alarm unit pushes the alarm information according to the preset pushing rules.

[0119] Step S6: According to the alarm information, match the related preplan to automatically execute the operation and maintenance script, complete the automatic operation and maintenance task. According to the weight scientific operation and maintenance score, complete the normal inspection to reduce the operation and maintenance workload, use the light operation linkage to illuminate the operation and maintenance area, create a "black light intelligent algorithm center", and reduce energy consumption.

[0120] Specifically, this step can be implemented by an operation and maintenance management module, through its internal pre-plan analysis unit, a mapping library of alarm type-pre-plan ID is built, when receiving alarm information, through parsing the alarm code, automatically matching and executing the predefined operation and maintenance script (such as restarting the service, migrating the container, adjusting the fan speed). Through the intelligent inspection unit, set the initial total score T=100 points for each inspection object, set different scores S and weights W for each type of inspection task, and scientifically evaluate the score according to the weight (when the inspection task matches the set condition, T=T-SxW; when it does not match, T=T), which is used as a quantitative indicator of device health. When T is lower than the preset threshold (such as 60 points), the system automatically marks the device as high risk and increases its inspection frequency, and at the same time, the computing task on the device is preferentially migrated to a node with higher health degree by the computing and power coordination module. Normalized inspection is performed by the operation and maintenance robot unit. Through the light and power linkage unit, a light-out center is created to illuminate only the operation and maintenance area, reducing energy consumption. Through the waste heat recovery unit, water circulation water-cooled air conditioners and water circulation indirect evaporation are used to recover waste heat, reducing electricity costs.

[0121] Step S7: Utilize wind and light power generation, and the generated electric energy is transmitted according to the scheduling instruction and stored by using energy storage.

[0122] Specifically, this step can be implemented by a micro-grid module, which utilizes light and wind energy during the day and wind energy at night for power generation. The generated electric energy is transmitted according to the scheduling instruction and stored by using the energy storage system to supply power for the intelligent computing center.

[0123] Step S8: Utilize large models to perform fusion analysis on computing and power data, and automatically push them to the knowledge base in the operation and maintenance management module for continuous improvement.

[0124] Specifically, this step can be implemented by a computing and power fusion module, which uses its internal large model analysis unit to input historical computing load, real-time power data, weather forecast (light, wind speed), and electricity price policy. Through a time series prediction model, the computing power demand curve, micro-grid power generation curve, and optimal power consumption strategy for a future period are predicted. The analysis results (such as suggesting starting batch training tasks at 14:00-15:00) are automatically pushed to the knowledge base of the operation and maintenance management module for continuous improvement, and are used as a decision reference for the computing and power coordination module.

[0125] Step S9: Real-time analysis of computing and power coordination strategy, realizing computing resource scheduling, power resource scheduling, and computing and power coordination scheduling, scheduling micro-grid electric energy according to the computing and power coordination strategy to ensure that the peak power consumption is reduced month by month.

[0126] Specifically, this step can be realized by an algorithm and power coordination module, which generates an algorithm and power coordination strategy in real time based on the information of the threshold alarm unit and the results of the large model analysis unit. First, the computing power task is classified: P0 (real-time inference, key scientific research) > P1 (model training) > P2 (data preprocessing, model optimization) > P3 (batch testing, non-urgent tasks), and this priority can be bound with the user's SLA (service level agreement). Then, through the peak and valley power adjustment unit, the micro-grid and energy storage power supplement are dispatched according to the peak and valley time strategy: when the current time Tn is in the peak period M, the micro-grid and energy storage power supplement are dispatched, and non-P0 priority tasks are suspended; when in the peak period H, the micro-grid and energy storage power supplement are dispatched, and P2 and below priority tasks are suspended; when in the normal period V, the micro-grid is used to charge the energy storage first, and after the energy storage is full, the power supply is supplemented, and P0 and P1 tasks are guaranteed, and the suspended tasks are preferentially run; when in the low valley period L, the micro-grid is dispatched to charge the energy storage, and all computing power tasks are run. At the same time, through the daily demand power adjustment unit, the monthly peak power HP of the previous 3 months is obtained, the critical value X is set, and the current peak power NP is compared with HP-X in real time: when NP≥HP-X, the micro-grid or energy storage power supplement is dispatched to reduce the current peak power; when NP<HP-X, the micro-grid power is input to the energy storage, and if the energy storage is full, the smart computing center power grid is connected to supplement.

[0127] Step S10: Realize the visual management of the smart computing center assets such as visual inventory, analysis, alarm, operation and maintenance, and disposal according to the algorithm and power fusion data.

[0128] Specifically, this step can be realized by a visual analysis module, which integrates various visual components using its internal visual layout unit to realize out-of-box large screen display; uses the asset map unit to display asset information, indicators, and states on the map at a glance, and reasonably allocates monitoring attention points; uses the indicator display unit to intuitively display alarm status and indicator data, and performs year-on-year analysis, distribution trend analysis, and multi-indicator aggregation analysis, thereby realizing the visual inventory, analysis, alarm, operation and maintenance, and disposal of the smart computing center assets.

[0129] The embodiment provides a kind of based on the resource management and control method of algorithm electricity cooperation of wisdom algorithm center, by building the complete technical framework covering monitoring object attribute setting, data acquisition and threshold alarm, automation operation and maintenance, micro-grid and energy storage management, algorithm electricity fusion analysis and collaborative scheduling, realizes the integrated intelligent monitoring and optimization scheduling of heterogeneous computing resources and green energy.The method utilizes large model to carry out multi-source data fusion and prediction, in combination with time-of-use electricity price strategy and task priority mechanism, dynamically generates algorithm electricity collaborative scheduling scheme, so as to effectively reduce the electricity cost of wisdom algorithm center, improve green power consumption efficiency under the premise of guaranteeing key business SLA, and realize operation and maintenance intelligentization and energy fine management through equipment health score and visual monitoring, help to create efficient, low-carbon "black light wisdom algorithm center".

[0130] Embodiment 2:

[0131] As Figure 3 shown, the embodiment provides a kind of based on the resource management and control system of algorithm electricity cooperation of wisdom algorithm center, and the system includes:

[0132] Data acquisition module 10 is used for real-time acquisition of the computing resource state data of wisdom algorithm center, the time-of-use electricity price data of external power grid, micro-grid state data and energy storage system state data;

[0133] Data fusion module 20 is connected with data acquisition module 10, for fusing computing resource state data, time-of-use electricity price data, micro-grid state data and energy storage system state data, to obtain fusion data;

[0134] Prediction module 30 is connected with data fusion module 20, for predicting fusion data using pre-trained large model, to generate computing resource demand prediction curve, micro-grid power generation capacity prediction curve in future set time window;

[0135] Strategy generation module 40 is connected with prediction module 30, for generating algorithm electricity collaborative scheduling strategy based on computing resource demand prediction curve, micro-grid power generation capacity prediction curve and preset computing task priority, with the optimization target of minimizing overall electricity cost;Wherein, algorithm electricity collaborative scheduling strategy includes collaborative control instruction to computing resource, micro-grid and energy storage system;

[0136] Strategy execution module 50 is connected with strategy generation module 40, for executing the allocation of computing resource, the adjustment of micro-grid output and the charge-discharge control of energy storage system according to algorithm electricity collaborative scheduling strategy, to realize the resource management and control of wisdom algorithm center based on algorithm electricity cooperation.

[0137] The system in the embodiment can execute the method in embodiment 1.

[0138] Embodiment 3:

[0139] As Figure 4 shown, the embodiment provides an electronic device, which comprises a memory 200 and a processor 100, the memory 200 stores a computer program, when the processor 100 runs the computer program stored in the memory 200, the processor 100 executes the intelligent computing center resource management and control method based on algorithm-electricity cooperation according to the embodiment 1.

[0140] It can be understood that the above implementation is only an exemplary implementation adopted for illustrating the principles of the present application, and the present application is not limited thereto. Various modifications and improvements can be made by those of ordinary skill in the art without departing from the spirit and essence of the present application, and these modifications and improvements are also considered to be within the scope of protection of the present application.

Claims

1. A resource management and control method for a smart computing center based on algorithm-electricity collaboration, characterized in that, The method comprises the following steps: Real-time collection of computing resource state data of the intelligent computing center, time-of-use electricity price data of an external power grid, micro-grid state data, and energy storage system state data; Fusion of the computing resource state data, the time-of-use electricity price data, the micro-grid state data, and the energy storage system state data to obtain fused data; Prediction of the fused data by using a pre-trained large model to generate a computing resource demand prediction curve and a micro-grid power generation capability prediction curve within a future set time window; Generation of an algorithm and electricity collaborative scheduling strategy based on the computing resource demand prediction curve, the micro-grid power generation capability prediction curve, and a preset computing task priority, with minimization of overall electricity cost as an optimization objective; According to the algorithm and electricity collaborative scheduling strategy, allocation of computing resources, adjustment of micro-grid output, and charging and discharging control of the energy storage system are performed to realize intelligent computing center resource management and control based on algorithm and electricity collaboration.

2. The intelligent computing center resource management and control method based on algorithm and electricity collaboration according to claim 1, wherein the computing resource state data is collected by an adaptive integration module; wherein the adaptive integration module is internally provided with a driver library and a software development kit of multiple brands of computing cards, and the performance indicators of different computing cards are uniformly mapped to a standardized data model through an interface abstraction layer.

3. The intelligent computing center resource management and control method based on algorithm and electricity collaboration according to claim 1, wherein after real-time collection of the computing resource state data of the intelligent computing center, the method further comprises: Comparing the computing resource state data with a preset threshold value; When the computing resource state data exceeds the preset threshold value, generating and pushing an alarm information; According to the alarm information, automatically matching and executing a corresponding predefined operation and maintenance script to eliminate abnormal computing resource state data before data fusion.

4. The intelligent computing center resource management and control method based on algorithm and electricity collaboration according to claim 1, wherein the time-of-use electricity price data comprises electricity price data corresponding to a peak time period, a high peak time period, a flat rate time period, and a low valley time period; The generation of the algorithm and electricity collaborative scheduling strategy specifically comprises: In the peak time period, a first strategy is generated, which comprises controlling the micro-grid and the energy storage system to supply power at maximum power and suspending or delaying the execution of computing tasks with a priority lower than a first preset threshold value; In the high peak time period, a second strategy is generated, which comprises controlling the micro-grid and the energy storage system to participate in power supply on demand and suspending or delaying the execution of computing tasks with a priority lower than a second preset threshold value; In the flat rate time period, a third strategy is generated, which comprises controlling the micro-grid power to charge the energy storage system to enable the energy storage system to participate in power supply after reaching a full charge state, while ensuring the execution of computing tasks with a priority higher than a third preset threshold value; In the low valley time period, a fourth strategy is generated, which comprises controlling the micro-grid power to charge the energy storage system and executing all computing tasks in the task queue.

5. The intelligent computing center resource management and control method based on algorithm and electricity collaboration according to claim 1, wherein ​ ​ ​ The method further comprises executing daily demand peak intervention while generating the algorithm-electricity collaborative scheduling strategy; The execution of the daily demand peak intervention specifically comprises: Obtaining a historical monthly electricity consumption peak HP and setting a critical value X; Real-time monitoring of the active power peak NP of the current external power grid; When NP≥(HP-X), generating an immediate discharge instruction to schedule the micro-grid and the energy storage system to discharge; When NP<(HP-X), controlling not to trigger the daily demand peak intervention or stopping the daily demand peak intervention.

6. The algorithm-electricity collaborative-based intelligent computing center resource management method according to claim 1, wherein the algorithm task priority is divided into four levels, which are: a first algorithm task priority P0 representing real-time reasoning and key scientific research tasks, a second algorithm task priority P1 representing model training tasks, a third algorithm task priority P2 representing data preprocessing and model optimization tasks, and a fourth algorithm task priority P3 representing batch testing and non-urgent tasks; wherein the priority order is: first algorithm task priority P0> second algorithm task priority P1> third algorithm task priority P2> fourth algorithm task priority P3; The first algorithm task priority P0, the second algorithm task priority P1, the third algorithm task priority P2, and the fourth algorithm task priority P3 are all bound to the service level agreement SLA of the user as a scheduling decision basis.

7. The algorithm-electricity collaborative-based intelligent computing center resource management method according to any one of claims 1 to 6, wherein according to the algorithm-electricity collaborative scheduling strategy, the allocation of algorithm resources specifically comprises: According to the algorithm-electricity collaborative scheduling strategy, determining a set of to-be-allocated algorithm equipment participating in scheduling; According to the intelligent inspection score of each device in the set of to-be-allocated algorithm equipment, the algorithm task is prioritized; wherein the device with higher intelligent inspection score is given priority to task allocation; Identify the unusable devices in the set of to-be-allocated algorithm equipment whose intelligent inspection score is lower than the preset risk threshold, and perform resource isolation or task migration operation on the identified unusable devices.

8. The algorithm-electricity collaborative-based intelligent computing center resource management method according to claim 7, wherein the intelligent inspection score is obtained by the following steps: Setting an initial health score for each algorithm device; Periodically or trigger-based detection of the algorithm device based on the preset inspection items; When the detection result of any inspection item is abnormal, the health score of the algorithm device is deducted according to the preset deduction rule and weight of the inspection item; The updated health score after deduction is taken as the intelligent inspection score of the algorithm device. The system comprises: a data acquisition module for real-time acquisition of algorithm resource state data of an intelligent computing center, time-of-use electricity price data of an external power grid, micro-grid state data, and energy storage system state data; a data fusion module connected with the data acquisition module for fusing the algorithm resource state data, the time-of-use electricity price data, the micro-grid state data, and the energy storage system state data to obtain fusion data; ​ 9. A resource management and control system for a smart computing center based on algorithm-electricity collaboration, characterized in that, ​ ​ ​ A prediction module connected with the data fusion module, configured to utilize a pre-trained large model to predict the fusion data, and generate a computing resource demand prediction curve and a micro-grid power generation capability prediction curve within a future setting time window; A strategy generation module connected with the prediction module, configured to generate an algorithm and electricity collaborative scheduling strategy based on the computing resource demand prediction curve, the micro-grid power generation capability prediction curve, and a preset computing task priority, with the optimization objective of minimizing the overall electricity cost; A strategy execution module connected with the strategy generation module, configured to perform allocation of computing resources, adjustment of micro-grid output, and charge and discharge control of the energy storage system according to the algorithm and electricity collaborative scheduling strategy, so as to realize algorithm and electricity collaborative resource management and control of the intelligent computing center.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes the algorithm and electricity collaborative resource management and control method of the intelligent computing center according to any one of claims 1-8.

Citation Information

Patent Citations

  • Green data center computing power demand prediction and energy consumption control method, system and device

    CN120508401A

  • Calculation and power collaboration system and method

    CN120691373A

  • Intelligent computing power availability analysis method, computer device, medium and product

    CN120951599A

  • Computing power task scheduling method and device based on new energy consumption and storage medium

    CN121055283A

  • Adaptive resource scaling system for multi-cloud data pipelines based on workflow latency

    DE202025103772U1

Cited By

  • A green electricity adaptive operation method and system of a containerized mobile computing power node

    CN122247985A