Cloud platform-based data center operation monitoring system and method
By acquiring equipment and power system data for adaptation analysis and compensation correction, the problem of power fluctuation impact in traditional computing power allocation methods has been solved, achieving efficient and stable task allocation and equipment selection, and improving system operating efficiency and task success rate.
Patent Information
- Application Number
- CN202511536735.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2045-10-27
AI Technical Summary
Traditional computing power allocation methods fail to effectively address the impact of power fluctuations on equipment operation, leading to decreased equipment performance and task execution failures. Furthermore, the lack of a scientific and reasonable assessment of equipment and task matching increases the need for manual intervention and system instability.
By acquiring equipment operation status, scenario data, and power system data, we can analyze the equipment operation adaptability and correct anomalies. Combined with power fluctuation impact analysis, we can dynamically adjust equipment matching assessment and task allocation to optimize task equipment selection.
It achieves efficient and low-latency task allocation, reduces resource waste, improves system stability and task success rate, reduces equipment failure risk, and provides accurate equipment selection basis.
Smart Images

Figure CN121036355B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of cloud platforms, and in particular to a data center operation monitoring system and method based on a cloud platform. BACKGROUND
[0002] In today's various computing systems, especially in complex computing scenarios involving power supply, such as power big data analysis, intelligent power grid dispatching control, and other tasks, the rationality of computing power allocation directly affects the system's operation efficiency, stability, and task execution success rate. Traditional computing power allocation methods often only consider static performance indicators of devices, such as the number of processor cores, memory size, etc., while ignoring the degree of adaptation between tasks and devices and the abnormal situations that may occur during actual operation of devices. This static evaluation method cannot adapt to the impact of power fluctuations on device operating status, as power fluctuations can cause device performance to decline, run abnormally, or even fail, thereby affecting task execution effectiveness. For example, during periods of unstable power supply, devices may experience slower running speeds, data processing errors, and other issues, but traditional methods cannot effectively evaluate and adjust these situations in a timely manner, and still follow fixed allocation strategies for task allocation, which can lead to task execution failure or low efficiency.
[0003] In addition, existing technologies typically only focus on device status at a single time point when evaluating the matching degree of devices and tasks, without fully considering the dynamic relationship between devices and tasks during task execution and the continuous impact of power operation on devices. This results in a large deviation between evaluation results and actual running scenarios, making it difficult to provide accurate and comprehensive basis for task device selection. In actual applications, due to the lack of scientific and reasonable evaluation and allocation mechanisms, a large amount of manual intervention is often required to adjust task allocation, which not only increases labor costs but also easily leads to human errors, reducing system operation efficiency and reliability. Therefore, there is an urgent need for an adaptive computing power allocation method for power fluctuations that can comprehensively consider the degree of adaptation between devices and tasks, compensation and correction of device operating abnormalities, and dynamic impact of task execution and power operation on devices to achieve efficient and low-latency task allocation and improve system operation stability and task execution success rate. SUMMARY
[0004] To overcome the defects and deficiencies of the prior art, the present application provides a data center operation monitoring system and method based on a cloud platform.
[0005] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:
[0006] In a first aspect, the present application provides a data center operation monitoring method based on a cloud platform, comprising the following steps:
[0007] Step S1, obtaining corresponding running conditions of each device of the data center and corresponding scene data, and obtaining power system data, and transmitting to a cloud platform;
[0008] Step S2, analyzing a scene device running adaptation degree through the corresponding device conditions and the corresponding scene data;
[0009] Step S3, compensating and correcting a device running exception through the corresponding running conditions of each device and power transmission data;
[0010] Step S4, performing device matching evaluation based on a compensation correction result of the device running exception and an analysis result of the device running adaptation degree;
[0011] Step S5, selecting a task device through a device running evaluation result.
[0012] In an implementation manner of the present application, the corresponding running conditions of each device of the data center include running data such as current, voltage and temperature of the device reflecting the device running conditions, and operation capability condition data such as the computing power and CPU occupation condition of the device reflecting the device computing power condition, the corresponding scene data includes task type condition data, task data amount condition and task time limit condition corresponding to the task to be processed, and the power system data includes data such as power system fluctuation condition and power system power running condition reflecting power supply quality, wherein the power system fluctuation condition is the output fluctuation of the output voltage of the power supply device, the output fluctuation of the voltage will interfere with the stability of the clock circuit of the operation device, cause the CPU / GPU operation period to lose synchronization, and cause instruction execution error or calculation delay, so it is necessary to analyze the running interference condition of the voltage fluctuation on the device.
[0013] In an implementation manner of the present application, the adaptation degree analysis in step S2 includes the following specific steps:
[0014] S21, obtaining the computing power of each operation device and the CPU occupation condition of the future task, obtaining the actual standard computing power condition of the device by multiplying the computing power of the device and the corresponding CPU occupation remaining proportion, quantifying the available computing resources of the device by combining the device computing power and the CPU occupation remaining proportion, avoiding over-allocation of tasks to high-load devices, and improving the accuracy of resource evaluation;
[0015] S22, evaluate the device operation exception through the operation data of the reaction device operation of the device, wherein the device operation exception evaluation process is: obtaining the device operation data corresponding to the time period, and obtaining the device operation exception of the corresponding device through the weighted sum of the average value of the standard deviation of the corresponding time period device operation data and the corresponding kind of safety data range; the device influence value is obtained by multiplying the device operation exception evaluation result by the influence coefficient of the device operation on the abnormal algorithm power; the deviation analysis of the dynamic operation data and the safety threshold is introduced, combined with the sensitive coefficient of the device to the abnormal algorithm power, the performance loss of the unreliable device is quantified, and the prediction ability of the fault risk is enhanced;
[0016] S23, the actual algorithm power of the corresponding device is obtained by multiplying the difference value of 1 minus the device influence value with the actual standard algorithm power of the device, the effective algorithm power of the device is dynamically adjusted through the correction value after eliminating the abnormal influence, and the task allocation considers the hardware health state, and the task interruption risk caused by the unstable device is reduced;
[0017] S24, the average calculation data amount of the corresponding task is obtained by dividing the task data amount by the task time limit, the adaptation degree analysis is carried out based on the actual algorithm power of the corresponding device in the future period and the average calculation data amount required by the corresponding task, the task demand is matched with the corrected algorithm power of the device, the adaptation degree is measured through the reciprocal of the standard deviation, the device with sufficient algorithm power and high stability is preferentially selected, and the task scheduling efficiency and reliability are optimized.
[0018] In an implementation manner of the present application, the compensation correction of the device operation exception in the step S3 comprises the following specific steps:
[0019] S31, acquire the output fluctuation of the power system and the power operation of the power system, and acquire the evaluation result of the equipment operation exception of the equipment, perform power fluctuation influence analysis based on the power fluctuation, wherein the power fluctuation influence analysis mode is: acquiring the power transmission in the time period, acquiring the power fluctuation based on the average value of the standard deviation of the voltage in the power transmission from the standard voltage, and obtaining the influence index of the power fluctuation on the equipment by multiplying the power fluctuation by the power fluctuation influence coefficient. Here, the influence of the equipment caused by the power fluctuation when the equipment is performing a task that needs to be processed is analyzed through the future power. This step quantifies the influence of the power fluctuation on the equipment by monitoring the output fluctuation and the running state of the power system, combining the equipment exception evaluation result, collecting the voltage deviation average value in the specified time period, reflecting the power fluctuation intensity, and multiplying the predefined influence coefficient to convert into the equipment influence index, so as to predict the equipment performance decline or risk that may be caused by future power fluctuation. The advantage is that: through the deviation analysis combined with the future power scene, the potential risk of the equipment is identified in advance to support preventive maintenance; avoid the limitation of fixed threshold judgment, more in line with the actual fluctuation characteristics, and improve the analysis reliability;
[0020] S32, compensate and correct the task completion process of the equipment operation exception of the equipment through the influence index of the power fluctuation on the equipment and the task data processing speed, and the compensation and correction process is: the compensation and correction coefficient is obtained by weighting and summing the influence index of the power fluctuation on the equipment and the task data processing speed after standardization, and the compensation and correction result of the equipment operation exception of the task completion process is obtained by multiplying the compensation and correction coefficient and the equipment operation exception of the equipment.
[0021] In an implementation manner of the present application, the equipment matching evaluation in the step S4 comprises the following specific contents:
[0022] S41, acquire the adaptation degree analysis result of each equipment and the task that needs to be processed, and the compensation and correction result of the corresponding equipment operation exception;
[0023] S42, weighting and summing the adaptation degree analysis result of the task that needs to be processed at each time point and the compensation and correction result of the corresponding equipment operation exception, and integrating the result in the time length and dividing by the corresponding time length to obtain the matching evaluation result of the equipment and the task that needs to be processed, that is, the matching of the task is performed through the adaptation of the equipment in the task process. Because the task and the power operation will affect the operation of the equipment, the influence of the task on the equipment and the original matching of the equipment are weighted and summed to obtain the final matching evaluation result of the equipment and the task in the process of the equipment operation;
[0024] S43, obtain the matching evaluation result of each device and the task to be processed.
[0025] In an implementation form of the present application, the selection of the task device by the device running evaluation result in step S5 comprises the following specific contents:
[0026] The obtained matching evaluation result of each device and the task to be processed is compared with the set matching threshold value respectively, the device corresponding to the matching evaluation result with the minimum absolute value of the difference from the matching threshold value is set as the task device, the task is sent to the task device for calculation and processing, and the device with the closest matching evaluation result to the threshold value is selected to execute the task, so that the absolute value comparison ensures the selection of the device most meeting the requirements, the success rate of the task is improved, manual intervention is reduced, and efficient and low-delay task allocation is realized.
[0027] In a second aspect, the present application further provides a cloud platform-based data center running monitoring system, comprising:
[0028] A data acquisition module acquires corresponding running conditions of each device of the data center and corresponding scene data, and acquires power system data at the same time, and transmits the data to the cloud platform.
[0029] An adaptation degree analysis module analyzes the scene device running adaptation degree based on the corresponding device conditions and the corresponding scene data.
[0030] A compensation correction module compensates and corrects the device running exception based on the corresponding running conditions of each device and the power transmission data.
[0031] A device matching evaluation module performs device matching evaluation based on the compensation correction result of the device running exception and the analysis result of the device running adaptation degree.
[0032] A device selection module selects the task device based on the device running evaluation result.
[0033] In a third aspect, the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores a computer program that can be called by the processor, and the processor executes the cloud platform-based data center running monitoring method by calling the computer program stored in the memory.
[0034] In a fourth aspect, the present application provides a computer readable storage medium storing instructions, when the instructions are run on a computer, the computer executes the cloud platform-based data center running monitoring method.
[0035] Compared with the prior art, the present application has the following advantages and beneficial effects:
[0036] Breaking away from the limitations of traditional methods that only consider static equipment performance indicators, this approach fully considers the compatibility between tasks and equipment, flexibly allocating computing power based on actual conditions. This avoids the impact of power fluctuations on equipment performance and task execution, thereby achieving efficient task allocation, reducing unnecessary resource waste, and significantly improving the overall system efficiency. By considering the dynamic impact of task execution and power operation on equipment, it can promptly assess and adjust equipment operating status, effectively addressing abnormal equipment operation caused by power fluctuations, reducing the risk of task execution failure, ensuring stable system operation, comprehensively evaluating the matching degree between equipment and tasks, and compensating for and correcting equipment operation anomalies. This provides an accurate and comprehensive basis for selecting task equipment, enabling tasks to be executed efficiently on more suitable equipment and greatly improving the success rate of task execution. Attached Figure Description
[0037] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0038] Fig. 1 This is a schematic diagram of the overall process of Embodiment 1 of the method of the present invention;
[0039] Fig. 2 This is a schematic diagram of step S2 of embodiment 1 of the method of the present invention;
[0040] Fig. 3 This is a schematic diagram of the structure of embodiment 2 of the system of the present invention. Detailed Implementation
[0041] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0042] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0043] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.
[0044] Example 1
[0045] like Figs. 1-2 As shown, this embodiment provides a data center operation monitoring method based on a cloud platform, which specifically includes the following steps:
[0046] In step S1, the corresponding running conditions of each device of the data center and the corresponding scene data are acquired, and power system data is acquired and transmitted to the cloud platform.
[0047] In this embodiment, the corresponding running conditions of each device of the data center include running data such as current, voltage and temperature of the device reflecting the running conditions of the device, and operation capability condition data such as computing power and CPU occupation of the device reflecting the computing power condition of the device, the corresponding scene data includes task type condition data, task data volume condition and task time limit condition corresponding to the tasks to be processed, and the power system data includes data such as fluctuation of the power system and power running condition of the power system reflecting power supply quality, wherein the fluctuation of the power system is the output fluctuation of the output voltage of the power supply device, and the power running condition includes the current and voltage condition of the power transmission line. The output fluctuation of the voltage will interfere with the stability of the clock circuit of the computing device, causing the CPU / GPU operation period to be out of synchronization, resulting in instruction execution error or calculation delay, so it is necessary to analyze the running interference of the voltage fluctuation on the device. For example, the running condition data of each device of the data center is acquired by installing high-precision sensors in the device to acquire the running data such as current, voltage and temperature of the device. These sensors monitor the device state in real time and transmit the data to the data acquisition system. The operation capability condition data such as computing power and CPU occupation of the device is collected by means of the monitoring software provided by the device. The monitoring software records and summarizes these data at regular intervals. The acquisition of the corresponding scene data needs to interact with the task initiator to acquire the task type condition data, task data volume condition and task time limit condition corresponding to the tasks to be processed through the interface. For the power system data, the fluctuation of the power system, i.e. the output fluctuation of the output voltage of the power supply device, is monitored in real time by installing a voltage monitor on the power transmission line. The current and voltage condition of the power transmission line in the power running condition is collected by installing intelligent electric meters and current transformers at the key nodes of the power transmission line. These devices transmit the monitored data to the power data management platform through the communication network.
[0048] In step S2, the scene device running adaptation degree is analyzed according to the corresponding device condition and the corresponding scene data.
[0049] In this embodiment, the adaptation degree analysis in step S2 includes the following specific steps:
[0050] S21, obtain the computing power of each computing device and the CPU occupation of future tasks, and obtain the actual standard computing power of the device by multiplying the computing power of the device and the corresponding CPU occupation remaining ratio, wherein the corresponding CPU occupation remaining ratio is 1 minus the CPU occupation ratio, by combining the device computing power and the CPU occupation remaining ratio, the available computing resources of the device are quantified, the high-load device is avoided to be excessively allocated tasks, and the accuracy of resource evaluation is improved;
[0051] S22, evaluate the device running exception through the running data of the reaction device running condition of the device, wherein the device running exception evaluation process is: obtaining the device running data of the corresponding time period, obtaining the device running exception of the corresponding device by weighted summation of the device running data of the corresponding time period and the average value of the standard deviation of the corresponding type safety data range, obtaining the device influence value by multiplying the device running exception evaluation result by the influence coefficient of the device running on the computing power exception, wherein the influence coefficient of the device running on the computing power exception is different for different devices, and the value is: taking value between 0-1, analyzing the ratio of the actual standard computing power of the device to the actual computing power, obtaining the influence coefficient of the device running on the computing power exception of the corresponding device, introducing the deviation analysis of the dynamic running data and the safety threshold, combining the sensitive coefficient of the device on the computing power exception, quantifying the performance loss of unreliable device, and enhancing the prediction ability of fault risk;
[0052] S23, obtain the actual computing power of the corresponding device by multiplying the difference value of 1 minus the device influence value and the actual standard computing power of the device, and dynamically adjust the effective computing power of the device by the correction value after eliminating the abnormal influence, so as to ensure that the hardware health state is considered when the task is allocated, and the risk of task interruption caused by unstable device is reduced;
[0053] S24, obtain the average computing data amount required by the corresponding task by dividing the task data amount by the task time limit, and perform adaptation degree analysis based on the actual computing power of the future period of the corresponding device and the average computing data amount required by the corresponding task, wherein the adaptation degree analysis process is: the reciprocal of the average value of the standard deviation of the actual computing power of the future period and the average computing data amount required by the task; match the task demand with the corrected computing power of the device, measure the adaptation degree by the reciprocal of the standard deviation, preferentially select the device with sufficient computing power and high stability, and optimize the task scheduling efficiency and reliability;
[0054] Step S3, compensate and correct the device running exception through the corresponding running condition of each device and the power transmission data;
[0055] In this embodiment, the compensation and correction of the device running exception in step S3 includes the following specific steps:
[0056] S31, acquire the output fluctuation of the power system and the power operation of the power system, and acquire the evaluation result of the equipment operation exception of the equipment, perform power fluctuation influence analysis based on the power fluctuation, wherein the power fluctuation influence analysis manner is: acquiring the power transmission in the time period, acquiring the power fluctuation based on the average value of the standard deviation of the voltage in the power transmission from the standard voltage, and obtaining the influence index of the equipment caused by the power fluctuation by multiplying the power fluctuation by the power fluctuation influence coefficient. Here, the influence of the equipment caused by the power fluctuation when the equipment is performing a task that needs to be processed is analyzed through the future power. This step quantifies the influence of the power fluctuation on the equipment by monitoring the output fluctuation and the running state of the power system, combining the equipment exception evaluation result, collecting the average value of the voltage deviation in a specified time period, reflecting the power fluctuation intensity, and multiplying a predefined influence coefficient to convert it into an equipment influence index, thereby predicting the equipment performance degradation or risk that may be caused by future power fluctuation. The advantage is that: by combining the deviation analysis with the future power scene, the potential risk of the equipment is identified in advance to support preventive maintenance; the limitation of fixed threshold judgment is avoided, which is more in line with the actual fluctuation characteristics and improves the analysis reliability;
[0057] S32, compensate and correct the task completion process of the equipment operation exception of the equipment based on the influence index of the power fluctuation on the equipment and the task data processing speed, wherein the compensation and correction process is: obtaining a compensation correction coefficient by standardizing and weighting the influence index of the power fluctuation on the equipment and the task data processing speed, and obtaining the compensation correction result of the equipment operation exception in the task completion process by multiplying the compensation correction coefficient by the equipment operation exception and adding the equipment operation exception. The standardization process of the present application is to divide by the standard value of the corresponding parameter. This step dynamically adjusts the equipment exception value based on the power fluctuation influence index and the task processing speed to optimize task execution. The specific method is: standardizing (eliminating dimension) the influence index and the task speed, weighting and summing to generate a correction coefficient, and finally dynamically compensating the equipment exception value. The advantage is that: real-time combination of power fluctuation and task demand, automatic optimization of compensation strategy, enhancement of equipment stability, weighted mechanism balancing power fluctuation and task efficiency, avoidance of compensation deviation caused by single factor, and improvement of overall operation robustness;
[0058] Step S4, perform equipment matching evaluation based on the compensation correction result of the equipment operation exception and the analysis result of the equipment operation adaptation degree;
[0059] In this embodiment, the equipment matching evaluation in step S4 includes the following specific contents:
[0060] S41, acquire the adaptation degree analysis result of each equipment and the task that needs to be processed, and the compensation correction result of the corresponding equipment operation exception;
[0061] S42, the matching evaluation result of each device and the task to be processed is obtained, the weighted sum result of the adaptation degree analysis result of the task to be processed at each time point and the compensation correction result of the equipment operation exception is obtained, and the result obtained by integrating in time length and dividing by the corresponding time length is set as the matching evaluation result of the equipment and the task to be processed, that is, the matching of the task is performed through the adaptation of the task and the equipment in the task process. Because the task is performed and the power is operated, the operation of the equipment is affected, so the influence of the task on the equipment and the original matching of the equipment are weighted and summed to obtain the final matching evaluation result of the equipment and the task;
[0062] S43, the matching evaluation result of each device and the task to be processed is obtained, S41 obtains the adaptation degree analysis result of each device and the task and the compensation correction result of the equipment operation exception, which provides a comprehensive and key data basis for subsequent evaluation, ensures that the evaluation can comprehensively consider the equipment adaptability and the compensation of the operation exception, S42 obtains the matching evaluation result by weighting and summing the adaptation degree analysis result and the compensation correction result at each time point, integrating and dividing by the time length. This way fully considers the influence of task performance and power operation on equipment operation, and the dynamic relationship between task and equipment in the running process is taken into account in the evaluation, so that the evaluation result is more in line with the actual running scene, avoiding the limitations of single time point or fixed factor evaluation. S43 obtains the matching evaluation result of each device and the task, which provides a clear and comprehensive basis for subsequent selection of task equipment, so that the most suitable equipment for executing the task can be accurately selected from a large number of devices, thereby improving the success rate of task execution, reducing manual intervention, realizing efficient and low-delay task allocation, and ensuring the stability and reliability of the whole system operation;
[0063] Step S5, selecting the task equipment through the equipment operation evaluation result;
[0064] In this embodiment, the task equipment is selected through the equipment operation evaluation result in step S5, which includes the following specific contents:
[0065] The matching evaluation result of each device and the task to be processed is obtained, and is compared with the set matching threshold value respectively, the device corresponding to the matching evaluation result with the smallest absolute value of the difference from the matching threshold value is set as the task equipment, and the task is sent to the task equipment for calculation and processing. The device with the closest matching evaluation result to the threshold value is selected to execute the task. The advantage is that the absolute value comparison ensures the selection of the most suitable device, improves the task success rate, reduces manual intervention, and realizes efficient and low-delay task allocation;
[0066] It should be noted that the setting of the weight parameter and the acquisition method of the matching threshold in the present application, the acquisition of the corresponding running conditions of each device of the historical data center and the corresponding scene data, and the acquisition of the historical power system data, the introduction of the historical data into each step of the present application for matching evaluation result analysis, and the acquisition of the distribution result with the highest accuracy of data processing result, the introduction of the obtained distribution result and the analysis result of the matching evaluation result into the matlab fitting software for data fitting to obtain the value meeting the highest distribution result accuracy;
[0067] It should be noted that the present embodiment has the following advantages and benefits: breaking away from the limitations of traditional static performance indicators of devices, fully considering the degree of adaptation of tasks and devices, flexibly allocating computing power according to actual conditions, avoiding the influence of device performance decline caused by power fluctuations on task execution, thereby realizing efficient task allocation, reducing unnecessary resource waste, significantly improving the overall operation efficiency of the system, considering the dynamic influence of task execution and power operation on devices, timely evaluating and adjusting the running state of the devices, effectively dealing with the abnormal operation of the devices caused by power fluctuations, reducing the risk of task execution failure, and ensuring stable operation of the system, comprehensively evaluating the matching degree of devices and tasks, and simultaneously compensating and correcting the abnormal operation of the devices, providing accurate and comprehensive basis for the selection of task devices, enabling tasks to be efficiently executed on more suitable devices, and greatly improving the success rate of task execution.
[0068] Embodiment 2
[0069] As shown in Fig. 3 The present embodiment provides a data center operation monitoring system based on a cloud platform, which is realized based on the data center operation monitoring method based on a cloud platform in embodiment 1, and includes: a data acquisition module that acquires the corresponding running conditions of each device of the data center and the corresponding scene data, and simultaneously acquires power system data and transmits the data to the cloud platform;
[0070] An adaptation degree analysis module that analyzes the scene device running adaptation degree based on the corresponding device conditions and the corresponding scene data;
[0071] A compensation correction module that compensates and corrects the abnormal operation of the devices based on the corresponding running conditions of each device and the power transmission data;
[0072] A device matching evaluation module that performs device matching evaluation based on the compensation correction result of the device abnormal operation and the analysis result of the device running adaptation degree;
[0073] A device selection module that selects task devices based on the device running evaluation result, and the specific steps of each module of the present system embodiment are the same as the specific steps of the method embodiment of embodiment 1, which will not be repeated here.
[0074] Embodiment 3
[0075] The electronic device of the embodiment of the present application comprises a processor and a memory, wherein the memory stores a computer program that can be invoked by the processor, and the processor executes the cloud platform-based data center operation monitoring method by invoking the computer program stored in the memory. It should be noted that all computer programs of the cloud platform-based data center operation monitoring method are implemented using C language.
[0076] Embodiment 4
[0077] The embodiment provides a computer readable storage medium, which stores an erasable computer program.
[0078] When the computer program runs on the computer device, the computer device executes the cloud platform-based data center operation monitoring method described above.
[0079] The above embodiments can be realized by software, hardware, firmware or any combination thereof, in whole or in part. When realized by software, the above embodiments can be realized in the form of a computer program product in whole or in part. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the flow or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another through a wired network or / and wireless network. The computer readable storage medium can be any available medium accessible by the computer or a data storage device such as a server, data center and the like containing one or more available medium sets. The available medium can be a magnetic medium (e.g. floppy disk, hard disk, magnetic tape), an optical medium (e.g. DVD) or a semiconductor medium. The semiconductor medium can be a solid state disk.
[0080] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in the present application can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0081] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0082] In several embodiments provided by the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. The unit division is merely a logical function division. There can be other division manners in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units can be indirect couplings or communication connections through some interfaces, devices or unit intermediaries, and can be electrical, mechanical or in other forms.
[0083] The unit described as a separate component can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0084] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit.
[0085] In the description of the specification, the description referring to the terms "one embodiment", "an example", "a specific example", and the like means that a specific feature, structure, material or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the present application. In the specification, the illustrative representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the described specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0086] The basic principles and main features of the present application and the advantages of the present application have been shown and described above. Those skilled in the art should understand that the present application is not limited by the above-described embodiments, and the above-described embodiments and descriptions in the specification are only to illustrate the principles of the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.
Claims
1. A data center operation monitoring method based on a cloud platform, characterized in that, Includes the following steps: Step S1: Obtain the corresponding operating status and scenario data of each device in the data center, and at the same time obtain the power system data and transmit it to the cloud platform; Step S2: Analyze the degree of scene device operation adaptation based on the corresponding device operation status and corresponding scene data; Step S3: Compensate and correct equipment operation anomalies based on the corresponding operating status of each device and power system data; Step S4: Based on the compensation and correction results of equipment operation anomalies and the analysis results of equipment operation adaptability, conduct equipment matching evaluation; Step S5: Select the task equipment based on the equipment operation evaluation results; The analysis of the degree of adaptation in step S2 includes the following specific steps: The computing power of each computing device and the CPU usage of future tasks are obtained. The actual standard computing power of the device is obtained by multiplying the computing power of the device by the corresponding remaining CPU usage ratio. The corresponding remaining CPU usage ratio is 1 minus the CPU usage ratio. The assessment of equipment operation anomalies is carried out by using the equipment operation data that reflects the equipment's operating status. The assessment process for equipment operation anomalies is as follows: obtain the equipment operation data for the corresponding time period, and obtain the corresponding equipment operation anomaly by weighted summing the average standard deviation of the equipment operation data for the corresponding time period and the corresponding type of safety data range. The actual computing power of the corresponding equipment is obtained by subtracting the equipment's impact value from 1 and multiplying the difference by the actual standard computing power of the equipment. The average amount of computational data required for the corresponding task is obtained by dividing the task data volume by the task time limit. The degree of adaptation is then analyzed based on the actual computing power of the corresponding device in the future cycle and the average amount of computational data required for the corresponding task. The compensation and correction for equipment malfunction in step S3 includes the following specific steps: The system acquires information on power system output fluctuations and power system operation, as well as assessment results of equipment operation anomalies. Based on the power fluctuation data, it performs power fluctuation impact analysis. The power fluctuation impact analysis method is as follows: acquire power transmission data within a time period, acquire power fluctuation data based on the average standard deviation of voltage from standard voltage in the power transmission data, and obtain the power fluctuation impact index by multiplying the power fluctuation data by the power fluctuation impact coefficient. The impact index of power fluctuations on equipment and the task data processing speed are used to compensate for and correct equipment operation anomalies during task completion. The compensation and correction process is as follows: the impact index of power fluctuations on equipment and the task data processing speed are standardized and then weighted and summed to obtain the compensation and correction coefficient. The product of the compensation and correction coefficient and the equipment operation anomaly is added to the equipment operation anomaly to obtain the compensation and correction result of the equipment operation anomaly during task completion.
2. The data center operation monitoring method based on a cloud platform according to claim 1, characterized in that, The device matching evaluation in step S4 includes the following specific contents: Obtain the results of the compatibility analysis between each device and the task to be processed, as well as the compensation and correction results for the corresponding device operation anomalies; The weighted sum of the results of the suitability analysis of the tasks to be processed at each time point and the corresponding compensation and correction results of equipment operation anomalies, and the result obtained by integrating over the time length and dividing by the corresponding duration, is set as the matching evaluation result between the equipment and the tasks to be processed.
3. The data center operation monitoring method based on a cloud platform according to claim 2, characterized in that, The selection of task equipment based on the equipment operation evaluation results in step S5 includes the following specific details: The matching evaluation results of each device and the task to be processed are obtained and compared with the set matching threshold. The device corresponding to the matching evaluation result with the smallest absolute value of the difference with the matching threshold is set as the task device, and the task is sent to the task device for calculation and processing.
4. The data center operation monitoring method based on a cloud platform according to claim 3, characterized in that, The adaptation analysis process is as follows: the reciprocal of the average standard deviation of the actual computing power in the future cycle and the average amount of computing data required by the task.
5. The data center operation monitoring method based on a cloud platform according to claim 1, characterized in that, The corresponding operational status of each device in the data center includes operational data reflecting the device's current, voltage, and temperature, as well as computing power and CPU usage data reflecting the device's computing power. The corresponding scenario data includes data on the type of task to be processed, the amount of task data, and the task time limit. The power system data includes data on power system fluctuations and power system operation status reflecting power supply quality. The power system fluctuations refer to the output voltage fluctuations of the power supply equipment, and the power operation status includes the current and voltage of the transmission lines.
6. A cloud-based data center operation monitoring system, implemented based on any one of claims 1-5, characterized in that, The system includes: The data acquisition module acquires the corresponding operating status and scenario data of each device in the data center, and also acquires power system data, which is then transmitted to the cloud platform. The compatibility analysis module analyzes the compatibility of scene and device operation based on the corresponding device operation status and scene data. The compensation and correction module compensates for and corrects equipment malfunctions by analyzing the corresponding operating conditions of each device and power system data. The equipment matching assessment module performs equipment matching assessment based on the compensation and correction results of equipment operation anomalies and the analysis results of equipment operation adaptability. The equipment selection module selects task equipment based on the results of equipment operation evaluation.
7. An electronic device, comprising: A processor and a memory, wherein the memory stores a computer program that can be called by the processor; characterized in that the processor executes the data center operation monitoring method based on a cloud platform as described in any one of claims 1-5 by calling the computer program stored in the memory.
Citation Information
Patent Citations
Distributed power supply system and method based on smart energy management platform
CN117767430A
Method for scheduling workloads of data center in electricity market environment, and device
WO2025082282A1