Intelligent operation and maintenance control system suitable for heterogeneous large-scale computing center
By introducing an intelligent operation and maintenance control system into heterogeneous large-scale computing centers, and combining resource occupancy status and data processing status for evaluation and optimization decisions, the problem of disordered operation and maintenance management in existing technologies has been solved, and the efficiency of operation and maintenance management and optimization has been improved.
Patent Information
- Application Number
- CN202510150938.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-02-11
AI Technical Summary
Existing heterogeneous large-scale computing center operation and maintenance management systems are unable to combine the resource occupancy status and data processing status of the computing center when handling computing tasks for operation and maintenance assessment and optimization decision analysis, resulting in disordered operation and maintenance management and low optimization efficiency.
An intelligent operation and maintenance control system was designed, including a computing power evaluation module, a processing analysis module, and an operation and maintenance management module. The system evaluates and optimizes the resource occupancy status, data processing status, and operation and maintenance status of computing tasks through evaluation coefficients, processing coefficients, and operation and maintenance coefficients, generates optimization signals, and sends them to the management personnel terminal.
It enables comprehensive evaluation and optimization decision analysis of the resource occupancy and data processing status of heterogeneous large-scale computing centers, thereby improving the efficiency of operation and maintenance management and optimization.
Smart Images

Figure CN120066916B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of computing power centers, and relates to intelligent operation and maintenance technology, and particularly relates to an intelligent operation and maintenance control system suitable for a heterogeneous large-scale computing power center. BACKGROUND
[0002] A heterogeneous large-scale computing power center refers to combining different types of computing units (such as CPUs, GPUs, FPGAs, etc.) in a computing system to fully exert the unique advantages of each processor, achieve higher computing performance and energy efficiency ratio, and the application scenarios of heterogeneous computing are very wide, including graphics rendering, deep learning, signal processing, data encryption, etc.
[0003] The patent for invention with the publication number CN110502390B discloses a university cloud computing center automatic operation and maintenance management system, which can timely find problems existing in the cloud computing center, realize fine distribution of operation and maintenance roles and effective tracing of operation and maintenance tasks, and improve the effect of operation and maintenance management of the cloud computing center; however, the operation and maintenance management system cannot combine the resource occupation state and data processing state of the computing center when processing computing tasks to perform operation and maintenance evaluation and optimization decision analysis, resulting in the problems of disorderly operation and maintenance and low optimization efficiency.
[0004] In view of the above technical problems, the present application provides a solution. SUMMARY
[0005] The application aims to provide an intelligent operation and maintenance control system suitable for a heterogeneous large-scale computing power center, which can solve the problem that operation and maintenance evaluation and optimization decision analysis cannot be performed in combination with the resource occupation state and data processing state of the computing center when processing computing tasks.
[0006] The technical problem to be solved by the application is how to provide an intelligent operation and maintenance control system suitable for a heterogeneous large-scale computing power center, which can combine the resource occupation state and data processing state of the computing center when processing computing tasks to perform operation and maintenance evaluation and optimization decision analysis.
[0007] The object of the application can be achieved by the following technical solutions.
[0008] The intelligent operation and maintenance control system suitable for a heterogeneous large-scale computing power center comprises a computing power evaluation module, a processing analysis module, an operation and maintenance management module and a database, the computing power evaluation module, the processing analysis module and the operation and maintenance management module are sequentially connected in communication, and the database is connected in communication with the computing power evaluation module, the processing analysis module and the operation and maintenance management module.
[0009] The computing power evaluation module is configured to evaluate and analyze the computing power resources of the heterogeneous large model computing power center: mark the computing nodes of the heterogeneous large model computing power center as analysis objects, generate an evaluation period, and obtain an evaluation coefficient PG of the computing task execution process when the heterogeneous large model computing power center performs a concurrent data computing task; and determine whether the computing resource occupation state of the computing task execution process meets the requirements through the evaluation coefficient PG.
[0010] The processing analysis module is configured to analyze the data computing processing state of the heterogeneous large model computing power center: obtain execution data ZX, parallel data BX, and energy consumption data NH of the computing task execution process and perform numerical calculation to obtain a processing coefficient CL of the computing task execution process; and determine whether the data computing processing state of the computing task execution process meets the requirements through the processing coefficient CL.
[0011] The operation and maintenance management module is configured to periodically analyze the operation and maintenance of the heterogeneous large model computing power center: obtain an operation and maintenance coefficient YW of the evaluation period at the end of the evaluation period, retrieve an operation and maintenance threshold YWmax through a database, compare the operation and maintenance coefficient YW of the evaluation period with the operation and maintenance threshold YWmax, determine that the operation and maintenance state of the heterogeneous large model computing power center meets the requirements if the operation and maintenance coefficient YW is less than the operation and maintenance threshold YWmax, determine that the operation and maintenance state of the heterogeneous large model computing power center does not meet the requirements if the operation and maintenance coefficient YW is greater than or equal to the operation and maintenance threshold YWmax, and perform optimization analysis on the heterogeneous large model computing power center.
[0012] Further, the evaluation coefficient PG of the computing task execution process includes: setting a plurality of evaluation time points in the computing task execution process, the time intervals between any two adjacent evaluation time points are equal, performing computing power occupation evaluation at the evaluation time points: obtaining the computing power resource occupation rate of the analysis object and marking it as the occupation value of the analysis object, summing and averaging all the occupation values of the analysis objects to obtain the occupation data of the evaluation time points; obtaining the occupation analysis value ZF and the concentration analysis value JF of the computing task execution process through all the occupation data of the evaluation time points; and obtaining the evaluation coefficient PG of the computing task execution process through numerical calculation of the occupation analysis value ZF and the concentration analysis value JF.
[0013] Further, at the end of the computing task execution process, the occupation data of all the evaluation time points are summed and averaged to obtain the occupation analysis value ZF, and the occupation data of all the evaluation time points are subjected to variance calculation to obtain the concentration analysis value JF.
[0014] Further, the specific process of determining whether the computing resource occupation state of the computing task execution process meets the requirements comprises: comparing the evaluation coefficient PG with the evaluation threshold PGmax by calling the evaluation threshold PGmax from the database; if the evaluation coefficient PG is less than the evaluation threshold PGmax, it is determined that the computing resource occupation state of the computing task execution process meets the requirements; if the evaluation coefficient PG is greater than or equal to the evaluation threshold PGmax, it is determined that the computing resource occupation state of the computing task execution process does not meet the requirements, and the computing task execution process is marked as an occupation abnormal process.
[0015] Further, the execution data ZX is the total amount of data processed by the computing task execution process, and the acquisition process of the parallel data BX comprises: acquiring the data processing amount completed by the analysis object in the computing task execution process and marking it as the processing value of the analysis object, and calculating the variance of the processing values of all analysis objects to obtain the parallel data BX. The energy consumption data NH is the sum of the energy consumption values of all analysis objects in the computing task execution process.
[0016] Further, the specific process of determining whether the data computing processing state of the computing task execution process meets the requirements comprises: comparing the processing coefficient CL of the computing task execution process with the processing threshold CLmax by calling the processing threshold CLmax from the database; if the processing coefficient CL is less than the processing threshold CLmax, it is determined that the data computing processing state of the computing task execution process meets the requirements; if the processing coefficient CL is greater than or equal to the processing threshold CLmax, it is determined that the data computing processing state of the computing task execution process does not meet the requirements, and the corresponding computing task execution process is marked as a processing abnormal process.
[0017] Further, the number of marked occupation abnormal processes and the number of marked processing abnormal processes are marked as occupation abnormal data ZY and processing abnormal data CY respectively, and the operation and maintenance coefficient YW of the evaluation period is obtained by numerically calculating the occupation abnormal data ZY and the processing abnormal data CY.
[0018] Further, the specific process of optimizing and analyzing the heterogeneous large model computing power center comprises: marking the computing task execution process that is simultaneously marked as an occupation abnormal process and a processing abnormal process as a coincident process, marking the ratio of the number of coincident processes in the evaluation period to the number of computing task execution processes as a coincidence coefficient, comparing the coincidence coefficient with the coincidence threshold by calling the coincidence threshold from the database; if the coincidence coefficient is less than the coincidence threshold, a hardware optimization signal is generated and sent to the mobile terminal of the manager; if the coincidence coefficient is greater than or equal to the coincidence threshold, a node configuration optimization signal is generated and sent to the mobile terminal of the manager.
[0019] The present application has the following advantages:
[0020] 1. The computing power assessment module can evaluate and analyze the computing power resources of heterogeneous large-scale model computing power centers. It can comprehensively analyze and calculate multiple resource usage parameters of computing nodes in heterogeneous large-scale model computing power centers when processing data computing tasks to obtain evaluation coefficients. The evaluation coefficients can be used to assess the computing power resource usage status and mark the computing task execution process when anomalies occur, providing data support for the operation and maintenance management analysis process.
[0021] 2. The processing and analysis module can analyze the data processing status of the heterogeneous large model computing center, calculate the processing coefficients by combining the processing parameters of the computing task execution process, and provide feedback on the data processing status of the computing task execution process based on the processing coefficients.
[0022] 3. The operation and maintenance management module can perform periodic operation and maintenance management analysis on heterogeneous large model computing power centers. When the operation and maintenance status does not meet the requirements, optimization analysis can be used to provide optimization directions for heterogeneous large model computing power centers and improve their optimization efficiency. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a system block diagram of Embodiment 1 of the present invention;
[0025] Figure 2 This is a flowchart of the method in Embodiment 2 of the present invention. Detailed Implementation
[0026] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] Example 1: As Figure 1 As shown, an intelligent operation and maintenance control system suitable for heterogeneous large-scale computing centers includes a computing power evaluation module, a processing and analysis module, an operation and maintenance management module, and a database. The computing power evaluation module, the processing and analysis module, and the operation and maintenance management module are connected to each other in sequence, and the database is also connected to the computing power evaluation module, the processing and analysis module, and the operation and maintenance management module.
[0028] The computing power assessment module is used to assess and analyze the computing power resources of heterogeneous large-scale model computing power centers. It marks the computing nodes of the heterogeneous large-scale model computing power centers as analysis objects, generates assessment periods, and sets several assessment time points during the execution of concurrent data computing tasks in the heterogeneous large-scale model computing power centers. The time interval between any two adjacent assessment time points is equal. At each assessment time point, computing power occupancy is assessed: the computing power resource occupancy rate of the analysis object is obtained and marked as the occupancy value of the analysis object. The occupancy values of all analysis objects are summed and averaged to obtain the occupancy data at each assessment time point. At the end of the computing task execution process, the occupancy data at all assessment time points are summed and averaged to obtain the occupancy analysis value ZF. The variance of the occupancy data at all assessment time points is calculated to obtain the ensemble analysis value JF.
[0029] The evaluation coefficient PG for the computation task execution process is obtained using the formula PG=k1×ZF+k2×JF, where k1 and k2 are proportional coefficients, and k1>k2>1. The evaluation threshold PGmax is retrieved from the database, and the evaluation coefficient PG is compared with the evaluation threshold PGmax: if the evaluation coefficient PG is less than the evaluation threshold PGmax, the computational resource occupancy status of the computation task execution process is determined to meet the requirements; if the evaluation coefficient PG is greater than or equal to the evaluation threshold PGmax, the computational resource occupancy status of the computation task execution process is determined to not meet the requirements, and the computation task execution process is marked as an abnormal occupancy process. The evaluation coefficient is obtained by comprehensively analyzing and calculating multiple resource occupancy parameters of the computing nodes in the heterogeneous large-scale model computing center when processing data computation tasks. The evaluation coefficient is used to assess the computational resource occupancy status, and the computation task execution process is marked when abnormalities occur, providing data support for the operation and maintenance management analysis process.
[0030] The processing and analysis module is used to analyze the data computing processing status of the heterogeneous large-scale model computing center: it acquires the execution data ZX, parallel data BX, and energy consumption data NH of the computing task execution process. The execution data ZX is the total amount of data processed during the computing task execution process. The acquisition process of parallel data BX includes: acquiring the amount of data processed by the analysis object during the computing task execution process and marking it as the processing value of the analysis object; calculating the variance of the processing values of all analysis objects to obtain the parallel data BX; and the energy consumption data NH is the sum of the energy consumption values of all analysis objects during the computing task execution process.
[0031] The processing coefficient CL for the computation task execution process is obtained using the formula CL=c1×BX+c2×NH-c3×ZX / SC, where c1, c2, and c3 are proportionality coefficients, and c1>c2>c3>1, and SC is the execution duration of the computation task. The processing threshold CLmax is retrieved from the database, and the processing coefficient CL is compared with the threshold CLmax: if the processing coefficient CL is less than the threshold CLmax, the data computation processing status of the computation task is deemed to meet the requirements; if the processing coefficient CL is greater than or equal to the threshold CLmax, the data computation processing status of the computation task is deemed to not meet the requirements, and the corresponding computation task execution process is marked as an abnormal process. The processing coefficient is calculated in conjunction with the processing parameters of the computation task execution process, and feedback on the data computation processing status of the computation task execution process is provided based on the processing coefficient.
[0032] The operation and maintenance management module is used to perform periodic operation and maintenance management analysis on heterogeneous large model computing power centers: at the end of the evaluation period, the number of marked abnormal processes and the number of marked abnormal processes are marked as occupied abnormal data ZY and handled abnormal data CY, respectively. The operation and maintenance coefficient YW of the evaluation period is obtained by formula YW=a1×ZY+a2×CY, where a1 and a2 are both proportional coefficients, and a1>a2>1.
[0033] The system retrieves the operation and maintenance threshold YWmax from the database and compares the operation and maintenance coefficient YW during the evaluation period with the operation and maintenance threshold YWmax. If the operation and maintenance coefficient YW is less than the operation and maintenance threshold YWmax, the operation and maintenance status of the heterogeneous large-scale computing center is deemed to meet the requirements. If the operation and maintenance coefficient YW is greater than or equal to the operation and maintenance threshold YWmax, the operation and maintenance status of the heterogeneous large-scale computing center is deemed to not meet the requirements, and optimization analysis is performed on the heterogeneous large-scale computing center. Computational task execution processes that are simultaneously marked as both occupancy and handling abnormal processes are marked as overlapping processes. The ratio of the number of overlapping processes to the number of computational task execution processes within the evaluation period is marked as the overlap coefficient. The system retrieves the overlap threshold from the database and compares the overlap coefficient with the overlap threshold. If the overlap coefficient is less than the overlap threshold, a hardware optimization signal is generated and sent to the administrator's mobile terminal. If the overlap coefficient is greater than or equal to the overlap threshold, a node configuration optimization signal is generated and sent to the administrator's mobile terminal. When the operation and maintenance status does not meet the requirements, optimization analysis provides optimization directions for the heterogeneous large-scale computing center, improving its optimization efficiency.
[0034] Example 2: Figure 2 As shown, an intelligent operation and maintenance control method applicable to heterogeneous large-scale computing centers includes the following steps:
[0035] Step one: the computing power resources of the heterogeneous large model computing power center are evaluated and analyzed: the computing nodes of the heterogeneous large model computing power center are marked as the analysis object, an evaluation period is generated, when the heterogeneous large model computing power center executes the concurrent data computing task, the evaluation coefficient PG of the computing task execution process is obtained, whether the computing power resource occupation state of the computing task execution process meets the requirements is judged through the evaluation coefficient PG;
[0036] Step two: the data computing processing state of the heterogeneous large model computing power center is analyzed: the execution data ZX, parallel data BX and energy consumption data NH of the computing task execution process are obtained and numerically calculated to obtain the processing coefficient CL of the computing task execution process, whether the data computing processing state of the computing task execution process meets the requirements is judged through the processing coefficient CL;
[0037] Step three: the periodic operation and maintenance management analysis of the heterogeneous large model computing power center is performed: the occupied heterogeneous data ZY and the heterogeneous data CY are obtained at the end of the evaluation period and are numerically calculated to obtain the operation and maintenance coefficient YW, whether the operation and maintenance state of the heterogeneous large model computing power center meets the requirements is judged through the operation and maintenance coefficient YW, and optimization analysis is performed when the requirements are not met.
[0038] The intelligent operation and maintenance control system suitable for the heterogeneous large-scale computing power center, when working, the computing nodes of the heterogeneous large model computing power center are marked as the analysis object, an evaluation period is generated, when the heterogeneous large model computing power center executes the concurrent data computing task, the evaluation coefficient PG of the computing task execution process is obtained, whether the computing power resource occupation state of the computing task execution process meets the requirements is judged through the evaluation coefficient PG; the execution data ZX, parallel data BX and energy consumption data NH of the computing task execution process are obtained and numerically calculated to obtain the processing coefficient CL of the computing task execution process, whether the data computing processing state of the computing task execution process meets the requirements is judged through the processing coefficient CL; the occupied heterogeneous data ZY and the heterogeneous data CY are obtained at the end of the evaluation period and are numerically calculated to obtain the operation and maintenance coefficient YW, whether the operation and maintenance state of the heterogeneous large model computing power center meets the requirements is judged through the operation and maintenance coefficient YW, and optimization analysis is performed when the requirements are not met.
[0039] The above content is only an example and description of the structure of the application, and those skilled in the art can make various modifications or supplements to the described specific embodiments or use similar ways to replace, as long as the modifications or supplements do not deviate from the structure of the application or exceed the scope defined by the present claims, and should belong to the protection scope of the application.
[0040] The above formulas are obtained by collecting a large amount of data for software simulation and selecting one formula close to the true value, and the coefficients in the formula are set by a person skilled in the art according to the actual situation; for example: the formula CL = c1BX + c2NH - c3ZX / SC; a plurality of sample data are collected by a person skilled in the art, and a corresponding processing coefficient is set for each sample data; the set processing coefficient and the collected sample data are substituted into the formula, any three formulas constitute a ternary linear equation group, the calculated coefficients are screened and the mean value is taken, and the values of c1, c2 and c3 are 3.85, 2.74 and 2.02 respectively;
[0041] The size of the coefficient is a specific value obtained by quantifying each parameter, which is convenient for subsequent comparison. The size of the coefficient depends on the number of sample data and the corresponding processing coefficient initially set by a person skilled in the art for each sample data. As long as it does not affect the proportional relationship between the parameter and the quantized value, the processing coefficient is directly proportional to the value of the parallel data.
[0042] In the description of the present specification, the description of the terms "one embodiment", "example", "specific example" and the like means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0043] The preferred embodiments of the application disclosed above are only used to help explain the application. The preferred embodiments do not describe all the details and limit the application to the specific embodiments. Obviously, many modifications and changes can be made according to the content of the present specification. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the application, so that a person skilled in the art can well understand and utilize the application. The application is limited by the claims and their entire scope and equivalents.
Claims
1. An intelligent operation and maintenance control system suitable for heterogeneous large-scale computing centers, characterized in that, It includes a computing power assessment module, a processing and analysis module, an operation and maintenance management module, and a database. The computing power assessment module, the processing and analysis module, and the operation and maintenance management module are sequentially connected to each other. The database is also connected to the computing power assessment module, the processing and analysis module, and the operation and maintenance management module. The computing power evaluation module is used to evaluate and analyze the computing power resources of the heterogeneous large model computing power center: marking the computing nodes of the heterogeneous large model computing power center as analysis objects, generating evaluation cycles, and obtaining the evaluation coefficient PG of the computing task execution process when executing concurrent data computing tasks in the heterogeneous large model computing power center. The evaluation coefficient PG is used to determine whether the computing resource occupancy status during the execution of the computing task meets the requirements; if the computing resource occupancy status does not meet the requirements, the execution process of the computing task is marked as an abnormal occupancy process. The processing and analysis module is used to analyze the data computing and processing status of the heterogeneous large model computing center: to obtain the execution data ZX, parallel data BX and energy consumption data NH of the computing task execution process and to perform numerical calculations to obtain the processing coefficient CL of the computing task execution process; The processing coefficient CL is used to determine whether the data processing status of the computation task execution process meets the requirements. If the data computation and processing status does not meet the requirements, the corresponding computation task execution process will be marked as an abnormal process. The number of marked abnormal processes and the number of marked abnormal processes are respectively marked as abnormal data ZY and abnormal data CY. The operation and maintenance coefficient YW of the evaluation cycle is obtained by numerically calculating the abnormal data ZY and abnormal data CY. The operation and maintenance management module is used to perform periodic operation and maintenance management analysis on the heterogeneous large model computing center: at the end of the evaluation period, the operation and maintenance coefficient YW of the evaluation period is obtained, the operation and maintenance threshold YWmax is retrieved from the database, and the operation and maintenance coefficient YW of the evaluation period is compared with the operation and maintenance threshold YWmax: if the operation and maintenance coefficient YW is less than the operation and maintenance threshold YWmax, it is determined that the operation and maintenance status of the heterogeneous large model computing center meets the requirements. If the operation and maintenance coefficient YW is greater than or equal to the operation and maintenance threshold YWmax, it is determined that the operation and maintenance status of the heterogeneous large model computing center does not meet the requirements, and optimization analysis is performed on the heterogeneous large model computing center. The specific process of optimizing the heterogeneous large-scale computing center includes: marking the execution process of computing tasks that are simultaneously marked as occupancy abnormal process and processing abnormal process as overlapping process; marking the ratio of the number of overlapping processes to the number of computing task execution processes within the evaluation period as the overlap coefficient; retrieving the overlap threshold from the database; comparing the overlap coefficient with the overlap threshold; if the overlap coefficient is less than the overlap threshold, generating a hardware optimization signal and sending the hardware optimization signal to the mobile terminal of the management personnel. If the overlap coefficient is greater than or equal to the overlap threshold, a node configuration optimization signal is generated and sent to the administrator's mobile terminal.
2. The intelligent operation and maintenance control system for heterogeneous large-scale computing centers according to claim 1, characterized in that, The process of obtaining the evaluation coefficient PG for the computation task execution process includes: setting several evaluation time points within the computation task execution process, with equal time intervals between any two adjacent evaluation time points; evaluating computing power occupancy at each evaluation time point; obtaining the computing power resource occupancy rate of the analysis object and marking it as the occupancy value of the analysis object; summing and averaging the occupancy values of all analysis objects to obtain the occupancy data at each evaluation time point; obtaining the occupancy analysis value ZF and the centralized analysis value JF for the computation task execution process through the occupancy data at all evaluation time points; and obtaining the evaluation coefficient PG for the computation task execution process by numerically calculating the occupancy analysis value ZF and the centralized analysis value JF.
3. The intelligent operation and maintenance control system for heterogeneous large-scale computing centers according to claim 2, characterized in that, At the end of the computation task execution process, the occupancy data at all evaluation time points are summed and averaged to obtain the occupancy analysis value ZF, and the variance of the occupancy data at all evaluation time points is calculated to obtain the ensemble analysis value JF.
4. The intelligent operation and maintenance control system for heterogeneous large-scale computing centers according to claim 3, characterized in that, The specific process for determining whether the computing resource occupancy status of the computing task execution process meets the requirements includes: retrieving the evaluation threshold PGmax from the database and comparing the evaluation coefficient PG with the evaluation threshold PGmax; if the evaluation coefficient PG is less than the evaluation threshold PGmax, the computing resource occupancy status of the computing task execution process is determined to meet the requirements; if the evaluation coefficient PG is greater than or equal to the evaluation threshold PGmax, the computing resource occupancy status of the computing task execution process is determined to not meet the requirements, and the computing task execution process is marked as an abnormal occupancy process.
5. The intelligent operation and maintenance control system for heterogeneous large-scale computing centers according to claim 4, characterized in that, The execution data ZX is the total amount of data processed during the execution of the computation task. The process of obtaining the parallel data BX includes: obtaining the amount of data processed by the analysis object during the execution of the computation task and marking it as the processing value of the analysis object; calculating the variance of the processing values of all analysis objects to obtain the parallel data BX; and the energy consumption data NH is the sum of the energy consumption values of all analysis objects during the execution of the computation task.
6. The intelligent operation and maintenance control system for heterogeneous large-scale computing centers according to claim 5, characterized in that, The specific process for determining whether the data processing status of the computation task execution process meets the requirements includes: retrieving the processing threshold CLmax from the database, comparing the processing coefficient CL of the computation task execution process with the processing threshold CLmax; if the processing coefficient CL is less than the processing threshold CLmax, the data processing status of the computation task execution process is determined to meet the requirements; if the processing coefficient CL is greater than or equal to the processing threshold CLmax, the data processing status of the computation task execution process is determined to not meet the requirements, and the corresponding computation task execution process is marked as an abnormal process.
Citation Information
Patent Citations
An automated operation and maintenance management system for university cloud computing centers
CN110502390B
Computing power pooling system for improving GPU (Graphics Processing Unit) utilization efficiency
CN115202836A
Operation and maintenance management monitoring method and device
CN118503043A