Intelligent operation and maintenance control system suitable for heterogeneous large-scale computing power center
By designing an intelligent operation and maintenance control system, combining computing power evaluation and processing analysis modules, the problem of inability to evaluate resource occupation and data processing status in the operation and maintenance management of heterogeneous large-scale computing power centers is solved, and more efficient operation and maintenance management and optimization analysis are achieved.
Patent Information
- Application Number
- CN202510150938.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The existing heterogeneous large-scale computing power center operation and maintenance management system cannot combine the resource occupation status and data processing status of the calculation task to conduct operation and maintenance evaluation and optimization decision analysis, resulting in disordered operation and maintenance management and low optimization efficiency.
An intelligent operation and maintenance control system is designed, including a computing power evaluation module, processing and analysis module and operation and maintenance management module. Through the coordinated work of these modules, the computing power resources and data calculation processing status of the heterogeneous large model computing power center can be evaluated and analyzed, and evaluation coefficients and processing coefficients can be generated to determine whether the resource occupation and processing status meet the requirements and perform optimization analysis.
Comprehensive evaluation and optimization analysis of the resource occupation and data processing status of heterogeneous large model computing power centers is realized, which improves the efficiency and effectiveness of operation and maintenance management, and ensures the normal execution of computing tasks and the rational utilization of resources.
Smart Images

Figure CN120066916A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computing power centers, relates to intelligent operation and maintenance technology, and specifically is an intelligent operation and maintenance control system applicable to heterogeneous large-scale computing power centers. Background Art
[0002] A heterogeneous large-scale computing power center refers to combining different types of computing units (such as CPUs, GPUs, FPGAs, etc.) in a computing system to give full play to the unique advantages of each processor and achieve higher computing performance and energy efficiency ratio. The application scenarios of heterogeneous computing are very extensive, including fields such as graphics rendering, deep learning, signal processing, and data encryption.
[0003] The invention patent with the publication number CN110502390B discloses an automated operation and maintenance management system for a college cloud computing center. This operation and maintenance management system can timely discover problems existing in the cloud computing center, realize the refined allocation of operation and maintenance roles and the effective traceability of operation and maintenance tasks, and improve the operation and maintenance management effect of the cloud computing center; however, this operation and maintenance management system cannot combine the resource occupancy status and data processing status when the computing center processes computing tasks for operation and maintenance evaluation and optimization decision-making analysis, resulting in problems such as disordered operation and maintenance management and low optimization efficiency.
[0004] In view of the above technical problems, this application proposes a solution. Summary of the Invention
[0005] The purpose of the present invention is to provide an intelligent operation and maintenance control system applicable to heterogeneous large-scale computing power centers, which is used to solve the problem that operation and maintenance evaluation and optimization decision-making analysis cannot be combined with the resource occupancy status and data processing status when the computing center processes computing tasks;
[0006] The technical problem that the present invention needs to solve is: how to provide an intelligent operation and maintenance control system applicable to heterogeneous large-scale computing power centers that can combine the resource occupancy status and data processing status when the computing center processes computing tasks for operation and maintenance evaluation and optimization decision-making analysis.
[0007] The purpose of the present invention can be achieved through the following technical solutions:
[0008] An intelligent operation and maintenance control system applicable to heterogeneous large-scale computing power centers includes a computing power evaluation module, a processing and analysis module, an operation and maintenance management module, and a database. The computing power evaluation module, the processing and analysis module, and the operation and maintenance management module are sequentially communicatively connected, and the database is communicatively connected to the computing power evaluation module, the processing and analysis module, and the operation and maintenance management module;
[0009] The computing power evaluation module is used to evaluate and analyze the computing power resources of the heterogeneous large model computing power center: mark the computing nodes of the heterogeneous large model computing power center as analysis objects, generate an evaluation period, and obtain the evaluation coefficient PG of the computing task execution process when the concurrent data computing task is executed in the heterogeneous large model computing power center; determine whether the computing resource occupancy status during the execution of the computing task meets the requirements through the evaluation coefficient PG;
[0010] The processing and analysis module is used to analyze the data computing and processing status of the heterogeneous large model computing power center: obtain the execution data ZX, parallel data BX, and energy consumption data NH during the execution of the computing task and perform numerical calculations to obtain the processing coefficient CL of the computing task execution process; determine whether the data computing and processing status during the execution of the computing task meets the requirements through the processing coefficient CL;
[0011] The operation and maintenance management module is used to perform periodic operation and maintenance management analysis on the heterogeneous large model computing power center: obtain the operation and maintenance coefficient YW of the evaluation period at the end of the evaluation period, retrieve the operation and maintenance threshold YWmax through the database, and compare the operation and maintenance coefficient YW of the evaluation period with the operation and maintenance threshold YWmax: if the operation and maintenance coefficient YW is less than the operation and maintenance threshold YWmax, it is determined that the operation and maintenance status of the heterogeneous large model computing power center meets the requirements; if the operation and maintenance coefficient YW is greater than or equal to the operation and maintenance threshold YWmax, it is determined that the operation and maintenance status of the heterogeneous large model computing power center does not meet the requirements, and optimize the analysis of the heterogeneous large model computing power center.
[0012] Furthermore, the process of obtaining the evaluation coefficient PG of the computing task execution process includes: setting several evaluation time points during the execution of the computing task, with equal time intervals between any two adjacent evaluation time points, and performing computing power occupancy evaluation at the evaluation time points: obtain the computing power occupancy rate of the analysis object and mark it as the occupancy value of the analysis object, sum and average the occupancy values of all analysis objects to obtain the occupancy data of the evaluation time point; obtain the occupancy analysis value ZF and the centralized analysis value JF of the computing task execution process through the occupancy data of all evaluation time points; obtain the evaluation coefficient PG of the computing task execution process through numerical calculations on the occupancy analysis value ZF and the centralized analysis value JF.
[0013] Furthermore, at the end of the computing task execution process, sum and average the occupancy data of all evaluation time points to obtain the occupancy analysis value ZF, and calculate the variance of the occupancy data of all evaluation time points to obtain the centralized analysis value JF.
[0014] Further, the specific process for determining whether the computing resource occupancy status during the execution of a computing task meets the requirements includes: retrieving the evaluation threshold PGmax from the database, and comparing the evaluation coefficient PG with the evaluation threshold PGmax. If the evaluation coefficient PG is less than the evaluation threshold PGmax, it is determined that the computing power resource occupancy status during the execution of the computing task meets the requirements. If the evaluation coefficient PG is greater than or equal to the evaluation threshold PGmax, it is determined that the computing power resource occupancy status during the execution of the computing task does not meet the requirements, and the execution process of the computing task is marked as an abnormal occupancy process.
[0015] Further, the execution data ZX is the total amount of data processed during the execution of the computing task. The process of obtaining the parallel data BX includes: obtaining the amount of data processed by the analysis object during the execution of the computing task and marking it as the processing value of the analysis object, calculating the variance of the processing values of all analysis objects to obtain the parallel data BX, and the energy consumption data NH is the sum of the energy consumption values of all analysis objects during the execution of the computing task.
[0016] Further, the specific process for determining whether the data calculation and processing status during the execution of the computing task meets the requirements includes: retrieving the processing threshold CLmax from the database, and comparing the processing coefficient CL during the execution of the computing task with the processing threshold CLmax. If the processing coefficient CL is less than the processing threshold CLmax, it is determined that the data calculation and processing status during the execution of the computing task meets the requirements. If the processing coefficient CL is greater than or equal to the processing threshold CLmax, it is determined that the data calculation and processing status during the execution of the computing task does not meet the requirements, and the corresponding execution process of the computing task is marked as an abnormal processing process.
[0017] Further, the marked numbers of the abnormal occupancy processes and the abnormal processing processes are respectively marked as the occupancy abnormal data ZY and the processing abnormal data CY, and the operation and maintenance coefficient YW of the evaluation period is obtained by performing numerical calculations on the occupancy abnormal data ZY and the processing abnormal data CY.
[0018] Further, the specific process for optimizing and analyzing the heterogeneous large model computing center includes: marking the execution processes of the computing tasks that are simultaneously marked as abnormal occupancy processes and abnormal processing processes as overlapping processes, marking the ratio of the number of overlapping processes within the evaluation period to the number of execution processes of the computing tasks as the overlapping coefficient, retrieving the overlapping threshold from the database, and comparing the overlapping coefficient with the overlapping threshold. If the overlapping coefficient is less than the overlapping threshold, a hardware optimization signal is generated and sent to the mobile terminal of the management personnel. If the overlapping coefficient is greater than or equal to the overlapping threshold, a node configuration optimization signal is generated and sent to the mobile terminal of the management personnel.
[0019] The present invention has the following beneficial effects:
[0020] 1. The computing power evaluation module can evaluate and analyze the computing power resources of the heterogeneous large model computing power center, comprehensively analyze and calculate multiple resource occupancy parameters of the computing nodes in the heterogeneous large model computing power center when processing data computing tasks to obtain an evaluation coefficient, evaluate the computing power resource occupancy status through the evaluation coefficient, mark the execution process of the computing task in case of anomalies, and provide data support for the operation and maintenance management analysis process;
[0021] 2. The processing and analysis module can analyze the data computing and processing status of the heterogeneous large model computing power center, calculate a processing coefficient by combining the processing parameters in the execution process of the computing task, and provide feedback on the data computing and processing status in the execution process of the computing task according to the processing coefficient;
[0022] 3. The operation and maintenance management module can perform periodic operation and maintenance management analysis on the heterogeneous large model computing power center, and give the optimization direction of the heterogeneous large model computing power center through optimization analysis when the operation and maintenance status does not meet the requirements, so as to improve its optimization efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 It is the system block diagram of Embodiment 1 of the present invention;
[0025] Figure 2 It is the method flow chart of Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0027] Embodiment 1: As Figure 1 shown, an intelligent operation and maintenance control system applicable to a heterogeneous large-scale computing power center includes a computing power evaluation module, a processing and analysis module, an operation and maintenance management module, and a database. The computing power evaluation module, the processing and analysis module, and the operation and maintenance management module are sequentially connected for communication, and the database is communicatively connected to the computing power evaluation module, the processing and analysis module, and the operation and maintenance management module.
[0028] The computing power evaluation module is used to evaluate and analyze the computing power resources of the heterogeneous large model computing power center: Mark the computing nodes of the heterogeneous large model computing power center as the analysis objects, generate an evaluation period. When performing concurrent data computing tasks in the heterogeneous large model computing power center, set several evaluation time points during the execution of the computing task. The time intervals between any two adjacent evaluation time points are equal. Perform computing power occupancy evaluation at the evaluation time points: Obtain the computing power resource occupancy rate of the analysis object and mark it as the occupancy value of the analysis object. Sum up and average the occupancy values of all analysis objects to obtain the occupancy data at the evaluation time point; At the end moment of the execution process of the computing task, sum up and average the occupancy data of all evaluation time points to obtain the occupancy analysis value ZF, and calculate the variance of the occupancy data of all evaluation time points to obtain the concentration analysis value JF.
[0029] Obtain the evaluation coefficient PG of the computing task execution process through the formula PG = k1×ZF + k2×JF, where both k1 and k2 are proportionality coefficients, and k1 > k2 > 1; Retrieve the evaluation threshold PGmax through the database, and compare the evaluation coefficient PG with the evaluation threshold PGmax: If the evaluation coefficient PG is less than the evaluation threshold PGmax, it is determined that the computing power resource occupancy status of the computing task execution process meets the requirements; If the evaluation coefficient PG is greater than or equal to the evaluation threshold PGmax, it is determined that the computing power resource occupancy status of the computing task execution process does not meet the requirements, and mark the computing task execution process as an occupancy abnormal process; Conduct comprehensive analysis and calculation on multiple resource occupancy parameters of the computing nodes of the heterogeneous large model computing power center during data computing tasks to obtain an evaluation coefficient, evaluate the computing power resource occupancy status through the evaluation coefficient, and mark the computing task execution process in case of anomalies, providing data support for the operation and maintenance management analysis process.
[0030] The processing and analysis module is used to analyze the data computing and processing status of the heterogeneous large model computing power center: Obtain the execution data ZX, parallel data BX, and energy consumption data NH of the computing task execution process. The execution data ZX is the total amount of data processed during the computing task execution process. The process of obtaining the parallel data BX includes: Obtain the amount of data processed by the analysis object during the computing task execution process and mark it as the processing value of the analysis object. Calculate the variance of the processing values of all analysis objects to obtain the parallel data BX. The energy consumption data NH is the sum of the energy consumption values of all analysis objects during the computing task execution process.
[0031] The processing coefficient CL of the computing task execution process is obtained through the formula CL = c1×BX + c2×NH - c3×ZX / SC, where c1, c2, and c3 are all proportionality coefficients, and c1 > c2 > c3 > 1, and SC is the duration of the computing task execution process; the processing threshold CLmax is retrieved from the database, and the processing coefficient CL of the computing task execution process is compared with the processing threshold CLmax: if the processing coefficient CL is less than the processing threshold CLmax, it is determined that the data calculation and processing status of the computing task execution process meets the requirements; if the processing coefficient CL is greater than or equal to the processing threshold CLmax, it is determined that the data calculation and processing status of the computing task execution process does not meet the requirements, and the corresponding computing task execution process is marked as a processing abnormal process; the processing coefficient is calculated by combining the processing parameters of the computing task execution process, and the data calculation and processing status of the computing task execution process is fed back according to the processing coefficient.
[0032] The operation and maintenance management module is used to perform periodic operation and maintenance management analysis on the heterogeneous large model computing power center: at the end of the evaluation period, the number of marked occupation abnormal processes and the number of marked processing abnormal processes are respectively marked as occupation abnormal data ZY and processing abnormal data CY, and the operation and maintenance coefficient YW of the evaluation period is obtained through the formula YW = a1×ZY + a2×CY, where a1 and a2 are both proportionality coefficients, and a1 > a2 > 1.
[0033] The operation and maintenance threshold YWmax is retrieved from the database, and the operation and maintenance coefficient YW of the evaluation period is compared with the operation and maintenance threshold YWmax: if the operation and maintenance coefficient YW is less than the operation and maintenance threshold YWmax, it is determined that the operation and maintenance status of the heterogeneous large model computing power center meets the requirements; if the operation and maintenance coefficient YW is greater than or equal to the operation and maintenance threshold YWmax, it is determined that the operation and maintenance status of the heterogeneous large model computing power center does not meet the requirements, and optimization analysis is performed on the heterogeneous large model computing power center: the computing task execution processes that are simultaneously marked as occupation abnormal processes and processing abnormal processes are marked as overlapping processes, the ratio of the number of overlapping processes in the evaluation period to the number of computing task execution processes is marked as the overlapping coefficient, the overlapping threshold is retrieved from the database, and the overlapping coefficient is compared with the overlapping threshold: if the overlapping coefficient is less than the overlapping threshold, a hardware optimization signal is generated and sent to the mobile terminal of the management personnel; if the overlapping coefficient is greater than or equal to the overlapping threshold, a node configuration optimization signal is generated and sent to the mobile terminal of the management personnel; when the operation and maintenance status does not meet the requirements, the optimization direction of the heterogeneous large model computing power center is given through optimization analysis to improve its optimization efficiency.
[0034] Embodiment 2: As Figure 2 shown, an intelligent operation and maintenance control method applicable to a heterogeneous large-scale computing power center includes the following steps:
[0035] Step 1: Evaluate and analyze the computing power resources of the heterogeneous large model computing power center: Mark the computing nodes of the heterogeneous large model computing power center as the analysis objects, generate an evaluation period. When performing concurrent data computing tasks in the heterogeneous large model computing power center, obtain the evaluation coefficient PG during the execution process of the computing task, and determine whether the occupancy status of the computing power resources during the execution process of the computing task meets the requirements through the evaluation coefficient PG;
[0036] Step 2: Analyze the data computing and processing status of the heterogeneous large model computing power center: Obtain the execution data ZX, parallel data BX, and energy consumption data NH during the execution process of the computing task and perform numerical calculations to obtain the processing coefficient CL of the execution process of the computing task, and determine whether the data computing and processing status during the execution process of the computing task meets the requirements through the processing coefficient CL;
[0037] Step 3: Conduct periodic operation and maintenance management analysis on the heterogeneous large model computing power center: Obtain the heterogeneous data ZY and the processed heterogeneous data CY at the end of the evaluation period and perform numerical calculations to obtain the operation and maintenance coefficient YW, and determine whether the operation and maintenance status of the heterogeneous large model computing power center meets the requirements through the operation and maintenance coefficient YW. When the requirements are not met, perform optimization analysis.
[0038] An intelligent operation and maintenance control system applicable to heterogeneous large-scale computing power centers, when working, marks the computing nodes of the heterogeneous large model computing power center as the analysis objects, generates an evaluation period. When performing concurrent data computing tasks in the heterogeneous large model computing power center, obtain the evaluation coefficient PG during the execution process of the computing task, and determine whether the occupancy status of the computing power resources during the execution process of the computing task meets the requirements through the evaluation coefficient PG; Obtain the execution data ZX, parallel data BX, and energy consumption data NH during the execution process of the computing task and perform numerical calculations to obtain the processing coefficient CL of the execution process of the computing task, and determine whether the data computing and processing status during the execution process of the computing task meets the requirements through the processing coefficient CL; Obtain the heterogeneous data ZY and the processed heterogeneous data CY at the end of the evaluation period and perform numerical calculations to obtain the operation and maintenance coefficient YW, and determine whether the operation and maintenance status of the heterogeneous large model computing power center meets the requirements through the operation and maintenance coefficient YW. When the requirements are not met, perform optimization analysis.
[0039] The above content is only an example and explanation of the structure of the present invention. Those skilled in the art of this technology make various modifications or supplements to the described specific embodiments or use similar methods for substitution. As long as they do not deviate from the structure of the invention or exceed the scope defined by this claim book, they should fall within the protection scope of the present invention.
[0040] The above formulas are all obtained by collecting a large amount of data for software simulation and selecting a formula close to the true value. The coefficients in the formula are set by those skilled in the art according to the actual situation. For example, the formula CL = c1×BX + c2×NH - c3×ZX / SC; those skilled in the art collect multiple groups of sample data and set corresponding processing coefficients for each group of sample data; substitute the set processing coefficients and the collected sample data into the formula, and any three formulas form a system of linear equations with three variables. Screen the calculated coefficients and take the average value to obtain the values of c1, c2, and c3 as 3.85, 2.74, and 2.02 respectively.
[0041] The magnitude of the coefficient is a specific value obtained by quantifying each parameter for subsequent comparison. Regarding the magnitude of the coefficient, it depends on the amount of sample data and the corresponding processing coefficients initially set by those skilled in the art for each group of sample data; as long as it does not affect the proportional relationship between the parameter and the quantified value, for example, the processing coefficient is directly proportional to the value of the parallel data.
[0042] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0043] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not elaborate on all the details, nor do they limit the present invention to only the specific implementation manners. Obviously, according to the content of this specification, many modifications and variations can be made. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the present invention, so that those skilled in the art in the relevant technical field can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. An intelligent operation and maintenance control system applicable to heterogeneous large-scale computing centers, characterized by: It includes a computing power evaluation module, a processing and analysis module, an operation and maintenance management module and a database, wherein the computing power evaluation module, the processing and analysis module and the operation and maintenance management module are communicatively connected in sequence, and the database is communicatively connected with the computing power evaluation module, the processing and analysis module and the operation and maintenance management module; The computing power evaluation module is used to evaluate and analyze the computing power resources of the heterogeneous large model computing power center: mark the computing nodes of the heterogeneous large model computing power center as analysis objects, generate an evaluation cycle, and obtain the evaluation coefficient PG of the computing task execution process when the heterogeneous large model computing power center executes concurrent data computing tasks; The evaluation coefficient PG is used to determine whether the computing resource occupancy status of the computing task execution process meets the requirements; The processing and analysis module is used to analyze the data computing and processing status of the heterogeneous large model computing center: obtain the execution data ZX, parallel data BX and energy consumption data NH of the computing task execution process and perform numerical calculations to obtain the processing coefficient CL of the computing task execution process; The processing coefficient CL is used to determine whether the data calculation processing status during the calculation task execution process meets the requirements; The operation and maintenance management module is used to perform periodic operation and maintenance management analysis on the heterogeneous large model computing power center: at the end of the evaluation period, the operation and maintenance coefficient YW of the evaluation period is obtained, the operation and maintenance threshold YWmax is retrieved through the database, and the operation and maintenance coefficient YW of the evaluation period is compared with the operation and maintenance threshold YWmax: if the operation and maintenance coefficient YW is less than the operation and maintenance threshold YWmax, it is determined that the operation and maintenance status of the heterogeneous large model computing power center meets the requirements; If the operation and maintenance coefficient YW is greater than or equal to the operation and maintenance threshold YWmax, it is determined that the operation and maintenance status of the heterogeneous large model computing power center does not meet the requirements, and the heterogeneous large model computing power center is optimized and analyzed.
2. The intelligent operation and maintenance control system applicable to heterogeneous large-scale computing centers according to claim 1 is characterized in that: The process of obtaining the evaluation coefficient PG of the computing task execution process includes: setting a number of evaluation time points in the computing task execution process, the time intervals between any two adjacent evaluation time points are equal, and performing computing power occupancy evaluation at the evaluation time points: obtaining the computing power resource occupancy rate of the analysis object and marking it as the occupancy value of the analysis object, summing up and averaging the occupancy values of all analysis objects to obtain the occupancy data at the evaluation time point; obtaining the occupancy analysis value ZF and the centralized analysis value JF of the computing task execution process through the occupancy data of all evaluation time points; and obtaining the evaluation coefficient PG of the computing task execution process by numerically calculating the occupancy analysis value ZF and the centralized analysis value JF.
3. The intelligent operation and maintenance control system applicable to heterogeneous large-scale computing centers according to claim 2 is characterized in that: At the end of the calculation task execution process, the occupancy data of all evaluation time points are summed and averaged to obtain the occupancy analysis value ZF, and the variance of the occupancy data of all evaluation time points is calculated to obtain the centralized analysis value JF.
4. The intelligent operation and maintenance control system applicable to heterogeneous large-scale computing centers according to claim 3 is characterized in that: The specific process of determining whether the computing resource occupancy status of the computing task execution process meets the requirements includes: retrieving the evaluation threshold PGmax through the database, and comparing the evaluation coefficient PG with the evaluation threshold PGmax: if the evaluation coefficient PG is less than the evaluation threshold PGmax, then it is determined that the computing resource occupancy status of the computing task execution process meets the requirements; if the evaluation coefficient PG is greater than or equal to the evaluation threshold PGmax, then it is determined that the computing resource occupancy status of the computing task execution process does not meet the requirements, and the computing task execution process is marked as an abnormal occupancy process.
5. The intelligent operation and maintenance control system applicable to heterogeneous large-scale computing centers according to claim 4 is characterized in that: The execution data ZX is the total amount of data processed during the execution of the computing task. The process of obtaining the parallel data BX includes: obtaining the amount of data processing completed by the analysis object during the execution of the computing task and marking it as the processing value of the analysis object, performing variance calculation on the processing values of all analysis objects to obtain the parallel data BX, and the energy consumption data NH is the sum of the energy consumption values of all analysis objects during the execution of the computing task.
6. The intelligent operation and maintenance control system applicable to heterogeneous large-scale computing centers according to claim 5 is characterized in that: The specific process of determining whether the data calculation processing status of the computing task execution process meets the requirements includes: retrieving the processing threshold CLmax through the database, and comparing the processing coefficient CL of the computing task execution process with the processing threshold CLmax: if the processing coefficient CL is less than the processing threshold CLmax, then it is determined that the data calculation processing status of the computing task execution process meets the requirements; if the processing coefficient CL is greater than or equal to the processing threshold CLmax, then it is determined that the data calculation processing status of the computing task execution process does not meet the requirements, and the corresponding computing task execution process is marked as a processing exception process.
7. The intelligent operation and maintenance control system applicable to heterogeneous large-scale computing centers according to claim 6 is characterized in that: The number of marks of the abnormal occupation process and the number of marks of the abnormal processing process are marked as abnormal occupation data ZY and abnormal processing data CY respectively. The operation and maintenance coefficient YW of the evaluation cycle is obtained by numerically calculating the abnormal occupation data ZY and the abnormal processing data CY.
8. The intelligent operation and maintenance control system applicable to heterogeneous large-scale computing centers according to claim 7 is characterized in that: The specific process of optimizing and analyzing the heterogeneous large model computing center includes: marking the computing task execution process that is marked as both the occupation exception process and the processing exception process as the overlapping process, marking the ratio of the number of overlapping processes to the number of computing task execution processes within the evaluation period as the overlapping coefficient, retrieving the overlapping threshold through the database, and comparing the overlapping coefficient with the overlapping threshold: if the overlapping coefficient is less than the overlapping threshold, a hardware optimization signal is generated and sent to the administrator's mobile terminal; if the overlapping coefficient is greater than or equal to the overlapping threshold, a node configuration optimization signal is generated and sent to the administrator's mobile terminal.
Citation Information
Patent Citations
An automated operation and maintenance management system for university cloud computing centers
CN110502390B
Computing power pooling system for improving GPU (Graphics Processing Unit) utilization efficiency
CN115202836A
Online service computing power optimization method and device based on cloud computing
CN117667602A
Operation and maintenance management monitoring method and device
CN118503043A
Electrical equipment operation and maintenance management system based on power supply and distribution project
CN118657339A