Method and device for intelligently diagnosing operation faults of computing power resources of intelligent computing center

By using pre-trained deep learning models to analyze computing power resource fault data in the intelligent computing center, the problem of low efficiency in computing power resource fault diagnosis is solved, and efficient and accurate fault detection and processing is achieved to ensure system stability.

CN120508427APending Publication Date: 2025-08-19DATACANVAS LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510622781.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

In the prior art, the computing power resources of intelligent computing centers are inefficient in fault diagnosis, and traditional methods rely on manual analysis or rule engines to efficiently process fault data.

Method used

The pre-trained deep learning model is used to automatically analyze the fault data of computing power resources, obtain fault repair suggestions and fault causes, reduce manual intervention, and improve diagnostic efficiency.

Benefits of technology

Automatically analyze computing resource failure data through deep learning models, reduce human judgment errors, improve fault diagnosis efficiency and quality, promptly discover potential problems, avoid problem expansion, and ensure stable operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508427A_ABST
    Figure CN120508427A_ABST
Patent Text Reader

Abstract

The invention provides a computing power resource operation fault intelligent diagnosis method and device of an intelligent computing center, and relates to the technical field of intelligent computing centers, intelligent computing centers and computing power infrastructures, and the method comprises the steps: S1, obtaining fault data of computing power resources in an operation process; step S2, analyzing the fault data by using a pre-trained deep learning model to obtain fault analysis information, the fault analysis information comprising fault repair suggestions and fault causes of the fault data; and S3, outputting the fault analysis information. According to the invention, the fault diagnosis efficiency and the computing power resource operation fault diagnosis operation and maintenance efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent computing centers, smart computing centers and computing power infrastructure, and specifically to a method and device for intelligent diagnosis of computing power resource operation failures in intelligent computing centers. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged.

[0003] An "Intelligent Computing Center" is a facility that uses large-scale heterogeneous computing resources, including general-purpose and intelligent computing power, to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model training, and model inference). The Intelligent Computing Center encompasses facilities, hardware, and software, and provides a full stack of capabilities, from bottom-level computing power to top-level application enablement.

[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.

[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform certain computing needs. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] Since the emergence of intelligent computing centers, improving the speed of computing resource fault diagnosis has been a pressing issue. Currently, when computing resources in intelligent computing centers experience faults, traditional methods rely on manual analysis or rule engines, making it difficult to efficiently process fault data. This results in very low computing resource operation fault diagnosis efficiency. Summary of the Invention

[0008] The embodiments of the present invention provide a method and device for intelligent diagnosis of computing power resource operation faults in an intelligent computing center, so as to solve the problem of low efficiency in diagnosing computing power resource operation faults in the prior art.

[0009] To solve the above problems, the present invention is achieved as follows:

[0010] In a first aspect, an embodiment of the present invention provides a method for intelligently diagnosing operational faults in computing resources of an intelligent computing center, comprising:

[0011] Step S1: Obtaining fault data of computing resources during operation;

[0012] Step S2: Analyze the fault data using a pre-trained deep learning model to obtain fault analysis information, where the fault analysis information includes fault repair suggestions and fault causes for the fault data;

[0013] Step S3: output the fault analysis information.

[0014] In one embodiment, the method further comprises:

[0015] Step S4: Acquire operation data of the computing resource, where the operation data of the computing resource includes operation information of the computing resource in the first time period;

[0016] Step S5: Using the pre-trained deep learning model, perform fault analysis on the operating data of the computing resource to obtain fault analysis information, where the fault analysis information is used to characterize fault information of the computing resource in a second time period, where the second time period is a time period after the first time period.

[0017] Step S6: Output the fault analysis information.

[0018] In one embodiment, before step S2, the method further includes:

[0019] Step S7: Obtain a first data set, where the first data set includes multiple data groups of the computing resources, each data group including a historical fault data and a set of corresponding label information, wherein the label information includes a fault repair suggestion and a fault cause for the corresponding historical fault data, wherein different data groups in the multiple data groups have different historical fault data;

[0020] Step S8: Train the initial deep learning model based on the first data set to obtain the pre-trained deep learning model.

[0021] In one embodiment, step S8 includes:

[0022] Step S81: Divide the first data set into a training data set and a validation data set, wherein the training data set includes a plurality of the data groups, and the validation data set includes a plurality of the data groups;

[0023] Step S82: training the initial deep learning model based on the training data set to obtain a trained deep learning model;

[0024] Step S83: verifying the trained deep learning model based on the verification data set to obtain a verification result;

[0025] Step S84: When the verification result indicates that the accuracy of the trained deep learning model in analyzing the historical fault data in the verification data set is greater than or equal to a preset threshold, determine that the trained deep learning model is consistent with the pre-trained deep learning model.

[0026] In one embodiment, the fault analysis information also includes urgency information of the fault data, and the urgency information includes an urgency level, an impact range, and a response requirement. The urgency level is used to indicate the processing priority of the fault data, the impact range includes functional impact and potential risks, and the response requirement is used to indicate the repair time limit of the fault data.

[0027] In one embodiment, the fault cause includes at least one of the following: network fault information, hardware fault information, and software fault information.

[0028] In a second aspect, an embodiment of the present invention further provides an intelligent diagnostic device for computing resource operation failures in an intelligent computing center, comprising:

[0029] The acquisition module is used to obtain fault data of computing resources during operation;

[0030] An analysis module, configured to analyze the fault data using a pre-trained deep learning model to obtain fault analysis information, wherein the fault analysis information includes a fault repair suggestion and a fault cause for the fault data;

[0031] An output module is used to output the fault analysis information.

[0032] In a third aspect, the present invention also provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the intelligent diagnosis method for operation failures of computing resources in an intelligent computing center as described in the first aspect above are implemented.

[0033] In a fourth aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the intelligent diagnosis method for computing power resource operation failures of an intelligent computing center as described in the first aspect above are implemented.

[0034] In a fifth aspect, the present invention further provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps of the intelligent diagnosis method for operation failures of computing resources in an intelligent computing center as described in the first aspect above.

[0035] In an embodiment of the present invention, fault data generated by computing power resources during operation is obtained; the fault data is analyzed using a pre-trained deep learning model to obtain fault analysis information, wherein the fault analysis information includes fault repair suggestions and fault causes for the fault data; and the fault analysis information is output. In this way, the fault data generated by computing power resources during operation is automatically analyzed using a deep learning model, thereby reducing the need for human intervention, avoiding human misjudgment, and improving the efficiency of fault diagnosis and the quality of fault handling. At the same time, potential problems can be quickly detected in the early stages of a fault and promptly notified to operation and maintenance personnel, thereby preventing the problem from expanding or affecting the normal operation of the system, thereby improving the efficiency of fault diagnosis for computing power resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0037] Figure 1 This is a flow chart of a method for intelligent diagnosis of computing resource operation failures in an intelligent computing center provided by an embodiment of the present invention;

[0038] Figure 2 This is a schematic diagram of an interface for intelligent diagnosis of computing resource operation failures in an intelligent computing center provided by an embodiment of the present invention;

[0039] Figure 3 This is a structural diagram of an intelligent diagnostic device for computing resource operation failures in an intelligent computing center provided by an embodiment of the present invention;

[0040] Figure 4 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0042] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data centers to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0043] The "computing power" (Computational Power, CP) mentioned in the present invention refers to: the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP=CP 通用 +CP 智能 +CP 超级 .

[0044] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capability of the computing power facility, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.

[0045] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in terms of data storage capacity, performance, security and reliability, and environmental friendliness. It is a comprehensive indicator for measuring a data center's data storage capacity, encompassing both external storage devices such as storage arrays and internal server storage. Storage capacity is commonly measured in exabytes (EB, 1EB = 2^60 bytes), while performance is commonly measured in IOPS / TB (Input / Output Operations Per Second / TB). Disaster recovery ratio is a key indicator of security and reliability.

[0046] The "computing power infrastructure" mentioned in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized calculation, storage, transmission and application of information.

[0047] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0048] The "computing power" mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.

[0049] The "general computing power" mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0050] The "intelligent computing power" mentioned in this invention refers to: a computing platform based on specialized chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various innovative artificial intelligence applications, such as natural language processing and machine vision.

[0051] The "supercomputing power" mentioned in the present invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, genetic analysis, etc.

[0052] The "intelligent computing center" described in this article refers to a facility that provides the computing power, data, and algorithms required for artificial intelligence applications (such as AI deep learning model development, model training, and model inference) by utilizing large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center encompasses facilities, hardware, and software, and can provide a full stack of capabilities, from bottom-level computing power to top-level application enablement.

[0053] The "intelligent computing center" mentioned in the present invention includes but is not limited to the "intelligent computing center".

[0054] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.

[0055] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0056] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide large-scale computing, storage and network services and other functions, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.

[0057] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.

[0058] In the prior art, since the emergence of intelligent computing centers, how to improve the fault diagnosis speed of computing resources has been an urgent problem to be solved. At present, when there is a fault in the computing resources of the intelligent computing center, the traditional method relies on manual analysis or rule engines, which makes it difficult to efficiently process fault data, resulting in a very low efficiency in fault diagnosis of computing resources. In an embodiment of the present invention, a deep learning model is used to automatically analyze the fault data generated by computing resources during operation, thereby reducing the need for manual intervention, avoiding human misjudgment, improving the efficiency of fault diagnosis and the quality of fault handling, and at the same time, it can quickly detect potential problems in the early stages of a fault and promptly notify the operation and maintenance personnel, which can prevent the problem from expanding or affecting the normal operation of the system, thereby improving the efficiency of fault diagnosis of computing resources.

[0059] For details, see Figure 1 , Figure 1 This is a flow chart of an intelligent diagnosis method for computing resource operation failures in an intelligent computing center provided by an embodiment of the present invention. Figure 1 As shown, the following steps are included:

[0060] Step S1: Obtain fault data of computing resources during operation.

[0061] In this step, the above-mentioned fault data can be abnormal information generated during the system operation, which can specifically include monitoring data (such as CPU usage, memory usage), log files (such as error logs, operation records), historical fault records (such as past fault types, repair solutions), etc.

[0062] Step S2: Analyze the fault data using a pre-trained deep learning model to obtain fault analysis information, where the fault analysis information includes fault repair suggestions and fault causes for the fault data.

[0063] In this step, the pre-trained deep learning model can be a neural network model trained with historical data, which can automatically learn the characteristics of fault data.

[0064] The above-mentioned fault repair suggestions can be solutions inferred by the pre-trained deep learning model based on the fault data. For example, the system may recommend a specific repair step or configuration adjustment operation based on past experience. This decision support can reduce the decision-making burden of operation and maintenance personnel and improve the accuracy of fault repair.

[0065] The above-mentioned fault cause can be the root cause inferred by the pre-trained deep learning model based on the fault data and determined through multi-dimensional analysis, such as hardware aging, software conflict, network delay and other problems.

[0066] Example causes include container misconfiguration; insufficient resources: Pods terminated due to GPU resource allocation failures or insufficient node resources (memory / CPU); dependent service anomalies: GPU driver installation errors, node GPU hardware compatibility issues, or Kubernetes API communication failures; image defects: container image corruption or version incompatibility with the current cluster environment. Troubleshooting recommendations include reviewing container logs; verifying GPU resource status: checking the node GPU driver for proper operation and confirming node labels; checking DaemonSet configuration: verifying volume mounts and environment variables for cluster compatibility; and rolling back or updating images.

[0067] Step S3: output the fault analysis information.

[0068] In this step, the output of the fault analysis information may be to present the fault repair suggestions and fault causes to the operation and maintenance personnel in an actionable form, such as alarm information, repair suggestion reports, or automatically executed repair instructions.

[0069] Specifically, an intelligent early warning mechanism can automatically trigger an alarm and notify operation and maintenance personnel via SMS, email or monitoring platform.

[0070] In an embodiment of the present invention, fault data generated by computing power resources during operation is obtained; the fault data is analyzed using a pre-trained deep learning model to obtain fault analysis information, wherein the fault analysis information includes fault repair suggestions and fault causes for the fault data; and the fault analysis information is output. In this way, the fault data generated by computing power resources during operation is automatically analyzed using a deep learning model, thereby reducing the need for human intervention, avoiding human misjudgment, and improving the efficiency of fault diagnosis and the quality of fault handling. At the same time, potential problems can be quickly detected in the early stages of a fault and promptly notified to operation and maintenance personnel, thereby preventing the problem from expanding or affecting the normal operation of the system, thereby improving the efficiency of fault diagnosis for computing power resources.

[0071] In one embodiment, the method further comprises:

[0072] Step S4: Acquire operation data of the computing resource, where the operation data of the computing resource includes operation information of the computing resource in the first time period;

[0073] Step S5: Using the pre-trained deep learning model, perform fault analysis on the operating data of the computing resource to obtain fault analysis information, where the fault analysis information is used to characterize fault information of the computing resource in a second time period, where the second time period is a time period after the first time period.

[0074] Step S6: Output the fault analysis information.

[0075] Specifically, the first time period may be the starting time range of the current fault analysis, such as the past hour, day, etc. The operation information may include but is not limited to monitoring data, log data, and configuration information, wherein the monitoring data may include CPU usage, memory usage, network latency, etc., the log data may include system logs, error logs, operation records, etc., and the configuration information may include hardware parameters, software versions, network topology, etc.

[0076] The above-mentioned fault analysis information may be the analysis results of the pre-trained deep learning model on the computing power resource failures that may occur in the second time period. Specifically, it may include the fault type, fault cause, probability of occurrence, development trend, and potential impacts such as service interruption and resource exhaustion. The above-mentioned second time period may refer to the time range after the first time period, such as the next hour or day.

[0077] Exemplarily, after analyzing the operating data of the first time period, the pre-trained deep learning model analyzes the possible memory overflow failure in the second time period, and the failure probability is 85%, which may affect service availability.

[0078] The above-mentioned output of the fault analysis information may be presenting the fault analysis information in an actionable form, such as alarm information, repair suggestions and visual reports, etc., wherein the visual report may include a fault probability trend graph, impact range analysis, etc.

[0079] In the above embodiment, by using the pre-trained deep learning model to perform fault analysis on the operating data of computing power resources, it is possible to analyze possible faults that may occur in the future, so that the operation and maintenance team can take preventive measures in advance to reduce the probability of faults, thereby reducing the system's downtime and losses, and further improving the efficiency of computing power resource operation fault diagnosis.

[0080] In one embodiment, before step S2, the method further includes:

[0081] Step S7: Obtain a first data set, where the first data set includes multiple data groups of the computing resources, each data group including a historical fault data and a set of corresponding label information, wherein the label information includes a fault repair suggestion and a fault cause for the corresponding historical fault data, wherein different data groups in the multiple data groups have different historical fault data;

[0082] Step S8: Train the initial deep learning model based on the first data set to obtain the pre-trained deep learning model.

[0083] Specifically, the first dataset may be a collection of historical fault data used to train a deep learning model. Each of the multiple datasets consists of historical fault data and corresponding label information, which is used to train the deep learning model to determine fault repair recommendations and fault causes. The label information may be manually annotated or automatically extracted metadata for the historical fault data.

[0084] In the above embodiment, by using historical fault data and label information to train the deep learning model, the deep learning model can learn from a large amount of historical data, achieve self-optimization, and adapt to the ever-changing system environment, thereby improving the accuracy of the deep learning model for fault diagnosis.

[0085] In one embodiment, step S8 includes:

[0086] Step S81: Divide the first data set into a training data set and a validation data set, wherein the training data set includes a plurality of the data groups, and the validation data set includes a plurality of the data groups;

[0087] Step S82: training the initial deep learning model based on the training data set to obtain a trained deep learning model;

[0088] Step S83: verifying the trained deep learning model based on the verification data set to obtain a verification result;

[0089] Step S84: When the verification result indicates that the accuracy of the trained deep learning model in analyzing the historical fault data in the verification data set is greater than or equal to a preset threshold, determine that the trained deep learning model is consistent with the pre-trained deep learning model.

[0090] Specifically, the first data set may be divided into a training data set and a validation data set according to a preset ratio. For example, if the first data set contains 100 data groups, the training data set may include 80 data groups and the validation data set may include 20 data groups.

[0091] The above-mentioned verification results may refer to the performance indicators of the trained deep learning model on the verification data set, such as the matching degree and accuracy of the model output results and the corresponding label information.

[0092] In the above embodiment, by dividing the deep learning model into a training data set and a verification data set for training, the generalization ability and reliability of the model can be improved. When the verification result indicates that the accuracy of the trained deep learning model in analyzing the historical fault data in the verification data set is greater than or equal to a preset threshold, it is determined that the trained deep learning model is consistent with the pre-trained deep learning model, thereby ensuring the accuracy of the deep learning model in actual scenarios, thereby further improving the accuracy of fault diagnosis.

[0093] In one embodiment, the fault analysis information also includes urgency information of the fault data, and the urgency information includes an urgency level, an impact range, and a response requirement. The urgency level is used to indicate the processing priority of the fault data, the impact range includes functional impact and potential risks, and the response requirement is used to indicate the repair time limit of the fault data.

[0094] Specifically, the urgency level can be a classification of fault severity, such as "high," "medium," or "low," indicating the priority of handling. The impact scope can describe the extent of the current fault's impact on the system, business, or resources, including functional impact and potential risks, such as "GPU-dependent service interruption (AI training / inference tasks)." The response requirement can be a time limit for fault repair, such as "intervention within 1 hour to prevent cluster GPU unavailability due to plug-in failure."

[0095] For example, Figure 2 This is a schematic diagram of an interface for intelligent diagnosis of computing resource operation failures in an intelligent computing center provided by an embodiment of the present invention. Figure 2As shown, the fault analysis information also includes urgency information of the fault data, and the urgency information specifically includes urgency level, impact range and response requirements.

[0096] In the above embodiment, the order of fault handling can be clarified by the emergency level, and the functional impact and potential risks can be combined to avoid missing key issues. The repair time limit can be set by responding to requirements to ensure that the fault is resolved in a timely manner, thereby improving the stability of computing resource operation and the efficiency of fault diagnosis.

[0097] In one embodiment, the fault cause includes at least one of the following: network fault information, hardware fault information, and software fault information.

[0098] Specifically, the aforementioned network fault information may be network-related issues that cause abnormal computing resource operation, typically involving abnormalities in network communication links, equipment, or configuration. The aforementioned hardware fault information may be physical device or resource issues that cause abnormal computing resource operation, typically involving overloaded hardware components or resources. The aforementioned software fault information may be software system-related issues that cause abnormal computing resource operation, typically involving abnormalities in applications, operating systems, middleware, or dependent components.

[0099] For example, the cause of the failure corresponding to the degradation of system performance may be a hardware failure, or may be the result of multiple factors such as a network problem and an application layer error (bug).

[0100] In the above embodiment, the deep learning model can attribute faults from a multi-dimensional and global perspective, perform hierarchical analysis of faults, and analyze them layer by layer from multiple dimensions such as hardware, software, and network to find the deepest causes of faults, avoiding the one-sidedness that may occur in traditional analysis methods, thereby improving the comprehensiveness and accuracy of fault diagnosis.

[0101] See Figure 3 , Figure 3 This is a structural diagram of an intelligent diagnostic device for computing resource operation failures in an intelligent computing center provided by an embodiment of the present invention. Figure 3 As shown, the intelligent computing center's computing power resource operation fault intelligent diagnosis device 300 includes:

[0102] The first acquisition module 301 is used to obtain fault data of computing resources during operation;

[0103] An analysis module 302 is configured to analyze the fault data using a pre-trained deep learning model to obtain fault analysis information, wherein the fault analysis information includes a fault repair suggestion and a fault cause for the fault data;

[0104] The first output module 303 is configured to output the fault analysis information.

[0105] In one embodiment, the intelligent computing center computing resource operation fault intelligent diagnosis device 300 further includes:

[0106] A second acquisition module is configured to acquire operation data of the computing resource, where the operation data of the computing resource includes operation information of the computing resource in a first time period;

[0107] an analysis module, configured to perform fault analysis on the operating data of the computing resource using the pre-trained deep learning model to obtain fault analysis information, wherein the fault analysis information is used to characterize fault information of the computing resource in a second time period, where the second time period is a time period after the first time period;

[0108] The second output module is used to output the fault analysis information.

[0109] In one embodiment, the intelligent computing center computing resource operation fault intelligent diagnosis device 300 further includes:

[0110] a third acquisition module, configured to acquire a first data set, the first data set comprising multiple data groups of the computing resources, the data groups comprising a historical fault data and a set of corresponding label information, the label information comprising a fault repair suggestion and a fault cause for the corresponding historical fault data, wherein different data groups in the multiple data groups have different historical fault data;

[0111] A training module is used to train the initial deep learning model based on the first data set to obtain the pre-trained deep learning model.

[0112] In one embodiment, the prediction module includes:

[0113] a division unit, configured to divide the first data set into a training data set and a validation data set, wherein the training data set includes a plurality of the data groups, and the validation data set includes a plurality of the data groups;

[0114] A training unit, configured to train the initial deep learning model based on the training data set to obtain a trained deep learning model;

[0115] A verification unit, configured to verify the trained deep learning model based on the verification data set to obtain a verification result;

[0116] A determination unit is used to determine that the trained deep learning model conforms to the pre-trained deep learning model when the verification result indicates that the accuracy of the trained deep learning model in analyzing the historical fault data in the verification data set is greater than or equal to a preset threshold.

[0117] In one embodiment, the fault analysis information also includes urgency information of the fault data, and the urgency information includes an urgency level, an impact range, and a response requirement. The urgency level is used to indicate the processing priority of the fault data, the impact range includes functional impact and potential risks, and the response requirement is used to indicate the repair time limit of the fault data.

[0118] In one embodiment, the fault cause includes at least one of the following: network fault information, hardware fault information, and software fault information.

[0119] The intelligent diagnosis device for computing power resource operation failures of an intelligent computing center provided by an embodiment of the present invention is capable of implementing the various processes of each embodiment of the intelligent diagnosis method for computing power resource operation failures of the above-mentioned intelligent computing center. The technical features correspond one to one and can achieve the same technical effects. To avoid repetition, they will not be described here.

[0120] It should be noted that the intelligent diagnosis device for computing power resource operation failures in the embodiment of the present invention can be a device, or a component, integrated circuit, or chip in an electronic device.

[0121] The embodiment of the present invention further provides an electronic device, see Figure 4 , Figure 4 The electronic device includes a memory 401, a processor 402, and a program or instruction stored in the memory 401 and executed by the processor 402. Figure 1 Any steps in the embodiment of the intelligent diagnosis method for computing power resource operation failures of the corresponding intelligent computing center and the same beneficial effects are achieved will not be repeated here.

[0122] The processor 402 may be a CPU, an ASIC, an FPGA, or a GPU.

[0123] Those skilled in the art will understand that all or part of the steps of the embodiment of the intelligent diagnosis method for computing power resource operation failures of the above-mentioned intelligent computing center can be completed through hardware related to program instructions, and the program can be stored in a readable medium.

[0124] The embodiment of the present invention further provides a readable storage medium, on which a computer program is stored, which can realize the above-mentioned Figure 1The corresponding steps in the embodiment of the intelligent diagnosis method for computing power resource operation failure of the intelligent computing center can achieve the same technical effect. To avoid repetition, they are not repeated here. The storage medium is such as read-only memory (ROM), random access memory (RAM), disk or optical disk.

[0125] The present invention also provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the above Figure 1 The computing power resources of the corresponding intelligent computing center run the various processes of the embodiment of the fault intelligent diagnosis method, and can achieve the same technical effect. To avoid repetition, they will not be repeated here.

[0126] The terms "first", "second" etc. in the embodiments of the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, the process, method, system, product or equipment comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or equipment. In addition, "and / or" is used in this application to represent at least one of the connected objects, for example A and / or B and / or C, which means comprising 7 situations including single A, single B, single C, and both A and B exist, both B and C exist, both A and C exist, and both A, B and C exist.

[0127] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0128] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of each embodiment of the present application.

[0129] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. An intelligent diagnosis method for computing resource operation failures in an intelligent computing center, characterized in that: include: Step S1: Obtaining fault data of computing resources during operation; Step S2: Analyze the fault data using a pre-trained deep learning model to obtain fault analysis information, where the fault analysis information includes fault repair suggestions and fault causes for the fault data; Step S3: output the fault analysis information.

2. The method according to claim 1, wherein The method further comprises: Step S4: Acquire operation data of the computing resource, where the operation data of the computing resource includes operation information of the computing resource in the first time period; Step S5: Using the pre-trained deep learning model, perform fault analysis on the operating data of the computing resource to obtain fault analysis information, where the fault analysis information is used to characterize fault information of the computing resource in a second time period, where the second time period is a time period after the first time period. Step S6: Output the fault analysis information.

3. The method according to claim 1, wherein Before step S2, the method further includes: Step S7: Obtain a first data set, where the first data set includes multiple data groups of the computing resources, each data group including a historical fault data and a set of corresponding label information, wherein the label information includes a fault repair suggestion and a fault cause for the corresponding historical fault data, wherein different data groups in the multiple data groups have different historical fault data; Step S8: Train the initial deep learning model based on the first data set to obtain the pre-trained deep learning model.

4. The method according to claim 3, wherein The step S8 comprises: Step S81: Divide the first data set into a training data set and a validation data set, wherein the training data set includes a plurality of the data groups, and the validation data set includes a plurality of the data groups; Step S82: training the initial deep learning model based on the training data set to obtain a trained deep learning model; Step S83: verifying the trained deep learning model based on the verification data set to obtain a verification result; Step S84: When the verification result indicates that the accuracy of the trained deep learning model in analyzing the historical fault data in the verification data set is greater than or equal to a preset threshold, determine that the trained deep learning model is consistent with the pre-trained deep learning model.

5. The method according to any one of claims 1 to 4, characterized in that The fault analysis information also includes urgency information of the fault data, which includes an urgency level, an impact range, and a response requirement. The urgency level is used to indicate the processing priority of the fault data, the impact range includes functional impact and potential risks, and the response requirement is used to indicate the repair time limit for the fault data.

6. The method according to any one of claims 1 to 4, characterized in that The fault cause includes at least one of the following: network fault information, hardware fault information, and software fault information.

7. An intelligent diagnostic device for computing resource operation failures in an intelligent computing center, characterized in that: include: The acquisition module is used to obtain fault data of computing resources during operation; An analysis module, configured to analyze the fault data using a pre-trained deep learning model to obtain fault analysis information, wherein the fault analysis information includes a fault repair suggestion and a fault cause for the fault data; An output module is used to output the fault analysis information.

8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the intelligent diagnosis method for computing power resource operation failures of an intelligent computing center as described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the intelligent diagnosis method for computing power resource operation failures of an intelligent computing center according to any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the intelligent diagnosis method for computing power resource operation failures of an intelligent computing center as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Method and system for identifying defects in metrology equipment

    CN122547628A