Equipment anomaly analysis method and related device

By using training exception points and resource load slices in GPU computing power hardware detection, the abnormal equipment is quickly screened through simulated operation process, solving the problems of long and high testing in the existing technology, and achieving rapid equipment flow and resource efficiency improvement.

CN120029863APending Publication Date: 2025-05-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311585524.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When detecting GPU computing power hardware, the test takes a long time, affects the rapid flow of equipment, and is costly, making it difficult to effectively screen abnormal equipment.

Method used

By obtaining the exception points of the first device during the training of the AI ​​model as the training exception points, a resource load slice and program slice with a minute limit for these exception points is constructed, and the running process of these slices is simulated on the second device to quickly filter out the exception devices.

Benefits of technology

It realizes rapid screening of abnormal equipment before the normal testing process, timely release and repair, improve computing resource efficiency, and reduce costs through short-term pre-screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029863A_ABST
    Figure CN120029863A_ABST
Patent Text Reader

Abstract

The invention discloses an equipment anomaly analysis method and a related device, which are applied to the field of artificial intelligence. According to the method provided by the invention, the abnormal point of the first equipment during artificial intelligence AI model training can be acquired as the training abnormal point, and the resource load slice and the program slice of the minute time limit are configured for the training abnormal point; the running process of the resource load slice and the program slice is simulated at the training abnormal point of the second equipment, whether the running process is abnormal or not is analyzed, and when running abnormity occurs in the running process, the second equipment can be determined to be abnormal equipment. Through the above mode, before a normal test process is carried out on the equipment, the training abnormal points of other equipment are firstly used, the operation process of the training abnormal points of the to-be-tested equipment is simulated through the resource load slices and the program slices of the minute time limit, the abnormal equipment is rapidly screened, timely put and maintained, and the computing power resource efficiency is improved; and the cost can be reduced and the efficiency can be increased through short-time-efficiency pre-screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a method for analyzing device anomalies and related devices. Background Art

[0002] With the rise of large artificial intelligence (AI) models, the scale of demand for heterogeneous computing power has also shown explosive growth. However, since the hardware production process of heterogeneous computing power itself cannot achieve a 100% yield rate, pre-delivery functional and performance testing is required before computing power resources are put into production environments.

[0003] The current normal testing process for the hardware detection of the graphics processing unit (GPU) computing power of the AI ​​large model of the device to be tested is to run the existing stress test program in the community, or to build a stress test program for simulated business operations. The program runs for days to test whether the card is running normally. However, the execution cycle is long, which is not conducive to the rapid circulation of equipment, and the cost of GPU computing power is high. Summary of the invention

[0004] The embodiments of the present application provide a method for analyzing equipment anomalies and related devices, which are used to quickly screen abnormal equipment before the normal testing process, promptly carry out maintenance, and improve the efficiency of computing resources. The short-term pre-screening can reduce costs and increase efficiency.

[0005] In view of this, the present application provides, on one hand, a method for analyzing device abnormalities, comprising:

[0006] Obtaining training anomaly points, where the training anomaly points include anomaly points of the first device during artificial intelligence (AI) model training;

[0007] Construct resource load slices and program slices for training anomalies. Resource load slices and program slices are slices with a time limit of minutes.

[0008] At the training abnormal point of the second device, the operation process is simulated according to the resource load slice and program slice, and the second device is the device to be tested for computing power hardware;

[0009] When the second device operates abnormally, the second device is determined to be an abnormal device.

[0010] In view of this, the present application provides, on the other hand, a device for analyzing device abnormalities, comprising:

[0011] An acquisition unit, used to acquire training abnormal points, where the training abnormal points include abnormal points of the first device during AI model training;

[0012] A construction unit is used to construct resource load slices and program slices for training abnormal points, where the resource load slices and program slices are slices with a time limit of minutes;

[0013] A simulation unit, used for simulating the running process according to resource load slicing and program slicing at a training abnormal point of a second device, where the second device is a device to be tested for computing power hardware;

[0014] The determination unit is used to determine that the second device is an abnormal device when the second device operates abnormally.

[0015] In one possible design, in another implementation of another aspect of the embodiment of the present application, the training anomaly point includes a plurality of first anomaly points, each first anomaly point corresponding to a resource load slice and a program slice;

[0016] The simulation unit is specifically used for:

[0017] Read a first abnormal point in sequence;

[0018] At the first abnormal point of the second device, the operation process is simulated according to the resource load slices and program slices corresponding to the first abnormal point.

[0019] In one possible design, in another implementation of another aspect of the embodiment of the present application, the resource load slice includes resource load, graphics processor GPU card temperature, GPU utilization and operator usage set, and the program slice includes neural network, training data, program execution action and kernel state slice;

[0020] The simulation unit is specifically used for:

[0021] At the training anomaly point of the second device, simulate the load usage based on the resource load, GPU card temperature, GPU utilization, and operator usage set;

[0022] The training program running process is simulated based on the neural network, training data, program execution actions and kernel state slices.

[0023] In a possible design, in another implementation of another aspect of the embodiment of the present application, the simulation unit is specifically used for:

[0024] At the training abnormal point of the second device, the resource load is simulated by increasing the amount of computing data;

[0025] Simulate GPU card temperature by adjusting fan speed;

[0026] Simulate GPU utilization by overlaying training data;

[0027] Reference the heterogeneous computing ecosystem library to obtain operators and use collections for simulation.

[0028] In a possible design, in another implementation of another aspect of the embodiment of the present application, the simulation unit is specifically used for:

[0029] Load the neural network and training data, and execute the training process;

[0030] The training process is completed according to program execution actions and kernel state slices.

[0031] In a possible design, in another implementation of another aspect of the embodiment of the present application, the device further includes a recording unit, and the recording unit is specifically used to:

[0032] The training anomalies, resource load slices and program slices are recorded in the anomaly repository, which uses a plug-in mode to save information.

[0033] Another aspect of the present application provides a computer device, comprising:

[0034] memories, transceivers, processors, and bus systems;

[0035] Wherein, the memory is used to store programs;

[0036] The processor is used to execute the program in the memory, including executing the above-mentioned methods;

[0037] The bus system is used to connect the memory and the processor so that the memory and the processor can communicate with each other.

[0038] Another aspect of the present application provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer is enabled to execute the above-mentioned methods.

[0039] Another aspect of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided by the above aspects.

[0040] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0041] The embodiment of the present application can obtain the abnormal points of the first device during AI model training as training abnormal points, configure resource load slices and program slices with a time limit of minutes for the training abnormal points, and simulate the operation process of resource load slices and program slices at the training abnormal points of the second device before the second device performs computing power hardware detection, and analyze whether the operation process is abnormal. When an operation abnormality occurs during the operation process, it can be determined that the second device is an abnormal device. Through the above method, before the normal test process of the device is carried out, the training abnormal points of other devices are used first, and the operation process of the training abnormal points of the device to be tested is simulated through resource load slices and program slices with a time limit of minutes, so as to quickly screen abnormal devices, put them into maintenance in time, and improve the efficiency of computing power resources. In addition, the short-term pre-screening can reduce costs and increase efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 Schematic diagram of the architecture of the abnormality analysis system in the embodiment of the present application;

[0043] Figure 2 A schematic diagram of a process flow of a device abnormality analysis method in an embodiment of the present application;

[0044] Figure 3 This is a schematic diagram of a pre-screening in an embodiment of the present application;

[0045] Figure 4 A schematic diagram of a pre-screening architecture in an embodiment of the present application;

[0046] Figure 5 A schematic diagram of a process of pre-screening anomalies in an embodiment of the present application;

[0047] Figure 6 This is a schematic diagram of the structure of a device abnormality analysis device in an embodiment of the present application;

[0048] Figure 7 A schematic diagram of the structure of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0049] The embodiments of the present application provide a method for analyzing equipment anomalies and related devices, which are used to quickly screen abnormal equipment before the normal testing process, promptly carry out maintenance, and improve the efficiency of computing resources. The short-term pre-screening can reduce costs and increase efficiency.

[0050] In the description and claims of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0051] The term "exemplary" as used herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein does not necessarily have to be construed as superior to or better than other embodiments.

[0052] In the embodiments of this application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.

[0053] In addition, for a better description of this application, numerous specific details are given in the following specific implementation manners. Those skilled in the art should understand that this application can still be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail in order to highlight the gist of this application.

[0054] With the rise of large artificial intelligence (AI) models, the scale of heterogeneous computing power required has also shown explosive growth. However, since the hardware production process of heterogeneous computing power itself cannot achieve a 100% yield rate, before the computing power resources are put into the production environment for use, functional and performance tests need to be carried out before delivery.

[0055] The traditional method is to run the existing stress test programs in the community or build stress test programs that simulate the operation of the business, and detect whether the card runs normally through the program running at the level of days. However, the execution time-consuming cycle is long, which is not conducive to the rapid turnover of equipment, and the cost of graphic processing unit (GPU) computing power is high.

[0056] Based on this, the embodiment of the present application provides a device anomaly analysis method, which obtains the anomaly points of the first device during AI model training as training anomaly points, configures resource load slices and program slices with a time limit of minutes for the training anomaly points, and simulates the operation process of resource load slices and program slices at the training anomaly points of the second device before the computing power hardware detection is performed on the second device, and analyzes whether the operation process is abnormal. When an operation abnormality occurs during the operation process, it can be determined that the second device is an abnormal device. Through the above method, before the normal test process of the device is carried out, the training anomaly points of other devices are used first, and the operation process of the training anomaly points of the device to be tested is simulated through resource load slices and program slices with a time limit of minutes, so as to quickly screen abnormal devices, put them into maintenance in time, and improve the efficiency of computing power resources. In addition, the short-term pre-screening can reduce costs and increase efficiency.

[0057] The embodiments of the present application are applied to the field of AI. Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0058] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models are also called large models and basic models. After fine-tuning, they can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0059] The AI ​​model training process of the embodiment of the present application is an application of machine learning (ML) / deep learning. Machine learning is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithmic complexity theory and other disciplines. It specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning generally include artificial neural networks (ANN), belief networks, reinforcement learning, transfer learning, inductive learning, and learning by teaching. The pre-trained model is the latest development in deep learning, integrating the above technologies.

[0060] The abnormal database in the embodiment of the present application is an application of a database, which can be simply regarded as an electronic file cabinet - a place to store electronic files, where users can add, query, update, delete, etc. the data in the files. The so-called "database" is a collection of data that is stored together in a certain way, can be shared with multiple users, has as little redundancy as possible, and is independent of the application program.

[0061] A database management system (DBMS) is a computer software system designed to manage databases, generally with basic functions such as storage, retrieval, security, and backup. Database management systems can be classified according to the database model they support, such as relational, extensible markup language (XML); or according to the type of computer they support, such as server clusters, mobile phones; or according to the query language used, such as structured query language (SQL), XQuery; or according to the performance focus, such as maximum scale, maximum operating speed; or other classification methods. Regardless of the classification method used, some DBMS can cross categories, for example, supporting multiple query languages ​​at the same time.

[0062] The device anomaly analysis method provided in the embodiment of the present application can be implemented by various electronic devices, for example, it can be implemented by a terminal device alone, or it can be implemented by a server and a terminal device in collaboration. For example, the terminal device performs the device anomaly analysis method described below alone, or the terminal device and the server jointly perform the device anomaly analysis method described below, the terminal device collects the anomaly points of multiple first devices during AI model training, and uses the anomaly points as an anomaly with a high probability of occurring in AI model training, and becomes a training anomaly point, and then sends the training anomaly point to the server, the server configures a resource load slice and a program slice with a time limit of minutes for each training anomaly point, and then sends the resource load slice and the program slice to the terminal device, the terminal device configures the resource load slice and the program slice to the second device to be detected, so that the second device simulates operation according to the resource load slice and the program slice at the corresponding training anomaly point, and the terminal device then confirms whether the second device is abnormal for the permission process, and determines that the second device is an abnormal device when an operation anomaly occurs.

[0063] The electronic device for device anomaly analysis provided in the embodiments of the present application may be various types of terminal devices or servers, wherein the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (context delivery networks, CDNs), and big data and artificial intelligence platforms; the terminal may be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.

[0064] Taking the server as an example, it can be a server cluster deployed in the cloud, opening artificial intelligence cloud services (AIaaS, AI as a Service) to objects. The AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service model is similar to an AI theme mall. All objects can access one or more artificial intelligence services provided by the AIaaS platform through an application programming interface.

[0065] The embodiments of the present application can also be applied to a separate terminal device. The terminal device collects abnormal points of multiple first devices during AI model training, configures resource load slices and program slices with a time limit of minutes for each training abnormal point, and then configures the resource load slices and program slices to the second device to be detected, so that the second device performs simulated operation according to the resource load slices and program slices at the corresponding training abnormal point. The terminal device then confirms whether the second device is abnormal based on the allowed process. If an operation abnormality occurs, it is determined that the second device is an abnormal device.

[0066] The following is an example of a server and a terminal device cooperating to implement the device abnormality analysis method provided in the embodiment of the present application. Figure 1 , Figure 1 The terminal device 11 is connected to the server 13 via the network 12 , which may be a wide area network or a local area network, or a combination of the two. The terminal device 11 is connected to the first device 14 and the second device 15 .

[0067] In some embodiments, the terminal device 11 collects the abnormal points of the first device 14 during the AI ​​model training as training abnormal points, and then sends the training abnormal points to the server 13. The first device 14 may include multiple devices (not shown in the figure). The server 13 configures resource load slices and program slices with a minute time limit based on the training abnormal points, and feeds back to the terminal device 11. The terminal device 11 configures the resource load slices and program slices to the second device 15 for simulated operation. When the second device 15 has an operational abnormality, it is determined that the second device 15 is an abnormal device.

[0068] The device anomaly analysis method provided in the embodiment of the present application will be described below in conjunction with the accompanying drawings. The executor of the following device anomaly analysis method takes the terminal device as an example, and can be specifically implemented by the terminal device by running the various computer programs mentioned above; of course, based on the understanding of the following text, it is not difficult to see that the device anomaly analysis method provided in the embodiment of the present application can also be implemented collaboratively by the terminal device and the server.

[0069] See also Figure 2 , Figure 2 The figure is a flow chart of a method for analyzing device abnormality provided by an embodiment of the present application, the method comprising:

[0070] Step 201. Obtain training anomaly points, where the training anomaly points include anomaly points of the first device during AI model training.

[0071] In one or more embodiments, the first device is a detected device, or a device in normal use. The first device may have an abnormality during the training of the AI ​​model. The abnormal operation position can be called an abnormal point. In the embodiment of the present application, the abnormal point can be called a training abnormal point, that is, the position where the abnormality occurs during the training of the AI ​​model. Among them, the way to judge the abnormality during the training of the AI ​​model can be determined from multiple aspects, such as abnormal fluctuations in resource performance, functional abnormalities in resource loads, and abnormalities in training programs. For abnormal fluctuations in resource performance, it is mainly executed from the perspective of utilization fluctuations. When the utilization fluctuation of the computing unit decreases by more than 10%, the video memory fluctuates by 50%, or the activity of the computing unit decreases by 10%, it is recorded as an abnormality in resource performance. For functional abnormalities in resource loads, resource loads include but are not limited to central processing unit (CPU) resources, memory resources, hard disk resources, etc., and functional abnormalities in resource loads can be determined from software stack errors, computing unit hardware errors, and excessive temperatures of heterogeneous cards. For abnormalities in training programs, it can be abnormal startup of training programs, performance degradation of training programs, and errors in running training programs.

[0072] Step 202: Construct resource load slices and program slices for training abnormal points. The resource load slices and program slices are slices with a time limit of minutes.

[0073] In one or more embodiments, after obtaining the training anomaly point, resource load slices and program slices can be configured based on the location of the training anomaly point. Resource load slices are used to simulate the resource load used for AI model training at the training anomaly point, and program slices are simulated training processes for AI model training. Resource load slices and program slices are both slices with a time limit of minutes, that is, the time required to simulate AI model training according to the resource load slices and program slices is only a time limit of minutes, that is, the simulation process of multiple training anomalies is only a superposition of time limits of minutes, and the total test time is at the hour level. Among them, each training anomaly point can be configured with one or more resource load slices and one or more program slices, that is, a training anomaly point can be configured with the same number of resource load slices and program slices.

[0074] Step 203. At the training abnormal point of the second device, simulate the running process according to resource load slicing and program slicing, and the second device is the device to be tested for computing power hardware.

[0075] In one or more embodiments, before putting the second device into the normal test process of computing power hardware detection, the terminal device can configure resource load slicing and program slicing to the second device, and instruct the second device to simulate the operation process at the training anomaly point. Accordingly, the second device performs AI model simulation training on the training process indicated by the program slice according to the resources of the resource load slice to obtain the simulated operation process of the training anomaly point.

[0076] Step 204: When the second device operates abnormally, determine that the second device is an abnormal device.

[0077] In one or more embodiments, the terminal device can analyze the simulated operation process of the second device to determine whether the simulated operation process is normal. If an abnormal situation occurs in the simulated operation process, it means that the second device is unqualified, that is, the second device is an abnormal device, then the second device may no longer need to invest in computing power hardware detection.

[0078] For example, the embodiment of the present application is a test library based on the abnormal case of prepositioning, which can be referred to Figure 3 The schematic diagram of pre-screening is shown, that is, after the equipment is powered on, pre-screening is performed based on the resource layer to screen out devices with normal hardware, so that before the equipment performs long-term test training, the faulty equipment is removed and sent back to the factory for repair, and then normalization testing is performed on normal equipment, and computing power is delivered to provide graphics processing unit (GPU) computing power resources. The product form can support mixed large models, video number recommendation model training, voice AI, visual CV, game AI and medical AI training computing services, and accelerate training efficiency.

[0079] The embodiment of the present application obtains the abnormal points of the first device during AI model training as training abnormal points, configures resource load slices and program slices with a time limit of minutes for the training abnormal points, and simulates the operation process of resource load slices and program slices at the training abnormal points of the second device before the second device performs computing power hardware detection, and analyzes whether the operation process is abnormal. When an operation abnormality occurs during the operation process, it can be determined that the second device is an abnormal device. Through the above method, before the normal test process of the device is carried out, the training abnormal points of other devices are used first, and the operation process of the training abnormal points of the device to be tested is simulated through resource load slices and program slices with a time limit of minutes, so as to quickly screen abnormal devices, put them into maintenance in time, and improve the efficiency of computing power resources. In addition, the short-term pre-screening can reduce costs and increase efficiency.

[0080] Optionally, in the above Figure 2On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, the training abnormal point includes multiple first abnormal points, each first abnormal point corresponds to a resource load slice and a program slice;

[0081] At the training abnormal point of the second device, the simulation operation process according to resource load slicing and program slicing includes:

[0082] Read a first abnormal point in sequence;

[0083] At the first abnormal point of the second device, the operation process is simulated according to the resource load slices and program slices corresponding to the first abnormal point.

[0084] In one or more embodiments, a method of simulating a training process is introduced. Each anomaly point collected by the terminal device can be referred to as a first anomaly point. After obtaining the first anomaly point, the first anomaly point can be saved through an anomaly point table, and the anomalies in the anomaly point table can be collectively referred to as training anomaly points. The terminal device can poll each first anomaly point in the anomaly table, and then configure resource load slices and program slices for the first anomaly point. The second device can perform AI model simulation training on the training process indicated by the program slice at the first anomaly point according to the resources of the resource load slice.

[0085] Secondly, in the embodiment of the present application, a method for simulating the training process is provided. Through the above method, the first abnormal point is polled in sequence to simulate the training of the second device, that is, to determine that the second device is a normal device, all training abnormal points need to be simulated and trained, which can improve the accuracy of device abnormality detection.

[0086] Optionally, in the above Figure 2 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, the resource load slice includes resource load, GPU card temperature, GPU utilization and operator usage set, and the program slice includes neural network, training data, program execution action and kernel state slice;

[0087] At the training abnormal point of the second device, the simulation operation process according to resource load slicing and program slicing includes:

[0088] At the training anomaly point of the second device, simulate the load usage based on the resource load, GPU card temperature, GPU utilization, and operator usage set;

[0089] The training program running process is simulated based on the neural network, training data, program execution actions and kernel state slices.

[0090] In one or more embodiments, a method of simulating a training process is introduced. Resource load slicing can simulate the load usage of the second device under the slice by indicating the utilization of the resource load, the temperature of the GPU card under the resource load, and the utilization of the GPU and the set of operator usage. Program slicing can simulate the load usage of the second device under the slice by indicating the neural network used for training, the training data used for training, the actions to be performed by the program during training, and the completion of the slice state preservation of the kernel state at different time points. Accordingly, the second device simulates the resource load, GPU card temperature, GPU utilization and the load usage of the operator usage set at the corresponding training abnormal point based on the resource load slicing and program slicing, and then uses the neural network to train the training data, and executes the program execution action during the training process, and completes the slice state preservation of the kernel state at different time points.

[0091] Secondly, in the embodiment of the present application, a method for simulating the training process is provided. Through the above method, resource load slicing and program slicing include multi-dimensional configurations, so that the simulated operation of the second device is more accurate and the simulation effect is improved.

[0092] Optionally, in the above Figure 2 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, at the training abnormal point of the second device, simulating the load usage according to the resource load, GPU card temperature, GPU utilization and operator usage set includes:

[0093] At the training abnormal point of the second device, the resource load is simulated by increasing the amount of computing data;

[0094] Simulate GPU card temperature by adjusting fan speed;

[0095] Simulate GPU utilization by overlaying training data;

[0096] Reference the heterogeneous computing ecosystem library to obtain operators and use collections for simulation.

[0097] In one or more embodiments, a method for simulating the usage of a load is introduced. There can be various ways to simulate resource usage according to this resource load slicing. In the embodiments of the present application, it can be to adjust the load by increasing the amount of computing data at the training anomaly point, that is, for the loads of CPU, memory, and disk input-output (IO) in the resource load, by increasing the amount of computing data, the load occupancy of CPU, memory, and disk IO is correspondingly increased. For different usage scenarios of the GPU, the temperature of the GPU card is also correspondingly different. Under normal circumstances, for the temperatures of different GPU cards, the fan speed of the second device will also be correspondingly different, that is, the higher the temperature of the GPU card, the faster the fan speed. In the embodiments of the present application, the required GPU card temperature can be simulated by adjusting the fan speed. For the GPU utilization rate, it can be simulated by adjusting the amount of training data, such as adjusting the GPU utilization rate by superimposing training data. For the set of operator usages, it can be obtained by importing a heterogeneous computing ecosystem library reference.

[0098] Exemplarily, the resource load slice can be configured with slices in multiple dimensions. The dimensions in Table 1 below include Time 1, Time 2, and Time 3, where each dimension is a minute time limit, and each dimension is configured with different CPU / memory / disk IO, GPU card temperature, GPU utilization rate, and set of operator usages. Among them, Float8, Float16, and BF16 are only examples.

[0099] Table 1

[0100]

[0101] Secondly, in the embodiments of the present application, a method for simulating the training process is provided. Through the above method, the resource load, GPU card temperature, GPU utilization rate, and computing power usage set are completely simulated, improving the simulation effect of the load usage.

[0102] Optionally, on the basis of the above Figure 2 corresponding respective embodiments, in another optional embodiment provided by the embodiments of the present application, simulating the running process of a training program according to a neural network, training data, program execution actions, and kernel state slices includes:

[0103] Loading the neural network and training data, and executing the training process;

[0104] Completing the training process according to the program execution actions and kernel state slices.

[0105] In one or more embodiments, a method of simulating the program running process is introduced. The second terminal device needs to first load the neural network and training data indicated by the program slice. The neural network can use a convolutional neural network (CNN), a recurrent neural network (RNN), and a generative adversarial neural network (GAN), etc. Among them, CNN, RNN and GAN are neural networks commonly used in AI training, which are open source and available, and can be downloaded directly from the AI ​​open source community. The second device can choose to directly configure the neural network or select any neural network as the model cornerstone, input the training data for AI model training, save the kernel slice state through kernel state slicing, and then complete the operation of non-training program automation through program execution actions, such as ending the training operation, to complete the training process.

[0106] Exemplarily, the program slice description of this solution is expanded as shown in Table 2 below. Based on Table 1, slices of the three dimensions of time 1, time 2 and time 3 are also configured. The neural networks used in training, such as CNN, RNN, GAN, etc., can be selected from different data sets of sample data. The extraction method can be through indicating the append point in the sample data, such as the path of the sample data and append point 1, the path of the sample data and append point 2, the path of the sample data and append point 3, that is, different data sets appended by the program are used as training data through different append points. The program execution action belongs to the operation of non-automatic execution of the training program. The typical example is the kill command in the semi-automatic state, which is used to illegally terminate the training program, etc.; the kernel state slice can use the kernel dump command at different time points to complete the dump error of the kernel execution state, that is, the kernel state is saved, which can be used to reproduce the kernel slice state under abnormal state.

[0107] Table 2

[0108]

[0109] Secondly, in the embodiment of the present application, a method for simulating the program running process is provided. Through the above method, the training process is completely simulated, the training simulation effect is improved, and the kernel slice state is saved, which is convenient for reproducing the kernel slice state under abnormal state.

[0110] Optionally, in the above Figure 2 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, after determining that the second device is an abnormal device, the method further includes:

[0111] The training anomalies, resource load slices and program slices are recorded in the anomaly repository, which uses a plug-in mode to save information.

[0112] In one or more embodiments, a method of recording anomalies is introduced. For each second device determined to be an abnormal device, the resource load slices and program slices simulated and used when the second device determines the abnormality, as well as the corresponding training abnormal points, can be saved to facilitate the reproduction of the abnormal state. For the constructed and analyzed abnormal points, a plug-in mode can be used for preservation in an embodiment of the present application, wherein the database used for preservation can be a relational database (MYSQL) for data with a small amount of data. For configuration-type data, the values ​​will be stored in a structured query language (SQL) table. For file-type data, the downloaded path will be stored in a SQL table. The file can be locked by loading the path to complete the file loading operation. The table structure is designed as follows, including the resource anomaly instance (case) table shown in Table 3 and the program anomaly case table shown in Table 4:

[0113] Table 3

[0114]

[0115] Table 4

[0116]

[0117] For each program exception case, there will be a table record with the same name in the resource exception case table, that is, the program exception case can be reproduced through the serial mode.

[0118] Exemplarily, the pre-screening architecture in the embodiment of the present application is as follows Figure 4 As shown in the figure, the architecture process is divided into three parts, namely, the abnormal monitoring center for running equipment, the scene construction center for abnormal cases, and the abnormal storage library. Among them, the abnormal monitoring center for running equipment realizes the source data collection of abnormal cases through this center, including resource performance abnormalities, resource load abnormalities, and training program abnormalities. Resource performance abnormalities can be the utilization rate of heterogeneous cards, the utilization rate of heterogeneous card video memory, and the activity of computing units at abnormal times. Resource load abnormalities can be software stack errors, computing unit hardware errors, and heterogeneous card temperatures at abnormal times. Training program abnormalities can be startup abnormalities, performance degradation, and operation errors. The scene construction center for abnormal cases realizes the incremental iteration accumulation of abnormalities through the plug-in mode. The purpose of the abnormal storage library is to record the operation conditions of the constructed abnormal plug-in configuration and the storage path of the plug-in.

[0119] A schematic diagram of a pre-processing abnormal screening process in the embodiment of the present application can be referred to Figure 5 As shown, step 501. For the device that is powered on, the kernel slices recorded in the program exception case table can be polled first; step 502. Execute the file download operation to load the kernel slices into the preparation environment, and then complete the download of the neural network and the loading of the training data through the download link; step 503. Complete the abnormal case replay, that is, after completing the data loading, join the resource list with the same name, complete the CPU, memory and disk load settings by superimposing the calculation data volume, and then execute the operator import loading, by superimposing the training data volume, and adjusting The fan speed makes the GPU utilization and temperature reach the values ​​recorded in the table, and finally executes the training calculation of the neural network according to the recorded time segment; Step 504. Determine whether the equipment is abnormal, that is, analyze whether it is abnormal through the training process, if so, execute step 505, if not, execute step 507; Step 505. Record the abnormality in the program abnormality case table; Step 506. Put the abnormal equipment into maintenance; Step 507. Determine whether the training abnormality point is polled, if so, execute step 508, if not, execute step 501; Step 508. Test the equipment for normal business execution.

[0120] Secondly, in the embodiment of the present application, a method for recording anomalies is provided. By recording the simulated resource load slices and program slices when determining anomalies in the anomaly storage library, the abnormal state can be easily reproduced, and the data can be saved by plug-in mode, which can save storage resources.

[0121] The following is a detailed description of the device abnormality analysis device in this application. Figure 6 , Figure 6 This is a schematic diagram of an embodiment of a device abnormality analysis device in an embodiment of the present application. The device abnormality analysis device 60 includes:

[0122] An acquisition unit 601 is used to acquire training abnormal points, where the training abnormal points include abnormal points of the first device during artificial intelligence AI model training;

[0123] A construction unit 602 is used to construct resource load slices and program slices for training abnormal points, where the resource load slices and program slices are slices with a time limit of minutes;

[0124] A simulation unit 603 is used to simulate the operation process according to resource load slicing and program slicing at the training abnormal point of the second device, where the second device is the device to be tested for computing power hardware;

[0125] The determining unit 604 is configured to determine that the second device is an abnormal device when the second device operates abnormally.

[0126] In the embodiment of the present application, a device for analyzing device abnormality is provided. Through the above device, before the normalization test of the device is performed, the training abnormal points of other devices are used first, and the operation process of the training abnormal points of the device to be tested is simulated through resource load slicing and program slicing with a time limit of minutes, so as to quickly screen abnormal devices, put them into maintenance in time, and improve the efficiency of computing resources. In addition, the short-term pre-screening can reduce costs and increase efficiency.

[0127] Optionally, in the above Figure 6 On the basis of the corresponding embodiment, in another embodiment of the device abnormality analysis device 60 provided in the embodiment of the present application,

[0128] The training abnormal point includes a plurality of first abnormal points, each of which corresponds to a resource load slice and a program slice;

[0129] The simulation unit 603 is specifically used for:

[0130] Read a first abnormal point in sequence;

[0131] At the first abnormal point of the second device, the operation process is simulated according to the resource load slices and program slices corresponding to the first abnormal point.

[0132] In an embodiment of the present application, a device anomaly analysis apparatus is provided. Through the above-mentioned apparatus, the first abnormal point is polled in sequence to perform simulation training on the second device, that is, to determine that the second device is a normal device, all training abnormal points need to be simulated and trained, which can improve the accuracy of device anomaly detection.

[0133] Optionally, in the above Figure 6 On the basis of the corresponding embodiment, in another embodiment of the device abnormality analysis device 60 provided in the embodiment of the present application,

[0134] Resource load slices include resource load, GPU card temperature, GPU utilization, and operator usage set. Program slices include neural network, training data, program execution action, and kernel state slices.

[0135] The simulation unit 603 is specifically used for:

[0136] At the training anomaly point of the second device, simulate the load usage based on the resource load, GPU card temperature, GPU utilization, and operator usage set;

[0137] The training program running process is simulated based on the neural network, training data, program execution actions and kernel state slices.

[0138] In an embodiment of the present application, a device abnormality analysis device is provided. Through the above device, resource load slicing and program slicing include multi-dimensional configurations, making the simulated operation of the second device more accurate and improving the simulation effect.

[0139] Optionally, in the above Figure 6 On the basis of the corresponding embodiment, in another embodiment of the device abnormality analysis device 60 provided in the embodiment of the present application,

[0140] The simulation unit 603 is specifically used for:

[0141] At the training abnormal point of the second device, the resource load is simulated by increasing the amount of computing data;

[0142] Simulate GPU card temperature by adjusting fan speed;

[0143] Simulate GPU utilization by overlaying training data;

[0144] Reference the heterogeneous computing ecosystem library to obtain operators and use collections for simulation.

[0145] In an embodiment of the present application, a device abnormality analysis device is provided. Through the above device, resource load, GPU card temperature, GPU utilization and computing power usage set are fully simulated to improve the load usage simulation effect.

[0146] Optionally, in the above Figure 6 On the basis of the corresponding embodiment, in another embodiment of the device abnormality analysis device 60 provided in the embodiment of the present application,

[0147] The simulation unit 603 is specifically used for:

[0148] Load the neural network and training data, and execute the training process;

[0149] The training process is completed according to program execution actions and kernel state slices.

[0150] In the embodiment of the present application, a device abnormality analysis apparatus is provided. Through the above-mentioned apparatus, the training process is completely simulated, the training simulation effect is improved, and the kernel slice state is saved, so as to facilitate the reproduction of the kernel slice state under abnormal state.

[0151] Optionally, in the above Figure 6 On the basis of the corresponding embodiment, in another embodiment of the device abnormality analysis device 60 provided in the embodiment of the present application,

[0152] The device 60 further includes a recording unit 605, which is specifically configured to:

[0153] The training anomalies, resource load slices and program slices are recorded in the anomaly repository, which uses a plug-in mode to save information.

[0154] In an embodiment of the present application, a device anomaly analysis apparatus is provided. Through the above-mentioned apparatus, the resource load slices and program slices simulated when the anomaly is determined are recorded in the anomaly storage library, which can facilitate the reproduction of the anomaly state, and save data by plug-in mode, which can save storage resources.

[0155] Figure 7 3 is a schematic diagram of a computer device structure provided by an embodiment of the present application. The computer device 300 may have relatively large differences due to different configurations or performances, and may include one or more CPUs 322 (for example, one or more processors) and memories 332, and one or more storage media 330 (for example, one or more mass storage devices) storing application programs 342 or data 344. Among them, the memory 332 and the storage medium 330 may be short-term storage or permanent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the computer device. Furthermore, the central processing unit 322 may be configured to communicate with the storage medium 330 to execute a series of instruction operations in the storage medium 330 on the computer device 300.

[0156] The computer device 300 may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input and output interfaces 358, and / or one or more operating systems 341, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0157] The steps performed by the terminal device in the above embodiment can be based on the Figure 7 The computer device structure shown.

[0158] A computer-readable storage medium is also provided in an embodiment of the present application, on which a computer program is stored. When the computer program is executed by a processor, the steps of the methods described in the above embodiments are implemented.

[0159] A computer program product is also provided in an embodiment of the present application, including a computer program, which, when executed by a processor, implements the steps of the methods described in the above embodiments.

[0160] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0161] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0162] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0163] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0164] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.

[0165] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for analyzing device abnormality, It is characterized in that include: Acquire training abnormal points, where the training abnormal points include abnormal points of the first device during artificial intelligence (AI) model training; Constructing resource load slices and program slices for the training anomaly points, wherein the resource load slices and the program slices are slices with a time limit of minutes; At the training abnormal point of the second device, simulating the running process according to the resource load slice and the program slice, the second device is a device to be tested for computing power hardware; When the second device operates abnormally, the second device is determined to be an abnormal device.

2. The method according to claim 1, It is characterized in that The training abnormal point includes a plurality of first abnormal points, each of which corresponds to one of the resource load slices and the program slice; At the training abnormal point of the second device, the simulation operation process according to the resource load slice and the program slice includes: Reading the first abnormal points one by one in sequence; At the first abnormal point of the second device, an operation process is simulated according to the resource load slice and the program slice corresponding to the first abnormal point.

3. The method according to claim 1, It is characterized in that The resource load slice includes resource load, GPU card temperature, GPU utilization and operator usage set, and the program slice includes neural network, training data, program execution action and kernel state slice; At the training abnormal point of the second device, the simulation operation process according to the resource load slice and the program slice includes: At the training abnormal point of the second device, simulating load usage according to the resource load, the GPU card temperature, the GPU utilization rate and the operator usage set; The training program running process is simulated according to the neural network, the training data, the program execution action and the kernel state slice.

4. The method according to claim 3, It is characterized in that At the training abnormal point of the second device, simulating load usage according to the resource load, the GPU card temperature, the GPU utilization rate and the operator usage set includes: At the training abnormal point of the second device, simulating the resource load by increasing the amount of calculation data; Simulating the GPU card temperature by adjusting the fan speed; Simulating the GPU utilization by superimposing the training data; The heterogeneous computing ecosystem library is referenced to obtain the operator usage collection for simulation.

5. The method according to claim 3, It is characterized in that The process of simulating the training program operation according to the neural network, the training data, the program execution action and the kernel state slice includes: Loading the neural network and the training data, and executing a training process; The training process is completed according to the program execution action and the kernel state slice.

6. The method according to claim 1, It is characterized in that After determining that the second device is an abnormal device, the method further includes: The training anomaly point, the resource load slice and the program slice are recorded in an anomaly repository, and the anomaly repository uses a plug-in mode to save information.

7. A device for analyzing equipment abnormality, It is characterized in that include: An acquisition unit, configured to acquire training abnormal points, wherein the training abnormal points include abnormal points of the first device during artificial intelligence (AI) model training; A construction unit, used to construct a resource load slice and a program slice for the training abnormal point, wherein the resource load slice and the program slice are slices with a time limit of minutes; A simulation unit, configured to simulate the operation process according to the resource load slice and the program slice at the training abnormal point of the second device, wherein the second device is a device to be tested for computing power hardware; A determination unit is used to determine that the second device is an abnormal device when the second device operates abnormally.

8. A computer device, It is characterized in that include: memories, transceivers, processors, and bus systems; Wherein, the memory is used to store programs; The processor is used to execute the program in the memory, including executing the method according to any one of claims 1 to 6; The bus system is used to connect the memory and the processor so that the memory and the processor can communicate with each other.

9. A computer-readable storage medium comprising instructions, which, when executed on a computer, causes the computer to perform the method according to any one of claims 1 to 6.

10. A computer program product, It is characterized in that When the computer program product is executed on a computer, the computer performs the method according to any one of claims 1 to 6.