Downtime analysis method of server
By obtaining system downtime screenshot images and hardware data when the server is downtime, and using a pre-built downtime analysis model, the cause of downtime is quickly and accurately identified, solving the problems of low identification accuracy and low manual analysis efficiency in the existing technology, and improving the fault diagnosis efficiency and security.
Patent Information
- Application Number
- CN202510732741.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, server downtime analysis lacks multi-dimensional correlation analysis, low recognition accuracy, limited applicable scenarios, cumbersome and inefficient manual analysis, making it difficult to quickly and accurately locate the cause of downtime.
By obtaining the system downtime screenshot image of the substrate management controller, extracting the color, texture and shape characteristics of the downtime, combining the hardware data before and after the downtime, input it into the pre-constructed downtime analysis model, extracting text features and combining the hardware data to output the downtime cause.
It realizes efficient and rapid diagnosis of the cause of downtime, shortens the time for fault diagnosis, reduces losses, avoids misjudgment, and ensures the safety of server operation.
Smart Images

Figure CN120256186A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and particularly to a method for analyzing server downtime. Background Art
[0002] In the related art, the health information data of the server can be obtained through the baseboard management controller and sample data can be constructed. Then, a monitoring model is used to analyze the sample data and output the monitoring result, and when a server failure exists in the monitoring result, it is displayed; or the downtime cause can be investigated by manually analyzing log files, checking the hardware status, etc.
[0003] However, in the related art, constructing sample data using health information data lacks multi-dimensional correlation analysis, has a low recognition accuracy, limited applicable scenarios, relies on manual processes which are cumbersome and inefficient, consumes a large amount of time and energy, and with the increase in the complexity of the server system, it is difficult to quickly and accurately locate the downtime cause only through logs and hardware status, and there is an urgent need for improvement. Summary of the Invention
[0004] The present invention provides a method for analyzing server downtime, so as to at least solve the problems in the related art, such as the lack of multi-dimensional correlation analysis in constructing sample data, low recognition accuracy, limited applicable scenarios, cumbersome manual analysis process, low efficiency, and difficulty in quickly and accurately locating the downtime cause with the increase in the complexity of the server system.
[0005] The present invention provides a method for analyzing server downtime, which is applied to the baseboard management controller. Wherein, the method includes the following steps: in the case of server downtime, obtain the system downtime screenshot image of the baseboard management controller, and determine at least one system downtime feature of the server according to the system downtime screenshot image; obtain the hardware data before downtime and the hardware data after downtime of the server; input the at least one system downtime feature, the hardware data before downtime and the hardware data after downtime into a pre-constructed downtime analysis model, extract the text features in the at least one system downtime feature, and output the downtime cause of the server in combination with the hardware data before downtime and the hardware data after downtime.
[0006] Optionally, in an embodiment of the present invention, the determining at least one system downtime feature of the server according to the system downtime screenshot image includes: based on the system downtime screenshot image, obtain at least one of the downtime color feature, downtime texture feature, and downtime shape feature in the system downtime screenshot image; based on at least one of the downtime color feature, downtime texture feature, and downtime shape feature, obtain the at least one system downtime feature.
[0007] Optionally, in an embodiment of the present invention, the step of inputting the at least one system downtime feature, the pre-downtime hardware data, and the post-downtime hardware data into a pre-constructed downtime analysis model, extracting the text features in the at least one system downtime feature, and outputting the downtime cause of the server in combination with the pre-downtime hardware data and the post-downtime hardware data includes: obtaining initial historical system downtime features corresponding to the at least one system downtime feature; screening out final historical system downtime features that meet preset similarity conditions from the initial historical system downtime features; and outputting the downtime cause based on at least one of the at least one system downtime feature, the final historical system downtime features, and the historical downtime causes corresponding to the final historical system downtime features, in combination with the pre-downtime hardware data and the post-downtime hardware data.
[0008] Optionally, in an embodiment of the present invention, the step of obtaining the system downtime screenshot image of the baseboard management controller includes: obtaining the display information of each hardware in the server; matching corresponding evaluation indicators according to the display information; and obtaining the system downtime screenshot image based on the evaluation indicators.
[0009] Optionally, in an embodiment of the present invention, before inputting the at least one system downtime feature, the pre-downtime hardware data, and the post-downtime hardware data into a pre-constructed downtime analysis model, it further includes: constructing a first downtime analysis model in the downtime analysis model based on a system downtime screenshot training set; constructing a second downtime analysis model in the downtime analysis model based on the system downtime screenshot training set, pre-downtime hardware training data, and post-downtime hardware training data; and obtaining the downtime analysis model according to the first downtime analysis model and the second downtime analysis.
[0010] Optionally, in an embodiment of the present invention, it further includes: generating visualization display information based on the downtime cause and displaying it to the user based on the visualization display information; and / or sending the downtime cause to a preset terminal to display the downtime cause using the preset terminal.
[0011] Optionally, in an embodiment of the present invention, before obtaining the system downtime screenshot image of the baseboard management controller, it further includes: obtaining at least one detection index of the server; determining whether the at least one detection index meets a preset downtime condition; and if the at least one detection index meets the preset downtime condition, determining that the server has experienced downtime.
[0012] Optionally, in an embodiment of the present invention, the obtaining of the pre-shutdown hardware data and the post-shutdown hardware data of the server includes: determining the shutdown moment according to the at least one detection index; and obtaining the pre-shutdown hardware data and the post-shutdown hardware data based on the shutdown moment.
[0013] The present invention also provides a shutdown analysis device for a server, which is applied to a baseboard management controller. The device includes: a first obtaining module, configured to obtain a system shutdown screenshot image of the baseboard management controller in the case of a server shutdown, and determine at least one system shutdown feature of the server according to the system shutdown screenshot image; a second obtaining module, configured to obtain the pre-shutdown hardware data and the post-shutdown hardware data of the server; and an output module, configured to input the at least one system shutdown feature, the pre-shutdown hardware data, and the post-shutdown hardware data into a pre-constructed shutdown analysis model, extract the text features in the system shutdown screenshot image, and output the shutdown reason of the server in combination with the pre-shutdown hardware data and the post-shutdown hardware data.
[0014] Optionally, in an embodiment of the present invention, the output module includes: a first obtaining unit, configured to obtain initial historical system shutdown features corresponding to the at least one system shutdown feature; a screening unit, configured to screen out final historical system shutdown features that meet a preset similarity condition from the initial historical system shutdown features; and an output unit, configured to output the shutdown reason in combination with the pre-shutdown hardware data and the post-shutdown hardware data based on at least one of the at least one system shutdown feature, the final historical system shutdown feature, and the historical shutdown reason corresponding to the final historical system shutdown feature.
[0015] Optionally, in an embodiment of the present invention, the first obtaining module includes: a second obtaining unit, configured to obtain the display information of each hardware in the server; a matching unit, configured to match corresponding evaluation indexes according to the display information; and a third obtaining unit, configured to obtain the system shutdown screenshot image based on the evaluation indexes.
[0016] Optionally, in an embodiment of the present invention, it further includes: a first construction module, configured to construct a first shutdown analysis model in the shutdown analysis model based on a system shutdown screenshot training set before inputting the at least one system shutdown feature, the pre-shutdown hardware data, and the post-shutdown hardware data into the pre-constructed shutdown analysis model; a second construction module, configured to construct a second shutdown analysis model in the shutdown analysis model based on the system shutdown screenshot training set, the pre-shutdown hardware training data, and the post-shutdown hardware training data; and a generation module, configured to obtain the shutdown analysis model according to the first shutdown analysis model and the second shutdown analysis.
[0017] Optionally, in an embodiment of the present invention, it further includes: a first display module, configured to generate visual display information based on the crash reason and display it to the user based on the visual display information; and / or, a second display module, configured to send the crash reason to a preset terminal to display the crash reason by using the preset terminal.
[0018] Optionally, in an embodiment of the present invention, it further includes: a third acquisition module, configured to acquire at least one detection index of the server before acquiring the system crash screenshot image of the baseboard management controller; a judgment module, configured to judge whether the at least one detection index meets a preset crash condition; a determination module, configured to determine that the server has crashed when the at least one detection index meets the preset crash condition.
[0019] Optionally, in an embodiment of the present invention, the second acquisition module includes: a determination unit, configured to determine the crash moment according to the at least one detection index; a fourth acquisition unit, configured to acquire the hardware data before the crash and the hardware data after the crash based on the crash moment.
[0020] Optionally, in an embodiment of the present invention, the first acquisition module includes: a fifth acquisition unit, configured to acquire at least one of the crash color feature, the crash texture feature, and the crash shape feature in the system crash screenshot image based on the system crash screenshot image; a generation unit, configured to obtain the at least one system crash feature based on at least one of the crash color feature, the crash texture feature, and the crash shape feature.
[0021] The present invention also provides a baseboard management controller, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above server crash analysis methods when executing the computer program.
[0022] The present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above server crash analysis methods are implemented.
[0023] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any one of the above server crash analysis methods are implemented.
[0024] Through the present invention, in the case of a server outage, it is possible to obtain the system outage screenshot image of the baseboard management controller, the hardware data before the outage, and the hardware data after the outage, and then input them into a pre-constructed outage analysis model. The corresponding system outage features are obtained from the system outage screenshot image, the text features are extracted, and in combination with the hardware data before the outage and the hardware data after the outage, the outage cause of the server is output. Therefore, it is possible to solve the problems of lack of multi-dimensional correlation analysis in constructing sample data, low recognition accuracy, limited applicable scenarios, cumbersome manual analysis process, low efficiency, and difficulty in quickly and accurately locating technical problems, achieving the technical effects of realizing efficient and rapid diagnosis by using the pre-constructed outage analysis model, shortening the fault diagnosis time, reducing losses, combining multi-modal data, avoiding misjudgment, and ensuring the safe operation of the server. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0026] Figure 1 It is a flowchart of a method for analyzing server outages provided by an embodiment of the present invention; Figure 2 It is a flowchart of the working principle of a method for analyzing server outages provided by an embodiment of the present invention; Figure 3 It is a flowchart of the working principle of a baseboard management controller provided by an embodiment of the present invention; Figure 4 It is a block diagram of a device for analyzing server outages provided by an embodiment of the present invention.
[0027] Reference Numerals: Wherein, 10 - a device for analyzing server outages; 100 - a first acquisition module, 200 - a second acquisition module, 300 - an output module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0029] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0030] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0031] An embodiment of the present invention provides a method for analyzing server downtime. The method will be described in detail in combination with the execution flow of the method for analyzing server downtime.
[0032] Specifically, Figure 1 FIG. is a flowchart of a method for analyzing server downtime according to an embodiment of the present invention.
[0033] As Figure 1 shown, the method for analyzing server downtime is applied to a baseboard management controller. Among them, the method includes the following steps: In step S101, in the case of server downtime, obtain a system downtime screenshot image of the baseboard management controller, and determine at least one system downtime feature of the server according to the system downtime screenshot image.
[0034] It should be noted that in the process of server operation in the embodiment of the present invention, downtime is a relatively common and serious problem, which will cause business interruption and economic losses. Therefore, when the server suddenly goes down, it is of great significance to be able to quickly analyze the cause of the server downtime locally and give effective diagnostic results and solution suggestions.
[0035] The baseboard management controller (abbreviated as BMC) is an independent management unit from the main CPU (Central Processing Unit) system in the server and has certain computing and storage capabilities. Therefore, the embodiment of the present invention can use the baseboard management controller to obtain a downtime screenshot image and perform analysis, providing a new idea for efficiently solving the problem of analyzing the cause of downtime.
[0036] As a possible implementation manner, the embodiments of the present invention can use a baseboard management controller to monitor the status of a server in real time. In the case of server downtime, the screenshot function of the baseboard management controller is triggered, and then the corresponding system downtime screenshot image is obtained, and at least one corresponding system downtime feature is determined according to the system downtime screenshot image.
[0037] Among them, in the embodiments of the present invention, the screenshot function of the baseboard management controller can be implemented through the OEM (Original Equipment Manufacturer) extension instruction of IPMI (Intelligent Platform Management Interface). For example, a downtime screenshot request is directly sent to the baseboard management controller using an IPMI command. Subsequently, the baseboard management controller accesses the graphics card frame buffer or directly reads the video memory data to obtain the corresponding system downtime screenshot image, and at least one corresponding system downtime feature is determined according to the system downtime screenshot image. Specifically, it can be set by those skilled in the art according to the actual situation, and the present invention does not make specific limitations.
[0038] In addition, in the automated scenario of the embodiments of the present invention, the screenshot function of the baseboard management controller can be configured to automatically execute a screenshot operation when the server experiences downtime without manual intervention, and then the corresponding system downtime screenshot image is obtained, and at least one corresponding system downtime feature is determined according to the system downtime screenshot image.
[0039] Exemplarily, the embodiments of the present invention can first confirm the model and function support of the baseboard management controller, and configure the network and access permissions of the baseboard management controller, so as to trigger the screenshot function of the baseboard management controller in the case of server downtime, obtain the corresponding system downtime screenshot image, and determine at least one corresponding system downtime feature according to the system downtime screenshot image.
[0040] Optionally, in an embodiment of the present invention, determining at least one system downtime feature of the server according to the system downtime screenshot image includes: obtaining at least one of the downtime color feature, downtime texture feature, and downtime shape feature in the system downtime screenshot image based on the system downtime screenshot image; and obtaining at least one system downtime feature based on at least one of the downtime color feature, downtime texture feature, and downtime shape feature.
[0041] It can be understood that in the embodiments of the present invention, the downtime color feature can be understood as when the server experiences downtime, a specific color area may appear on the screen. For example, a red warning area or a blue screen phenomenon, etc. The present invention does not make specific limitations. Furthermore, the embodiments of the present invention can extract the corresponding downtime color feature through the system downtime screenshot image, so as to determine whether the server has a blue screen or other color abnormalities.
[0042] In some embodiments, the embodiments of the present invention can obtain the corresponding downtime color feature from the system downtime screenshot image, so as to obtain the corresponding system downtime feature.
[0043] Exemplarily, the embodiments of the present invention can obtain the corresponding downtime color feature through a color analysis kernel, or can also obtain the corresponding downtime color feature through machine learning or deep learning. Specifically, it can be set by those skilled in the art according to the actual situation. The present invention does not make specific limitations.
[0044] The downtime texture feature can be understood as when the server interface freezes, the texture on the screen (such as background patterns, icon details, etc., the present invention does not make specific limitations) may no longer be updated or may appear abnormally, or some regular textures (such as grid lines, gradient backgrounds, etc., the present invention does not make specific limitations) are damaged. The present invention does not make specific limitations. Therefore, the embodiments of the present invention can extract the corresponding downtime texture feature through the system downtime screenshot image, so as to determine whether the server state is normal.
[0045] In some embodiments, the embodiments of the present invention can obtain the corresponding downtime texture feature from the system downtime screenshot image, so as to obtain the corresponding system downtime feature.
[0046] Exemplarily, the embodiments of the present invention can obtain the corresponding downtime texture feature through a texture analysis kernel, or can also obtain the corresponding downtime color feature through machine learning or deep learning. Specifically, it can be set by those skilled in the art according to the actual situation. The present invention does not make specific limitations.
[0047] Furthermore, in the embodiments of the present invention, the downtime shape feature can be understood as when the server experiences downtime, the system may display specific error icons or prompt messages. For example, warning icons, error prompt icons, etc. The present invention does not make specific limitations. Therefore, the embodiments of the present invention can extract the corresponding downtime shape feature through the system downtime screenshot image, so as to determine whether the server has encountered an error or a warning.
[0048] In some embodiments, the embodiments of the present invention can obtain the corresponding downtime shape feature from the system downtime screenshot image, so as to obtain the corresponding system downtime feature.
[0049] Exemplarily, in the embodiments of the present invention, the corresponding downtime shape features can be obtained through a shape analysis kernel, or the corresponding downtime shape features can be obtained through machine learning or deep learning. Specifically, those skilled in the art can set according to the actual situation, and the present invention does not make specific limitations.
[0050] In the embodiments of the present invention, by extracting the downtime color features, downtime texture features, and downtime shape features, the reasons for the server downtime can be accurately identified, thereby improving the accuracy of downtime reason judgment. By comprehensively considering the downtime color features, downtime texture features, and downtime shape features, important fault information can be avoided from being omitted, enhancing the comprehensiveness of fault diagnosis, adapting to different downtime scenarios, and thus improving the fault handling efficiency.
[0051] Optionally, in an embodiment of the present invention, before obtaining the system downtime screenshot image of the baseboard management controller, it further includes: obtaining at least one detection index of the server; determining whether the at least one detection index meets a preset downtime condition; if the at least one detection index meets the preset downtime condition, it is determined that the server has downtime.
[0052] It can be understood that the detection indexes of the server in the embodiments of the present invention can include but are not limited to the CPU power state, memory error counter, system heartbeat signal, the time for the BIOS (Basic Input / Output System) to synchronously load devices with the baseboard management controller during the startup phase, etc. The present invention does not make specific limitations.
[0053] In the actual execution process, in the embodiments of the present invention, the detection indexes of the server can be obtained first, and when the detection indexes do not meet certain downtime conditions, it is determined that the server has downtime. Among them, the certain downtime conditions can be set by those skilled in the art according to the actual situation, and the present invention does not make specific limitations.
[0054] Exemplarily, in the embodiments of the present invention, it can be determined whether the server has downtime by detecting the system heartbeat signal of the server and the time for the BIOS to synchronously load devices with the baseboard management controller during the startup phase. For example, when the system heartbeat signal is interrupted for more than 30 seconds, meeting certain downtime conditions, it is determined that the server has downtime; when the BIOS synchronously loads hardware configuration information (such as the FRU (Field Replaceable Unit) storage device list, the present invention does not make specific limitations) with the baseboard management controller during the startup phase and the time for loading devices is more than 20 minutes, meeting certain downtime conditions, the baseboard management controller directly triggers the screenshot function.
[0055] Embodiments of the present invention can determine whether a server has crashed based on the detection metrics of the server, comprehensively evaluate the operating status of the server from multiple dimensions, avoid misjudgment caused by a single factor, and achieve accurate determination of crash events. In addition, by analyzing these detection metrics, the root cause of the failure can be more accurately located, improving the efficiency of troubleshooting.
[0056] Optionally, in an embodiment of the present invention, obtaining a system crash screenshot image of the baseboard management controller includes: obtaining the display information of each hardware in the server; matching the corresponding evaluation metrics according to the display information; and obtaining the system crash screenshot image based on the evaluation metrics.
[0057] It can be understood that in the embodiments of the present invention, the display methods of different hardware in the server are different. For example, the display method of the BIOS interface is text mode, while the display method of the operating system GUI (Graphical User Interface) is graphical mode. Therefore, the embodiments of the present invention can dynamically adjust the decoding algorithm of the baseboard management controller according to different display information. For example, character dot matrix recognition is used in text mode, and RGB (Red Green Blue) pixel data is parsed in graphical mode.
[0058] Furthermore, in the text mode of the embodiments of the present invention, the character recognition accuracy or OCR (Optical Character Recognition) speed is used as the evaluation metric; in the image mode, the color restoration degree and resolution can be used as the evaluation metrics, and then the corresponding system crash screenshot image is obtained.
[0059] Exemplarily, when the embodiments of the present invention use the baseboard management controller to directly access the video memory area of the graphics card through the PCIe (Peripheral Component Interconnect Express) interface or a dedicated management bus (such as the Video Graphics Array (VGA) shared channel, etc., which is not specifically limited in the present invention) to obtain the system crash screenshot image, it can be compatible with the display methods of different hardware (such as integrated graphics cards, independent GPUs, etc., which are not specifically limited in the present invention).
[0060] Furthermore, the embodiments of the present invention can match corresponding evaluation metrics for different display information. For example, the image resolution is the standard resolution (such as 640×480, etc., which is not specifically limited in the present invention) to reduce storage and transmission overhead; for a low-brightness environment, the baseboard management controller can apply gamma correction (such as etc., which is not specifically limited in the present invention) to optimize the image color, thereby improving the image readability.
[0061] In an embodiment of the present invention, by obtaining display information and matching evaluation indicators, a system downtime screenshot image can be obtained, realizing dynamic adaptation of hardware types and fault characteristics, reducing invalid screenshot data, thereby reducing storage space occupancy and analysis burden, and improving the intelligence and automation of downtime analysis.
[0062] In step S102, the pre-downtime hardware data and post-downtime hardware data of the server are obtained.
[0063] It can be understood that in the embodiment of the present invention, the pre-downtime hardware data and post-downtime hardware data may but are not limited to include CPU data, such as CPU usage rate, CPU temperature, etc., which are not specifically limited in the present invention; memory data, such as memory usage, etc., which are not specifically limited in the present invention; storage device data, such as disk capacity, I / O performance, etc., which are not specifically limited in the present invention; network interface data, etc., which are not specifically limited in the present invention.
[0064] Furthermore, in the embodiment of the present invention, the pre-downtime hardware data and post-downtime hardware data can be obtained through system instructions, or through hardware management tools, or through third-party monitoring software, etc. Specifically, it can be set by those skilled in the art according to the actual situation, and the present invention does not make specific limitations.
[0065] Furthermore, in the embodiment of the present invention, the obtained pre-downtime hardware data and post-downtime hardware data can be saved in the flash memory of the baseboard management controller, or stored in an external memory card. Specifically, it can be set by those skilled in the art according to the actual situation, and the present invention does not make specific limitations. Among them, in the embodiment of the present invention, the flash memory or the external memory card can allocate an independent partition (such as a capacity of 10MB to 50MB, which is not specifically limited in the present invention) to prevent log overwriting; or a cyclic storage strategy can be adopted to automatically overwrite the oldest file when the space is insufficient, which is not specifically limited in the present invention.
[0066] In addition, it should be noted that in the embodiment of the present invention, the system downtime screenshot image can also be saved in the flash memory of the baseboard management controller, or in an external memory card. Specifically, it can be set by those skilled in the art according to the actual situation, and the present invention does not make specific limitations. At the same time, the corresponding downtime analysis model is also saved in the storage space; a knowledge database for fault diagnosis, for model reference, etc., which are not specifically limited in the present invention.
[0067] Optionally, in an embodiment of the present invention, obtaining the pre-downtime hardware data and post-downtime hardware data of the server includes: determining the downtime moment according to at least one detection indicator; obtaining the pre-downtime hardware data and post-downtime hardware data based on the downtime moment.
[0068] It can be understood that in the embodiments of the present invention, the downtime moment can be used to divide the hardware data into pre-downtime hardware data and post-downtime hardware data, and the present invention does not make specific limitations. Among them, in the embodiments of the present invention, the downtime moment can be determined by using detection indicators.
[0069] In some embodiments, the embodiments of the present invention can use detection indicators to determine the downtime moment, and then determine the pre-downtime hardware data and post-downtime hardware data according to the downtime moment.
[0070] Exemplarily, in the embodiments of the present invention, when the system heartbeat signal is interrupted for more than 30 seconds, or when the BIOS synchronizes the hardware configuration information with the baseboard management controller during the startup phase and the loading device takes more than 20 minutes, it is determined that the server has crashed, and then the corresponding downtime moment is determined, and the pre-downtime hardware data from 5 minutes to 30 seconds before the downtime moment is obtained. For some key detection indicators, such as CPU temperature, etc., the present invention does not make specific limitations. After the crash occurs, the post-downtime hardware data is obtained at a frequency of 1 second / time. For some secondary detection indicators, such as fan speed, etc., the present invention does not make specific limitations. After the crash occurs, the post-downtime hardware data is obtained at a frequency of 5 seconds / time.
[0071] The embodiments of the present invention divide the hardware data into pre-downtime hardware data and post-downtime hardware data according to the downtime moment determined by the detection indicators. By accurately positioning the downtime moment, the pertinence of data collection is improved, the detection indicators are dynamically adjusted, and the correlation analysis of multi-dimensional detection indicators is carried out to enhance the ability of fault reproduction and root cause location.
[0072] Optionally, in an embodiment of the present invention, before inputting at least one system crash feature, pre-downtime hardware data, and post-downtime hardware data into a pre-constructed crash analysis model, it further includes: constructing a first crash analysis model in the crash analysis model based on the system crash screenshot training set; constructing a second crash analysis model in the crash analysis model based on the system crash screenshot training set, pre-downtime hardware training data, and post-downtime hardware training data; and obtaining the crash analysis model according to the first crash analysis model and the second crash analysis.
[0073] It can be understood that in the embodiments of the present invention, the first crash analysis model can be understood as an OCR model and plays an important role in the crash analysis model.
[0074] Among them, in terms of information extraction by the OCR model, since the system downtime features can include but are not limited to various text information, such as error messages, system logs, process names, etc., the present invention does not make specific limitations. The OCR engine can recognize the text information in the system downtime features and convert it into text features that can be processed by a computer. For example, when the server crashes due to a specific software failure, the system downtime features can display the corresponding error code and prompt message, and are accurately extracted through the OCR model, providing key data for subsequent analysis and assisting the second downtime analysis model, such as an AI (Artificial Intelligence) model analysis. Combining the text information extracted by the OCR with other data (such as the pre-downtime hardware training data and the post-downtime hardware training data, which are not specifically limited in the present invention) can provide richer features for the AI model.
[0075] The second downtime analysis model can be understood as an AI model. Among them, the AI model can include but is not limited to various algorithm libraries, such as long short-term memory networks, fault tree models, etc., which are not specifically limited in the present invention. It uses the system downtime screenshot training set, pre-downtime hardware training data, and post-downtime hardware training data to perform more accurate downtime cause analysis.
[0076] In the actual execution process, before inputting at least one system downtime feature, pre-downtime hardware data, and post-downtime hardware data into the pre-constructed downtime analysis model according to the embodiments of the present invention, an OCR model can be constructed using the system downtime screenshot training set, and an AI model can be constructed according to the system downtime screenshot training set, pre-downtime hardware training data, and post-downtime hardware training data, and then a downtime analysis model can be obtained.
[0077] In addition, in the embodiments of the present invention, the system downtime screenshot training set can include a knowledge database that stores common error information, downtime information, and corresponding fault causes and solutions, etc., for the AI model to call and refer to when analyzing data. Among them, the knowledge database is stored using a relational database, which is convenient for data query and management. When the AI model performs downtime cause analysis, it first queries the knowledge database, and based on similar error information and downtime scenarios, quickly obtains possible fault causes and solutions, improving the analysis efficiency and accuracy.
[0078] The OCR engine of the embodiments of the present invention adopts a deep learning algorithm based on a convolutional neural network, which can efficiently and accurately recognize the text in the image; during the training process of the AI model, a large amount of historical downtime data is used for iterative training, continuously optimizing the model parameters, and improving the accuracy of fault diagnosis.
[0079] In step S103, at least one system downtime feature, the hardware data before downtime, and the hardware data after downtime are input into a pre-constructed downtime analysis model, the text features in at least one system downtime feature are extracted, and the downtime cause of the server is output by combining the hardware data before downtime and the hardware data after downtime.
[0080] In some embodiments, the embodiments of the present invention can use at least one system downtime feature, the hardware data before downtime, and the hardware data after downtime as input data and input them into a pre-constructed downtime analysis model. When at least one system downtime feature is input, the OCR model will be immediately called and analyzed. The OCR engine incorporates image processing and text extraction algorithms to extract text information such as text, error codes, and keywords in at least one system downtime feature, and obtain the text features in at least one system downtime feature. The AI model then comprehensively analyzes the text features, the hardware data before downtime, and the hardware data after downtime to more accurately determine whether the downtime is caused by software conflicts or hardware failures, realizing the positioning of the downtime cause and giving repair suggestions.
[0081] The embodiments of the present invention improve the accuracy of fault location, reduce the risk of misjudgment, and utilize the OCR model to automatically extract text features, shortening the fault troubleshooting time by combining system downtime features, the hardware data before downtime, and the hardware data after downtime. By associating text features with known fault patterns and combining the hardware data before downtime and the hardware data after downtime, irrelevant components can be quickly excluded, avoiding item-by-item troubleshooting, covering complex fault scenarios, and enhancing the comprehensiveness of fault diagnosis.
[0082] Optionally, in an embodiment of the present invention, inputting at least one system downtime feature, the hardware data before downtime, and the hardware data after downtime into a pre-constructed downtime analysis model, extracting the text features in the system downtime screenshot image, and outputting the downtime cause of the server by combining the hardware data before downtime and the hardware data after downtime includes: obtaining the initial historical system downtime features corresponding to at least one system downtime feature; screening the final historical system downtime features that meet the preset similarity conditions from the initial historical system downtime features; and outputting the downtime cause by combining at least one of at least one system downtime feature, the final historical system downtime features, and the historical downtime causes corresponding to the final historical system downtime features, the hardware data before downtime, and the hardware data after downtime.
[0083] It can be understood that the embodiments of the present invention can obtain the initial historical system downtime features similar to at least one system downtime feature from the knowledge database based on at least one system downtime feature, and screen the final historical system downtime features that meet certain similarity conditions from them. Among them, the certain similarity conditions can be set by those skilled in the art according to the actual situation, and the present invention does not make specific limitations.
[0084] In some embodiments, embodiments of the present invention may first screen out final historical system downtime characteristics that meet certain similarity conditions from at least one system downtime characteristic, obtain the corresponding historical downtime reasons for the final historical system downtime characteristics, and then output the downtime reasons based on at least one system downtime characteristic, the final historical system downtime characteristic, the corresponding historical downtime reason, the hardware data before downtime, and the hardware data after downtime.
[0085] Exemplarily, embodiments of the present invention may obtain initial historical system downtime characteristics with a similarity greater than 50% to the system downtime characteristic from a knowledge database based on the system downtime characteristic. In the case of a large amount of initial historical system downtime characteristic data, further screen the initial historical system downtime characteristics to screen out final historical system downtime characteristics with a similarity greater than 70% to the system downtime characteristic, and determine the corresponding historical downtime reasons for the final historical system downtime characteristics. Then, based on the system downtime characteristic, the final historical system downtime characteristic, the historical downtime reason, the hardware data before downtime, and the hardware data after downtime, output the downtime reason.
[0086] Embodiments of the present invention screen historical system downtime characteristics through system downtime characteristics, implement similar case matching, quickly locate the historical root cause and repair solutions, thereby improving the accuracy and efficiency of root cause analysis, reducing repeated analysis, combining the hardware data before downtime and the hardware data after downtime, and performing multi-modal data fusion to reduce the misjudgment rate.
[0087] Optionally, in an embodiment of the present invention, it further includes: generating visualization display information based on the downtime reason and displaying it to the user based on the visualization display information; and / or sending the downtime reason to a preset terminal to display the downtime reason using the preset terminal.
[0088] In some embodiments, embodiments of the present invention may perform visualization display through a baseboard management controller management interface, generate corresponding visualization display information for the downtime reason, and display it to the user.
[0089] Exemplarily, when there is a hardware failure in embodiments of the present invention, such as the CPU temperature exceeding the limit, a corresponding CPU temperature curve graph may be generated by the baseboard management controller and displayed to the user. In addition, embodiments of the present invention may also generate a corresponding offline root cause analysis report, including screenshots of the fault scene, a hardware data comparison table, and repair suggestions, which are not specifically limited in the present invention.
[0090] In some embodiments, embodiments of the present invention may send the downtime reason to a preset terminal and then display the downtime reason. Among them, the preset terminal can be set by those skilled in the art according to the actual situation, and the present invention is not specifically limited.
[0091] Exemplarily, embodiments of the present invention can send the reasons for downtime in real time through enterprise WeChat, emails, etc., or send the offline root cause analysis report as content to the corresponding users, and the present invention does not make specific limitations.
[0092] Embodiments of the present invention can visually display the reasons for downtime to users. By intuitively presenting the reasons for downtime, downtime information can be quickly transmitted, the response time can be shortened, and the efficiency of fault handling can be improved. Using multi-dimensional data cross-verification can enhance the accuracy of fault diagnosis and better adapt to complex scenarios.
[0093] Next, the working principle of the downtime analysis method for the server proposed in the embodiments of the present invention will be introduced in conjunction with an embodiment.
[0094] Among them, Figure 2 is a flowchart of the working principle of the downtime analysis method for the server provided according to an embodiment of the present invention.
[0095] Step S201: Determine that the server is down.
[0096] Among them, embodiments of the present invention can determine whether the server is down according to detection indicators. For example, when the detection indicators meet certain downtime conditions, it is determined that the server is down; otherwise, the server is not down.
[0097] Step S202: Collect the system downtime screenshot image and the hardware data before and after downtime.
[0098] Step S203: Image preprocessing.
[0099] Among them, embodiments of the present invention can match the corresponding evaluation indicators according to the display information of different hardware, and then preprocess the image to obtain the system downtime screenshot image, and obtain the corresponding downtime color features, downtime texture features, and downtime shape features therefrom.
[0100] Step S204: Identify text features.
[0101] Among them, embodiments of the present invention can use the OCR model in the pre-constructed downtime analysis model to extract the text features in the system downtime features.
[0102] Step S205: Analyze the reasons for downtime.
[0103] Among them, embodiments of the present invention can output the reasons for the server downtime by using the AI model in the pre-constructed downtime analysis model based on the text features, the hardware data before downtime, and the hardware data after downtime.
[0104] Step S206: Perform visual display.
[0105] Among them, in the embodiments of the present invention, visualization can be performed through the baseboard management controller management interface, generating corresponding visualization display information for the downtime cause and presenting it to the user; or the downtime cause can be sent to a preset terminal to further display the downtime cause.
[0106] In addition, Figure 3 FIG. 5 is a flowchart of the working principle of the baseboard management controller according to an embodiment of the present invention.
[0107] Step S301: Connect to the server.
[0108] Among them, in the embodiments of the present invention, the baseboard management controller can be used to monitor the status of the server in real time, and in the case of server downtime, trigger the screenshot function of the baseboard management controller, thereby obtaining the corresponding system downtime screenshot image and obtaining the corresponding system downtime characteristics therefrom.
[0109] Step S302: Determine the functions of the baseboard management controller.
[0110] Among them, in the embodiments of the present invention, the baseboard management controller includes a storage module and a program processing module.
[0111] The storage module is used to store the downtime analysis model, system downtime screenshot image, system downtime characteristics, pre-downtime hardware data, post-downtime hardware data, knowledge database, etc., which are not specifically limited in the present invention.
[0112] The program processing module is used to process the system downtime screenshot image, image preprocessing, knowledge database query, and processing and analysis of the downtime analysis model, etc., which are not specifically limited in the present invention.
[0113] Step S303: Visualization interface.
[0114] Among them, in the embodiments of the present invention, the downtime cause can be visually displayed.
[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.
[0116] According to the server downtime analysis method provided by the embodiments of the present invention, in the case of server downtime, a system downtime screenshot image of the baseboard management controller, pre-downtime hardware data, and post-downtime hardware data can be obtained, and corresponding system downtime features can be obtained from the system downtime screenshot image, and then input into a pre-constructed downtime analysis model to extract text features, and combined with the pre-downtime hardware data and the post-downtime hardware data, the downtime cause of the server is output, achieving efficient and rapid diagnosis by using the pre-constructed downtime analysis model, shortening the fault diagnosis time, reducing losses, combining multi-modal data, avoiding misjudgment, and ensuring the technical effect of the safe operation of the server. Thus, the technical problems in the related art, such as the lack of multi-dimensional correlation analysis in constructing sample data, low recognition accuracy, limited applicable scenarios, cumbersome manual analysis process, low efficiency, and difficulty in quickly and accurately positioning, are solved.
[0117] An embodiment of the present invention also provides a server downtime analysis device.
[0118] Figure 4 It is a block diagram of the server downtime analysis device provided by the embodiments of the present invention.
[0119] As Figure 4 shown, the server downtime analysis device 10 is applied to the baseboard management controller. Among them, the server downtime analysis device 10 includes: a first acquisition module 100, a second acquisition module 200, and an output module 300.
[0120] Among them, the first acquisition module 100 is used to acquire a system downtime screenshot image of the baseboard management controller in the case of server downtime, and determine at least one system downtime feature of the server according to the system downtime screenshot image.
[0121] The second acquisition module 200 is used to acquire pre-downtime hardware data and post-downtime hardware data of the server.
[0122] The output module 300 is used to input at least one system downtime feature, pre-downtime hardware data, and post-downtime hardware data into a pre-constructed downtime analysis model, extract text features in the system downtime screenshot image, and output the downtime cause of the server in combination with the pre-downtime hardware data and the post-downtime hardware data.
[0123] Optionally, in an embodiment of the present invention, the output module 300 includes: a first acquisition unit, a screening unit, and an output unit.
[0124] Among them, the first acquisition unit is used to acquire initial historical system downtime features corresponding to at least one system downtime feature.
[0125] A screening unit for screening out the final historical system downtime features that meet the preset similarity conditions from the initial historical system downtime features.
[0126] An output unit for outputting the downtime cause by combining the pre-downtime hardware data and the post-downtime hardware data based on at least one of the at least one system downtime feature, the final historical system downtime feature, and the historical downtime cause corresponding to the final historical system downtime feature.
[0127] Optionally, in an embodiment of the present invention, the first acquisition module 100 includes: a second acquisition unit, a matching unit, and a third acquisition unit.
[0128] Among them, the second acquisition unit is used to acquire the display information of each hardware in the server.
[0129] The matching unit is used to match the corresponding evaluation indicators according to the display information.
[0130] The third acquisition unit is used to acquire the system downtime screenshot image based on the evaluation indicators.
[0131] Optionally, in an embodiment of the present invention, it further includes: a first construction module, a second construction module, and a generation module.
[0132] Among them, the first construction module is used to construct the first downtime analysis model in the downtime analysis model based on the system downtime screenshot training set before inputting at least one system downtime feature, the pre-downtime hardware data, and the post-downtime hardware data into the pre-constructed downtime analysis model.
[0133] The second construction module is used to construct the second downtime analysis model in the downtime analysis model based on the system downtime screenshot training set, the pre-downtime hardware training data, and the post-downtime hardware training data.
[0134] The generation module is used to obtain the downtime analysis model according to the first downtime analysis model and the second downtime analysis.
[0135] Optionally, in an embodiment of the present invention, it further includes: a first display module, and / or, a second display module.
[0136] Among them, the first display module is used to generate visual display information based on the downtime cause and display it to the user based on the visual display information.
[0137] And / or, the second display module is used to send the downtime cause to a preset terminal to display the downtime cause by using the preset terminal.
[0138] Optionally, in an embodiment of the present invention, it further includes: a third acquisition module, a judgment module, and a determination module.
[0139] Among them, the third acquisition module is used to acquire at least one detection index of the server before acquiring the system downtime screenshot image of the baseboard management controller.
[0140] The judgment module is used to judge whether at least one detection index meets the preset downtime condition.
[0141] The determination module is used to determine that the server has crashed when at least one detection index meets the preset downtime condition.
[0142] Optionally, in an embodiment of the present invention, the second acquisition module 200 includes: a determination unit and a fourth acquisition unit.
[0143] Among them, the determination unit is used to determine the downtime moment according to at least one detection index.
[0144] The fourth acquisition unit is used to acquire the hardware data before downtime and the hardware data after downtime based on the downtime moment.
[0145] Optionally, in an embodiment of the present invention, the first acquisition module 100 includes: a fifth acquisition unit and a generation unit.
[0146] Among them, the fifth acquisition unit is used to acquire at least one of the downtime color feature, the downtime texture feature, and the downtime shape feature in the system downtime screenshot image based on the system downtime screenshot image.
[0147] The generation unit is used to obtain at least one system downtime feature based on at least one of the downtime color feature, the downtime texture feature, and the downtime shape feature.
[0148] For the description of the features in the corresponding embodiments of the server downtime analysis device, reference can be made to the relevant descriptions in the corresponding embodiments of the server downtime analysis method, which will not be elaborated here one by one.
[0149] According to the server downtime analysis device proposed in the embodiment of the present invention, in the case of server downtime, it can acquire the system downtime screenshot image of the baseboard management controller, the hardware data before downtime, and the hardware data after downtime, and obtain the corresponding system downtime features from the system downtime screenshot image, and then input them into a pre-constructed downtime analysis model to extract text features, and combine the hardware data before downtime and the hardware data after downtime to output the downtime reason of the server. It can achieve efficient and rapid diagnosis by using the pre-constructed downtime analysis model, shorten the fault diagnosis time, reduce losses, combine multi-modal data, avoid misjudgment, and ensure the technical effect of the safe operation of the server. Thus, it solves the technical problems in the related art, such as the lack of multi-dimensional correlation analysis in constructing sample data, the low recognition accuracy, the limited applicable scenarios, the cumbersome manual analysis process, the low efficiency, and the difficulty in quickly and accurately positioning.
[0150] An embodiment of the present invention further provides a baseboard management controller, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-described embodiments of the server downtime analysis method.
[0151] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described embodiments of the server downtime analysis method when running.
[0152] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROMs for short), random access memories (RAMs for short), external hard drives, magnetic disks, or optical discs that can store computer programs.
[0153] An embodiment of the present invention further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the server downtime analysis method.
[0154] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the server downtime analysis method.
[0155] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0156] The above has introduced in detail a server downtime analysis method provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A method for analyzing server downtime, characterized in that Applied to a baseboard management controller, wherein the method comprises the following steps: In the case of a server outage, obtain a system outage screenshot image of the baseboard management controller, and determine at least one system outage feature of the server according to the system outage screenshot image; Obtain the pre-outage hardware data and post-outage hardware data of the server; Input the at least one system outage feature, the pre-outage hardware data, and the post-outage hardware data into a pre-constructed outage analysis model, extract the text features in the at least one system outage feature, and output the outage cause of the server in combination with the pre-outage hardware data and the post-outage hardware data.
2. The method for analyzing server downtime according to claim 1, wherein The determining at least one system outage feature of the server according to the system outage screenshot image includes: Based on the system outage screenshot image, obtain at least one of the outage color feature, outage texture feature, and outage shape feature in the system outage screenshot image; Based on at least one of the outage color feature, outage texture feature, and outage shape feature, obtain the at least one system outage feature.
3. The method for analyzing the downtime of the server according to claim 1, wherein The inputting the at least one system outage feature, the pre-outage hardware data, and the post-outage hardware data into a pre-constructed outage analysis model, extracting the text features in the at least one system outage feature, and outputting the outage cause of the server in combination with the pre-outage hardware data and the post-outage hardware data includes: Obtain the initial historical system outage features corresponding to the at least one system outage feature; Screen the final historical system outage features that meet the preset similarity conditions from the initial historical system outage features; Based on at least one of the at least one system outage feature, the final historical system outage feature, and the historical outage cause corresponding to the final historical system outage feature, in combination with the pre-outage hardware data and the post-outage hardware data, output the outage cause.
4. The method for analyzing server downtime according to claim 1, wherein The obtaining the system outage screenshot image of the baseboard management controller includes: Obtain the display information of each hardware in the server; Match the corresponding evaluation indexes according to the display information; Based on the evaluation indexes, obtain the system outage screenshot image.
5. The method for analyzing server downtime according to claim 1, characterized in that, Before inputting the at least one system outage feature, the pre-outage hardware data, and the post-outage hardware data into a pre-constructed outage analysis model, it further includes: Based on a system outage screenshot training set, construct a first outage analysis model in the outage analysis model; Based on the system outage screenshot training set, pre-outage hardware training data, and post-outage hardware training data, construct a second outage analysis model in the outage analysis model; Obtain the outage analysis model according to the first outage analysis model and the second outage analysis.
6. The method for analyzing server downtime according to claim 1, wherein It further includes: Generate visualization display information based on the outage cause, and display it to the user based on the visualization display information; And / or, send the outage cause to a preset terminal to display the outage cause by using the preset terminal.
7. The method for analyzing server downtime according to claim 1, wherein Before obtaining the system outage screenshot image of the baseboard management controller, it further includes: Obtain at least one detection index of the server; Determine whether the at least one detection index meets a preset downtime condition; If the at least one detection index meets the preset downtime condition, it is determined that the server has crashed.
8. The method for analyzing server downtime according to claim 7, wherein The obtaining of the pre-downtime hardware data and the post-downtime hardware data of the server includes: Determine the downtime moment according to the at least one detection index; Based on the downtime moment, obtain the pre-downtime hardware data and the post-downtime hardware data.
9. A baseboard management controller, characterized in that, Includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the server downtime analysis method according to any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to be used to implement the server downtime analysis method according to any one of claims 1-8.
Citation Information
Patent Citations
Context-based operation and maintenance fault root cause positioning method and device, equipment and medium
CN110309009A
Crash recovery method and device, equipment and storage medium
CN115756924A
Server downtime automatic processing method and device, data processing unit and medium
CN119046038A
Reduce recurring issues and incidents by remediation artificial intelligence
US20250165330A1