Fault diagnosis method and system in startup stage of server
By establishing a communication connection between BMC and BIOS during the server boot stage, initializing the KCS module and setting up the IPMI command format, creating independent threads and fault prediction models, the problem of unintuitive BIOS status information monitoring and data transmission in the existing technology is solved, and rapid fault location and prediction are achieved.
Patent Information
- Application Number
- CN202510442151.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
The existing technology cannot monitor BIOS status information in real time during the server boot stage, cannot verify the correctness of data transmission, and the information display is not intuitive, resulting in difficulty in positioning the fault and consuming a lot of manpower and time.
By establishing a communication connection between BMC and BIOS, initializing the KCS module, setting up IPMI command format, creating independent threads, reading and translating status information codes, establishing a historical database and preprocessing and annotating, creating a boot fault prediction model, collecting and optimizing model parameters in real time.
It realizes intuitive display of BIOS status information and real-time verification of data transmission, improves fault positioning efficiency, quickly identify fault patterns through model prediction and provides solutions, and continuously optimizes fault prediction capabilities.
Smart Images

Figure CN120336060A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of server fault diagnosis, and particularly relates to a fault diagnosis method and system during the server startup phase. Background Art
[0002] During the operation of a server, when a fault occurs in the server, the corresponding log can be viewed to locate the fault information. If the server has not entered the system and a problem occurs during the startup phase, the server cannot enter the system at this time, and the fault information cannot be viewed through the log. At this time, the hardware-related personnel continuously debug to find the cause of the fault and then solve it. This method requires a large amount of human and time costs.
[0003] Therefore, a method for locating fault information during the server startup phase has emerged. It reports the running status information to the BMC by the BIOS during the entire running phase, and the BMC can monitor in real time which stage the BIOS is running to, and at the same time, the BIOS running status is displayed in real time on the BMC page. When the server freezes due to a fault during the startup process, it can be directly seen on the BMC web page at which stage the server is stuck, and the cause of the fault can be quickly located.
[0004] However, the above method for locating fault information during the startup phase needs to locate the fault during the startup phase by reading the PC pointer of the server chip. This requires professional operation and the displayed information is not intuitive enough, and it is impossible to monitor the status information sent during the startup phase in real time and verify the correctness of data transmission. Therefore, how to improve the existing method for locating fault information during the startup phase to monitor the status information sent during the startup phase in real time, verify the correctness of data transmission, analyze the data from the startup phase, compare it with the status information code agreed with the BIOS in advance, and translate the corresponding information code into an intuitive language description and display it on the web is a technical problem that needs to be solved urgently at present. Summary of the Invention
[0005] The purpose of the present invention is to provide a fault diagnosis method and system during the server startup phase to improve the existing method for locating fault information during the startup phase, so as to monitor the status information sent during the startup phase in real time, verify the correctness of data transmission, analyze the data from the startup phase, compare it with the status information code agreed with the BIOS in advance, and translate the corresponding information code into an intuitive language description and display it on the web.
[0006] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:
[0007] A fault diagnosis method during the server startup phase includes the following steps:
[0008] S1: After the BMC is started, initialize the BMC, and at the same time initialize the channel for transmitting information with the BIOS, and create OEM IPMI commands and the agreed information format for communicating with the BIOS;
[0009] S2: Create two independent threads, including a thread for receiving IPMI information and a thread for parsing IPMI information, and run the two independent threads;
[0010] S3: Read the status information code from the memory of the BMC web page, and translate the status information code into the corresponding text description according to the agreement with the BIOS, and display it on the web page;
[0011] S4: The BMC obtains the error code reported by the BIOS during the startup of the server, saves the system status information when the failure occurs, and establishes a historical database, which includes the error code reported by the historical BIOS, the fault log, the cause of the fault, and its corresponding solution;
[0012] S5: Perform data preprocessing on the historical database, perform data annotation on the preprocessed historical database, create a startup fault prediction model, and input the annotated historical database into the startup fault prediction model for model training;
[0013] S6: Real-time collect the error code reported by the BIOS when the server starts up, and input it into the trained startup fault prediction model, and the startup fault prediction model outputs the fault prediction result during the server startup stage;
[0014] S7: Based on the fault prediction result, call the historical data, obtain the fault log, the cause of the fault, and the corresponding solution corresponding to the fault prediction result, execute the corresponding solution in the server, and obtain the specified feedback data of the server after executing the solution;
[0015] S8: Input the specified feedback data into the startup fault prediction model to iteratively optimize the parameters of the startup fault prediction model.
[0016] Preferably, the specific process of step S1 is as follows:
[0017] S11: Establish a communication connection between the BMC and the BIOS;
[0018] S12: Initialize the KCS module;
[0019] S13: Perform data transmission between the BMC and the BIOS;
[0020] S14: Set the command format of the Intelligent Platform Management Interface IPMI.
[0021] Preferably, the specific process of step S11 is as follows:
[0022] Connect the BMC and the BIOS for communication based on the LPC bus, connect the BMC to the LPC controller of the bridge chip. The BIOS accesses the registers of the devices under the LPC of the BMC through the address in the LPC bus domain, maps the address of the LPC bus to the IO space of the LPC controller of the bridge chip, and the BIOS accesses the device registers under the BMC LPC by accessing the IO space.
[0023] Preferably, in step S12, the KCS module belongs to a sub-module under the BMC LPC controller. The initialization content of the KCS module includes: channel selection, clock enabling, and the address of register mapping. The register is completed by the driver of the BMC. After the KCS module is initialized, the KCS module starts to work normally. The BIOS performs access. The initialization of the KCS module is completed by the driver in the BMC kernel.
[0024] Preferably, the data transfer between the BMC and the BIOS in step S13 includes setting the register mapping address of the KCS. The BISO complies with the specified specifications of the Intelligent Platform Management Interface (IPMI) and reads and writes data according to the status of the registers in the protocol specifications. The BIOS sends an IPMI command to the BMC. If the BMC successfully receives it, it returns the specified information to the BIOS.
[0025] Preferably, for the receiving IPMI information thread in step S2, when the BIOS sends IPMI information through the KCS channel, the BMC reads the information through the KCS driver and passes it to the IPMI information parsing thread;
[0026] The parsing IPMI information thread first determines whether the format of the IPMI information is correct, verifies whether the netfn and cmd are consistent with the agreement, reads the status information code of the BIOS to determine whether it is in the agreed status information codes. If so, it saves it in the memory and returns the completion code to the BIOS. If not, it directly returns the error code to the BIOS;
[0027] The system status information when a fault occurs in step S4 includes timestamp, hardware configuration, and operating system version.
[0028] Preferably, the data preprocessing process in step S5 is as follows:
[0029] S51: Remove the redundant data and error data in the historical database, and process the missing values and outliers;
[0030] S52: Convert the error codes reported by the BIOS into a unified format, which includes hexadecimal data or decimal data, and standardize the fault logs.
[0031] S53: Extract specified key features from the standardized historical data. The key features include frequently occurring BIOS-reported error codes, fault occurrence times, and environmental factors.
[0032] Preferably, the specific process of processing missing values and outliers in step S51 is as follows:
[0033] S511: Calculate the mean and standard deviation of the data in the historical database, and calculate the difference between the data points in the historical database and the mean of the historical database.
[0034] S512: Compare the result of dividing the difference by the standard deviation with a preset threshold, and replace the data whose result is greater than the threshold or less than the threshold with specified data.
[0035] In a second aspect, a fault diagnosis system during the server startup stage is provided, which is used to implement any one of the described fault diagnosis methods during the server startup stage. It includes an initialization module, a BMC, a thread creation module, an information code reading module, a data preprocessing module, a model creation module, a startup fault prediction model, a data collection module, and an optimization module.
[0036] The initialization module is connected to the BMC, the BMC is connected to the thread creation module and the information code reading module, the data preprocessing module is respectively connected to the BMC and the model creation module, the model creation module is connected to the startup fault prediction model, and the startup fault prediction model is connected to the optimization module.
[0037] The initialization module is used to initialize the BMC after the BMC starts, and at the same time initialize the channel for transmitting information with the BIOS, and create OEM IPMI commands and agreed information formats for communicating with the BIOS.
[0038] The thread creation module is used to create two independent threads, including a thread for receiving IPMI information and a thread for parsing IPMI information, and run the two independent threads.
[0039] The information code reading module is used to read the status information code from the memory of the BMC web page, and translate the status information code into the corresponding text description according to the agreement with the BIOS, and display it on the web page.
[0040] The BMC is used to obtain the error codes reported by the BIOS during the server startup process, save the system status information when a fault occurs, and establish a historical database, where the database includes the error codes reported by the historical BIOS, fault logs, fault causes, and their corresponding solutions;
[0041] The data preprocessing module is used to perform data preprocessing on the historical database and perform data annotation on the preprocessed historical database;
[0042] The model creation module is used to create a startup fault prediction model and input the annotated historical database into the startup fault prediction model for model training;
[0043] The data acquisition module is used to collect in real time the error codes reported by the BIOS when the server starts up and input them into the trained startup fault prediction model;
[0044] The startup fault prediction model is used to output the fault prediction result during the server startup phase, call the historical data based on the fault prediction result, obtain the fault log, fault cause, and the corresponding solution corresponding to the fault prediction result, execute the corresponding solution on the server, and obtain the specified feedback data of the server after executing the solution;
[0045] The optimization module is used to input the specified feedback data into the startup fault prediction model and perform iterative optimization on the parameters of the startup fault prediction model.
[0046] The beneficial effects of the present invention include:
[0047] The fault diagnosis method and system during the server startup phase provided by the present invention initialize the channel for the BMC to transmit information with the BIOS, create OEM IPMI commands and information formats for communicating with the BIOS; create two independent threads; read the status information code from the memory of the BMC web page; establish a historical database, perform data preprocessing and annotation, and create a startup fault prediction model; collect in real time the error codes reported by the BIOS when the server starts up and input them into the startup fault prediction model, and the startup fault prediction model outputs the fault prediction result during the server startup phase; call the historical data based on the fault prediction result, obtain the fault log, fault cause, and the corresponding solution corresponding to the fault prediction result, and obtain the feedback data; input the feedback data into the startup fault prediction model and perform iterative optimization on the parameters of the startup fault prediction model. The above process directly sees the text description of the BIOS status, and different server models customize different status information codes.
[0048] First, by establishing a communication connection between the BMC and the BIOS, initializing the KCS module, transferring data between the BMC and the BIOS, and setting the command format of the Intelligent Platform Management Interface (IPMI), the intuitive display of the text description of the BIOS status can be achieved.
[0049] Secondly, by preprocessing the data in the historical database, annotating the preprocessed historical database, creating a startup failure prediction model, inputting the annotated historical database into the startup failure prediction model for model training, real-time collecting the error codes reported by the BIOS when the server starts up, and inputting them into the trained startup failure prediction model, the startup failure prediction model outputs the failure prediction results during the server startup phase. The constructed and trained AI model can quickly identify the potential patterns of failures and output the failure prediction results, improving the efficiency of failure handling.
[0050] Thirdly, by removing redundant data and error data in the historical database and processing missing values and outliers, converting the error codes reported by the BIOS into a unified format, which includes hexadecimal data or decimal data, and standardizing the fault logs, and extracting specified key features from the standardized historical data, where the key features include frequently occurring BIOS-reported error codes, fault occurrence times, and environmental factors, an effective data basis for subsequent model training is provided, improving the efficiency of model training.
[0051] Finally, by executing the corresponding solution in the server and obtaining the specified feedback data after the server executes the solution, inputting the specified feedback data into the startup failure prediction model, and iteratively optimizing the parameters of the startup failure prediction model, the startup failure prediction model can be continuously optimized to adapt to and predict new startup failures, continuously improving the failure prediction ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic flowchart of the failure diagnosis method during the server startup phase of the present invention.
[0053] Figure 2 It is a schematic flowchart of the data preprocessing of the present invention.
[0054] Figure 3 It is a schematic diagram of the system architecture of the failure diagnosis during the server startup phase of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The following further describes the present invention in detail with reference to the attached Figures 1 to 3 Drawings:
[0056] Embodiment 1
[0057] See the appendix Figure 1 As shown, a fault diagnosis method during the server startup phase includes the following steps:
[0058] S1: After the BMC starts up, initialize the BMC, and at the same time initialize the channel for transmitting information with the BIOS, and create an OEM IPMI command and a predefined information format for communicating with the BIOS;
[0059] S2: Create two independent threads, including an IPMI information receiving thread and an IPMI information parsing thread, and run the two independent threads;
[0060] S3: Read the status information code from the memory of the BMC web page, and translate the status information code into the corresponding text description according to the agreement with the BIOS, and display it on the web page;
[0061] S4: The BMC obtains the error code reported by the BIOS during the server startup process, saves the system status information when the fault occurs, and establishes a historical database, which includes the error codes reported by the historical BIOS, fault logs, fault causes, and their corresponding solutions;
[0062] S5: Perform data preprocessing on the historical database, perform data annotation on the preprocessed historical database, create a startup fault prediction model, and input the annotated historical database into the startup fault prediction model for model training;
[0063] S6: Real-time collect the error codes reported by the BIOS when the server starts up, and input them into the trained startup fault prediction model, and the startup fault prediction model outputs the fault prediction result during the server startup phase;
[0064] S7: Based on the fault prediction result, call the historical data, obtain the fault log, fault cause, and the corresponding solution corresponding to the fault prediction result, execute the corresponding solution on the server, and obtain the specified feedback data after the server executes the solution;
[0065] S8: Input the specified feedback data into the startup fault prediction model to iteratively optimize the parameters of the startup fault prediction model.
[0066] In the above solution, by initializing the channel for the BMC and the BIOS to transmit information, an OEM IPMI command and information format for communicating with the BIOS are created; two independent threads are created; the BMC web page is used to read the status information code from memory; a historical database is established, and data preprocessing and annotation are performed to create a power-on failure prediction model; the error codes reported by the BIOS during server startup are collected in real time and input into the power-on failure prediction model, and the power-on failure prediction model outputs the failure prediction result during the server startup phase; based on the failure prediction result, historical data is called to obtain the failure log, failure cause, and corresponding solution corresponding to the failure prediction result, and feedback data is obtained; the feedback data is input into the power-on failure prediction model to iteratively optimize the parameters of the power-on failure prediction model, so that the text description of the BIOS status can be directly viewed on the web page, which is more intuitive, different status information codes can be customized according to different server models, and at the same time, the server failure handling ability is improved through the power-on failure prediction model.
[0067] Embodiment 2
[0068] On the basis of Embodiment 1, the specific process of step S1 is as follows:
[0069] S11: Establish a communication connection between the BMC and the BIOS;
[0070] S12: Initialize the KCS module;
[0071] S13: Perform data transmission between the BMC and the BIOS;
[0072] S14: Set the command format of the Intelligent Platform Management Interface IPMI.
[0073] The communication between the BMC and the BIOS is mainly connected through the LPC bus. Generally, the LPC controller of the BMC is connected to the LPC controller of the bridge chip. The LPC controller of the bridge chip is the master, and the BMC side is the slave. In this way, the BIOS can access the registers of the devices under the LPC of the BMC through the address in the LPC bus domain. The address of the LPC bus is mapped to the IO space of the bridge chip. In this way, the BIOS can access the registers of the devices under the BMC LPC by accessing the IO space to achieve communication.
[0074] The KCS module is a sub-module under the BMC LPC controller. The LPC controller interface of the BMC is very rich, and the KCS is just one of them. The initialization of this module, including channel selection, clock enabling, and the addresses of register mapping, these registers are all completed by the driver of the BMC. That is, after the KCS module has been initialized as necessary, the KCS module can work normally, and at this time, the BIOS can access it. The initialization work is completed by the driver in the BMC kernel.
[0075] The KCS register address determines the address of the register used for reading and writing data. Its final virtual address is the base address in the IO space plus the corresponding offset, that is, LPC_IO_BASE + offset. The generally mapped addresses include the addresses of the data and command registers. The specific IO space base address and offset address can be viewed in the chip manual.
[0076] BISO needs to comply with the specification of the Keyboard Controller Style (KCS) Interface of IPMI and read and write data according to the status of the registers in the protocol specification. The BMC cannot communicate with the BIOS actively. The BIOS sends IPMI commands to the BMC, and if the BMC successfully receives them, it will send a return message to the BIOS.
[0077] The full name of IPMI is Intelligent Platform Management Interface. It is a complete hardware management specification for servers and other systems (such as storage devices, networks, and communication devices). Through this specification, users can use the IPMI protocol to monitor the physical health characteristics of servers, such as temperature, voltage, fan working status, power status, etc.
[0078] Each IPMI command contains the following parts: NetFn, Lun, and Cmd data. The implementation of each IPMI command is different. Some commands do not require the data part, some do, some commands do not return data, and some do. The NetFn, Lun, and CMd of each command are defined by the protocol, and it also provides commands defined by OEMs themselves.
[0079] In the response data, three bytes are required, namely CompletionCode, NetFn, and Cmd, which are the content to be returned by each command. The first byte is the CompletionCode, which represents the status of the command sent to the BMC. If the command is successfully received there, the CompletionCode is 0, otherwise there is a problem.
[0080] By establishing a communication connection between the BMC and the BIOS, initializing the KCS module, transferring data between the BMC and the BIOS, and setting the command format of the Intelligent Platform Management Interface IPMI, the intuitive display of the text description of the BIOS status can be achieved.
[0081] In this embodiment, the specific process of step S11 is as follows:
[0082] Connect the BMC and the BIOS for communication based on the LPC bus, connect the BMC to the LPC controller of the bridge chip. The BIOS accesses the registers of the devices under the LPC of the BMC end through the address in the LPC bus domain, maps the address of the LPC bus to the IO space of the LPC controller of the bridge chip, and the BIOS accesses the device registers under the BMC LPC by accessing the IO space.
[0083] In step S12, the KCS module belongs to a sub-module under the BMC LPC controller. The initialization content of the KCS module includes: channel selection, clock enable, and the address of register mapping. The register is completed by the driver of the BMC. After the KCS module is initialized, the KCS module works normally. The BIOS performs access. The initialization of the KCS module is completed by the driver in the BMC kernel.
[0084] In step S13, the data transmission between the BMC and the BIOS includes setting the register mapping address of the KCS. The BISO complies with the specified specifications of the Intelligent Platform Management Interface (IPMI) and reads and writes data according to the status of the registers in the protocol specifications. The BIOS sends an IPMI command to the BMC. If the BMC successfully receives it, it returns the specified information to the BIOS.
[0085] Embodiment 3
[0086] Based on Embodiment 1 or Embodiment 2, for the receiving IPMI information thread in step S2, when the BIOS sends IPMI information through the KCS channel, the BMC reads the information through the KCS driver and passes it to the IPMI information parsing thread;
[0087] The parsing IPMI information thread first determines whether the format of the IPMI information is correct, verifies whether the netfn and cmd are consistent with the agreement, reads the status information code of the BIOS to determine whether it is in the agreed status information code. If so, it saves it in the memory and returns the completion code to the BIOS. If not, it directly returns the error code to the BIOS;
[0088] The system status information when a fault occurs in step S4 includes timestamp, hardware configuration, and operating system version.
[0089] See Figure 2 , the data preprocessing process in step S5 is as follows:
[0090] S51: Remove the redundant data and error data in the historical database, and process the missing values and outliers;
[0091] S52: Convert the error codes reported by the BIOS into a unified format, which includes hexadecimal data or decimal data, and standardize the fault logs.
[0092] S53: Extract specified key features from the standardized historical data. The key features include frequently occurring BIOS-reported error codes, fault occurrence times, and environmental factors.
[0093] The above process of data preprocessing provides an effective data basis for subsequent model training and improves the efficiency of model training.
[0094] In this embodiment, the specific process of processing missing values and outliers in step S51 is as follows:
[0095] S511: Calculate the mean and standard deviation of the data in the historical database, and calculate the difference between the data points in the historical database and the mean of the historical database.
[0096] S512: Compare the result of dividing the difference by the standard deviation with a preset threshold, and replace the data whose result is greater than or less than the threshold with specified data.
[0097] The above process of processing missing values and outliers avoids the influence of missing values and outliers in the data on the subsequent model training process and efficiency.
[0098] A fault diagnosis system during the server startup phase, used to implement any one of the described fault diagnosis methods during the server startup phase. See Figure 3 , including an initialization module, BMC, thread creation module, information code reading module, data preprocessing module, model creation module, startup fault prediction model, data acquisition module, and optimization module. The initialization module is connected to the BMC, the BMC is connected to the thread creation module and the information code reading module, the data preprocessing module is respectively connected to the BMC and the model creation module, the model creation module is connected to the startup fault prediction model, and the startup fault prediction model is connected to the optimization module.
[0099] The initialization module is used to initialize the BMC after the BMC starts, initialize the channel for transmitting information with the BIOS at the same time, and create OEM IPMI commands and agreed information formats for communicating with the BIOS; the thread creation module is used to create two independent threads, including a thread for receiving IPMI information and a thread for parsing IPMI information, and run the two independent threads; the information code reading module is used to read the status information code from the memory of the BMC web page, and translate the status information code into the corresponding text description according to the agreement with the BIOS, and display it on the web page; the BMC is used to obtain the error code reported by the BIOS during the startup process of the server, save the system status information when the fault occurs, and establish a historical database, and the database includes the error code reported by the historical BIOS, the fault log, the cause of the fault and its corresponding solution.
[0100] The data preprocessing module is used to preprocess the historical database and perform data annotation on the preprocessed historical database; the model creation module is used to create a startup fault prediction model, and input the annotated historical database into the startup fault prediction model for model training; the data collection module is used to collect in real time the error code reported by the BIOS when the server starts up, and input it into the trained startup fault prediction model; the startup fault prediction model is used to output the fault prediction result during the server startup stage, call the historical data based on the fault prediction result, obtain the fault log, the cause of the fault and the corresponding solution corresponding to the fault prediction result, execute the corresponding solution on the server, and obtain the specified feedback data after the server executes the solution, and the optimization module is used to input the specified feedback data into the startup fault prediction model to iteratively optimize the parameters of the startup fault prediction model.
[0101] In summary, the fault diagnosis method and system provided by the present invention during the server startup stage.
Claims
1. A fault diagnosis method during the server startup phase, characterized in that It includes the following steps: S1: After the BMC is started, initialize the BMC. At the same time, initialize the channel for transmitting information with the BIOS, and create OEM IPMI commands for communicating with the BIOS and the agreed information format. S2: Create two independent threads, including a thread for receiving IPMI information and a thread for parsing IPMI information, and run the two independent threads. S3: Read the status information code from the memory of the BMC web page, and translate the status information code into the corresponding text description according to the agreement with the BIOS, and display it on the web page. S4: The BMC obtains the error code reported by the BIOS during the startup process of the server, saves the system status information when the fault occurs, and establishes a historical database. The database includes the error codes reported by the historical BIOS, fault logs, fault causes, and their corresponding solutions. S5: Perform data preprocessing on the historical database, perform data annotation on the preprocessed historical database, create a power-on fault prediction model, and input the annotated historical database into the power-on fault prediction model for model training. S6: Real-time collect the error codes reported by the BIOS when the server is powered on, and input them into the trained power-on fault prediction model. The power-on fault prediction model outputs the fault prediction result during the server startup phase. S7: Based on the fault prediction result, call the historical data, obtain the fault log, fault cause, and the corresponding solution corresponding to the fault prediction result, execute the corresponding solution in the server, and obtain the specified feedback data of the server after executing the solution. S8: Input the specified feedback data into the power-on fault prediction model to iteratively optimize the parameters of the power-on fault prediction model.
2. The fault diagnosis method during the server startup phase according to claim 1, wherein The specific process of step S1 is as follows: S11: Establish a communication connection between the BMC and the BIOS. S12: Initialize the KCS module. S13: Perform data transmission between the BMC and the BIOS. S14: Set the command format of the Intelligent Platform Management Interface IPMI.
3. A fault diagnosis method during the server startup phase according to claim 2, characterized in that The specific process of step S11 is as follows: Connect the BMC and the BIOS for communication based on the LPC bus, connect the BMC to the LPC controller of the bridge chip. The BIOS accesses the registers of the devices under the LPC of the BMC end through the address in the LPC bus domain, maps the address of the LPC bus to the IO space of the LPC controller of the bridge chip, and the BIOS accesses the device registers under the BMC LPC by accessing the IO space.
4. A fault diagnosis method during the server startup phase according to claim 2, characterized in that, In step S12, the KCS module is a sub-module under the BMC LPC controller. The initialization content of the KCS module includes: channel selection, clock enabling, and the address of register mapping. The register is completed by the driver of the BMC. After the KCS module is initialized, the KCS module works normally, and the BIOS performs access. The initialization of the KCS module is completed by the driver in the BMC kernel.
5. A fault diagnosis method during the server startup phase according to claim 2, characterized in that, The data transfer between the BMC and the BIOS in step S13 includes setting the KCS register mapping address. The BISO complies with the specified specifications of the Intelligent Platform Management Interface (IPMI) and reads and writes data according to the status of the registers in the protocol specifications. The BIOS sends an IPMI command to the BMC. If the BMC successfully receives it, it returns the specified information to the BIOS.
6. A fault diagnosis method during the server startup phase according to claim 1, characterized in that, In the receiving IPMI information thread in step S2, when the BIOS sends IPMI information through the KCS channel, the BMC reads the information through the KCS driver and passes it to the IPMI information parsing thread. The parsing IPMI information thread first determines whether the format of the IPMI information is correct, verifies whether the netfn and cmd are consistent with the agreement, reads the status information code of the BIOS, and determines whether it is in the agreed status information code. If so, it saves it in the memory and returns the completion code to the BIOS. If not, it directly returns the error code to the BIOS. The system status information at the time of a fault in step S4 includes the timestamp, hardware configuration, and operating system version.
7. A fault diagnosis method during the server startup phase according to claim 1, characterized in that, The data preprocessing process in step S5 is as follows: S51: Remove redundant data and error data in the historical database, and process missing values and outliers. S52: Convert the error codes reported by the BIOS into a unified format, and the converted format includes hexadecimal data or decimal data, and standardize the fault logs. S53: Extract specified key features from the standardized historical data, and the key features include frequently occurring BIOS-reported error codes, fault occurrence times, and environmental factors.
8. A fault diagnosis method during the server startup phase according to claim 7, characterized in that, The specific process of processing missing values and outliers in step S51 is as follows: S511: Calculate the mean and standard deviation of the data in the historical database, and calculate the difference between the data points in the historical database and the mean of the historical database. S512: Compare the result of dividing the difference by the standard deviation with a preset threshold, and replace the data whose result is greater than or less than the threshold with the specified data.
9. A fault diagnosis system during the server startup phase, which is used to implement a fault diagnosis method during the server startup phase described in any one of claims 1-8, characterized in that, It includes an initialization module, a BMC, a thread creation module, an information code reading module, a data preprocessing module, a model creation module, a power-on fault prediction model, a data acquisition module, and an optimization module. The initialization module is connected to the BMC, the BMC is connected to the thread creation module and the information code reading module, the data preprocessing module is respectively connected to the BMC and the model creation module, the model creation module is connected to the power-on fault prediction model, and the power-on fault prediction model is connected to the optimization module. The initialization module is used to initialize the BMC after the BMC starts, initialize the channel for transmitting information with the BIOS at the same time, and create the OEM IPMI command and the agreed information format for communicating with the BIOS. The thread creation module is used to create two independent threads, including a receiving IPMI information thread and a parsing IPMI information thread, and run the two independent threads. The information code reading module is used to read the status information code from the memory of the BMC web page, translate the status information code into the corresponding text description according to the agreement with the BIOS, and display it on the web page; The BMC is used to obtain the error code reported by the BIOS during the server startup process, save the system status information when the failure occurs, and establish a historical database, which includes the error codes reported by the historical BIOS, fault logs, fault causes and their corresponding solutions; The data preprocessing module is used to perform data preprocessing on the historical database and perform data annotation on the preprocessed historical database; The model creation module is used to create a power-on failure prediction model, and input the annotated historical database into the power-on failure prediction model for model training; The data acquisition module is used to collect in real time the error codes reported by the BIOS when the server is powered on and input them into the trained power-on failure prediction model; The power-on failure prediction model is used to output the failure prediction result during the server power-on stage, call the historical data based on the failure prediction result, obtain the fault log, fault cause and the corresponding solution corresponding to the failure prediction result, execute the corresponding solution on the server, and obtain the specified feedback data of the server after executing the solution; The optimization module is used to input the specified feedback data into the power-on failure prediction model to iteratively optimize the parameters of the power-on failure prediction model.
Citation Information
Cited By
Multi-BIOS (Basic Input / Output System) starting switching method and equipment, storage medium and computer program product
CN121433987A
Multi-bios boot switching method, device, storage medium and computer program product
CN121433987B