A method for a BMC to acquire server error information
By mapping the host serial port to the BMC using CPLD in the server, the server stage can be determined and serial port debugging can be performed. This solves the hardware and software complexity problems of error information collection in low-end servers, simplifies error information collection, and improves the reliability and stability of the system.
Patent Information
- Application Number
- CN202411872732.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-12-18
AI Technical Summary
Existing technologies struggle to efficiently collect error information in low-end servers, resulting in complex hardware support and software deployment. Furthermore, serial port debugging conflicts with error information collection, leading to design complexity and low resource utilization efficiency.
By mapping the host serial port to the BMC based on CPLD, the CPLD is used to assist in judging the server stage and performing serial port debugging. During the power-on startup stage, the BMC collects startup data, and during the operation stage, data collection and error warning are performed based on the serial port input or the BMC.
It simplifies the server error information collection process, requires no additional hardware support or complex software deployment, and improves the reliability and stability of the server.
Smart Images

Figure CN119782020B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of fault diagnosis technology, and in particular to a method for obtaining server error information using a BMC. Background Technology
[0002] In modern server management, collecting and processing error information is a crucial step in ensuring normal system operation. However, existing technologies typically rely on RAS architectures or shared memory mechanisms, which have high hardware requirements, especially requiring high-end CPUs, making them difficult to implement on low-end servers. Furthermore, using additional monitoring software increases system design complexity and consumes server computing resources. Simultaneously, the conflict between serial port debugging and error information collection makes it difficult for existing methods to coordinate the two effectively. These factors present numerous challenges to existing solutions in terms of hardware compatibility, design simplification, and resource utilization.
[0003] At present, the relevant technologies have technical problems such as high requirements for hardware support and software deployment, complex design and low efficiency in server error information collection. Summary of the Invention
[0004] This application provides a method for obtaining server error information by BMC. It uses CPLD to map the host serial port to BMC, and uses CPLD to assist in judging the server stage and performing serial port debugging. During the power-on startup stage, the BMC collects startup data. During the operation stage, data is collected and error warnings are given based on serial port input or the BMC. This simplifies the server error information collection process and achieves the technical effect of not requiring additional hardware support or complex software deployment.
[0005] This application provides a method for BMC to obtain server error information, including:
[0006] Based on the CPLD, the host serial port is mapped to the BMC. The CPLD is connected to a serial port interface for host serial port debugging. Depending on the server stage, the CPLD assists in system entry judgment and serial port debugging, and performs server data collection and error warning. Specifically, if the server stage is the power-on startup stage, the host serial port input is empty, and the BMC collects startup data. If the server stage is the running stage, data collection is performed based on the serial port input or the BMC.
[0007] The proposed method for obtaining server error information using a BMC involves first mapping the host serial port to the BMC using a CPLD. The CPLD assists in determining the server stage and performing serial port debugging. During the power-on startup stage, the BMC collects startup data. During the runtime stage, data collection and error warnings are performed based on serial port input or the BMC. This method simplifies the server error information collection process and eliminates the need for additional hardware support or complex software deployment. Attached Figure Description
[0008] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly described below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0009] Figure 1 A schematic diagram illustrating the principle of a BMC for obtaining server error information, provided in an embodiment of this application;
[0010] Figure 2 This is a schematic diagram of a BMC monitoring and collection process for obtaining server error information, provided in an embodiment of this application. Detailed Implementation
[0011] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below.
[0012] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of this application will be provided in conjunction with the accompanying drawings. The described embodiments should not be considered as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] In the following description, references to "some embodiments" describe a subset of all possible embodiments. However, it is understood that "some embodiments" can be the same or different subsets of all possible embodiments and can be combined with each other without conflict. The terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. The terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only.
[0014] This application provides a method for BMC to obtain server error information, such as... Figure 1 As shown, the method includes:
[0015] Step S100: Based on the CPLD, the host serial port is mapped to the BMC. The CPLD is connected to a serial port connector for host serial port debugging. Specifically, in this scheme, the CPLD (Complex Programmable Logic Device) first establishes correct connections with the host serial port, the BMC (Baseboard Management Controller), and the serial port connector, adhering to electrical specifications. Then, it completes initialization programming by loading configuration data written in a schematic design tool or hardware description language and downloaded via a download cable, thus enabling the mapping function. When the host serial port outputs a signal, the CPLD's internal logic circuitry enhances its driving capability with a signal buffer and allocates paths using a multiplexer, ensuring that the signal is transmitted simultaneously to the serial port connector and the BMC with integrity and correct timing. The serial port connector is used for host serial port debugging. During debugging, technicians connect debugging equipment, and the CPLD controls the signal path switching to ensure normal debugging communication without affecting BMC data collection and host system operation. The data obtained by BMC through mapping is of great significance. During the power-on startup phase, it can collect information such as BIOS (Basic Input / Output System) or UEFI (Unified Extensible Firmware Interface) boot logs to discover hardware and configuration problems. During the operation phase, it collects system operation data, which, after analysis and processing, can provide early warning of faults and performance issues, thereby improving server reliability and stability.
[0016] In one possible implementation, based on a CPLD, the host serial port is mapped to the BMC. The CPLD is connected to a serial port for host serial debugging. Step S100 further includes step S110, determining the internal interaction logic between the server host and the BMC. Specifically, this involves studying the needs and relationships between the server host and the BMC (Baseboard Management Controller) in data transmission and control signal interaction. For example, it clarifies which host status information needs to be promptly transmitted to the BMC during server operation, and how the BMC sends control commands to the host (such as input switching commands in serial debugging scenarios). It considers the interaction methods between the two under different server operating modes (such as normal operation, startup phase, and fault state), and how to ensure the accuracy, integrity, and real-time performance of data. Through comprehensive analysis of these factors, a clear internal interaction logic between the server host and the BMC is identified, providing a theoretical framework for subsequent hardware configuration and programming.
[0017] Step S120: Based on the internal interaction logic, perform logic programming and configure the programmable interconnect array. Specifically, according to the determined internal interaction logic, begin logic programming operations. Use a professional hardware description language (such as VHDL or Verilog) to convert the interaction logic into code that the programmable logic device (CPLD) can understand and execute. Configure the programmable interconnect array in the CPLD (Complex Programmable Logic Device). The programmable interconnect array is the key part of the CPLD (Complex Programmable Logic Device) to realize various logic functions. By reasonably setting its internal connection relationships and logic gate circuit combinations, the CPLD can process input and output signals according to the predetermined interaction logic. For example, according to logic needs, configure specific signal path connections, determine which signals need to be buffered, multiplexed, or logically operated, so as to realize the correct mapping of host serial port output signals to BMC (Baseboard Management Controller) and serial port interface, as well as the switching control of serial port input signals, etc., to ensure that the logic functions of the entire system are correctly realized.
[0018] Step S130: Initialize the CPLD based on the programmable interconnect array. Specifically, after configuring the programmable interconnect array, the CPLD (Complex Programmable Logic Device) needs to be initialized. The compiled configuration data is downloaded to the CPLD chip via a download cable. The initialization process enables the various logic units, registers, and other components inside the CPLD to be initially set according to the configuration data and enter the predetermined working state. The CPLD has the ability to process signals according to the pre-designed logic function and can accurately execute the interaction between the host and the BMC (Baseboard Management Controller) during server operation. For example, when the server powers on, the CPLD can correctly transmit the BIOS (Basic Input / Output System) boot log information output from the host serial port to both the BMC and the serial port (if debugging is required) according to the initialization settings. It can also switch the serial port input according to subsequent instructions (such as control signals from the BMC), thereby realizing the normal operation and data interaction management of the entire server system at the hardware level.
[0019] Step S200: Based on the server phase, the CPLD assists in system entry judgment and serial port debugging, and performs server data collection and error warning. Specifically, server operation is divided into power-on startup and operation phases. The POST End signal is used to assist in the CPLD (Complex Programmable Logic Device) judgment phase. After the signal is pulled high on GPIO (General Purpose Input / Output), it is transmitted from the CPLD to the BMC. During serial port debugging, flag information and IPMI (Intelligent Platform Management Interface) commands are first defined. The user sends a command to select the serial port usage status. If it is used, the CPLD switches the input to the serial port and changes the flag, and the BMC stops receiving data but records the activity. If it is not used, the input is switched to the BMC (Baseboard Management Controller) and the flag is changed. Automatic switching occurs after half an hour of inactivity. During the startup phase, the BMC receives startup logs. During the operation phase, when debugging is not performed, the BMC receives system logs and other data. Before error warning, abnormal feature words are extracted, coefficients are determined, and a database is built. During operation, the BMC traverses the database for matching. If the threshold is exceeded, a serious alarm and guidance are issued. If the threshold is not exceeded, a normal alarm log is reported. This achieves error warning and effective server management, improving reliability and stability.
[0020] In one possible implementation, such as Figure 2As described above, according to the server stage, the CPLD is assisted in performing system entry judgment and serial port debugging, and server data collection and error warning are executed. Step S200 further includes step S210, determining whether the server host has entered the system, wherein the judgment method is based on querying the POST End signal. Specifically, during the operation of the server, accurately determining whether the server host has entered the system is an important basis for subsequent operations. The judgment is made by querying the POST End signal. POST, or Power-On Self-Test, is a series of hardware detection and initialization operations automatically executed by the computer system after power-on. The POST End signal marks the end of the POST process. When the server host completes all hardware detection and initialization and is ready to enter the operating system loading stage, the BIOS (Basic Input / Output System) or UEFI will send a POST End signal. As a key identifier, the signal plays a triggering role in the state transition in the entire server system. It indicates that the server host has completed the basic hardware preparation work and is about to enter the system operation stage or has already entered the system operation stage, providing an important judgment basis for different subsequent operations.
[0021] Step S220: If the POST End signal is detected, the GPIO of the CPLD auxiliary connection will exchange the POST End signal with the BMC. If the system has not been entered, the server stage is the power-on startup stage; if the system has been entered, the server stage is the running stage. Specifically, once the POST End signal is detected, the CPLD (Complex Programmable Logic Device) begins to play a crucial role in auxiliary connectivity. Through its internal logic circuitry and connection to GPIO, the CPLD processes the received POST End signal and interacts with the BMC (Baseboard Management Controller). The CPLD acts as a bridge for signal conversion and transmission, accurately identifying the POST End signal from the BIOS / UEFI and converting it into a suitable signal format for transmission between the BMC and the host. This signal is then transmitted to the BMC via GPIO, ensuring that the BMC can promptly obtain the server host's status information. This allows the BMC to adjust its data collection and management strategies based on the server's current stage (running or power-on startup). For example, during the running stage, the BMC may perform more real-time monitoring and data collection tasks, while during the power-on startup stage, the BMC primarily focuses on collecting critical information during the startup process, such as BIOS / UEFI boot logs. The server stage is clearly divided into the power-on startup stage and the running stage based on whether the POST End signal is detected. If the POST End signal is not detected... If the POST End signal is detected, the server is currently in the power-on startup phase. During this phase, the host system mainly performs hardware initialization and self-test. At this time, the host serial port input is usually empty. The BMC focuses on collecting startup data generated during this phase, such as detailed log information output by the BIOS / UEFI during startup. This data is of great significance for analyzing whether there are hardware failures or configuration problems during the server startup process. If the POST End signal is detected, it indicates that the server host has entered the system operation phase. All hardware components of the server have been initialized and are ready to run the operating system and applications. The BMC can adjust its working mode according to this status. For example, during the operation phase, the BMC may strengthen real-time monitoring of system operation data and prepare to deal with various possible system events, such as error warnings and performance monitoring, to ensure the stability and reliability of the server during operation. The server phase judgment mechanism based on the POST End signal provides a basic guarantee for the orderly operation and management of the entire server system.
[0022] In one possible implementation, based on the server stage, the CPLD assists in performing system entry judgment and serial port debugging, and executes server data collection and error warning. Step S200 further includes step S230, which involves defining serial port flag information, determining whether the host serial port is being used, performing serial port debugging, and modifying the status of the serial port flag information. Specifically, firstly, in order to effectively manage the usage of the host serial port, it is necessary to define serial port flag information. The flag information is a status indicator used to record whether the host serial port is currently in use, and this information can be obtained by different processes in the system, thereby achieving global monitoring of the serial port usage status.
[0023] Step S240, wherein determining whether to use the host serial port includes: Specifically, when determining whether to use the host serial port, this flag information is used as an important basis. By continuously querying the status of this flag information, the system can understand the serial port occupancy in real time, and thus determine the subsequent operation process, ensuring that serial port resources can be reasonably allocated and effectively utilized under different needs.
[0024] Step S250: If the serial port is to be used, a first command is generated and sent to the CPLD to switch the input of the host serial port to the serial port port and adjust the serial port flag information to the used state. Specifically, when it is determined that the host serial port needs to be used for debugging or other operations, the system will generate a first command. The first command is an instruction specifically designed for the CPLD (Complex Programmable Logic Device) to precisely control the switching of the serial port input. After generating the command, the system sends it to the CPLD. After receiving the command, the CPLD quickly executes the switching operation according to its internal preset logic circuit, switching the input of the host serial port to the serial port port. External debugging devices (such as serial port debugging terminals) can then establish a direct connection with the host serial port through the serial port port to achieve bidirectional data transmission, facilitating various debugging tasks for technicians, such as system fault diagnosis and software program debugging. In order to maintain the consistency and accuracy of the serial port status information, the system will adjust the serial port flag information to the used state so that other processes can know in a timely manner that the serial port is currently occupied, avoiding interference from other operations to the serial port debugging process.
[0025] Step S260: If the serial port is not used, a second command is generated and sent to the CPLD to switch the host serial port input to the BMC and adjust the serial port flag information to an unused state. Specifically, when it is determined that the host serial port is not needed, the system generates a second command, which is also sent to the CPLD (Complex Programmable Logic Device). The CPLD performs the opposite operation according to this command, that is, switches the host serial port input to the BMC. This operation allows the BMC to continue to collect data output from the host serial port normally, such as system operation logs and status information, ensuring that the BMC's (Baseboard Management Controller) monitoring function of the server system is not affected. The system adjusts the serial port flag information to an unused state, indicating to the entire system that the serial port is idle and available for other possible operations. This achieves flexible switching and efficient utilization of serial port resources under different demand scenarios, optimizing the overall operating efficiency of the server system.
[0026] In one possible implementation, if the port is not in use, a second command is generated and sent to the CPLD to switch the host serial port input to the BMC and adjust the serial port flag information to an unused state. Step S260 further includes step S261, acquiring serial port activity information, wherein the serial port activity information includes at least the generation timestamp of the last message. Specifically, the system needs to acquire serial port activity information, which is an important data source for the entire serial port debugging and management process. Serial port activity information covers various dynamic data during the use of the serial port, including at least the generation timestamp of the last message. This timestamp accurately records the moment when the last data transmission of the serial port was completed. It acts like a time stamp, providing a key basis for subsequent judgment of whether the serial port is in an idle state. By continuously monitoring and recording serial port activity information, the system can grasp the usage status of the serial port in real time, laying the foundation for reasonable management of serial port resources.
[0027] Step S262: Set the debugging constraint time zone. Specifically, the system will set a debugging constraint time zone, which is a predefined time range rule. The purpose of setting the debugging constraint time zone is to provide a reasonable time limit for the serial port debugging process, so as to avoid affecting other normal operations or resource consumption of the server system due to excessive debugging time. For example, depending on the server's workload, performance requirements, and actual application scenarios, the administrator may set the debugging constraint time zone to a specific duration such as half an hour or one hour. The set value will serve as an important criterion for subsequent judgment on whether the serial port debugging should end, ensuring that the serial port debugging activity is carried out within a reasonable time range and improving the overall operating efficiency of the server system.
[0028] Step S263: By identifying the serial port activity information and combining it with the real-time time node, the interval time zone is determined, and it is judged whether the debugging constraint time zone is met. Specifically, the system identifies the acquired serial port activity information, paying particular attention to the generation timestamp of the last piece of information, and calculates the interval time zone between the two in combination with the current real-time time node. The interval time zone represents the length of time elapsed from the end of the last serial port activity to the current moment. Then, the calculated interval time zone is compared with the preset debugging constraint time zone to determine whether the requirements of the debugging constraint time zone are met. If the interval time zone is greater than or equal to the debugging constraint time zone, it means that the serial port has been inactive for a long time, and the debugging work may have been completed or it is in an idle state, requiring further processing; conversely, if the interval time zone is less than the debugging constraint time zone, it means that the serial port is still within the normal debugging usage time range, and the system continues to maintain the current state and continuously monitor the serial port activity.
[0029] Step S264: If the debugging constraint time zone is met, a debugging termination command is generated. Specifically, when the determination result is that the debugging constraint time zone is met, the system will automatically generate a debugging termination command. The command is a control signal used to notify the system to perform a series of operations to end the current serial port debugging process. The generation of the debugging termination command marks the system's preparation for transitioning from serial port debugging mode to normal operation mode. It ensures that serial port debugging activities are terminated at the appropriate time, avoiding unnecessary resource waste and potential system risks, while ensuring that the server system can promptly return to normal working state and continue to perform other important tasks.
[0030] Step S265: Based on the debug termination command, control the CPLD to switch inputs and synchronize the status of the serial port flag information. Specifically, based on the generated debug termination command, the system controls the CPLD (Complex Programmable Logic Device) to perform an input switching operation. The CPLD switches the host serial port input from the current debug connection (e.g., serial port) to the BMC (Baseboard Management Controller) according to the command, enabling the BMC to resume collecting data output from the host serial port and restore normal monitoring of the server system. To maintain the system's accurate understanding of the serial port status, the system synchronously updates the status of the serial port flag information, changing it to an unused state. Other processes can immediately understand that the serial port has finished debugging and is available for other operations when querying the serial port flag information, achieving consistency and efficiency in serial port resource management and ensuring coordinated operation between various components of the server system.
[0031] In one possible implementation, based on the server stage, the CPLD is assisted in performing system entry judgment and serial port debugging, and server data collection and error warning are executed. Step S200 further includes step S270, acquiring server operation records and mining abnormal feature words, wherein the abnormal feature words are identified by error coefficients and error types. Specifically, in order to effectively determine server errors, the system needs to acquire server operation records. The operation records contain various data generated by the server during operation, such as system logs, application operation logs, hardware status information, etc., comprehensively reflecting the server's working status. By analyzing the operation records, the system uses specific data mining algorithms and techniques to mine abnormal feature words. Abnormal feature words are key information fragments closely related to possible server errors, and each abnormal feature word is identified by an error coefficient and an error type. The error coefficient is used to quantify the severity of the error represented by the feature word. For example, a feature word with a high error coefficient may indicate that the server has a serious hardware failure risk, while the error type clarifies the category to which the error belongs, such as hardware error, software error, network error, etc. This helps the system to more accurately locate and understand the problems that the server may face.
[0032] Step S280: Integrate and consolidate the abnormal feature words to construct an error feature library. Specifically, the mined abnormal feature words need to be integrated and consolidated to construct a complete error feature library. Each abnormal feature word and its related error coefficients and error types are organized according to certain rules and structures to form a database that the system can quickly query and compare. The error feature library is like a knowledge base for error diagnosis, storing error feature patterns that may occur in servers under various operating conditions. By continuously accumulating and improving this library, the system can cover more types of server error situations, improve the ability to identify different errors, provide a comprehensive and accurate basis for subsequent error judgment, and enhance the effectiveness and reliability of server error management.
[0033] Step S290: Based on the error feature library, perform error determination on the server data. Specifically, when the server continues to run and generates new data, the system performs error determination on this server data based on the pre-built error feature library. Specifically, the system compares the real-time acquired server data with the abnormal feature words in the error feature library one by one. The system traverses each information segment in the server data, checking whether it matches a certain abnormal feature word in the error feature library. If a matching abnormal feature word is found, based on its corresponding error coefficient and error type, the system can quickly determine whether the server has an error and the severity and type of the error. For example, if an abnormal feature word related to hardware overheating is detected in the server data, and its error coefficient is high, the system can promptly determine that the server has a serious hardware problem and take corresponding measures according to the error type, such as issuing an alarm to notify the administrator or activating a hardware protection mechanism, thereby achieving real-time monitoring and effective handling of server errors and ensuring the stable operation of the server.
[0034] In one possible implementation, server operation records are obtained, and abnormal feature words are mined. These abnormal feature words are identified by error coefficients and error types. Step S270 further includes step S271, traversing the abnormal feature words and mining a first error coefficient for each individual feature word. Specifically, the mined abnormal feature words are traversed. For each individual feature word, the system uses specific data mining techniques and algorithms to mine its corresponding first error coefficient. This requires in-depth analysis of factors such as the frequency of occurrence of the single feature word in the server operation data, its contextual association, and its similarity to known error patterns. For example, if a single feature word frequently appears in the logs when the server malfunctions and is associated with a typical hardware failure, it may be assigned a high first error coefficient. Through such detailed analysis and calculation of each single feature word, the system can initially quantify the potential error severity represented by each feature word, laying the foundation for a more comprehensive and accurate determination of the error coefficient.
[0035] Step S272: Determine feature word groups based on the synchronization relationships of the abnormal feature words. Specifically, after completing the preliminary analysis of single feature words, feature word groups are determined based on the synchronization relationships between abnormal feature words. Synchronization relationships refer to the simultaneous occurrence or close correlation of certain feature words in server operation data under specific time sequences or logical relationships. For example, when a server experiences a network failure, the feature words "network connection interruption" and "high packet loss rate" may appear simultaneously. There is a clear logical correlation between them, so they can be combined into a feature word group. The system identifies these feature word combinations with synchronization relationships through the analysis and pattern recognition of a large amount of server operation data, thereby forming more representative and targeted feature word groups. These feature word groups can more comprehensively reflect the complex error situations that may exist in the server.
[0036] Step S273: For the feature word group, mine the second error coefficient. Specifically, for the determined feature word group, the system again uses data mining methods to mine its corresponding second error coefficient. The process of mining the second error coefficient is similar to mining the first error coefficient of a single feature, but the factors considered are more complex. This is because there are interrelationships between multiple feature words in the feature word group. In addition to considering the attributes of each feature word itself, it is also necessary to comprehensively analyze factors such as the overall probability of the feature word group appearing in server error scenarios, the degree of impact on server operation, and the degree of matching with other known error combination patterns. For example, a feature word group containing multiple key hardware failure-related feature words often appears when the server has serious hardware problems, so it will be assigned a high second error coefficient to reflect its important indicative role in server errors.
[0037] Step S274: Determine the error coefficient based on the first error coefficient and the second error coefficient. Specifically, based on the first and second error coefficients already mined, the system determines the final error coefficient for each abnormal feature word (including single feature words and feature words in feature word groups). The determination method may involve weighted summation, logical judgment, or other comprehensive calculation methods. For example, different weights can be set for the first and second error coefficients according to the importance of feature words and feature word groups in server error determination, and then they can be weighted and summed to obtain the final error coefficient. The determined final error coefficient can more comprehensively and accurately reflect the true severity of the server error represented by the abnormal feature word, providing a more reliable quantitative basis for subsequent server data error determination based on the error feature library, thereby achieving more accurate server error management and processing.
[0038] In one possible implementation, based on the first error coefficient and the second error coefficient, the error coefficient is determined. Step S274 further includes step S2741, collecting server data, traversing the error feature library for matching, and determining the error matching result. Specifically, the system first continuously collects various data generated by the server during operation. The data covers various aspects such as server hardware status information, system operation logs, and application operation records. The collected data is the basis for subsequent error determination. The system traverses the pre-built error feature library, comparing each information fragment in the collected server data with the abnormal feature words in the error feature library one by one. During the matching process, it checks whether the keywords, data patterns, and numerical ranges in the data match the abnormal feature words. For example, if the server data contains a specific error code related to hardware failure or an abnormal system status description, the system will search for the corresponding abnormal feature words in the error feature library. Through a comprehensive and detailed matching process, it determines whether there is content in the server data that matches the abnormal situation defined in the error feature library, thereby obtaining the error matching result. The result will reflect whether the server may have an error and what kind of error may exist.
[0039] Step S2742: If the error matching result is not empty and is greater than a preset coefficient threshold, a first alarm command is generated and operational guidance information is output. Specifically, when the error matching result is not empty, it indicates that a possible error has been detected in the server data, and the severity of the error needs to be further determined. The system compares the error coefficient corresponding to the error matching result with the preset coefficient threshold. If the error coefficient is greater than the preset coefficient threshold, it indicates that the detected error is relatively serious and may have a significant impact on the normal operation of the server. In this case, the system will generate a first alarm command, which will trigger a series of emergency measures. For example, the first alarm command may cause the server to sound an alarm or send an emergency notification to the administrator, such as an email or SMS, to ensure that the administrator can be informed of the serious problem on the server in a timely manner. At the same time, the system will also output operational guidance information. This information is generated based on the pre-analysis and handling experience of this type of error and aims to help the administrator quickly locate the problem and take effective solutions. For example, it provides possible fault cause analysis, recommended repair steps, or links to relevant technical documents, so that the administrator can quickly respond to serious error situations, reduce server downtime, and reduce the impact on business.
[0040] Step S2743: If the error matching result is not empty and is less than a preset coefficient threshold, an alarm log is generated and reported. The preset coefficient threshold is a critical value for defining the severity of the error. Specifically, if the error matching result is not empty, but the error coefficient is less than the preset coefficient threshold, the detected error is relatively minor. Although it will not immediately pose a serious threat to the normal operation of the server, it still needs to be recorded and monitored. The system will generate an alarm log, which records detailed information about the error, including specific abnormal feature words in the error matching result, the corresponding error coefficient, the time the error was discovered, and server-related context information (such as the application currently running, system load, etc.). After generating the alarm log, the system will report it to the server management system or related monitoring platform. Administrators can subsequently view these alarm logs to retrospectively analyze the server's operating status, understand the minor errors that occurred on the server over a period of time, and promptly identify potential problem trends to take preventative measures in advance, preventing small problems from gradually accumulating into serious failures, thereby ensuring the long-term stable operation of the server.
[0041] Step S300: If the server stage is the power-on startup stage, the host serial port input is empty, and the BMC collects startup data. If the server stage is the running stage, data collection is performed based on the serial port input or the BMC. Specifically, when the server is in the power-on startup stage, the host serial port input is empty, and the BMC focuses on collecting startup data, such as BIOS or UEFI startup logs. This data helps troubleshoot problems during hardware initialization and self-test. After entering the running stage, if serial port debugging is required, the serial port is connected to a debugging device, and the host serial port input is switched to this position. Although the BMC pauses regular data collection, it records serial port debugging operations. If debugging is not required, the host serial port input is switched to the BMC, and the BMC executes commands to collect system running logs and other data. At the same time, it performs anomaly checks and judges the server's running status based on the collected data compared with a preset range or mode, providing a basis for error warnings, thereby achieving effective data management and status monitoring of the server at different stages.
[0042] This application embodiment uses a CPLD to map the host serial port to the BMC. The CPLD assists in determining the server stage and performing serial port debugging. During the power-on startup stage, the BMC collects startup data. During the operation stage, data collection and error warnings are performed based on the serial port input or the BMC. This achieves the technical effect of simplifying the server error information collection process without the need for additional hardware support or complex software deployment.
[0043] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A method for obtaining server error information using a BMC, characterized in that, The method includes: Based on CPLD, the host serial port is mapped to BMC, wherein the CPLD is connected to a serial port interface, which is used for serial port debugging of the host. Based on the server phase, with the assistance of CPLD, system entry judgment and serial port debugging are performed, and server data collection and error warning are executed; If the server stage is the power-on startup stage, the host serial port input is empty, and the BMC collects startup data; if the server stage is the running stage, data is collected based on the serial port input or the BMC. Perform serial port debugging, including: Define serial port flag information, determine whether to use the host serial port, perform serial port debugging, and modify the status of the serial port flag information; The process of determining whether to use the host serial port includes: If not in use, generate the first command, send it to the CPLD, switch the input of the host serial port to the serial port port, and adjust the serial port flag information to the use state; If not in use, generate a second command and send it to the CPLD to switch the host serial port input to the BMC and adjust the serial port flag information to a non-use state; Obtain serial port activity information, wherein the serial port activity information includes at least the generation timestamp of the last message; Set the debugging constraint time zone; By identifying the serial port activity information and combining it with real-time time nodes, the interval time zone is determined, and it is judged whether the debugging constraint time zone is met. Here, combining the real-time time nodes means focusing on the generation timestamp of the last piece of information and combining it with the current real-time time node. If the time zone for the debugging constraints is met, a debugging termination command is generated. Based on the debug termination command, the CPLD is controlled to switch inputs and the status of the serial port flag information is synchronized.
2. The method for obtaining server error information by a BMC as described in claim 1, characterized in that, Configure CPLD, including: Determine the internal interaction logic between the server host and the BMC; Based on the aforementioned internal interaction logic, the logic is programmed and a programmable interconnect array is configured. The CPLD is initialized based on the programmable interconnect array.
3. The method for obtaining server error information by a BMC as described in claim 1, characterized in that, The system entry determination includes: To determine whether the server host has entered the system, the POST End signal is used as the determination method. POST stands for Power-On Self-Test, and the POST End signal marks the end of the POST process. If the POST End signal is detected, the POST End signal is exchanged to the BMC based on the GPIO of the CPLD auxiliary connection. If the system has not been entered, the server stage is the power-on startup stage; if the system has been entered, the server stage is the running stage.
4. The method for obtaining server error information by a BMC as described in claim 1, characterized in that, Before executing error warnings, the following should be included: Obtain server operation records and mine abnormal feature words, wherein the abnormal feature words are identified by error coefficient and error type; Integrate and consolidate the aforementioned abnormal feature words to construct an error feature library; Based on the aforementioned error signature database, errors in server data are determined.
5. A method for obtaining server error information using a BMC as described in claim 4, characterized in that, The abnormal feature words are identified by error coefficients, including: Traverse the abnormal feature words and, for each single feature word, mine the first error coefficient; Based on the synchronization relationship of the abnormal feature words, feature word groups are determined; For the aforementioned feature word groups, a second error coefficient is mined; The error coefficient is determined based on the first error coefficient and the second error coefficient.
6. A method for obtaining server error information using a BMC as described in claim 4, characterized in that, Execute error warnings, including: Collect server data, traverse the error feature database for matching, and determine the error matching results; If the error matching result is not empty and is greater than the preset coefficient threshold, generate the first alarm command and output the operation guidance information; If the error matching result is not empty and is less than the preset coefficient threshold, an alarm log is generated and reported. The preset coefficient threshold is the critical value for defining the severity of the error.
Citation Information
Patent Citations
Serial port information collection device and method and server
CN114579400A
Method and device for mining abnormal rules
CN115659232A