Diagnostic data management method and device, equipment and storage medium
By classifying, compressing and remotely storing BMC operation data, the problem of insufficient BMC local storage space is solved, efficient data management and fault diagnosis are achieved, and fault location efficiency is improved.
Patent Information
- Application Number
- CN202510896446.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional BMC data management methods cannot effectively cope with the massive log retention requirements in high-concurrency fault scenarios, resulting in core dump file sizes far exceeding local storage space, leading to historical data loss and delayed fault diagnosis.
A differentiated compression strategy is used to classify and compress BMC operation data, and the compressed data is transmitted to the remote storage module. The remote storage module is used for synchronization and disaster recovery, solving the BMC local storage space limitation.
It improves the compression rate and transmission efficiency of operating data, ensures data integrity and reliability, achieves efficient fault diagnosis and location, reduces storage requirements and improves fault location efficiency.
Smart Images

Figure CN120711085A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a diagnostic data management method, apparatus, device and storage medium. Background Art
[0002] With the rapid development of cloud computing, artificial intelligence, and edge computing technologies, the scale of data centers is growing exponentially. As a core component for server hardware management, the Baseboard Management Controller (BMC) undertakes critical functions such as hardware status monitoring, remote control, and fault diagnosis. Currently, the number of servers in most data centers has exceeded 40 million. The BMC corresponding to a single server can generate over 200 pieces of sensor data per second, and the average annual storage volume of fault event logs reaches 2-5TB. This poses a severe challenge to traditional BMC data management methods. When a BMC fails or crashes, a complete coredump file (core dump file) must be generated to record the real-time status of the BMC at the time of the failure or crash. However, BMCs primarily use local storage to record operational data. This storage of logs and other data is limited by the BMC's own limited storage space and is unable to cope with the massive log retention requirements in high-concurrency fault scenarios. The size of a complete coredump file often far exceeds the BMC's local storage capacity. Currently, BMC can alleviate storage pressure through a local circular storage strategy. However, this strategy will continuously overwrite old data, which will lead to the loss of historical data and serious consequences such as delayed fault diagnosis.
[0003] It can be seen that how to help BMC achieve complete and effective operation data management is a problem that technical personnel in this field need to solve. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a diagnostic data management method, apparatus, device, and storage medium, which can help the BMC solve the problem of operating data management.
[0005] To solve the above technical problems, an embodiment of the present invention provides a diagnostic data management method, comprising:
[0006] Acquiring operation data of the baseboard management controller; the operation data includes at least one of process status parameters, hardware parameters and log data of the baseboard management controller;
[0007] Classify the operating data based on their importance, and divide all the operating data into several data groups corresponding to several preset categories;
[0008] Using a preset compression strategy to compress the data in each data group; wherein the preset compression strategy includes a plurality of preset compression methods corresponding to the plurality of data groups;
[0009] Transmit all compressed running data to the remote storage module.
[0010] In some embodiments, after compressing the data in each data group using a preset compression strategy, the method further includes:
[0011] Determine the time identifier corresponding to each operation data, and define the time identifier as the timestamp corresponding to the operation data;
[0012] Get the server identifier of the server to which the baseboard management controller belongs;
[0013] Determine whether each operating data has a corresponding fault type. If so, define the fault type as a fault type label corresponding to the operating data; wherein the fault type is the type of fault that occurs during the operation of the baseboard management controller;
[0014] Use timestamp, server identifier, and fault type label as index fields to build an index for the compressed running data.
[0015] In some embodiments, transmitting all compressed running data to a remote storage module includes:
[0016] Configure priorities for all compressed operation data based on their importance;
[0017] Sort all compressed running data in descending order of priority to generate a transmission queue;
[0018] All compressed running data are transmitted in sequence according to the transmission queue.
[0019] In some embodiments, the remote storage module includes an edge node module and a cloud disaster recovery module, the physical distance between the edge node module and the baseboard management controller is less than the physical distance between the cloud disaster recovery module and the baseboard management controller, and the edge node module and the baseboard management controller are in the same server architecture;
[0020] Transmit all compressed running data to the remote storage module, including:
[0021] All compressed running data is transmitted to the edge node module so that the edge node module can transmit all compressed running data to the cloud disaster recovery module.
[0022] In some embodiments, the process status parameters of the baseboard management controller include at least one of a process identifier, stack information, and the number of file descriptors; the hardware parameters of the baseboard management controller include at least one of a central processing unit occupancy rate, a memory usage rate, a power status, a register status, and a firmware version; the log data of the baseboard management controller includes at least one of an operating system log, a network parameter, and driver information; and obtaining the operating data of the baseboard management controller includes:
[0023] All operating data of the baseboard management controller is collected in an asynchronous writing manner.
[0024] In some embodiments, after transmitting all compressed running data to the remote storage module, the method further includes:
[0025] The remote storage module is used to perform fault diagnosis on the baseboard management controller based on all compressed operating data.
[0026] In some embodiments, the specific process of the remote storage module performing fault diagnosis on the baseboard management controller based on all compressed operating data includes:
[0027] Inputting all compressed operation data into a pre-built rule engine so that the rule engine can determine whether there is any abnormal process operation of the baseboard management controller based on the process status parameters and hardware parameters in the operation data;
[0028] The abnormal process operation includes the current running process reaching the process crash threshold and the current running process having a memory leak;
[0029] If the rule engine determines that a process of the baseboard management controller is running abnormally, the abnormal process causing the abnormal process is determined and the abnormal process is restarted;
[0030] Inputting all compressed operating data into a pre-built machine learning model so that the machine learning model can determine whether the baseboard management controller has a crash risk based on process status parameters, hardware parameters, and log data in the operating data;
[0031] If the machine learning model determines that the baseboard management controller is at risk of crashing, the baseboard management controller is controlled to be reset.
[0032] To solve the above technical problems, an embodiment of the present invention further provides a diagnostic data management device, comprising:
[0033] A data acquisition unit, configured to acquire operating data of the baseboard management controller; the operating data includes at least one of a process status parameter, hardware parameter, and log data of the baseboard management controller;
[0034] A classification unit, configured to classify the operation data based on their importance, and divide all the operation data into a number of data groups corresponding to a number of preset categories;
[0035] A compression unit, configured to compress the data in each data group using a preset compression strategy; wherein the preset compression strategy includes a plurality of preset compression modes corresponding to the plurality of data groups;
[0036] The transmission unit is used to transmit all compressed running data to the remote storage module.
[0037] To solve the above technical problems, an embodiment of the present invention further provides an electronic device, including:
[0038] memory for storing computer programs;
[0039] The processor is configured to execute the computer program to implement the steps of the aforementioned diagnostic data management method.
[0040] To solve the above technical problems, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the diagnostic data management method as described above are implemented.
[0041] It can be seen from the above technical solution that after all the acquired operation data are classified, a differentiated compression strategy is used to compress each operation data, and then all the compressed operation data are transmitted to the remote storage module to realize remote management of the operation data of the baseboard management controller; the beneficial effect of the present invention is that the differentiated compression strategy is used to improve the compression rate of the operation data, and the operation data can be quickly transmitted even when the data volume is very large. The differentiated compression strategy and remote storage are collaboratively designed, and the synchronization and disaster recovery of all operation data are realized through remote storage, breaking through the storage limitations of the baseboard management controller itself, and using the remote storage module to ensure the integrity of the operation data, providing an efficient and reliable data management foundation for large-scale and even ultra-large-scale data centers. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0043] Figure 1 A flowchart of a diagnostic data management method provided by an embodiment of the present invention;
[0044] Figure 2 A schematic diagram of a diagnostic data processing flow provided by an embodiment of the present invention;
[0045] Figure 3 A schematic diagram of a fault diagnosis architecture of a remote storage module provided by an embodiment of the present invention;
[0046] Figure 4 A schematic structural diagram of a diagnostic data management device provided by an embodiment of the present invention;
[0047] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0049] The terms "including" and "having," as used in the present description and accompanying drawings, and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements and may include steps or elements that are not listed.
[0050] In order to enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0051] Next, a diagnostic data management method provided by an embodiment of the present invention is described in detail. Figure 1 As shown, Figure 1 A flowchart of a diagnostic data management method provided by an embodiment of the present invention; the diagnostic data management method includes:
[0052] S11: Acquire operation data of a baseboard management controller; the operation data includes at least one of a process status parameter, hardware parameter, and log data of the baseboard management controller;
[0053] It is not difficult to understand that in order to effectively record the real-time status of the BMC when it fails or crashes, this application will obtain the operating data of the baseboard management controller. The operating data includes the process status parameters, hardware parameters and log data of the baseboard management controller. In order to ensure a comprehensive reflection of the BMC status, a preferred embodiment is to obtain all process status parameters, hardware parameters and log data so that subsequent operators can integrate the real-time operating status of the BMC based on these operating data. These operating data can not only help operators perform fault diagnosis and fault location when the BMC fails or crashes, but also can realize diagnostic analysis of the BMC by analyzing and processing these operating data, thereby promptly determining whether the BMC has failed or crashed. The specific type and implementation method of the operating data are not specifically limited in this application. The frequency and method of obtaining the operating data are not specifically limited in this application. The operating data can be obtained in real time during the operation of the BMC, or it can be obtained periodically at a certain frequency, or it can be obtained only when an abnormality is detected in the BMC. The abnormality of the BMC includes the abnormality detected by the BMC itself and the abnormality detected by the monitoring system outside the BMC.
[0054] S12: Classifying the operating data based on their importance, and dividing all the operating data into a number of data groups corresponding to a number of preset categories;
[0055] S13: compressing the data in each data group using a preset compression strategy; wherein the preset compression strategy includes a plurality of preset compression methods corresponding to the plurality of data groups;
[0056] Understandably, to reduce the storage pressure of diagnostic data, after acquiring operational data, it is compressed. Especially for heterogeneous computing applications, the complex BMC hardware architecture results in a large amount of operational data requiring storage. Compression can effectively reduce the required storage volume for operational data. However, different types of operational data have varying levels of importance, requiring different compression strategies. For more important data, data integrity is prioritized, so lossless compression algorithms with lower compression ratios that fully preserve data integrity can be selected. For less important data, reducing storage volume is more important, so lossy compression algorithms with higher compression ratios can be selected. Therefore, after acquiring operational data, it can be categorized by importance into several data groups. Each data group can then be compressed using a different pre-defined compression method, achieving tiered compression for different data groups. By adopting a dynamic priority compression strategy, compression levels are divided according to data type, balancing data integrity and storage volume.
[0057] It should be noted that this application does not impose any special restrictions on the specific number and implementation methods of the preset categories. Multiple different categories can be divided according to different levels of importance. After the division according to the level of importance, further division can be performed according to the data type, semantic features, etc. of the running data. It is not limited to the division by importance. For example, differentiated compression strategies can be implemented according to the data type and the importance of the data. This application does not impose any special restrictions on the preset compression strategies and the specific types and implementation methods of each preset compression method. They can be selected and adjusted according to actual application conditions. The preset compression method can be implemented using a specific compression algorithm such as the LZ4 algorithm, the Zstandard algorithm, the LZMA (Lempel-Ziv-Markov chain Algorithm) algorithm, or Huffman coding. It can also be implemented using dimensional compression strategies such as intelligent compression using the semantic information or context association of the data, sampling compression based on the sampling time window, and extracting key features or statistical summaries of the data. There are also many options for determining the importance of operating data. This application does not make any special restrictions here. Generally speaking, more important data will be in the form of structured data with a clear data model and a fixed format. Therefore, the importance of the data can be judged by detecting the data format.
[0058] As a specific embodiment, all operational data is divided into two data groups based on their importance. The first data group uses primary compression (lossless compression). For core operational data such as process stacks and register states, a lossless compression algorithm (such as LZMA, LZ4, or Zstandard) is used to maintain a compression rate of 50%-70%, ensuring data integrity and preserving complete diagnostic information. The second data group uses secondary compression (lossy compression). For auxiliary log data such as system logs and network parameters, such as time series data generated by servers or BMC systems, various devices in the server architecture, and server or BMC applications during operation, a time-window-based sampling compression strategy is used. This strategy uses a Huffman coding compression algorithm combined with timestamp differential compression, effectively increasing the compression rate by 60%. When performing time-window-based sampling compression, only operational data corresponding to key timestamps, such as operation times and event occurrence times, is retained, while redundant entries are discarded, and duplicate, invalid, or minor data entries are eliminated, further streamlining data and reducing storage costs.
[0059] S14: Transmitting all compressed running data to the remote storage module.
[0060] As is readily understood, to avoid the storage space limitations of local BMC storage, this application chooses to remotely transmit compressed operational data to a remote storage module. This remote storage module, independent of local physical storage space, can achieve very large storage capacity through clustering, distributed architecture, or cloud services, enabling cross-regional collaboration and assisting operators in remote BMC diagnosis. Furthermore, multiple replica backups and distributed storage architectures can be used to ensure data reliability and security, enabling redundant backup and disaster recovery of operational data. This application does not specifically limit the specific transmission method or protocol. Blocked transmission can be used to improve data transmission efficiency. Specifically, a lightweight transmission protocol such as MQTT (Message Queuing Telemetry Transport) can be used. A dedicated transport protocol stack is used for low-overhead communication. Header compression technology is used to reduce protocol overhead (approximately 30% of standard TCP / IP). An adaptive bitrate adjustment mechanism is used to dynamically switch QoS (Quality of Service) levels based on network quality. Furthermore, the transmission of operational data can be encrypted using methods such as the national secret SM4 algorithm. The encryption key can be dynamically generated by the HSM (Hardware Security Module) in the BMC. This application does not specifically limit the specific type and implementation of the remote storage module. It can be implemented through a remote host computer, cloud services, and other methods. There are also various options for connecting the BMC and the remote storage module, which are not specifically limited in this application.
[0061] It should be noted that, in order to ensure the effective transmission and storage of operating data, before acquiring operating data, the remote storage module and the BMC will pre-confirm the feasibility of data transmission. If effective data transmission can be performed between the remote storage module and the BMC, step S11 is entered to acquire operating data. If effective data transmission cannot be performed between the remote storage module and the BMC, the remote storage module and / or the BMC will execute a corresponding alarm strategy so that the operator can promptly troubleshoot the problem and implement the diagnostic data management process, avoiding resource waste caused by invalid operations. This confirmation of the feasibility of data transmission can be actively initiated by the remote storage module and / or the BMC. The specific method of confirming the feasibility of data transmission is not specifically limited in this application. Taking the remote storage module as an example, a preset register can be pre-set in the BMC. The preset register is used to directly indicate whether the BMC currently supports data transmission. The remote storage module only needs to obtain the status of this preset register to determine the feasibility of operating data acquisition. Furthermore, if the remote storage module determines that effective data transmission is not possible between itself and the BMC, it can issue an instruction to trigger the BMC's hardware reset mechanism, or instruct the operator to trigger the BMC's hardware reset mechanism, and wait for the BMC to restart and then re-determine whether data transmission is supported.
[0062] Furthermore, to ensure the reliability of operational data, this application adopts a distributed storage architecture for operational data storage. On the one hand, after acquiring operational data, the operational data can be stored locally on the BMC, establishing a local cache mechanism for the BMC. A data snapshot of the last 24 hours is retained locally on the BMC, and a cyclic write strategy is used to prevent storage overflow. At the same time, a backup BMC is added to the BMC, and data redundancy is achieved through a dual-BMC design. Data synchronization occurs between the primary and backup BMCs. When the primary BMC fails or crashes, it automatically switches to the backup BMC for use, and the two synchronize fault logs via the LOG partition (log partition) of NAND Flash (a non-volatile storage medium). On the other hand, the BMC is connected to a remote storage module. For example, if the remote storage module is implemented using cloud storage, the BMC connects to it through a cloud storage interface, uploads compressed operational data to cloud storage via HTTP (HyperText Transfer Protocol), RESTful API (Representational State Transfer based Application Programming Interface), and other methods. The AOF persistence mechanism is used to ensure transmission reliability.
[0063] The present invention provides an optimized storage and analysis method for crash core dump (coredump) data for BMC, which is suitable for embedded management scenarios such as servers, network devices, and edge computing devices, and is used for fault diagnosis of embedded systems. It can be directly applied in BMC, and can also be used as a software framework for configuration management of other devices. A layered data processing method for BMC crash diagnosis is implemented, which solves the storage and transmission problems caused by the large size of coredump files through layered compression strategies, block transmission protocols, etc. The storage requirements can be reduced by more than 90%, and the efficiency of fault location can be effectively improved by 60%. It solves the problems of local storage being unfeasible, remote transmission being inefficient, and key information extraction being difficult due to the large size of coredump files when BMC crashes. The storage and transmission methods of diagnostic data when BMC crashes are optimized to achieve an innovative data management solution that integrates intelligent layered compression and secure remote storage. In the management of diagnostic data, dynamic layered compression is used to improve compression rates. Differentiated compression strategies are designed based on the semantic characteristics of operational data (such as stack traces and sensor time series data). When transmitting compressed operational data, bandwidth-aware transmission is further employed, dynamically adjusting data transmission strategies based on network bandwidth and establishing a QoS guarantee mechanism to achieve 90% effective bandwidth utilization in a 10Gbps network environment. Secure distributed storage is used to store operational data, ensuring the integrity and privacy of cross-data center log storage through blockchain verification and federated learning technologies. The collaborative application of layered compression and remote storage provides an efficient and reliable BMC crash diagnostic data management infrastructure for hyperscale data centers.
[0064] In some embodiments, after compressing the data in each data group using a preset compression strategy, the method further includes:
[0065] Determine the time identifier corresponding to each operation data, and define the time identifier as the timestamp corresponding to the operation data;
[0066] Get the server identifier of the server to which the baseboard management controller belongs;
[0067] Determine whether each operating data has a corresponding fault type. If so, define the fault type as a fault type label corresponding to the operating data; wherein the fault type is the type of fault that occurs during the operation of the baseboard management controller;
[0068] Use timestamp, server identifier, and fault type label as index fields to build an index for the compressed running data.
[0069] It's easy to understand that, in order to facilitate operators in quickly extracting key information from operational data and enabling timely fault analysis and location of BMC failures or crashes, an index can be constructed for the compressed operational data after compression. The specific index fields can be set and adjusted based on actual application scenarios and are not limited to the timestamp, server identifier, and fault type tag in this embodiment. The time identifier corresponding to the operational data refers to the time attribute of the operational data. For state parameters such as process status parameters and hardware parameters, the time identifier is the time when the operational data was acquired. For event-triggered data such as log data, the time identifier is the time when the event corresponding to the operational data occurred. The server identifier of the server to which the baseboard management controller belongs refers to the server ID. The BMC is set in the server architecture and is used for hardware status monitoring and remote management of the server. For a remote storage module, which can simultaneously store compressed operational data of multiple BMCs, adding the server ID as an index field facilitates operators to promptly locate abnormal BMCs based on this index. The fault type corresponding to the operating data refers to the type of fault that occurs when the BMC is operating. In this embodiment, the fault that occurs when the BMC is operating refers to an abnormality that the BMC itself can detect. At this time, the BMC can still continue to operate and can still effectively obtain operating data to transmit to the remote storage module. The operating data in this case is the operating data corresponding to the BMC when the fault occurs, so there is a corresponding fault type. That is, the compressed operating data stored in the remote storage module includes the operating data corresponding to the normal operation of the BMC and the operating data corresponding to the fault that the BMC can detect. By using the fault type label as an index field, it can help operators use the faults that have occurred to conduct further diagnosis and analysis for the BMC, determine whether these BMC faults may cause other abnormalities, and help operators locate the fault in a timely manner.
[0070] It should be noted that this application does not make any special restrictions on the specific type and implementation method of the constructed index. It can be implemented by generating a hash index, etc., associating the timestamp, server ID and possible fault type label of the BMC to support fast retrieval. This application does not make any special restrictions on the specific acquisition method of the timestamp, server identifier and fault type label. The BMC can directly use the connection between itself and the server to query the server identifier. If the BMC has a fault that it can detect itself, the BMC will generally generate a corresponding alarm signal. Therefore, it is possible to determine whether the operating data has a corresponding fault by directly detecting whether the BMC currently has an alarm signal when obtaining the operating data, and when the alarm signal exists, directly determine the type of fault based on the alarm manifestation of the alarm signal or the corresponding alarm code and description information, so as to configure the corresponding fault type label for the operating data according to the type of fault and the preset correspondence between the label.
[0071] Furthermore, for a BMC, multiple devices within it will generate multiple log data for the same event. Therefore, after acquiring the operational data, redundant data within the operational data can be identified and removed to achieve intelligent filtering of redundant data. This application does not specifically limit the specific method for determining redundant data. For example, for log data, data with the same data type and the same timestamp can be determined as redundant data. Alternatively, a machine learning model can be established to identify duplicate features of each operational data item. When the duplicate features of two data items reach a certain threshold, one of the data items can be determined as redundant data. By identifying redundant data and removing it from the operational data, the storage space occupied by the operational data can be further reduced (by approximately 75%). When storing compressed operational data in the remote storage module, the compressed operational data can be categorized and stored according to the fault type (e.g., hardware fault, firmware error) for operational data with fault type tags. Tampering of the stored data can also be prevented by attaching digital signatures.
[0072] Specifically, by indexing the operation data, operators can quickly identify the operation data corresponding to each BMC and each operation time from the remote storage module based on the index, thereby improving the processing speed of the operator's diagnosis, analysis and fault location.
[0073] In some embodiments, transmitting all compressed running data to a remote storage module includes:
[0074] Configure priorities for all compressed operation data based on their importance;
[0075] Sort all compressed running data in descending order of priority to generate a transmission queue;
[0076] All compressed running data are transmitted in sequence according to the transmission queue.
[0077] It is understandable that in order to ensure the quality of service, when transmitting compressed operating data, it is also possible to further configure the priority of each operating data according to the importance of the operating data, and then transmit the high-priority data to the remote storage module first during transmission. In the remote storage module, operating data of different priorities can also be stored in a hierarchical manner, and operating data with high priority is stored in a storage area with higher security and reliability. This application does not make any special restrictions on the configuration method of the priority of operating data. Generally speaking, the priority of log data is greater than the priority of registers and firmware information in hardware parameters, the priority of registers and firmware information in hardware parameters is greater than the priority of process status parameters, and the priority of process status parameters is greater than the priority of other hardware monitoring data in hardware parameters. The importance of operating data can also be determined based on the business criticality corresponding to the operating data, and the importance of different operating data can also be flexibly defined according to the actual application scenario. For example, the priority of operating data that represents abnormal records of the CPU (Central Processing Unit) is relatively high.
[0078] Specifically, by giving priority to the transmission of critical data with higher importance, the real-time and reliability of key services in the BMC are guaranteed. A clear transmission order is set to optimize resource allocation and improve the efficiency of the entire data transmission process.
[0079] In some embodiments, the remote storage module includes an edge node module and a cloud disaster recovery module, the physical distance between the edge node module and the baseboard management controller is less than the physical distance between the cloud disaster recovery module and the baseboard management controller, and the edge node module and the baseboard management controller are in the same server architecture;
[0080] Transmit all compressed running data to the remote storage module, including:
[0081] All compressed running data is transmitted to the edge node module so that the edge node module can transmit all compressed running data to the cloud disaster recovery module.
[0082] It's easy to understand that in order to further improve data transmission efficiency and ensure the storage reliability of compressed operational data, multi-level disaster recovery can be achieved by building a distributed storage architecture. First, a local cache layer is established using the BMC's internal memory, such as the BMC's onboard NOR Flash (Nor-type Flash Memory). The local cache layer only stores the on-site operational data corresponding to the three most recent BMC crashes (including restart cause and type information), and its capacity is optimized to 128KB. Second, in the server architecture where the BMC is located, the Rack Management Controller (RMC) is used as an edge node module to establish an edge node layer. Compressed operational data is first transmitted to the edge node layer via lightweight MQTT protocols and other methods. The shorter physical distance can increase transmission speed, and the edge node module can provide larger storage space and support breakpoint resume and data verification. Finally, the cloud-based disaster recovery layer is implemented through the cloud-based storage module. The edge node layer can transmit the compressed running data to the cloud-based disaster recovery module at the appropriate time. The cloud-based disaster recovery module can provide the largest storage space to ensure the storage of all compressed running data. Blockchain technology is used to implement data sharding storage across data centers, ensuring the immutability and global traceability of crashed data. This application does not specifically limit the specific types and implementation methods of edge node modules and cloud-based disaster recovery modules. Edge node modules are generally implemented using devices deployed on the network edge side of the server architecture, physically close to the cloud-based disaster recovery module, shortening the data transmission distance.
[0083] Furthermore, when a BMC crashes, a distributed storage architecture can be leveraged to implement disaster recovery mechanisms. For example, during the first restart after a BMC crash, basic operating data can be loaded from the local cache layer. During the second recovery restart, incremental cloud data can be supplemented through edge node modules. This dual-stage data recovery accelerates BMC crash recovery. Fault injection testing can also be performed on the BMC by implementing simulated crash scenarios during the uboot (Universal Boot Loader) phase to verify the effectiveness of the remote storage module's data integrity protection mechanism.
[0084] Specifically, by setting up a hierarchical storage architecture, after obtaining the BMC's operating data, the transmission efficiency can be effectively improved through hierarchical transmission. The hierarchical distributed architecture can realize the redundant design of compressed storage data, improve the reliability and security of storage data, and ensure the effective implementation of disaster recovery.
[0085] In some embodiments, the process status parameters of the baseboard management controller include at least one of a process identifier, stack information, and the number of file descriptors; the hardware parameters of the baseboard management controller include at least one of a central processing unit occupancy rate, a memory usage rate, a power status, a register status, and a firmware version; the log data of the baseboard management controller includes at least one of an operating system log, a network parameter, and driver information; and obtaining the operating data of the baseboard management controller includes:
[0086] All operating data of the baseboard management controller is collected in an asynchronous writing manner.
[0087] It's understandable that to achieve comprehensive BMC diagnostic analysis and ensure the accuracy and reliability of operations like fault location, it's necessary to collect multi-dimensional operational data in real time during BMC operation. The process status parameters in this operational data effectively represent the state of processes within the BMC. A process in the BMC is a program instance that performs specific tasks during BMC operation. These process status parameters include the process identifier, stack information, and the number of file descriptors. The process identifier (PID) is an integer that uniquely identifies a running process. Stack information refers to the state and data of the stack area in memory during program execution. A file descriptor is a non-negative integer (such as 0, 1, or 2) that identifies I / O (input / output) resources such as open files, sockets, pipes, and devices. Each process has its own file descriptor table. These process status parameters can be obtained by accessing the / proc / pid / directory, for example. Log data is obtained from the BMC's system logs, including operating system logs (such as Linux OS logs), network parameters, driver information, and custom event logs. Among them, operating system logs are records generated by the operating system kernel, services, and components during operation, and are used to describe system status, events, errors, user operations, and other information in the BMC. Network parameters are used to record key information related to network connections, protocol status, device performance, and abnormal events, and are used to monitor network health, troubleshoot connection failures, analyze traffic patterns, or locate security incidents. Driver information refers to records generated by the operating system when loading, running, or managing hardware drivers, and is used to describe the interaction status between the driver and the hardware device and the operating system kernel. Custom event logs are logs that users or developers actively design and record based on specific needs, and are used to track non-standardized, personalized events in the system, application, or business, such as firmware upgrade failure records.
[0088] It should be noted that baseboard management controller hardware parameters include hardware monitoring data and register and firmware information. Hardware monitoring data refers to real-time monitoring data for the BMC and the hardware devices in its server architecture. This includes the BMC's CPU utilization, memory usage, sensor detection results such as temperature sensors in the server architecture where the BMC resides, and the server architecture's power status. Memory usage can be obtained by accessing the / proc / meminfo directory. Register and firmware information includes the status of BMC registers and the BMC firmware version, and can be obtained through the BMC's internal IPMI (Intelligent Platform Management Interface) interface.
[0089] To improve the reliability of operational data acquisition, an asynchronous write mechanism is used to collect and acquire operational data. This improves efficiency and prevents data loss due to BMC crashes. The BMC's memory buffer temporarily stores collected raw data. This asynchronous write mechanism ensures that data has been received or temporarily stored in the buffer before writing it to a storage medium (such as a hard drive, memory, or cache). A "write success" signal is returned to the caller without waiting for the data to actually be written. A background thread or process then completes the actual physical write to the remote storage module. In actual applications, operational data such as BMC configuration files and BIOS (Basic Input / Output System) parameters can also be obtained as needed.
[0090] Specifically, by acquiring multi-dimensional data from the BMC and coordinating it with real-time monitoring data from hardware monitoring devices (such as sensors), operators can integrate on-site BMC operation information, monitor the BMC process level through process status parameters, and analyze the process status, number of file descriptors, and resource usage in the / proc / pid directory in real time. In conjunction with the intelligent collection of diagnostic data, a diagnostic analysis module is established to achieve real-time monitoring of the BMC and help operators effectively respond to BMC crashes in a timely manner.
[0091] See also Figure 2 As shown, Figure 2 A schematic diagram of a diagnostic data processing flow provided by an embodiment of the present invention; in some embodiments, after transmitting all compressed operating data to a remote storage module, the process further includes:
[0092] The remote storage module is used to perform fault diagnosis on the baseboard management controller based on all compressed operating data.
[0093] It is easy to understand that in addition to faults that the BMC itself can detect, the remote storage module can also analyze and process all compressed operating data to help the BMC diagnose and analyze anomalies that the BMC itself cannot detect. Anomalies that the BMC cannot detect include BMC crashes, process crashes, etc. When the BMC crashes, the BMC is no longer able to operate normally, and fault diagnosis operations such as fault location cannot be performed. At this time, the remote storage module can use all the compressed operating data it has received to perform diagnostic analysis and detect anomalies that the BMC itself has not discovered from the operating data. The specific method for the remote storage module to perform fault diagnosis can be set and adjusted according to actual application requirements, and this application does not make any special restrictions here.
[0094] Specifically, the combination of layered compression and remote storage enables effective management and backup of BMCs, particularly diagnostic data during BMC crashes. This addresses issues such as large data volumes, low storage efficiency, and insufficient remote analysis capabilities in BMC fault diagnosis. Layered compression technology optimizes data storage efficiency, while remote storage enables multi-node data synchronization and disaster recovery. Remote storage modules are used for fault diagnosis, supporting rapid fault location and recovery in the event of a BMC crash.
[0095] See also Figure 3 As shown, Figure 3 A schematic diagram of a fault diagnosis architecture of a remote storage module provided in an embodiment of the present invention; in some embodiments, the specific process of the remote storage module performing fault diagnosis on a baseboard management controller based on all compressed operating data includes:
[0096] Inputting all compressed operation data into a pre-built rule engine so that the rule engine can determine whether there is any abnormal process operation of the baseboard management controller based on the process status parameters and hardware parameters in the operation data;
[0097] The abnormal process operation includes the current running process reaching the process crash threshold and the current running process having a memory leak;
[0098] If the rule engine determines that a process of the baseboard management controller is running abnormally, the abnormal process causing the abnormal process is determined and the abnormal process is restarted;
[0099] Inputting all compressed operating data into a pre-built machine learning model so that the machine learning model can determine whether the baseboard management controller has a crash risk based on process status parameters, hardware parameters, and log data in the operating data;
[0100] If the machine learning model determines that the baseboard management controller is at risk of crashing, the baseboard management controller is controlled to be reset.
[0101] It is understood that in order to achieve automated diagnosis of the BMC based on compressed operating data, a rule engine and / or machine learning model can be pre-configured in the remote storage module. The rule engine pre-defines a fault rule base, which includes pre-defined one-to-one matching operating data and known fault modes. By inputting the operating data into the rule engine, abnormal events can be matched in real time based on the received operating data. Known fault modes include process crash thresholds and memory leak modes. The process crash threshold is a rule or condition set to determine whether a process has crashed abnormally. When the operating status or resource usage of a process reaches or exceeds the corresponding preset conditions, the system will determine that the process has reached the process crash threshold, determine that the corresponding process has crashed, and trigger the corresponding processing mechanism. The memory leak mode refers to the memory space dynamically allocated during the program's operation, which is not properly released after use, resulting in a gradual decrease in available memory, which may eventually lead to performance degradation, program crashes, or system instability. Specifically, the memory usage of each process can be determined through process status parameters. If the memory usage of a process shows a state of continuous growth over a long period of time, it can be determined that the process has a memory leak. Machine learning models can predict BMC CPU and / or memory trends based on operational data by training LSTM networks (Long Short-Term Memory Networks), and identify potential crash risks by predicting patterns, trends, and developments in CPU and memory resource usage over time. This application does not specifically limit the specific types and implementations of the rule engine and machine learning models. Other methods can also be used in the remote storage module to implement BMC crash prediction and fault analysis.
[0102] Furthermore, the remote storage module can also set up a corresponding connected visual interface to display the results of the diagnostic analysis through the Web (WorldWide Web) terminal and other means, supporting log keyword filtering, sensor data timing analysis and fault tree derivation. After the diagnostic analysis, the diagnostic results are set to trigger an automated response operation. For example, if there is an abnormal operation of the process, the process can be controlled to restart. If there is a risk of crashing the BMC, the BMC can be controlled to reset. The specific operations of the automated response can be set according to the actual application situation and the type of diagnostic results, and this application does not make any special restrictions here. When there are abnormalities such as process operation abnormalities and / or BMC crash risks, the operator can also be notified through SNMPTrap (Simple Network Management Protocol Trap) or email.
[0103] Specifically, the remote storage module can use multi-dimensional operation data and process-level monitoring data to perform diagnostic analysis on the BMC by setting up a rule engine and / or a machine learning model, effectively identify possible anomalies that the BMC cannot identify itself, predict the risk of BMC crash, and realize diagnostic analysis of the BMC while achieving diagnostic data synchronization and disaster recovery, further ensuring the reliable operation of the BMC.
[0104] See also Figure 4 As shown, Figure 4 A schematic structural diagram of a diagnostic data management device provided by an embodiment of the present invention; to solve the above technical problems, an embodiment of the present invention further provides a diagnostic data management device, comprising:
[0105] The data acquisition unit 11 is used to acquire the operation data of the baseboard management controller; the operation data includes at least one of the process status parameters, hardware parameters and log data of the baseboard management controller;
[0106] The classification unit 12 is used to classify the operation data based on the importance of the operation data, and divide all the operation data into a plurality of data groups corresponding to a plurality of preset categories;
[0107] A compression unit 13 is configured to compress the data in each data group using a preset compression strategy; wherein the preset compression strategy includes a plurality of preset compression methods corresponding to the plurality of data groups;
[0108] The transmission unit 14 is used to transmit all compressed running data to the remote storage module.
[0109] In some embodiments, further comprising:
[0110] A timestamp definition unit, used to determine the acquisition time of each operation data and define the acquisition time as the timestamp corresponding to the operation data;
[0111] A server identifier definition unit, configured to obtain a server identifier of a server to which a baseboard management controller belongs;
[0112] A fault type label definition unit is used to determine whether each operating data has a corresponding fault type, and if so, define the fault type as a fault type label corresponding to the operating data; wherein the fault type is the type of fault that occurs during the operation of the baseboard management controller;
[0113] The index building unit is used to use the timestamp, server identifier, and fault type label as index fields to build an index for the compressed running data.
[0114] In some embodiments, the transmission unit 14 includes:
[0115] A priority configuration unit, configured to configure a priority for all compressed operation data based on the importance of the operation data;
[0116] A queue generating unit, configured to sort all compressed running data in descending order of priority to generate a transmission queue;
[0117] The priority transmission unit is used to transmit all compressed running data in sequence according to the order of the transmission queue.
[0118] In some embodiments, the remote storage module includes an edge node module and a cloud disaster recovery module. The physical distance between the edge node module and the baseboard management controller is less than the physical distance between the cloud disaster recovery module and the baseboard management controller. The edge node module and the baseboard management controller are located in the same server architecture. The transmission unit 14 includes:
[0119] The transmission sub-unit is used to transmit all the compressed running data to the edge node module, so that the edge node module transmits all the compressed running data to the cloud disaster recovery module.
[0120] In some embodiments, the process status parameter of the baseboard management controller includes at least one of a process identifier, stack information, and the number of file descriptors; the hardware parameter of the baseboard management controller includes at least one of a central processing unit occupancy rate, a memory usage rate, a power status, a register status, and a firmware version; the log data of the baseboard management controller includes at least one of an operating system log, a network parameter, and driver information; and the data acquisition unit 11 includes:
[0121] The data acquisition subunit is used to collect all the operating data of the baseboard management controller in an asynchronous writing manner.
[0122] In some embodiments, further comprising:
[0123] The diagnosis and analysis unit is used to transmit all compressed operating data to the remote storage module, and then use the remote storage module to perform fault diagnosis on the baseboard management controller based on all compressed operating data.
[0124] In some embodiments, the remote storage module includes:
[0125] A first abnormality determination unit is configured to input all compressed operating data into a pre-built rule engine, so that the rule engine determines whether a process operating abnormality exists in the baseboard management controller based on process status parameters and hardware parameters in the operating data; wherein the process operating abnormality includes the current running process reaching a process crash threshold and the current running process having a memory leak;
[0126] A process restart unit, configured to determine an abnormal process causing the abnormal process operation and restart the abnormal process if the rule engine determines that the baseboard management controller has a process operation abnormality;
[0127] a second abnormality determination unit, configured to input all compressed operating data into a pre-built machine learning model, so that the machine learning model determines whether there is a crash risk of the baseboard management controller based on process status parameters, hardware parameters, and log data in the operating data;
[0128] The controller reset unit is configured to control the baseboard management controller to reset if the machine learning model determines that the baseboard management controller is at risk of crashing.
[0129] For descriptions of features of the diagnostic data management device provided in the embodiments of the present invention, reference may be made to the relevant descriptions of the embodiments of the diagnostic data management method, which will not be described in detail here.
[0130] See also Figure 5 As shown, Figure 5 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. To solve the above technical problems, an embodiment of the present invention further provides an electronic device, including:
[0131] Memory 60, for storing computer programs;
[0132] The processor 61 is configured to execute a computer program to implement the steps of the aforementioned diagnostic data management method.
[0133] The processor 61 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 61 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 61 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 61 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing content required to be displayed on the display screen. In some embodiments, the processor 61 may also include an artificial intelligence (AI) processor for handling computational operations related to machine learning.
[0134] The memory 60 may include one or more computer-readable storage media, which may be non-transitory. The memory 60 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 60 is at least used to store the following computer program 601, wherein, after the computer program is loaded and executed by the processor 61, it can implement the relevant steps of the diagnostic data management method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 60 may also include an operating system 602 and data 603, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 602 may include Windows, Unix, Linux, etc. The data 603 may include but is not limited to the data in the diagnostic data management method, etc.
[0135] In some embodiments, the electronic device may further include a display screen 62 , an input / output interface 63 , a communication interface 64 , a power supply 65 , and a communication bus 66 .
[0136] Those skilled in the art will understand that Figure 5 The structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown in the figure.
[0137] For descriptions of features in the electronic device provided by the embodiments of the present invention, reference may be made to the relevant descriptions of the embodiments of the diagnostic data management method, which will not be described in detail here.
[0138] It is understandable that if the diagnostic data management method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the current technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and executes all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, a magnetic disk, or an optical disk, and other media that can store program code.
[0139] To solve the above technical problems, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the diagnostic data management method as described above are implemented.
[0140] For descriptions of features in the computer-readable storage medium provided in the embodiments of the present invention, reference may be made to the relevant descriptions of the embodiments of the diagnostic data management method, which will not be described in detail here.
[0141] An embodiment of the present invention further provides a computer program product, including a computer program / instruction, which implements the steps of the diagnostic data management method described in the above embodiment when executed by a processor.
[0142] For descriptions of features in the computer program product provided by the embodiments of the present invention, reference may be made to the relevant descriptions of the embodiments of the diagnostic data management method, which will not be detailed here.
[0143] The above describes in detail a diagnostic data management method, apparatus, device, and storage medium provided by an embodiment of the present invention. The various embodiments in the specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between the various embodiments can be referenced to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method section.
[0144] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0145] The above is a detailed introduction to a diagnostic data management method, device, equipment and storage medium provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified in several ways, and these improvements and modifications also fall within the scope of protection of the present invention.
Claims
1. A diagnostic data management method, characterized in that: include: Acquiring operation data of a baseboard management controller; the operation data including at least one of process status parameters, hardware parameters and log data of the baseboard management controller; Classifying the operating data based on the importance of the operating data, and dividing all the operating data into a plurality of data groups corresponding to a plurality of preset categories; Using a preset compression strategy to compress the data in each of the data groups; wherein the preset compression strategy includes a plurality of preset compression methods corresponding to the plurality of data groups; Transmit all the compressed running data to the remote storage module.
2. The diagnostic data management method according to claim 1, wherein: After compressing the data in each of the data groups using a preset compression strategy, the method further includes: Determine a time identifier corresponding to each of the operation data, and define the time identifier as a timestamp corresponding to the operation data; Obtaining a server identifier of a server to which the baseboard management controller belongs; Determine whether each of the operating data has a corresponding fault type, and if so, define the fault type as a fault type label corresponding to the operating data; wherein the fault type is a type of fault that occurs when the baseboard management controller is running; The timestamp, server identifier, and fault type tag are used as index fields to construct an index for the compressed operating data.
3. The diagnostic data management method according to claim 1, wherein: Transmitting all the compressed running data to the remote storage module includes: assigning priorities to all the compressed operating data based on the importance of the operating data; Sorting all the compressed running data in descending order of priority to generate a transmission queue; All the compressed running data are transmitted in sequence according to the order of the transmission queue.
4. The diagnostic data management method according to claim 1, wherein: The remote storage module includes an edge node module and a cloud disaster recovery module, the physical distance between the edge node module and the baseboard management controller is smaller than the physical distance between the cloud disaster recovery module and the baseboard management controller, and the edge node module and the baseboard management controller are in the same server architecture; Transmitting all the compressed running data to the remote storage module includes: All the compressed running data are transmitted to the edge node module, so that the edge node module transmits all the compressed running data to the cloud disaster recovery module.
5. The diagnostic data management method according to claim 1, wherein: The process status parameter of the baseboard management controller includes at least one of a process identifier, stack information, and the number of file descriptors; the hardware parameter of the baseboard management controller includes at least one of a CPU occupancy rate, a memory usage rate, a power status, a register status, and a firmware version; The log data of the baseboard management controller includes at least one of an operating system log, network parameters and driver information; Obtain baseboard management controller operation data, including: All operating data of the baseboard management controller is collected in an asynchronous writing manner.
6. The diagnostic data management method according to any one of claims 1 to 5, characterized in that: After transmitting all the compressed running data to the remote storage module, the method further includes: The remote storage module is used to perform fault diagnosis on the baseboard management controller based on all the compressed operating data.
7. The diagnostic data management method according to claim 6, wherein: The specific process of the remote storage module performing fault diagnosis on the baseboard management controller based on all the compressed operating data includes: Inputting all the compressed operating data into a pre-built rule engine, so that the rule engine determines whether there is any process operation abnormality in the baseboard management controller based on the process status parameters and hardware parameters in the operating data; The process operation abnormality includes the current running process reaching the process crash threshold and the current running process having a memory leak; If the rule engine determines that a process of the baseboard management controller is running abnormally, determining the abnormal process that causes the abnormal process, and restarting the abnormal process; Inputting all the compressed operating data into a pre-built machine learning model, so that the machine learning model determines whether the baseboard management controller has a crash risk based on process status parameters, hardware parameters, and log data in the operating data; If the machine learning model determines that the baseboard management controller has a crash risk, the baseboard management controller is controlled to reset.
8. A diagnostic data management device, characterized in that: include: A data acquisition unit, configured to acquire operating data of a baseboard management controller; the operating data comprising at least one of process status parameters, hardware parameters, and log data of the baseboard management controller; a classification unit, configured to classify the operation data based on the importance of the operation data, and divide all the operation data into a plurality of data groups corresponding to a plurality of preset categories; A compression unit, configured to compress the data in each of the data groups using a preset compression strategy; wherein the preset compression strategy includes a plurality of preset compression modes corresponding to the plurality of data groups; The transmission unit is used to transmit all the compressed running data to the remote storage module.
9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to execute the computer program to implement the steps of the diagnostic data management method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the diagnostic data management method according to any one of claims 1 to 7 are implemented.