Log collection method and device, electronic equipment and storage medium

By dispersedly collecting log files of the EBOF expansion cabinet based on fault alarm information when the storage system fails, the problem of lack of flexibility and dynamic adjustment of log collection methods in the existing technology is solved, and resource conservation and system stability are improved.

CN120492301APending Publication Date: 2025-08-15INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510637636.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, the log file collection method of the EBOF expansion cabinet lacks flexibility and dynamic adjustment capabilities, resulting in increased system resource consumption, and centralized management methods are susceptible to central node failures, increasing the risk of data leakage.

Method used

When the storage system fails, the log collection request is determined based on the fault alarm information, and the request is issued to the chassis management module. The distributed storage expansion cabinet controller collects and uploads the log files to be uploaded to the corresponding node to ensure that the log files are stored in the target directory for users to view.

Benefits of technology

It improves the flexibility and dynamic adjustment capabilities of log collection, reduces system resource consumption, ensures system stability and reduces data leakage risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492301A_ABST
    Figure CN120492301A_ABST
Patent Text Reader

Abstract

The invention discloses a log collection method and device, electronic equipment and a storage medium, and relates to the technical field of computers, and the method comprises the steps that under the condition that a storage system breaks down, a user determines a log collection request based on fault alarm information of the storage system, and issues the log collection request to a case management module in a storage main cabinet through a server, and the case management module receives the log collection request, and issues the log collection request to the storage extension cabinet controllers corresponding to all nodes in the storage main cabinet, so as to trigger to collect a to-be-uploaded log file matched with demand information in the log collection request in the storage extension cabinet controllers. And collecting the corresponding to-be-uploaded log file based on the demand information in the log collection request, thereby solving the technical problem that the log collection mode in the related technology is lack of flexibility and dynamic adjustment capability and increases the system resource consumption, and achieving the technical effects of improving the log collection flexibility and dynamic adjustment capability and reducing the system resource consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a log collection method, device, electronic device, and storage medium. Background Art

[0002] In the storage system field, Ethernet-attached Bunch of Flash (EBOF) expansion cabinets, with their superior performance and flexible scalability, have become a core component of data storage infrastructure. EBOF expansion cabinet log files fully record critical system operation data. Whether quickly locating the root cause of system failures or implementing performance tuning strategies during daily operations and maintenance, these log files play an irreplaceable and critical role in ensuring the efficient and stable operation of storage systems.

[0003] In the related art, the log file collection method of the EBOF expansion cabinet is: collecting all log files of the EBOF expansion cabinet based on a simple trigger mechanism such as fixed time interval triggering or simple event triggering, which lacks flexibility and dynamic adjustment capabilities and increases system resource consumption. Summary of the Invention

[0004] The present application provides a log collection method, device, electronic device and storage medium to at least solve the problem that the log collection method in the related art lacks flexibility and dynamic adjustment capability, thereby increasing system resource consumption.

[0005] This application provides a log collection method, which is applied to a chassis management module in a storage main cabinet, including:

[0006] receiving a log collection request, wherein, in the event of a storage system failure, a user determines a log collection request based on the storage system's fault alarm information, and sends the log collection request to the chassis management module through the server, the log collection request including at least one required information of a target log type, a target log file, and a target storage expansion cabinet;

[0007] Sending a log collection request to the storage expansion cabinet controllers corresponding to all nodes in the storage main cabinet, so that the storage expansion cabinet controllers determine the log files to be uploaded based on the demand information in the log collection request and upload the log files to their corresponding nodes;

[0008] Collect the log files to be uploaded from all nodes and store them in the target directory for users to view.

[0009] The present application also provides a log collection device, which is applied to a chassis management module in a storage main cabinet, comprising:

[0010] A receiving module is configured to receive a log collection request, wherein, in the event of a storage system failure, a user determines a log collection request based on the storage system's fault alarm information, and sends the log collection request to the chassis management module via the server, the log collection request including at least one required information of a target log type, a target log file, and a target storage expansion cabinet;

[0011] A sending module is used to send the log collection request to the storage expansion cabinet controllers corresponding to all nodes in the storage main cabinet, so that the storage expansion cabinet controllers determine the log files to be uploaded based on the demand information in the log collection request and upload the log files to their corresponding nodes;

[0012] The collection module is used to collect log files to be uploaded from all nodes and store the log files to be uploaded in the target directory for users to view.

[0013] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned log collection methods when executing the computer program.

[0014] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned log collection methods are implemented.

[0015] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned log collection methods when executed by a processor.

[0016] Through this application, in the event of a storage system failure, the user determines a log collection request based on the storage system's fault alarm information, and sends the log collection request to the chassis management module in the storage main cabinet through the server. The chassis management module receives the log collection request and sends the log collection request to the storage expansion cabinet controllers corresponding to all nodes in the storage main cabinet, thereby triggering the collection of log files to be uploaded in the storage expansion cabinet controllers that match the demand information in the log collection request. Collecting the corresponding log files to be uploaded based on the demand information in the log collection request solves the technical problem that the log collection method in the related art lacks flexibility and dynamic adjustment capabilities, which increases system resource consumption, and achieves the technical effect of improving log collection flexibility and dynamic adjustment capabilities, and reducing system resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 A schematic diagram of the structure of the log collection system provided in an embodiment of the present application;

[0019] Figure 2 A flow chart of a log collection method provided in an embodiment of the present application;

[0020] Figure 3 A flow chart of another log collection method provided in an embodiment of the present application;

[0021] Figure 4 Schematic diagram of the log collection request issuance process provided in the embodiment of the present application;

[0022] Figure 5 A schematic diagram of the log collection process provided in an embodiment of the present application;

[0023] Figure 6 A block diagram of the log collection device provided in an embodiment of the present application;

[0024] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0026] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0027] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0028] In the storage system field, Just a Bunch of Disks / Drives (JBOD), Just a Bunch of Flash (JBOF), and EBOF are three common storage expansion technologies. JBOD primarily relies on the Open Storage Engineering Suite (OSES) firmware for basic management of the JBOD expansion cabinet. It combines multiple independent disks to easily expand storage capacity, with a relatively simple management model. JBOF, managed by Platform Services for X86 (PSX), offers advanced functionality compared to JBOD and can meet storage requirements in specific scenarios. EBOF expansion cabinets are equipped with a complete independent system, including firmware for various hardware management chips, such as the Basic Input / Output System (BIOS), Baseboard Management Controller (BMC), and Complex Programmable Logic Device (CPLD). These firmware work together to provide EBOF expansion cabinets with more powerful capabilities.

[0029] With its superior performance and flexible scalability, the EBOF expansion cabinet has become a core component of data storage infrastructure. Its log files fully record critical system operation data. Whether quickly locating the root cause of a system failure or implementing performance tuning strategies during daily operations and maintenance, these log files play an irreplaceable and critical role, ensuring the efficient and stable operation of the storage system.

[0030] Because EBOF expansion cabinets are independent and complete systems, the content and scenarios of EBOF expansion cabinet log collection are more complex when an EBOF expansion cabinet fails. Related technologies use simple triggering techniques, such as fixed time interval triggering or simple event triggering, to collect all EBOF expansion cabinet log files. This lacks flexibility and dynamic adjustment capabilities, and can lead to frequent triggering of log file collection tasks at unnecessary times, increasing system resource consumption.

[0031] Furthermore, the log collection method for EBOF expansion cabinets in related technologies uses centralized management. Under this centralized management approach, any problem with the central node will affect the operation of the entire log collection system. As the volume of logs increases, the processing capacity of the central node may become a bottleneck, and the centralized storage of all log files increases the risk of data leakage.

[0032] In response to the above problems, the embodiments of the present application provide a log collection method, device, electronic device and storage medium, which is applied to the chassis management module in the storage main cabinet, including: receiving a log collection request, wherein, in the event of a storage system failure, the user determines the log collection request based on the fault alarm information of the storage system, and sends the log collection request to the chassis management module through the server, the log collection request including the target log type, the target log file and at least one requirement information of the target storage expansion cabinet; sending the log collection request to the storage expansion cabinet controller corresponding to all nodes in the storage main cabinet, so that the storage expansion cabinet controller determines the log file to be uploaded based on the requirement information in the log collection request, and uploads the log file to be uploaded to its corresponding node; collecting the log files to be uploaded from all nodes, and storing the log files to be uploaded in the target directory for the user to view. The method provided by the above scheme solves the technical problem that the log collection method in the related art lacks flexibility and dynamic adjustment capability, which increases system resource consumption, by collecting the corresponding log files to be uploaded based on the requirement information in the log collection request in the event of a storage system failure, and achieves the technical effect of improving log collection flexibility and dynamic adjustment capability, and reducing system resource consumption. By distributing and storing the log files to be uploaded in different nodes, the stability of the system operation is ensured and the risk of data leakage is reduced.

[0033] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the log collection method depends, the specific application environment architecture or specific hardware architecture is described here.

[0034] The log collection method, device, electronic device and storage medium provided in the embodiments of the present application are suitable for collecting log files of storage expansion cabinets. Figure 1As shown, this is a structural diagram of the log collection system on which the present application is based, and the log collection system includes a server and a storage system, wherein the storage system includes a storage main cabinet and a storage expansion cabinet. The storage system includes at least one storage expansion cabinet. In the event of a failure in the storage system, the user determines a log collection request based on the fault alarm information of the storage system, and sends the log collection request to the storage main cabinet through the server. The storage main cabinet receives the log collection request and sends the log collection request to the chassis management module in the storage main cabinet. The chassis management module sends the log collection request to the storage expansion cabinet controller corresponding to all nodes in the storage main cabinet, so that the storage expansion cabinet controller determines the log file to be uploaded based on the demand information in the log collection request, and uploads the log file to be uploaded to its corresponding node. The chassis management module collects the log files to be uploaded from all nodes, and stores the log files to be uploaded in the target directory for the user to view.

[0035] The embodiment of the present application provides a log collection method, which is applied to the chassis management module in the storage main cabinet. The chassis management module is the storage system chassis management component (Enclosure Management, abbreviated as: EN). Figure 2 A flow chart of the log collection method provided in the embodiment of the present application is shown as follows: Figure 2 As shown, the log collection method includes the following steps:

[0036] Step S201, receiving a log collection request, wherein, in the event of a storage system failure, the user determines a log collection request based on the storage system's fault alarm information, and sends the log collection request to the chassis management module through the server. The log collection request includes at least one requirement information of the target log type, the target log file, and the target storage expansion cabinet.

[0037] Among them, the user sends a log collection request to the storage main cabinet through the server, the storage main cabinet receives the log collection request, and sends the log collection request to the chassis management module. The chassis management module detects the log collection request, that is, the chassis management module receives the log collection request, traverses all nodes in the storage main cabinet, and sends the log collection request to the storage expansion cabinet controllers corresponding to all nodes.

[0038] It is understandable that the storage system supports the collection of EBOF expansion cabinet log files to facilitate troubleshooting and locating problems when EBOF expansion cabinets have problems. The maximum capacity of each log file can be 100 MB.

[0039] The target log type indicates the type of log files to be collected, the target log file indicates which log files to collect, and the target storage enclosure indicates the storage enclosure from which log files to be collected. The target log type can be default, dumps, or livedump. The target log file can be a dump file specified using the -f parameter, and the target storage enclosure can be specified using the enclosure_id (storage enclosure identifier).

[0040] Exemplarily, the log collection request includes a target log type and a target storage expansion cabinet. If the target log type is a default log type, the log collection request is:

[0041] #Trigger regular log collection (default type)

[0042] mcsop triggerenclosuredump-enclosure<enclosure_id>

[0043] In another exemplary embodiment, the log collection request includes a target log file and a target storage expansion cabinet. The target log file is a dump file. Then, the log collection request is:

[0044] #Collect the specified dump file

[0045] mcsop triggerenclosuredump-dump-file <filename>-enclosure<enclosure_id>

[0046] In another example, the log collection request includes a target log type and a target storage expansion cabinet. If the target log type is a livedump type, the log collection request is:

[0047] #Trigger live memory dump (livedump)

[0048] mcsop triggerenclosuredump-livedump-enclosure<enclosure_id>

[0049] It should be noted that the storage expansion cabinet in the embodiment of the present application is an EBOF expansion cabinet.

[0050] Step S202: Send the log collection request to the storage expansion cabinet controllers corresponding to all nodes in the storage main cabinet, so that the storage expansion cabinet controllers determine the log files to be uploaded based on the demand information in the log collection request and upload the log files to their corresponding nodes.

[0051] The log file of the storage expansion cabinet is generated by the storage expansion cabinet controller, and thus the log file of the storage expansion cabinet needs to be obtained from the storage expansion cabinet controller.

[0052] The storage expansion enclosure controller corresponding to a node is the storage expansion enclosure controller to which the node is connected.

[0053] The chassis management module traverses all nodes in the storage main cabinet and sends log collection requests to the storage expansion cabinet controllers corresponding to all nodes in turn. After receiving the log collection request, the storage expansion cabinet controller begins to prepare the log files to be uploaded for the storage expansion cabinet generated by the storage expansion cabinet controller itself based on the log collection request, ensuring that the log files to be uploaded are ready. After the log files to be uploaded are ready, the storage expansion cabinet controller uploads the log files to be uploaded to the first directory of the node of the storage main cabinet connected to the storage expansion cabinet controller through the connection link between the storage main cabinet and the storage expansion cabinet controller for storage, where the first directory is / dumps / triggerenclosuredumps.

[0054] Step S203: Collect the log files to be uploaded from all nodes and store them in the target directory for users to view.

[0055] The chassis management module collects log files to be uploaded from the / dumps / triggerenclosuredumps directory of all nodes and stores them in the chassis management module's target directory. Users can view the corresponding log files in the chassis management module's target directory to determine the cause of the current storage system failure. The target directory can be / dumps / .

[0056] The log collection method provided by the embodiment of the present application collects the corresponding log files to be uploaded based on the demand information in the log collection request in the event of a storage system failure. The log collection task is performed only when necessary, that is, when the storage system fails. The log collection method collects the corresponding log files to be uploaded based on the demand information in the log collection request, which solves the technical problem that the log collection method in the related art lacks flexibility and dynamic adjustment capability, and increases the consumption of system resources. The method achieves the technical effect of improving the flexibility and dynamic adjustment capability of log collection and reducing the consumption of system resources. By distributing and storing the log files to be uploaded in different nodes, the stability of the system operation is guaranteed and the risk of data leakage is reduced. The log collection command is issued by the storage main cabinet, specifying the target log type, target log file and at least one demand information in the target storage expansion cabinet, ensuring that the EBOF expansion cabinet controllers corresponding to all nodes can be detected during the log collection process, so that the EBOF expansion cabinet controllers prepare and upload the corresponding log files, and the chassis management module collects the corresponding log files, thereby realizing unified management and efficient collection of the log files of the EBOF expansion cabinet.

[0057] The embodiment of the present application provides a log collection method, which is applied to a chassis management module in a storage main cabinet. Figure 3 A flow chart of the log collection method provided in the embodiment of the present application is shown as follows: Figure 3 As shown, the log collection method includes the following steps:

[0058] Step S301: Receive a log collection request. In the event of a storage system failure, the user determines a log collection request based on the storage system's fault alarm information and sends the log collection request to the chassis management module via the server. The log collection request includes at least one required information about the target log type, target log file, and target storage expansion cabinet. For details, see Figure 2 Step S201 of the illustrated embodiment will not be described in detail here.

[0059] Step S302: Send the log collection request to the storage expansion cabinet controllers corresponding to all nodes in the storage main cabinet, so that the storage expansion cabinet controllers determine the log files to be uploaded based on the demand information in the log collection request and upload the log files to their corresponding nodes.

[0060] Specifically, the above step S302 includes:

[0061] Step S3021: Determine the current node to be traversed in the storage main cabinet.

[0062] Step S3022, obtaining the status information of the current node to be traversed. If the status information of the current node to be traversed indicates that the current node to be traversed is offline, skipping the current node to be traversed and returning to the step of determining the current node to be traversed in the storage main cabinet.

[0063] It is understandable that if the current node to be traversed is in an offline state, the current node to be traversed is skipped and the next node among the untraversed nodes in the storage main cabinet is traversed.

[0064] Step S3023: If the status information of the current node to be traversed indicates that the current node to be traversed is in an online state, determine the storage expansion cabinet controller corresponding to the current node to be traversed, send the log collection request to the storage expansion cabinet controller corresponding to the current node to be traversed, and return to execute the step of determining the current node to be traversed in the storage main cabinet until all nodes in the storage main cabinet are traversed.

[0065] The chassis management module traverses all nodes in the entire storage main cabinet using the en_csm_vrt_statesave_prepare_dumps_run function and sends log collection requests to the node-visible EBOF expansion cabinet controller, i.e., the storage expansion cabinet controller corresponding to the node, using the en_csm_agent_task_start_statesave_force function. The log collection request may include one or more of the chassis index and controller index corresponding to the aforementioned target storage expansion cabinet, the target log type, and the target log file.

[0066] Figure 4 Schematic diagram of the log collection request sending process provided by the embodiment of this application. Figure 4 As shown, a storage main cabinet including four nodes is used as an example for description. The user sends a log collection request, that is, the user sends the log collection request to the EN through the server. At this time, the log collection progress of each node group is initialized to 0%.

[0067] EN detects and distributes requests, initializes the node variable to 1, that is, the current traversed node is the first node, and determines whether the node variable is less than or equal to 4. If the node variable is less than or equal to 4, it determines whether the current Node is online, that is, whether the currently traversed node is online. If the currently traversed node is not online, that is, offline, it updates the node variable and continues to traverse the next node.

[0068] If the currently traversed node is online, the system determines whether the storage expansion cabinet controller corresponding to the currently traversed node is online. If so, the system sends a log collection request to the storage expansion cabinet controller corresponding to the currently traversed node, instructing the storage expansion cabinet controller to prepare the log files to be uploaded and upload them to the corresponding node. If the storage expansion cabinet controller corresponding to the currently traversed node is offline, the system reports an alarm.

[0069] After sending the log collection request to the storage expansion cabinet controller corresponding to the currently traversed node, the node variable is updated and the next node is traversed until the node variable is greater than 4. The current process ends and the sending of the log collection request is completed.

[0070] Step S303: Collect the log files to be uploaded from all nodes and store them in the target directory for users to view. Figure 2 Step S203 of the illustrated embodiment will not be described in detail here.

[0071] The log collection method provided by the embodiments of the present application directly skips offline nodes by determining the node status (online or offline), avoiding unnecessary operations on invalid nodes and thus improving the system's operating efficiency. By issuing log collection requests only to online nodes, resources are properly allocated to available nodes, reducing meaningless computing and communication overhead, improving the efficiency and accuracy of log file collection for EBOF expansion cabinets, and providing a strong guarantee for the stable operation of the storage system.

[0072] In some optional implementations, step S303 includes:

[0073] Step a1: After sending the log collection request to the storage expansion cabinet controllers corresponding to all nodes in the storage main cabinet, wait for a preset time period and detect all nodes.

[0074] Among them, the log collection method in the related art is prone to problems such as log missing during the collection process, resulting in incomplete collected log information.

[0075] This application sets a delay waiting and resource detection mechanism during the log collection process to ensure that the uploaded log files are uploaded and the collected log information is complete.

[0076] The preset time period is set by a technician and is not specifically limited here. For example, the preset time period is 30 seconds.

[0077] Step a2: If at least one of all nodes stores the corresponding log file to be uploaded and the log file to be uploaded is not empty, then wait for a preset time period, collect the log file to be uploaded from all nodes, and store the log file to be uploaded in the target directory.

[0078] All nodes are checked, that is, whether the first directory of all nodes stores corresponding non-empty log files to be uploaded. If there is a corresponding log file to be uploaded stored in the first directory of at least one node and the log file to be uploaded is not empty, then delay and wait for a preset time period to ensure that all log files to be uploaded are uploaded to the corresponding nodes and start the log collection process.

[0079] The log collection method provided in the embodiments of the present application ensures that all log files to be uploaded on all nodes have sufficient time to complete the upload process by setting a delay of a preset time period (e.g., 30 seconds) and performing a check on all nodes. This effectively avoids the problem of missing log files due to incomplete upload.

[0080] In some optional implementations, step S303 includes:

[0081] Step b1: traverse and access all nodes to collect log files to be uploaded belonging to the first storage expansion cabinet controller in all nodes.

[0082] Each storage expansion cabinet includes a first storage expansion cabinet controller and a second storage expansion cabinet controller. A log file for each storage expansion cabinet is jointly generated by the first storage expansion cabinet controller and the second storage expansion cabinet controller corresponding to the storage expansion cabinet. The first storage expansion cabinet controller and the second storage expansion cabinet controller included in each storage expansion cabinet are connected to different nodes in the storage main cabinet.

[0083] It can be understood that the log file to be uploaded belonging to the first storage expansion cabinet controller is the log file to be uploaded uploaded to the corresponding node by the first storage expansion cabinet controller of the storage expansion cabinet.

[0084] Step b2: traverse and access all nodes to collect log files to be uploaded belonging to the second storage expansion cabinet controller in all nodes.

[0085] In the process of traversing and visiting all nodes, if there is at least one node whose status information indicates that the node is offline, the offline alarm information of the node is reported, the node is skipped and the log files to be uploaded in the next node are continued to be collected. If there is at least one node where the log files to be uploaded do not exist, the node is skipped and the log files to be uploaded in the next node are continued to be collected.

[0086] It should be noted that after all the log files to be uploaded in any node are collected, all the log files to be uploaded in the node are cleared.

[0087] Step b3: After completing the traversal of all nodes, determine whether the collected log files to be uploaded match the requirement information in the log collection request.

[0088] After completing the traversal of all nodes, that is, after completing steps b1 and b2, it is determined whether the collected log files to be uploaded match the requirement information in the log collection request, that is, whether the collected log files to be uploaded meet the log collection requirement.

[0089] In step b4, if the collected log files to be uploaded match the requirement information in the log collection request, the log files to be uploaded are stored in the target directory.

[0090] In step b5, if the collected log files to be uploaded do not match the requirement information in the log collection request, a log collection alarm message is reported. The log collection alarm message is used to indicate that the collected log files do not match the requirement information in the log collection request.

[0091] In related technologies, node offline problems are prone to occur, resulting in incomplete collected log information. This application reports log collection alarm information when the node is offline and the collected logs are incomplete, so that the user can determine that the log collection is incomplete based on the log collection alarm information, and then collect the corresponding uncollected logs again to ensure that complete logs are collected.

[0092] The log collection alarm information must clearly indicate which controller of the EBOF expansion enclosure that matches the requirements in the log collection request has not had its log files collected.

[0093] The log collection method provided in the embodiment of the present application performs a requirement matching judgment on the collected log files to be uploaded after completing the traversal of all nodes. This step ensures that the log files ultimately collected meet the requirement information in the log collection request, improving the accuracy and reliability of log collection. In the event of a mismatch, a log collection alarm message is promptly reported, clearly indicating the log files that have not been collected and match the requirement information in the log collection request. This alarm mechanism helps to quickly locate problems and take appropriate measures, further enhancing the robustness of the system.

[0094] In some optional implementations, the log collection method further includes:

[0095] Step c1, in the process of collecting log files to be uploaded from all nodes, if the time for collecting the log files to be uploaded from any node exceeds the preset time threshold and the collection of the log files to be uploaded of the node is not completed, then stop collecting the log files to be uploaded of the node, determine that the log file collection has failed, report the log file collection failure alarm information, and set an interruption flag file, which includes the collection interruption position of the log files to be uploaded.

[0096] The preset time threshold is set by a technician and is not specifically limited here. For example, the preset time threshold is 300 seconds.

[0097] If the time taken to collect log files from any node exceeds 300 seconds and the log files are not collected from that node, the log file collection is considered to have failed, a log file collection failure alarm is reported, and an interrupt flag file is set. The log file collection failure alarm can be recorded in the / var / log / enclosure / error_log directory.

[0098] In step c2, after the user collects the failure alarm information based on the log file and locates and repairs the abnormality of the storage system, the user continues to collect the log files to be uploaded based on the interruption mark file.

[0099] If the log collection time for any node exceeds a preset threshold, the storage system will generate a corresponding alarm log, possibly due to insufficient directory space on the node, an abnormal node restart or failure, or a storage device anomaly. After the user locates and repairs the storage system anomaly based on the alarm log, they can continue collecting log files to be uploaded based on the interruption flag file. In this embodiment, the alarm log is an alarm indicating that log file collection failed.

[0100] That is to say, the log collection method in the embodiment of the present application supports breakpoint resumption during the log collection process.

[0101] Furthermore, when the storage system anomaly is a storage device anomaly, the log file collection failure alarm information indicates the storage device anomaly. After the user locates and repairs the storage device anomaly based on the alarm log under the storage system, the log files to be uploaded continue to be collected based on the interrupt flag file.

[0102] When the storage system anomaly is an abnormal restart or failure of a node, the log file collection failure alarm information is used to indicate that the node has an abnormal restart or failure. After the user locates and repairs the node anomaly based on the alarm log under the storage system, the log files to be uploaded continue to be collected based on the interruption flag file.

[0103] The log collection method provided in the embodiment of the present application improves the flexibility of log collection and ensures the stability and reliability of the log collection process by providing corresponding fault handling mechanisms for abnormal situations such as collection timeouts, such as resuming transmission from breakpoints, recording alarm logs, and continuing collection after recovery.

[0104] In some optional implementations, the log collection method further includes:

[0105] Step d1, in the process of collecting log files to be uploaded from all nodes, obtain the total number of nodes in each node group and the number of collected nodes, wherein all nodes in the storage main cabinet are divided into multiple node groups.

[0106] It can be understood that the collected node is a node that has collected all the log files to be uploaded stored therein.

[0107] Step d2: For any node group, based on the total number of nodes in the node group and the number of collected nodes, determine the log collection progress of the node group.

[0108] The log collection progress of the node group = the number of nodes in the node group that have collected data / the total number of nodes in the node group × 100%. The initial collection progress of each node group is 0%.

[0109] Step d3: determining the average log collection progress of all node groups based on the log collection progress of each node group in all node groups.

[0110] Add the log collection progress of all node groups and divide it by the total number of node groups to determine the average log collection progress of all node groups.

[0111] In step d4, the log collection progress of each node group and the average log collection progress of all node groups are reported to the server, so that the user can view the log collection progress of each node group and the average log collection progress of all node groups through the server to determine the log collection status.

[0112] It is understandable that, in the process of collecting the log files to be uploaded from all nodes, the log collection progress of each node group and the average log collection progress of all node groups are obtained and updated in real time.

[0113] The monitoring and updating of log collection progress in related technologies is not timely and accurate, which brings additional management burden to system administrators. The log collection method provided by the embodiment of the present application updates the log collection progress of each node group in real time during the log collection process based on the ratio of the number of collected nodes in each node group to the total number of nodes, and provides statistical information on the average log collection progress, so that users can understand the log collection status in a timely manner.

[0114] In some optional implementations, step S303 includes:

[0115] In step e1, for any log file to be uploaded, a preset hash function is used to determine a first hash value of the log file to be uploaded.

[0116] The preset hash function is the Message-Digest Algorithm 5 (MD5).

[0117] Step e2: Obtain a second hash value of the log file to be uploaded from the storage expansion cabinet controller to which the log file to be uploaded belongs.

[0118] It is understandable that the second hash value of the log file to be uploaded stored in the storage expansion cabinet controller is pre-calculated and stored using a preset hash function for the log file to be uploaded, and is used to compare with the first hash value to determine the integrity of the log file to be uploaded.

[0119] Step e3: Compare the first hash value and the second hash value.

[0120] In step e4, when the first hash value and the second hash value are consistent, it is determined that the log file to be uploaded is complete.

[0121] The first hash value and the second hash value are consistent, indicating that the log file to be uploaded has not been tampered with or damaged, and its integrity has been verified.

[0122] In step e5, when all the log files to be uploaded are complete, all the log files to be uploaded are compressed into a log file package to be uploaded, and the log file package to be uploaded is stored in the target directory.

[0123] The log file package to be uploaded can be enclosure_ <id> _ <timestamp>.tar.gz.

[0124] The log collection method provided by the embodiment of the present application ensures the integrity and accuracy of the log files through integrity verification. By compressing all complete log files to be uploaded into one file package, the storage space occupied is reduced, which facilitates subsequent analysis and management.

[0125] In some optional implementations, step S303 includes:

[0126] Step f1, collects encrypted log files to be uploaded from all nodes, wherein the storage expansion cabinet controller determines the log files to be uploaded based on the demand information in the log collection request, encrypts the log files to be uploaded using an encryption algorithm, and uploads the encrypted log files to be uploaded to their corresponding nodes.

[0127] In step f2, the encrypted log file to be uploaded is re-encrypted at the storage level, and the re-encrypted log file to be uploaded is stored in the target directory for the user to view.

[0128] It is understood that after performing storage-level decryption and decryption corresponding to the encryption algorithm on the re-encrypted log file to be uploaded, the user obtains the original log file to be uploaded, ensuring that only authorized users can access the log file to be uploaded. Users who can decrypt the re-encrypted log file to be uploaded and obtain the original log file to be uploaded are authorized users.

[0129] The log collection method provided in the embodiment of the present application uses secondary encryption. Even if the encrypted log file is obtained by unauthorized users, they cannot decrypt the log file content, thereby protecting the data from illegal access, reducing the risk of data leakage, and improving overall data security and system reliability.

[0130] In order to make the log collection process of this application clearer, Figure 5 Describe the log collection process. Figure 5 The flow chart for log collection provided in the embodiment of the present application is as follows: Figure 5 As shown in the figure, the storage system includes a storage expansion cabinet, the storage main cabinet includes 4 nodes, each storage expansion cabinet includes two storage expansion cabinet controllers, and both storage expansion cabinet controllers have log files that meet the requirements of the log collection request. The process includes:

[0131] If all nodes pass the test, a delay of 30 seconds is performed, and the node variable is initialized to 1. The storage expansion cabinet controller variable is initialized to 1, indicating that the first node is currently being traversed and the log file to be uploaded for the first storage expansion cabinet controller (canister) in the storage expansion cabinet is being retrieved. If all nodes pass the test as described above, it means that at least one of all nodes stores the corresponding log file to be uploaded and that the log file to be uploaded is not empty.

[0132] Determine whether the current storage expansion cabinet controller variable is less than or equal to 2. If so, determine whether the current node variable is less than or equal to 4. If the node variable is less than or equal to 4, determine whether the currently traversed node is online. If the currently traversed node is not online, update the node variable (Node++), report an alarm, update the log collection progress, and return to the step of determining whether the current node variable is less than or equal to 4. If the node variable is greater than 4, update the storage expansion cabinet controller variable (canister++) to 1, and return to the step of determining whether the current storage expansion cabinet controller variable is less than or equal to 2.

[0133] If the currently traversed node is online, determine whether the current node has a log file, that is, determine whether the currently traversed node stores the log file to be uploaded of the first storage expansion cabinet controller in the storage expansion cabinet. If the currently traversed node does not store the log file to be uploaded of the first storage expansion cabinet controller in the storage expansion cabinet, update the node variable and determine whether the updated node variable is greater than 4. If the updated node variable is not greater than 4, update the log collection progress and return to the step of determining whether the current node variable is less than or equal to 4. If the updated node variable is greater than 4, report an alarm to determine that log collection is incomplete.

[0134] If the currently traversed node stores the log file to be uploaded of the first storage expansion cabinet controller in the storage expansion cabinet, then collect the log, that is, collect the log file to be uploaded in the currently traversed node, and judge whether the log collection is completed within the preset time threshold, that is, judge whether the time for collecting the log file to be uploaded in the currently traversed node exceeds the preset time threshold; if the log collection is not completed within the preset time threshold, set the interruption flag file and report an alarm; after the user handles the storage system exception, continue to collect the log file to be uploaded in the currently traversed node based on the interruption flag file.

[0135] If log collection is completed within the preset time threshold, log cleaning is performed on the current node, that is, the log file to be uploaded of the first storage expansion cabinet controller in the currently traversed node is cleaned up.

[0136] Initialize the node variable to 1, update the storage expansion cabinet controller variable, that is, canister++, to indicate that the first node is currently traversed and the log file to be uploaded of the second storage expansion cabinet controller in the storage expansion cabinet is obtained, update the log collection progress, and return to the step of determining whether the current storage expansion cabinet controller variable is less than or equal to 2.

[0137] If the current storage expansion cabinet controller variable is greater than 2, log collection is determined to be complete, and the log files are compressed and packaged and stored in a specific directory.

[0138] The log collection method provided in the embodiment of the present application achieves precise positioning of log collection by sending a log collection request from the storage main cabinet, specifying the target log type, target log file, and at least one required information in the target storage expansion cabinet. After the EN detects the log collection request, it traverses the nodes and sends the log collection request to the storage expansion cabinet controller in sequence, ensuring the comprehensiveness of log collection. After receiving the log collection request, the EBOF expansion cabinet controller prepares and uploads the log file to be uploaded, and realizes efficient transmission of the log file to be uploaded through the connection link between the storage main cabinet and the EBOF expansion cabinet controller.

[0139] It should be noted that the log collection method provided in the embodiment of the present application is also applicable to the log collection of multi-port JBOD products and the log collection of other storage expansion cabinets.

[0140] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0141] The embodiment of the present application also provides a log collection device, which is applied to the chassis management module in the storage main cabinet, such as Figure 6 As shown, the log collection device includes:

[0142] Receiving module 601 is used to receive a log collection request. In the event of a storage system failure, the user determines a log collection request based on the storage system's fault alarm information, and sends the log collection request to the chassis management module through the server. The log collection request includes at least one requirement information of the target log type, the target log file, and the target storage expansion cabinet.

[0143] The sending module 602 is used to send the log collection request to the storage expansion cabinet controller corresponding to all nodes in the storage main cabinet, so that the storage expansion cabinet controller determines the log file to be uploaded based on the demand information in the log collection request and uploads the log file to be uploaded to its corresponding node.

[0144] The collection module 603 is used to collect the log files to be uploaded from all nodes and store the log files to be uploaded in a target directory for users to view.

[0145] In some optional implementations, the sending module 602 includes:

[0146] The first determining unit is configured to determine a current node to be traversed in the storage main cabinet.

[0147] The first acquisition unit is used to obtain the status information of the current node to be traversed. If the status information of the current node to be traversed indicates that the current node to be traversed is in an offline state, the current node to be traversed is skipped and the step of determining the current node to be traversed in the storage main cabinet is returned to.

[0148] The second determination unit is used to determine the storage expansion cabinet controller corresponding to the current node to be traversed if the status information of the current node to be traversed indicates that the current node to be traversed is in an online state, send the log collection request to the storage expansion cabinet controller corresponding to the current node to be traversed, and return to execute the step of determining the current node to be traversed in the storage main cabinet until all nodes in the storage main cabinet are traversed.

[0149] In some optional implementations, the collection module 603 includes:

[0150] The first waiting unit is configured to delay and wait for a preset time period after sending the log collection request to the storage expansion cabinet controllers corresponding to all nodes in the storage main cabinet, and detect all nodes.

[0151] The second waiting unit is used to delay waiting for a preset time period if there is at least one node among all nodes that stores the corresponding log file to be uploaded and the log file to be uploaded is not empty, collect the log files to be uploaded from all nodes, and store the log files to be uploaded in the target directory.

[0152] In some optional implementations, the collection module 603 includes:

[0153] The first collecting unit is configured to traverse and access all nodes, and collect log files to be uploaded belonging to the first storage expansion cabinet controller in all nodes.

[0154] The second collecting unit is configured to traverse and access all nodes, and collect log files to be uploaded belonging to the second storage expansion cabinet controller in all nodes.

[0155] In the process of traversing and visiting all nodes, if there is at least one node whose status information indicates that the node is offline, the offline alarm information of the node is reported, the node is skipped and the log files to be uploaded in the next node are continued to be collected. If there is at least one node where the log files to be uploaded do not exist, the node is skipped and the log files to be uploaded in the next node are continued to be collected.

[0156] The first judgment unit is used to judge whether the collected log files to be uploaded match the requirement information in the log collection request after completing the traversal of all nodes.

[0157] The first storage unit is configured to store the log files to be uploaded in a target directory if the collected log files to be uploaded match the requirement information in the log collection request.

[0158] The first reporting unit is used to report log collection alarm information if the collected log files to be uploaded do not match the requirement information in the log collection request. The log collection alarm information is used to indicate that the collected log files do not match the requirement information in the log collection request.

[0159] Each storage expansion cabinet includes a first storage expansion cabinet controller and a second storage expansion cabinet controller.

[0160] In some optional implementations, the log collection device further includes:

[0161] The third determination unit is used to stop collecting the log files to be uploaded of the node if the time for collecting the log files to be uploaded of the node from any node exceeds a preset time threshold and the collection of the log files to be uploaded of the node is not completed during the process of collecting the log files to be uploaded from all nodes, determine that the log file collection has failed, report the log file collection failure alarm information, and set an interruption flag file, which includes the collection interruption position of the log files to be uploaded.

[0162] The third collecting unit is configured to continue collecting log files to be uploaded based on the interruption flag file after the user locates and repairs the abnormality of the storage system after collecting the failure alarm information based on the log file.

[0163] In some optional implementations, the log collection device further includes:

[0164] The second acquisition unit is used to acquire the total number of nodes in each node group and the number of collected nodes in the process of collecting log files to be uploaded from all nodes, wherein all nodes in the storage main cabinet are divided into multiple node groups.

[0165] The fourth determining unit is configured to determine, for any node group, the log collection progress of the node group based on the total number of nodes in the node group and the number of collected nodes.

[0166] The fifth determining unit is configured to determine an average log collection progress of all node groups based on the log collection progress of each node group in all node groups.

[0167] The second reporting unit is used to report the log collection progress of each node group and the average log collection progress of all node groups to the server, so that the user can view the log collection progress of each node group and the average log collection progress of all node groups through the server to determine the log collection status.

[0168] In some optional implementations, the collection module 603 includes:

[0169] The sixth determining unit is configured to determine, for any log file to be uploaded, a first hash value of the log file to be uploaded by using a preset hash function.

[0170] The third obtaining unit is configured to obtain a second hash value of the log file to be uploaded from the storage expansion cabinet controller to which the log file to be uploaded belongs.

[0171] The first comparison unit is configured to compare the first Hash value and the second Hash value.

[0172] The seventh determining unit is configured to determine that the log file to be uploaded is complete when the first hash value and the second hash value are consistent.

[0173] The second storage unit is used to compress all the log files to be uploaded into a log file package to be uploaded when all the log files to be uploaded are complete, and store the log file package to be uploaded in the target directory.

[0174] For descriptions of features in the embodiments corresponding to the log collection device, please refer to the relevant descriptions of the embodiments corresponding to the log collection method, and will not be repeated here.

[0175] The embodiment of the present application also provides an electronic device, such as Figure 7 As shown, it includes a processor 701 and a memory 702, wherein the memory 702 stores a computer program, and the processor 701 is configured to run the computer program to execute the steps in any of the above-mentioned log collection method embodiments.

[0176] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned log collection method embodiments when running.

[0177] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0178] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned log collection method embodiments are implemented.

[0179] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned log collection method embodiments are implemented.

[0180] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0181] The above is a detailed introduction to a log collection method, device, electronic device and storage medium provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.< / timestamp> < / id> < / filename>

Claims

1. A log collection method, characterized in that: The chassis management module used in the storage main cabinet includes: receiving a log collection request, wherein, in the event of a storage system failure, a user determines a log collection request based on the storage system's fault alarm information, and sends the log collection request to the chassis management module through the server, the log collection request including at least one requirement information of a target log type, a target log file, and a target storage expansion cabinet; Sending the log collection request to the storage expansion cabinet controllers corresponding to all nodes in the storage main cabinet, so that the storage expansion cabinet controllers determine the log files to be uploaded based on the demand information in the log collection request and upload the log files to their corresponding nodes; Collect the log files to be uploaded from all nodes and store them in the target directory for users to view.

2. The log collection method according to claim 1, characterized in that: The sending of the log collection request to the storage expansion cabinet controllers corresponding to all nodes in the storage main cabinet includes: Determine the current node to be traversed in the storage main cabinet; Obtaining status information of the current node to be traversed; if the status information of the current node to be traversed indicates that the current node to be traversed is in an offline state, skipping the current node to be traversed and returning to the step of determining the current node to be traversed in the storage main cabinet; If the status information of the node to be traversed currently indicates that the node to be traversed currently is in an online state, the storage expansion cabinet controller corresponding to the node to be traversed currently is determined, the log collection request is sent to the storage expansion cabinet controller corresponding to the node to be traversed currently, and the step of determining the node to be traversed currently in the storage main cabinet is returned to be executed until all nodes in the storage main cabinet are traversed.

3. The log collection method according to claim 1, wherein: The step of collecting the log files to be uploaded from all nodes and storing the log files to be uploaded in a target directory includes: After sending the log collection request to the storage expansion cabinet controllers corresponding to all nodes in the storage main cabinet, delaying and waiting for a preset time period to detect all nodes; If there is at least one node among all the nodes that stores the corresponding log file to be uploaded and the log file to be uploaded is not empty, then the process will be delayed for a preset period of time, the log file to be uploaded will be collected from all the nodes, and the log file to be uploaded will be stored in the target directory.

4. The log collection method according to claim 1, wherein: The step of collecting the log files to be uploaded from all nodes and storing the log files to be uploaded in a target directory includes: Traverse and access all nodes to collect log files to be uploaded belonging to the first storage expansion cabinet controller in all nodes; Traverse all nodes and collect log files to be uploaded belonging to the second storage expansion cabinet controller in all nodes; In the process of traversing and visiting all nodes, if there is at least one node whose status information indicates that the node is offline, then report the offline alarm information of the node, skip the node and continue to collect the log files to be uploaded in the next node; if there is at least one node where the log files to be uploaded do not exist, then skip the node and continue to collect the log files to be uploaded in the next node; After completing the traversal of all nodes, determine whether the collected log files to be uploaded match the requirement information in the log collection request; If the collected log files to be uploaded match the requirement information in the log collection request, the log files to be uploaded are stored in the target directory; If the collected log files to be uploaded do not match the requirement information in the log collection request, a log collection alarm message is reported, where the log collection alarm message is used to indicate that the log files that match the requirement information in the log collection request have not been collected; Each storage expansion cabinet includes a first storage expansion cabinet controller and a second storage expansion cabinet controller.

5. The log collection method according to claim 1, wherein: The method further comprises: In the process of collecting the log files to be uploaded from all nodes, if the time for collecting the log files to be uploaded from any node exceeds a preset time threshold and the collection of the log files to be uploaded of the node is not completed, the collection of the log files to be uploaded of the node is stopped, the log file collection failure is determined, the log file collection failure alarm information is reported, and an interruption flag file is set, wherein the interruption flag file includes the collection interruption position of the log files to be uploaded; After the user collects the failure alarm information based on the log file and locates and repairs the abnormality of the storage system, the user continues to collect the log files to be uploaded based on the interruption mark file.

6. The log collection method according to claim 1, wherein: The method further comprises: In the process of collecting the log files to be uploaded from all nodes, obtaining the total number of nodes in each node group and the number of collected nodes, wherein all nodes in the storage main cabinet are divided into multiple node groups; For any node group, determine the log collection progress of the node group based on the total number of nodes in the node group and the number of collected nodes; Based on the log collection progress of each node group in all node groups, determine the average log collection progress of all node groups; The log collection progress of each node group and the average log collection progress of all node groups are reported to the server, so that users can view the log collection progress of each node group and the average log collection progress of all node groups through the server to determine the log collection status.

7. The log collection method according to claim 1, wherein: The step of storing the log file to be uploaded in the target directory includes: For any log file to be uploaded, using a preset hash function, determine a first hash value of the log file to be uploaded; Obtaining a second hash value of the log file to be uploaded from the storage expansion cabinet controller to which the log file to be uploaded belongs; comparing the first hash value and the second hash value; If the first hash value and the second hash value are consistent, determining that the log file to be uploaded is complete; When all the log files to be uploaded are complete, all the log files to be uploaded are compressed into a log file package to be uploaded, and the log file package to be uploaded is stored in the target directory.

8. A log collection device, characterized in that: The chassis management module used in the storage main cabinet includes: a receiving module configured to receive a log collection request, wherein, in the event of a storage system failure, a user determines a log collection request based on the storage system's failure alarm information, and sends the log collection request to the chassis management module via the server, wherein the log collection request includes at least one required information of a target log type, a target log file, and a target storage expansion cabinet; a sending module, configured to send the log collection request to the storage expansion cabinet controllers corresponding to all nodes in the storage main cabinet, so that the storage expansion cabinet controllers determine the log files to be uploaded based on the demand information in the log collection request, and upload the log files to be uploaded to the corresponding nodes; The collection module is used to collect the log files to be uploaded from all nodes and store the log files to be uploaded in a target directory for users to view.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the log collection method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the log collection method according to any one of claims 1 to 7 are implemented.