Server service data management system, method and server

By introducing a time-to-live mechanism and a buffer control engine, data is written to the storage hardware after the storage component recovers, solving the business interruption problem caused by the failure of all paths of multi-path software when the storage is silent, and realizing business continuity and data integrity.

CN120670210BActive Publication Date: 2025-11-21INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511178361.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-11-21
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

The problem of upper-layer business interruption caused by the failure of all paths in multipath software when input/output silence occurs in storage.

Method used

An output/output liveness time mechanism and a buffer control engine are introduced to temporarily store data to storage hardware when a path fails, and write the data back after the storage recovers, thus avoiding business interruption caused by the failure of all paths.

Benefits of technology

During the silent period of input/output storage, repeatedly retry and temporarily store timed-out output/output requests to avoid business interruption caused by failure of all paths, improve user experience, and ensure business continuity and data integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670210B_ABST
    Figure CN120670210B_ABST
Patent Text Reader

Abstract

The application discloses a kind of server service data management system, method and server, it is related to computer storage technical field.System includes the storage component for storing service data and the management component of management multiple access paths, service host includes storage hardware, kernel driver engine and buffer control engine.Kernel driver engine is according to the data request generated by upper layer service, by management component selection access path, and record processing duration;When processing duration exceeds survival time threshold value, determine request failure and send buffer instruction to buffer control engine, simultaneously feedback request completion information to upper layer.Buffer control engine stores data to storage hardware temporarily, write again after storage component recovers, solve the technical problem that related technology storage occurs input / output silence, multiple path software all path failure leads to upper layer service interruption, reach the technical effect that guarantee service does not interrupt when storage input / output silence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer storage technology, and in particular to a server business data management system, method and server. Background Technology

[0002] Compared to conventional storage, commercial storage offers advantages such as massive data management, high security, and business continuity. It is equipped with multi-pathing software deployed on the business host, aggregating multiple physical paths of the same storage volume into a single multi-pathing device. This enables automatic path selection and second-level failover, ensuring seamless business operation. In actual business operations, data read / write between the host and storage is completed through input / output requests. The multi-pathing software is responsible for selecting the optimal path for these input / output requests and handling failover.

[0003] Basic input / output processing flow as follows Figure 1 As shown: After receiving a request, the multipath input / output calls the multipath routing interface to select the optimal path, which is then sent to the storage system via the host bus adapter card driver. Successful requests are counted; failures result in retrying or rerouting based on error information, sending the request again via a new path until completion. However, when input / output silence occurs within the storage (e.g., storage controller restart or heartbeat interruption), the input / output request will fail regardless of which link the multipath software selects, causing the multipath software to determine that no path is available, ultimately leading to upper-layer service interruption. Summary of the Invention

[0004] This application provides a server business data management system, method, and server to at least solve the problem in related technologies where the failure of all paths in multipath software leads to the interruption of upper-layer business when the storage experiences input / output silence.

[0005] This application provides a server business data management system, including: a storage component for storing business data; a management component for managing multiple access paths of the storage component; and a business host, which includes storage hardware, a kernel driver engine, and a buffer control engine. The kernel driver engine, based on at least one data processing request generated by the upper-layer business request from the user, sends the data processing request to the management component. The management component responds to the data processing request and selects the target access path of the storage component for the business host. The kernel driver engine records the actual processing time of the data processing request. If the actual processing time is greater than or equal to a time-to-live (TTL) threshold, it determines that the data processing request has failed and sends a buffering instruction to the buffer control engine, feeding back completion information of the data processing request to the upper layer. The buffer control engine responds to the buffering instruction, stores the data corresponding to the data processing request in the storage hardware, and writes the data from the storage hardware back to the storage component after the storage component recovers.

[0006] This application also provides a server, including the aforementioned server business data management system.

[0007] This application also provides a server business data management method, which is applied to the business host of the aforementioned server business data management system. The business host is configured to perform the following steps: Execute the kernel driver engine, which, based on at least one data processing request generated by the upper-layer business according to the user request, sends the data processing request to the management component. The management component responds to the data processing request and selects the target access path of the storage component for the business host; Execute the kernel driver engine, which records the actual processing time of the data processing request. If the actual processing time is greater than or equal to a time-to-live threshold, it determines that the data processing request has failed, sends a buffering instruction to the cache control engine, and feeds back the completion information of the data processing request to the upper layer; Execute the buffer control engine, which responds to the buffering instruction and stores the data corresponding to the data processing request in the storage hardware. After the storage component recovers, it writes the data from the storage hardware back into the storage component.

[0008] This application also provides a non-volatile computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the above-described server business data management method.

[0009] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described server business data management method.

[0010] This application, by introducing an output / output lifetime mechanism and a buffer control engine, sends a buffer command to the buffer control engine when all paths fail, storing the data processing request data in the storage hardware. After the storage component recovers, the data from the storage hardware is written back to the storage component, and completion information for the data processing request is fed back to the upper layer. This allows for repeated retries and temporary storage of timed-out output / output requests during storage output / output silence, preventing business interruptions caused by the failure of all paths and improving the user experience. Therefore, it solves the problem of upper-layer business interruption caused by the failure of all paths in multi-path software when storage input / output silence occurs, achieving the technical effect of ensuring uninterrupted business during storage input / output silence. Attached Figure Description

[0011] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart of multi-path software input / output processing provided for related technologies;

[0013] Figure 2 This is a schematic diagram of the structure of a server business data management system provided in an embodiment of this application;

[0014] Figure 3 This is a schematic diagram of the persistent memory space structure provided in the embodiments of this application;

[0015] Figure 4 A flowchart illustrating a server business data management method provided in an embodiment of this application. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0017] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0018] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] Figure 2 This is a schematic diagram of the structure of the server business data management system provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the server business data management system 10 specifically includes: a storage component 100, a management component 200, and a business host 300.

[0020] The storage component 100 is used to store business data; the management component 200 is used to manage multiple access paths of the storage component 100; the business host 300 includes storage hardware 301, a kernel driver engine 302, and a buffer control engine 303. The kernel driver engine 302 sends at least one data processing request generated by the upper-layer business according to the user request to the management component 200. The management component 200 responds to the data processing request and selects the target access path of the storage component 100 for the business host 300. The kernel driver engine 302 records the actual processing time of the data processing request. If the actual processing time is greater than or equal to the survival time threshold, it determines that the data processing request has failed and sends a buffering instruction to the buffer control engine to feed back the completion information of the data processing request to the upper layer. The buffer control engine 303 responds to the buffering instruction and stores the data corresponding to the data processing request in the storage hardware 301. After the storage component 100 recovers, it writes the data in the storage hardware 301 back into the storage component 100.

[0021] Storage component 100 is the hardware or software system part used to store business data, responsible for actually saving and managing the data. Management component 200 is responsible for managing multiple access paths of storage component 100, implementing path selection and scheduling, and ensuring efficient and reliable access to data requests. Business host 300 is a server running business applications, including components such as storage hardware 301, kernel driver engine 302, and buffer control engine 303. Storage hardware 301 is the physical storage device inside business host 300 used for temporary data caching, such as persistent memory. Kernel driver engine 302 is responsible for receiving data processing requests generated by upper-layer business, calling management component 200 to select a path to execute the request, and recording the actual processing time of the request and determining whether the request timed out or failed. Buffer control engine 303 receives buffering instructions sent by kernel driver engine 302, stores data that was not successfully written to storage component 100 in storage hardware 301 first, and is responsible for writing the data back to storage component 100 after storage component 100 recovers. The time-to-live threshold is a preset maximum allowed data processing time; exceeding this time is considered a data request processing failure. The target access path is the optimal physical or logical path to the storage component 100 selected by the management component 200 based on path status and performance. The buffer instruction is a command issued by the kernel driver engine 302, which instructs the buffer control engine 303 to perform data caching operations.

[0022] It should be noted that the data processing request in this application embodiment can be an input / output request.

[0023] Understandably, when storage component 100 malfunctions or experiences delayed input / output responses, a time-to-live threshold is introduced to determine if a request has timed out. Timed-out request data is temporarily cached in the storage hardware 301 of the business host 300 to prevent business interruption due to storage access failures. Once storage component 100 recovers, the buffer control engine 303 writes back the cached data, ensuring data integrity and consistency, thereby guaranteeing the continuous and stable operation of upper-layer services. The write-back process involves writing data temporarily stored in the cache or memory to the backend storage device.

[0024] In this embodiment of the application, the management component 200 includes a routing interface, a completion interface, and a retry interface. The management component 200 calls the routing interface to select an access path, and the management component 200 calls the completion interface to determine the processing result of the data processing request. If the current access path cannot process the data processing request, the retry interface is called to reselect an access path carrying an unused tag.

[0025] The routing interface, called by the management component 200, selects the optimal access path to process data requests based on a specific algorithm. The completion interface, also called by the management component 200, confirms the processing result of the data request and determines its success. The retry interface, called by the management component 200, selects an unused access path when the current access path fails to process the data request, ensuring the request can be processed. The access path is the data transmission path connecting the business host 300 and the storage component 100; multiple paths may exist for redundancy and load balancing. The unused tag indicates whether an access path has already been tried, ensuring the retry interface prioritizes untried paths, improving retry efficiency and success rate.

[0026] It is understood that, through the routing, completion, and retry interfaces of the management component 200, this application embodiment realizes dynamic management and fault tolerance of storage access paths, ensuring that when a certain path is unavailable, it can be switched to an unused backup path in a timely manner, thereby improving the success rate of data requests and the reliability of the system.

[0027] This embodiment adds an input / output lifetime parameter to the multipath kernel module. When the processing time of an input / output operation does not exceed the preset lifetime, the multipath input / output module will continuously attempt to write that input / output to storage. Even during storage input / output silence periods, when all paths of the volume are unavailable, multipath will not immediately return an input / output failure to the multipath input / output framework. Here, a volume is a logical storage unit provided by the storage component 100 to the service host 300, connected to the host through multiple access paths, and managed by multipath software to ensure access reliability.

[0028] Specifically, the upper-layer business sends input / output requests to the multipath device. The multipath input / output framework calls the path selection interface of the multipath software to select the optimal path to process the input / output, and then calls the input / output completion interface of the multipath software. The multipath kernel driver then determines whether the input / output was successful.

[0029] If input / output processing fails, the multipath kernel driver notifies the multipath input / output framework to retry. The multipath input / output framework calls the input / output retry interface of the multipath software. The multipath kernel driver marks the last used path as "used", traverses all unused paths in the volume, and selects the next optimal path to reprocess the input / output according to the path selection algorithm.

[0030] In this embodiment of the application, the kernel driver engine 302 records the time when the data processing request starts processing as the first time and the time when the last access path cannot process the data processing request as the second time, and calculates the actual processing time based on the first time and the second time.

[0031] Understandably, by recording the start and end times to calculate the actual processing time, it is possible to determine whether the data request exceeds the lifespan threshold, and thus decide whether to retry or trigger the buffer, thereby ensuring the continuity of upper-layer business and the success rate of requests.

[0032] Specifically, the multipath kernel module obtains the current kernel time since system startup as the first moment and records this time in the input / output object. At the same time, it selects an optimal path to process the input / output request based on the current path selection algorithm.

[0033] During processing, if all paths of the volume have been tried and the request still cannot be processed, the current kernel time since system startup is obtained as the second time point, and the time difference between the first time point and the second time point is calculated. This time difference is the actual processing time.

[0034] In this embodiment of the application, after the last access path of the kernel driver engine 302 cannot process the data processing request, if the actual processing time is less than the survival time threshold, the used tag of the access path is cleared, and a reselection instruction is sent to the management component 200. The management component 200 responds to the reselection instruction and reselects the access path.

[0035] The survival time threshold is a preset maximum allowed data request processing time; if this time is exceeded, the request is considered to have failed.

[0036] Understandably, even if all access paths are temporarily unavailable, as long as the actual processing time of the data processing request does not exceed the lifespan threshold, the system can still continue to process the request by clearing the path label and reselecting an available path, thereby avoiding interruption of upper-layer business and improving the reliability and continuity of data requests.

[0037] During input / output processing, the multipath kernel driver continuously follows the process of "determining whether the input / output was successful—retrying if it failed." Specifically, the multipath input / output framework calls the input / output completion interface of the multipath software to determine success; if it fails, it calls the input / output retry interface to select a new unused path to continue processing the input / output. When all paths on the volume have been tried but the input / output is still unsuccessful, if the actual processing time is less than the preset input / output lifetime threshold, the "used" label for all paths is cleared, and the management component 200 reselects the optimal path based on the path selection algorithm to continue the retry process until the input / output succeeds or a timeout occurs.

[0038] In this embodiment, the service host 300 has multiple volumes mounted on it, and the persistent memory space of the storage hardware 301 is allocated multiple namespaces. The namespaces are associated with the volumes, and the data corresponding to the data processing requests that failed to process on the volume is cached in the namespace.

[0039] The persistent memory space is a memory area in storage hardware 301 used for long-term data storage, preserving data even in the event of power failure or system malfunction. The namespace is an independent storage area allocated separately for each volume within the persistent memory, used to cache data corresponding to failed data requests on the volume.

[0040] In this embodiment of the application, caching means that when a data processing request is not successfully processed on the volume, the corresponding data is temporarily stored in the namespace to ensure that the data is not lost and can be rewritten to the volume later.

[0041] Understandably, by allocating an independent namespace for each volume in the persistent memory of storage hardware 301, isolated caching of failed data is achieved. Even if the storage component 100 experiences input / output quiescence or temporary path unavailability, the business host 300 can still securely store incomplete data in the namespace, preventing data loss or data order corruption. Furthermore, because the namespace corresponds one-to-one with the volume, cache management is more granular, enabling rapid and accurate data write-back after storage recovery, thereby ensuring business continuity and data consistency.

[0042] To ensure that cached input / output data is not lost, persistent memory is used in this embodiment. When the business host 300 mounts a volume, the multipath kernel driver allocates an independent namespace for each volume in persistent memory and associates this namespace with each volume. These namespaces correspond to specific address spaces in persistent memory and are used to cache input / output data that has not yet been successfully processed on that volume. Whenever a failed input / output occurs on a volume, the data is written to the corresponding volume's namespace for persistent storage, ensuring that data is not lost during brief unavailability of storage component 100 or path anomalies. Furthermore, since each volume corresponds to an independent namespace, cache management and write-back can be precise to the volume level. After storage recovers, cached data can be efficiently written back to the storage device in its original order, thereby ensuring business continuity and data consistency.

[0043] In this embodiment, the storage hardware 301 is allocated a common address space, which is used to store metadata corresponding to the data processing request.

[0044] Metadata describes the data attributes in a data processing request, including the volume where the data resides, logical address, size, checksum, sequence number, etc., and is used to manage and write back cached data.

[0045] Understandably, centralized management of cached metadata for all volumes facilitates quick data location and write-back, ensuring data order.

[0046] To efficiently manage the input / output cached data of different volumes, this embodiment of the application also allocates a common address space in persistent memory to store the input / output metadata information of all volumes. Each piece of metadata records the data location, length, verification information, and global sequence number of the corresponding input / output, which is used to accurately locate and verify the integrity of the data during write-back or recovery. This metadata is managed through a linked list structure, enabling scan and write-back operations to process each input / output sequentially in the original commit order. Figure 3This illustration demonstrates the address space layout for caching input / output data in persistent memory according to an embodiment of this application. The entire address space is divided into multiple contiguous blocks. The leftmost block is the "I / O (Input / Output) Metadata Address Space," used to store metadata information corresponding to each input / output operation, including data location, checksum, and volume-related identifiers. Following this are several namespace blocks, labeled name_space_1, name_space_2, up to name_space_n. Each namespace is associated with a specific storage volume and stores the actual cached content of unprocessed input / output data on that volume. Through this structure, the system can establish a one-to-one correspondence between volumes and cache spaces, ensuring that unprocessed input / output data for each volume can be accurately mapped to the specified namespace. Simultaneously, the input / output metadata address space tracks and manages this cached data, achieving ordered data storage and efficient access. This embodiment of the application, combining the volume-corresponding namespace and the public metadata linked list, forms a data structure in the persistent memory space that supports both independent volume caching and unified management.

[0047] The metadata information for input / output data is organized as follows:

[0048] struct buffer_meta{

[0049] uint volume_id;

[0050] ulong lba;

[0051] void* data;

[0052] size_t size;

[0053] ulong crc;

[0054] ulong seq_id;

[0055] struct buffer_meta list;

[0056] }

[0057] Here, volume_id represents the ID (Identifier) ​​of the volume to which the input / output needs to be written, lba represents the original logical block address of the input / output, data represents the data content corresponding to the input / output, size represents the length of the data corresponding to the input / output, CRC (Cyclic Redundancy Check) represents the checksum of the input / output data, seq_id represents the global sequence number of the input / output (because some input / output writes need to be guaranteed in order), and List represents a pointer to the metadata of the input / output, pointing to the next metadata object.

[0058] In this embodiment of the application, the persistent memory space stores metadata through a data linked list. The structure of the data linked list includes multiple padding bits. The first padding bit is filled with the address space of the metadata, and the padding bits after the first padding bit are filled with the namespace.

[0059] In this context, a data linked list is a structure for organizing metadata. Each metadata object contains a pointer to the next metadata object, enabling sequential traversal and management. Padding bits are fields in the data linked list used to store specific information.

[0060] Understandably, by using a data linked list to store metadata in persistent memory space, the cached data corresponding to each data processing request can be managed in an orderly manner and located quickly, ensuring that cached data is not lost when system exceptions occur or volume processing fails; storing the address space and namespace information of metadata separately can clearly distinguish the mapping relationship between the data storage location and the volume to which it belongs.

[0061] In this embodiment, persistent memory space is used to cache metadata information of input / output data, which is organized and managed using a linked list structure. Specifically, each linked list consists of multiple fields, where the first field stores the address information of the metadata object in persistent memory, and subsequent fields record the namespace information associated with the storage volume corresponding to the metadata. This organization ensures orderly storage and fast access to cached data.

[0062] In this embodiment, the buffer control engine 303 starts a scanning thread. The scanning thread traverses the metadata address space at preset intervals and writes the metadata back to the storage device. The scanning thread is a worker thread in the buffer control engine 303 used to periodically check and process cached data, performing scanning operations at preset time intervals.

[0063] Understandably, by adding a buffer control engine 303 module to the multipath kernel driver, unsuccessfully processed inputs / outputs are temporarily stored. Once the stored inputs / outputs become available again, the buffer control engine 303 writes the cached inputs / outputs back to the storage device sequentially. After writing the failed inputs / outputs to the buffer module cache, the multipath kernel immediately reports the completion of input / output processing to the upper layer, making the upper-layer application unaware of input / output failures, thereby improving the continuity of host services.

[0064] In this embodiment, the buffer control engine 303 is specifically used for: scanning threads traversing the data linked list and sorting it from smallest to largest according to the labels in the metadata, so that the metadata is arranged in the order issued by the upper-layer application; for each metadata object, obtaining the metadata address in the corresponding namespace; reading the metadata and reference check code from the storage component 100 according to the metadata address, comparing the check code stored in the metadata object with the reference check code, and if they are inconsistent, rewriting the metadata; for each metadata, writing the metadata to the storage component 100 according to the target access path selected by the management component 200; if they are consistent, releasing the metadata and the corresponding memory space in the namespace after the metadata is successfully written to the storage component 100.

[0065] Understandably, by sorting metadata by label, the order of data written back to storage component 100 is strictly consistent with the order of data distributed by the upper-layer application, thereby effectively maintaining the consistency and correctness of business data. Using checksums for comparison allows for the timely detection and correction of potential data anomalies or corruption before data is written back, ensuring data integrity and reliability. The system only performs write-back operations on data with inconsistent checksums, avoiding unnecessary write overhead and significantly improving storage write efficiency. After successful data writing, the corresponding metadata and occupied memory space are released promptly, achieving efficient memory management and resource reclamation. Furthermore, the management component 200 intelligently selects the optimal access path for data writing, which not only improves the stability and fault tolerance of write-back operations but also effectively balances the load in a multi-path environment.

[0066] In this embodiment, a scanning thread is activated in the buffer control engine 303. This thread traverses the input / output metadata address space at preset time intervals, attempting to write back the input / output data in the persistent cache to the storage device. The specific process is as follows: The scanning thread first traverses the input / output metadata linked list and sorts each input / output metadata object in ascending order according to its label, so that the input / output is arranged in the order issued by the upper-layer application. Then, for each input / output metadata object, the address of the input / output in the corresponding namespace is obtained according to the volume_id and data attributes, and the data is read from the storage device to calculate the reference check code. The calculated reference check code is then compared with the check code stored in the input / output metadata. If they match, the input / output is successfully written to the storage device. If they do not match, it means that the data has changed and needs to be rewritten. The buffer control engine 303 selects the optimal path according to the volume path selection algorithm and submits the input / output to the multi-path input / output framework for writing to the storage device. After the input / output is successfully written to the storage device, the buffer control engine 303 releases the corresponding metadata and the input / output data space in the namespace.

[0067] In summary, the embodiments of this application can optimize the input / output retry mechanism to continuously attempt to write to storage before the input / output reaches its lifespan; implement an input / output caching function in the buffer control engine 303 to write unsuccessfully processed input / output to persistent memory to prevent upper-layer business input / output failure due to excessively long storage input / output quiescent time; and implement an input / output write-back function in the buffer control engine 303 to rewrite the input / output data in persistent memory to the storage device after the storage input / output recovers, ensuring data integrity and business continuity.

[0068] The server business data management system of this application embodiment includes a storage component for storing business data, a management component for managing multiple access paths of the storage component, and a business host comprising storage hardware, a kernel driver engine, and a buffer control engine. The kernel driver engine generates a data processing request based on a user request and sends it to the management component, which then selects the target access path for the storage component. The kernel driver engine records the actual processing time of the request; if it exceeds a time-to-live threshold, processing is deemed a failure, and a buffering instruction is sent to the buffer control engine, while simultaneously feeding back completion information to the upper layer. The buffer control engine stores the requested data in the storage hardware according to the instruction and writes it back to the storage component after the storage component recovers. This solves the technical problem in related technologies where the failure of all paths in multi-path software leads to upper-layer business interruption when storage input / output is silent, achieving the technical effect of ensuring uninterrupted business when storage input / output is silent.

[0069] Embodiments of this application also provide a server, including the server business data management system described above.

[0070] Embodiments of this application provide a server business data management method, such as... Figure 4 As shown, the method is applied to the business host of the aforementioned server business data management system, and the business host is configured to perform the following steps:

[0071] In step S101, the kernel driver engine is executed. The kernel driver engine sends the data processing request to the management component based on at least one data processing request generated by the upper-layer business of the user request. The management component responds to the data processing request and selects the target access path of the storage component for the business host.

[0072] Understandably, by using the kernel driver engine to send data processing requests generated by upper-layer business logic to the management component, effective connection between business requests and storage access is achieved; the management component selects the optimal access path according to the policy, thereby improving the efficiency and reliability of data transmission.

[0073] In step S102, the kernel driver engine is executed. The kernel driver engine records the actual processing time of the data processing request. If the actual processing time is greater than or equal to the survival time threshold, it is determined that the data processing request has failed. A buffer instruction is sent to the cache control engine to report the completion information of the data processing request to the upper layer.

[0074] Understandably, by recording the actual processing time of data processing requests, it is possible to promptly determine whether a request has timed out or failed, ensuring that the system can respond quickly to abnormal situations. When a request fails to process, the data is sent to the buffer control engine for caching to avoid data loss, and at the same time, completion information is fed back to the upper layer so that the upper-layer business can be aware of the processing result.

[0075] In step S103, the buffer control engine is executed. The buffer control engine responds to the buffer instruction and stores the data corresponding to the data processing request into the storage hardware. After the storage component returns to normal, the data from the storage hardware is written back into the storage component.

[0076] Understandably, by temporarily storing failed data in the storage hardware, the buffer control engine can ensure that data is not lost during storage component failures; and write the data back after the storage component recovers, thus achieving reliable data persistence; at the same time, upper-layer services are unaware of storage failures, improving the continuity and stability of the system.

[0077] According to the server business data management method proposed in this application, when the kernel driver engine is executed, it generates a data processing request based on the user request and sends it to the management component, which then selects the target access path for the storage component. The kernel driver engine records the actual processing time of the request; if it exceeds the lifespan threshold, it determines that the processing has failed, sends a buffering instruction to the buffer control engine, and feeds back completion information to the upper layer. When the buffer control engine is executed, it stores the data in the storage hardware and writes it back to the storage component after the storage component recovers. This solves the technical problem of upper-layer business interruption caused by the failure of all paths in multi-path software when the storage input / output is silent, achieving the technical effect of ensuring uninterrupted business when the storage input / output is silent.

[0078] This application also provides a non-volatile computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the above-described server business data management method.

[0079] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described server business data management method.

[0080] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0081] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0082] The above provides a detailed description of a server business data management system provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A server business data management system, characterized in that, include: A storage component, wherein the storage component is used to store business data; A management component for managing multiple access paths to the storage component; The service host includes storage hardware, a kernel driver engine, and a buffer control engine, wherein... The kernel driver engine generates at least one data processing request based on the upper-layer business request from the user, and sends the data processing request to the management component. The management component responds to the data processing request and selects the target access path of the storage component for the business host. The kernel driver engine records the actual processing time of the data processing request. If the actual processing time is greater than or equal to the lifespan threshold, it determines that the data processing request has failed and sends a buffering instruction to the cache control engine, feeding back the completion information of the data processing request to the upper layer. After the last access path of the kernel driver engine cannot process the data processing request, if the actual processing time is less than the lifespan threshold, it clears the used tag of the access path and issues a reselection instruction to the management component. The management component responds to the reselection instruction and reselects the access path. The buffer control engine responds to the buffer instruction. When a data processing request failure is detected, it stores the data corresponding to the data processing request in the storage hardware. After the storage component recovers, it writes the data from the storage hardware back into the storage component. The buffer control engine is used for: The scanning thread traverses the data linked list, sorting it from smallest to largest according to the labels in the metadata, ensuring that the metadata is arranged in the order issued by the upper-layer application. The data linked list describes the location, size, and verification information of the data corresponding to the data processing request in the storage hardware. The metadata describes the data attributes of the data processing request, including a sequence number to identify the data order, the address of the corresponding namespace, and a reference checksum, facilitating the management and write-back of the data corresponding to the data processing request. For the metadata object, the metadata address in the corresponding namespace is obtained. Based on the metadata address, the metadata and reference checksum are read from the storage component. The checksum stored in the metadata object is compared with the reference checksum. If they do not match, the metadata is rewritten. For the metadata, the target access path is selected by the management component to write the metadata to the storage component. If they match, after the metadata is successfully written to the storage component, the metadata and the corresponding memory space in the namespace are released. The memory space is the persistent memory space in the storage hardware of the business host, used to retain data in case of power failure or system abnormalities.

2. The server business data management system according to claim 1, characterized in that, The management component includes a routing interface, a completion interface, and a retry interface. The management component calls the routing interface to select an access path, and the management component calls the completion interface to determine the processing result of the data processing request. If the current access path cannot process the data processing request, the retry interface is called to reselect an access path carrying an unused tag.

3. The server business data management system according to any one of claims 1-2, characterized in that, The kernel driver engine records the time when the data processing request starts processing as the first time and the time when the last access path cannot process the data processing request as the second time, and calculates the actual processing time based on the first time and the second time.

4. The server business data management system according to claim 1, characterized in that, The service host has multiple volumes mounted on it, and the persistent memory space of the storage hardware is allocated multiple namespaces. The namespaces are associated with the volumes, and the data corresponding to the data processing requests that failed to process on the volume is cached in the namespaces.

5. The server business data management system according to claim 4, characterized in that, The storage hardware is allocated a public address space, which is used to store metadata of the data corresponding to the data processing request.

6. The server business data management system according to claim 5, characterized in that, The persistent memory space stores the metadata through a data linked list, wherein the structure of the data linked list includes multiple padding bits, the first padding bit of the multiple padding bits fills the address space of the metadata, and the padding bits after the first padding bit fill the namespace.

7. The server business data management system according to claim 6, characterized in that, The buffer control engine starts a scanning thread, which traverses the metadata address space at preset intervals and writes the metadata back to the storage device.

8. A server, characterized in that, Includes the server business data management system as described in any one of claims 1-7.

9. A server business data management method, characterized in that, The method is applied to the business host of the server business data management system according to any one of claims 1-7, wherein the business host is configured to perform the following steps: The kernel driver engine is executed. Based on at least one data processing request generated by the upper-layer business of the user request, the kernel driver engine sends the data processing request to the management component. The management component responds to the data processing request and selects the target access path of the storage component for the business host. The kernel driver engine is executed. The kernel driver engine records the actual processing time of the data processing request. If the actual processing time is greater than or equal to the lifespan threshold, it is determined that the data processing request has failed. A buffering instruction is sent to the cache control engine, and the completion information of the data processing request is fed back to the upper layer. After the last access path of the kernel driver engine cannot process the data processing request, if the actual processing time is less than the lifespan threshold, the used tag of the access path is cleared, and a reselection instruction is issued to the management component. The management component responds to the reselection instruction and reselects the access path. The buffer control engine is executed. Responding to the buffering instruction, the buffer control engine stores the data corresponding to the data processing request in the storage hardware. After the storage component recovers, the data from the storage hardware is written back to the storage component. Specifically, the buffer control engine is used to: scan a thread to traverse a data list, sorting it from smallest to largest according to the labels in the metadata, ensuring the metadata is arranged in the order issued by the upper-layer application. The data list describes the location, size, and verification information of the data corresponding to the data processing request in the storage hardware; the metadata describes the data attributes of the data processing request, including a sequence number identifying the data order, the address of the corresponding namespace, and so on. The system includes a reference checksum to facilitate the management and write-back of the data corresponding to the data processing request; for the metadata object, the system obtains the metadata address in the corresponding namespace; based on the metadata address, the system reads the metadata and the reference checksum from the storage component, compares the checksum stored in the metadata object with the reference checksum, and if they do not match, rewrites the metadata; for the metadata, the system selects the target access path according to the management component and writes the metadata to the storage component; if they match, after the metadata is successfully written to the storage component, the system releases the metadata and the corresponding memory space in the namespace, where the memory space is the persistent memory space in the storage hardware of the business host, used to retain data in the event of power failure or system abnormality.

Citation Information

Patent Citations

  • Path fault isolation processing method and device for multi-path software

    CN117221090A

  • Data storage method, system and equipment and medium

    CN119806408A