Request Processing Method, Apparatus, Storage Medium, and Electronic Device

Through the new access driver method, the connection between the virtual machine and the distributed storage backend is simplified, resource utilization efficiency is improved, and the problems of complex and inefficient connection process are solved.

CN115016892BActive Publication Date: 2025-07-25BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210622921.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-01
Publication Date
2025-07-25
Estimated Expiration
2042-06-01

AI Technical Summary

Technical Problem

The connection process between virtual machines and distributed storage backends is complicated and resource usage is inefficient.

Method used

The new access target driver is adopted to assume the connection role of the general block layer and the distributed storage backend, and the target request of the virtual machine is transmitted to the distributed storage backend through the target driver, and the callback is completed.

Benefits of technology

Simplifies the connection process between virtual machines and distributed storage backends and improves resource usage efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115016892B_ABST
    Figure CN115016892B_ABST
Patent Text Reader

Abstract

The present invention discloses a request processing method, apparatus, storage medium, and electronic device. The method includes: obtaining a target request of a virtual machine, where the target request is used to perform read / write operations or control operations on a distributed storage system; transmitting the target request to a general block layer; sending the target request to the distributed storage system by an accessed target driver; and performing a callback on the target request after the sending of the target request is completed. The present invention solves the technical problems of complex connection process between a virtual machine and a distributed storage backend or low resource utilization efficiency after connection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular, to a request processing method, apparatus, storage medium, and electronic device. Background Art

[0002] In the prior art, in the process of a virtual machine connecting to a distributed storage system, methods of docking the distributed storage backend based on the iscsi protocol and docking the distributed storage backend based on the Sheep dog+virtio-blk protocol are usually used.

[0003] However, the method of the above iscsi protocol is complex in steps during the hot migration of the virtual machine, and the method of Sheep dog+virtio-blk has low efficiency in using system resources.

[0004] This application aims to provide a new method for connecting a virtual machine to a distributed storage backend. Summary of the Invention

[0005] Embodiments of the present invention provide a request processing method, apparatus, storage medium, and electronic device to at least solve the technical problems of complex connection process between a virtual machine and a distributed storage backend or low resource utilization efficiency after connection.

[0006] According to one aspect of the embodiments of the present invention, a request processing method is provided, including: obtaining a target request of a virtual machine, where the target request is used to perform a read / write operation or a control operation on a distributed storage system; transmitting the target request to a general block layer; sending the target request to the distributed storage system by an accessed target driver; and performing a callback on the target request after the target request is sent.

[0007] According to another aspect of the embodiments of the present invention, a request processing apparatus is provided, including: an obtaining module, configured to obtain a target request of a virtual machine, where the target request is used to perform a read / write operation or a control operation on a distributed storage system; a transmitting module, configured to transmit the target request to a general block layer; a sending module, configured to send the target request to the distributed storage system by an accessed target driver; and a callback module, configured to perform a callback on the target request after the target request is sent.

[0008] As an optional example, the apparatus further includes: a declaration module, configured to declare a data access interface and a control flow interface of the target driver before sending the target request to the distributed storage system by the accessed target driver, where the data access interface includes a read operation interface and a write operation interface of the target driver, and the data access interface includes a control operation interface.

[0009] As an alternative example, the sending module includes: a first sending unit configured to send the target request to a thread pool by the target drive; an allocation unit configured to allocate the target drive to a target core in the thread pool; and a second sending unit configured to send the target request to the distributed storage system by the target core.

[0010] As an alternative example, the second sending unit includes: a storage subunit configured to store the book target request in a request queue of the target core in the order in which the target core receives the request; and a sending subunit configured to send the target request to the distributed storage system when the target request is at the head of the request queue.

[0011] As an alternative example, the callback module includes: a first callback unit configured to perform a callback on the target request after performing a read / write operation or a control operation on the distributed storage system according to the target request.

[0012] As an alternative example, the callback module includes: a second callback unit configured to perform a callback on the target request after transmitting the target request to a general block layer.

[0013] As an alternative example, the callback module further includes: an adding unit configured to add a layer of encapsulation layer on the side of the distributed storage system during the callback of the target request.

[0014] As an alternative example, the apparatus further includes: a receiving module configured to receive a returned error code when the read / write operation or the control operation fails after sending the target request to the distributed storage system; a processing module configured to randomly select a time point from a target interval to retry the target request when the error code indicates that the target request exceeds a processing speed limit or a queue limit; to retry the target request multiple times within a predetermined duration when the error code indicates that the target request exceeds a duration threshold and is not responded to or when the backend for receiving the target request is unreachable; and to return an error message indicating that the target request response fails when the error code indicates that the logical address for receiving the target request or the corresponding device does not exist.

[0015] According to another aspect of an embodiment of the present invention, there is also provided a storage medium storing a computer program, wherein the computer program, when run by a processor, executes the above-mentioned request processing method.

[0016] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned request processing method through the computer program.

[0017] In the embodiments of the present invention, a method is adopted to obtain a target request of a virtual machine, where the target request is used to perform read / write operations or control operations on a distributed storage system; transmit the target request to a general block layer; send the target request to the distributed storage system by an accessed target driver; and perform a callback on the target request after the target request is sent. Since in the above method, a new accessed target driver is used to undertake the role of connecting the general block layer and the distributed storage backend, after the target request of the virtual machine is transmitted to the general block layer, the target driver is used to transmit the target request to the distributed storage backend and complete the callback, thereby proposing a new method for connecting the virtual machine and the distributed storage backend, and further solving the technical problems of complex connection process between the virtual machine and the distributed storage backend or low resource utilization efficiency after connection. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0019] Figure 1 is a flowchart of an optional request processing method according to an embodiment of the present invention;

[0020] Figure 2 is an architecture diagram of an optional request processing method according to an embodiment of the present invention;

[0021] Figure 3 is a system schematic diagram of an optional request processing method according to an embodiment of the present invention;

[0022] Figure 4 is a system schematic diagram of another optional request processing method according to an embodiment of the present invention;

[0023] Figure 5 is a structural schematic diagram of an optional request processing device according to an embodiment of the present invention;

[0024] Figure 6 is a schematic diagram of an optional electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] Distributed storage backend: refers to the server nodes and processes that receive and actually store business data in a distributed system;

[0028] Distributed storage frontend: a software module in a distributed block storage system used to connect virtual machines and backend nodes

[0029] Qemu: a set of emulated processors.

[0030] Virtio protocol: is an I / O para-virtualization solution, a set of general I / O device virtualization programs, and an abstraction of a set of general I / O devices in a para-virtualized Hypervisor. It provides a communication framework and programming interface between upper-layer applications and various Hypervisor virtualized devices (such as KVM, Xen, VMware, etc.), reduces compatibility problems brought by cross-platform, and greatly improves the development efficiency of driver programs.

[0031] Spdk: an application software acceleration library used to accelerate the use of NVMe SSD as the backend storage;

[0032] Sheepdog protocol: specifically refers to a protocol for data and control flow communication that interacts with the qemu backend in this article. It customizes the data exchange format between the distributed storage front-end and back-end through a fixed-length header and its internal fields.

[0033] Spdk bdev: The general block layer after abstracting devices such as nvme in Spdk, similar to the general block layer in the Linux operating system;

[0034] Spdk vhost user target protocol: Specifically refers to a protocol for data and control flow communication between a storage front-end and a virtual machine, characterized by avoiding data copying from the virtual machine to its host machine.

[0035] Gateway: In this article, it specifically refers to a module that interacts between the distributed block storage module and the qemu backend, which can be considered as the front-end of the distributed block storage.

[0036] EBS: Elastic block store, a scalable distributed block storage system

[0037] Tablet Server: An EBS module responsible for converting read and write requests of Kylin bdev into read and write of internal pages;

[0038] Kylin Master: An EBS module responsible for managing cloud disks, managing Tablet Servers, and managing EBS service metadata;

[0039] Brpc: A remote procedure call (RPC) framework

[0040] According to the first aspect of the embodiments of the present invention, a request processing method is provided. Optionally, as Figure 1 shown, the above method includes:

[0041] S102, obtaining a target request of a virtual machine, where the target request is used to perform read and write operations or control operations on a distributed storage system;

[0042] S104, transmitting the target request into the general block layer;

[0043] S106, sending the target request to the distributed storage system by an accessed target driver;

[0044] S108, after the target request is sent, performing a callback on the target request.

[0045] Optionally, the distributed storage system in this embodiment may be a distributed storage backend, referring to server nodes and processes that receive and actually store service data in a distributed system.

[0046] Optionally, the method in this embodiment can be applied to the connection between a virtual machine running on a physical machine and a distributed storage system. The requests of the virtual machine are sent to the vhost process of SPDK via the vhost-user protocol through the QEMU backend, and the vhost process issues the IO requests to the distributed storage backend through a customized bdev layer. QEMU is a processor emulator that can emulate an operating system. SPDK is an application acceleration library for accelerating the use of NVMe SSDs as backend storage, and NVMe SSD is a solid-state drive storage medium. The bdev can be the target driver in this embodiment.

[0047] The bdev can be called Kylin bdev. It is the driver for the Elastic Block Store (EBS) docking client of Elastic Sustainable Storage. It converts the read and write requests of the client for a certain offset of the cloud disk into the read and write requests inside the EBS, thus docking the functions of the underlying read and write stream (IO stream) and control stream of the EBS. During deployment, in the virtual host scenario, Kylin bdev and the virtual host are deployed together on the HOST master node.

[0048] Figure 2 is the framework diagram of this embodiment. Figure 2 In this embodiment, the Elastic Block Store (EBS) gateway of Elastic Sustainable Storage adopts the SPDK vhost architecture. Data is transferred between QEMU and SPDK vhost through shared memory. SPDK provides a good module-level abstraction. By writing the corresponding bdev for Kylin, that is, the target driver, the docking between the vhost and the EBS storage cluster can be completed. The written Kylin bdev includes functions such as volume mounting, unmounting, and cloning snapshots. By accessing the bdev, the bdev can connect the general block layer and the distributed storage backend. After the requests of the virtual machine are transmitted from the virtual machine to the general block layer, they are transmitted to the distributed storage backend through the bdev.

[0049] This embodiment uses a newly accessed target driver to undertake the role of connecting the general block layer and the distributed storage backend. After the target requests of the virtual machine are transmitted to the general block layer, the target requests are transmitted to the distributed storage backend through the target driver and the callback is completed, thus proposing a new method for connecting the virtual machine and the distributed storage backend, and further solving the technical problems of complex connection process between the virtual machine and the distributed storage backend or low resource utilization efficiency after connection.

[0050] As an optional example, before the target driver transmits the target requests to the distributed storage system, the above method further includes:

[0051] Declare the target-driven data access interface and control flow interface, where the data access interface includes the target-driven read operation interface and write operation interface, and the data access interface includes a control operation interface.

[0052] Optionally, in this embodiment, before using the target drive, the interfaces of the target drive need to be declared. The main job of Kylin bdev is to complete the conversion of the storage protocol, convert the general block device read and write into a protocol adapted to the EBS storage cluster, and send it to the backend storage cluster. It mainly provides two data access-related interfaces:

[0053] 1) Write(vid, offset, size, data): Write data with a length of size to a specific offset of the cloud disk;

[0054] 2) Read(vid, offset, size, data): Read data with a length of size from a specific offset of the cloud disk into data;

[0055] Kylin bdev directly interacts with the Tablet Server and Kylin Master in the EBS storage cluster. The former is mainly the data stream, and the latter is mainly the metadata operation. The operation frequency is relatively low compared to the data stream and uses brpc communication. The control flow-related interfaces include:

[0056] kylin_mount(const char* vol, void* ctx);

[0057] kylin_unmount(const char* vol, void* ctx);

[0058] void kylin_gateway_deinit(void);

[0059] uint64_t kylin_get_vdi_size(const char* vol);

[0060] uint64_t kylin_resize_vdi(const char* vol, uint64_t new_size);

[0061] These interfaces can be called by kylin bdev. The above brpc is a remote process scheduling framework.

[0062] The Spdk framework can pass the data and length of the IO request in the structure of iovec, and the vhost user interface mainly implements several interfaces and data structures of bdev.

[0063]

[0064]

[0065] The above data structure realizes the registration to the SPDK bdev framework, which can be understood as adding a new type of block device driver to the kernel, enabling the corresponding operation entry to be found when creating a bdev of this type later.

[0066] The IO stream of the new bdev realizes the IO data reading and writing functions through the following customized spdk_bdev_fn_table.

[0067]

[0068] The most important interface among them is submit_request, which realizes passing the IO issued by the virtual machine to the backend storage (here, the encapsulated interface can be called to achieve this). After the IO is completed, it is then notified to the host.

[0069] The specific interface functions and their function descriptions are shown in Table 1 below:

[0070] Table 1

[0071]

[0072] On this basis, some basic operations of the bdev module can also be provided, such as creating a bdev module, in order to achieve compatibility with the control flow (mainly rpc.py) in the SPDK framework. The main functions include:

[0073] Interfaces such as bdev_kylin_create, bdev_kylin_delete, bdev_kylin_resize.

[0074] Through the creation and declaration of the above interfaces, when in use, the target request can be sent to the distributed storage system through the bdev.

[0075] As an optional example, the sending of the target request by the accessed target driver to the distributed storage system includes:

[0076] Sending the target request from the target driver to the thread pool;

[0077] Allocating the target driver to the target core in the thread pool;

[0078] Sending the target request from the target core to the distributed storage system.

[0079] As an alternative example, the sending of the target request by the target core to the distributed storage system includes:

[0080] Storing the book target request in the request queue of the target core in the order in which the target core receives the request;

[0081] In the case where the target request is at the head of the request queue, sending the target request to the distributed storage system.

[0082] Optionally, in this embodiment, a thread model is proposed. The Spdk vhost user bdev framework supports the Spdkthread polling mechanism. In each execution of the polling process, the IO request can be sent to the tablet server through the encapsulated brpc interface. The default callback of Brpc may be executed on any processor core. In the callback, the actual operation to be executed is sent to a specific core through spdk_thread_send_msg, and let it execute the actual callback spdk_bdev_io_complete on the Spdk thread. This callback is the callback in the spdk vhost user / bdev framework. The overall architecture is as Figure 3 , multiple requests are distributed to multiple core request queues in a cyclic manner, and the core queue sends the requests to the table server. When the vhost process starts, it usually binds multiple cores. When submitting, the IO requests are distributed by the vhost blockcontroller core to the spdk threads corresponding to other processing cores in turn. Each spdk thread is associated with a request queue. When it receives these requests, it sends these requests to different distributed storage backends through the read and write interfaces encapsulated by the old gateway in the order of arrival.

[0083] After the distributed storage backend finishes executing the IO, as Figure 4 shown, the common callback function will put the completed io_task into a completion queue (done queue) corresponding to the spdk thread that processed it before. There is also a periodic poller on this spdk thread, which is responsible for batch sending the requests in the completion queue to the vhost blockcore to notify the bdev common layer which IOs have been completed currently.

[0084] As an alternative example, after the target request is transmitted to the general block layer, the callback of the target request includes:

[0085] After performing read / write operations or control operations on the distributed storage system according to the target request, the target request is called back.

[0086] As an optional example, the above method further includes:

[0087] After transmitting the target request to the general block layer, the target request is called back.

[0088] Optionally, in this embodiment, the high-performance scenario usually includes an asynchronous process, and asynchrony inevitably involves callbacks. The places where the gateway in this embodiment needs to consider callbacks are:

[0089] 1. After the request submitted by the vhost controller is completed, the corresponding callback (bdev_callback) needs to be executed to tell the framework;

[0090] 2. After the asynchronous read / write of the backend distributed cluster is completed, the corresponding callback (backend_storage_sdk_callback) will also be executed.

[0091] Therefore, the call and callback process on the complete IO path is as follows:

[0092] bdev send io request--->backend storage sdk api---->backend storagesdk callback--->bdev_io_cal lback

[0093] As an optional example, the above method further includes:

[0094] During the process of calling back the target request, a layer of encapsulation layer is added on the side of the distributed storage system.

[0095] During the callback process, the callback of the distributed storage backend is usually implemented in C++; while bdev_callback is in bdev and is implemented in C. The two cannot be directly called. If the callback is implemented according to the call order listed above, there will be a situation where the spdk bdev io callback is called in the distributed backend storage, which will cause the distributed storage backend to depend on and link the spdk code and libraries, making the originally independent distributed storage backend and spdk become deeply coupled, resulting in the need to sort out and design the callback process on the IO path, and at the same time ensure compatibility with different language implementations of the backend. To avoid this dependency and reduce coupling, a layer of wrapper can be added:

[0096] On the backend storage sdk side:

[0097] Backend storage sdk api ----> Backend storage sdk callback ---> Wrapper

[0098] / * The implementation mode of this wrapper on the SPDK bdev side is as follows:

[0099] Wrapper(ctx) {

[0100] Bdev io_callback(ctx);

[0101] } * /

[0102] Then the backend storage sdk side can perceive the address of this wrapper. When initializing the bdev, the SPDK bdev side will tell the backend storage sdk layer the address of this wrapper function. The implementation process is as follows:

[0103] 1. Add a global wrapper() function address pointer inside the backend storage sdk side;

[0104] 2. The backend storage sdk adds an interface to set the address of this wrapper function: register_callback();

[0105] 3. The backend storage sdk exports this register_callback() and compiles a C library that can be called externally;

[0106] 4. The SPDK bdev layer calls register_callback(bdev io_callback) in this C library;

[0107] 5. On the backend storage sdk side, execute the IO callback in the following way:

[0108] Backend storage sdk api ----> Backend storage sdk callback ---> (*wrapper)()

[0109] As an optional example, the above method further includes:

[0110] After sending the target request to the distributed storage system, in the case where the read / write operation or the control operation fails to be executed according to the target request, receive the returned error code;

[0111] In the case where the error code indicates that the target request exceeds the processing speed limit or queue limit, randomly select a time point from the target interval to retry the target request;

[0112] In the case where the error code indicates that the target request exceeds the duration threshold and is not responded to, or the backend for receiving the target request is unreachable, retry the target request multiple times within a predetermined duration;

[0113] In the case where the error code indicates that the logical address for receiving the target request or the corresponding device does not exist, return an error message indicating that the target request response fails.

[0114] In a distributed scenario, multiple servers are usually deployed at the backend. When writing to the backend distributed cluster, due to reasons such as network anomalies, excessive pressure, node downtime, disk failures, etc., the write may fail, or there may be a situation where data transmission from the front end to the backend goes wrong.

[0115] For these errors, the backend sdk should be able to return an error code to the customized bdev layer, and the bdev performs corresponding error handling and retries according to the specific error type.

[0116] 1. Error code transfer mechanism

[0117] Transfer method:

[0118] 1.1 After the read / write request sent to the backend is executed, the status and msg representing the execution result (success or failure) can be obtained through rpc callback;

[0119] 1.2 The status is transferred to the bdev layer through the transfer mechanism between the callback mechanism and the bdev layer (the above figure)

[0120] 1.3 Based on the IO ctx given to the backend, confirm the bdev_io that obtains the execution result through the transfer mechanism between the callback mechanism and the bdev layer;

[0121] 1.4 Call spdk_bdev_io_complete() to return the execution result to the bdev general layer:

[0122] If the execution is successful, return SPDK_BDEV_IO_STATUS_SUCCESS to the upper layer; otherwise return SPDK_BDEV_IO_STATUS_FAILED;

[0123] 2. Error Handling

[0124] The status above indicates the result after the IO is submitted to the backend storage node and executed. Its common values include:

[0125] OK: Indicates that the IO request is successfully completed at the backend;

[0126] Busy: Indicates exceeding the processing speed limit or queue limit of the backend;

[0127] TimeOut: Indicates that after the request is sent, no response is received within the specified timeout period;

[0128] Unavailable: Indicates that the backend expected to receive the IO request is unreachable (unable to access);

[0129] NoResource: Indicates that the logical address expected to receive the request or the corresponding device does not exist;

[0130] 3. Retry Strategy

[0131] Based on the above error codes, there are different retry strategies:

[0132] Busy: Retry the request by delaying a random time within a certain range;

[0133] TimeOut: Retry the request several times by delaying for a period of time;

[0134] Unavailable: Retry the request several times by delaying for a period of time;

[0135] Ok: Do not retry and directly return SPDK_BDEV_IO_STATUS_SUCCESS to the general bdev layer

[0136] NoResource: Do not retry and directly return SPDK_BDEV_IO_STATUS_FAILED to the general bdev layer.

[0137] Based on the comprehensive requirements of the bdev access model, thread model, and callback mechanism analyzed above, the key implementation of the main interfaces can be obtained. The main interfaces on the Spdk bdev side are as follows:

[0138] bdev_kylin_initialize

[0139] The main process is as follows:

[0140] 1. Obtain the configurations of the backend storage cluster, bdev spdk thread core binding, and cloud data directory;

[0141] 2. Implement one gateway_callback_wrapper mentioned above, call the callback registration function provided in the backend library, and register it to gateway_callback_wrapper;

[0142] 3. Call the gateway initialization function provided by the backend library;

[0143] 4. Call spdk_io_device_register to register the newly implemented bdev (&kylin_if) above;

[0144] 5. If the created bdev is recorded in the metadata directory, restore it by creating the bdev and vhost controller

[0145] bdev_kylin_create

[0146] The main process is as follows:

[0147] Define and apply for a kylin_disk data structure that includes spdk_bdev, spdk_poller, and task management;

[0148] Set the bdev name according to the passed-in parameters;

[0149] Call the xuanwu_mount() interface provided by the backend library to mount the disk;

[0150] Call the xuanwu_get_vdi_size() interface provided by the backend library to obtain the capacity of the disk and record it in memory;

[0151] Call the xuanwu_get_block_size() provided by the backend library to obtain the block size of the disk and record it in memory;

[0152] Set the fn_table of the spdk_bdev member in kylin_disk to kylin_fn_table to associate the customized operations of kylinbdev with its data structure;

[0153] Call spdk_bdev_register to register a new bdev to the framework: If this bdev is created for the first time, write information such as the capacity and speed limit of the disk to the metadata directory;

[0154] bdev_kylin_submit_request

[0155] The main sending process of the IO request is as follows:

[0156] As the interaction interface between the spdk bdev framework and the customized bdev, bdev_kylin_submit_request comes with parameters such as spdk_bdev_io. The main process of submitting an IO request is as follows:

[0157] 1. Construct an io_task;

[0158] 2. Sequentially select a thread from the spdk threads used to process IOs, and record the core_id of the core where it is located;

[0159] 3. Send the above io_task to the spdk_thread corresponding to core_id;

[0160] 4. This spdk_thread calls the xuanwu_write() or xuanwu_read() interface encapsulated by the backend according to the IO read / write type;

[0161] The main callback process of the IO request is as follows:

[0162] After the IO of the distributed storage backend is completed, the previously registered gateway_callback_wrapper will be called;

[0163] Inside the gateway_callback_wrapper, the completed io_task is sent to the only bdev thread through spdk_thread_send_msg;

[0164] The only bdev thrad calls spdk_bdev_io_complete according to the completion status in the io_task to notify the upper-layer framework that the current IO has been completed;

[0165] bdev_kylin_resize

[0166] The core implementation process of "bdev_kylin_resize" is as follows:

[0167] 1. Register "bdev_kylin_resize" to the rpc.py framework through SPDK_RPC_REGISTER

[0168] 2. Obtain the size of the volume to be set from the input parameters, compare it with the current one, and if it is not larger than the current one, report a parameter error and then exit;

[0169] 3. Call the xuanwu_resize() interface provided by the backend storage to modify the volume size;

[0170] 4. Notify the bdev framework through spdk_bdev_notify_blockcnt_change to update the disk size seen by the virtual machine.

[0171] The interfaces that need to be encapsulated and provided on the distributed storage side include:

[0172] int xuanwu_gateway_init(const char* conf_dir);

[0173] void xuanwu_write(const char* vol, char* buf, int32_t size, int64_t off, void* ctx);

[0174] void xuanwu_read(const char* vol, char* buf, int32_t size, int64_t off, void* ctx);

[0175] void xuanwu_mount(const char* vol, void* ctx);

[0176] void xuanwu_unmount(const char* vol, void* ctx);

[0177] void xuanwu_gateway_deinit(void);

[0178] uint64_t xuanwu_get_vdi_size(const char* vol);

[0179] uint64_t xuanwu_resize(const char* vol, uint64_t);

[0180] The functions of the above interfaces are consistent with their names and are already supported by the distributed storage backend. However, the backend storage may be implemented in C++, with different interface parameters and usage. All that needs to be done is to add a layer of wrapper on the distributed storage backend to encapsulate the above functions into C language interfaces that meet the requirements of the above interfaces and provide the corresponding library for bdev to use. This achieves the goal of decoupling.

[0181] Through testing, using the method of this embodiment, in terms of the peak performance of a single volume: the maximum random write IOPS increased from 80.5K to 150K; the random read increased from 90.2K to 169K; the sequential write throughput increased from 678.4MB / s to 1359MiB / s; in terms of the peak performance of a single virtual machine: the highest (in the case of 4 volumes) random write is 139.5K IOPS, random read is 157.2K IOPS, write throughput is 933MiB / s, and read throughput is 1182MiB / s. While the highest random write of the old gateway is 96.26K IOPS, random read is 119.8K IOPS, write throughput is 860MiB / s, and read throughput is 1240MiB / s; in terms of the peak performance of a single gateway: improvements can be seen under 6 / 8 / 12 / 16 cores, with the highest random write being 191.7K IOPS; random read is 188.5K IOPS, write throughput is 1508MiB / s, and read throughput is 3456MiB / s. While the old gateway is 111.8K IOPS / 118.6KIOS / 948MiB / s and 2139MB / s respectively. And the performance can be seen to improve as the number of cores used increases from 6 / 8 / 12 / 16, indicating that its performance can improve with the expansion of the number of cores used.

[0182] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0183] According to another aspect of the embodiments of the present application, a request processing device is further provided, as Figure 5 shown, including:

[0184] An acquisition module 502, configured to acquire a target request of a virtual machine, where the target request is used to perform read / write operations or control operations on a distributed storage system;

[0185] A transmission module 504, configured to transmit the target request to a general block layer;

[0186] A sending module 506, configured to send the target request to the distributed storage system by an accessed target driver;

[0187] A callback module 508, configured to perform a callback on the target request after the target request is sent.

[0188] Optionally, the distributed storage system in this embodiment may be a distributed storage backend, which refers to the server nodes and processes that receive and actually store service data in a distributed system.

[0189] Optionally, the method in this embodiment can be applied to the connection between a virtual machine running on a physical machine and a distributed storage system. The requests of the virtual machine are sent to the vhost process of spdk through the vhost-user protocol via the qemu backend. This vhost process distributes the IO requests to the distributed storage backend through a customized bdev layer. qemu is a processor emulator that can emulate an operating system. spdk is an application software acceleration library for accelerating the use of NVMe SSD as the backend storage, and NVMe SSD is a solid-state drive storage medium. bdev can be the target driver in this embodiment.

[0190] bdev can be called Kylin bdev. It is the driver for the elastic and sustainable storage EBS docking client, which converts the read and write operations of the client at a certain offset of the cloud disk into read and write requests inside EBS, thus docking the functions of the underlying read and write stream (IO stream) and control stream of EBS. During deployment, in the virtual host scenario, Kylin bdev and the virtual host are deployed together on the HOST master node.

[0191] Figure 2 is the framework diagram of this embodiment. Figure 2 In [the figure], for the elastic and sustainable storage EBS gateway, the spdkvhost architecture is adopted. Data is transmitted between qemu and spdkvhost through shared memory. spdk provides a good module-level abstraction. By writing the corresponding bdev for Kylin, that is, the target driver, the docking between vhost and the EBS storage cluster can be completed. The written Kylin bdev includes functions such as volume mounting, unmounting, and cloning snapshots. By accessing the bdev, the bdev can connect the general block layer and the distributed storage backend. After the requests of the virtual machine are transmitted from the virtual machine to the general block layer, they are transmitted to the distributed storage backend through the bdev.

[0192] This embodiment uses a newly accessed target driver to undertake the role of connecting the general block layer and the distributed storage backend. After the target requests of the virtual machine are transmitted to the general block layer, the target requests are transmitted to the distributed storage backend through the target driver and the callback is completed, thus proposing a new method for connecting the virtual machine and the distributed storage backend, and further solving the technical problems of complex connection process between the virtual machine and the distributed storage backend or low resource usage efficiency after connection.

[0193] For other examples of this embodiment, please refer to the above examples and will not be elaborated here.

[0194] Figure 6 It is a block diagram of an optional electronic device according to an embodiment of the present application. As Figure 6 shown, it includes a processor 602, a communication interface 604, a memory 606, and a communication bus 608. Among them, the processor 602, the communication interface 604, and the memory 606 complete communication with each other through the communication bus 608. Among them,

[0195] The memory 606 is used to store computer programs;

[0196] The processor 602, when executing the computer program stored on the memory 606, realizes the following steps:

[0197] Obtain a target request of a virtual machine, where the target request is used to perform read / write operations or control operations on a distributed storage system;

[0198] Transmit the target request to the general block layer;

[0199] Send the target request to the distributed storage system by an accessed target driver;

[0200] After the target request is sent, perform a callback on the target request.

[0201] Optionally, in this embodiment, the above communication bus may be a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used for communication between the above electronic device and other devices.

[0202] The memory may include RAM, or may also include non-volatile memory, for example, at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0203] As an example, the above memory 606 may but is not limited to include the acquisition module 502, the transmission module 504, the sending module 506, and the callback module 508 in the above request processing device. In addition, it may also include but is not limited to other module units in the above request processing device, which will not be elaborated in this example.

[0204] The above-mentioned processor can be a general-purpose processor, including but not limited to: CPU (Central Processing Unit, central processing unit), NP (Network Processor, network processor), etc.; it can also be a DSP (Digital Signal Processing, digital signal processor), ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), FPGA (Field - Programmable Gate Array, field-programmable gate array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0205] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be elaborated here.

[0206] Those of ordinary skill in the art can understand that Figure 6 The structure shown is only for illustration. The device for implementing the above request processing method can be a terminal device, and the terminal device can be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a mobile Internet device (Mobile Internet Devices, MID), a PAD and other terminal devices. Figure 6 It does not limit the structure of the above electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 6 or have a different configuration from that shown in Figure 6 Those shown.

[0207] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: a flash drive, a ROM, a RAM, a magnetic disk or an optical disc, etc.

[0208] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is run by a processor, it executes the steps in the above request processing method.

[0209] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program. This program can be stored in a computer-readable storage medium, and the storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disc, etc.

[0210] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0211] If the integrated unit in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0212] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0213] In the several embodiments provided by this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0214] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0215] In addition, in each embodiment of the present invention, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0216] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for request processing, characterized in that, Including: Obtaining a target request of a virtual machine, where the target request is used to perform read / write operations or control operations on a distributed storage system; Transmitting the target request to a general block layer; Sending the target request to the distributed storage system by an accessed target driver; Performing a callback on the target request after the target request is sent successfully; Wherein, before the accessed target driver sends the target request to the distributed storage system, the method further includes: declaring a data access interface and a control flow interface of the target driver, where the data access interface includes a read operation interface and a write operation interface of the target driver, the data access interface includes a control operation interface, the write operation interface is used to write data with a length of size to a specific offset of a cloud disk, the read operation interface is used to read data with a length of size from a specific offset of the cloud disk into data, and the target driver is used to convert read / write operations on a general block into a protocol adapted to the distributed storage system and send it to the distributed storage system.

2. The method according to claim 1, wherein The sending the target request to the distributed storage system by the accessed target driver includes: Sending the target request from the target driver to a thread pool; Allocating the target driver to a target core in the thread pool; Sending the target request to the distributed storage system by the target core.

3. The method according to claim 2, characterized in that, The sending the target request to the distributed storage system by the target core includes: Storing the book target request in a request queue of the target core in the order in which the target core receives the request; When the target request is at the head of the request queue, sending the target request to the distributed storage system.

4. The method according to claim 1, characterized in that, After the target request is transmitted to the general block layer, performing a callback on the target request includes: Performing a callback on the target request after performing read / write operations or control operations on the distributed storage system according to the target request.

5. The method according to claim 4, characterized in that The method further includes: Performing a callback on the target request after the target request is transmitted to the general block layer.

6. The method according to claim 5, wherein The method further includes: During the callback of the target request, adding a layer of encapsulation layer on the side of the distributed storage system.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: After the target request is sent to the distributed storage system, when the read / write operation or the control operation performed according to the target request fails, receiving a returned error code; When the error code indicates that the target request exceeds the processing speed limit or queue limit, randomly selecting a time point from a target interval to retry the target request; When the error code indicates that the target request exceeds the duration threshold and is not responded, or the backend for receiving the target request is unreachable, retrying the target request multiple times within a predetermined duration; When the error code indicates that the logical address for receiving the target request or the corresponding device does not exist, returning an error message indicating that the target request response fails.

8. A device for request processing, characterized in that, Including: An acquisition module, configured to acquire a target request of a virtual machine, where the target request is used to perform a read / write operation or a control operation on a distributed storage system; A transmission module, configured to transmit the target request to a general block layer; A sending module, configured to send the target request to the distributed storage system by an accessed target driver; A callback module, configured to perform a callback on the target request after the sending of the target request is completed; Wherein, before the target request is sent to the distributed storage system by the accessed target driver, it further includes: declaring a data access interface and a control flow interface of the target driver, where the data access interface includes a read operation interface and a write operation interface of the target driver, the data access interface includes a control operation interface, the write operation interface is used to write data with a length of size to a specific offset of a cloud disk, the read operation interface is used to read data with a length of size from a specific offset of the cloud disk into data, and the target driver is used to convert the read / write of a general block into a protocol adapted to the distributed storage system and send it to the distributed storage system.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is run by a processor, it executes the method described in any one of claims 1 to 7.

10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method described in any one of claims 1 to 7 through the computer program.

Citation Information

Patent Citations

  • Method and device for processing read and write request

    CN108008911A

  • Cloud platform virtual machine I / O acceleration method, device and system

    CN114020406A