A message distribution method, device and storage medium

By creating multi-IO request/response queues and publish/subscribe queues in the host and NVME hard disk, the shortcomings of existing message distribution systems in concurrency capabilities, throughput, hardware link transmission performance, etc. are solved, and efficient and low-power message distribution effect is achieved.

CN116225742BActive Publication Date: 2025-05-13SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310239212.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-09
Publication Date
2025-05-13
Estimated Expiration
2043-03-09

AI Technical Summary

Technical Problem

The existing message distribution systems have shortcomings in concurrency capabilities, throughput, hardware link transmission performance, processor usage, reliability, power consumption and latency, resulting in low system performance.

Method used

By creating multiple IO request/response queues, publish queues and subscription queues in the host and NVME hard disk, and establishing corresponding mapping tables, high concurrent and asynchronous message publish/subscribe command submission is achieved, making full use of the pipeline parallel execution characteristics of the hardware circuit.

Benefits of technology

It improves the concurrency capability of business processes, the throughput of the software protocol stack, the transmission performance of the hardware link, reduces the processor usage, improves the overall performance of the system, and improves the reliability of the message protocol, reduces power consumption and delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116225742B_ABST
    Figure CN116225742B_ABST
Patent Text Reader

Abstract

The invention discloses a message distribution method, comprising the following steps: creating a corresponding IO request queue and a response queue in a host according to the identifier of each business process; creating a publishing queue and a subscription queue corresponding to each business process in an NVME hard disk; in response to the business process sending a message subscription request to the NVME hard disk, constructing a first IO write command and writing it into the corresponding IO request queue, and writing the first IO write command into the NVME hard disk; in response to the business process sending a message publishing request to the NVME hard disk, constructing a second IO write command by using a message type to be published, a corresponding publishing queue and a message content, writing it into the corresponding IO request queue, and writing the second IO write command into the NVME hard disk, forwarding the message content and the message type to a subscription queue corresponding to the message type to be published; in response to detecting that there is a non-empty subscription queue, reporting the message content in the non-empty subscription queue to the corresponding business process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of message processing, and in particular to a message distribution method, device and storage medium. Background Art

[0002] The message distribution system is a data distribution application system based on the message queue protocol, such as Kafka, RabbitMQ, mosquitto, etc. Messages are sent from a source address such as host memory to one or more destination addresses such as the local machine or the host memory of the network peer. Common message queue protocols include network socket protocol, publish-subscribe protocol, advanced message queue protocol, etc.

[0003] The usual method of system module interaction is to build a message service cluster and run a specific message communication protocol (such as the publish-subscribe protocol) middleware on the message server. Each system module sets the subscribed message type during the initialization phase. When the business process needs to send a message during the system operation phase, the message is first sent to the message server, and then the message server forwards the message to all target modules that have subscribed to the message. The target module receives the message and executes the corresponding processing flow.

[0004] This approach can effectively implement message communication interaction between modules, but the following problems still exist:

[0005] The business process concurrency capability is weak: Since the protocol stack cannot allocate a separate input and output queue for each business process, multiple business processes share the same queue, resulting in the need for mutually exclusive synchronization operations between processes, thereby reducing concurrency performance.

[0006] The software protocol stack has a low throughput: Since the software protocol stack has fewer queues and a lower queue depth, such as message communications that rely on network protocols, it generally only provides a queue number and depth that are proportional to the number of hardware lines and the maximum transmission unit.

[0007] The hardware link transmission performance is low: Since the hardware link DMA channel is limited by the data throughput of the software protocol stack, such as the maximum transmission unit of the network protocol, the buffer capacity, etc., the single DMA data volume is limited, resulting in low hardware link performance.

[0008] High system processor utilization: Since all message communication protocols are implemented by software, such as the socket transmission protocol based on the network protocol stack, and lack a dedicated high-speed hardware DMA data transmission channel, when the message communication interactions between modules are too frequent, a large amount of processor resources of the message server will be occupied, resulting in high server processor utilization.

[0009] The overall system performance is low: The increase in processor utilization further limits the data transmission performance of the message communication middleware and may cause other modules that require processor resources, such as network communication or memory read and write performance to decline.

[0010] The message protocol has low reliability: Since the message lacks a hardware error recovery method, when changes in the hardware and software environment during data transmission cause message communication abnormalities, message loss may occur, resulting in low transmission reliability.

[0011] The message protocol consumes high power: Due to the lack of an automatic switching mechanism for software and hardware power consumption states, the system is in full-speed operation for a long time without data transmission, such as the survival heartbeat detection process of the network protocol stack, resulting in high system power consumption.

[0012] The message protocol has a high latency: Since the message protocol stack has many software layers, such as a message distribution system based on a network protocol, the protocol stack contains multiple layers from the application layer to the physical layer. The layered data format conversion causes a high transmission latency. Summary of the invention

[0013] In view of this, in order to overcome at least one aspect of the above problems, an embodiment of the present invention proposes a message distribution method, comprising the following steps:

[0014] In the host, create corresponding IO request queues and response queues according to the identifier of each business process;

[0015] Create a publishing queue and a subscription queue corresponding to each of the business processes in the NVME hard disk, respectively, and establish a first mapping table recording the interrupt number of each of the business processes, the mapping relationship between the publishing queue and the subscription queue, and return the addresses of the publishing queue and the subscription queue to the host to establish a second mapping table recording the identification of the business process, the IO request queue and the response queue, and the mapping relationship between the publishing queue and the subscription queue in the host;

[0016] In response to the business process sending a message subscription request to the NVME hard disk, a first IO write command is constructed using the message type to be subscribed and the address of the subscription queue determined according to the second mapping table and written into the corresponding IO request queue, and the first IO write command is written to the NVME hard disk to establish a third mapping table in the NVME hard disk that records the mapping relationship between the message type and the subscription queue;

[0017] In response to the business process sending a message publishing request to the NVME hard disk, a second IO write command is constructed using the message type to be published, the corresponding publishing queue and the message content, and the second IO write command is written into the corresponding IO request queue and the second IO write command is written into the NVME hard disk, and the message content and the message type are forwarded to the subscription queue corresponding to the message type to be published according to the third mapping table;

[0018] In response to detecting that there is a non-empty subscription queue, the message content in the non-empty subscription queue is reported to the corresponding business process according to the first mapping table.

[0019] In some embodiments, an initialization process is also included, and the initialization process includes:

[0020] In the host, a management process is used to execute a host initialization process to create a management request queue and a management response queue, and the addresses of the management request queue and the management response queue are written into the NVME hard disk, and a status register of the NVME hard disk is updated to indicate that the host initialization is complete;

[0021] The NVME hard disk detects that the status register indicates that the host is initialized, applies for the memory of the first mapping table and the third mapping table and initializes them to empty, and updates the status register to indicate that the NVME hard disk is initialized;

[0022] In response to the host detecting that the status register indicates that the NVME hard disk initialization is complete, the initialization process ends.

[0023] In some embodiments, a publishing queue and a subscription queue corresponding to each of the business processes are respectively created in the NVME hard disk, and a first mapping table recording the interrupt number of each of the business processes, the mapping relationship between the publishing queue and the subscription queue is established, and the publishing queue and the subscription queue are returned to the host to establish a second mapping table recording the mapping relationship between the business process identifier, the IO request queue and the response queue, and the publishing queue and the subscription queue in the host, further comprising:

[0024] In the host, an interrupt number is assigned to each of the service processes, and a publish queue and a subscribe queue creation command is constructed using the interrupt number of the service process as a parameter, and the creation command is written into the management request queue;

[0025] The NVME hard disk reads the create command in the management request queue through DMA;

[0026] The publishing queue and the subscription queue are created according to the creation command, and a first mapping table is established to record the mapping relationship between the interruption number, the publishing queue and the subscription queue.

[0027] In some embodiments, it also includes:

[0028] Constructing response information with the addresses of the publishing queue and the subscription queue as parameters and writing the response information into the management response queue;

[0029] The host obtains and parses the response information from the management response queue to obtain the addresses of the publishing queue and the subscription queue, and then establishes a second mapping table recording the mapping relationship between the business process identifier, the IO request queue and the response queue, and the publishing queue and the subscription queue.

[0030] In some embodiments, in response to the business process sending a message subscription request to the NVME hard disk, a first IO write command is constructed using the message type to be subscribed and the address of the subscription queue determined according to the second mapping table and written into the corresponding IO request queue, and the first IO write command is written to the NVME hard disk to establish a third mapping table in the NVME hard disk that records the mapping relationship between the message type and the subscription queue, further comprising:

[0031] Obtaining the message type, data block address, and callback function address to be subscribed that are passed in by the business process, and establishing a fourth mapping table that records the mapping relationship between the message type, the data block address, and the callback function address;

[0032] Obtaining the address of the subscription queue corresponding to the business process in the second mapping table;

[0033] Constructing a first IO write command and configuring a data pointer of the first IO write command to be an address corresponding to the message type, and SLBA to be the address of the subscription queue;

[0034] Obtaining the IO request queue corresponding to the business process in the second mapping table and writing the first IO write command into the corresponding IO request queue;

[0035] The NVME hard disk obtains the first IO write command from the IO request queue and parses to obtain the message type to be subscribed and the address of the subscription queue, thereby establishing a third mapping table recording the mapping relationship between the message type and the subscription queue;

[0036] Constructing response information with the address of the subscription queue as a parameter and writing the response information to the response queue corresponding to the business process;

[0037] The host obtains the response information from the response queue corresponding to the business process and ends the message subscription.

[0038] In some embodiments, in response to detecting that there is a non-empty subscription queue, reporting the message content in the non-empty subscription queue to the corresponding business process according to the first mapping table further includes:

[0039] Querying the first mapping table to obtain an interrupt number corresponding to the non-empty subscription queue;

[0040] Update the length and message type of the message content in the non-empty subscription queue to the status register and notify the host of the length and message type of the message content by triggering the MSI interrupt;

[0041] querying the fourth mapping table to obtain a data block address according to the message type, and querying the second mapping table to obtain a subscription queue address corresponding to the corresponding business process;

[0042] Construct an IO read command and configure the data pointer of the IO read command to be the address corresponding to the data block, SLBA is the address of the subscription queue address corresponding to the corresponding business process, and NLB is the length of the message content;

[0043] According to the second mapping table, the IO read command is sent to the IO request queue corresponding to the corresponding business process and the NVME hard disk is notified to execute the IO read command so that the message content in the non-empty subscription queue is moved to the address corresponding to the data block corresponding to the corresponding business process.

[0044] In some embodiments, it also includes:

[0045] The non-empty address of the subscription queue is used as a parameter to construct a response message and write the response message into the response queue corresponding to the corresponding business process;

[0046] The host obtains the response information from the response queue corresponding to the corresponding business process and ends the message reporting;

[0047] The fourth mapping table is queried according to the message type to obtain the address of the callback function so as to process the message content in the address corresponding to the data block through the callback function.

[0048] In some embodiments, in response to the business process sending a message publishing request to the NVME hard disk, a second IO write command is constructed using the message type to be published, the corresponding publishing queue and the message content, and the corresponding IO request queue is written and the second IO write command is written to the NVME hard disk, and the message content and message type are forwarded to the subscription queue corresponding to the message type to be published according to the third mapping table, further comprising:

[0049] Obtain the message type, message content, and data block address to be published that are passed in by the business process;

[0050] Obtaining the address of the subscription queue corresponding to the business process in the second mapping table;

[0051] Constructing a second IO write command and configuring a data pointer of the second IO write command to be an address corresponding to the message type and an address corresponding to the message content, and SLBA is the address of the publishing queue;

[0052] Obtaining the IO request queue corresponding to the business process in the second mapping table and writing the second IO write command into the corresponding IO request queue;

[0053] The NVME hard disk obtains the second IO write command from the IO request queue and parses to obtain the address of the message type to be published and the address of the message content, executes DMA to obtain the message type and the message content to be published according to the address of the message type to be published and the address of the message content, and transmits them to the corresponding publishing queue;

[0054] Forwarding the message type and the message content in the corresponding publishing queue to the subscription queue corresponding to the message type to be published according to the third mapping table;

[0055] Constructing response information with the address of the publishing queue as a parameter and writing the response information to the response queue corresponding to the business process;

[0056] The host obtains the response information from the response queue corresponding to the business process and ends the message publishing.

[0057] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention further provides a computer device, including:

[0058] at least one processor; and

[0059] A memory storing a computer program executable on the processor, wherein the processor executes the steps of any one of the message distribution methods described above when executing the program.

[0060] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any one of the message distribution methods described above are performed.

[0061] The present invention has one of the following beneficial technical effects: The scheme proposed by the present invention is based on the host-side message publishing / subscribing command submission process of the NVMe protocol, and by introducing multiple IO request / response queues, high-concurrency, asynchronous message publishing / subscribing command submission is achieved. At the same time, the hard disk-side message publishing / subscribing command submission process based on the NVMe protocol, by introducing multiple IO request / response queues, publishing / subscribing queues, and message forwarding modules, fully utilizes the pipeline parallel execution characteristics of the hardware circuit, which not only reduces the processor usage, but also improves the transmission performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can be obtained based on these drawings without paying creative work.

[0063] Figure 1 A flow chart of a message distribution method provided by an embodiment of the present invention;

[0064] Figure 2 A schematic diagram of the module hierarchy structure of a system provided by an embodiment of the present invention;

[0065] Figure 3 A schematic diagram of the interaction flow of software and hardware modules of the message distribution method provided by an embodiment of the present invention;

[0066] Figure 4 A schematic diagram of a queue structure provided for an embodiment of the present invention;

[0067] Figure 5 A schematic diagram of a mapping table structure provided for an embodiment of the present invention;

[0068] Figure 6 An initialization flow chart provided for an embodiment of the present invention;

[0069] Figure 7 A queue creation flow chart provided for an embodiment of the present invention;

[0070] Figure 8 A message subscription flow chart provided for an embodiment of the present invention;

[0071] Fig. 9 A message reporting flow chart provided for an embodiment of the present invention;

[0072] Fig.10 A message publishing flow chart provided for an embodiment of the present invention;

[0073] Fig.11A schematic diagram of the structure of a computer device provided by an embodiment of the present invention;

[0074] Fig.12 A schematic diagram of the structure of a computer-readable storage medium provided for an embodiment of the present invention. DETAILED DESCRIPTION

[0075] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the embodiments of the present invention are further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0076] It should be noted that all expressions using "first" and "second" in the embodiments of the present invention are for distinguishing two non-identical entities with the same name or non-identical parameters. It can be seen that "first" and "second" are only for the convenience of expression and should not be understood as limitations on the embodiments of the present invention. The subsequent embodiments will not explain this one by one.

[0077] In an embodiment of the present invention, SLBA is the starting logical block address of the device side included in the IO command of the NVMe protocol. The device divides the address space in units of logical blocks. When the host submits the IO command, it specifies the logical block address, which indicates which address in the device space is to be read and written. NLB is the number of logical blocks on the device side included in the IO command of the NVMe protocol. When the host submits the IO command, it specifies the number of logical blocks in the IO command, which indicates how many logical blocks of data are to be read and written from the SLBA of the device space.

[0078] According to one aspect of the present invention, an embodiment of the present invention provides a message distribution method, such as Figure 1 As shown, it may include the steps of:

[0079] S1, create corresponding IO request queues and response queues in the host according to the identifier of each business process;

[0080] S2, respectively create a publishing queue and a subscription queue corresponding to each of the business processes in the NVME hard disk, and establish a first mapping table recording the interrupt number of each of the business processes, the mapping relationship between the publishing queue and the subscription queue, and return the addresses of the publishing queue and the subscription queue to the host to establish a second mapping table recording the identification of the business process, the IO request queue and the response queue, and the mapping relationship between the publishing queue and the subscription queue in the host;

[0081] S3, in response to the business process sending a message subscription request to the NVME hard disk, using the message type to be subscribed and the address of the subscription queue determined according to the second mapping table to construct a first IO write command and write it into the corresponding IO request queue, and write the first IO write command to the NVME hard disk to establish a third mapping table in the NVME hard disk that records the mapping relationship between message types and subscription queues;

[0082] S4, in response to the business process sending a message publishing request to the NVME hard disk, using the message type to be published, the corresponding publishing queue and the message content to construct a second IO write command and write it into the corresponding IO request queue and write the second IO write command to the NVME hard disk, and forwarding the message content and message type to the subscription queue corresponding to the message type to be published according to the third mapping table;

[0083] S5: In response to detecting that there is a non-empty subscription queue, report the message content in the non-empty subscription queue to the corresponding business process according to the first mapping table.

[0084] The scheme proposed by the present invention is based on the host-side message publishing / subscribing command submission process of the NVMe protocol, and realizes high-concurrency, asynchronous message publishing / subscribing command submission by introducing multiple IO request / response queues. At the same time, the hard disk-side message publishing / subscribing command submission process based on the NVMe protocol, by introducing multiple IO request / response queues, publishing / subscribing queues, and message forwarding modules, fully utilizes the pipeline parallel execution characteristics of the hardware circuit, which not only reduces the processor utilization rate, but also improves the transmission performance.

[0085] In some embodiments, the present invention divides the system into two parts: a host and an NVMe hard disk (taking SSD as an example), and the host is connected to multiple NVMe SSDs through a PCIe interface.

[0086] The module hierarchy of the system is as follows Figure 2 As shown, the host includes a management process, several business processes, a user link library, and an NVMe driver module, which is used to concurrently submit message publish / subscribe requests of the business processes to the NVMe SSD through the multi-IO queue mechanism of the NVMe protocol.

[0087] NVMe SSD includes hardware and firmware modules for executing the publish / subscribe process of messages through SSD hardware.

[0088] Combine the following Figure 3 The software and hardware module interaction process of the NVMe SSD-based message distribution system acceleration method illustrates the functions of each module of the host and the NVMe SSD:

[0089] Host side:

[0090] 1) Management process, further divided into:

[0091] a) Framework initialization, used to initialize the message distribution system;

[0092] b) Resource allocation, which is used to create business processes and allocate resources to them.

[0093] 2) Business processes are used to perform specific functional businesses and are further divided into:

[0094] a) Data block address, used to fill in the data address of the message to be published / subscribed;

[0095] b) Callback function pointer, which is used to process the callback function address of the subscription message.

[0096] 3) User link library, further divided into:

[0097] a) Link library initialization, used to initialize the message type-callback function address-data block address mapping table structure and call the initialization interface of the NVMe driver;

[0098] The message type-callback function address-data block address mapping table (the fourth mapping table) is defined as follows:

[0099] Using the operating system link library mechanism, a mapping table is generated for each business process. Each table entry stores the message type subscribed by the process, the address of the message processing callback function, and the data address of the received message.

[0100] b) Queue acquisition, used to call the queue creation interface of the NVMe driver;

[0101] c) Message subscription, used to provide a message subscription interface to the business process;

[0102] d) Message publishing, used to provide a message publishing interface to the business process;

[0103] e) Message reading, used to read the message data to be processed after receiving the interrupt notification of the NVMe driver;

[0104] f) Queue release: Used to call the queue deletion interface of the NVMe driver.

[0105] 4)NVMe drivers are further divided into:

[0106] a) NVMe host initialization. It is used to implement the host initialization process (create management request / response queues) of the NVMe protocol through the user-mode NVMe driver of the operating system, such as the Linux SPDK framework, and initialize the process identifier---IO request / response queue address---SSD publish / subscribe queue address mapping table (second mapping table), which is defined as follows: each table entry stores the business process identifier, the host IO request / response queue address, and the SSD publish / subscribe queue address, so that each business process has a unique IO request / response queue and SSD publish / subscribe queue;

[0107] b) Queue creation, used to create IO request / response queues for business processes in the host memory;

[0108] c) IO request / response queue reading and writing, used to read and write the IO request / response queue of a business process;

[0109] d) Interrupt processing, used to receive the subscription interrupt request of NVMe SSD and forward it to the business process;

[0110] e) Queue deletion, used to delete the IO request / response queue of the business process.

[0111] The functions of each SSD module are defined as follows:

[0112] 1) Hardware, further divided into:

[0113] a) Status register, which is used as the shared memory of the host and SSD firmware for synchronization between the host and SSD firmware;

[0114] b) Queue address register, used for host and SSD to configure IO request / response queue and SSD publish / subscribe queue address. Hardware can read and write queue data from this address in combination with queue Head and Tail register values.

[0115] c) Queue Head and Tail registers 1 to n. Used to record the head and tail pointers of each IO request / response queue and SSD publish / subscribe queue. The hardware reads data from the queue through this pointer;

[0116] d) Queue reading and writing, used to read and write IO request / response queues and SSD publish / subscribe queue data;

[0117] e) Message forwarding, used to distribute message data to all business processes that have subscribed to this message type;

[0118] f) Interrupt request, used to notify the host IO response queue or SSD subscription queue that data is to be read;

[0119] g) Cache management, used to periodically synchronize and cache SSD publish / subscribe queue data to Flash to avoid power loss;

[0120] 2) Firmware, further divided into:

[0121] a) NVMe device initialization, which is used to execute the device initialization process of the NVMe protocol (create NVMe namespace, enumerate NVMe devices, obtain hardware status), and initialize the following two mapping tables:

[0122] i. Message type - (SSD subscription queue address set) mapping table (third mapping table)

[0123] Each table entry stores the message type and the SSD subscription queue address of the process that subscribes to the message type.

[0124] ii. Process interrupt number - SSD subscription queue address mapping table (first mapping table)

[0125] Each table entry stores the subscription notification interrupt number of the business process and the process SSD subscription queue address.

[0126] b) Cache queue creation, which is used to divide the cache / Flash into fixed-size partitions and select one from the partitions as the SSD publish / subscribe queue corresponding to the business process one by one;

[0127] c) Cache queue deletion, used to release the mapping relationship between the business process and the SSD publish / subscribe queue;

[0128] d) Abnormal notification, used to notify the host when an error occurs during firmware operation;

[0129] The system can connect multiple SSDs to transfer data using multiple PCIe links, multiplying data transfer performance.

[0130] In some embodiments, Figure 4 The management request / response queue, IO request / response queue, and SSD publish / subscribe queue structure diagram shown in the figure, the last 64 bits of each data page are the address of the next data page; the first data page is the message control field, including the message type, whether to broadcast, whether to encrypt, checksum, the address of the next data page, etc. The SSD hardware uses this field to perform more control operations on the message forwarding process. Figure 5 The schematic diagram of the mapping table structure shown in the figure uses the operating system dynamic link library mechanism to enable each process to have a copy of the message type-callback function address-data block address mapping table, thereby avoiding the synchronization overhead of multiple processes accessing the same table.

[0131] In some embodiments, an initialization process is also included, and the initialization process includes:

[0132] In the host, a management process is used to execute a host initialization process to create a management request queue and a management response queue, and the addresses of the management request queue and the management response queue are written into the NVME hard disk, and a status register of the NVME hard disk is updated to indicate that the host initialization is complete;

[0133] The NVME hard disk detects that the status register indicates that the host is initialized, applies for the memory of the first mapping table and the third mapping table and initializes them to empty, and updates the status register to indicate that the NVME hard disk is initialized;

[0134] In response to the host detecting that the status register indicates that the NVME hard disk initialization is complete, the initialization process ends.

[0135] Specifically, Figure 6 In the initialization flowchart shown in FIG. 1 , when initializing, perform the following steps: Figure 3 1)a to 1)d):

[0136] a) The management process executes the framework initialization process, in which the link library initialization process is called;

[0137] b) The user library executes the library initialization process, in which the NVMe host initialization process is called.

[0138] c) The NVMe driver creates a management request / response queue, configures the queue address to the SSD hardware register, and updates the SSD hardware status register to indicate that the NVMe host initialization is complete;

[0139] d) After the SSD firmware polls the SSD status register to indicate "host-side initialization completed", it executes the device-side initialization process of the NVMe protocol, and then updates the SSD status register to indicate that the device initialization is completed.

[0140] In some embodiments, a publishing queue and a subscription queue corresponding to each of the business processes are respectively created in the NVME hard disk, and a first mapping table recording the interrupt number of each of the business processes, the mapping relationship between the publishing queue and the subscription queue is established, and the publishing queue and the subscription queue are returned to the host to establish a second mapping table recording the mapping relationship between the business process identifier, the IO request queue and the response queue, and the publishing queue and the subscription queue in the host, further comprising:

[0141] In the host, an interrupt number is assigned to each of the service processes, and a publish queue and a subscribe queue creation command is constructed using the interrupt number of the service process as a parameter, and the creation command is written into the management request queue;

[0142] The NVME hard disk reads the create command in the management request queue through DMA;

[0143] The publishing queue and the subscription queue are created according to the creation command, and a first mapping table is established to record the mapping relationship between the interruption number, the publishing queue and the subscription queue.

[0144] In some embodiments, it also includes:

[0145] Constructing response information with the addresses of the publishing queue and the subscription queue as parameters and writing the response information into the management response queue;

[0146] The host obtains and parses the response information from the management response queue to obtain the addresses of the publishing queue and the subscription queue, and then establishes a second mapping table recording the mapping relationship between the business process identifier, the IO request queue and the response queue, and the publishing queue and the subscription queue.

[0147] Specifically, Figure 7 As shown in the queue creation flowchart, when creating a queue, perform the following steps: Figure 3 2)a to 2)h) of the above:

[0148] a) The management process executes the resource allocation process, in which the queue acquisition process of the user link library is called;

[0149] b) The user link library executes the queue acquisition process, in which the queue creation process of the NVMe driver is called;

[0150] c) The NVMe driver creates a one-to-one IO request / response queue for the business process;

[0151] d) The NVMe driver updates the process ID and IO request / response queue address to the process ID---IO request / response queue address---SSD publish / subscribe queue address mapping table;

[0152] e) The NVMe driver submits an SSD queue creation command to the management request queue, and the command includes the process interrupt number;

[0153] f) After the SSD hardware reads the SSD queue creation command from the host's management request queue, it notifies the SSD firmware to execute the cache queue creation process;

[0154] g) The SSD firmware executes the cache queue creation process and creates an SSD publish / subscribe queue for the business process;

[0155] h) SSD firmware update process interrupt number—SSD subscription queue address mapping table.

[0156] Among them, ① you can write the MSI configuration register so that the SSD hardware can send an interrupt to the host through the MSI method after the message forwarding is completed, thereby notifying the host that there are messages to be processed; ② you can refer to the custom command process of the NVMe protocol and build an SSD queue creation command with the interrupt number as a parameter.

[0157] In some embodiments, in response to the business process sending a message subscription request to the NVME hard disk, a first IO write command is constructed using the message type to be subscribed and the address of the subscription queue determined according to the second mapping table and written into the corresponding IO request queue, and the first IO write command is written to the NVME hard disk to establish a third mapping table in the NVME hard disk that records the mapping relationship between the message type and the subscription queue, further comprising:

[0158] Obtaining the message type, data block address, and callback function address to be subscribed that are passed in by the business process, and establishing a fourth mapping table that records the mapping relationship between the message type, the data block address, and the callback function address;

[0159] Obtaining the address of the subscription queue corresponding to the business process in the second mapping table;

[0160] Constructing a first IO write command and configuring a data pointer of the first IO write command to be an address corresponding to the message type, and SLBA to be the address of the subscription queue;

[0161] Obtaining the IO request queue corresponding to the business process in the second mapping table and writing the first IO write command into the corresponding IO request queue;

[0162] The NVME hard disk obtains the first IO write command from the IO request queue and parses to obtain the message type to be subscribed and the address of the subscription queue, thereby establishing a third mapping table that records the mapping relationship between the message type and the subscription queue;

[0163] Constructing response information with the address of the subscription queue as a parameter and writing the response information to the response queue corresponding to the business process;

[0164] The host obtains the response information from the response queue corresponding to the business process and ends the message subscription.

[0165] Specifically, Figure 8 In the message subscription flow chart shown in the figure, when subscribing to a message, perform the following steps: Figure 3 3)a to 3)e) of the above:

[0166] a) The business process calls the user link library message subscription process and passes in the message type, data block and callback function address;

[0167] b) The user link library updates the message type-callback function address-data block address mapping table, indicating that after the process receives this type of message, it must process the data contained in the data block address through the callback function;

[0168] c) The user link library message subscription process builds an IO write command, configures the data pointer as the memory page address corresponding to the message type, SLBA as the SSD subscription queue address, NLB as 1, and calls the NVMe driver's IO request / response queue read / write interface to submit the IO write command;

[0169] d) The NVMe driver updates the queue Tail register of the SSD hardware and notifies the SSD hardware to fetch IO commands;

[0170] e) After the SSD hardware reads the IO command, it adds the SSD subscription queue address to the message type-(SSD subscription queue address set) mapping table, indicating that the hardware wants to forward the message to the SSD subscription queue.

[0171] In some embodiments, in response to detecting that there is a non-empty subscription queue, reporting the message content in the non-empty subscription queue to the corresponding business process according to the first mapping table further includes:

[0172] Querying the first mapping table to obtain an interrupt number corresponding to the non-empty subscription queue;

[0173] Update the length and message type of the message content in the non-empty subscription queue to the status register and notify the host of the length and message type of the message content by triggering the MSI interrupt;

[0174] querying the fourth mapping table to obtain a data block address according to the message type, and querying the second mapping table to obtain a subscription queue address corresponding to the corresponding business process;

[0175] Construct an IO read command and configure the data pointer of the IO read command to be the address corresponding to the data block, SLBA is the address of the subscription queue address corresponding to the corresponding business process, and NLB is the length of the message content;

[0176] According to the second mapping table, the IO read command is sent to the IO request queue corresponding to the corresponding business process and the NVME hard disk is notified to execute the IO read command so that the message content in the non-empty subscription queue is moved to the address corresponding to the data block corresponding to the corresponding business process.

[0177] In some embodiments, it also includes:

[0178] The non-empty address of the subscription queue is used as a parameter to construct a response message and write the response message into the response queue corresponding to the corresponding business process;

[0179] The host obtains the response information from the response queue corresponding to the corresponding business process and ends the message reporting;

[0180] The fourth mapping table is queried according to the message type to obtain the address of the callback function so as to process the message content in the address corresponding to the data block through the callback function.

[0181] Specifically, Fig. 9 In the message reporting flow chart shown in FIG. 1 , when reporting a message, perform the following steps: Figure 3 5)a to 5)c) of the above:

[0182] a) The SSD hardware queries the process interrupt number-SSD subscription queue address mapping table to obtain the process interrupt number corresponding to the non-empty SSD subscription queue;

[0183] b) The SSD hardware updates the message type and message data length to the status register, and then notifies the NVMe driver of a message to be processed by a business process through the MSI interrupt;

[0184] c) After the NVMe driver reads the SSD hardware status register, it notifies the user link library that there is a message to be processed;

[0185] d) The user link library queries the message type-callback function address-data block address mapping table, constructs an IO read command to transmit the message data to the data block address, and then executes the callback function to process the message data.

[0186] In some embodiments, in response to the business process sending a message publishing request to the NVME hard disk, a second IO write command is constructed using the message type to be published, the corresponding publishing queue and the message content, and the corresponding IO request queue is written and the second IO write command is written to the NVME hard disk, and the message content and message type are forwarded to the subscription queue corresponding to the message type to be published according to the third mapping table, further comprising:

[0187] Obtain the message type, message content, and data block address to be published that are passed in by the business process;

[0188] Obtaining the address of the subscription queue corresponding to the business process in the second mapping table;

[0189] Constructing a second IO write command and configuring a data pointer of the second IO write command to be an address corresponding to the message type and an address corresponding to the message content, and SLBA is the address of the publishing queue;

[0190] Obtaining the IO request queue corresponding to the business process in the second mapping table and writing the second IO write command into the corresponding IO request queue;

[0191] The NVME hard disk obtains the second IO write command from the IO request queue and parses to obtain the address of the message type to be published and the address of the message content, executes DMA to obtain the message type and the message content to be published according to the address of the message type to be published and the address of the message content, and transmits them to the corresponding publishing queue;

[0192] Forwarding the message type and the message content in the corresponding publishing queue to the subscription queue corresponding to the message type to be published according to the third mapping table;

[0193] Constructing response information with the address of the publishing queue as a parameter and writing the response information to the response queue corresponding to the business process;

[0194] The host obtains the response information from the response queue corresponding to the business process and ends the message publishing.

[0195] Specifically, Fig.10 As shown in the message publishing flow chart, when subscribing to a message, perform the following steps: Figure 3 4)a to 4)e) of the above:

[0196] a) The business process calls the link library message publishing process and passes in the message type and data block address to be published;

[0197] b) The user link library constructs an IO write command, configures the data pointer as a linked list of data block memory page addresses containing the message type and message content, SLBA as the SSD release queue address, NLB as the actual number of data block memory pages, and then calls the NVMe driver's IO request / response queue read / write interface to submit the IO write command;

[0198] c) The NVMe driver updates the queue Tail register of the SSD hardware and notifies the SSD hardware to fetch IO commands;

[0199] d) After the SSD hardware reads the IO command, it queries the message type-(SSD subscription queue address set) mapping table to obtain the SSD subscription queue addresses that have subscribed to messages of this type;

[0200] e) The SSD hardware copies the message data via DMA to all SSD subscription queues that have subscribed to messages of this type.

[0201] The technical solution of the present invention proposes a message distribution system acceleration method, which defines a management process, a business process, a user link library, an NVMe driver, a management request / response queue, an IO request / response queue, an SSD publish / subscribe queue, a mapping table, a message forwarding engine, etc. It is only for understanding the specific implementation mode of the present invention, and is not used to limit the present invention. Any optimization made without departing from the spirit and scope of the present invention, especially the design optimization of the interaction process between the NVMe host and the SSD, each queue, the mapping table, and the message forwarding process, are within the protection scope of the present invention.

[0202] The solution proposed by the present invention can bring the following beneficial effects:

[0203] Strong business process concurrency capability: Since the protocol stack allocates a separate input and output queue for each business process, multiple business processes do not share the same queue, so there is no need for mutually exclusive synchronization operations between processes, thereby improving concurrency performance.

[0204] The software protocol stack has a higher throughput rate: Since the software protocol stack has a large number of queues and a deep queue, compared with the message distribution system that relies on network protocols, multi-queue concurrency and asynchronous transmission make the software protocol stack have a higher throughput rate.

[0205] High hardware link transmission performance: Since the hardware link DMA channel is no longer limited by the data throughput of the software protocol stack, such as the network protocol maximum transmission unit, buffer capacity, synchronization overhead, etc., the hardware link DMA performance is high.

[0206] Low system processor utilization: Since the message forwarding protocol is entirely implemented by the hardware message forwarding module, the software only needs to continuously submit IO command requests to the input / output queue, so the host processor utilization is low.

[0207] The overall system performance is higher: The reduction in processor utilization further helps improve the data transmission performance of other modules in the system that require processor resources, such as network communication or memory reading and writing, so the overall system performance is higher.

[0208] The message protocol has high reliability: Since the message has the PCIe hardware error recovery method, when the message communication is abnormal due to changes in the software and hardware environment during the data transmission process, there will be no message loss, so the data transmission reliability is high.

[0209] The message protocol has low power consumption: Because it has an automatic switching mechanism for the hardware power consumption state of PCIe devices and has multiple power consumption levels, the system will automatically switch to a low power consumption state when there is no data transmission for a long time, so the system power consumption is low.

[0210] The message protocol has lower latency: Since the message protocol stack has fewer software layers, compared with message distribution based on network protocols (including multiple layers from the application layer to the physical layer), the transmission delay caused by layered data format conversion is reduced.

[0211] Based on the same inventive concept, according to another aspect of the present invention, Fig.11 As shown, an embodiment of the present invention further provides a computer device 501, including:

[0212] at least one processor 520; and

[0213] The memory 510 stores a computer program 511 that can be run on a processor. When the processor 520 executes the program, the processor 520 performs the steps of any of the above message distribution methods.

[0214] Based on the same inventive concept, according to another aspect of the present invention, Fig.12 As shown, an embodiment of the present invention further provides a computer-readable storage medium 601, which stores a computer program 610. When the computer program 610 is executed by a processor, the steps of any of the above message distribution methods are performed.

[0215] Finally, it should be noted that a person skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods.

[0216] Furthermore, it should be appreciated that the computer-readable storage medium (eg, memory) herein may be either a volatile memory or a nonvolatile memory, or may include both volatile and nonvolatile memory.

[0217] It will also be appreciated by those skilled in the art that various exemplary logic blocks, modules, circuits and algorithm steps described in conjunction with the disclosure herein can be implemented as electronic hardware, computer software or a combination of the two. In order to clearly illustrate this interchangeability of hardware and software, a general description has been given to the functions of various schematic components, blocks, modules, circuits and steps. Whether this function is implemented as software or hardware depends on specific applications and the design constraints imposed on the entire system. Those skilled in the art can implement the function in various ways for each specific application, but this implementation decision should not be interpreted as causing a departure from the disclosed scope of the embodiments of the present invention.

[0218] The above are exemplary embodiments disclosed in the present invention, but it should be noted that various changes and modifications may be made without departing from the scope disclosed in the embodiments of the present invention as defined in the claims. The functions, steps and / or actions of the method claims according to the disclosed embodiments described herein do not need to be performed in any particular order. In addition, although the elements disclosed in the embodiments of the present invention may be described or required in individual form, they may also be understood as multiple unless explicitly limited to the singular.

[0219] It should be understood that, as used herein, the singular forms "a", "an" are intended to include the plural forms as well, unless the context clearly supports an exception. It should also be understood that, as used herein, "and / or" refers to any and all possible combinations including one or more of the associated listed items.

[0220] The serial numbers of the embodiments disclosed in the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0221] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0222] A person skilled in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the disclosure of the embodiments of the present invention (including the claims) is limited to these examples; under the concept of the embodiments of the present invention, the technical features in the above embodiments or different embodiments can also be combined, and there are many other changes in different aspects of the embodiments of the present invention as above, which are not provided in detail for the sake of simplicity. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention should be included in the protection scope of the embodiments of the present invention.

Claims

1. A message distribution method, characterized in that: The following steps are involved: In the host, create corresponding IO request queues and response queues according to the identifier of each business process; Create a publishing queue and a subscription queue corresponding to each of the business processes in the NVME hard disk, respectively, and establish a first mapping table recording the interrupt number of each of the business processes, the mapping relationship between the publishing queue and the subscription queue, and return the addresses of the publishing queue and the subscription queue to the host to establish a second mapping table recording the identification of the business process, the IO request queue and the response queue, and the mapping relationship between the publishing queue and the subscription queue in the host; In response to the business process sending a message subscription request to the NVME hard disk, a first IO write command is constructed using the message type to be subscribed and the address of the subscription queue determined according to the second mapping table and written into the corresponding IO request queue, and the first IO write command is written to the NVME hard disk to establish a third mapping table in the NVME hard disk that records the mapping relationship between the message type and the subscription queue; In response to the business process sending a message publishing request to the NVME hard disk, a second IO write command is constructed using the message type to be published, the corresponding publishing queue and the message content, and the second IO write command is written into the corresponding IO request queue and the second IO write command is written into the NVME hard disk, and the message content and the message type are forwarded to the subscription queue corresponding to the message type to be published according to the third mapping table; In response to detecting that there is a non-empty subscription queue, the message content in the non-empty subscription queue is reported to the corresponding business process according to the first mapping table.

2. The method according to claim 1, characterized in that It also includes an initialization process, which includes: In the host, a management process is used to execute a host initialization process to create a management request queue and a management response queue, and the addresses of the management request queue and the management response queue are written into the NVME hard disk, and a status register of the NVME hard disk is updated to indicate that the host initialization is complete; The NVME hard disk detects that the status register indicates that the host is initialized, applies for the memory of the first mapping table and the third mapping table and initializes them to empty, and updates the status register to indicate that the NVME hard disk is initialized; In response to the host detecting that the status register indicates that the NVME hard disk initialization is complete, the initialization process ends.

3. The method according to claim 2, characterized in that Create a publishing queue and a subscription queue corresponding to each of the business processes in the NVME hard disk, respectively, and establish a first mapping table recording the interrupt number of each of the business processes, the mapping relationship between the publishing queue and the subscription queue, and return the publishing queue and the subscription queue to the host to establish a second mapping table recording the mapping relationship between the business process identifier, the IO request queue and the response queue, and the publishing queue and the subscription queue in the host, further comprising: In the host, an interrupt number is assigned to each of the service processes, and a publish queue and a subscription queue creation command is constructed using the interrupt number of the service process as a parameter, and the creation command is written into the management request queue; The NVME hard disk reads the create command in the management request queue through DMA; The publishing queue and the subscription queue are created according to the creation command, and a first mapping table is established to record the mapping relationship between the interruption number, the publishing queue and the subscription queue.

4. The method according to claim 3, characterized in that Also includes: Constructing response information with the addresses of the publishing queue and the subscription queue as parameters and writing the response information into the management response queue; The host obtains and parses the response information from the management response queue to obtain the addresses of the publishing queue and the subscription queue, and then establishes a second mapping table recording the mapping relationship between the business process identifier, the IO request queue and the response queue, and the publishing queue and the subscription queue.

5. The method according to claim 1, characterized in that In response to the business process sending a message subscription request to the NVME hard disk, a first IO write command is constructed using the message type to be subscribed and the address of the subscription queue determined according to the second mapping table and written into the corresponding IO request queue, and the first IO write command is written to the NVME hard disk to establish a third mapping table in the NVME hard disk that records the mapping relationship between the message type and the subscription queue, further comprising: Obtaining the message type, data block address, and callback function address to be subscribed that are passed in by the business process, and establishing a fourth mapping table that records the mapping relationship between the message type, the data block address, and the callback function address; Obtaining the address of the subscription queue corresponding to the business process in the second mapping table; Constructing a first IO write command and configuring a data pointer of the first IO write command to be an address corresponding to the message type, and SLBA to be the address of the subscription queue; Obtaining the IO request queue corresponding to the business process in the second mapping table and writing the first IO write command into the corresponding IO request queue; The NVME hard disk obtains the first IO write command from the IO request queue and parses to obtain the message type to be subscribed and the address of the subscription queue, thereby establishing a third mapping table that records the mapping relationship between the message type and the subscription queue; Constructing response information with the address of the subscription queue as a parameter and writing the response information to the response queue corresponding to the business process; The host obtains the response information from the response queue corresponding to the business process and ends the message subscription.

6. The method according to claim 5, characterized in that In response to detecting that there is a non-empty subscription queue, reporting the message content in the non-empty subscription queue to the corresponding business process according to the first mapping table, further comprising: Querying the first mapping table to obtain an interrupt number corresponding to the non-empty subscription queue; Update the length and message type of the message content in the non-empty subscription queue to the status register and notify the host of the length and message type of the message content by triggering the MSI interrupt; querying the fourth mapping table to obtain a data block address according to the message type, and querying the second mapping table to obtain a subscription queue address corresponding to the corresponding business process; Construct an IO read command and configure the data pointer of the IO read command to be the address corresponding to the data block, SLBA is the address of the subscription queue address corresponding to the corresponding business process, and NLB is the length of the message content; According to the second mapping table, the IO read command is sent to the IO request queue corresponding to the corresponding business process and the NVME hard disk is notified to execute the IO read command so that the message content in the non-empty subscription queue is moved to the address corresponding to the data block corresponding to the corresponding business process.

7. The method according to claim 6, characterized in that Also includes: The non-empty address of the subscription queue is used as a parameter to construct a response message and write the response message into the response queue corresponding to the corresponding business process; The host obtains the response information from the response queue corresponding to the corresponding business process and ends the message reporting; The fourth mapping table is queried according to the message type to obtain the address of the callback function so as to process the message content in the address corresponding to the data block through the callback function.

8. The method according to claim 1, characterized in that In response to the business process sending a message publishing request to the NVME hard disk, a second IO write command is constructed using the message type to be published, the corresponding publishing queue and the message content, and the corresponding IO request queue is written and the second IO write command is written to the NVME hard disk, and the message content and the message type are forwarded to the subscription queue corresponding to the message type to be published according to the third mapping table, further comprising: Obtain the message type, message content, and data block address to be published that are passed in by the business process; Obtaining the address of the subscription queue corresponding to the business process in the second mapping table; Constructing a second IO write command and configuring a data pointer of the second IO write command to be an address corresponding to the message type and an address corresponding to the message content, and SLBA is the address of the publishing queue; Obtaining the IO request queue corresponding to the business process in the second mapping table and writing the second IO write command into the corresponding IO request queue; The NVME hard disk obtains the second IO write command from the IO request queue and parses to obtain the address of the message type to be published and the address of the message content, executes DMA to obtain the message type and the message content to be published according to the address of the message type to be published and the address of the message content, and transmits them to the corresponding publishing queue; Forwarding the message type and the message content in the corresponding publishing queue to the subscription queue corresponding to the message type to be published according to the third mapping table; Constructing response information with the address of the publishing queue as a parameter and writing the response information to the response queue corresponding to the business process; The host obtains the response information from the response queue corresponding to the business process and ends the message publishing.

9. A computer device comprising: at least one processor; as well as A memory storing a computer program executable on the processor, wherein the processor executes the steps of the method according to any one of claims 1 to 8 when executing the program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are performed.

Citation Information

Patent Citations

  • System(s) and method(s) for multiple sender support in low latency FIFO messaging using TCP / IP protocol

    CN104639597A

  • Message queue architecture and interface for a multi-application platform

    US11277369B1