Fault-tolerant method and system for DMA efficient data transmission

By initializing the IO pool and hardware queue in the DMA system, using a ring queue to cache IO data and perform verification, the problem of high probability of reading incorrect data by DMA cache is solved, and the effect of improving DMA transmission availability, reliability and data consistency is achieved.

CN120086058AActive Publication Date: 2025-06-03CHENGDU HUARUI SHUXIN TECHNOLOGY CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510567728.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

In the prior art, the probability of reading incorrect data from the DMA cache is high, resulting in problems such as data reading and writing timeout, IO bandwidth reduction and data inconsistency.

Method used

By initializing the IO pool and hardware queue, setting the input cache queue IQ and output cache queue OQ, cache IO data using a ring queue, and verify the data, including data queue verification, IO data reference verification and data ID verification. For incorrect data, reread and reread verification are performed by setting a memory barrier and clearing the CPU cache.

Benefits of technology

It reduces the probability of reading incorrect data from the DMA cache, improves the availability, reliability and data consistency of DMA transmission, and avoids the problem of data inconsistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086058A_ABST
    Figure CN120086058A_ABST
Patent Text Reader

Abstract

The invention relates to the field of DMA (direct memory access) data transmission, in particular to a fault-tolerant method and system for DMA efficient data transmission, and solves the problem of the probability of reading error data from a DMA cache in the prior art. The method comprises the following steps: connecting a host equipment driving interface and a DMA; the DMA is arranged in a peripheral driver and comprises an IO pool for caching IO data and a hardware queue, and the DMA obtains the IO data and verifies the data; according to the method, the DMA transmission efficiency, availability, reliability and data consistency are improved through DMA annular buffering, CPU balancing and other modes; transforming and strengthening IO data ID (Identity) of the data of the issued peripheral; and the correctness of the data returned by the peripheral through the DMA is checked, so that the probability of reading error data from the DMA cache is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of DMA data transmission, and particularly to a fault-tolerant method and system for efficient DMA data transmission. Background Art

[0002] DMA (Direct Memory Access) is a function in a computer system. DMA is often used for communication between hardware devices such as network cards, RAID cards, hard disks, etc. and software operating systems. DMA data transmission is a direct interaction between a peripheral device and the host memory without passing through the CPU. When a software operating system processes data buffered by DMA, it can generally transmit data at high speed and correctly read and write data. When directly performing read and write operations on data between a peripheral device and the host through a DMA buffer. During the read process, the host reads data from the main memory mapped by DMA, and the main memory data needs to be cached by the CPU's cache before it can be processed by the CPU.

[0003] The load condition of the CPU will affect the situation where the cache fails to update data in a timely manner, thereby affecting the correctness of DMA data reading, especially when the CPU load is relatively large. What DMA updates is the data in the main memory. When the CPU load is relatively large, the old data copy remains in the CPU's cache, which may result in reading old data from the cache and incorrect data from the DMA cache.

[0004] Except for X86, most current CPUs adopt a weak consistency model, and the hardware does not automatically maintain cache consistency. There are phenomena such as incorrect memory refreshing and CPU cache invalidation when DMA writes data. If the peripheral devices are devices such as hard disks, network devices, and RAID cards, such phenomena of DMA will cause problems such as data read and write timeouts, reduced IO bandwidth, and data inconsistency.

[0005] There is a probability of reading incorrect data from the DMA cache. In the driver programs of many peripheral devices, a data verification function is enabled. When it is detected that the data is incorrect, the peripheral device is often directly shut down or restarted. Although these methods can reduce the possibility of data errors to a certain extent, due to the imperfect verification method, some data has been incorrect but not detected, which will cause serious errors and impacts. Moreover, the behaviors of directly shutting down and restarting the peripheral device greatly reduce the product experience, stability, efficiency, etc.

[0006] There is an urgent need for a new type of fault-tolerant method and system for efficient DMA data transmission that can solve the above problems. Summary of the Invention

[0007] The present invention provides a fault-tolerant method and system for efficient DMA data transmission, which solves the problem of the probability of reading incorrect data from the DMA cache in the prior art.

[0008] The technical solution of the present invention is implemented as follows: A fault-tolerant method for efficient DMA data transmission, including the following: Step 1: Initialize the system: Initialize the IO pool and the hardware queue. According to the number of CPUs and the maximum number of IOs M that the peripheral can accept, the length of the IO pool is less than M, and set the number of hardware queues N, where N is less than the number of CPUs; Initialize the input buffer queue IQ and the output buffer queue OQ of each hardware queue. The queue lengths of the input buffer queue IQ and the output buffer queue OQ are X, and X = MIN(M / N, 32). Initialize the PI and CI of the input buffer queue IQ and the output buffer queue OQ; where PI refers to the producer index, representing the position of the circular queue where data should be written currently, and CI refers to the consumer index, representing the position of the circular queue where the consumer should read data currently. Step 2: The IO pool receives IO data from the host device driver interface: If there is free space in the IO pool, save the IO data; if there is no free space, return the information that the queue is busy to the device driver interface. Step 3: The hardware queue obtains IO data from the IO pool as IO metadata, and saves the data index, ID, and the IO metadata of the IO pool to the input buffer queue IQ, and modifies PI. Step 4: The peripheral processes the IO metadata: The peripheral receives the IO metadata according to the change of PI, obtains the IO metadata according to the IO metadata, and processes it; after processing, modify the CI of the input buffer queue IQ, write the processing result to the IO pool queue and the output buffer queue OQ, and modify the PI of the output buffer queue OQ. Step 5: Output buffer queue OQ: After receiving the change of PI, the output buffer queue OQ processes the IO response message sent by the peripheral, accesses the buffer pointed to by CI in the output buffer queue OQ. The IO pool index saved in this buffer is used to find the corresponding IO data according to the IO pool index, compare and obtain the processed IO data and IO metadata, and verify the data, that is, data queue verification, IO data reference verification, and data ID verification. Step 6: Modify the CI of the output buffer queue OQ, complete the processing of the current IO data, release the space of the IO data in the IO pool and the hardware queue, and return the processed IO data to the device driver interface.

[0009] Specifically, in Step 3: According to the idle degree of each hardware queue, select a hardware queue to obtain IO metadata, and judge whether the ID in the original IO metadata can uniquely determine the IO metadata in the hardware queue, that is, whether the ID can uniquely identify the IO metadata in the input buffer queue IQ and the output buffer queue OQ; if the ID of the IO metadata cannot be uniquely determined, modify it: The first method: Generate a self-incrementing ID based on the hardware queue length X, do mapping inside the hardware queue, record the mapping relationship between request_ID and data VID+request_ID, and modify the ID to make the ID of the IO metadata unique in a separate hardware queue; The second method: Analyze the IO metadata data structure sent by the device driver interface, and use the unused space in the IO metadata data structure to regenerate the data VID. The length of the input cache queue IQ and the output cache queue OQ of the hardware queue are both less than 32. Only 5 bytes of space are needed to generate an automatically growing VID. Using 5 bytes of space can ensure the uniqueness of the ID. The 5-byte space can come from the reserved data area of ​​the IO metadata, or the original ID can be modified to generate a new IO data ID using the original ID + 5 bytes of space.

[0010] If the data verification fails in step 5, then go to step 7 to set a memory barrier: if the CI in the OP queue is verified to have a data verification error, then set a memory barrier to maintain the data consistency between the DMA cache data and the CPU cache data, re-read the DMA cache data, and perform data verification; if the verification data is correct, then go to step 6; if the verification fails, then go to step 8; Step 8: Clear the CPU cache: re-read the DMA cache data and perform data verification; if the verified data is correct, proceed to step 6; if the correct data is still not read, then record the CI value of the output cache queue OQ unreachable data in the hardware queue, classify the IO metadata as unreachable data, and proceed to step 9; Step 9: Modify the CI of the output buffer queue OQ: notify the peripheral device to complete the current processing, but do not notify the device driver interface of the failure of the current processing; set the timeout processing function of the unreached data, which is asynchronous processing, that is, after asynchronously waiting for the set time, notify the device driver interface that the unreached data has timed out; wait for data to time out; Step 10: The output buffer queue OQ continues to wait for new data: within the timeout period, the output buffer queue OQ is waited for whether new data arrives; if there is no new data, the device driver interface is notified that the IO data processing has timed out; if there is new data, then step 11 is entered; Step 11: The output buffer queue OQ receives new data from the peripheral: that is, the PI of the output buffer queue OQ changes, then check whether there is a record of unreached data in the hardware queue, if there is no unreached data, then go to step 5; if there is unreached data, then go to step 12; Step 12: Determine the position of the CI of the unarrived data pointing to the output buffer queue OQ: Whether the data at this position has been updated. If it has not been updated or the data check is incorrect, then notify the device driver interface that the IO data processing has timed out, delete the record of the unarrived data in the hardware queue, and complete the processing of the unarrived data; If there is data update and the data verification is correct, then go to Step 13; Step 13: Access the data at the CI position of the unarrived data, complete the current IO data processing, and return the IO data processing result to the device driver interface; Regardless of whether the data is correct, the hardware queue deletes the record of the unarrived data; Resume the processing of new data in the output buffer queue OQ, that is: go to Step 5.

[0011] A fault-tolerant system for efficient DMA data transmission, connecting the host device driver interface and the peripheral for transmitting IO data; It includes a DMA respectively connected to the device driver interface and the peripheral; The DMA is set in the peripheral driver and includes an IO pool for caching IO data and a hardware queue, receives IO data from the host device driver interface, caches the IO data through the DMA, and transmits the IO data to the peripheral through the DMA. After the peripheral finishes processing the IO data, it returns the processing result through the DMA; The hardware queue mainly includes an output buffer queue OQ and an input buffer queue IQ inside; The IO pool is connected to the host through the device driver interface and transmits the buffered IO data to the input buffer queue IQ of the hardware queue; The peripheral receives IO data from the input buffer queue IQ and writes the processed IO data into the output buffer queue OQ; The DMA obtains the IO data and verifies the data.

[0012] The IO pool caches the IO data from the host device driver interface in a circular queue manner; The unit of the IO pool circular queue is the request of the data structure memory space; The request includes a data ID, IO metadata, and the IO metadata includes IO data type, IO data buffer corresponding DMA physical address information, and IO data time information; When there are idle resources in the input buffer queue IQ of the hardware queue, the information in the IO pool is sent to the input buffer queue IQ for caching the data index, original data, and DMA data type information of the IO pool; The original data includes a data VID and a request_ID; The input buffer queue IQ includes two pointer indexes PI and CI respectively pointing to the circular queue.

[0013] The peripheral obtains the original data from the input buffer queue IQ, obtains the complete IO data information through the DMA physical address information and processes it; The peripheral judges whether there is new data coming according to the change of PI, processes the newly arrived IO data, and modifies CI after processing.

[0014] The output cache queue OQ caches the IO pool index information, IO metadata, and IO data processing results returned after the peripheral processes the IO completion; the output cache queue OQ adopts a circular queue method, and the output cache queue OQ includes two pointer indexes respectively pointing to the PI and CI of the output cache queue OQ.

[0015] The device driver interface can be a host storage layer protocol or a host network layer protocol.

[0016] Improve the availability, reliability, and data consistency of DMA transmission, especially for PCIE-based peripheral cards. In the case of continuous high pressure of IO data, high CPU load, and IO data bursts, when the latest data in the DMA buffer cannot be accessed in time, data inconsistency problems can be avoided. The present invention improves DMA by providing a DMA circular cache with hardware queues as units, with each CPU corresponding to a hardware queue; by uniquely identifying each issued data IO, and by means of techniques such as setting memory barriers for DMA data rereading, clearing CPU cache rereading, and readback, to reduce the problems caused by the inability of the CPU's Cache to be updated in time, and to improve data fault tolerance without affecting the original data processing efficiency of the original RAID card driver, and to improve the reliability, availability, and data consistency of peripheral drivers such as RAID cards, network cards, and hard disks.

[0017] The present invention has the following beneficial effects: Transform the data sent to the peripheral, and strengthen the IO data ID without destroying the original data structure; perform correctness verification on the data returned by the peripheral through DMA, and reduce the probability of reading incorrect data from the DMA cache; For incorrect data, by setting memory barriers and clearing the CPU cache, reread and reread verify the DMA data, provide strong fault tolerance, and ensure data correctness; Add a method for readback of the DMA buffer to avoid data loss; Improve the DMA transmission efficiency, availability, reliability, and data consistency through DMA circular buffering and CPU balancing. Brief Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 : Schematic diagram of the DMA high-speed transmission framework; Figure 2 : DMA High-Speed Data Transmission and Fault-Tolerant Flowchart; Figure 3 : Schematic Diagram of Data Processing. Specific Embodiment

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0021] Combined with Figure 1 Schematic Diagram of DMA High-Speed Transmission Framework, Figure 2 DMA High-Speed Data Transmission and Fault-Tolerant Flowchart, and Figure 3 Schematic Diagram of Data Processing, a fault-tolerant system for efficient DMA data transmission disclosed by the present invention is connected to a host device driver interface and a peripheral for transmitting IO data; it includes a DMA respectively connected to the device driver interface and the peripheral; The DMA is set in the peripheral driver and includes an IO pool and a hardware queue for caching IO data, receives IO data from the host device driver interface, caches the IO data through the DMA, and transmits the IO data to the peripheral through the DMA. After the peripheral processes the IO data, the processing result is returned through the DMA; the hardware queue mainly includes an output cache queue OQ and an input cache queue IQ inside; the IO pool is connected to the host through the device driver interface and transmits the buffered IO data to the input cache queue IQ of the hardware queue; the peripheral receives the IO data from the input cache queue IQ and writes the processed IO data into the output cache queue OQ; the output cache queue OQ returns the IO data to the host; the DMA obtains the IO data and verifies the data. In addition to data queue verification and IO data reference verification, it also verifies whether the data ID is correct.

[0022] The IO pool caches IO data from the host device driver interface in a circular queue manner; the unit of the IO pool circular queue is a request for the data structure memory space; the request includes a data ID, IO metadata, and the IO metadata includes an IO data type, the DMA physical address information corresponding to the IO data buffer, and the IO data time information; when there are idle resources in the input buffer queue IQ of the hardware queue, the information in the IO pool is sent to the input buffer queue IQ for caching the data index, original data, and DMA data type information of the IO pool; the original data includes a data VID and a request_ID; the request_ID is used by the device driver interface to identify the uniqueness of the IO data, that is, the original ID of the IO data, and the data VID is an additional data ID added in the present invention to ensure data uniqueness. The data VID and the request_ID together form the data ID; the input buffer queue IQ includes two pointer indexes, PI and CI, respectively pointing to the circular queue.

[0023] The peripheral obtains the original data from the input buffer queue IQ, obtains the complete IO data information through the DMA physical address, and processes it; the peripheral determines whether there is new data according to the change of PI, and processes the newly arrived IO data. After the processing is completed, CI is modified; according to the maximum cache quantity that the peripheral can process and the maximum cache quantity that the host driver device can process, it is determined whether PI and CI are stored in the peripheral register or the DMA cache.

[0024] The output buffer queue OQ caches the IO pool index information, IO metadata, and IO data processing results returned after the peripheral finishes processing the IO; the output buffer queue OQ adopts a circular queue manner. The output buffer queue OQ includes two pointer indexes, PI and CI, respectively pointing to the output buffer queue OQ. According to the physical characteristics of the peripheral, it is determined whether PI and CI are stored in the peripheral register or the DMA cache.

[0025] The device driver interface can be a host storage layer protocol or a host network layer protocol.

[0026] A fault-tolerant method for DMA efficient data transmission using the above system includes the following: Step 1: Initialize the system: Initialize the IO pool and the hardware queue. According to the number of CPUs and the maximum number of IOs M that the peripheral can receive, the length of the IO pool is less than M, and the number of hardware queues N is set, where N is less than the number of CPUs; Initialize the input buffer queue IQ and the output buffer queue OQ of each hardware queue. The queue lengths X of the input buffer queue IQ and the output buffer queue OQ are X = MIN(M / N, 32). Initialize the PI and CI of the input buffer queue IQ and the output buffer queue OQ. Among them, CI refers to the consumer index, and PI refers to the producer index. They are two key pointer variables. They are jointly used to manage and track the status of data in the circular queue cache, so as to efficiently utilize the array space.

[0027] PI is used to indicate which position in the circular queue the producer should write data to currently. It always points to the next available write position in the circular queue. The producer writes data to the position pointed to by the PI pointer in the circular queue cache. After writing the data, the producer modifies the PI pointer to make it point to the next available write position. Usually, the PI pointer moves cyclically according to the size of the circular queue, that is, when PI reaches the end of the queue, it will "wrap around" to the beginning of the queue.

[0028] CI is used to indicate which position in the circular queue the consumer should read data from currently. It always points to the next available read position in the circular queue. The consumer reads data from the position pointed to by the CI pointer in the circular queue cache. After reading the data, the consumer modifies the CI pointer to make it point to the next available read position. Similarly, the CI pointer also moves cyclically according to the size of the circular queue. When the peripheral uses the circular queue to send data to the host, the peripheral is the producer and the host is the consumer. When the host uses the circular queue to send data to the peripheral, the host is the producer and the peripheral is the consumer.

[0029] Step 2: The IO pool receives IO data from the host device driver interface: If there is free space in the IO pool, save the IO data; if there is no free space, return the information that the queue is busy to the device driver interface; Step 3: The hardware queue obtains the IO data as IO metadata from the IO pool, and saves the data index, ID, and this IO metadata of the IO pool to the input cache queue IQ, and modifies PI; Step 4: The peripheral processes the IO metadata: The peripheral receives the IO metadata according to the change of PI, obtains the IO metadata according to the IO metadata and processes it; after processing, modify CI of the input cache queue IQ, write the processing result to the IO pool queue and the output cache queue OQ, and modify PI of the output cache queue OQ; Step 5: The output cache queue OQ: After receiving the change of PI, the output cache queue OQ processes the IO response message sent by the peripheral, accesses the cache pointed to by CI in the output cache queue OQ, the IO pool index saved in this cache, finds the corresponding IO data according to the IO pool index, compares and obtains the processed IO data and IO metadata, and verifies the data, that is, data queue verification, IO data reference verification and data ID verification; Step 6: Modify the CI of the output buffer queue OQ, complete the IO data processing, release the space of the IO data in the IO pool and the hardware queue, and return the processed IO data to the device driver interface.

[0030] Specifically, step 3 includes: selecting a hardware queue to obtain IO metadata according to the idleness of each hardware queue, and determining whether the ID in the original IO metadata can uniquely identify the IO metadata in the queue of the hardware queue, that is, whether the ID can uniquely identify the IO metadata in the input buffer queue IQ and the output buffer queue OQ; if the ID of the IO metadata cannot be uniquely determined, modifying it: The first method: Generate a self-incrementing ID based on the hardware queue length X, do mapping inside the hardware queue, record the mapping relationship between request_ID and data VID+request_ID, and modify the ID to make the ID of the IO metadata unique in a separate hardware queue; The second method: Analyze the IO metadata data structure sent by the device driver interface, and use the unused space in the IO metadata data structure to regenerate the data VID. The length of the input cache queue IQ and the output cache queue OQ of the hardware queue are both less than 32. Only 5 bytes of space are needed to generate an automatically growing VID. Using 5 bytes of space can ensure the uniqueness of the ID. The 5-byte space can come from the reserved data area of ​​the IO metadata, or the original ID can be modified to generate a new IO data ID using the original ID + 5 bytes of space.

[0031] If the data verification fails in step 5, then go to step 7 to set a memory barrier: if the CI in the OP queue is verified to have a data verification error, then set a memory barrier to maintain the data consistency between the DMA cache data and the CPU cache data, re-read the DMA cache data, and perform data verification; if the verification data is correct, then go to step 6; if the verification fails, then go to step 8; Step 8: Clear the CPU cache: re-read the DMA cache data and perform data verification; if the verified data is correct, proceed to step 6; if the correct data is still not read, then record the CI value of the output cache queue OQ unreachable data in the hardware queue, classify the IO metadata as unreachable data, and proceed to step 9; Step 9: Modify the CI of the output buffer queue OQ: notify the peripheral device to complete the current processing, but do not notify the device driver interface of the failure of the current processing; set the timeout processing function of the unreached data, which is asynchronous processing, that is, after asynchronously waiting for the set time, notify the device driver interface that the unreached data has timed out; wait for data to time out; Step 10: The output cache queue OQ continues to wait for new data: Wait within the timeout period to check if new data arrives in the output cache queue OQ; if there is no new data, then notify the device driver interface that the IO data processing has timed out; if new data arrives, then proceed to Step 11; Step 11: The output cache queue OQ receives new data from the peripheral: That is, if the PI of the output cache queue OQ changes, then check if there are records of unprocessed data in the hardware queue. If there is no unprocessed data, then proceed to Step 5; if there is unprocessed data, then proceed to Step 12; Step 12: Determine if the CI of the unprocessed data points to the position in the output cache queue OQ: Whether the data at this position has been updated. If it has not been updated or the data check is incorrect, then notify the device driver interface that the IO data processing has timed out, delete the record of the unprocessed data in the hardware queue, and complete the processing of the unprocessed data; if the data has been updated and the data verification is correct, then proceed to Step 13; Step 13: Access the data at the CI position of the unprocessed data, complete the current IO data processing, and return the IO data processing result to the device driver interface; regardless of whether the data is correct, delete the record of the unprocessed data in the hardware queue to avoid getting stuck in a loop of processing unprocessed data; regardless of whether the data is correct, the unprocessed data is only processed once, and the record of the unprocessed data will be deleted from the hardware queue; resume the processing of new data in the output cache queue OQ, that is: proceed to Step 5.

[0032] The present invention improves the availability, reliability, and data consistency of DMA transmission, especially for PCIE-based peripheral cards. In the case of continuous high pressure of IO data, high CPU load, and IO data bursts, when the latest data in the DMA buffer cannot be accessed in time, data inconsistency problems are avoided. The present invention provides a DMA circular cache with a hardware queue as a unit, with each CPU corresponding to a hardware queue to improve DMA; by uniquely identifying each issued data IO, and using technical means such as setting memory barriers for DMA data rereading, clearing CPU cache rereading, and readback, to reduce problems caused by the CPU cache not being updated in time, and improve data fault tolerance without affecting the original data processing efficiency of the original RAID card driver, and improve the reliability, availability, and data consistency of peripheral drivers such as RAID cards, network cards, and hard disks.

[0033] The present invention has the following beneficial effects: Transform the data sent to the peripheral, and strengthen the IO data ID without destroying the original data structure; perform correctness verification on the data returned by the peripheral through DMA, reducing the probability of reading incorrect data from the DMA cache; For incorrect data, by setting memory barriers and clearing the CPU cache, the DMA data is reread and reread-verified to provide strong fault tolerance and ensure data correctness; Add a method for reading back the DMA buffer to avoid data loss; Improve the DMA transfer efficiency, availability, reliability, and data consistency through methods such as DMA circular buffering and CPU balancing.

[0034] Of course, without departing from the spirit and essence of the present invention, those skilled in the art should be able to make various corresponding changes and modifications according to the present invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims of the present invention.

Claims

1. A fault-tolerant method for DMA efficient data transmission, characterized in that: Includes the following: Step 1: Initialize the system: Initialize the IO pool and hardware queues, set the number of hardware queues N according to the number of CPUs and the maximum number of IOs M accepted by the peripherals, where the length of the IO pool is less than M, and N is less than the number of CPUs; initialize the input cache queue IQ and output cache queue OQ of each hardware queue, the queue length of the input cache queue IQ and the output cache queue OQ is X, X=MIN(M / N,32), initialize the PI and CI of the input cache queue IQ and the output cache queue OQ; where PI refers to the producer index, representing the position of the ring queue where the data should be written at present, and CI refers to the consumer index, representing the position of the ring queue where the consumer should read the data at present; Step 2: The IO pool receives IO data: The IO data comes from the host driver device driver interface. If the IO pool has free space, the IO data is saved; if the IO pool has no free space, the queue busy information is returned; Step 3: The hardware queue obtains IO data from the IO pool as IO metadata, saves the data index, ID, and IO metadata of the IO pool to the input buffer queue IQ, and modifies the PI; Step 4: The peripheral processes the IO metadata: The peripheral receives the IO metadata according to the change of PI, obtains the IO metadata and processes it according to the IO metadata; After the processing is completed, modify the CI of the input buffer queue IQ, write the processing results to the IO pool queue and the output buffer queue OQ, and modify the PI of the output buffer queue OQ; Step 5: Output buffer queue OQ: After receiving the PI change, the output buffer queue OQ processes the IO response message sent by the peripheral, and accesses the cache pointed to by CI in the output buffer queue OQ. The cache stores the IO pool index, finds the corresponding IO data according to the IO pool index, compares the processed IO data and IO metadata, and verifies the data, namely, data queue verification, IO data reference verification, and data ID verification; Step 6: Modify the CI of the output buffer queue OQ, complete the IO data processing, release the space of the IO data in the IO pool and the hardware queue, and return the processed IO data to the device driver interface.

2. The fault-tolerant method for DMA efficient data transmission according to claim 1, characterized in that: Specifically, step 3 includes: selecting a hardware queue to obtain IO metadata according to the idleness of each hardware queue, and determining whether the ID in the original IO metadata can uniquely identify the IO metadata in the queue of the hardware queue, that is, whether the ID can uniquely identify the IO metadata in the input buffer queue IQ and the output buffer queue OQ; if the ID of the IO metadata cannot be uniquely determined, modifying it: The first method: Generate a self-incrementing ID based on the hardware queue length X, do mapping inside the hardware queue, record the mapping relationship between request_ID and data VID+request_ID, and modify the ID to make the ID of the IO metadata unique in a separate hardware queue; The second method: Analyze the IO metadata data structure sent by the device driver interface, and use the unused space in the IO metadata data structure to regenerate the data VID. The length of the input cache queue IQ and the output cache queue OQ of the hardware queue are both less than 32. Only 5 bytes of space are needed to generate an automatically growing VID. Using 5 bytes of space can ensure the uniqueness of the ID. The 5-byte space comes from the reserved data area of ​​the IO metadata, or the original ID is modified, and the original ID + 5 bytes of space are used to generate a new IO data ID.

3. The fault-tolerant method for DMA efficient data transmission according to claim 2, characterized in that: If the data verification fails in step 5, step 7 is performed to set a memory barrier: if the CI in the OP queue is verified to have a data verification error, a memory barrier is set to maintain the data consistency between the DMA cache data and the CPU cache data, and the DMA cache data is re-read and the data is verified; If the verification data is correct, then go to step 6; if the verification fails, then go to step 8; Step 8: Clear the CPU cache: re-read the DMA cache data and perform data verification; if the verified data is correct, proceed to step 6; if the correct data is still not read, then record the CI value of the output cache queue OQ unreachable data in the hardware queue, classify the IO metadata as unreachable data, and proceed to step 9; Step 9: Modify the CI of the output buffer queue OQ: notify the peripheral device to complete this processing, but do not notify the device driver interface that this processing failed; set the timeout processing function of the unreached data, which is asynchronous processing, that is, after asynchronously waiting for the set time, notify the device driver interface that the unreached data has timed out; wait for data to time out; Step 10: The output buffer queue OQ continues to wait for new data: wait for the output buffer queue OQ to have new data within the timeout period; if there is no new data, then notify the device driver interface that the IO data processing has timed out; if there is new data, then go to step 11; Step 11: The output buffer queue OQ receives new data from the peripheral device: that is, the PI of the output buffer queue OQ changes, then check whether there is a record of unreached data in the hardware queue, if there is no unreached data, then go to step 5; if there is unreached data, then go to step 12; Step 12: Determine the CI of the unreachable data pointing to the output buffer queue OQ position: whether the data at this position has been updated; if not updated or the data check is incorrect, then notify the device driver interface that the IO data processing has timed out, delete the record of the unreachable data in the hardware queue, and complete the processing of the unreachable data; if the data has been updated and the data check is correct, then go to step 13; Step 13: Access data at the CI position of the unreached data, complete the IO data processing, and return the IO data processing result to the device driver interface; regardless of whether the data is correct, the hardware queue deletes the record of the unreached data; resume the new data processing in the output cache queue OQ, that is: enter step 5.

4. A fault-tolerant system for DMA efficient data transmission, connecting a host device driver interface and a peripheral device for transmitting IO data; characterized in that: Including DMA connected to device driver interface and peripherals respectively; The DMA is set in the peripheral driver, including an IO pool and a hardware queue for caching IO data, receiving IO data from the host device driver interface, caching IO data, and transmitting IO data to the peripheral. After the peripheral processes the IO data, it returns the processing result; The hardware queue mainly includes an output buffer queue OQ and an input buffer queue IQ; The IO pool is connected to the host through the device driver interface, and transmits the buffered IO data to the input buffer queue IQ of the hardware queue; the peripheral receives the IO data from the input buffer queue IQ, and writes the processed IO data into the output buffer queue OQ; The DMA acquires IO data and verifies the data.

5. A fault-tolerant system for DMA efficient data transmission according to claim 4, characterized in that: The IO pool caches IO data from the host device driver interface in a circular queue manner; The unit of the IO pool circular queue is a request for data structure memory space; the request includes data ID and IO metadata; the IO metadata includes IO data type, IO data buffer and IO data time information; the IO data buffer corresponds to DMA physical address information; When the input buffer queue IQ has free resources, the hardware queue sends the information in the IO pool to the input buffer queue IQ: data index, original data and DMA data type information for caching the IO pool; the original data includes data VID and request_ID; The input buffer queue IQ includes two pointer indexes PI and CI pointing to the circular queues respectively.

6. A fault-tolerant system for DMA efficient data transmission according to claim 5, characterized in that: The peripheral obtains the original data from the input buffer queue IQ, obtains the complete information of the IO data through the DMA physical address information and processes it; The peripheral device determines whether new data has arrived according to the change of PI, processes the newly arrived IO data, and modifies CI after the processing is completed.

7. A fault-tolerant system for DMA efficient data transmission according to claim 6, characterized in that: The output buffer queue OQ caches the IO pool index information, IO metadata and IO data processing results returned after the peripheral device completes the IO processing; The output cache queue OQ adopts a circular queue mode, and the output cache queue OQ includes two pointer indexes PI and CI pointing to the output cache queue OQ respectively.

8. A fault-tolerant system for DMA efficient data transmission according to any one of claims 5 to 7, characterized in that: The device driver interface is a host storage layer protocol or a host network layer protocol.

Citation Information

Patent Citations

  • System and method for improving direct memory access (DMA) efficiency of multi-data buffer

    CN102541779A

  • Data transmission method based on DMA

    CN104123250A

  • DMA data transmission control system

    CN116225534A

  • Efficient data sampling processing method and computer equipment

    CN118502892A

  • DMA engine controller and control method thereof, electronic equipment and storage medium

    CN118860925A