Data Transmission Method and Device

By dividing the dma buffer logic into multiple sub-blocks and combining multi-threading technology, the problems of low space utilization and unsatisfactory transmission in data transmission between PCIe host and PCIe board equipment are solved, and more efficient data transmission is achieved.

CN111666228BActive Publication Date: 2025-07-18NEW H3C SEMICON TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010395016.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-12
Publication Date
2025-07-18
Estimated Expiration
2040-05-12

AI Technical Summary

Technical Problem

In the prior art, the data transmission between the PCIe host and the PCIe board and card device has the problem that the dma buffer space configuration is not flexible enough, the space utilization rate is low, and the data exchange transmission rate is not ideal enough.

Method used

By dividing the dma buffer logic into multiple sub-blocks, and combining multi-threading technology, PCIe host and PCIe board equipment obtain and update the head and tail values of the sub-blocks respectively, to achieve flexible writing and reading of data and improve data transmission efficiency.

Benefits of technology

On the premise of sharing a physical dma buffer space, multiple types of data operations can be performed simultaneously without affecting each other, improving data transmission efficiency, and solving the problems of low space utilization and unsatisfactory transmission rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111666228B_ABST
    Figure CN111666228B_ABST
Patent Text Reader

Abstract

The present application provides a data transmission method and apparatus. The method is applied to a CPU included in a PCIe host, and the method includes: obtaining a block semaphore, where the block semaphore is used to indicate a sub-block in a dma buffer block included in a PCIe board device; determining a first sub-block from the dma buffer block according to the block semaphore; obtaining, from defined global variables, a head value stored in a head address field included in the first sub-block and a base address of the first sub-block; obtaining, from the dma buffer block, a tail value stored in a tail address field included in the first sub-block; when the head value of the first sub-block is equal to the tail value, updating the head value according to the length value of data to be written; and writing the data to be written into a sub-buffer field included in the first sub-block through a PCIe bus according to the base address and the updated head value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a data transmission method and apparatus. Background Art

[0002] Peripheral Component Interconnect Express (PCle) is a bus and interface standard, that is, a device connection method of point-to-point serial connection. When each device transmits data, it establishes its own dedicated transmission channel to avoid interference from other devices. Direct Memory Access (DMA) is a data exchange mode that directly accesses data from memory without going through the CPU, and it is an important technology for solving data interaction between memory and external chips. The research and application of data transmission methods combining the advantages of both are also gradually carried out, and currently, there are more and more communication devices (such as routers, switches) based on PCIe and DMA.

[0003] Currently, by running an application program on a PCIe board device (such as a C-programmable Task Optimized Processor (ctop)), the data of a PCIe host (such as a host cpu) can be quickly transferred to the PCIe board device. The brief process of data transmission between devices is as Figure 1 shown Figure 1 is a block diagram of the data transmission process through PCIe in the prior art.

[0004] The cpu writes data into the dma buffer through the PCIe channel, and the application program running on the ctop continuously reads the value at the specified position of the dma buffer to determine whether to perform data transmission. If data transmission can be performed, the data read by the application program is first temporarily stored in the cmem, and then the data is stored at the specified position in the emem.

[0005] Assume that a dma buffer space with a physical size of 16K is used as the data transfer buffer space, as Figure 2As shown in the figure, dma addr represents the starting address of the dma buffer. Each time the cpu operates on the data, it offsets the starting address by the length of head; while ctop operates on the data by offsetting the starting address by the length of tail each time. Head addr represents the offset used to store the data when the cpu operates. When the value of head reaches the buffer length, it is reset to 0; tail addr represents the offset used to store the data when ctop operates. When the value of tail reaches the buffer length, it is reset to 0. Therefore, after each data transfer operation is completed, the values of head and tail are the same.

[0006] During the data transfer process, the application program continuously reads the values of head and tail stored in head addr and tail addr. If the value of head is greater than the value of tail, it means that the cpu has written data into the dma buffer, and the length of this data is the difference between the value of head and the value of tail. Ctop reads the data to complete the data transfer.

[0007] The data transfer method given in the prior art only simply realizes the function of data exchange between the PCIe host and the PCIe board device, but it has the disadvantages of inflexible space configuration, low space utilization rate, and less than ideal data exchange and transfer rate.

[0008] First, each data exchange operation occupies the entire dma buffer exclusively and locks the dma buffer when in use. In this way, only a single operation can be performed at the same time. Second, the dma buffer space is wasted greatly. Since only 128B or 256B of space is used at a time during the write operation, however, the space occupied by each operation on the dma buffer alone is much larger than the space required during the write operation; the space occupied during the read operation is also excessive, resulting in serious resource waste. The above reasons lead to an unsatisfactory data exchange and transfer rate. Summary of the Invention

[0009] In view of this, the present application provides a data transfer method and device to solve the disadvantages in the prior art, such as inflexible space configuration of the dma buffer, low space utilization rate, and less than ideal data exchange and transfer rate.

[0010] In a first aspect, the present application provides a data transfer method, which is applied to the cpu included in the PCIe host. The PCIe host is connected to the PCIe board device through the PCIe bus. The method includes:

[0011] Obtain a block semaphore, where the block semaphore is used to indicate a sub-block in a dma buffer included in the PCIe board device, and the dma buffer includes multiple sub-blocks;

[0012] According to the block semaphore, determine a first sub-block from the multiple sub-blocks included in the dma buffer;

[0013] From the defined global variables, obtain the head value stored in the head address field included in the first sub-block and the base address of the first sub-block;

[0014] From the dma buffer, obtain the tail value stored in the tail address field included in the first sub-block;

[0015] When the head value of the first sub-block is equal to the tail value, update the head value according to the length value of the data to be written;

[0016] According to the base address and the updated head value, write the data to be written into the sub-buffer field included in the first sub-block through the PCIe bus.

[0017] In a second aspect, the present application provides a data transmission method, which is applied to a ctop included in a PCIe board device. The PCIe board device further includes a cpu core and a dma buffer, and the dma buffer includes multiple sub-blocks. The PCIe board device is connected to a PCIe host through a PCIe bus. The method includes:

[0018] Allocate a sub-block included in the dma buffer to a thread running in the cpu core;

[0019] For each sub-block, obtain the head value stored in the head address field included in the sub-block and the tail value stored in the tail address field included in the sub-block;

[0020] When the head value is not equal to the tail value, obtain the data transmitted by the PCIe host through the PCIe bus from the sub-buffer field included in the sub-block, and the size of the data is the difference between the head value and the tail value.

[0021] In a third aspect, the present application provides a data transmission device, which is applied to a cpu included in a PCIe host. The PCIe host is connected to a PCIe board device through a PCIe bus. The device includes:

[0022] An acquisition unit, configured to acquire a block semaphore, where the block semaphore is used to indicate a sub-block in a dma buffer included in the PCIe board device, and the dma buffer includes multiple sub-blocks;

[0023] A determination unit, configured to determine a first sub-block from the multiple sub-blocks included in the dma buffer according to the block semaphore;

[0024] The acquisition unit is further configured to acquire, from a defined global variable, a head value stored in a head address field included in the first sub-block and a base address of the first sub-block;

[0025] The acquisition unit is further configured to acquire, from the dma buffer, a tail value stored in a tail address field included in the first sub-block;

[0026] An update unit, configured to update the head value according to a length value of data to be written when the head value of the first sub-block is equal to the tail value;

[0027] A write unit, configured to write the data to be written into a sub-buffer field included in the first sub-block through the PCIe bus according to the base address and the updated head value.

[0028] In a fourth aspect, the present application provides a data transmission device, where the device is applied to a ctop included in a PCIe board device, the PCIe board device further includes a cpu core and a dma buffer, the dma buffer includes multiple sub-blocks, the PCIe board device is connected to a PCIe host through a PCIe bus, and the device includes:

[0029] An allocation unit, configured to allocate a sub-block included in the dma buffer to a thread running in the cpu core;

[0030] An acquisition unit, configured to, for each sub-block, acquire a head value stored in a head address field included in the sub-block and a tail value stored in a tail address field included in the sub-block;

[0031] The acquisition unit is further configured to, when the head value is not equal to the tail value, acquire data transmitted by the PCIe host through the PCIe bus from a sub-buffer field included in the sub-block, and a size of the data is a difference between the head value and the tail value.

[0032] Therefore, by applying a data transmission method and apparatus provided in this application, a PCIe host and a PCIe board device are connected through a PCIe bus. The CPU obtains a chunk semaphore for indicating multiple sub-chunks in a DMA buffer included in the PCIe board device. According to the chunk semaphore, the CPU determines a first sub-chunk from the multiple sub-chunks included in the DMA buffer; from the defined global variables, the CPU obtains a head value and a base address stored in a head address field included in the first sub-chunk. From the DMA buffer, the CPU obtains a tail value stored in a tail address field included in the first sub-chunk. When the head value of the first sub-chunk is equal to the tail value, the CPU updates the head value according to the length value of the data to be written. According to the base address and the updated head value, through the PCIe bus, the CPU writes the data to be written into a sub-buffer field included in the first sub-chunk.

[0033] In the foregoing manner, each operation only requires the PCIe host to update the data, and the program on the PCIe board device actively reads the data in the DMA buffer, reducing the scheduling overhead of the dual CPUs; on the premise of sharing a physical DMA buffer space, the DMA buffer is logically divided into N chunks and the multi-threading technology is combined to operate different positions of the DMA buffer so that multiple types of data or read / write operations can be performed simultaneously without affecting each other, improving the data transmission efficiency by N times; solving the disadvantages in the prior art that the configuration of the DMA buffer space is not flexible enough, the space utilization rate is low, and the data exchange and transmission rate is not ideal enough. Description of the Drawings

[0034] Figure 1 It is a block diagram of the process of transmitting data through PCIe in the prior art;

[0035] Figure 2 It is a schematic diagram of the structure of the DMA buffer in the prior art;

[0036] Figure 3 It is a flowchart of a data transmission method provided by an embodiment of this application;

[0037] Figure 4 It is a schematic diagram of the structure of the DMA buffer after chunking provided by an embodiment of this application;

[0038] Figure 5 It is a flowchart of another data transmission method provided by an embodiment of this application;

[0039] Figure 6 It is a block diagram of the process of transmitting data through PCIe provided by an embodiment of this application;

[0040] Figure 7Structural diagram of a data transmission device provided by an embodiment of the present application

[0041] Figure 8 Another structural diagram of a data transmission device provided by an embodiment of the present application. Detailed implementation manners

[0042] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0043] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0044] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0045] The data transmission method provided by the embodiment of the present application will be described in detail below. Refer to Figure 3 , Figure 3 It is a flowchart of a data transmission method provided by an embodiment of the present application. This method is applied to the CPU included in the PCIe host, and the PCIe host is connected to the PCIe board device through the PCIe bus, and specifically includes the following steps.

[0046] Step 310: Obtain a block semaphore, where the block semaphore is used to indicate a sub-block in the dma buffer included in the PCIe board device, and the dma buffer includes multiple sub-blocks.

[0047] Specifically, in the embodiments of the present application, the CPU allocates a base address for the DMA buffer and configures the size of the DMA buffer. Without affecting the operation performance, the PCIe host logically divides the physical space of the DMA buffer included in the PCIe board device into blocks. Each sub-block after division has the same size of storage space. Each sub-block includes a sub-buffer field, a head address field, a tail address field, an error code address field, and a reply address field.

[0048] Taking a DMA buffer with a size of 16K, logically divided into 4 blocks and running 4 threads as an example, the structure of the DMA buffer after division is as Figure 4 shown.

[0049] In the DMA buffer, each sub-block has the same structure. Taking the first sub-block as an example, dma addr_1 is the base address of the sub-block (also known as the starting address); head addr_1 is used to store the offset relative to the base address (this base address is the base address of the DMA buffer) after each operation of the PCIe host on the DMA buffer. When this value reaches the length of the DMA buffer, it is reset to 0; tail addr_1 is used to store the offset relative to the base address (this base address is the base address of the DMA buffer) after each operation of the PCIe board device on the DMA buffer. When this value reaches the length of the DMA buffer, it is also reset to 0.

[0050] The CPU defines a global variable, which is used to store the base address of each sub-block and the values of each address field included in each sub-block. The foregoing content stored in the global variable is used by the PCIe host.

[0051] The CPU also writes the base address of the DMA buffer, the configured available size of the DMA buffer, the base address of the sub-block, and the values of each address field included in each sub-block into the address register together. The CPU also sets the status register to 1 to indicate that the initialization of the PCIe host is completed.

[0052] It can be understood that the address register and the status register can be accessed by the PCIe board device. In this way, the PCIe board device can clarify the current status of the PCIe host and the configuration of the PCIe host for the DMA buffer.

[0053] When the PCIe host needs to transmit data to the PCIe board device, the PCIe host first needs to select a sub-block from the DMA buffer as the target sub-block and write the data to be transmitted into the target sub-block.

[0054] Specifically, the CPU defines a variable in a task that uses the DMA buffer. After performing an operation on the variable, the CPU obtains a block semaphore. The block semaphore is used to indicate the sub-blocks in the DMA buffer included in the PCIe board device.

[0055] Step 320: Determine a first sub-block from the multiple sub-blocks included in the DMA buffer according to the block semaphore.

[0056] Specifically, according to the block semaphore, the CPU locates a certain sub-block (for example, the first sub-block) in the DMA buffer. It can be understood that according to different tasks, the block semaphore after the operation can point to different sub-blocks in the DMA buffer, so that the tasks can be more evenly hashed on multiple sub-blocks.

[0057] Step 330: Obtain the head value stored in the head address field included in the first sub-block and the base address of the first sub-block from the defined global variables.

[0058] Specifically, according to the description of the foregoing step 310, the CPU obtains the head value stored in the head address field (for example, head addr_1) included in the first sub-block and the base address of the first sub-block (for example, dmaaddr_1) from the defined global variables.

[0059] Step 340: Obtain the tail value stored in the tail address field included in the first sub-block from the DMA buffer.

[0060] Specifically, the CPU accesses the tail address field (for example, tail addr_1) included in the first sub-block and obtains the tail value therefrom.

[0061] Step 350: When the head value of the first sub-block is equal to the tail value, update the head value according to the length value of the data to be written.

[0062] Specifically, the CPU determines whether the head value is equal to the tail value. If the head value is equal to the tail value, the CPU updates the head value according to the length value of the data to be written.

[0063] Among them, the size of the increase in the head value is equal to the size of the data transmitted this time and serves as the offset relative to the base address when the next PCIe host transmits data to the DMA buffer. The value of the head can also be a signal to notify the PCIe board device, that is, the PCIe host has data to transmit.

[0064] If the head value is not equal to the tail value, the CPU repeats step 340. In this way, it can be ensured that the data involved will not be overwritten by the current data when the previous operation is not completed.

[0065] Step 360: According to the base address and the updated head value, write the data to be written into the sub-buffer field included in the first sub-block through the PCIe bus.

[0066] Specifically, after the CPU updates the head value, using the base address of the first sub-block and the updated head value, through the PCIe bus, starting from the base address offset by the head value as the starting address, write the data to be written into the sub-buffer field included in the first sub-block.

[0067] Therefore, by applying a data transmission method provided in the present application, the PCIe host and the PCIe board device are connected through the PCIe bus. The CPU obtains a chunk semaphore for indicating multiple sub-blocks in the dma buffer included in the PCIe board device. According to the chunk semaphore, the CPU determines the first sub-block from the multiple sub-blocks included in the dma buffer. From the defined global variables, the CPU obtains the head value and the base address stored in the head address field included in the first sub-block. From the dma buffer, the CPU obtains the tail value stored in the tail address field included in the first sub-block. When the head value of the first sub-block is equal to the tail value, according to the length value of the data to be written, the CPU updates the head value. According to the base address and the updated head value, through the PCIe bus, the CPU writes the data to be written into the sub-buffer field included in the first sub-block.

[0068] In the foregoing manner, each operation only requires the PCIe host to update the data, and the program on the PCIe board device actively reads the data in the dma buffer, reducing the scheduling overhead of both CPUs; on the premise of sharing a physical dma buffer space, the dma buffer is logically divided into N blocks and combined with multi-threading technology to operate different positions of the dma buffer so that multiple types of data or read / write operations can be performed simultaneously without affecting each other, improving the data transmission efficiency by N times; solving the disadvantages in the prior art such as inflexible configuration of the dma buffer space, low space utilization rate, and less-than-ideal data exchange and transmission rate.

[0069] Optionally, after the foregoing step 360, the following process is further included:

[0070] First, the CPU stores the updated head value in the global variable into the head address field included in the first sub-block. Then, the CPU periodically retrieves the response value stored in the response address field (e.g., response addr_1) included in the first sub-block. The CPU determines whether the response value is set to 1. When the response value is 1, the CPU determines that the current data transfer has been completed. When the response value is not 1, the CPU retrieves the response value stored in the response address field of the first sub-block again.

[0071] The CPU assigns a lock to each of the multiple sub-blocks, and this lock is used to ensure that data will not be overwritten when multiple tasks on the PCIe host operate on multiple sub-blocks simultaneously.

[0072] When the PCIe host obtains data from the PCIe board device, the PCIe board device writes data from the base address of the dma buffer block, and then the CPU reads data from the base address of the dma buffer block.

[0073] Optionally, since PCI Multithread involves multiple threads operating on related functions in parallel, such as reading counters, operating on table entries, etc. Therefore, it is necessary to ensure data consistency and avoid the situation where the same piece of data is operated on by multiple threads simultaneously at the same moment. On this basis, it is necessary to add an operation of a mutex lock to the module using PCI Multithread.

[0074] The following will be illustrated by taking the transaction of operating on table entries using PCI Multithread as an example:

[0075] (1) When using PCI Multithread to operate on table entries, n threads can operate on different table entries simultaneously. However, in order to avoid the situation where multiple threads operate on the same table entry simultaneously, which may damage data consistency, it is necessary to add a mutex lock to the transaction of operating on table entries: that is, lock the entire table before operating on the table entry.

[0076] (2) As described in (1), although conflicts are avoided, parallel operations cannot be performed when operating on different table entries in the same table, which fails to achieve the purpose of performance optimization. Therefore, it is necessary to use a mutex lock with a smaller granularity, a mutex lock for one table entry.

[0077] (3) A unique semaphore needs to be specified when locking. When locking the entire table, the unique value, the Struct ID of the table, can be used as the semaphore; but when locking one table entry, different table entries need to specify different semaphores, and this semaphore needs to be unique for each table entry.

[0078] Based on the above problems, for different types of tables, the calculation of semaphores is also different. For a direct table (Table), its search key (Key) is unique for each table entry. The value of the Key can be used for n-thread hashing as the semaphore of the mutex lock. For a hash table, its first hash value uniquely determines its position in the main table. Therefore, the first hash value can be used for hashing.

[0079] After calculating the above hashing, locking with the hash value can not only avoid data inconsistency problems caused by a table entry being operated on by multiple threads simultaneously, but also avoid conflicts in the dma channel.

[0080] The data transmission method provided in the embodiments of the present application will be described in detail below. Refer to Figure 5 , Figure 5 which is a flowchart of another data transmission method provided in the embodiments of the present application. This method is applied to the ctop included in the PCIe board device. The PCIe board device also includes a cpu core and a dma buffer. The dma buffer includes multiple sub-blocks. The PCIe board device is connected to the PCIe host through the PCIe bus, and specifically includes the following steps.

[0081] Step 510: Allocate a sub-block included in the dma buffer to the thread running in the cpu core.

[0082] Specifically, according to the description of the foregoing embodiments, the PCIe host performs relevant configurations during the process of implementing data transmission. Then, the PCIe board device also performs adaptive configurations so that it can implement data transmission together with the PCIe host.

[0083] Further, according to the first number of cpu cores included in the PCIe board device (for example, 4 cpus, taking cpu16 - 19 as an example), the ctop replicates the second number of threads equal to the first number of cpu cores (for example, forks 4 threads), and binds the PCI multithread program of each replicated thread to a cpu core, that is, each thread exclusively uses one cpu.

[0084] Ctop reads the value of the status register in a loop and determines whether the value of the status register is set to 1. When the status register is set to 1, ctop allocates a section of dam buffer for each thread, and the size is a sub-block of the available dmabuffer configured by the PCIe host. Ctop obtains the base address of the dma buffer from the address register. For each thread, based on the base address and the identifier of the cpu core, ctop determines the corresponding sub-block allocated for each thread. Ctop performs initialization processing on each sub-block and increments the count register by 1. After all sub-blocks are initialized, ctop clears the count register and sets the status register to 0, where the value of the status register is used to indicate that the initialization of the PCIe board device is completed.

[0085] Further, the process of ctop determining the corresponding sub-block allocated for each thread according to the base address and the identifier of the cpu core is specifically as follows: For each thread, ctop performs an offset operation of n*4k on the base address according to the identifier of the CPU core. Ctop determines the sub-block allocated for the thread according to the base address after the offset operation and the sub-block size configured by the PCIe host, where each sub-block includes a sub-buffer field, a head address field, a tail address field, an error code address field, and a response address field.

[0086] Ctop performs an offset operation of n*K (where n≤N, N is the number of sub-blocks, n is an integer; K is the quotient of the size of the dma buffer and the number of sub-blocks) on the base address to obtain the base address of each sub-block. In the embodiment of the present application, K = 4k. Then, ctop reads the size of a sub-block of the available dma buffer configured by the PCIe host from the address register. In this way, ctop determines the sub-block allocated for each thread.

[0087] Further, the process of ctop performing initialization processing on each sub-block is specifically as follows: Ctop obtains the values of each address field included in the corresponding sub-block from the address register (for example, the values of each address field in the dma_addr field, head_addr field, tail_addr field, error_code_addr field, and response_addr field. The values of each field are configured by the PCIe host). Ctop stores the values of each address field included in the obtained sub-block into the corresponding address fields included in the sub-block allocated for the thread.

[0088] Step 520, for each of the sub-blocks, obtain the head value stored in the head address field included in the sub-block and the tail value stored in the tail address field included in the sub-block.

[0089] Specifically, for each sub-block, taking the first sub-block as an example, ctop obtains the head value of the head_addr_1 field included in the first sub-block and the tail value of the tail_addr_1 field.

[0090] Step 530: When the head value is not equal to the tail value, obtain the data transmitted by the PCIe host through the PCIe bus from the sub-buffer field included in the sub-block, and the size of the data is the difference between the head value and the tail value.

[0091] Specifically, ctop determines whether the head value is equal to the tail value. If the head value is not equal to the tail value, ctop determines that the PCIe host has written the data of the first sub-block. Ctop obtains the data transmitted by the PCIe host through the PCIe bus from the sub-buffer field included in the first sub-block. Among them, the size of the data is the difference between the head value and the tail value.

[0092] As Figure 6 shown, Figure 6 is a block diagram of the data transmission process provided by the embodiment of the present application through PCIe. Among them, the dma buffer is divided into multiple sub-blocks, and each sub-block is used by a thread running in a cpu core.

[0093] Further, ctop updates the base address of the first sub-block according to the obtained data length value. Among them, the base address of the first sub-block is the sum of the tail value of the first sub-block and the data length value, and is used as the starting address for the next PCIe board device to read data in the dma buffer. Then, ctop sets the reply value stored in the response_addr_1 field included in the first sub-block to 1, and this reply value is used to enable the PCIe host to determine that the current data transmission is completed.

[0094] Therefore, by applying a data transmission method provided by the present application, the PCIe host and the PCIe board device are connected through the PCIe bus. The PCIe board device further includes a cpu core and a dma buffer, and the dma buffer includes multiple sub-blocks. Ctop allocates a sub-block included in the dma buffer to a thread running in the cpu core. For each sub-block, ctop obtains the head value stored in the head address field included in the sub-block and the tail value stored in the tail address field included in the sub-block. When the head value is not equal to the tail value, ctop obtains the data transmitted by the PCIe host through the PCIe bus from the sub-buffer field included in the sub-block, where the size of the data is the difference between the head value and the tail value.

[0095] In the foregoing manner, each operation only requires the PCIe host to update data. The program on the PCIe board device actively reads the data in the dma buffer, reducing the scheduling overhead of the CPUs on both sides. On the premise of sharing a physical dma buffer space, the dma buffer is logically divided into N blocks and the multi-threading technology is combined to operate different positions of the dma buffer so that multiple types of data or read / write operations can be carried out simultaneously without affecting each other, improving the data transmission efficiency by N times. It solves the disadvantages in the prior art such as the inflexible configuration of the dma buffer space, low space utilization rate, and less-than-ideal data exchange and transmission rate.

[0096] Optionally, when the PCIe host obtains data from the PCIe board device, ctop writes data from the base address of the dma buffer, and the CPU reads data from the base address of the dma buffer block.

[0097] Based on the same inventive concept, an embodiment of the present application further provides a data transmission device corresponding to the data transmission method described above. Refer to Figure 3 FIG. Figure 7 , Figure 7 which is a structural diagram of a data transmission device provided by an embodiment of the present application. The device is applied to the CPU included in the PCIe host. The PCIe host is connected to the PCIe board device through a PCIe bus. The device includes:

[0098] An obtaining unit 710, configured to obtain a block semaphore, where the block semaphore is used to indicate a sub-block in a dma buffer included in the PCIe board device, and the dma buffer includes multiple sub-blocks;

[0099] A determining unit 720, configured to determine a first sub-block from multiple sub-blocks included in the dma buffer according to the block semaphore;

[0100] The obtaining unit 710 is further configured to obtain a head value stored in a head address field included in the first sub-block and a base address of the first sub-block from a defined global variable;

[0101] The obtaining unit 710 is further configured to obtain a tail value stored in a tail address field included in the first sub-block from the DMA Buffer;

[0102] An updating unit 730, configured to update the head value according to a length value of data to be written when the head value of the first sub-block is equal to the tail value;

[0103] A write unit 740, configured to write the data to be written into a sub-buffer field included in the first sub-block through the PCIe bus according to the base address and the updated head value.

[0104] Optionally, the device further includes: a storage unit (not shown in the figure), configured to store the updated head value into a head address field included in the first sub-block;

[0105] The obtaining unit 710 is further configured to obtain a reply value stored in a reply address field included in the first sub-block;

[0106] A determination unit (not shown in the figure), configured to determine that the current data transmission is completed when the reply value is 1.

[0107] Optionally, the device further includes: a configuration unit (not shown in the figure), configured to allocate a base address for the dma buffer and configure the size of the dma buffer;

[0108] The write unit 740 is further configured to write the allocated base address and the configured size of the dma buffer into an address register;

[0109] The configuration unit (not shown in the figure) is further configured to set a status register to 1, and the value of the status register is used to indicate that the PCIe host initialization is completed.

[0110] Optionally, the configuration unit (not shown in the figure) is further configured to allocate a lock for each of the multiple sub-blocks.

[0111] Optionally, the device further includes: a read unit (not shown in the figure), configured to read data from the base address of the dma buffer block when the PCIe host obtains data from the PCIe board device.

[0112] Therefore, by applying a data transmission device provided in the present application, a PCIe host and a PCIe board device are connected through a PCIe bus. The device obtains a chunk semaphore for indicating multiple sub-chunks in a dma buffer included in the PCIe board device. According to the chunk semaphore, the device determines a first sub-chunk from the multiple sub-chunks included in the dma buffer. From the defined global variables, the device obtains a head value and a base address stored in a head address field included in the first sub-chunk. From the dma buffer, the device obtains a tail value stored in a tail address field included in the first sub-chunk. When the head value of the first sub-chunk is equal to the tail value, according to the length value of the data to be written, the device updates the head value. According to the base address and the updated head value, through the PCIe bus, the device writes the data to be written into a sub-buffer field included in the first sub-chunk.

[0113] In the foregoing manner, each operation only requires the PCIe host to update the data, and the program on the PCIe board device actively reads the data in the dma buffer, reducing the scheduling overhead of both CPUs; on the premise of sharing a physical dma buffer space, the dma buffer is logically divided into N chunks and the multi-threading technology is combined to operate different positions of the dma buffer so that multiple types of data or read / write operations can be performed simultaneously without affecting each other, improving the data transmission efficiency by N times; solving the disadvantages in the prior art that the dma buffer space configuration is not flexible enough, the space utilization rate is low, and the data exchange and transmission rate is not ideal enough.

[0114] Based on the same inventive concept, an embodiment of the present application further provides a data transmission device corresponding to the data transmission method described above. Refer to Figure 5 FIG. Figure 8 , Figure 8 FIG. is a structural diagram of another data transmission device provided in an embodiment of the present application. The device is applied to a ctop included in a PCIe board device. The PCIe board device further includes a CPU core and a dma buffer. Among them, the dma buffer includes multiple sub-chunks. The PCIe board device is connected to a PCIe host through a PCIe bus. The device includes:

[0115] An allocation unit 810, configured to allocate a sub-chunk included in the dma buffer to a thread running in the CPU core;

[0116] An acquisition unit 820, configured to, for each sub-chunk, acquire a head value stored in a head address field included in the sub-chunk and a tail value stored in a tail address field included in the sub-chunk;

[0117] The obtaining unit 820 is further configured to, when the head value is not equal to the tail value, obtain the data transmitted by the PCIe host through the PCIe bus from the sub-buffer field included in the sub-block, where the size of the data is the difference between the head value and the tail value.

[0118] Optionally, the device further includes: a replication unit (not shown in the figure), configured to replicate a second number of threads equal to the first number of CPU cores according to the first number of CPU cores included in the PCIe board device, and bind each replicated thread to one of the CPU cores;

[0119] The obtaining unit 820 is further configured to, when the status register is set to 1, obtain the base address of the dma buffer from the address register;

[0120] A determination unit (not shown in the figure), configured to, for each thread, determine the corresponding sub-block allocated to each thread according to the base address and the identifier of the CPU core;

[0121] An initialization unit (not shown in the figure), configured to perform an initialization process on each sub-block and increment the count register by 1;

[0122] A configuration unit (not shown in the figure), configured to, when all sub-blocks are initialized, clear the count register and set the status register to 0, where the value of the status register is used to indicate that the initialization of the PCIe board device is completed.

[0123] Optionally, the determination unit (not shown in the figure) is specifically configured to, for each thread, perform an offset operation of n*K on the base address according to the identifier of the CPU core;

[0124] Determine the sub-block allocated to the thread according to the base address after the offset operation and the sub-block size configured by the PCIe host, where the sub-block includes a sub-buffer field, a head address field, a tail address field, an error code address field, and a reply address field;

[0125] where n≤N, N is the number of sub-blocks, n is an integer; K is the quotient of the size of the dma buffer and the number of sub-blocks.

[0126] Optionally, the initialization unit (not shown in the figure) is specifically configured to obtain the values of the respective address fields included in the corresponding sub-block from the address register;

[0127] Correspondingly store the values of the respective address fields included in the obtained sub-block into the respective address fields included in the sub-block allocated to the thread.

[0128] Optionally, the device further includes: an update unit (not shown in the figure), configured to update the base address of the sub-block, where the base address of the sub-block is the sum of the tail value of the sub-block and the length value of the data;

[0129] A storage unit (not shown in the figure), configured to set the reply value stored in the reply address field included in the sub-block to 1, where the reply value is used to enable the PCIe host to determine the completion of the current data transmission.

[0130] Optionally, the device further includes: a write unit (not shown in the figure), configured to write data from the base address of the dma buffer when the PCIe host obtains data from the PCIe board device.

[0131] Therefore, by applying a data transmission device provided in the present application, the PCIe host and the PCIe board device are connected through a PCIe bus. The PCIe board device further includes a cpu core and a dma buffer, and the dma buffer includes multiple sub-blocks. The device allocates a sub-block included in the dma buffer to a thread running in the cpu core. For each sub-block, the device obtains the head value stored in the head address field included in the sub-block and the tail value stored in the tail address field included in the sub-block. When the head value is not equal to the tail value, the device obtains the data transmitted by the PCIe host through the PCIe bus from the sub-buffer field included in the sub-block, where the size of the data is the difference between the head value and the tail value.

[0132] In the foregoing manner, each operation only requires the PCIe host to update the data, and the program on the PCIe board device actively reads the data in the dma buffer, reducing the scheduling overhead of both cpus; on the premise of sharing a physical dma buffer space, the dma buffer is logically divided into N blocks and the multi-thread technology is combined to operate different positions of the dma buffer, so that multiple types of data or read / write operations can be performed simultaneously without affecting each other, improving the data transmission efficiency by N times; solving the disadvantages in the prior art that the dma buffer space configuration is not flexible enough, the space utilization rate is low, and the data exchange and transmission rate is not ideal.

[0133] The implementation processes of the functions and roles of each unit in the above device are specifically described in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.

[0134] For the apparatus embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions in the method embodiments. The apparatus embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0135] For the embodiments of the data transmission apparatus, since the method content involved is basically similar to the foregoing method embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the method embodiments.

[0136] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the scope of protection of this application.

Claims

1. A data transmission method, characterized in that, The method is applied to the CPU included in a PCIe host, and the PCIe host is connected to a PCIe board device via a PCIe bus. The method includes: Obtain a chunk semaphore, where the chunk semaphore is used to indicate a sub - chunk in a dma buffer included in the PCIe board device; wherein, the PCIe host logically chunks the dma buffer so that the dma buffer includes multiple sub - chunks; Determine a first sub - chunk from the multiple sub - chunks included in the dma buffer according to the chunk semaphore; Obtain the head value stored in the head address field included in the first sub - chunk and the base address of the first sub - chunk from defined global variables; Obtain the tail value stored in the tail address field included in the first sub - chunk from the dma buffer; When the head value of the first sub - chunk is equal to the tail value, update the head value according to the length value of the data to be written; Write the data to be written into the sub - buffer field included in the first sub - chunk via the PCIe bus according to the base address and the updated head value; 2. The method according to claim 1, wherein The method further includes: Store the updated head value into the head address field included in the first sub - chunk; Obtain the reply value stored in the reply address field included in the first sub - chunk; When the reply value is 1, determine that the current data transmission is completed; 3. The method according to claim 1, wherein Before obtaining the chunk semaphore, the method further includes: Allocate a base address for the dma buffer and configure the size of the dma buffer; Write the allocated base address and the configured size of the dma buffer into an address register; Set the status register to 1, where the value of the status register is used to indicate that the initialization of the PCIe host is completed; 4. The method according to claim 1, wherein The method further includes: Allocate a lock for each of the multiple sub - chunks; 5. The method according to claim 1, wherein The method further includes: When the PCIe host obtains data from the PCIe board device, read the data from the base address of the dma buffer block; 6. A data transmission method, characterized in that, The method is applied to the ctop included in a PCIe board device, and the PCIe board device further includes a CPU core and a dma buffer. The dma buffer includes multiple sub - chunks obtained by the PCIe host logically chunking the dma buffer, The PCIe board device is connected to the PCIe host via a PCIe bus. The method includes: Allocate a sub - chunk included in the dma buffer for a thread running in the CPU core; For each sub - chunk, obtain the head value stored in the head address field included in the sub - chunk and the tail value stored in the tail address field included in the sub - chunk; When the head value is not equal to the tail value, obtain the data transmitted by the PCIe host through the PCIe bus from within the sub-buffer field included in the sub-block, where the size of the data is the difference between the head value and the tail value.

7. The method according to claim 6, wherein Before allocating a sub-block included in the dma buffer to a thread running in the cpu core, the method further includes: According to the first number of cpu cores included in the PCIe board device, replicate a second number of threads equal to the first number of cpu cores, and bind each replicated thread to one of the cpu cores; When the status register is set to 1, obtain the base address of the dma buffer from the address register; For each thread, determine the corresponding sub-block allocated to each thread according to the base address and the identifier of the cpu core; Perform initialization processing on each sub-block and increment the count register by 1; When all sub-blocks are initialized, clear the count register and set the status register to 0, where the value of the status register is used to indicate that the initialization of the PCIe board device is complete.

8. The method according to claim 7, wherein The determining the corresponding sub-block allocated to each thread according to the identifier of the cpu core specifically includes: For each thread, perform an offset operation of n*K on the base address according to the identifier of the cpu core; According to the base address after the offset operation and the sub-block size configured by the PCIe host, determine the sub-block allocated to the thread, where the sub-block includes a sub-buffer field, a head address field, a tail address field, an error code address field, and a reply address field; Where n≤N, N is the number of sub-blocks, n is an integer; K is the quotient of the size of the dma buffer and the number of sub-blocks.

9. The method according to claim 8, wherein The performing initialization processing on each sub-block specifically includes: Obtain the values of the respective address fields included in the corresponding sub-block from the address register; Store the values of the respective address fields included in the obtained sub-block into the respective address fields included in the sub-block allocated to the thread.

10. The method according to claim 6, wherein After obtaining the data transmitted by the PCIe host through the PCIe bus from within the sub-buffer field included in the sub-block, the method further includes: Update the base address of the sub-block, where the base address of the sub-block is the sum of the tail value of the sub-block and the length value of the data; Set the reply value stored in the reply address field included in the sub-block to 1, where the reply value is used to enable the PCIe host to determine that the current data transmission is complete.

11. The method according to claim 6, wherein The method further includes: When the PCIe host obtains data from the PCIe board device, write the data from the base address of the dma buffer.

12. A data transmission device, characterized in that, The apparatus is applied to the cpu included in the PCIe host, the PCIe host is connected to the PCIe board device through the PCIe bus, and the apparatus includes: An acquisition unit for acquiring a chunk semaphore, where the chunk semaphore is used to indicate a sub - chunk in a dma buffer included in the PCIe board device; wherein, the PCIe host logically chunks the dma buffer so that the dma buffer includes multiple sub - chunks; A determination unit for determining a first sub - chunk from the multiple sub - chunks included in the dma buffer according to the chunk semaphore; The acquisition unit is further configured to acquire, from a defined global variable, a head value stored in a head address field included in the first sub - chunk and a base address of the first sub - chunk; The acquisition unit is further configured to acquire, from the dma buffer, a tail value stored in a tail address field included in the first sub - chunk; An update unit for updating the head value according to the length value of the data to be written when the head value of the first sub - chunk is equal to the tail value; A write unit for writing the data to be written into a sub - buffer field included in the first sub - chunk through the PCIe bus according to the base address and the updated head value; 13. A data transmission device, characterized in that, The device is applied to a ctop included in a PCIe board device, the PCIe board device further includes a cpu core and a dma buffer, the dma buffer includes multiple sub - chunks obtained by logically chunking the dma buffer by a PCIe host, the PCIe board device is connected to the PCIe host through a PCIe bus, and the device includes: An allocation unit for allocating a sub - chunk included in the dma buffer to a thread running in the cpu core; An acquisition unit for, for each sub - chunk, acquiring a head value stored in a head address field included in the sub - chunk and a tail value stored in a tail address field included in the sub - chunk; The acquisition unit is further configured to, when the head value is not equal to the tail value, acquire data transmitted by the PCIe host through the PCIe bus from a sub - buffer field included in the sub - chunk, and the size of the data is the difference between the head value and the tail value.

Citation Information

Patent Citations

  • Multiple core processor device with multithreading

    CN107980118A

  • Method for initiatively realizing data exchange with CPU by peripheral

    CN108388529A

  • DMA controller based on PCIE protocol and DMA data transmission method

    CN110046114A