A PCIe-P2P communication device and method
Through the unified management module, the memory resources and DMA engine of the PCIe network are planned, and the read operation is converted into write operation is solved, which solves the problem of too many read operations in the PCIe network and improves network efficiency and bandwidth utilization.
Patent Information
- Application Number
- CN202510352026.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-25
AI Technical Summary
There are too many read operation requests in the existing PCIe switched network, resulting in inefficiency of the network, especially when accessing big data, and inappropriate use of DMA leads to frequent read requests.
Through the unified management module, the memory resources and DMA engine of the PCIe network are uniformly planned, and the read operations of the user program are converted into write operations, and the DMA engine and buffer closest to the source are preferred, reducing read operations and improving bandwidth utilization.
It significantly improves the available bandwidth and network bandwidth utilization of the PCIe network, reduces the frequency of read operations, and improves P2P communication performance.
Smart Images

Figure CN119862143B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of data communication and integrated circuit data processing, and particularly to a PCIe-P2P communication device and method based on unified resource planning. Background Art
[0002] The communication mode of the PCIe bus is direct shared memory read and write. In a PCIe switching system, it is necessary to pre-set the space mapping relationship, that is, each EP of the PCIe maps a section of its own memory to the PCIe address space. At the same time, it is also necessary to map the PCIe address space to its own memory space. After setting the space mapping relationship, when EP1 writes a message to the mapped memory space of EP2, the message can be sent to the memory space of EP2.
[0003] For PCIe-based P2P communication, the sender needs to know the address of the operable memory space of the peer, and a mechanism to notify the peer EP that the operation has been completed is required, so that the peer EP and the local EP can perform the next processing.
[0004] DMA, the full name is Direct Memory Access, that is, direct memory access. DMA transfer copies data from one address space to another address space, providing high-speed data transfer between peripherals and memories or between memories.
[0005] Due to the design reasons of the PCIe specification, the efficiency of PCIe write operations is much greater than that of read operations. Therefore, how to reduce the read operations in software operations is the key factor to improve the efficiency of PCIe.
[0006] In the existing PCIe switching network, the following problems exist:
[0007] (1) The read and write operations of PCIe coexist, especially in a network with large data access, this phenomenon is more obvious; due to efficiency issues, the network is filled with a large number of READ request operations, that is, read operation requests;
[0008] (2) In a PCIe network accelerated by using DMA, due to improper use of conventional DMA, there are often many PCIe read request operations. Summary of the Invention
[0009] In view of this, to solve the above problems existing in the prior art, embodiments of the present invention provide a PCIe-P2P communication device and method. Specifically, the following technical solutions are provided:
[0010] On the one hand, the present invention provides a PCIe-P2P communication device, and the communication device includes:
[0011] A user layer module, a unified management module, and a driver layer module; the communication device is connected to multiple EP devices through a PCIe network;
[0012] The user layer module is used to apply for a communication path to the unified management module;
[0013] The unified management module is used to collect information of PCIe devices, and select and process P2P channels; the unified management module uniformly manages the allocation of user data buffer resources and the allocation of DMA engines; the unified management module plans the read operation of the user layer module as a write operation on the destination data side EP1 to the read data side EP2; the user data buffer includes a send buffer allocated for EP1 and a receive buffer allocated for EP2;
[0014] The driver layer module receives the communication operation descriptor or register operation control data sent by the unified management module, and is used to control DMA data transfer and operate multiple EP devices.
[0015] Preferably, when performing P2P communication, the destination data side EP1 executes the following process:
[0016] S11. EP1 applies to the unified management module to establish an EP1-EP2 path;
[0017] S12. The unified management module allocates the best DMA engine, and allocates user data buffers for EP1 and EP2 to form a communication operation descriptor;
[0018] S13. The unified management module sends the communication operation descriptor to EP1;
[0019] S14. After receiving the communication operation descriptor, EP1 starts a DMA operation, starts the allocated DMA engine, and constructs a DMA descriptor;
[0020] S15. Execute data transfer through the allocated DMA engine, and transfer the data in the send buffer to the receive buffer; the data in the send buffer is the data from EP1 that EP2 needs to read.
[0021] Preferably, when performing P2P communication, the read data side EP2 executes the following process:
[0022] S21. Apply to the unified management module for a path between EP1 and itself;
[0023] S22. The unified management module checks whether a path has been allocated. If a path has been allocated, it sends the result to EP2. If no path has been allocated, it allocates a path and sends the allocation result to EP2;
[0024] S23. The unified management module sends a startable communication notice to EP1 and EP2;
[0025] S24. After receiving the start communication notice, EP2 establishes a polling thread and prepares to receive data;
[0026] S25. When EP1 finishes sending data, EP2 polls all the data and completes the data reading operation.
[0027] Preferably, in S12, the specific method of allocating user data buffers for EP1 and EP2 is as follows:
[0028] The unified management module records the memory information in the collected PCIe device information as M{Addr1{start, end, bdf1}, …, Addrp{start, end, bdfp}}, where p represents the total number of EP devices in the PCIe network, bdf refers to the BDF address of the EP device in PCIe, Addrp represents the memory on EPp, start represents the start address of the corresponding memory, and end represents the end address of the corresponding memory;
[0029] Find all free memory in M to form a free memory set, and respectively select the free memory closest to EP1 and the free memory closest to EP2 in the free memory set;
[0030] Allocate a send buffer H1 in the free memory closest to EP1 and allocate a receive buffer H2 in the free memory closest to EP2;
[0031] If there is no free memory, return allocation failure.
[0032] Preferably, the method for judging the closest distance is: judge the distance based on the number of links spanned between the free memory and the target node, and the smaller the number of links, the closer the distance; the target nodes are EP1 and EP2.
[0033] Preferably, in S11, the data structure of the path is: {DMA engine, send buffer, receive buffer, data length}, and the data length is the length of the data to be transmitted between EP1 and EP2.
[0034] Preferably, in 12, the priority for allocating the best DMA engine is:
[0035] Priority one: preferentially select the DMA engine of EP1;
[0036] Priority two: if priority one does not exist or is unavailable, then select the DMA engine on the switching device;
[0037] Priority Three: If neither Priority One nor Priority Two exists or is unavailable, select the DMA engine on an EP device that is not EP1;
[0038] If neither Priority One, Priority Two, nor Priority Three exists, return a DMA engine allocation failure signal.
[0039] Preferably, in S12, the specific method for allocating the optimal DMA engine is as follows:
[0040] S121. The unified management module collects PCIe networking topology information and records it as , where k represents the cost between two nodes in the PCIe network, m and n respectively represent the node numbers, and n represents the total number of PCIe network nodes;
[0041] S122. Aggregate to obtain a DMA set, denoted as D{DMA1,..., DMAc}, where c represents the number of DMAs, and c is less than n;
[0042] S123. In the set T, filter out the cost set C of the DMAs in the set D that are associated with the destination data party, denoted as C{C1,..., CI}, where I is a positive integer, and its maximum value is 2c; if the set C is empty, return a DMA engine allocation failure signal;
[0043] S124. Filter out the node with the minimum cost in the set C and denote it as Cmin, and set the DMA on the node with the number min as the optimal DMA engine.
[0044] Preferably, the communication operation descriptor includes whether it is a sender, the allocated DMA engine, the send buffer, and the receive buffer.
[0045] Preferably, after returning the DMA engine allocation failure signal, perform data transmission between the destination data party EP1 and the data reading party EP2 according to the PCIe read communication process without DMA control.
[0046] On the other hand, the present invention also provides a PCIe-P2P communication method, which is applied to the communication device as described above. The method includes:
[0047] Step 1. The user side applies to the unified management module for a communication path to read data from the data reading party EP2 to the destination data party EP1;
[0048] Step 2. The unified management module allocates a DMA engine based on the application of the user side.
[0049] Step 3: The unified management module allocates resources for the user data buffer based on the application from the user side. The user data buffer includes a transmission buffer allocated for EP1 and a reception buffer allocated for EP2.
[0050] Step 4: The unified management module establishes an EP1 - EP2 path.
[0051] Step 5: After the path is established, EP1 starts a DMA operation, and transfers the data to be read from the transmission buffer of the destination data party to the reception buffer of the data reading party through the allocated DMA engine.
[0052] Step 6: After the data transfer is completed, it ends.
[0053] Preferably, if Step 2 fails, the data transfer between EP1 of the destination data party and EP2 of the data reading party is performed according to the PCIe read communication process without DMA control.
[0054] Compared with the prior art, the technical solution of the present invention converts the communication requirements between user programs into Write operations after planning channels through the unified management layer, minimizing the READ operations in the PCIe network and improving the bandwidth utilization rate of the entire network. Therefore, compared with the PCIe P2P communication in the prior art, this solution has the following advantages:
[0055] (1) In the user program operation, the READ operation is converted into a Write transaction operation of PCIe DMA, minimizing the READ operations in the PCIe network and greatly increasing the available bandwidth of the PCIe network.
[0056] (2) The unified management layer makes a unified and reasonable plan for the memory resources of the entire PCIe network, reducing memory allocation conflicts, and trying to select the allocation method closest to the source, which can greatly save the path length of PCIe packets propagating in the PCIe network, thereby improving the bandwidth utilization rate of the entire PCIe network.
[0057] This solution is particularly applicable to some environments based on PCIe bus networking, where there is P2P communication processing based on the DMA engine, etc. It can greatly reduce the READ transactions in the network, thereby improving the available bandwidth of the entire network. Based on the memory planning of the entire network, the memory buffer resources are reasonably planned. Through these two measures, the P2P transmission performance of the PCIe network will be greatly improved. Description of the Drawings
[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0059] Figure 1 Schematic diagram of the system architecture of the embodiment of the present invention;
[0060] Figure 2 Schematic diagram of the user P2P communication data processing flow of the embodiment of the present invention;
[0061] Figure 3 Schematic diagram of the execution method flow of the embodiment of the present invention. Specific embodiments
[0062] The following will describe the embodiments of the present invention in detail with reference to the drawings. It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0063] Those skilled in the art should be aware that the following specific embodiments or specific implementation manners are a series of optimized setting manners listed by the present invention to further explain the specific invention content, and these setting manners can be combined with each other or used in association with each other, unless the present invention clearly states that some or a specific embodiment or implementation manner cannot be associated or used together with other embodiments or implementation manners. At the same time, the following specific embodiments or implementation manners are only used as the most optimized setting manners, rather than as an understanding of limiting the protection scope of the present invention.
[0064] Embodiment 1:
[0065] To solve the above-mentioned disadvantages existing in the existing PCIe switched network P2P communication, this solution proposes a PCIe-P2P communication method and device with high execution speed based on unified resource planning. This solution centrally plans the DMA and cache used for P2P communication through a centralized control module. For P2P communication, the following principles are followed in this embodiment:
[0066] (1) Allocate buff (i.e., user data buffer) resources and select the DMA engine through a unified management module; in this embodiment, the user data buffer includes a send buffer and a receive buffer, which are respectively allocated to the destination data party and the data reading party;
[0067] (2)For the user's read operation, it is planned to perform a write operation at the destination data side to the data side that needs to be read; in this way, all the user's read operations can generate PCIe write transactions.
[0068] (3)For the case where there is a DMA operation, when there are DMA engines on both EPs (i.e., PCIe endpoint devices that support P2P), following the principle of being closest to the source, select the DMA engine on the EP where the data source is located. In this way, it can be ensured that all DMA operations are DMA write operations, thus avoiding read operations and saving data operation resources.
[0069] In this embodiment, the system framework of this solution is as Figure 1 shown. The top layer is the user program layer, that is, the user space. Next are the unified management layer, the driver layer, and the hardware device layer, where:
[0070] (1)The user program is mainly responsible for applying for a communication path to the unified management module. Here, it can be represented by the user layer module. The user program is also each user end and is included in the user layer module.
[0071] (2)The unified management module is mainly responsible for collecting PCIe device information, including topology information, Bar address (i.e., the range of operable memory address space), and DMA information, etc. At the same time, this module is also responsible for unified memory application and release, DMA engine selection, and P2P channel selection processing. The Bar addresses of the main memory and the peripheral EP collected during initialization are unified into a memory pool, and subsequent allocation and release manage and allocate resources from this memory pool. The DMA engine selection follows the principle of being closest to the data source. After determining the DMA engine, select an idle communication channel or queue in the selected DMA engine as the P2P channel.
[0072] (3)The driver layer is mainly responsible for operating the DMA to transfer data, receiving the processing operations of the unified management module above and operating the specific hardware device below.
[0073] (4)The hardware device layer mainly includes various PCIe EP devices.
[0074] Next, continue to combine Figure 2 , and illustrate the P2P communication process of this solution. In this embodiment, take the example that the EP2 user program wants to access the data of EP1 (that is, the EP2 user program wants to read the data of EP1):
[0075] First, in this embodiment, the unified management module can set the selection of the DMA engine according to the priority order:
[0076] (1)First, select the DMA engine on the data source (i.e., the destination data side) EP1;
[0077] (2) If (1) does not exist or is unavailable, then select the DMA on the PCIe Switch; here, preferably, the selection of the DMA engine can be based on the lowest cost or the nearest distance method;
[0078] (3) If both (1) and (2) are unavailable or do not exist, then select the DMA on other EP devices that are not the data source (i.e., the destination data party); at this time, if there are multiple EP devices, the screening can also follow the lowest cost or the nearest distance method;
[0079] (4) If it still does not exist, then it fails, that is, there is no recommended DMA engine available, then at this time, execute according to the PCIe read communication process controlled by non-DMA. The PCIe read communication process controlled by non-DMA belongs to the conventional data reading process in the art and will not be elaborated here.
[0080] Second, in this embodiment, the data memory buffer allocation method of the unified management module is set to the method of finding the nearest data memory cache, and the judgment of the nearest distance is determined by the number of links spanned between the cache and the target node. The smaller the number, the closer the distance. The specific setting process is as follows:
[0081] (1) First, find all free P2P caches in the memory pool of the unified management module;
[0082] (2) If it does not exist, then feedback that the allocation fails; if the allocation fails, then this communication fails;
[0083] (3) Otherwise, perform the cache selection for the sending end (i.e., the destination data party EP1 in this embodiment), and based on the selection of the DMA engine in the previous step, select the cache closest to the DMA engine and allocate it to the sending end;
[0084] (4) Then perform the cache selection for the receiving end, and at this time, select the cache closest to the receiving end.
[0085] In this embodiment, the unified management layer makes a unified and reasonable planning for the memory resources of the entire PCIe network, reduces memory allocation conflicts, and tries to select the allocation method closest to the source, which can greatly save the path length of PCIe messages propagated in the PCIe network, thereby improving the bandwidth utilization rate of the entire PCIe network.
[0086] Third, in this embodiment, the specific data execution process for the EP1 side is as follows:
[0087] (1) The EP1 client first applies to the unified management module to establish a connection between EP1 and EP2. In a preferred embodiment, the components of the connection are set as {DMA engine, transmit buffer, receive buffer, data length}, where the data length refers to the length of the data to be transmitted between EP1 and EP2.
[0088] (2) When selecting a DMA engine, a more preferred method in this embodiment is as follows: The unified management module records the PCIe networking topology information collected as , where kmn represents the cost (i.e., cost) between two nodes in the PCIe network, m and n are the node numbers in the PCIe network, m, n are positive integers, and n represents the total number of nodes. Here, nodes include nodes such as EP and Switch in the PCIe network. First, select the DMA set denoted as set D {DMA1,..., DMAc}, where c is a positive integer and c is less than the number of nodes n. Then, screen out the cost set of the DMA in the DMA set associated with the destination node h from T, denoted as C {C1,..., Cl}, where I is a positive integer and its maximum value is 2c. Select the node with the minimum cost in C, denoted as Cmin. At this time, the DMA on the node numbered min is the DMA engine to be selected. If the set C is empty, the selection fails, and there is no valid DMA for this communication, and the selection fails.
[0089] (4) The unified management module records the memory information in the PCIe device information collected as M {Addr1 {start, end, bdf1},..., Addrp {start, end, bdfp}}, where p is a positive integer representing the total number of EPs in the PCIe network, bdf refers to the BDF address of the EP device in PCIe, and Addrp is the memory on EPp. Allocate a transmit buffer H1 on Addr1 and a receive buffer H2 on Addr2. If there is no free memory, the allocation fails. At this time, due to the buffer allocation failure, P2P communication cannot be performed, and a buffer allocation failure signal is returned. In a more preferred embodiment, after obtaining the memory information set M, first screen out the free memory set. If the memory corresponding to EP1 and / or EP2 is not free, that is, not in the free memory set, then allocate the buffer in the nearest distance manner. At this time, allocate a transmit buffer H1 for the data sender in the free memory closest to the data sender, and allocate a receive buffer H2 for the data receiver in the free memory closest to the data receiver. The judgment of the nearest distance is determined by the number of links crossed between the buffer and the target node. The smaller the number, the closer the distance.
[0090] More preferably here, the buffer is allocated preferentially on the memories of the data sender and the data receiver themselves. If there is no free memory in these two EP devices, then it is allocated in other free memories in the closest distance manner. It can also be considered that the memories of the data sender and the data receiver are the two memories closest to each of them respectively, that is, their links are both 0. If there is only one remaining free memory and this free memory is sufficient to allocate two buffers, then they can be set in the same free memory. If its memory size is not enough for allocation, then allocation failure is returned.
[0091] (5) The unified management module returns the processing result (i.e., the communication operation descriptor) to the EP1 client. This processing result includes whether it is the sender, the allocated DMA engine device (i.e., the optimal DMA engine), the send buffer, the receive buffer, etc.; among them, to notify EP1 whether it is the sender, a method such as identifying through a specific field can be adopted. For example, setting a certain field in a certain signal to 1 indicates that EP1 is the sender, and setting this field to 0 indicates that it is the receiver.
[0092] (6) If the EP1 client discovers that it is selected as the sender in this operation, it starts a DMA operation to the EP1 driver.
[0093] Here, it needs to be further pointed out that if EP1 discovers that it is not designated as the sender but as the receiver in this operation, then it follows the read operation process of EP2. At this time, it degrades to an unoptimized read operation process (i.e., the conventional read operation process); if EP1 discovers that it is not designated as the sender nor as the receiver in this operation, then an exception is reported at this time and the communication fails.
[0094] (7) After receiving the start of the DMA operation, the EP1 driver starts the hardware DMA, that is, starts the allocated DMA engine and constructs a DMA descriptor; the DMA descriptor here carries information such as the send buffer, the receive buffer, and the data length;
[0095] (8) After receiving the command, the allocated DMA engine performs data transfer, transfers the data in the source buffer (i.e., the send buffer) to the receive buffer (i.e., buffer H2 of EP2 in this embodiment), and triggers through the receiver EP2 doorbell. The doorbell trigger method is a conventional method in this field and will not be elaborated here.
[0096] IV. The processing process for the EP2 side is as follows:
[0097] (1) The EP2 client first applies to the unified management module to establish a connection with EP1.
[0098] After receiving the above request, the unified management module first checks whether the EP1-EP2 path has been allocated. If not, it turns to the path allocation process of EP1 for path allocation. If the EP1-EP2 path has been allocated, it directly sends the EP1-EP2 path information back to the EP2 client;
[0099] The unified management module simultaneously notifies the sending end (i.e., EP1 in this embodiment) that it can start communicating and synchronously notifies the EP2 client that it can start communicating;
[0100] After receiving the start communication notification information, the EP2 client establishes a polling thread and prepares to receive data;
[0101] After EP1 finishes sending data, EP2 polls all the data and performs subsequent service processing.
[0102] In this embodiment, it details how to convert the READ operation in the client operation into a Write transaction operation of PCIe DMA. This method effectively reduces the READ operations of the PCIe network and greatly increases the available bandwidth of the PCIe network.
[0103] Embodiment 2:
[0104] Combined with Figure 3 As shown, in another implementation manner of this solution, it can be implemented by an executable communication method. The method includes:
[0105] Step 1: The client applies to the unified management module for a communication path to read data from the data reading party EP2 to the destination data party EP1;
[0106] Step 2: The unified management module allocates the DMA engine based on the client's application;
[0107] Step 3: The unified management module allocates user data buffer resources based on the client's application. The user data buffer includes a sending buffer allocated for EP1 and a receiving buffer allocated for EP2;
[0108] Step 4: The unified management module establishes the EP1-EP2 path;
[0109] Step 5: After the path is established, EP1 starts the DMA operation and transfers the data to be read from the sending buffer of the destination data party to the receiving buffer of the data reading party through the allocated DMA engine;
[0110] Step 6: After the data transfer is completed, it ends.
[0111] The execution of this method can be used in the communication device in Embodiment 1.
[0112] In addition, it can be understood that any process or method description represented in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present solution includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of the present solution belong. The processor executes the various methods and processes described above. For example, the method embodiments in the present solution can be implemented as a software program, which is tangibly included in a machine-readable medium, such as a memory. In some embodiments, part or all of the software program can be loaded and / or installed via the memory and / or communication interface. When the software program is loaded into the memory and executed by the processor, one or more steps of the methods described above can be executed. Alternatively, in other embodiments, the processor can be configured to execute one of the above methods in any other suitable manner (e.g., by means of firmware).
[0113] The logic and / or steps represented in the flowchart or described in other ways herein can be specifically implemented in any readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices.
[0114] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the technical field to which the present invention pertains within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A PCIe-P2P communication device, characterized in that, The communication device includes: a user layer module, a unified management module, and a driver layer module; the communication device is connected to multiple EP devices through a PCIe network; the user layer module is used to apply for a communication path to the unified management module; the unified management module is used to collect information of PCIe devices, and select and process P2P channels; the unified management module uniformly manages the allocation of user data buffer resources and the allocation of the DMA engine; the unified management module plans the read operation of the user layer module as a write operation on the destination data side EP1 to the read data side EP2; the user data buffer includes a transmission buffer allocated for EP1 and a reception buffer allocated for EP2; the driver layer module receives the communication operation descriptor or register operation control data sent by the unified management module, and is used to control DMA data transfer and operate multiple EP devices; When performing P2P communication, the process of the destination data side EP1 is as follows: S11. EP1 applies to the unified management module to establish an EP1-EP2 path; S12. The unified management module allocates the best DMA engine, and allocates user data buffers for EP1 and EP2 to form a communication operation descriptor; the user data buffer allocation method is: The unified management module records the memory information in the collected PCIe device information as M{Addr1{start, end, bdf1}, …, Addrp{start, end, bdfp}}, where p represents the total number of EP devices in the PCIe network, bdf refers to the BDF address of the EP device in PCIe, Addrp represents the memory on EPp, start represents the start address of the corresponding memory, and end represents the end address of the corresponding memory; Find all free memories in M to form a free memory set, and respectively select the free memory closest to EP1 and the free memory closest to EP2 in the free memory set; allocate a transmission buffer H1 in the free memory closest to EP1, and allocate a reception buffer H2 in the free memory closest to EP2; S13. The unified management module sends the communication operation descriptor to EP1; S14. After receiving the communication operation descriptor, EP1 starts a DMA operation, starts the allocated DMA engine, and constructs a DMA descriptor; S15. Execute data transfer through the allocated DMA engine, and transfer the data in the transmission buffer to the reception buffer; the data in the transmission buffer is the data from EP1 that EP2 wants to read.
2. The communication device according to claim 1, wherein When performing P2P communication, the process of the read data side EP2 is as follows: S21. Apply to the unified management module for a path with EP1; S22. The unified management module checks whether a path has been allocated. If a path has been allocated, the result is sent to EP2. If no path has been allocated, a path is allocated and the allocation result is sent to EP2; S23. The unified management module sends a communication start notification to EP1 and EP2; S24. After receiving the communication start notification, EP2 establishes a polling thread and prepares to receive data; S25. After the EP1 data transmission is completed, EP2 polls all the data and completes the data reading operation.
3. The communication device according to claim 1, wherein In the S12, the specific method for allocating user data buffers for EP1 and EP2 further includes: if there is no free memory, return a failure signal for allocation.
4. The communication device according to claim 3, characterized in that, The method for judging the closest distance is: judging the distance based on the number of links spanned between the free memory and the target nodes, and the smaller the number of links, the closer the distance; the target nodes are EP1 and EP2.
5. The communication device according to claim 1, characterized in that, In the S11, the data structure of the path is: {DMA engine, transmit buffer, receive buffer, data length}, and the data length is the length of the data to be transmitted between EP1 and EP2.
6. The communication device according to claim 1, wherein In the S12, the priority for allocating the best DMA engine is: Priority one: preferentially select the DMA engine of EP1; Priority two: if Priority one does not exist or is unavailable, then select the DMA engine on the switching device; Priority three: if both Priority one and Priority two do not exist or are unavailable, then select the DMA engine on the EP device other than EP1; If none of Priority one, Priority two, and Priority three exist, return a failure signal for DMA engine allocation.
7. The communication device according to claim 6, wherein In the S12, the specific method for allocating the best DMA engine is: S121. The unified management module collects PCIe networking topology information and records it as T = , where k represents the cost between two nodes in the PCIe network, m and n respectively represent the node numbers, and n represents the total number of PCIe network nodes; S122. Aggregate to obtain a DMA set, denoted as D{DMA1,..., DMAc}, where c represents the number of DMAs, and c is less than n; S123. Screen out from the set T the cost set C of the DMAs in the set D associated with the destination data party, denoted as C{C1,..., CI}, where I is a positive integer, and its maximum value is 2c; if the set C is empty, return a failure signal for DMA engine allocation; S124. Screen out the node with the minimum cost in the set C and denote it as Cmin, and set the DMA on the node numbered min as the best DMA engine.
8. The communication device according to claim 1, characterized in that, The communication operation descriptor includes whether it is a sender, the allocated DMA engine, the transmit buffer, and the receive buffer.
9. A PCIe-P2P communication method, characterized in that, The method is applied to the communication device according to any one of claims 1-8, and the method includes: Step 1. The user side applies to the unified management module for a communication path to read data from the data reading party EP2 to the destination data party EP1; Step 2. The unified management module allocates a DMA engine based on the application of the user side; Step 3. The unified management module allocates user data buffer resources based on the application of the user side, and the user data buffer includes a transmit buffer allocated for EP1 and a receive buffer allocated for EP2; Step 4. The unified management module establishes an EP1-EP2 path; Step 5. After the path is established, EP1 starts a DMA operation, and transfers the data to be read from the transmit buffer of the destination data party to the receive buffer of the data reading party through the allocated DMA engine; Step 6. After the data transfer is completed, end.
Citation Information
Patent Citations
DMA device for nodes in multi-computer system and communication method
CN101539902A
Method and system for transmitting data between computing nodes and electronic equipment
CN109828843A