Video memory sharing method and device based on CXL protocol and RDMA protocol
By adopting the CXL and RDMA protocol video memory sharing method in multi-user and multi-graphics card scenarios, the problems of low utilization rate and high transmission delay in graphics card video memory sharing are solved, and efficient video memory sharing and computing rendering effects are achieved.
Patent Information
- Application Number
- CN202510112417.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
In the multi-user and multi-graphics card scenarios, the shared use of graphics card memory in the prior art has problems such as low utilization rate and high transmission delay.
The video memory sharing method based on the computing fast link (CXL) protocol and the remote direct memory access (RDMA) protocol is adopted. By building a global video memory unified management interface and lock-free ring queue on the local server, efficient management of video memory resources and rapid processing of graphics card commands are achieved, and users can directly access the video memory address of the local server through the RDMA protocol.
It improves the shared usage utilization rate of graphics card graphics card memory, reduces the latency of memory access management and transmission delay, and improves the computing throughput and memory utilization.
Smart Images

Figure CN120045135A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of graphics card virtualization. More specifically, it relates to a video memory sharing method and device based on the Compute Express Link (CXL) protocol and the Remote Direct Memory Access (RDMA) protocol. Background Technique
[0002] Graphics card virtualization technology generally refers to a graphics card resource sharing technology that multiplexes the graphics card on a server for use by virtual machines through software reuse / hardware resource partitioning and other methods. By extracting the display, computing, and other resources of the hardware graphics card and projecting them into the corresponding virtual machines, users can use the graphics card for accelerated rendering in the virtual machines, improve the performance of the virtual machines, and ultimately break through the indivisible barrier of the physical structure of the graphics card to achieve the purpose of sharing the physical server graphics card resources by multiple virtual machines.
[0003] Currently, the key difficulties in graphics card virtualization lie in achieving shared access and management of the graphics card video memory on the basis of virtualizing the graphics card hardware unit. Existing technologies usually simplify the GPU virtualization scenario. The physical server is responsible for allocating and initializing the video memory of the virtual machine graphics card. The virtual machine writes the computed and rendered graphic frames into the specified video memory, and then the physical server or the virtual machine copies and transfers the specified video memory to the remote user. In multi-user and multi-graphics card scenarios, there are still difficulties such as low utilization rate and high transmission latency in the shared use of the graphics card video memory. Summary of the Invention
[0004] Aiming at the defects of the existing technology, the purpose of this application is to provide a video memory sharing method and device based on the CXL protocol and the RDMA protocol, aiming to solve the problems of low utilization rate and high transmission latency in the shared use of the graphics card video memory in the existing technology in multi-user and multi-graphics card scenarios.
[0005] To achieve the above object, in the first aspect, this application provides a video memory sharing method based on the CXL protocol and the RDMA protocol, including:
[0006] Construct a global video memory unified management interface on the local server based on the Compute Express Link (CXL) protocol, and the global video memory unified management interface is used to achieve global unified management of video memory resources;
[0007] Share the video memory of the local server to the user terminal through the Remote Direct Memory Access (RDMA) protocol.
[0008] This application constructs a unified access entry for local video memory based on CXL technology, realizes efficient connectivity of local video memory, reduces the access management latency of local video memory, improves computing throughput, increases the shared utilization rate of video memory of the graphics card, uses RDMA technology, enables users to directly access the video memory address of the local server through the network, realizes efficient sharing of the video memory of the graphics card between the local server and remote users, and reduces the transmission latency of shared use of the video memory of the graphics card.
[0009] A video memory sharing method based on CXL protocol and RDMA protocol provided by the present invention, the global video memory unified management interface uses a double-layer hash table to implement mapping, wherein the upper-layer hash table records the mapping relationship from the video memory address to the corresponding graphics processing unit (GPU) of the graphics card, and the lower-layer hash table records the offset within the CXL video memory page from the address.
[0010] This application realizes fast address indexing through the mapping method of a double-layer hash table, reducing system coupling.
[0011] A video memory sharing method based on CXL protocol and RDMA protocol provided by the present invention, the method further includes:
[0012] Construct a lock-free circular queue on the local server, the lock-free circular queue is used to quickly process graphics card commands, and construct a stable and efficient interconnection path between the local processor and the video memory.
[0013] This application constructs a lock-free circular queue on the local server, quickly processes graphics card commands, reduces synchronization overhead, constructs a stable and efficient interconnection path between the server processor and the video memory, and improves the work efficiency of intensive computing and rendering of the local server and the utilization rate of the video memory.
[0014] A video memory sharing method based on CXL protocol and RDMA protocol provided by the present invention, the lock-free circular queue includes multiple channels, and in the initialization process, a video memory block with a fixed length and granularity alignment is allocated to each channel, and the writing order of multiple channels is alternating, and the reading order is in sequence number order.
[0015] A video memory sharing method based on CXL protocol and RDMA protocol provided by the present invention, the method further includes:
[0016] Construct a lock-free First In-First Out (FIFO) queue on the user terminal, and the lock-free FIFO queue is used to dynamically transmit video memory data to the user terminal.
[0017] This application constructs a lock-free FIFO queue on the user terminal to ensure the continuity of the user's screen and overcome the jitter of the RDMA network transmission speed.
[0018] A video memory sharing method based on the CXL protocol and the RDMA protocol provided by the present invention. The storage address of the queue elements of the lock-free FIFO queue is fixed during initialization. The queue read / write start address is not transmitted in the elements. The queue elements include data count and video memory cut-off pointer.
[0019] In a second aspect, the present application provides a video memory sharing device based on the CXL protocol and the RDMA protocol, including:
[0020] A construction module, configured to construct a global video memory unified management interface on a local server based on the Compute Express Link (CXL) protocol, and the global video memory unified management interface is used to implement global unified management of video memory resources;
[0021] A sharing module, configured to share the video memory of the local server to a user terminal through the Remote Direct Memory Access (RDMA) protocol.
[0022] In a third aspect, the present application provides an electronic device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is configured to execute the video memory sharing method based on the CXL protocol and the RDMA protocol described in the first aspect or any possible implementation manner of the first aspect.
[0023] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program. When the computer program runs on a processor, it causes the processor to execute the video memory sharing method based on the CXL protocol and the RDMA protocol described in the first aspect or any possible implementation manner of the first aspect.
[0024] In a fifth aspect, the present application provides a computer program product. When the computer program product runs on a processor, it causes the processor to execute the video memory sharing method based on the CXL protocol and the RDMA protocol described in the first aspect or any possible implementation manner of the first aspect.
[0025] It can be understood that the beneficial effects of the above second aspect to fifth aspect can refer to the relevant descriptions in the first aspect, and will not be elaborated here.
[0026] Generally speaking, compared with the prior art, the above technical solution conceived by the present application has the following beneficial effects:
[0027] (1) Build a unified access entry for local video memory based on CXL technology to achieve efficient connectivity of local video memory, reduce the access management latency of local video memory, improve computing throughput, increase the shared utilization rate of the video card's video memory, use RDMA technology to enable users to directly access the video memory address of the local server through the network, and achieve efficient sharing of the video card's video memory between the local server and remote users, reducing the transmission latency of the shared use of the video card's video memory.
[0028] (2) Implement fast address indexing through a double-layer hash table mapping method to reduce system coupling.
[0029] (3) Build a lock-free circular queue on the local server to quickly process video card commands, reduce synchronization overhead, and build a stable and efficient interconnection path between the server processor and the video memory to improve the work efficiency of intensive computing and rendering on the local server and the utilization rate of the video memory.
[0030] (4) Build a lock-free FIFO queue on the user terminal to ensure the continuity of the user's screen and overcome the jitter of the RDMA network transmission speed. Description of the Drawings
[0031] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0032] Figure 1 is a schematic flowchart of a video memory sharing method based on the CXL protocol and the RDMA protocol provided by an embodiment of the present application;
[0033] Figure 2 is a schematic diagram of the working principle of the global video memory unified management interface provided by an embodiment of the present application;
[0034] Figure 3 is a schematic diagram of the principle of a lock-free circular queue provided by an embodiment of the present application;
[0035] Figure 4 is a schematic diagram of the principle of a lock-free FIFO queue provided by an embodiment of the present application;
[0036] Figure 5 is a general framework diagram of a video memory sharing method based on the CXL protocol and the RDMA protocol provided by an embodiment of the present application;
[0037] Figure 6 is a schematic structural diagram of a video memory sharing device based on the CXL protocol and the RDMA protocol provided by an embodiment of the present application;
[0038] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0039] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0040] The term "and / or" in this document is an association relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The symbol " / " in this document represents an "or" relationship between associated objects. For example, A / B represents A or B.
[0041] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner.
[0042] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more. For example, a plurality of processing units refers to two or more processing units, etc.; a plurality of elements refers to two or more elements, etc.
[0043] Next, in combination with Figures 1 - 5 An off-chip memory sharing method based on the CXL protocol and the RDMA protocol provided in the embodiments of the present application will be introduced.
[0044] Figure 1 It is a schematic flowchart of an off-chip memory sharing method based on the CXL protocol and the RDMA protocol provided by an embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0045] Step 100, construct a global off-chip memory unified management interface on the local server based on the Compute Express Link (CXL) protocol. The global off-chip memory unified management interface is used to achieve global unified management of off-chip memory resources;
[0046] The CXL protocol is a new type of open interconnect technology standard. It aims to achieve high-speed and efficient interconnection between a central processing unit (CPU) and a GPU, a field programmable gate array (FPGA), or other accelerators to meet the needs of high-performance heterogeneous computing. The main advantages of the CXL protocol lie in its high compatibility and memory consistency.
[0047] This application constructs a unified local video memory access entry based on CXL technology, that is, a global video memory unified management interface is built in the local server to achieve the global unified management of video memory resources, thereby realizing the efficient connection of local video memory, reducing the access management latency of local video memory, and improving the computing throughput.
[0048] Step 110, share the video memory of the local server to the user terminal through the remote direct memory access (RDMA) protocol.
[0049] RDMA is an efficient network communication technology that allows computers in a network to directly access the memory of another computer without excessive intervention from the operating system and CPU.
[0050] This application uses the RDMA access mechanism between the local server video memory and the user terminal to enable users to directly access the video memory address of the local server through the network, realizing the efficient sharing of the video memory of the graphics card between the local server and remote users.
[0051] In an embodiment of this application, the RDMA network video memory interface design for the user terminal video memory access is as shown in the following table:
[0052] Table 1 Main RDMA network video memory access interfaces
[0053]
[0054] The Open / Close operation opens / closes the video memory resources according to user configuration. The Alloc / Free operations are used to allocate and release video memory blocks. The Read / Write operations read and write data of a set size to a specified address through the RDMA network.
[0055] A video memory sharing method based on the CXL protocol and the RDMA protocol provided by this application constructs a unified local video memory access entry based on CXL technology, realizes the efficient connection of local video memory, reduces the access management latency of local video memory, improves the computing throughput, improves the sharing utilization rate of the video memory of the graphics card, uses the RDMA technology to enable users to directly access the video memory address of the local server through the network, realizes the efficient sharing of the video memory of the graphics card between the local server and remote users, and reduces the sharing transmission latency of the video memory of the graphics card.
[0056] In some embodiments, the global video memory unified management interface in step 100 is implemented with a double-layer hash table for mapping. The upper-layer hash table records the mapping relationship from the video memory address to the corresponding graphics processing unit (GPU) of the graphics card, and the lower-layer hash table records the offset within the CXL video memory page for the address.
[0057] Figure 2 is a schematic diagram of the working principle of the global video memory unified management interface provided by the embodiments of the present application. As Figure 2 shown, this interface is used to achieve global unified management of video memory resources, responsible for managing the allocation of global video memory pages, and dividing and managing the video memory pool by pages. The global video memory unified management interface uses a double-layer hash table to implement video memory mapping. The upper-layer hash table records the mapping relationship from page_id to the corresponding GPU (gpu id) of the graphics card; the lower-layer hash table records the offset cxloffset within the CXL video memory page for page_id. Through this hierarchical mapping method, fast address indexing is achieved, reducing system coupling.
[0058] The global unified address occupies a total of 64 bits and consists of two parts: the high-order page_id occupies 52 bits, and the low-order page_offset occupies 12 bits. The page_id is issued by the global video memory unified management interface and is used to identify the video memory page; the page_offset represents the offset within the video memory page, ultimately achieving fast addressing from the global unified address to the CXL-managed video memory address.
[0059] In some embodiments, the method further includes:
[0060] Step 120: Build a lock-free circular queue on the local server. The lock-free circular queue is used to quickly process graphics card commands and build a stable and efficient interconnection path between the local processor and the video memory.
[0061] The present application builds a lock-free circular queue on the local server to quickly process graphics card commands, reducing the synchronization overhead, building a stable and efficient interconnection path between the server processor and the video memory, and improving the intensive computing and rendering work efficiency of the local server and the video memory utilization rate.
[0062] In some embodiments, the lock-free circular queue in step 120 includes multiple channels. During the initialization process, a video memory block with a fixed length and granularity alignment is allocated for each channel. The multi-channel writing order is alternating, and the reading order is sequential by number.
[0063] Figure 3 is a schematic diagram of the principle of the lock-free circular queue provided by the embodiments of the present application. As Figure 3 shown, the physical server writes the graphics card driver commands into the lock-free circular queue through a unified interface. To implement the concurrent function of the graphics card commands, the constructed queue includes multiple channels.Figure 3 Taking four channels as an example in China.
[0064] As needed, in the initialization process, a video memory block with a fixed length and aligned granularity (usually 64 bytes) is allocated for each channel. The video memory blocks within a channel are used to write graphics card commands. Four video memory blocks form a lock-free circular queue. When actually writing graphics card commands concurrently in multiple threads, the commands are written alternately in the order of the dotted lines in the figure.
[0065] As Figure 3 shown, thread a writes to the first position (1) of the first line of the video memory block, thread b writes to the first position (2) of the second line of the video memory block, and so on. When thread a writes for the second time, it writes from the second position (5) of the first cache line. The command execution proceeds in the order of the circular queue 1 - 2 - 3 - 4 - n.
[0066] Through this lock-free circular queue, the synchronization overhead can be reduced when multiple threads concurrently execute graphics card commands, improving the rendering work efficiency and video memory utilization rate.
[0067] In some embodiments, the method further includes:
[0068] Step 130: Construct a lock-free first-in-first-out (FIFO) queue on the user terminal. The lock-free FIFO queue is used to dynamically transfer video memory data to the user terminal.
[0069] To ensure the continuity of the user's screen and overcome the jitter of the RDMA network transmission speed, a lock-free FIFO queue is designed on the user terminal side to store the read frame data. In chronological order, each frame of data is guaranteed to be first-in-first-out. Finally, the frame data parsing and decoding module on the user side parses and decodes the video memory frame data and then uses it for virtual machine screen display.
[0070] In some embodiments, the storage address of the queue elements of the lock-free FIFO queue in step 130 is fixed during initialization. The queue read and write start addresses are not transmitted in the elements. The queue elements include a data count and a video memory end pointer.
[0071] Figure 4 is a schematic diagram of the principle of the lock-free FIFO queue provided by the embodiments of the present application. As Figure 4 shown, a unidirectional lock-free FIFO queue is constructed on the user terminal. Since the storage address addr of the queue elements is fixed during initialization, the queue read and write start addresses do not need to be transmitted in the elements. The designed queue elements include two parts: a data count cnt and a video memory end pointer end.
[0072] The data count cnt represents the number of data blocks contained in a single element, and the video memory end pointer end represents the end address of the last data block in the element. Then the read / write start address is ptr = addr+(0, 1, 2, …cnt-1), and the read / write end address is ptr = addr+end.
[0073] The feature of this FIFO queue is that it considers the real-time change in the amount of video memory data transferred during screen updates. To avoid bandwidth waste caused by setting queue elements of a fixed size, it is designed to be compatible with data of dynamically changing sizes by pointing queue element blocks to data sub-blocks, ensuring the continuity of the user's screen and overcoming the jitter in the transmission speed of the RDMA network.
[0074] Figure 5 It is the overall framework diagram of the video memory sharing method based on the CXL protocol and the RDMA protocol provided by the embodiments of this application. As Figure 5 shown, the overall structure of the shared video memory includes a physical server video memory management module and remote user video memory access, and the latter can be adjusted according to the number of remote users. The physical server video memory management module consists of a global video memory unified management interface and a lock-free circular queue module, and the latter includes a lock-free FIFO queue module.
[0075] The physical server video memory management module constructs a CXL video memory resource pool within the physical server, provides unified address management, allocation, and usage interfaces for the local graphics card driver, writes graphics card driver commands into the lock-free circular queue through a unified video memory address, and the CXL protocol is responsible for sequentially writing the driver commands into the graphics card hardware and executing them, saving the final frame data in the set video memory, and waiting for remote users to obtain the frame data through the RDMA network.
[0076] This application deploys CXL on the local server, constructs a lock-free circular queue for quickly processing graphics card commands, reduces the synchronization overhead, constructs a stable and efficient interconnection path between the processor and the video memory, and improves the local server's intensive computing and rendering work efficiency and video memory utilization rate. An RDMA access mechanism is introduced between remote users and local video memory to construct a lock-free unidirectional video memory FIFO queue, avoiding the overhead of lock operations and reducing the transmission delay. Thus, a lock-free video memory sharing design method based on CXL / RDMA is realized.
[0077] Figure 6 It is the structural schematic diagram of the video memory sharing device based on the CXL protocol and the RDMA protocol provided by the embodiments of this application. As Figure 6 shown, the device includes a construction module 610 and a sharing module 620, where:
[0078] A building block 610 is used to build a global video memory unified management interface on a local server based on the Compute Express Link (CXL) protocol. The global video memory unified management interface is used to achieve the global unified management of video memory resources.
[0079] A sharing module 620 is used to share the video memory of the local server to a user terminal through the Remote Direct Memory Access (RDMA) protocol.
[0080] It should be understood that the above device is used to execute the method in the above embodiment. For the corresponding program module in the device, its implementation principle and technical effect are similar to the description in the above method. The working process of the device can refer to the corresponding process in the above method and will not be elaborated here.
[0081] Based on the method in the above embodiment, Figure 7 An example of the physical structure diagram of an electronic device is shown as Figure 7 As shown, an embodiment of the present application provides an electronic device, which may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 complete mutual communication through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute the video memory sharing method based on the CXL protocol and the RDMA protocol in the above embodiment.
[0082] In addition, when the logical instructions in the above memory 730 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the video memory sharing method based on the CXL protocol and the RDMA protocol described in various embodiments of the present application.
[0083] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a processor, it enables the processor to execute the video memory sharing method based on the CXL protocol and the RDMA protocol in the above embodiment.
[0084] Based on the method in the above embodiments, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, it causes the processor to execute the video memory sharing method based on the CXL protocol and the RDMA protocol in the above embodiments.
[0085] It can be understood that the processor in the embodiment of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0086] The method steps in the embodiment of the present application may be implemented in a hardware manner or by a processor executing software instructions. The software instructions may be composed of corresponding software modules. The software modules may be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in an ASIC.
[0087] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0088] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.
[0089] Those skilled in the art can easily understand that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A video memory sharing method based on CXL protocol and RDMA protocol, characterized in that: include: Building a global unified graphics memory management interface on a local server based on a computing fast link CXL protocol, wherein the global unified graphics memory management interface is used to implement global unified management of graphics memory resources; The video memory of the local server is shared with the user terminal based on the Remote Direct Memory Access (RDMA) protocol.
2. The video memory sharing method based on the CXL protocol and the RDMA protocol according to claim 1, characterized in that: The global unified video memory management interface uses a double-layer hash table to achieve mapping, wherein the upper hash table records the mapping relationship between the video memory address and the corresponding graphics card GPU, and the lower hash table records the offset of the address to the CXL video memory page.
3. The video memory sharing method based on the CXL protocol and the RDMA protocol according to claim 1, characterized in that: The method further comprises: A lock-free circular queue is constructed on the local server, and the lock-free circular queue is used to quickly process graphics card commands and to construct a stable and efficient interconnection path between the local processor and the graphics memory.
4. The video memory sharing method based on the CXL protocol and the RDMA protocol according to claim 3, characterized in that: The lock-free circular queue includes multiple channels. The initialization process allocates a fixed-length, granularity-aligned video memory block to each channel. The writing order of multiple channels is alternating, and the reading order is in sequence number order.
5. The video memory sharing method based on the CXL protocol and the RDMA protocol according to claim 1, characterized in that: The method further comprises: A lock-free first-in first-out FIFO queue is constructed on the user terminal, and the lock-free FIFO queue is used to dynamically transmit video memory data to the user terminal.
6. The video memory sharing method based on the CXL protocol and the RDMA protocol according to claim 5, characterized in that: The queue element storage address of the lock-free FIFO queue is fixed during initialization, the queue read and write start address is not transmitted in the element, and the queue element includes a data count and a video memory cutoff pointer.
7. A video memory sharing device based on CXL protocol and RDMA protocol, characterized in that: include: A construction module is used to construct a global unified graphics memory management interface based on the computing fast link CXL protocol on a local server, wherein the global unified graphics memory management interface is used to implement global unified management of graphics memory resources; The sharing module is used to share the video memory of the local server with the user terminal based on the remote direct memory access RDMA protocol.
8. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the video memory sharing method based on the CXL protocol and the RDMA protocol as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor, the processor is enabled to execute the video memory sharing method based on the CXL protocol and the RDMA protocol as described in any one of claims 1 to 6.
10. A computer program product, characterized in that When the computer program product runs on a processor, the processor is enabled to execute the video memory sharing method based on the CXL protocol and the RDMA protocol as described in any one of claims 1 to 6.
Citation Information
Cited By
CXL pooling shared memory-oriented read-write consistency guarantee method
CN121116664A