Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

20 results about "Remote memory access" patented technology

Parallel communication system for chips

The application relates to a chip parallel communication system. The system comprises a plurality of nodes, each node comprising a GMAC unit, and a processor unit, an MC unit and a NIC unit connected to the GMAC unit; the GMAC unit is configured to, in response to receiving an access request, transmit a first remote memory access request to the NIC unit and a second remote memory access request to the MC unit; the NIC unit is configured to transmit the first remote memory access request to a corresponding external node; the MC unit is configured to read or write data from a memory coupled to the MC unit according to at least one of a local memory access request and the second remote memory access request, and in the case of reading out data, transmit the read-out data to a corresponding processor unit through the GMAC unit; by configuring the GMAC unit and the processor unit, the MC unit and the NIC unit connected to the GMAC unit, the application reduces the communication delay between the nodes, thereby improving the data volume communication parallel computing performance, and improving the system performance and efficiency.
Owner:SOUTHERN POWER GRID DIGITAL GRID RESEARCH INSTITUTE CO LTD

Numa system page cache copy management method and system

ActiveCN122132334BRemote memory accessTerm memory
The application provides a NUMA system page cache copy management method and system. The application first generates a read-only copy page in the local memory of an access node through the linkage of target file page cache identification, access subject NUMA node information acquisition and policy copy judgment, thereby avoiding cross-node remote memory access. Then, in combination with local copy priority mapping and page fault triggered by unmapping when process node scheduling is changed, the dynamic adaptation of process scheduling behavior and page mapping relationship is realized. Finally, in the read page fault interrupt process, the page mapping update is completed based on the copy strategy, and the association relationship between the node and the access task is recorded. Without damaging the file page cache semantics, the remote access overhead of the multi-NUMA node system file page cache can be greatly reduced, the problem of continuous remote access caused by process node drift can be solved, and the overall file access performance and running stability in the concurrent scenario can be significantly improved.
Owner:CHINA UNICOM DIGITAL TECNOLOGY CO LTD

Method for direct remote memory access for graphics processors, computing device, computer-readable storage medium and computer program product

The present application relates to a method for direct remote memory access of a graphics processor, a computing device, a computer readable storage medium and a computer program product. The method comprises: obtaining page granularity of source device memory and target device memory to be communicated respectively; in response to determining that at least one of the source device memory and the target device memory is not aligned to a predetermined size page granularity of PCIE, constructing a virtual address space mapped to a physical address space of the device unaligned memory at the predetermined size page granularity, so that the mapped virtual address space meets the requirements of a PCIE page table; and communicating between the memories of the source device and the target device in a direct remote memory access manner based on the constructed virtual address space. The present application can avoid configuring additional memory page aligned physical memory as an intermediate cache, and can simplify memory call operations, and significantly reduce transmission delay.
Owner:SHANGHAI BIREN TECH CO LTD

MCM-gpu adaptive last level cache structure and cache switching method thereof

The application discloses an MCM-GPU adaptive last-level cache structure arranged in a GPU module, comprising a Tag Array and a Date Array, wherein the Date Array is used for storing data, and the Tag Array is used for checking whether the data corresponding to an address is in the cache; the application further comprises a local memory access queue used for storing a memory access request of the current GPU module, a remote memory access queue used for storing a memory access request of other GPU modules, an LLC architecture change flag register used for storing an LLC architecture change flag indicating whether the architecture organization mode of the current last-level cache needs to be changed, and an LLC architecture flag register used for storing an LLC architecture flag indicating whether the current last-level cache is switched to a private last-level cache design or a shared last-level cache design. The application can support the dynamic switching of the shared last-level cache and the private last-level cache, can adaptively select the last-level cache architecture organization mode according to the configuration during program running, can meet the program memory access requirement, and can improve the performance of the MCM-GPU.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

A method and system for improving the speed of writing shared main memory critical resources in parallel from cores based on a new generation sunway many-core processor

ActiveCN116909741BAvoid locking operationshigh speedSupercomputerComputer architecture
The application relates to a method and system for improving the speed of writing shared main memory critical resources in parallel by a new generation Shenwei many-core processor, which comprises the following steps: a slave core applies for a data space on its private local data memory; critical resource data in the main memory is copied to the respective private local data memory; each slave core performs read-write operation; each slave core initiates a reduction operation through a remote memory access (RMA) channel, wherein the reduction operation refers to performing certain aggregation function operation on the critical resource data in the private local data memory of the plurality of slave cores to obtain a final result; and the critical resource data in the private local data memory after the reduction operation is written back to the main memory through a direct memory access (DMA) channel. The method can effectively improve the speed of reading and writing the shared main memory critical resources by the slave core of the Shenwei many-core processor, and improve the performance and efficiency of the supercomputer.
Owner:SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN +1

Server-free remote memory access performance optimization method based on cross-process memory tracking

PendingCN122086606Aeliminate overheadAchieve shared awarenessResource allocationMemory systemsPathPingRemote memory access
The invention discloses a server-free remote memory access performance optimization method based on cross-process memory tracking, and belongs to the technical field of computer memory management. The method comprises the following steps: designing a memory state table as a unified perception layer for remote memory access of homologous server-free containers, and realizing cross-process tracking of memory page states among different containers; a memory state table query process is embedded into hardware page table traversal, page state query is completed by hardware acceleration, and extra overhead on a memory access key path is eliminated; the remote memory access is merged and optimized based on the memory state table, cached local memory pages are reused in the container creation stage, the same remote memory access requests are merged in the operation stage, and the page missing exception frequency and the network bandwidth contention are reduced. According to the method, extra overhead caused by remote memory access in a server-free environment is effectively reduced, page missing abnormity and network bandwidth contention in a concurrent scene are greatly reduced, and the running performance and the system expandability of memory-intensive server-free applications are remarkably improved.
Owner:HUAZHONG UNIV OF SCI & TECH

Data packet out-of-order recovery method based on selective confirmation

PendingCN121508749AError preventionData packRemote memory access
The invention relates to the technical field of remote memory access, in particular to a selective confirmation-based data packet out-of-order recovery method, which comprises the following steps of: when a data packet is received, judging whether a hole region exists in a bitmap, whether the hole region is a newly generated hole region and whether all hole regions can be cleared; and feeding back a selective acknowledgement message or an acknowledgement response frame according to a judgment result. In order to solve the problem that the RoCE in the prior art needs to repeatedly transmit subsequent data packets for multiple times in the retransmission process, in the embodiment, the data packets needing to be retransmitted are indicated on the receiving end side in a bitmap maintenance mode. In the process of receiving the data packets, filling is sequentially carried out according to the received serial numbers of the data packets, the holes needing to be retransmitted are indicated, the sending end retransmits the specific data packets according to the fed back hole information, the backspacing process does not need to be completely executed, and therefore the transmission efficiency is improved.
Owner:SHANGHAI INTELLIGENT COMPUTING TECHNOLOGY CO LTD

Aggregating small remote memory access requests

ActiveCN117951051BRemote memory accessEngineering
The present disclosure relates to aggregating small remote memory access requests. A network interface card (NIC) receives a stream of commands, respective commands including memory operation requests, each request associated with a destination NIC. The NIC asynchronously buffers the requests into queues based on the destination NIC, each queue specific to a corresponding destination NIC. When a first queue of requests reaches a threshold, the NIC aggregates the first queue of requests into a first packet and sends the first packet to the destination NIC. The NIC receives a plurality of packets, a second packet including memory operation requests, each request associated with a same destination NIC and a destination core. The NIC asynchronously buffers the requests of the second packet into queues based on the destination core, each queue specific to a corresponding destination core. When a second queue of requests reaches a threshold, the NIC aggregates the second queue of requests into a third packet and sends the third packet to the destination core.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Memory allocation method and electronic device

ActiveCN120892214BResource allocationComputer networkRemote memory access
The application discloses a memory allocation method and an electronic device, and relates to the technical field of computers, which comprises the following steps: identifying semantic information of an object, dynamically binding the semantic information with a cache area and a memory allocation strategy, improving the flexibility of cache replacement, and realizing differentiated management and control of cache resources. The selection of the memory access node is optimized by fusing remote delay, shared domain characteristics and the semantic information of the object. The management of the L3 cache area, the memory pool and the memory access node in the related art is in an independent optimization state. Through semantic-driven partition binding and cross-node decision-making, the cache hit rate is improved and the delay of cross-node memory access is reduced. The problems of low cache replacement flexibility, low cache utilization and high remote memory access delay are solved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Method and apparatus for implementing remote memory access on basis of ethernet

PCT designated stageWO2026067640A1TransmissionRemote memory accessIp address
Disclosed in the present application are a method, apparatus and system for implementing remote memory access on the basis of Ethernet, and an electronic device and a medium. The method comprises the following steps: at a transport layer, using a source port number field, a destination port number field and an RDMA base transport header to encapsulate a packet generated by an application layer; enabling a packet header of the packet at a network layer to comprise a source IP address field, a destination IP address field, a type of service field, a total length field and a protocol field, wherein the type of service field, the total length field and the protocol field are used for representing a frame priority, a frame length and an upper-layer protocol, respectively; and at a data link layer, using the source IP address field and the destination IP address field to replace a source MAC address field and a destination MAC address field, respectively, so as to address the packet. The present application can reduce communication overheads and improve the communication performance.
Owner:THEO END (SHENZHEN) COMPUTING TECHNOLOGY CO LTD

An ai-based carbon data intelligent calculation operating system

This invention relates to the field of electronic digital data processing technology and discloses an AI-based intelligent computing operating system for energy and carbon data. The system includes an energy and carbon tensor management module, a bypass scheduling module, and a mapping processing module. The energy and carbon tensor management module allocates a non-paged, continuous physical space in memory to store time-series energy consumption data streams. The bypass scheduling module acquires the server's memory access topology graph and identifies the local memory nodes corresponding to different cores. The mapping processing module extracts topology boundary nodes based on the computational topology correlation matrix and generates kernel-state topology boundary cut sets. The operating system implements physical mirror multi-projection mapping based on the topology graph and cut sets, establishing redirection paths to convert remote memory access into local addressing access. This invention eliminates hardware cache consistency probe signals through a mirroring mechanism, solving the memory access latency bottleneck in large-scale graph computation and ensuring computational determinism.
Owner:SINRIDIGITALCITYTECCO LTD

System level cache with configurable partitioning

ActiveUS12625813B2Memory systemsRemote memory accessData memory
A data processing apparatus includes one or more cache configuration data stores, a coherence manager, and a shared cache. The coherence manager is configured to track and maintain coherency of cache lines accessed by local caching agents and one or more remote caching agents. The cache lines include local cache lines accessed from a local memory region and remote cache lines accessed from a remote memory region. The shared cache is configured to store local cache lines in a first partition and to store remote cache lines in a second partition. The sizes of the first and second partitions are determined based on values in the one or more cache configuration data stores and may or not overlap. The cache configuration data stores may be programmable by a user or dynamically programmed in response to local memory and remote memory access patterns.
Owner:ARM LTD

Super-large scale network security feature library loading method based on far-end memory access

PendingCN121984741ASecuring communicationHigh level techniquesRemote memory accessTerm memory
The invention discloses a super-large scale network security feature library loading method based on remote memory access, which relates to the field of network security, and comprises the following steps: acquiring a network security feature library and network security monitoring equipment; constructing a far-end memory access cluster formed by combining a plurality of computing nodes, and configuring a local memory access module and a far-end memory access module; performing logic division on the network security feature data in the network security feature library to obtain logic fragments and mapping the logic fragments into corresponding computing nodes; acquiring operation monitoring data corresponding to the network security feature library and the network security monitoring equipment; respectively performing access scheduling analysis on the network security monitoring equipment and dynamic expansion analysis on the network security feature library according to the obtained operation monitoring data, and generating dynamic adjustment information; performing self-adaptive adjustment on the access identification process of each computing node according to the obtained dynamic adjustment information; according to the invention, the timeliness and flexibility of the loading process are improved.
Owner:HUNAN DATACOM INFORMATION TECH SERVICE CO LTD

Network interface instruction processing method and apparatus for multi-path remote memory access

PendingCN122420262AQuality of serviceRemote memory access
This invention discloses a network interface instruction processing method and apparatus for multi-path remote memory access, belonging to the field of high-performance computing. The method includes state maintenance of a multi-path instruction queue, read request scheduling of the multi-path instruction queue, on-chip caching of the multi-path instructions, and network transmission request scheduling of the multi-path instruction queue. The apparatus includes a state maintenance controller, a read request scheduler, an on-chip cache device, a network transmission request scheduler, and a computer-readable storage medium. This invention can meet the high-speed processing requirements of thousands of multi-path remote memory access instructions, ensuring fair scheduling and efficient response for all remote memory accesses. It aims to solve the technical problems of low throughput, high latency, uneven cache resource allocation, poor scheduling fairness, and insufficient quality of service in large-scale node interconnection scenarios when network interfaces process massive numbers of multi-path remote memory access instructions.
Owner:NAT UNIV OF DEFENSE TECH

Remote memory access function test method and device, equipment and storage medium

PendingCN121691125ATransmissionComputer hardwareRemote memory access
The invention discloses a remote memory access function test method, device and equipment and a storage medium, and relates to the technical field of network testing, and the method comprises the steps: generating a test message according with a remote memory access protocol specification based on a test script, and sending the test message to a switch; capturing an actual transmission message corresponding to the test message from a mirror image port of the switch; based on the actual transmission message, extracting a header encapsulation field and a remote memory access message; according to the method, the header encapsulation field and the remote memory access message in the message are extracted and separated, the encapsulation compliance and the transmission logic correctness are independently and collaboratively verified, and the test report is generated. Full-process automatic and high-coverage-rate verification of the remote direct memory access function from network packaging to semantic transmission is achieved, and the accuracy and execution efficiency of the testing process are remarkably improved.
Owner:SHENZHEN FENGRUNDA TECH CO LTD

Distributed query method, apparatus, system and storage medium

The application discloses a distributed query method, device, system and storage medium, comprising: obtaining a query operation request of a user; obtaining key-value pair data according to the query operation request of the user, and outputting a correct multicast sent data packet after processing; obtaining a batch processed data stream operation request according to the key-value pair data; performing remote memory access processing on the batch processed data stream operation request to obtain a correct multicast sent data packet; processing the correct multicast sent data packet to obtain a rate limited data packet; sending the rate limited data packet to a remote memory; and tracking a transmission state of the data packet. The application optimizes data transmission and query processing between nodes in a distributed architecture.
Owner:BEIJING UNIV OF POSTS & TELECOMM

A dynamic resource perception and adaptive computing task scheduling method and system for heterogeneous edge computing devices

This invention relates to the fields of distributed computing and edge computing technology, providing a dynamic resource awareness and adaptive computing task scheduling method and system for heterogeneous edge computing devices. The method collects node runtime hardware status parameters through a sliding time window, including storage resource access conflict indicators, storage resource residency status, computing unit operating status, and data transmission link status. Based on these parameters, it performs node availability determination, excluding nodes from the scheduling candidate set when there is a risk of resource conflict or transmission bottleneck. A distributed scheduling structure is constructed based on the candidate node set, and tasks are adaptively partitioned and matched according to node resource status. When system load increases, a constraint control mechanism reduces the task allocation ratio of high-resource nodes and limits their upper limit, while increasing the participation of low-load nodes to control task distribution across multiple nodes. When a node malfunctions, intermediate state data is extracted and migrated to other nodes for continued execution. Furthermore, under network conditions that meet bandwidth and latency requirements, a remote memory access mechanism enables collaborative use of cross-node storage resources, thereby expanding computing resource capabilities.
Owner:董根源 +1

Multi-host remote memory access

An apparatus and method for efficiently performing remote memory access requests among multiple processing nodes. In various implementations, a computing system has a first node and a second node. Each of these nodes has a corresponding virtual address space in the computing system, and each node assigns subdivisions of the virtual address spaces to multiple clients of the node. The nodes assign a subset of a first virtual address space of a first client in the first node to remote data stored in a second virtual address space of a second client in the second node. Remote presence check (RPC) circuits of the nodes assign the second virtual address space to a subset of a network physical address (NPA) space. The RPC circuits of the nodes use assignments to the NPA space to verify that address mappings in TLBs are still available prior to routing memory access requests.
Owner:ADVANCED MICRO DEVICES INC +1

NFS file system I / O processing method based on NUMA architecture

The invention discloses an NFS (Network File System) I / O (Input / Output) processing method based on an NUMA (Non Uniform Memory Access) architecture, which is characterized in that a corresponding configuration mechanism of a network port and a memory node is introduced under the NUMA architecture, so that an NFS server can identify a NODE to which a connection link belongs in a client mounting stage, and a working process is selected and awakened from a corresponding nfsd process pool based on the NODE attribute in an I / O request stage; therefore, the front-end network data and the rear-end processing thread are guaranteed to be located in the same NUMA node, localization of data processing and memory access is achieved, the problem of cross-node memory access caused by random scheduling of the nfsd process in the prior art is effectively avoided, high delay and bandwidth bottleneck caused by remote memory access are eliminated, and the service life of the NUMA node is prolonged. The I / O data is transmitted and copied in the same node, the I / O throughput rate and response speed of the NFS server in a high-concurrency scene are remarkably improved, and the CPU cache hit rate and the resource utilization rate are improved.
Owner:HUNAN TONGYOU FEIJI TECH CO LTD

Numa system page cache copy management method and system

PendingCN122132334AMemory adressing/allocation/relocationRemote memory accessTerm memory
This application provides a method and system for managing page cache replicas in a NUMA system. First, by linking target file page cache identification, NUMA node information acquisition of the accessing entity, and policy-based replica judgment, a read-only replica page is generated in the local memory of the accessing node, avoiding cross-node remote memory access. Then, by combining local replica priority mapping and page fault triggering during process node scheduling changes, dynamic adaptation between process scheduling behavior and page mapping relationships is achieved. Finally, during the read page fault interruption process, page mapping is updated based on the replica strategy, and the association between nodes and access tasks is recorded. This significantly reduces the overhead of remote access to file page caches in multi-NUMA node systems without disrupting the semantics of file page caches, while simultaneously solving the problem of continuous remote access caused by process node drift, significantly improving the overall file access performance and operational stability in concurrent scenarios.
Owner:CHINA UNICOM DIGITAL TECNOLOGY CO LTD