Memory access method, apparatus, device, storage medium, and product
Patent Information
- Application Number
- CN202611232376.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本申请的主要目的在于提供一种内存访问方法、装置、设备、存储介质及产品,旨在解决降低了内存访问的效率的技术问题
与相关技术中,内核态的内存分配完全不受用户态NUMA策略的约束,设备驱动、DMA缓冲区等仍可能在任意NUMA节点分配内存,产生不可控的跨路流量,且ARM CMN(Coherent Mesh Network)网络和CXL(Compute Express Link)控制器为授权IP,无法直接修改其内部逻辑,并进行推广,降低了内存访问的效率相比,本申请应用于访问优化模块,所述访问优化模块至少包括访问分析模块和热点缓存模块,由于访问优化模块独立与ARMCMN(Coherent Mesh Network)网络和CXL(Compute Express Link)控制器,所以本申请在接收到CMN网络发送的内存访问请求,且内存访问请求对应的目标物理地址不属于本地内存地址范围的情况下,会将所述内存访问请求转发至目标模块,在目标模块为访问分析模块时,且目标物理地址所处的统计区域为热点区域,直接对热点缓存模块进行内存访问,不需要进行软件层面修改,优化模块被放置在CMN与CXL控制器之间的必经通路上,无论跨路请求来自用户态应用还是内核态驱动,一律在硬件层面被截获,提高了内存访问的效率。
Smart Images

Figure CN122817138A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multiprocessor technology, and in particular to a memory access method, apparatus, device, storage medium and product. Background Technology
[0002] In the field of ARM servers, optimization of NUMA (Non-Uniform Memory Access) architecture usually involves optimizing the software layer, such as the operating system and application software. For example, by using tools such as numactl and mbind to formulate user-space NUMA strategies, the memory of the application is bound to the local memory of the node where the application resides as much as possible to reduce the latency of cross-path access.
[0003] However, since kernel-mode memory allocation is completely unconstrained by user-mode NUMA policies, device drivers, DMA buffers, etc., may still allocate memory on any NUMA node, generating uncontrollable cross-path traffic. Furthermore, the ARM CMN (CoherentMesh Network) network and CXL (Compute Express Link) controller are licensed IPs, and their internal logic cannot be directly modified and promoted, which reduces the efficiency of memory access. Summary of the Invention
[0004] The main objective of this application is to provide a memory access method, apparatus, device, storage medium, and product, which aims to solve the technical problem of reduced memory access efficiency.
[0005] To achieve the above objectives, this application proposes a memory access method, the memory access method comprising: If a memory access request is received from the CMN network and the target physical address corresponding to the memory access request is not within the range of local memory addresses, the memory access request is forwarded to the target module. If the target module is the access analysis module, determine whether the statistical region where the target physical address is located is a hotspot region; If the statistical area is a hotspot area, memory access is performed through the hotspot caching module.
[0006] In one embodiment, the step of accessing memory through the hotspot caching module if the statistical region is the hotspot region includes: If the statistical region is the hotspot region, the memory access request is decomposed to obtain target parameters, which include at least group index and row label; The memory access request is sent to the hot spot cache module, and based on the group index and the line label, it is determined whether the cache line in the target cache group is hit in the cache group of the hot spot cache module, and a second determination result is obtained. Based on the second judgment result, memory access is performed through the hotspot caching module.
[0007] In one embodiment, the target parameter includes an inline offset, and the step of accessing memory through the hotspot caching module based on the second determination result includes: If a match is found, the cache line in the target cache group is accessed, the first data corresponding to the offset in the line is obtained, the first data is returned to the CMN network, and the first age position of the target cache group is updated. If a cache miss occurs, a replacement cache group is selected based on the second age bit of each cache group, and the second data in the replacement cache group is accessed and returned to the CMN network.
[0008] In one embodiment, the step of accessing the second data in the replacement cache group includes: Determine the dirty status of the replacement cache group; Based on the dirty status, process the third data in the replacement cache group to obtain the target replacement cache group; The memory access request is forwarded to the CXL controller module, and the second data returned by the CXL controller module based on the memory access request is stored in the replacement cache group, and the second data in the replacement cache group is accessed.
[0009] In one embodiment, the step of processing the third data in the replacement cache group based on the dirty bit state to obtain the target replacement cache group includes: If the dirty bit status is valid, the third data in the replacement cache group is written back to the original memory address through the CXL controller module to obtain the target replacement cache group.
[0010] In one embodiment, the step of sending the memory access request to the hotspot cache module includes: Update the historical address queue based on the target physical address; Based on the access addresses in the historical address queue, calculate the first step length and obtain the second step length before the historical address queue is updated; Based on the first step length and the second step length, the prefetching mode is determined; Based on the prefetch mode, the prefetch depth is determined, and a prefetch request is generated based on the prefetch depth and the target physical address. The prefetch request is forwarded to the CXL controller module, and the fifth data returned by the CXL controller module based on the prefetch request is stored in the cache group of the hot spot cache module.
[0011] Furthermore, to achieve the above objectives, this application also proposes a memory access device, the memory access device comprising: The forwarding module is used to forward the memory access request to the target module when it receives a memory access request sent by the CMN network and the target physical address corresponding to the memory access request is not within the range of local memory addresses. The judgment module is used to determine whether the statistical region where the target physical address is located is a hotspot region if the target module is the access analysis module. The access module is used to access memory through the hotspot caching module if the statistical area is the hotspot area.
[0012] In addition, to achieve the above objectives, this application also proposes a memory access device, the device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the memory access method as described above.
[0013] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the memory access method described above.
[0014] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the memory access method described above.
[0015] One or more technical solutions proposed in this application have at least the following technical effects: In contrast to related technologies, kernel-mode memory allocation is completely unconstrained by user-mode NUMA policies. Device drivers, DMA buffers, etc., can still allocate memory on any NUMA node, generating uncontrollable cross-path traffic. Furthermore, the ARM CMN (Coherent Mesh Network) and CXL (Compute Express Link) controllers are licensed IPs, making it impossible to directly modify their internal logic and promote them, thus reducing memory access efficiency. This application addresses this issue by applying an access optimization module, which includes at least an access analysis module and a hotspot caching module. Because the access optimization module is independent of the ARM CMN (Coherent Mesh Network) and CXL (Compute Express Link) controllers, this application addresses the issue by applying an access optimization module. The CMN controller is used for memory access requests. When this application receives a memory access request from the CMN network and the target physical address of the memory access request is not within the local memory address range, it will forward the memory access request to the target module. When the target module is an access analysis module and the statistical area where the target physical address is located is a hotspot area, memory access is directly performed on the hotspot cache module without the need for software modifications. The optimization module is placed on the necessary path between the CMN and CXL controllers. Regardless of whether the cross-path request comes from a user-space application or a kernel-space driver, it is intercepted at the hardware level, which improves the efficiency of memory access. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating an embodiment of the memory access method of this application. Figure 2 This is the overall flowchart of the memory access method in this application; Figure 3 This is a flowchart illustrating Embodiment 2 of the memory access method of this application; Figure 4 This is a schematic diagram of the module structure of the memory access device according to an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the memory access method in the embodiments of this application.
[0019] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0020] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0021] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0022] The main solution of this application embodiment is: when a memory access request is received from the CMN network and the target physical address corresponding to the memory access request does not belong to the local memory address range, the memory access request is forwarded to the target module; if the target module is the access analysis module, it is determined whether the statistical area where the target physical address is located is a hotspot area; if the statistical area is the hotspot area, memory access is performed through the hotspot caching module.
[0023] Since kernel-mode memory allocation is completely unconstrained by user-mode NUMA policies, device drivers, DMA buffers, etc., may still allocate memory on any NUMA node, generating uncontrollable cross-path traffic. Furthermore, the ARM CMN (Coherent Mesh Network) network and CXL (Compute Express Link) controller are licensed IPs, making it impossible to directly modify their internal logic and promote them, thus reducing the efficiency of memory access.
[0024] When this application receives a memory access request sent by the CMN network, and the target physical address corresponding to the memory access request is not within the local memory address range, it forwards the memory access request to the target module. When the target module is an access analysis module, and whether the statistical area where the target physical address is located is a hotspot area, it directly accesses the memory of the hotspot cache module without the need for software modifications. The optimization module is placed on the necessary path between the CMN and the CXL controller. Regardless of whether the cross-path request comes from a user-space application or a kernel-space driver, it is intercepted at the hardware level, thus improving the efficiency of memory access.
[0025] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or access optimization module capable of performing the above functions. The following description uses an access optimization module as an example to illustrate this embodiment and the subsequent embodiments.
[0026] Based on this, embodiments of this application provide a memory access method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the memory access method of this application.
[0027] In this embodiment, the memory access method includes steps S10 to S30: Step S10: When a memory access request is received from the CMN network, and the target physical address corresponding to the memory access request does not belong to the local memory address range, the memory access request is forwarded to the target module. It should be noted that the execution entity in this embodiment is the access optimization module, which includes an access analysis module and a hotspot caching module. CMN network refers to the Coherent Mesh Network in the ARM architecture, an on-chip network structure connecting the CPU core, various levels of cache, memory controller, and internal interconnect components, responsible for efficiently transmitting data and maintaining consistency information within the chip. A memory access request refers to a transaction request initiated by the CPU core or device to perform read or write operations on the system's physical memory, containing key information such as the target physical address and operation type (read / write). The target physical address is the specific address value in the system's physical address space pointed to by the memory access request, used to locate the exact location of the data to be accessed in physical memory. The local memory address range refers to the address range corresponding to the physical memory on the same NUMA node (i.e., the same socket slot) as the CPU issuing the current memory access request. For the CPU on that node, accessing addresses within this range is considered local access, with low latency. The target module refers to the overall top-level entity of the access optimization module (i.e., numa_optimizer_top), which integrates sub-modules such as access pattern analyzer, hot spot cache, intelligent prefetch engine and bypass control logic.
[0028] Understandably, when the access optimization module receives a memory access request sent from the CMN network through its standardized interface connected to the CMN network, it first parses the target physical address carried in the request. Then, it precisely compares the target physical address with the local memory address range managed by the current CPU socket. If the comparison result determines that the target physical address does not belong to the local memory address range, it confirms that the request is a remote memory access across NUMA nodes. At this time, the NUMA optimization module forwards the memory access request to its internal subsequent processing pipeline, namely the target module (top-level optimization entity), for the next step of processing, instead of directly allowing it to be sent to the CXL controller for remote memory.
[0029] Because this application intercepts all memory access requests transmitted via the CMN network in real time at the hardware level, and accurately filters out cross-path requests whose target physical address does not belong to the current local memory address range through the address comparison mechanism, and forwards the request to a specially designed target module for unified processing, all accesses that would otherwise generate high-latency cross-path traffic are centrally intercepted and controlled before entering the CXL link.
[0030] Step S20: If the target module is the access analysis module, determine whether the statistical area where the target physical address is located is a hotspot area; It should be noted that the access analysis module is responsible for real-time monitoring of passing memory requests, statistically analyzing access frequency at a fixed granularity, and identifying hotspots. Statistical regions refer to several fixed-size statistical units within the access analysis module, divided into physical address spaces using address hashing units. These units can be granular at 4KB (aligned with the operating system page size), and each region has an independent access counter to accumulate the number of accesses within its covered address range. Hotspot regions refer to specific statistical regions whose access counts reach or exceed a preset hardware threshold within a statistical period, indicating that this region is experiencing high-frequency cross-path memory access. When a memory access request is confirmed and forwarded to the access analysis module, the access optimization module first extracts the target physical address from the request, maps the target physical address to the corresponding statistical region using its internally integrated address hashing unit, and identifies whether the current statistical region is a high-frequency access hotspot region.
[0031] Step S30: If the statistical area is the hotspot area, memory access is performed through the hotspot caching module.
[0032] As can be understood, the hotspot cache module refers to the hardware entity of the hotspot cache within the optimization module. It adopts a 16-way set-associative structure, has a configurable capacity (default 8MB), and a cache line size of 64 bytes. It is used to temporarily store copies of memory data identified as hotspot areas and has the capability to maintain cache consistency and manage write-back policies. When the access analysis module determines that the statistical region where the target physical address is located is a hotspot region, the memory access request is further forwarded to the hotspot cache module.
[0033] In one feasible implementation, step S20 includes: If the statistical region is the hotspot region, the memory access request is decomposed to obtain target parameters, which include at least group index and row label; It's important to note that the set index refers to the middle bit of the target physical address, used to locate the specific cache set to which the target data belongs in the set-associative cache. The tag refers to the high-order bit of the target physical address, used to locate the cache set found through the set index. After a memory access request is confirmed to have a target physical address belonging to a hotspot region, the hotspot cache module does not directly access the SRAM memory. Instead, it first directs the request to the internal address resolution unit. Upon receiving the target physical address in the request, this unit, according to the cache module's preset organizational structure parameters (i.e., 16-way set-associative, each set containing several cache lines), divides the complete physical address from low to high order into multiple independent fields, including the block offset, set index, and tag.
[0034] The memory access request is sent to the hot spot cache module, and based on the group index and the line label, it is determined whether the cache line in the target cache group is hit in the cache group of the hot spot cache module, and a second determination result is obtained. As can be understood, a cache group refers to a collection of storage units within the hotspot cache module, organized according to a group-associative structure. Each cache group contains a fixed number of cache paths (which can be set to 16 paths). A cache line refers to the smallest unit of data storage in the hotspot cache module, with a size of 64 bytes, aligned with the CPU cache line. Each cache path stores one cache line. After address decomposition is completed and the group index and line label are obtained, the memory access request is sent to the lookup pipeline of the hotspot cache module. The hotspot cache module first uses the group index field to locate the corresponding cache group in its internal storage array. Then, the hotspot cache module compares the line label with the labels stored in all 16 cache paths within the cache group in parallel, while checking whether the valid bit of each path is valid. If the line label matches and the valid bit is valid in a certain cache path, it is determined as a hit. If none of the 16 cache paths meet the above conditions, it is determined as a miss. This determination result is output as the second determination result.
[0035] Based on the second judgment result, memory access is performed through the hotspot caching module.
[0036] It should be noted that after obtaining the second judgment result, the hot spot caching module divides the processing flow into two parallel hardware execution paths based on the judgment value, reads the data, and simultaneously returns the data to the request initiator to complete this memory access.
[0037] In one feasible implementation, the target parameter includes an inline offset, and the step of accessing memory through the hotspot caching module based on the second determination result includes: If a match is found, the cache line in the target cache group is accessed, the first data corresponding to the offset in the line is obtained, the first data is returned to the CMN network, and the first age position of the target cache group is updated. Understandably, the in-line offset refers to the low-order field in the target physical address, used to specify the starting byte position of the data to be read within the located 64-byte cache line, determining which segment of continuous data to extract from the cache line data body. The age bit refers to the age status bit (or bit group) maintained by the hotspot cache module for each cache path within each cache group, used to implement the tree-based pseudo-LRU replacement strategy. When the second judgment result is a hit, the hot spot cache module starts the data reading pipeline. First, it locates the target cache group through the group index field and locks the specific cache path (Way) within the group where the row label matches and the valid bit is valid. Then, it extracts the inline offset field from the target physical address and reads the required length of first data from the 64-byte cache row data body stored in the cache path according to the position indicated by the inline offset. After the data reading is completed, the hot spot cache module sends the first data to the CMN network through the return path and sends it back to the CPU core that issued the request via the CMN. At the same time, the module performs a cache group status update operation, refreshes the first age bit corresponding to the currently hit path in the target cache group, and marks it as recently accessed (for example, updating the age bit along the path in a tree pseudo-LRU).
[0038] If a cache miss occurs, a replacement cache group is selected based on the second age bit of each cache group, and the second data in the replacement cache group is accessed and returned to the CMN network.
[0039] It should be noted that the replacement cache group refers to the cache group selected by the hot cache module from multiple candidate cache groups based on the second age comparison result after a cache miss occurs, which is used to take over the new data loaded from remote memory. When the second judgment result is a miss, the hot spot cache module does not return data to the requester. Instead, it first initiates the replacement target selection process, reads the second age bit maintained by each path in the current target cache group, and selects the cache path with the oldest age (i.e., the one that has not been accessed for the longest time) from all 16 cache paths in the cache group through hardware comparison logic such as tree pseudo-LRU. The storage location corresponding to this path is the replacement cache group (i.e., the selected replacement target path). After selecting the replacement cache group, the CXL controller initiates a read request to the remote physical memory to obtain the complete 64-byte cache line data (i.e., the second data) corresponding to the target physical address. After the remote data is returned, the module performs two parallel operations. On the one hand, it writes the second data to the selected replacement cache group (updating the data storage body, line tag, and valid bit of this path). On the other hand, it returns the second data to the CMN network through the response channel, and the CMN finally delivers it to the requester to complete this memory access transaction.
[0040] In one feasible implementation, the step of accessing the second data in the replacement cache group includes: Determine the dirty status of the replacement cache group; As can be understood, the dirty status refers to one of the status flags maintained by the hot spot cache module for each cache line. It is used to indicate whether the data content of the cache line has been modified (i.e., a write operation has occurred) after being written to the local cache, but has not yet been synchronously written back to the remote physical memory. The hot spot cache module first reads the dirty status flag corresponding to the replacement cache group, obtains the current value of the flag bit through hardware comparison logic, and determines its dirty status.
[0041] Based on the dirty status, process the third data in the replacement cache group to obtain the target replacement cache group; It should be noted that the hot spot cache module performs differentiated hardware processing on the third data based on the binary judgment result of the dirty bit status. After processing, the replacement cache group is transformed into the target replacement cache group where the data has been safely cleaned, ensuring that the second data to be written can safely overwrite the old data without causing any loss of modified data.
[0042] The memory access request is forwarded to the CXL controller module, and the second data returned by the CXL controller module based on the memory access request is stored in the replacement cache group, and the second data in the replacement cache group is accessed.
[0043] It is understandable that the CXL controller module refers to the interface hardware entity between the optimization module and the CXL interconnect link. It is responsible for sending the remote memory access request processed by the optimization module to the CXL protocol layer, and interacting with the memory controller of the remote socket through the CXL link. Finally, it sends the data returned from the remote end back to the access optimization module. After completing the dirty data processing of the replacement cache group and obtaining the target replacement cache group, the hot spot cache module forwards the current memory access request to the CXL controller module. After receiving the request, the CXL controller module initiates a read transaction to the memory controller of the remote socket slot through the CXL interconnect protocol to obtain the complete 64-byte cache line data (i.e., the second data) corresponding to the target physical address. After the remote data is returned to the CXL controller module, it is passed to the hot spot cache module. The hot spot cache module then performs two operations. On the one hand, it writes the second data into the prepared target replacement cache group (i.e., the original replacement cache group), updates the valid bit and row label of the cache path, and sets the dirty bit status to "clean" (because the data is consistent with the remote memory). On the other hand, it directly accesses the second data in the replacement cache group after it has been written, and returns the data to the CMN network through the response path, and finally delivers it to the request initiator.
[0044] In one feasible implementation, the step of processing the third data in the replacement cache group based on the dirty bit state to obtain the target replacement cache group includes: If the dirty bit status is valid, the third data in the replacement cache group is written back to the original memory address through the CXL controller module to obtain the target replacement cache group.
[0045] It should be noted that a dirty status means that the dirty flag corresponding to the replacement cache group is set to "1" (i.e., dirty status), indicating that the third data stored in the cache line has been modified by a local write operation since it was last loaded from the remote physical memory, and the modified content has not yet been written back to the remote physical memory. After the hot spot cache module determines that the dirty bit status of the selected replacement cache group is valid (i.e., the dirty bit is "1"), it immediately initiates the dirty data write-back process. The hot spot cache module reads the complete 64 bytes of third data from the replacement cache group and encapsulates the third data together with its corresponding original memory address into a remote memory write request. This request is sent to the CXL interconnect link through the CXL controller module and finally received by the memory controller of the remote socket slot and written to the original address location in physical memory. After the CXL controller module confirms that the remote write transaction is successful (i.e., the third data has been reliably written to the remote physical memory), the original third data in the replacement cache group is safely cleaned up. The copy in the remote memory and the old copy in the cache are now in a consistent state. At this point, the replacement cache group becomes a safe and rewritable available state, which is the target replacement cache group, used to receive the second data newly loaded from the remote memory.
[0046] Specifically, refer to Figure 2 , Figure 2 A flowchart is provided, illustrating the complete execution path of the access optimization module in handling CMN memory requests. When a request arrives, the system first determines if the module is enabled. If disabled, it is directly bypassed and forwarded to CXL; if enabled, it enters the access pattern analyzer for address hashing and hotspot detection. If a non-hotspot region is identified, the request is directly forwarded to CXL; if a hotspot is found, it attempts to read data from the hotspot cache (with a delay of 1-3ns). While the cache is being searched, the intelligent prefetch engine works in parallel, detecting whether the current access follows a fixed-step or sequential pattern. Once a match is found, subsequent data is prefetched in the background and written to the cache. Finally, after all operations are completed, the relevant statistics registers are updated.
[0047] In this embodiment, when a cache miss occurs, the CXL controller module retrieves complete second data from remote memory and simultaneously completes the writing to the target replacement cache group and the data return to the requester. This ensures that although the current request undergoes a remote access (approximately 100-150ns, which is better than the original 300ns due to merging and prefetching optimization), the data written is already residing in the local hot spot cache at the end of this transaction. Thus, when subsequent access requests for the same address or the same cache line arrive, they can be directly hit from the on-chip cache (approximately 5-10ns) without needing to repeatedly access remote memory through the CXL controller module.
[0048] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 Before the step of sending the memory access request to the hotspot cache module, the memory access method further includes steps S01 to S05: Step S01: Update the historical address queue based on the target physical address; It should be noted that the access optimization module also has a smart prefetch engine. This engine internally maintains a hardware register queue with a depth of 2 (i.e., a capacity of 2) to store the historical address queue of the two most recent memory access requests. Simultaneously or after a memory access request completes cache lookup (whether it's a cache hit or miss) and data return processing, the hotspot caching module passes the target physical address carried by the current request to the smart prefetch engine. Upon receiving this address, the smart prefetch engine sends it to its internal historical address queue for an update operation. The update rule follows a first-in-first-out (FIFO) shift mechanism. The original second historical address (i.e., the penultimate access address) in the queue is removed and discarded, while the original first historical address (i.e., the most recent access address) is moved to the second address position. The current target physical address is stored as the latest access address in the first address position of the queue. After the update, the historical address queue always contains the current access address and the immediately preceding access address.
[0049] Step S02: Calculate the first step length based on the access addresses in the historical address queue, and obtain the second step length before the historical address queue is updated; Understandably, the first step length refers to the difference between two adjacent access addresses in the historical address queue after the update (i.e., the current access address and the previous access address). The second step length refers to the difference between two existing adjacent addresses in the queue (i.e., the previous access address and the access address two years ago) before the current historical address queue update. After the intelligent prefetch engine completes the update of the historical address queue (i.e., moving the current target physical address into the queue), the engine immediately starts the step length calculation pipeline. The engine uses an internal hardware subtractor to extract the current access address and the previous access address from the updated historical address queue, subtracts the previous access address from the current access address, and calculates the first step length in real time. At the same time, the engine reads the second step length from the internal step length latch register. This second step length is the previous step length value (i.e., the difference between the previous access and the access address two years ago) that was calculated and saved before the current queue update.
[0050] Step S03: Determine the prefetching mode based on the first step length and the second step length; It should be noted that after obtaining the first step length and the second step length, the intelligent prefetch engine immediately sends these two values to the internal pattern detector for comparison and judgment in order to identify the prefetch pattern.
[0051] Step S04: Based on the prefetch mode, determine the prefetch depth, and based on the prefetch depth and the target physical address, generate a prefetch request; Understandably, prefetch depth refers to the number of cache lines that the intelligent prefetch engine decides to read from the CXL controller before the current target physical address after detecting a regular access pattern. The value ranges from 1 to 4 cache lines. After the pattern detector determines the prefetch pattern, the intelligent prefetch engine enters the prefetch strategy decision-making stage. The engine first determines the prefetch depth value for this prefetch operation based on the determined prefetch pattern type and its internal adaptive adjustment logic.
[0052] By capturing access patterns through an intelligent prefetch engine, subsequent data is loaded from remote memory into the local hotspot cache before the CPU core actually issues the next read request. When the core actually needs this data, the request is returned directly from the cache, completely avoiding cross-path access latency (reducing the hit path latency from 300ns to 5-10ns). It is also highly effective for x265 encoding scenarios with sequential access intensity because the sequential access step size of x265 is completely fixed. The prefetch engine will hardly generate invalid prefetches (i.e., prefetched but unused data, so the cache hit rate can reach more than 85%).
[0053] Step S05: Forward the prefetch request to the CXL controller module, and store the fifth data returned by the CXL controller module based on the prefetch request into the cache group of the hot spot cache module.
[0054] It should be noted that after the intelligent prefetch engine generates a prefetch request, it forwards the request to the CXL controller module. Upon receiving the prefetch request, the CXL controller module initiates a background read transaction with the memory controller of the remote socket via the CXL interconnect protocol to obtain the complete 64-byte cache line data (i.e., the fifth data) corresponding to the prefetch target physical address. After the remote data is returned to the CXL controller module, it is then passed to the hot spot cache module. After receiving the fifth data, the hot spot cache module obtains the group index and row label by address decomposition based on the target physical address carried in the prefetch request. It uses the group index to locate the corresponding cache group and selects a cache path within that cache group according to the replacement strategy (such as pseudo-LRU). The fifth data is written to that cache path, and the row label, valid bit (set to valid), and dirty bit status (set to "clean") of that path are updated simultaneously.
[0055] In one feasible implementation, step S03 includes: If the first step length and the second step length are equal, and the first step length is not equal to a preset value, then it is determined that the fixed step length prefetching mode is in effect. It is understandable that the preset value refers to a fixed reference value pre-set within the intelligent prefetch engine, which can be set to 0. Fixed-step prefetch mode refers to an access pattern type identified by the intelligent prefetch engine where the step size of the current memory access sequence has a constant non-zero value and is not equal to 64 bytes. This is common in scenarios where memory is read at fixed element intervals (such as array column access, traversal of specific structure members, or non-contiguous row span access in a matrix). After obtaining the first and second step sizes, the pattern detector of the intelligent prefetch engine executes precise conditional judgment logic. The detector first compares the first and second step sizes for equality to confirm whether the two step sizes are completely consistent. If they are consistent, it further determines whether the value of the first step size is not equal to the preset value (i.e., 64 bytes). If the two step sizes are equal and the first step size is not equal to 64 bytes, the engine confirms that the current access sequence belongs to an access behavior that is not 64 bytes but has a constant interval pattern. At this time, the internal pattern status is determined to be fixed-step prefetch mode.
[0056] When the fixed step size prefetch mode is in effect and the first step size is equal to the preset byte length, it is determined that the sequential prefetch mode is in effect.
[0057] It should be noted that sequential prefetching mode refers to a special case pattern further identified by the intelligent prefetching engine within the framework of the generalized fixed-step prefetching mode. Its core characteristic is that the first step length is exactly equal to the preset byte length (i.e., 64 bytes, the size of a standard cache line), indicating that the program is traversing adjacent address spaces sequentially, either increasing or decreasing, in cache line units. Once the intelligent prefetching engine has determined that it is in fixed-step prefetching mode, the engine's pattern detector does not stop there. Instead, it further compares the currently acquired first step length value with the preset byte length (i.e., 64 bytes) for precise equality. If the comparison result is equal, meaning the first step length is exactly 64 bytes, it indicates that the fixed step length of the current access sequence is equal to the size of a standard cache line, and the program is traversing the address space sequentially, in cache line units. At this point, the engine further refines and updates its internal status flag from the generalized fixed-step prefetching mode to sequential prefetching mode.
[0058] In one feasible implementation, the step of determining whether the statistical region where the target physical address is located is a hotspot region includes: Map the target physical address to a statistical region, and update the access count and third age bit of the statistical region; Understandably, while memory access requests complete cache lookup and prefetching, the access analysis module independently executes the statistical information update process. First, it extracts the high-order bits of the target physical address and maps them to one of the 256 pre-divided statistical regions (each region covers a 4KB granular address space) through the address hash unit. After locating the corresponding statistical region, the module performs two parallel update operations: first, it increments the 32-bit access counter corresponding to the region by 1 (distinguishing between read and write accesses); second, it refreshes the third age bit of the region, marking the statistical region as "recently accessed". For example, in the LRU age bit array, the age value of the region is reset to the latest state, while the age bits of other regions are incremented accordingly to reflect the relative decrease in their access freshness.
[0059] Because this application updates the access counter and the third age bit in parallel within the same hardware cycle, the access analysis module can continuously track the access frequency and timing freshness of each 4KB address region with periodic precision, ensuring that the least used cold regions can be prioritized for elimination when the statistical region resources need to be reallocated.
[0060] If the updated access count is greater than or equal to the preset hotspot threshold, then the statistical area is determined to be a hotspot area.
[0061] It should be noted that the preset hotspot threshold refers to the access count threshold configured internally by the access analysis module through the control register (HOTSPOT_THRESHOLD) to determine whether a statistical region is a hotspot region. The value ranges from 4 to 32 times (the default value is 8). After updating the access count and refreshing the third age digit of the statistical region corresponding to the target physical address, the access analysis module compares the updated access count with the preset hotspot threshold pre-configured internally (set through the HOTSPOT_THRESHOLD register, with a value range of 4 to 32 times and a default value of 8). If the updated access count is greater than or equal to the preset hotspot threshold, the current statistical region is determined to be a hotspot region. Conversely, if the access count has not yet reached the threshold, the statistical region is not considered a hotspot region and awaits further accumulated access counts.
[0062] In one feasible implementation, the target module further includes a CXL controller module, and the step of forwarding the memory access request to the target module includes: Determine the current working mode; Understandably, the access optimization module also has bypass control logic. When the access optimization module receives any memory access request (whether it is a cross-path request from the CMN network or an internally generated prefetch request) and enters the front end of the processing pipeline, the bypass control logic first performs a working mode determination operation. The bypass control logic reads the enable bit and abnormal status bit in the internal status register (MODULE_STATUS) and combines it with the current value of the forced bypass register (BYPASS_FORCE) to comprehensively determine whether the module should be in normal mode or bypass mode.
[0063] If the operating mode is normal mode, then the memory access request is forwarded to the access analysis module; It should be noted that when the bypass control logic determines the current operating mode as normal, the NUMA optimization module forwards the currently received memory access requests (including cross-path memory access requests from the CMN network and prefetch requests generated by the internal intelligent prefetch engine) to the access analysis module through the internal data path.
[0064] If the operating mode is bypass mode or exception protection mode, the memory access request is forwarded to the CXL controller module.
[0065] It is understandable that bypass mode refers to a working state where the access optimization module is set to full pass-through. In this mode, all memory access requests are forwarded directly from input to output without any optimization processing. Anomaly protection mode refers to a protective working state automatically triggered by the bypass control logic when a timeout, state machine freeze, or other abnormal situation occurs within the access optimization module. Its behavior is the same as bypass mode (all requests are directly pass-through). When the bypass control logic determines the current working mode as either bypass mode or anomaly protection mode, the module does not forward memory access requests to any optimization pipeline stages such as the access analysis module or hotspot cache module. Instead, it directly forwards the memory access requests to the CXL controller module through an internal pure combinational logic multiplexer (MUX). This forwarding path bypasses the access analysis module, hotspot cache module, and intelligent prefetch engine, effectively removing the NUMA optimization module from the data path at the hardware level, forming a direct connection between the CMN network and the CXL controller. In anomaly protection mode, the bypass control logic also sets the exception flag in the STATUS register to record the abnormal event while forwarding the request. After this forwarding, the request reaches the CXL controller module directly, and the CXL controller sends it to the memory controller of the remote socket for normal processing via the CXL interconnect protocol.
[0066] In this embodiment, after the normal mode is determined, the memory access request is forwarded to the access analysis module with extremely low latency through a pure combinational logic multiplexer without software intervention or complex protocol conversion. This ensures that the request does not incur any additional time overhead when entering the optimization pipeline, and that even when optimization is enabled, the forwarding latency of the normal mode is still much lower than the memory latency of cross-path access itself (300ns), and the basic latency of request processing will not be significantly increased by introducing the optimization module.
[0067] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the memory access method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0068] This application also provides a memory access device, please refer to... Figure 4 The memory access device includes: The forwarding module 10 is used to forward the memory access request to the target module when it receives a memory access request sent by the CMN network and the target physical address corresponding to the memory access request is not within the range of local memory addresses. The judgment module 20 is used to determine whether the statistical area where the target physical address is located is a hotspot area if the target module is the access analysis module; Access module 30 is used to access memory through the hotspot cache module if the statistical area is the hotspot area.
[0069] The memory access device provided in this application, employing the memory access method in the above embodiments, can solve the technical problem of memory access. Compared with the prior art, the beneficial effects of the memory access device provided in this application are the same as those of the memory access method provided in the above embodiments, and other technical features in the memory access device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0070] This application provides a memory access device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the memory access method in Embodiment 1 above.
[0071] The following is for reference. Figure 5The diagram illustrates a structural schematic of a memory access device suitable for implementing embodiments of this application. The memory access device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, tablets, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The memory access device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0072] like Figure 5 As shown, the memory access device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the memory access device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the memory access device to communicate wirelessly or wiredly with other devices to exchange data. Although memory access devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0073] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0074] The memory access device provided in this application, employing the memory access method in the above embodiments, can solve the technical problem of memory access. Compared with the prior art, the beneficial effects of the memory access device provided in this application are the same as those of the memory access method provided in the above embodiments, and other technical features in this memory access device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0075] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0076] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0077] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the memory access method in the above embodiments.
[0078] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0079] The aforementioned computer-readable storage medium may be included in a memory access device or may exist independently without being assembled into a memory access device.
[0080] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by a memory access device, the memory access device: upon receiving a memory access request from the CMN network, and if the target physical address corresponding to the memory access request does not fall within the local memory address range, forwards the memory access request to a target module; if the target module is the access analysis module, determines whether the statistical region where the target physical address is located is a hotspot region; if the statistical region is a hotspot region, performs memory access through the hotspot caching module.
[0081] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0082] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0083] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0084] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described memory access method, thereby solving the technical problem of memory access. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the memory access method provided in the above embodiments, and will not be repeated here.
[0085] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the memory access method described above.
[0086] The computer program product provided in this application can solve the technical problem of memory access. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the memory access method provided in the above embodiments, and will not be repeated here.
[0087] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.
Claims
1. A memory access method, characterized in that, The memory access method is applied to an access optimization module, which includes at least an access analysis module and a hotspot caching module, and includes: If a memory access request is received from the CMN network and the target physical address corresponding to the memory access request is not within the range of local memory addresses, the memory access request is forwarded to the target module. If the target module is the access analysis module, determine whether the statistical region where the target physical address is located is a hotspot region; If the statistical area is a hotspot area, memory access is performed through the hotspot caching module.
2. The memory access method as described in claim 1, characterized in that, If the statistical region is the hotspot region, the step of accessing memory through the hotspot caching module includes: If the statistical region is the hotspot region, the memory access request is decomposed to obtain target parameters, which include at least group index and row label; The memory access request is sent to the hot spot cache module, and based on the group index and the line label, it is determined whether the cache line in the target cache group is hit in the cache group of the hot spot cache module, and a second determination result is obtained. Based on the second judgment result, memory access is performed through the hotspot caching module.
3. The memory access method as described in claim 2, characterized in that, The target parameter includes the inline offset, and the step of accessing memory through the hotspot caching module based on the second judgment result includes: If a match is found, the cache line in the target cache group is accessed, the first data corresponding to the offset in the line is obtained, the first data is returned to the CMN network, and the first age position of the target cache group is updated. If a cache miss occurs, a replacement cache group is selected based on the second age bit of each cache group, and the second data in the replacement cache group is accessed and returned to the CMN network.
4. The memory access method as described in claim 3, characterized in that, The step of accessing and replacing the second data in the cache group includes: Determine the dirty status of the replacement cache group; Based on the dirty status, process the third data in the replacement cache group to obtain the target replacement cache group; The memory access request is forwarded to the CXL controller module, and the second data returned by the CXL controller module based on the memory access request is stored in the replacement cache group, and the second data in the replacement cache group is accessed.
5. The memory access method as described in claim 4, characterized in that, The step of processing the third data in the replacement cache group based on the dirty status to obtain the target replacement cache group includes: If the dirty bit status is valid, the third data in the replacement cache group is written back to the original memory address through the CXL controller module to obtain the target replacement cache group.
6. The memory access method as described in claim 2, characterized in that, The step of sending the memory access request to the hotspot cache module includes the following: Update the historical address queue based on the target physical address; Based on the access addresses in the historical address queue, calculate the first step length and obtain the second step length before the historical address queue is updated; Based on the first step length and the second step length, the prefetching mode is determined; Based on the prefetch mode, the prefetch depth is determined, and a prefetch request is generated based on the prefetch depth and the target physical address. The prefetch request is forwarded to the CXL controller module, and the fifth data returned by the CXL controller module based on the prefetch request is stored in the cache group of the hot spot cache module.
7. A memory access device, characterized in that, The device includes: The forwarding module is used to forward the memory access request to the target module when it receives a memory access request sent by the CMN network and the target physical address corresponding to the memory access request is not within the range of local memory addresses. The judgment module is used to determine whether the statistical region where the target physical address is located is a hotspot region if the target module is an access analysis module. The access module is used to access memory through the hotspot caching module if the statistical area is the hotspot area.
8. A memory access device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the memory access method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the memory access method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the memory access method as described in any one of claims 1 to 6.