Cache resource allocation method and device
By allocating cache space in the cache according to the data type, the problem of a large number of data requests in the processor competing for the same cache space is solved, and the processing efficiency of the processor is improved.
Patent Information
- Application Number
- CN202311700116.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-08
- Publication Date
- 2025-06-10
AI Technical Summary
A large amount of data requests compete for the same cache space in the processor, resulting in a degradation of processor processing performance.
The data type indicated by the data request is divided into hot data, cold data or address translation data, and the corresponding cache space is allocated in the buffer according to the data type.
The fine-grained management of the cache space in the processor is realized, which avoids the problem of large amounts of data requests competing for the same cache space, and improves the processor's processing efficiency of data requests.
Smart Images

Figure CN120123256A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data storage, and in particular to a cache resource allocation method and device. Background Art
[0002] The third-level cache (L3 cache) is a cache integrated in the processor. The L3 cache can be used to implement high-speed data buffering between the processor and the main memory. Usually, in order to efficiently utilize the cache space of the L3 cache, identity documents (IDs) are often set for different software threads, and corresponding cache space in the L3 cache is configured for different IDs. However, a business flow processed by the same software thread may have a large number of requests, which will compete for the cache space in the L3 cache, resulting in reduced processing performance of the processor. Summary of the invention
[0003] The present application provides a cache resource allocation method and device to solve the problem that a large number of data requests compete for the same cache space in a processor, resulting in reduced processor processing performance.
[0004] In order to achieve the above-mentioned purpose, this application adopts the following technical solution.
[0005] In a first aspect, the present application provides a cache resource allocation method. The cache resource allocation method is executed by a processor, or a physical device that supports the implementation of the cache resource allocation method, for example, the physical device includes a chip system. The cache resource allocation method includes: the processor obtains a data request, and allocates cache space in a cache for the data request based on the data type indicated by the data request. Among them, the data type includes: hot data, cold data or address conversion data, the cache space matches the data type, and the address conversion data is used to indicate the correspondence between the virtual address and the physical address.
[0006] In the present application, the data types indicated by the data request are divided into multiple types, such as hot data, cold data or address conversion data, and cache space is allocated to the data request according to the data type. Since different types of data, such as the hot data, cold data and address conversion data mentioned above, have different management granularities, allocating corresponding cache space to different types of data is conducive to the fine-grained management of the cache space in the processor. This avoids the problem that a large number of requests may compete for the same cache space in the processor under the same business flow, and improves the processor's processing efficiency for data requests.
[0007] In conventional technology, data requests and instruction requests are managed separately, and corresponding cache space is allocated to the data requests and instruction requests; in contrast, the present application refines the data requests, divides the data types indicated by the data requests into hot data, cold data, and address conversion data, and allocates corresponding storage space to the aforementioned data requests indicating different data types, which is conducive to realizing fine-grained management of the cache space in the processor, and further improves the processor's processing efficiency for data requests.
[0008] In a possible implementation, the processor allocates cache space in the cache for the data request, including: the processor determines the cache space corresponding to the data type from an allocation strategy, and then allocates the cache space in the cache for the data request. The allocation strategy includes: a correspondence between the cache space in the cache and the data type.
[0009] In the present application, since the allocation strategy indicates the correspondence between the cache space in the cache and the data type, the processor can accurately determine the cache space corresponding to the data request from the allocation strategy. Then, the cache space in the cache is allocated to the data request, so that data requests indicating different data types are allocated to different cache spaces, avoiding the problem of a large number of data requests competing for the same cache space, and improving the processing efficiency of the processor.
[0010] In a possible example, the cache space in the cache allocated by the processor for the data request is used to store the to-be-accessed data indicated by the data request, where the type of the to-be-accessed data is hot data, cold data, or address conversion data.
[0011] In one possible implementation, the processor determines the cache space corresponding to the data type from the allocation strategy, including: the processor assigns a tag of the data request to a third value, where the third value indicates the data type, and then determines the cache space corresponding to the tag from the allocation strategy.
[0012] In the present application, the processor assigns a value, such as a third value, to the tag in the data request so that the third value of the tag indicates the data type. Then, the processor allocates corresponding cache space to the data request according to the tag and allocation strategy carried by the data request, thereby improving the efficiency of allocating cache space to the data request.
[0013] In a possible example, the processor defines a reserved field in a data request as a tag (such as a hot / coldtype field) and assigns a value to the tag so that the value corresponding to the tag indicates the data type.
[0014] In a possible implementation, the above cache resource allocation method further includes: The processor parses the data request to obtain a virtual address, and the virtual address indicates the access location of the data to be accessed. The processor determines whether the data to be accessed is hot data or cold data from the page table according to the virtual address. If the data to be accessed is hot data, it determines that the data type is hot data. If the data to be accessed is cold data, it determines that the data type is cold data. Among them, the page table includes the relationship between the virtual address and hot data / cold data.
[0015] In this application, since the page table indicates the correspondence between the virtual address and hot data or cold data, the processor determines whether the data to be accessed is hot data or cold data from the page table according to the virtual address in the data request. Furthermore, the processor can accurately determine that the data type is hot data or cold data, improving the accuracy of allocating the corresponding cache space for the data request, avoiding the problem that a large number of data requests compete for the same cache space, and improving the processing efficiency of the processor.
[0016] In a possible implementation, the page table includes a type field, and the type field is a first value or a second value. The first value indicates that the data to be accessed is hot data, and the second value indicates that the data to be accessed is cold data.
[0017] In a possible example, the type field is a hot / cold type field. If the value of the hot / cold type field is 1, it means that the data to be accessed is hot data; if the value of the hot / cold type field is 0, it means that the data to be accessed is cold data.
[0018] In a possible implementation, the above cache resource allocation method further includes: The processor obtains the access information of the page table. If the access information meets the condition, it determines that the page table is a hot page, and the hot page indicates that the data to be accessed is hot data. If the access information does not meet the condition, it determines that the page table is a cold page, and the cold page indicates that the data to be accessed is cold data.
[0019] In this application, the processor determines that the page table is a hot page or a cold page through the access information of the page table. Since the hot page indicates that the data to be accessed is hot data and the cold page indicates that the data to be accessed is cold data. Furthermore, when the processor receives a data request, it can directly determine from the page table (hot page or cold page) that the data request corresponding to the data to be accessed is cold data / hot data, so as to quickly obtain the data type, improving the efficiency of the processor to allocate cache space according to the data type.
[0020] In a possible example, the above access information includes the number of accesses or the access frequency, etc., and the above condition is whether the number of accesses or the access frequency is greater than or equal to a threshold.
[0021] Exemplarily, the processor determines hot pages or cold pages based on the access information of the page table. Further, the processor adjusts the page table attributes of the hot pages / cold pages, that is, the reserved field in the page table is defined as the hot / cold type field and this field is assigned a value, so that different values of the hot / cold type field represent different data types.
[0022] In a second aspect, the present application provides a cache resource allocation device. The cache resource allocation device is applied to a processor or a physical device that supports the implementation of the cache resource allocation method, for example, the physical device includes a chip system. The cache resource allocation device includes various modules for executing the cache resource allocation method in the first aspect or any possible implementation manner of the first aspect. Exemplarily, the cache resource allocation device includes: an acquisition module and an allocation module. Among them, the acquisition module is used to acquire a data request.
[0023] The allocation module is used to allocate cache space for the data request in the cache according to the data type indicated by the data request. Among them, the data type includes: hot data, cold data or address translation data, and the cache space matches the data type; the address translation data is used to indicate the correspondence between the virtual address and the physical address.
[0024] For more detailed implementation content of the cache resource allocation device, reference may be made to the description of any implementation manner in the above first aspect and the content of the following specific implementation manners, which will not be elaborated here.
[0025] In a third aspect, the present application provides a chip. The chip includes: an interface circuit and a control circuit. The interface circuit is used to receive signals from other devices outside the processor and transmit them to the control circuit, or send signals from the control circuit to other devices outside the processor. For example, the interface circuit acquires a data request transmitted by other devices, and the control circuit and the interface circuit cooperate through a logic circuit or execute code instructions to implement the method in the first aspect or any possible implementation manner in the first aspect.
[0026] In a fourth aspect, the present application provides a processor. The processor is used to execute a computer program or instruction to implement the first aspect or any possible implementation manner in the first aspect.
[0027] Fifth aspect, the present application provides a computing device. The computing device includes a memory and a processor. The memory is used to store computer instructions, and the processor is used to call and run the computer instructions from the memory to execute the method in the first aspect or any possible implementation manner of the first aspect. In some optional implementation manners, the processor in the computing device provided in the fifth aspect may include one or more chips provided in the third aspect. These chips are interconnected through a bus and are used to implement the method in the first aspect or any possible implementation manner of the first aspect. For example, the bus may include but is not limited to: a data bus, an address bus, or other types of buses, etc. Alternatively, the computing device provided in the fifth aspect includes the processor in the fourth aspect.
[0028] Sixth aspect, the present application provides a computer-readable storage medium. The storage medium stores a computer program or instructions. When the computer program or instructions are executed by a processing device, the method in the first aspect or any possible implementation manner in the first aspect is implemented.
[0029] Seventh aspect, the present application provides a computer program product. The computer program product includes a computer program or instructions. When the computer program or instructions are executed by a processing device, the method in the first aspect or any possible implementation manner in the first aspect is implemented.
[0030] For the beneficial effects of the above second aspect to seventh aspect, reference may be made to the first aspect or any possible implementation manner in the first aspect, and details are not described herein. Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Schematic diagram for allocating cache resources according to software thread ID;
[0032] Figure 2 Schematic diagram of a data storage system provided by the present application;
[0033] Figure 3 Schematic diagram of the structure of a processor provided by the present application;
[0034] Figure 4 Schematic diagram of the flow of a cache resource allocation method provided by the present application;
[0035] Figure 5 Schematic diagram of the flow of a data type determination method provided by the present application;
[0036] Figure 6 Schematic diagram of the division of a cache provided by the present application;
[0037] Figure 7Schematic flowchart of the method for allocating cache resources in the L3 cache provided by this application;
[0038] Figure 8 Schematic structure of the cache resource allocation device provided by this application Figure 1 ;
[0039] Fig. 9 Schematic structure of the cache resource allocation device provided by this application Figure 2 ;
[0040] Fig.10 Schematic diagram of the structure of a computing device provided by this application. Detailed implementation manners
[0041] For ease of understanding, first, the technical terms involved in this application are introduced.
[0042] The cache is used to implement high-speed data buffering between the processor and the memory. The cache in the processor may include a level 1 cache (L1 cache), a level 2 cache (L2 cache), and an L3 cache. Among them, the read / write speeds of the L1 cache, L2 cache, and L3 cache decrease in turn, while the cache capacity usually gradually increases.
[0043] A processing element (PE) refers to the basic unit in the processor that executes calculation or processing instructions. In a multi-core processor or some processors with specific architectures, a PE can refer to a single core or an execution unit.
[0044] MPAM (memory system resource partitioning and monitoring) is a technology for memory or cache system resource management. MPAM realizes improving the utilization rate and management efficiency of memory or cache through resource partitioning, resource monitoring, QoS (quality of service management), and performance optimization.
[0045] SMMU (system memory management unit) is a hardware component used to manage memory or cache access between devices. The SMMU is usually used for the interaction between the processor and external devices. Especially in a virtualization environment, when multiple devices share the same physical memory, the SMMU can isolate and manage memory access.
[0046] The MSC (memory system component) refers to a chip or module used for data transfer and control between the processor and the cache. The MSC usually consists of multiple sub-modules, such as a memory controller, a cache, a bus interface, etc., for managing and optimizing the cache.
[0047] The device memory controller (DMC) is responsible for coordinating data exchange between the processor and the memory to ensure efficient memory access and data transfer. For example, the DMC is responsible for loading data from the memory into the L3 cache, or writing data back from the L3 cache to the cache.
[0048] The memory management unit (MMU) is used to implement address translation, that is, to convert virtual addresses into physical addresses for accessing data and instructions in the memory. The MMU is mainly responsible for performing the conversion of virtual addresses to physical addresses and sending the converted addresses to the memory controller for access. The MMU is usually implemented by hardware circuits and includes components such as a translation lookaside buffer (TLB) and a page table. The page table includes the mapping relationship between virtual addresses and physical addresses, and a page table includes multiple virtual addresses. The TLB is a cache used to store the mapping relationship between virtual addresses and physical addresses that the processor has recently used.
[0049] With the continuous progress of manufacturing technology, the processor has entered the 3-nanometer manufacturing stage. The processor can accommodate more transistors per unit area, so the cache that can be deployed by the processor is also increasing continuously. How to effectively utilize the cache in the processor has become an urgent problem to be solved.
[0050] For the above problems, two possible solution examples are provided below.
[0051] Example 1: MPAM sets tags / identifiers (such as PARTID or PMG) for different software threads, and then configures different cache resources for different IDs.
[0052] As Figure 1 shown, Figure 1 is a schematic diagram of allocating cache resources according to the software thread ID. The L3 cache is N-way set associative. The user tags different threads and controls the logic. For example, thread 1 corresponds to PARTID 1, and the L3 cache resources available for requests corresponding to PARTID 1 are 0 - M way (M < N). Thread 2 corresponds to PARTID 2, and the L3 cache resources available for requests corresponding to PARTID 2 are M - N way.
[0053] When thread 1 runs on PE1, the memory access request A corresponding to thread 1 is tagged with corresponding tag information through MPAM, such as PARTID being 1. The memory access request A needs to apply for cache space on the L3 cache. MPAM determines, from the above control logic, to allocate the 0-Mway in the L3 cache resources when PARTID is 1 by comparing the tag information carried by the memory access request A, and then applies for the corresponding cache space for the memory access request A on the 0-M way.
[0054] When thread 2 runs on PE2, the memory access request B corresponding to thread 2 is tagged with corresponding tag information through MPAM, such as PARTID being 2. The memory access request B needs to apply for cache space on the L3 cache. MPAM determines, from the above control logic, to allocate the M-Nway in the L3 cache resources when PARTID is 2 by comparing the tag information carried by the memory access request B, and then applies for the corresponding cache space for the memory access request B on the M-N way.
[0055] Example 2: MPAM tags different PARTIDs for data requests / instruction requests through a processor (such as a CPU (central processing unit)) / SMMU, and then configures different cache resources for different PARTIDs.
[0056] Compared with Example 1, the granularity in Example 2 is changed from thread / business flow to data request / instruction request. The overall process of Example 2 is the same as that of Example 1 and will not be elaborated here.
[0057] However, both the allocation of cache resources with the business flow as the granularity in Example 1 and the allocation of cache resources with the data request / instruction request as the granularity in Example 2 have the problem of relatively coarse granularity, which may result in a large number of requests competing for the same cache space, and the processor cannot obtain data or instructions in time, leading to a reduction in the processing performance of the processor.
[0058] Based on this, the present application provides a cache resource allocation method, which can be applied to a computing device or a chip (such as a processor) in a computing device. The cache resource allocation method includes: obtaining a data request, and allocating cache space for the data request in a cache according to the data type indicated by the data request. Wherein, the data type includes: hot data, cold data or address translation data, the cache space matches the data type, and the address translation data is used to indicate the correspondence between a virtual address and a physical address.
[0059] In the present application, the data types indicated by data requests are classified into multiple types, such as hot data, cold data, or address translation data, and cache spaces are allocated for the data requests according to the data types. Since different types of data, such as the above-mentioned hot data, cold data, and address translation data, have different management granularities, allocating respective corresponding cache spaces for different types of data is conducive to achieving fine-grained management of the cache space in the processor. It avoids the problem that there may be a large number of requests competing for the same cache space in the processor under the same service flow, and improves the processing efficiency of the processor for data requests.
[0060] In the conventional technology, data requests and instruction requests are managed separately, and corresponding cache spaces are allocated for data requests and instruction requests. In contrast, the present application refines data requests, classifies the data types indicated by data requests into hot data, cold data, and address translation data, and allocates corresponding storage spaces for data requests indicating different data types, which is conducive to achieving fine-grained management of the cache space in the processor and further improving the processing efficiency of the processor for data requests.
[0061] The above method can be applied to Figure 2 the data storage system shown in Figure 2 FIG. 10 is a schematic diagram of a data storage system provided by the present application. The data storage system includes a computing device 200 and a storage system 220. In Figure 2 the application scenario shown, a user accesses and stores data through an application program. The computer running these application programs can be referred to as a "computing device". The computing device 200 can be a physical machine or a virtual machine. Physical computing devices include, but are not limited to, desktop computers, servers, laptop computers, and mobile devices. In some alternative cases, the computing device 200 may also be referred to as a client, a user terminal, or a request terminal, etc., which is not limited herein.
[0062] In one possible example, the computing device 200 accesses the storage system 220 through a network 230 to access and store data. For example, the network 230 may include a switch 210.
[0063] In another possible example, the computing device 200 can also communicate with the storage system 220 through a wired connection, such as a universal serial bus (USB), a peripheral component interconnect express (PCIe) bus, a unified bus (UB or Ubus), etc.
[0064] Figure 2The storage system 220 shown may be a centralized storage system. The characteristic of a centralized storage system is that there is a unified entry point through which all data from external devices must pass. This entry point is the engine 221 of the centralized storage system. The engine 221 is the most core component in the centralized storage system, and many advanced functions of the storage system are implemented therein.
[0065] As Figure 2 shown, there may be one or more controllers in the engine 221. Figure 2 Taking the case where the engine 221 includes one controller as an example. In one possible example, if the engine 221 has multiple controllers, there may be a mirror channel between any two controllers to implement the function of mutual backup between any two controllers, thereby avoiding the unavailability of the entire storage system 220 caused by hardware failures.
[0066] The engine 221 also includes a front-end interface 2211 and a back-end interface 2214. The front-end interface 2211 is used to communicate with the computing device 200, thereby providing data access services for the computing device 200. The back-end interface 2214 is used to communicate with hard disks to expand the capacity of the storage system 220. Through the back-end interface 2214, the engine 221 can connect more hard disks, thereby forming a very large storage resource pool.
[0067] In terms of hardware, as Figure 2 shown, the controller at least includes a processor 2212 and a memory 2213. The processor 2212 is a CPU, which is used to process data requests from outside the storage system 220 (servers or other storage systems), and is also used to process requests generated inside the storage system 220. Exemplarily, when the processor 2212 receives a write data request sent by the computing device 200 through the front-end interface 2211, it will temporarily save the data in these write data requests in the memory 2213. When the total amount of data in the memory 2213 reaches a certain threshold, the processor 2212 sends the data stored in the memory 2213 to at least one of the mechanical hard disk (harddisk drive, HDD) 2221, mechanical hard disk 2222, solid state drive (SSD) 2223 or other hard disk 2224 through the back-end port for persistent storage.
[0068] The memory 2213 refers to the internal memory that directly exchanges data with the processor. It can read and write data at any time and is very fast. It serves as the temporary data memory for the operating system or other running programs. The memory includes at least two types of memories. For example, the memory can be either a random access memory or a read only memory (ROM). For instance, the random access memory can be a dynamic random access memory (DRAM) or a storage class memory (SCM). DRAM is a semiconductor memory, and like most random access memories (RAM), it belongs to a volatile memory device. SCM is a composite storage technology that combines the characteristics of traditional storage devices and memories. The storage class memory can provide faster read and write speeds than hard disks, but its access speed is slower than that of DRAM, and its cost is also lower than that of DRAM. However, DRAM and SCM are only exemplary illustrations in this embodiment. The memory can also include other random access memories, such as static random access memory (SRAM), etc. For the read only memory, for example, it can be a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), etc.
[0069] In addition, the memory 2213 can also be a dual in-line memory module or a dual in-line memory module (DIMM), that is, a module composed of dynamic random access memory (DRAM), or it can be an SSD. In practical applications, multiple memories 2213 and different types of memories 2213 can be configured in the controller. The number and type of the memories 2213 are not limited in this embodiment. In addition, the memory 2213 can be configured to have a power retention function. The power retention function means that when the system experiences a power failure and then powers on again, the data stored in the memory 2213 will not be lost. The memory with the power retention function is called a non-volatile memory.
[0070] The software program is stored in the memory 2213, and the processor 2212 runs the software program in the memory 2213 to manage the hard disk. For example, the hard disk is abstracted as a storage resource pool, and the storage resource pool is provided to the server in the form of a logical unit number (LUN) for use. Here, the LUN is actually the hard disk seen on the server. Of course, some centralized storage systems are also file servers themselves and can provide shared file services for the server.
[0071] As Figure 2 shown, in this system, the engine 221 may not have a hard disk slot. The hard disk needs to be placed in the hard disk enclosure 222, and the back-end interface 2214 communicates with the hard disk enclosure 222. The back-end interface 2214 exists in the engine 221 in the form of an adapter card. Two or more back-end interfaces 2214 can be used simultaneously on one engine 221 to connect multiple hard disk enclosures. Alternatively, the adapter card can also be integrated on the motherboard. In this case, the adapter card can communicate with the processor 2212 through the PCIe bus.
[0072] It should be noted that Figure 2 only one engine 221 is shown, but in actual applications, the storage system may include two or more engines 221, and redundancy or load balancing is performed among multiple engines 221.
[0073] The hard disk enclosure 222 includes a control unit 2225 and several hard disks. The control unit 2225 can have various forms. In one case, the hard disk enclosure 222 belongs to an intelligent disk enclosure, such as Figure 2As shown, the control unit 2225 includes a CPU and a memory. The CPU is used to perform operations such as address conversion and data reading and writing. The memory is used to temporarily store data to be written to the hard disk or data read from the hard disk to be sent to the controller. In another case, the control unit 2225 is a programmable electronic component, such as a data processing unit (DPU). The DPU has the generality and programmability of the CPU, but is more specialized and can operate efficiently on network data packets, storage requests or analysis requests. The DPU is distinguished from the CPU by a high degree of parallelism (requiring the processing of a large number of requests). Optionally, the DPU here can also be replaced by a graphics processing unit (GPU), an embedded neural-network processing unit (NPU), or other processing chips. Usually, the number of control units 2225 can be one, or two or more. The functions of the control unit 2225 can be offloaded to the network card 2226. In other words, in this embodiment, the hard disk enclosure 222 does not have a control unit 2225 inside, but the network card 2226 is used to complete data reading and writing, address conversion, and other computing functions. At this time, the network card 2226 is a smart network card. It can include a CPU and a memory. The CPU is used to perform operations such as address conversion and data reading and writing. The memory is used to temporarily store data to be written to the hard disk or data read from the hard disk to be sent to the controller. It can also be a programmable electronic component, such as a DPU. There is no ownership relationship between the network card 2226 and the hard disk in the hard disk enclosure 222, and the network card 2226 can access any hard disk in the hard disk enclosure 222 (such as Figure 2 the mechanical hard disk 2221, the mechanical hard disk 2222, the solid-state hard disk 2223, and the other hard disk 2224 shown), so it is more convenient to expand the hard disk when the storage space is insufficient.
[0074] According to the type of communication protocol between the engine 221 and the hard disk enclosure 222, the hard disk enclosure 222 may be a serial attached small computer system interface (SAS) hard disk enclosure, or an NVMe (Non-Volatile Memory express) hard disk enclosure, as well as other types of hard disk enclosures. The SAS hard disk enclosure adopts the SAS 3.0 protocol, and each enclosure supports 25 SAS hard disks. The engine 221 is connected to the hard disk enclosure 222 through an on-board SAS interface or a SAS interface module. The NVMe hard disk enclosure is more like a complete computer system, and the NVMe hard disks are inserted into the NVMe hard disk enclosure. The NVMe hard disk enclosure is then connected to the engine 221 through a remote direct memory access (RDMA) port.
[0075] In an alternative implementation, the storage system 220 is a centralized storage system with integrated disk control. The storage system 220 does not have the above-mentioned hard disk enclosure 222, and the engine 221 is used to manage multiple hard disks connected through hard disk slots. The functions of the hard disk slots can be implemented by the backend interface 2214.
[0076] In another alternative implementation, Figure 2 The storage system 220 shown is a distributed storage system. The distributed storage system includes a computing device cluster and a storage device cluster. The computing device cluster includes one or more computing devices, and the computing devices can communicate with each other. The computing device can be a type of computing device, such as a server, a desktop computer, or a controller of a storage array, etc. Hardware-wise, the computing device can include a processor, a memory, and a network card, etc. Among them, the processor is a CPU, which is used to process data requests from outside the computing device, or data requests generated inside the computing device. Exemplarily, when the processor receives a write data request sent by a user, it will temporarily store the data in these write data requests in the memory. When the total amount of data in the memory reaches a certain threshold, the processor sends the data stored in the memory to the storage device for persistent storage. In addition, the processor is also used to calculate or process data, such as metadata management, deduplication, data compression, virtualized storage space, and address translation, etc. In one example, any computing device can access any storage device in the storage device cluster through the network. The storage device cluster includes multiple storage devices. A storage device includes one or more controllers, a network card, and multiple hard disks, and the network card is used to communicate with the computing device.
[0077] In yet another alternative implementation, Figure 2The storage system 220 shown may refer to a server. For example, the server is used to provide computing resources. In the case of a single server, it may include multiple processors or processor cores, and each processor or processor core can be a computing resource. Therefore, a single server can provide multiple computing resources. For example, the server may refer to an application server, a file server, or the like.
[0078] It should be noted that the above are only examples of the application scenarios or systems provided in this embodiment, and should not be construed as a limitation to this application.
[0079] To solve the problem that the above requests compete for the same cache space, resulting in a reduction in the processing efficiency of the processor 2212. In this application, when the data type indicated by the data request processed by the processor 2212 is hot data, cold data, or address translation data, corresponding storage spaces are respectively allocated to reduce the request competition for the cache space and avoid the reduction in the processing efficiency of the processor 2212.
[0080] Figure 3 It is a schematic structural diagram of a processor provided in this application. Figure 3 The processor 310 in provides a possible hardware implementation manner for the above-mentioned processor 2212. Figure 3 The memory 320 in can implement the functions of the above-mentioned memory 2213.
[0081] Exemplarily, the processor 310 may be, but is not limited to, a processor with neural network processing capabilities such as a CPU, an NPU, or a GPU, or may also be a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. This application does not limit this.
[0082] The processor 310 includes multiple processor cores, such as core 311, core 312, and core 313, a memory controller 314, and multiple levels of caches, such as a first-level cache (L1 cache or L1), a second-level cache (L2 cache or L2), and a last-level cache (LLC) for data interaction with the memory 320. In the Figure 3 embodiment provided, the LLC refers to the L3 cache, and one of the multiple processor cores may include an L1 cache, an L2 cache, and a PE. It should be noted that Figure 3 This is only an example provided for the embodiments of this application. In some possible examples, the processor 310 may further include more levels of caches or fewer levels of caches. In this article, without causing misunderstanding, the N-level cache can be represented by LN cache, where N is a positive integer. For example, the first-level cache can be represented by L1 cache, the second-level cache can be represented by L2 cache, and the third-level cache can be represented by L3 cache. If the processor in this article is further provided with more levels of caches, such as a fourth-level cache, this fourth-level cache can be represented by L4 cache. In other examples, if the processor is further provided with a "heterogeneous cache", for example, a certain manufacturer proposes a "2.5-level cache", this "2.5-level cache" can also be represented by L2.5 cache.
[0083] In a possible scenario, the L2 cache can be shared by multiple CPUs or installed on the motherboard. The L1 cache can be divided into an instruction cache and a data cache so that the CPU can read instructions and data in the L1 cache simultaneously.
[0084] An MPAM tag module is also provided inside the PE in the processor core. This MPAM tag module is used to tag data requests according to the data type indicated by the data request. For example, when the PE receives an instruction request, a tag, such as 1, can be written in the reserved field of the instruction request. When the PE receives a data request, the page table is then queried based on the virtual address carried in the data request to determine the type field corresponding to the virtual address in the page table. If the type field indicates that the data to be accessed corresponding to the virtual address is hot data, the value of the tag, such as 2, is then written in the reserved field of the data request. The foregoing content is only an example provided for this application and should not be construed as a limitation to this application.
[0085] An MPAM control module is also provided in the management unit of the L3 cache. The MPAM control module is used to allocate corresponding cache space for data requests according to the data type.
[0086] The DMC 214 is used to manage and plan data transfer from memory to the processing unit. It can be a separate chip or integrated into the chip of the processing unit. The DMC 214 is a bus circuit controller that controls the internal memory 320 of the processor 310 and is used to manage and plan data transfer from the memory 320 to the processor core. Through the DMC 214, data can be exchanged between the memory 320 and the processor core. The DMC can be a separate chip and is connected to the processor core through the system bus. Those skilled in the art will know that the memory controller can also be integrated into the processor 310, or built into the north bridge, or can be an independent memory controller chip. The specific location and form of the memory controller are not limited in this embodiment. In practical applications, the memory controller can control the necessary logic to write data into the memory 320 or read data from the memory 320. The memory controller 314 can be a memory controller in a processor system such as a general-purpose processor, a dedicated accelerator, a GPU, an FPGA, or an embedded processor.
[0087] Exemplarily, the processor 310 can access the memory 320 at high speed through the memory controller 314 and perform read and write operations on any storage unit (such as a memory page) in the memory 320.
[0088] Such as Figure 3 shown, the memory 320 is used to store application data (as well as instructions or data required to run the application), page tables, etc.
[0089] It should be noted that Figure 3 is only an example of a processor structure and should not be construed as a limitation of this application. In other embodiments of this application, the processor may further include more cores, or may further include a TLB.
[0090] The following describes in detail the implementation manner of the embodiments of this application with reference to the accompanying drawings.
[0091] Figure 4 is a schematic flowchart of a cache resource allocation method provided by this application. This cache resource allocation method can be applied to Figure 2 the data processing system shown, such as this cache resource allocation method can be implemented by Figure 2 the processor 2212 shown or Figure 3 the processor 310 in Figure 3 shown. Here, taking the cache resource allocation method implemented by the processor 310 shown in
[0092] S410. The processor 310 obtains a data request.
[0093] The processor 310 obtains a data request sent by other devices or other hardware in the computing device including the processor 310. The data indicated by the data request is the data to be accessed that the processor 310 needs to process or use.
[0094] Exemplarily, the control unit in the processor 310 obtains the data request.
[0095] In a possible scenario, the processor 310 can generate a data request according to a user instruction.
[0096] For example, a virtual machine running on the processor 310 generates a data request according to a user's trigger operation (such as a click or swipe operation, etc.) to process or use the data to be accessed.
[0097] S420. The processor 310 determines the data type indicated by the data request.
[0098] Among them, the data type may include hot data, cold data, or address translation data. This address translation data is used to indicate the correspondence between the virtual address and the physical address.
[0099] It should be noted that the above address translation data can also be called the page table used by the MMU during the address translation process. The following content will be described in terms of page table data.
[0100] In a possible implementation manner, the processor 310 can first determine that the data type is application data or page table data. Among them, the application data includes hot data and cold data.
[0101] Exemplarily, the processor 310 can determine the data type according to the virtual address carried by the data request.
[0102] The processor 310 passes the virtual address carried by the data request to the internal address decoding unit. The address decoding unit determines whether the virtual address belongs to the address storing the page table or the address storing the application data according to the specified address range or specific identifier. The application data and the page table are stored in different regions of the memory. If the virtual address belongs to the address storing the page table, the data type is page table data. If the virtual address belongs to the address storing the application data, the data type is application data.
[0103] Furthermore, the processor 310 differentiates between hot data and cold data in the application data.
[0104] Regarding the above differentiation of application data into hot data or cold data, a possible implementation manner is provided below.
[0105] As Figure 5 shown, Figure 5Schematic flowchart of the data type determination method provided for this application. The method may include the following steps S510 - S540.
[0106] S510. The processor 310 parses the data request to obtain a virtual address.
[0107] This virtual address indicates the access location of the data to be accessed, that is, the physical address where the data to be accessed is stored in the memory.
[0108] Since the data request adopts a fixed format. For example, the data request includes a virtual address and the processing means for the data to be accessed corresponding to the virtual address, etc. This processing means can be access, read, etc. Furthermore, the processor 310 can read the virtual address from the specified field in the data request, thereby achieving parsing the data request to obtain the virtual address.
[0109] S520. The processor 310 determines whether the data to be accessed is hot data or cold data according to the virtual address from the page table.
[0110] Wherein, the page table includes the relationship between the virtual address and hot data or cold data.
[0111] Exemplarily, the page table can be stored in Figure 3 the memory 320 therein, or in the TLB in the processor 310. For the format of the page table, see Table 1 below.
[0112] Table 1
[0113]
[0114] Among them, the above-mentioned IGNORED represents reserved bits, which can also be referred to as reserved fields, that is, bits / fields that are not used. PBHA (page base hardware attribute) is an optional, implementation-defined feature. It allows software to set up to four bits in the page table, which are then propagated to memory through transactions and can be used in memory to control memory components. hot / cold type is used to represent the type of data to be accessed corresponding to the virtual address. For example, when the hot / cold type value is the first value, it indicates that the data to be accessed is hot data; when the value corresponding to hot / cold type is the second value, it indicates that the data to be accessed is cold data. UXN is used to control whether the memory space of the page table can be executed in the non-privileged state of el0. PXN controls whether the memory space corresponding to the page table can execute code in the privileged state. Contiguous indicates whether the address at which the page table is executed is continuous and is used for some optimizations related to the TLB. DBM (dirty bit management) represents dirty bit management and is used to indicate whether the page table has been modified and needs to be written back to memory. RES0 represents a reserved field and is usually set to 0. Output address represents the output address, which usually refers to the field in the page table used to store the physical address, thus mapping between the virtual address and the physical address of the page table attributes.
[0115] It should be noted that the content shown in Table 1 above is only an example and should not be construed as a limitation of this application. The values corresponding to each field in Table 1 above are not filled in. In actual situations, the processor can assign values according to needs. In other embodiments of this application, the page table may also include more or fewer fields. For example, the page table can also include GP (guarded page) or virtual address. GP represents the protected page table field and is used to implement the page table protection mechanism. For example, when GP is set to 1, it indicates that this page table is a protected page table, which can prevent illegal access or accidental modification of sensitive data.
[0116] In this application, the processor 310 inserts a hot / cold type field into the page table, so that the page table has a corresponding relationship between the virtual address and the hot / cold type field. Then, according to the virtual address carried in the data request, it can quickly determine whether the data to be accessed is cold data or hot data, thereby improving the efficiency of determining the data type indicated by the data request.
[0117] In a possible implementation manner, after the processor 310 obtains the page table from the memory 320 or the TLB, it queries the value of the type field corresponding to the virtual address from the page table. Exemplarily, this type field is the cold / hot data field.
[0118] For example, the cold / hot data type field may be the hot / coldtype field in Table 1 above.
[0119] In a possible scenario, if the type field corresponding to the virtual address in the page table is the first value, it indicates that the data to be accessed is hot data. If the type field corresponding to the virtual address in the page table is the second value, it indicates that the data to be accessed is cold data. For example, the above first value may be 1, the second value may be 0, etc., and the present application does not limit this.
[0120] For the content of assigning values (the first value / the second value) to the type field in the page table, reference can be made to the following embodiments and will not be elaborated here.
[0121] S530. If the data to be accessed is hot data, the processor 310 determines that the data type indicated by the data request is hot data.
[0122] If the value of the hot / coldtype field in Table 1 above is 1, the data to be accessed is hot data, which further indicates that the data to be accessed for this data request, such as access or read, is hot data, so the data type is hot data.
[0123] The value of 1 for the above hot / coldtype field is only an example and should not be construed as a limitation to the present application.
[0124] S540. If the data to be accessed is cold data, the processor 310 determines that the data type indicated by the data request is cold data.
[0125] If the value of the hot / coldtype field in Table 1 above is 0, the data to be accessed is cold data, which further indicates that the data to be accessed for this data request, such as access or read, is cold data, so the data type is cold data.
[0126] The value of 0 for the above hot / coldtype field is only an example and should not be construed as a limitation to the present application.
[0127] In the present application, since the page table indicates the correspondence between the virtual address and hot data or cold data, the processor 310 determines whether the data to be accessed is hot data or cold data from the page table according to the virtual address in the data request. Further, the processor 310 can accurately determine that the data type is hot data or cold data, improving the accuracy of allocating the corresponding cache space for the data request, avoiding the problem that a large number of data requests compete for the same cache space, and improving the processing efficiency of the processor 310.
[0128] Please continue to refer to Figure 4 , the cache resource allocation method provided in this embodiment further includes the following step S430.
[0129] S430. The processor 310 allocates cache space in the cache for the data request according to the data type indicated by the data request.
[0130] Among them, the cache space matches the data type indicated by the data request.
[0131] In a possible implementation, the processor 310 determines the cache space corresponding to the data type indicated by the data request from the allocation policy, and allocates the cache space in the cache for the data request.
[0132] Exemplarily, the processor 310 allocates corresponding cache space for data requests indicating different data types, such as hot data, cold data, or page table data, according to the allocation policy.
[0133] It should be noted that the cache space allocated by the processor 310 for data requests indicating different data types does not overlap. In other words, the address ranges of the cache space allocated for data requests indicating the data types of hot data, cold data, or page table data do not cross. The above three sections of cache space can be adjacent or set at intervals. The cache space can be the cache space included in the cache of the processor 310. For example, it can be the cache space included in one or more of L1 cache, L2 cache, and L3 cache.
[0134] In a possible example, the cache space allocated by the processor 310 for data requests of the data type of hot data is the largest, the cache space allocated for data requests of the data type of cold data is the second largest, and the cache space allocated for data requests of the data type of page table data is the smallest. The above content is only an example provided by this application and should not be construed as a limitation of this application. In other examples of this application, the size, area, or address of the cache space allocated for data requests of the data types of hot data, cold data, or page table data can be set by the user according to their needs.
[0135] The user can define the allocation policy to implement setting the size, area, or address of the cache space allocated for the data types of hot data, cold data, or page table data. The allocation policy can be stored in memory, L1 cache, L2 cache, or L3 cache. The allocation policy indicates the correspondence between the cache space in the cache and the data type. After the processor 310 reads the allocation policy from memory, L1 cache, L2 cache, or L3 cache, it queries the cache space corresponding to the data type from the allocation policy. For example, the processor 310 queries the area or address of the cache space corresponding to the data type of hot data, and then allocates the cache space in the corresponding area or address for the data request.
[0136] In this application, since the allocation policy indicates the correspondence between the cache space in the buffer and the data type, the processor 310 can accurately determine the cache space corresponding to the data type from the allocation policy. Furthermore, the cache space in the buffer is allocated for the data request, so that data requests of different data types are allocated to different cache spaces, avoiding the problem that a large number of data requests compete for the same cache space, and improving the processing efficiency of the processor 310.
[0137] For example, the user sets the above allocation policy through the kernel interface provided by the MPAM module. The MPAM module can be an MPAM label module, an MPAM control module, or an MPAM label module and an MPAM control module.
[0138] The following Table 2 provides an example of an allocation policy.
[0139] Table 2
[0140] Data Types Cache space Page table data 0-Away Hot Data A-Bway Cold Data B-Cway
[0141] It should be noted that the above intervals include the right boundary, and the values of A, B, and C increase in sequence. In other words, the three storage spaces of 0 - Away, A - Bway, and B - Cway do not overlap. The above Table 2 is only an example and should not be construed as a limitation of this application. In other embodiments of this application, the above three cache spaces may not be adjacent.
[0142] In a possible example, the processor 310 can further apply for a specific storage space in the cache space. Exemplarily, the processor 310 can allocate a storage space that meets the size of the data to be accessed indicated by the data request for the data request according to the size of the data to be accessed indicated by the data request.
[0143] In a possible scenario, taking 0 - Away in the above 0 - Away, A - B way, and B - C way as an example for illustration, 0 - Away can represent the 0 - A cache groups or 0 - A cache lines in a cache group, and this application does not limit this. For the description of cache groups or cache lines, reference can be made to the following Figure 6 content.
[0144] As Figure 6 shown, Figure 6 is a schematic diagram of the division of a cache provided by this application. The cache 600 can refer to Figure 2 any level of cache, such as L1 cache, L2 cache, or L3 cache, etc. Here, it is illustrated by taking the cache 600 as L3 cache. The cache 600 includes multiple cache groups, such as Figure 3The cache groups 1 to 8 shown each include one or more cache lines.
[0145] As Figure 6 shown, cache group 3 includes 128 cache lines. Assuming that the line size of each cache line is 64 bytes (B), the storage capacity of this cache group 3 is 64 B × 128 = 8 kilobytes (KB), and the storage capacity of cache 600 is 8 KB × 8 = 64 KB.
[0146] In one possible example, cache space 0 - Away can correspond to Figure 6 cache groups 1 - 3 in Figure 6 cache space A - Bway can correspond to Figure 6 cache groups 3 - 5 in
[0147] In another possible example, cache space 0 - Away can correspond to Figure 6 cache lines 0 - 1 in Figure 6 cache space A - B way can correspond to Figure 6 the third cache line in
[0148] Exemplarily, prior to - be - accessed data may be stored in each cache line for the processor 310 to access.
[0149] Regarding the content of the cache space corresponding to the data type indicated by the data request determined by the processor 310 from the allocation policy, a possible embodiment is provided below, and an example of allocating cache resources in L3 cache among L1 cache, L2 cache, and L3 cache is used for illustration. As Figure 7 shown, Figure 7 is a schematic flowchart of the method for allocating cache resources in L3 cache provided by this application. The method includes the following steps S710 and S720.
[0150] S710. The processor 310 assigns a third value to the tag of the data request according to the data type indicated by the data request.
[0151] This tag can indicate the data type indicated by the data request.
[0152] In one possible implementation manner, the MPAM tag module in the processor 310 assigns a third value to the tag (PARTID or PMG) of the data request according to the data type indicated by the data request determined by the PE.
[0153] Exemplarily, the MPAM tag module can define a tag in the reserved field of a data request and assign a value to the tag, so that the data request carries the tag, and the value of the tag can indicate the data type.
[0154] For example, after the PE determines that the data type is page table data, the MPAM tag module assigns the value 1 to the tag of the data request.
[0155] After the PE determines that the data type is hot data, the MPAM tag module assigns the value 2 to the tag of the data request.
[0156] After the PE determines that the data type is cold data, the MPAM tag module assigns the value 3 to the tag of the data request.
[0157] S720, the processor 310 determines the cache space corresponding to the tag from the allocation policy.
[0158] The processor 310 can obtain the user-defined allocation policy from the memory, L1 cache, L2 cache or L3 cache, and then determine the cache space in the L3 cache corresponding to the tag from the allocation policy.
[0159] The allocation policy of this embodiment shows the correspondence between the data type indicated by the data request and the tag, and the correspondence between the tag and the cache space in the cache. The user can set the above allocation policy through the kernel interface provided by the MPAM module.
[0160] For the allocation policy in this embodiment, an example as shown in Table 3 below is provided.
[0161] Table 3
[0162] Data Types Label Cache space Page table data 1 0-Away Hot Data 2 A-Bway Cold Data 3 B-Cway
[0163] It should be noted that the above intervals include the right boundary, and the values of A, B, and C increase in sequence. In other words, the three cache spaces of 0 - Away, A - Bway, and B - Cway do not overlap. The above Table 3 is only an example and should not be construed as a limitation of the present application. In other embodiments of the present application, the above three cache spaces may not be adjacent.
[0164] For the content of 0 - Away, A - B way, and B - C way, reference can be made to the above Table 2 and Figure 6 the expressions shown, which will not be elaborated here.
[0165] In a possible implementation manner, the MPAM control module in the L3 cache allocates a corresponding cache space for the data request according to the tag carried by the data request.
[0166] Exemplarily, if the value of the tag in the data request is 1, the cache space allocated for the data request is cache space 0 - Away.
[0167] If the value of the tag in the data request is 2, the cache space allocated for the data request is cache space A - B way.
[0168] If the value of the tag in the data request is 3, the cache space allocated for the data request is cache space B - C way.
[0169] In this application, the processor 310 defines the reserved field in the data request as a tag and assigns a value to the tag, so that the data request carries a tag indicating the data type. Then, the MPAM control module in the L3 cache allocates a corresponding cache space for the data request according to the tag carried by the data request and the allocation policy, which can improve the efficiency of allocating cache space for the data request.
[0170] In a possible embodiment, the memory controller 314 in the processor 310 queries the page table according to the virtual address in the data request to obtain the physical address corresponding to the virtual address, and then loads / writes the data to be accessed stored in the physical address into the cache space.
[0171] Taking the hot data among the page table data, hot data or cold data as an example for illustration, the data request carries a virtual address a, and the virtual address corresponds to the data to be accessed.
[0172] The memory controller 314 queries the physical address a of the data to be accessed corresponding to the virtual address a from the page table, and then reads the data to be accessed from the physical address a and writes it into the cache space 0 - Away in the L3 cache.
[0173] In a possible embodiment, the assignment of the hot / coldtype field carried in the above page table can be determined in the following way.
[0174] The processor 310 obtains the access information of the page table, and then determines that the page table is a hot page when the access information meets the conditions; determines that the page table is a cold page when the access information does not meet the conditions. Thus, the processor 310 modifies the page table attributes in the hot page / cold page. For example, defines the reserved field in the hot page as a type field and sets the type field to hottype or 0, etc., and defines the reserved field in the cold page as a type field and sets the type field to cold or 1, etc.
[0175] Among them, the above access information includes the number of accesses or access frequency, etc. The hot page indicates that the data to be accessed is hot data, and the cold page indicates that the data to be accessed is cold data.
[0176] In a possible implementation, the processor 310 may obtain access information of the page table from a register. The register records the access count of each page table and the total access count of all page tables, and this register is located inside the processor 310.
[0177] The above condition may be whether the access count reaches a threshold a, or whether the access frequency reaches a threshold b.
[0178] In a possible example, when the access count meets the threshold a, it is determined that the page table is a hot page, and then the data to be accessed indicated by the virtual address in the hot page is all hot data. When the access count does not meet the threshold a, it is determined that the page table is a cold page, and then the data to be accessed indicated by the virtual address in the cold page is all cold data.
[0179] In another possible example, when the access frequency meets the threshold b, it is determined that the page table is a hot page, and then the data to be accessed indicated by the virtual address in the hot page is all hot data. When the access frequency does not meet the threshold b, it is determined that the page table is a cold page, and then the data to be accessed indicated by the virtual address in the cold page is all cold data.
[0180] In a possible scenario, a hot and cold page statistical method may be adopted to determine whether the above page table is a hot page or a cold page. For example, the hot and cold page statistical methods include: SPE (statistical profiling extension), ARC (adaptive replacement cache), etc.
[0181] In this application, the processor 310 determines a hot page or a cold page through the access information of the page table, and then adjusts the page table attributes of the hot page / cold page, that is, defines a hot / cold type field in the page table and assigns a value to this field. When the processor 310 receives a data request, it can directly determine from the page table that the data to be accessed corresponding to the data request is cold data / hot data, so that the data type can be obtained quickly, and the efficiency of the processor 310 to allocate cache space according to the data type for the data request is improved.
[0182] In other embodiments of this application, the processor 310 may also obtain an instruction request, and then allocate a cache space in the cache for the instruction request according to the cache space determined corresponding to the instruction request in the allocation policy.
[0183] In a possible example, based on Table 2 above, the allocation policy further includes the correspondence between data instructions and the cache space in the cache.
[0184] In another possible example, based on Table 3 above, the allocation policy further includes the correspondence between data instructions and tags, and the correspondence between tags and the cache space in the cache.
[0185] Exemplarily, the CU in the processor 310 obtains an instruction request, and then the PE processes the instruction request, such as tagging a value on the reserved field of the instruction request.
[0186] The processor 310 determines the cache space corresponding to the tag in the instruction request from the allocation policy, and then allocates cache space for the instruction request.
[0187] For more content on the processor 310 allocating cache space for instruction requests, reference may be made to the above content on the processor 310 allocating cache space for data requests, which will not be elaborated here.
[0188] It can be understood that in order to implement the functions in the above embodiments, the processor includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the units and method steps of each example described in the embodiments disclosed in the present application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application scenarios and design constraints of the technical solution.
[0189] In the above text, in combination with Figures 2 to 7 ..., the cache resource allocation method provided by the present application is described in detail. Next, in combination with Figure 8 ..., Figure 8 The structural schematic of the cache resource allocation device provided by the present application Figure 1 ..., the cache resource allocation device provided by the present application will be described. The cache resource allocation device 800 can be used to implement the functions of the processor 310 in the above method embodiments, and thus can also achieve the beneficial effects possessed by the above method embodiments.
[0190] Such as Figure 8 shown, the cache resource allocation device 800 includes an acquisition module 810 and an allocation module 820. The cache resource allocation device 800 is used to implement the functions of the processor 310 in the corresponding method embodiments described above. In a possible example, the specific process of the cache resource allocation device 800 for implementing the above cache resource allocation method includes the following process: Figures 2 to 7 The acquisition module 810 is used to acquire a data request.
[0191] The allocation module 820 is used to allocate cache space for the data request in the cache according to the data type indicated by the data request. Wherein, the data type includes: hot data, cold data or address conversion data, the cache space matches the data type, and the address conversion data is used to indicate the correspondence between the virtual address and the physical address.
[0192]
[0193] To further implement the functions in the method embodiments shown above Figures 2 to 7 in this application, a cache resource allocation device is also provided, as shown in Fig. 9 shown below, Fig. 9 is a structural schematic diagram of the cache resource allocation device provided by this application Figure 2 . The cache resource allocation device 800 further includes: a determination module 830 and an identification module 840.
[0194] The determination module 830 is configured to parse a data request to obtain a virtual address, and the virtual address is used to indicate the access location of the data to be accessed. Determine whether the data to be accessed is hot data or cold data from the page table according to the virtual address. The page table includes: the relationship between the virtual address and hot data or cold data. If the data to be accessed is hot data, determine the data type as hot data; if the data to be accessed is cold data, determine the data type as cold data.
[0195] The identification module 840 is configured to obtain the access information of the page table; if the access information meets the conditions, determine that the page table is a hot page, and the hot page indicates that the data to be accessed is hot data; if the access information does not meet the foregoing conditions, determine that the page table is a cold page, and the cold page indicates that the data to be accessed is cold data.
[0196] It should be understood that the cache resource allocation device 800 in the embodiments of the present invention and this application can be implemented by a CPU, or by an ASIC, or by a programmable logic device (PLD). The above PLD can be a complex programmable logical device (CPLD), FPGA, generic array logic (GAL), or any combination thereof. When the cache resource allocation device 800 is implemented by software Figures 2 to 7 in any of the cache resource allocation methods shown above, the cache resource allocation device 800 and its various modules can also be software modules.
[0197] Figure 8 Or Fig. 9 The cache resource allocation device 800 provided is only an example provided in this embodiment. In some cases, the cache resource allocation device 800 may include more or fewer software units, and this application does not limit this. For a more detailed description of the above cache resource allocation device 800, it can be directly obtained by referring to the relevant descriptions in the above Figures 2 to 7 shown embodiments, and will not be elaborated here.
[0198] Exemplarily, when the cache resource allocation device 800 is implemented by hardware, the hardware can be implemented by a chip, such as the aforementioned processor 310 and processor 2212, etc. The chip includes an interface circuit and a control circuit. The interface circuit is used to receive signals from other devices outside the processor and transmit them to the control circuit, such as obtaining a data request or an instruction request, or sending the signals from the control circuit to other devices outside the processor.
[0199] The control circuit is used to implement the method of any possible implementation manner in the above embodiments through a logic circuit or by executing code instructions. The beneficial effects can be referred to the description of any possible implementation manner in the above embodiments, and will not be elaborated here.
[0200] This application also provides a computing device. As Fig.10 shown, Fig.10 FIG. 1000 is a schematic structural diagram of a computing device provided by this application. The computing device 1000 includes: a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other through the bus 1002. The computing device 1000 can be a server or a terminal device, and the computing device 1000 may include the aforementioned processor 310. It should be noted that this application does not limit the number of processors and memories in the computing device 1000.
[0201] The bus 1002 can be, but is not limited to: a PCIe bus, a universal serial bus (USB), or an inter-integrated circuit (I2C) bus, an EISA (extended industry standard architecture) bus, a UB, a CXL (compute express link), a CCIX (cache coherent interconnect for accelerators), etc. The bus 1002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Fig.10 only one line is shown in FIG. 1000, but it does not mean that there is only one bus or one type of bus. The bus 1002 can include a path for transmitting information between various components (for example, the memory 1006, the processor 1004, the communication interface 1008) of the computing device 1000.
[0202] The processor 1004 can include a CPU, a GPU, an NPU, a microprocessor (MP), a DSP, an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.
[0203] The memory 1006 may include volatile memory, such as RAM. The memory 1006 may also include non-volatile memory, such as ROM, flash memory, HDD, or SSD.
[0204] The executable program code is stored in the memory 1006, and the processor 1004 executes the executable program code to implement the functions of the aforementioned acquisition module and allocation module respectively, so as to implement the above cache resource allocation method. That is, the memory 1006 stores instructions for executing the cache resource allocation method.
[0205] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 1000 and other devices or communication networks.
[0206] The embodiment of the present application also provides a computer program product containing instructions. The computer program product may be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute the cache resource allocation method.
[0207] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the cache resource allocation method.
[0208] The method steps in the embodiments of the present application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions may be composed of corresponding software modules, and the software modules may be stored in RAM, flash memory, ROM, PROM, EPROM, EEPROM, registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in an ASIC. Additionally, the ASIC may be located in a network device or a terminal device. Of course, the processor and the storage medium may also exist as discrete components in a computing device.
[0209] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in the form of a computer program product in whole or in part. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are executed in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, or other programmable devices. The computer program or instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it may also be an optical medium, such as a DVD; or it may be a semiconductor medium, such as an SSD.
[0210] In various embodiments of the present application, if there is no special description and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be cross-referenced. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships. The various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application. The magnitudes of the serial numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic.
Claims
1. A cache resource allocation method, characterized in that, the method includes: Obtain a data request; Allocate cache space for the data request in a cache according to the data type indicated by the data request; wherein, the data type includes: hot data, cold data or address translation data, the cache space matches the data type, and the address translation data is used to indicate the correspondence between a virtual address and a physical address.
2. The method according to claim 1, characterized in that, allocating cache space for the data request in the cache includes: Determine the cache space corresponding to the data type from an allocation policy; the allocation policy includes: the correspondence between the cache space in the cache and the data type; Allocate the cache space in the cache for the data request.
3. The method according to claim 2, characterized in that, determining the cache space corresponding to the data type from the allocation policy includes: Assign a third value to the tag of the data request, and the third value indicates the data type; Determine the cache space corresponding to the tag from the allocation policy.
4. The method according to any one of claims 1 to 3, characterized in that, the method further includes: Parse the data request to obtain a virtual address, and the virtual address indicates the access location of the data to be accessed; Determine whether the data to be accessed is hot data or cold data from a page table according to the virtual address; the page table includes: the relationship between the virtual address and hot data or cold data; If the data to be accessed is hot data, determine that the data type is the hot data; If the data to be accessed is cold data, determine that the data type is the cold data.
5. The method according to claim 4, characterized in that, the page table includes a type field, and the type field is a first value or a second value, the first value indicates that the data to be accessed is hot data, and the second value indicates that the data to be accessed is cold data.
6. The method according to claim 4 or 5, characterized in that, the method further includes: Obtain the access information of the page table; If the access information meets the conditions, determine that the page table is a hot page, and the hot page indicates that the data to be accessed is hot data; If the access information does not meet the conditions, determine that the page table is a cold page, and the cold page indicates that the data to be accessed is cold data.
7. A cache resource allocation device, characterized in that, the device includes: An acquisition module, configured to acquire a data request; An allocation module, configured to allocate cache space for the data request in a cache according to the data type indicated by the data request; wherein, the data type includes: hot data, cold data or address translation data, the cache space matches the data type, and the address translation data is used to indicate the correspondence between a virtual address and a physical address.
8. A chip, characterized in that, including a control circuit and an interface circuit, the interface circuit is used to obtain a data request, and the control circuit and the interface circuit cooperate to execute the method according to any one of claims 1 to 6.
9. A processor, characterized in that, The processor is used to execute computer programs or instructions to implement the method described in any one of claims 1 to 6 above.
10. A computing device, characterized in that it includes a processor and a memory; the processor is used to execute the instructions stored in the memory, so that the computing device executes the method described in any one of claims 1 to 6.
11. A computer-readable storage medium, characterized in that the storage medium stores computer programs or instructions, and when the computer programs or instructions are executed by a processing device, the method described in any one of claims 1 to 6 is implemented.