Data processing method, control chip, CXL device, host and storage medium
Patent Information
- Application Number
- CN202311568646.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
In storage systems, data transmission efficiency between different memories is low, especially when the host CPU needs to copy data from one memory to the host cache and then copy it from the host cache to another memory, resulting in low cache utilization and slow data copying speed.
By performing a data processing method in the processor, when determining the data blocks that need to be copied, search for data blocks in the target cache area, and copy the data blocks into the cache line when not found, setting two marks as valid so that the cache line can be mapped to two memories at the same time.
The shared cache line of data blocks between two memories is realized, which improves the utilization rate of cache and the speed and efficiency of data copying.
Smart Images

Figure CN120029522A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to data storage technology, and more particularly to a data processing method, a control chip, a CXL device, a host, and a storage medium. Background Art
[0002] CXL (Compute Express Link) technology is a new type of high-speed interconnection technology that supports high-bandwidth, low-latency data transmission, has better flexibility and scalability, and can achieve mixed use of different types of hardware devices. The application scenarios of CXL technology are very broad, including data centers, artificial intelligence, and processor interconnection. In the data center field, CXL technology can interconnect different computing and storage resources to improve system performance and efficiency; in the field of artificial intelligence, CXL technology can enable accelerators such as GPUs and FPGAs to better collaborate with the main processor and increase the speed of AI model training and reasoning; in terms of processor interconnection, CXL technology can achieve interconnection between processors from different manufacturers, improving the overall performance and flexibility of the system.
[0003] With the promotion of CXL technology applications, improving the data transmission efficiency between different memories in the storage system has become an urgent problem to be solved. Summary of the invention
[0004] The present disclosure provides a data processing method, which is applied to a processor capable of accessing a cache, a first memory, and a second memory. The method includes: when determining that a data block stored in a first address of the first memory needs to be copied to a second address of the second memory, performing the following data copy operation:
[0005] Searching for the data block in the target cache area to which the first address is mapped;
[0006] If the data block is not found, copy the data block from the first memory to a first cache line in the target cache area, set the value of a first tag of the first cache line to a tag value Tag1V corresponding to the first address, and set the first tag to be valid;
[0007] In the case where the second address is also mapped to the target cache area, the value of the second tag of the first cache line is set to the tag value Tag2V corresponding to the second address, and the second tag is set to be valid.
[0008] An embodiment of the present disclosure further provides a control chip of a CXL device, comprising a processor and a cache, wherein the processor can access the cache and a first memory and a second memory connected to the control chip, wherein the processor is configured to execute the data processing method described in any embodiment of the present disclosure.
[0009] The embodiment of the present disclosure further provides a CXL device, comprising the control chip described in any embodiment of the present disclosure and a first memory and a second memory connected to the control chip.
[0010] An embodiment of the present disclosure further provides a host, including a processor and a cache, wherein the processor can access the cache and a first memory and a second memory connected to the host, and the processor is configured to execute the data processing method described in any embodiment of the present disclosure.
[0011] The embodiment of the present disclosure further provides a non-transitory computer storage medium storing a computer program, which, when executed by a processor, can implement the data processing method of the embodiment of the present disclosure.
[0012] The embodiment of the present disclosure further provides a computer product, including a computer program. When the computer program is executed by a processor, the data processing method of the embodiment of the present disclosure can be implemented.
[0013] Compared with the related art, the above-mentioned embodiment of the present disclosure, when determining that it is necessary to copy the data block stored in the first address of the first memory to the second address of the second memory, when the first address and the second address are both mapped to the same target cache area, sets two tags for the first cache line in the target cache area, sets the values of the two tags to the tag value corresponding to the first address and the tag value of the second address respectively and sets the two tags to valid, so that the first cache line can be mapped to the first memory and the second memory at the same time, and the data blocks in the first memory and the second memory can share one cache line, that is, the two storage devices in the storage system can share the cache, thereby improving the utilization of the cache, as well as the speed and efficiency of data copying.
[0014] Other features and advantages of the present disclosure will be described in the following description, and partly become apparent from the description, or be understood by implementing the present disclosure. Other advantages of the present disclosure can be realized and obtained by the schemes described in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings are used to provide an understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation on the technical solution of the present disclosure.
[0016] Figure 1 is a schematic diagram of a system architecture involved in an embodiment of the present disclosure;
[0017] Figure 2 is a flow chart of a data copy operation according to an embodiment of the present disclosure;
[0018] Figure 3A is a schematic diagram of the structure of a CXL device according to an embodiment of the present disclosure;
[0019] Figure 3B is a schematic diagram of a host and a memory connected to the host according to an embodiment of the present disclosure;
[0020] Figure 4 is a schematic diagram of a cache line according to an embodiment of the present disclosure;
[0021] Figure 5 is a schematic diagram of mapping of a CXL memory, a host memory and a first cache group in a cache according to an embodiment of the present disclosure;
[0022] Figure 6 is a schematic diagram of mapping between a CXL memory, a host memory and a second cache group in a cache according to an embodiment of the present disclosure;
[0023] Figure 7 is a schematic diagram of mapping of a CXL memory, a host memory and a third cache group in a cache according to an embodiment of the present disclosure;
[0024] Figure 8 is a schematic diagram of mapping of a CXL memory, a host memory and a fourth cache group in a cache in an embodiment of the present disclosure;
[0025] Fig. 9 The figure is a flow chart of a data copy method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0026] The present disclosure describes multiple embodiments, but the description is exemplary rather than restrictive, and it is apparent to those skilled in the art that there may be more embodiments and implementations within the scope of the embodiments described in the present disclosure. Although many possible feature combinations are shown in the drawings and discussed in the specific embodiments, many other combinations of the disclosed features are possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with any other feature or element in any other embodiment, or may replace any other feature or element in any other embodiment.
[0027] The present disclosure includes and contemplates combinations of features and elements known to those of ordinary skill in the art. The embodiments, features, and elements disclosed in the present disclosure may also be combined with any conventional features or elements to form a unique invention scheme defined by the claims. Any features or elements of any embodiment may also be combined with features or elements from other invention schemes to form another unique invention scheme defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in the present disclosure may be implemented individually or in any appropriate combination. Therefore, except for the limitations made according to the attached claims and their equivalents, the embodiments are not subject to other limitations. In addition, various modifications and changes may be made within the scope of protection of the attached claims.
[0028] In addition, when describing representative embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not rely on the specific order of the steps described herein, the method or process should not be limited to the steps in the specific order described. As will be appreciated by those of ordinary skill in the art, other orders of steps are also possible. Therefore, the specific order of the steps set forth in the specification should not be interpreted as a limitation to the claims. In addition, the claims for the method and / or process should not be limited to the steps performed in the order written, and those skilled in the art can easily understand that these orders can be changed and still remain within the spirit and scope of the disclosed embodiments.
[0029] CXL (Compute Express Link) is developed based on PCIE 5.0 and runs on the PCIE physical layer. It supports high-bandwidth, low-latency data transmission, has better flexibility and scalability, and can achieve mixed use of different types of hardware devices. Based on the characteristics of memory access,
[0030] The CXL standard supports a variety of use cases through the following three protocols:
[0031] CXL.io: This protocol leverages the broad industry adoption and familiarity of PCIe as the underlying communications protocol.
[0032] CXL.cache: This protocol is designed for more specific applications and enables accelerators to efficiently access and cache host memory to optimize performance.
[0033] CXL.memory: This protocol enables a host (e.g., a processor) to access device-attached memory using load / store commands.
[0034] Together, the three protocols facilitate the consistent sharing of memory resources between computing devices, such as CPU hosts and AI accelerators. The CXL Consortium has identified three main types of devices that will adopt the new interconnect:
[0035] Type 1 CXL devices (CXL Type 1 Device) support CXL.io and CXL.cache, such as accelerators such as smart NICs, which usually lack local memory. Through CXL, these devices can communicate with the host processor's DDR memory, etc.
[0036] Type 2 CXL devices (CXL Type2 Device) support CXL.io, CXL.cache and CXL.memory. For example, GPU, ASIC and FPGA are equipped with DDR or HBM memory devices (also called accelerators). Through the CXL protocol, the host's memory can be used by the accelerator, and the accelerator's memory can also be used by the host CPU. They are co-located in the same cache coherent domain, which helps to improve heterogeneous workloads.
[0037] Type 3 CXL devices (CXL Type3 Device) support CXL.io and CXL.memory protocols, such as memory devices connected to the host through CXL, which can provide additional bandwidth and capacity to the host processor. The type of device-side memory is independent of the host's memory.
[0038] As mentioned above, with the promotion of CXL technology applications, improving the data transmission efficiency between different memories in storage systems (such as heterogeneous storage systems) has become an urgent problem to be solved. For example, in the scenario of copying data between a memory that communicates with a processor through a memory interface and a memory in a CXL device that communicates with the processor, since the host CPU needs to copy the data in one memory to the host cache (Cache), and then copy it from the host cache to another memory, and a cache line of the host cache can only be mapped to Local DRAM or CXL DRAM, the utilization rate of the host cache is not high, and the cache needs to be cached twice during copying, resulting in slow data copying speed and low copying efficiency between the two memories. The above-mentioned memory that communicates with the processor through the memory interface can be a DRAM that communicates with the processor through a memory interface, which can be called a host-side memory or a host memory, etc., and is recorded as Local DRAM in this article; the above-mentioned memory in the CXL device can be a DRAM in a CXL device (such as a type 2 or type 3 CXL device), which can be called a device-side memory, and is recorded as CXL DRAM in this article.
[0039] In addition to data copying between Local DRAM and CXL DRAM, the above problem also exists in other scenarios of copying data from a memory connected to the host to another memory connected to the host.
[0040] To this end, an embodiment of the present disclosure provides a data processing method, which is applied to a processor that can access a cache, a first memory, and a second memory. The method includes: when it is determined that a data block stored in a first address of the first memory needs to be copied to a second address of the second memory, Figure 1 As shown, the following data copy operations can be performed:
[0041] S210: searching for the data block in the target cache area to which the first address is mapped;
[0042] S220: if the data block is not found, copy the data block from the first memory to the first cache line in the target cache area, set the value of the first tag of the first cache line to the tag value Tag1V corresponding to the first address, and set the first tag to be valid;
[0043] S230: When the second address is also mapped to the target cache area, the value of the second tag of the first cache line is set to the tag value Tag2V corresponding to the second address, and the second tag is set to be valid.
[0044] The processor of the disclosed embodiment of the present invention may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a microprocessor, etc., or other conventional processors, etc.; the processor may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), discrete logic or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components; or a combination of the above devices. That is, the processor of the above embodiment may be any processing device or device combination that implements the methods, steps and logic block diagrams disclosed in the embodiments of the present invention. If the embodiments of the present invention are partially implemented in software, the instructions for the software may be stored in a suitable non-volatile computer-readable storage medium, and one or more processors may be used to execute the instructions in hardware to implement the methods of the embodiments of the present invention.
[0045] The first memory and the second memory in the text are used to represent the two memories involved in the copy operation, respectively, to distinguish them from other possible memories. The first cache line and the second cache line, the third cache line, etc. mentioned later represent the cache lines where the data blocks are written when the corresponding operation is performed, such as the idle cache lines. When all the cache lines of the cache group have data blocks, the data blocks of some cache lines can be replaced according to the set algorithm and then new data blocks can be written.
[0046] In this embodiment, the processor copies the data in the first memory to the second memory in units of data blocks. The size of a data block is equal to the size of a cache line, such as 64 types but not limited thereto, and may also be 32 types, 128 types, 256 types, etc. This embodiment takes the copying of a data block as an example. In the copying operation, the data copied once may include multiple data blocks. At this time, the copying of each data block may be implemented according to the above steps. The data copied once may also be smaller than a data block. Since the minimum unit of the cache is a data block, it is also necessary to copy the data block where the data is located.
[0047] In this embodiment, the cache and processor are both cache and processor on the host side. The first memory and the second memory are two different memories connected to the processor, which can be a memory on the host side or a memory on the device side, and the present disclosure does not limit their types. In an example of the present embodiment, the first memory is a memory connected to the processor through a memory interface, or a memory in a CXL device connected to the processor through a CXL interface; the second memory can also be a memory connected to the processor through a memory interface, or a memory in a CXL device connected to the processor through a CXL interface (the CXL device can be a CXL device of type 2 or type 3 including a memory). For example, the first memory is a DRAM (denoted as LocalDRAM) connected to the processor through a memory interface, and the second memory is a DRAM (denoted as CXL DRAM) in a CXL device; for another example, the first memory is a CXL DRAM, and the second memory is a Local DRAM; the first memory and the second memory can also be different Local DRAMs, or different CXL DRAMs, and so on. The present disclosure does not limit the type of memory, and in addition to DRAM, it can also be other types of memory such as SSD.
[0048] The above method of the present disclosure is not limited to being applied to the processor of the host. In other embodiments of the present disclosure, the above cache and processor are cache and processor of the device end communicating with the host, such as cache and processor in a CXL device, and the cache and processor can be integrated on the control chip of the CXL device. The first memory and the second memory are two different memories connected to the control chip in the CXL device, and the connection between the control chip and the memory can be a CXL interface or other types of memory interfaces.
[0049] In an example of this embodiment, the memory address includes a tag field, an index field and an offset field. The tag in the tag field can be represented by the area number of the storage area to which the address belongs, and the index in the index field can represent the position of the address segment in the storage area. One address segment can store one data block, and the index can be represented by the block number. For example, a storage area includes 4 address segments, and the block number can be 00, 01, 10 and 11, and the value of the index can also be one of 00, 01, 10 and 11. When the number of address segments in the storage area is equal to the number of cache groups in the cache, the index of the index field in the memory address corresponds to the index of the cache group in the cache one by one, and the target cache group to which the memory address is mapped can be determined according to the index in the index field. The offset in the offset field can represent the position of the data to be accessed in the data block determined by the tag and the index, such as which of the 64 bytes included in a data block is the first byte of the data to be accessed. Assume that the first address of the first memory is composed of a tag Tag1V, an index Index1V and an offset Offset1V; the second address of the second memory is composed of a tag Tag2V, an index Index2V and an offset Offset2V; where Tag1V, Tag2V are the values of the tag in the address, Index1V, Index1V are the values of the index in the address, and Offset1, Offset2 are the values of the offset in the address. In addition to caching a data block, each cache line in the host cache is also set with two tags and two valid bits, see Figure 4 The initial values of the two marks, namely the first mark and the second mark, can be set to a default invalid value. One of the two valid bits indicates whether the first mark is valid, and the other indicates whether the second mark is valid.
[0050] In this embodiment, the data block stored in the first address of the first memory is copied to the second address of the second memory, which means that the first address and the second address store the same data block. If the first address and the second address are both mapped to the same cache group, the same data blocks stored in the two memories can share a cache line. In the case where the index in the address corresponds one-to-one with the index of the cache group, when Index1V in the first address is equal to Index2V in the second address, a cache line can be selected in the cache group (i.e., the target cache area) with index index1V to cache the data block, and the value of the first tag of the cache line is set to the value Tag1V of the tag in the first address (i.e., the tag value corresponding to the first address), and the value of the second tag is set to the value Tag2V of the tag in the second address (i.e., the tag value corresponding to the second address).
[0051] In an exemplary embodiment, setting the first tag to be valid may include: setting the valid bit (referred to as the first valid bit) corresponding to the first cache line and the first tag to a value indicating that the tag is valid; setting the second tag to be valid may include: setting the valid bit (referred to as the second valid bit) corresponding to the first cache line and the second tag to a value indicating that the tag is valid. Since the data in the memory and the cache may be inconsistent due to the modification operation, the first valid bit of a cache line indicates that the first tag is valid, which means that the data block stored in the cache is consistent with the data block stored in the corresponding address of the first memory, and the first valid bit indicates that the first tag is invalid, which means that the data block stored in the cache line is inconsistent with the data block stored in the corresponding address of the first memory; and the second valid bit of a cache line indicates that the second tag is valid, which means that the data block stored in the cache line is consistent with the data block stored in the corresponding address of the second memory, and the second valid bit indicates that the second tag is invalid, which means that the data block stored in the cache line is inconsistent with the data block stored in the corresponding address of the second memory. The initial states of the first valid bit and the second valid bit can be set to values representing that the corresponding tags are invalid.
[0052] The situation in which the data block is not found in the above step S210 represents a cache miss. For example, when searching for a cache group with an index of Index1 in the cache according to the tag Tag1V of the first address, the index Index1V and the offset Offset1V, if the first tags of all cache lines in the cache group are not equal to Tag1V, it is determined that the data block is not found, and the cache miss occurs. If a cache line with a first tag equal to Tag1V is found in the cache group, it is determined that the data block is found, that is, a cache hit occurs.
[0053] In this embodiment, the data blocks in the two memories can share a cache line, thereby improving the utilization rate of the cache. In addition, during the data copying process, in the case of a cache miss, it only needs to be cached once, and there is no need to write the data block to the two cache lines in the cache. This can speed up the data copying between the two storage devices and improve the operation efficiency.
[0054] In an exemplary embodiment, after searching for a data block in the target cache area, the data processing method may further include: when the data block is found and the second address is also mapped to the target cache area, setting the value of the second tag of the first cache line to Tag2V and setting the second tag to valid; wherein, finding the data block means finding a cache line in the target cache area whose first tag is valid and whose value is equal to Tag1V. In this embodiment, finding the data block is a cache hit, indicating that the data block stored in the first address of the first memory has been written into the cache, so there is no need to read the data block from the first memory and write it into the cache. When the second address is also mapped to the target cache area (such as the value of the index in the second address is equal to the index of the cache group where the data block is located), it is only necessary to set the second tag of the cache line to the tag in the second address and set the second tag to valid (such as setting the second tag to 1), so that the cached data block is the data block stored in the second address of the second memory.
[0055] In an exemplary embodiment, the data copy operation further includes: searching for a cache line in the target cache area whose second tag is valid and whose value is equal to Tag2V, and invalidating the second tag of the cache line found. When performing the data copy operation, the data block previously stored in the second address of the second memory may have been written into the cache, so the cache line in the target cache area whose second tag is valid and whose value is equal to Tag2V is searched to determine whether the data block stored in the second address of the second memory has been cached. If found, the second tag of the cache line found is invalidated, so that at the same time, only one cache line in the cache has a valid second tag and a value equal to Tag2V, so as to meet the data consistency requirement. If the cache line with a valid second tag and a value equal to Tag2V is not found, the value of the second tag of the first cache line can be directly set to Tag2V and set to valid. The operation of searching for a cache line in the target cache area whose second tag is valid and whose value is equal to Tag2V can be performed before or after the operation of setting the value of the second tag of the first cache line to Tag2V and setting the second tag to valid, and the present disclosure is not limited to this.
[0056] In an exemplary embodiment, after the value of the second tag of the first cache line is set to Tag2V and the second tag is set to valid, the data copy operation is completed; the method further includes: before the data block is replaced from the first cache line or the first cache line is released, the data block is copied to the second address of the second memory. That is to say, when the operation of copying the data block stored in the first address of the first memory to the second address of the second memory is executed, the data block does not have to be written to the second address of the second memory immediately, and the data copy operation can be terminated first. In this way, when the data block to be copied is cache hit, the entire copy operation only needs to perform the modification of the first tag and the second valid bit of the first cache line, which greatly speeds up the data copy speed and improves the processing efficiency. After the data copy is completed, the cached data block can be modified in response to the system software, and before determining to replace the data block from the first cache line or release the first cache line, the data block is copied to the second address of the second memory to reduce the access to the memory.
[0057] In an exemplary embodiment, after completing the data copy operation, the data processing method may further include: receiving a request to modify the data block stored in the first address of the first memory; finding a first cache line in the target cache area in which the value of the first tag is equal to Tag1V and both the first tag and the second tag are valid; setting the first tag of the first cache line to invalid; and writing the modified value of the data block to a second cache line in the target cache area, setting the value of the first tag of the second cache line to Tag1V, and setting the first tag to valid; wherein the second cache line is different from the first cache line.
[0058] In another exemplary embodiment, after completing the data copy operation, the data processing method may further include: receiving a request to modify the data block that has been copied to the second address of the second memory; finding a first cache line in the target cache area in which the value of the second tag is equal to Tag2V and both the first tag and the second tag are valid; setting the second tag of the first cache line to invalid; and writing the modified value of the data block to a third cache line in the target cache area, setting the value of the second tag of the third cache line to Tag2V and setting the second tag to valid; the third cache line is different from the first cache line.
[0059] In the above two embodiments, the request to modify the data block may be sent by the system software or other host. In the case where the data block to be modified is in the cache and is shared by two memories, the modification of the data block in one of the memories cannot affect the consistency between the data block in the cache and the data block in the other memory. Therefore, it is necessary to invalidate one of the flags of the cache line where the data block is located (the other flag remains valid at this time), and select another cache line to store the modified value of the data block. In addition, in the above two embodiments, the initial value of the second flag bit of the second cache line indicates that the second flag is invalid, and there is no need to invalidate the second flag through additional operations; similarly, the initial value of the first flag bit of the third cache line indicates that the first flag is invalid, and there is no need to invalidate the first flag through additional operations.
[0060] In an exemplary embodiment, a multi-way group associative mapping method is used between the address of the first memory and the address of the cache, and the target cache area to which the first address is mapped is a cache group in the cache to which the first address is mapped; or a full associative mapping method is used between the address of the first memory and the address of the cache, and the target cache area to which the first address is mapped is all cache lines in the cache;
[0061] A multi-way group associative mapping method is used between the address of the second memory and the address of the cache, and the target cache area to which the second address is mapped is a cache group in the cache to which the second address is mapped; or a fully associative mapping method is used between the address of the second memory and the address of the cache, and the target cache area to which the second address is mapped is all cache lines in the cache.
[0062] The fully associative mapping method can be regarded as a special multi-way group associative mapping method, that is, a multi-way group associative mapping method when the number of cache groups is equal to 1. This cache group includes all cache lines in the cache. It can also be said that the memory address is mapped to all cache lines in the cache, and there is no need to use an index to determine the target cache group. The memory address under the fully associative mapping method may not have an index field set, and the memory address may be composed of a tag field and an offset field, but the data processing method of the disclosed embodiment can still be applicable.
[0063] In an exemplary embodiment, at least two mapping modes may be used between the address of the first memory and the address of the cache:
[0064] The first type: a multi-way group-connected mapping method is used between the address of the first memory and the address of the cache, and the first address includes a tag, an index, and an offset; wherein the value of the tag in the first address is equal to Tag1V, and the target cache area mapped to the first address is a cache group whose index value in the cache is equal to the index value in the first address;
[0065] The second type: a fully associative mapping method is adopted between the address of the first memory and the address of the cache, and the first address includes a tag and an offset; wherein: the value of the tag in the first address is equal to Tag1V, and the target cache area mapped to the first address is all cache lines in the cache.
[0066] There are at least two mapping modes between the address of the second memory and the address of the cache:
[0067] The first type: a multi-way group-connected mapping method is used between the address of the second memory and the address of the cache, and the second address includes a tag, an index, and an offset; wherein the value of the tag in the second address is equal to Tag2V, and the target cache area mapped to the second address is a cache group whose index value in the cache is equal to the index value in the second address;
[0068] The second type: a fully associative mapping method is adopted between the address of the second memory and the address of the cache, and the second address includes a tag and an offset; wherein: the value of the tag in the second address is equal to Tag2V, and the target cache area mapped to the second address is all cache lines in the cache.
[0069] The mapping method between the address of the first memory and the address of the cache may be the same as or different from the mapping method between the address of the second memory and the address of the cache. When a multi-way group associative mapping method is used between the memory address and the cache address, the tag and index in a memory address (such as the first address and the second address) can locate the storage address of a data block. When a fully associative mapping method is used between the memory address and the cache address, the tag in a memory address can locate the storage address of a data block. The offset value in the above-mentioned first address and second address represents the offset of the data to be copied in the data block.
[0070] In an exemplary embodiment, the data copy operation may also include: when the second address is mapped to another cache area different from the target cache area, copying the data block to a fourth cache line in the other cache area, setting the value of the second tag of the fourth cache line to the tag value Tag2V corresponding to the second address, and setting the second tag to valid. Because the second tag of the first cache line and the first tag of the fourth cache line can be invalid by default, no additional operation is required. In this embodiment, when the first address and the second address are mapped to different target cache groups, the copied data block is not shared in the cache. At this time, two cache lines can be selected to cache the data block, one cache line is mapped to the data block in the first memory, and the other cache line is mapped to the data block in the second memory. By setting the corresponding tag to valid, it can be determined which memory the cache line has a mapping relationship with.
[0071] In an exemplary embodiment, the data processing method may further include: when it is determined that the data block stored in the first address of the first memory needs to be added to the cache, the following operations are performed: searching for the data block in the target cache area to which the first address is mapped; and, if the data block is not found, copying the data block from the first memory to the first cache line in the target cache area, setting the value of the first tag of the first cache line to the tag value Tag1V corresponding to the first address, and setting the first tag of the cache line to valid. The second tag may be invalid by default, but may also be set to invalid. This embodiment may be that when the processor reads data in the memory, the data block in the memory is written to the cache line, but the present disclosure is not limited to this, and the cache line is not shared at this time.
[0072] Some CXL devices (such as Type 2 CXL devices and Type 3 CXL devices) have a processor, a cache, and multiple memories, so the above data processing method can be applied to the CXL devices.
[0073] An embodiment of the present disclosure further provides a control chip of a CXL device, comprising a processor and a cache, wherein the processor can access the cache and a first memory and a second memory connected to the control chip, and the processor is configured to execute the data processing method described in any embodiment of the present disclosure.
[0074] An embodiment of the present disclosure further provides a CXL device, comprising the control chip described in any embodiment of the present disclosure and a first memory and a second memory connected to the control chip. Figure 3AThe figure is a schematic diagram of the structure of the CXL device of the present embodiment. The first memory and the second memory in the figure are both DRAM chips as examples, but the present disclosure is not limited to this. The processor in the control chip can access the cache, the first memory and the second memory. By executing the data processing method of the embodiment of the present disclosure, the utilization rate of the cache in the CXL device and the speed and efficiency of data copying between different memories can be improved.
[0075] An embodiment of the present disclosure further provides a host, including a processor and a cache, wherein the processor can access the cache and a first memory and a second memory connected to the host, and the processor can execute a data processing method as described in any embodiment of the present disclosure. The first memory is a memory connected to the processor via a memory interface, or a memory in a CXL device connected to the processor via a CXL interface; the second memory is a memory connected to the processor via a memory interface, or a memory in a CXL device connected to the processor via a CXL interface. Figure 3B In the example shown, the first memory is a memory that communicates with the processor through a memory interface, that is, a local memory, and the second memory is a memory in a CXL device that is connected to the processor through a CXL interface.
[0076] The embodiments of the present disclosure also provide a non-transitory computer storage medium storing a computer program. When the computer program is executed by a processor, the data processing method described in any embodiment of the present disclosure can be implemented.
[0077] The embodiment of the present disclosure also provides a computer product, including a computer program, and when the computer program is executed by a processor, the data processing method can be implemented.
[0078] The following uses examples in applications to specifically illustrate the embodiments of the present disclosure. In this example, the first memory is CXL DRAM, such as the DRAM in a Type 2 CXL device, and the second memory is Local DRAM, such as the DRAM connected to the host through a memory interface. The address of the CXL DRAM is mapped to the address of the cache using a multi-way group-connected mapping method, and the address of the Local DRAM is mapped to the address of the cache using a multi-way group-connected mapping method. For the convenience of explanation, it is assumed that the cache on the host side includes four cache groups, and each cache group includes two cache lines;
[0079] Assume that CXL DRAM is divided into four areas, and the area number is represented by Tag1; each area includes four blocks, each block can store a data block, and the block number is represented by Index1; Local DRAM is divided into four areas, and the area number is represented by Tag2; each area also includes four blocks, and the block number is represented by Index2. The value Tag1V of Tag1 and the value Tag2V of Tag2 can be 00, 01, 10, 11; the value Index1V of Index1 and the value Index2V of Index2 can be 00, 01, 10, 11. In this example, the addresses of Local DRAM and CXL DRAM are both represented by tag + index + offset, where the tag is the area number, the index is the block number and corresponds to the index of the cache group one by one. The offset can indicate the starting position of the data in the data block, such as the byte number. In this embodiment, data copy takes a data block of the size of a cache line as the minimum unit. According to the area number and block number in each address, a data block can be located, and the data block is called the data block stored in the address. The indexes of the four cache groups of the cache (also called group numbers) are represented by Index, and the value of IndexV can be 00, 01, 10, 11. Each cache group has two cache lines, and each cache line has two tags and two valid bits, namely: the first tag, the second tag, the first valid bit, and the second valid bit. Figure 5 When the first valid bit is 1, it indicates that the first flag is valid, and when the second valid bit is 1, it indicates that the second flag is valid. The initial values of the first flag bit and the second flag bit are both 0, indicating that the corresponding flag is invalid.
[0080] The cache line in the cache can be shared by CXL DRAM and Local DRAM, that is, the addresses of CXL DRAM and Local DRAM can be mapped to the same cache line. Assume that the address of CXL DRAM is called the first address, which is composed of Tag1V, Index1V and the first offset value; the address of Local DRAM is called the second address, which is composed of Tag2V, Index2V and the second offset value. Figure 5 As shown (the offset in the address is omitted in the figure), the address of the index value Index1V=00 of the CXL DRAM (distributed in the 4 areas of the CXL DRAM) and the address of the index value Index2V=00 of the Local DRAM (distributed in the 4 areas of the Local DRAM) are both mapped to the cache group with the index value IndexV=00; Figure 6 As shown, the address of the index value Index1V=01 in the CXLDRAM and the address of the index value Index2V=01 in the Local DRAM are both mapped to the cache group of the index value IndexV=01. Figure 7As shown, the address of the index value Index1V=10 in the CXL DRAM and the address of the index value Index2V=10 in the Local DRAM are both mapped to the cache group of the index value IndexV=10. Figure 8 As shown, the address of the index value Index1V=11 in the CXLDRAM and the address of the index value Index2V=11 in the Local DRAM are both mapped to the cache group of the index value Index=11.
[0081] In the process of copying data between the first memory and the second memory, when it is necessary to copy a data block stored in a first address of the first memory to a second address of the second memory, the following steps may be performed, such as: Fig. 9 As shown:
[0082] S901: The processor determines that the data block stored in the first address (Tag1V=00, Index1V=00) of the CXL DRAM needs to be copied to the second address (Tag2V=10, Index2V=00) of the Local DRAM;
[0083] S902: Search for the data block in the two cache lines (Cacheline0, Cacheline1) of the target cache group (IndexV=00) mapped to the first address of the CXL DRAM; if not found, execute S903; if found, execute S904;
[0084] In this step, if there is a cache line in the cache group with IndexV=00 whose first tag is equal to Tag1V and is valid, the data block is found; if there is no cache line whose first tag is equal to Tag1V and is valid, the data block is not found.
[0085] S903: copy the data block stored in the first address to a cache line (Cacheline0 or Cacheline1) of the target cache group, set the value of the first tag of the cache line to the tag value Tag1V corresponding to the first address, that is, set the value of the first tag to 00, and set the first valid bit to 1;
[0086] S904: Set the value of the second tag of Cacheline0 or Cacheline1 to the tag value Tag2V corresponding to the second address, that is, set the value of the second tag to 10, and set the second valid bit to 1;
[0087] S905: Determine whether the second tag of another cache line (Cacheline1 or Cacheline0) in the target cache group (IndexV=00) is equal to Tag2V and valid: if yes, execute step S906; if no, execute step S907;
[0088] S906, setting the second flag bit of the other cache line to 0;
[0089] S907: End the copy operation on the data block.
[0090] After completing the copy operation on the data block, if the cache line (Cacheline0 or Cacheline1) storing the data block needs to be replaced or released, the data block can be first copied from Cacheline0 or Cacheline1 to the second address (Tag2V=10, Index2=00) of LocalDRAM.
[0091] After completing the copy operation on the data block, if it is necessary to modify the data block stored in the first address of the first memory, the following steps may be performed:
[0092] S1001: The processor receives a modification request: modify the data block stored in the first address (Tag1V=00, Index1V=00) of the CXL DRAM;
[0093] S1002: Find a cache line whose first tag value is equal to Tag1V (that is, the value of the first tag is equal to 00) and whose first valid bit and second valid bit are both 1 in the two cache lines (Cacheline0, Cacheline1) of the Cache target cache group (IndexV=00) mapped to the first address (Tag1V=00, Index1V=00) of the CXL DRAM, and assume that the cache line is Cacheline0;
[0094] S1003: Set the first valid bit of Cacheline0 to 0; write the modified value of the data block to Cacheline1; set the value of the first tag of Cacheline1 to Tag1V, that is, 00; set the first valid bit of Cacheline1 to 1, and end the modification operation.
[0095] After completing the copy operation on the data block, if it is necessary to modify the data block stored in the second address of the second memory, the following steps may be performed:
[0096] S1101: The processor receives a modification request: the data block in the second address (Tag2V=00, Index2V=00) of the Local DRAM is modified;
[0097] S1102: Search for a cache line whose second tag value is equal to Tag2V (i.e., the second tag value is equal to 00) and whose first valid bit and second valid bit are both 1 in the two cache lines (Cacheline0, Cacheline1) of the Cache target cache group (IndexV=00) mapped to the second address (Tag2V=00, Index2V=00) of the Local DRAM, and whose cache line is assumed to be Cacheline0;
[0098] S1103: Set the second significant bit of Cacheline0 to 0; write the modified value of the data block to Cacheline1; set the value of the second tag of Cacheline1 to Tag2V, that is, 00; set the second significant bit of Cacheline1 to 1, and end the modification operation.
[0099] Those skilled in the art will appreciate that the functional modules / units in all or some steps, systems, and devices disclosed in the above method can be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
Claims
1. A data processing method, applied to a processor capable of accessing a cache, a first memory, and a second memory, the method include: When it is determined that the data block stored in the first address of the first memory needs to be copied to the second address of the second memory, the following data copy operation is performed: Searching for the data block in the target cache area to which the first address is mapped; If the data block is not found, copy the data block from the first memory to a first cache line in the target cache area, set the value of a first tag of the first cache line to a tag value Tag1V corresponding to the first address, and set the first tag to be valid; In the case where the second address is also mapped to the target cache area, the value of the second tag of the first cache line is set to the tag value Tag2V corresponding to the second address, and the second tag is set to be valid.
2. The method according to claim 1, Features: After searching for the data block in the target cache area, the method further includes: when the data block is found and the second address is also mapped to the target cache area, setting the value of the second tag of the first cache line to Tag2V and setting the second tag to valid; wherein, finding the data block means finding a cache line in the target cache area whose first tag is valid and has a value equal to Tag1V.
3. The method according to claim 1 or 2, Features: The data copy operation further includes: searching for a cache line in the target cache area whose second tag is valid and whose value is equal to Tag2V, and invalidating the second tag of the found cache line.
4. The method according to claim 1 or 2, Features: After setting the value of the second tag of the first cache line to Tag2V and setting the second tag to valid, the data copy operation is completed; The method further includes: before replacing the data block from the first cache line or releasing the first cache line, copying the data block to a second address of the second memory.
5. The method according to claim 1, It is characterized in that After completing the data copy operation, the method further includes: receiving a request to modify the data block stored in the first address of the first memory; Finding in the target cache area a first cache line in which the value of the first tag is equal to Tag1V and both the first tag and the second tag are valid; Setting the first tag of the first cache line to invalid; and, writing the modified value of the data block to the second cache line of the target cache area, setting the value of the first tag of the second cache line to Tag1V, and setting the first tag to valid; wherein the second cache line is different from the first cache line.
6. The method according to claim 1, It is characterized in that After completing the data copy operation, the method further includes: receiving a request to modify the data block at a second address of the second memory; Finding in the target cache area a first cache line where the value of the second tag is equal to Tag2V and both the first tag and the second tag are valid; Setting the second tag of the first cache line to invalid; and, writing the modified value of the data block to a third cache line of the target cache area, setting the value of the second tag of the third cache line to Tag2V, and setting the second tag to valid; wherein the third cache line is different from the first cache line.
7. The method according to claim 1, Features: The setting the first tag to be valid includes: setting a valid bit corresponding to the first tag and set for the first cache line to a value indicating that the tag is valid; The setting the second tag to be valid includes: setting a valid bit corresponding to the second tag set for the first cache line to a value indicating that the tag is valid.
8. The method according to claim 1, Features: A multi-way group-associative mapping method is used between the address of the first memory and the address of the cache, and the target cache area to which the first address is mapped is a cache group in the cache to which the first address is mapped; or a full-associative mapping method is used between the address of the first memory and the address of the cache, and the target cache area to which the first address is mapped is all cache lines in the cache; A multi-way group-associated mapping method is used between the address of the second memory and the address of the cache, and the target cache area to which the second address is mapped is a cache group in the cache to which the second address is mapped; Alternatively, a fully associative mapping method is adopted between the address of the second memory and the address of the cache, and the target cache area to which the second address is mapped is all cache lines in the cache.
9. The method according to claim 1, Features: A multi-way group-associated mapping method is used between the address of the first memory and the address of the cache, and a multi-way group-associated mapping method is used between the address of the second memory and the address of the cache; Both the first address and the second address include a tag, an index and an offset; wherein: the value of the tag in the first address is equal to Tag1V, and the value of the tag in the second address is equal to Tag2V; the target cache area mapped to the first address is a cache group whose index value in the cache is equal to the value of the index in the first address, and the target cache group mapped to the second address is a cache group whose index value in the cache group is equal to the value of the index in the second address.
10. The method according to claim 1, It is characterized in that The method further includes: when it is determined that the data block stored at the first address of the first memory needs to be added to the cache, performing the following operations: Searching for the data block in the target cache area to which the first address is mapped; If the data block is not found, the data block is copied from the first memory to a cache line in the target cache area, the value of the first tag of the cache line is set to the tag value Tag1V corresponding to the first address, and the first tag of the cache line is set to be valid.
11. A control chip of a CXL device, comprising a processor and a cache, wherein the processor can access the cache and a first memory and a second memory connected to the control chip, It is characterized in that The processor is configured to execute the data processing method according to any one of claims 1 to 10.
12. A CXL device, It is characterized in that The invention comprises the control chip as claimed in claim 11 and a first memory and a second memory connected to the control chip.
13. A host, comprising a processor and a cache, wherein the processor can access the cache and a first memory and a second memory connected to the host, It is characterized in that The processor can execute the data processing method according to any one of claims 1 to 10.
14. The host according to claim 12, Features: The first memory is a memory communicating with the processor via a memory interface, or a memory in a CXL device communicating with the processor via a CXL interface; The second memory is a memory communicating with the processor through a memory interface, or a memory in a CXL device communicating with the processor through a CXL interface.
15. A non-transitory computer storage medium storing a computer program, It is characterized in that When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 10 can be implemented.
16. A computer product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 10 can be implemented.
Citation Information
Cited By
Processor and data caching method thereof, electronic equipment and storage medium
CN121029641A