Memory management method and device and related equipment
By setting prediction identifiers at the memory area level and updating them in real time, the problem of large storage resource overhead in bandwidth compression technology is solved, the prediction accuracy and robustness are improved, and more efficient memory management is achieved.
Patent Information
- Application Number
- CN202410035657.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-11
AI Technical Summary
The existing bandwidth compression technology has problems such as large storage resource overhead, low prediction accuracy, poor robustness and lack of dynamic adjustment mechanism in memory management.
By setting prediction identifiers in units of memory areas, the data compression state in the memory area is determined based on the prediction identifier, and the prediction identifier is updated in real time to reduce the storage resource overhead of metadata cache and improve prediction accuracy.
While ensuring prediction accuracy, the overhead of on-chip or off-chip storage resources is reduced, and the efficiency and flexibility of bandwidth compression technology is improved.
Smart Images

Figure CN120295743A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of memory access technologies, and in particular, to a memory management method, apparatus, and related devices. Background Art
[0002] Bandwidth compression technology is an effective method for solving the memory bandwidth bottleneck of computer systems. By compressing data, the bandwidth compression technology enables the memory system to transfer data through fewer pins and fewer memory chips, thereby increasing the bandwidth. The mainstream compression methods mainly perform data compression through efficient compression algorithms to improve the bandwidth and capacity of the main memory.
[0003] However, the memory compression technology itself has overheads, including metadata overheads. That is, bandwidth compression requires using metadata to describe the compressed data, and the size of the metadata is usually related to the compression algorithm and the compression ratio; latency overheads, that is, bandwidth compression requires transmitting the compressed data and metadata between the memory controller and the processor, which will increase the transmission latency. In addition, decompression also requires a certain amount of time, which will increase the latency of accessing the memory. Therefore, when designing the bandwidth compression technology, factors such as the compression ratio, metadata size, transmission latency, and decompression latency need to be comprehensively considered to balance the effects and side effects of bandwidth compression.
[0004] Currently, the methods of bandwidth compression mainly use metadata to manage the compressed data. Since the scale of the metadata corresponding to all memory data is large, it is generally stored in an independent space such as Double Data Rate (DDR) memory. The role of the Metadata Cache (MC) set on the System on Chip (SOC) (such as in the memory controller) is mainly to reduce the number of accesses to the main memory and the latency. Since the metadata is usually smaller than the data, the metadata can be cached in the cache, that is, the metadata cache MC, to reduce the number of accesses to the main memory. In addition, the metadata cache MC can also reduce the latency of accessing the main memory because reading data from the cache is faster than reading data from the main memory. By reducing the number of accesses to the main memory and the latency, the metadata cache can improve the performance and energy efficiency of the memory system.
[0005] However, the metadata cache MC also needs to occupy a certain amount of storage space to store the cached metadata, and due to the large number of memory blocks and on-chip caches in the main memory, the capacity overhead of the metadata cache may be very large. For example, a 16GB main memory system requires 32MB of metadata, which may occupy a large amount of storage space. Summary of the Invention
[0006] The embodiments of this application provide a memory management method, apparatus, and related devices, which reduce the overhead of the bandwidth compression technology on storage resources.
[0007] In a first aspect, an embodiment of the present application provides a memory management method, which may include: receiving a memory access request for a target memory unit, and determining an identifier of a target memory area to which the target memory unit belongs; querying a target prediction identifier corresponding to the identifier of the target memory area from a status prediction table, and determining a predicted compression status of data in the target memory unit based on the target prediction identifier; the status prediction table includes a mapping relationship between the identifier of the memory area and the prediction identifier; wherein, the memory area includes a plurality of memory units, and the prediction identifier is used to predict whether the data in the plurality of memory units is in a compressed state; reading target data in the target memory unit based on the predicted compression status, and determining an actual compression status of the target data; updating the target prediction identifier based on the predicted compression status and the actual compression status.
[0008] Embodiments of the present application provide a memory management method, which is applied to the scenario of bandwidth compression technology. By setting corresponding prediction identifiers in units of memory regions and determining the predicted compression states of all memory cells within the memory region based on the prediction identifiers, the compression states of a larger range of data can be predicted with fewer prediction identifiers, thereby reducing the storage resource overhead caused by bandwidth compression. Further, after reading a target data within the memory region based on the predicted compression state, the actual compression state of the target data can be obtained, and finally, the prediction identifier corresponding to the memory region can be updated according to the predicted compression state and the actual compression state of the target data to improve the accuracy of compression prediction in units of memory regions, thereby reducing the storage resource overhead of the bandwidth compression technology on the premise of ensuring the prediction accuracy. Specifically, corresponding prediction identifiers are set in units of memory regions in the state prediction table, and the prediction identifiers are used to uniformly predict whether the data in all memory cells within the memory region is in a compressed state, that is, multiple memory cells within the memory region (such as multiple memory cells of the memory access granularity size) share the same prediction identifier, thus greatly saving the metadata originally required for each of the multiple memory cells. When applied to an on-chip system or an off-chip memory system, the storage resource overhead on the chip or off-chip can be greatly reduced. Further, since multiple memory cells share the same prediction identifier, and embodiments of the present application are based on the assumption (theory) that the compression states of the data stored in the same memory region usually have a certain continuity, in other words, there is a strong correlation, especially a strong correlation with the compression state of the data read most recently (such as the previous time). Therefore, embodiments of the present application can perform a reasonable update (such as updating the prediction state and / or prediction probability) on the current prediction identifier based on the relationship between the actual compression state and the prediction state of the data read last time, so that the prediction identifier can more accurately predict the compression state of the data to be read next time. Although, in embodiments of the present application, the prediction identifier cannot completely accurately represent the actual compression state of the data, it can predict the data compression states of a large number of memory cells with the prediction identifiers of a small number of memory regions, and the prediction identifier of the memory region where the data is located can be updated in real time with each data reading to ensure the prediction accuracy, thereby achieving accurate prediction of the compression states of flexible-size regions with a small amount of storage resources. In addition, embodiments of the present application can also flexibly support the metadata cache (MetadataCache) structure of previous work, and can also cooperate with the MetadataCache to achieve more accurate compressed data management when there is enough space. Compared with the prior art, embodiments of the present application can greatly reduce the data space overhead of a memory access system such as an on-chip system (SOC) or an off-chip memory system on the premise of ensuring the prediction accuracy of the compression state (adjusted in real time according to the actual prediction result).
[0009] In a possible implementation, the size of the memory unit is the size of the memory access granularity; the method further includes: receiving a memory access request, where the memory access request includes a starting address and an access length; determining a target memory unit based on the starting address and the access length, where the target memory unit includes one or more memory units. In the embodiments of the present application, when a request for accessing a certain address in the memory issued by the processor carries the starting address and the access length of the data to be accessed, based on this, the memory unit corresponding to the memory access address space can be determined (essentially determining the physical address in the memory), and finally, the memory area to which it belongs is determined according to the determined target memory unit. It can be understood that the target memory unit may correspond to one memory unit or may correspond to multiple memory units, depending on the length of the data to be accessed. The embodiments of the present application do not make specific limitations on this.
[0010] In a possible implementation, the type of the prediction identifier includes at least two groups of identifiers, where each group of identifiers is used to indicate the prediction probability that the data is in a compressed state or an uncompressed state, and the prediction probabilities corresponding to the same group of identifiers are the same, and the prediction probabilities corresponding to different groups of identifiers are different.
[0011] In the embodiments of the present application, the type of the prediction identifier may include at least two groups of identifiers, where each group of identifiers is used to indicate the compressed state and the uncompressed state of the memory area under the same prediction probability, and the prediction probabilities corresponding to different groups are different. For example, one group of identifiers is 0 and 3. Specifically, 0 is used to indicate that the prediction probability that the data in the memory area is in an uncompressed state is 100%, and 3 is used to indicate that the prediction probability that the data in the memory area is in a compressed state is 100%; another example is that another group of identifiers is 1 and 2, where 1 is used to indicate that the prediction probability that the data in the memory area is in an uncompressed state is 50%, and 2 is used to indicate that the prediction probability that the data in the memory area is in a compressed state is 50%. Optionally, the prediction identifier may further include more groups of identifiers. For example, in addition to the above 0, 1, 2, and 3, there may be another group of identifiers 4 and 5, where 4 is used to indicate that the prediction probability that the data in the memory area is in an uncompressed state is 80%, and 5 is used to indicate that the prediction probability that the data in the memory area is in a compressed state is 80%. The prediction identifier in the embodiments of the present application includes multiple groups of identifiers, which can greatly improve the accuracy and robustness of the prediction. It can be understood that the embodiments of the present application do not make specific limitations on the specific type and quantity of the prediction identifier.
[0012] In a possible implementation, the at least two groups of identifiers include a first group of identifiers and a second group of identifiers. The first group of identifiers are respectively an uncompressed identifier and a compressed identifier, and the second group of identifiers are respectively a suspected uncompressed identifier and a suspected compressed identifier. The prediction probability corresponding to the first group of identifiers is higher than the prediction probability corresponding to the second group of identifiers. In the embodiments of the present application, after determining the target prediction identifier corresponding to the target memory unit through the mapping relationship between the identifier in the memory area of the state prediction table and the prediction identifier, the compression state of the data in the target memory unit is determined based on the target prediction identifier. For example, if the prediction probability that the target prediction identifier currently indicates that the data is in an uncompressed state is 100% (i.e., corresponding to the uncompressed identifier), the predicted compression state of the target memory unit is determined to be an uncompressed state according to the preset prediction rule. Another example is that if the prediction probability that the target prediction identifier currently indicates that the data is in a compressed state is 100% (i.e., corresponding to the compressed identifier), the predicted compression state of the target memory unit is determined to be a compressed state according to the preset prediction rule. Still another example is that if the prediction probability that the target prediction identifier currently indicates that the data is in an uncompressed state is 50% (i.e., corresponding to the suspected uncompressed identifier), the current actual requirements of the memory access system (such as a system-on-chip SOC or an off-chip memory management system, etc.) need to be further determined according to the preset prediction rule to determine whether the predicted compression state of the target memory unit is a compressed state or an uncompressed state. For example, the compression state corresponding to the current target prediction identifier is determined according to the current requirements of the memory access system for bandwidth sensitivity or latency sensitivity.
[0013] In a possible implementation, determining the predicted compression state of the data in the target memory unit based on the target prediction identifier includes: if the target prediction identifier is the compressed identifier, determining that the predicted compression state of the data in the target memory unit is a compressed state; if the target prediction identifier is the uncompressed identifier, determining that the predicted compression state of the data in the target memory unit is an uncompressed state. In the embodiments of the present application, the predicted compression state of the target memory unit is predicted based on the target prediction identifier. When the prediction identifier includes four identifiers: an uncompressed identifier, a suspected uncompressed identifier, a suspected compressed identifier, and a compressed identifier, for the uncompressed identifier or the compressed identifier among the above four identifiers, the compression or uncompression state of the memory area can be directly predicted based on these two identifiers.
[0014] In a possible implementation, the method is applied to a memory access system; determining the predicted compression state of the data in the target memory unit based on the target prediction identifier includes: if the current target prediction identifier is the suspected compression identifier or the suspected non-compression identifier, then determining the relationship between the current latency sensitivity and bandwidth sensitivity of the memory access system; when the memory access system is more sensitive to bandwidth, determining the predicted compression state of the data in the target memory unit as the compression state; when the memory access system is more sensitive to latency, determining the predicted compression state of the data in the target memory unit as the non-compression state. In the embodiments of the present application, on the premise that the prediction identifiers include the non-compression identifier, the suspected non-compression identifier, the suspected compression identifier, and the compression identifier, among the above four identifiers, the compression states corresponding to the non-compression identifier or the compression identifier can be set to be directly determinable, while the compression states corresponding to the suspected non-compression identifier or the suspected compression identifier are set to require further consideration of the actual requirements of the current memory access system (such as a system on chip (SOC) or an off-chip memory management system, etc.) to be determined. For example, when the target prediction identifier is the suspected compression identifier or the suspected non-compression identifier, the actual requirements of the memory access system for the system are further determined, and then the compression state of the data in the target memory unit is determined based on the target prediction identifier.
[0015] In a possible implementation, updating the target prediction identifier based on the predicted compression state and the actual compression state includes: if both the predicted compression state and the actual compression state are non-compression states, and the target prediction identifier is currently the non-compression identifier, then keep the target prediction identifier unchanged; if both the predicted compression state and the actual compression state are non-compression states, and the target prediction identifier is currently the suspected non-compression identifier or the suspected compression identifier, then adjust the target prediction identifier in the direction of the non-compression identifier. Since in the embodiments of the present application, the predicted identifier corresponding to the memory area in the state prediction table is updated based on the actual compression state of the read data (i.e., the actual compression state), that is, the actual compression state (i.e., compressed or non-compressed) of each piece of data read within a certain memory area can be used to reversely update the overall prediction result of this memory area, that is, the predicted identifier. When the predicted result of the compression state of the data is the same as the actual result, it means that the current prediction direction is correct. Then, if the current target prediction identifier is already a compression identifier or a non-compression identifier with a prediction probability of 100%, the current target prediction identifier can be kept unchanged. If the current target prediction identifier has a prediction probability less than 1 (such as 50% or 80%, etc.), it can be continuously adjusted (or strengthened) in the direction of the currently determined actual compression state or non-compression state. For example, if 0, 1, 2, 3 represent the non-compression identifier, the suspected non-compression identifier, the suspected compression identifier, and the compression identifier respectively, then adjusting (or strengthening) in the direction of the compression identifier means adding 1 to the current predicted identifier, and adjusting (strengthening) in the direction of the non-compression identifier means subtracting 1 from the current predicted identifier.
[0016] In a possible implementation, updating the target prediction identifier based on the predicted compression state and the actual compression state includes: if the predicted compression state is a non-compression state and the actual compression state is a compression state, then adjust the target prediction identifier in the direction of the compression identifier. In the embodiments of the present application, when the predicted result of the data compression state is inconsistent with the actual result, it is adjusted in the opposite direction of the current target prediction identifier. For example, if 0, 1, 2, 3 represent the non-compression identifier, the suspected non-compression identifier, the suspected compression identifier, and the compression identifier respectively, then adjusting (strengthening) in the direction of the compression identifier means adding 1 to the current predicted identifier.
[0017] In a possible implementation, updating the target prediction identifier based on the predicted compression state and the actual compression state includes: if the predicted compression state is the compression state and the actual compression state is the non-compression state, adjusting the target prediction identifier in the direction of the non-compression identifier. In the embodiments of the present application, when the predicted result of the data compression state is inconsistent with the actual result, it is adjusted in the opposite direction of the current target prediction identifier. For example, if 0, 1, 2, and 3 respectively represent the non-compression identifier, the suspected non-compression identifier, the suspected compression identifier, and the compression identifier, then adjusting (strengthening) in the direction of the non-compression identifier means subtracting 1 from the current prediction identifier.
[0018] In a possible implementation, updating the target prediction identifier based on the predicted compression state and the actual compression state includes: if both the predicted compression state and the actual compression state are the compression state, and the current target prediction identifier is the compression identifier, keeping the target prediction identifier unchanged; if both the predicted compression state and the actual compression state are the compression state, and the current target prediction identifier is the suspected non-compression identifier or the suspected compression identifier, adjusting the target prediction identifier in the direction of the compression identifier. In the embodiments of the present application, when the predicted result of the data compression is consistent with the actual result, it is also necessary to further determine whether the current prediction identifier is consistent with the above actual result. If they are consistent, the target prediction identifier remains unchanged. If they are inconsistent, the current prediction identifier needs to be strengthened in the opposite direction of the current actual result.
[0019] In a possible implementation, the method further includes: creating a state prediction table, and establishing a mapping relationship between the identifier of each memory area and the corresponding prediction identifier in the state prediction table; setting the initial identifier of each prediction identifier to be used to indicate the non-compression state. In the embodiments of the present application, a state prediction table is also established, and a mapping relationship between the identifier of each memory area and the corresponding prediction identifier is established in the table. Since when establishing these mapping relationships, an initial identifier needs to be set for the prediction identifier corresponding to each memory area, considering that the situation of each memory area is not clear initially, the most primitive reading method can be adopted first, that is, reading in the non-compression mode. Reading in the non-compression mode will read according to the complete physical address of the target memory unit, so as to reduce the probability of rereading due to the lack of prediction basis in the initial situation.
[0020] In a possible implementation, the reading of the target data in the target memory cell based on the predicted compression state includes: if the predicted compression state is the compressed state, parsing the first physical address corresponding to the target memory cell according to the rule of compressed data, and reading the data stored in the first physical address. In the embodiments of the present application, when the data of the target memory cell is predicted to be in the compressed state, the physical address corresponding to the target cell is parsed according to the rule of compressed data. Since the physical address corresponding to the data will necessarily become shorter after being compressed, when the data is predicted to be in the compressed state, the physical address read is smaller than the physical address corresponding to the original data.
[0021] In a possible implementation, the reading of the target data in the target memory cell based on the predicted compression state includes: if the predicted compression state is the uncompressed state, parsing the second physical address corresponding to the target memory cell according to the rule of uncompressed state, and reading the data stored in the second physical address. In the embodiments of the present application, when the data of the target memory cell is predicted to be in the uncompressed state, the physical address corresponding to the target cell is parsed according to the rule of uncompressed data. Since the physical address corresponding to the data is the same as the original data if the data is not compressed, when the data is predicted to be in the uncompressed state, the physical address read is the same as the physical address corresponding to the original data.
[0022] In a possible implementation, the determining of the actual compression state of the target data includes: after reading the target data, obtaining a compression state flag carried in the target data, where the compression state flag is used to indicate whether the target data is compressed data; determining the actual compression state of the target data based on the compression state flag. In the embodiments of the present application, the data in the memory cell itself carries a compression state flag (such as metadata) for indicating whether the data has been compressed, that is, the actual reading of the data in the memory cell can obtain the true situation of whether the data is compressed data. That is, in the embodiments of the present application, the predicted compression state is the state predicted based on the state prediction table, and the actual compression state is the state determined based on the metadata of the data itself.
[0023] In a possible implementation, the method further includes: when the predicted compression state is the compressed state and the actual compression state is the uncompressed state, rereading the target data based on the uncompressed state. In the embodiments of the present application, when the predicted compression state is the compressed state but the actual compression state is the uncompressed state, it will cause the data to be read according to the compression rule (for example, only reading part of the high bits or part of the low bits), which will cause the problem of failure to read the target data. Therefore, when it is found that there is a problem with data reading, the target data needs to be reread according to the rule of the uncompressed state.
[0024] In a possible implementation, the method is applied to a memory access system; the reading of the target data in the target memory unit based on the predicted compression state includes: when the predicted compression state is an uncompressed state and the actual compression state is a compressed state, reading the first compressed data and the second compressed data stored in the target memory unit; the memory access system includes a prefetch cache; the method further includes: decompressing the first data and the second data; reading the decompressed first data according to the memory access request, and storing the decompressed second data into the prefetch cache. In the embodiment of the present application, when it is predicted based on the state prediction table that the currently to-be-accessed data is in an uncompressed state, but the actually to-be-accessed data is in a compressed state, then at this time, it may cause unnecessary data to be read during data reading. However, considering the continuity of data access, this unnecessary redundant data can also be synchronously read into the prefetch cache (such as an on-chip cache) to avoid waste of data reading bandwidth. When a memory access request is issued again next time, the on-chip cache can be searched first and the data can be read from the cache, reducing the bandwidth resource overhead between the memory access end and the accessed memory end (on-chip system and memory). That is, the data can be read from the on-chip cache within the memory access system such as an SOC, or the upper-level memory in a multi-level memory system, instead of reading the data through the bandwidth between on-chip and off-chip, or the bandwidth between the upper-level or lower-level memories.
[0025] In a possible implementation, if the target memory unit includes two adjacent memory units, and the two adjacent memory units include a high-order memory unit and a low-order memory unit; the reading of the target memory unit according to the rule of compressed data includes: reading the low-order memory unit; the reading of the target memory unit according to the rule of normal data includes: reading the high-order memory unit and the low-order memory unit. In the embodiment of the present application, when a certain data is compressed, for example, 128B of data (i.e., the original data involves two memory units) is compressed to 64B, and assuming that each memory unit can store 64B of data, the compressed data can be stored at the low 64B address position, that is, stored in the low-order memory unit among the two adjacent memory units to achieve data compression; correspondingly, when the data needs to be read, only the low 64B needs to be read, that is, the data of one memory unit needs to be read. When a certain data does not need to be compressed, that is, when writing normally, for example, 128B of data is written into two memory units with a size of 64B each; correspondingly, when the data needs to be read, the complete data bits are read according to the rule of normal uncompressed data, that is, the high-order memory unit and the low-order memory unit need to be read simultaneously.
[0026] In a possible implementation, the method further includes: receiving a write request for data to be written, determining whether the data to be written needs to be compressed, where the size of the data to be written is M bits; if the data to be written needs to be fully compressed, compressing the M bits of the data to be written and writing the compressed data into the corresponding memory unit; if only a part of the data to be written needs to be compressed, compressing the N bits of the data to be written, writing the compressed data into the first memory unit, and writing the remaining (M - N) bits of the data to be written into the second memory unit; both M and N are integers greater than 1, and N is less than M; updating the compression status flag bits in the first memory unit and the second memory unit to a compression flag and a non-compression flag respectively. In the embodiments of the present application, when a write request for data is received, it is necessary to first determine whether the data to be written needs to be compressed. If it needs to be compressed, it is compressed first and then written into the memory. If it does not need to be compressed, it can be directly written into the memory. Among them, when compression is required, it can be divided into two cases: full compression or partial compression. When full compression is required, the data to be written is compressed as a whole and then written into the memory unit. When only partial compression is required, the part of the data to be written that needs to be compressed is compressed and written into the corresponding memory unit, while the part of the data to be written that does not need to be compressed is written without compression into another memory unit. Further, for the data to be written stored in different memory units, different compression status flag bits can be set for different memory units, that is, for the same data to be written written into different memory units, independent compression flag bits can be set. For example, for 128B of data to be written (assuming the memory unit size is 64B), if 64B of it needs to be compressed and the other 64B does not need to be compressed, the 64B of data that needs to be compressed is compressed and written as 32B of data into the memory unit, while the other 64B of data that does not need to be compressed is directly written into an adjacent memory unit. In summary, in the embodiments of the present application, partial compression of data can be achieved through the hybrid compression granularity mechanism, that is, more fine-grained and precise compression can be achieved for the same data.
[0027] In a possible implementation, the compression of the M bits of the data to be written includes: splitting the M bits of the data to be written into K-bit data and (M-K)-bit data respectively, and compressing the K-bit data and the (M-K)-bit data respectively. In the embodiments of the present application, when all the data to be written needs to be compressed, it is also possible to locally split the data to be written and compress the split data bits respectively. For example, the mixed granularity includes compressing 128B of data to 64B, or compressing two 64B into 32B respectively. Correspondingly, decompression also supports decompressing 64B into 128B, or decompressing two independent 32B into 64B respectively. The embodiments of the present application can effectively increase the compression ratio by supporting mixed granularity compression.
[0028] In a possible implementation, the method further includes: calculating the current prediction accuracy rate of the state prediction table based on the predicted compression state and the actual compression state of multiple memory access data; when the data to be written needs to be compressed and the prediction accuracy rate is lower than a preset threshold, controlling the data to be written to be written into the corresponding memory unit in an uncompressed state; when the data to be written needs to be compressed and the prediction accuracy rate is higher than a preset threshold, controlling the data to be written to be written into the corresponding memory unit in a compressed state. In the embodiments of the present application, the current prediction accuracy rate of the state prediction table can be calculated based on multiple prediction results of the state prediction table. For example, when the predicted compression state is consistent with the actual compression state, it means that the prediction is accurate and the prediction accuracy rate increases accordingly. When the predicted compression state is inconsistent with the actual compression state, it means that the prediction is incorrect and the prediction accuracy rate decreases accordingly. When the cumulative prediction accuracy rate is lower than a preset accuracy rate threshold, the prediction function can be controlled to be turned off, that is, when a certain data to be written needs to be compressed and written, it is controlled not to be compressed and directly written in an uncompressed state, so as to avoid repeated reading and wasting resources and bandwidth due to inaccurate prediction. When the cumulative prediction accuracy rate is higher than a preset accuracy rate threshold, the prediction function can be controlled to be turned on, that is, when a certain data to be written needs to be compressed and written, it is controlled to be compressed and written in a compressed state, so as to reduce the bandwidth resources between on-chip and off-chip through the bandwidth compression technology when the prediction accuracy rate is guaranteed.
[0029] In a second aspect, an embodiment of the present application provides a memory management device, which may include:
[0030] A predictor for managing a state prediction table, where the state prediction table includes a mapping relationship between an identifier of a memory area and a prediction identifier; wherein, the memory area includes a plurality of memory units, and the prediction identifier is used to predict whether the data in the plurality of memory units is in a compressed state;
[0031] A controller, configured to receive a memory access request for a target memory cell and determine an identifier of a target memory region to which the target memory cell belongs;
[0032] The predictor is further configured to query a target prediction identifier corresponding to the identifier of the target memory region from the status prediction table, and determine a predicted compression status of data in the target memory cell based on the target prediction identifier;
[0033] The controller is further configured to read target data in the target memory cell based on the predicted compression status and determine an actual compression status of the target data;
[0034] The predictor is further configured to update the target prediction identifier based on the predicted compression status and the actual compression status.
[0035] In a possible implementation, the size of the memory cell is the memory access granularity size; the controller is further configured to:
[0036] Receive a memory access request, where the memory access request includes a starting address and an access length;
[0037] Determine a target memory cell based on the starting address and the access length, where the target memory cell includes one or more memory cells.
[0038] In a possible implementation, the type of the prediction identifier includes at least two groups of identifiers, where each group of identifiers is used to indicate a prediction probability that the data is in a compressed state or a non-compressed state, and the prediction probabilities corresponding to the same group of identifiers are the same, and the prediction probabilities corresponding to different groups of identifiers are different.
[0039] In a possible implementation, the at least two groups of identifiers include a first group of identifiers and a second group of identifiers, the first group of identifiers are respectively a non-compressed identifier and a compressed identifier, and the second group of identifiers are respectively a suspected non-compressed identifier and a suspected compressed identifier; the prediction probability corresponding to the first group of identifiers is higher than the prediction probability corresponding to the second group of identifiers.
[0040] In a possible implementation, the predictor is specifically configured to:
[0041] If the target prediction identifier is the compressed identifier, determine that the predicted compression status of the data in the target memory cell is the compressed state;
[0042] If the target prediction identifier is the non-compressed identifier, determine that the predicted compression status of the data in the target memory cell is the non-compressed state.
[0043] In a possible implementation, the device is applied to a memory access system; the predictor is specifically configured to:
[0044] If the target prediction flag is currently the suspected compression flag or the suspected non - compression flag, then determine the relationship between the current latency sensitivity and bandwidth sensitivity of the memory access system;
[0045] When the memory access system is more sensitive to bandwidth, determine that the predicted compression state of the data in the target memory unit is the compression state;
[0046] When the memory access system is more sensitive to latency, determine that the predicted compression state of the data in the target memory unit is the non - compression state.
[0047] In a possible implementation, the predictor is specifically configured to:
[0048] If both the predicted compression state and the actual compression state are non - compression states, and the target prediction flag is currently the non - compression flag, then keep the target prediction flag unchanged;
[0049] If both the predicted compression state and the actual compression state are non - compression states, and the target prediction flag is currently the suspected non - compression flag or the suspected compression flag, then adjust the target prediction flag in the direction of the non - compression flag.
[0050] In a possible implementation, the predictor is specifically configured to:
[0051] If the predicted compression state is the non - compression state and the actual compression state is the compression state, then adjust the target prediction flag in the direction of the compression flag.
[0052] In a possible implementation, the predictor is specifically configured to:
[0053] If the predicted compression state is the compression state and the actual compression state is the non - compression state, then adjust the target prediction flag in the direction of the non - compression flag.
[0054] In a possible implementation, the predictor is specifically configured to:
[0055] If both the predicted compression state and the actual compression state are compression states, and the target prediction flag is currently the compression flag, then keep the target prediction flag unchanged;
[0056] If both the predicted compression state and the actual compression state are compression states, and the target prediction flag is currently the suspected non - compression flag or the suspected compression flag, then adjust the target prediction flag in the direction of the compression flag.
[0057] In a possible implementation, the predictor is further configured to:
[0058] Create a status prediction table, and establish a mapping relationship between the identifiers of each memory area and the corresponding prediction identifiers in the status prediction table;
[0059] Set the initial identifier of each prediction identifier to indicate the uncompressed state.
[0060] In a possible implementation manner, the controller is specifically configured to:
[0061] If the predicted predicted compression state is the compression state, parse the first physical address corresponding to the target memory unit according to the rule of compressed data, and read the data stored in the first physical address.
[0062] In a possible implementation manner, the controller is specifically configured to:
[0063] If the predicted predicted compression state is the uncompressed state, parse the second physical address corresponding to the target memory unit according to the rule of the uncompressed state, and read the data stored in the second physical address.
[0064] In a possible implementation manner, the controller is specifically configured to:
[0065] After reading the target data, obtain the compression state identifier carried in the target data, where the compression state identifier is used to indicate whether the target data is compressed data;
[0066] Determine the actual compression state of the target data based on the compression state identifier.
[0067] In a possible implementation manner, the controller is further configured to:
[0068] When the predicted compression state is the compression state and the actual compression state is the uncompressed state, reread the target data based on the uncompressed state.
[0069] In a possible implementation manner, the device is applied to a memory access system; the controller is specifically configured to:
[0070] When the predicted compression state is the uncompressed state and the actual compression state is the compression state, read the compressed first data and second data stored in the target memory unit;
[0071] The memory access system includes a prefetch cache; the device further includes: a decompression module, configured to decompress the first data and the second data;
[0072] The controller is further configured to read the decompressed first data according to the memory access request, and store the decompressed second data into the prefetch cache.
[0073] In a possible implementation, if the target memory unit includes two adjacent memory units, and the two adjacent memory units include a high-order memory unit and a low-order memory unit;
[0074] The controller is specifically configured to: if reading the target memory unit according to the rule of compressed data, read the low-order memory unit;
[0075] The controller is specifically configured to: if reading the target memory unit according to the rule of normal data, read the high-order memory unit and the low-order memory unit.
[0076] In a possible implementation, the device further includes a compression module; the controller is further configured to:
[0077] Receive a write request for the data to be written, and determine whether the data to be written needs to be compressed, where the size of the data to be written is M bits;
[0078] The compression module is configured to: if the data to be written needs to be fully compressed, compress the M bits of the data to be written; the controller is further configured to: write the compressed data into the corresponding memory unit;
[0079] The compression module is further configured to: if a part of the data to be written needs to be compressed, compress N bits of the data to be written; the controller is further configured to: write the compressed data into the first memory unit, and write the remaining (M - N) bits of the data to be written into the second memory unit; both M and N are integers greater than 1, and N is less than M;
[0080] The controller is further configured to: update the compression status flag bits in the first memory unit and the second memory unit to a compression flag and a non-compression flag, respectively.
[0081] In a possible implementation, the compression module is specifically configured to:
[0082] Split the M bits of the data to be written into K bits of data and (M - K) bits of data, and compress the K bits of data and the (M - K) bits of data respectively.
[0083] In a possible implementation, the device further includes: a dynamic switch module, configured to:
[0084] Obtain the current prediction accuracy rate of the status prediction table calculated based on the predicted compression status and the actual compression status of multiple memory access data;
[0085] When the data to be written needs to be compressed and the prediction accuracy rate is lower than a preset threshold, the compression module is controlled to turn off the compression function, so as to write the data to be written into the corresponding memory unit in an uncompressed state;
[0086] When the data to be written needs to be compressed and the prediction accuracy rate is higher than a preset threshold, the compression module is controlled to turn on the compression function, so as to compress the data to be written and write it into the corresponding memory unit.
[0087] In a third aspect, an embodiment of the present application provides a system on chip (SOC), which may include the memory management device provided by any one of the implementation manners in the second aspect above.
[0088] In a fourth aspect, the present application provides a semiconductor chip, which may include the memory management device provided by any one of the implementation manners in the second aspect above.
[0089] In a fifth aspect, the present application provides a chip system, which includes the memory management device provided by any one of the implementation manners in the second aspect above. In a possible design, the chip system further includes a memory, and the memory is used to store program instructions and data necessary or related during the operation of the chip system. The chip system may be composed of chips or may include chips and other discrete devices.
[0090] In a sixth aspect, the present application provides an electronic device, which has the function of implementing any one of the memory management methods in the first aspect above. This function may be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0091] In a seventh aspect, the present application provides an electronic device, which includes the memory management device provided by any one of the implementation manners in the second aspect above. In a possible design, the electronic device further includes a memory, and the memory is used to store program instructions and data necessary or related during the operation of the electronic device. The chip system may be composed of chips or may include chips and other discrete devices.
[0092] In an eighth aspect, the present application provides a memory management device, which has the function of implementing any one of the memory management methods in the first aspect above. This function may be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0093] In a ninth aspect, the present application provides a terminal device, which includes a memory management device, and the memory management device is the electronic device provided by any of the implementation manners in the first aspect above. The terminal device may further include a memory, which is used to be coupled with the memory management device and store necessary program instructions and data of the terminal device. The terminal device may further include a communication interface for the terminal device to communicate with other devices or communication networks.
[0094] In a tenth aspect, the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by multiple electronic devices, the power supply management method flow described in any one of the second aspect above is implemented.
[0095] In an eleventh aspect, an embodiment of the present application provides a computer program, which includes instructions, and when the computer program is executed by multiple electronic devices, the electronic devices can execute the power supply management method flow described in any one of the second aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] Figure 1A It is a schematic diagram of the system architecture of a multi-level management solution.
[0097] Figure 1B It is a schematic diagram of the system architecture of an on-chip metadata optimization solution.
[0098] Figure 1C It is a schematic diagram of the system architecture of a multi-level metadata management solution.
[0099] Figure 2A It is a schematic diagram of the structure of a system on chip (SOC) provided by an embodiment of the present application.
[0100] Figure 2B It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application.
[0101] Figure 2C It is a schematic diagram of the structure of another electronic device provided by an embodiment of the present application.
[0102] Figure 2D It is a schematic diagram of the structure of yet another electronic device provided by an embodiment of the present application.
[0103] Figure 3A It is a schematic diagram of the division of a physical memory space provided in an embodiment of the present application.
[0104] Figure 3B It is a schematic diagram of the division of a memory area provided by an embodiment of the present application.
[0105] Figure 3C It is another schematic diagram of the division of a memory area provided by an embodiment of the present application.
[0106] Figure 3D Another schematic diagram of memory area division provided by the embodiment of the present application.
[0107] Figure 4A A schematic diagram of a state prediction table provided by the embodiment of the present application.
[0108] Figure 4B Another schematic diagram of a state prediction table provided by the embodiment of the present application.
[0109] Figure 5A A schematic diagram of the state transition of a prediction identifier provided in the embodiment of the present application.
[0110] Figure 5B A schematic diagram of the state transition of a prediction identifier provided in the embodiment of the present application.
[0111] Figure 6 A schematic diagram of the process of a controller processing read and write requests provided by the embodiment of the present application.
[0112] Figure 7 A schematic diagram of the process of a memory management method provided by the embodiment of the present application. Detailed implementation manners
[0113] Next, the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application. Terms such as "first", "second", "third", and "fourth" in the specification and claims of the present application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. The mention of "embodiment" in this article means that a specific feature, structure, or characteristic described in conjunction with the embodiment may be included in at least one embodiment of the present application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0114] As used in this specification, terms such as "component", "module", "system", etc. are used to denote computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and the computing device can be components. One or more components can reside in a process and / or an execution thread, and components can be located on one computer and / or distributed between two or more computers. In addition, these components can execute from various computer-readable media storing various data structures. Components can communicate, for example, through local and / or remote processes via signals having one or more data packets (such as data from two components interacting with another component among a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0115] First, some terms used in this application are explained to facilitate understanding by those skilled in the art.
[0116] (1) Integrated Circuit (IC) is a kind of microelectronic device or component. Using a certain process, the transistors, resistors, capacitors, inductors and other components and wirings required in a circuit are interconnected and fabricated on a small piece or a few small pieces of semiconductor wafers or dielectric substrates, and then encapsulated in a package to form a micro-structure with the required circuit functions; that is, an IC chip is to place an integrated circuit formed by a large number of microelectronic components (transistors, resistors, capacitors, etc.) on a plastic substrate to make a chip.
[0117] (2) Read Modify Write (RMW) is a basic operation on a memory, which can include a read operation, a write operation, or a read-modify-write operation. The read-modify-write operation means first reading the data in the memory, modifying it, and then writing it back to the memory.
[0118] (3) Double Data Rate (DDR), that is, DDR SDRAM = Double Data Rate Synchronous Dynamic Random Access Memory, usually referred to as DDR. Among them, SDRAM is the abbreviation of Synchronous Dynamic Random Access Memory, that is, synchronous dynamic random access memory. The "synchronous" here means that the memory operation requires a synchronous clock, and the sending of internal commands and data transmission are both based on it. DDR is a storage device that loses data when powered off and requires periodic refreshing to maintain data integrity. The DDR SDRAM subsystem includes three parts: a DDR controller, a DDR PHY, and DRAM memory chips.
[0119] (4) Static Random-Access Memory (SRAM) is a type of random access memory. The so-called "static" means that as long as this memory remains powered on, the data stored in it can be constantly maintained. In contrast, the data stored in Dynamic Random-Access Memory (DRAM) needs to be updated periodically. However, when the power supply stops, the data stored in SRAM will still disappear (referred to as volatile memory), which is different from ROM or flash memory that can store data even after power-off.
[0120] (5) Nand-flash memory is a type of non-volatile memory based on NAND technology. Compared with traditional Flash memory, NAND Flash has higher storage density, lower power consumption, and longer lifespan.
[0121] (6) Robustness refers to the ability of a system or algorithm to remain stable and reliable under different circumstances. Specifically, when faced with some unexpected or abnormal situations, a system or algorithm with strong robustness can maintain its functions and performance without crashing or failing due to these abnormal situations.
[0122] (7) Metadata, also known as intermediary data or relay data, is data about data. It is mainly information describing the properties of data and is used to support functions such as indicating storage locations, historical data, resource search, and file records. Metadata can be regarded as a kind of electronic directory. To achieve the purpose of compiling a directory, it is necessary to describe and collect the content or characteristics of the data, and then achieve the purpose of assisting data retrieval. The metadata in the embodiments of this application can be used to record information related to whether the data has been compressed. For example, the metadata records that a certain data is compressed data or uncompressed data.
[0123] (8) Bit (bit), a computer can convert 0 and 1 into signals in the circuit for calculation. A bit is the smallest unit for a computer to store and move data, and it has only 2 values, 0 and 1. Its abbreviation is the lowercase letter "b".
[0124] (9) Byte, Byte is the English name for byte. Its abbreviation is the capital letter "B". Since it is called a byte, it must be related to characters. English characters are usually one byte, that is, 1B. Chinese characters usually exceed 2 bytes due to the character set. Conversion relationship: 8 bits is equal to 1 byte. One byte is equal to eight bits.
[0125] First, for the convenience of understanding the embodiments of the present application, the technical problems to be specifically solved by the present application are further analyzed and proposed. Currently, the implementation of bandwidth compression technology includes various technical solutions. The following are three exemplary ones, among which,
[0126] Solution 1: Multi-level management solution
[0127] As Figure 1A shown, Figure 1A is a schematic diagram of the system architecture of a multi-level management solution. This multi-level management solution includes the following parts:
[0128] (1) Use a metadata cache to store the metadata of 64 consecutive datablocks.
[0129] (2) Memory write-back requests will be completed in the write buffer.
[0130] (3) Allow the compressed memory block to be of any size in the [0, 64) byte spectrum, and 7-bit metadata / block is required.
[0131] The disadvantages of this Solution 1:
[0132] (1) Rely on a high compression ratio, resulting in a very large amount of metadata size, and 7 bits are required for each cacheline.
[0133] (2) Determining the DDR access address completely depends on the MetadataCache, resulting in a large consumption of on-chip cache space.
[0134] (3) Additional write cache space is required to increase the compression ratio.
[0135] (4) Lack of a dynamic adjustment mechanism.
[0136] Solution 2: Off-chip metadata optimization solution
[0137] As Figure 1B shown, Figure 1B is a schematic diagram of the system architecture of the off-chip metadata optimization solution. In this off-chip metadata optimization solution, the storage of off-chip metadata is optimized, and the metadata is stored in the extra space saved by compression, and a method to solve the conflict with the original data is proposed. It is mainly divided into the following parts:
[0138] (1) By restricting the size of the compressed data, reserve space to store the metadata and the compressed data together, and record it with a specific marker.
[0139] (2) When the metadata tag conflicts with the original data, use a Line Inversion Table (LIT) to record the complement of the conflicting data.
[0140] (3) A set of solutions to address the limited space of the Line Inversion Table (LIT).
[0141] Disadvantages of Solution 2:
[0142] (1) It does not solve the problem of high overhead in managing on-chip metadata.
[0143] (2) When the DDR has independent flag bits, there is no need to merge and store off-chip metadata.
[0144] (3) The Line Inversion Table (LIT) module has poor scalability, and if an attacker knows the design of the flag bits and the size of the LIT, there are security risks and it is vulnerable to malicious attacks.
[0145] (4) Lack of a dynamic adjustment mechanism.
[0146] Solution 3: Multilevel Metadata Management Solution
[0147] As Figure 1C shown, Figure 1C is a schematic diagram of the system architecture of the multilevel metadata management solution. To reduce the storage overhead of on-chip metadata, this work proposes a set of multilevel metadata management solutions, mainly including the following key points:
[0148] (1) Adopt a three-level scheme of global predictor, page predictor, and row predictor (metadata storage).
[0149] (2) This multilevel metadata management solution can achieve better performance than a 1MB metadata cache with an overhead of 368KB.
[0150] Disadvantages of Solution 3:
[0151] (1) Although the accuracy of the predictor has been improved, the absolute accuracy is still relatively low, resulting in a still large space overhead.
[0152] (2) The three-level prediction linkage mechanism is complex and the prediction robustness is poor.
[0153] (3) Lack of a dynamic adjustment mechanism.
[0154] In summary, the above three bandwidth compression technologies in the prior art each have their own defects and cannot achieve a balance between prediction accuracy and space overhead. For example, there are problems such as excessive space overhead, complex prediction mechanisms, poor robustness, and inflexible mechanisms. Therefore, the technical problems to be solved by this application may include one or more of the following aspects: providing a bandwidth compression scheme that can improve prediction accuracy while minimizing the overhead of storage space (such as on-chip storage space), reducing prediction complexity, enhancing prediction robustness, and improving dynamic flexibility of adjustment.
[0155] Based on the above, please refer to Figure 2A , Figure 2A FIG. is a schematic structural diagram of a system on a chip (SOC) provided by an embodiment of this application. The SOC 10 may specifically include a predictor 101 and a controller 102. Please refer to Figure 2B , Figure 2B FIG. is a schematic structural diagram of an electronic device provided by an embodiment of this application. The electronic device 01 may include the above SOC 10. Optionally, it may further include a memory 20. The above SOC 10 or electronic device 01 may be located in any electronic device, such as various devices like a computer, a computer, a mobile phone, a tablet, a smart wearable device, etc. The above SOC 10 or electronic device 01 may specifically also be a chip or a chipset or a circuit board equipped with a chip or a chipset, and the chip or the chipset or the circuit board equipped with a chip or a chipset may work under the drive of necessary software. The above SOC 10 in the embodiments of this application or the electronic device 01 including the SOC 10 may be applied in a memory system, including a general server, a high-performance (HPC) server, an artificial intelligence (AI) server, a mobile device, a terminal device, or a memory control module of a terminal device, etc. The SOC 10 or electronic device 01 in the embodiments of this application deployed in relevant scenarios can significantly reduce the memory access bandwidth and the storage resource overhead. It should be noted that the memory management device in this application may be the above system on a chip SOC 10, or the above electronic device 01, or a module in the above SOC 10 or electronic device 01. That is, for the structure and function of the memory management device in this application, reference may be made to the relevant descriptions of the SOC 10 or electronic device 01 in the embodiments of this application, and details will not be described hereinafter.
[0156] In a possible implementation, the memory management method or device in the embodiments of the present application can also be applied to an off-chip memory management system. For example, in a multi-level memory management system, any one of the memory management methods or devices in the present application is set in one or more levels of memory to perform memory management on the next-level memory that uses bandwidth compression technology. For details, reference can be made to the relevant embodiments in which the memory management method or device in the embodiments of the present application is applied to a system-on-chip, which will not be elaborated here. Optionally, the memory management method or device in the present application can also be applied to the memory management between a storage access end and a storage to-be-accessed end that uses bandwidth compression technology, that is, the memory management method in the present application is executed by the storage access end to manage the storage to-be-accessed end, thereby reducing the overhead of the bandwidth compression technology on storage resources.
[0157] Further, please refer to Figure 2C , Figure 2C FIG. is a schematic structural diagram of another electronic device provided by an embodiment of the present application. The SOC 10 in the electronic device 01 may further include some or all of the processor 100, prefetch cache 103, decompression module 104, and compression module 105 shown in Figure 2C . Among them,
[0158] The SOC 10 may be an integrated circuit with a dedicated target, which includes a complete system and all the content of the embedded software. Exemplarily, the SOC 10 in the embodiments of the present application may specifically include some or all of the following modules:
[0159] Predictor 101: It can be used to set and store the state prediction table in the embodiments of the present application, including setting the identifier of the memory area and the initial mapping relationship between prediction identifiers. Further, it also includes updating the prediction identifiers based on the predicted state and actual state of whether the memory access data is compressed data. For example, the predictor 101 performs binary classification prediction (compressed, uncompressed) based on the BiModal algorithm. Exemplarily, the process of the predictor 101 updating the prediction identifiers may include that after the predictor 101 obtains the access address, it divides the access address into blocks (i.e., corresponding to the memory area), and the block granularity can be flexibly set, such as 1, 2, 4 KB, etc. Taking the block granularity of 4 KB as an example, a hashing operation of taking the remainder of the divided address is performed to obtain a two-digit confidence register, which is also the prediction identifier in the embodiments of the present application (for example, a compression identifier, an uncompressed identifier, a suspected compression identifier, a suspected uncompressed identifier), and is used to predict the compression state of the data. Exemplarily, when the confidence is 3, it is predicted to be in the compressed state; when the confidence is 0, it is predicted to be in the uncompressed state. When the confidence is 1 or 2, the data compression state can be further predicted according to the memory (such as DDR) bandwidth requirement and latency requirement. After updating the confidence, it is written back to the register corresponding to the hash table, and all read requests within this granularity can be updated.
[0160] Controller 102: It can be used for various memories in the electronic device 01, such as the memory 20, the prefetch cache 103, etc., and is responsible for data exchange, data reading and writing, and memory allocation and management between the CPU and the memory 20. Optionally, the controller 102 can be a logic function module with corresponding functions, such as a logic state machine, or can be implemented by executing corresponding software on relevant hardware. For example, the controller 102 receives and parses read commands or write commands sent by the CPU, etc., and according to the logical address of the data carried in the read command or write command, etc., resolves the logical address of the data into a physical address according to the fixed address mapping relationship, and then finds the position of the data to be read or written, and finally sends a control signal corresponding to the command to the memory 20. The specific functions involved in the memory access process of the controller 102 applying bandwidth compression and compression state prediction in the embodiments of the present application will be introduced in subsequent embodiments.
[0161] Processor (Central Processing Unit, CPU) 100: The processor is the core part of the SOC and is responsible for executing various computing and control tasks. Exemplarily, the CPU 100 can run an operating system or application programs to control multiple hardware or software components connected to the processor 100, and can process various data and perform operations. Exemplarily, the processor 100 can load instructions or data from the external memory 30 into the memory 20 (please refer to Figure 2C ), and initiate various memory access requests to the memory 20 when needed (such as sending a read data request, a write data request, a read-modify-write request, etc. to the memory 20). Optionally, the processor 100 may include one or more processing units (also referred to as processing cores). For example, the processor 100 may include a central processing unit (CPU), an application processor (AP), a modulation and demodulation processing unit, a graphics processing unit (GPU), an image signal processor (ISP), a video codec unit, a digital signal processor (DSP), a baseband processing unit, and a neural-network processing unit (NPU), etc., one or more of them. Among them, different processing units can be independent devices or integrated in one or more devices. Optionally, a memory may also be provided in the processor 100 for storing instructions and data. In some embodiments, the memory in the processor 100 is a cache memory (Cache). The Cache can save the instructions or data that the processor 100 has just used or recycled. If the processor 100 needs to use the instruction or data again, it can be directly called from the Cache. This avoids repeated accesses, reduces the waiting time of the processor 100, and thus improves the efficiency of the system.
[0162] Compression Module (Compression Engine, CE) 105: It is used to compress the data that needs to be compressed. Optionally, the compression module 105 in the embodiments of the present application can support mixed granularity compression, effectively increasing the compression ratio. Exemplarily, the mixed granularity includes compressing 128B data to 64B, and compressing two 64B data into 32B respectively. Correspondingly, decompression supports decompressing 64B into 128B, and decompressing two independent 32B data into 64B respectively.
[0163] The decompression module (Decompression Engine, DCE) 104 is used to decompress the compressed data. Optionally, the decompression module 104 in the embodiments of the present application can support mixed granularity compression, effectively increasing the compression ratio. Exemplarily, the mixed granularity includes compressing 128B data to 64B, and compressing two 64B data into 32B respectively. Correspondingly, decompression supports decompressing 64B into 128B, and decompressing two independent 32B data into 64B respectively.
[0164] The prefetch buffer (PB) 103: can be used to store the extra data obtained due to compression. Optionally, the structure of the prefetch buffer 103 can reuse the existing cache structure. For example, if the memory access request hits in the PB, the data is directly read from the PB and returned to the request. Assume that the original memory access request is 64 bytes, and the state read from the memory unit is the compressed state, then the extra fetched data can be placed in the PB 103.
[0165] Further, please refer to Figure 2D , Figure 2D which is a schematic structural diagram of another electronic device provided by the embodiments of the present application. The electronic device 01 may further include a dynamic switch (Dynamic Adjustment, DA) module 106 as shown in Figure 2D :
[0166] The dynamic switch module 106: is used to dynamically turn on / off the compression feature. Specifically, the embodiments of the present application can make a decision to turn off the compression feature according to the prediction accuracy rate. Optionally, the data compression ratio can also be used to make a decision to turn on the compression feature. For example, when the prediction accuracy rate is lower than 80%, the compression feature is turned off, and the data of the write request is no longer compressed. Another example is that when the compression ratio is higher than 80%, the compression feature is turned on, and the data of the write request is compressed. It should be noted that when the compression feature is turned off, the data in the memory may still be in the compressed state, so the data path of the read request is still the same as when the compression is turned on.
[0167] It can be understood that in addition to the above main functional modules, the SoC 101 may also include other functional modules such as a security module, a power management unit, a touch controller, and various interface controllers. These modules cooperate with each other to provide various functions and performances of the SOC, which will not be listed one by one here. In summary, the SOC in the embodiments of the present application is a chip integrating multiple functional modules, and these functional modules cooperate with each other to provide various functions such as computing, graphics, communication, multimedia, and sensing for the SOC. It can be understood that the internal structure and functions of the SOC in different application scenarios or different electronic devices may be different, and the embodiments of the present application do not make specific limitations thereto.
[0168] Further, please refer to Figure 2C or Figure 2D , the electronic device 01 further includes a memory 20; optionally, the electronic device 01 may further include an external memory 30, where
[0169] The memory 20 is usually a power-off volatile memory, and the content stored thereon will be lost when powered off. The memory 20 in this application refers to a readable and writable operating memory, which can also be called an internal memory in this application, abbreviated as memory (Memory) or main memory. Its function is to temporarily store the operation data in the processor 100, and exchange data with the external memory 30 or other external memories, and can be used as a storage medium for temporary data of the operating system or other running programs. For example, the operating system running on the processor 100 transfers the data to be operated from the memory 20 to the processor 100 for operation, and when the operation is completed, the processor 100 then transmits the result. Since the operation of all programs needs to be first loaded into the memory 20, and then the processor 100 can load and run, therefore, the performance of the memory 20 has a great impact on the operation performance of the processor 100, and determines whether the electronic device 01 itself or the electronic device where the electronic device 01 is located can operate normally, stably and efficiently.
[0170] The internal memory, that is, the memory 20, may include one or more of dynamic random access memory (DRAM), static random access memory (SRAM), synchronous dynamic random access memory (SDRAM), etc. Among them, DRAM further includes double data rate synchronous dynamic random access memory (Double Data Rate Synchronous Dynamic Random Access Memory, DDR SDRAM), abbreviated as DDR, double data rate synchronous dynamic random access memory of the second generation (DDR2), double data rate synchronous dynamic random access memory of the third generation (DDR3), low power double data rate synchronous dynamic random access memory of the fourth generation (Low Power Double Data Rate 4, LPDDR4), and low power double data rate synchronous dynamic random access memory of the fifth generation (Low Power Double Data Rate 5, LPDDR5), etc.
[0171] The external memory 30 is usually a non-volatile memory, and the content stored in it will not be lost after power-off. The external memory 30 in this application may include a read-only memory (ROM) for storing system information and a boot program, and a readable and writable external memory (such as Flash) for storing programs and data. Its function is to store instructions and data for a long time. For example, the system information includes system files such as the Linux kernel and the Android operating system; the programs may include system-built-in application programs (such as an application market, a wallet application, a security center, etc.) that come with the electronic device 01 when it leaves the factory and application programs downloaded and installed by users later (such as social applications, video applications, mobile payment applications, game applications, etc.); the data may include system data related to system operation (such as configuration file data, log file data, cache data, etc.), and data generated during the user's use process (such as fingerprint data, chat history data, photo and video data, etc.).
[0172] The external memory 30 may include one or more of a one-time programmable read-only memory (OTPROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a mask ROM, a flash ROM, a universal flash storage (UFS), a Flash memory (such as NAND flash, NOR flash, etc.), a hard disk drive, or a solid state drive (SSD).
[0173] It can be understood that the structure schematically shown in the embodiments of this application does not constitute a specific limitation on the SOC 10 or the electronic device 01. In some other embodiments of this application, the SOC 10 or the electronic device 01 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware.
[0174] Based on the above Figures 2A - 2D For any possible structure of the SOC 10 or the electronic device 01, the memory management functions specifically implemented by the SOC 10 or the electronic device 01 in the embodiments of this application are introduced as follows:
[0175] The predictor 101 is used to manage the status prediction table, and the status prediction table includes the mapping relationship between the identifiers of memory regions and prediction identifiers; wherein, the memory region includes a plurality of memory cells, and the prediction identifier is used to predict whether the data in the plurality of memory cells is in a compressed state;
[0176] The controller 102 is used to receive a memory access request for a target memory cell and determine the identifier of the target memory region to which the target memory cell belongs;
[0177] The predictor 101 is further configured to query a target prediction identifier corresponding to the identifier of the target memory area from the state prediction table, and determine a predicted compression state of the data in the target memory cell based on the target prediction identifier;
[0178] The controller 102 is further configured to read target data in the target memory cell based on the predicted compression state, and determine an actual compression state of the target data;
[0179] The predictor 101 is further configured to update the target prediction identifier based on the predicted compression state and the actual compression state.
[0180] Specifically, in the embodiments of the present application, the memory 20 may be divided into multiple memory areas (such as Figure 2B memory area 1, memory area 2, memory area 3... memory area N shown in
[0181] Please refer to Figure 2B, multiple memory regions (such as memory region 1, memory region 2, memory region 3, ……, memory region N) are divided for the physical space of the memory. Each memory region includes multiple memory units (memory unit 1, memory unit 2, memory unit 3, ……, memory unit M). That is, the physical memory space of the memory 20 is divided into multiple sub-physical memory spaces (i.e., corresponding memory regions), and the memory regions are further divided into multiple physical address segments with finer granularity (i.e., corresponding memory units). For multiple memory regions in the memory 20, the predictor 101 establishes a corresponding state prediction table, and creates and stores the correspondence between the identifiers of each memory region and the prediction identifiers in the state prediction table (that is, each memory region corresponds to a prediction identifier, and each memory region itself also has an identifier for identifying the region. Therefore, the identifier of each memory region corresponds to a prediction identifier). When the system issues a memory access request for a certain target address (usually the base start address + access length), first determine the memory region to which the target address (i.e., the target memory unit) requested to be read by the memory access request belongs, then find the mapped prediction identifier through the identifier of this memory region, and predict whether the data in the target memory unit is compressed based on this prediction identifier. Finally, read the data in the target unit according to the rule corresponding to the prediction result (i.e., the predicted compression state). For example, if it is predicted as compressed data, read the data according to the compressed physical address, and if it is predicted as uncompressed data, read the data according to the complete address. After the target data is read, since the target data carries the corresponding state compression identifier (i.e., the identifier indicating whether the data is actually compressed). Therefore, further, in the embodiment of the present application, by comparing the compression state (i.e., the actual compression state) corresponding to the actual compression state identifier and the compression state (i.e., the predicted compression state) corresponding to the above prediction identifier, it is judged whether the prediction result is consistent with the actual result, and then the target prediction identifier corresponding to the target memory region is updated based on the judgment result, so that the state prediction mechanism is more accurate.
[0182] Optionally, when the judgment results are consistent, it means that the current prediction direction (i.e., the compression direction or the non-compression direction) is accurate. Therefore, the current prediction identifier can be kept unchanged, or further adjusted in the current prediction direction (i.e., increasing the prediction probability in this direction); when the judgment results are inconsistent, it means that the current prediction direction is wrong, and then the current prediction identifier needs to be adjusted in the opposite direction of the prediction direction (i.e., reducing the prediction probability in this direction and / or increasing the prediction probability in the opposite direction).
[0183] In the embodiments of the present application, by uniformly setting the same prediction identifier for the entire memory area containing multiple memory cells, it is used to predict the compression state of the data in all the memory cells in the memory area. Different from the prior art where a compression state identifier is set for each memory cell, it saves on-chip storage space. Further, on this basis, by comparing the predicted compression state before reading the data with the actual compression state obtained after reading, and then updating the prediction identifier of the corresponding memory area according to the comparison result, that is, predicting the compression state of the data in the entire memory area based on the compression state of the local memory cells in the memory area, and updating the prediction identifier of whether the memory area is compressed or not according to the actually read content, making the prediction accuracy higher. It reduces the on-chip storage space for metadata (prediction identifier) and improves the bandwidth utilization rate between the on-chip system and the memory DDR.
[0184] Exemplarily, please refer to Figure 3A , Figure 3A which is a schematic diagram of the physical memory space division provided in the embodiments of the present application. Since physical memory is usually managed in units of pages, and virtual memory ultimately needs to be mapped to physical memory, there is also a corresponding concept of virtual pages in the virtual memory space. The mapping of memory is usually carried out in units of pages. The continuous virtual memory seen by a process (a process of the system or an application) may be discontinuous in physical memory. For example, in Figure 3A above, virtual page 2 and virtual page 3 in process 1 are continuous in the virtual memory space of process 1, but the physical memory pages they are mapped to are discontinuous (corresponding to physical page 1 and physical page 3 respectively). Another example is that in process 2, virtual memory page 5 is mapped to physical memory page 5, and virtual memory page 6 is mapped to physical memory page 7. Physical page 1, physical page 2, and physical page 3 in the physical memory space are the memory being used by process 1, and physical page 5 and physical page 7 are the memory being used by process 2. These complex and specific memory mapping details can be managed by the memory management subsystem running on the processor or controller.
[0185] Optionally, please refer to Figure 3B , Figure 3B which is a schematic diagram of the memory area division provided in the embodiments of the present application. In this Figure 3BIn [the situation], a physical memory page can be divided into a memory region, that is, one memory region corresponds to one memory page. For example, when the granularity of a Cache Line in the System on Chip (SOC) 10 is 64 bytes (Byte), and the size of a physical page is 4KB (Kilobyte), at this time, a physical page can be divided into 64 memory units of 64 - byte size. Since 1KB = 1024B, then 4KB / 64B = 64 memory units, that is, a memory region (with a size of 4K) includes 64 memory units. Among them, a Cache Line is the minimum unit for data transfer between the Central Processing Unit (CPU) and the memory (main memory). Taking the Cache Line size as 64 bytes (64Byte) as an example, the CPU will fetch 64 - byte data continuously from the memory when transferring data.
[0186] Optionally, please refer to Figure 3C , Figure 3C which is another schematic diagram of memory region division provided by the embodiments of this application. In this Figure 3C situation, one or more physical pages can be divided into a memory region, and the size of each memory region can be the same or different. For example, in Figure 3C this situation, memory region 1 includes 2 physical memory pages (physical page 1 and physical page 2), memory region 2 includes 3 physical memory pages (physical page 3, physical page 4, and physical page 5), memory region 3 includes 2 physical memory pages (physical page 6 and physical page 7), and memory region 4 includes 1 physical memory page (physical page 8). Correspondingly, the sizes of each memory region are 8K, 12K, 8K, and 4K respectively, and they include 128 memory units, 192 memory units, and 64 memory units respectively. That is, in the embodiments of this application, the size of each memory region is not specifically limited. Optionally, please refer to Figure 3D , Figure 3D which is yet another schematic diagram of memory region division provided by the embodiments of this application. In Figure 3D this situation, for example, a 4K memory page can include 32 memory units with the granularity of cacheLine, and the size of each memory unit is 64Byte. The high 64 - bit and low 64 - bit of 128Byte data can be stored between two adjacent memory units respectively.
[0187] In a possible implementation manner, the predictor 101 is further configured to create a state prediction table, and establish a mapping relationship between the identifiers of each memory region and the corresponding prediction identifiers in the state prediction table; set the initial identifier of each prediction identifier to indicate the non - compressed state. Please refer to Figure 4A , Figure 4AA schematic diagram of a state prediction table provided by an embodiment of the present application. For example, the memory area identifiers are 01, 02, 03, 04, 05, ……, N respectively, and the initial identifiers of the corresponding prediction identifiers are all 0. Assume that at this time, 0 is used to predict that the data in the corresponding memory area is uncompressed data. In the embodiment of the present application, a state prediction table is also established, and a mapping relationship is established between the identifiers of each memory area and the corresponding prediction identifiers in the table. Since when establishing these mapping relationships, an initial identifier needs to be set for the prediction identifier corresponding to each memory area, considering that the situation of each memory area is not clear initially, the most primitive reading method can be adopted first, that is, reading in an uncompressed manner. Reading in an uncompressed manner will read according to the complete physical address of the target memory unit to reduce the probability of rereading caused by the lack of prediction basis in the initial situation. Further, please refer to Figure 4B , Figure 4B A schematic diagram of another state prediction table provided by an embodiment of the present application. After multiple rounds of prediction for each memory area in this state prediction table, the prediction identifiers (for example, including four cases of 0, 1, 2, and 3, which will be described in detail in subsequent embodiments) will be updated based on different prediction situations and actual compression situations to update the predicted compression state corresponding to each memory area in real time, thereby improving the accuracy and robustness of the prediction.
[0188] In a possible implementation manner, the type of the prediction identifier includes at least two groups of identifiers. Among them, each group of identifiers is used to indicate the prediction probability that the data is in a compressed state or an uncompressed state, and the prediction probabilities corresponding to the same group of identifiers are the same, while the prediction probabilities corresponding to different groups of identifiers are different. Specifically, the type of the prediction identifier can include at least two groups of identifiers. Among them, each group of identifiers is used to indicate the compressed state and uncompressed state of the memory area under the same prediction probability, and the prediction probabilities corresponding to different groups are different. For example, one group of identifiers is 0 and 3. Specifically, 0 is used to indicate that the prediction probability that the data in the memory area is in an uncompressed state is 100%, and 3 is used to indicate that the prediction probability that the data in the memory area is in a compressed state is 100%; another example is that another group of identifiers is 1 and 2, where 1 is used to indicate that the prediction probability that the data in the memory area is in an uncompressed state is 50%, and 2 is used to indicate that the prediction probability that the data in the memory area is in a compressed state is 50%. Optionally, the prediction identifier can also include more groups of identifiers. For example, in addition to the above 0, 1, 2, 3, there can be another group of identifiers 4 and 5, where 4 is used to indicate that the prediction probability that the data in the memory area is in an uncompressed state is 80%, and 5 is used to indicate that the prediction probability that the data in the memory area is in a compressed state is 80%. The prediction identifier in the embodiment of the present application includes multiple groups of identifiers, which can greatly improve the accuracy and robustness of the prediction. It can be understood that the specific type and quantity of the prediction identifier in the embodiment of the present application are not specifically limited.
[0189] In a possible implementation, the at least two sets of identifiers include a first set of identifiers and a second set of identifiers. The first set of identifiers are respectively an uncompressed identifier and a compressed identifier, and the second set of identifiers are respectively a suspected uncompressed identifier and a suspected compressed identifier. The prediction probability corresponding to the first set of identifiers is higher than the prediction probability corresponding to the second set of identifiers. Specifically, after the predictor 101 determines the target prediction identifier corresponding to the target memory cell through the mapping relationship between the identifier in the memory area of the state prediction table and the prediction identifier, it further determines whether the data in the target memory cell is in a compressed state based on the target prediction identifier. For example, if the current prediction probability that the target prediction identifier indicates that the data is in an uncompressed state is 100% (i.e., corresponding to the uncompressed identifier), then according to the preset prediction rule, it is determined that the predicted compression state of the target memory cell is the uncompressed state. Another example is that if the current prediction probability that the target prediction identifier indicates that the data is in a compressed state is 100% (i.e., corresponding to the compressed identifier), then according to the preset prediction rule, it is determined that the predicted compression state of the target memory cell is the compressed state. Still another example is that if the current prediction probability that the target prediction identifier indicates that the data is in an uncompressed state is 50% (i.e., corresponding to the suspected uncompressed identifier), then according to the preset prediction rule, it is necessary to further determine the current actual requirements of the memory access system (such as the on-chip system SOC or the off-chip memory management system, etc.) to determine whether the predicted compression state of the target memory cell is the compressed state or the uncompressed state. For example, according to the current requirements of the memory access system for bandwidth sensitivity or latency sensitivity, the compression state corresponding to the current target prediction identifier is determined.
[0190] In some possible embodiments, the predictor 101 predicting the predicted compression state of the data in the target memory cell based on the prediction identifier may specifically include the following situations:
[0191] (1) If the target prediction identifier is the compressed identifier among the four types (uncompressed identifier and compressed identifier, suspected uncompressed identifier and suspected compressed identifier), it is determined that the predicted compression state of the data in the target memory cell is the compressed state.
[0192] When the prediction identifier includes four identifiers: an uncompressed identifier (corresponding prediction probability is 100%), a suspected uncompressed identifier (corresponding prediction probability is 50%), a suspected compressed identifier (corresponding prediction probability is 50%), and a compressed identifier (corresponding prediction probability is 100%), for the uncompressed identifier or the compressed identifier among the above four identifiers, the compression or uncompression state of the memory area can be directly predicted based on these two identifiers. It can be understood that the compressed identifier or the uncompressed identifier are two prediction results under a prediction probability of 100%. That is, when the prediction identifier is the compression state identifier under a 100% prediction probability, it can be determined that the predicted compression state is the compressed state.
[0193] (2) If the target prediction identifier is an uncompressed identifier among the four types (uncompressed identifier and compressed identifier, suspected uncompressed identifier and suspected compressed identifier), determine that the predicted compression state of the data in the target memory unit is the uncompressed state.
[0194] Similarly, when the prediction identifiers include four types: uncompressed identifier (corresponding prediction probability is 100%), suspected uncompressed identifier (corresponding prediction probability is 50%), suspected compressed identifier (corresponding prediction probability is 50%), and compressed identifier (corresponding prediction probability is 100%), the compression or uncompression state of this memory area can be directly predicted based on these two identifiers. It can be understood that the compressed identifier and the uncompressed identifier are two prediction results under a prediction probability of 100%. That is, when the prediction identifier represents the uncompressed state identifier under a 100% prediction probability, the predicted compression state can be determined to be the uncompressed state.
[0195] (3) If the target prediction identifier is a suspected compressed identifier or a suspected uncompressed identifier among the four types (uncompressed identifier and compressed identifier, suspected uncompressed identifier and suspected compressed identifier), judge the relationship between the current latency sensitivity and bandwidth sensitivity of the SOC; when the SOC is more sensitive to bandwidth, determine that the predicted compression state of the data in the target memory unit is the compressed state.
[0196] When the prediction identifiers include four types: uncompressed identifier (corresponding prediction probability is 100%), suspected uncompressed identifier (corresponding prediction probability is 50%), suspected compressed identifier (corresponding prediction probability is 50%), and compressed identifier (corresponding prediction probability is 100%), the uncompressed identifier or the compressed identifier can be set for direct judgment, while the suspected uncompressed identifier or the suspected compressed identifier can be set to require further combination with the actual requirements of the current SOC to determine. For example, when the target prediction identifier is a suspected compressed identifier or a suspected uncompressed identifier, further determine the current SOC's demand for bandwidth or latency. If the current SOC is more sensitive to bandwidth, determine that the predicted compression state of the data in the target memory unit is the compressed state.
[0197] (4) If the target prediction identifier is a suspected compressed identifier or a suspected uncompressed identifier among the four types (uncompressed identifier and compressed identifier, suspected uncompressed identifier and suspected compressed identifier), judge the relationship between the current latency sensitivity and bandwidth sensitivity of the SOC; when the SOC is more sensitive to latency, determine that the predicted compression state of the data in the target memory unit is the uncompressed state.
[0198] When the prediction identifiers include four types of identifiers: non-compressed identifier (corresponding prediction probability is 100%), suspected non-compressed identifier (corresponding prediction probability is 50%), suspected compressed identifier (corresponding prediction probability is 50%), and compressed identifier (corresponding prediction probability is 100%), the non-compressed identifier or the compressed identifier can be set for direct judgment, while the suspected non-compressed identifier or the suspected compressed identifier is set to require further combination with the actual requirements of the current SOC to be determined. For example, when the target prediction identifier is a suspected compressed identifier or a suspected non-compressed identifier, further determine the current SOC's demand for bandwidth or latency. If the current SOC is more sensitive to latency, then determine the predicted compression state of the data in the target memory unit as the non-compressed state.
[0199] After the predictor 101 predicts the compression state of the data to be accessed (i.e., the predicted compression state), the corresponding data to be accessed can be read through the corresponding access rules or access methods. The following is an exemplary description of how the controller 102 reads the target data in the target memory unit based on the predicted compression state. Specifically, it can include the following two cases:
[0200] (1) The prediction result is the compressed state
[0201] If the determined predicted compression state is the compressed state, the controller 102 specifically parses the first physical address corresponding to the target memory unit according to the rules of compressed data and reads the data stored in the first physical address. In the embodiment of the present application, when the data in the target memory unit is predicted to be in the compressed state, the physical address corresponding to the target unit is parsed according to the rules of compressed data. Since the physical address corresponding to the data will inevitably become shorter after being compressed, when the data is predicted to be in the compressed state, the physical address read is smaller than the physical address corresponding to the original data.
[0202] In a possible implementation, if the target memory unit includes two adjacent memory units, and the two adjacent memory units include a high-order memory unit and a low-order memory unit; reading the target memory unit according to the rule of compressed data includes: reading the low-order memory unit. Optionally, when the cacheline granularity is 64B, the compressed data in the embodiments of the present application can be placed in the lower 64 bits, that is, the compressed data is stored in the lower 64 bits of two consecutive cachelines, and there is a bit in each cacheline in the off-chip memory to indicate whether the data is actually in a compressed state or an uncompressed state. In the embodiments of the present application, when compressing a certain data, for example, compressing 128B of data (i.e., the original data involves two memory units) to 64B, and assuming that each memory unit can store 64B of data, the compressed data can be stored at the address position of the lower 64B, that is, stored in the low-order memory unit of two adjacent memory units to achieve data compression; correspondingly, when the data needs to be read, only the lower 64B needs to be read, that is, only the data of one memory unit needs to be read.
[0203] (2) The prediction result is an uncompressed state
[0204] If the determined predicted compression state is an uncompressed state, the controller 102 specifically parses the second physical address corresponding to the target memory unit according to the rule of the uncompressed state and reads the data stored in the second physical address. In the embodiments of the present application, when the data of the target memory unit is predicted to be in an uncompressed state, the physical address corresponding to the target unit is parsed according to the rule of uncompressed data. Since the physical address corresponding to the data if it is not compressed is the same as the original data, when the data is predicted to be in an uncompressed state, the physical address read is the same as the physical address corresponding to the original data.
[0205] In a possible implementation, if the target memory unit includes two adjacent memory units, and the two adjacent memory units include a high-order memory unit and a low-order memory unit; reading the target memory unit according to the rule of normal data includes: reading the high-order memory unit and the low-order memory unit. In the embodiments of the present application, when no compression processing is required for a certain data, that is, when writing normally, for example, writing 128B of data into two memory units with a size of 64B; correspondingly, when the data needs to be read, the complete data bits are read according to the rule of normal uncompressed data, that is, the high-order memory unit and the low-order memory unit need to be read simultaneously.
[0206] Further, after the controller 102 reads the target data, it is necessary to determine the actual compression state of the target data according to the read target data. In a possible implementation, after the controller 102 reads the target data, it obtains the compression state flag carried in the target data, and the compression state flag is used to indicate whether the target data is compressed data; then, the controller 102 can determine the actual compression state of the target data based on the compression state flag. In the embodiments of the present application, the data in the memory unit itself carries a compression state flag (such as metadata) for indicating whether the data is compressed, that is, actually reading the data in the memory unit can obtain the true situation of whether the data is compressed data. That is, in the embodiments of the present application, the predicted compression state is the state predicted based on the state prediction table, and the actual compression state is the state determined based on the metadata of the data itself.
[0207] Still further, when the predicted compression state determined by the predictor 101 is the compression state and the actual compression state of the data is the non-compression state, the controller 102 can also reread the target data based on the non-compression state. In the embodiments of the present application, when the predicted compression state is the compression state, but the actual compression state is the non-compression state, it will cause the data to be read according to the compression rule when reading the data (for example, only reading some high bits or some low bits), and at this time, the problem of failure to read the target data will occur. Therefore, when it is found that there is a problem with data reading, it is necessary to reread the target data according to the rule of the non-compression state.
[0208] In a possible implementation, the controller 102 reads data in the target memory cell based on the predicted compression state. Specifically, when the predicted compression state is the uncompressed state and the actual compression state is the compressed state, the controller reads the first compressed data and the second compressed data stored in the target memory cell. The SOC includes a prefetch cache. The apparatus further includes: a decompression module 104, configured to decompress the first data and the second data; and the controller 102 is further configured to read the decompressed first data according to the memory access request and store the decompressed second data into the prefetch cache. Optionally, when the storage space of the system on chip (SOC) is sufficient, the prefetch cache (PB) may be set to cache the on-chip data, and the data cached in the PB is the decompressed data; when the storage space of the SOC is insufficient, the prefetch cache PB may not be set. The solution of the embodiments of the present application can be implemented for both of these two cases. In the embodiments of the present application, when it is predicted based on the state prediction table that the currently to-be-accessed data is in the uncompressed state, but the actually to-be-accessed data is in the compressed state, it may cause unnecessary data to be read when reading the data. However, considering the continuity of data access, the unnecessary redundant data can also be synchronously read into the on-chip cache to avoid waste of the data read bandwidth. When the next memory access request is issued, the on-chip cache can be searched first and the data can be read from the cache, reducing the bandwidth resource overhead between the SOC and the memory. That is, the data can be directly read from the on-chip cache within the SOC, without the need to read the data through the bandwidth between the on-chip and off-chip.
[0209] After the controller 102 reads the target data in the memory 20, the predictor 101 may update the target prediction identifier based on the predicted compression state and the actual compression state. In the embodiments of the present application, the types of the set prediction identifiers include at least two groups of identifiers. Each group of identifiers is used to indicate the prediction probability that the data is in the compressed state or the uncompressed state, and the prediction probabilities corresponding to the same group of identifiers are the same, and the prediction probabilities corresponding to different groups of identifiers are different. Further optionally, the at least two groups of identifiers include a first group of identifiers and a second group of identifiers. The first group of identifiers are respectively an uncompressed identifier and a compressed identifier, and the second group of identifiers are respectively a suspected uncompressed identifier and a suspected compressed identifier; the prediction probability corresponding to the first group of identifiers is higher than the prediction probability corresponding to the second group of identifiers.
[0210] Exemplarily, please refer to Figure 5A , Figure 5A which is a schematic diagram of the state transition of a prediction identifier provided in the embodiments of the present application. In Figure 5A , the prediction identifier may transition between an uncompressed identifier, a suspected uncompressed identifier, a suspected compressed identifier, and a compressed identifier, and from Figure 5AIt can be seen that the state transition occurs between two adjacent identifiers, and the identifiers at both ends can remain unchanged. Correspondingly, please refer to Figure 5B , Figure 5B which is a schematic diagram of the state transition of a predicted identifier provided in an embodiment of the present application. In Figure 5B , the non-compressed identifier, the suspected non-compressed identifier, the suspected compressed identifier, and the compressed identifier are set to 0, 1, 2, and 3 respectively. Therefore, the predicted identifier can transition between these identifiers. Similarly, it can also transition between adjacent ones, and only the two endpoints can remain unchanged.
[0211] Exemplarily, the above predictor 101 can determine whether the data in the cacheline requested by the previous (e.g., the most recent in time) memory access is predicted correctly (correct prediction cases: compressed data → compressed, non-compressed data → non-compressed; preset incorrect cases: compressed data → non-compressed, non-compressed data → compressed). If the prediction is correct, there is no need for a second re-read. If the prediction is incorrect, a second re-read may be required. In a possible implementation, the predicted identifier includes two groups of identifiers, that is, four cases under two prediction probabilities. For example, it includes 0, 1, 2, 3, where "0" represents the non-compressed identifier (e.g., the prediction probability is 100%), "1" represents the suspected non-compressed identifier (e.g., the prediction probability is 50%), "2" represents the suspected compressed identifier (e.g., the prediction probability is 50%), and "3" represents the strongly compressed identifier (e.g., the prediction probability is 100%). That is, the first group of identifiers includes 0 and 3, the second group of identifiers includes 1 and 2, and the prediction probability corresponding to the first group of identifiers is 100%, and the prediction probability corresponding to the second group of identifiers is 50%. Specifically, as shown in Table 1 below:
[0212] Table 1
[0213]
[0214] Further, based on the fact that the predicted identifier includes the four identifiers in Table 1 above, the following describes how the predictor 101 updates the predicted identifier in the state prediction table based on the predicted compression state of the target data and the actual compression state of the read target data. Specifically, it can include the following six adjustment and update strategies for the predicted identifier:
[0215] (1) If both the predicted compression state and the actual compression state are non-compressed states, and the target predicted identifier is currently the non-compressed identifier, then keep the target predicted identifier unchanged.
[0216] Specifically, when the compression state of the target data predicted by the predictor 101 for the target memory cell is the uncompressed state (for example, at this time, the target prediction flag may currently be an uncompressed flag, a suspected uncompressed flag, or a suspected compressed flag), and when the controller 102 determines that the actual compression state of the target data is also the uncompressed state after reading the target data, that is, when the state predicted by the predictor is consistent with the actual compression state of the data, it can be considered that this prediction is accurate. At this time, it is necessary to further combine the prediction flag corresponding to the identifier of the memory area (target memory area) corresponding to the data, and comprehensively determine how to update the prediction flag currently. The reason is that although the predicted compression state determined by the predictor 101 is the corresponding compression state or uncompressed state, the prediction flag can correspond to multiple flags (for example, the four flags in Table 1 above), that is, the prediction flag not only includes the prediction flag for predicting the compression or uncompressed state, but can further include the prediction flags for the compression state or uncompressed state under different prediction probabilities. In this way, the robustness of the data compression state prediction can be greatly improved. Because multiple prediction flags can have multiple prediction states for whether a memory area is compressed under different prediction probabilities, such as predicted to be 100% compressed, 80% compressed, 50% compressed, 30% compressed, 100% uncompressed, 80% uncompressed, 50% uncompressed, 30% compressed, etc. In this way, the prediction result will not change violently, avoiding crashing or failing due to certain abnormal situations. Based on the above, in the embodiment of the present application, the prediction flag corresponding to the memory area in the state prediction table is updated based on the actual compression state of the read data, that is, the actual compression state (i.e., compressed or uncompressed) of each read data inside a memory area can be used to feedback and update the overall prediction result of the memory area, that is, the prediction flag. When the prediction result of the compression state of the data is the same as the actual result, it means that the current prediction direction is correct. If the current target prediction flag is already a compression flag or uncompressed flag with a prediction probability of 100%, then the current target prediction flag can be kept unchanged. If the current target prediction flag is a prediction probability less than 1 (such as 50%, 60%, 80%, etc.), then it can continue to be strengthened in the direction of the currently determined actual compression state or uncompressed state. For example, if 0, 1, 2, 3 represent the uncompressed flag, the suspected uncompressed flag, the suspected compressed flag, and the compressed flag respectively, then strengthening in the direction of the compressed flag means adding 1 to the current prediction flag, and strengthening in the direction of the uncompressed flag means subtracting 1 from the current prediction flag. Therefore, in the embodiment of the present application, if the predicted compression state and the actual compression state are both the uncompressed state, and the target prediction flag is currently the uncompressed flag, then the target prediction flag remains unchanged.
[0217] (2) If both the predicted compression state and the actual compression state are non-compressed states, and the target prediction flag is currently a suspected non-compressed flag or a suspected compressed flag, then adjust the target prediction flag in the direction of the non-compressed flag.
[0218] Specifically, when the compression state of the target data predicted by the predictor 101 for the target memory cell is a non-compressed state (for example, at this time the target prediction flag may currently be a non-compressed flag, a suspected non-compressed flag, or a suspected compressed flag), and when the controller 102 determines that the actual compression state of the target data is also a non-compressed state after reading the target data, that is, when the state predicted by the predictor is consistent with the actual compression state of the data, it can be considered that this prediction is accurate. At this time, it is necessary to further combine the prediction flag corresponding to the identifier of the memory area (target memory area) corresponding to this data to comprehensively determine how to update the prediction flag currently; and when the target prediction flag is currently a suspected non-compressed flag or a suspected compressed flag, then adjust the target prediction flag in the direction of the non-compressed flag (i.e., the actual result). For example, if 0, 1, 2, 3 represent the non-compressed flag, the suspected non-compressed flag, the suspected compressed flag, and the compressed flag respectively, then the meaning of strengthening in the direction of the non-compressed flag is to subtract 1 based on the current prediction flag.
[0219] (3) If the predicted compression state is a non-compressed state and the actual compression state is a compressed state, then adjust the target prediction flag in the direction of the compressed flag.
[0220] Specifically, the predictor 101 is specifically used to adjust the current target prediction flag in the opposite direction of the current target prediction flag when the prediction result of the data compression state is inconsistent with the actual result. That is, if the predicted compression state is a non-compressed state (for example, in the case where the target prediction flag is currently a non-compressed flag, a suspected non-compressed flag, or a suspected compressed flag), and the actual compression state is a compressed state, that is, the prediction result is inconsistent with the actual result, then adjust the current target prediction flag in the direction of the compressed flag (i.e., the actual result). For example, if 0, 1, 2, 3 represent non-compressed, suspected non-compressed flag, suspected compressed flag, and compressed flag respectively, then the meaning of strengthening in the direction of the compressed flag is to add 1 based on the current prediction flag.
[0221] (4) If the predicted compression state is a compressed state and the actual compression state is a non-compressed state, then adjust the target prediction flag in the direction of the non-compressed flag.
[0222] Specifically, when the prediction result of the data compression state is inconsistent with the actual result, the predictor 101 is specifically configured to adjust the current target prediction identifier in the opposite direction of the current target prediction identifier. That is, if the predicted compression state is the compression state (for example, when the current target prediction identifier is the suspected non-compression identifier, the suspected compression identifier, or the compression identifier), and the actual compression state is the non-compression state, that is, the prediction result is inconsistent with the actual result, then the current target prediction identifier is adjusted in the direction of the non-compression identifier (i.e., the actual result). For example, if 0, 1, 2, and 3 represent the non-compression identifier, the suspected non-compression identifier, the suspected compression identifier, and the compression identifier respectively, then the meaning of strengthening in the direction of the non-compression identifier is to subtract 1 from the current prediction identifier.
[0223] (5) If both the predicted compression state and the actual compression state are the compression state, and the current target prediction identifier is the compression identifier, then the target prediction identifier remains unchanged.
[0224] Specifically, when the data compression prediction result is consistent with the actual result, the predictor 101 also needs to determine whether the current prediction identifier is consistent with the above actual result. If they are consistent, the target prediction identifier remains unchanged. If they are inconsistent, the current prediction identifier needs to be strengthened in the opposite direction of the current actual result. That is, if both the predicted compression state and the actual compression state are the compression state (for example, when the current target prediction identifier is the suspected non-compression identifier, the suspected compression identifier, or the compression identifier), and the current target prediction identifier is the compression identifier, then the target prediction identifier remains unchanged. For example, if 0, 1, 2, and 3 represent the non-compression identifier, the suspected non-compression identifier, the suspected compression identifier, and the compression identifier respectively, then currently the target prediction identifier continues to remain 3 unchanged.
[0225] (6) If both the predicted compression state and the actual compression state are the compression state, and the current target prediction identifier is the suspected non-compression identifier or the suspected compression identifier, then the target prediction identifier is adjusted in the direction of the compression identifier.
[0226] Specifically, when the data compression prediction result is consistent with the actual result, it is also necessary to determine whether the current prediction flag is consistent with the above actual result. If they are consistent, the target prediction flag remains unchanged. If they are not consistent, the current prediction flag needs to be strengthened in the opposite direction of the current actual result. That is, if both the predicted compression state and the actual compression state are in the compressed state (for example, when the target prediction flag is currently a suspected non-compressed flag, a suspected compressed flag, or a compressed flag), and the target prediction flag is currently a suspected non-compressed flag or a suspected compressed flag, then the target prediction flag is adjusted towards the compressed flag (i.e., the actual result). For example, if 0, 1, 2, and 3 represent non-compressed flag, suspected non-compressed flag, suspected compressed flag, and compressed flag respectively, then strengthening towards the compressed flag means adding 1 to the current prediction flag.
[0227] Exemplarily, for a certain memory page (say, the target memory page), the compression state flag of the target memory page is stored in the state prediction table (such as a hash table) in the predictor. Assuming the initial flag is 0, after the first memory access operation for this memory page, if it is verified that the prediction is correct, that is, the prediction flag of the target memory area to which this target memory page belongs continues to be maintained as 0 (i.e., non-compressed state). When the second memory access operation is performed on other target memory pages in this target memory area next time and it is verified that the prediction is incorrect, then not only is it necessary to reread the data twice, but also the prediction flag needs to be updated. The specific update situation of the prediction flag in the state prediction table can be seen in Table 2 below.
[0228] Table 2
[0229]
[0230]
[0231] When it is verified that the actual compression state is the non-compressed state (that is, the currently accessed memory page is actually in the non-compressed state), and the determined predicted compression state is the non-compressed state, and the prediction flag corresponding to the memory area to which this memory page belongs is currently = 0, then the target prediction flag remains 0 unchanged.
[0232] When it is verified that the actual compression state is the non-compressed state (that is, the currently accessed memory page is actually in the non-compressed state), and the determined predicted compression state is the non-compressed state, and the prediction flag corresponding to the memory area to which this memory page belongs is currently < 3 and ≠ 0 (i.e., the prediction flag is 1 or 2), then the target prediction flag is determined to be -1.
[0233] When it is verified that the actual compression state is the uncompressed state (i.e., the currently accessed memory page is actually in the uncompressed state), and the determined predicted compression state is the compressed state, and the predicted identifier corresponding to the memory area to which this memory page belongs is currently > 0 (i.e., the predicted identifier is 1 or 2 or 3), then the target predicted identifier is decreased by 1.
[0234] When it is verified that the actual compression state is the compressed state (i.e., the currently accessed memory page is actually in the compressed state), and the determined predicted compression state is the uncompressed state, and the predicted identifier corresponding to the memory area to which this memory page belongs is currently < 3 (i.e., the predicted identifier is 0 or 1 or 2), then the target predicted identifier is increased by 1.
[0235] When it is verified that the actual compression state is the compressed state (i.e., the currently accessed memory page is actually in the compressed state), and the determined predicted compression state is the compressed state, and the predicted identifier corresponding to the memory area to which this memory page belongs is currently = 3, then the target predicted identifier remains unchanged at 3.
[0236] When it is verified that the actual compression state is the compressed state (i.e., the currently accessed memory page is actually in the compressed state), and the determined predicted compression state is the compressed state, and the predicted identifier corresponding to the memory area to which this memory page belongs is currently > 0 and ≠ 3 (i.e., the predicted identifier is 1 or 2), then the target predicted identifier is increased by 1.
[0237] Exemplarily, after determining the compression state (i.e., the actual compression state) of a hit CacheLine, the change process of the current predicted identifier (i.e., corresponding to the actual compression state) and the adjusted predicted identifier can be specifically referred to Table 3 below.
[0238] Table 3
[0239]
[0240] The embodiments of the present application mainly solve the strong dependence of the bandwidth compression technology on the metadata cache and the pain point that the bandwidth compression technology has a large demand for on-chip storage space. Considering the deficiencies of the existing methods, a state machine of a fine-grained compression state predictor based on the BiModal prediction algorithm is proposed, which can accurately predict the compression state of a flexible-sized area. Further, the embodiments of the present application can flexibly support the metadata cache (MetadataCache) structure of the previous work, and can cooperate with the (Metadata Cache) to achieve more accurate compressed data management when there is enough space. That is, while ensuring the prediction accuracy, the space overhead is reduced. Further, the method in the present application can improve the robustness of the prediction, that is, it will not cause the problem of excessive oscillation and too low accuracy of the prediction due to a single prediction error.
[0241] In the embodiments of the present application, bandwidth compression technology is involved in compressing and storing relevant data. Further, the prediction identifier of the memory area is also used to predict the data compression state. During this process, when it comes to writing data into the memory, it is first necessary to determine whether the data to be written needs to be compressed. If compression is required, compression processing is performed according to the preset compression rules. Among them, the preset compression rules in the embodiments of the present application may include various compression methods such as complete compression or hybrid compression.
[0242] In a possible implementation manner, the memory management device further includes a compression module 105; the controller 102 is further configured to receive a write request for the data to be written, and determine whether the data to be written needs to be compressed, and the size of the data to be written is M bits; the compression module 105 is configured to compress the M bits of the data to be written if all of the data to be written needs to be compressed; the controller 102 is further configured to write the compressed data into the corresponding memory cell; the compression module 105 is further configured to compress N bits of the data to be written if a part of the data to be written needs to be compressed; the controller 102 is further configured to write the compressed data into the first memory cell, and write the remaining (M - N) bits of the data to be written into the second memory cell; the controller 102 is further configured to: update the compression status flag bits in the first memory cell and the second memory cell to a compression flag and a non-compression flag respectively. In the embodiments of the present application, when a write request for data is received, it is necessary to first determine whether the data to be written needs to be compressed. If compression is required, it is first compressed and then written into the memory. If compression is not required, it can be directly written into the memory; among them, when compression is required, it can be divided into the cases of complete compression or partial compression. When complete compression is required, the data to be written is compressed as a whole and then written into the memory cell. When only partial compression is required, the part of the data to be written that needs to be compressed is compressed and written into the corresponding memory cell, while the part of the data to be written that does not need to be compressed is written without compression and written into another memory cell; further, the data to be written stored in different memory cells can be set with different compression status flag bits for different memory cells, that is, for the same data to be written written into different memory cells, independent compression flag bits can be set. For example, for 128B of data to be written (assuming the memory cell size is 64B), if 64B of it needs to be compressed and the other 64B does not need to be compressed, the 64B of data that needs to be compressed is compressed and written as 32B of data into the memory cell, while the other 64B of data that does not need to be compressed is directly written into an adjacent another memory cell. In summary, in the embodiments of the present application, partial compression of data can be achieved through the hybrid compression granularity mechanism, that is, more fine-grained and accurate compression can be achieved for the same data.
[0243] In a possible implementation, when compressing the M bits of the data to be written, the compression module 105 is specifically configured to: split the M bits of the data to be written into K-bit data and (M-K)-bit data respectively, and compress the K-bit data and the (M-K)-bit data respectively. In the embodiments of the present application, when all the data to be written needs to be compressed, it is also possible to locally split the data to be written and compress the split data bits respectively. For example, the hybrid granularity includes compressing 128B of data to 64B, or compressing two 64B into 32B respectively. Correspondingly, decompression also supports decompressing 64B into 128B, or decompressing two independent 32B into 64B respectively. By supporting hybrid granularity compression, the compression ratio is effectively increased in the embodiments of the present application.
[0244] Furthermore, the memory management device in the embodiments of the present application further includes: a dynamic switch module 106 for: obtaining the current prediction accuracy rate of the state prediction table calculated based on the predicted compression state and the actual compression state of multiple memory access data; when the data to be written needs to be compressed and the prediction accuracy rate is lower than a preset threshold, then control the compression module to turn off the compression function for writing the data to be written into the corresponding memory unit in an uncompressed state; when the data to be written needs to be compressed and the prediction accuracy rate is higher than a preset threshold, then control the compression module to turn on the compression function for compressing the data to be written and writing it into the corresponding memory unit. In the embodiments of the present application, the current prediction accuracy rate of the state prediction table can be calculated based on multiple prediction results of the state prediction table. For example, when the predicted compression state is consistent with the actual compression state, it means that the prediction is accurate and the prediction accuracy rate increases accordingly. When the predicted compression state is inconsistent with the actual compression state, it means that the prediction is incorrect and the prediction accuracy rate decreases accordingly. When the cumulative prediction accuracy rate is lower than the preset accuracy rate threshold, the prediction function can be controlled to be turned off, that is, when a certain data to be written needs to be compressed and written, it is controlled not to be compressed but directly written in an uncompressed state, avoiding repeated reading and wasting resources and bandwidth due to inaccurate prediction. When the cumulative prediction accuracy rate is higher than the preset accuracy rate threshold, the prediction function can be controlled to be turned on, that is, when a certain data to be written needs to be compressed and written, it is controlled to be compressed and written in a compressed state, so as to reduce the bandwidth resources between on-chip and off-chip through the bandwidth compression technology when the prediction accuracy rate is guaranteed.
[0245] Please refer to Figure 6 , Figure 6A schematic diagram of the process for a controller to process read and write requests provided by an embodiment of the present application. The embodiment of the present application comprehensively considers the deficiencies of existing methods and proposes a new lightweight and scalable compressed data management scheme with a BiModal predictor as the core path. Among them, in Figure 6 exemplarily, taking the logic state (or state machine) in the controller as the main body, for the read process, it mainly includes a compressed idle state, a compressed search state, a compressed parsing state (read request), a read compressed data state (read transaction), a repeated read state, a memory allocation state, a read memory allocation state, and a transaction fully processed state; exemplarily, taking the logic state machine in the controller as the main body, for the write process, it mainly includes a compressed idle state, a compressed search state, a compressed parsing state (write request), a read compressed data state (write transaction), a write compressed data state, and a transaction fully processed state.
[0246] Taking the case where the processor issues a DDR memory access request and supports two lengths of 64B and 128B as an example, the access to DDR after being processed by the decompression module or the compression module is described exemplarily. Among them, it mainly includes that the decompression module or the compression module parses the physical memory address according to the address and length of the access request, and according to the compressed state or the uncompressed state, and then proceeds with the relevant processes of the access. Taking the read request as an example with a granularity of 128B or 64B, the specific processes of processing 128B and 64B read requests in the embodiment of the present application are described exemplarily:
[0247] Scenario 1: 128B read process
[0248] When the processor issues a 128B DDR request and the controller 102 receives a 128B read request in the compressed idle state, it first queries whether the data exists in the prefetch buffer (PrefetchBuffer) 103 through the compressed search state. If the complete 128B data exists, the request ends. If partial 64B data exists, the remaining request is processed as a 64B read. If the query is missing, the controller 102 determines the compression state of the data through the compressed parsing state of the read request (i.e., according to the prediction result in the predictor 101). If the prediction is compressed data, it accesses the low 64B address position in 64B length through the read compressed data state and the memory allocation state, otherwise it accesses the complete 128B data through the transaction fully processed state. If the prediction fails and the data itself is in an uncompressed state, it is necessary to repeatedly read the high 64B address position through the repeated read state and update the state prediction table in the predictor 101. If the prediction fails and the data is in a compressed state, it puts the redundant data into the prefetch buffer (PrefetchBuffer) 103 through the read memory allocation state and updates the state prediction table in the predictor 101.
[0249] Scenario 2: 64B read process
[0250] The processor 100 issues a 64B DDR request. After the controller 102 compresses the 64B read request received in the idle state, it first queries the prefetch buffer 103 through the compression lookup state to check if the data exists. If the data exists, the request ends; if the query is missing, and the memory access request is to read the lower 64B of 128B, that is, directly access the address position of the lower 64B. If the controller 102 determines through the compression parsing state of the read request that the data is compressed data, it places the redundant data in the prefetch buffer 103 and updates the state prediction table in the predictor 101. If the query is missing, and the memory access request is to read the upper 64B of 128B, the controller 102 determines the compression state of the data through the compression parsing state (i.e., according to the state prediction table in the predictor 101). If the prediction is compressed data, it accesses the address position of the lower 64B in 64B length through the read compressed data state and the memory allocation state, otherwise it accesses the data of the upper 64B. If the prediction fails and the data itself is in an uncompressed state, it is necessary to repeatedly read the address position of the upper 64B through the repeated read state and update the predictor. If the prediction fails and the data is in a compressed state, it places the redundant data in the prefetch buffer 103 through the read memory allocation state, repeatedly accesses the lower 64B, and updates the predictor.
[0251] Please refer to Figure 6 , Figure 6 FIG. is a schematic flowchart of a process for processing read and write requests provided by an embodiment of the present application. Taking the write request as 128B granularity or 64B granularity as an example, the specific process of processing 128B and 64B write requests in the embodiment of the present application is described exemplarily as follows:
[0252] Scenario 1: 128B write process
[0253] The controller 102 determines the compression state of the data in the compression parsing state of the write request (i.e., through the compression module 105) and writes to the DDR according to the compression state. If it can be compressed, the controller 102 controls to compress the data and write it to the lower 64B through the write compressed data state (or through the read compressed data state and the write compressed data state); if it cannot be compressed, the controller 102 controls to write the data completely to 128B through the practical complete processing state.
[0254] It should be noted that in the description of the above embodiments, the numerical serial numbers (one), (two), (three), etc. do not represent a strict time execution order or sequence. That is, the serial numbers do not limit the execution order of the process or the sequence of state occurrences. Moreover, the embodiments in the above different scenarios do not limit each other. The above different embodiments can be executed independently of each other or in combination with each other, and will not be elaborated in detail here one by one.
[0255] In general-purpose servers, as the number of cores continues to increase, memory bandwidth becomes a performance bottleneck. Using the device or method provided by the embodiments of the present application, DDR data compression can be achieved with extremely low additional overhead. In the scenario of general-purpose CPU servers, the specific requirements for peripheral modules are as follows.
[0256] (1) For processor cores
[0257] Traditional data compression requires compressing 128B to 64B, so the processor core needs to send a 128B read / write request to obtain performance benefits. The present application can utilize the prefetch buffer (PrefetchBuffer) to avoid DDR memory access under a 64B request and still obtain performance benefits. The increase in the 128B read / write requests of the processor core will amplify the effect of the compression technology. The specific sources of 128B read requests include consecutive accesses and consecutive prefetching, etc.
[0258] (2) For the coherence directory
[0259] To avoid the negative benefits brought by excessive read-write, the embodiments of the present application can rely on the directory to record the compression state of on-chip data.
[0260] The technical effects of the embodiments of the present application are mainly reflected in performance improvement and space overhead savings. The embodiments of the present application conduct experiments on the SPEC2017 test set for the general-purpose server scenario and compare the benefits of the scheme based on metadata (MetadataCache) (MC Baseline) and the present scheme. Among them, the overhead of the compression method is shown in Table 4 below.
[0261] Table 4
[0262]
[0263] As can be seen from Table 4 above, compared with the performance benefits of the non-compression mode. The present application can achieve a performance effect that exceeds the scheme based on the metadata cache with large overhead. The main improvement of the embodiments of the present application over the prior art lies in proposing a compression management scheme for MetadataCache that does not rely on a large amount of on-chip space. Through the accurate prediction of the BiModal predictor and the perfect handling of speculation failures, the effect beyond the metadatacache is achieved. In addition, the embodiments of the present application also propose a hybrid-granularity compression and decompression mechanism. By supporting independent compression of 64B to 32B and an independent compression flag bit, the negative benefits of read-write are greatly reduced.
[0264] Please refer to Figure 7 , Figure 7It is a schematic flowchart of a memory management method provided by an embodiment of the present application. This memory management method can be applied to a system on chip (SOC), an off-chip memory management system, a memory access system, an electronic device including an SOC, or a memory management module in an electronic device, etc. This method may include the following steps S701 - step S705, where
[0265] Step S701: Manage a status prediction table, where the status prediction table includes a mapping relationship between an identifier of a memory area and a prediction identifier; wherein, the memory area includes a plurality of memory cells, and the prediction identifier is used to predict whether the data in the plurality of memory cells is in a compressed state.
[0266] Step S702: Receive a memory access request for a target memory cell, and determine the identifier of the target memory area to which the target memory cell belongs.
[0267] Step S703: Query the target prediction identifier corresponding to the identifier of the target memory area from the status prediction table, and determine the predicted compression state of the data in the target memory cell based on the target prediction identifier.
[0268] Step S704: Read the target data in the target memory cell based on the predicted compression state, and determine the actual compression state of the target data.
[0269] Step S705: Update the target prediction identifier based on the predicted compression state and the actual compression state.
[0270] In a possible implementation manner, the size of the memory cell is the memory access granularity size; the method further includes: receiving a memory access request, where the memory access request includes a starting address and an access length; determining a target memory cell based on the starting address and the access length, and the target memory cell includes one or more memory cells.
[0271] In a possible implementation manner, the type of the prediction identifier includes at least two groups of identifiers, where each group of identifiers is used to indicate the prediction probability that the data is in a compressed state or a non-compressed state, and the prediction probabilities corresponding to the same group of identifiers are the same, and the prediction probabilities corresponding to different groups of identifiers are different.
[0272] In a possible implementation manner, the at least two groups of identifiers include a first group of identifiers and a second group of identifiers. The first group of identifiers are respectively a non-compressed identifier and a compressed identifier, and the second group of identifiers are respectively a suspected non-compressed identifier and a suspected compressed identifier; the prediction probability corresponding to the first group of identifiers is higher than the prediction probability corresponding to the second group of identifiers.
[0273] In a possible implementation, determining the predicted compression state of the data in the target memory unit based on the target prediction identifier includes: if the target prediction identifier is the compression identifier, determining that the predicted compression state of the data in the target memory unit is the compressed state; if the target prediction identifier is the non-compression identifier, determining that the predicted compression state of the data in the target memory unit is the non-compressed state.
[0274] In a possible implementation, the method is applied to a memory access system; determining the predicted compression state of the data in the target memory unit based on the target prediction identifier includes: if the target prediction identifier is currently the suspected compression identifier or the suspected non-compression identifier, determining the relationship between the current latency sensitivity and bandwidth sensitivity of the memory access system; when the memory access system is more sensitive to bandwidth, determining that the predicted compression state of the data in the target memory unit is the compressed state; when the memory access system is more sensitive to latency, determining that the predicted compression state of the data in the target memory unit is the non-compressed state.
[0275] In a possible implementation, updating the target prediction identifier based on the predicted compression state and the actual compression state includes: if both the predicted compression state and the actual compression state are non-compressed states, and the target prediction identifier is currently the non-compression identifier, keeping the target prediction identifier unchanged; if both the predicted compression state and the actual compression state are non-compressed states, and the target prediction identifier is currently the suspected non-compression identifier or the suspected compression identifier, adjusting the target prediction identifier towards the non-compression identifier.
[0276] In a possible implementation, updating the target prediction identifier based on the predicted compression state and the actual compression state includes: if the predicted compression state is the non-compressed state and the actual compression state is the compressed state, adjusting the target prediction identifier towards the compression identifier.
[0277] In a possible implementation, updating the target prediction identifier based on the predicted compression state and the actual compression state includes: if the predicted compression state is the compressed state and the actual compression state is the non-compressed state, adjusting the target prediction identifier towards the non-compression identifier.
[0278] In a possible implementation, updating the target prediction identifier based on the predicted compression state and the actual compression state includes: if both the predicted compression state and the actual compression state are compression states, and the target prediction identifier is currently a compression identifier, then keep the target prediction identifier unchanged; if both the predicted compression state and the actual compression state are compression states, and the target prediction identifier is currently the suspected non-compression identifier or the suspected compression identifier, then adjust the target prediction identifier in the direction of the compression identifier.
[0279] In a possible implementation, the method further includes: creating a state prediction table, establishing a mapping relationship between the identifiers of each memory area and the corresponding prediction identifiers in the state prediction table; setting the initial identifier of each prediction identifier to indicate a non-compression state.
[0280] In a possible implementation, reading the target data in the target memory unit based on the predicted compression state includes:
[0281] If the predicted predicted compression state is a compression state, then parse the first physical address corresponding to the target memory unit according to the rules of compressed data, and read the data stored in the first physical address.
[0282] In a possible implementation, reading the target data in the target memory unit based on the predicted compression state includes: if the predicted predicted compression state is a non-compression state, then parse the second physical address corresponding to the target memory unit according to the rules of the non-compression state, and read the data stored in the second physical address.
[0283] In a possible implementation, determining the actual compression state of the target data includes: after reading the target data, obtaining the compression state identifier carried in the target data, where the compression state identifier is used to indicate whether the target data is compressed data; determining the actual compression state of the target data based on the compression state identifier.
[0284] In a possible implementation, the method further includes: when the predicted compression state is a compression state and the actual compression state is a non-compression state, then reread the target data based on the non-compression state.
[0285] In a possible implementation, the method is applied to a memory access system; the reading of the target data in the target memory cell based on the predicted compression state includes: when the predicted compression state is a non-compressed state and the actual compression state is a compressed state, reading the first data and the second data stored in the target memory cell after compression; the memory access system includes a prefetch cache; the method further includes: decompressing the first data and the second data; reading the decompressed first data according to the memory access request, and storing the decompressed second data into the prefetch cache.
[0286] In a possible implementation, if the target memory cell includes two adjacent memory cells, and the two adjacent memory cells include a high-order memory cell and a low-order memory cell; the reading of the target memory cell according to the rule of compressed data includes: reading the low-order memory cell; the reading of the target memory cell according to the rule of normal data includes: reading the high-order memory cell and the low-order memory cell.
[0287] In a possible implementation, the method further includes: receiving a write request for data to be written, determining whether the data to be written needs to be compressed, and the size of the data to be written is M bits; if the data to be written needs to be fully compressed, compressing the M bits of the data to be written, and writing the compressed data into the corresponding memory cell; if a part of the data to be written needs to be compressed, compressing the N bits of the data to be written, writing the compressed data into the first memory cell, and writing the remaining (M - N) bits of the data to be written into the second memory cell; both M and N are integers greater than 1, and N is less than M; updating the compression state flag bits in the first memory cell and the second memory cell to a compression flag and a non-compression flag, respectively.
[0288] In a possible implementation, the compressing of the M bits of the data to be written includes: splitting the M bits of the data to be written into K-bit data and (M - K)-bit data respectively, and compressing the K-bit data and the (M - K)-bit data respectively.
[0289] In a possible implementation, the method further includes: calculating the current prediction accuracy rate of the state prediction table based on the predicted compression state and the actual compression state of multiple memory access data; when the data to be written needs to be compressed and the prediction accuracy rate is lower than a preset threshold, controlling the data to be written to be written into the corresponding memory cell in a non-compressed state; when the data to be written needs to be compressed and the prediction accuracy rate is higher than a preset threshold, controlling the data to be written to be written into the corresponding memory cell in a compressed state.
[0290] It should be noted that for the specific process of the memory management method described in the embodiments of the present application, reference may be made to the relevant descriptions in the above Figures 2A - 6 application embodiments described in, which will not be elaborated here.
[0291] The embodiments of the present application further provide a computer-readable storage medium, wherein the computer-readable storage medium may store a program, and when the program is executed by an electronic device, it includes some or all of the steps recorded in any one of the above method embodiments.
[0292] The embodiments of the present application further provide a computer program, the computer program includes instructions, and when the computer program is executed by an electronic device, the electronic device can execute some or all of the steps of any one of the memory management methods.
[0293] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not elaborated in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0294] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps may be implemented in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0295] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above unit division is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0296] The units described as separate components above may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0297] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0298] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc., specifically, the processor in the computer device) to execute all or part of the steps of the above-mentioned method in each embodiment of the present application. Among them, the aforementioned storage medium can include: USB flash drives, mobile hard disks, magnetic disks, optical discs, read-only memory (ROM), or random access memory (RAM), etc., various media that can store program codes.
[0299] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A memory management method, characterized in that, The method includes: Receiving a memory access request for a target memory unit and determining an identifier of a target memory region to which the target memory unit belongs; Querying a target prediction identifier corresponding to the identifier of the target memory region from a status prediction table, and determining a predicted compression status of data in the target memory unit based on the target prediction identifier; the status prediction table includes a mapping relationship between the identifier of the memory region and the prediction identifier; wherein, the memory region includes multiple memory units, and the prediction identifier is used to predict whether the data in the multiple memory units is in a compressed state; Reading target data in the target memory unit based on the predicted compression status and determining an actual compression status of the target data; Updating the target prediction identifier based on the predicted compression status and the actual compression status.
2. The method according to claim 1, wherein The type of the prediction identifier includes at least two groups of identifiers, wherein each group of identifiers is used to indicate a prediction probability that the data is in a compressed state or a non-compressed state, and the prediction probabilities corresponding to the same group of identifiers are the same, and the prediction probabilities corresponding to different groups of identifiers are different.
3. The method according to claim 2, wherein The at least two groups of identifiers include a first group of identifiers and a second group of identifiers, the first group of identifiers are respectively a non-compressed identifier and a compressed identifier, and the second group of identifiers are respectively a suspected non-compressed identifier and a suspected compressed identifier; the prediction probability corresponding to the first group of identifiers is higher than the prediction probability corresponding to the second group of identifiers.
4. The method according to claim 3, wherein The determining the predicted compression status of the data in the target memory unit based on the target prediction identifier includes: If the target prediction identifier is the compressed identifier, determining that the predicted compression status of the data in the target memory unit is the compressed state; If the target prediction identifier is the non-compressed identifier, determining that the predicted compression status of the data in the target memory unit is the non-compressed state.
5. The method according to claim 3 or 4, characterized in that, The method is applied to a memory access system; the determining the predicted compression status of the data in the target memory unit based on the target prediction identifier includes: If the target prediction identifier is currently the suspected compressed identifier or the suspected non-compressed identifier, determining the relationship between the current latency sensitivity and bandwidth sensitivity of the memory access system; When the memory access system is more sensitive to bandwidth, determining that the predicted compression status of the data in the target memory unit is the compressed state; When the memory access system is more sensitive to latency, determining that the predicted compression status of the data in the target memory unit is the non-compressed state.
6. The method according to any one of claims 3-5, characterized in that, The updating the target prediction identifier based on the predicted compression status and the actual compression status includes: If both the predicted compression status and the actual compression status are non-compressed states, and the target prediction identifier is currently the non-compressed identifier, keeping the target prediction identifier unchanged; If both the predicted compression status and the actual compression status are non-compressed states, and the target prediction identifier is currently the suspected non-compressed identifier or the suspected compressed identifier, adjusting the target prediction identifier in the direction of the non-compressed identifier.
7. The method according to any one of claims 3-6, characterized in that, The updating the target prediction identifier based on the predicted compression status and the actual compression status includes: If the predicted compression state is the uncompressed state and the actual compression state is the compressed state, adjust the target prediction identifier in the direction of the compression identifier.
8. The method according to any one of claims 3-7, characterized in that, Updating the target prediction identifier based on the predicted compression state and the actual compression state includes: If the predicted compression state is the compressed state and the actual compression state is the uncompressed state, adjust the target prediction identifier in the direction of the uncompressed identifier.
9. The method according to any one of claims 3-8, characterized in that, Updating the target prediction identifier based on the predicted compression state and the actual compression state includes: If both the predicted compression state and the actual compression state are the compressed state, and the target prediction identifier is currently the compression identifier, keep the target prediction identifier unchanged; If both the predicted compression state and the actual compression state are the compressed state, and the target prediction identifier is currently the suspected uncompressed identifier or the suspected compression identifier, adjust the target prediction identifier in the direction of the compression identifier.
10. The method according to any one of claims 1-9, characterized in that, The method further includes: Create a state prediction table, and establish a mapping relationship between the identifiers of each memory area and the corresponding prediction identifiers in the state prediction table; Set the initial identifier of each prediction identifier to indicate the uncompressed state.
11. The method according to any one of claims 1-10, characterized in that, Reading the target data in the target memory unit based on the predicted compression state includes: If the predicted compression state is the compressed state, parse the first physical address corresponding to the target memory unit according to the rules of compressed data, and read the data stored in the first physical address.
12. The method according to any one of claims 1-11, characterized in that, Reading the target data in the target memory unit based on the predicted compression state includes: If the predicted compression state is the uncompressed state, parse the second physical address corresponding to the target memory unit according to the rules of the uncompressed state, and read the data stored in the second physical address.
13. The method according to any one of claims 1-12, characterized in that, Determining the actual compression state of the target data includes: After reading the target data, obtain the compression state identifier carried in the target data, and the compression state identifier is used to indicate whether the target data is compressed data; Determine the actual compression state of the target data based on the compression state identifier.
14. The method according to any one of claims 1-13, characterized in that, The method further includes: When the predicted compression state is the compressed state and the actual compression state is the uncompressed state, reread the target data based on the uncompressed state.
15. The method according to any one of claims 1-14, characterized in that, The method is applied to a memory access system; reading the target data in the target memory unit based on the predicted compression state includes: When the predicted compression state is the uncompressed state and the actual compression state is the compressed state, read the first data and the second data stored in the target memory unit after compression; The memory access system includes a prefetch cache; the method further includes: Decompress the first data and the second data; Read the decompressed first data according to the memory access request, and store the decompressed second data in the prefetch cache.
16. The method according to any one of claims 1-15, characterized in that, The method further includes: Receive a write request for the data to be written, and determine whether the data to be written needs to be compressed. The size of the data to be written is M bits; If all the data to be written needs to be compressed, compress M bits of the data to be written and write the compressed data into the corresponding memory unit; If a part of the data to be written needs to be compressed, compress N bits of the data to be written, write the compressed data into the first memory unit, and write the remaining (M - N) bits of data in the data to be written into the second memory unit; both M and N are integers greater than 1, and N is less than M; Update the compression status flag bits in the first memory unit and the second memory unit to the compression flag and the non-compression flag, respectively.
17. The method according to claim 16, wherein The compression of M bits of the data to be written includes: Split the M bits of the data to be written into K-bit data and (M - K)-bit data respectively, and compress the K-bit data and the (M - K)-bit data respectively.
18. The method according to any one of claims 1 to 17, characterized in that, The method further includes: Calculate the current prediction accuracy rate of the status prediction table based on the predicted compression status and the actual compression status of multiple memory access data; When the data to be written needs to be compressed and the prediction accuracy rate is lower than the preset threshold, control the data to be written to be written into the corresponding memory unit in an uncompressed state; When the data to be written needs to be compressed and the prediction accuracy rate is higher than the preset threshold, control the data to be written to be written into the corresponding memory unit in a compressed state.
19. A memory management device, characterized in that, It includes: A predictor for managing a status prediction table, where the status prediction table includes a mapping relationship between the identifier of a memory area and a prediction identifier; wherein, the memory area includes multiple memory units, and the prediction identifier is used to predict whether the data in the multiple memory units is in a compressed state; A controller for receiving a memory access request for a target memory unit and determining the identifier of the target memory area to which the target memory unit belongs; The predictor is further configured to query the target prediction identifier corresponding to the identifier of the target memory area from the status prediction table, and determine the predicted compression status of the data in the target memory unit based on the target prediction identifier; The controller is further configured to read the target data in the target memory unit based on the predicted compression status and determine the actual compression status of the target data; The predictor is further configured to update the target prediction identifier based on the predicted compression status and the actual compression status.
20. The device according to claim 19, wherein The type of the prediction identifier includes at least two groups of identifiers, where each group of identifiers is used to indicate the prediction probability that the data is in a compressed state or an uncompressed state, and the prediction probabilities corresponding to the same group of identifiers are the same, and the prediction probabilities corresponding to different groups of identifiers are different.
21. The device according to claim 20, characterized in that, The at least two groups of identifiers include a first group of identifiers and a second group of identifiers. The first group of identifiers are respectively a non-compression identifier and a compression identifier, and the second group of identifiers are respectively a suspected non-compression identifier and a suspected compression identifier; the prediction probability corresponding to the first group of identifiers is higher than the prediction probability corresponding to the second group of identifiers.
22. The device according to claim 21, wherein, The predictor is specifically configured to: If the target prediction identifier is the compression identifier, determine that the predicted compression status of the data in the target memory unit is the compression state; If the target prediction identifier is the uncompressed identifier, determine that the predicted compression state of the data in the target memory cell is the uncompressed state.
23. The device according to claim 21 or 22, characterized in that, The device is applied to a memory access system; the predictor is specifically configured to: If the current target prediction identifier is the suspected compression identifier or the suspected uncompressed identifier, determine the relationship between the current latency sensitivity and bandwidth sensitivity of the memory access system; When the memory access system is more sensitive to bandwidth, determine that the predicted compression state of the data in the target memory cell is the compressed state; When the memory access system is more sensitive to latency, determine that the predicted compression state of the data in the target memory cell is the uncompressed state.
24. The device according to any one of claims 21-23, characterized in that, The predictor is specifically configured to: If both the predicted compression state and the actual compression state are the uncompressed state, and the current target prediction identifier is the uncompressed identifier, keep the target prediction identifier unchanged; If both the predicted compression state and the actual compression state are the uncompressed state, and the current target prediction identifier is the suspected uncompressed identifier or the suspected compression identifier, adjust the target prediction identifier in the direction of the uncompressed identifier.
25. The device according to any one of claims 21-24, characterized in that, The predictor is specifically configured to: If the predicted compression state is the uncompressed state and the actual compression state is the compressed state, adjust the target prediction identifier in the direction of the compressed identifier.
26. The device according to any one of claims 21-25, characterized in that, The predictor is specifically configured to: If the predicted compression state is the compressed state and the actual compression state is the uncompressed state, adjust the target prediction identifier in the direction of the uncompressed identifier.
27. The device according to any one of claims 21-26, characterized in that, The predictor is specifically configured to: If both the predicted compression state and the actual compression state are the compressed state, and the current target prediction identifier is the compressed identifier, keep the target prediction identifier unchanged; If both the predicted compression state and the actual compression state are the compressed state, and the current target prediction identifier is the suspected uncompressed identifier or the suspected compression identifier, adjust the target prediction identifier in the direction of the compressed identifier.
28. The device according to any one of claims 19-27, characterized in that, The predictor is further configured to: Create a state prediction table, and establish a mapping relationship between the identifiers of each memory area and the corresponding prediction identifiers in the state prediction table; Set the initial identifier of each prediction identifier to indicate the uncompressed state.
29. The device according to any one of claims 19-28, characterized in that, The controller is specifically configured to: If the predicted compression state is the compressed state, resolve the first physical address corresponding to the target memory cell according to the rules of compressed data, and read the data stored in the first physical address.
30. The device according to any one of claims 19-29, characterized in that The controller is specifically configured to: If the predicted compression state is the uncompressed state, resolve the second physical address corresponding to the target memory cell according to the rules of the uncompressed state, and read the data stored in the second physical address.
31. The device according to any one of claims 19-30, characterized in that, The controller is specifically configured to: After reading the target data, obtain the compression state identifier carried in the target data, where the compression state identifier is used to indicate whether the target data is compressed data; Determine the actual compression state of the target data based on the compression state identifier.
32. The device according to any one of claims 19-31, characterized in that, The controller is further configured to: When the predicted compression state is the compression state and the actual compression state is the non - compression state, the target data is reread based on the non - compression state.
33. The device according to any one of claims 19 - 32, characterized in that, The device is applied to a memory access system; the controller is specifically configured to: When the predicted compression state is the non - compression state and the actual compression state is the compression state, read the first compressed data and the second compressed data stored in the target memory unit. The memory access system includes a pre - fetch cache; the device further includes: a decompression module for decompressing the first data and the second data. The controller is further configured to read the decompressed first data according to the memory access request and store the decompressed second data into the pre - fetch cache.
34. The device according to any one of claims 19-33, characterized in that, The device further includes a compression module; the controller is further configured to: Receive a write request for the data to be written, and determine whether the data to be written needs to be compressed. The size of the data to be written is M bits. The compression module is configured to, if the data to be written needs to be fully compressed, compress the M bits of the data to be written; the controller is further configured to: write the compressed data into the corresponding memory unit. The compression module is further configured to, if a part of the data to be written needs to be compressed, compress the N bits of the data to be written; the controller is further configured to: write the compressed data into the first memory unit and write the remaining (M - N) bits of the data to be written into the second memory unit; both M and N are integers greater than 1, and N is less than M. The controller is further configured to: update the compression status flag bits in the first memory unit and the second memory unit to the compression flag and the non - compression flag respectively.
35. The device according to claim 34, characterized in that, The compression module is specifically configured to: Split the M bits of the data to be written into K - bit data and (M - K) - bit data respectively, and compress the K - bit data and the (M - K) - bit data respectively.
36. The device according to any one of claims 19-35, characterized in that The device further includes: A dynamic switch module for: Obtaining the current prediction accuracy rate of the state prediction table calculated based on the predicted compression state and the actual compression state of multiple memory access data. When the data to be written needs to be compressed and the prediction accuracy rate is lower than a preset threshold, control the compression module to turn off the compression function to write the data to be written into the corresponding memory unit in the non - compression state. When the data to be written needs to be compressed and the prediction accuracy rate is higher than a preset threshold, control the compression module to turn on the compression function to compress the data to be written and write it into the corresponding memory unit.
37. A computer-readable storage medium, characterized in that, The computer - readable medium is used to store program code, and when the program code is executed by an electronic device, the method described in any one of claims 1 - 18 above is implemented.
38. A computer program, characterized in that, The computer program includes instructions, and when the instructions are executed by an electronic device, the electronic device is caused to execute the method described in any one of claims 1 - 18.
Citation Information
Cited By
Cache access method, controller and cache prediction system
CN121144223A