Memory control device and application processor for controlled utilization and performance of an input / output device and method for actuating the memory control device
The memory controller dynamically adjusts address translation schemes based on real-time memory resource utilization to improve load and performance by aligning with user usage patterns.
Patent Information
- Application Number
- DE102019103114
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-02-12
- Filing Date
- 2019-02-08
- Publication Date
- 2025-10-30
- Estimated Expiration
- 2039-02-08
AI Technical Summary
Existing memory controllers fail to adaptively optimize address translation schemes to match varying user usage patterns, leading to suboptimal memory device load and performance.
A memory controller that dynamically evaluates and selects the most efficient address translation scheme in real-time based on memory resource utilization, using a processing circuit to translate system addresses into memory addresses and adjust the scheme according to usage patterns.
Enhances memory device load and performance by adaptively optimizing address translation, aligning with user patterns and maximizing memory resource utilization.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] Exemplary embodiments of inventive ideas relate to a memory control device and / or an application processor. For example, at least some embodiments relate to a memory control device for controlling the utilization and performance of an input / output (I / O) device, an application processor (AP) comprising a memory control device, and / or a method for actuating the memory control device.
[0002] A memory controller, or AP, can be used in an electronic system, such as a data processing system, and can exchange various signals with different peripheral devices. For example, the memory controller can control a volatile memory device, such as dynamic random access memory (DRAM), or a non-volatile memory device, such as flash memory or resistance-based memory, and can translate a system address received from a host processor requesting access to the memory device into a memory address suitable for that device. The memory controller can use an address translation scheme to translate the system address into the memory address.
[0003] US 2013 / 0132704A1 discloses a memory system that maps physical addresses to device addresses in a way that reduces power consumption. The system includes circuitry for deriving efficiency measures for memory utilization and selects between different address allocation schemes to improve efficiency. The address allocation schemes can be tailored to a specific memory configuration or a specific mix of active applications or application threads. Schemes tailored to a specific mix of applications or application threads can be applied each time the given mix is executed and updated for further optimization. Some embodiments simulate the presence of a disruptive thread to distribute memory addresses across the available banks, thereby reducing the likelihood of disruption by subsequently introduced threads. SUMMARY
[0004] Exemplary embodiments of the inventive ideas include a storage control device with which the utilization and performance of a storage device and other input / output devices can be increased, an application processor comprising the storage control device, and / or a method for actuating the storage control device.
[0005] According to some embodiments of the inventive ideas, a storage control device is specified for controlling a storage device.The memory control device includes a processing circuit designed to translate a first address received from a host processor into a second address associated with the storage device, based on an address translation scheme. The address translation scheme is selected from a plurality of address translation schemes based on memory resource utilization, and the memory resource utilization of each of the plurality of address translation schemes is evaluated based on a plurality of memory addresses generated using the plurality of address translation schemes, such that the processing circuit evaluates the memory resource utilization of each of the plurality of address translation schemes during the translation of the first address into the second address based on the first address translation scheme.
[0006] According to other embodiments of the inventive ideas, an application processor is specified which comprises: a host processor designed to provide an access request and a first address;and a memory control device designed to perform address translation in order to translate the first address into a second address associated with a memory address, based on a first address translation scheme selected from a plurality of address translation schemes, during a first period calculating a resource utilization score that specifies a memory resource utilization for each of the plurality of address translation schemes while performing the address translation based on the first address translation scheme, and during a second period translating the first address into the second address based on a second address translation scheme such that the memory resource utilization of the second address translation scheme is the largest among the plurality of address translation schemes during the first period, the second period following the first period.
[0007] According to yet another embodiment of the inventive ideas, a method for actuating a storage device is specified in order to translate a system address into a storage address associated with a storage device.In some embodiments, the method involves translating the system address to the memory address based on an address translation scheme selected from a plurality of possible translation scheme candidates; calculating resource utilization scores for the plurality of possible translation scheme candidates while translating the first address to the second address based on the first address translation scheme; selecting the next address translation scheme with the highest resource utilization from the plurality of possible translation scheme candidates based on the resource utilization scores; changing the address translation scheme to the next translation scheme; and translating the system address to the memory address based on the next translation scheme.
[0008] According to other embodiments, a data processing system is specified which comprises: at least one proprietary (IP) block; an input / output device designed to access the at least one IP block; and a control device designed to sequentially derive, by performing machine learning based on system addresses received from the at least one IP block, a plurality of address translation schemes and a resource utilization assessment for each of the plurality of address translation schemes, and, based on an address translation scheme selected from the plurality of address translation schemes based on the resource utilization assessment, to translate the system addresses received from the at least one IP block into a device address for the input / output device.
[0009] The invention is defined in the attached independent claims. Further developments of the invention are specified in the dependent claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Examples of the inventive ideas will become more understandable from the following detailed description in conjunction with the accompanying drawings: Fig. Figure 1 is a block diagram of a data processing system comprising a storage control device, according to an embodiment of the inventive ideas; Fig. Figure 2 is a diagram that illustrates an example of an address translation scheme; Fig. 3A and Fig. 3B are diagrams illustrating an example of a storage device of Fig. 1; Fig. Figure 4 is a block diagram of a storage control device according to an embodiment of the inventive ideas; Fig. 5 is a block diagram of a first counter block of Fig. 4; Fig. Figure 6 is a flowchart of a method for actuating a storage control device according to an embodiment of the inventive ideas; Fig. 7 is a flowchart of a procedure for calculating a resource utilization assessment in Fig. 6; Fig. Figure 8 is a diagram illustrating evaluation criteria for calculating a resource utilization assessment and a procedure for calculating values of the evaluation criteria in a specific example; Fig. 9A and Fig. 9B are diagrams of procedures for updating a resource utilization assessment according to exemplary implementations; Fig. 10A to Fig. 10E are diagrams illustrating a method for dynamically changing an address translation scheme reflecting a usage pattern, using a storage control device according to an embodiment of the inventive ideas; Fig. Figure 11 is a graph showing memory latency in relation to a memory page state; Fig. Figure 12 is a block diagram of a storage control device according to some embodiments of the inventive ideas; Fig. Figure 13 is a block diagram of a storage control device according to some embodiments of the inventive ideas; Fig. Figure 14 is a diagram of an example procedure for setting an address translation scheme based on machine learning; and Fig. Figure 15 is a block diagram of an application processor according to an embodiment of the inventive ideas. DETAILED DESCRIPTION
[0011] Fig. Figure 1 is a block diagram of a data processing system comprising a storage control device, according to an embodiment of the inventive ideas.
[0012] As in Fig. As shown in Figure 1, a data processing system 1000 can be incorporated into various types of electronic devices, such as a laptop computer, a smartphone, a tablet personal computer (PC), a drone, a personal digital assistant (PDA), an enterprise digital assistant (EDA), a digital camera, a portable multimedia playback device (PMP), a handheld game console, a mobile internet device, a multimedia device, a wearable computer, an Internet of Things (IoT) device, an Internet of Everything (IoE) device, an e-book, a smart home device, a medical device, and a device used for driving a vehicle.
[0013] The data processing system 1000 can include a storage control unit 100, a storage device 200, and a processor 300. The data processing system 1000 can also include various types of input / output (I / O) devices and intellectual property (IP) components (or IP blocks). In some embodiments, the storage control unit 100 and the processor 300 can be integrated on a single semiconductor chip. For example, the storage control unit 100 and the processor 300 can form an application processor (AP) implemented as a system-on-a-chip (SoC).
[0014] Processor 300 is an IP that requests access to storage device 200. Processor 300 can, for example, include a central processing unit (CPU), a graphics processing unit (GPU), and a display controller, and can be referred to as the master IP. Processor 300 can send an access request RQ (e.g., a write request or a read request for data DATA) and a system address SA to storage controller 100 via a system bus.
[0015] The storage device 200 can include volatile memory and / or non-volatile memory. If the storage device 200 includes volatile memory, it can include memory such as Double Data Rate (DDR) Synchronous Dynamic Random Access Memory (SDRAM), Low Energy DDR (LPDDR) SDRAM, Graphics DDR (GDDR) SDRAM, and Rambus DRAM (RDRAM). However, embodiments of the inventive ideas are not limited to these. For example, the storage device 200 can include non-volatile memory such as Flash memory, Magnetic RAM (MRAM), Ferroelectric RAM (FRAM), Phase Change RAM (PRAM), and Resistance-Based RAM (ReRAM).
[0016] The storage device 200 can have a plurality of memories M1 to Mm. Each of the memories M1 to Mm represents a physically and logically classified storage resource. Each of the memories M1 to Mm can have a storage rank, a storage bank, a storage row (or page), and a storage column (or column), which are referred to below as rank, bank, row, and column, or it can include logically classified regions.
[0017] The memory device 200 can be a semiconductor package containing at least one memory chip, or a memory module in which multiple memory chips are mounted on a module board. The memory device 200 can be embedded in a system-on-a-chip (SoC).
[0018] The memory control unit 100 is an interface that controls access to the memory device 200 according to the type of memory device 200 (e.g., flash memory or DRAM). The memory control unit 100 can control the memory device 200 so that data DATA is written to the memory device 200 in response to a write request received from the processor 300, or read from the memory device 200 in response to a read request received from the processor 300. Based on the access request RQ and the system address SA of the processor 300, the memory control unit 100 generates various signals, e.g., an instruction CMD and a memory address MA, to control the memory device 200, and provides these signals to the memory device 200.
[0019] The memory control unit 100 can translate the system address SA received from the processor 300 into the memory address MA corresponding to the memory device 200. System address SA refers to an address structure recognized by the processor 300, and memory address MA refers to an address structure, such as a rank, bank, row, or column, recognized by the memory device 200. The memory control unit 100 can translate the system address SA into the memory address MA based on a configured address translation scheme. An address translation scheme is defined with reference to Fig. 2 described.
[0020] Fig. Figure 2 is a diagram that illustrates an example of an address translation scheme.
[0021] As in Fig. As shown in Figure 2, the system address SA and the memory address MA can contain a plurality of bits. Here, MSB denotes a most significant bit and LSB denotes a least significant bit. As a non-restrictive example, the system address SA can be expressed as a hexadecimal code (e.g., 0X80000000), which has eight code values and can contain 32 bits. The 32 bits of the system address SA can be allocated to a rank signal RK, a bank signal BK, a row signal R, and a column signal C of the memory address MA. A rank, a bank, a row, and a column corresponding to the memory address MA can be selected in the memory device 200 according to the rank signal RK, the bank signal BK, the row signal R, and the column signal C of the memory address MA.
[0022] The address translation scheme can be defined as a method for mapping bits contained in the system address SA to the rank signal RK, the bank signal BK, the row signal R, and the column signal C. As in Fig. As shown in Figure 2, for example, some bits or a single bit from the system address SA can be assigned to a signal from the rank signal RK, bank signal BK, row signal R, and column signal C, or they can be subjected to an operation (e.g., an exclusive OR operation) with other bits, and the result of the operation can be assigned to a signal from the rank signal RK, bank signal BK, row signal R, and column signal C. The address translation scheme can be represented by the position of a bit that is assigned to a signal from the rank signal RK, bank signal BK, row signal R, and column signal C, and a hash function that uses the bit. The address translation scheme can be varied by selecting the bit position and the hash function.
[0023] It will be revisited Fig. 1 Reference is made where it is shown that the storage control device 100 can have a processing circuit and a memory.
[0024] The processing circuit can be, among other things, a processor, a central processing unit (CPU), a control unit, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a system-on-chip (SoC), a programmable logic unit, a microprocessor, or any device capable of performing operations in a defined manner.
[0025] The processing circuit can be configured as a purpose-built computer to implement the functions of an address translator 110 and an evaluation module 120, either through a layout design or the execution of computer-readable commands stored in a memory (not shown).
[0026] The address translator 110 can have a plurality of address translation modules, e.g., a first address translation module MD1, a second address translation module MD2, and a third address translation module MD3, which are implemented according to different address translation schemes (or address mapping schemes). The address translation schemes of the first to third address translation modules MD1 to MD3 can each be set (or alternatively, specified), and the first to third address translation modules MD1 to MD3 can be implemented in hardware. However, embodiments of the inventive ideas are not limited to these.In some embodiments, the first to third address translation modules MD1 to MD3 can be implemented by a processing circuit that executes software to transform the processor into a purpose-built processor, and address translation schemes can be designed during the operation of the data processing system 1000. Even if the address translator 110 is in . Fig. 1. While the first to third address translation module MD1 to MD3 is included, embodiments of the inventive ideas are not limited to these. The address translator 110 can include at least two address translation modules.
[0027] Each of the first three address translation modules, MD1 to MD3, can perform an address translation at system address SA and generate the memory address MA. However, the memory address MA output by an address translation module selected from the first three address translation modules, MD1 to MD3, can be sent to memory device 200. In other words, the first three address translation modules, MD1 to MD3, can be translation modules that qualify as address translation modules applied to memory device 200, and only one selected from the first three address translation modules, MD1 to MD3, can actually perform an address translation. For example, if the first address translation module, MD1, is selected, only the memory address MA output by the first address translation module, MD1, would be sent to memory device 200.In other words, address translation operations of the second and third address translation modules MD2 and MD3 can be background operations.
[0028] The evaluation module 120 can evaluate the memory resource utilization of each of the first three address translation modules (MD1 to MD3) in real time. Specifically, the evaluation module 120 can analyze the memory addresses (MA) output by each of the first three address translation modules (MD1 to MD3) and evaluate the memory resource utilization of each of these modules. For example, the evaluation module 120 can calculate a resource utilization score based on at least two values: a rank, a bank, and a row selected from the memory address (MA) output by the first address translation module (MD1). This score indicates the memory resource utilization of the first address translation module (MD1). In some embodiments, the evaluation module 120 can calculate the resource utilization score based on a single rank, a bank, and a row.The evaluation module 120 can calculate resource utilization scores for the second and third address translation modules, MD2 and MD3, respectively, in the same manner as described above. Therefore, the evaluation module 120 can evaluate the memory resource utilization of each of the first through third address translation modules, MD1 through MD3. A resource utilization score can be calculated and updated in real time. A procedure for calculating a resource utilization score is described below with reference to [reference to relevant section]. Fig. 8 to Fig. 10E described.
[0029] Based on the most recent results of an evaluation of memory resource utilization for each of the first three address translation modules (MD1 to MD3) during a system restart or idle period, the memory control unit 100 can select the address translation module with the highest memory resource utilization from among the first three address translation modules (MD1 to MD3) and can use the selected address translation module as the address translation module of the address translator 110. For example, the first address translation module (MD1) can perform an address translation while the memory resource utilization of the first three address translation modules (MD1 to MD3) is evaluated simultaneously.If the evaluation indicates that the memory resource utilization of the second address translation module (MD2) is highest up to a system restart, an address translation module used in address translator 110 can be switched from the first address translation module (MD1) to the second address translation module (MD2) during the system restart. Once the system restart is complete, the second address translation module (MD2) can perform the address translation. This process can be repeated.
[0030] Since the storage device 200 has a plurality of memory locations M1 to Mm, the page conflict rate (or page hit rate) and memory utilization of the storage device 200 can vary depending on the address translation scheme used in the memory control unit 100, i.e., depending on which of the memory locations M1 to Mm the system address SA is mapped to. Even if an address translation scheme that can maximize memory utilization under various conditions is selected through many performance tests during the manufacture of the memory control unit 100, the standardized address translation scheme may not be suitable for the usage patterns of different users. According to a user's usage pattern, i.e.,Depending on the type of applications available in the data processing system 1000, which are commonly executed by the user, the pattern of the system address SA received from the processor 300 can be changed, i.e., positions of bits that are frequently switched among all the bits contained in the system address SA can be changed.
[0031] Meanwhile, the result of the evaluation of memory resource utilization with respect to a plurality of address translation schemes in the memory control unit 100 can vary with the system address SA pattern. Accordingly, in one or more embodiments, the memory control unit 100 can adaptively operate an address translation scheme according to the system address SA pattern by evaluating the memory resource utilization in real time with respect to a plurality of address translation schemes and by dynamically changing an address translation scheme based on the evaluation result. As a result, the utilization and performance of the memory device 200 can be increased.
[0032] Fig. 3A and Fig. 3B are diagrams illustrating an example of a storage device of Fig. 1.
[0033] As in Fig. As shown in Figure 3A, a storage device 200a can have memory locations M1 to M16, the first through 16th. Memory locations M1 to M8 can form a first rank RK0, and memory locations M9 to M16 can form a second rank RK1. A command and address signal C / A can be sent from the storage controller 100 to the storage device 200a. The memory address MA contained in the command and address signal C / A can include a rank signal, a bank signal, a row signal, and a column signal, as described above. The storage controller 100 can exchange data containing a plurality of bits, e.g., data[63:0] containing 64 bits, with the storage device 200a via a data bus (or data channel). It is assumed that a write command is received from the storage controller 100.
[0034] Data[63:0] can be sent by eight bits to the first to eighth memory locations M1 to M8 and to the ninth to 16th memory locations M9 to M16. One of the first and second ranks RK0 and RK1 can be selected according to a rank signal of the memory address MA, and data[63:0] can be sent to the selected rank.
[0035] The first memory M1 from Fig. 3A is referred to Fig. 3B is described. The first memory M1 can contain banks BK0 to BK7, the first through eighth. Data [7:0] can be sent to a selected bank via a selector MUX. The selector MUX can be activated based on a bank at memory address MA. One of the first through eighth banks, BK0 to BK7, can be selected according to a bank signal at memory address MA, and data [7:0] can be sent to the selected bank.
[0036] Each of the first to eighth banks, BK0 to BK7, can have a memory cell array 21, a row decoder 22, a read / write circuit 23, and a column decoder 24. The memory cell array 21 can have a first to eighth row, R0 to R7, and a first to eighth column, C0 to C7. The row decoder 22 and the column decoder 24 can select a region (e.g., a page) in the memory cell array 21, according to a row and a column of the memory address MA, into which data [7:0] is to be written. The read / write circuit 23 can write data [7:0] into the selected region.
[0037] The time it takes to send and receive data between the memory controller 100 and the storage device 200, and the time it takes to write data to a rank, bank, and row selected by the memory address MA, can affect memory latency. If successively received memory addresses MA select different ranks or different banks, writing to the ranks and banks can be performed in parallel. Furthermore, if the successively received memory addresses MA select the same rank, bank, and row, the time can be reduced compared to when the successively received memory addresses MA select the same rank and bank but different rows.As stated above, if the memory addresses MA select different ranks or banks, or the same rank, bank and row, memory utilization can be increased and latency can be reduced.
[0038] Thus, the memory control unit 100 can increase memory utilization by real-time evaluation of memory resource utilization with reference to a plurality of address translation schemes based on at least two of a rank, a bank and a row selected from the memory address MA issued according to the address translation schemes, and by dynamically using an address translation scheme that has the highest memory resource utilization.
[0039] Fig. Figure 4 is a block diagram of a storage control device according to an embodiment of the inventive ideas.
[0040] As in Fig. As shown in Figure 4, a storage control unit 100a can include the address translator 110, the evaluation module 120 and a query queue 130.
[0041] A queue index QI can be set in request queue 130. Access requests originating from processor 300 (of Fig. 1) Access requests received from Processor 300 can be queued according to the order in which they are received. The type of access request QC (e.g., Q_1 and Q_2) and the system address SA (e.g., SA_1 and SA_2) corresponding to each access request received from Processor 300 can be queued. For example, the type of access request QC can be a write request or a read request. Queued requests can be processed sequentially or non-sequentially, depending on their position in the queue. As above, with reference to Fig. As described in section 1, the address translator 110 can have multiple address translation modules, for example, the first to third address translation modules MD1 to MD3. The first to third address translation modules MD1 to MD3 can each generate first to third memory addresses MA1, MA2, and MA3 based on one and the same system address SA, which differ from each other. The first to third memory addresses MA1 to MA3 can be provided at the evaluation module 120 and at a selector 115.
[0042] The selector 115 can select the memory address MA from the first to third memory addresses MA1 to MA3 according to a selection signal SEL and can assign the memory address MA to the storage device 200 (from Fig. 1) Output. The selection signal SEL can indicate which of the first to third address translation modules MD1 to MD3 is currently selected as the address translation module of address translator 110.
[0043] The evaluation module 120 can set the window size to the number of access requests received from the processor 300, i.e., a number N of access requests (where N is an integer of at least 2), and can evaluate the memory resource utilization of each of the first to third address translation modules MD1 to MD3 per N access request. In some embodiments, the evaluation module 120 can calculate an average of the ranks selected by each of the first to third address translation modules MD1 to MD3, the number of selected banks, and an average of the rows selected in the same rank and bank as the first evaluation value, second evaluation value, and third evaluation value, respectively, and can calculate a resource utilization score based on the first to third evaluation values.
[0044] For example, if the window size is set to 16, the evaluation module 120 can count the number of ranks, benches, and rows each time the number of access requests in the request queue 130 reaches 16, and can calculate and update a resource utilization assessment based on these numbers. This is explained in detail below with reference to Fig. 8 to Fig. 9B described.
[0045] The evaluation module 120 can have multiple counter blocks, e.g., a first to third counter block CB1, CB2, and CB3, and a logic circuit LC. The first to third counter blocks CB1 to CB3 can each receive the first to third memory addresses MA1 to MA3, which are output by the first to third address translation modules MD1 to MD3. The first counter block CB1 is used as an example with reference to Fig. 5 described.
[0046] Fig. 5 is a block diagram of a first counter block of Fig. 4.
[0047] As in Fig. As shown in Figure 5, the first counter block B1 can contain a rank counter 121, a bank counter 122, and a row counter 123. The rank counter 121, the bank counter 122, and the row counter 123 can count ranks, banks, and rows, respectively, selected from first memory addresses MA1, according to the set window size. For example, if the window size is set to 16, then the number of ranks, banks, and rows specified by 16 first memory addresses MA1 can be counted each time the number of items in the query queue reaches 16.
[0048] Rank counter 121 can count the number of ranks specified by rank signals RK from a plurality of first memory addresses MA1 and generate a rank counter value CV_RK. The rank counter value CV_RK can include a counter value for each of the ranks, e.g., a first rank and a second rank.
[0049] The bank counter 122 can count the number of banks specified by the rank signals RK and the bank signals BK of a plurality of first memory addresses MA1 and generate a bank counter value CV_BK. For example, if a memory device has two ranks and eight banks in each of the ranks, then the bank counter value CV_BK can include a counter value for each of the banks selected from a total of 16 banks, i.e., for the first through eighth banks of the first rank and the first through eighth banks of the second rank.
[0050] The row counter 123 can count the number of rows specified by the rank signals RK, the bank signals BK, and the row signals R of a plurality of first memory addresses MA1, and generate a row count value CV_R. For example, if there are eight rows in a bank within a rank, then the row count value CV_R can include a count value for each of the rows selected from 128 rows, i.e., for eight rows in each of the 16 banks.
[0051] The functionality of the second and third counter blocks CB2 and CB3 is similar to that of the first counter block CB1. Therefore, a repeated description is omitted.
[0052] It will be revisited Fig. Reference is made to Figure 4, which shows that the logic circuit LC can calculate a resource utilization score for each of the first to third address translation modules MD1 to MD3 based on counter values received from the first to third counter blocks CB1 to CB3. In some embodiments, the logic circuit LC can select an address translation module with the highest resource utilization score from the first to third address translation modules MD1 to MD3. In one embodiment, the processor 300 (or a CPU of a data processing system equipped with a memory controller) can compare resource utilization between the first to third address translation modules MD1 to MD3 and, based on the comparison result, select an address translation module with the highest resource utilization score.
[0053] Fig. Figure 6 is a flowchart of a method for actuating a storage control device according to an embodiment of the inventive ideas. In detail, it shows Fig. 6. A method for setting an address translation scheme using a memory control device. The one described in Fig. The 6 methods shown can be implemented in the storage control unit 100 of Fig. 1 and in the storage control unit 100a of Fig. 4 can be carried out. Accordingly, the measures referred to in Fig. 1 and Fig. The descriptions given in section 4 are applied to the procedure discussed herein.
[0054] As in Fig. As shown in Figure 6, the memory control unit 100, 100a can perform an address translation from a plurality of address translation schemes in a single operation or step S110, based on a preset address translation scheme. The plurality of address translation schemes can be selected (e.g., in advance) and stored in the memory control unit or on an external data carrier. In some embodiments, the plurality of address translation schemes can be implemented by hardware, software, or a combination thereof. A default value can be set so that one of the address translation schemes is used for address translation during the first system startup. Accordingly, the memory control unit 100, 100a can perform an address translation based on the preset address translation scheme during the first system startup.
[0055] In step S120, the memory control unit 100, 100a can calculate a resource utilization assessment for each of the address translation schemes. The resource utilization assessment for each address translation scheme can be calculated and updated for each configured window size. The memory control unit 100, 100a can perform steps S110 and S120 simultaneously. Furthermore, in some embodiments, the memory control unit 100, 100a can perform step S12 as a background operation.
[0056] The S120 work step is described in detail with reference to Fig. 7 described.
[0057] Fig. 7 is a flowchart of a procedure for calculating a resource utilization assessment in Fig. 6.
[0058] As in Fig. As shown in Figure 7, the storage control unit 100, 100a can count a rank signal, a bank signal and / or a series signal in step S121, which are output by each of a plurality of address translation modules using the address translation schemes.
[0059] In step S122, the storage control unit 100, 100a can calculate a first to third evaluation value based on count values. The first evaluation value can relate to a rank selection, the second evaluation value to a bank selection, and the third evaluation value to a row selection. For example, the first evaluation value can specify an average of selected ranks in the window size, the second evaluation value specifies the number of selected banks, and the third evaluation value specifies an average of selected rows in the same ranks and the same banks.
[0060] In step S123, the storage control unit 100, 100a can weight each of the first three evaluation values. When calculating a resource utilization assessment, the weightings of the first three evaluation values do not have to be equal. Therefore, an evaluation value with a higher weighting can be given a higher weighting.
[0061] In step S124, the storage control unit 100, 100a can calculate the resource utilization assessment based on weighted values. For example, the storage control unit can perform an arithmetic operation on the weighted values to calculate the resource utilization assessment. The procedure is described in detail with reference to Fig. 8 described.
[0062] It will be revisited Fig. Reference is made to Section 6, which shows that the memory controller 100, 100a can select an address translation scheme in step S130. For example, the memory controller 100, 100a can select the address translation scheme with the highest resource utilization rating from among the majority of address translation schemes. In some embodiments, apart from the memory controller 100, 100a, another processor in a data processing system equipped with the memory controller 100, 100a, e.g., a main processor, can select the address translation scheme with the highest resource utilization rating. Information about the selected address translation scheme can be stored and updated at a memory location, e.g., a register, within the memory controller 100, 100a or in non-volatile memory (or storage medium) outside the memory controller 100, 100a.For example, if the data processing system enters a period of inactivity, the information about the selected address translation scheme can be stored and updated in non-volatile memory or at the storage location within the memory control unit 100, 100a.
[0063] Afterwards, the storage control unit 100, 100a can perform an address translation based on the selected address translation scheme in step S140. In other words, the address translation scheme can be changed.
[0064] If the address translation scheme is changed during system runtime, data stored in storage device 200 may be moved to a location determined by the changed address translation scheme. This data relocation can take a significant amount of time, and therefore the address translation scheme may be changed if the estimated time required for the data processing system to enter and remain in idle mode is longer than the time required for the data relocation itself.
[0065] Alternatively, the address translation scheme can be changed to a selected address translation scheme upon system restart. Upon system restart, a bootloader or startup program applies the updated address translation scheme in non-volatile memory when accessing the system, and therefore the memory control unit 100, 100a can perform an address translation based on the changed address translation scheme, which differs from the address translation scheme used before the system restart.
[0066] Meanwhile, the storage control unit 100, 100a can repeatedly perform operations S120 to S140. Accordingly, the storage control unit 100, 100a can perform an address translation based on an address translation scheme selected based on a user's usage pattern.
[0067] Fig. Figure 8 is a diagram illustrating evaluation criteria for calculating a resource utilization assessment and a procedure for calculating values of the evaluation criteria in a specific example.
[0068] Fig. Figure 8 shows evaluation criteria for calculating a resource utilization score for the first address translation module MD1 based on the first memory address MA1 output by the first address translation module MD1. It is assumed that the window size is 16, the number of ranks in a memory device is 2, the number of banks in each rank is 8, and the number of rows in each bank is 8. It is also assumed that the first address translation module MD1 (more precisely, each first memory address output by the first address translation module MD1) is assigned queue indices QI from 1 to 16 in response to the individual system addresses received with access requests, as shown in Figure 8. Fig. 8 shown, selecting a rank, a bank and a row.
[0069] As in Fig. As shown in Figure 8, the evaluation criteria can include a first evaluation value EV1, corresponding to a rank selection of a memory address; a second evaluation value EV2, corresponding to a bank selection of the memory address; and a third evaluation value EV3, corresponding to a row selection of the memory address. As described above, the first evaluation value EV1 indicates an average of ranks selected (or used) by the first address translation module MD1; the second evaluation value EV2 indicates the number of banks selected by the first address translation module MD1; and the third evaluation value EV3 indicates an average of rows selected by the first address translation module MD1 in the same ranks and the same banks.
[0070] With reference to the first evaluation value EV1, the distribution of Rank 0 and Rank 1 can be analyzed, and it can be calculated how many ranks were selected on average within the window size. For example, if only Rank 0 or Rank 1 is selected, then it can be seen that on average only one rank is selected, and 1 can be a minimum value for the first evaluation value EV1. Conversely, if both Rank 0 and Rank 1 are selected eight times, then it can be seen that on average two ranks are selected, and 2 can be a maximum value for the first evaluation value EV1. If Rank 0 is selected 13 times, as in Fig. If 8 is shown, and Rank 1 is selected three times, or vice versa, it can be calculated that an average of 1.38 ranks are selected. The closer the first evaluation value EV1 is to the highest value 2, the more the memory utilization can increase.
[0071] The first evaluation value EV1 can be specifically defined as equation 1: EV1=2*(1−(Max(PR0,PR1)−0.5)), where PR0 is the ratio of rank 0 selections to the total number of selection events (e.g., 16), and PR1 is the ratio of rank 1 selections to the total number of selection events. Max(PR0,PR1) represents the highest value between PR0 and PR1. However, equation 1 is merely an example formula for calculating the first evaluation value EV1 and can be modified in various ways.
[0072] Regarding the second evaluation value, EV2, the number of banks selected within the window size can be calculated based on the analysis of banks selected in both rank 0 and rank 1. For example, since there are eight banks in each rank and the window size is 16, the highest possible value for EV2 can be 16, and the lowest possible value can be 1. The closer EV2 is to its highest value, the greater the potential increase in memory usage.
[0073] As in Fig. As shown in Figure 8, Bank 1, Bank 3 and Bank 7 are selected in Rank 0 and Bank 0 is selected in Rank 1, and therefore the second evaluation value EV2 can be 4.
[0074] Regarding the third evaluation value, EV3, an average can be calculated for rows contained in the same ranks and banks. The smallest possible value for EV3 is 1 (if a row is selected 16 times in the same rank and bank). Since the window size is 16, the largest possible value for EV3 is 16. The closer EV3 is to its smallest value, the greater the potential increase in memory usage.
[0075] As in Fig. As shown in 8, 2 is the number of rows selected in rank 0 and bank 1, 5 is the number of rows selected in rank 0 and bank 3, 6 is the number of rows selected in rank 0 and bank 7, and 3 is the number of rows selected in rank 1 and bank 0. Accordingly, the third evaluation value EV3 can be 4.
[0076] The first evaluation value EV1, the second evaluation value EV2, and the third evaluation value EV3 can be calculated as described above. However, with regard to increasing storage utilization, the first evaluation value EV1 may have the highest significance, and the third evaluation value EV3 may have the lowest. Accordingly, when calculating a resource utilization assessment based on the first three evaluation values EV1 to EV3, each of these values can be weighted differently. The first evaluation value EV1 can be weighted most heavily, and the third evaluation value EV3 can be weighted least heavily.
[0077] As described above, the closer the first and second evaluation values (EV1 and EV2) are to the highest values, and the closer the third evaluation value (EV3) is to the lowest value, the greater the increase in memory utilization. Proximity to the respective highest value can be calculated for both the first and second evaluation values (EV1 and EV2), and proximity to the lowest value can be calculated for the third evaluation value (EV3). A weight, which is set differently for each evaluation criterion, can be multiplied by the calculated proximity, and the results of these multiplications can be added to calculate the resource utilization score.
[0078] For example, a proximity PXY1 to a highest value Max_EV1 of the first evaluation value EV1 can be defined as in equation 2: PXY1=1(Max_EV1−EV1) / (Max_EV1−Min_EV1). where Min_EV1 is the smallest value of the first evaluation value EV1.
[0079] A proximity PXY2 to a highest value Max_EV2 of the second evaluation value EV2 is defined similarly to equation 2 and can therefore be derived from equation 2.
[0080] A proximity of PXY3 to a smallest value Min_Ev3 of the third evaluation value EV3 can be defined as in equation 3: PXY3=(Max_EV3−EV3) / (Max_EV3−Min_EV3).
[0081] For example, in Fig. 8. The proximity PXY1 of the first evaluation value EV1 can be calculated as 0.38, the proximity PXY2 of the second evaluation value EV2 can be calculated as 0.2, and the proximity PXY3 of the third evaluation value EV3 can be calculated as 0.8.
[0082] If a first weight for the first evaluation value EV1 is set to 0.50, a second weight for the second evaluation value EV2 is set to 0.29, and a third weight for the third evaluation value EV3 is set to 0.21, a resource utilization rating RUV of the first address translation module MD1 can be calculated as (0.38*0.5+0.2*0.29+0.8*0.21)* 100=41.6.
[0083] The procedure for calculating the first to third evaluation values EV1 to EV3 and the procedure for calculating the resource utilization assessment RUV based on the first to third evaluation values EV1 to EV3 were developed with reference to Fig. 8 is described, but embodiments of the inventive ideas are not limited therein. The method for calculating the first to third evaluation values EV1 to EV3, the method for calculating the resource utilization rating RUV, and specific figures can be modified in various ways. For example, the storage resource utilization is calculated based on at least two evaluation values.
[0084] Fig. 9A and Fig. 9B are diagrams of procedures for updating a resource utilization assessment according to exemplary implementations.
[0085] As in Fig. 9A and Fig. As shown in Figure 9B, the first address translation module, MD1, can perform an address translation during a period T1 starting at time t1. A system restart or idle period can begin at time t2. As described above, a resource utilization score can be calculated for each configured window size. If the window size is set to 16, a resource utilization score can be calculated every time 16 access requests (e.g., Q1_1 to Q16_1, Q1_2 to Q16_2, Q1_3 to Q16_3, or Q1_4 to Q16_4) are received. For example, four resource utilization scores can be calculated: a first resource utilization score RUV1, a second resource utilization score RUV2, a third resource utilization score RUV3, and a fourth resource utilization score RUV4.
[0086] As in Fig. As shown in Figure 9A, the final resource utilization score (RUV) considered when an address translation scheme is changed can be an average of the first four resource utilization scores (RUV1 to RUV4). In other words, each time a resource utilization score is generated, an average of a new resource utilization score and an old resource utilization score can be stored as the resource utilization score. Thus, a user's usage pattern can be accumulated and reflected in the resource utilization score.
[0087] As in Fig. As shown in Figure 9B, the storage control unit 100, 100a can, in some other embodiments, use a moving average to determine the resource utilization rating. For example, an average of a desired (or alternatively, a predetermined) number of recent resource utilization ratings from a plurality of resource utilization ratings calculated during the period T1 can be stored as the latest resource utilization rating RUV. An average of at least two resource utilization ratings is stored in Fig. 9B is stored as the last resource utilization rating (RUV), but embodiments of the inventive ideas are not limited to this. The number of resource utilization ratings reflected in the last resource utilization rating (RUV) can be changed. Thus, a user's usage pattern within a limited period can be reflected in a resource utilization rating.
[0088] Fig. 10A to Fig. 10E are diagrams illustrating a method for dynamically changing an address translation scheme reflecting a usage pattern, using a memory control device according to an embodiment of the inventive ideas.
[0089] As in Fig. As shown in Figure 10A, the positions of bits assigned to ranks and a hash function differ between the first three address translation modules, MD1 to MD3. In other words, different address translation schemes are applied to each of the first three address translation modules, MD1 to MD3. For example, a rank determination scheme, Rank intlv, which specifies the position of the bit within the system address SA that is assigned to the bank, and a bank determination scheme, Bank intlv, which specifies a hash function, can be set for each of the first three address translation modules, MD1 to MD3. Here, BW denotes a bit width of the system address SA.
[0090] The address translation schemes for the first to third address translation modules MD1 to MD3 can be selected and stored on a non-volatile data carrier. For example, the 1000 data processing system can be used by Fig. 1. They may also have non-volatile memory (or storage medium), and the address translation schemes for the first to third address translation modules MD1 to MD3 may be stored in non-volatile memory before a system startup.
[0091] As in Fig. As shown in Figure 10B, the first address translation module, MD1, can be selected by default upon initial system startup, and MD1 can then perform an address translation. The resource utilization rating (RUV) of each of the first three address translation modules, MD1 to MD3, can be calculated in real time based on the configured window size. The resource utilization rating (RUV) of the second address translation module, MD2, is 67.13, making it the highest. Therefore, a symbol indicating that MD2 has the highest resource utilization rating (RUV) can be updated and stored in non-volatile memory during a period of inactivity. Afterward, MD2 can perform an address translation when the system is restarted.
[0092] As in Fig. As shown in Figure 10C, the second address translation module, MD2, can perform an address translation when the system is restarted a second time. The resource utilization rating (RUV) of each of the first three address translation modules, MD1 to MD3, can be calculated in real time based on the window size. The resource utilization rating (RUV) of the third address translation module, MD3, is 64.99 and is therefore the highest. Thus, a symbol indicating that the resource utilization rating (RUV) of the third address translation module, MD3, is the highest can be updated and stored in non-volatile memory during an idle period. Afterward, the third address translation module, MD3, can perform an address translation when the system is restarted.
[0093] As in Fig. As shown in Figure 10D, the third address translation module, MD3, can perform an address translation when the system is started for the third time. During this process, the resource utilization rating (RUV) can be calculated for each of the address translation modules from MD1 to MD3. The resource utilization rating (RUV) of the third address translation module, MD3, is 81 and is therefore the highest. Thus, a symbol indicating that the resource utilization rating (RUV) of the third address translation module, MD3, is the highest can be updated during an idle period and stored in non-volatile memory. As shown in Figure 10D. Fig. As shown in Figure 10E, the third address translation module MD3 can thus perform address translation continuously when the system is started a fourth time.
[0094] Fig. Figure 11 is a graph showing memory latency in relation to a memory page state.
[0095] As in Fig. As shown in Figure 11, memory latency can be improved by increasing the page hit rate or decreasing the page conflict rate. As shown in Fig. As shown in Figure 11, the latency difference between a case where every memory traffic is a page hit and a case where every memory traffic is a page conflict is at most 37.5 nanoseconds (ns). As described above, if an address translation scheme is improved using a memory control device and a method for actuating the memory control device, then, according to one embodiment of the inventive ideas, the page hit rate can be increased (or alternatively maximized) taking into account a user's usage pattern, thus improving the memory latency.
[0096] Fig. Figure 12 is a block diagram of a storage control device according to some embodiments of the inventive ideas.
[0097] As in Fig. As shown in Figure 12, a memory control unit 100b can include an address translator 110b, a read-only memory (ROM) 150, the evaluation module 120, and the query queue 130. The operations of the evaluation module 120 and the query queue 130 are the same as those described with reference to Fig. 4 were described. Therefore, redundant descriptions are avoided, and descriptions focus on the differences between the storage control unit 100a and the storage control unit 100b.
[0098] A plurality of address translation modules, i.e., the first to third address translation modules MD1 to MD3, can be implemented by a processing circuit that executes software or firmware stored in ROM 150 (or a non-volatile memory region). The address translator 110b can be implemented by dedicated hardware for executing the first to third address translation modules MD1 to MD3.
[0099] When a system starts, the first three address translation modules, MD1 to MD3, can be executed by the address translator 110b, thus enabling address translation. As described above, the address translator 110b can perform address translation based on a preset address translation module, such as the first address translation module, MD1, which is selected by default, and can also execute the second and third address translation modules, MD2 and MD3, as background operations. The first memory address, MA1, generated by executing the first address translation module, MD1, can be output to a storage device as memory address MA. The second memory address, MA2, can be generated while the second address translation module, MD2, is executing. The third memory address, MA3, can be generated while the third address translation module, MD3, is executing.In some embodiments, the first to third memory addresses MA1 to MA3 can be generated sequentially. Each of the first to third memory addresses MA1 to MA3 can be provided at a corresponding counter block CB1 to CB3, representing the first to third counter block.
[0100] The evaluation module 120 can calculate a resource utilization score for each of the first three address translation modules, MD1 to MD3. Information about an address translation module with the highest resource utilization score can be stored and updated in non-volatile memory or in a register within the address translator 110b.
[0101] Fig. Figure 13 is a block diagram of a storage control device according to some embodiments of the inventive ideas. Fig. Figure 14 is a diagram of an example procedure for setting an address translation scheme based on machine learning.
[0102] As in Fig. As shown in Figure 13, a memory controller 100c can include an address translator 110c, a machine learning logic 160, an evaluation module 120c, and a query queue 130. The evaluation module 120c can include a counter block CB and a logic circuit LC. The address translator 110c can be implemented by dedicated hardware and can perform address translation based on a predefined address translation scheme. Although not shown, an address translation module that uses the predefined address translation scheme can be implemented by a processing circuit that executes software stored in ROM (not shown) or can be implemented by hardware.
[0103] The system address SA can be provided to the machine learning logic 160, as can the address translator 110c, which performs the address translation. The machine learning logic 160 can perform machine learning and analyze a pattern change of the system address SA. Furthermore, the machine learning logic 160 can derive a plurality of other address translation schemes sequentially. For example, the machine learning logic 160 can analyze a change pattern of the individual bits contained in a large amount of the system address SA (e.g., a pattern in which each bit undergoes data toggling) and can derive a plurality of address translation schemes from the analysis result. Even if the machine learning logic 160 is in Fig. While the storage control unit 100c contains component 13, embodiments of the inventive ideas are not limited to this. The machine learning logic 160 can be implemented as a component separate from the storage control unit 100c and can be integrated into a data processing system (e.g., the data processing system 1000 of [unclear text]). Fig. 1) be included.
[0104] As in Fig. As shown in Figure 14, the address translator 110c can perform address translation based on a predefined address translation scheme DTS. The evaluation module 120c can evaluate memory resource utilization with respect to the predefined address translation scheme DTS. For example, the evaluation module 120c can calculate a resource utilization score as described above.
[0105] Meanwhile, the machine learning logic 160 can analyze a pattern of the system address SA and derive an address translation scheme, e.g., a first address translation scheme TS1, based on the analysis result in a background operation. While the address translator 110c generates the memory address MA by performing an address translation based on the preset address translation scheme DTS, it can also generate a different memory address based on the first address translation scheme TS1 in a background operation. The evaluation module 120c can calculate a resource utilization assessment with respect to the first address translation scheme TS1 based on the memory address that was generated according to the first address translation scheme TS1.
[0106] The machine learning logic 160 can derive multiple address translation schemes sequentially based on accumulated data, such as the system address SA, used for training. A resource utilization score can be calculated for each of the address translation schemes. Furthermore, the machine learning logic 160 can derive an address translation scheme not only based on the system address SA, but also on the calculated resource utilization scores. In other words, the machine learning logic 160 can derive a new address translation scheme that can improve memory utilization based on input data, such as the system address SA, and output data, such as resource utilization scores calculated with respect to the system address SA.
[0107] If the resource utilization score of a newly derived address translation scheme is higher than that of a previously derived address translation scheme or the default address translation scheme DTS, information about the newly derived address translation scheme and its resource utilization score can be stored. This information can be stored in a non-volatile storage device. For example, if the resource utilization score of a second address translation scheme TS2, among a plurality of address translation schemes derived from machine learning logic 160, is higher than that of the default address translation scheme DTS and the first address translation scheme TS1, information about the second address translation scheme TS2 and its resource utilization score can be stored.If the resource utilization rating of a fourth address translation scheme TS4, derived by machine learning logic 160, is higher than those of the default address translation scheme DTS and the previously derived address translation schemes, e.g., the first to third address translation schemes TS1 to TS3, then information about the fourth address translation scheme TS4 and its resource utilization rating can be stored. If the system is subsequently restarted at time t2 or enters a downtime period, an address translation scheme can be changed. Of the resource utilization ratings with respect to the first to fifth address translation schemes TS1 to TS5, i.e., memory resource utilization evaluated during time t1, the resource utilization rating of the fourth address translation scheme TS4 is the highest. Thus, after time t2, i.e.,h. During a period T2, the address translator 110c performs an address translation based on the fourth address translation scheme TS4. During the same period T2, the machine learning logic 160 can also derive a new address translation scheme through training.
[0108] Even though the calculation of the resource utilization assessment is performed by the evaluation module 120c in some of the embodiments, the embodiments of the inventive ideas are not limited to this, and the machine learning logic 160 can calculate the resource utilization assessments. For example, the machine learning logic 160 can include the evaluation module 120c.
[0109] Fig. 15 is a block diagram of an AP according to an embodiment of the inventive ideas.
[0110] As in Fig. As shown in Figure 15, an AP 2000 can have at least one IP. For example, the AP 2000 can have a processor 210, a ROM 220, a neural processing unit (NPU) 230, an embedded memory 240, a memory control device 250, a memory interface (I / F) 260, and a display I / F 270, all interconnected via a system bus 280. The AP 2000 can also have other elements, such as I / O devices, and some elements that are shown in Figure 15 can be included in the system bus 280. Fig. Items shown in 15 do not necessarily have to be included in AP 2000.
[0111] The advanced microcontroller bus architecture (AMBA) protocol Advanced RISC Machine (ARM) can be used as a standard specification for the AP 2000 system bus 280. Bus types of the AMBA protocol can include Advanced High-Performance Bus (AHB), Advanced Peripheral Bus (APB), Advanced Extensible Interface (AXI), AXI4, and AXI Coherency Extensions (ACE). In addition, other types of protocols, such as uNetwork, CoreConnect, or Open Core Protocol, optimized for a specific system, can be used.
[0112] The processor 210 can control all operations of the AP 2000. For example, software for management operations for various IPs in the AP 2000 can be loaded into a storage device 251, which is provided outside the AP 2000, and / or into the embedded memory 240, and the processor 210 can perform various management operations by executing the loaded software.
[0113] A bootloader can be stored in ROM 220. Additionally, various settings related to the AP 2000 can be stored in ROM 220.
[0114] The NPU 230 is an operational unit that performs a deep learning operation (or neural network operation). The NPU 230 can perform a deep learning operation when a deep learning-based application is running on the AP 2000, thus ensuring performance. For example, if the memory controller 250 derives an address translation scheme through machine learning, as described above, then the NPU 230 can perform an operation.
[0115] The storage I / F 260 can provide a connection between the AP 2000 and a storage medium 261 located outside the AP 2000. The storage medium 261 can contain non-volatile memory. The storage medium 261 can include cells such as NAND or NOR flash memory cells, which retain data when the power is interrupted, or various types of non-volatile memory, such as MRAM, ReRAM, FRAM, or phase-change memory (PCM).
[0116] The display I / F 270 can provide image data or updated image data to a display module 271 under the control of the processor 210 or a GPU (not shown).
[0117] The memory device 251 can be implemented using volatile memory. In some embodiments, the memory device 251 can include volatile memory, such as DRAM and / or static RAM (SRAM).
[0118] As described above in various embodiments, the memory control unit 250 can control the memory device 251 and translate an address, e.g., a system address, received along with an access request from another IP (e.g., the processor 210), into a memory address corresponding to the memory device 251. The memory control unit 250 can evaluate and assess each memory address that has undergone address translation according to a plurality of address translation schemes in real time with respect to memory resource utilization and continuously update these assessments during system operation. When a system enters a sleep period or is restarted, the memory control unit 250 can change an address translation scheme (e.g.,into an address translation scheme that has the highest rating), so that an address translation scheme suitable for a user pattern can be used.
[0119] The memory control unit 250 can incorporate machine learning logic and derive an address translation scheme that increases (or alternatively maximizes) memory utilization through machine learning based on the results of evaluating memory utilization with respect to input system addresses and address translation schemes. In some alternative embodiments, the machine learning logic can be implemented separately from the memory control unit 250.
[0120] To access the storage medium 261 and the display module 271, the memory interface 260 and / or the display interface 270 can also translate a received address into an address structure suitable for a device to be accessed. As described above with reference to the memory control unit 250, the memory interface 260 and / or the display interface 270 can evaluate and assess a plurality of pre-selected address translation schemes (or device mapping schemes) in real time with respect to a device utilization and can change an address translation scheme based on these assessments. As a result, an address translation scheme can be used that is improved (or alternatively optimized) for an individual user's usage pattern. A method for selecting an address translation scheme such as one described above can also be applied to other I / O devices.
[0121] According to exemplary embodiments of the inventive ideas, a plurality of device addresses generated based on a plurality of address translation schemes can be evaluated and assessed in real time, the resource utilization of each address translation scheme can be calculated, and address translation can be performed using an address translation scheme that exhibits a desired resource utilization (or alternatively, a maximum resource utilization), so that an address translation scheme can be used that is improved (or alternatively, optimized) for a user's usage pattern. As a result, the utilization and performance of a storage device and other I / O devices can be increased.
[0122] According to one or more embodiments, the units and / or devices described above, including elements of the storage control unit 100, such as the address translator 110, the evaluation module 120 and sub-elements thereof, such as the address translation modules MD1-MD3, the counter blocks CB1-CB3 and the logic circuit LC and the selector 115, can be implemented using hardware, a combination of hardware and software or a non-volatile storage medium that stores software executable to perform the functions thereof.
[0123] Hardware can be implemented using a processing circuit, such as one or more processors, one or more central processing units (CPUs), one or more control devices, one or more arithmetic logic units (ALUs), one or more digital signal processors (DSPs), one or more microcomputers, one or more field programmable gate arrays (FPGAs), one or more system-on-chips (SoCs), one or more programmable logic units (PLUs), one or more microprocessors, one or more application-specific integrated circuits (ASICs), or any other device or devices capable of responding to and executing instructions in a defined manner.
[0124] Software can consist of a computer program, program code, commands, or any combination thereof, used to instruct or configure a hardware device, independently or collectively, to operate as desired. The computer program and / or program code may include a program or computer-readable commands, software components, software modules, data files, data structures, etc., which can be implemented by one or more hardware devices, such as one or more of the hardware devices mentioned above. Examples of program code include both machine code, produced by a compiler, and higher-level program code, executed by an interpreter.
[0125] If a hardware device is, for example, a computer processing device (e.g., one or more processors, CPUs, control units, ALUs, DSPs, microcomputers, microprocessors, etc.), then the computer processing device can be designed to execute program code by performing arithmetic, logical, and input / output operations according to the program code. Once the program code has been loaded into a computer processing device, the computer processing device can be programmed to execute the program code, thereby transforming the computer processing device into a computer processing device for a specific purpose. In a more concrete example, once the program code has been loaded into the processor, it is programmed to execute the program code and the corresponding operations, thereby transforming the processor into a processor for a specific purpose.In another example, the hardware device could be an integrated circuit that has been adapted to be a processing circuit for a specific purpose (e.g., an ASIC).
[0126] A hardware device, such as a computer processing device, can run an operating system (OS) and one or more software applications that run on the OS. The computer processing device can also access, store, edit, process, and generate data in response to the execution of the software. For the sake of simplicity, one or more embodiments may be explained using the same computer processing device; however, those skilled in the art will recognize that a hardware device can have multiple processing elements and multiple types of processing elements. For example, a hardware device can have multiple processors or one processor and a control unit. Furthermore, other processing configurations are possible, such as parallel processors.
[0127] Software and / or data can be permanently or temporarily embodied in any type of storage medium, including, but not limited to, any machine, component, physical or virtual equipment, or computer storage medium or device that can supply or interpret commands or data to a hardware device. The software can also be distributed across networked computer systems, so that the software is stored and executed in a distributed manner. In particular, for example, software and data can be stored on one or more computer-readable recording media, including physical or non-volatile computer-readable storage media, as discussed herein.
[0128] Storage media may also include one or more storage devices on units and / or devices according to one or more embodiments. The one or more storage devices may be physical or non-volatile computer-readable storage media, such as read / write memory (RAM), read-only memory (ROM), a permanent mass storage device (such as a hard disk drive), and / or any other data storage mechanism capable of storing and recording data. The one or more storage devices may be designed to store computer programs, program code, instructions, or any combination thereof for one or more operating systems and / or for implementing the embodiments described herein. The computer programs, program code, instructions, or any combination thereof may be...The computer programs, program code, instructions, or any combination thereof can also be loaded into one or more computer processing units from a separate computer-readable storage medium using a control mechanism. Such a separate computer-readable storage medium can include a Universal Serial Bus (USB) flash drive, a memory stick, a Blu-ray / DVD / CD-ROM drive, a memory card, and / or similar computer-readable storage media. Instead of loading from a computer-readable storage medium, the computer programs, program code, instructions, or any combination thereof can be loaded into one or more storage devices and / or computer processing units from a remote data storage device via a network interface.Furthermore, the computer programs, program code, instructions, or any combination thereof can be loaded by a remote computing system designed to transmit and / or distribute the computer programs, program code, instructions, or any combination thereof, over a network, into one or more storage devices and / or one or more processors. The remote computing system can transmit and / or distribute the computer programs, program code, instructions, or any combination thereof via a wired interface, an air interface, and / or any other similar medium.
[0129] The one or more hardware devices, storage media, computer programs, program code, instructions, or any combination thereof may be specifically designed and constructed for the purposes of the exemplary embodiments, or may be known devices that are altered and / or modified for the purposes of the exemplary embodiments.
Claims
[1] Storage control device designed to control a storage device, the storage control device comprising (100; 100a; 100b; 100c; 250): a processing circuit designed to to translate a first address (SA) received from a host processor (300; 210) into a second address (MA) associated with the storage device (200; 200a; 251) based on an address translation scheme, wherein the address translation scheme is selected from a plurality of address translation schemes (TS1 to TS5) based on memory resource utilization; and to evaluate the memory resource utilization of each of the plurality of address translation schemes (TS1 to TS5) based on a plurality of memory addresses (MA1, MA2, MA3) generated using the plurality of address translation schemes (TS1 to TS5) in such a way that the processing circuit evaluates the memory resource utilization of each of the plurality of address translation schemes (TS1 to TS5) during the translation of the first address (SA) to the second address (MA) based on the first address translation scheme (TS1). [2] Memory control device according to claim 1, wherein the processing circuit is designed to calculate the memory resource utilization on the basis of at least two of a first evaluation value (EV1), a second evaluation value (EV2) and a third evaluation value (EV3), and wherein the first evaluation value (EV1), the second evaluation value (EV2) and the third evaluation value (EV3) are calculated on the basis of a rank selection, a bank selection or a row selection of the plurality of memory addresses (MA1, MA2, MA3) in each of the plurality of address translation schemes (TS1 to TS5). [3] Storage control device according to claim 2, wherein the processing circuit is designed to calculate the storage resource utilization (RUV) on the basis of a sum of a first value and a second value, wherein the first value is obtained by a first weighting of the first evaluation value (EV1) and the second value is obtained by a second weighting of the second evaluation value (EV2), wherein the first weighting is higher than the second weighting. [4] Storage control device according to one of claims 1 to 3, wherein the processing circuit is designed to to perform an address translation based on an initial address translation scheme, to evaluate the memory resource utilization during an initial period (T1) associated with the first address translation scheme, and to perform the address translation based on a second address translation scheme during a second period (T2), wherein the second address translation scheme is selected based on the memory resource utilization (RUV) calculated during the first period (T1). [5] Storage control device according to claim 4, wherein the processing circuit is designed to change the address translation scheme from the first address translation scheme to the second address translation scheme in response to a restart of the storage control device (100; 100a; 100b; 100c; 250) after the first period (T1). [6] Storage control device according to claim 4, wherein the processing circuit is designed to change the address translation scheme from the first address translation scheme to the second address translation scheme in response to the fact that a rest period occurs after the first period (T1). [7] Storage control device according to any one of claims 1 to 6, wherein the processing circuit is designed to to perform address translation using different address translation schemes (TS1 to TS5) to generate the majority of memory addresses (MA1, MA2, MA3), by counting ranks, banks and rows associated with the plurality of memory addresses (MA1, MA2, MA3) to generate counter values (CV_R, CV_BK, CV_RK), where the ranks, banks and rows are selected by the plurality of memory addresses (MA1, MA2, MA3), and Based on the counter values (CV_R, CV_BK, CV_RK), a utilization rating (RUV) is calculated, which indicates the storage resource utilization. [8] Memory control device according to claim 7, wherein the processing circuit is designed to count the ranks, banks and rows with respect to N memory addresses (MA1, MA2, MA3), where N is an integer greater than or equal to 2. [9] Memory control device according to any one of claims 1 to 8, wherein the processing circuit is designed to perform machine learning based on a plurality of first addresses (SA, SA_1, SA_2) in order to derive the plurality of address translation schemes (TS1 to TS5) sequentially, wherein the plurality of first addresses (SA, SA_1, SA_2) are received from the host processor (300; 210). [10] Storage control device according to claim 9, wherein the processing circuit is designed to derive the plurality of address translation schemes (TS1 to TS5) based on a change pattern of bits contained in the plurality of first addresses (SA, SA_1, SA_2). [11] Application processor comprising: a host processor (300; 210) designed to provide an access request (RQ; Q1_1 to Q16_1) and a first address (SA); and a storage control unit (100; 100a; 100b; 100c; 250) designed to to perform an address translation in order to translate the first address into a second address associated with a memory address (MA) based on a first address translation scheme selected from a plurality of address translation schemes (TS1 to TS5), During a first period (T1), a resource utilization score (RUS) is calculated that indicates the memory resource utilization of each of the plurality of address translation schemes (TS1 to TS5) while performing address translation based on the first address translation scheme, and during a second period (T2), address translation is performed to translate the first address to the second address based on a second address translation scheme, such that the memory resource utilization of the second address translation scheme is highest among the plurality of address translation schemes (TS1 to TS5) during the first period (T1), with the second period (T2) following the first period (T1). [12] Application processor according to claim 11, wherein the memory control device (100; 100a; 100b; 100c; 250) is designed to to recalculate a resource utilization score (RUS) for each of the majority of address translation schemes (TS1 to TS5) during the second period (T2), and to perform the address translation during a third period after the second period (T2) based on a third address translation scheme, because the memory resource utilization of the third address translation scheme was the highest of the majority of address translation schemes (TS1 to TS5) during the second period (T2), with the third period coming after the second period (T2). [13] Application processor according to claim 11 or 12, wherein the memory control device (100; 100a; 100b; 100c; 250) is designed to calculate the resource utilization evaluation (RUV) for the plurality of address translation schemes (TS1 to TS5) based on a plurality of memory addresses (MA1, MA2, MA3) generated according to a selected address translation scheme (TS1 to TS5). [14] Application processor according to any one of claims 11 to 13, wherein the memory control device (100; 100a; 100b; 100c; 250) is designed to calculate the resource utilization evaluation (RUV) based on a first evaluation value (EV1) calculated according to a rank selection, a second evaluation value (EV2) calculated according to a bank selection, and a third evaluation value (EV3) calculated according to a row selection. [15] Application processor according to any one of claims 11 to 14, wherein the memory control device (100; 100a; 100b; 100c; 250) is designed to to perform address translation using various address translation schemes (TS1 to TS5) to generate a plurality of memory addresses (MA1, MA2, MA3), and to calculate the Resource Utilization Evaluation (RUE) for each of the majority of address translation schemes (TS1 to TS5) based on the majority of memory addresses. [16] Application processor according to claim 15, wherein the memory control device is designed to to generate counter values (CV_R, CV_BK, CV_RK) by counting ranks, banks, and rows selected based on the plurality of memory addresses (MA1, MA2, MA3); and to calculate the resource utilization assessment (RUV) based on the counter values (CV_R, CV_BK, CV_RK). [17] Application processor according to any one of claims 11 to 16, wherein the memory control device (100; 100a; 100b; 100c; 250) is designed to to compare resource utilization assessments (RUVs) calculated in relation to the majority of address translation schemes (TS1 to TS5), and to select the second address translation scheme based on resource utilization assessments, so that the memory resource utilization of the second address translation scheme is the highest among the majority of address translation schemes (TS1 to TS5). [18] Application processor according to any one of claims 11 to 17, wherein information indicating that the memory resource utilization (RUV) of the second address translation scheme is highest among the plurality of address translation schemes is stored in a memory device, and the memory control device (100; 100a; 100b; 100c; 250) is designed to switch from the first address translation scheme to the second address translation scheme based on the information loaded from the memory device after the first period (T1) upon restart. [19] Method for actuating a storage control device for translating a system address (SA) into a storage address (MA) associated with a storage device (200; 200a; 251), the method comprising: Translating the system address (SA) to the memory address (MA) based on an address translation scheme selected from a plurality of possible translation schemes (TS1 to TS5); Calculating resource utilization assessments for the majority of eligible translation schemes (TS1 to TS5) during the translation of the first address (SA) to the second address (MA) based on the first address translation scheme (TS1); Select the next address translation scheme with the highest resource utilization rating from the majority of eligible translation schemes (TS1 to TS5) based on the resource utilization ratings; Changing the address translation scheme to the next address translation scheme; and Translating the system address (SA) to the memory address (MA) based on the nearest address translation scheme. [20] Method according to claim 19, wherein the calculation of the resource utilization assessment (RUA) comprises: Generating a plurality of possible memory addresses (MA1, MA3, MA3) corresponding to a plurality of system addresses, based on each of the plurality of possible translation schemes (TS1 to TS5); Generating a monitoring result by monitoring a change in memory resources associated with different memory addresses from the majority of eligible ones; and Calculating the resource utilization assessment based on the monitoring results.
Citation Information
Patent Citations
Memory controller and method for tuned address mapping
US20130132704A1