Method and related device for access control by means of multi-stage memory mapping queue

The instability of MLC flash memory is addressed by using a multi-stage memory mapping queue, which reduces the complexity and cost of the hardware architecture and improves the operational correctness and performance of the memory device.

CN114995745BActive Publication Date: 2025-10-28SILICON MOTION INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210204015.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-01
Filing Date
2022-03-02
Publication Date
2025-10-28
Estimated Expiration
2042-03-02

AI Technical Summary

Technical Problem

MLC flash memory has instability issues in memory devices, which leads to increased hardware architecture complexity and cost, and existing management mechanisms are difficult to solve effectively.

Method used

A multi-stage memory-mapped queue is used for access control. The multi-stage memory-mapped queue is shared by the processing circuit and the secondary processing circuit, which realizes message queuing and correct operation of the memory device, and reduces the complexity of the hardware architecture.

Benefits of technology

Ensure that memory devices operate correctly under various conditions, reduce costs, and improve the real-time response and overall performance of host devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114995745B_ABST
    Figure CN114995745B_ABST
Patent Text Reader

Abstract

This invention provides a method and related apparatus for access control of a memory device using a multi-stage memory-mapped queue. The method includes: receiving a first host instruction from a host device; and in response to the first host instruction, using processing circuitry within a controller to send a first operation instruction to a non-volatile memory via control logic circuitry of the controller, and triggering a first set of secondary processing circuitry within the controller to operate and interact through the multi-stage memory-mapped queue for accessing first data for the host device. The processing circuitry and the first set of secondary processing circuitry share the multi-stage memory-mapped queue, and the multi-stage memory-mapped queue is used as multiple link message queues associated with multiple stages for message queuing within a link processing architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to memory control, and more particularly to a method and related apparatus for access control using a multi-phase memory-mapped queue, such as a memory device, its memory controller, and a system-on-chip (SoC) integrated circuit (IC). Background Technology

[0002] Advances in memory technology have enabled the widespread use of various portable and non-portable memory devices, such as memory cards compliant with SD / MMC, CF, MS, and XD standards, and embedded storage devices compliant with UFS and eMMC standards. Improving access control for these memory devices has always been a problem that needs to be solved in this field.

[0003] NAND flash memory can include single-level cell (SLC) and multiple-level cell (MLC) flash memory. In SLC flash memory, each transistor, acting as a memory cell, can have either a charge value corresponding to logic 0 or 1. In contrast, in MLC flash memory, the storage capacity of each transistor acting as a memory cell can be fully utilized. Transistors in MLC flash memory can be driven by higher voltages than those in SLC flash memory, and different voltage levels can be used to record at least two bits of information (e.g., 00, 01, 11, or 10). Theoretically, the recording density of MLC flash memory can be at least twice that of SLC flash memory, which is why NAND flash memory manufacturers prefer MLC flash memory.

[0004] The low cost and large capacity of MLC flash memory mean it is more likely to be used in memory devices than SLC flash memory. However, MLC flash memory does have instability issues. To ensure that access control of flash memory in a memory device meets the required specifications, the flash memory controller can be equipped with certain management mechanisms to properly manage data access.

[0005] However, even memory devices with the aforementioned management mechanisms may have certain drawbacks. For example, multiple functional blocks within the hardware architecture, such as hardware engines, can be implemented to perform predetermined processing operations. Since more reliable results can be obtained by increasing the number of hardware engines, the number of message queues may increase accordingly. This can make the hardware architecture very complex, leading to increased costs. Therefore, a novel approach and related architecture are needed to avoid these problems without or with minimal side effects. Summary of the Invention

[0006] One object of the present invention is to provide a method and related devices, such as a memory device, its memory controller, a SoC IC, etc., for access control using a multi-stage memory mapping queue, in order to solve the above-mentioned problems.

[0007] At least one embodiment of the present invention discloses a method for access control of a memory device using a multi-phase memory-mapped queue. This method is applicable to a controller of the memory device, which includes the controller and a non-volatile (NV) memory comprising at least one NV memory element. The method includes: receiving a first host instruction from a host device, wherein the first host instruction indicates access to first data at a first logical address; and responding to the first host instruction, sending a first operation instruction to the NV memory via a control logic circuit of the controller using a processing circuit within the controller, and triggering a first set of secondary processing within the controller. The circuit operates and interacts through the multi-stage memory-mapped queue to access the first data for the host device, wherein the first operation instruction carries a first physical address associated with the first logical address to indicate a storage location in the non-volatile memory, and the processing circuit and the first set of secondary processing circuits share the multi-stage memory-mapped queue and use the multi-stage memory-mapped queue as multiple chained message queues associated with multiple stages for message queuing of a chained processing architecture including the processing circuit and the first set of secondary processing circuits.

[0008] In addition to the methods described above, the present invention also provides a memory device comprising a non-volatile (NV) memory and a controller. The non-volatile (NV) memory is used to store information, and includes at least one non-volatile memory element. The controller is coupled to the non-volatile memory and is used to control the operation of the memory device. The controller includes a processing circuit, a plurality of secondary processing circuits, and a multi-stage memory mapping queue. The processing circuit is used to control the controller according to a plurality of host instructions from a host device, allowing the host device to access the non-volatile memory through the controller. The plurality of secondary processing circuits are used to operate as a plurality of hardware engines. The multi-stage memory mapping queue, coupled to the processing circuit and the plurality of secondary processing circuits, is used to perform message queuing for the processing circuit and the plurality of secondary processing circuits. For example: the controller receives a first host instruction from the host device, wherein the first host instruction indicates access to first data at a first logical address, and the first host instruction is one of a plurality of host instructions; and in response to the first host instruction, the controller uses the processing circuitry to send a first operation instruction to the non-volatile memory through a control logic circuitry of the controller, and triggers a first group of secondary processing circuits among the plurality of secondary processing circuits to operate and interact through the multi-stage memory mapping queue for accessing the first data for the host device, wherein the first operation instruction carries a first physical address associated with the first logical address for indicating a storage location in the non-volatile memory, and the processing circuitry and the first group of secondary processing circuits share the multi-stage memory mapping queue, and use the multi-stage memory mapping queue as a plurality of chained message queues associated with a plurality of stages respectively, for message queuing for a chained processing architecture including the processing circuitry and the first group of secondary processing circuits.

[0009] In addition to the methods described above, the present invention also provides a controller for a memory device, the memory device including the controller and a non-volatile memory, the non-volatile memory including at least one non-volatile memory element, the controller including a processing circuit, a plurality of secondary processing circuits, and a multi-stage memory mapping queue. The processing circuit is used to control the controller according to a plurality of host instructions from a host device, so as to allow the host device to access the non-volatile memory through the controller. The plurality of secondary processing circuits are used to operate as a plurality of hardware engines. The multi-stage memory mapping queue is coupled to the processing circuit and the plurality of secondary processing circuits for queuing messages for the processing circuit and the plurality of secondary processing circuits. For example: the controller receives a first host instruction from the host device, wherein the first host instruction indicates access to first data at a first logical address, and the first host instruction is one of a plurality of host instructions; and in response to the first host instruction, the controller uses the processing circuitry to send a first operation instruction to the non-volatile memory through a control logic circuitry of the controller, and triggers a first group of secondary processing circuits among the plurality of secondary processing circuits to operate and interact through the multi-stage memory mapping queue for accessing the first data for the host device, wherein the first operation instruction carries a first physical address associated with the first logical address for indicating a storage location in the non-volatile memory, and the processing circuitry and the first group of secondary processing circuits share the multi-stage memory mapping queue, and use the multi-stage memory mapping queue as a plurality of chained message queues associated with a plurality of stages respectively, for message queuing for a chained processing architecture including the processing circuitry and the first group of secondary processing circuits.

[0010] The method and related apparatus of this invention ensure that the memory device operates correctly under various conditions. Examples of the aforementioned apparatus include controllers and memory devices. Furthermore, the method and related apparatus provided by this invention can solve related technical problems without side effects or with a low likelihood of causing side effects, thus reducing related costs. Additionally, the method and related apparatus provided by this invention, by utilizing a multi-stage memory mapping queue, can ensure real-time response of the memory device to the host device, thereby improving overall performance. Attached Figure Description

[0011] Figure 1 This is a schematic diagram of an electronic device according to an embodiment of the present invention.

[0012] Figure 2 The lower half of the diagram illustrates, according to an embodiment of the present invention, a method for performing a memory device such as... Figure 1This is a multi-phase queue control scheme for access control of memory devices, where, for better understanding, Figure 2 The upper half of the diagram illustrates a single-stage queue control scheme.

[0013] Figure 3 An embodiment of the present invention illustrates a remote update control scheme for the outgoing queue tail (OQT) of the method.

[0014] Figure 4 A one-stage inspection and control scheme of the method is illustrated according to an embodiment of the present invention.

[0015] Figure 5 A hybrid control scheme of the method is illustrated according to an embodiment of the present invention.

[0016] Figure 6 An embodiment of the present invention is illustrated as follows: Figure 2 The diagram shows an initial state before a series of operations in the multi-stage queue control scheme.

[0017] Figure 7 An embodiment of the present invention is illustrated as follows: Figure 2 The first intermediate state after one or more first operations in the series of operations described above in the multi-stage queue control scheme shown.

[0018] Figure 8 An embodiment of the present invention is illustrated as follows: Figure 2 The second intermediate state following one or more second operations in the series of operations described above in the multi-stage queue control scheme shown.

[0019] Figure 9 An embodiment of the present invention is illustrated as follows: Figure 2 The third intermediate state following one or more third operations in the series of operations described above in the multi-stage queue control scheme shown.

[0020] Figure 10 An embodiment of the present invention is illustrated as follows: Figure 2 The fourth intermediate state following one or more fourth operations in the series of operations described above in the multi-stage queue control scheme shown.

[0021] Figure 11 An embodiment of the present invention is illustrated as follows: Figure 2 The fifth intermediate state following one or more fifth operations in the series of operations described above in the multi-stage queue control scheme shown.

[0022] Figure 12 A queue splitting and merging control scheme of the method is illustrated according to an embodiment of the present invention.

[0023] Figure 13 An embodiment of the present invention is illustrated as follows: Figure 12 The diagram shows some implementation details of the queue splitting and merging control scheme.

[0024] Figure 14 An embodiment of the present invention illustrates a multi-memory-domain attribute control scheme of the method.

[0025] Figure 15 A queue splitting and merging control scheme of the method is illustrated according to another embodiment of the present invention.

[0026] Figure 16A A first part of a flowchart illustrating the method according to an embodiment of the present invention is shown.

[0027] Figure 16B The second part of the flowchart illustrating this method is shown.

[0028] Symbol Explanation

[0029] 10: Electronic Systems

[0030] 50: Main unit

[0031] 52: Processor

[0032] 54: Power Supply Circuit

[0033] 100: Memory device

[0034] 110: Memory controller

[0035] 112: Microprocessor

[0036] 112M: Read-Only Memory (ROM)

[0037] 112C: Program Code

[0038] 114: Control Logic Circuit

[0039] 115: Engine Circuit

[0040] 116: Random Access Memory (RAM)

[0041] 116T: Temporary Table

[0042] 118: Transmission Interface Circuit

[0043] 120: Non-volatile (NV) memory

[0044] 122-1, 122-2~122-N: Non-volatile (NV) memory elements

[0045] 122T: Non-temporary table

[0046] MMRBUF,MMRBUF(0),MMRBUF(1)

[0047] MMRBUF(2), MMRBUF(3): Memory-mapped ring buffers

[0048] IQH0, IQH1, IQH2, IQH3: Passed to the head of the queue

[0049] OQT0, OQT1, OQT2, OQT3: Output queue tail

[0050] S10, S11, S12, S12A, S12B, S12C

[0051] S12D, S12E, S12F, S16, S17, S17A

[0052] S17B, S17C, S17D, S17E, S17F: Steps Detailed Implementation

[0053] Embodiments of the present invention provide a method and related apparatus for access control of a memory device using a multi-phase memory-mapped queue, such as a multi-phase memory-mapped message queue. The apparatus may represent any application-specific integrated circuit (ASIC) product in which the multi-phase memory-mapped queue (such as a messaging mechanism between one or more processors / processor cores and multiple hardware engines) is implemented, but the invention is not limited thereto. The apparatus may include at least a portion (e.g., a portion or all) of an electronic device having an integrated circuit (IC) in which the multi-phase memory-mapped queue is implemented. For example, the apparatus may include a part of the electronic device, such as the memory device, its controller, etc. Alternatively, the apparatus may include the entire electronic device. Another example is that the apparatus may include a system-on-chip (SoC) integrated circuit, such as a SoC including the controller. Furthermore, certain control schemes of the present invention provide the following features:

[0054] (1) The multi-stage memory-mapped queue may contain multiple linked message queues implemented within a single memory-mapped ring buffer, wherein each of the multiple linked message queues is associated with a stage of multiple stages corresponding to the multiple linked message queues, and the multi-stage memory-mapped queue is transparent to each processor or engine for en-queuing or de-queuing operations.

[0055] (2) The architecture of the multi-stage memory-mapped queue can be varied. For example, the method and apparatus of the present invention can change one or more message flows of the multi-stage memory-mapped queue by splitting and merging message flows at any given stage.

[0056] (3) This multi-stage memory-mapped queue can be set up or established through flexible multi-memorydomain access attribute configuration. For example, for various data structures that require the implementation of a memory-mapped queue, the method and apparatus of this invention can dynamically configure memory access attributes associated with memory domains, such as doorbells, queue entries, queue bodies, and data buffers, wherein memory domain access attributes can be configured on a per-queue base or per-message base basis; and

[0057] (4) For memory-mapped queue distribution, the method and apparatus of the present invention can configure any incoming queue or outgoing queue in the engine as a linked or unlinked queue, wherein a request message queue can generate multiple linked messages and multiple completion messages to multiple linked outgoing queues and multiple completion outgoing queues.

[0058] One or more of the features listed above can be combined, but the present invention is not limited thereto. By using this multi-stage memory-mapped queue, the method and related apparatus of the present invention can solve related technical problems without side effects or with a low probability of causing side effects.

[0059] Figure 1This is a schematic diagram of an electronic system 10 according to an embodiment of the present invention, wherein the electronic system 10 includes a host device 50 and a memory device 100. The host device 50 may include at least one processor (e.g., one or more processors), collectively referred to as processor 52, and the host device 50 may further include a power supply circuit 54 coupled to processor 52. Processor 52 is configured to control the operation of the host device 50, and the power supply circuit 54 is configured to provide power to processor 52 and memory device 100, and to output one or more drive voltages to memory device 100. Memory device 100 may be configured to provide storage space to host device 50 and to obtain one or more drive voltages from host device 50 as power for memory device 100. Examples of host device 50 may include (but are not limited to): a multifunction mobile phone, a wearable device, a tablet computer, and personal computers such as desktop computers and laptop computers. Examples of memory device 100 may include (but are not limited to): a portable memory device (e.g., a memory card conforming to SD / MMC, CF, MS, or XD specifications), a solid-state drive (SSD), and various embedded storage devices such as embedded storage devices conforming to the Universal Flash Storage (UFS) standard or the embedded MMC (eMMC) standard. According to this embodiment, memory device 100 may include a controller such as a memory controller 110 and a non-volatile (NV) memory 120, wherein the controller is configured to control the operation of memory device 100 and access NV memory 120, and NV memory 120 is configured to store information. NV memory 120 may include at least one NV memory element (e.g., one or more NV memory elements), such as multiple NV memory elements 122-1, 122-2… and 122-N, where "N" may represent a positive integer greater than 1. For example, NV memory 120 can be flash memory, and multiple NV memory elements 122-1, 122-2, ..., 122-N can be multiple flash memory chips or multiple flash memory dies, but the present invention is not limited thereto. Furthermore, Figure 1 The electronic device 10, memory device 100, and controller such as memory controller 110 in the architecture shown can be examples of the electronic device, memory device, and controller with IC as described above.

[0060] like Figure 1As shown, the memory controller 110 may include a processing circuit such as a microprocessor 112, a storage unit such as a read-only memory (ROM) 112M, a control logic circuit 114, an engine circuit 115 (e.g., a digital signal processing (DSP) engine), a random access memory (RAM) 116, and a transmission interface circuit 118, wherein the above components can be interconnected via a bus. The engine circuit 115 may include multiple secondary processing circuits, and these multiple secondary processing circuits may be hardware functional blocks, which may be referred to as hardware engines. These hardware engines, such as engines #1, #2, etc., may include a direct memory access (DMA) engine, a compression engine, a decompression engine, an encoding engine (e.g., an encoder), a decoding engine (e.g., a decoder), a randomization engine (e.g., a randomizer), and a derandomization engine (e.g., a derandomizer), but the present invention is not limited thereto. For example, the encoding engine (e.g., the encoder) and the decoding engine (e.g., the decoder) can be integrated into the same module (such as an encoding and decoding engine), the randomization engine (e.g., the randomizer) and the derandomization engine (e.g., the derandomizer) can be integrated into the same module (such as a randomization and derandomization engine), and / or the compression engine and the decompression engine can be integrated into the same module (such as a compression and decompression engine). According to some embodiments, one or more sub-circuits of the engine circuit 115 can be integrated into the control logic circuit 114.

[0061] In engine circuit 115, a data protection circuit including the encoding engine (e.g., the encoder) and the decoding engine (e.g., the decoder) can be configured to protect data and / or perform error correction. Specifically, it can be configured to perform encoding and decoding operations separately, and the randomization engine (e.g., the randomizer) and the derandomization engine (e.g., the derandomizer) can be configured to perform randomization and derandomization operations respectively. Additionally, the compression engine and the decompression engine can be configured to perform compression and decompression operations respectively. Furthermore, the DMA engine can be configured to perform DMA operations. For example, during a data write operation requested by host device 50, the DMA engine can perform a DMA operation on a first host-side memory region of a memory in host device 50 via transfer interface circuit 118 to receive data from the first host-side memory region via transfer interface circuit 118. For example, during a data read as requested by host device 50, the DMA engine can perform DMA operations on a second host-side memory region (which may be the same as or different from the first host-side memory region) of the memory in host device 50 via transfer interface circuit 118 to transfer (e.g., return) data to the second host-side memory region via transfer interface circuit 118.

[0062] RAM 116 is implemented using static random access memory (SRAM), but the invention is not limited thereto. RAM 116 can be configured to provide internal storage space to memory controller 110. For example, RAM 116 can serve as a buffer memory to buffer data. In particular, RAM 116 can contain a memory region (e.g., a predetermined memory region) used as a multi-stage memory mapping queue (MPMMQ), which can be an example of the aforementioned multi-stage memory mapping queue, but the invention is not limited thereto. For example, the multi-stage memory mapping queue (MPMMQ) can be implemented in another memory within memory controller 110. Additionally, in this embodiment, ROM 112M is used to store program code 112C, and microprocessor 112 is used to execute program code 112C to control access to NV memory 120. Note that in some examples, program code 112C can be stored in RAM 116 or any form of memory. In addition, the transmission interface circuit 118 may conform to a specific communication standard (such as Serial Advanced Technology Attachment (SATA), Universal Serial Bus (USB), Peripheral Component Interconnect Express (PCIe), Embedded Multimedia Card (eMMC), or Universal Flash Storage (UFS)), and may communicate in accordance with that specific communication standard.

[0063] In this embodiment, the host device 50 can access the memory device 100 by sending a host instruction and a corresponding logical address to the memory controller 110. The memory controller 110 receives the host instruction and the logical address, converts the host instruction into a memory operation instruction (hereinafter referred to as an operation instruction), and controls the NV memory to read, write / program, etc., memory cells (e.g., data pages) with physical addresses in the NV memory 120, wherein the physical address may be associated with the logical address. When the memory controller 110 performs an erase operation on any NV memory element 122-n (the symbol "n" can represent any integer in the interval [1, N]) among a plurality of NV memory elements 122-1, 122-2, ... and 122-N, at least one block among the plurality of blocks of the NV memory device 122-n will be erased, wherein each of the plurality of blocks may contain a plurality of pages (e.g., data pages), and access operations (e.g., read or write) may be performed on one or more pages.

[0064] Some implementation details regarding the internal control of the memory device 100 can be further described below. According to some embodiments, the processing circuitry, such as the microprocessor 112, can control the memory controller 110 according to a plurality of host instructions from the host device 50, allowing the host device 50 to access the NV memory 120 via the memory controller 110. The memory controller 110 can store data for the host device 50 into the NV memory 120, read the stored data in response to a host instruction from the host device 50 (e.g., one of the plurality of host instructions), and provide the data read from the NV memory 120 to the host device 50. In the NV memory 120, such as flash memory, the aforementioned at least one NV memory element (e.g., a plurality of NV memory elements 122-1, 122-2, ..., and 122-N) can include a plurality of blocks, such as a first set of physical blocks in NV memory element 122-1, a second set of physical blocks in NV memory element 122-2, ..., and an Nth set of physical blocks in NV memory element 122-N. The memory controller 110 can be designed to properly manage the plurality of blocks, such as these groups of physical blocks (i.e., the first group to the Nth group of physical blocks; which may be referred to as the N physical block groups).

[0065] The memory controller 110 can record, maintain, and / or update block management information for block management in at least one table, such as at least one temporary table (e.g., one or more temporary tables) in RAM 116 and at least one non-temporary table (e.g., one or more non-temporary tables) in NV memory 120, wherein the at least one temporary table can be collectively referred to as temporary table 116T, and the at least one non-temporary table can be collectively referred to as non-temporary table 122T. Temporary table 116T may contain temporary versions of at least a portion (e.g., a portion or all) of non-temporary table 122T. For example, a non-temporary table 122T may contain at least one logical-to-physical (L2P) address mapping table (e.g., one or more L2P address mapping tables) to record multiple logical addresses (e.g., logical block addresses (LBAs) indicating multiple logical blocks and logical page addresses (LPAs) indicating multiple logical pages in any of the multiple logical blocks) and multiple physical addresses (e.g., physical block addresses (PBAs) indicating multiple physical blocks and physical page addresses indicating multiple physical pages in any of the multiple physical blocks). The temporary table 116T may contain a temporary version of at least one sub-table (e.g., one or more sub-tables) of the aforementioned at least one L2P address mapping table, wherein the memory controller 110 (e.g., microprocessor 112) may perform bidirectional address translation between the host-side storage space (e.g., logical address) of the host device 50 and the device-side storage space (e.g., physical address) of the NV memory 120 within the memory device 100, in order to access data for the host device 50. For better understanding, the non-temporary table 122T may be illustrated in the NV memory element 122-1, but the invention is not limited thereto. For example, the non-temporary table 122T may be stored in one or more of the NV memory elements 122-1, 122-2, ... and 122-N. Additionally, when needed, the memory controller 110 may back up the temporary table 116T to a non-temporary table 122T in the NV memory 120 (e.g., one or more NV memory elements 122-1, 122-2, ... and 122-N), and the memory controller 110 may load at least a portion (e.g., partially or entirely) of the non-temporary table 122T into RAM 116 to become the temporary table 116T for quick reference.

[0066] Figure 2The lower half of the diagram illustrates, according to an embodiment of the present invention, a method for performing a memory device such as... Figure 1 This is a multi-phase queue control scheme for access control of memory devices, where, for better understanding, Figure 2 The upper half of the diagram illustrates a single-stage queue control scheme. This method can be applied to... Figure 1 The architecture shown includes, for example, electronic device 10, memory device 100, memory controller 110, microprocessor 112, engine circuitry 115, and a multi-stage memory-mapped queue (MPMMQ). The MPMMQ within memory controller 110 can be implemented using a memory-mapped ring buffer (MMRBUF), which may reside within RAM 116, but the invention is not limited thereto.

[0067] like Figure 2 As shown in the upper part, a message chain can be implemented using a processor, multiple engines, and multiple memory-mapped ring buffers. For the case where the memory-mapped ring buffer count is four, the multiple memory-mapped ring buffers can comprise four memory-mapped ring buffers: MMRBUF(0), MMRBUF(1), MMRBUF(2), and MMRBUF(3). Each of these memory-mapped ring buffers can be considered a single-stage memory-mapped ring buffer. As the engine count increases, the memory-mapped ring buffer count can increase accordingly. Therefore, when the engine count is greater than one hundred, the memory-mapped ring buffer count is also greater than one hundred, which may lead to increased costs.

[0068] like Figure 2As shown in the lower part, the method and apparatus of the present invention can utilize a single memory-mapped ring buffer, such as a memory-mapped ring buffer MMRBUF, to replace the multiple memory-mapped ring buffers such as memory-mapped ring buffers MMRBUF(0), MMRBUF(1), etc., wherein the memory-mapped ring buffer MMRBUF can be regarded as a multi-stage memory-mapped ring buffer. The memory-mapped ring buffer MMRBUF can contain multiple sub-queues {SQ(x)} corresponding to multiple stages {Phase(x)}, such as sub-queues SQ(0), SQ(1), SQ(2), and SQ(3) corresponding to stages Phase(0), Phase(1), Phase(2), and Phase(3) respectively (referred to as "sub-queue SQ(0) of Phase(0)," "sub-queue SQ(1) of Phase(1)," "sub-queue SQ(2) of Phase(2)," and "sub-queue SQ(3) of Phase(3)" for brevity). Please note that the symbol "x" can represent any non-negative integer falling within the interval [0, (X-1)], and the symbol "X" can represent the subqueue count of the plurality of subqueues {SQ(x)}. As the engine count increases, the subqueue count X can increase accordingly, but the memory-mapped ring buffer count can be very small, more specifically, equal to 1. Therefore, when the engine count is greater than one hundred, the memory-mapped ring buffer count is very limited, thus saving related costs.

[0069] Although the local buffers in the memory-mapped ring buffers MMRBUF(0), MMRBUF(1), MMRBUF(2), and MMRBUF(3) can be described using the same or similar shading patterns as the subqueues SQ(0), SQ(1), SQ(2), and SQ(3) to indicate the storage location of messages output from the processor and engine, the local buffers in the memory-mapped ring buffers MMRBUF(0), MMRBUF(1), MMRBUF(2), and MMRBUF(3) are not equivalent to the subqueues SQ(0), SQ(1), SQ(2), and SQ(3). Note that in Figure 2In the single-stage queue control scheme shown in the upper part, each memory-mapped ring buffer (MMRBUF(0), MMRBUF(1), MMRBUF(2), and MMRBUF(3)) always exists as a whole. Furthermore, each of the memory-mapped ring buffers (MMRBUF(0), MMRBUF(1), MMRBUF(2), and MMRBUF(3)) operates independently. For example, an operation on one of the memory-mapped ring buffers (MMRBUF(0), MMRBUF(1), MMRBUF(2), and MMRBUF(3)) will not affect the operation on another of the memory-mapped ring buffers (MMRBUF(0), MMRBUF(1), MMRBUF(2), and MMRBUF(3)).

[0070] In this multi-stage queue control scheme, processors and engines #1, #2, etc., can respectively represent... Figure 1 The processing circuitry in the illustrated architecture includes, for example, a microprocessor 112, and secondary processing circuitry includes, for example, a hardware engine (or engine #1, #2, etc.), but the invention is not limited thereto. According to certain embodiments, the processor and engines #1, #2, etc., in this multi-stage queue control scheme can be replaced by any combination of one or more processors / processor cores and / or hardware engines, such as a combination of one processor core and multiple hardware engines, a combination of multiple processor cores and multiple hardware engines, a combination of multiple processor cores, a combination of multiple hardware engines, etc.

[0071] Some implementation details of the design rules for a multi-stage memory-mapped queue (MPMMQ) (e.g., a memory-mapped ring buffer (MMRBUF)) can be described as follows. According to certain embodiments, for any of the processors or engines #1, #2, etc., in this multi-stage queue control scheme, the associated outgoing queue (e.g., one of the plurality of sub-queues {SQ(x)}) will never overflow. This is guaranteed because an entry is only queued into the outgoing queue (e.g., input to the outgoing queue) when at least one entry has been dequeued from the incoming queue (e.g., output from the incoming queue). That is, the processor or any of the engines can operate according to the following rules:

[0072] (1) The outgoing queue tail (OQT) of the processor or any engine must never cross the incoming queue head (IQH) of the processor or any engine; and

[0073] (2) The head of the incoming queue IQH(x) must never exceed the tail of the outgoing queue OQT(x-1) of an upstream processor or engine. Regarding the former of the two rules above: if x > 0, the tail of the outgoing queue OQT(x) of engine #x will never exceed the head of the incoming queue IQH(x) of engine #x; otherwise, when x = 0, the tail of the outgoing queue OQT(0) of the processor will never exceed the head of the incoming queue IQH(0) of the processor. In particular, the processor or any engine can directly compare the local tail of the outgoing queue and the head of the incoming queue, such as the tail of the outgoing queue OQT(x) and the head of the incoming queue IQH(x). Furthermore, regarding the latter of the two rules above: if x > 1, then the head of the incoming queue IQH(x) of engine #x will never exceed the tail of the outgoing queue OQT(x-1) of the upstream engine #(x-1); if x = 1, then the head of the incoming queue IQH(1) of engine #1 will never cross the tail of the outgoing queue OQT(0) of the upstream processor; otherwise, when x = 0, the head of the incoming queue IQH(0) of this processor will never cross the tail of the outgoing queue OQT(X-1) of the upstream engine #(X-1). In particular, the processor or either engine can, according to respectively in Figures 3 to 5 The operation is performed using any of the embodiments shown. In subsequent embodiments, the input queue head IQH(x), such as input queue head (abbreviated as IQH) IQH(0), IQH(1), etc., and the output queue tail OQT(x), such as output queue tail (abbreviated as OQT) OQT(0), OQT(1), etc., can be written as IQH IQHx, such as IQH IQH0, IQH1, etc., and OQT OQTx, such as OQT OQT0, OQT1, etc., for the sake of brevity.

[0074] Figure 3 According to an embodiment of the present invention, a remote update control scheme for the outgoing queue tail is illustrated. The processor or either engine can remotely write the OQT to a downstream engine in the chain, and the downstream engine in the linked architecture can compare the remotely updated OQT (e.g., the OQT written by the processor or either engine) as the incoming queue tail with its local IQH. For example:

[0075] (1) The processor can remotely write OQT OQT0 to downstream engine #1 in the chain, and downstream engine #1 in the chain architecture can compare the remotely updated OQT OQT0 as the tail of the incoming queue with the local IQH IQH1 of downstream engine #1.

[0076] (2) Engine #1 can remotely write OQT OQT1 to downstream engine #2 in the chain, and downstream engine #2 in the chain architecture can compare the remotely updated OQT OQT1 as the tail of the incoming queue with the local IQH IQH2 of downstream engine #2.

[0077] (3) Engine #2 can remotely write OQT OQT2 to downstream engine #3 in the chain, and downstream engine #3 in the chain architecture can compare the remotely updated OQT OQT2 as the tail of the incoming queue with the local IQH IQH3 of downstream engine #3; and

[0078] (4) Engine #3 can remotely write OQT OQT3 to downstream processors in the chain, and the downstream processors in the chain architecture can compare the remotely updated OQT OQT3 as the tail of the incoming queue with the local IQH IQH0 of the downstream processor.

[0079] However, this invention is not limited thereto. For the sake of simplicity, similar content in this embodiment will not be repeated here.

[0080] Figure 4 A one-stage check control scheme for this method is illustrated according to an embodiment of the present invention. For each outgoing message sent to the outgoing queue in the chain, there is a stage field, and a unique stage field value is associated with each processor or engine. Specifically, the processor or any engine can read incoming queue entries from an upstream engine in the chain until the stage field value changes, wherein this reading can be solicited (e.g., via an interrupt) or unsolicited. Figure 4 As shown, the phase field value of each outgoing message sent to an outgoing queue such as subqueue SQ(0) can be equal to "01" (denoted as "Phase='01'" for simplicity), the phase field value of each outgoing message sent to an outgoing queue such as subqueue SQ(1) can be equal to "10" (denoted as "Phase='10'" for simplicity), the phase field value of each outgoing message sent to an outgoing queue such as subqueue SQ(2) can be equal to "11" (denoted as "Phase='11'" for simplicity), and the phase field value of each outgoing message sent to an outgoing queue such as subqueue SQ(3) can be equal to "00" (denoted as "Phase='00'" for simplicity), but the present invention is not limited thereto. For the sake of simplicity, similar content in this embodiment will not be repeated here.

[0081] Figure 5 A hybrid control scheme for the method is illustrated according to an embodiment of the present invention. The processor or any of the engines can be based on... Figure 3The OQT remote update control scheme and Figure 4 The phase check control scheme shown operates using one or more control schemes. For example, certain operations of the processor and engines #1, #2, and #3 may conform to... Figure 4 The OQT remote update control scheme is shown, and all engines in the tire processor and engines #1, #2, and #3 can be updated according to... Figure 4 The phased inspection and control scheme shown is implemented, but the present invention is not limited thereto. For the sake of simplicity, similar content in this embodiment will not be repeated here.

[0082] Figure 6 An embodiment of the present invention is illustrated as follows: Figure 2 The diagram shows an initial state before a series of operations in this multi-stage queue control scheme. For example... Figure 6 As shown, all IQHs and all OQTs are set to the same address pointer (e.g., a starting address) in the memory-mapped ring buffer MMRBUF. This initial state is a logical state where there are no outstanding messages (e.g., requests or completions) yet to be processed in any queue or engine, and the processor retains all entries. In this state, there are no head or tail pointers that can be moved except for the processor's OQTs (e.g., OQT OQT0).

[0083] Figure 7 An embodiment of the present invention is illustrated as follows: Figure 2 This is a first intermediate state following one or more of the first operations in the series of operations described above for the multi-stage queue control scheme. In these one or more first operations, the processor enqueues four entries to an outgoing queue, such as a subqueue SQ(0), which is passed to engine #1.

[0084] Figure 8 An embodiment of the present invention is illustrated as follows: Figure 2 The multi-stage queue control scheme shown is a second intermediate state following one or more of the second operations in the above series of operations. In the one or more second operations, engine #1 de-queues three entries from the incoming queue such as subqueue SQ(0) to obtain the three entries, and queues two entries into the outgoing queue such as subqueue SQ(1) passed to engine #2.

[0085] Figure 9 An embodiment of the present invention is illustrated as follows: Figure 2This is a third intermediate state following one or more third operations in the series of operations described above for the multi-stage queue control scheme. In one or more third operations, the processor queues twelve entries to an outgoing queue such as subqueue SQ(0) sent to engine #1, engine #1 dequeues ten entries from the incoming queue such as subqueue SQ(0) to obtain these ten entries, and queues eight entries to an output queue such as subqueue SQ(1) sent to engine #2, and engine #2 dequeues eight entries from the incoming queue such as subqueue SQ(1) to obtain these eight entries, and queues five entries to an output queue such as subqueue SQ(2) sent to engine #3.

[0086] Figure 10 An embodiment of the present invention is illustrated as follows: Figure 2 The above-described series of operations of the multi-stage queue control scheme is shown as a fourth intermediate state following one or more fourth operations. In one or more fourth operations, engine #3 dequeues four entries from an incoming queue such as subqueue SQ(2) to obtain these four entries, and queues three entries to an outgoing queue such as subqueue SQ(3) passed to the processor.

[0087] Figure 11 An embodiment of the present invention is illustrated as follows: Figure 2 The illustrated multi-stage queue control scheme represents a fifth intermediate state following one or more fifth operations in the aforementioned series of operations. In these one or more fifth operations, the processor dequeues two entries from an incoming queue, such as a subqueue SQ(3), to retrieve those two entries. For example, the processor or either engine may perform additional operations within the aforementioned series of operations.

[0088] Figure 12 A queue splitting and merging control scheme according to an embodiment of the present invention is illustrated. For better understanding, certain changes to sub-queues and related local message flows can be described as follows: Figure 12The diagram is explained using a logical view. For example, subqueue SQ(0) can be viewed as a request queue Q10, and subqueue SQ(3) can be viewed as a completion queue Q17. Since engine #2 is split into Y engines (e.g., the symbol "Y" can represent a positive integer greater than one), such as three engines #2.0, #2.1, and #2.2, the local message flow between engines #1 and #2 can logically be split into three local message flows. Therefore, subqueue SQ(1) can logically be split into corresponding subqueues such as link request queues Q11 to Q13, where some queue entries of link request queues Q11 to Q13 can be labeled as "0", "1", and "2" respectively to indicate that they are on the three local message flows from engine #1 to the three engines #2.0, #2.1, and #2.2. Similarly, since engine #2 is split into Y engines, such as three engines #2.0, #2.1, and #2.2, the local message flow between engines #2 and #3 can logically be split into three local message flows. Therefore, sub-queue SQ(2) can be logically split into corresponding sub-queues, such as link request queues Q14 to Q16. Some queue entries in link request queues Q14 to Q16 can be marked as "0", "1", and "2" respectively to indicate their location in the three local message flows from the three engines #2.0, #2.1, and #2.2 to engine #3. For the sake of simplicity, similar content will not be repeated here in this embodiment.

[0089] Figure 13 An embodiment of the present invention is illustrated as follows: Figure 12 The following are some implementation details of the queue splitting and merging control scheme. Sub-queues SQ(0) and SQ(3), which are unrelated to queue splitting and merging, can be regarded as request queue Q10 and completion queue Q17, respectively. For the purpose of implementing queue splitting and merging, sub-queue SQ(1) can be regarded as a combination of linked request queues Q11 to Q13, and sub-queue SQ(2) can be regarded as a combination of linked request queues Q14 to Q16. The respective queue entries of linked request queues Q11 to Q13, such as queue entries with shaded patterns corresponding to sub-queue SQ(1) marked as "0", "1", and "2", can be arranged in a predetermined order (e.g., the order of the three local message flows from engine #1 to the three engines #2.0, #2.1, and #2.2). Furthermore, the respective queue entries of the link request queues Q14 to Q16, such as queue entries with shaded patterns corresponding to sub-queue SQ(2) marked as "0", "1", and "2" respectively, can be arranged in a predetermined order (e.g., the order of the three local message flows from the three engines #2.0, #2.1, and #2.2 to engine #3). For the sake of simplicity, similar content will not be repeated here in this embodiment.

[0090] According to some embodiments, in the case of a 1:Y split (e.g., engine #2 is split into Y engines), queue entries before the split phase and after the merge phase can be spaced (Y-1). Therefore, the queue tails and heads in these phases represent an increase of Y. Additionally, for split queues, the queue tails and heads should also represent an increase of Y, but for any two split queues corresponding to the same sub-queue SQ(x), the respective offsets of these two split queues are different from each other. Figure 13 Taking the architecture shown as an example, queue entries with a shaded pattern corresponding to sub-queue SQ(0) and queue entries with a shaded pattern corresponding to sub-queue SQ(3) can be spaced (3-1) = 2 (e.g., there are two empty queue elements between two consecutive queue entries in these queue entries), and any OQT or IQH among OQT(0), OQT(1.0), OQT(1.1), OQT(1.2), OQT(2.0), OQT(2.1), OQT(2.2) and OQT(3) (referred to as "OQT0", "OQT1.0", "OQT1.1", "OQT1.2", "OQT2.0", "OQT2.1", "OQT2.2" and "OQT3") can be used. Any of the following IQH values ​​(IQH(0), IQH(1), IQH(2.0), IQH(2.1), IQH(2.2), IQH(3.0), IQH(3.1), and IQH(3.2) (labeled as "IQH0", "IQH1", "IQH2.0", "IQH2.1", "IQH2.2", "IQH3.0", "IQH3.1", and "IQH3.2", respectively) can be increased by an increment of three. For the sake of simplicity, similar content will not be repeated in this embodiment.

[0091] In this queue splitting and merging control scheme, engine #2 can be considered as an example of a split processing / secondary processing circuit among all processing / secondary processing circuits, but the invention is not limited thereto. According to some embodiments, the number of split processing / secondary processing circuits and / or the number of split processing / secondary processing circuits can be varied. For example, the split processing / secondary processing circuit can be the processor or another engine among engines #1, #2, etc. As another example, there may be more than one split processing / secondary processing circuit.

[0092] Figure 14According to an embodiment of the present invention, a multi-memory-domain attribute control scheme of the method is illustrated. A set of memory domains can be mapped using plane addresses, and the way a processing / secondary processing circuit (such as a processor / processor core, a hardware (HW) engine, etc.) accesses attributes for any two memory domains in the set of memory domains may differ from each other. For example, the set of memory domains may include:

[0093] (1) A doorbell register domain (referred to as "register domain" for simplicity), in which processing / secondary processing circuitry (e.g., processor core or hardware engine) can access the doorbell register (referred to as "doorbell" for simplicity) in the doorbell register domain for linking messages.

[0094] (2) A message header domain (also called a queue entry domain, since message headers form queue entries), wherein processing / secondary processing circuitry (e.g., processor cores or hardware engines) can access queue entries such as message headers (e.g., a request message header and a completion message header) in the message header domain, and this memory domain for these queue entries is configured on a per-queue basis (denoted as "per-queue" for brevity) by a per-queue register. For example, the memory of the message header domain can be any of a coherent domain level-one cache (L1$), a coherent domain level-two cache (L2$), or any of a non-coherent data / message memory.

[0095] (3) A message body domain, wherein processing / secondary processing circuitry (e.g., a processor core or hardware engine) can access the message body (e.g., a request message body and a completion message body) within the message body domain, and the message body domain for a message should be set or configured per message (denoted as "per message" for brevity) by means of a message header. For example, the memory of the message body domain may be any one of a co-homogeneous domain L1 cache (L1$), a co-homogeneous domain L2 cache (L2$), etc., or any one of a non-co-homogeneous data / message memory; and

[0096] (4) A message data buffer, wherein processing / secondary processing circuitry (e.g., processor core or hardware engine) can access message data buffers (e.g., multiple source data buffers and multiple destination data buffers) in the message data buffer, and the message data buffer shall be set or configured on a per SGL elementbase (denoted as "per SGL element" for brevity) in a message header or a message body;

[0097] However, the present invention is not limited thereto. For example, the request message header may include a completion queue number (denoted as "Completion Q#" for brevity) pointing to the relevant completion queue, a request message body address pointing to the request message body, a completion message body address pointing to the completion message body, an SGL address pointing to the multiple source data buffers, and an SGL address pointing to the multiple destination data buffers; and the request message body may include the SGL address pointing to the multiple source data buffers and the SGL address pointing to the multiple destination data buffers.

[0098] Figure 15 A queue splitting and merging control scheme of the method is illustrated according to another embodiment of the present invention. Figure 13 Compared to the architecture shown, certain local message flows can be further divided into more local message flows, and the associated sub-queues can also be divided accordingly. For example, sub-queue SQ(0) can be regarded as a combination of request queues Q10, Q20, and Q30; sub-queue SQ(1) can be regarded as a combination linking request queues Q11-Q13, Q21, and Q31; and sub-queue SQ(2) can be regarded as a combination linking request queues Q14-Q16 and completion queues Q22 and Q32. Sub-queue SQ(3), which is unrelated to queue splitting and merging, can be regarded as completion queue Q17. For the sake of simplicity, similar content will not be repeated in this embodiment.

[0099] Figure 16A as well as Figure 16B According to an embodiment of the present invention, a first part and a second part of a flowchart of the method are respectively drawn, wherein nodes A and B can indicate Figure 16A and Figure 16B This method connects the various local workflows. It can be applied to... Figure 1The architecture shown (e.g., electronic device 10, memory device 100, memory controller 110, and microprocessor 112) can be executed by the memory controller 110 (e.g., microprocessor 112) of memory device 100. During a first type of access operation (e.g., one of data read and data write) requested by host device 50, the processing circuitry such as microprocessor 112 and a group of engines corresponding to the first type of access operation in a first set of secondary processing circuitry such as engines #1, #2, etc. can share a multi-stage memory-mapped queue (MPMMQ) (e.g., memory-mapped ring buffer MMRBUF) and use the multi-stage memory-mapped queue (MPMMQ) as multiple linked message queues associated with multiple stages (e.g., multiple sub-queues {SQ(x)} corresponding to the multiple stages {Phase(x)}, configured for the first type of access operation) for message queuing for a linked processing architecture including the processing circuitry and the first set of secondary processing circuitry. Additionally, during a second type of access operation (e.g., data read and data write) requested by the host device 50, the processing circuitry such as the microprocessor 112 and a second set of secondary processing circuitry such as engines #1, #2, etc., corresponding to the second type of access operation, can share a multi-stage memory-mapped queue (MPMMQ) (e.g., a memory-mapped ring buffer MMRBUF), and use the MPMMQ as multiple other linked message queues associated with multiple other stages (e.g., multiple sub-queues {SQ(x)} corresponding to the multiple stages {Phase(x)}, configured for the second type of access operation), for message queuing for another linked processing architecture including the processing circuitry and the second set of secondary processing circuitry.

[0100] In step S10, the memory controller 110 (e.g., microprocessor 112) may determine whether a host instruction (e.g., one of a plurality of host instructions) has been received. If yes, proceed to step S11; otherwise, proceed to step S10.

[0101] In step S11, the memory controller 110 (e.g., microprocessor 112) determines whether the host instruction (i.e., the host instruction just received and detected in step S10) is a host read instruction. If yes, proceed to step S12; otherwise, proceed to step S16. The host read instruction may indicate that data is to be read from a first logical address.

[0102] In step S12, in response to the host read instruction (i.e., the host read instruction just received as detected in steps S10 and S11), the memory controller 110 (e.g., microprocessor 112) can send a first operation instruction (e.g., one of the above operation instructions) such as a read instruction to the NV memory 120 via the control logic circuit 114, and trigger a first set of engines to operate and interact via a multi-stage memory mapping queue (MPMMQ, e.g., memory mapping ring buffer MMRBUF) to read data for the host device 50 (e.g., data to be read as requested by the host device 50). The first operation instruction, such as the read instruction, may carry a first physical address associated with the first logical address to indicate a storage location within the NV memory 120, and the first physical address may be determined by the memory controller 110 (e.g., microprocessor 112) based on the above-mentioned at least one L2P address mapping table.

[0103] For example, the first set of engines may include the derandomization engine (e.g., the derandomizer), the decoding engine (e.g., the decoder), the decompression engine, and the DMA engine, and the processing circuitry such as the microprocessor 112 and certain secondary processing circuitry such as these engines may form a configuration similar to... Figure 2 The lower half shows a link processing architecture for a multi-stage queue control scheme, wherein the processor and engines #1, #2, etc. in this multi-stage queue control scheme can respectively represent the microprocessor 112 and these engines arranged in the order listed above, but the present invention is not limited thereto. According to some embodiments, the first group of engines and the associated link processing architecture can be varied.

[0104] In step S12A, the memory controller 110 (e.g., microprocessor 112) may use the control logic circuit 114 to read the NV memory 120, specifically, to read from a location in the NV memory 120 (e.g., a physical address within a range of physical addresses starting from the first physical address) to obtain read data from the NV memory 120.

[0105] In step S12B, the memory controller 110 (e.g., microprocessor 112) may use the derandomization engine (e.g., the derandomizer) to derandomize the read data, such as a derandomization operation, to generate derandomized data.

[0106] In step S12C, the memory controller 110 (e.g., microprocessor 112) may use the decoding engine (e.g., the decoder) to decode the derandomized data, such as a decoding operation, to generate decoded data (e.g., error-corrected data, such as data corrected based on parity data).

[0107] In step S12D, the memory controller 110 (e.g., microprocessor 112) may use the decompression engine to decompress the decoded data, such as a decompression operation, to produce decompressed data as partial data for use in data to be sent (e.g., returned) to the host device 50.

[0108] In step S12E, the memory controller 110 (e.g., microprocessor 112) can use the DMA engine to perform DMA operations on the second host-side memory region of the memory in the host device 50 via the transfer interface circuit 118, so as to send (e.g., return) the partial data of the data (e.g., the data to be read as requested by the host device 50) to the second host-side memory region via the transfer interface circuit 118.

[0109] In step S12F, the memory controller 110 (e.g., microprocessor 112) may check whether the entire data read of the data (e.g., the data to be read as requested by the host device 50) has been completed. If yes, proceed to step S10; otherwise, proceed to step S12A.

[0110] For data reading performed via a loop including sub-steps (e.g., steps S12A to S12F) of step S12, a multi-stage memory-mapped queue (MPMMQ) containing multiple linked message queues for data reading (e.g., the multiple sub-queues {SQ(x)} configured for data reading) is implemented within a single memory-mapped ring buffer such as a memory-mapped ring buffer (MMRBUF), and each of the multiple linked message queues is associated with one of the multiple stages (e.g., the multiple stages {Phase(x)} configured for data reading), wherein the multi-stage memory-mapped queue is transparent to each of the processing circuitry (e.g., microprocessor 112) and the set of secondary processing circuitry corresponding to the data reading (e.g., the first set of engines such as the derandomization engine, the decoding engine, the decompression engine, and the DMA engine) for en-queuing or de-queuing operations. Under the control of the processing circuit and at least one circuit (e.g., one or more circuits) in the set of secondary processing circuits corresponding to data reading, the multiple sub-queues {SQ(x)} corresponding to the multiple phases {Phase(x)} can be configured to have dynamically adjusted queue lengths, such as the respective queue lengths of sub-queues SQ(0), SQ(1), etc., for use as the multiple linked message queues for data reading.

[0111] In step S16, the memory controller 110 (e.g., microprocessor 112) can determine whether the host instruction (i.e., the host instruction just received and detected in step S10) is a host write instruction. If yes, proceed to step S17; otherwise, proceed to step S18. The host write instruction can indicate writing data to a second logical address.

[0112] In step S17, in response to the host write instruction (i.e., the host write instruction just received as detected in steps S10 and S16), the memory controller 110 (e.g., microprocessor 112) can send a second operation instruction (e.g., another operation instruction among the above operation instructions) such as a write instruction to the NV memory 120 via the control logic circuit 114, and trigger a second set of engines to operate and interact via a multi-stage memory-mapped queue (MPMMQ, e.g., memory-mapped ring buffer (MMRBUF)) to write data to the host device 50 (e.g., data to be written as requested by the host device 50), wherein the second operation instruction such as the write instruction may carry a second physical address associated with the second logical address to indicate a storage location within the NV memory 120, and the second physical address may be determined by the memory controller 110 (e.g., microprocessor 112).

[0113] For example, the second set of engines may include the DMA engine, the compression engine, the encoding engine (e.g., the encoder), and the randomization engine (e.g., the randomizer), and the processing circuitry such as the microprocessor 112 and certain secondary processing circuitry such as these engines can form a configuration similar to... Figure 2 The lower half shows a link processing architecture for a multi-stage queue control scheme, wherein the processors and engines #1, #2, etc., in this multi-stage queue control scheme can respectively represent the microprocessor 112 and these engines arranged in the order listed above, but the present invention is not limited thereto. According to some embodiments, the second group of engines and the associated link processing architecture can be varied.

[0114] In step S17A, the memory controller 110 (e.g., microprocessor 112) can use the DMA engine to perform DMA operations on the first host-side memory region of the memory in the host device 50 via the transfer interface circuit 118, so as to receive partial data of data (e.g., data to be written as requested by the host device 50) from the first host-side memory region via the transfer interface circuit 118 as received data.

[0115] In step S17B, the memory controller 110 (e.g., microprocessor 112) may use the compression engine to compress the received data, such as performing a compression operation, to produce compressed data.

[0116] In step S17C, the memory controller 110 (e.g., microprocessor 112) may use the encoding engine (e.g., the encoder) to encode the compressed data, such as an encoding operation, to produce encoded data (e.g., a combination of the received data and its co-occurring data).

[0117] In step S17D, the memory controller 110 (e.g., microprocessor 112) may use the randomization engine (e.g., the randomizer) to randomize the encoded data, such as performing randomization operations, to generate randomized data.

[0118] In step S17E, the memory controller 110 (e.g., microprocessor 112) can be programmed using control logic circuitry 114, specifically, the randomized data is programmed into the NV memory 120.

[0119] In step S17F, the memory controller 110 (e.g., microprocessor 112) may check whether the overall data writing of the data (e.g., the data to be written as requested by the host device 50) has been completed. If yes, proceed to step S10; otherwise, proceed to step S17A.

[0120] For data writing performed via a loop including sub-steps (e.g., steps S17A to S17F) of step S17, a multi-stage memory-mapped queue (MPMMQ) containing multiple linked message queues for data writing (e.g., the multiple sub-queues {SQ(x)} configured for data writing) is implemented within a single memory-mapped ring buffer such as a memory-mapped ring buffer (MMRBUF), and each of the multiple linked message queues is associated with one of the multiple stages (e.g., the multiple stages {Phase(x)} configured for data writing), wherein the multi-stage memory-mapped queue is transparent to each of the processing circuitry (e.g., microprocessor 112) and the set of secondary processing circuitry corresponding to the data writing (e.g., the second set of engines such as the DMA engine, the compression engine, the encoding engine, and the randomization engine) for queuing or dequeuing operations. Under the control of the processing circuit and at least one circuit (e.g., one or more circuits) in the set of secondary processing circuits corresponding to data writing, the multiple sub-queues {SQ(x)} corresponding to the multiple phases {Phase(x)} can be configured to have dynamically adjusted queue lengths, such as the respective queue lengths of sub-queues SQ(0), SQ(1), etc., for use as the multiple linked message queues for data writing.

[0121] In step S18, the memory controller 110 (e.g., microprocessor 112) may perform other processing. For example, when the host instruction (e.g., the host instruction just received as detected in step S10) is a different instruction from either the host read instruction or the host write instruction, the memory controller 110 (e.g., microprocessor 112) may perform related operations.

[0122] To better understand this method, we can use... Figure 16A as well as Figure 16B The workflow shown is for illustrative purposes only, but the invention is not limited thereto. According to some embodiments, one or more steps may be performed... Figure 16A as well as Figure 16B Add, delete, or modify within the workflow shown.

[0123] According to certain embodiments, assuming X = 5, some relationships between the plurality of subqueues {SQ(x)} and the associated local message flows can be described as follows:

[0124] (1) An initial local message flow between the processing circuit (e.g., microprocessor 112) and a first primary processing circuit (e.g., the derandomization engine) in the group of secondary processing circuits (e.g., the first group of engines) corresponding to data reading can be passed through a subqueue SQ(0) corresponding to the phase (0).

[0125] (2) A first local message flow between the first primary processing circuit (e.g., the derandomization engine) and a second secondary processing circuit (e.g., the decoding engine) in the group of secondary processing circuits (e.g., the first group of engines) corresponding to data reading can be passed through the subqueue SQ(1) corresponding to the phase (1).

[0126] (3) A second local message flow between the second-level processing circuit (e.g., the decoding engine) and a third-level processing circuit (e.g., the decompression engine) in the group of secondary processing circuits (e.g., the first group of engines) corresponding to data reading can be transmitted through a sub-queue SQ(2) corresponding to the phase (2).

[0127] (4) A third local message flow between the third-level processing circuit (e.g., the decompression engine) and a fourth-level processing circuit (e.g., the DMA engine) in the group of secondary processing circuits (e.g., the first group of engines) corresponding to data reading can be transmitted through a sub-queue SQ(3) corresponding to phase (3); and

[0128] (5) A fourth local message flow between the fourth secondary processing circuit (e.g., the DMA engine) and the processing circuit (e.g., the microprocessor 112) can be transmitted through a subqueue SQ (4) corresponding to the phase (4);

[0129] However, the present invention is not limited thereto. For the sake of brevity, similar contents in these embodiments will not be repeated here.

[0130] According to certain embodiments, assuming X = 5, some relationships between the plurality of subqueues {SQ(x)} and the associated local message flows can be described as follows:

[0131] (1) An initial local message flow between the processing circuit (e.g., microprocessor 112) and a primary processing circuit (e.g., the DMA engine) in the group of secondary processing circuits (e.g., the second group of engines) corresponding to data writing can be transmitted through a sub-queue SQ(0) corresponding to the phase (0).

[0132] (2) A first local message flow between the first primary processing circuit (e.g., the DMA engine) and a second secondary processing circuit (e.g., the compression engine) in the group of secondary processing circuits (e.g., the second group of engines) corresponding to data writing can be transmitted through a sub-queue SQ(1) corresponding to the phase (1).

[0133] (3) A second local message flow between the second-level processing circuit (e.g., the compression engine) and a third-level processing circuit (e.g., the encoding engine) in the group of secondary processing circuits (e.g., the second group of engines) corresponding to data writing can be transmitted through a sub-queue SQ(2) corresponding to the phase (2).

[0134] (4) A third local message flow between the third-level processing circuit (e.g., the encoding engine) and a fourth-level processing circuit (e.g., the randomization engine) in the group of secondary processing circuits (e.g., the second group of engines) corresponding to data writing can be transmitted through a subqueue SQ(3) corresponding to phase (3); and

[0135] (5) A fourth local message flow between the fourth secondary processing circuit (e.g., the randomization engine) and the processing circuit (e.g., the microprocessor 112) can be transmitted through a subqueue SQ (4) corresponding to the phase (4).

[0136] However, the present invention is not limited thereto. For the sake of brevity, similar contents in these embodiments will not be repeated here.

[0137] One of the many advantages of this invention is that the method and apparatus of this invention can reduce the memory size requirement. For example, by sharing multiple linked queues from only one queue of the same size, the memory requirement can be reduced to 1 / X when comparing an X-stage queue to X single-stage queues. This is important because in most applications, queue entries reside in a coherent memory domain, with each entry being a cache line. Furthermore, the method and apparatus of this invention can merge multiple multi-stage memory-mapped queues, particularly by adding splitting and merging at a synchronization point at a certain stage of the message. This increases the parallel operation of multiple engines, thereby reducing the latency of the overall message flow. Since multiple memory domains typically exist in a complex SoC design, and different memory domains can have different memory attributes, management can be very complex. The method and apparatus of this invention can provide a well-designed, high-performance architecture. To implement memory-mapped queues, four data structures are required for the doorbell register, queue entries, message body, and data buffer. Depending on the application, SoC device operating status, quality of service (QoS), and performance requirements, each data structure may need to be dynamically mapped to a memory domain. The method and apparatus of this invention can utilize a general-purpose distributed memory queuing system to dynamically configure memory access attributes associated with each domain to achieve QoS and performance goals. Furthermore, for a complex SoC device, flexible or even dynamic configuration of message flow may be necessary, particularly depending on specific application, QoS and performance requirements, and device status. The method and apparatus of this invention can provide flexibility for linking one engine to multiple engines and sending completion messages at different levels / stages of the message flow. This can improve overall performance.

[0138] The above description is only a preferred embodiment of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the present invention.

Claims

1. A method for access control of a memory device using a multi-stage memory mapping queue, the method being a controller applicable to the memory device, the memory device including the controller and a non-volatile memory, the non-volatile memory including at least one non-volatile memory element, the method comprising: Receive a first host instruction from a host device, wherein the first host instruction indicates accessing first data at a first logical address; and In response to the first host instruction, a processing circuit within the controller sends a first operation instruction to the non-volatile memory via a control logic circuit of the controller, and triggers a first group of secondary processing circuits within the controller to operate and interact through the multi-stage memory mapping queue for accessing the first data for the host device. The first operation instruction carries a first physical address associated with the first logical address to indicate a storage location in the non-volatile memory. The processing circuit and the first group of secondary processing circuits share the multi-stage memory mapping queue and use the multi-stage memory mapping queue as multiple link message queues associated with multiple stages for message queuing of a link processing architecture including the processing circuit and the first group of secondary processing circuits.

2. The method as described in claim 1, characterized in that... The multi-stage memory-mapped queue, which contains the multiple linked message queues, is implemented within a single memory-mapped ring buffer.

3. The method as described in claim 1, characterized in that... Each of the multiple linked message queues is associated with one of the multiple stages.

4. The method as described in claim 1, characterized in that... For each of the processing circuits and the first group of secondary processing circuits, the multi-stage memory-mapped queue is transparent for queuing or dequeuing operations.

5. The method as described in claim 1, characterized in that... Under the control of at least one of the processing circuits and the first group of secondary processing circuits, multiple sub-queues corresponding to the multiple stages are configured with dynamically adjusted queue lengths for use as the multiple linked message queues.

6. The method as described in claim 5, characterized in that: A first local message flow between a primary processing circuit and a secondary processing circuit in the first group of secondary processing circuits passes through a first sub-queue corresponding to a first stage in the plurality of sub-queues; and A second local message flow between the second-level processing circuit and a third-level processing circuit in the first group of secondary processing circuits passes through a second sub-queue corresponding to a second stage in the plurality of sub-queues.

7. The method as described in claim 6, characterized in that: A partial message flow between the processing circuit and the first-level processing circuit passes through another sub-queue in the plurality of sub-queues that corresponds to another stage.

8. The method as described in claim 1, characterized in that... The first set of secondary processing circuits includes a direct memory access engine, a decoding engine, and a derandomization engine.

9. The method as described in claim 1, characterized in that... The first set of secondary processing circuits includes a direct memory access engine, an encoding engine, and a randomization engine.

10. The method as described in claim 1, characterized in that, Also includes: The host device receives a second host instruction, wherein the second host instruction indicates accessing second data at a second logical address; as well as In response to the second host instruction, the processing circuit within the controller sends a second operation instruction to the non-volatile memory and triggers a second set of secondary processing circuits within the controller to operate and interact through the multi-stage memory mapping queue for accessing the second data for the host device. The second operation instruction carries a second physical address associated with the second logical address to indicate another storage location in the non-volatile memory. The processing circuit and the second set of secondary processing circuits share the multi-stage memory mapping queue and use the multi-stage memory mapping queue as multiple other link message queues associated with multiple other stages for message queuing for another link processing architecture including the processing circuit and the second set of secondary processing circuits.

11. A system-on-a-chip integrated circuit that operates according to the method of claim 1, wherein the system-on-a-chip integrated circuit includes the controller.

12. A memory device comprising: A non-volatile memory for storing information, wherein the non-volatile memory includes at least one non-volatile memory element; and A controller, coupled to the non-volatile memory, is used to control the operation of the memory device, wherein the controller includes: A processing circuit is used to control the controller according to a plurality of host instructions from a host device, so as to allow the host device to access the non-volatile memory through the controller. Multiple secondary processing circuits are used to operate as multiple hardware engines; as well as A multi-stage memory-mapped queue, coupled to the processing circuit and the plurality of secondary processing circuits, is used to queue messages for the processing circuit and the plurality of secondary processing circuits. in: The controller receives a first host instruction from the host device, wherein the first host instruction indicates access to first data at a first logical address, and the first host instruction is one of a plurality of host instructions; as well as In response to the first host instruction, the controller uses the processing circuitry to send a first operation instruction to the non-volatile memory through a control logic circuitry of the controller, and triggers a first group of secondary processing circuits among the plurality of secondary processing circuits to operate and interact through the multi-stage memory mapping queue for accessing the first data for the host device. The first operation instruction carries a first physical address associated with the first logical address to indicate a storage location in the non-volatile memory. The processing circuitry and the first group of secondary processing circuits share the multi-stage memory mapping queue and use the multi-stage memory mapping queue as a plurality of link message queues associated with the plurality of stages for message queuing of a link processing architecture including the processing circuitry and the first group of secondary processing circuits.

13. A controller for a memory device, the memory device including the controller and a non-volatile memory, the non-volatile memory including at least one non-volatile memory element, the controller comprising: A processing circuit is used to control the controller according to a plurality of host instructions from a host device, so as to allow the host device to access the non-volatile memory through the controller. Multiple secondary processing circuits are used to operate as multiple hardware engines; as well as A multi-stage memory-mapped queue, coupled to the processing circuit and the plurality of secondary processing circuits, is used to queue messages for the processing circuit and the plurality of secondary processing circuits. in: The controller receives a first host instruction from the host device, wherein the first host instruction indicates access to first data at a first logical address, and the first host instruction is one of a plurality of host instructions; as well as In response to the first host instruction, the controller uses the processing circuitry to send a first operation instruction to the non-volatile memory through a control logic circuitry of the controller, and triggers a first group of secondary processing circuits among the plurality of secondary processing circuits to operate and interact through the multi-stage memory mapping queue for accessing the first data for the host device. The first operation instruction carries a first physical address associated with the first logical address to indicate a storage location in the non-volatile memory. The processing circuitry and the first group of secondary processing circuits share the multi-stage memory mapping queue and use the multi-stage memory mapping queue as a plurality of link message queues associated with the plurality of stages for message queuing of a link processing architecture including the processing circuitry and the first group of secondary processing circuits.

Citation Information

Patent Citations

  • System and method for processing and arbitrating submission and completion queues

    CN110088725A

  • Packet processing device

    CN1467965A