A performance analysis system and analysis method for a memory

CN119649891BActive Publication Date: 2025-06-03合肥康芯威存储技术有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510174973.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-03
Estimated Expiration
2045-02-18

Smart Images

  • Figure CN119649891B_ABST
    Figure CN119649891B_ABST
Patent Text Reader

Abstract

The present invention provides a performance analysis system and method for a memory. The performance analysis system includes a test board, to which a memory and a logic analyzer are connected. The logic analyzer is communicatively connected to the flash interface of the memory. A processing module is disposed on the test board and is communicatively connected to the storage interface of the memory. And a host is communicatively connected to the processing module and the logic analyzer. The host is used to send different test instructions to the memory through the processing module so that the memory performs different operations. Among them, the logic analyzer is used to obtain the instruction stream and response data sent by the firmware of the memory to the flash interface when the memory performs different operations, upload the instruction stream and response data to the host, and perform performance analysis on different memories. Through the performance analysis system and method for a memory provided by the present invention, the differences in the firmware of different memories can be analyzed to improve the performance of the firmware.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of storage, and particularly to a performance analysis system and method for a memory. Background Art

[0002] Memories are widely used in terminal devices such as mobile phones, tablet computers, televisions, and set-top boxes. Currently, there are many manufacturers producing memories in the market. For different manufacturers, the firmware of the memories they design is also different, and the quality of the firmware directly affects the performance of the memory. How to quickly analyze the differences in the firmware of memories produced by different manufacturers to improve the performance of the firmware is an urgent problem to be solved currently. Summary of the Invention

[0003] The purpose of the present invention is to provide a performance analysis system and method for a memory, which can analyze the differences in the firmware of the memory to improve the performance of the firmware.

[0004] To solve the above technical problems, the present invention is implemented through the following technical solutions:

[0005] The present invention provides a performance analysis system for a memory, including:

[0006] A test board, on which a memory is connected. A plurality of detection pins are set on the test board, and the detection pins are communicatively connected to the test pins of the memory;

[0007] A logic analyzer, which is communicatively connected to the flash interface of the memory through the detection pins;

[0008] A processing module, which is arranged on the test board and is communicatively connected to the storage interface of the memory; and

[0009] A host, which is communicatively connected to the processing module and the logic analyzer. The host is used to send different test instructions to the memory through the processing module to make the memory perform different operations;

[0010] Wherein, the logic analyzer is used to obtain the instruction stream and response data sent by the firmware to the flash interface when the memory performs different operations, upload the instruction stream and response data to the host, and perform performance analysis on different memories.

[0011] In an embodiment of the present invention, the host is further used to send initialization instructions to a plurality of memories respectively to obtain the instruction stream of the internal firmware and the corresponding boot time when the plurality of memories perform initialization operations.

[0012] In an embodiment of the present invention, the host is further configured to compare and analyze the instruction streams of the memory with the shortest boot time with those of other memories, so as to optimize the firmware in the directions of reducing unknown commands, adopting a more efficient storage mode, and optimizing the data reading method.

[0013] In an embodiment of the present invention, the host is further configured to send multi-page read instructions and single-page read instructions to multiple memories respectively, so as to obtain the instruction streams and corresponding read performance values of the internal firmware of the multiple memories when performing multi-page read operations and single-page read operations.

[0014] In an embodiment of the present invention, the host is further configured to compare and analyze the instruction streams of the memory with the best read performance value with those of other memories, so as to optimize the firmware in the directions of reducing unknown commands, adopting a more efficient storage mode, optimizing the data access method, and utilizing the cache and prefetch mechanisms.

[0015] In an embodiment of the present invention, the host is further configured to send multi-page write instructions and single-page write instructions to multiple memories respectively, so as to obtain the instruction streams and corresponding write performance values of the internal firmware of the multiple memories when performing multi-page write operations and single-page write operations.

[0016] In an embodiment of the present invention, the host is further configured to compare and analyze the instruction streams of the memory with the best write performance value with those of other memories, so as to optimize the firmware in the directions of reducing unknown commands, optimizing the write process to reduce read and erase operations, adopting a more efficient storage mode, and optimizing the data write strategy to reduce the number of programming times.

[0017] In an embodiment of the present invention, the host is further configured to send partial erase instructions and full disk erase instructions to multiple memories respectively, so as to obtain the instruction streams and corresponding erase performance values of the internal firmware of the multiple memories when performing partial erase operations and full disk erase operations. The host is further configured to compare and analyze the instruction streams of the memory with the best erase performance value with those of other memories, so as to optimize the firmware.

[0018] In an embodiment of the present invention, the host is further configured to send queue instructions to multiple memories respectively, so as to obtain the instruction execution order and execution duration of the multiple memories when performing queue operations. The host is further configured to compare and analyze the instruction execution order of the memory with the shortest execution duration with the instruction execution orders of other memories, so as to optimize the firmware.

[0019] The present invention also provides a method for analyzing the performance of a memory, including:

[0020] Construct a test environment for the memory;

[0021] Send different test instructions to the memory so that the memory performs different operations;

[0022] Obtain the instruction stream and response data sent by the firmware of the memory to the flash interface when the memory performs different operations;

[0023] Upload the instruction stream and response data and perform performance analysis on them.

[0024] As described above, the present invention provides a performance analysis system and method for a memory. By monitoring the execution status of the NAND Interface of different memories, it is possible to analyze and identify the behavioral differences of the firmware, and quickly locate and analyze the differential behaviors of the firmware of different storage manufacturers. By analyzing the differential behaviors of the firmware, it is possible to identify the deficiencies in the firmware, improve the overall system performance, and quickly take optimization measures to improve the overall system performance.

[0025] Of course, it is not necessary for any product implementing the present invention to achieve all the above advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 It is a schematic diagram of a performance analysis system for a memory in an embodiment of the present invention;

[0028] Figure 2 It is a flowchart of a performance analysis method for a memory in an embodiment of the present invention.

[0029] In the figure: 100, test board; 200, processing module; 300, detection pin; 400, logic analyzer; 500, host; 600, power supply module; 700, transmission module; 800, interface module; 900, memory module; 1000, memory. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0031] Please refer to Figure 1, the present invention provides a test system for a memory, which can analyze the performance of memories produced by different manufacturers, quickly analyze the differences in the firmware of memories produced by different manufacturers, and improve the firmware of the memories according to the differences. The test system may include a test board 100, a processing module 200 (CPU), detection pins 300 (TPx), a logic analyzer 400, a host 500, a power supply module 600 (POWER), a transmission module 700 (USB), an interface module 800 (UART), and a memory module 900 (DRAM).

[0032] In one embodiment, the test board 100 may be a circuit board (PCB). The processing module 200, the detection pins 300, the power supply module 600, the transmission module 700, the interface module 800, and the memory module 900 may be integrated on the test board 100.

[0033] In one embodiment, the processing module 200 may be responsible for executing instructions in a computer program and controlling other hardware devices to work cooperatively. The processing module 200 may obtain instructions, decode instructions, and execute instructions to complete its main functions.

[0034] In one embodiment, the detection pins 300 may be test points on the test board 100. The number of detection pins 300 may be the same as the number of signal channels of the flash memory interface (NAND Interface) of the memory 1000. For example, the number of signal channels of the flash memory interface inside the memory 1000 is 26. At this time, the number of detection pins 300 is also 26, which are classified as TP1~TP26 according to the serial numbers. Each detection pin 300 may be correspondingly connected to a signal channel of a flash memory interface of the memory 1000.

[0035] Among them, the signal channels of the flash memory interface of the memory 1000 refer to the protocols and signal channels used for communication with the NAND flash memory inside the memory 1000. The signal channels of the flash memory interface of the memory 1000 process data operations inside the memory 1000, such as reading, writing, and erasing. The signals and protocols of the flash memory interface may include an instruction bus, an address bus, a data bus, control signals, etc. The instruction bus can be used to send instructions to the NAND flash memory, such as read, write, erase, and other operations. The instruction bus usually includes a group of instruction lines for transmitting specific operation instructions. The address bus can be used to transmit the address information of the data so that the NAND flash memory knows the location of the data to be operated on. The data bus can be used for actual data transmission. The data bus usually works in cooperation with the instruction bus and the address bus to transmit the stored data. The control signals may include signals such as write enable (WE), read enable (RE), etc., for controlling the operation state of the NAND flash memory.

[0036] In one embodiment, a logic analyzer 400 (LA) is a tool for capturing and analyzing digital signals. The logic analyzer 400 can be electrically connected to multiple detection pins 300 on the test board 100, and thus can monitor and record the timing waveforms of the NAND Interface of the memory 1000. The logic analyzer 400 can not only capture the signals of the NAND Interface of the memory 1000, but also analyze the relationships between the signals, thereby analyzing the protocol behavior of the memory 1000. For example, analyzing read and write operations of the data bus, address selection of the address bus, and enabling and disabling of control signals, etc.

[0037] In one embodiment, the logic analyzer 400 can also be connected to the storage interface (eMMC Interface) between the memory 1000 and the processing module 200 to obtain various instructions and data sent by the processing module 200 to the memory 1000. The eMMC Interface is the protocol and signal channel for communication between the memory 1000 and the processing module 200.

[0038] Among them, the eMMC Interface can be responsible for processing external instructions and data requests to ensure correct data transmission between the processing module 200 and the memory 1000. It processes the instructions sent by the processing module 200 and converts them into operations that can be executed internally. The signals and protocols of the storage interface can include instruction (CMD) signals, clock (CLK) signals, data (DAT) lines, power supply and ground (VCC, GND) lines, etc. Instruction signals can be used to send instructions from the processing module 200 to the memory 1000, and these instructions can include read and write operations, initialization, status query, etc. The data lines can be used for the actual transmission of data. There are usually multiple data lines to support different data widths (such as 1-bit, 4-bit or 8-bit). The clock signal can be used to synchronize data transmission to ensure that instructions and data are transmitted in a predetermined timing sequence. The power supply and ground wires can supply power to the eMMC memory chips.

[0039] In one embodiment, the host 500 can be electrically connected to the logic analyzer 400 to obtain the data parsed by the logic analyzer 400.

[0040] In one embodiment, the power supply module 600 can supply power to the processing module 200.

[0041] In one embodiment, the transmission module 700 can be connected between the host 500 and the processing module 200. The host 500 can send various test instructions and corresponding test data to the processing module 200 through the transmission module 700. The processing module 200 can transmit the test information of the memory 1000 to the host 500 through the transmission module 700.

[0042] In one embodiment, the interface module 800 can be electrically connected to the processing module 200. The interface module 800 can be an interface for serial communication. The processing module 200 can perform serial communication with an external device through the interface module 800.

[0043] In one embodiment, the memory module 900 can be electrically connected to the processing module 200. The memory module 900 can be used to temporarily store the currently running programs and data. The processing module 200 can quickly access this data to improve computing performance.

[0044] In one embodiment, when testing the performance of a memory 1000 produced by a certain manufacturer, a test environment for the memory 1000 can be built through the cooperation of the processing module 200, the logic analyzer 400, and the host 500.

[0045] Specifically, the memory 1000 can be communicatively connected to the test board 100 first, so that the NAND Interface of the memory 1000 is correspondingly connected to the detection pins 300. After that, the power supply module 600 can supply power to the test board 100 to enable the processing module 200 and other modules on it to start working. Subsequently, the processing module 200 can also be communicatively connected to the host 500 through the transmission module 700. The host 500 can select a specific firmware or operating system image (SOC Image) and burn it onto the storage module (not shown in the figure) of the test board 100 through the transmission module 700. After the burning is completed, the test board 100 is powered on again to start the newly burned image.

[0046] Furthermore, after the memory 1000 is installed on the test board 100, multiple signal lines of the logic analyzer 400 can be sequentially communicatively connected to the detection pins 300 of the test board 100, and some other signal lines can also be communicatively connected to the eMMC Interface of the memory 1000 to build a test environment for the memory 1000.

[0047] In one embodiment, after the test environment for the memory 1000 is built, the host 500 can send different test instructions and corresponding test data to the memory 1000 through the processing module 200 to make the memory 1000 perform different operations. At this time, the logic analyzer 400 can obtain the test instructions and corresponding test data sent by the host 500 from the eMMC Interface of the memory 1000. The logic analyzer 400 can also obtain the response data of the internal NAND Interface of the memory 1000 when it performs different operations through the detection pins 300.

[0048] In one embodiment, the host 500 may send an initialization instruction to the memory 1000 through the processing module 200 to perform an initialization process on the memory 1000 and obtain the power-on duration of the memory 1000.

[0049] Specifically, after the power module 600 is powered on, the host 500 may send an initialization instruction to the memory 1000 through the processing module 200 to restart the memory 1000. The initialization instruction may include CMD 0, CMD 8, ACMD 41, CMD 2, CMD 3, etc.

[0050] Among them, the host 500 may send a CMD 0 instruction to the memory 1000 to reset the memory 1000 to the standby state (IDLE). The CMD 0 instruction can be used to reset the memory 1000 and clear any potential error states. Subsequently, the host 500 may send a CMD 8 instruction to query whether the memory 1000 supports the high-capacity SD card mode. If the memory 1000 supports it, this instruction will return a corresponding response. Subsequently, the host 500 may send an ACMD 41 instruction to start the initialization process of the memory 1000 and request the memory 1000 to enter the initialization state. The host 500 may also send a CMD 2 instruction to request the memory 1000 to send the content of its CID (Card Identification) register for identifying the memory 1000. Finally, the host 500 may also send a CMD 3 instruction to request the memory 1000 to set the corresponding relative address (RCA).

[0051] Further, after receiving CMD 0, the internal firmware of the memory 1000 parses this instruction and places the memory 1000 in the IDLE state. At the same time, the memory 1000 may continue to execute other received initialization instructions, such as setting the relative address, checking the functions supported by the device, etc. Subsequently, the firmware may instruct the main controller of the memory 1000 to perform an initialization operation on the NAND flash, including configuring the control register and data transfer path of the NAND flash, etc. At this time, it is necessary to ensure that the main controller sends correct instructions through the NAND Interface, such as reading the identification information of the NAND and setting the block erase parameters, etc.

[0052] Furthermore, during the initialization process of the memory 1000, the logic analyzer 400 can monitor the protocol data of the eMMC Interface and the NAND Interface, capture the instruction stream and responses during the initialization process, and record the sending time and receiving time of each instruction for subsequent analysis. The logic analyzer 400 can upload the captured data such as instructions and responses to the host 500. The host 500 can analyze this data to ensure that the instructions and responses of the memory 1000 comply with the protocol requirements. Meanwhile, during the monitoring process of the logic analyzer 400, it is necessary to capture the timestamp of each instruction executed by the memory 1000.

[0053] Furthermore, during the initialization process of the memory 1000, the system generates various logs, including the initialization timestamp, the system service startup time, etc. The host 500 can check these logs, extract the timestamps related to the initialization of the memory 1000 and the system startup, and calculate the boot time. The boot time is expressed as the duration from the power-on time of the memory 1000 to the time when the memory 1000 is fully started.

[0054] In one embodiment, the host 500 can send a multi-page read instruction (CMD 18) to the memory 1000 through the processing module 200 to read multiple data pages from the memory 1000 and calculate the multi-page read performance value of the memory 1000.

[0055] Specifically, after the memory 1000 completes the initialization operation, the host 500 can send a CMD 18 instruction to the memory 1000. The memory 1000 will read data from the specified address and continuously transmit the data until it receives a stop instruction (such as CMD12). The multi-page read instruction can include parameters such as the starting data block address and the number of data blocks to be read.

[0056] Furthermore, after receiving the CMD 18 instruction, the firmware inside the memory 1000 starts to parse the content of the multi-page read instruction. After the firmware confirms the legality of the multi-page read instruction and the correctness of the parameters, it forwards the multi-page read instruction to the main controller. The main controller performs the corresponding data read operation through the NAND Interface according to the instruction requirements.

[0057] Furthermore, the logic analyzer 400 can monitor the communication data on the eMMC Interface and the NAND Interface in real time. When the CMD 18 instruction is sent, the logic analyzer 400 will capture the instruction and its response data. Meanwhile, the logic analyzer 400 will also capture the instruction stream sent by the internal main controller to the NAND flash through the NAND Interface and the response data of the NAND flash.

[0058] Furthermore, the logic analyzer 400 can transmit the captured data to the host 500 for processing. The host 500 can parse the data to extract key timestamp and data volume information during the execution of the CMD 18 instruction. The host 500 can also calculate the multi-page read speed and instruction execution time of the memory 1000 based on the key timestamp and data volume information.

[0059] In one embodiment, the host 500 can send a single-page read instruction (CMD 17) to the memory 1000 through the processing module 200 to read a single data page from the memory 1000 and calculate the single-page read performance value of the memory 1000.

[0060] Specifically, after the memory 1000 completes the initialization operation, the host 500 can send the CMD 17 instruction to the memory 1000, and the memory 1000 will read data from the specified address. Among them, the single-page read instruction can include parameters such as the starting address of the data block and the size of the data block.

[0061] Furthermore, after the memory 1000 receives the CMD 17 instruction, its internal firmware begins to parse the content of the multi-page read instruction. After the firmware confirms the legality of the multi-page read instruction and the correctness of the parameters, it forwards the multi-page read instruction to the main controller. The main controller performs the corresponding data read operation through the NAND Interface according to the instruction requirements.

[0062] Furthermore, the logic analyzer 400 can monitor the communication data on the eMMC Interface and the NAND Interface in real time. When the CMD 17 instruction is issued, the logic analyzer 400 will capture the instruction and its response data. At the same time, the logic analyzer 400 will also capture the instruction stream sent by the internal main controller to the NAND flash through the NAND Interface and the response data of the NAND flash.

[0063] Furthermore, the logic analyzer 400 can transmit the captured data to the host 500 for processing. The host 500 can parse the data to extract key timestamp and data volume information during the execution of the CMD 17 instruction. The host 500 can also calculate the single-page read speed and instruction execution time of the memory 1000 based on the key timestamp and data volume information.

[0064] In one embodiment, the host 500 can send a single-page write instruction (CMD 24) and corresponding test data to the memory 1000 through the processing module 200 to write test data for a single data page into the NAND flash of the memory 1000 and calculate the single-page write performance value of the memory 1000.

[0065] Specifically, after the memory 1000 completes the initialization operation, the host 500 can send a CMD 24 instruction and corresponding test data to the memory 1000, and the memory 1000 will write the test data to the specified storage location in the NAND flash. Among them, the single-page write instruction can include parameters such as the starting address of the written data and the size of the data block.

[0066] Further, after the memory 1000 receives the CMD 24 instruction, the firmware inside it starts to parse the content of the single-page write instruction to confirm its legality and parameters. Subsequently, the firmware can forward the CMD 24 instruction and the test data to the main controller. According to the instruction requirements, the main controller issues a write command through the NAND Interface to write the data to the NAND flash. After the writing is completed, the main controller can update the internal cache status of the memory 1000 and feedback the result to the host 500.

[0067] Furthermore, the logic analyzer 400 can monitor the communication data on the eMMC Interface and the NAND Interface in real time. When the CMD 24 instruction is issued, the logic analyzer 400 will capture the instruction and its response data. At the same time, the logic analyzer 400 will also record the transmission situation of the CMD 24 instruction and its data on the eMMC Interface, as well as the instruction stream sent by the main controller through the NAND Interface and the response of the NAND flash.

[0068] Furthermore, the logic analyzer 400 can transmit the captured data to the host 500 for processing. The host 500 can parse the data to extract the key timestamp and data volume information during the execution of the CMD 24 instruction. The host 500 can also calculate the single-page write speed and instruction execution time of the memory 1000 based on the key timestamp and data volume information.

[0069] In one embodiment, the host 500 can issue a single-page write instruction (CMD 25) and corresponding test data to the memory 1000 through the processing module 200, so as to write the test data of multiple data pages to the NAND flash of the memory 1000 and calculate the multi-page write performance value of the memory 1000.

[0070] Specifically, after the memory 1000 completes the initialization operation, the host 500 can send a CMD 25 instruction and corresponding test data to the memory 1000, and the memory 1000 will write the test data to the specified storage location in the NAND flash. Among them, the multi-page write instruction can include parameters such as the starting address of the written data and the size of the data block.

[0071] Further, after the memory 1000 receives the CMD 25 instruction, the firmware inside it starts to parse the content of the single-page write instruction to confirm its legality and parameters. Subsequently, the firmware can forward the CMD 25 instruction and the test data to the main controller. According to the instruction requirements, the main controller issues a write command through the NAND Interface to write the data into the NAND flash. After the writing is completed, the main controller can update the internal cache status of the memory 1000 and feedback the result to the host 500.

[0072] Furthermore, the logic analyzer 400 can monitor the communication data on the eMMC Interface and the NAND Interface in real time. When the CMD 25 instruction is issued, the logic analyzer 400 will capture the instruction and its response data. At the same time, the logic analyzer 400 will also record the transmission situation of the CMD 25 instruction and its data on the eMMC Interface, as well as the instruction stream sent by the main controller through the NAND Interface and the response of the NAND flash.

[0073] Furthermore, the logic analyzer 400 can transmit the captured data to the host 500 for processing. The host 500 can parse the data to extract the key timestamps and data volume information during the execution of the CMD 24 instruction. The host 500 can also calculate the multi-page write speed and instruction execution time of the memory 1000 based on the key timestamps and data volume information.

[0074] In one embodiment, the host 500 can sequentially send an erase start address instruction (CMD 35), an erase end address instruction (CMD 36), a partial erase instruction (CMD 37), and a full erase instruction (CMD 38) to the memory 1000 through the processing module 200 to perform partial erase or full erase on the data in the data blocks of the NAND flash of the memory 1000, and calculate the data block erase performance value and full erase performance value of the memory 1000.

[0075] Specifically, after the memory 1000 completes the initialization operation, the host 500 can send the CMD 35 instruction, the CMD 36 instruction, and the CMD 37 instruction to the memory 1000. Among them, the CMD 35 instruction may include the start address of the data block to be erased. The CMD 36 instruction may include the end address of the data block to be erased.

[0076] Further, after the memory 1000 receives the CMD 35 instruction, CMD 36 instruction, and CMD 37 instruction, the firmware inside it starts to parse the content of the above instructions to confirm their legality and parameters. Subsequently, the firmware can forward the corresponding erase command to the main controller. According to the requirements of the erase command, the main controller issues an instruction stream through the NAND Interface to erase the data in the specified storage area.

[0077] Furthermore, the logic analyzer 400 can monitor the communication data on the eMMC Interface and NAND Interface in real time. The logic analyzer 400 can record the sending of the CMD 35 instruction, CMD 36 instruction, and CMD 37 instruction and their execution status on the eMMC Interface and NAND Interface.

[0078] Furthermore, the logic analyzer 400 can transmit the captured data to the host 500 for processing. The host 500 can parse the data to extract the key timestamps and data volume information during the execution of the CMD 35 instruction, CMD 36 instruction, and CMD 37 instruction. Among them, the data volume information can be determined according to the start address and end address of the erased data block. The host 500 can also calculate the data block erase performance value and instruction execution time of the memory 1000 based on the key timestamps and data volume information.

[0079] Furthermore, after the memory 1000 completes the initialization operation, the host 500 can also send the CMD 38 instruction to the memory 1000. Among them, the CMD 38 instruction can perform a full erase operation (full disk erase). After the memory 1000 receives the CMD 38 instruction, the firmware inside it starts to parse the content of the above instructions to confirm their legality and parameters. Subsequently, the firmware can forward the corresponding erase command to the main controller. According to the requirements of the erase command, the main controller issues an instruction stream through the NAND Interface to perform a full disk erase on the data in the storage area.

[0080] Furthermore, the logic analyzer 400 can monitor the communication data on the eMMC Interface and NAND Interface in real time. The logic analyzer 400 can record the sending of the CMD 38 instruction and its execution status on the eMMC Interface and NAND Interface.

[0081] Furthermore, the logic analyzer 400 can transmit the captured data to the host 500 for processing. The host 500 can parse the data to extract the key timestamps and data volume information during the execution of the CMD 38 instruction. Among them, the data volume information can be all the data stored in the memory 1000. The host 500 can also calculate the full disk erasure performance value and instruction execution time of the memory 1000 based on the key timestamps and data volume information.

[0082] In one embodiment, after the memory 1000 executes multi-page read instructions, single-page read instructions, multi-page write instructions, single-page write instructions, partial erasure instructions, and full disk erasure instructions, the host 500 can also be used to count the power consumption data of the memory 1000 during the execution of the above operations.

[0083] In one embodiment, the host 500 can simultaneously send queue instructions to the memory 1000 through the processing module 200 to monitor the instruction execution order of the memory 1000. Among them, the queue instructions can include read queue instructions (CMD Q44), write queue instructions (CMD Q45), commit queue instructions (CMD Q46), build request instructions (CMD Q47), and queue management instructions (CMD Q48). The CMD Q44 instruction can be used to build a series of read queue instructions to define the target address and data of the operation. The CMD Q45 instruction can be used to build a series of write queue instructions to define the target address and data of the operation. The CMD Q46 instruction can be used to build a commit queue instruction to inform the memory 1000 that it can start processing the commands in the previous queue. The CMD Q47 instruction can be used to build an instruction to request the queue processing result. The CMD Q48 instruction can be used to build an instruction to manage the queue, such as clearing or modifying the commands in the queue.

[0084] Specifically, after the memory 1000 completes the initialization operation, the host 500 can sequentially send the CMD Q44 instruction, CMD Q45 instruction, CMD Q46 instruction, CMD Q47 instruction, and CMD Q48 instruction to the memory 1000 through the eMMC Interface. Subsequently, the firmware inside the memory 1000 can receive and parse the queue instructions from the host 500. According to the received read and write tasks, the firmware guides the main controller to optimize the sorting of the tasks to maximize the operation efficiency. The main controller performs the corresponding read or write operations through the NAND Interface based on the optimized order.

[0085] Furthermore, the logic analyzer 400 can monitor in real time the data such as the instruction streams and responses on the eMMC Interface and the NAND Interface. The logic analyzer 400 can record the sending of CMD Q44 instructions, CMD Q45 instructions, CMD Q46 instructions, CMD Q47 instructions, CMD Q48 instructions and their execution statuses on the eMMC Interface and the NAND Interface.

[0086] Furthermore, the logic analyzer 400 can transmit the captured data to the host 500 for processing. The host 500 can parse the data to obtain the instruction execution order and the corresponding execution duration when the memory 1000 executes the queue instructions.

[0087] Furthermore, when multiple memories 1000 execute the queue instructions, the host 500 can obtain the memory 1000 with the shortest execution duration. At this time, it can be considered that the instruction execution order inside this memory 1000 is the optimal order. Subsequently, the host 500 can compare and analyze the instruction execution order of this memory 1000 with that of other memories 1000 to optimize the firmware.

[0088] In one embodiment, after the above tests, the host 500 has completed the performance test of one memory 1000. Subsequently, the host 500 can also test other different memories 1000 to obtain the performance values of multiple memories 1000. Subsequently, the differences in the performance values of multiple memories 1000 can be analyzed to analyze the advantages and disadvantages of the behavioral differences of the firmware of multiple memories 1000, so as to specifically optimize the behavior of the firmware and improve the performance of the memory 1000 to a certain extent.

[0089] In one embodiment, the comparative analysis of two memories 1000 is taken as an example for illustration. The two memories 1000 can be classified as competitor 1 and competitor 2. When the host 500 issues the same instructions to competitor 1 and competitor 2, the execution statuses on the NAND Interface of competitor 1 and competitor 2 are also different.

[0090] In one embodiment, referring to Table 1, Table 1 shows multiple execution states on the NAND Interface of Competitor 1 and Competitor 2 when performing the initialization operation. As can be seen from Table 1, both Competitor 1 and Competitor 2 will execute certain unknown instructions (Unknown CMD) internally when performing the initialization operation. Among them, the unknown instruction can be expressed as that the instruction is not in the JEDEC standard (such as JESD84 - B51). When Competitor 1 and Competitor 2 perform the initialization operation, the logic analyzer 400 can monitor the NAND Interface in real time, and the host 500 can monitor the instruction stream executed inside the memory 1000 through the logic analyzer 400 to find the direction for optimizing the firmware. The specific optimization directions can include reducing unknown commands, adopting a more efficient storage mode, and optimizing the data reading method.

[0091] Table 1: Multiple execution states of Competitor 1 and Competitor 2 when performing the initialization operation.

[0092]

[0093] Furthermore, when the memory 1000 performs the initialization operation, it may internally process some unknown instructions (Unknown CMD), which will increase unnecessary time consumption. When the memory 1000 is working, in some cases, Unknown CMD may be generated due to incorrect instruction parsing or transmission problems. These unknown instructions cannot be correctly processed by the memory 1000, resulting in time waste. At this time, the firmware or instruction parsing mechanism can be improved to ensure that only valid instructions are sent and the generation of invalid instructions is avoided, thereby reducing the boot time.

[0094] Furthermore, compared with the XLC (multi - level cell) mode, the SLC (single - level cell) mode has a faster data reading speed and a lower write latency. XLC usually refers to multi - level cell technologies such as MLC and TLC. Although these technologies have more advantages in terms of unit storage capacity, they are inferior to SLC in terms of speed and reliability. In scenarios that require high - performance data reading (such as during the initialization operation), the firmware can be improved to preferentially select the SLC mode to improve the reading speed and reduce the boot time.

[0095] Furthermore, sequential read is usually more efficient than random read because sequential read reduces the addressing and data scheduling overhead inside the memory. At this time, the data access mode of the firmware can be optimized so that more data reading operations are performed in the sequential read manner, improving the reading speed and thus shortening the boot time.

[0096] In one embodiment, please refer to Table 2, which shows multiple execution states on the NANDInterface of Competitor 1 and Competitor 2 when performing a write operation. As can be seen from Table 2, Competitor 1 and Competitor 2 will execute certain Unknown CMDs internally when performing a write operation. When Competitor 1 and Competitor 2 perform a write operation, the logic analyzer 400 can monitor the NANDInterface in real time, and the host 500 can monitor the instruction stream executed inside the memory 1000 through the logic analyzer 400 to find the optimization direction of the firmware. The optimization direction may include reducing unknown instructions, optimizing the write process to reduce read and erase operations, adopting a more efficient storage mode, and optimizing the data write strategy to reduce the number of programming times to increase the write speed.

[0097] Table 2: Multiple execution states of competitor 1 and competitor 2 when performing write operations.

[0098]

[0099] Furthermore, when writing data, if an unknown instruction is parsed inside the memory 1000, processing delays or even erroneous operations may occur. Through strict management and optimization of the instruction set, it is ensured that only valid instructions are sent during the writing process, and the occurrence of invalid or unknown instructions is avoided, thereby reducing the time consumption caused thereby. At this point, the firmware or instruction parsing mechanism can be improved to ensure that only valid instructions are sent and the generation of invalid instructions is avoided.

[0100] Furthermore, when writing data, the memory 1000 needs to reduce the action of reading data. For example, in some memory architectures, the write operation may involve the process of first reading data and then writing it. This read operation will increase additional time overhead. By optimizing the firmware write process, unnecessary read operations can be minimized or avoided to reduce the overall time consumption. This may require adjusting the firmware to ensure that the write operation is performed directly without involving redundant read processes.

[0101] Furthermore, when writing data, the memory 1000 needs to reduce the erase action. For example, in NAND Flash, an erase operation (Erase) is usually required before writing new data, and this step consumes a lot of time. By improving the management strategy of the firmware, such as erasing free blocks in advance or using a more efficient erase algorithm, the erase operation that needs to be performed immediately during the writing process is reduced, thereby improving the writing speed.

[0102] Further, when writing data, the memory 1000 preferably performs the write operation in a single-level storage mode. Among them, the single-level cell (SLC) mode stores only one bit of data in each storage cell. Compared with the multi-level storage (such as MLC, TLC) mode, it has a faster write speed and higher reliability. In scenarios where high write speed is required, the firmware can preferentially select the SLC mode for writing to maximize the write speed and performance.

[0103] Further, when writing data, the memory 1000 needs to reduce the number of programming operations and concentrate on writing large chunks of data instead of scattered small data. Specifically, frequent small data writes will cause the memory to perform a large number of programming operations, which not only affects the write efficiency but also may increase the wear of the memory. The firmware can batch process the data, merge multiple small data blocks into a large data block for centralized writing, reduce the number of programming operations, thereby improving the write performance and extending the memory life.

[0104] In one embodiment, when optimizing the read speed of the memory 1000, the host 500 can monitor the instruction stream for the read operation inside the memory 1000 through the logic analyzer 400 to find the optimization direction of the firmware. The optimization direction can include reducing invalid commands, adopting a more efficient storage mode, optimizing the data access method, using cache and prefetch mechanisms, etc., to improve the read speed.

[0105] Further, when reading data, if an unknown instruction is parsed inside the memory 1000, it may cause processing delays and even require the command to be resent, which will reduce the read efficiency. By optimizing the firmware and the command parsing process, ensure that only valid and correct commands are used in the read operation to avoid generating unknown instructions, thereby reducing unnecessary time consumption.

[0106] Further, when reading critical data or data that requires high performance, the SLC mode is preferentially used to improve the read speed.

[0107] Further, through the optimization at the firmware level, the data access mode is preferably optimized to sequential reading as much as possible. For example, merge scattered read requests and perform reading in the order of the physical storage layout, which can maximize the read bandwidth of the memory.

[0108] Further, in some cases, the data can be pre-loaded into the cache to reduce the read latency. Utilize the prefetch and cache mechanisms of the firmware to pre-load the data that is expected to be read into the high-speed cache of the memory 1000, so as to quickly access it when needed. The cache mechanism can significantly reduce the data transfer latency between the memory 1000 and the host 500.

[0109] Furthermore, some memories may perform redundant internal operations (such as redundant checks and internal data verification) during the reading process, and these operations will increase the reading latency. By optimizing the internal operation process, reducing or combining unnecessary internal reading operations, the overall reading speed can be improved. For example, the reading speed can be accelerated by optimizing the calculation method of the error correction code (ECC) or reducing the frequency of redundant data verification.

[0110] It can be seen that in the above solution, by monitoring the execution status of the NAND Interface of different memories, the behavioral differences of the firmware can be analyzed and identified, and the differential behaviors of the firmware of different storage manufacturers can be quickly located and analyzed. By analyzing the differential behaviors of the firmware, the deficiencies in the firmware can be identified, the overall system performance can be improved, and optimization measures can be quickly taken to improve the overall system performance.

[0111] Please refer to Figure 2 , the present invention also provides a method for analyzing the performance of a memory, which can be applied to the above performance analysis system to analyze the performance of different memories 1000. The steps of this performance analysis method correspond one by one to the above performance analysis system. The performance analysis method may include the following steps:

[0112] Step S10: Build a test environment for the memory;

[0113] Step S20: Send different test instructions to the memory to make the memory perform different operations;

[0114] Step S30: Obtain the instruction stream and response data sent by the firmware of the memory to the flash interface when the memory performs different operations;

[0115] Step S40: Upload the instruction stream and response data and perform performance analysis on them.

[0116] The embodiments of the present invention disclosed above are only used to help explain the present invention. The embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principle and practical application of the present invention, so that those skilled in the art in the relevant technical field can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A memory performance analysis system, characterized in that: include: A test board, on which a memory is connected, wherein a plurality of detection pins are set on the test board, and the detection pins are communicatively connected with the test pins of the memory; A logic analyzer is communicatively connected to the flash memory interface of the memory through the detection pin; A processing module is disposed on the test board and is communicatively connected to the storage interface of the memory; as well as A host, which is in communication with the processing module and the logic analyzer, and is used to send different test instructions to the memory through the processing module so that the memory performs different operations; The logic analyzer is used to obtain the instruction stream and response data sent by the firmware of the memory to the flash memory interface when the memory performs different operations, upload the instruction stream and response data to the host, and perform performance analysis on different memories; The host is also used to send initialization instructions to the multiple memories respectively, so as to obtain the instruction streams of the internal firmware and the corresponding boot time of the multiple memories when the initialization operations are performed; The host is also used to compare and analyze the instruction stream of the memory with the shortest boot time with the instruction streams of other memories, so as to optimize the firmware from the direction of reducing unknown commands, adopting a more efficient storage mode and optimizing the data reading method; wherein, the unknown command refers to an instruction that is not in the JEDEC standard; the more efficient storage mode refers to performing write operations in a single-level storage mode; the optimized data reading method refers to giving priority to performing read operations in a single-level storage mode, and the read operation is performed in a sequential read manner.

2. The memory performance analysis system according to claim 1, characterized in that: The host is also used to send multi-page read instructions and single-page read instructions to multiple memories respectively, so as to obtain the instruction stream and corresponding read performance value of the internal firmware of the multiple memories when performing multi-page read operations and single-page read operations.

3. The memory performance analysis system according to claim 2, characterized in that: The host is also used to compare and analyze the instruction stream of the memory with the best read performance value with the instruction stream of other memories, so as to optimize the firmware from the direction of reducing unknown commands, adopting a more efficient storage mode, optimizing data access methods, and utilizing cache and pre-fetch mechanisms.

4. The memory performance analysis system according to claim 1, characterized in that: The host is also used to send multi-page write instructions and single-page write instructions to multiple memories respectively, so as to obtain the instruction stream and corresponding write performance value of the internal firmware of the multiple memories when performing multi-page write operations and single-page write operations.

5. The memory performance analysis system according to claim 4, characterized in that: The host is also used to compare and analyze the instruction stream of the memory with the best write performance value with the instruction stream of other memories, so as to optimize the firmware from the direction of reducing unknown commands, optimizing the write process to reduce read and erase operations, adopting a more efficient storage mode, and optimizing the data write strategy to reduce the number of programming times.

6. The memory performance analysis system according to claim 1, characterized in that: The host is also used to send partial erase instructions and full erase instructions to multiple memories respectively to obtain the instruction streams and corresponding erase performance values ​​of the internal firmware of the multiple memories when executing partial erase operations and full erase operations. The host is also used to compare and analyze the instruction stream of the memory with the best erase performance value with the instruction streams of other memories to optimize the firmware.

7. The memory performance analysis system according to claim 1, characterized in that: The host is also used to send queue instructions to multiple memories respectively to obtain the instruction execution order and execution time of the multiple memories when performing queue operations. The host is also used to compare and analyze the instruction execution order of the memory with the shortest execution time with the instruction execution order of other memories to optimize the firmware.

8. A memory performance analysis method, characterized in that: include: Build a test environment for storage; Sending different test instructions to the memory so that the memory performs different operations; Acquire the instruction stream and response data sent by the firmware of the memory to the flash memory interface when the memory performs different operations; Uploading the instruction stream and response data, and performing performance analysis on them; Sending initialization instructions to the multiple memories respectively to obtain the instruction streams of the internal firmware and the corresponding boot durations of the multiple memories when the multiple memories perform initialization operations; The instruction stream of the memory with the shortest boot time is compared and analyzed with the instruction streams of other memories to optimize the firmware from the direction of reducing unknown commands, adopting a more efficient storage mode and optimizing the data reading method; wherein, the unknown command refers to the instruction that is not in the JEDEC standard; the more efficient storage mode refers to performing write operations in a single-level storage mode; the optimized data reading method refers to giving priority to performing read operations in a single-level storage mode, and the read operation is performed in a sequential read manner.

Citation Information

Patent Citations

  • System and method for predicting and improving boot-up sequence

    CN105051684A

  • Test system and test method of memory

    CN119091952A

  • System, Method and Computer-Readable Medium for Dynamically Configuring an Operational Mode in a Storage Controller

    US20150286438A1