Memory chip and memory system

By adopting a parallel data transmission scheme of multiple memory blocks and alignment circuits in the storage system, the problems of low data transmission rate and high power consumption between the dynamic random access memory chip and the logic circuit are solved, and more efficient data transmission and cost-reducing effect are achieved.

CN120255791APending Publication Date: 2025-07-04ETRON TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510015254.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-13
Filing Date
2025-01-06
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In existing storage systems, the data transmission rate between the dynamic random access memory chip and the logic circuit is low, and the use of a faster data rate or a wide data bus will increase power consumption and clock delay, resulting in a storage wall effect and affecting system efficiency.

Method used

Multiple memory blocks and alignment circuits are used to transmit data in parallel, cancel the parallel serial circuit and serial parallel circuit, and use the direct send/receive bus and alignment circuit for data transmission to reduce clock delay and power consumption.

Benefits of technology

It improves the data transmission rate of the memory chip, reduces power consumption and cost, reduces the memory wall problem between the memory chip and logic circuit, and enhances the bandwidth of the input/output bus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255791A_ABST
    Figure CN120255791A_ABST
Patent Text Reader

Abstract

The invention discloses a memory chip. The memory chip comprises a plurality of memory blocks, an input / output bus and a plurality of alignment circuits. Each storage block outputs or receives a data group in parallel. The plurality of alignment circuits are respectively corresponding to the plurality of memory blocks. Data sets of a memory block are transmitted to a corresponding alignment circuit, and then the corresponding alignment circuit simultaneously transmits the data sets to the input / output bus in parallel or transmits the data sets from the input / output bus to the corresponding alignment circuit. The corresponding alignment circuits simultaneously transmit the data sets to the memory blocks in parallel. There is no parallel-serial circuit and serial-parallel circuit between the input / output bus and each memory block. According to the invention, the power consumption, the access delay and the cost of the direct interface wide bus storage chip can be reduced, and the bandwidth and the data transmission rate of the input / output bus of the direct interface wide bus storage chip can be increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a memory chip and a memory system, and more particularly to a memory chip and a memory system that can simultaneously transfer wide-bus data between a logic circuit and a memory chip in parallel to reduce the power consumption, access latency, and cost of the memory chip, and increase the bandwidth and data transfer rate of the input / output bus of the memory chip. Background Art

[0002] Currently, memory systems used in high-performance computing or artificial intelligence systems typically include dynamic random access memory chips and logic circuits. Due to the stack structure of the dynamic random access memory chips, the size of the dynamic random access memory chips cannot keep up with the size of the logic circuits. Therefore, the memory-wall effect occurs, resulting in a reduction in the data transfer rate between the logic circuit and the dynamic random access memory chips. To overcome the memory-wall effect, the prior art typically uses a faster data rate (e.g., from double data rate three (DDR3) to double data rate fourth (DDR4) or double data rate fifth (DDR5)) to transfer data between the dynamic random access memory chips and the logic circuit, or uses the wide data bus of the logic circuit and the wide data bus of the dynamic random access memory chips (e.g., High Bandwidth Memory (HBM)) to transfer data between the dynamic random access memory chips and the logic circuit. However, the faster data rate has some disadvantages (e.g., more expensive testers, smaller noise margin, etc.), and the wide data bus of the logic circuit and the wide data bus of the dynamic random access memory chips also have some disadvantages (e.g., higher power, larger die area, expensive Through-Silicon Via process, etc.). Moreover, whether it is the faster data rate of the aforementioned dynamic random access memory chips or the wide data bus of the dynamic random access memory chips, serial-to-parallel circuits and parallel-to-serial circuits are required, and the serial-to-parallel circuits and the parallel-to-serial circuits will increase the clock delay and power consumption.

[0003] Please refer to Figure 1 , Figure 1 which is a schematic diagram illustrating a memory system 10 disclosed in the prior art. As Figure 1 shown, the memory system 10 includes a memory 20 and a logic circuit 30, where the memory 20 is a dynamic random access memory. As Figure 1As shown, the memory 20 includes a cell array 21, a serial-to-parallel circuit 22, and a parallel-to-serial circuit 23; the logic circuit 30 includes a physical layer 31 and a controller 32, and the physical layer 31 further includes a parallel-to-serial circuit 312 and a serial-to-parallel circuit 314. In addition, the logic circuit 30 further includes other functional circuits (not shown in Figure 1 ), where the other functional circuits may include a central processing unit, a digital signal processor, a peripheral interface, etc. As Figure 1 shown, when the logic circuit 30 writes data into the memory 20, the serial-to-parallel circuit 314 can receive data (e.g., N-bit data) in parallel from the controller 32, convert the N-bit data into several groups of Q-bit data, where Q is less than N, and transmit the several groups of Q-bit data to the parallel-to-serial circuit 23; the parallel-to-serial circuit 23 can receive the several groups of Q-bit data from the serial-to-parallel circuit 314, convert the several groups of Q-bit data into the N-bit data, and transmit the N-bit data in parallel to the cell array 21. In addition, when the logic circuit 30 reads data from the controller 20, the serial-to-parallel circuit 22 can receive data (e.g., the N-bit data) in parallel from the cell array 21, convert the N-bit data into the several groups of Q-bit data, and transmit the several groups of Q-bit data to the parallel-to-serial circuit 312; the parallel-to-serial circuit 312 can receive the several groups of Q-bit data from the serial-to-parallel circuit 22, convert the several groups of Q-bit data into the N-bit data, and transmit the N-bit data in parallel to the controller 32.

[0004] Please refer to Figure 2A 、 2B . Figure 2A 、 2B are timing diagrams of the logic circuit 30 writing data into the memory 20. As Figure 2A shown, taking the logic circuit 30 writing 8-bit data D0 - D7 into the memory 20 as an example, when the logic circuit 30 writes 8-bit data D0 - D7 into the memory 20, the register of the serial-to-parallel circuit 314 (not shown in Figure 1 ) can serially transmit the 8-bit parallel data D0 - D7 to the parallel-to-serial circuit 23 using three signals clk1, clk2, clk3. For example, when clk1 = 1, clk2 = 1, clk3 = 1, the serial-to-parallel circuit 314 transmits data D0 to the parallel-to-serial circuit 23, and when clk1 = 1, clk2 = 1, clk3 = 0, the serial-to-parallel circuit 314 transmits data D1 to the parallel-to-serial circuit 23, and so on. Therefore, the serial-to-parallel circuit 314 starts transmitting data D0 at time T0 and finally transmits data D7 at time T4.

[0005] As Figure 2B shown, similarly, the register of the parallel-to-serial circuit 23 (not shown in Figure 1 ) can also serially process the 8-bit serial data D0 - D7 from the serial-to-parallel circuit 314 using the clock signals clk1, clk2, clk3. AsFigure 2B As shown, when clk1 = 1, clk2 = 1, and clk3 = 1, the serial-to-parallel circuit 23 receives the data D0 from the parallel-to-serial circuit 314. When clk1 = 1, clk2 = 1, and clk3 = 0, the serial-to-parallel circuit 23 receives the data D1 from the parallel-to-serial circuit 314, and so on. Therefore, the serial-to-parallel circuit 23 starts receiving the data D0 at time T0 and finally receives the data D7 at time T4. Among them, between time T0 and time T4, there is a delay of 4 clocks for the clock signal clk3. That is to say, after waiting for 4 clock delays, the serial-to-parallel circuit 23 will start to parallel-transmit the 8-bit data D0 - D7 to the cell array 21.

[0006] Although the prior art can reduce the 4-clock delay (for example, reduce it to 3.5 clock delays) by optimizing the storage system 10, the serial-to-parallel conversion program executed by the above serial-to-parallel circuit 23 and the serial-to-parallel conversion program executed by the above parallel-to-serial circuit 314 will require additional power, transmission delay, and die areas, resulting in low efficiency of the storage system 10. Therefore, how to reduce power consumption, transmission delay, and die areas is an important problem that designers of storage systems need to solve. Summary of the Invention

[0007] An embodiment of the present invention provides a storage chip. The storage chip includes a plurality of storage blocks, an input / output bus, and a plurality of alignment circuits. Each storage block outputs or receives a data group in parallel. The plurality of alignment circuits respectively correspond to the plurality of storage blocks. The data group of a storage block is transmitted to a corresponding alignment circuit, and then the corresponding alignment circuit simultaneously transmits the data group in parallel to the input / output bus, or transmits the data group from the input / output bus to the corresponding alignment circuit, and then the corresponding alignment circuit simultaneously transmits the data group in parallel to the storage block. There is no parallel-to-serial circuit and serial-to-parallel circuit between the input / output bus and each storage block.

[0008] In an embodiment of the present invention, each alignment circuit includes a plurality of first transceivers, the plurality of first transceivers are connected to the input / output bus through a direct send / receive bus, and the width of the input / output bus is equal to the width of the data group received or output by each storage block.

[0009] In an embodiment of the present invention, the data groups of the plurality of storage blocks are output to the input / output bus in a predetermined order.

[0010] In an embodiment of the present invention, data groups of each memory block share a common row address, and column addresses of data groups of each memory block are different from each other.

[0011] In an embodiment of the present invention, a column address of a data group of each memory block is generated inside the memory chip or received from a memory controller outside the memory chip.

[0012] In an embodiment of the present invention, multiple data groups of the multiple memory blocks are output to the input / output bus within one bit switching cycle. The bit switching cycle includes multiple phases, and a data group of each memory block is output to the input / output bus at a corresponding phase of the bit switching cycle.

[0013] In an embodiment of the present invention, the multiple phases of the bit switching cycle include 2 N phases, a period of a clock signal of the memory chip is equal to the bit switching cycle divided by 2 N-1 , and N is an integer not less than 1.

[0014] In an embodiment of the present invention, the number of the multiple phases of the bit switching cycle is set in a mode register in the memory chip.

[0015] In an embodiment of the present invention, the memory chip further includes multiple data lines and multiple groups of sense amplifiers. The multiple groups of sense amplifiers are coupled to the multiple data lines, wherein a memory block corresponds to a group of sense amplifiers, and the group of sense amplifiers is disposed between the memory block and the corresponding alignment circuit.

[0016] In an embodiment of the present invention, the multiple memory blocks include a first memory block and a second memory block; the multiple groups of sense amplifiers include a first group of sense amplifiers coupled to the multiple data lines and a second group of sense amplifiers coupled to the multiple data lines; the first group of sense amplifiers corresponds to the first memory block, and a first data group is simultaneously and parallelly transmitted between the first group of sense amplifiers and the input / output bus through an alignment circuit corresponding to the first memory block; the second group of sense amplifiers corresponds to the second memory block, and a second data group is simultaneously and parallelly transmitted between the second group of sense amplifiers and the input / output bus through an alignment circuit corresponding to the second memory block; a width of the input / output bus is equal to a width of the first data group and a width of the second data group.

[0017] In an embodiment of the present invention, the width of a double data rate physical layer interface (DDR PHY Interface, Dfi) bus of a physical layer within a logic circuit is equal to the sum of the widths of the first data group and the second data group, wherein the double data rate physical layer interface bus is coupled between a controller within the logic circuit and the physical layer, the controller is further coupled to an Advanced eXtensible Interface bus outside the logic circuit, and the logic circuit is coupled to the input / output bus of the memory chip.

[0018] In an embodiment of the present invention, the memory chip further includes a plurality of bit lines, a third group of sense amplifiers, and a fourth group of sense amplifiers. The third group of sense amplifiers is coupled to the plurality of bit lines and disposed between the first storage block and the first group of sense amplifiers. The fourth group of sense amplifiers is coupled to the plurality of bit lines and disposed between the second storage block and the second group of sense amplifiers.

[0019] In an embodiment of the present invention, the memory chip further includes a first bit switch group and a second bit switch group. The first bit switch group is located between the first group of sense amplifiers and the third group of sense amplifiers. The second bit switch group is located between the second group of sense amplifiers and the fourth group of sense amplifiers.

[0020] An embodiment of the present invention provides a memory system. The memory system includes a first memory chip and a logic circuit. The first memory chip includes a plurality of first storage blocks, an input / output bus, and a plurality of alignment circuits. Each first storage block outputs or receives a data group in parallel. The plurality of alignment circuits respectively correspond to the plurality of first storage blocks. A data group of a first storage block is transmitted to a corresponding alignment circuit, and then the corresponding alignment circuit simultaneously transmits the data group in parallel to the input / output bus, or transmits the data group from the input / output bus to the corresponding alignment circuit, and then the corresponding alignment circuit simultaneously transmits the data group in parallel to the first storage block. There is no serial-to-parallel circuit and parallel-to-serial circuit between the input / output bus and each first storage block. The logic circuit has a physical layer, wherein the logic circuit is located outside the first memory chip and electrically connected to the input / output bus of the first memory chip, and there is a serial-to-parallel circuit and a parallel-to-serial circuit within the physical layer.

[0021] In an embodiment of the present invention, the physical layer further includes a plurality of second transceivers, and the plurality of second transceivers are electrically connected to the serial-to-parallel circuit and the parallel-to-serial circuit.

[0022] In an embodiment of the present invention, the plurality of first storage blocks includes 2 N first storage blocks, N is an integer not less than 1, the serial-to-parallel circuit is a 2 N :1 serial-to-parallel circuit, and the parallel-to-serial circuit is a 1:2 N parallel-to-serial circuit.

[0023] In an embodiment of the present invention, the width of a double data rate physical layer interface bus of the physical layer is equal to the sum of the widths of the data groups of each first storage block in the first storage chip, and the double data rate physical layer interface bus is coupled between a controller in the logic circuit and the physical layer.

[0024] In an embodiment of the present invention, each alignment circuit in the first storage chip includes a plurality of first transceivers, the plurality of first transceivers are connected to the input / output bus through a direct transmit / receive bus, and the width of the input / output bus is equal to the width of the data group received or output by each first storage block.

[0025] In an embodiment of the present invention, the plurality of data groups of the plurality of first storage blocks are output to the input / output bus in a predetermined order.

[0026] In an embodiment of the present invention, the data groups of each first storage block share a common row address, and the column addresses of the data groups of each first storage block are different from each other.

[0027] In an embodiment of the present invention, the column address of the data group of each first storage block is generated inside the first storage chip or received from a storage controller outside the first storage chip.

[0028] In an embodiment of the present invention, the plurality of data groups of the plurality of first storage blocks are output to the input / output bus within one bit switching cycle, the bit switching cycle includes a plurality of stages, and the data group of each first storage block is output to the input / output bus at a corresponding stage of the bit switching cycle.

[0029] In an embodiment of the present invention, the plurality of stages of the bit switching cycle includes 2 N stages, the period of a clock signal of the first storage chip is equal to the bit switching cycle divided by 2 N-1 , and N is an integer not less than 1.

[0030] In an embodiment of the present invention, the number of the plurality of stages of the bit switching cycle is set in a mode register in the first storage chip.

[0031] In an embodiment of the present invention, the storage system further includes a second storage chip. The second storage chip includes a plurality of second storage blocks, an input / output bus, and a plurality of alignment circuits. Each second storage block outputs or receives a data group in parallel. The plurality of alignment circuits respectively correspond to the plurality of second storage blocks. A data group of a second storage block in the second storage chip is transmitted to a corresponding alignment circuit, and then the corresponding alignment circuit simultaneously transmits the data group in parallel to the input / output bus, or transmits the data group of the second storage block in the second storage chip from the input / output bus to the corresponding alignment circuit, and then the corresponding alignment circuit simultaneously transmits the data group in parallel to the second storage block in the second storage chip. There is no serial-to-parallel circuit and parallel-to-serial circuit between the input / output bus of the second storage chip and each second storage block of the second storage chip. A plurality of data groups of a plurality of first storage blocks of the first storage chip are output to the input / output bus of the first storage chip within one bit switching cycle, and the bit switching cycle includes 2 N phases. A data group of each first storage block of the first storage chip is output to the input / output bus of the first storage chip at a corresponding phase of the bit switching cycle, and the period of a clock signal of the first storage chip is equal to the bit switching cycle divided by 2 N-1 , and N is an integer not less than 1. A plurality of data groups of a plurality of second storage blocks of the second storage chip are output to the input / output bus of the second storage chip within the bit switching cycle. A data group of each second storage block of the second storage chip is output to the input / output bus of the second storage chip at a corresponding phase of the bit switching cycle, and the period of a clock signal of the second storage chip is equal to the period of the clock signal of the first storage chip. Description of the Drawings

[0032] Figure 1 is a schematic diagram of a storage system disclosed in the prior art.

[0033] Figure 2A , 2B is a timing schematic diagram of a logic circuit writing data into a memory.

[0034] Figure 3 is a schematic diagram of a storage system disclosed in the first embodiment of the present invention.

[0035] Figure 4 is a schematic diagram of two transceiver structures disclosed in another embodiment of the present invention.

[0036] Figure 5 is a timing schematic diagram for comparing a conventional storage system and the storage system of the present invention.

[0037] Figure 6 It is a schematic diagram showing that the area of the memory in the present invention is smaller than that of the conventional memory, and the area of the physical layer in the present invention is also smaller than that of the physical layer in the conventional logic circuit.

[0038] Figure 7 It is a schematic diagram showing that the data width of the memory disclosed in another embodiment of the present invention changes according to a control signal.

[0039] Figure 8 、 9 It is a schematic diagram of the memory disclosed in different embodiments of the present invention.

[0040] Figure 10 It is a schematic diagram of the storage system disclosed in another embodiment of the present invention.

[0041] Figure 11 It is a schematic diagram showing the relationship between the bit switch, read data and signal TAU of the memory cell array at the rising edge of the clock, the bit switch, read data and signal TAU of the memory cell array at the falling edge of the clock, the clock signal, column address, data DQ and signal DQS during the read cycle of the memory chip.

[0042] Figure 12 It is a schematic diagram showing the relationship between the bit switch and write data of the memory cell array at the rising edge of the clock, the bit switch and write data of the memory cell array at the falling edge of the clock, the clock signal, column address, data DQ and signal DQS during the write cycle of the memory chip.

[0043] Figure 13 It is a schematic diagram showing that when the column address is generated from the internal direct interface wide bus counter in the memory chip, the remaining column addresses other than the starting column address are not cared about.

[0044] Figure 14 It is a schematic diagram showing the relationship between a bit switch cycle and a clock cycle when the memory chip has different numbers of subsystems.

[0045] Figure 15 It is a schematic diagram of a memory chip with 4 subsystems disclosed in another embodiment of the present invention.

[0046] Figure 16 It is a schematic diagram showing the relationship between the clock signal, bit switch and data of the clock rising edge 1 memory cell array, the clock falling edge 1 memory cell array, the clock rising edge 2 memory cell array and the clock falling edge 2 memory cell array during the read / write cycle of the memory chip.

[0047] Figure 17It is a schematic diagram showing the relationship between the clock signal and data (phases) when the storage chip has different numbers of subsystems.

[0048] Figure 18 It is a schematic diagram showing the relationship between the clock signal and the read data when the storage chip has 4 subsystems.

[0049] Figure 19 It is a schematic diagram showing the relationship between the clock signal and the write data when the storage chip has 4 subsystems.

[0050] Among them, the reference numerals are explained as follows:

[0051] 10, 100, 1000 Storage systems

[0052] 20, 101 Memories

[0053] 21 Cell array

[0054] 22, 314 Serial-to-parallel circuit

[0055] 23, 312 Parallel-to-serial circuit

[0056] 30 Logic circuit

[0057] 31, 103, 1004 Physical layer

[0058] 32, 105 Controller

[0059] 1011 First alignment circuit

[0060] 102 Logic circuit

[0061] 1031 Second alignment circuit

[0062] 1002, 1502 Storage chips

[0063] 10022 Clock rising edge memory cell array

[0064] 10024 Clock falling edge memory cell array

[0065] 10026 First alignment circuit

[0066] 10028 Second alignment circuit

[0067] 1042 Alignment circuit

[0068] 10030, 15038 First direct transfer data bus

[0069] 10032, 15040 Second direct transfer data bus

[0070] 10034, 15046 First direct receive data bus

[0071] 10036, 15048 Second direct receive data bus

[0072] 15022 Clock rising edge 1 memory cell array

[0073] 15024 Clock falling edge 1 memory cell array

[0074] 15026 Clock rising edge 2 memory cell array

[0075] 15028 Clock falling edge 2 memory cell array

[0076] 15030 First alignment circuit

[0077] 15032 Second alignment circuit

[0078] 15034 Third alignment circuit

[0079] 15036 Fourth alignment circuit

[0080] 15042 Third direct transfer data bus

[0081] 15044 Fourth direct transfer data bus

[0082] 15050 Third direct receive data bus

[0083] 15052 Fourth direct receive data bus

[0084] A0 - A7 Column address

[0085] BS0 - BS15 Bit switch

[0086] Dfi_wrdata Double data rate physical layer interface read data

[0087] Dfi_rddata Double data rate physical layer interface write data

[0088] Dqt0 - Dqt15, DQ Data

[0089] DQS Signal

[0090] f1 Clock falling edge 1 f2 Clock falling edge 2 CS Control signal

[0091] FP First pad

[0092] FCP First control pad

[0093] SP Second pad

[0094] SCP Second Control Pad

[0095] TR1 and TR2 Transceivers

[0096] Time t0 to t15

[0097] FPN and SPN Pads

[0098] DLSA First Sense Amplifier

[0099] BLSA Second Sense Amplifier

[0100] Storage Blocks B0 - B3

[0101] XCLK and CLK Clock Signals

[0102] XADDR Column Address Detailed Implementation Manner

[0103] Please refer to Figure 3 , Figure 3 is a schematic diagram of the storage system 100 disclosed in the first embodiment of the present invention. As Figure 3 shown, the storage system 100 includes a memory 101 and a logic circuit 102, where the memory 101 can be a Dynamic Random Access Memory (DRAM), a Static Random Access Memory (SRAM), a flash memory, or other memories, and the logic circuit 102 can be an artificial intelligence chip or a System on a Chip (SOC). In addition, in an embodiment of the present invention, the memory 101 may include a base DRAM chip and a plurality of DRAM chips stacked on the base DRAM chip. In addition, the logic circuit 102 can be coupled to other devices or processors through an Advanced eXtensible Interface (AXI) bus, where the AXI bus is a bus protocol that is part of the Advanced Microcontroller Bus Architecture (AMBA) 3.0 protocol. The AXI bus includes a write data bus and a read data bus. In addition, the operation method of the AXI bus is well known to those skilled in the art and will not be elaborated here.

[0104] Memory 101 includes a first alignment circuit 1011 and a plurality of first pads FP, wherein the first alignment circuit 1011 is used to align data with respect to the memory 101, and the first alignment circuit 1011 includes a plurality of transceivers. That is, the first alignment circuit 1011 is used to transmit the data simultaneously or receive the data simultaneously (for example, transmit the data at the same clock, or receive the data at the same clock, that is, the plurality of transceivers of the first alignment circuit 1011 can transmit the data in parallel or receive the data in parallel). On the other hand, the logic circuit 102 includes a physical layer 103 and a controller 105, wherein the physical layer 103 is electrically connected to the controller 105 through a double data rate physical layer interface (DDR PHY Interface, DFI) bus. The double data rate physical layer interface bus includes a plurality of line pairs, and the plurality of line pairs includes a plurality of write lines and a plurality of read lines. In addition, the physical layer 103 includes a second alignment circuit 1031 and a plurality of second pads SP, wherein the second alignment circuit 1031 is used to align the data, and the second alignment circuit 1031 also includes a plurality of transceivers. That is, the second alignment circuit 1031 is used to transmit the data simultaneously or receive the data simultaneously (for example, transmit the data at the same clock, or receive the data at the same clock, that is, the plurality of transceivers of the second alignment circuit 1031 can transmit the data in parallel or receive the data in parallel).

[0105] In this embodiment of the present invention, the first alignment circuit 1011 and the second alignment circuit 1031 can align the data and transmit the data in parallel, or can align and receive the data in parallel. In the memory 101 and the physical layer 103, without the need for traditional serial-to-parallel circuits and parallel-to-serial circuits, the data can be transmitted between the memory 101 and the logic circuit 102. Therefore, the controller (or storage controller) 105 can use the plurality of line pairs, the second alignment circuit 1031, the plurality of second pads SP, the plurality of first pads FP, and the first alignment circuit 1011 to access the data with respect to the memory 101 in parallel. The number of the plurality of first pads FP can be equal to the number of the plurality of write lines (or the number of the plurality of read lines) in the plurality of line pairs. Additionally, the number of the plurality of second pads SP can be equal to the number of the plurality of write lines (or the number of the plurality of read lines) in the plurality of line pairs.

[0106] For example, as Figure 3As shown, the number of multiple first pads FP or the number of multiple second pads SP is equal to N, and the data may be N-bit data RD read from the cell array of the memory 101 or N-bit data WD written to the cell array of the memory 101. When the logic circuit 102 reads N-bit data RD from the cell array of the memory 101 in parallel, the first alignment circuit 1011 receives N-bit data RD from the cell array of the memory 101 in parallel, and simultaneously transmits N-bit data RD to the second alignment circuit 1031 in parallel through the multiple first pads FP and the multiple second pads SP. After the second alignment circuit 1031 receives N-bit data RD in parallel, the second alignment circuit 1031 transmits N-bit data RD to the controller 105 in parallel through the multiple read lines of the multiple line pairs of the double data rate physical layer interface bus. On the other hand, when the logic circuit 102 writes N-bit data WD to the cell array of the memory 101 in parallel, the second alignment circuit 1031 receives N-bit data WD from the controller 105 in parallel through the multiple write lines of the multiple line pairs of the double data rate physical layer interface bus. Then, the second alignment circuit 1031 can transmit N-bit data WD to the first alignment circuit 1011 in parallel without passing through a conventional parallel-to-serial circuit and a serial-to-parallel circuit. After the first alignment circuit 1011 receives N-bit data WD, the first alignment circuit 1011 writes N-bit data WD to the cell array of the memory 101 in parallel.

[0107] In addition, each of the first alignment circuit 1011 and the second alignment circuit 1031 includes multiple transceivers, where each transceiver of the first alignment circuit 1011 is coupled to a corresponding pad of the multiple first pads FP, and each transceiver of the second alignment circuit 1031 is coupled to a corresponding pad of the multiple second pads SP. Please refer to Figure 4 . Figure 4 FIG. is a schematic structural diagram of two transceivers TR1 and TR2 disclosed in another embodiment of the present invention, where each transceiver of the first alignment circuit 1011 (not shown in Figure 4 ) may be the transceiver TR1, and each transceiver of the second alignment circuit 1031 (not shown in Figure 4 ) may be the transceiver TR2. In addition, the components of the transceiver TR1 and the transceiver TR2 are well known to those skilled in the art and will not be described in detail herein. In addition, the coupling relationship between the components of the transceiver TR1 and the transceiver TR2 can be referred to Figure 4, which will not be elaborated here. When the write enable signal W_EN is enabled and the read enable signal R_EN is disabled, the transceiver TR2 transmits one bit of data WD_N of the N-bit data WD to the transceiver TR1 through a first pad FPN and a second pad SPN. On the other hand, when the write enable signal W_EN is disabled and the read enable signal R_EN is enabled, the transceiver TR1 transmits one bit of data RD_N of the N-bit data RD to the transceiver TR2 through the first pad FPN and the second pad SPN. Since the write enable signal W_EN and the read enable signal R_EN are common signals to the first alignment circuit 1011 and the second alignment circuit 1031, the first alignment circuit 1011 can simultaneously transmit the N-bit data RD in parallel or receive the N-bit data WD in parallel, and the second alignment circuit 1031 can simultaneously transmit the N-bit data WD in parallel or receive the N-bit data RD in parallel.

[0108] In another embodiment of the present invention, a first write enable signal and a first read enable signal are signals for the first alignment circuit 1011, and a second write enable signal and a second read enable signal are signals for the second alignment circuit 1031, wherein the first write enable signal and the first read enable signal respectively correspond to the second write enable signal and the second read enable signal.

[0109] The first alignment circuit 1011 and the second alignment circuit 1031 do not need to be connected to a traditional parallel-to-serial circuit and a serial-to-parallel circuit. The first alignment circuit 1011 can simultaneously transmit the N-bit data RD to the second alignment circuit 1031 in parallel or receive the N-bit data WD from the second alignment circuit 1031 in parallel. Similarly, the second alignment circuit 1031 can simultaneously receive the N-bit data RD from the first alignment circuit 1011 in parallel or transmit the N-bit data WD to the first alignment circuit 1011 in parallel. In addition, as Figure 4 shown, the present invention is not limited to each transceiver of the first alignment circuit 1011 being the transceiver TR1 and each transceiver of the second alignment circuit 1031 being the transceiver TR2. That is to say, each transceiver of the first alignment circuit 1011 and each transceiver of the second alignment circuit 1031 can be other transceiver circuits, buffers, or registers.

[0110] Please refer to Figure 5 . Figure 5 is a timing schematic diagram for comparing a traditional storage system and the storage system 100. For example, as Figure 5As shown in (a), when a conventional logic circuit reads 8-bit serial data D0 - D7 from a conventional memory, the conventional memory must use three clocks clk1, clk2, clk3 to form 8 states (for example, data D0 corresponds to the state clk1 = 1, clk2 = 1, clk3 = 1, data D1 corresponds to the state clk1 = 1, clk2 = 1, clk3 = 0... etc.), so that the 8-bit serial data D0 - D7 can be converted into a parallel state. Therefore, a controller of the conventional logic circuit can only start to receive the parallel data D0 - D7 in parallel until a time T4.

[0111] However, as Figure 5 shown in (b), since there is no need to connect a conventional parallel-to-serial circuit and a serial-to-parallel circuit, the data D0 - D7 are transmitted in parallel by the first alignment circuit of the memory 101 at the same time, and the second alignment circuit of the controller 105 can start to receive the parallel data D0 - D7 at a time T0. Therefore, compared with the conventional memory system, the present invention can save 4 clock delays. In addition, the operation method of writing 8-bit data D0 - D7 is similar to the operation method described above and will not be elaborated here.

[0112] Please refer to Figure 3 again. As Figure 3 shown, the controller 105 is further coupled to the physical layer 103 through a plurality of control lines. The physical layer 103 further includes a plurality of second control pads SCP, the memory 101 further includes a plurality of first control pads FCP, and the plurality of first control pads FCP are electrically connected to the plurality of second control pads SCP. Therefore, the controller 105 can use the plurality of control lines, the plurality of second control pads SCP, and the plurality of first control pads FCP to transmit control signals CS, etc. to the memory 101. In addition, Figure 3 only three first control pads, three second control pads and three control lines are shown, but the present invention is not limited thereto. In addition, the plurality of control lines and the plurality of line pairs between the physical layer 103 and the controller 105 are all included in the double data rate physical layer interface bus, where the double data rate physical layer interface bus defines the signals, timing parameters and programmable parameters required for communication between the physical layer 103 and the controller 105. Therefore, the control signals CS, etc. are defined by the double data rate physical layer interface bus and may include, for example, a write enable signal, a read enable signal, and a chip select signal. In addition, the operation method of the double data rate physical layer interface bus is well known to those skilled in the art and will not be elaborated here. In addition, in another embodiment of the present invention, the logic circuit 102 may further include a system circuit (not shown in Figure 3In (China), the system circuit may include other peripheral interfaces. The controller (or storage controller) 105 communicates with the system circuit via the Advanced eXtensible Interface bus (AXI bus). For example, the controller 105 may transmit N-bit data RD to the system circuit via the Advanced eXtensible Interface bus, or receive N-bit data WD from the system circuit via the Advanced eXtensible Interface bus for other devices or processors.

[0113] In addition, a plurality of first pads FP can be electrically connected to a plurality of second pads SP through wire bonds, metal bridges, flip-chips, micro-bumps, or other bonding techniques. In addition, in other embodiments of the present invention, since the plurality of first pads FP are electrically connected to the plurality of second pads SP, the plurality of first pads FP and the plurality of second pads SP are not coupled to an environment outside the storage system 100. Therefore, the plurality of first pads FP and the plurality of second pads SP do not need to include a conventional electrostatic discharge protection circuit, and the sizes of the plurality of first pads FP and the plurality of second pads SP can be reduced.

[0114] In other embodiments of the present invention, the second alignment circuit 1031 of the physical layer 103 can be applied to different data widths, depending on the data width of the Advanced eXtensible Interface bus. However, in other embodiments of the present invention, both the second alignment circuit 1031 of the physical layer 103 and the first alignment circuit 1011 of the memory 101 can be applied to different data widths simultaneously, depending on the data width of the Advanced eXtensible Interface bus. For example, when the logic circuit 102 is applied to a memory with a Q-bit data width, the controller 105 can notify the physical layer 103 to adjust the second alignment circuit 1031 so that the second alignment circuit 1031 only uses Q read lines out of the plurality of wire pairs to transmit Q-bit data to the controller 105 (or only uses Q write lines out of the plurality of wire pairs to receive Q-bit data from the controller 105), where Q is a positive integer greater than 1 and less than N. Therefore, the physical layer 103 and the controller 105 can be applied to different system circuits and different memories with different data widths.

[0115] Since the first alignment circuit 1011 and the second alignment circuit 1031 become smaller and simpler, and the conventional parallel-to-serial and serial-to-parallel circuits are omitted from the memory 101 and the physical layer 103, the write / read speed of the memory 101 is significantly increased, the area of the memory 101 is smaller than that of a conventional memory, and the area of the physical layer 103 is also smaller than the area of the physical layer in a conventional logic circuit (such as Figure 6As shown, the memory wall problem between the memory 101 and the logic circuit 102 is also reduced. In addition, the physical layer 103 can receive signals such as Dfi cke, Dfi CK / CKB, Dfi BA, Dfi address, Dfi cs, Dfi ras, Dfi cas, Dfi we, Dfi wrdata, Dfi wrdata mask, Dfi wrdata valid, etc. from the controller 105 through the double data rate physical layer interface bus (DFI bus), and transmit signals such as Dfi rddata, Dfi rddata valid, etc. to the controller 105. The signals such as Dfi cke, Dfi CK / CKB, Dfi BA, Dfi address, Dfi cs, Dfi ras, Dfi cas, Dfi we, Dfi wrdata, Dfi wrdata mask, Dfi wrdata valid, etc., and the signals such as Dfi rddata, Dfi rddata valid, etc. are defined in the double data rate physical layer interface specification, so they will not be elaborated here. In addition, the physical layer 103 can transmit signals such as CKE, CK / CKB, BA, Addr, CSB, RASB, CASB, WEB, DQ, DM, DQS / DQSB, etc. to the memory 101. The signals such as CKE, CK / CKB, BA, Addr, CSB, RASB, CASB, WEB, DQ, DM, DQS / DQSB, etc. are also defined in the double data rate physical layer interface specification, so they will not be elaborated here either. Therefore, even if the memory 101 and the logic circuit 102 are manufactured by heterogeneous processes, multiple first pads FP can still be electrically connected to multiple second pads SP. For example, the transistors of the memory 101 can be planar transistors or trench transistors used in current memory technologies (such as dynamic random access memory or high bandwidth memory technology), while the transistors of the logic circuit 102 can be three-dimensional transistors (such as triple-gate transistors, fin field-effect transistors (FinFETs), or gate-all-around transistors). However, in other embodiments of the present invention, the memory 101 and the logic circuit 102 are manufactured by homogeneous processes. That is to say, the transistors of the memory 101 and the logic circuit 102 can adopt planar transistors or trench transistors, triple-gate transistors, fin field-effect transistors, gate-all-around transistors, or other transistors.In addition, since the memory 101 and the logic circuit 102 employ the first alignment circuit 1011 and the second alignment circuit 1031 instead of the conventional parallel-to-serial circuit and serial-to-parallel circuit, the power consumption of the memory 101 and the logic circuit 102 can be saved, the latency of accessing the memory 101 is reduced, and the area cost of the memory 101 and the logic circuit 102 is decreased. Therefore, the read / write window margins of the storage system 100 can be improved.

[0116] In addition, please refer to Figure 7 . Figure 7 FIG. is a schematic diagram showing that the data width of the memory disclosed in another embodiment of the present invention changes according to a control signal. For example (but not limited thereto), the memory 101 includes a memory cell array (abbreviated as cell array), M second sense amplifiers BLSA (such as bit line sense amplifiers), and N first sense amplifiers DLSA (such as data line sense amplifiers), wherein the number of connections between the M second sense amplifiers BLSA and the N first sense amplifiers DLSA can be changed according to a control signal (such as SB0-SB4 in Table 1), the second sense amplifiers BLSA are located between the cell array and the first sense amplifiers DLSA, the first sense amplifiers DLSA are located between the second sense amplifiers BLSA and the first alignment circuit 1011, wherein the first alignment circuit 1011 includes a plurality of transceivers, and the first alignment circuit 1011 is located between the first sense amplifiers DLSA and the input / output data bus (not shown in Figure 7 ), N is a positive integer not greater than M, and the input / output data bus is coupled to a plurality of first pads FP.

[0117] In one embodiment, the control signal is stored in a register (not shown in Figure 7 ) of the memory 101, such as a mode register. In addition, the second sense amplifiers BLSA are coupled to the bit lines (not shown in Figure 7 ) of the memory 101, and the first sense amplifiers DLSA are coupled to the data lines (not shown in Figure 7 ) of the memory 101. The N first sense amplifiers DLSA are electrically connected to all or part of the M second sense amplifiers BLSA through a plurality of bit switches, and the plurality of bit switches can be selected or enabled according to the above control signal.

[0118] As shown in Table 1 and Figure 7 FIG., for example, when the control signals SB0-SB4 are 0 / 0 / 0 / 0 / 1, the second sense amplifiers will pass through bit switches (not shown in Figure 7, a selected group of bit switches, such as 128 or fewer bit switches selected by control signals SB0 - SB4 (0 / 0 / 0 / 0 / 1) based on a given column address, transfer 128 bits of data to the first sense amplifier. That is, 128 - bit data can be read from the cell array of memory 101 through electrical connection of all or part of the second sense amplifiers and the first sense amplifier (e.g., through 128 connected second sense amplifiers and 128 first sense amplifiers), or 128 - bit data can be electrically connected through all or part of the second sense amplifiers and the first sense amplifier (e.g., through 128 connected second sense amplifiers and 128 first sense amplifiers) and written into the cell array of memory 101 by the first alignment circuit 1011. That is to say, when the 128 - bit data is read from the cell array of memory 101, the multiple transceivers of the first alignment circuit 1011 receive the 128 - bit data in parallel from the 128 first sense amplifiers and then transfer it to the input / output data bus of memory 101; or when the 128 - bit data is written into the cell array of memory 101, the multiple transceivers of the first alignment circuit 1011 receive the 128 - bit data in parallel from the input / output data bus and then transfer it to the 128 first sense amplifiers. In other words, when the 128 - bit data is read from the cell array of memory 101, part of the second sense amplifiers BLSA (e.g., the 128 connected second sense amplifiers) output the 128 - bit data to the first sense amplifier DLSA (e.g., the 128 first sense amplifiers), then the first sense amplifier DLSA outputs the 128 - bit data in parallel to the first alignment circuit 1011, and then the first alignment circuit 1011 transfers the 128 - bit data in parallel through the input / output data bus; or when the 128 - bit data is written into the cell array of memory 101, the input / output data bus transfers the 128 - bit data in parallel to the first alignment circuit 1011, then the first alignment circuit 1011 outputs the 128 - bit data in parallel to the 128 first sense amplifiers, and then the 128 first sense amplifiers output the 128 - bit data in parallel to the connected second sense amplifiers (e.g., the 128 connected second sense amplifiers BLSA). In addition, according to the 128 - bit data width output in parallel by the first sense amplifier, the data width of memory 101 (i.e., the width of the input / output data bus of memory 101) is equal to 128 bits width.At this time, since the data width of the memory 101 is equal to 128-bit width, according to the control signals SB0 - SB4, the data width of the write line (or read line) of the double data rate physical layer interface bus (DFI bus) coupled to the physical layer 103 is also equal to or set to 128-bit width, and the data width of the controller 105 and the write data bus (or read data bus) of the advanced extensible interface bus are also both equal to 128-bit width. For example. Figure 7 As shown, when the logic circuit 102 is included in a computing system, this computing system has a system interface bus (i.e., the advanced extensible interface bus) including a read data bus and a write data bus. According to the control signals SB0 - SB4 (0 / 0 / 0 / 0 / 1) input to the controller 105, the widths of both the read data bus and the write data bus are equal to 128-bit width. In addition, the width of the double data rate physical layer interface bus is selectively adjusted according to the control signals SB0 - SB4 (0 / 0 / 0 / 0 / 1) input to the physical layer 103.

[0119] Similarly, as shown in Table 1 and Figure 7 When the control signals SB0 - SB4 are 0 / 0 / 0 / 1 / 0, 256 of the M second sense amplifiers are electrically connected to 256 first sense amplifiers through another group of selected bit switches (e.g., 256 or fewer bit switches selected based on a given column address), so according to the 256 first sense amplifiers, the data width of the memory 101 is equal to 256-bit width; when the control signals SB0 - SB4 are 0 / 0 / 0 / 1 / 1, 512 of the M second sense amplifiers are electrically connected to 512 first sense amplifiers through another group of selected bit switches (e.g., 512 or fewer bit switches selected based on a given column address), so according to the 512 first sense amplifiers, the data width of the memory 101 is equal to 512-bit width; when the control signals SB0 - SB4 are 0 / 0 / 1 / 0 / 0, 1024 of the M second sense amplifiers are electrically connected to 1024 first sense amplifiers through another group of selected bit switches (e.g., 1024 or fewer bit switches selected based on a given column address), so according to the 1024 first sense amplifiers, the data width of the memory 101 is equal to 1024-bit width; when the control signals SB0 - SB4 are 0 / 0 / 0 / 0 / 0, 64 of the M second sense amplifiers are electrically connected to 64 first sense amplifiers through another group of selected bit switches (e.g., 64 or fewer bit switches selected based on a given column address), so according to the 64 first sense amplifiers, the data width of the memory 101 is equal to 64-bit width.

[0120] That is to say, since there is no need to connect the traditional serial-to-parallel circuit and parallel-to-serial circuit, and the first sense amplifier can be directly connected to the first alignment circuit, the parallel output data width of the first sense amplifier is equal to the data width of the memory 101. The data width of the memory 101 is equal to the data width of the write line (or read line) of the double data rate physical layer interface bus (DFI bus) of the physical layer 103, and is also equal to the data width of the controller 105 and the write data bus (or read data bus) data width of the advanced extensible interface bus. From another perspective, when the data width of the write data bus (or read data bus) of the advanced extensible interface bus in the application environment of the memory 101 and the logic circuit 102 is confirmed, by setting the control signals SB0 - SB4, the data width of the memory 101 can be made equal to the data width of the write line (or read line) of the double data rate physical layer interface bus (DFI bus) of the physical layer 103, and is also equal to the data width of the controller 105 and the write data bus (or read data bus) data width of the advanced extensible interface bus. In addition, the present invention is not limited to the memory 101 including M second sense amplifiers, nor is it limited to Figure 7 the configuration of the control signals SB0 - SB4 shown. In addition, the present invention is not limited to the number of the control signals SB0 - SB4, that is to say, the present invention can have more or fewer control signals than the number of the control signals SB0 - SB4.

[0121]

[0122] Table 1

[0123] In addition, please refer to Figure 8 。 Figure 8 FIG. is a schematic diagram of the memory 801 disclosed in another embodiment of the present invention. The difference between the memory 801 and the memory 101 is that the memory 801 includes 4 memory blocks B0 - B3, and each memory block of the memory blocks B0 - B3 is the cell array of the memory 101. However, the present invention is not limited to the memory 801 including 4 memory blocks B0 - B3 (that is to say, the memory 801 may include multiple memory blocks). In addition, for simplicity, the M second sense amplifiers BLSA and the N first sense amplifiers DLSA are not shown in Figure 8 。

[0124] As shown in Table 2 and Figure 8As shown, when the control signals SB0 - SB4 are 0 / 0 / 0 / 1 / 0, 256 second sense amplifiers (not shown) of a specific storage block of the memory 801 can be electrically connected to 256 first sense amplifiers (not shown) according to the control signals SB0 - SB4. Thus, through the 256 connected second sense amplifiers and the 256 first sense amplifiers, 256 - bit data can be read out from the specific storage block of the memory 801 by the first alignment circuit 1011, or through the 256 connected second sense amplifiers and the 256 first sense amplifiers, 256 - bit data can be written by the first alignment circuit 1011 into the specific storage block of the memory 801. Additionally, the specific storage block of the memory 801 can be selected by other signals, such as a block selection signal. That is, as shown in Table 2, according to the 256 first sense amplifiers, the data width of the selected storage block of the memory 801 can be adjusted to be equal to 256. Moreover, since the four storage blocks B0 - B3 are independent of each other, the data width of the memory 801 (i.e., the width of the input / output data bus of the memory 801) is also equal to 256. Furthermore, in other embodiments of the present invention, according to the control signals SB0 - SB4 (0 / 0 / 0 / 1 / 0), the data width of the controller 105 and the data width of the write line (or read line) of the double data rate physical layer interface bus are both equal to 256.

[0125] Furthermore, for the other data widths of each storage block of the memory 801 corresponding to the control signals SB0 - SB4 (0 / 0 / 1 / 0 / 0), (0 / 0 / 0 / 1 / 1), (0 / 0 / 0 / 0 / 1), (0 / 0 / 0 / 0 / 0), and the other data widths of the memory 801, reference can be made to Table 2, so they will not be elaborated here. Moreover, the present invention is not limited to Figure 8 the configuration of the control signals SB0 - SB4 shown.

[0126]

[0127]

[0128] Table 2

[0129] Furthermore, please refer to Figure 9 . Figure 9FIG. 901 is a schematic diagram of the memory 901 disclosed in another embodiment of the present invention. The difference between the memory 901 and the memory 801 is that the storage blocks B0 and B1 are included in the block group BG0, and the storage blocks B2 and B3 are included in the block group BG1. However, the present invention is not limited to the block group BG0 including the storage blocks B0 and B1, and the block group BG1 including the storage blocks B2 and B3. For example, all the blocks B0, B1, B2, and B3 can be grouped into a block group BGX.

[0130] Taking the block group BG0 as an example, a first group of sense amplifiers is coupled to the data lines, and a second group of sense amplifiers is coupled to the data lines. The first group of sense amplifiers corresponds to the storage block B0 and is used to output a plurality of first data in parallel. The second group of sense amplifiers corresponds to the storage block B1 and is used to output a plurality of second data in parallel. The first group of sense amplifiers and the second group of sense amplifiers are the first sense amplifiers (i.e., DLSA) mentioned previously. In addition, a third group of sense amplifiers is coupled to the bit lines and is arranged between the storage block B0 and the first group of sense amplifiers. A fourth group of sense amplifiers is coupled to the bit lines and is arranged between the storage block B1 and the second group of sense amplifiers. The third group of sense amplifiers and the fourth group of sense amplifiers are the second sense amplifiers (i.e., BLSA) mentioned previously.

[0131] Therefore, as shown in Table 3 and Figure 9As shown, when the control signals SB0 - SB4 are 0 / 1 / 0 / 1 / 0, according to the control signals SB0 - SB4, 128 second sense amplifiers of each memory block corresponding to a specific block group (such as block group BG0) are electrically connected to 128 first sense amplifiers of each memory block corresponding to the specific block group. Therefore, through the 256 connected second sense amplifiers and the 256 first sense amplifiers, 256 - bit data can be read out from the specific block group by the first alignment circuit 1011 (because the first alignment circuit 1011 can read 128 - bit data of the 256 - bit data from the memory block of the specific block group through 128 connected second sense amplifiers and 128 first sense amplifiers corresponding to one memory block, and the first alignment circuit 1011 can read the other 128 - bit data of the 256 - bit data from the other memory block of the specific block group through another 128 connected second sense amplifiers and another 128 first sense amplifiers corresponding to another memory block), or the 256 - bit data can be written into the specific block group by the first alignment circuit 1011 through the 256 connected second sense amplifiers and the 256 first sense amplifiers (because the first alignment circuit 1011 can write 128 - bit data of the 256 - bit data into the memory block of the specific block group through 128 connected second sense amplifiers and 128 first sense amplifiers corresponding to one memory block, and the first alignment circuit 1011 can write the other 128 - bit data of the 256 - bit data into the other memory block of the specific block group through another 128 connected second sense amplifiers and another 128 first sense amplifiers corresponding to another memory block). That is to say, as shown in Table 3, according to the 128 first sense amplifiers, the data width of each memory block of the specific block group is limited to be equal to 128. In addition, since the memory blocks B0, B1 are included in the block group BG0, the data width of the memory 901 (that is, the width of the input / output data bus of the memory 901) is equal to the sum of the data widths of all memory blocks of the specific block group (that is, 128 + 128 = 256). And compared with Figure 8 , the available memory blocks will be reduced by half.

[0132] Since there is no need to connect traditional serial-to-parallel circuits and parallel-to-serial circuits, and the first sense amplifier can be directly connected to the first alignment circuit, the data width of the memory 101 is equal to the data width of the write line (or read line) of the double data rate physical layer interface bus (DFI bus) of the physical layer 103, and is also equal to the data width of the controller 105 and the write data bus (or read data bus) data width of the advanced extensible interface bus. From another perspective, when the data width of the write data bus (or read data bus) of the advanced extensible interface bus in the application environment of the memory 101 and the logic circuit 102 is confirmed, by setting the control signals SB0 - SB4, the data width of the memory 101 can be made equal to the data width of the write line (or read line) of the double data rate physical layer interface bus (DFI bus) of the physical layer 103, and is also equal to the data width of the controller 105 and the write data bus (or read data bus) data width of the advanced extensible interface bus. In addition, for the other data widths of each memory block of the memory 901 corresponding to the control signals SB0 - SB4 (0 / 1 / 0 / 0 / 0), (0 / 1 / 0 / 0 / 1), (0 / 1 / 0 / 1 / 1), (0 / 0 / 0 / 0 / 0), reference can be made to Table 3, so it will not be elaborated here. In addition, the present invention is not limited to Figure 9 the configuration of the control signals SB0 - SB4 shown.

[0133]

[0134]

[0135] Table 3

[0136] Please refer to Figure 10 , Figure 10 FIG. Figure 10 is a schematic diagram of a storage system 1000 disclosed in another embodiment of the present invention, where the storage system 1000 includes a storage chip 1002 and a physical layer 1004, and the storage chip 1002 is an advanced direct interface wide bus (DWB) memory, and the storage chip 1002 has a combination of a clock rising edge subsystem (clock rising edge memory cell array 10022 (such as memory block 0)) and a clock falling edge subsystem (clock falling edge memory cell array 10024 (such as memory block 1)) to increase the bandwidth or performance of some specific address applications, and each switching cycle includes 2 stages. In addition, the physical layer 1004 is included in a logic circuit (not shown in Figure 10 ), and the logic circuit can refer to Figure 7The logic circuit 102 therein. The storage system 1000 can be employed in a storage chip where each memory cell array has a different data width (such as 32, 64, 128, 256, 512... bits), and the storage system 1000 takes a data width of 128 bits as an example. The rising-edge memory cell array 10022 can simultaneously transfer 128-bit data to a first alignment circuit 10026 or receive 128-bit data from the first alignment circuit 10026 in parallel at the rising edge of the clock signal CLK, and the falling-edge memory cell array 10024 can also simultaneously transfer 128-bit data to a second alignment circuit 10028 or receive 128-bit data from the second alignment circuit 10028 in parallel at the falling edge of the clock signal CLK, where the rising-edge memory cell array 10022 samples instructions, addresses, data, and operations at the rising edge of the clock signal CLK, and the falling-edge memory cell array 10024 samples instructions, addresses, data, and operations at the falling edge of the clock signal CLK. Additionally, the first alignment circuit 10026 and the second alignment circuit 10028 have a simple double data rate physical layer interface read / write data alignment function (which can be achieved through the rising edge of the clock or through a replicated data path of the rising edge of the clock), where the functions of the first alignment circuit 10026 and the second alignment circuit 10028 are the same as those of the first alignment circuit 1011 shown in Figure 7 and thus will not be elaborated here. For example, each alignment circuit in the first alignment circuit 10026 and the second alignment circuit 10028 includes a plurality of transceivers. As shown in Figure 10 , since the rising-edge memory cell array 10022 operates at the rising edge of the clock signal CLK and the falling-edge memory cell array 10024 operates at the falling edge of the clock signal CLK, the width of the input / output bus (not shown in Figure 10 ) of the storage chip 1002 is 128 bits. Therefore, the number of pins of the storage chip 1002 will not increase significantly. Additionally, as shown in Figure 10 , the storage chip 1002 further includes a first direct transfer data bus 10030, a second direct transfer data bus 10032, a first direct receive data bus 10034, and a second direct receive data bus 10036, where the first alignment circuit 10026 transfers 128-bit data to the input / output bus through the first direct transfer data bus 10030 or receives 128-bit data from the input / output bus through the first direct receive data bus 10034, and the second alignment circuit 10028 transfers 128-bit data to the input / output bus through the second direct transfer data bus 10032 or receives 128-bit data from the input / output bus through the second direct receive data bus 10036. Just as shown in Figure 7 , inFigure 10 In an embodiment, there is no parallel-to-serial circuit and serial-to-parallel circuit between the memory cell array 10022 at the rising edge of the clock (or its corresponding data line sense amplifier) and the input / output bus (or input / output pad), and there is also no parallel-to-serial circuit and serial-to-parallel circuit between the memory cell array 10024 at the falling edge of the clock (or its corresponding data line sense amplifier) and the input / output bus (or the input / output pad).

[0137] In an embodiment of the present invention, 128-bit data from the memory cell array 10022 at the rising edge of the clock and 128-bit data from the memory cell array 10024 at the falling edge of the clock share the same row address. Additionally, since the memory cell array 10022 at the rising edge of the clock can transfer 128-bit data to the first alignment circuit 10026 or receive 128-bit data from the first alignment circuit 10026 at the rising edge of the clock signal CLK, and the memory cell array 10024 at the falling edge of the clock can also transfer 128-bit data to the second alignment circuit 10028 or receive 128-bit data from the second alignment circuit 10028 at the falling edge of the clock signal CLK, the memory cell array 10022 at the rising edge of the clock and the memory cell array 10024 at the falling edge of the clock can have their respective corresponding row addresses, where the corresponding row addresses can be specified according to the requirements of the application program (the following are several examples):

[0138] 1) The memory cell array 10022 at the rising edge of the clock can store even column address data, while the memory cell array 10024 at the falling edge of the clock stores odd column address data (or vice versa).

[0139] 2) Or different address assignment methods, such as the memory cell array 10022 at the rising edge of the clock stores the column address corresponding to memory block 0 (that is, the column address B2 = 0 or B5 = 0 is assigned to the memory cell array 10022 at the rising edge of the clock), while the memory cell array 10024 at the falling edge of the clock stores the column address corresponding to memory block 1 (that is, the column address B2 = 1 or B5 = 1 is assigned to the memory cell array 10024 at the falling edge of the clock). That is to say, any assignment method of the column address of the rising-edge clock subsystem / falling-edge clock subsystem falls within the scope of the present invention.

[0140] Since the clock rising-edge memory cell array 10022 can transfer 128-bit data to the input / output bus of the memory chip 1002 or receive 128-bit data from the input / output bus of the memory chip 1002 through the first alignment circuit 10026 at the clock rising edge of the clock signal CLK, and the clock falling-edge memory cell array 10024 can also transfer 128-bit data to the input / output bus of the memory chip 1002 or receive 128-bit data from the input / output bus of the memory chip 1002 through the second alignment circuit 10028 at the clock falling edge of the clock signal CLK, a pair of alignment circuits 1042 included in the physical layer 1004 requires a 2:1 serial-to-parallel circuit or a 1:2 parallel-to-serial circuit (or a 2:1 multiplexer) to process 256-bit data (e.g., receiving the 256-bit data from the memory chip 1002 or transferring the 256-bit data to the memory chip 1002). For example, in the read cycle of the memory chip 1002, during a clock cycle (including a clock rising edge and a clock falling edge), the alignment circuit 1042 can combine the 128-bit data from the clock rising-edge memory cell array 10022 and the 128-bit data from the clock falling-edge memory cell array 10024 to generate 256-bit double data rate physical layer interface read data Dfi_rddata and transfer it to Figure 7 the controller 105 shown and then to Figure 7 the advanced extensible interface bus shown. And in the write cycle of the memory chip 1002, during a clock cycle (including a clock rising edge and a clock falling edge), the alignment circuit 1042 can receive 256-bit double data rate physical layer interface write data Dfi_wrdata from the advanced extensible interface bus through Figure 7 the controller 105 shown to transfer 128-bit data to the clock rising-edge memory cell array 10022 (memory block 0) at the clock rising edge of the clock cycle and transfer 128-bit data to the clock falling-edge memory cell array 10024 (memory block 1) at the clock falling edge of the clock cycle. Therefore, the alignment circuit 1042 still requires a 2:1 serial-to-parallel circuit, a 1:2 parallel-to-serial circuit, and 256 transceivers.

[0141] In addition to the 2:1 serial-to-parallel circuit and the 1:2 parallel-to-serial circuit, the alignment circuit 1042 also requires an additional simple read / write data alignment function (which can be achieved through the signal DQS, or the clock signal CLK, or other means), where the function of the alignment circuit 1042 and Figure 7The functions of the second alignment circuit 1031 shown are the same, so they will not be elaborated here. Therefore, the memory chip 1002 can maintain the same data width (i.e., 128 bits) and the same bit width (i.e., 128 bits) in the memory cell array 10022 (memory block 0) at the rising edge of the clock and the memory cell array 10024 (memory block 1) at the falling edge of the clock. However, the physical layer 1004 can Figure 7 receive 256-bit data from or transfer 256-bit data to the controller 105 shown Figure 7 to the controller 105 shown. Therefore, without increasing the number of pins of the memory chip 1002, the bus width of the controller 105 or the logic circuit 102 is doubled.

[0142] Next, the read cycle and write cycle of the memory chip 1002 will be described in detail. In Figure 11 the embodiment shown, one clock cycle is equal to a one-bit switching cycle with two phases, and one bit switching cycle can range from 1 nanosecond (ns) to nanoseconds (ns) depending on the array architecture.

[0143] Read cycle:

[0144] As Figure 11 shown, after the memory chip 1002 receives a read instruction / address, the memory cell array 10022 at the rising edge of the clock can sample the read instruction / address using the clock signal XCLK to turn on the bit switches BS0, BS2, BS4, BS6 at the rising edge of the clock signal XCLK (corresponding to time t0, time t2, time t4, time t6). In an embodiment of the present invention, the turning on of the bit switches BS0, BS2, BS4, BS6 corresponds to stage 1 of the bit switching cycle or the clock cycle. Additionally, the memory cell array 10022 at the rising edge of the clock can also generate a signal TAU in stage 1. Similarly, the memory cell array 10024 at the falling edge of the clock can sample the read instruction / address using the clock signal XCLK to turn on the bit switches BS1, BS3, BS5, BS7 at the falling edge of the clock signal XCLK (corresponding to time t1, time t3, time t5, time t7). In an embodiment of the present invention, the turning on of the bit switches BS1, BS3, BS5, BS7 corresponds to stage 2 of the bit switching cycle or the clock cycle. Additionally, the memory cell array 10024 at the falling edge of the clock can also generate another signal TAU in stage 2. Additionally, as Figure 11As shown, data Dqt0 (corresponding to time t0 and the even column address A0 in column address XADDR), data Dqt2 (corresponding to time t2 and the even column address A2 in column address XADDR), data Dqt4 (corresponding to time t4 and the even column address A4 in column address XADDR), and data Dqt6 (corresponding to time t6 and the even column address A6 in column address XADDR) can be read from the rising-edge memory cell array 10022 respectively after the bit switches BS0, BS2, BS4, and BS6 are turned on. Similarly, as Figure 11 shown, data Dqt1 (corresponding to time t1 and the odd column address A1 in column address XADDR), data Dqt3 (corresponding to time t3 and the odd column address A3 in column address XADDR), data Dqt5 (corresponding to time t5 and the odd column address A5 in column address XADDR), and data Dqt7 (corresponding to time t7 and the odd column address A7 in column address XADDR) can be read from the falling-edge memory cell array 10024 respectively after the bit switches BS1, BS3, BS5, and BS7 are turned on. Additionally, signal TAU (phase 1) and signal TAU (phase 2) can be combined into a single signal DQS, where signal DQS tracks the read data of each phase of the rising-edge memory cell array 10022 and the falling-edge memory cell array 10024. Additionally, as Figure 11 shown, data Dqt0, data Dqt1, data Dqt2, data Dqt3, data Dqt4, data Dqt5, data Dqt6, and data Dqt7 respectively correspond to column address A0, column address A1, column address A2, column address A3, column address A4, column address A5, column address A6, and column address A7. The rising-edge memory cell array 10022 and the falling-edge memory cell array 10024 share the same row address, instruction, and control signal, but have slightly different row addresses according to different applications. For example, in the bit switch cycle or the clock cycle, data Dqt0 from the even column address A0 in phase 1 and data Dqt1 from the odd column address A1 in phase 2 share the same row address.

[0145] Write cycle:

[0146] As Figure 12As shown, after the storage chip 1002 receives a write instruction / address, the clock rising edge memory cell array 10022 can also sample the write instruction / address using the clock signal XCLK to turn on the bit switches BS0, BS2, BS4, and BS6 at the clock rising edge of the clock signal XCLK (corresponding to time t0, time t2, time t4, and time t6, where time t0, time t2, time t4, and time t6 correspond to stage 1 of the bit switch cycle or the clock cycle). Similarly, the clock falling edge memory cell array 10024 can also sample the write instruction / address using the clock signal XCLK to turn on the bit switches BS1, BS3, BS5, and BS7 at the clock falling edge of the clock signal XCLK (corresponding to time t1, time t3, time t5, and time t7, where time t1, time t3, time t5, and time t7 correspond to stage 2 of the bit switch cycle or the clock cycle). Additionally, as Figure 12 shown, data Dqt0 (corresponding to time t0 and the even column addresses A0 in the column address XADDR), data Dqt2 (corresponding to time t2 and the even column addresses A2 in the column address XADDR), data Dqt4 (corresponding to time t4 and the even column addresses A4 in the column address XADDR), and data Dqt6 (corresponding to time t6 and the even column addresses A6 in the column address XADDR) can be written to the clock rising edge memory cell array 10022 respectively after the bit switches BS0, BS2, BS4, and BS6 are turned on. Similarly, as Figure 12 shown, data Dqt1 (corresponding to time t1 and the odd column addresses A1 in the column address XADDR), data Dqt3 (corresponding to time t3 and the odd column addresses A3 in the column address XADDR), data Dqt5 (corresponding to time t5 and the odd column addresses A5 in the column address XADDR), and data Dqt7 (corresponding to time t7 and the odd column addresses A7 in the column address XADDR) can be written to the clock falling edge memory cell array 10024 respectively after the bit switches BS1, BS3, BS5, and BS7 are turned on. Additionally, as Figure 12 shown, the signal DQS can sample the data DQ at the rising edge and falling edge of the clock signal XCLK and perform a write function on the corresponding bit switches.

[0147] Additionally, Figure 11 and / or Figure 12 the required column address can be generated from a direct interface wide bus (DWB) controller (such as the controller 105 shown in Figure 7 ) or generated from an internal direct interface wide bus counter in the storage chip 1002. For example, as Figure 13As shown in the upper part, in a burst read operation or a burst write operation, there is a starting address (e.g., an even column address A0) and a burst length (e.g., the burst length = 8). Therefore, the addresses after the starting address can be generated from an internal direct interface wide bus counter in the memory chip 1002. The burst length can be directly applied from the AXI instruction to the controller, and then the controller can issue a direct interface wide bus instruction according to the burst length. Additionally, in another embodiment of the present invention, as Figure 13 shown in the lower part, the direct interface wide bus controller (e.g., Figure 7 the controller 105 shown) can generate all column addresses A0 - A7.

[0148] Although the memory chip 1002 can provide multiple data transfer rates with two phases within one bit switching cycle, the memory chip 1002 can also provide normal data access with a single data transfer rate and normal bandwidth (e.g., accessing 128 bits of data at each clock rising edge and no data access at each clock falling edge, or accessing 128 bits of data at each clock falling edge and no data access at each clock rising edge) to save power (as Figure 7 shown in the memory 101). The switching between the multiple data transfer rates and normal data access of the memory chip 1002 is set in a mode register within the memory chip 1002.

[0149] In addition, the memory chip 1002 can be extended to have multiple (more than two) data transfer rates or multiple phases within one bit switching cycle. Please refer to Figure 14 , as Figure 14 (a) shows. As mentioned above, the memory chip 1002 includes two subsystems or two memory cell arrays. One clock cycle is equal to one bit switching cycle, and one bit switching cycle has two bit switching phases (phase 1, phase 2), where the two bit switching phases are used to continuously activate two memory cell arrays within one bit switching cycle. On the other hand, as Figure 14 (b) shows, if a memory chip includes four subsystems or four memory cell arrays, then two clock cycles are equal to one bit switching cycle, and one bit switching cycle has four bit switching phases (phase 1, phase 2, phase 3, phase 4), where the four bit switching phases are used to continuously activate four memory cell arrays within one bit switching cycle.

[0150] Next, please refer to Figure 15 , Figure 15Schematic diagram of a storage chip 1502 with 4 subsystems disclosed in another embodiment of the present invention, wherein the 4 subsystems include a rising-edge 1 memory cell array 15022 (e.g., memory block 0), a falling-edge 1 memory cell array 15024 (e.g., memory block 1), a rising-edge 2 memory cell array 15026 (e.g., memory block 2), and a falling-edge 2 memory cell array 15028 (e.g., memory block 3) to increase the bandwidth or performance of some specific address applications. Each switching cycle contains 4 phases, so the frequency of the clock signal of 1502 is twice the frequency of the clock signal CLK of the storage chip 1002. Additionally, the storage chip 1502 further includes a first alignment circuit 15030, a second alignment circuit 15032, a third alignment circuit 15034, a fourth alignment circuit 15036, a first direct data transfer bus 15038, a second direct data transfer bus 15040, a third direct data transfer bus 15042, a fourth direct data transfer bus 15044, a first direct data reception bus 15046, a second direct data reception bus 15048, a third direct data reception bus 15050, and a fourth direct data reception bus 15052. The functions of the first alignment circuit 15030, the second alignment circuit 15032, the third alignment circuit 15034, and the fourth alignment circuit 15036 are the same as those of the first alignment circuit 10026, so they will not be elaborated here. Similarly, the functions of the first direct data transfer bus 15038, the second direct data transfer bus 15040, the third direct data transfer bus 15042, and the fourth direct data transfer bus 15044 are the same as those of the first direct data transfer bus 10030, so they will not be elaborated here. Additionally, the functions of the first direct data reception bus 15046, the second direct data reception bus 15048, the third direct data reception bus 15050, and the fourth direct data reception bus 15052 are the same as those of the first direct data reception bus 10034, so they will not be elaborated here either. Additionally, similar to the rising-edge memory cell array 10022 and the falling-edge memory cell array 10024, the rising-edge 1 memory cell array 15022, the falling-edge 1 memory cell array 15024, the rising-edge 2 memory cell array 15026, and the falling-edge 2 memory cell array 15028 also share the same row address, instruction, and control signal, but have slightly different row addresses according to different applications.It is again worth noting that there is no parallel-to-serial circuit and serial-to-parallel circuit between the memory cell array 15022 (or its corresponding data line sense amplifier) and the input / output bus (or input / output pad) at the rising edge of clock 1, between the memory cell array 15024 (or its corresponding data line sense amplifier) and the input / output bus (or input / output pad) at the falling edge of clock 1, between the memory cell array 15026 (or its corresponding data line sense amplifier) and the input / output bus (or input / output pad) at the rising edge of clock 2, and between the memory cell array 15028 (or its corresponding data line sense amplifier) and the input / output bus (or input / output pad) at the falling edge of clock 2.

[0151] Because in the memory chip 1502, each bit switching cycle has 4 phases, where each bit switching cycle is equal to two clock cycles. In the read cycle or write cycle of the memory chip 1502, the relationship between the clock signals XCLK, r1 (corresponding to phase 1), f1 (corresponding to phase 2), r2 (corresponding to phase 3), f2 (corresponding to phase 4), data (phase 1), data (phase 2), data (phase 3), and data (phase 4) can be referred to Figure 16 , where t0 to t15 are time, r1 and r2 are the rising edges of clock 1 and clock 2 respectively, f1 and f2 are the falling edges of clock 1 and clock 2 respectively, Dqt0 to Dqt15 are data, and BS0 to BS15 are bit switches.

[0152] For example, in a bit switch cycle between t0 and t3, at t0&r1 (corresponding to stage 1), the bit switch BS0 (which may include multiple sub-bit switches) is turned on and the data Dqt0 (or 128-bit data) in the memory cell array 15022 at the rising edge of the read clock is read into the first alignment circuit 15030 (or 128-bit data is written from the first alignment circuit 15030 to the memory cell array 15022 at the rising edge of the clock); at t1&f1 (corresponding to stage 2), the bit switch BS1 (which may include multiple sub-bit switches) is turned on and the data Dqt1 (or 128-bit data) in the memory cell array 15024 at the falling edge of the read clock is read into the second alignment circuit 15032 (or 128-bit data is written from the second alignment circuit 15032 to the memory cell array 15024 at the falling edge of the clock). Additionally, at t2&r2 (corresponding to stage 3), the bit switch B2 (which may include multiple sub-bit switches) is turned on and the data Dqt2 (or 128-bit data) in the memory cell array 15026 at the rising edge of the second read clock is read into the third alignment circuit 15034 (or 128-bit data is written from the third alignment circuit 15034 to the memory cell array 15022 at the rising edge of the second clock). At t3&f2 (corresponding to stage 3), the bit switch B3 (which may include multiple sub-bit switches) is turned on and the data Dqt3 (or 128-bit data) in the memory cell array 15028 at the falling edge of the second read clock is read into the fourth alignment circuit 15036 (or 128-bit data is written from the fourth alignment circuit 15036 to the memory cell array 15028 at the falling edge of the second clock). Additionally, Figure 16 The other bit switch cycles shown have the same operations as the bit switch cycle between t0 and t3, so they will not be elaborated here.

[0153] Figure 17 It is a schematic diagram showing the relationship between the clock signal and the data phase when the memory chip includes different numbers of subsystems, where each subsystem or memory cell array shares the same row address, instruction, and control signal, but has slightly different row addresses according to different applications. Such as Figure 17As shown, when the memory chip includes 2 subsystems or memory cell arrays, each bit switching cycle (equal to one clock cycle) has 2 phases and the corresponding physical layer has 2:1 serial-to-parallel circuits and 1:2 parallel-to-serial circuits (or 2:1 multiplexers) to process 256 bits of data (when each subsystem outputs / receives 128 bits of data); when the memory chip includes 4 subsystems or memory cell arrays, each bit switching cycle (equal to two clock cycles) has 4 phases and the corresponding physical layer has 4:1 serial-to-parallel circuits and 1:4 parallel-to-serial circuits (or 4:1 multiplexers) to process 512 bits of data (when each subsystem outputs / receives 128 bits of data); when the memory chip includes 8 subsystems or memory cell arrays, each bit switching cycle (equal to four clock cycles) has 8 phases and the corresponding physical layer has 8:1 serial-to-parallel circuits and 1:8 parallel-to-serial circuits (or 8:1 multiplexers) to process 1024 bits of data (when each subsystem outputs / receives 128 bits of data); and when the memory chip includes 16 subsystems or memory cell arrays, each bit switching cycle (equal to eight clock cycles) has 8 phases and the corresponding physical layer has 8:1 serial-to-parallel circuits and 1:8 parallel-to-serial circuits (or 8:1 multiplexers) to process 2048 bits of data (when each subsystem outputs / receives 128 bits of data), etc. Of course, in Figure 17 any embodiment, the physical layer may further include the alignment circuit as described above.

[0154] Next, please refer to Figure 18 and Figure 19 , where Figure 18 is a schematic diagram showing the relationship between the clock signal CLK and the read data when the memory chip 1502 has 4 subsystems, and Figure 19 is a schematic diagram showing the relationship between the clock signal CLK and the write data when the memory chip 1502 has 4 subsystems. As Figure 18 and Figure 19 shown, each bit switching cycle includes 4 phases and a bit switching cycle (where for example each bit switching cycle is equal to 2.5 ns) is equal to two clock cycles (for example each clock cycle is equal to 1.25 ns), where the read data respectively from the 4 subsystems will be read within one bit switching cycle. Additionally, as Figure 18 shown, RR’ read data represents reusing the read data and resampling with the clock of the data signal DQ to achieve better alignment. Similarly, as Figure 19 shown, 4 write data will be respectively written into the 4 subsystems within one bit switching cycle.

[0155] Thus, the direct interface wide bus (DWB) memory chip disclosed by the present invention can have 8, 16... subsystems. Therefore, when the memory chip 1502 (direct interface wide bus (DWB) memory chip) has 8, 16... subsystems, the 8, 16... subsystems respectively have 8, 16... stages corresponding to each bit switching cycle, respectively correspond to each bit switching cycle equal to 4, 8... clock cycles tCK, and also correspond to the clock cycle tCK = bit switching cycle / 4, bit switching cycle / 8.... Additionally, taking the bit switching cycle equal to 2.5 ns and the width of the input / output bus of the memory chip 1502 equal to 128 bits as an example, Table 4 below illustrates the relationship between the number of subsystems (1, 2, 4, 8, 16), the clock cycle tCK, the data transfer rate, and the bandwidth of the 128-bit input / output bus:

[0156]

[0157] Table 4

[0158] Additionally, in another embodiment of the present invention, the memory system can include a direct interface wide bus (DWB) memory chip and a single data rate (SDR) memory chip. For example, the direct interface wide bus memory chip has 8 subsystems, the single data rate memory chip has 4 subsystems, and the operation of the direct interface wide bus memory chip can refer to Figure 17 and the operation of the single data rate memory chip can also refer to Figure 17 (however, the single data rate memory chip only operates on the rising edge (or falling edge) of the clock (800 MHz) of the 8 subsystems). Therefore, the memory system can simultaneously operate the direct interface wide bus memory chip and the single data rate memory chip on the clock (800 MHz) of the 8 subsystems. Additionally, the characteristics of the memory system can refer to Table 5, where for example, a bit switching cycle is equal to 5 ns and the clock frequency of the clock of the 8 subsystems is 800 MHz.

[0159] Direct interface wide bus (DWB) memory chip Clock frequency (MHz) 800 Data transfer rate (Mbps), DDR 1600 Phase 8 Bit switching period (ns) 5 DVW = Bit switching period / Phase (ns) 0.625 Single data rate (SDR) memory chip Clock frequency (MHz) 800 Data transfer rate (Mbps), SDR 800 Phase 4 Bit switching period (ns) 5 DVW = Bit switching period / Phase (ns) 1.25

[0160] Table 5

[0161] In summary, a direct interface wide bus memory chip with multiple subsystems can utilize each subsystem to transfer a set of data in parallel to the input / output bus of the direct interface wide bus memory chip (or receive the set of data in parallel from the input / output bus). In one bit switching cycle, a set of data can be read from each subsystem to the physical layer, or multiple sets of data can be separately written from the physical layer to the multiple subsystems. Therefore, the present invention can increase the data transfer rate of the direct interface wide bus memory chip (and also the data transfer rate between the physical layer and the controller) without increasing the width of the input / output bus or the number of input / output pins of the direct interface wide bus memory chip. Therefore, compared with the prior art, the present invention can not only reduce the power consumption, access latency and cost of the direct interface wide bus memory chip, but also increase the bandwidth and data transfer rate of the input / output bus of the direct interface wide bus memory chip.

[0162] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A storage chip, characterized in that Comprising: A plurality of storage blocks, wherein each storage block outputs or receives a data group in parallel; An input / output bus; and A plurality of alignment circuits, respectively corresponding to the plurality of storage blocks; Wherein a data group of one storage block is transmitted to a corresponding alignment circuit, and then the corresponding alignment circuit simultaneously transmits the data group in parallel to the input / output bus, or transmits the data group from the input / output bus to the corresponding alignment circuit, and then the corresponding alignment circuit simultaneously transmits the data group in parallel to the storage block; Wherein there is no serial-to-parallel circuit and parallel-to-serial circuit between the input / output bus and each storage block.

2. The storage chip according to claim 1, wherein Each alignment circuit includes a plurality of first transceivers, the plurality of first transceivers are connected to the input / output bus through a direct transmission / reception bus, and the width of the input / output bus is equal to the width of the data group received or output by each storage block.

3. The storage chip according to claim 1, characterized in that The plurality of data groups of the plurality of storage blocks are output to the input / output bus in a predetermined order.

4. The storage chip according to claim 3, wherein The data groups of each storage block share a common row address, and the column addresses of the data groups of each storage block are different from each other.

5. The storage chip according to claim 4, characterized in that The column address of the data group of each storage block is generated inside the storage chip or received from a storage controller outside the storage chip.

6. The storage chip according to claim 3, wherein The plurality of data groups of the plurality of storage blocks are output to the input / output bus within one bit switching cycle, the bit switching cycle includes a plurality of stages, and the data group of each storage block is output to the input / output bus at a corresponding stage of the bit switching cycle.

7. The storage chip according to claim 6, characterized in that The multiple stages of the bit-switching period include 2 N stages, and the period of a clock signal of the memory chip is equal to the bit-switching period divided by 2 N-1 , and N is an integer not less than 1.

8. The storage chip according to claim 6, characterized in that The number of stages of the bit switching cycle is set in a mode register in the storage chip.

9. The storage chip according to claim 1, characterized in that Further comprising: A plurality of data lines; and A plurality of groups of sense amplifiers, coupled to the plurality of data lines, wherein the storage block corresponds to a group of sense amplifiers, and the group of sense amplifiers is disposed between the storage block and the corresponding alignment circuit.

10. The storage chip according to claim 9, wherein: The plurality of storage blocks include a first storage block and a second storage block; The plurality of groups of sense amplifiers include a first group of sense amplifiers coupled to the plurality of data lines and a second group of sense amplifiers coupled to the plurality of data lines; The first group of sense amplifiers corresponds to the first storage block, and a first data group is simultaneously transmitted in parallel between the first group of sense amplifiers and the input / output bus through an alignment circuit corresponding to the first storage block; The second group of sense amplifiers corresponds to the second storage block, and a second data group is simultaneously transmitted in parallel between the second group of sense amplifiers and the input / output bus through an alignment circuit corresponding to the second storage block; The width of the input / output bus is equal to the width of the first data group and the width of the second data group.

11. The storage chip according to claim 10, characterized in that The width of a double data rate physical layer interface bus of a physical layer within a logic circuit is equal to the sum of the widths of the first data group and the second data group, wherein the double data rate physical layer interface bus is coupled between a controller within the logic circuit and the physical layer, the controller is further coupled to a high performance interface bus outside the logic circuit, and the logic circuit is coupled to the input / output bus of the memory chip.

12. The storage chip according to claim 10, characterized in that Further comprising: A plurality of bit lines; A third group of sense amplifiers, coupled to the plurality of bit lines and disposed between the first memory block and the first group of sense amplifiers; and A fourth group of sense amplifiers, coupled to the plurality of bit lines and disposed between the second memory block and the second group of sense amplifiers.

13. The storage chip according to claim 12, wherein Further comprising: A first bit switch group, located between the first group of sense amplifiers and the third group of sense amplifiers; and A second bit switch group, located between the second group of sense amplifiers and the fourth group of sense amplifiers.

14. A storage system, characterized in that Comprising: A first memory chip, comprising: A plurality of first memory blocks, wherein each first memory block outputs or receives a data group in parallel; An input / output bus; and A plurality of alignment circuits, corresponding to the plurality of first memory blocks respectively; Wherein a data group of a first memory block is transmitted to a corresponding alignment circuit, and then the corresponding alignment circuit simultaneously transmits the data group in parallel to the input / output bus, or transmits the data group from the input / output bus to the corresponding alignment circuit, and then the corresponding alignment circuit simultaneously transmits the data group in parallel to the first memory block; Wherein there is no serial-to-parallel circuit and parallel-to-serial circuit between the input / output bus and each first memory block; And A logic circuit, having a physical layer, wherein the logic circuit is located outside the first memory chip and electrically connected to the input / output bus of the first memory chip, and there is a serial-to-parallel circuit and a parallel-to-serial circuit within the physical layer.

15. The storage system according to claim 14, wherein The physical layer further comprises a plurality of second transceivers, and the plurality of second transceivers are electrically connected to the serial-to-parallel circuit and the parallel-to-serial circuit.

16. The storage system according to claim 15, wherein The multiple first storage blocks include 2 N first storage blocks, where N is an integer not less than 1, the serial-to-parallel circuit is a 2 N :1 serial-to-parallel circuit, and the parallel-to-serial circuit is a 1:2 N parallel-to-serial circuit.

17. The storage system according to claim 15, characterized in that The width of a double data rate physical layer interface bus of the physical layer is equal to the sum of the widths of the data groups of each first memory block in the first memory chip, and the double data rate physical layer interface bus is coupled between a controller within the logic circuit and the physical layer.

18. The storage system according to claim 14, wherein Each alignment circuit in the first memory chip comprises a plurality of first transceivers, the plurality of first transceivers are connected to the input / output bus through a direct transmit / receive bus, and the width of the input / output bus is equal to the width of the data group received or output by each first memory block.

19. The storage system according to claim 14, wherein The plurality of data groups of the plurality of first memory blocks are output to the input / output bus in a predetermined order.

20. The storage system according to claim 19, wherein The data groups of each first memory block share a common row address, and the column addresses of the data groups of each first memory block are different from each other.

21. The storage system according to claim 20, wherein A column address of a data group of each first storage block is generated inside the first storage chip or received from a storage controller outside the first storage chip.

22. The storage system according to claim 19, wherein Data groups of the multiple first storage blocks are output to the input / output bus within one bit switching cycle. The bit switching cycle includes multiple phases, and a data group of each first storage block is output to the input / output bus at a corresponding phase of the bit switching cycle.

23. The storage system according to claim 22, wherein The multiple phases of the bit-switching period include 2 N phases, the period of a clock signal of the first memory chip being equal to the bit-switching period divided by 2 N-1 , and N is an integer not less than 1.

24. The storage system according to claim 22, wherein The number of phases of the bit switching cycle is set in a mode register in the first storage chip.

25. The storage system according to claim 14, wherein Also included are: A second storage chip, including: Multiple second storage blocks, where each second storage block outputs or receives a data group in parallel; An input / output bus; And Multiple alignment circuits, respectively corresponding to the multiple second storage blocks; Wherein a data group of a second storage block in the second storage chip is transmitted to a corresponding alignment circuit, and then the corresponding alignment circuit simultaneously transmits the data group in parallel to the input / output bus, or transmits the data group of the second storage block in the second storage chip from the input / output bus to the corresponding alignment circuit, and then the corresponding alignment circuit simultaneously transmits the data group in parallel to the second storage block in the second storage chip; Wherein there is no serial-to-parallel circuit and parallel-to-serial circuit between the input / output bus of the second storage chip and each second storage block of the second storage chip; Among them, multiple data groups of multiple first storage blocks of the first storage chip are output to the input / output bus of the first storage chip within one bit switching cycle, and the bit switching cycle includes 2 N phases. The data group of each first storage block of the first storage chip is output to the input / output bus of the first storage chip at a corresponding phase of the bit switching cycle, and the period of a clock signal of the first storage chip is equal to the bit switching cycle divided by 2 N -1 , and N is an integer not less than 1; and Wherein data groups of the multiple second storage blocks of the second storage chip are output to the input / output bus of the second storage chip within the bit switching cycle, a data group of each second storage block of the second storage chip is output to the input / output bus of the second storage chip at a corresponding phase of the bit switching cycle, and the period of a clock signal of the second storage chip is equal to the period of the clock signal of the first storage chip.