A computer memory array employing a memory bank and an integrated serializer / deserializer circuit to support serialization / deserialization of read / write data in burst read / write mode, and related methods

The integrated serializer/deserializer circuit in high-density memory arrays addresses bandwidth and power consumption issues by enabling simultaneous operations across multiple banks and sub-banks, enhancing performance through serialized data handling and dedicated bit lines.

JP2025520284APending Publication Date: 2025-07-03MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024569368
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-23
Filing Date
2023-05-11
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing high-density memory arrays face challenges in maintaining bandwidth while reducing power consumption and access latency, as dividing memory into banks and sub-banks can lead to glitches and reduced bandwidth due to frequent switching between banks.

Method used

Implementing an integrated serializer/deserializer circuit that converts parallel data streams into serialized streams for burst operations, allowing simultaneous writing and reading across multiple banks without reducing bandwidth, and using dedicated bit lines for sub-banks to improve access frequency and reduce power consumption.

Benefits of technology

The solution effectively doubles read access bandwidth and maintains write bandwidth by avoiding glitches, while reducing power consumption and access latency, thus optimizing memory performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520284000001_ABST
    Figure 2025520284000001_ABST
Patent Text Reader

Abstract

A computer memory array employing a memory bank and an integrated serializer / deserializer circuit to support serialization / deserialization of read / write data in burst read / write mode, and related methods, are disclosed. The memory array can include a serialization circuit configured to convert a parallel data stream of read data received from separately switched memory banks into a single serialized read data stream in burst read mode. The memory array can also include a deserialization circuit configured to convert a received serialized write data stream on an input bus for write operations into separate parallel write data streams to be simultaneously written to the memory banks in burst write mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a computer memory system, and more particularly, to a high-density memory array that is divided into separate memory banks.

Background Art

[0002] A processor-based system includes a memory system to support read and write operations from a central processing unit (CPU) or other processor. Memory can be used not only for data storage but also for storing program code for storing instructions to be executed. Such a processor-based system has conventionally employed both cache memory and non-cache memory, and may also be referred to as "main memory" or "system memory". For example, a CPU can access an on-chip local private cache memory. In a processor-based system, multiple CPUs can also access a shared cache memory. A processor-based system also employs a main memory or system memory that includes memory storage units (i.e., memory bit cells) across the entire physical address space of the processor-based system. Each of these different types of memory generally employs a memory array including memory bit cells organized in a matrix structure for storing data. Memory rows including memory bit cells in each column are accessed to read data words from the memory or to write data words. Memory bit cells can be provided with various memory technologies such as static random access memory (RAM) (SRAM) bit cells and dynamic RAM (DRAM) bit cells.

[0003] In the memory of a processor-based system, it is becoming increasingly important to be able to provide a larger density of memory arrays with increased bandwidth (i.e., reduced access latency). However, as memory density increases, it can potentially degrade power, performance, and / or area (PPA) requirements. As memory density increases, it consumes more semiconductor die area than a smaller memory employing the same memory bit cell technology. Since the overall access latency of the memory is based on the access time to the memory bit cell that is furthest from the support side access circuitry, as memory density increases, the access latency can also increase compared to a smaller cache memory. Also, as memory density increases, it has bit lines of extended length coupled to the support side access circuitry (e.g., read sense circuitry) to reach the increased number of memory row circuits of the memory array. Increasing the length of the bit line increases the capacitance on the bit line and thus increases the access latency. The manufacturing design rules and associated manufacturing processes may limit the overall length of the bit lines in the memory array, thus effectively limiting the density of the memory array despite the trade-off of allowing an increase in access latency.

[0004] One way to improve the read and write access performance of a high density memory array is to increase the size of the transistors of its memory bit cells. Larger transistors support a larger gate voltage, resulting in a larger drive current and enabling the bit lines to be charged and discharged more quickly for read and write access. However, increasing the size of the transistors within the memory bit cells of the memory array increases the area of the overall memory array. Larger sized transistors also have more leakage current, increasing the overall power consumption by the memory array.

[0005] Alternatively, to improve the PPA of a memory with a larger density, the memory array may be divided into a plurality of memory banks of a smaller size. For example, FIG. 1 shows a memory array 100 including first and second memory sub-arrays 102(1) and 102(2), which are two memory sub-arrays. The second memory sub-array 102(2) has the same structure as the first memory sub-array 102(1), although it is shown in more detail in FIG. 1. Taking the second memory sub-array 102(2) as an example, to reduce bit errors, the second memory sub-array 102(2) is further divided into two memory banks 104(1) and 104(2) in an interleaved bank arrangement. In this example, the first memory bank 104(1) is configured to store data with odd memory addresses, and the second memory bank 104(2) is configured to store data with even memory addresses. The total memory capacity of the small-sized memory banks 104(1) and 104(2) is designed to be the desired overall memory capacity for the memory sub-array 102(2). By dividing the memory sub-array 102(2) into separate memory banks 104(1) and 104(2), the distance between the outermost memory bit cells in each memory bank 104(1) and 104(2) and the local access circuit 106 (e.g., column multiplexing circuit and column sense amplifier) is shortened, and the access latency can be reduced. Also, each memory bank 104(1) and 104(2) has dedicated bit lines that are shorter in length than if the memory sub-array 102(2) were not divided into separate memory banks 104(1) and 104(2). By shortening the length of the bit lines, the capacitance of the bit lines is reduced, and as a result, the charging time of the bit lines is shortened, reducing the memory access latency and power consumption. The memory banks 104(1) and 104(2) share the local access circuit 106 to save area and power consumption, but in return, only one of the memory banks 104(1) and 104(2) can be accessed at a given time.

[0006] Each memory bank 104(1), 104(2) of the memory sub-array 102(2) in FIG. 1 can be designed to be accessed at a lower frequency in order to further reduce power. However, in order to avoid data contention on the common input bus 110 and output bus 112, only one of the two memory banks 104(1), 104(2) within the memory sub-array 104(2) can be accessed at a time by the global input / output (I / O) circuit 108. The memory controller may be configured to control the global control circuit 114 to control the input multiplexer 116 and transfer the write data WD from the input bus 110 to the selected memory banks 104(1), 104(2) in a "ping-pong" manner to control. In this way, a two-fold access frequency can be effectively achieved, which is equivalent to a non-bank type memory sub-array operating at twice the lower frequency. However, although the shared input bus 110 for the memory banks 104(1), 104(2) only writes the write data WD to the selected active memory banks 104(1), 104(2), the memory address is received by each memory bank 104(1), 104(2), thereby activating the circuits of both the active and non-active memory banks 104(1), 104(2), which will increase power consumption. Also, when controlling the input multiplexer 116 to switch the writing of the write data WD between different memory banks 104(1), 104(2), write glitches of the writing memory may occur. By adding a delay to the timing of the write access to the memory banks 104(1), 104(2) and the switching of the input multiplexer 116 for transferring the write data WD from the input bus 110 to the selected memory banks 104(1), 104(2), the write glitches of the writing memory can be avoided, but at the cost of a reduction in the write bandwidth of the memory sub-array 102(2).

[0007] The memory controller may be configured to control the global control circuit 114 to control the output multiplexer 118 and control the transfer of the read data RD from the selected memory banks 104(1), 104(2) to the output bus 112. Since only one of the memory banks 104(1), 104(2) can be selected at a time, only the read data RD from the selected memory banks 104(1), 104(2) can be asserted on the output bus 112 at a given point in time. The global control circuit 114 may be configured to further request the read data RD from each of the memory banks 104(1), 104(2) in a ping-pong manner. However, read memory glitches may occur when controlling the output multiplexer 118 to switch the transfer of the read data RD from different memory banks 104(1), 104(2). By adding a delay to the timing of the read access to the memory banks 104(1), 104(2) and the switching of the output multiplexer 118 for asserting the read data RD from the corresponding memory banks 104(1), 104(2) on the output bus 112, the read memory glitches can be avoided, but at the cost of a reduction in the read bandwidth of the memory subarray 102(2).

[0008] Another way to reduce power consumption in the memory sub-array 102(2) that is divided into two memory banks 104(1) and 104(2) of FIG. 1 is to further divide each of the memory banks 104(1) and 104(2) into a plurality of memory sub-banks. By dividing the memory banks 104(1) and 104(2) into memory sub-banks, dedicated bit lines may be provided for each memory sub-bank. Thus, the length of the bit lines of the inner memory sub-banks located closer to the local access circuit 106 is further shorter than when the memory bank is not divided into memory sub-banks. Thereby, the memory sub-banks can be accessed at a higher frequency rate, and the access performance can be improved. However, a part of the power consumption savings realized by the sub-banking of the memory is offset by the increased power of the data input bus 110 that conveys such write data WD for the accessed memory banks 104(1) and 104(2), even though the write data WD can be written only to one of the memory sub-banks within the memory banks 104(1) and 104(2). The power consumption and area may also be increased in the memory array 100 where the memory bank is further divided into memory sub-banks due to the circuit (e.g., a column multiplexing circuit in the local access circuit 106 having more inputs for dedicated bit lines from each memory sub-bank) for managing a more complex addressing scheme for individually addressing each of the memory sub-banks.

[0009] In a memory system, it is desirable to find a method to realize the banking of the memory and / or the sub-banking of the memory in the memory array without incurring the cost of reducing the memory bandwidth, while maintaining the area and reducing the power consumption.

Summary of the Invention

[0010] Exemplary aspects disclosed herein include a computer memory array employing memory banks and an integrated serializer / deserializer circuit to support serialization / deserialization of read / write data in burst read / write mode. Related methods are also disclosed. The memory array is divided into a plurality of memory banks to divide the memory capacity among the plurality of memory banks. For example, the memory array may be divided into two memory banks, and each memory bank is configured to store data at even and odd memory locations, respectively. Banked memory shortens the total length of the bit lines and the distance between the outermost memory bit cells and the access circuit in a given memory bank, improving memory access latency and reducing power consumption. Typically, since only one memory bank within the memory array can be accessed at a time for a read operation to avoid read data contention on the output bus, the memory controller may be configured to switch back and forth between different memory banks for read operations. However, it may be necessary to lower the frequency (i.e., data rate) of each memory bank to avoid read glitches resulting from switching back and forth between each memory bank. In an exemplary aspect, to avoid the need to lower the frequency of the memory banks while avoiding or reducing read glitches and still realizing the advantage of reduced power consumption due to banked memory, the memory array includes a serialization circuit. The serialization circuit is configured to convert a parallel data stream of read data received from separately switched memory banks into a single serialized read data stream provided on the output bus in burst read mode. In an example of a memory array having two memory banks, the serialization circuit can provide a serialized read data stream on the output bus over two consecutive clock cycles in high-bandwidth burst read mode, thereby effectively doubling the output data rate on the output bus and increasing the read access bandwidth (e.g., bits per second) of the memory array.The serialization circuit may further include a circuit that reduces or avoids read glitches on the output bus when switching between memory banks so as to achieve an improvement in the output data rate. The serialization circuit can also be configured to operate in a normal non-burst read mode in which the read data from the addressed memory bank is coupled to the output bus without serializing the read data from multiple memory banks.

[0011] In another exemplary aspect, the memory array also includes a deserialization circuit configured to convert a received serialized write data stream on the input bus for write operations into separate parallel write data streams so as to be simultaneously written to the memory banks in burst write mode. The deserialization circuit can be configured to write each separate data stream to its respective memory bank at half the frequency of the input bus, thereby lengthening the time to switch between memory banks to write the write data to multiple memory banks and reducing or avoiding glitches in the write data. In an example where the memory array includes two memory banks, the deserialization circuit can be configured to store the write data received over two clock cycles in burst write mode. As an example, separate write lines for activating the write operation of each memory bank can be set up by the second clock cycle so that separate write data streams can be simultaneously written to each memory bank. In this way, since the deserialization circuit can write the parallelized write data streams to the memory banks simultaneously, the overall write bandwidth of the memory array does not decrease from the frequency of the input bus. In this way, the write bandwidth of the memory array need not be decreased from the frequency of the input bus in order to write data between the memory banks that are switched within the memory array. The deserialization circuit can also be configured to operate in a normal non-burst write mode in which the received write data is written to only one addressed memory bank at a time.

[0012] In another exemplary aspect, the memory banks within the memory array may be further divided into memory sub-banks in order to reduce access latency and power consumption. By dividing the memory banks into memory sub-banks, each memory sub-bank is provided with a dedicated bit line, thereby increasing the access frequency and improving the access performance. Separate dedicated bit lines for each memory sub-bank are independently coupled to a common access circuit that can be controlled based on the accessed memory sub-bank. Since not all memory bit cells are coupled to the same bit line, the power consumption by the memory bit cells is reduced. In this way, the sub-banking of the memory can further reduce the access latency and lower the dynamic energy consumption. In another exemplary aspect, the bit lines to the outer memory sub-banks within a memory bank can be realized as a fly-out as a flying bit line that extends to a different metal layer than the bit lines within the inner memory sub-banks in order to avoid the need to provide a memory bit cell circuit specialized for each of the memory bit cells within the inner memory sub-banks.

[0013] In this regard, in one exemplary aspect, a memory array is provided. The memory array includes a first output bus, a first memory bank including a first read output, a second memory bank including a second read output, and a first read driver circuit. The first read driver circuit is clocked by a source clock signal. The first read driver circuit is configured to access first read data stored at a first memory address within the first memory bank asserted on the first read output and to access second read data stored at a second memory address within the second memory bank asserted on the second read output. The memory array also includes a serialization circuit. The serialization circuit is configured to assert the first read data on the first output bus in response to a first clock cycle of the source clock signal and to assert the second read data on the first output bus after the assertion of the first read data on the first output bus in response to a second clock cycle of the source clock signal following immediately after the first clock cycle of the source clock signal.

[0014] In this regard, in another exemplary aspect, a method for serializing read data from a plurality of memory banks within a memory array is provided. The method includes receiving a source clock signal. The method also includes accessing first read data stored at a first memory address within a first memory bank in a first memory array based on the source clock signal. The method also includes asserting the first read data on a first read output of the first memory bank. The method also includes accessing second read data stored at a second memory address within a second memory bank in the first memory array. The method also includes asserting the second read data on a second read output of the second memory bank. The method also includes asserting the first read data on a first output bus in response to a first clock cycle of the source clock signal. The method also includes asserting the second read data on the first output bus after the assertion of the first read data on the first output bus in response to a second clock cycle of the source clock signal that immediately follows the first clock cycle of the source clock signal.

[0015] In this regard, in another exemplary aspect, a memory array is provided. The memory array includes a first input bus, a first memory bank, a second memory bank, a first write output coupled to the first memory bank, and a second write output coupled to the second memory bank. The memory array also includes a first write driver circuit clocked by a source clock. The first write driver circuit is configured to assert a first write data stream on the first input bus to begin writing from a first memory address within the memory array. The memory array also includes a deserialization circuit. The deserialization circuit is configured to receive the first write data stream from the first input bus in response to a first clock cycle of the source clock signal. The deserialization circuit is also configured to demultiplex the first write data and the second write data from the first write data stream. The deserialization circuit is also configured to assert the first write data of the first write data stream on the first write output to be written to a first memory address within the first memory bank in response to a second clock cycle of the source clock signal following immediately after the first clock cycle of the source clock signal. The deserialization circuit is also configured to assert the second write data of the first write data stream on the second write output to be written to a second memory address based on the first memory address within the second memory bank in response to the second clock cycle of the source clock signal.

[0016] In this regard, in another exemplary aspect, a method is provided for deserializing write data from an input bus to be written to a plurality of memory banks within a memory array. The method includes receiving a source clock signal. The method also includes asserting, based on the source clock signal, a first write data stream on a first input bus to begin writing from a first memory address within the memory array. The method also includes receiving, in response to a first clock cycle of the source clock signal, the first write data stream from the first input bus. The method also includes demultiplexing first write data and second write data from the first write data stream. The method also includes, in response to a second clock cycle of the source clock signal that immediately follows the first clock cycle of the source clock signal, asserting, on a first write output coupled to the first memory bank within the memory array, the first write data of the first write data stream to be written to the first memory address within the first memory bank. The method also includes, in response to the second clock cycle of the source clock signal, asserting, on a second write output coupled to the second memory bank within the memory array, the second write data of the first write data stream to be written to a second memory address based on the first memory address within the second memory bank.

[0017] Those skilled in the art will understand the scope of the present disclosure and its further aspects upon reading the following detailed description of the preferred embodiments in conjunction with the accompanying drawings. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate some aspects of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

[0019] Exemplary aspects disclosed herein include a computer memory array employing memory banks and an integrated serializer / deserializer circuit to support serialization / deserialization of read / write data in burst read / write mode. Related methods are also disclosed. The memory array is divided into a plurality of memory banks to divide the memory capacity among the plurality of memory banks. For example, the memory array may be divided into two memory banks, and each memory bank is configured to store data at even and odd memory locations, respectively. Banked memory shortens the total length of the bit lines and the distance between the outermost memory bit cells and the access circuit in a given memory bank, improving memory access latency and reducing power consumption. Typically, only one memory bank within the memory array can be accessed at a time for the operations of a read operation to avoid read data contention on the output bus, so the memory controller may be configured to switch back and forth between different memory banks for read operations. However, it may be necessary to lower the frequency (i.e., data rate) of each memory bank to avoid read glitches resulting from the back-and-forth switching between each memory bank. In an exemplary aspect, to avoid the need to lower the frequency of the memory banks while avoiding or reducing read glitches and still realizing the advantage of reduced power consumption due to banked memory, the memory array includes a serialization circuit. The serialization circuit is configured to convert a parallel data stream of read data received from separately switched memory banks into a single serialized read data stream provided on the output bus in burst read mode. In an example of a memory array having two memory banks, the serialization circuit can provide a serialized read data stream on the output bus over two consecutive clock cycles in high-bandwidth burst read mode, thereby effectively doubling the output data rate on the output bus and increasing the read access bandwidth (e.g., bits per second) of the memory array.The serialization circuit can further include a circuit that reduces or avoids read glitches on the output bus when switching between memory banks, enabling an improvement in the output data rate. The serialization circuit can also be configured to operate in a normal non-burst read mode in which the read data from the addressed memory bank is coupled to the output bus without serializing the read data from multiple memory banks.

[0020] In another exemplary aspect, the memory array also includes a deserialization circuit configured to convert a received serialized write data stream on the input bus for write operations into separate parallel write data streams to be simultaneously written to the memory banks in burst write mode. The deserialization circuit can be configured to write each separate data stream to its respective memory bank at half the frequency of the input bus, thereby lengthening the time to switch between memory banks to write the write data to multiple memory banks and reducing or avoiding write glitches. In an example where the memory array includes two memory banks, the deserialization circuit can be configured to store the write data received over two clock cycles in burst write mode. As an example, separate write lines for activating the write operations of each memory bank can be set up by the second clock cycle to enable separate write data streams to be simultaneously written to each memory bank. In this way, the deserialization circuit can cause the parallelized write data streams to be simultaneously written to the memory banks, so that the overall write bandwidth of the memory array does not decrease from the frequency of the input bus. In this way, the write bandwidth of the memory array need not be decreased from the frequency of the input bus to write data between the memory banks that are switched within the memory array. The deserialization circuit can also be configured to operate in a normal non-burst write mode in which the received write data is written to only one addressed memory bank at a time.

[0021] In this regard, FIG. 2 is a diagram of an exemplary processor-based system 100 that includes a processor 202 and a memory system 208. Beginning with FIG. 4 and discussed in more detail below, the memory system 208 may include a memory array that includes an integrated serialization circuit configured to convert parallel data streams of read data received from separately switched memory banks into a single serialized read data stream provided on an output bus in burst read mode. Also, beginning with FIG. 4 and discussed in more detail below, the memory system 208 may include a memory array that also includes a deserialization circuit configured to convert a received serialized write data stream on an input bus for write operations into separate parallel write data streams to be simultaneously written to the memory banks in burst write mode. Before such exemplary serialization and deserialization circuits, first, the exemplary processor-based system 200 of FIG. 2 having such a memory system 208, as well as the exemplary memory array of FIG. 3 that does not include an integrated serialization circuit and a deserialization circuit, are described below.

[0022] In this regard, referring to FIG. 2, the processor 202 of the processor-based system 200 includes one or more respective CPUs 204(0) to 204(N), where "N" is a positive integer representing the number of CPUs included in the processor 202. The processor 202 may be packaged in an integrated circuit (IC) chip 206. The CPUs 204(0) to 204(N) within the processor 202 are configured to issue memory requests (i.e., data read requests and data write requests) to the memory system 208. The memory system 208 includes a cache memory system 210 and a system memory 212. The system memory 212 is a memory that can be fully addressed by the physical address (PA) space of the processor-based system 200. For example, the system memory 212 may be a dynamic random access memory (DRAM) provided in separate DRAM chips. The cache memory system 210 within the memory system 208 includes one or more cache memories 214(1) to 214(X), where "X" is a positive integer representing the number of cache memories included in the processor 202. The cache memories 214(2) to 214(X) may exist at different levels within the processor-based system 200 and are logically arranged between the CPUs 204(0) to 204(N) and the system memory 212. The memory controller 216 controls access to the system memory 212. For example, the CPUs 204(0) to 204(N) as the requesting devices can issue a data request 218 for reading data in response to processing a load instruction. The data request 218 includes the target address of the data to be read from the memory. Taking the CPU 204(0) as an example, if the requested data is not in the private cache memory 214(1) which can be considered as a level 1 (L1) cache memory (i.e., a cache miss for the cache memory 214(1)), the private cache memory 214(1) sends the data request 218, via the interconnect bus 220 in this example, to a shared cache memory 214(X) which can be a level 3 (L3) cache memory shared by all of the CPUs 204(0) to 204(N).The requested data within the data request 218 is ultimately satisfied either in the cache memories 214(1) to 214(X) or, if not contained in any of the cache memories 214(1) to 214(X), in the system memory 212.

[0023] The cache memories 214(1) to 114(X) and / or the system memory 212 within the memory system 208 of FIG. 2 can employ a memory array including memory bit cells generally organized in a matrix structure to store data. Memory rows containing memory bit cells in each column are accessed to read data words from the memory or to write data words to the memory. The memory bit cells can be provided with various memory technologies such as static random access memory (RAM) (SRAM) bit cells and dynamic RAM (DRAM). As another example, the system memory 112 may be implemented with multiple large-sized memory arrays using multiple memory banks to improve memory performance and power consumption. By providing multiple memory banks in the memory within the memory system 208, the memory capacity can be divided among different arrays that can be independently accessed by the CPUs 204(1) to 204(N). Also, providing separate memory banks in the memory system 208 means that the power consumption can be managed independently for each memory bank from other memory banks. Thus, for example, if software or data is persistent in one part of the memory but not in another, these separate parts of the memory may be divided into separate memory banks, thereby saving power without affecting other parts that may be fully powered for memory access while each is separately powered down or put into an idle state.

[0024] FIG. 3 is an exemplary memory system 300 that may be provided in the memory of the memory system 208 within the processor-based system 200 of FIG. 2. The memory system 300 includes a memory array 302 that is a four-column multiplexed (CM4) memory array. It should be noted that the memory array 302 can also be a memory sub-array that divides the memory array. In this example, the memory array 302 is not further divided into memory sub-arrays. The memory array 302 is divided into a first memory bank 304(1) and a second memory bank 304(2), and can store data at odd and even memory positions, respectively. A memory bank is a local unit of memory storage (e.g., memory bit cells), and its read and write accesses are controlled by a memory controller. A memory bank includes a memory bit cell array and an access circuit (e.g., a write driver, a sense amplifier, a column multiplexer, a charge circuit, and a write assist circuit) used to address the bit cell array for read and write operations. When the memory is divided into multiple memory banks, the multiple memory banks usually share some common support-side access circuits, so usually only one memory bank can be accessed at a time to avoid data contention. A memory bank may be further divided into multiple memory sub-banks. A memory sub-bank is a division of memory bit cells from a memory bank. Memory sub-banks within a memory bank share a common access circuit with another memory sub-bank within that memory bank. Thus, at a given point in time, only one memory sub-bank within a memory bank may be accessible.

[0025] In this regard, referring to FIG. 3, the first memory bank 304(1) includes a plurality of first memory row circuits 306(0) to 306(X), each of which includes a plurality of first memory bit cell circuits 308(0) to 308(X). "X + 1" is the number of the memory row circuits 306(0) to 306(X). Each of the memory row circuits 306(0) to 306(X) includes a respective set of the memory bit cell circuits 308(0) to 308(X). The sets of the memory bit cell circuits 308(0) to 308(X) in each of the memory row circuits 306(0) to 306(X) within the first memory bank 304(1) are arranged into memory column circuits 310(0) to 310(Y). Thus, each set of the memory bit cell circuits 308(0) to 308(X) includes "Y + 1" memory bit cells. The first memory bank 304(1) has interleaved memory column circuits 310(0) to 310(Y) configured to store interleaved data words A, B, C, D according to interleaved memory column circuits labeled A1, B1, C1, D1, ..., A4, B4, C4, D4. In this regard, the data words A1 to A4 are interleaved among the memory column circuits 310(0) to 310(Y) corresponding to the memory column circuits A1, A2, A3, A4. By interleaving the storage of data words within the memory array, the bit error rate (BER) can be reduced.

[0026] Similar to the first memory bank 304(1), the second memory bank 304(2) includes a plurality of first memory row circuits 316(0) to 316(X), each of which includes a plurality of second memory bit cell circuits 318(0) to 318(X). Each memory row circuit 316(0) to 316(X) includes a respective set of memory bit cell circuits 318(0) to 318(X). The sets of memory bit cell circuits 318(0) to 318(X) within each memory row circuit 316(0) to 316(X) in the second memory bank 304(2) are organized into memory column circuits 320(0) to 320(Y), and thus each set of memory bit cell circuits 318(0) to 318(X) includes "Y + 1" memory bit cells. The second memory bank 304(2) has interleaved memory column circuits 320(0) to 320(Y) configured to store interleaved data words E, F, G, H according to interleaved memory column circuits labeled E1, F1, G1, H1, ..., E4, F4, G4, H4. In this regard, data words E1 to E4 are interleaved among the memory column circuits 320(0) to 320(Y) corresponding to memory column circuits E1, E2, E3, E4.

[0027] Note that since the first and second memory banks 304(1), 304(2) are coupled to the shared memory column access circuits 314(0) to 314(3), only one of the first and second memory banks 304(1), 304(2) is accessed at a time.

[0028] When the first memory bank 304(1) is accessed (i.e., addressed) in response to a memory operation (read or write operation), the word line WL1 is activated for the selected memory row circuits 306(0) to 306(X) according to the decoded memory address 330 on the input bus 332 for the memory access operation. In a memory write operation, the write data 334 from the input bus 332 is coupled to the bit lines BL of the selected memory row circuits 306(0) to 306(X) and written to the corresponding memory bit cell circuits 308(0) to 308(X) for the selected memory row circuits 306(0) to 306(X). In a memory read operation, a column selection CS1 is generated for each of the plurality of column multiplexer circuits 312(0) to 312(3) coupled to the respective bit lines coupled to each of the memory bit cell circuits 308(0) to 308(X) in the respective memory column circuits 310(0) to 310(Y) representing the interleaved bits from the selected memory row circuits 306(0) to 306(X). Each column multiplexer circuit 312(0) to 312(3) couples one of the coupled bit lines BL from its coupled memory column circuits 310(0) to 310(Y) and provides the corresponding bit to the respective memory column access circuits 314(0) to 314(3) (e.g., sense amplifiers). The memory column access circuits 314(0) to 314(3) provide the read data word 336 from the multiplexed bit lines BL in the first memory bank 304(1) onto the shared output bus 338. In this way, the column multiplexer circuits 312(0) to 312(3) are controlled to multiplex the selected bits from the interleaved data words in the selected memory row circuits 306(0) to 306(X) to the respective memory column access circuits 314(0) to 314(3).For example, in a memory read operation, when it is desired to select interleaved data words A1 to A4 from the selected memory row circuits 306(0) to 306(X), the column multiplexer circuits 312(0) to 312(3) are controlled by the column selection CS1 and multiplex the bits A1 to A4 on the respective bit lines BL1 to BL4 from the selected memory row circuits 306(0) to 306(X) to the respective memory column access circuits 314(0) to 314(3). Therefore, in this example, the first memory bank 304(1) is configured in a 4-bit column multiplexing (CM4) arrangement.

[0029] When the second memory bank 304(2) is accessed (i.e., addressed) in response to a memory operation (read or write operation), the word line WL2 is activated for the selected memory row circuits 316(0) to 316(X) according to the decoded memory address 330 on the input bus 332 for the memory access operation. In a memory write operation, the write data 334 from the input bus 332 is coupled to the bit lines BL of the selected memory row circuits 316(0) to 316(X) and written to the corresponding memory bit cell circuits 318(0) to 318(X) for the selected memory row circuits 316(0) to 316(X). In a memory read operation, a column selection CS2 is generated for a plurality of column multiplexer circuits 322(0) to 322(3) coupled to respective bit lines coupled to each of the memory bit cell circuits 318(0) to 318(X) in each of the memory column circuits 320(0) to 320(Y) representing the interleaved bits from the selected memory row circuits 316(0) to 316(X). Each column multiplexer circuit 322(0) to 322(3) couples one of the coupled bit lines from its coupled memory column circuits 320(0) to 320(Y) and provides the corresponding bit shared with the first memory bank 304(1) to the respective memory column access circuits 314(0) to 314(3) (e.g., sense amplifiers). In this way, the column multiplexer circuits 322(0) to 322(3) are controlled to multiplex the selected bits from the interleaved data words in the selected memory row circuits 316(0) to 316(X) to the respective memory column access circuits 314(0) to 314(3). The memory column access circuits 314(0) to 314(3) provide the read data word 336 from the multiplexed bit lines BL of the second memory bank 304(2) onto the shared output bus 338.For example, in a memory read operation, when it is desired to select interleaved data words E1 to E4 from the selected memory row circuits 316(0) to 316(X), the column multiplexer circuits 322(0) to 322(3) are controlled by the column selection CS1 and multiplex the bits E1 to E4 on the respective bit lines BL5 to BL8 from the selected memory row circuits 316(0) to 316(X) to the respective memory column access circuits 314(0) to 314(3). Therefore, the second memory bank 304(2) is also configured in the CM4 arrangement in this example.

[0030] By dividing the memory array 302 into the first and second memory banks 304(1) and 304(2), the distance between the outermost memory bit cell circuits (for example, the memory bit cell circuits 308(X) and 308(X) of each memory bank 304(1) and 304(2)) and each column multiplexer circuit 312(1) to 312(4), 322(1) to 322(4) and the shared memory column access circuits 314(0) to 314(3) is shortened, and the access latency can be reduced. Also, each memory bank 304(1) and 304(2) has dedicated bit lines BL with a shorter length than when the memory array 302 is not divided into separate first and second memory banks 304(1) and 304(2). By shortening the length of the bit lines, the capacitance of the bit lines is reduced, and as a result, the charging time of the bit lines is shortened, reducing the memory access latency and power consumption. The first and second memory banks 304(1) and 304(2) share the memory column access circuits 314(0) to 314(3) to save area and power consumption, but in return, only one of the first and second memory banks 304(1) and 304(2) can be accessed at a given point in time.

[0031] Each memory bank 304(1), 304(2) of the memory array 302 in FIG. 3 can be designed to be accessed at a lower frequency in order to further reduce power. However, in order to avoid data contention on the shared input bus 332 and output bus 338, only one of the first and second memory banks 304(1), 304(2) can be accessed at a time. The memory controller may be configured to control the transfer of the write data 334 from the input bus 332 to the selected memory bank 304(1), 304(2) in a "ping-pong" manner. Thereby, a two-fold access frequency can be effectively realized, which is equivalent to a non-bank type memory sub-array operating at twice the lower frequency. However, the shared input bus 332 for the first and second memory banks 304(1), 304(2) increases power consumption because the decoded memory address 330 is received by each of the first and second memory banks 304(1), 304(2), thereby activating the circuits of both the active and non-active memory banks 304(1), 304(2), even though the write data 334 is written only to the selected memory bank 304(1), 304(2) where the write data 334 is active. Also, when controlling the switching of the writing of the write data 334 between the first memory bank 304(1) and the second memory bank 304(2), write memory glitches may occur. By adding a delay to the timing of the write access to the first and second memory banks 304(1), 304(2) and the switching for transferring the write data 334 from the input bus 332 to the selected first or second memory bank 304(1), 304(2), the write memory glitches can be avoided, but at the cost of reducing the write bandwidth of the memory array 302.

[0032] The memory controller may also be configured to control the transfer of the read data 336 from the selected first or second memory bank 304(1), 304(2) to the output bus 338 in a "ping-pong" manner. Since only one memory bank 304(1), 304(2) can be selected at a time, at a given point in time, only the read data 336 from the selected memory bank 304(1), 304(2) can be asserted on the output bus 338. The read memory operation may be controlled to simply request the read data 336 to the first memory bank 304(1) or the second memory bank 304(2) in a ping-pong manner at a time. However, when controlling the assertion of the read data 336 from the different first and second memory banks 304(1), 304(2) to the shared memory column access circuits 314(0) to 314(3), read memory glitches may occur. By adding a delay to the timing of the read access to the first and second memory banks 304(1), 304(2) and the switching to assert the read data 336 from the corresponding first and second memory banks 304(1), 304(2) to the shared memory column access circuits 314(0) to 314(3) so as to be asserted on the output bus 338, the read memory glitches can be avoided, but at the cost of reducing the read bandwidth of the memory array 302.

[0033] In this regard, in order to avoid the need to reduce the frequency of memory banks in the memory array while avoiding or reducing read glitches and still realizing the advantage of power consumption reduction due to memory banking, an exemplary memory array 400 is provided in FIG. 4. The memory array 400 of FIG. 4 is divided into first and second memory sub-arrays 402(1), 402(2). Exemplary details of the second memory sub-array 402(2) are shown in FIG. 4 and will be described later, but such details are also applicable to the first memory sub-array 402(1). The second memory sub-array 402(2) is divided into first and second memory banks 404(1), 404(2), which can be similar to the first and second memory banks 304(1), 304(2) of FIG. 3. In this example, the first odd memory bank 404(1) is configured to store data at odd memory addresses, and the second even memory bank 404(2) is configured to store data at even memory addresses. The first and second memory banks 404(1), 404(2) share access to a local access circuit 406 (e.g., a write driver circuit, a shared memory column access circuit (e.g., a sense amplifier)) to control the assertion of the read data 408R(1), 408R(2) read from the selected memory banks 404(1), 404(2) in the memory read operation to the respective read outputs 410R(1), 410R(2) coupled to the respective memory banks 404(1), 404(2). The local access circuit 406 also controls the transfer of the write data 408W(1), 408W(2) written to the selected memory banks 404(1), 404(2) in the memory write operation to the respective first and second write outputs 410W(1), 410W(2) coupled to the respective memory banks 404(1), 404(2).

[0034] Memory array 400 includes a global control circuit 412 shared between a first memory sub-array 402(1) and a second memory sub-array 402(2). The global control circuit 412 controls the transfer of read data 408R(1), 408R(2) from one of the memory banks 404(1), 404(2) to the shared output bus 414. Since only one memory bank 404(1), 404(2) can be selected at a time, at a given point in time, only the read data 408R(1), 408R(2) from the selected memory bank 404(1), 404(2) can be asserted on the output bus 414R. The global control circuit 412 can also be configured to request the read data 408R(1), 408R(2) from each memory bank 404(1), 404(2) in a ping-pong fashion. As will be discussed in more detail below, in order to avoid or reduce the need to lower the frequency of the memory banks 404(1), 404(2) while avoiding read glitches due to the switching of the read data 408R(1), 408R(2) from the selected memory bank 404(1), 404(2) onto the output bus 414R, the memory array 402 includes a serialization circuit 416S. The serialization circuit 416S is clocked by a read switching clock signal 417, which is either the source clock within the memory array 400 or generated by a serialization clock generation circuit 421 based on the source clock signal 419 in burst read mode. The read switching clock signal 417 can be considered as the source clock signal for the serialization circuit 416S. The source clock signal 419 can be used to clock not only the input and output buses 414W, 414R but also other circuits within the memory array 400 for read and write operations. The serialization circuit 416S is configured to convert the parallel data streams of the read data 408R(1), 408R(2) received from the separately switched first and second memory banks 404(1), 404(2) into a single serialized read data stream 408R asserted on the output bus 414R in burst read mode controlled by a burst detection circuit 418.The read data 408R(1) and 408R(2) are generated as the results of the read driver circuit 415R starting read access at respective memory addresses in the respective memory banks 404(1) and 404(2), and are generated as respective read outputs 410R(1) and 410R(2). The read driver circuit 415R is controlled at a frequency based on the source clock signal 419.

[0035] In an example of the memory sub-array 402(2) having two memory banks 404(1), 404(2), the serialization circuit 416S can provide, in a high-bandwidth burst read mode, a serialized read data stream 408R of the read data 408R(1), 408R(2) onto the output bus 414R over two consecutive clock cycles. Thereby, the output data rate on the output bus 414R is effectively doubled, increasing the read access bandwidth (e.g., bits per second) of the memory sub-array 402(2) and the memory array 400. As will be discussed in more detail below, in order to assert the serialized read data stream 408R on the output bus 414R in burst read mode, the serialization circuit 416S is first configured to assert the first read data 408R(1) on the output bus 414R in response to the rising edge of the read switching clock signal 417 in the first clock cycle. Next, the serialization circuit 416S is configured to sequentially assert the second read data 408R(2) on the output bus 414R following the first read data 408R(1) in response to the rising edge of the read switching clock signal 417 in the second clock cycle following immediately after the first clock cycle. The read switching clock signal 417 can be generated to be twice the frequency of the source clock signal 419 in burst read mode such that the rising edge of the read switching clock signal 417 in the second clock cycle responds to the falling edge of the first clock cycle of the source clock signal 419 in this example. In this way, the deserialization circuit 416S is configured to assert the serialized read data stream 408R of the first and second read data 408R(1), 408R(2) onto the output bus 414R without bubbles in consecutive clock cycles of the read switching clock signal 417 at a bandwidth much higher than the source clock signal 419.

[0036] Also, as will be discussed in more detail below, the serialization circuit 416S within the memory array 400 of FIG. 4 may also include a circuit that reduces or avoids read glitches on the output bus 414R during switching between the memory banks 404(1), 404(2). For example, as will be discussed in more detail below, the serialization clock generation circuit 421 may be configured to generate the read switching clock signal 417 in a manner such that its clock pulses do not overlap. Since the read switching clock signal 417 controls the switching of the deserialization circuit 416S as to which of the received first read data 408R(1) and second read data 408R(2) is transferred to the output bus 414S, by controlling the clock pulses of the read switching clock signal 417 so as not to overlap with a desired margin, before the first read data 408R(1) or the second read data 408R(2) can be latched from the output bus 414R, the first read data 408R(1) or the second read data 408R(2) is output on the output bus 414S in a way that the first read data 408R(1) or the second read data 408R(2) interferes with each other, and the read data glitch caused by the assertion of the first read data 408R(1) or the second read data 408R(2) on the output bus 414S can be reduced or avoided. Thereby, the memory array 400 can achieve an improvement in the output data rate at the output bus 414R.

[0037] The serialization circuit 416S can be held in a burst read mode such that subsequent read data 408R(1), 408R(2) continues to be asserted as a serialized read data stream 408R on the output bus 414R based on the read switching clock signal 417. Also, as will be discussed in more detail below, the serialization circuit 416S can be configured to operate in a normal non-burst read mode when detected by the burst detection circuit 418. In the non-burst read mode, the read data 408R(1), 408R(2) received from the addressed memory banks 404(1), 404(2) in response to the read operation is coupled to the output bus 414R without serializing the read data 408R(1), 408R(2). In this way, the serialization circuit 416S can be configured to assert the received read data 408R(1), 408R(2) on the output bus 414R, for example, based on the source clock signal 419 when received. For example, in the non-burst read mode, the serialization circuit 416S can be configured to assert the received read data 408R(1), 408R(2) on the output bus 414R based on the rising edge of the source clock signal 419, or in the non-burst read mode, based on the rising edge of the read switching clock signal 417 whose frequency is reduced to that of the source clock signal 419 etc. when received.

[0038] Continuing to refer to FIG. 4, and as will be discussed in more detail below, the memory array 400 of this example may also include a deserialization circuit 416D. The deserialization circuit 416D is configured to convert the serialized write data stream 408W received from the common input bus 414W for write operations into separate parallel first and second write data 408W(1), 408W(2) so as to be simultaneously written to the memory banks 404(1), 404(2) in burst write mode. The serialized write data stream 408W to be written is asserted by a write request from the write driver circuit 415W to write the data in the serialized write data stream 408W starting from the specified memory addresses included in the memory banks 404(1), 404(2). The write driver circuit 415W is controlled at a frequency based on the source clock signal 419. The deserialization circuit 416D may be configured to demultiplex the first and second write data 408W(1), 408W(2) from the serialized write data stream 408W. Next, the deserialization circuit 416D asserts the separate write data 408W(1), 408W(2) in the write data stream 408W to the respective memory banks 404(1), 404(2) at half the frequency of the input bus 414W, thereby increasing the time to switch between the memory banks 404(1), 404(2) to write the separate write data 408W(1), 408W(2) to the respective memory banks 404(1), 404(2), and reducing or avoiding write data glitches.

[0039] In an example of the memory sub-array 402(2) including two memory banks 404(1) and 404(2), the deserialization circuit 416D may be configured to store the write data 408W received over consecutive first and second clock cycles of the source clock signal 419 in burst write mode. Alternatively, the deserialization circuit 416D may be configured to store the write data 408W received over consecutive first and second clock cycles of the write clock signal 417W in burst write mode generated by the deserialization clock generation circuit 423. For example, the deserialization clock generation circuit 423 may be configured to generate the write clock signal 417W at the same frequency as the source clock signal 419. In one example, separate write lines for activating the write operations of each memory bank 404(1) and 404(2) may be set up by the second clock cycle of the source clock signal 419 or the write clock signal 417W, and then separate write data 408W(1) and 408W(2) may be simultaneously asserted and written to each memory bank 404(1) and 404(2). For example, the deserialization circuit 416S may be configured to assert the first write data 408W(1) on the first write output 410 to be written to the first memory bank 404(1) in response to the second clock cycle of the source clock signal 419 or the write clock signal 417W following immediately after the first clock cycle of the source clock signal 419 or the write clock signal 417W. The deserialization circuit 416S may also assert the second write data 408W(1) on the second write output 410 to be written to the second memory bank 404(1) and, in parallel with the first write data 408W(1), to be written to each of the respective memory banks 404(1) and 404(2) in response to the second clock cycle of the source clock signal 419 or the write clock signal 417W.In this way, the deserialization circuit 416D enables the serialized write data 408W(1) and 408W(2) to be written into the memory banks 404(1) and 404(2) simultaneously, so the overall write bandwidth of the memory subarray 402(2) and the memory array 400 does not decrease from the frequency of the input bus 414W. In this way, the write bandwidth of the memory subarray 402(2) and the memory array 400 does not need to be decreased from the frequency of the input bus 414W to write the write data 408W(1) and 408W(2) between the switched memory banks 404(1) and 404(2).

[0040] The deserialization circuit 416S can also be maintained in the burst write mode, whereby subsequent write data 408W(1) and 408W(2) from the subsequently received serialized write data stream 408W are continuously converted into parallel first and second write data 408W(1) and 408W(2) so as to be written into the first and second memory banks 404(1) and 404(2) in parallel based on the source clock signal 419. Also, as will be discussed in more detail below, the deserialization circuit 416D can be configured to operate in the normal non-burst write mode when detected by the burst detection circuit 418. In the non-burst write mode, the write data 408W is written only into one addressed memory bank 404(1) or 404(2) at a time. For example, as will be discussed in more detail below, the deserialization circuit 416S can be configured to write the received write data in the write data stream 408W into one memory bank 404(1) or 404(2) at a time when it is received in the non-burst write mode.

[0041] Note that the memory banks 404(1) and 404(2) of the memory array 400 in FIG. 4 can be further divided into memory sub - banks respectively. For example, the memory bank 404(1) may be divided into two memory sub - banks 420(1)(1) and 420(1)(2). The memory bank 404(2) may be divided into two memory sub - banks 420(2)(1) and 420(2)(2). A memory sub - bank is a division of memory bit cells from a memory bank. Memory sub - banks within a memory bank share a common access circuit with another memory sub - bank within that memory bank. Thus, at a given point in time, only one memory sub - bank within a memory bank may be accessible. The serialization circuit 416S and deserialization circuit 416D described above may be configured to serialize read data and parallelize write data for different memory sub - banks 420(1)(1), 420(1)(2), 420(2)(1), 420(2)(2) of the respective memory banks 404(1) and 404(2).

[0042] FIG. 5 is a flowchart showing an exemplary serialization process 500 of a serialization circuit 416S that converts parallel data streams of first and second read data 408R(1) and 408R(2) received from separately switched memory banks 404(1) and 404(4) within the memory array 400 of FIG. 4 into a single serialized read data stream 408R provided on the output bus 414R in burst read mode.

[0043] In this regard, as shown in FIG. 5, the serialization process 500 includes receiving a source clock signal 419 (block 502 in FIG. 5). Receiving the source clock signal 419 may also include generating a read switching clock signal 417R. The serialization process 500 also includes the read driver circuit 415R accessing first read data 408R(1) stored at a first memory address in a first memory bank 404(1) within the first memory array 400 (block 504 in FIG. 5). The serialization process 500 also includes the first memory bank 404(1) asserting the first read data 408R(1) on a first read output 410R(1) of the first memory bank 404(1) (block 506 in FIG. 5). The serialization process 500 also includes the read driver circuit 415R accessing second read data 408R(2) stored at a second memory address in a second memory bank 404(2) within the first memory array 400 (block 508 in FIG. 5). The serialization process 500 also includes the second memory bank 404(2) asserting the second read data 408R(2) on a second read output 410R(2) of the second memory bank 404(2) (block 510 in FIG. 5). The serialization process 500 also includes the serialization circuit 416S asserting the first read data 408R(1) on a first output bus 414R in response to a first clock cycle of the source clock signal 419 (or a read switching clock cycle 417R) (block 512 in FIG. 5). The serialization process 500 also includes the serialization circuit 416S asserting the second read data 408R(2) on the first output bus 414R after the assertion of the first read data 408R(1) on the first output bus 414R in response to a first clock cycle (e.g., a second clock cycle of the source clock signal 419 or a first clock cycle of the read switching clock cycle 417R) (block 514 in FIG. 5).

[0044] As described above with respect to the memory array 400 of FIG. 4, the deserialization circuit 416D can also be provided within the memory array 400 to convert the received serialized write data stream 408W received on the input bus 414W for write operation into separate parallel first and second write data streams 408W(1), 408W(2) so as to be simultaneously written into the respective memory banks 404(1), 404(2) in burst write mode. In this regard, FIG. 6 is a flowchart showing an exemplary deserialization process 600 of the deserialization circuit 416D within the memory array 40 of FIG. 4, which converts the received serialized write data stream 408W received on the input bus 414W for write operation into separate parallel first and second write data streams 408W(1), 408W(2) so as to be simultaneously written into the respective memory banks 404(1), 404(2) in burst write mode.

[0045] In this regard, as shown in FIG. 6, the deserialization process 600 includes receiving a source clock signal 419 (block 602 in FIG. 6). Receiving the source clock signal 419 may also include generating a read switching clock signal 417R. The deserialization process 600 also includes the write driver circuit 415W of FIG. 4 asserting that a first write data stream 408W on a first input bus 414W begins to be written from a first memory address of the memory array 400 (block 604 in FIG. 6). The deserialization process 600 also includes the deserialization circuit 416D receiving the first write data stream 408W from the first input bus 414W in response to a first clock cycle of the source clock signal 419 (block 606 in FIG. 6). The deserialization process 600 also includes the deserialization circuit 416D demultiplexing the first write data stream 408W into a first write data 408W(1) and a second write data 408W(2) (block 608 in FIG. 6). The deserialization process 600 also includes, in response to a second clock cycle of the source clock signal 419 following immediately after the first clock cycle of the source clock signal 419, asserting the first write data 408W(1) of the first write data stream 408W on a first write output 410W(1) coupled to a first memory bank 404(1) within the memory array 400 so as to be written to a first memory address within the first memory bank 404(1) (block 610 in FIG. 6). The deserialization process 600 also includes, in response to the second clock cycle of the source clock signal 419, asserting the second write data 408W(2) of the first write data stream 408W on a second write output 410W(2) coupled to a second memory bank 404(2) within the memory array 400 so as to be written to a second memory address based on a first memory address within the second memory bank 404(2) (block 612 in FIG. 6).

[0046] The serialization circuit 416S and the deserialization circuit 416D may be used together or separately in the memory array 400 of FIG. 4. The serialization circuit 416S and the deserialization circuit 416D may further be implemented in different circuit forms.

[0047] In this regard, FIG. 7 is a circuit diagram of an exemplary serialization circuit 716S that may be provided as the serialization circuit 416S within the memory array 400 of FIG. 4, for example. The serialization circuit 716S of FIG. 7 is discussed in this example with reference to the memory array 400 of FIG. 4. The serialization circuit 716D is configured to serialize the read data 408R(1), 408R(2) from each of the four memory sub-banks 404(1)(1), 404(1)(2), 404(2)(1), 404(2)(2). As shown in FIG. 7, the serialization circuit 716S includes four multiplexer circuits 700(1)(1), 700(1)(2), 700(2)(1), 700(2)(2). The multiplexer circuits 700(1)(1), 700(1)(2) are for serializing the first read data 408R(1) from the respective memory sub-banks 420(1)(1), 420(1)(2) of the first memory bank 404(1). The multiplexer circuits 700(2)(1), 700(2)(2) are for serializing the second read data 408R(2) from the respective memory sub-banks 420(2)(1), 420(2)(2) of the second memory bank 404(2).

[0048] Continuing to refer to FIG. 7, the multiplexer circuits 700(1)(1), 700(1)(2) each include respective clock inputs 702(1)(1), 702(1)(2) configured to receive the read switching clock signal 417R of FIG. 4. The multiplexer circuits 700(2)(1), 700(2)(2) also each include respective clock inputs 702(2)(1), 702(2)(2) configured to receive the read switching clock signal 417R of FIG. 4. The multiplexer circuits 700(1)(1), 700(1)(2) are each configured to be coupled to the bit lines BL and the complement bit lines BLB from the respective memory sub-banks 420(1)(1), 420(1)(2) to receive the first read data 408R(1). The multiplexer circuits 700(2)(1), 700(2)(2) are each configured to be coupled to the bit lines BL and the complement bit lines BLB from the respective memory sub-banks 420(2)(1), 420(2)(2) to receive the second read data 408R(2). In response to the first clock cycle of the read switching clock signal 417R, when accessed, the multiplexer circuits 700(1)(1), 700(1)(2) are configured to pass the first read data 408R(1) from the respective memory sub-banks 420(1)(1), 420(1)(2) of the first memory bank 404(2) to the output 704. In response to the second clock cycle immediately following the read switching clock signal 417R, when accessed, the multiplexer circuits 700(2)(1), 700(2)(2) are configured to pass the second read data 408R(2) from the respective memory sub-banks 420(2)(1), 420(2)(2) of the second memory bank 404(2) to the output 704. FIG. 8 is a signal diagram showing the read switching clock signal 417R generated by the serialization clock generation circuit 421 of FIG. 4. As shown in FIG. 8, in this example, the read switching clock signal 417R is based on twice the frequency of the source clock signal 419.In this example, the read switching clock signal 417R is generated by the serialization clock generation circuit 421 to have non-overlapping pulses so that the multiplexer circuits 700(1)(1), 700(1)(2) and the multiplexer circuits 700(2)(1), 700(2)(2) do not switch at the same time point, or do not switch at a close time that causes glitches in the output 704. In this way, glitches in the first and second read data 408R(1), 408R(2) on the output 704 are reduced or avoided.

[0049] In this way, the multiplexer circuits 700(1)(1), 700(1)(2), 700(2)(1), and 700(2)(2) can be controlled to multiplex the first and second read data 408R(1) and 408R(2) received separately from the first and second memory banks 404(1) and 404(2) into a serialized read data stream 408R on the output 704. Only one of the multiplexer circuits among the multiplexer circuits 700(1)(1) and 700(1)(2) for the memory sub-banks 420(1)(1) and 420(1)(2) becomes active at a time and will have valid first read data 408R(1) from the first memory bank 404(1). The multiplexer circuits 700(1)(1), 700(1)(2), 700(2)(1), and 700(2)(2) multiplex the read data 408R(1) and 408R(2) to the output 704 based on the read switching clock signal 417R in this example. Only one of the multiplexer circuits among the multiplexer circuits 700(2)(1) and 700(2)(2) for the memory sub-banks 420(2)(1) and 420(2)(2) becomes active at a time and will have valid second read data 408R(2) from the second memory bank 404(2). As shown in FIG. 7, the multiplexer circuits 700(1)(1) and 700(1)(2) for the first memory bank 404(1) are configured to respond to the rising edge 800(1) of the first clock cycle 802(1) of the read switching clock signal 417R. The multiplexer circuits 700(2)(1) and 700(2)(2) for the second memory bank 404(4) are configured to respond to the rising edge 800(2) of the second clock cycle 802(2) of the read switching clock signal 417R that follows immediately after the first clock cycle 802(1). Therefore, as described above, thereby, the multiplexer circuits 700(1)(1), 700(1)(2), 700(2)(1), and 700(2)(2) can multiplex the first and second read data 408R(1) and 408R(2) respectively on the output 704 based on the frequency of the read switching clock signal 417R.

[0050] Alternatively, the multiplexer circuits 700(1)(1) and 700(1)(2) for the first memory bank 404(1) may be configured to respond to the rising edge 804(1) of the first clock cycle 806(1) of the source clock signal 419. The multiplexer circuits 700(2)(1) and 700(2)(2) for the second memory bank 404(4) may be configured to respond to the falling edge 804(2) of the first clock cycle 806(1) of the source clock signal 419. In this example, the frequency of the source clock signal 419 is half the frequency of the read switching clock signal 417R. The switching of the multiplexer circuits 700(1)(1) and 700(1)(2) based on the read switching clock signal 417R can be considered to be based on the source clock signal 419. Thus, as described above, thereby, the multiplexer circuits 700(1)(1), 700(1)(2), 700(2)(1), and 700(2)(2) can multiplex the first and second read data 408R(1) and 408R(2) respectively on the output 704 based on the frequency of the read switching clock signal 417R.

[0051] Continuing to refer to FIG. 7, the output 704 is coupled to a read latch circuit 706 that latches either the first or second read data 408R(1) or 408R(2). In this example, the read latch circuit 706 is a cross-coupled inverter circuit 708(1) and 708(2) that holds the data on the output 704 until the data on the output 704 changes. Additional buffer circuits 712(1) and 712(2) (e.g., two serially-coupled inverter circuits) may be coupled to the output 704 of the read latch circuit 706 for the purpose of adding delay and / or achieving a voltage domain shift to provide the multiplexed read data 408R(1) and 408R(2) as a serialized read data stream 408R to the output bus 414R shown in FIG. 4.

[0052] Note that the description of the multiplexer circuits 700(1)(1), 700(1)(2), 700(2)(1), and 700(2)(2) of the serialization circuit 716S in FIG. 7 above relates to the case where the memory array 400 is in the burst read mode. The burst read mode can be detected by a burst detection circuit such as the burst detection circuit 418 in FIG. 4 as an example and transmitted to the serialization circuit 716S. As described above, the multiplexer circuits 700(1)(1), 700(1)(2), 700(2)(1), and 700(2)(2) of the serialization circuit 716S in FIG. 7 may also be configured in the non-burst read mode. In the non-burst read mode, the read data 408R(1) and 408R(2) received from the memory banks 404(1) and 404(2) specified by the address in response to the read operation are coupled to the output bus 414R without necessarily being serialized in consecutive clock cycles of the read switching clock signal 417D or actually the source clock signal 419. In the non-burst read mode, the serialization circuit 416S can be configured to assert the read data 408R(1) and 408R(2) on the output bus 414R when received, for example, based on the source clock signal 419. For example, in the non-burst read mode, the serialization circuit 416S can be configured to assert the read data 408R(1) and 408R(2) on the output bus 704 when received based on the rising edge of the source clock signal 419 or, in the non-burst read mode, the rising edge of the read switching clock signal 417 whose frequency is reduced to that of the source clock signal 419 or the like.

[0053] FIG. 9 is a circuit diagram of an exemplary deserialization circuit 916D provided in the memory array 900. The memory array 900 may be included in the memory array 400 of FIG. 4. The deserialization circuit 916D will be described with reference to FIGS. 4 and 9. The deserialization circuit 916D multiplexes the received serialized write data stream 408W for write operation into separate parallel first and second write data 408W(1), 408W(2) so as to be simultaneously written into separate memory banks 404(1), 404(2) in burst write mode.

[0054] As shown in FIG. 9, the first memory bank 404(1) includes "X + 1" memory row circuits 906(0) to 906(X), each including a plurality of memory row bit cells 908(0)(0) to 908(X)(Y) organized in respective memory column circuits 910(0) to 910(Y). In this example, the memory row bit cells 908(0)(0) to 908(X)(Y) are SRAM bit cells that support bit lines BL and complementary bit lines BLB. The second memory bank 404(2) includes "X + 1" memory row circuits 916(0) to 916(X), each including a plurality of memory row bit cells 918(0)(0) to 918(X)(Y) organized in respective memory column circuits 920(0) to 920(Y). In this example, the memory row bit cells 918(0)(0) to 918(X)(Y) are SRAM bit cells that support bit lines BL and complementary bit lines BLB. The memory banks 404(1) and 404(2) are further divided into separate memory sub-banks 420(1)(1) to 420(1)(2), 420(2)(1) to 420(2)(2) as previously described in FIG. 4 above. In this example, the memory row bit cells 908(0)(0) to 908(X)(Y), 918(0)(0) to 918(X)(Y) within each of the separate memory sub-banks 420(1)(1) to 420(1)(2), 420(2)(1) to 420(2)(2) share separate bit lines BL and complementary bit lines BLB along the memory sub-bank boundaries. The outer memory sub-banks 420(1)(2) and 420(2)(2) may employ separate "floating" bit lines BL and BLB, as described in another example of FIG. 11 that may be employed in the memory array 900 of FIG. 9. The bit lines BL and complementary bit lines BLB for the first memory bank 404(1) are coupled to the first column multiplexer circuit 930(1). The bit lines BL and complementary bit lines BLB for the second memory bank 404(2) are coupled to the second column multiplexer circuit 930(2).The write driver circuits 415W(1) and 415W(2) are provided for each of the memory banks 404(1) and 404(2), and write the respective write data 408W(1) and 408W(2) at each memory address designated for write operation to each of the memory banks 404(1) and 404(2) of the memory row circuits 906(0) to 906(X) and 916(0) to 916(X). The write data 408W(1) and 408W(2) written to the memory array 900 are received by the write data stream 408W.

[0055] Continuing to refer to FIG. 9, in the burst write mode, as described above with respect to FIG. 4, a deserialization circuit 916D is provided to deserialize the write data stream 408W into respective separate parallel first and second write data 408W(1) and 408W(2) that can be simultaneously written to separate memory banks 404(1) and 404(2). The deserialization circuit 416D of FIG. 4 can be the deserialization circuit 916D of FIG. 9. In this example, the deserialization circuit 416D includes first and second write latch circuits 932(1) and 932(2). The first and second write latch circuits 932(1) and 932(2) may be, for example, flip-flops. The first and second write latch circuits 932(1) and 932(2) are clocked by respective write clock signals 417W(1) and 417W(2) generated by a deserialization clock generation circuit 423 based on a source clock signal 419, as shown in the signal diagram of FIG. 10. Both the first and second write latch circuits 932(1) and 932(2) are coupled to an input bus 414W that carries the write data stream 408W. As an example, assuming that the write data stream 408W consists of data D0 and D1, as shown by the Din signal in FIG. 10, the first write latch circuit 932(1) is configured to latch the first write data 408W(1)D0 in response to the first clock cycle 434(1) of the write clock signal 417W(1), also as shown in FIG. 10. The second write latch circuit 932(2) is configured to latch the second write data 408W(2)D1 in response to the second clock cycle 434(2) of the write clock signal 417W(2), also as shown in FIG. 10. The second clock cycle 434(2) of the write clock signal 417W(2) follows immediately after the first clock cycle 434(1) of the write clock signal 417W(1).In this way, the first and second write latch circuits 932(1) and 932(2) are configured to sequentially store the write data when their respective first and second write data 408W(1) and 408W(2) arrive at the write data stream 408W on the input bus 414W.

[0056] The deserialization circuit 916D, and more specifically its first and second latch circuits 932(1), 932(2), are configured to assert the multiplexed and latched first and second write data 408W(1), 408W(2) to their respective multiplexer circuits 936(1), 936(2). The first and second clock cycles 434(1), 434(2) of the write clock signals 417W(1), 417W(2) can be controlled to equalize the clock to q (a certain amount) (data) minimum and maximum setup times and minimum and maximum latch times to avoid glitches in the first and second write data 408W(1), 408W(2) for their respective multiplexer circuits 936(1), 936(2). The deserialization circuit 916D has the function of asserting the first write data 408W(1) and writing it to either the first memory bank 404(1) or the second memory bank 404(2), and also has the function of asserting the second write data 408W(2) and writing it to either the first memory bank 404(1) or the second memory bank 404(2). For example, if the memory bank 404(1) is configured to store data for odd memory addresses and the memory bank 404(2) is configured to store data for even memory addresses, in this example the multiplexer circuits 936(1), 936(2) may be controlled to direct the first write data 404W(1) to the first memory bank 404(1) when the data is written to an odd memory address or to the second memory bank 404(2) when the data is written to an even memory address. Similarly, the multiplexer circuits 936(1), 936(2) may be controlled to direct the second write data 404W(2) to the first memory bank 404(1) when the data is written to an odd memory address or to the second memory bank 404(2) when the data is written to an even memory address in this example.

[0057] Note that it should be noted that the multiplexer circuits 936(1) and 936(2) in FIG. 9 are optional. In this example, the multiplexer circuits 936(1) and 936(2) are provided so that the deserialization circuit 916D has a function of asserting the first write data 408W(1) to either the first memory bank 404(1) or the second memory bank 404(2) and writing it there. Also, as described above, the multiplexer circuits 936(1) and 936(2) enable the function of asserting the second write data 408W(2) to either the first memory bank 404(1) or the second memory bank 404(2) and writing it there. However, it should be noted that the deserialization circuit 916D does not require the inclusion of the multiplexer circuits 936(1) and 936(2). The deserialization circuit 916D can be configured to write the first write data 408W(1) to the first memory bank 404(1) and the second write data 408W(2) to the second memory bank 404(2), or vice versa.

[0058] As shown in the signal diagram of FIG. 10, in this example, the first and second write latch circuits 932(1) and 932(2) of the deserialization circuit 916D assert the clock cycles of the first and second write data 408W(1) and 408W(2) to the multiplexer circuits 936(1) and 936(2), and are configured to be written to both memory banks 404(1) and 404(2) simultaneously. This is achieved by activating the write lines WL of both memory banks 404(1) and 404(2) in response to the second clock cycle 434(2) of the write clock signal 417W(2). The write lines are shown as WL1 and WL2 for each memory bank 404(1) and 404(2), but it should be noted that only one write line can be activated at a time for each memory bank 404(1) and 404(2) to select the respective memory row circuits 906(0) to 906(X) and 916(0) to 916(X) where the first and second write data 408W(1) and 408W(2) are written. In this way, both the data D0 for the first write data 408W(1) and the data D1 for the second write data 408W(2) are stored in the respective first and second write latch circuits 932(1) and 932(2). Then, in burst write mode, both the latched data D0 of the first write data 408W(1) and the latched data D1 of the second write data 408W(2) can be written to the respective memory banks 404(1) and 404(2) simultaneously. The burst read mode can be detected by a burst detection circuit such as the burst detection circuit 418 of FIG. 4 as an example and transmitted to the deserialization circuit 916S. The burst detection circuit can set the burst write mode based on the burst enable signal BURST_EN shown in FIG. 10. In this example, the first and second write clock signals 432(1) and 432(2) have a frequency twice that of the source clock signal 419.The deserialization circuit 916D can continue to write the multiplexed and separated write data 408W(1), 408(2) from the continuously received write data stream 408W in consecutive clock cycles (e.g., D2, D3, D4, D5, etc. as shown in FIG. 10).

[0059] Note that the deserialization circuit in FIG. 9 can also be set to the non-burst write mode. As shown in FIG. 10, in the non-burst write mode 440, only the first write latch circuit 932(1) is clocked at half the frequency of the source clock signal 419 by the first write clock signal 432(1), and writes the write data 408W(1), 408W(2) from the write data stream 408W sequentially to the addressed memory banks 404(1), 404(2).

[0060] FIG. 11 is an exemplary memory system 1100 having a memory array 1102 including memory banks 1103 in an exemplary CM4 interleaved arrangement similar to the memory banks 304(1), 304(2) of the memory system 300 in FIG. 3. The serialization and / or deserialization circuits and processes described above in FIGS. 4-10 can be provided and / or executed in the memory array 1102 of FIG. 11. As will be described later, the memory system 1100 includes memory banks 1103 with a higher memory density than the memory banks 304(1), 304(2) of FIG. 3, for example, with reduced power consumption per memory unit. In this regard, the memory array 1102 of FIG. 11 includes a first inner memory sub-bank 1104(1) and a first outer memory sub-bank 1104(2). The memory bank may be further divided into a plurality of memory sub-banks. A memory sub-bank is a division of memory bit cells from a memory bank. Memory sub-banks within a memory bank share a common access circuit with another memory sub-bank within that memory bank. Thus, at a given point in time, only one memory sub-bank within a memory bank may be accessible.

[0061] The memory sub-bank 1104(1) is coupled to the column multiplexer circuits 1112(0) to 1112(3) to multiplex the data bits from the first memory sub-bank 1104(1) to the respective memory column access circuits 1114(0) to 1114(3). The memory column access circuits 1114(0) to 1114(3) are, in this example, sense amplifier circuits and can sense the memory state regarding the signals on the respective bit lines BL(1)(0) to BL(1)(Y) multiplexed from the memory states from the respective column multiplexer circuits 1112(0) to 1112(3). However, the memory array 1102 of FIG. 11 also includes a second memory sub-bank 1104(2) that is also coupled to the column multiplexer circuits 1112(0) to 1112(3) to multiplex the data bits from the bit lines BL(2)(0) to BL(2)(Y) from the second memory sub-bank 1104(3) to the respective memory column access circuits 1114(0) to 1114(3). Also, the memory column access circuits 1114(0) to 1114(3) can sense the memory state regarding the signals on the respective bit lines BL(2)(0) to BL(2)(Y) multiplexed from the memory states from the respective column multiplexer circuits 1112(0) to 1112(3). In this way, the memory density of the memory array 1102 having two memory sub-banks 1104(1) and 1104(2) sharing the common column multiplexer circuits 1112(0) to 1112(3) and the memory column access circuits 1114(0) to 1114(3) is increased. However, even when the second memory sub-bank 1104(2) is added, it is desirable that the length of the bit lines of the first memory sub-bank 1104(1) does not increase.

[0062] In this regard, as shown in FIG. 11, the first inner memory sub-bank 1104(1) is arranged closer to the column multiplexer circuits 1112(0) to 1112(3) and the memory column access circuits 1114(0) to 1114(3) than the second outer memory sub-bank 1104(2). The inner memory sub-bank 1104(1) includes X + 1 memory row circuits 1106(0) to 1106(X), each including a plurality of memory bit cell circuits 1108(0)(0) to 1108(X)(Y). For example, the memory row circuit 1106(0) includes Y + 1 memory bit cell circuits 1108(0)(0) to 1108(0)(Y). The memory row circuit 1106(X) includes Y memory bit cell circuits 1108(X)(0) to 1108(X)(Y). As a non-limiting example, the memory bit cell circuits 1108(0)(0) to 1108(X)(Y) may be static random access memory (SRAM) bit cells employing 6 transistors (6-T) or more transistor counts. As another example, the memory bit cell circuits 1108(0)(0) to 1108(X)(Y) can also be dynamic random access memory (DRAM) bit cells. The arrangement of the memory bit cell circuits 1108(0)(0) to 1108(X)(Y) is such that one memory bit cell circuit 1108()(0) to 1108()(Y) from each of the memory row circuits 1106(0) to 1106(X) is arranged in each of the same memory column circuits 1110(0) to 1110(Y). In FIG. 3, only the memory column circuits 1110(0) and 1110(Y) are labeled. For example, the inner memory sub-bank 1104(1) may be provided with 256 memory column circuits 1110(0) to 1110(255).

[0063] Continuing to refer to FIG. 11, the inner memory sub-bank 1104(1) includes Y first bit lines BL(1)(0) to BL(1)(Y) respectively coupled to the memory bit cell circuits 1108(0)(0) to 1108(X)(Y) within each of the memory row circuits 1106(0) to 1106(X). The first bit lines BL(1)(0) to BL(1)(Y) can be precharged to write data to the memory bit cell circuits 1108(0)(0) to 1108(X)(Y) of the selected memory row circuits 1106(0) to 1106(Y) under the control of the activation of the word line WL by the memory driver circuit 1118 for the selected memory row circuits 1106(0) to 1106(X) according to the decoded memory address 1116. Although only one WL is shown in FIG. 11, it should be noted that in each of the memory row circuits 1106(0) to 1106(X), a separate WL is provided for each of the memory row circuits 1106(0) to 1106(X) coupled to the respective memory bit cell circuits 1108(0)(0) to 1108(X)(Y). Only one of the WLs for a given memory row circuits 1106(0) to 1106(X) is activated to select the memory row circuits 1106(0) to 1106(X) for a memory access operation. The memory bit cell circuits 1108(0)(0) to 1108(X)(Y) of the selected memory row circuits 1106(0) to 1106(X) can also assert data on the respective bit lines BL(1)(0) to BL(1)(Y) for the memory read operation provided by the column multiplexer circuits 1112(0) to 1112(3) and the memory column access circuits 1114(0) to 1114(3).

[0064] As described above, in order to increase the memory density of the memory array 1102, the second outer memory sub-bank 1104(2) is also included in the memory system 1100. The outer memory sub-bank 1104(2) is arranged at a location farther from the column multiplexer circuits 1112(0) to 1112(3) and the memory column access circuits 1114(0) to 1114(3) than the inner memory sub-bank 1104(1). Similar to the inner memory sub-bank 1104(1), the outer memory sub-bank 1104(2) has X + 1 memory row circuits 1126(0) to 1126(X), each including a plurality of memory bit cell circuits 1128(0)(0) to 1128(X)(Y). For example, the memory row circuit 1126(0) includes Y + 1 memory bit cell circuits 1128(0)(0) to 1128(0)(Y). The memory row circuit 1126(X) includes Y memory bit cell circuits 1128(X)(0) to 1128(X)(Y). As a non-limiting example, the memory bit cell circuits 1128(0)(0) to 1128(X)(Y) may be SRAM bit cells employing 6 transistors (6-T) or more transistor counts. As another example, the memory bit cell circuits 1128(0)(0) to 1128(X)(Y) can also be DRAM bit cells. The arrangement of the memory bit cell circuits 1128(0)(0) to 1128(X)(Y) is such that one memory bit cell circuit 1128()(0) to 1128()(Y) from each of the memory row circuits 1126(0) to 1126(X) is arranged in each of the same memory column circuits 1130(0) to 1130(Y). In FIG. 11, only the memory column circuits 1130(0) and 1130(Y) are labeled. For example, the outer memory sub-bank 1104(2) may be provided with 256 memory column circuits 1130(0) to 1130(255).

[0065] Continuing to refer to FIG. 11, the outer memory sub-bank 1104(2) also includes Y second bit lines BL(2)(0) to BL(2)(Y) respectively coupled to the memory bit cell circuits 1128(0)(0) to 1128(X)(Y) within each of the memory row circuits 1126(0) to 1126(X). The second bit lines BL(2)(0) to BL(2)(Y) can be precharged to write data to the memory bit cell circuits 1128(0)(0) to 1128(X)(Y) of the selected memory row circuits 1126(0) to 1126(Y) under the control of activation of the word line WL by the memory driver circuit 1118 for the selected memory row circuits 1126(0) to 1126(X) according to the decoded memory address 1116. Note that in each of the memory row circuits 1126(0) to 1126(X), a separate WL is provided for each of the memory row circuits 1126(0) to 1126(X) coupled to the respective memory bit cell circuits 1128(0)(0) to 1128(X)(Y). Only one of the WLs of the given memory row circuits 1126(0) to 1126(X) in the outer memory sub-bank 1104(2) and the memory row circuits 1106(0) to 1106(X) in the inner memory sub-bank 1104(1) is activated to select either the memory row circuits 1126(0) to 1126(X) or the memory row circuits 1106(0) to 1106(X) for a memory access operation. The memory bit cell circuits 1128(0)(0) to 1128(X)(Y) of the selected memory row circuits 1126(0) to 1126(X) can also assert data to the respective bit lines BL(2)(0) to BL(2)(Y) for the memory read operation provided by the column multiplexer circuits 1112(0) to 1112(3) and the memory column access circuits 1114(0) to 1114(3).

[0066] The first and second memory sub-banks 1104(1) and 1104(2) are designed to store interleaved data words A, B, C, D according to an interleaved memory column circuit labeled A1, B1, C1, D1, ..., A4, B4, C4, D4. Therefore, the memory array 1102 is also configured in a CM4 interleaved arrangement. Thus, in this example, there are four columns of multiplexer circuits 1112(0) to 1112(3) to support the CM4 interleaved arrangement. The number of column multiplexer circuits 1112(0) to 1112(3) can be two or more to match the interleaving scheme.

[0067] When the inner memory sub-bank 1104(1) and the outer memory sub-bank 1104(2) are accessed in response to a memory read operation, the word line WL is activated for the selected memory row circuits 1106(0) to 1106(X), 1126(0) to 1126(X) according to the decoded memory address 1116 for the memory access operation. The column selection CS1 is generated for each column multiplexer circuit 1112(0) to 1112(3) coupled to the respective first and second bit lines BL(1)(0) to BL(1)(Y), BL(2)(0) to BL(2)(Y) coupled to the respective memory bit cell circuits 1108(0)(0) to 1108(0)(Y), 1128(0)(0) to 1128(0)(Y) in the respective memory column circuits 1110(0) to 1110(Y), 1130(0) to 1130(Y) representing the interleaved bits from the selected memory row circuits 1106(0) to 1106(X), 1126(0) to 1126(X). Each column multiplexer circuit 1112(0) to 1112(3) couples one of the coupled first and second bit lines BL(1)(0) to BL(1)(Y), BL(2)(0) to BL(2)(Y) from the coupled memory column circuits 1110(0) to 1110(Y), 1130(0) to 1130(Y) to the respective multiplexed outputs 1120(0) to 1120(3) to provide the corresponding bits to the respective memory column access circuits 1114(0) to 1114(3) (e.g., sense amplifiers). In this way, the column multiplexer circuits 1112(0) to 1112(3) are controlled to multiplex the selected bits from the interleaved data words in the selected memory row circuits 1106(0) to 1106(X), 1126(0) to 1126(X) according to the respective memory column access circuits 1114(0) to 1114(3). The memory column access circuits 1114(0) to 1114(3) are configured to provide the bits of the data output word 1124 on the respective column outputs 1122(0) to 1122(3) for the memory read operation.

[0068] For example, in a memory read operation, when it is desired to select interleaved data words A1 to A4 from the selected memory row circuits 1106(0) to 1106(X), the column multiplexer circuits 1112(0) to 1112(3) are controlled by a column selection CS1 and select bits A1 to A4 on respective first bit lines BL(1)(0), BL(1)(3), BL(1)(7), BL(1)(11) from the selected memory row circuits 1106(0) to 1108(X) and multiplex them onto respective multiplexed outputs 1120(0) to 1120(3) to respective memory column access circuits 1114(0) to 1114(3). The memory column access circuits 1114(0) to 1114(3) are configured to provide signals indicating the read bits on the first bit lines BL(1)(0), BL(1)(3), BL(1)(7), BL(1)(11) as data output words 1124 onto respective column outputs 1122(0) to 1122(3).

[0069] As shown in FIG. 11 and as described above, the first and second bit lines BL(1)(0) to BL(1)(Y), BL(2)(0) to BL(2)(Y) are provided for each of the memory column circuits 1110(0) to 1110(Y), 1130(0) to 1130(Y) of the respective memory sub-banks 1104(1), 1104(2). The length of the first bit lines BL(1)(0) to BL(1)(Y) can be extended to provide bit lines for each of the memory column circuits 1130(0) to 1130(Y) of the outer memory sub-bank 1104(2). For example, the first bit lines BL(1)(0) to BL(1)(Y) for the inner memory sub-bank 1104(1) can extend within or above the first metal layer (e.g., M0 or M2) within the memory bit cell circuits 1108(0)(0) to 1108(X)(Y). Extending the length of the first bit lines BL(1)(0) to BL(1)(Y) increases the capacitance of the first bit lines BL(1)(0) to BL(1)(Y), which undesirably degrades the memory performance with respect to the memory array 1102.

[0070] Therefore, in order to avoid the need to lengthen the first bit lines BL(1)(0) to BL(1)(Y) of the inner memory sub-bank 1104(1) to provide bit lines for the outer memory sub-bank 1104(2), the second bit lines BL(2)(0) to BL(2)(Y) for the outer memory sub-bank 1104(2) in the memory system 1100 of FIG. 11 are provided as separate bit lines. The second bit lines BL(2)(0) to BL(2)(Y) for the outer memory sub-bank 1104(2) are separate bit lines and are coupled to their respective column multiplexer circuits 1112(0) to 1112(3) separately from the first bit lines BL(1)(0) to BL(1)(Y) for the inner memory sub-bank 1104(1). However, a path must be provided between the second bit lines BL(2)(0) to BL(2)(Y) for the outer memory sub-bank 1104(2) and the column multiplexer circuits 1112(0) to 1112(3). The memory bit cell circuits 1108(0)(0) to 1108(X)(Y) can be redesigned to accommodate the coupling of the second bit lines BL(2)(0) to BL(2)(Y) within an additional metal wire routing path in the first metal layer that houses the first bit lines BL(1)(0) to BL(1)(Y) for the inner memory sub-bank 1104(1), and this coupling extends through the inner memory sub-bank 1104(2) alongside BL(1)(0) to BL(1)(Y) to the column multiplexer circuits 1112(0) to 1112(3). However, manufacturing limitations may prevent or may be undesirable to change the cell design of all the memory bit cell circuits 1108(0)(0) to 1108(X)(Y) to accommodate the coupling of the first bit lines BL(1)(0) to BL(1)(Y) to the first metal layer and provide an additional metal wire routing path for the second bit lines BL(2)(0) to BL(2)(Y) alongside the first bit lines BL(1)(0) to BL(1)(Y).

[0071] In this regard, as shown in FIG. 11, in order to avoid both the need to lengthen the first bit lines BL(1)(0) to BL(1)(Y) of the inner memory sub-bank 1104(1) and extend them to the outer memory sub-bank 1104(2), jumper cell circuits 1132(0) to 1132(Y) are provided. In this example, as the jumper cell circuits 1132(0) to 1132(Y), the outermost memory bit cell circuits 1108(X)(0) to 1108(X)(Y) adjacent to the outer memory sub-bank 1104(2) are provided. The jumper cell circuits 1132(0) to 1132(Y) are each coupled to the second bit lines BL(2)(0) to BL(2)(Y) for the outer memory sub-bank 1104(2) and are also each coupled to the first bit lines BL(1)(0) to BL(1)(Y) for the inner memory sub-bank 1104(1). The jumper cell circuits 1132(0) to 1132(Y) each include respective metal interconnections 1134(0) to 1134(Y) that couple the respective second bit lines BL(2)(0) to BL(2)(Y) for the outer memory sub-bank 1104(2) to the respective flying bit lines FBL(0) to FBL(Y). In this example, the metal interconnections 1134(0) to 1134(Y) of each of the jumper cell circuits 1132(0) to 1132(2) couple each of the second bit lines BL(2)(0) to BL(2)(Y) of the first metal layer for the outer memory sub-bank 1104(2) to the respective flying bit lines FBL(0) to FBL(Y) of the second metal layer ML2 (e.g., M4). For example, the second metal layer ML2 may be disposed in a metal layer higher than the first metal layer ML1 in which the first bit lines BL(1)(0) to BL(1)(Y) for the inner memory sub-bank 1104(1) are disposed. In this way, the flying bit lines FBL(0) to FBL(Y) can be coupled to the respective column multiplexer circuits 1112(1) to 1112(3) by "jumping over" the first metal layer ML1 in which the first bit lines BL(1)(0) to BL(1)(Y) for the inner memory sub-bank 1104(1) are disposed vertically.

[0072] FIG. 12 is a block diagram of an exemplary processor-based system 1200 that includes a processor 1202 configured to execute computer instructions for execution. The processor-based system also includes a memory system 1204 that includes one or more memory arrays, each including a plurality of memory banks, and an integrated serialization circuit configured to convert parallel data streams of read data received from separately switchable memory banks into a single serialized read data stream for providing on an output bus in burst read mode, and / or a deserialization circuit configured to convert a received serialized write data stream on an input bus for write operations into separate and parallel write data streams for simultaneous writing onto the memory banks in burst write mode. The memory system 1204 includes, in this example, an instruction cache 1206, a data cache 1208, and a system memory 1210. Any memory within the memory system 1204 of FIG. 12 may include, by way of non-limiting example, the memory arrays 400, 900, 1102 of FIGS. 4, 9, and 11.

[0073] Continuing to refer to FIG. 12, the processor-based system 1200 may be one or more circuits included in an electronic substrate card such as a printed circuit board (PCB), a server, a personal computer, a desktop computer, a laptop computer, a personal digital assistant (PDA), a computing pad, a mobile device, or other device, and can represent, for example, a server or the user's computer. The processor 1202 represents one or more general-purpose processing circuits such as a microprocessor, a central processing unit. The processor 1202 includes an instruction processing circuit 1209 configured to execute processing logic in computer instructions for performing the operations and steps discussed herein. The processor 1202 also includes an instruction cache 1206 for temporary high-speed access memory storage of instructions. Instructions fetched or prefetched from memory such as system memory 1210 via the system bus 1212 are stored in the instruction cache 1206. The processor 1202 also includes a data cache 1208 for temporary high-speed access memory storage of data from system memory 1210 via the system bus 1212.

[0074] Processor 1202 and system memory 1210 are coupled to system bus 1212 and can interconnect peripheral devices included in processor-based system 1200. As is well known, processor 1202 communicates with these other devices by exchanging address, control, and data information via system bus 1212. For example, processor 1202 can communicate a burst transaction request to memory controller 1214 within system memory 1210 as an example of a slave device. Although not shown in FIG. 12, a plurality of system buses 1212 may be provided, where each system bus has a different structure. In this example, memory controller 1214 is configured to provide memory access requests to memory array 1216 within system memory 1210. Memory array 1216 is composed of an array of memory bit cells for storing data. System memory 1210 may be, by way of non-limiting example, read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM), and static memory (e.g., flash memory, static random access memory (SRAM), etc.).

[0075] Other devices may be connected to the system bus 1212. As shown in FIG. 12, these devices may include, by way of example, system memories 1210, one or more input devices 1218, one or more output devices 1220, a modem 1222, and one or more display controllers 1224. The input device 1218 can include any type of input device including, but not limited to, input keys, switches, voice processors, etc. The output device 1220 can include any type of output device including, but not limited to, voice, video, other visual indicators, etc. The modem 1222 can be any device configured to enable data exchange with the network 1226. The network 1226 can be any type of network including, but not limited to, wired or wireless networks, private or public networks, local area networks (LANs), wireless local area networks (WLANs), wide area networks (WANs), BLUETOOTH (trademark) networks, and the Internet. The modem 1222 can be configured to support any desired type of communication protocol. The processor 1202 can also be configured to access the display controller 1224 via the system bus 1212 and control the information transmitted to one or more displays 1228. The display 1228 can include any type of display including, but not limited to, cathode ray tubes (CRTs), liquid crystal displays (LCDs), plasma displays, etc.

[0076] When executed by a processor such as processor 1202, the processor-based system 1200 of FIG. 12 performs serialization of read data from the memory system 1204 by converting a parallel data stream of read data received from separately switched memory banks into a single serialized read data stream to be provided on the output bus in burst read mode, and / or deserialization of write data transmitted to the memory system 1204 by converting a received serialized write data stream on the input bus for write operations into separate parallel write data streams to be simultaneously written onto the memory banks in burst write mode. A set of instructions 1230 can be included. The instructions 1230 can be stored in the system memory 1210, the processor 1202, and / or the instruction cache 1206 as examples of non-transitory computer-readable media 1232. The instructions 1230 can also exist fully or at least partially within the system memory 1210 and / or within the processor 1202 during their execution. The instructions 1230 can further be transmitted or received via the network 1226 when the network 1226 transmits or receives via the modem 1222 to or from the non-transitory computer-readable media 1232, or other examples such as the input device 1218 via the network 1226.

[0077] The non-transitory computer-readable medium 1232 is shown as a single medium in the exemplary embodiment, but the term "computer-readable medium" should be taken to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated cache and server) that store one or more sets of instructions. The term "computer-readable medium" should also be taken to include any medium that is capable of storing, encoding, or carrying a set of instructions for execution by a processing device and that causes a processing device to execute any one or more of the methodologies of the embodiments disclosed herein. Thus, the term "computer-readable medium" should be taken to include, without limitation, solid-state memory, optical media, and magnetic media.

[0078] The embodiments disclosed herein include various steps. The steps of the embodiments disclosed herein may be formed by hardware components or may be embodied as machine-executable instructions, which may be used to cause a general-purpose processor or a special-purpose processor programmed with the instructions to execute the steps. Alternatively, the steps may be executed by a combination of hardware and software.

[0079] The embodiments disclosed herein may include a computer program product that can be provided as software, which includes a machine-readable medium (or computer-readable medium) storing instructions that can be used to program a computer system (or other electronic device) to execute a process according to the embodiments disclosed herein. A machine-readable medium includes any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer). For example, machine-readable media include machine-readable storage media (e.g., ROM, random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, etc.).

[0080] Unless stated otherwise specifically, and as is apparent from the foregoing discussion, throughout this specification, discussions using terms such as "processing," "computing," "determining," "displaying," etc., refer to actions and processes of a computer system, or similar electronic computing device, that manipulate and transform data represented as physical (electronic) quantities within the registers of the computer system, and memory, into other data similarly represented as physical quantities within the memory, registers, or other such information storage devices, transmission devices, or display devices of the computer system.

[0081] The algorithms and displays presented in this specification are not inherently related to any particular computer or other device. Various systems may be used with programs in accordance with the teachings of this specification, or it may prove convenient to construct more specialized devices to perform the required method steps. The structures required for these various systems will become apparent from the above description. Additionally, the embodiments described in this specification are not described with reference to any particular programming language. It should be appreciated that various programming languages may be used to implement the teachings of the embodiments as described herein.

[0082] Those skilled in the art will further understand that various exemplary logical blocks, modules, circuits, and algorithms described in conjunction with the embodiments disclosed herein may be implemented as electronic hardware, instructions stored in memory or another computer-readable medium, or executed by a processor or other processing device or a combination of both. The components of the distributed antenna system described herein may be employed, by way of example, in any circuit, hardware component, integrated circuit (IC), or IC chip. The memory disclosed herein may be of any type and size and may be configured to store any desired type of information. For the sake of clarity in explaining this interchangeability, various exemplary components, blocks, modules, circuits, and steps have been generally described above from the perspective of their functionality. How such functionality is implemented depends on the particular application, design choices, and / or design constraints imposed on the overall system. Skilled artisans can implement the described functionality in various ways for each particular application, but such implementation decisions should not be construed as causing a departure from the scope of the present embodiments.

[0083] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein can be implemented or executed using a processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. Further, a controller may be a processor. The processor may be a microprocessor, but in the alternative, may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration).

[0084] The embodiments disclosed herein may be embodied in hardware and in instructions stored in hardware, and the instructions may be present, for example, in RAM, flash memory, ROM, electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, a hard disk, a removable disk, a CD-ROM, or any other form of computer readable medium known in the art. An illustrative memory medium is coupled to the processor such that the processor can read information from, and write information to, the memory medium. In the alternative, the memory medium may be integral to the processor. The processor and the memory medium may reside in an ASIC. The ASIC may reside in a remote station. In the alternative, the processor and the memory medium may reside as discrete components in a remote station, base station, or server.

[0085] Note that the operation steps described in any of the exemplary embodiments in this specification are explained for the purpose of providing examples and discussions. The operations described can be performed in many different orders other than the order shown. Furthermore, the operations described as a single operation step may actually be executed in a plurality of different steps. In addition, one or more operation steps discussed in the exemplary embodiments may be combined. Also, those skilled in the art will understand that information and signals can be represented using any of a variety of techniques and technologies. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referred to throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields, or particles, optical fields or particles, or any combination thereof.

[0086] Unless explicitly stated otherwise, it is never intended that any method described in this specification be construed as requiring that its steps be performed in a particular order. Thus, when a method claim does not actually recite the order in which its steps are to be followed, or when the steps are not specifically stated in the claim or specification to be limited to a particular order, it is never intended that a particular order be inferred.

[0087] It will be apparent to those skilled in the art that various modifications and variations can be made without departing from the spirit and scope of the invention. Since modifications, combinations, sub - combinations, and variations of the disclosed embodiments incorporating the spirit and substance of the invention can occur to those skilled in the art, the invention should be construed to include all within the scope of the appended claims and their equivalents.

Claims

1. a first output bus, a first memory bank including a first read output, a second memory bank including a second read output, a first read driver circuit clocked by a source clock signal, accessing first read data stored at a first memory address in the first memory bank, which is asserted on the first read output, accessing second read data stored at a second memory address in the second memory bank, which is asserted on the second read output and a first read driver circuit configured to perform the above; a serialization circuit, asserting the first read data on the first output bus in response to a first clock cycle of the source clock signal, and asserting the second read data on the first output bus after the assertion of the first read data on the first output bus in response to a second clock cycle of the source clock signal following immediately after the first clock cycle of the source clock signal and a serialization circuit configured to perform the above a memory array including the above.

2. The memory array according to claim 1, wherein the first read driver circuit is further configured to access third read data stored at a third memory address in the first memory bank, which is asserted on the first read output, access fourth read data stored at a fourth memory address in the second memory bank, which is asserted on the second read output and the serialization circuit is further configured to assert the third read data on the first output bus in response to a third clock cycle of the source clock signal following immediately after the second clock cycle of the source clock signal, and asserting the fourth read data on the first output bus after the assertion of the third read data on the first output bus in response to a fourth clock cycle of the source clock signal following immediately after the third clock cycle of the source clock signal a further configured memory array.

3. The memory array according to claim 1, wherein ​ Receiving the source clock signal of the first frequency; Generating a read switching clock signal at a second frequency higher than the first frequency in burst read mode; Further comprising a serialization clock generation circuit configured to perform the above; The serialization circuit includes a serialization clock input configured to receive the read switching clock signal; The serialization circuit is configured to: Assert the first read data on the first output bus in response to a first clock cycle of the read switching clock signal; Assert the second read data on the first output bus after the assertion of the first read data on the first output bus in response to a second clock cycle of the read switching clock signal following immediately after the first clock cycle of the read switching clock signal; Further configured to perform the above; Memory array.

4. The memory array according to claim 3, wherein the serialization clock generation circuit is further configured to generate the read switching clock signal including non-overlapping clock pulses.

5. The memory array according to claim 3, wherein the second frequency of the read switching clock signal is twice the first frequency of the source clock signal.

6. The memory array according to claim 1, comprising: A first input bus; A first write output coupled to the first memory bank; A second write output coupled to the second memory bank; A first write driver circuit clocked by the source clock signal, configured to: Assert a first write data stream on the first input bus to start writing at a third memory address within the memory array; A first write driver circuit configured as above; A deserialization circuit, configured to: Receive the first write data stream from the first input bus in response to a third clock cycle of the source clock signal; Demultiplex the first write data and the second write data from the first write data stream; In response to a fourth clock cycle of the source clock signal following immediately after the third clock cycle of the source clock signal, assert on the first write output the first write data of the first write data stream to be written to the third memory address in the first memory bank; In response to the fourth clock cycle of the source clock signal, assert on the second write output the second write data of the first write data stream to be written to a fourth memory address based on the third memory address in the second memory bank; A deserialization circuit configured to perform the above; A memory array further including the above. **Claim 7** The memory array according to claim 1, further comprising a memory column access circuit, The first memory bank includes A first metal layer, A second metal layer different from the first metal layer, the second metal layer including a plurality of flying bit lines each coupled to the memory column access circuit, A first memory sub-bank, A plurality of first memory row circuits each including a plurality of first memory bit cell circuits respectively arranged in each of the plurality of first memory column circuits of the plurality of first memory column circuits, A plurality of first bit lines arranged on the first metal layer and respectively coupled to one of the plurality of first memory column circuits of the plurality of first memory column circuits and the memory column access circuit The first memory sub-bank including the above, A second memory sub-bank, A plurality of second memory row circuits each including a plurality of second memory bit cell circuits respectively arranged in each of the plurality of second memory column circuits of the plurality of second memory column circuits, A plurality of second bit lines respectively arranged on the first metal layer and respectively coupled to one of the plurality of second memory column circuits of the plurality of second memory column circuits, The second memory sub-bank including the above, A first jumper row circuit including a plurality of first jumper cell circuits respectively coupled to one of the plurality of second bit lines of one of the plurality of second memory column circuits in the first metal layer and one of the plurality of first flying bit lines of the plurality of first flying bit lines in the second metal layer The memory array further including the above. **Claim 8** A method for serializing read data from a plurality of memory banks within a memory array, comprising: Receiving a source clock signal; Based on the source clock signal, accessing first read data stored at a first memory address in a first memory bank within a first memory array; Asserting the first read data on a first read output of the first memory bank; Accessing second read data stored at a second memory address in a second memory bank within the first memory array; Asserting the second read data on a second read output of the second memory bank; In response to a first clock cycle of the source clock signal, asserting the first read data on a first output bus; In response to a second clock cycle of the source clock signal following immediately after the first clock cycle of the source clock signal, after the assertion of the first read data on the first output bus, asserting the second read data on the first output bus; A method comprising the steps of: **Claim 9** The method according to claim 8, further comprising: Based on the source clock signal, asserting a first write data stream on a first input bus to begin writing at a third memory address within the memory array; In response to a third clock cycle of the source clock signal, receiving the first write data stream from the first input bus; Demultiplexing first write data and second write data from the first write data stream; In response to a fourth clock cycle of the source clock signal following immediately after the third clock cycle of the source clock signal, asserting the first write data of the first write data stream on a first write output coupled to the first memory bank within the memory array to be written to the third memory address within the first memory bank; In response to the fourth clock cycle of the source clock signal, assert on a second write output coupled to a second memory bank within the memory array such that second write data of the first write data stream is written to a fourth memory address based on the third memory address within the second memory bank A method further comprising. **Claim 10** A first input bus, A first memory bank, A second memory bank, A first write output coupled to the first memory bank, A second write output coupled to the second memory bank, A first write driver circuit clocked by a source clock, the first write driver circuit configured to assert on the first input bus such that a first write data stream begins to be written at a first memory address within a memory array A deserialization circuit, In response to a first clock cycle of the source clock signal, receive the first write data stream from the first input bus, Multiplex and separate first write data and second write data from the first write data stream, In response to a second clock cycle of the source clock signal following immediately after the first clock cycle of the source clock signal, assert on the first write output such that the first write data of the first write data stream is written to the first memory address within the first memory bank, In response to the second clock cycle of the source clock signal, assert on the second write output such that the second write data of the first write data stream is written to a second memory address based on the first memory address within the second memory bank A deserialization circuit configured to perform A memory array including. **Claim 11** The memory array according to claim 10, wherein The first write driver circuit is Further configured to assert on the first input bus such that a second write data stream begins to be written at a third memory address within the memory array And, The deserialization circuit is In response to a third clock cycle of the source clock signal, receiving the second write data stream from the first input bus; Demultiplexing third write data and fourth write data from the second write data stream; In response to a fourth clock cycle of the source clock signal following immediately after the third clock cycle of the source clock, asserting, on the first write output, the third write data of the second write data stream to be written to the third memory address in the first memory bank; In response to the fourth clock cycle of the source clock signal, asserting, on the second write output, the fourth write data of the second write data stream to be written to a fourth memory address based on the third memory address in the second memory bank; Further configured to perform; A memory array. **Claim 12** The memory array according to claim 10, A deserialization clock generation circuit, Receiving the source clock signal of a first frequency; Generating a first write clock signal at a second frequency based on the first frequency in burst write mode; Generating a second write clock signal at a second frequency based on the first frequency in burst write mode; A deserialization clock generation circuit configured to perform; Further comprising; The deserialization circuit, In response to a first clock cycle of the write signal, receiving the first write data stream from the first input bus; Demultiplexing the first write data and the second write data from the first write data stream; In response to a first clock cycle of the second write clock signal following immediately after a first clock cycle of the first write clock signal, asserting, on the first write input, the first write data of the first write data stream to be written to the first memory address in the first memory bank; In response to the first clock cycle of the second write clock signal, assert on the first write input such that the second write data of the first write data stream is written to the second memory address in the first memory bank configured to perform a memory array

13. The memory array according to claim 10, wherein the deserialization circuit in response to the second clock cycle of the source clock and the first memory address being an even memory address, assert on the first write output such that the first write data is written to the first memory address; in response to the second clock cycle of the source clock and the second memory address being an odd memory address, assert on the second write output such that the second write data is written to the second memory address; in response to the second clock cycle of the source clock and the first memory address being an odd memory address, assert on the first write output such that the first write data is written to the first memory address; in response to the second clock cycle of the source clock and the second memory address being an even memory address, assert on the second write output such that the second write data is written to the second memory address further comprising a memory array

14. The memory array according to claim 10, further comprising a memory column access circuit, wherein the first memory bank a first metal layer; a second metal layer different from the first metal layer, the second metal layer including a plurality of flying bit lines each coupled to the memory column access circuit; a first memory sub-bank, a plurality of first memory row circuits each including a plurality of first memory bit cell circuits respectively disposed in each of the plurality of first memory column circuits of the plurality of first memory column circuits; a plurality of first bit lines disposed on the first metal layer and respectively coupled to one of the plurality of first memory column circuits of the plurality of first memory column circuits and the memory column access circuit including a first memory sub-bank; a second memory sub-bank, A plurality of second memory row circuits each including a plurality of second memory bit cell circuits respectively arranged in each of the plurality of second memory column circuits; A plurality of second bit lines respectively arranged in the first metal layer and respectively coupled to one of the plurality of second memory column circuits; A second memory sub-bank including: A first jumper row circuit including a plurality of first jumper cell circuits respectively coupled to one of the plurality of second bit lines in one of the plurality of second memory column circuits in the first metal layer and one of the plurality of first flying bit lines in the second metal layer; A memory array further including:

15. A method for deserializing write data from an input bus to be written into a plurality of memory banks in a memory array, the method comprising: Receiving a source clock signal; Asserting, based on the source clock signal, a first write data stream on a first input bus to begin writing at a first memory address in the memory array; Receiving the first write data stream from the first input bus in response to a first clock cycle of the source clock signal; Demultiplexing first write data and second write data from the first write data stream; Asserting, in response to a second clock cycle of the source clock signal following immediately after the first clock cycle of the source clock signal, the first write data of the first write data stream on a first write output coupled to a first memory bank in the memory array to be written at the first memory address in the first memory bank; Asserting, in response to the second clock cycle of the source clock signal, the second write data of the first write data stream on a second write output coupled to a second memory bank in the memory array to be written at a second memory address based on the first memory address in the second memory bank; A method comprising: