Multi-layer packaged chip and clock control method thereof

Through the clock control method between the main chip and the memory chip, the clock synchronization of the multi-layer stacked chips is achieved, which solves the communication cost and data synchronization problems between the multi-layer stacked chips and reduces communication delay and power consumption.

CN120315536BActive Publication Date: 2025-09-09BEIJING QINGYUN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510772473.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-09
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

How to achieve clock synchronization between multi-layer stacked chips to reduce communication costs and data synchronization issues.

Method used

The main chip notifies the memory chip to prepare the clock, and the memory chip performs clock arbitration and phase synchronization. The main chip detects the phase difference and generates a reference clock. The memory chip adjusts the working clock signal according to the reference clock to achieve clock synchronization between multiple chips.

Benefits of technology

It reduces the communication delay between memory chip layers and the resource usage of asynchronous FIFO, reduces the communication cost between multi-layer stacked chips, and realizes dynamic adjustment of clock domain division, reducing the cost and dynamic power consumption of chip clock domain division.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120315536B_ABST
    Figure CN120315536B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-layer stacked chip and a clock control method thereof, wherein the memory chip can prepare a corresponding clock according to the current function to be run, and through clock training, achieve clock synchronization between multiple memory chips that need to be run synchronously, thereby reducing the communication delay between memory chip layers and the resource occupation of asynchronous FIFOs, thereby reducing the communication cost between multi-layer stacked chips. In addition, the memory chip can identify information such as the current function and / or operating mode to be run based on notifications, and then perform clock arbitration and distribution, so that the division of the on-chip clock domain can be dynamically adjusted according to the different functions being run, realizing automatic division and control of the clock domain, thereby reducing the cost of chip clock domain division and dynamic power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of integrated circuits, and in particular to a multi-layer stacked chip and a clock control method thereof. Background Art

[0002] With the advent of the era of information data explosion, the market demand for memory continues to grow. Currently, multi-chip stacking technology is generally used to reduce the space occupied by multi-chip packaging, minimize the size of memory devices, and achieve high storage density in devices of the same size.

[0003] Among them, how to achieve clock synchronization between multi-layer stacked chips and thereby reduce the communication cost between the multi-layer stacked chips is one of the hot issues that technical personnel in this field urgently need to solve. Summary of the Invention

[0004] The object of the present invention is to provide a multi-layer stacked chip and a clock control method thereof, thereby further reducing the communication cost and data synchronization problem between the multi-layer stacked chips.

[0005] To achieve the above-mentioned object, the present invention provides a clock control method for a multi-layer stacked chip, wherein the multi-layer stacked chip has a three-dimensionally stacked main chip and multiple memory chips, and the clock control method comprises:

[0006] The main chip notifies each memory chip that needs to work synchronously to prepare a clock according to the current function that needs to be run;

[0007] Each memory chip that needs to work synchronously performs clock arbitration according to the received notification, thereby generating corresponding clock signals, control signals and multi-chip synchronization signals;

[0008] Each of the memory chips that need to work synchronously opens its coupled clock loop according to the control signal, and then feeds the clock signal back to the master chip;

[0009] The master chip detects the phase difference between the clock signals of the memory chips that need to work synchronously, and generates a reference clock according to the detection result;

[0010] Each of the memory chips that need to work synchronously adjusts its internal working clock signal according to the reference clock until the working clock signals of the memory chips that need to work synchronously achieve phase synchronization;

[0011] Each of the memory chips that needs to work synchronously receives instructions and works synchronously according to the multi-chip synchronization signal and the phase-synchronized working clock signal.

[0012] Optionally, the master chip performs function status decoding to obtain the current function that needs to be run, and then notifies each memory chip that needs to work synchronously to prepare a clock.

[0013] Optionally, the multiple memory chips are hybrid-bonded to each other and to the main chip via through-silicon vias, and the clock loop is arranged in the through-silicon via hybrid bonding layer.

[0014] Optionally, the clock loop includes a main clock path and a backup clock path, the backup clock path is used to back up the main clock path, and the clock control method also includes: when the main clock path coupled to the memory chip is offset or the silicon through via fails, the memory chip switches the backup clock path to open.

[0015] Optionally, the step of adjusting the internal clock signal of each memory chip that needs to work synchronously according to the reference clock includes:

[0016] performing through silicon via impedance calibration on the reference clock;

[0017] Coarsely adjusting the phase alignment between the internal working clock signal and the calibrated reference clock;

[0018] Send a corresponding training sequence to perform phase synchronization training on the working clock signal, and capture the jumping edge of the trained working clock signal during the training process to calculate the phase difference between the trained working clock signal and the calibrated reference clock, and then perform phase compensation on the trained working clock signal until the trained working clock signal and the calibrated reference clock achieve phase synchronization.

[0019] Optionally, after receiving the notification, each memory chip identifies the current function and / or operating mode that needs to be run according to the notification, and performs clock arbitration and allocation based on the identification result, thereby only turning on the clock domain where the functional modules required for the current function and / or operating mode are located.

[0020] Optionally, when clock arbitration and distribution are performed for each memory chip, clock arbitration and distribution are performed in descending order of priority of the test mode, user register, and user function.

[0021] Optionally, the functional module includes at least two of the following (1) to (5):

[0022] (1) a write module, which is used to write data into the memory chip, wherein the write module is located in a write clock domain, and when the instruction is a write instruction, the clock control method only turns on the write clock domain;

[0023] (2) a read module, which is used to read data from the memory chip, wherein the read module is located in a read clock domain, and when the instruction is a read instruction, the clock control method only turns on the read clock domain;

[0024] (3) a built-in self-test module, which is used to perform a built-in self-test on the memory chip to detect bad pixels in the memory chip, the built-in self-test module is located in a logic test control clock domain, and when the instruction is a built-in self-test instruction, the clock control method only turns on the logic test control clock domain;

[0025] (4) a redundant repair module, which is used to perform redundant repair on bad points in the memory chip, wherein the redundant repair module is located in a redundant repair clock domain, and when the instruction is a redundant repair instruction, the clock control method only turns on the redundant repair clock domain;

[0026] (5) A wafer test module, which is used to perform wafer testing on the memory chip after the memory chip is manufactured and before it is packaged. The wafer test module is located in a test mode clock domain. When the instruction is a wafer test instruction, the clock control method only turns on the test mode clock domain.

[0027] Based on the same inventive concept, the present invention also provides a multi-layer stacked chip, wherein the multi-layer stacked chip has a three-dimensionally stacked main chip and multiple memory chips, wherein:

[0028] The master chip is used to notify each memory chip that needs to work synchronously to prepare a clock according to the current function that needs to be run, detect the phase difference between the clock signals fed back by each memory chip that needs to work synchronously, and generate a reference clock based on the detection result;

[0029] The memory chip has:

[0030] Multiple functional modules are used to implement corresponding functions;

[0031] a clock arbitration and distribution circuit, configured to receive the notification, perform clock arbitration according to the notification, generate corresponding clock signals, control signals, and multi-chip synchronization signals, and enable the corresponding functional modules to operate according to the multi-chip synchronization signal and the phase-synchronized working clock signals after the working clock signals of the memory chips to be operated synchronously are phase-synchronized;

[0032] A clock adjustment circuit is coupled to the clock arbitration and distribution circuit and the corresponding clock loop, and is used to open the clock loop to which it is coupled according to the control signal to feed back the clock signal generated by the clock arbitration and distribution circuit to the main chip, and to receive the reference clock through the clock loop, and to adjust the working clock signal inside the memory chip according to the reference clock until the working clock signals of each of the memory chips that need to work synchronously are phase synchronized.

[0033] Optionally, the main chip includes:

[0034] A function status decoder, used for decoding the function status to obtain the current function to be run;

[0035] a notification module coupled to the functional status decoder and the clock arbitration and distribution circuit of each memory chip, and configured to notify each memory chip that needs to work synchronously to prepare a clock according to a decoding result of the functional status decoder;

[0036] a phase detector coupled to the clock loop to which each memory chip is coupled and used to detect a phase difference between clock signals fed back by each memory chip that needs to work synchronously;

[0037] A reference clock source is coupled to the phase detector and the clock adjustment circuit of each memory chip, and is used to generate a corresponding reference clock according to the detection result of the phase detector.

[0038] Optionally, the multiple memory chips are hybrid-bonded to each other and to the main chip via through-silicon vias, and the clock loop is arranged in the through-silicon via hybrid bonding layer.

[0039] Optionally, the clock loop includes a main clock path and a backup clock path, the backup clock path is used to back up the main clock path, and the main chip is also used to: when the main clock path of the memory chip is offset or the silicon through via fails, the clock arbitration and distribution circuit of the memory chip switches the backup clock path to open.

[0040] Optionally, the clock adjustment circuit includes:

[0041] a TSV driving unit coupled to the clock arbitration and distribution circuit and a corresponding clock loop, and configured to open the clock loop according to the control signal to feed back the clock signal generated by the clock arbitration and distribution circuit to the master chip, receive the reference clock through the clock loop, and perform through-silicon via impedance calibration on the reference clock;

[0042] A digitally controlled delay line is coupled to the TSV driving unit and the clock arbitration and distribution circuit, and is used to first coarsely adjust the phase alignment of the working clock signal output by the delay line with the calibrated reference clock, and then send a corresponding training sequence to the delay line to perform phase synchronization training on the working clock signal output by the delay line, and capture the jump edge of the trained clock signal during the training process to calculate the phase difference between the trained working clock signal and the calibrated reference clock, and then perform phase compensation on the trained working clock signal until the trained working clock signal achieves phase synchronization with the calibrated reference clock, and sends the phase-synchronized working clock signal to the clock arbitration and distribution circuit.

[0043] Optionally, the multiple functional modules within each memory chip are located in different clock domains; and the clock arbitration and distribution circuit includes:

[0044] an identification module, configured to identify the current function and / or operation mode to be run according to the notification after receiving the notification;

[0045] An arbitration and distribution module is coupled to the identification module, the clock adjustment circuit and each of the functional modules, and is used to perform clock arbitration and distribution according to the identification result of the identification module, thereby only turning on the clock domain where the functional modules required for the current function and / or operating mode are located.

[0046] Optionally, when performing clock arbitration and distribution, the arbitration and distribution module further performs clock arbitration and distribution in the order of priority of the test mode, the user register, and the user function from high to low.

[0047] Optionally, each of the memory chips that need to work synchronously receives instructions and works synchronously according to the multi-chip synchronization signal and the phase-synchronized working clock signal, and the functional module includes at least two of the following (1) to (5):

[0048] (1) a write module, which is used to write data into the memory chip. The write module is located in the write clock domain. When the instruction is a write instruction, the arbitration and allocation module only turns on the write clock domain.

[0049] (2) a read module, which is used to read data from the memory chip. The read module is located in the read clock domain. When the instruction is a read instruction, the arbitration and allocation module only turns on the read clock domain.

[0050] (3) a built-in self-test module, which is used to perform a built-in self-test on the memory chip to detect bad pixels in the memory chip, wherein the built-in self-test module is located in a logic test control clock domain, and when the instruction is a built-in self-test instruction, the arbitration and distribution module only turns on the logic test control clock domain;

[0051] (4) a redundancy repair module, which is used to perform redundancy repair on bad points in the memory chip. The redundancy repair module is located in a redundancy repair clock domain. When the instruction is a redundancy repair instruction, the arbitration and distribution module only turns on the redundancy repair clock domain.

[0052] (5) A wafer test module, which is used to perform wafer testing on the memory chip after the memory chip completes wafer manufacturing and before packaging. The wafer test module is located in the test mode clock domain. When the instruction is a wafer test instruction, the arbitration and allocation module only turns on the test mode clock domain.

[0053] Compared with the prior art, the technical solution of the present invention has at least one of the following beneficial effects:

[0054] 1. Its memory chip can prepare the corresponding clock according to the current function that needs to be run, and through clock training, it can achieve clock synchronization between multiple memory chips that need to run synchronously, reducing the communication delay between memory chip layers and the resource occupation of asynchronous FIFO, thereby reducing the communication cost between multi-layer stacked chips.

[0055] 2. Its memory chip can identify the current function and / or operating mode that needs to be run based on the notification, and then perform clock arbitration and distribution, so that the division of the on-chip clock domain can be dynamically adjusted according to the different functions being run, realizing automatic division and control of the clock domain, thereby reducing the cost and dynamic power consumption of the chip clock domain division. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Those skilled in the art will appreciate that the accompanying drawings are provided for a better understanding of the present invention and do not constitute any limitation on the scope of the present invention.

[0057] Figure 1 The figure is a flow chart of a clock control method for a multi-layer packaged chip according to an embodiment of the present invention.

[0058] Figure 2 Schematic diagram of the packaging structure of a multi-layer stacked chip according to an embodiment of the present invention.

[0059] Figure 3 FIG. 1 is a schematic diagram of the internal structure of a main chip in a multi-layer packaged chip according to an embodiment of the present invention.

[0060] Figure 4 FIG. 1 is a schematic diagram of the internal structure of a memory chip in a multi-layer package chip according to an embodiment of the present invention. DETAILED DESCRIPTION

[0061] In the following description, a large number of specific details are given to provide a more thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, some technical features known in the art are not described to avoid confusion with the present invention. It should be understood that the present invention can be implemented in different forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, providing these embodiments will make the disclosure thorough and complete and fully convey the scope of the present invention to those skilled in the art. The same reference numerals throughout represent the same elements. It should be understood that when an element is referred to as being "connected to" or "coupled to" another element, it can be directly connected to the other element, or there can be intervening elements. Conversely, when an element is referred to as being "directly connected to" another element, there are no intervening elements. When used herein, the singular forms "a," "an," and "said / the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "comprising" is used to identify the presence of certain features, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups. As used herein, the term "and / or" includes any and all combinations of the relevant listed items.

[0062] Please refer to Figure 1 and Figure 2 One embodiment of the present invention provides a clock control method for a multi-layer stacked chip. The multi-layer stacked chip comprises a three-dimensionally stacked main chip 10 and n+1 memory chips 30-3n, where n is an integer and is greater than or equal to 1. The main chip 10 is located at the bottom layer, and the n+1 memory chips 30-3n are sequentially stacked above the main chip 10.

[0063] In addition, the master chip 10 can implement interface conversion between a master device (i.e., a device accessing the memory, not shown) and n+1 memory chips 30-3n, and complete address decoding and data format conversion (e.g., data bit width) between the master device and the n+1 memory chips 30-3n, as well as converting operation instructions such as read, write, and refresh issued by the master device into signals that can be recognized by the n+1 memory chips 30-3n, thereby implementing the necessary control (including control of address signals, data signals, and various instruction signals) of the master device's access to the n+1 memory chips 30-3n for refresh operations, read and write operations, etc., so that the master device can access (or "use," "operate") the storage resources (i.e., the corresponding storage units) on the n+1 memory chips 30-3n according to user needs.

[0064] The host device may include any type of processing device with computing processing capabilities, such as a central processing unit (CPU), a digital signal processor (DSP), a network processor, an application processor (AP), a field programmable gate array (FPGA), a dedicated processor, etc. The processing device may be configured to execute instructions or software (including code, operating system or application, etc.) that can be executed by one or more computers, firmware or a combination thereof.

[0065] The n+1 memory chips 30-3n can each be a DRAM or any other suitable type of memory die structure. The DRAM can be any suitable type, such as synchronous DRAM (SDRAM) or wide I / O DRAM. The n+1 memory chips 30-3n can each be implemented as an unbuffered dual in-line memory module (UDIMM), a registered DIMM (RDIMM), a load-reduced DIMM (LRDIMM), a fully buffered DIMM (FBDIMM), a small outline DIMM (SODIMM), or the like.

[0066] In other embodiments, the main chip may also be a buffer chip with a memory controller inside, and the main device is arranged outside the main chip. When the main device is integrated into a logic chip, the logic chip is stacked under the main chip.

[0067] In one example, n+1 memory chips 30-3n are bonded to each other and to the main chip 10 using through-silicon via (TSV) hybrid bonding (HB). In this example, a layer of TSV hybrid bonding is provided between each two adjacent layers of chips, resulting in a total of n+1 TSV hybrid bonding layers 20-20n.

[0068] Please continue to refer to Figure 1 and Figure 2 、 Figure 4 The clock control method of the multi-layer packaged chip of this embodiment includes:

[0069] S1, the main chip 10 notifies each memory chip (which is at least two of the memory chips 30-3n, defined as 3i-3j, ji≥2) that needs to work synchronously according to the current function to be run to prepare the clock;

[0070] S2: Each memory chip 3i-3j that needs to work synchronously performs clock arbitration based on the received notification, thereby generating corresponding clock signals clki-clkj, working clock signals clkwi-clkwj, control signals, and multi-chip synchronization signals. The multi-chip synchronization signals of each memory chip 3i-3j correspond to the same instruction (e.g., a write instruction), thereby enabling the memory chips 3i-3j to jointly implement the function of the same instruction (e.g., writing all data input by the same write instruction).

[0071] S3, each memory chip 3i-3j that needs to work synchronously turns on its coupled clock loop according to the control signal, and then feeds back the clock signals clki-clkj generated by it to the main chip 10;

[0072] S4, the master chip 10 detects the phase difference between the clock signals clki-clkj of the memory chips 3i-3j that need to work synchronously, and generates a reference clock clkref according to the detection result;

[0073] S5, each memory chip 3i-3j that needs to work synchronously adjusts its internal working clock signal clkwi-clkwj according to the reference clock clkref until the clock signals clkwi-clkwj of each memory chip 3i-3j that need to work synchronously are phase-synchronized. The initial value of the working clock signal of each memory chip can be equal to the clock signal sent to the master chip 10. For example, the initial value of the working clock signal clkwi of the memory chip 3i is equal to clki.

[0074] S6, each memory chip 3i~3j that needs to work synchronously receives the instruction to be executed and works synchronously according to the multi-chip synchronization signal generated by it and the phase-synchronized clock signals clkwi~clkwj;

[0075] S7, the instruction is executed and the memory chips 3i-3j that need to work synchronously turn off their clock signals clki-clkj and working clock signals clkwi-clkwj, and wait for the next function to be executed.

[0076] In an example, see Figure 1In step S1, the master chip 10 decodes the function status of the instruction to be executed to obtain the current function to be executed (for example, write data or refresh data). Then, the master chip 10 notifies each memory chip 3i-3j to prepare a clock to execute the instruction.

[0077] In one example, combine Figure 1 and Figure 2 The clock loop 200 coupled to each memory chip 30~3n is set in the corresponding silicon through-hole hybrid bonding layer 20~20n, and includes a main clock path 200a and a backup clock path 200b. The backup clock path 200b is used to back up the main clock path 200a.

[0078] Alternatively, refer to Figure 2 The main clock path 200a and backup clock path 200b in each TSV hybrid bonding layer are constructed through the TSVs and rewiring in the TSV hybrid bonding layer. The location of the rewiring is arranged according to the transmission path of the clock signal, and the present invention does not specifically limit this. For example, in the TSV hybrid bonding layers 20-20n, the main clock path 200a in the lower layer is on the left side of the chip, and the main clock path 200a in the upper layer is on the right side of the chip. Thus, the main clock paths 200a in the TSV hybrid bonding layers 20-20n together form a serpentine routing structure. Similarly, the backup clock paths 200b in the TSV hybrid bonding layers 20-20n together form a serpentine routing structure.

[0079] This example clock control method also includes: Under normal circumstances, each memory chip 30-3n activates the main clock path 200a in the clock loop 200 according to the control signal. However, if the main clock path 200a deviates or a through-silicon via (TSV) fails, the memory chip 200b switches to activate the backup clock path 200b in the clock loop 200 according to the control signal, thereby ensuring chip functionality.

[0080] In one example, combine Figure 1 and Figure 4 In step S2, after receiving the notification from the master chip 10, each memory chip 3i-3j identifies the current function and / or operating mode to be run based on the notification. Based on the identification results, clock arbitration and allocation are performed, and only the clock domains containing the functional modules required to run the current function and / or operating mode are enabled. This achieves finer-grained dynamic division and control of clock domains, reducing power consumption.

[0081] Optionally, combine Figure 1 and Figure 4In step S2, when each memory chip 3i~3j performs clock arbitration and distribution, if the instruction is illegal or conflicting, clock arbitration and distribution (i.e., automatic clock control) are performed in the order of test mode (which is also a running mode), user register, and user function priority from high to low.

[0082] For example, please combine Figure 1 and Figure 4 Taking the memory chip 3i as an example, the functional module 300 in each memory chip 30-3n includes at least two of the following (1)-(5):

[0083] (1) A write module 303b, which is used to write data into the memory array (ARRAY) 300 of the memory chip 3i. The write module 303b is located in the write clock domain. When the instruction to be executed is a write instruction (i.e., the function that the user needs to execute is to write data), in step S2, only the write clock domain is turned on. Thus, in step S6, only the functional modules 303 in the write clock domain, such as the write module 303b, are enabled to work, and the functional modules 303 in other clock domains are disabled.

[0084] (2) A read module 303a, which is used to read data from the memory array 300 of the memory chip 3i. The read module 303a is located in the read clock domain. When the instruction to be executed is a read instruction (i.e., the function to be executed by the user is to read data), in step S2, only the read clock domain is turned on. Thus, in step S6, only the functional modules 303 in the read clock domain, such as the read module 303a, are enabled to operate, and the functional modules 303 in other clock domains are disabled.

[0085] (3) A built-in self-test (BIST) module 303c, which is used to perform a built-in self-test on the memory array 300 of the memory chip 3i to detect error bits in the memory array 300 of the memory chip. The built-in self-test module 303c is located in a logic test control (LTC) clock domain. When the instruction to be run is a built-in self-test instruction (i.e., the function to be run by the user is a built-in self-test), in step S2, only the logic test control clock domain is turned on. Thus, in step S6, only the built-in self-test module 303c and other functional modules 303 in the logic test control clock domain are enabled to operate, and the functional modules 303 in other clock domains are disabled.

[0086] (4) Redundancy repair (RDN) module 303d, which is used to perform redundancy repair on bad points in the memory chip 3i. The redundancy repair module is located in the redundancy repair clock domain. When the instruction to be run is a redundancy repair instruction (i.e., the function to be run by the user is redundancy repair), in step S2, only the redundancy repair clock domain is turned on. Thus, in step S6, only the functional modules 303 in the redundancy repair clock domain, such as the redundancy repair module 303d, are enabled to operate, and the functional modules 303 in other clock domains are disabled.

[0087] (5) Wafer test module 303e, which is used to perform wafer test on the memory chip after wafer manufacturing and before packaging, or to perform relevant tests on the memory chip using a probe at any required appropriate stage before or after delivery. The wafer test module 303e is located in the test mode (TMODE) clock domain. When the instruction to be run is a wafer test instruction (that is, the function that the user needs to run is wafer test or probe CP test), only the test mode clock domain is turned on in step S2, thereby enabling only the wafer test module 303e and other functional modules 303 in the wafer test clock domain to work in step S6, and the functional modules 303 in other clock domains are disabled.

[0088] Therefore, according to the current function that needs to be run, the working clock of the functional module that implements the current function can be switched, and at the same time, the enablement of each functional module that implements the current function can be controlled, and the remaining functional modules can be disabled, thereby achieving finer-grained dynamic division and control of the clock domain and reducing power consumption.

[0089] The following takes the memory chip 3i as an example to explain in detail the steps of adjusting the internal working clock signals clkwi~clkwj of the memory chips 3i~3j that need to work synchronously in step S5 according to the reference clock clkref. Figure 1 and Figure 4 The step of adjusting the internal working clock signal clkwi of the memory chip 3i according to the reference clock clkref includes:

[0090] First, a through-silicon-via impedance calibration is performed on the reference clock clkref to obtain a calibrated reference clock clkref';

[0091] Then, the phase of the internal working clock signal clkwi is roughly adjusted to be aligned with the calibrated reference clock clkref', wherein the initial value of the working clock signal clkwi may be equal to clki;

[0092] Next, a corresponding training sequence is sent to perform phase synchronization training on its working clock signal clkwi, and during the training process, the transition edge of the trained working clock signal clkwi is captured to calculate the phase difference between the trained working clock signal clkwi and the calibrated reference clock clkref', and then the trained working clock signal clkwi is phase compensated until the trained working clock signal clkwi and the calibrated reference clock clkref' are phase synchronized.

[0093] When the working clock signals clkwi-clkwj in each memory chip 3i-3j are phase-synchronized with the calibrated reference clock clkref′, the memory chips 3i-3j are synchronized.

[0094] Thus, the clock transmission of the TSV path is corrected to achieve reliable driving of the clock signal transmitted on the clock loop 200, ensuring the integrity and reliability of the clock signal on the clock path, and further ensuring the reliability of the clock synchronization of the memory chips 3i~3j that need to work synchronously.

[0095] Based on the same invention concept, please refer to Figures 2 to 4 This embodiment also provides a multi-layer stacked chip that can implement clock control between multiple chips using the clock control method for a multi-layer stacked chip of this embodiment. The multi-layer stacked chip comprises a three-dimensionally stacked main chip 10 and n+1 memory chips 30-3n, where n ≥ 1 and is an integer. The main chip 10 is located at the bottom layer, and the n+1 memory chips 30-3n are sequentially stacked above the main chip 10.

[0096] In an example, see Figure 2 Through-silicon-via (TSV) hybrid bonding (HB) is used between the n+1 memory chips 30-3n and between the memory chips and the main chip 10. This means that the multi-layer packaged chip in this example also includes n+1 TSV hybrid bonding layers 20-20n, each of which is disposed between each two adjacent layers of chips in the multi-layer packaged chip.

[0097] Each TSV hybrid bonding layer is equipped with a corresponding clock loop 200. This clock loop 200 includes a primary clock path 200a and a backup clock path 200b. The backup clock path 200b is used to back up the primary clock path 200a. Each TSV hybrid bonding layer 20-20n is also equipped with a corresponding clock loop 200. This clock loop 200 includes a primary clock path 200a and a backup clock path 200b. The backup clock path 200b is used to back up the primary clock path 200a.

[0098] Optionally, the main clock path 200a and the backup clock path 200b in each layer of the through-silicon-via hybrid bonding layer are constructed by through-silicon vias and rewiring in the through-silicon-via hybrid bonding layer. The location of the rewiring is arranged according to the transmission path of the clock signal, and the present invention does not specifically limit this. For example, in the through-silicon-via hybrid bonding layers 20~20n, the main clock path 200a of the lower layer is on the left side of the chip, and the main clock path 200a of the upper layer is on the right side of the chip. Thus, the main clock paths 200a in the through-silicon-via hybrid bonding layers 20~20n together constitute a serpentine routing structure. Similarly, the backup clock paths 200b in the through-silicon-via hybrid bonding layers 20~20n together constitute a serpentine routing structure.

[0099] Under normal circumstances, clock signals are transmitted between each memory chip 30-3n and the main chip 10 through the main clock path 200a in the clock loop 200. However, if the main clock path 200a deviates or the through-silicon via (TSV) fails, the backup clock path 200b of the clock loop 200 is switched to transmit the clock signal, thereby ensuring chip functionality.

[0100] In this embodiment, please combine Figure 1 and Figure 2 The main chip 10 is used to notify each memory chip (at least two of the memory chips 30~3n, defined as 3i~3j, ji≥2) that needs to work synchronously according to the current function to be run to prepare the clock (to achieve Figure 1 Step S1 in the embodiment) detects the phase difference between the clock signals clki~clkj fed back by the clock loop 200 coupled to the memory chips 3i~3j that need to work synchronously, and generates the reference clock clkref according to the detection result (implementing Figure 1 4).

[0101] The n+1 memory chips 30-3n can each be a DRAM or any other suitable type of memory die structure, which is not specifically limited in the present invention. The DRAM can be any suitable type, such as synchronous DRAM (SDRAM) or wide I / O DRAM. The n+1 memory chips 30-3n can each be implemented as an unbuffered dual in-line memory module (UDIMM), a registered DIMM (RDIMM), a load-reduced DIMM (LRDIMM), a fully buffered DIMM (FBDIMM), a small outline DIMM (SODIMM), or the like.

[0102] The master chip 10 can also implement interface conversion between a master device (i.e., a device accessing memory, not shown) and n+1 memory chips 30-3n, complete address decoding and data format conversion (e.g., data bit width) between the master device and the n+1 memory chips 30-3n, and convert read, write, refresh, and other operation instructions issued by the master device into signals recognizable by the n+1 memory chips 30-3n. This implements the necessary control (including control of address signals, data signals, and various command signals) of the master device's access to the n+1 memory chips 30-3n for refresh operations, read and write operations, and other operations. This allows the master device to access (or "use," "operate") the storage resources (i.e., the corresponding storage units) on the n+1 memory chips 30-3n according to user needs.

[0103] The host device may include any type of processing device with computing processing capabilities, such as a central processing unit (CPU), a digital signal processor (DSP), a network processor, an application processor (AP), a field programmable gate array (FPGA), a dedicated processor, etc. The processing device may be configured to execute instructions or software (including code, operating system or application, etc.) that can be executed by one or more computers, firmware or a combination thereof.

[0104] In other embodiments, the main chip 10 may also be a buffer chip with a memory controller inside, and the main device is set outside the main chip 10. When the main device is integrated into a logic chip, the logic chip is stacked under the main chip 10.

[0105] Please refer to Figure 3 In one example, the main chip 10 includes a functional status decoder 101, a notification module 102, a phase detector 103 and a reference clock source 104. The functional status decoder 101, the notification module 102, the phase detector 103 and the reference clock source 104 can all be integrated into a memory controller (not shown) inside the main chip 10.

[0106] Among them, please refer to Figure 2 、 Figure 3 and Figure 4 The function state decoder 101 is used to decode the function state of the current instruction to obtain the current function to be run. The notification module 102 is coupled to the function state decoder 101 and the clock arbitration and distribution circuit 301 of each memory chip 30-3n, and is used to notify the clock arbitration and distribution circuit 301 of each memory chip 3i-3j that needs to work synchronously in the memory chips 30-3n to prepare the clocks clki-clkj according to the decoding result of the function state decoder 101 (that is, the function state decoder 101 and the notification module 102 are used to implement Figure 1 Step S1 in ( ). Wherein, 0≤i≤j≤n, and i, j, and n are all integers.

[0107] Please refer to Figures 2 to 4 The phase detector 103 is coupled to the clock circuit 200 connected to each memory chip 30-3n and is used to detect phase differences between the clock signals clki-clkj fed back from the clock circuit 200 connected to each memory chip 3i-3j to be operated synchronously. The reference clock source 104 is coupled to the phase detector 103 and the clock adjustment circuit 302 of each memory chip 30-3n and is used to generate a corresponding reference clock clkref based on the detection result of the phase detector to provide to each memory chip 3i-3j to be operated synchronously.

[0108] Please refer to Figure 1 、 Figure 2 and Figure 4 In this embodiment, each memory chip 30 to 3n includes a memory array 300, a clock arbitration and distribution circuit 301, a clock adjustment circuit 302, and multiple functional modules 303. The multiple functional modules 303 can implement multiple different functions.

[0109] The functions of the clock arbitration and distribution circuit 301 and the clock adjustment circuit 302 are described in detail below by taking the memory chip 3i (0≤i≤n, and i is an integer) as an example.

[0110] Please refer to Figure 1 、 Figure 2 and Figure 4 In the memory chip 3i, the clock arbitration and distribution circuit 301 is coupled to the main chip 10, and is used to receive notifications sent by the main chip 10, and perform clock arbitration according to the notification to generate corresponding clock signals clki, control signals (unmarked) and multi-chip synchronization signals (unmarked), wherein the clock signal clki is sent to the main chip 10 through the clock loop 200 in the silicon through-via bonding layer below it on the one hand, and is sent to the clock adjustment circuit 302 on the other hand to generate the working clock signal clkwi. The clock arbitration and distribution circuit 301 is also used to enable the corresponding functional module 303 in the memory chip 3i to work according to the phase-synchronized working clock signal clkwi output by the multi-chip synchronization signal and the clock adjustment circuit 302 after the working clock signals clkwi~clkwj of the memory chips 3i~3j that need to work synchronously are phase-synchronized. That is, for the memory chip 3i, the clock arbitration and distribution circuit 301 in the memory chip 3i is used to implement Figure 1 Step S2 in the above example enables the corresponding functional module 303 to implement Figure 1In step S6, the functional module 303 enabled by its clock arbitration and distribution circuit 301 is used to receive instructions and work synchronously with the functional modules 303 in other memory chips 3i+1~3j that need to work synchronously.

[0111] In an example, see Figure 4 In the memory chip 3i, the clock arbitration and distribution circuit 301 may include an identification module 301a and an arbitration and distribution module 301b.

[0112] The identification module 301a is used to receive notifications sent by the main chip 10 and identify the current function and / or operating mode to be run based on the received notification. The identification module 301a can be any suitable design such as a decoding circuit, a state control circuit, or a register, and the present invention specifically limits this.

[0113] The arbitration and distribution module 301b is coupled to the identification module 301a, the clock adjustment circuit 302, and each functional module 303, and is used to perform clock arbitration and distribution based on the identification result of the identification module 301a, thereby generating corresponding clock signals clki, control signals (unlabeled), and multi-chip synchronization signals. After achieving phase synchronization with the operating clock signals clkwi-clkwj of the memory chips 3i-3j that require synchronous operation, the arbitration and distribution module 301b activates only the clock domains where the functional modules 303 required for the current function and / or operating mode are located, based on the multi-chip synchronization signals and the phase-synchronized operating clock signals clkwi output by the clock adjustment circuit 302. The module then generates corresponding enable signals en to enable the functional modules 303 in these clock domains. If an instruction is illegal or conflicting during clock arbitration and distribution, the arbitration and distribution module 301b performs clock arbitration and distribution in descending order of priority: test mode, user register, and user function.

[0114] Optionally, each memory chip 3i~3j that needs to work synchronously receives instructions and works synchronously according to the multi-chip synchronization signal and the phase-synchronized working clock signals clkwi~clkwj, and the functional module 303 in the memory chip 3i includes at least two of the following (1)~(5):

[0115] (1) A write module 303b, which is used to write data into the memory array (ARRAY) 300 of the memory chip 3i. The write module 303b is located in the write clock domain. When the instruction to be executed is a write instruction (i.e., the function that the user needs to execute is to write data), and after the phase synchronization of the working clock signals clkwi~clkwj of the memory chips 3i~3j that need to work synchronously is achieved, the arbitration and distribution module 301b only turns on the write clock domain when distributing the clock (i.e., distributing the working clock signal clkwi), thereby only enabling the functional modules 303 in the write clock domain, such as the write module 303b, to work. The functional modules 303 in other clock domains are disabled by the arbitration and distribution module 301b.

[0116] (2) A read module 303a, which is used to read data from the memory array 300 of the memory chip 3i. The read module 303a is located in the read clock domain. When the instruction to be executed is a read instruction (i.e., the function that the user needs to execute is to read data), and after the phase synchronization of the working clock signals clkwi~clkwj of the memory chips 3i~3j that need to work synchronously is achieved, the arbitration and distribution module 301b only turns on the read clock domain when distributing the clock (i.e., distributing the working clock signal clkwi), thereby only enabling the functional modules 303 in the read clock domain, such as the read module 303a, to work. The functional modules 303 in other clock domains are disabled by the arbitration and distribution module 301b.

[0117] (3) A built-in self-test (BIST) module 303c is used to perform a built-in self-test on the memory array 300 of the memory chip 3i to detect error bits in the memory array 300 of the memory chip. The built-in self-test module 303c is located in the logic test control (LTC) clock domain. When the instruction to be run is a built-in self-test instruction (i.e., the function that the user needs to run is a built-in self-test), and after the phase synchronization of the working clock signals clkwi~clkwj of the memory chips 3i~3j that need to work synchronously is achieved, the arbitration and distribution module 301b only turns on the built-in self-test clock domain when distributing the clock (i.e., distributing the working clock signal clkwi), thereby only enabling the built-in self-test module 303 and other functional modules 303 in the built-in self-test clock domain to work, and the functional modules 303 in other clock domains are disabled by the arbitration and distribution module 301b.

[0118] (4) Redundancy repair (RDN) module 303d, which is used to perform redundant repair on bad points in the memory chip 3i. The redundant repair module is located in the redundant repair clock domain. When the instruction to be run is a redundant repair instruction (i.e., the function that the user needs to run is redundant repair), and after the phase synchronization of the working clock signals clkwi~clkwj of the memory chips 3i~3j that need to work synchronously is achieved, the arbitration and distribution module 301b only turns on the redundant repair clock domain when distributing the clock (i.e., distributing the working clock signal clkwi), thereby only enabling the functional modules 303 in the redundant repair clock domain, such as the redundant repair module 303c, to work. The functional modules 303 in other clock domains are disabled by the arbitration and distribution module 301b.

[0119] (5) Wafer test module 303e, which is used to perform wafer test on the memory chip 3i after wafer manufacturing and before packaging of the memory chip 3i, or to perform relevant tests on the memory chip 3i using a probe at any required appropriate stage before or after leaving the factory. The wafer test module 303e is located in the test mode (TMODE) clock domain. When the instruction to be run is a wafer test instruction (that is, the function that the user needs to run is wafer test or CP probe test), and after the phase synchronization of the working clock signals clkwi~clkwj of the memory chips 3i~3j that need to work synchronously is achieved, the arbitration and distribution module 301b only turns on the wafer test clock domain when distributing the clock (that is, distributing the working clock signal clkwi), and then only enables the functional modules 303 such as the wafer test module 303e in the wafer test clock domain to work, and the functional modules 303 in other clock domains are disabled by the arbitration and distribution module 301b.

[0120] Therefore, the arbitration and allocation module 301b can switch the working clock clkwi of the functional module that implements the current function according to the current function that needs to be run (for example, when data needs to be read, clkwi switches to the read clock, when data needs to be written, clkwi switches to the write clock, and when refresh is required, clkwi switches to the refresh clock), and at the same time control the enablement of each functional module that implements the current function and disable the remaining functional modules, thereby realizing finer-grained dynamic division and control of the clock domain and reducing power consumption.

[0121] Please continue to refer to Figure 2 and Figure 4 The clock adjustment circuit 302 of the memory chip 3i is coupled to its clock arbitration and distribution circuit 301 and the corresponding clock loop 200, and is used to open the clock loop 200 to which it is coupled according to the control signal sent by the clock arbitration and distribution circuit 301, so as to feed back the clock signal clki generated by the clock arbitration and distribution circuit 301 to the main chip 10 (i.e., to achieve Figure 1Step S3 in the process), and receiving the reference clock clkref sent by the master chip 10 through the opened clock loop, and adjusting the phase of the working clock signal clkwi of the memory chip 3i according to the reference clock clkref, until the working clock signals clkwi~clkwj of the memory chips 3i~3j that need to work synchronously are phase synchronized (i.e., achieving Figure 1 , step S5 in the process).

[0122] Alternatively, refer to Figure 4 The clock adjustment circuit 302 includes a TSV driving unit 302 a and a digitally controlled delay line 302 b.

[0123] The TSV driving unit 302a is coupled to the clock arbitration and distribution circuit 301, the digital control delay line 302b and the corresponding clock loop 200, and is used to open the clock loop 200 coupled thereto according to the control signal sent by the clock arbitration and distribution circuit 301, so as to feed back the clock signal clki generated by the clock arbitration and distribution circuit 301 to the main chip 10 (i.e., to achieve Figure 1 ), and receiving the reference clock clkref sent by the master chip 10 through the clock loop 200, and performing silicon via impedance calibration on the reference clock clkref to eliminate the influence of corresponding parameter changes on the silicon via impedance and delay, wherein the parameters include but are not limited to manufacturing process, operating voltage and / or temperature.

[0124] The digitally controlled delay line 302b is coupled to the TSV driving unit 302a and the clock arbitration and distribution circuit 301, and is used to first roughly adjust the phase alignment of the working clock signal clkwi output by the delay line with the calibrated reference clock clkref', and then send a corresponding training sequence to the delay line to perform phase synchronization training on the working clock signal clkwi output by the delay line, and capture the jump edge of the trained working clock signal clkwi during the training process to calculate the phase difference between the trained working clock signal clkwi and the calibrated reference clock clkref', and then perform phase compensation on the trained working clock signal clkwi until the trained working clock signal clkwi is phase-synchronized with the calibrated reference clock clkref', and the phase-synchronized working clock signal clkwi is sent to the clock arbitration and distribution circuit 301, so that the clock arbitration and distribution circuit 301 can distribute the working clock signal clkwi to the corresponding functional module 303, thereby realizing the corresponding function. That is, the TSV driving unit 302a and the digitally controlled delay line 302b are jointly implemented Figure 1 Step S5 in .

[0125] The initial value of the working clock signal of each memory chip may be equal to the clock signal it sends to the master chip 10. For example, the initial value of the working clock signal clkw0 of the memory chip 30 may be equal to the clock signal clk0 it sends to the master chip 10. The initial value of the working clock signal clkwi of the memory chip 3i may be equal to the clock signal clki it sends to the master chip 10. The initial value of the working clock signal clkwj of the memory chip 3j may be equal to the clock signal clkj it sends to the master chip 10. The initial value of the working clock signal clkwn of the memory chip 3n may be equal to the clock signal clkn it sends to the master chip 10.

[0126] It should be understood that, in addition to the above-mentioned circuit modules, the main chip 10 and each memory chip 30~3n of this embodiment may also have other circuit modules to implement other functions. For example, the main chip 10 or the memory chips 30~3n may also include circuits for performing error correction (such as ECC correction) on the memory chip, clock and frequency control circuits (such as phase-locked loops PLLs), circuits for managing power consumption and temperature, first-in-first-out queue registers (FIFOs), and any other required circuits. These circuits are not the focus of the present invention and are therefore not described in detail here.

[0127] In addition, the division of the various circuit modules within the main chip 10 and each memory chip 30~3n of this embodiment is schematic, mainly a logical function division. In actual implementation, there may be other division methods. In the embodiments of the present application, the various circuit modules within the main chip 10 or each memory chip can be integrated into a physical module, or each circuit module can exist physically separately, or two or more circuit modules can be integrated into a physical module. The above-mentioned integrated circuit modules can be implemented in whole or in part through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in accordance with the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., an SSD).

[0128] In summary, in the multi-layer stacked chip and its clock control method provided by the present invention, its memory chip can prepare the corresponding clock according to the current function to be run, and through clock training, achieve clock synchronization between multiple memory chips that need to run synchronously, reducing the communication delay between memory chip layers and the resource occupation of asynchronous FIFOs, thereby reducing the communication cost between multi-layer stacked chips. In addition, its memory chip can identify information such as the current function and / or operating mode to be run based on notifications, and then perform clock arbitration and allocation, so that the division of its on-chip clock domain can be dynamically adjusted according to the different functions being run, realizing automatic division and control of the clock domain, thereby reducing the cost of chip clock domain division and dynamic power consumption.

[0129] The above description is only a description of the preferred embodiment of the present invention and does not limit the scope of the present invention. Any changes and modifications made by ordinary technicians in the field of the present invention based on the above disclosure are within the scope of protection of the technical solution of the present invention.

Claims

1. A clock control method for a multi-layer packaged chip, characterized in that: The multi-layer stacked chip comprises a three-dimensionally stacked main chip and multiple memory chips, and the clock control method comprises: The main chip notifies each memory chip that needs to work synchronously to prepare a clock according to the current function that needs to be run; Each memory chip that needs to work synchronously performs clock arbitration according to the received notification, thereby generating corresponding clock signals, control signals and multi-chip synchronization signals; Each of the memory chips that need to work synchronously opens its coupled clock loop according to the control signal, and then feeds the clock signal back to the master chip; The master chip detects the phase difference between the clock signals of the memory chips that need to work synchronously, and generates a reference clock according to the detection result; Each of the memory chips that need to work synchronously adjusts its internal working clock signal according to the reference clock until the working clock signals of the memory chips that need to work synchronously achieve phase synchronization; Each of the memory chips that needs to work synchronously receives instructions and works synchronously according to the multi-chip synchronization signal and the phase-synchronized working clock signal.

2. The clock control method according to claim 1, wherein: The master chip performs function status decoding to obtain the current function that needs to be run, and then notifies each memory chip that needs to work synchronously to prepare a clock.

3. The clock control method according to claim 1, wherein: The multiple memory chips are hybrid bonded to each other and to the main chip via through silicon vias, and the clock loop is arranged in the through silicon via hybrid bonding layer.

4. The clock control method according to claim 3, wherein: The clock loop includes a main clock path and a backup clock path, wherein the backup clock path is used to back up the main clock path. The clock control method further includes: when the main clock path coupled to the memory chip is offset or the silicon through via fails, the memory chip switches the backup clock path to open.

5. The clock control method according to claim 3, wherein: The step of adjusting the internal working clock signal of each memory chip that needs to work synchronously according to the reference clock includes: performing through silicon via impedance calibration on the reference clock; Coarsely adjusting the phase alignment between the internal working clock signal and the calibrated reference clock; Send a corresponding training sequence to perform phase synchronization training on the working clock signal, and capture the jumping edge of the trained working clock signal during the training process to calculate the phase difference between the trained working clock signal and the calibrated reference clock, and then perform phase compensation on the trained working clock signal until the trained working clock signal and the calibrated reference clock achieve phase synchronization.

6. The clock control method according to any one of claims 1 to 5, wherein: After receiving the notification, each memory chip identifies the current function and / or operating mode that needs to be run according to the notification, and performs clock arbitration and allocation based on the identification result, thereby only turning on the clock domain where the functional modules required for the current function and / or operating mode are located.

7. The clock control method according to claim 6, wherein: When each memory chip performs clock arbitration and distribution, the clock arbitration and distribution are performed in the order of the priorities of the test mode, the user register, and the user function from high to low.

8. The clock control method according to claim 6, wherein: The functional modules include at least two of the following (1) to (5): (1) a write module, which is used to write data into the memory chip, wherein the write module is located in a write clock domain, and when the instruction is a write instruction, the clock control method only turns on the write clock domain; (2) a read module, which is used to read data from the memory chip, wherein the read module is located in a read clock domain, and when the instruction is a read instruction, the clock control method only turns on the read clock domain; (3) a built-in self-test module, which is used to perform a built-in self-test on the memory chip to detect bad pixels in the memory chip, the built-in self-test module is located in a logic test control clock domain, and when the instruction is a built-in self-test instruction, the clock control method only turns on the logic test control clock domain; (4) a redundant repair module, which is used to perform redundant repair on bad points in the memory chip, wherein the redundant repair module is located in a redundant repair clock domain, and when the instruction is a redundant repair instruction, the clock control method only turns on the redundant repair clock domain; (5) A wafer test module, which is used to perform wafer testing on the memory chip after the memory chip is manufactured and before it is packaged. The wafer test module is located in a test mode clock domain. When the instruction is a wafer test instruction, the clock control method only turns on the test mode clock domain.

9. A multi-layer stacked chip, comprising a three-dimensionally stacked main chip and multiple memory chips, characterized in that: The master chip is used to notify each memory chip that needs to work synchronously to prepare a clock according to the current function that needs to be run, detect the phase difference between the clock signals fed back by each memory chip that needs to work synchronously, and generate a reference clock based on the detection result; The memory chip has: Multiple functional modules are used to implement corresponding functions; a clock arbitration and distribution circuit, configured to receive the notification, perform clock arbitration according to the notification, generate corresponding clock signals, control signals, and multi-chip synchronization signals, and enable the corresponding functional modules to operate according to the multi-chip synchronization signal and the phase-synchronized working clock signals after the working clock signals of the memory chips to be operated synchronously are phase-synchronized; A clock adjustment circuit is coupled to the clock arbitration and distribution circuit and the corresponding clock loop, and is used to open the clock loop to which it is coupled according to the control signal to feed back the clock signal generated by the clock arbitration and distribution circuit to the main chip, and to receive the reference clock through the clock loop, and to adjust the working clock signal inside the memory chip according to the reference clock until the working clock signals of each of the memory chips that need to work synchronously are phase synchronized.

10. The multi-layer packaged chip according to claim 9, wherein: The main chip includes: A function status decoder, used for decoding the function status to obtain the current function to be run; a notification module coupled to the functional status decoder and the clock arbitration and distribution circuit of each memory chip, and configured to notify each memory chip that needs to work synchronously to prepare a clock according to a decoding result of the functional status decoder; a phase detector coupled to the clock loop to which each memory chip is coupled and used to detect a phase difference between clock signals fed back by each memory chip that needs to work synchronously; A reference clock source is coupled to the phase detector and the clock adjustment circuit of each memory chip, and is used to generate a corresponding reference clock according to the detection result of the phase detector.

11. The multi-layer packaged chip according to claim 9, wherein: The multiple memory chips are hybrid bonded to each other and to the main chip via through silicon vias, and the clock loop is arranged in the through silicon via hybrid bonding layer.

12. The multi-layer packaged chip according to claim 11, wherein: The clock loop includes a main clock path and a backup clock path. The backup clock path is used to back up the main clock path. The main chip is also used to: when the main clock path of the memory chip is offset or the silicon through via fails, the clock adjustment circuit of the memory chip switches the backup clock path to open.

13. The multi-layer packaged chip according to claim 11, wherein: The clock adjustment circuit includes: a TSV driving unit coupled to the clock arbitration and distribution circuit and a corresponding clock loop, and configured to open the clock loop according to the control signal to feed back the clock signal generated by the clock arbitration and distribution circuit to the master chip, receive the reference clock through the clock loop, and perform through-silicon via impedance calibration on the reference clock; A digitally controlled delay line is coupled to the TSV driving unit and the clock arbitration and distribution circuit, and is used to first coarsely adjust the phase alignment of the working clock signal output by the delay line with the calibrated reference clock, and then send a corresponding training sequence to the delay line to perform phase synchronization training on the working clock signal output by the delay line, and capture the jump edge of the trained clock signal during the training process to calculate the phase difference between the trained working clock signal and the calibrated reference clock, and then perform phase compensation on the trained working clock signal until the trained working clock signal achieves phase synchronization with the calibrated reference clock, and sends the phase-synchronized working clock signal to the clock arbitration and distribution circuit.

14. The multi-layer packaged chip according to any one of claims 9 to 13, wherein: The multiple functional modules inside each memory chip are located in different clock domains; The clock arbitration and distribution circuit includes: an identification module, configured to identify the current function and / or operation mode to be run according to the notification after receiving the notification; An arbitration and distribution module is coupled to the identification module, the clock adjustment circuit and each of the functional modules, and is used to perform clock arbitration and distribution according to the identification result of the identification module, thereby only turning on the clock domain where the functional modules required for the current function and / or operating mode are located.

15. The multi-layer packaged chip according to claim 14, wherein: When performing clock arbitration and distribution, the arbitration and distribution module also performs clock arbitration and distribution in the order of the priorities of the test mode, the user register, and the user function from high to low.

16. The multi-layer packaged chip according to claim 14, wherein: Each of the memory chips that need to work synchronously receives instructions and works synchronously according to the multi-chip synchronization signal and the phase-synchronized working clock signal, and the functional module includes at least two of the following (1) to (5): (1) a write module, which is used to write data into the memory chip. The write module is located in the write clock domain. When the instruction is a write instruction, the arbitration and allocation module only turns on the write clock domain. (2) a read module, which is used to read data from the memory chip. The read module is located in the read clock domain. When the instruction is a read instruction, the arbitration and allocation module only turns on the read clock domain. (3) a built-in self-test module, which is used to perform a built-in self-test on the memory chip to detect bad pixels in the memory chip, wherein the built-in self-test module is located in a logic test control clock domain, and when the instruction is a built-in self-test instruction, the arbitration and distribution module only turns on the logic test control clock domain; (4) a redundancy repair module, which is used to perform redundancy repair on bad points in the memory chip. The redundancy repair module is located in a redundancy repair clock domain. When the instruction is a redundancy repair instruction, the arbitration and distribution module only turns on the redundancy repair clock domain. (5) A wafer test module, which is used to perform wafer testing on the memory chip after the memory chip completes wafer manufacturing and before packaging. The wafer test module is located in the test mode clock domain. When the instruction is a wafer test instruction, the arbitration and allocation module only turns on the test mode clock domain.

Citation Information

Patent Citations

  • Method for achieving low power consumption of three-dimensional measurement chip

    CN105676995A

  • Time sequence synchronization method based on 3D integrated chip

    CN114125338A