Multi-layer Stacked Packaging Method and Related Equipment for Edge AI Computing and Storage Chips

By building a cooling channel after the chip is cut in a multi-chip stacked package, the problem of chip overheating is solved, and high-performance, low-power edge AI computing chips are realized.

CN119677118BActive Publication Date: 2025-06-10SHENZHEN OSCOO TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510189405.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

In multi-chip stacked packages, chips are close to each other and the heat dissipation channel becomes longer, making it difficult to dissipate heat inside the chip, which easily causes the chip to overheat.

Method used

The cooling efficiency is improved by building a cooling channel after chip cutting. The specific steps include: obtaining the functional requirements of the preset edge AI computing chip, determining the chip architecture and sequence arrangement; TSV etching is performed through deep reactive ion etching technology, cutting to form independent cutting chips, and building cooling channels on these chips, and finally multi-chip integration and packaging.

Benefits of technology

It effectively solves the problem of chip overheating, improves heat dissipation efficiency, and realizes high-performance, low-power consumption edge AI computing and storage chips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119677118B_ABST
    Figure CN119677118B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-chip stacking packaging method and related equipment for an edge AI computing and storage chip, comprising the following steps: obtaining the functional requirements of a preset edge AI computing and storage chip, and determining the chip architecture of the edge AI computing and storage chip and the corresponding sequential arrangement according to the functional requirements; performing TSV etching on a preset wafer through deep reactive ion etching technology to obtain a wafer with a via array; cutting the wafer with the via array to obtain independent diced chips; constructing cooling channels for the independent diced chips to obtain cooling function chips; performing multi-chip integration and packaging on the cooling function chips based on the chip architecture and the corresponding sequential arrangement to obtain an edge AI computing and storage chip; solving the technical problem that the stacked chips are close to each other, the heat dissipation channels become longer, resulting in difficult heat dissipation inside the chips and easy overheating of the chips.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of edge AI computing and storage chips, and particularly to a multi-layer stacked packaging method and related equipment for edge AI computing and storage chips. Background Art

[0002] Edge AI computing has developed rapidly in recent years, posing increasingly high requirements for low power consumption, high performance, and miniaturized computing power. When traditional von Neumann architecture chips process large-scale data, especially complex AI-related calculations, due to the frequent transfer of data between the processor and the memory, serious power consumption and latency problems occur, which is known as the "memory wall" bottleneck. To break through this bottleneck, the technology of computing in memory emerged. It closely combines storage units and computing units, performs calculations inside the storage units, thereby significantly reducing data transfer, lowering power consumption and latency, and improving computing efficiency. This is crucial for edge AI devices such as autonomous driving, smartphones, and Internet of Things devices, as these devices usually have very strict requirements for power consumption and real-time performance.

[0003] However, with the continuous increase in the complexity of edge AI applications, the demand for computing power has also increased sharply. The improvement of the computing power of a single chip is restricted by the slowdown of Moore's Law and power consumption limitations. Therefore, multi-chip stacked packaging technology has become the key to further improving computing power. By integrating multiple chips in a package through a three-dimensional stacking method, the integration density and interconnection bandwidth of the chips can be significantly improved, thereby achieving higher computing power density and lower system power consumption. However, multi-chip stacking also brings new challenges, such as heat dissipation problems. The stacked chips are close to each other, and the heat dissipation channels become longer, resulting in difficulty in dissipating the heat inside the chips, easily causing the chips to overheat, and affecting performance and reliability.

[0004] Therefore, how to effectively solve the heat dissipation problem in multi-chip stacked packaging has become the key to realizing high-performance and low-power edge AI computing and storage chips. The "multi-layer stacked packaging method for an edge AI computing and storage chip" proposed in the above patent precisely aims at this problem and attempts to improve the heat dissipation efficiency by constructing a cooling channel after chip dicing, thereby realizing the manufacture of high-performance edge AI computing and storage chips. This is of great significance for promoting the development of edge AI technology and meeting the growing demand for computing power. Summary of the Invention

[0005] The main object of the present invention is to provide a multi-layer stacked packaging method and related equipment for edge AI computing and storage chips, which solves the technical problem that the stacked chips are close to each other, the heat dissipation channels become longer, resulting in difficulty in dissipating the heat inside the chips, and easily causing the chips to overheat.

[0006] To achieve the above object, the present invention provides a multi-layer stacked packaging method for edge AI computing and storage chips, including the following steps:

[0007] Obtain the functional requirements of a preset edge AI computing and storage chip, and determine the chip architecture of the edge AI computing and storage chip and the corresponding sequential arrangement based on the functional requirements;

[0008] Perform TSV etching on a preset wafer through deep reactive ion etching technology to obtain a wafer with a via array;

[0009] Cut the wafer with the via array to obtain independent diced chips; and construct cooling channels for the independent diced chips to obtain chips with cooling functions;

[0010] Perform multi-chip integration and packaging on the chips with cooling functions based on the chip architecture and the corresponding sequential arrangement to obtain an edge AI computing and storage chip.

[0011] Further, the obtaining the functional requirements of a preset edge AI computing and storage chip, and determining the chip architecture of the edge AI computing and storage chip and the corresponding sequential arrangement based on the functional requirements includes:

[0012] Obtain the target application scenario of a preset AI computing and storage chip, and determine the functional requirements of the AI computing and storage chip based on the target application scenario; wherein, the functional requirements include throughput, latency, and power consumption;

[0013] Perform hardware architecture search based on the functional requirements through a preset neural network search algorithm to obtain a candidate architecture set;

[0014] Screen the candidate architecture set through a preset power consumption evaluation model to obtain a low-power architecture;

[0015] Perform on-chip storage architecture design on the low-power architecture based on the functional requirements and the target application scenario to determine the chip architecture of the edge AI computing and storage chip and the corresponding sequential arrangement; wherein, the on-chip storage architecture design includes determining the type, capacity, bit width, structure of the memory, and the data interaction method with the computing unit to meet the data access mode and bandwidth requirements of the target scenario application.

[0016] Further, the performing TSV etching on a preset wafer through deep reactive ion etching technology to obtain a wafer with a via array includes:

[0017] Perform thickness measurement on a preset wafer based on optical interferometry to obtain a wafer thickness map, and perform thickness uniformity analysis on the wafer thickness map to obtain a wafer with uniform thickness;

[0018] The thickness-uniform wafer is subjected to TSV initial etching using deep reactive ion etching technology to obtain an initially etched wafer, and the sidewalls of the initially etched wafer are passivated to obtain a passivated wafer;

[0019] A barrier layer is deposited on the passivated wafer by atomic layer deposition technology to obtain a barrier layer wafer, and the barrier layer wafer is lithographically patterned to obtain a patterned wafer;

[0020] Deep hole etching is performed on the patterned wafer based on the Bosch etching process to obtain a deep hole wafer;

[0021] An insulating layer is deposited on the deep hole wafer by chemical vapor deposition technology to obtain an insulating layer wafer, and the insulating layer wafer is chemically mechanically polished to obtain a through-hole array wafer.

[0022] Further, the through-hole array wafer is cut to obtain independent diced chips; and a cooling channel is constructed for the independent diced chips to obtain a cooling function chip, including:

[0023] The internal structure of the through-hole array wafer is scanned by a preset three-dimensional X-ray tomography technique to obtain a three-dimensional structure distribution map, and a cutting path is planned for the through-hole array wafer based on the three-dimensional structure distribution map to obtain a cutting path;

[0024] By multi-wavelength laser cutting technology, the through-hole array wafer is layer-cut based on the cutting path to obtain preliminarily diced chips, and the edges of the preliminarily diced chips are chamfered to obtain independent diced chips;

[0025] The surface of the independent diced chip is atomically planarized by atomic layer etching technology to obtain a planarized chip;

[0026] A cooling channel pattern is constructed for the planarized chip by microfluidic channel design technology to obtain a microchannel structure, and the microchannel structure is anisotropically etched to obtain a three-dimensional flow channel network;

[0027] The three-dimensional flow channel network is subjected to superhydrophilic surface modification to obtain a surface wetting property chip, and a nanoscale functional coating is deposited on the surface wetting property chip to obtain a hydrodynamic chip;

[0028] Using microscale heat transfer analysis technology, the heat transfer performance of the hydrodynamic chip is evaluated to obtain thermal resistance and temperature rise distribution data. If the thermal resistance and temperature rise distribution data are within the preset safe operating range, a cooling function chip is obtained;

[0029] If not, perform in-depth analysis on the thermal resistance and temperature rise distribution data through a preset genetic algorithm to obtain optimized cooling channel structure parameters, and optimize the hydrodynamic chip based on the optimized cooling channel structure parameters to obtain a cooling function chip.

[0030] Further, perform internal structure scanning on the through-hole array wafer through a preset three-dimensional X-ray tomography technique to obtain a three-dimensional structure distribution map, and plan a cutting path for the through-hole array wafer based on the three-dimensional structure distribution map to obtain a cutting path, including:

[0031] Perform phase-contrast X-ray tomography scanning on the through-hole array wafer to obtain an original diffraction image, and perform phase reconstruction on the original diffraction image to obtain a multi-layer diffraction pattern;

[0032] Perform digital reconstruction on the multi-layer diffraction pattern through a spatial frequency domain filtering technique to obtain density distribution data, and perform three-dimensional voxel reconstruction on the density distribution data to obtain a voxel distribution map;

[0033] Perform three-dimensional reconstruction on the voxel distribution map through an inverse projection algorithm to obtain a three-dimensional structure distribution map, and plan a cutting area for the three-dimensional structure distribution map to obtain cutting area parameters; wherein, the cutting area parameters include cutting critical points, stress-sensitive areas, and material dividing lines;

[0034] Utilize topology optimization technology to plan a cutting path for the through-hole array wafer based on the cutting area parameters to obtain a cutting path.

[0035] Further, perform multi-chip integration and packaging on the cooling function chip based on the chip architecture and the corresponding sequential arrangement of the chip architecture to obtain an edge AI computing and storage chip, including:

[0036] Arrange the cooling function chips based on the chip architecture and the corresponding sequential arrangement of the chip architecture to obtain arranged chips;

[0037] Interconnect the arranged chips using micro-bump bonding technology to obtain bump-interconnected chips, and perform underfill on the bump-interconnected chips to obtain underfilled chips;

[0038] Stack the underfilled chips through a high-density three-dimensional stacking technology to obtain a three-dimensional chip stack;

[0039] Package the three-dimensional chip stack to obtain a WLP package chip, and perform wire bonding on the WLP package chip to obtain a wire-bonded chip;

[0040] The RDL of the wire-bonded chip is constructed by the redistribution layer technology to obtain an RDL chip, and the RDL chip is electrically tested to obtain electrical test results;

[0041] If there are RDL chips in the electrical test results that do not meet the preset standards, the RDL chips that do not meet the preset standards are removed to obtain qualified RDL chips;

[0042] Based on the system-level packaging technology, the qualified RDL chips are system-level packaged to obtain edge AI computing and storage chips.

[0043] Further, the interconnected chips are obtained by using the micro-bump bonding technology to interconnect the arranged chips, and the underfilled chips are obtained by underfilling the interconnected chips with micro-bumps, including:

[0044] The UBM metal layer is sputtered on the arranged chips to obtain UBM structure chips;

[0045] The micro-bumps are grown on the UBM structure chips by ultrasonic-assisted electroplating technology to obtain micro-bump chips;

[0046] The chips with micro-bumps are aligned to obtain aligned chips with micro-bumps, and the aligned chips with micro-bumps are pre-bonded at low temperature to obtain preliminarily bonded chips;

[0047] Based on the thermocompression bonding technology, the preliminarily bonded chips are thermocompression bonded to obtain interconnected chips with micro-bumps;

[0048] The interconnected chips with micro-bumps are underfilled by capillary action to obtain filled interconnected chips with micro-bumps, and the filled interconnected chips with micro-bumps are cured to obtain underfilled chips.

[0049] The present invention also provides a multi-layer stacked packaging device for edge AI computing and storage chips, including:

[0050] An acquisition module, configured to acquire the functional requirements of the preset edge AI computing and storage chips, and determine the chip architecture of the edge AI computing and storage chips and the corresponding sequential arrangement based on the functional requirements;

[0051] A via module, configured to perform TSV etching on a preset wafer by deep reactive ion etching technology to obtain a wafer with a via array;

[0052] A cutting module, configured to cut the wafer with the via array to obtain independent cut chips; and construct cooling channels for the independent cut chips to obtain chips with cooling functions;

[0053] A packaging module is used to perform multi-chip integration and packaging on the cooling function chips based on the chip architecture and the corresponding sequential arrangement of the chip architecture, so as to obtain an edge AI computing and storage chip.

[0054] The present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0055] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0056] The present invention provides a multi-layer stacking packaging method for an edge AI computing and storage chip, including the following steps: obtaining the functional requirements of a preset edge AI computing and storage chip, and determining the chip architecture of the edge AI computing and storage chip and the corresponding sequential arrangement based on the functional requirements; performing TSV etching on a preset wafer through a deep reactive ion etching technique to obtain a through-hole array wafer; cutting the through-hole array wafer to obtain independent cut chips; and constructing cooling channels for the independent cut chips to obtain cooling function chips; performing multi-chip integration and packaging on the cooling function chips based on the chip architecture and the corresponding sequential arrangement to obtain an edge AI computing and storage chip; through the above technical means, the technical problem that the stacked chips are close to each other, the heat dissipation channels become longer, resulting in difficult heat dissipation inside the chips and easy overheating of the chips is solved. The technical effect that multi-chip stacking packaging can shorten the interconnection distance between chips, thereby improving the chip interconnection bandwidth, reducing data transmission delay, and enhancing the overall performance of the system is achieved. Description of the Drawings

[0057] Figure 1 is a schematic diagram of the steps of the multi-layer stacking packaging method for an edge AI computing and storage chip in an embodiment of the present invention;

[0058] Figure 2 is a structural block diagram of the multi-layer stacking packaging device for an edge AI computing and storage chip in an embodiment of the present invention;

[0059] Figure 3 is a structural schematic block diagram of a computer device in an embodiment of the present invention.

[0060] The realization, functional characteristics, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. Detailed Embodiments

[0061] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0062] As Figure 1 shown, Figure 1 is a schematic diagram of the steps of a multi-layer stacked packaging method for an edge AI computing and storage chip in an embodiment of the present invention;

[0063] An embodiment of the present invention provides a multi-layer stacked packaging method for an edge AI computing and storage chip, including the following steps:

[0064] Step S1, obtaining the functional requirements of a preset edge AI computing and storage chip, and determining the chip architecture of the edge AI computing and storage chip and the corresponding sequential arrangement of the chip architecture based on the functional requirements.

[0065] Specifically, to implement the step of "obtaining the functional requirements of a preset edge AI computing and storage chip and determining the chip architecture of the edge AI computing and storage chip and the corresponding sequential arrangement of the chip architecture", it is first necessary to clarify the functional requirements of the target edge AI application scenario. For example, in the field of autonomous driving, the edge AI chip needs to process a large amount of data from sensors such as cameras and radars in real time and perform complex calculations such as object detection and path planning. Therefore, it requires high computing power, low latency, and low power consumption. In the field of smart home, the edge AI chip may focus more on functions such as voice recognition and image processing, with relatively lower requirements for computing power, but more sensitive to cost and power consumption. Therefore, for different application scenarios, specific performance index requirements need to be clarified, such as computing power (TOPS), power consumption (W), latency (ms), memory size, and so on. After clarifying the functional requirements, it is necessary to determine the architecture of the edge AI computing and storage chip according to these requirements. For example, if high computing power is required, a multi-core processor architecture or a dedicated AI accelerator architecture can be adopted; if low power consumption is required, a memory-computation-integrated architecture can be adopted to reduce data transfer; if a large amount of data needs to be processed, it is necessary to consider increasing the on-chip storage capacity. At the same time, it is also necessary to determine the arrangement order of each module in the chip architecture. This is closely related to subsequent stacking and packaging. A reasonable arrangement order can optimize the interconnection between chips, reduce the signal transmission distance, and lower power consumption and latency. For example, the computing unit and the storage unit can be placed close to each other to reduce the overhead of data transfer. In addition, the sequential arrangement of the chip architecture also needs to consider the requirements of heat dissipation design. For example, modules with higher heat generation can be placed close to the cooling channel for more effective heat dissipation. For example, in the autonomous driving scenario, the convolutional neural network accelerator for processing image data usually has a relatively high power consumption, so it can be placed close to the cooling channel, while the control unit with relatively lower power consumption can be placed away from the cooling channel. In the smart home scenario, the chip for voice recognition has a relatively low power consumption, so it can be placed in the center of the stack, while other modules are placed on the periphery. In this way, according to different application scenarios and functional requirements, the optimal chip architecture and stacking order can be designed to achieve the best balance between performance and power consumption.

[0066] Step S2: Through deep reactive ion etching technology, perform TSV etching on a preset wafer to obtain a wafer with a via array.

[0067] Specifically, the core of the step of "performing TSV etching on a preset wafer through deep reactive ion etching technology to obtain a wafer with a via array" lies in using deep reactive ion etching (DRIE) technology to fabricate a large number of TSVs (Through-Silicon Vias) on the wafer, ultimately forming a wafer with a via array. Deep reactive ion etching is an advanced dry etching technology that combines the advantages of chemical etching and physical etching. It can etch deep holes with high aspect ratios and regular shapes on silicon wafers, which is the key technology required for manufacturing TSVs. The preset wafer usually refers to a silicon wafer that has undergone some pretreatment steps, such as cleaning, oxidation, etc., to prepare for subsequent TSV etching. The DRIE process typically uses the Bosch process, which is a cyclic etching method that alternates between etching cycles and passivation cycles. In the etching cycle, reactive ions in the plasma etch the silicon to form vias; in the passivation cycle, a protective film is deposited on the etched surface to prevent the sidewalls from being etched. This alternating etching and passivation cycle can ensure the etching of TSVs with high aspect ratios and maintain the perpendicularity and smoothness of the sidewalls. In actual operation, first, a layer of photoresist needs to be coated on the wafer surface, and then the pre-designed TSV pattern is transferred to the photoresist using lithography technology. Next, using deep reactive ion etching technology, with the photoresist pattern as a mask, the exposed silicon is etched. After etching is completed, the remaining photoresist is removed, and the required TSV via array can be obtained. During the etching process, etching parameters such as etching time, gas flow rate, power, etc. need to be precisely controlled to ensure that the size and shape of the TSVs meet the design requirements. In addition, the etched TSVs need to be cleaned and inspected to remove residues and ensure the quality of the TSVs. The finally obtained wafer with a via array contains a large number of TSVs, which can be used for electrical connections between different levels inside the chip, such as connecting different chips, connecting the chip and the substrate, etc. Through TSV technology, three-dimensional integration of the chip can be achieved, improving the integration and performance of the chip. For example, in high-performance autonomous driving AI chips, high-speed interconnections need to be established between different functional modules inside the chip, and TSV technology can provide high-bandwidth and low-latency interconnection channels. A large number of TSVs are etched on the wafer to form a dense via array for connecting modules such as processor cores, memory, and high-speed interfaces. This makes data transmission inside the chip more efficient, thereby improving the overall performance of the chip. Similarly, in low-power application scenarios such as smart homes, TSV technology can also be used for interconnections inside the chip, thereby improving the integration of the chip and reducing power consumption. Therefore, TSV etching technology is crucial for the manufacturing of various edge AI compute-in-memory chips, providing a key connection channel for the three-dimensional integration of the chips.

[0068] Step S3: Cut the through-hole array wafer to obtain independent cut chips; and construct cooling channels for the independent cut chips to obtain cooling function chips.

[0069] Specifically, for the step of "cutting the through-hole array wafer to obtain independent cut chips; and constructing cooling channels for the independent cut chips to obtain cooling function chips", first, the wafer with completed TSV through-holes is cut into individual independent chip units. These independent cut chips already have a preset TSV through-hole array for subsequent chip interconnection. Then, cooling channels are constructed on these independent cut chips, which is the key point of this patented method and an important feature differentiating it from traditional packaging methods. The construction method of the cooling channels can adopt different methods according to specific requirements. For example, microchannel technology can be used to etch tiny fluid channels on the back or inside of the chip for the circulation of the coolant to carry away the heat generated by the chip. Or, micro heat sinks or heat pipes and other heat dissipation structures can be bonded to the chip surface to enhance the heat dissipation effect. After constructing the cooling channels, the independent cut chips are transformed into chips with cooling functions, providing a good heat dissipation foundation for subsequent multi-chip stacking. This way of constructing cooling channels after cutting can more effectively control the temperature of each chip and avoid the problem of uneven heat dissipation caused by the mutual influence between chips after stacking. For example, in the autonomous driving scenario, due to the need to process a large amount of image data, the AI computing chips have high power consumption and large heat generation. By constructing independent cooling channels on each chip, the temperature of each chip can be effectively controlled to ensure the stable operation of the chips under high load. In the smart home scenario, although the power consumption of the chips is relatively low, heat will still be generated during long-term operation. By pre-constructing cooling channels, the service life of the chips can be effectively extended and the reliability of the device can be improved. Therefore, this method of constructing cooling channels after cutting is crucial for improving the heat dissipation efficiency of multi-chip stacking packaging, can effectively solve the heat dissipation problem of stacked chips, and thus achieve high-performance and low-power edge AI computing and storage chips.

[0070] Step S4: Based on the chip architecture and the corresponding sequential arrangement, perform multi-chip integration and packaging on the cooling function chips to obtain edge AI computing and storage chips.

[0071] Specifically, the step of "performing multi-chip integration and packaging on the cooling function chips based on the chip architecture and the corresponding sequential arrangement of the chip architecture to obtain an edge AI computing and storage chip" stacks and integrates the independent chips with cooling functions prepared in the previous steps according to the pre-designed chip architecture and sequence, and finally packages them into a complete edge AI computing and storage chip. In this process, the stacking order of the chips is crucial, as it directly affects the interconnection efficiency and heat dissipation performance between the chips. According to the previously determined chip architecture, chips with different functions (such as computing units, storage units, control units, etc.) are stacked together in a specific order, and interconnected using the previously reserved TSV vias. For example, the computing unit and the storage unit can be placed close to each other to reduce data transmission distance and latency, while the high-power chips are placed close to the cooling channels to improve heat dissipation efficiency. After the chip stacking is completed, packaging is required to protect the chips and provide external interfaces. The packaging process includes selecting suitable packaging materials, designing the packaging structure, and performing packaging tests, etc. The packaging materials need to have good thermal conductivity and electrical insulation performance to ensure the stable operation of the chips. The packaging structure needs to consider factors such as the size of the chips, the number of pins, and the heat dissipation requirements. For example, in the autonomous driving scenario, to achieve higher computing power, multiple AI accelerator chips can be stacked together and tightly integrated with the storage chips to form a high-performance computing unit. At the same time, to meet the low-power requirements, advanced packaging technologies such as 2.5D or 3D packaging technologies can be used to reduce the interconnection distance between the chips and lower the power consumption. In the smart home scenario, since the computing power requirement is relatively low, a computing chip can be stacked with multiple storage chips to form a compact system-on-chip (SoC). In this way, according to different application scenarios and functional requirements, the chip architecture and stacking order can be flexibly selected, and appropriate packaging technologies can be adopted to finally obtain an edge AI computing and storage chip that meets specific requirements.

[0072] In a specific embodiment, the obtaining of the functional requirements of the preset edge AI computing and storage chip and determining the chip architecture of the edge AI computing and storage chip and the corresponding sequential arrangement based on the functional requirements includes:

[0073] Obtaining the target application scenario of the preset AI computing and storage chip, and determining the functional requirements of the AI computing and storage chip based on the target application scenario; wherein, the functional requirements include throughput, latency, and power consumption;

[0074] Performing a hardware architecture search based on the functional requirements through a preset neural network search algorithm to obtain a candidate architecture set;

[0075] Screening the candidate architecture set through a preset power consumption evaluation model to obtain a low-power architecture;

[0076] Based on the functional requirements and the target application scenario, perform on-chip storage architecture design for the low-power architecture to determine the chip architecture of the edge AI computing and storage chip and the corresponding sequential arrangement of the chip architecture; wherein, the on-chip storage architecture design includes determining the type, capacity, bit width, structure of the memory, and the data interaction method with the computing unit to meet the data access mode and bandwidth requirements of the target scenario application.

[0077] Specifically, implementing the step of "obtaining the functional requirements of a preset edge AI computing and storage chip and determining the chip architecture of the edge AI computing and storage chip and the corresponding sequential arrangement of the chip architecture" covers the complete process from determining the application scenario to the final chip architecture design. First, it is necessary to clarify the target application scenario of the preset AI computing and storage chip, which is the basis for all subsequent steps. Different application scenarios have huge differences in the functional requirements of the chip. For example, in the field of autonomous driving, the edge AI chip needs to process a large amount of data from multiple sensors such as cameras, radars, and lidar in real time, perform complex calculations such as object detection, path planning, and decision-making control. Therefore, it has extremely high requirements for the throughput and latency of the chip, and at the same time, it also needs to take into account low power consumption to extend the battery life. In the field of smart home, the main tasks of the edge AI chip may be speech recognition, image processing, simple control logic, etc. The requirements for throughput are relatively low, but more attention is paid to low cost and low power consumption to extend the battery life. Therefore, the first step is to clarify the functional requirements of the chip according to the specific application scenario, including throughput (such as the number of image frames or voice data that can be processed per second), latency (such as the time from data input to result output), and power consumption (such as the average power consumption or peak power consumption of the chip). Next, based on the determined functional requirements, a preset neural network search algorithm is used for hardware architecture search. The goal of this step is to find a set of candidate architectures that can meet the functional requirements from a huge design space. The neural network search algorithm can automatically explore different hardware architectures, such as different types and numbers of computing units, interconnection methods, memory hierarchies, etc., and evaluate the performance of each architecture. Commonly used neural network search algorithms include reinforcement learning, evolutionary algorithms, Bayesian optimization, etc. Through these algorithms, a set of candidate architectures with excellent performance can be efficiently searched. After obtaining the candidate architecture set, further screening is required to determine the final low-power architecture. This is because even if the functional requirements are met, there may be significant differences in power consumption among different architectures. To select the optimal architecture, a preset power consumption evaluation model is used to evaluate the candidate architecture set. The power consumption evaluation model can estimate the power consumption of each architecture according to the specific parameters of the architecture, such as the type and number of computing units, the access frequency and power consumption of the memory, the length and power consumption of the interconnection lines, etc. By comparing the power consumption of different architectures, the architecture with the lowest power consumption can be selected as the final low-power architecture. Finally, based on the functional requirements and the target application scenario, the on-chip storage architecture of the selected low-power architecture is designed. The design of the on-chip storage architecture is crucial for the overall performance and power consumption of the chip. In this step, it is necessary to determine the type, capacity, bit width, structure of the memory, and the data interaction method with the computing unit. The type of memory can be selected from SRAM, DRAM, HBM, etc. The capacity needs to be determined according to the data volume of the application scenario. The bit width affects the data transfer rate. The structure can be single-port, multi-port, or shared memory, etc.In addition, it is necessary to design efficient data interaction methods, such as DMA (Direct Memory Access) controllers, cache mechanisms, etc., to minimize the overhead of data transfer and improve data access efficiency. For example, in the autonomous driving scenario, to meet the requirements of high throughput and low latency, it may be necessary to adopt a large-capacity, high-bandwidth HBM memory and design a dedicated data path to achieve high-speed data transmission. In the smart home scenario, due to the relatively small amount of data, a small-capacity, low-power SRAM memory can be adopted, and a simple bus structure can be used for data interaction. Through the optimized design of the on-chip storage architecture, the performance of the chip can be maximized and the power consumption can be reduced. Through the above steps, the chip architecture of the edge AI computing and storage chip and the corresponding sequential arrangement of the chip architecture are finally determined, which provides important guidance for subsequent chip design, manufacturing, and packaging. This chip architecture design method driven by application scenarios and functional requirements can effectively improve the performance and efficiency of the chip, reduce power consumption and cost, and thus better meet the requirements of different application scenarios.

[0078] In a specific embodiment, the TSV etching of a preset wafer by the deep reactive ion etching technology to obtain a wafer with a via array includes:

[0079] Measuring the thickness of a preset wafer based on optical interferometry to obtain a wafer thickness map, and performing thickness uniformity analysis on the wafer thickness map to obtain a wafer with uniform thickness;

[0080] Performing TSV initial etching on the wafer with uniform thickness by the deep reactive ion etching technology to obtain an initially etched wafer, and performing sidewall passivation treatment on the initially etched wafer to obtain a passivated wafer;

[0081] Depositing a barrier layer on the passivated wafer by atomic layer deposition technology to obtain a wafer with a barrier layer, and performing photolithographic patterning on the wafer with a barrier layer to obtain a patterned wafer;

[0082] Performing deep hole etching on the patterned wafer based on the Bosch etching process to obtain a wafer with deep holes;

[0083] Depositing an insulating layer on the wafer with deep holes by chemical vapor deposition technology to obtain a wafer with an insulating layer, and performing chemical mechanical polishing on the wafer with an insulating layer to obtain a wafer with a via array.

[0084] Specifically, the step of "performing TSV etching on a preset wafer through deep reactive ion etching technology to obtain a wafer with a via array" describes how to etch TSV vias on the wafer using deep reactive ion etching technology to form a wafer with a via array. First, optical interferometry is used to measure the thickness of the preset wafer to obtain a wafer thickness map. Optical interferometry is a high-precision thickness measurement method that can accurately measure the thickness of different positions on the wafer. Then, thickness uniformity analysis is performed on the wafer thickness map to screen out wafers with uniform thickness, resulting in wafers with uniform thickness. Thickness uniformity is crucial for subsequent etching processes as it ensures the depth and uniformity of etching. Next, the wafer with uniform thickness is subjected to initial TSV etching using deep reactive ion etching technology to obtain an initially etched wafer. Deep reactive ion etching is a dry etching technology that can precisely control the depth and shape of etching. The purpose of the initial etching is to form the initial shape of the TSV vias on the wafer surface. Then, sidewall passivation treatment is performed on the initially etched wafer to obtain a passivated wafer. Sidewall passivation can prevent corrosion and collapse of the sidewalls during the etching process, improving the quality of the TSV vias. After that, a barrier layer is deposited on the passivated wafer through atomic layer deposition (ALD) technology to obtain a wafer with a barrier layer. ALD technology can precisely control the deposition thickness and uniformity to form a dense barrier layer. The barrier layer can prevent subsequent etching processes from damaging the wafer. Then, photolithographic patterning is performed on the wafer with the barrier layer to obtain a patterned wafer. Photolithographic patterning is the process of transferring a pre-designed TSV via pattern onto the wafer surface, providing a template for subsequent deep hole etching. Next, deep hole etching is performed on the patterned wafer based on the Bosch etching process to obtain a wafer with deep holes. The Bosch etching process is a cyclic etching process that can achieve high aspect ratio etching, thus forming deep and narrow TSV vias. Deep hole etching is a key step in TSV manufacturing as it determines the final size and shape of the TSV vias. Finally, an insulating layer is deposited on the wafer with deep holes through chemical vapor deposition (CVD) technology to obtain a wafer with an insulating layer. CVD technology can deposit an insulating layer on the inner wall of the TSV vias for isolating the TSV vias from the surrounding circuits. Then, chemical mechanical polishing (CMP) is performed on the wafer with the insulating layer to obtain the final wafer with a via array. CMP can polish the wafer surface flat, remove the excess insulating layer, and expose the openings of the TSV vias for subsequent metal filling and interconnection. For example, in high-performance autonomous driving AI chips, high-density TSV vias are required to achieve high-speed data transmission between different modules inside the chip. Through the above process flow, TSV via arrays with high aspect ratio, high density, and high reliability can be manufactured to meet the requirements of high-performance chips. Even in low-power application scenarios such as smart homes, high-quality TSV via arrays can also improve the chip integration and performance and reduce power consumption.Therefore, the TSV etching process is crucial for the manufacturing of various edge AI computing-in-memory chips.

[0085] In a specific embodiment, the through-hole array wafer is cut to obtain independent cut chips; and a cooling channel is constructed for the independent cut chips to obtain a cooling function chip, including:

[0086] The internal structure of the through-hole array wafer is scanned by a preset three-dimensional X-ray tomography technique to obtain a three-dimensional structure distribution map, and a cutting path is planned for the through-hole array wafer based on the three-dimensional structure distribution map to obtain a cutting path;

[0087] The through-hole array wafer is cut layer by layer based on the cutting path by a multi-wavelength laser cutting technique to obtain a preliminary cut chip, and the edge of the preliminary cut chip is chamfered to obtain an independent cut chip;

[0088] The surface of the independent cut chip is planarized at the atomic level by an atomic layer etching technique to obtain a planarized chip;

[0089] A cooling channel pattern is constructed for the planarized chip by a microfluidic channel design technique to obtain a microchannel structure, and the microchannel structure is anisotropically etched to obtain a three-dimensional flow channel network;

[0090] The surface of the three-dimensional flow channel network is modified to have superhydrophilicity to obtain a surface wetting characteristic chip, and a nanoscale functional coating is deposited on the surface wetting characteristic chip to obtain a hydrodynamic chip;

[0091] The heat transfer performance of the hydrodynamic chip is evaluated by a microscale heat flow analysis technique to obtain thermal resistance and temperature rise distribution data. If the thermal resistance and temperature rise distribution data are within a preset safe operating range, a cooling function chip is obtained;

[0092] If not, the thermal resistance and temperature rise distribution data are deeply analyzed by a preset genetic algorithm to obtain optimized cooling channel structure parameters, and the hydrodynamic chip is optimized based on the optimized cooling channel structure parameters to obtain a cooling function chip.

[0093] Specifically, to implement the step of "cutting the wafer with a through-hole array to obtain independent cut chips; and constructing cooling channels for the independent cut chips to obtain chips with cooling functions", it is first necessary to precisely cut the wafer with a through-hole array and construct an efficient cooling channel on this basis. First, use the preset three-dimensional X-ray tomography technology to scan the internal structure of the wafer with a through-hole array to obtain a three-dimensional structure distribution map of the TSV through-holes inside the wafer. This step is crucial because the cutting path must avoid the TSV through-holes to prevent damage to these key interconnect structures. Based on the obtained three-dimensional structure distribution map, perform cutting path planning to determine the optimal cutting path to ensure that the cutting process does not affect the function and performance of the chip. Then, use the multi-wavelength laser cutting technology to perform layer-by-layer cutting on the wafer with a through-hole array according to the planned cutting path. The multi-wavelength laser cutting technology can achieve high-precision and high-efficiency cutting with less damage to the wafer. After cutting, the preliminary cut chips are obtained, and the edges of these chips need to be chamfered to remove burrs and sharp edges to obtain independent cut chips, preparing for the subsequent construction of cooling channels. Next, construct cooling channels for the independent cut chips. First, use atomic layer etching technology to perform surface atomic-level planarization on the independent cut chips to obtain planarized chips. The atomic layer etching technology can precisely remove the atomic layer on the chip surface to achieve extremely high surface flatness, which is very important for the subsequent construction of microfluidic channels. Then, use microfluidic channel design technology to construct a cooling channel pattern on the surface of the planarized chip to form a microchannel structure. The microfluidic channel design technology can precisely control the size and shape of the microchannels to achieve the best cooling effect. Then, perform anisotropic etching on the microchannel structure to form a three-dimensional flow channel network. Anisotropic etching can precisely control the etching direction and depth, thereby forming a flow channel network with a complex three-dimensional structure to improve the cooling efficiency. To further improve the cooling performance, perform superhydrophilic surface modification on the constructed three-dimensional flow channel network to make it have excellent surface wetting characteristics. The superhydrophilic surface can promote the flow of the coolant in the flow channel, reduce the flow resistance, and improve the heat dissipation efficiency. Then, deposit a nano-scale functional coating on the chip with surface wetting characteristics to further optimize the hydrodynamic characteristics to obtain a hydrodynamic chip. Finally, use microscale heat flow analysis technology to evaluate the heat transfer performance of the hydrodynamic chip to obtain data on the thermal resistance and temperature rise distribution of the chip. Compare these data with the preset safe operating range. If the thermal resistance and temperature rise distribution data are within the safe operating range, it is considered that the construction of the chip with cooling functions is completed. If not, the cooling channel structure needs to be optimized. At this time, use the preset genetic algorithm to deeply analyze the thermal resistance and temperature rise distribution data, find the key parameters affecting the cooling performance, and obtain the optimized cooling channel structure parameters.Based on these optimized parameters, the hydrodynamic chip is optimized, the cooling channels are reconstructed, and the heat transfer performance is evaluated again until the data of thermal resistance and temperature rise distribution reach the preset safe operating range, and finally a cooling function chip is obtained. For example, in high-performance autonomous driving AI chips, due to the high power consumption of the chips, the heat dissipation problem is particularly prominent. The microfluidic cooling system constructed through the above process flow can effectively reduce the chip temperature and ensure the stability and reliability of the chip. In low-power application scenarios such as smart homes, although the heat dissipation pressure is relatively small, an efficient cooling system can still extend the chip life and improve the reliability of the system. Therefore, constructing an efficient cooling system is crucial for various edge AI computing and storage chips.

[0094] In a specific embodiment, the internal structure of the via array wafer is scanned by a preset three-dimensional X-ray tomography technique to obtain a three-dimensional structure distribution map, and a cutting path is planned for the via array wafer based on the three-dimensional structure distribution map, including:

[0095] Performing phase-contrast X-ray tomography scanning on the via array wafer to obtain an original diffraction image, and performing phase reconstruction on the original diffraction image to obtain a multi-layer diffraction pattern;

[0096] Digitally reconstructing the multi-layer diffraction pattern by a spatial frequency domain filtering technique to obtain density distribution data, and performing three-dimensional voxel reconstruction on the density distribution data to obtain a voxel distribution map;

[0097] Performing three-dimensional reconstruction on the voxel distribution map by an inverse projection algorithm to obtain a three-dimensional structure distribution map, and planning a cutting area for the three-dimensional structure distribution map to obtain cutting area parameters; wherein, the cutting area parameters include cutting critical points, stress-sensitive areas, and material demarcation lines;

[0098] Using topology optimization technology, planning a cutting path for the via array wafer based on the cutting area parameters to obtain a cutting path.

[0099] Specifically, for the step of "scanning the internal structure of the through-hole array wafer by a preset three-dimensional X-ray tomography technique to obtain a three-dimensional structure distribution map, and planning a cutting path for the through-hole array wafer based on the three-dimensional structure distribution map to obtain a cutting path", the core lies in using advanced imaging and analysis techniques to accurately determine the internal structure of the wafer and plan a safe cutting path, thereby avoiding damaging the critical TSV through-hole structure during the cutting process. First, perform phase-contrast X-ray tomography scanning on the through-hole array wafer. Phase-contrast X-ray tomography is a non-destructive imaging technique that uses the phase difference generated when X-rays pass through a sample to reconstruct the internal structure of the sample, and is particularly suitable for imaging materials with small density differences, such as silicon and TSV filling materials. After scanning, the original diffraction images are obtained. These images contain information about the internal structure of the wafer, but further processing is required to obtain a clear three-dimensional structure. Next, perform phase reconstruction on the original diffraction images to obtain a multi-layer diffraction pattern. Phase reconstruction is the process of converting the phase information in the original diffraction images into data that can be used to reconstruct the image. It can improve the contrast and resolution of the image, making the structure of the TSV through-holes more clearly visible. Then, perform digital reconstruction on the multi-layer diffraction pattern through spatial frequency domain filtering technology to obtain density distribution data. Spatial frequency domain filtering technology can remove noise and artifacts in the image and improve the quality of the image. Based on the density distribution data, perform three-dimensional voxel reconstruction to obtain a voxel distribution map. A voxel is the smallest unit in three-dimensional space, and the voxel distribution map can visually display the density distribution of different materials inside the wafer. Next, perform three-dimensional reconstruction on the voxel distribution map through an inverse projection algorithm to obtain a three-dimensional structure distribution map. The inverse projection algorithm is a commonly used three-dimensional reconstruction algorithm that can convert two-dimensional projection data into a three-dimensional image. The obtained three-dimensional structure distribution map clearly shows information such as the position, size, and orientation of the TSV through-holes inside the wafer. Based on this three-dimensional structure distribution map, perform cutting area planning to determine the cutting boundaries and key areas, and obtain cutting area parameters. These parameters include cutting critical points, which are the key positions that the cutting path must avoid; stress-sensitive areas, which are the areas where stress is easily generated during the cutting process; and material boundaries, which are the boundaries between different materials. Finally, use topology optimization technology to plan a cutting path for the through-hole array wafer based on the cutting area parameters to obtain the final cutting path. Topology optimization technology is a mathematical optimization method that can find the optimal cutting path, which can not only avoid critical structures, minimize the stress generated by cutting, but also ensure cutting efficiency. For example, in high-performance autonomous driving AI chips, the density of TSV through-holes is very high, and the planning of the cutting path must be very precise to avoid damaging these through-holes. Through the above process flow, the position and orientation of the TSV through-holes can be accurately determined, and a safe cutting path can be planned to ensure the yield and performance of the chip.Even in low-power application scenarios such as smart homes, accurate cutting path planning can improve chip production efficiency and reduce manufacturing costs. Therefore, accurate cutting path planning is crucial for the manufacture of various edge AI computing and storage chips.

[0100] In a specific embodiment, the cooling function chip is integrated and packaged into multiple chips based on the chip architecture and the sequence arrangement corresponding to the chip architecture to obtain an edge AI computing and storage chip, including:

[0101] Arranging the cooling function chips based on the chip architecture and the sequence arrangement corresponding to the chip architecture to obtain arranged chips;

[0102] The arranged chips are interconnected by using micro-bump bonding technology to obtain bump interconnected chips, and the bump interconnected chips are bottom-filled to obtain bottom-filled chips;

[0103] Stacking the bottom-filled chips by high-density three-dimensional stacking technology to obtain a three-dimensional chip stack;

[0104] Packaging the three-dimensional chip stack to obtain a WLP packaged chip, and wire bonding the WLP packaged chip to obtain a wire-bonded chip;

[0105] Performing RDL construction on the wire bonding chip by using redistribution layer technology to obtain an RDL chip, and performing electrical testing on the RDL chip to obtain electrical test results;

[0106] If the electrical test results show that there are RDL chips that do not meet the preset standards, the RDL chips that do not meet the preset standards are removed to obtain qualified RDL chips;

[0107] The qualified RDL chip is packaged at the system level based on the system-level packaging technology to obtain an edge AI computing and storage chip.

[0108] Specifically, the step of "performing multi-chip integration and packaging on the cooling function chips based on the chip architecture and the corresponding sequential arrangement to obtain the edge AI computing and storage chip" describes how to integrate and package multiple cooling function chips into a complete edge AI computing and storage chip. First, according to the pre-designed chip architecture and the corresponding sequential arrangement, the cooling function chips are arranged to obtain the arranged chips. The chip architecture defines the connection relationships and functional divisions between different chips, while the sequential arrangement determines the specific positions of the chips in the package. This step lays the foundation for subsequent interconnection and stacking. Next, the arranged chips are interconnected using the microbump bonding technology to form the bump-interconnected chips. The microbump bonding technology is a high-density interconnection technology that realizes electrical connection by forming tiny bumps between the chips, enabling high-bandwidth and low-latency communication between the chips. Then, the bump-interconnected chips are underfilled to obtain the underfilled chips. The purpose of underfilling is to protect the microbumps, enhance the mechanical strength of the chips, and improve the reliability of the chips. After that, the underfilled chips are stacked using the high-density three-dimensional stacking technology to obtain the three-dimensional chip stack. The three-dimensional stacking technology can vertically stack multiple chips together, thus significantly improving the integration and performance of the chips. The stacked chips are interconnected through microbumps to form a tightly connected whole. Then, the three-dimensional chip stack is packaged to obtain the WLP (wafer-level packaging) packaged chips. The WLP packaging technology can directly perform packaging on the wafer, which can reduce the size of the chips and lower the packaging cost. Next, wire bonding is performed on the WLP packaged chips to obtain the wire-bonded chips. Wire bonding is the process of connecting the pads on the chip to the external circuit, which provides an interface for the chip to communicate with the outside world. To further improve the interconnection density and performance of the chips, RDL (redistribution layer) construction is performed on the wire-bonded chips using the RDL technology to obtain the RDL chips. The RDL technology can build high-density metal lines on the chip surface for connecting different functional modules inside the chip and providing more I / O interfaces for the chip. After constructing the RDL layer, electrical testing needs to be performed on the RDL chips to verify the quality and reliability of the RDL and obtain the electrical test results. If the electrical test results show that there are RDL chips that do not meet the preset standards, these unqualified chips are removed, and only the qualified RDL chips are retained. Finally, system-level packaging is performed on the qualified RDL chips based on the system-in-package (SiP) technology to obtain the final edge AI computing and storage chip. The SiP technology can integrate multiple chips with different functions in one package to form a complete system. Through SiP packaging, the size of the chips can be further reduced, and the integration and performance of the chips can be improved. For example, in a high-performance autonomous driving AI chip, multiple processor cores, memory chips, and interface chips need to be integrated together to achieve complex computing and data processing functions.Through the above multi-chip integration and packaging processes, chips with different functions can be integrated into a compact package to form a high-performance and low-power edge AI computing and storage chip. In low-power application scenarios such as smart homes, multi-chip integration technology can integrate processors, memories, sensors, etc. together to form a highly integrated intelligent terminal. Therefore, multi-chip integration and packaging technology is crucial for the implementation of various edge AI computing and storage chips.

[0109] In a specific embodiment, the method of using the micro-bump bonding technology to interconnect the arranged chips to obtain bump-interconnected chips and perform underfill on the bump-interconnected chips to obtain underfilled chips includes:

[0110] Sputter the UBM metal layer on the arranged chips to obtain UBM-structured chips;

[0111] Grow micro-bumps on the UBM-structured chips through ultrasonic-assisted electroplating technology to obtain micro-bump chips;

[0112] Align the micro-bump chips to obtain aligned micro-bump chips, and perform low-temperature pre-bonding on the aligned micro-bump chips to obtain preliminarily bonded chips;

[0113] Perform thermocompression bonding on the preliminarily bonded chips based on thermocompression bonding technology to obtain bump-interconnected chips;

[0114] Perform underfill on the bump-interconnected chips using capillary action to obtain filled bump-interconnected chips, and perform curing treatment on the filled bump-interconnected chips to obtain underfilled chips.

[0115] Specifically, the step of "using the microbump bonding technology to interconnect the arranged chips to obtain a bump-interconnected chip and perform underfill on the bump-interconnected chip to obtain an underfilled chip" details the process of using the microbump bonding technology to achieve chip interconnection and subsequent underfill. First, sputter the UBM (UnderBump Metallization) metal layer on the arranged chips to form a UBM-structured chip. The UBM layer is usually composed of multiple layers of metals, such as titanium, nickel, and gold. It serves as the basis for microbumps, providing good electrical conductivity and adhesion. Then, grow microbumps on the UBM-structured chip through ultrasonic-assisted electroplating technology to obtain a microbump chip. The introduction of ultrasonic waves can improve the uniformity and efficiency of electroplating, thus obtaining high-quality microbumps. These microbumps will serve as the bridges for electrical connection between chips. Next, perform chip alignment to ensure that the microbumps on the chips to be bonded are precisely aligned with the corresponding pads on the target chip, obtaining an aligned microbump chip. The accuracy of chip alignment directly affects the quality and reliability of bonding. After alignment, perform low-temperature pre-bonding to obtain a preliminarily bonded chip. Low-temperature pre-bonding can generate a preliminary connection between chips, providing a basis for subsequent thermocompression bonding. Then, perform thermocompression bonding on the preliminarily bonded chip based on the thermocompression bonding technology to obtain a bump-interconnected chip. Thermocompression bonding is carried out under high temperature and high pressure, which can firmly connect the microbumps with the pads on the target chip to form a reliable electrical connection. So far, the interconnection between chips is completed. To further enhance the mechanical strength and reliability of the chips, it is necessary to perform underfill on the bump-interconnected chip. Use capillary action to perform underfill on the bump-interconnected chip to obtain a filled bump-interconnected chip. Capillary action can make the filling material automatically fill the gaps between chips, completely wrapping the microbumps. Then, perform curing treatment on the filled bump-interconnected chip to cure the filling material and obtain the final underfilled chip. The cured filling material can effectively protect the microbumps, prevent them from being affected by the external environment, and improve the mechanical strength and shock resistance of the chips. For example, in high-performance autonomous driving AI chips, multiple processor cores, memory chips, and interface chips need to be interconnected through the microbump bonding technology to achieve high-speed data transmission and collaborative work. Underfill can protect these microbumps and ensure that the chips can still work stably under harsh environments such as high temperature and high vibration. Even in low-power application scenarios such as smart homes, the microbump bonding and underfill technologies can improve the chip integration and reliability, and extend the service life of the product. Therefore, the microbump bonding and underfill technologies are crucial for the manufacturing of various edge AI compute-in-memory chips.

[0116] The above described the multi-layer stacked packaging method of the edge AI computing and storage chip in the embodiments of the present invention. Next, the multi-layer stacked packaging device of the edge AI computing and storage chip in the embodiments of the present invention will be described. Please refer to Figure 2 One embodiment of the multi-layer stacked packaging device of the edge AI computing and storage chip in the embodiments of the present invention includes:

[0117] An acquisition module 21, configured to acquire the functional requirements of a preset edge AI computing and storage chip, and determine the chip architecture of the edge AI computing and storage chip and the corresponding sequential arrangement according to the functional requirements;

[0118] A via module 22, configured to perform TSV etching on a preset wafer through deep reactive ion etching technology to obtain a wafer with a via array;

[0119] A cutting module 23, configured to cut the wafer with the via array to obtain independent cut chips; and construct cooling channels for the independent cut chips to obtain cooling function chips;

[0120] A packaging module 24, configured to perform multi-chip integration and packaging on the cooling function chips based on the chip architecture and the corresponding sequential arrangement to obtain an edge AI computing and storage chip.

[0121] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to that described in the above method embodiment, and details will not be repeated here.

[0122] Refer to Figure 3 In addition, an embodiment of the present invention further provides a computer device, and its internal structure can be as Figure 3 shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.

[0123] Those skilled in the art can understand that Figure 3 the structure shown in

[0124] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0125] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium provided by the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0126] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, apparatus, article or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, apparatus, article or method including the element.

[0127] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token.

Claims

1. A multi-layer stacking packaging method for edge AI computing and storage chips, characterized in that: The following steps are involved: Obtaining the functional requirements of a preset edge AI computing and storage chip, and determining the chip architecture of the edge AI computing and storage chip and the order of arrangement corresponding to the chip architecture based on the functional requirements; Through deep reactive ion etching technology, TSV etching is performed on the preset wafer to obtain a through-hole array wafer; Cutting the through hole array wafer to obtain independently cut chips; and constructing cooling channels for the independently cut chips to obtain cooling function chips; Based on the chip architecture and the sequence arrangement corresponding to the chip architecture, the cooling function chip is integrated and packaged into multiple chips to obtain an edge AI computing and storage chip; The method of performing TSV etching on a preset wafer by deep reactive ion etching technology to obtain a through hole array wafer includes: Based on the optical interferometry method, a thickness measurement is performed on a preset wafer to obtain a wafer thickness map, and a thickness uniformity analysis is performed on the wafer thickness map to obtain a wafer with uniform thickness; Performing TSV initial etching on the wafer with uniform thickness by using deep reactive ion etching technology to obtain an initial etched wafer, and performing sidewall passivation treatment on the initial etched wafer to obtain a passivated wafer; Depositing a barrier layer on the passivation wafer by atomic layer deposition technology to obtain a barrier layer wafer, and performing photolithography patterning on the barrier layer wafer to obtain a patterned wafer; Performing deep hole etching on the patterned wafer based on a Bosch etching process to obtain a deep hole wafer; Depositing an insulating layer on the deep hole wafer by chemical vapor deposition technology to obtain an insulating layer wafer, and performing chemical mechanical polishing on the insulating layer wafer to obtain a through hole array wafer; The through hole array wafer is cut to obtain independent cut chips; and cooling channels are constructed for the independent cut chips to obtain cooling function chips, including: Scanning the internal structure of the through-hole array wafer by a preset three-dimensional X-ray tomography technology to obtain a three-dimensional structure distribution map, and planning a cutting path for the through-hole array wafer based on the three-dimensional structure distribution map to obtain a cutting path; Using a multi-wavelength laser cutting technology, the through hole array wafer is cut in layers based on the cutting path to obtain preliminary cut chips, and the preliminary cut chips are chamfered to obtain independent cut chips; The independently cut chip is subjected to atomic-level flattening of the surface by atomic layer etching technology to obtain a flattened chip; Using microfluidic channel design technology to construct a cooling channel pattern on the flattened chip to obtain a microchannel structure, and anisotropically etching the microchannel structure to obtain a three-dimensional channel network; Performing super-hydrophilic surface modification on the three-dimensional flow channel network to obtain a surface wetting characteristic chip, and performing nano-scale functional coating deposition on the surface wetting characteristic chip to obtain a fluid dynamics chip; Using micro-scale heat flow analysis technology, the heat transfer performance of the fluid dynamics chip is evaluated to obtain thermal resistance and temperature rise distribution data. If the thermal resistance and temperature rise distribution data are within a preset safe working range, a cooling function chip is obtained; If not, the thermal resistance and temperature rise distribution data are deeply analyzed by a preset genetic algorithm to obtain optimized cooling channel structural parameters, and the fluid dynamics chip is optimized based on the optimized cooling channel structural parameters to obtain a cooling function chip; The method comprises: scanning the internal structure of the through hole array wafer by a preset three-dimensional X-ray tomography technology to obtain a three-dimensional structure distribution map, and planning a cutting path for the through hole array wafer based on the three-dimensional structure distribution map to obtain a cutting path, including: Performing phase contrast X-ray tomography scanning on the through hole array wafer to obtain an original diffraction image, and performing phase reconstruction on the original diffraction image to obtain a multi-layer diffraction spectrum; Digitally reconstructing the multi-layer diffraction spectrum by spatial frequency domain filtering technology to obtain density distribution data, and performing three-dimensional voxel reconstruction on the density distribution data to obtain a voxel distribution map; The voxel distribution map is three-dimensionally reconstructed by a reverse projection algorithm to obtain a three-dimensional structure distribution map, and a cutting area is planned for the three-dimensional structure distribution map to obtain cutting area parameters; wherein the cutting area parameters include a cutting critical point, a stress sensitive area and a material boundary line; The topology optimization technology is used to plan the cutting path of the through hole array wafer based on the cutting area parameters to obtain a cutting path.

2. The multi-layer stacking packaging method of the edge AI computing and storage chip according to claim 1 is characterized in that: The obtaining of the functional requirements of the preset edge AI computing and storage chip, and determining the chip architecture of the edge AI computing and storage chip and the sequence arrangement corresponding to the chip architecture based on the functional requirements, includes: Obtain a preset target application scenario of the AI ​​computing and storage chip, and determine the functional requirements of the AI ​​computing and storage chip based on the target application scenario; wherein the functional requirements include throughput, latency, and power consumption; Performing a hardware architecture search based on the functional requirements by using a preset neural network search algorithm to obtain a candidate architecture set; The candidate architecture set is screened by a preset power consumption evaluation model to obtain a low power consumption architecture; Based on the functional requirements and the target application scenarios, an on-chip storage architecture is designed for the low-power architecture to determine the chip architecture of the edge AI computing and storage chip and the corresponding sequence arrangement of the chip architecture; wherein the on-chip storage architecture design includes determining the type, capacity, bit width, structure of the memory and the data interaction method with the computing unit to meet the data access mode and bandwidth requirements of the target scenario application.

3. The multi-layer stacking packaging method of the edge AI computing and storage chip according to claim 1 is characterized in that: The method of integrating and packaging the cooling function chip into multiple chips based on the chip architecture and the sequence arrangement corresponding to the chip architecture to obtain an edge AI computing and storage chip includes: Arranging the cooling function chips based on the chip architecture and the sequence arrangement corresponding to the chip architecture to obtain arranged chips; The arranged chips are interconnected by using micro-bump bonding technology to obtain bump interconnected chips, and the bump interconnected chips are bottom-filled to obtain bottom-filled chips; Stacking the bottom-filled chips by high-density three-dimensional stacking technology to obtain a three-dimensional chip stack; Packaging the three-dimensional chip stack to obtain a WLP packaged chip, and wire bonding the WLP packaged chip to obtain a wire-bonded chip; Performing RDL construction on the wire bonding chip by using redistribution layer technology to obtain an RDL chip, and performing electrical testing on the RDL chip to obtain electrical test results; If the electrical test results show that there are RDL chips that do not meet the preset standards, the RDL chips that do not meet the preset standards are removed to obtain qualified RDL chips; The qualified RDL chip is packaged at the system level based on the system-level packaging technology to obtain an edge AI computing and storage chip.

4. The multi-layer stacking packaging method of the edge AI computing and storage chip according to claim 3 is characterized in that: The method comprises: interconnecting the arranged chips by using micro-bump bonding technology to obtain a bump interconnection chip, and bottom-filling the bump interconnection chip to obtain a bottom-filled chip, comprising: Sputtering a UBM metal layer on the arranged chip to obtain a UBM structure chip; Growing micro-bumps on the UBM structure chip by ultrasonic assisted electroplating technology to obtain a micro-bump chip; Performing chip alignment on the micro-bump chip to obtain an aligned micro-bump chip, and performing low-temperature pre-bonding on the aligned micro-bump chip to obtain a preliminary bonded chip; Performing thermal compression bonding on the preliminary bonded chip based on thermal compression bonding technology to obtain a bump interconnection chip; The bump interconnect chip is bottom-filled by utilizing capillary action to obtain a filled bump interconnect chip, and the filled bump interconnect chip is solidified to obtain a bottom-filled chip.

5. A multi-layer stacking packaging device for edge AI computing and storage chips, characterized in that: include: An acquisition module, used to acquire the functional requirements of a preset edge AI computing and storage chip, and determine the chip architecture of the edge AI computing and storage chip and the order of arrangement corresponding to the chip architecture based on the functional requirements; A through-hole module is used to perform TSV through-holes on a preset wafer by deep reactive ion etching technology to obtain a through-hole array wafer; A cutting module is used to cut the through hole array wafer to obtain independently cut chips; and to construct cooling channels for the independently cut chips to obtain cooling function chips; A packaging module, used for integrating and packaging the cooling function chip into multiple chips based on the chip architecture and the sequence arrangement corresponding to the chip architecture to obtain an edge AI computing and storage chip; The method of performing TSV etching on a preset wafer by deep reactive ion etching technology to obtain a through hole array wafer includes: Based on the optical interferometry method, a thickness measurement is performed on a preset wafer to obtain a wafer thickness map, and a thickness uniformity analysis is performed on the wafer thickness map to obtain a wafer with uniform thickness; Performing TSV initial etching on the wafer with uniform thickness by using deep reactive ion etching technology to obtain an initial etched wafer, and performing sidewall passivation treatment on the initial etched wafer to obtain a passivated wafer; Depositing a barrier layer on the passivation wafer by atomic layer deposition technology to obtain a barrier layer wafer, and performing photolithography patterning on the barrier layer wafer to obtain a patterned wafer; Performing deep hole etching on the patterned wafer based on a Bosch etching process to obtain a deep hole wafer; Depositing an insulating layer on the deep hole wafer by chemical vapor deposition technology to obtain an insulating layer wafer, and performing chemical mechanical polishing on the insulating layer wafer to obtain a through hole array wafer; The through hole array wafer is cut to obtain independent cut chips; and cooling channels are constructed for the independent cut chips to obtain cooling function chips, including: Scanning the internal structure of the through-hole array wafer by a preset three-dimensional X-ray tomography technology to obtain a three-dimensional structure distribution map, and planning a cutting path for the through-hole array wafer based on the three-dimensional structure distribution map to obtain a cutting path; Using a multi-wavelength laser cutting technology, the through hole array wafer is cut in layers based on the cutting path to obtain preliminary cut chips, and the preliminary cut chips are chamfered to obtain independent cut chips; The independently cut chip is subjected to atomic-level flattening of the surface by atomic layer etching technology to obtain a flattened chip; Using microfluidic channel design technology to construct a cooling channel pattern on the flattened chip to obtain a microchannel structure, and anisotropically etching the microchannel structure to obtain a three-dimensional channel network; Performing super-hydrophilic surface modification on the three-dimensional flow channel network to obtain a surface wetting characteristic chip, and performing nano-scale functional coating deposition on the surface wetting characteristic chip to obtain a fluid dynamics chip; Using micro-scale heat flow analysis technology, the heat transfer performance of the fluid dynamics chip is evaluated to obtain thermal resistance and temperature rise distribution data. If the thermal resistance and temperature rise distribution data are within a preset safe working range, a cooling function chip is obtained; If not, the thermal resistance and temperature rise distribution data are deeply analyzed by a preset genetic algorithm to obtain optimized cooling channel structural parameters, and the fluid dynamics chip is optimized based on the optimized cooling channel structural parameters to obtain a cooling function chip; The method comprises: scanning the internal structure of the through hole array wafer by a preset three-dimensional X-ray tomography technology to obtain a three-dimensional structure distribution map, and planning a cutting path for the through hole array wafer based on the three-dimensional structure distribution map to obtain a cutting path, including: Performing phase contrast X-ray tomography scanning on the through hole array wafer to obtain an original diffraction image, and performing phase reconstruction on the original diffraction image to obtain a multi-layer diffraction spectrum; Digitally reconstructing the multi-layer diffraction spectrum by spatial frequency domain filtering technology to obtain density distribution data, and performing three-dimensional voxel reconstruction on the density distribution data to obtain a voxel distribution map; The voxel distribution map is three-dimensionally reconstructed by a reverse projection algorithm to obtain a three-dimensional structure distribution map, and a cutting area is planned for the three-dimensional structure distribution map to obtain cutting area parameters; wherein the cutting area parameters include a cutting critical point, a stress sensitive area and a material boundary line; The topology optimization technology is used to plan the cutting path of the through hole array wafer based on the cutting area parameters to obtain a cutting path.

6. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Three-dimensional stacked packaging structure containing micro-channel heat dissipation structure and packaging method of three-dimensional stacked packaging structure

    CN114551385A