Chip architecture and related method
By adopting a stacked packaging method in the chip architecture, and using the dedicated processing units of the first chip and the second chip to perform flexible allocation of data processing tasks, the shortcomings of the existing chip architecture in high computing requirements and complex application scenarios are solved, and efficient performance improvement and power consumption optimization are achieved.
Patent Information
- Application Number
- CN202510080806.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2020-02-28
- Publication Date
- 2025-06-10
AI Technical Summary
The existing chip architecture has shortcomings in meeting users' high computing needs and complex application scenarios, and the traditional SOC integrated chip architecture has problems such as insufficient computing power, excessive power consumption and insufficient storage bandwidth due to the limited chip area.
The chip architecture adopts a stacked package, including a first chip and a second chip, the first chip includes a general purpose processor, a bus and at least one dedicated processing unit, and the second chip includes a dedicated processing unit with partially the same computing function, and realizes flexible allocation and processing of data processing tasks through interconnection between chips.
Without increasing the product volume, improve product performance, meet users' high power consumption and computing needs, enhance the chip's computing power, solve the problem of insufficient computing power, and optimize power consumption performance.
Smart Images

Figure CN120127083A_ABST
Abstract
Description
[0001] This application is a divisional application. The application number of the original application is 202080001141.9, the original application date is February 28, 2020, and the entire content of the original application is incorporated herein by reference. Technical Field
[0002] This application relates to the field of chip technology, and in particular, to a chip architecture and related methods. Background Art
[0003] The system-on-chip (SOC) of intelligent terminals has been continuously evolving and developing along with the progress of semiconductor technology, evolving from 28nm to 7nm, and even 5nm. The technology enables machine vision, such as camera algorithms, and neural network algorithms to be integrated into the intelligent terminal SOC chip, and the power consumption of the intelligent terminal SOC chip does not exceed the current battery power supply capacity and heat dissipation capacity.
[0004] However, with the slowdown of Moore's Law, the rhythm of technology evolution has lengthened, and the cost invested in each generation of technology has increased step by step, bringing great challenges to product competitiveness. Therefore, the traditional SOC integrated chip architecture is gradually difficult to meet the current requirements, and the technical bottlenecks faced by its traditional SOC integrated chip architecture are mainly problems such as insufficient computing power, excessive power consumption, and insufficient memory bandwidth of the chip due to limited chip area. The existing chip stacking technology uses through-silicon vias for interconnection to achieve a stacked architecture between chips, which can be used to separate and decouple devices such as memory, analog, or input / output (IO) from the main chip, that is, devices such as memory, analog, or input / output (IO) become independent chips, achieving the purpose of using different process technologies for the two chips and the independence of technology evolution. Moreover, the stacked chip architecture will not significantly increase the volume of the product. However, with the rapid development of current Internet media and neural network algorithms, and the increasing demand for audio and video on intelligent terminals by users, the requirements for the computing power, power consumption, and memory bandwidth of the SOC integrated chip on intelligent terminals have also been greatly improved. In the existing stacking solutions, the chip only splits the analog and input / output devices into independent chips by expanding the stacked memory, and then the resulting stacked chip architecture cannot meet the increasingly high computing requirements and complex and changing application scenarios of users, and the power consumption performance of the split-and-then-stacked chip architecture is not optimized enough.
[0005] Therefore, how to provide an efficient chip structure that can meet the increasingly high power consumption and computing requirements of users without increasing the product volume, while improving the product performance and meeting the flexibility of task processing, is an urgent problem to be solved. Summary of the Invention
[0006] Embodiments of the present application provide a chip architecture and related methods to improve product performance without increasing the product volume and meet the flexibility of task processing.
[0007] In a first aspect, embodiments of the present application provide a data processing device, which may include a first chip and a second chip stacked and packaged; the first chip includes a general-purpose processor, a bus, and at least one first dedicated processing unit (DPU), the general-purpose processor and the at least one first dedicated processing unit are connected to the bus, and the general-purpose processor is used to generate data processing tasks; the second chip includes a second dedicated processing unit, and the second dedicated processing unit has at least partially the same computing function as one or more of the at least one first dedicated processing units, and at least one of the one or more first dedicated processing units and the second dedicated processing unit can process at least a part of the data processing tasks based on the computing function; wherein, the first chip and the second chip are interconnected through an inter-chip interconnect line.
[0008] With the device provided in the first aspect, when the general-purpose processor in the first chip generates a data processing task, because the second dedicated processing unit has at least partially the same computing function as one or more of the at least one first dedicated processing units, the second dedicated processing unit can process at least a part of the data processing tasks based on the at least partially same computing function. Therefore, the data processing device can flexibly allocate data processing tasks to the first dedicated processing unit and the second dedicated processing unit according to the needs of data tasks to meet the user's computing requirements. For example, when the amount of data processing tasks is small, it can be allocated to the first dedicated processing unit alone, or to the second dedicated processing unit alone, etc.; for another example, when the amount of data processing tasks is large, it can be allocated to the first dedicated processing unit and the second dedicated processing unit at the same time. Moreover, the stacked chip architecture in this data processing device can also, without significantly increasing the overall chip volume, have the second dedicated processing unit assist the first dedicated processing unit in processing data processing tasks, enhance the computing power of the first dedicated processing unit, greatly alleviate and solve the demand for improving the chip computing power due to the rapid development of current algorithms, and at the same time avoid the problem that the chip cannot complete data processing tasks due to insufficient computing power of the first dedicated processing unit. Therefore, this data processing device can meet the user's increasingly high power consumption and computing requirements without increasing the product volume, and at the same time improve the product performance and meet the flexibility of task processing.
[0009] In a possible implementation, the second dedicated processing unit has the same computing function as one or more of the at least one first dedicated processing unit. Implementing the embodiments of the present application, since the second dedicated processing unit has the same computing function as one or more of the at least one first dedicated processing unit, the second dedicated processing unit can process the same data processing tasks as one or more of the first dedicated processing units. For example, when one or more of the first dedicated processing units cannot execute data processing tasks in a timely manner, the second dedicated processing unit can assist one or more of the first dedicated processing units in executing data processing tasks to meet the increasingly high computing requirements of users.
[0010] In a possible implementation, the general-purpose processor includes a central processing unit (CPU). Implementing the embodiments of the present application, the general-purpose processor in the first chip can be a central processing unit, which can serve as the operation and control core of the chip system, generate data processing tasks, and at the same time is the final execution unit for information processing and program operation to meet the basic computing requirements of users.
[0011] In a possible implementation, each of the one or more first dedicated processing units and the second dedicated processing unit includes at least one of a graphics processing unit (GPU), an image signal processor (ISP), a digital signal processor (DSP), or a neural network processing unit (NPU). Implementing the embodiments of the present application, the one or more first dedicated processing units and the second dedicated processing unit can be respectively used to execute data processing tasks of different data types so that the intelligent terminal can adapt to the increasing different types of data processing requirements of users.
[0012] In a possible implementation, the inter-chip interconnect line includes at least one of a through-silicon via (TSV) interconnect line or a wire bonding interconnect line. Implementing the embodiments of the present application, various efficient inter-chip interconnect lines can be used for the inter-chip interconnect line, such as a through-silicon via (TSV) interconnect line, a wire bonding interconnect line, and so on. For example, as a through-chip via interconnect technology, TSV has a small aperture, low latency, and flexible configuration of the inter-chip data bandwidth, which improves the overall computing efficiency of the chip system. Using the TSV through-silicon via technology, a non-bump bonding structure can also be achieved, and adjacent chips with different properties can be integrated together. The wire bonding interconnect line of the stacked chip reduces the length of the inter-chip interconnect line and effectively improves the working efficiency of the device itself.
[0013] In a possible implementation, the device further includes a third chip, which is stacked and packaged with the first chip and the second chip. The third chip is connected to at least one of the first chip or the second chip through the inter-chip interconnection line. The third chip includes at least one of a memory, a power transmission circuit module, an input / output circuit module, or an analog module. Implementing the embodiments of the present application, the first chip can be stacked and packaged together with the third chip. For example, stacking the memory with the first chip can partially solve the problem of insufficient memory bandwidth while increasing the computing power of the chip. Stacking one or more of the power transmission circuit module, the input / output circuit module, or the analog module with the first chip can achieve the purpose of separating and decoupling the analog and logical calculations of the SOC chip while increasing the computing power of the chip, and continue to meet the increasingly high demands of chip evolution and business scenarios for chips.
[0014] In a possible implementation, the inter-chip interconnection line is connected between the one or more first dedicated processing units and the second dedicated processing unit. The second dedicated processing unit is configured to obtain at least a part of the data processing task from the one or more first dedicated processing units. Implementing the embodiments of the present application, the second chip is connected to the one or more first dedicated processing units in the first chip and can obtain at least a part of the data processing task from the one or more first dedicated processing units. Therefore, the second dedicated processing unit can be allocated by the one or more first dedicated processing units to execute at least a part of the data processing task, thereby meeting the flexibility of task processing.
[0015] In a possible implementation, the inter-chip interconnection line is connected between the second dedicated processing unit and the bus. The second dedicated processing unit is configured to obtain at least a part of the data processing task from the general-purpose processor through the bus. Implementing the embodiments of the present application, the second chip is connected to the bus in the first chip and can obtain at least a part of the data processing task from the general-purpose processor. Therefore, the second dedicated processing unit can be allocated by the general-purpose processor to execute at least a part of the data processing task alone or together with the one or more first dedicated processing units, thereby meeting the flexibility of task processing.
[0016] In a possible implementation, the general-purpose processor is configured to send startup information to the second dedicated processing unit via the inter-chip interconnect line; the second dedicated processing unit is configured to, in response to the startup information, transition from the waiting state to the startup state and process at least a part of the data processing task based on the computing function. Implementing the embodiments of the present application, the general-purpose processor can allocate the computing power of the second dedicated processing unit. For example, when the general-purpose processor in the first chip sends startup information to the second dedicated processing unit via the inter-chip interconnect line, the second chip can transition from the waiting state to the startup state to execute the data processing task. Among them, the power consumption in the waiting state is lower than that in the startup state. Therefore, when the general-purpose processor does not send startup information to the second dedicated processing unit, the second dedicated processing unit will always be in the waiting state, thereby effectively controlling the power consumption of the stacked chips.
[0017] In a possible implementation, the general-purpose processor is configured to send startup information to the second dedicated processing unit via the inter-chip interconnect line when the computing power of the one or more first dedicated processing units does not meet the requirements. Implementing the embodiments of the present application, when one or more first dedicated processing units in the first chip are executing a data processing task and the computing power of the one or more first dedicated processing units is insufficient, the second chip in the stacked chips can receive the startup information sent by the general-purpose processor and, according to the startup information, transition the second chip from the waiting state to the startup state to assist the one or more first dedicated processing units in the first chip to execute the data processing task, so as to enhance or supplement the computing power of the one or more first dedicated processing units and avoid the chip being unable to complete the data processing task of the target data due to the insufficient computing power of the one or more first dedicated processing units. Therefore, when the computing power of the one or more first dedicated processing units meets the computing requirements, the second chip can always be in the waiting state, and there is no need to start the startup state of the second chip. Only the dedicated processing units in the first chip are made to execute the data processing task, reducing the overall power consumption of the chip.
[0018] In a possible implementation, the one or more first dedicated processing units are configured to send startup information to the second dedicated processing unit via the inter-chip interconnect line; the second dedicated processing unit is configured to, in response to the startup information, transition from a waiting state to a startup state and process at least a part of the data processing task based on the computing function. Implementing the embodiments of the present application, the one or more first dedicated processing units can allocate the computing power of the second dedicated processing unit. For example, when one or more first dedicated processing units in the general-purpose processor of the first chip send startup information to the second dedicated processing unit via the inter-chip interconnect line, the second chip can transition from a waiting state to a startup state to execute the data processing task. Among them, the power consumption in the waiting state is lower than that in the startup state. Therefore, when the one or more first dedicated processing units do not send startup information to the second dedicated processing unit, the second dedicated processing unit will always be in the waiting state, thereby effectively controlling the power consumption of the stacked chips.
[0019] In a possible implementation, the one or more first dedicated processing units are configured to send startup information to the second dedicated processing unit via the inter-chip interconnect line when the computing power of the one or more first dedicated processing units does not meet the requirements. Implementing the embodiments of the present application, when one or more first dedicated processing units in the first chip are executing a data processing task and the computing power of the one or more first dedicated processing units is insufficient, the second chip in the stacked chips can receive the startup information sent by the one or more first dedicated processing units and, according to the startup information, transition the second chip from a waiting state to a startup state to assist the one or more first dedicated processing units of the first chip in executing the data processing task, so as to enhance or supplement the computing power of the one or more first dedicated processing units, avoid the situation where the chip cannot complete the data processing task of the target data due to insufficient computing power of the one or more first dedicated processing units, and meet the computing requirements of the user. Therefore, when the computing power of the one or more first dedicated processing units meets the computing requirements, the second chip can always be in the waiting state, and there is no need to start the startup state of the second chip, and only the dedicated processing units in the first chip are used to execute the data processing task, reducing the overall power consumption of the chip.
[0020] In a second aspect, an embodiment of the present application provides a data processing method, which is characterized by including: generating a data processing task through a general-purpose processor in a first chip, where the first chip includes the general-purpose processor, a bus, and at least one first dedicated processing unit (DPU), and the general-purpose processor and the at least one first dedicated processing unit are connected to the bus; processing at least a part of the data processing task through one or more of the at least one first dedicated processing unit and at least one of the second dedicated processing units in a second chip package, where the second dedicated processing unit has at least partially the same computing function as one or more of the at least one first dedicated processing unit, and the first chip and the second chip are stacked and packaged and connected to each other through an inter-chip interconnect line.
[0021] In a possible implementation, the processing at least a part of the data processing task through one or more of the at least one first dedicated processing unit and at least one of the second dedicated processing units in a second chip package includes: sending startup information to the second dedicated processing unit through the general-purpose processor via the inter-chip interconnect line; in response to the startup information, the second dedicated processing unit transitions from a waiting state to a startup state and processes at least a part of the data processing task based on the computing function.
[0022] In a possible implementation, the sending startup information to the second dedicated processing unit through the general-purpose processor via the inter-chip interconnect line includes: when the computing power of one or more of the at least one first dedicated processing unit does not meet the requirements, sending startup information to the second dedicated processing unit through the general-purpose processor via the inter-chip interconnect line.
[0023] In a possible implementation, the processing at least a part of the data processing task through one or more of the at least one first dedicated processing unit and at least one of the second dedicated processing units in a second chip package includes: sending startup information to the second dedicated processing unit through one or more of the at least one first dedicated processing unit via the inter-chip interconnect line; the second dedicated processing unit is configured to, in response to the startup information, transition from a waiting state to a startup state and process at least a part of the data processing task based on the computing function.
[0024] In a possible implementation, the process of the one or more first dedicated processing units sending startup information to the second dedicated processing unit through the inter-chip interconnecting line includes: when the computing power of the one or more first dedicated processing units does not meet the requirements, the one or more first dedicated processing units send startup information to the second dedicated processing unit through the inter-chip interconnecting line.
[0025] In a third aspect, an embodiment of the present application provides a chip system. The chip system includes any device for supporting data processing involved in the first aspect above. The chip system may be composed of chips, or may include chips and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] To clearly illustrate the technical solutions in the embodiments of the present application, the following describes the drawings required to be used in the embodiments of the present application.
[0027] Figure 1 It is a diagram of a traditional von Neumann chip architecture provided by an embodiment of the present application.
[0028] Figure 2A It is a schematic diagram of a data processing architecture provided by an embodiment of the present application.
[0029] Figure 2B It is a schematic diagram of a stacked package chip architecture provided by an embodiment of the present application.
[0030] Figure 2C It is a schematic diagram of a stacked package chip architecture in actual application provided by an embodiment of the present application.
[0031] Figure 2D It is a kind of the above-mentioned Figure 2C Schematic diagram of the interaction between the second dedicated processing unit and the first dedicated processing unit in the stacked package chip shown.
[0032] Figure 2E It is another schematic diagram of a stacked package chip architecture in actual application provided by an embodiment of the present application.
[0033] Figure 2F It is a kind of the above-mentioned Figure 2E Schematic diagram of the interaction between the second dedicated processing unit and the first dedicated processing unit in the stacked package chip shown.
[0034] Figure 2G It is another schematic diagram of a stacked package chip architecture provided by an embodiment of the present application.
[0035] Figure 3 It is a schematic diagram of the process flow of a data processing method provided by an embodiment of the present application. Detailed implementation manners
[0036] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.
[0037] The terms "first", "second", "third", "fourth", etc. in the specification, claims and drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices.
[0038] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appearing at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0039] The terms "component", "module", "system", etc. used in this specification are used to represent computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, an application running on a computing device and the computing device can both be components. One or more components may reside in a process and / or an execution thread, and the components may be located on one computer and / or distributed between two or more computers. In addition, these components may execute from various computer-readable media on which various data structures are stored. The components may communicate, for example, through local and / or remote processes according to signals having one or more data packets (for example, data from two components interacting with another component between a local system, a distributed system, and / or a network, such as data interacting with other systems through signals on the Internet).
[0040] The "connection" involved in this embodiment means that communication, data interaction, energy transfer, etc. can be carried out between the connected modules, units or devices. It can be a direct connection or an indirect connection through other devices or modules. For example, it can be connected through some wires, conductors, media, interfaces, devices or units, because it can be regarded as an electrical connection or coupling in a broad sense.
[0041] First, some terms in this application are explained to facilitate understanding by those skilled in the art.
[0042] (1) Through Silicon Via (TSV), also known as via silicon through, is a vertical interconnection that penetrates a silicon wafer or chip. TSV can be used to implement 3D integrated circuit (IC) packaging, following Moore's Law. For example, TSV can stack multiple chips, and its design concept comes from a printed circuit board (PCB). Specifically, small holes are drilled in the chip (the process can be divided into two types: via first and via last), and metal is filled from the bottom of the small hole. For example, holes (vias) are drilled in the silicon wafer forming the chip by etching or laser, and then filled with conductive materials such as copper, polysilicon, tungsten, etc. This technology can effectively improve the integration and performance of the system at a lower cost.
[0043] (2) System on Chip (SOC), also known as a system-on-a-chip, is an integrated circuit with a dedicated purpose, which contains a complete functional circuit system and includes all the content of the embedded software. There are multiple different functional components on the SOC, which will be introduced later.
[0044] (3) Stacked structure refers to a type of system packaging. Among them, system packaging can be divided into three types: adjacent structure, stacked structure, and buried structure. The use of a stacked structure can increase the packaging density in the three-dimensional direction and can be applied to different levels or grades of packaging. For example: Package on Package (PoP), Package-in-Package (PiP), chip or die stacking, chip and wafer stacking. Chip stack packaging is common in various terminal products. Its advantage is that standard chips, wire bonding, and subsequent packaging can be achieved using existing equipment and processes, but it has limitations on the thickness of the entire package and cannot be too large. Currently, it is already possible to achieve vertical installation of up to 8 dies in a single package, and its thickness is less than 1.2 mm, which requires that each die in the stacked package be a thin wafer, thin substrate, low wire bond arc, low mold cover height, etc.
[0045] (4) Wafer refers to the silicon wafer used for the production of silicon semiconductor integrated circuits. Since its shape is circular, it is also called a silicon wafer.
[0046] (5) Die is an integrated circuit product that contains various circuit element structures and has specific electrical functions after being processed and cut on a wafer. Among them, a die can be made into a chip after being packaged.
[0047] (6) The Graphics Processing Unit (GPU), also known as the display core, visual processor, or display chip, is a microprocessor specialized for performing image and graphics-related operations on personal computers, workstations, game consoles, and some mobile devices (such as tablets, smartphones, etc.). The GPU reduces the device's dependence on the CPU and undertakes some of the work originally done by the CPU. Especially in 3D graphics processing, the core technologies adopted by the GPU include hardware geometry transformation and lighting processing (T&L), cubic environment texture mapping and vertex blending, texture compression and bump mapping, dual-texture quad-pixel 256-bit rendering engine, etc. Among them, the hardware geometry transformation and lighting processing technology can be regarded as the hallmark of the GPU.
[0048] (7) The Digital Signal Processor (DSP) generally refers to a chip or processor that executes digital signal processing technology. Digital signal processing technology is the technology of converting analog information (such as sound, video, and pictures) into digital information. They may also be used to process this analog information and then output it as analog information. It can also be based on digital signal processing theory, hardware technology, and software technology to study digital signal processing algorithms and their implementation methods.
[0049] (8) The Neural-network Processing Unit (NPU) is a processor for processing neural models. It can be regarded as a component (or subsystem), and sometimes it can also be called an NPU coprocessor. Generally, it adopts an architecture of "data-driven parallel computing" and is particularly good at processing massive multimedia data such as videos and images.
[0050] (9) Printed circuit boards (PCBs) are providers of electrical connections for electronic components. According to the number of circuit board layers, they can be divided into single-sided boards, double-sided boards, four-layer boards, six-layer boards, and other multi-layer circuit boards.
[0051] (10) Static Random-Access Memory (SRAM) is a type of random-access memory. The so-called "static" means that as long as this memory remains powered on, the data stored in it can be constantly maintained. In contrast, the data stored in Dynamic Random-Access Memory (DRAM) needs to be updated periodically. However, when the power supply stops, the data stored in SRAM will still disappear (referred to as volatile memory), which is different from ROM or flash memory that can store data even after power-off.
[0052] (11) Input / Output (I / O) generally refers to the input and output of data between internal devices and external memories or other peripheral devices, which is the communication between the information processing system of internal devices (such as calculators) and the external world (which may be humans or another information processing system). Input is the signal or data received by the system, and output is the signal or data sent from it. This term can also be used as part of an action; to "run I / O" is to perform the operation of running input or output.
[0053] (12) The von Neumann architecture, also known as the Princeton architecture, is a memory architecture that combines the program instruction memory and the data memory. The program instruction storage address and the data storage address point to different physical locations of the same memory, so the widths of program instructions and data are the same.
[0054] (13) The arithmetic unit refers to the smallest arithmetic unit of the arithmetic unit in the CPU and is a hardware structure. Currently, operations are all completed through logic circuits formed by tiny components such as electronic circuits, and the signals processed are high and low level signals, that is, binary signals.
[0055] Please refer to the appendix Figure 1 , Figure 1 is a traditional von Neumann chip architecture diagram provided by an embodiment of the present application. As Figure 1 shown, it includes an arithmetic unit, a controller, a memory, an input device, and an output device.
[0056] The traditional von Neumann chip architecture is difficult to meet the current requirements of chip computing power, power consumption, and chip storage bandwidth. Therefore, the chip stacking technology can be used to solve some problems such as insufficient chip computing power, excessive power consumption, and insufficient chip storage bandwidth caused by limited chip area. Regarding the technology of chip stacking technology (taking the SOC chip of a mobile intelligent terminal as an example), the power transmission circuit, I / O, or radio frequency circuit in the SOC chip can be separately split onto another chip, and the decoupling of circuits with different functions can be achieved. An embodiment of the present application can also decouple the input / output (I / O) module or the dedicated processing unit from the functions of the main chip as an independent chip. In this way, the main chip will no longer have the functions of the decoupled independent chip, and the main chip and the independent chip need to work simultaneously, and the power consumption performance is not optimized enough.
[0057] In view of the above analysis, a data processing device provided by an embodiment of the present application can enhance the computing power of the chip to process data, reduce the overall power consumption of the chip, and consider the product volume problem. To facilitate the understanding of the embodiments of the present application, the architecture of one of the data processing devices on which the embodiments of the present application are based will be described below. Please refer to the appendix Figure 2A , Figure 2A is a schematic diagram of a data processing architecture provided by an embodiment of the present application.Figure 2A The architecture shown mainly takes stacked chips as the main body and is described from the perspective of data processing.
[0058] It should be noted that, in a broad sense, a chip can be a packaged chip or an unpackaged chip, that is, a bare die. The "first chip", "second chip", "main chip", etc. involved in the chip structures for stacked packaging in this application can all be understood as unpackaged chips, that is, bare dies (dies), which refer to integrated circuit products that contain various circuit element structures on a cut silicon wafer and have specific electrical functions. Therefore, the chips that need to be stacked and packaged in the embodiments of this application are all bare dies that have not been packaged.
[0059] As Figure 2A shown, when the user finishes using the recording function of the intelligent terminal, the intelligent terminal can process the recorded audio file through the stacked chips and then store it in the memory. Among them, the stacked chip system in the intelligent terminal includes a first chip 001 and a second chip 002. The first chip 001 can be a system-on-chip (SOC). The first chip 001 includes a general-purpose processor 011, a bus 00, and at least one first dedicated processing unit (Domain Processing Unit, DPU) 021. The general-purpose processor 011 and the at least one first dedicated processing unit 021 are connected to the bus. Among them, the at least one first dedicated processing unit 021 can be sequentially denoted as: DPU_1 to DPU-n; the second chip 002 includes a second dedicated processing unit 012 (which can be denoted as: DPU-A); the first chip 001 and the second chip 002 are connected through an inter-chip interconnect line.
[0060] The first chip 001 is used to process data and generate data processing tasks, and can also send startup information to the second chip 002. The second chip 002 is used to convert from the waiting state to the startup state when receiving the startup information, and execute part or all of the data processing tasks through the second dedicated processing unit.
[0061] Among them, the general-purpose processor 011 (for example: central processing unit, CPU) in the first chip 001 serves as the operation and control core of the chip system and is the final execution unit for information processing and program operation. In the field of mobile terminals, the general-purpose processor 011 generally may include the Advanced RISC Machines (ARM) series, which may include one or more core processing units. In the embodiments of the present application, the general-purpose processor 011 may be used to generate a data processing task for target data. Optionally, the general-purpose processor 011 may also select one or more first dedicated processing units 021 from at least one first dedicated processing unit 021 for at least a part of the data processing task. Optionally, the general-purpose processor 011 may also be used to process simple data processing tasks without the assistance of at least one first dedicated processing unit 021 and the second dedicated processing unit 012.
[0062] At least one first dedicated processing unit 021 in the first chip 001 may process at least a part of the data processing task based on a computing function. For example, the data processing task of image recognition may be executed by a graphics processing unit GPU. Optionally, at least one first dedicated processing unit 021 may include at least one of a graphics processing unit GPU, an image signal processor ISP, a digital signal processor DSP, or a neural network processing unit NPU. Optionally, all the first dedicated processing units DPU in at least one first dedicated processing unit 021 may work simultaneously, or only one may work.
[0063] The bus 00 in the first chip 001, also known as the internal bus or board-level bus or microcomputer bus. It can be used to connect the various functional components within the chip to form a complete chip system, and can also be used to transmit various data signals, control commands, etc., to assist in the communication between the various functional devices. For example, it can connect the general-purpose processor 011 and at least one first dedicated processing unit 021, so that the general-purpose processor 011 can control one or more first dedicated processing units 021 in at least one first dedicated processing unit 021 to execute data processing tasks.
[0064] The second dedicated processing unit 012 in the second chip 002 has at least partially the same computing functions as one or more of the at least one first dedicated processing unit 021, and at least one of the one or more first dedicated processing units 021 and the second dedicated processing unit 012 can process at least a part of the data processing task based on the computing functions. Among them, that the second dedicated processing unit 012 has at least partially the same computing functions as the one or more first dedicated processing units 021 may mean that the computing functions of the second dedicated processing unit 012 and the computing functions of the one or more first dedicated processing units 021 can be at least partially the same, and this part of the same functions is the common functions shared by the second dedicated processing unit 012 and the one or more first dedicated processing units 021. This part of the common functions can be used to calculate the data processing task, and the data processing task can be assigned to the second dedicated processing unit 012 and the one or more first dedicated processing units 021 for processing. For example, the computing power can be shared by the second dedicated processing unit 012 and the one or more first dedicated processing units 021, and a detailed introduction to this will be given in the following lecture. Of course, the computing functions of the second dedicated processing unit 012 and the one or more first dedicated processing units 021 can also be completely the same. For example, the second dedicated processing unit 012 and the one or more first dedicated processing units 021 can execute the same data processing task through partially the same operation logic, operation method, and / or operation purpose. For example: one or more of the first dedicated processing units 021 can include the computing ability of a convolutional neural network, and the second dedicated processing unit 012 can also include at least part of the computing ability of a convolutional neural network, such as most of the ability. Another example: in neural operations, one or more of the first dedicated processing units 021 can convert a piece of speech into several keywords through a speech recognition method, and the second dedicated processing unit 012 can convert a piece of speech into a string of characters including the keywords through a speech recognition method. Another example, one or more of the first dedicated processing units 021 can include the image operation processing ability of ISP to generate a captured image. For example, this ability can include white balance, noise reduction, pixel calibration, image sharpening, and gamma calibration. The second dedicated processing unit 012 can also include most of the previous image operation processing ability of ISP to generate a captured image. For example, this ability can include white balance, noise reduction, pixel calibration, or image sharpening.
[0065] Optionally, the second dedicated processing unit 012 in the second chip 002 has the same computing function as one or more of the at least one first dedicated processing unit 021, and can be used to execute all processing tasks of the acquired data processing task. Among them, the second dedicated processing unit 012 and the one or more first dedicated processing units 021 having the same computing function may include: the second dedicated processing unit 012 and the one or more first dedicated processing units 021 may have exactly the same operation logic and / or operation mode to execute data processing tasks. For example: when one of the at least one first dedicated processing unit 021 includes a unit of a parallel matrix operation array, the second dedicated processing unit 012 may also include a unit for implementing a parallel matrix operation array, and the algorithm type may also be consistent with the algorithm type used by the first dedicated processing unit 021.
[0066] Among them, the inter-chip interconnection line includes any one of through-silicon via (TSV) interconnection lines and wire bonding interconnection lines. For example, as an inter-chip through-hole interconnection technology, TSV has a small aperture, low latency, flexible configuration of inter-chip data bandwidth, and improves the overall computing efficiency of the chip system. Using TSV technology, a non-bulging bonding structure can be achieved, and adjacent chips with different properties can be integrated together. The wire bonding interconnection line of the stacked chip reduces the length of the inter-chip interconnection line and effectively improves the working efficiency of the device itself.
[0067] It should be noted that the connection signals of the inter-chip interconnection line may include data signals and control signals. Among them, the digital signal may be used to transmit the target data, and the control signal may be used to allocate the data processing tasks of the target data. The present application does not make specific limitations in this regard.
[0068] It should also be noted that Figure 2A the data processing architecture in is only an exemplary implementation manner in the embodiments of the present application. The data processing architecture in the embodiments of the present application includes but is not limited to the above data processing architecture. For example: the first chip 001 may also be stacked and packaged with multiple second chips 002 into a single chip, and the second dedicated processing units included in the multiple second chips 002 respectively have at least partially the same computing function as at least one first dedicated processing unit 021 included in the first chip 001.
[0069] It should also be noted that the stacked chip system may be configured in different data processing devices, corresponding to different main control forms in different data processing devices. The embodiments of the present application do not limit the form of the main control. For example, servers, laptops, smartphones, in-vehicle TVs, and so on.
[0070] Secondly, based on the aboveFigure 2A The provided data processing architecture. In the embodiments of the present application, two chip stacking architectures in the data processing device architecture are described. Please refer to the appendix Figure 2B , appendix Figure 2B is a schematic diagram of a stacked package chip architecture provided by the embodiments of the present application. As Figure 2B shown in the stacked package chip architecture, the first chip and the second chip are interconnected through inter-chip interconnect lines, mainly with the SOC chip as the main chip, and are described from the perspective of data processing.
[0071] As Figure 2B shown in the chip architecture, it includes: a first chip 001 and a second chip 002; the first chip 001 includes a general-purpose processor 011 (such as: CPU, or optionally a microcontroller), at least one first dedicated processing unit 021 (DPU_1, DPU_2... DPU_n) and a bus 00, and also includes: a memory 031, an analog module 041, an input / output module 051, etc.; the second chip 002 is a computing power stacked chip connected to the first dedicated processing unit DPU_1, and the second chip 002 includes a second dedicated processing unit 012 (i.e.: DPU_A).
[0072] Among them, the general-purpose processor 011 in the first chip 001 can be a CPU, which is used to generate data processing tasks.
[0073] Optionally, the general-purpose processor 011 in the first chip 001 is further used to allocate data processing tasks to one or more of the at least one first dedicated processing unit 021, and / or the second dedicated processing unit 012 in the second chip. The general-purpose processor 011 in the first chip 001 can flexibly allocate data processing tasks to the first dedicated processing unit and the second dedicated processing unit according to the needs of the data task, so as to meet the increasingly high computing requirements of users. For example: when the amount of data processing tasks is small, it can be separately allocated to the first dedicated processing unit; for another example: when the amount of data processing tasks is large, it can be allocated to the first dedicated processing unit and the second dedicated processing unit at the same time, or separately allocated to the second dedicated processing unit, etc. It can meet the flexibility of task processing without increasing the product volume.
[0074] Optionally, the inter-chip interconnecting line is connected between the second dedicated processing unit and the bus. The general-purpose processor 011 in the first chip 001 can also be used to send startup information to the second dedicated processing unit through the inter-chip interconnecting line, so that the second dedicated processing unit can, in response to the startup information, convert from the waiting state to the startup state and process at least a part of the data processing task based on its computing function. Among them, the power consumption in the waiting state is lower than that in the startup state. Therefore, when the general-purpose processor does not send startup information to the second dedicated processing unit, the second dedicated processing unit will always be in the waiting state, thereby effectively controlling the power consumption of the stacked chips.
[0075] Optionally, the inter-chip interconnecting line is connected between the second dedicated processing unit and the bus. The general-purpose processor 011 in the first chip 001 can also be used to send startup information to the second dedicated processing unit through the inter-chip interconnecting line when the computing power of one or more first dedicated processing units does not meet the requirements. When one or more first dedicated processing units 021 in the first chip are executing a data processing task and the computing power of one or more first dedicated processing units 021 is insufficient, the second chip 002 in the stacked chips can receive the startup information sent by the general-purpose processor 011, and convert the second chip 002 from the waiting state to the startup state according to the startup information, and assist the one or more first dedicated processing units 021 in the first chip 001 to execute the data processing task, so as to enhance or supplement the computing power of the one or more first dedicated processing units 021, and avoid the situation that the chip cannot complete the data processing task of the target data due to the insufficient computing power of one or more first dedicated processing units 021.
[0076] It should be noted that according to the computing power of the one or more first dedicated processing units 021, it is predicted whether the one or more first dedicated processing units 021 can complete the execution of the data processing task within a preset time; if the one or more first dedicated processing units 021 cannot complete the execution of the data processing task within the preset time, it is determined that the one or more first dedicated processing units 021 have insufficient computing power when executing the data processing task.
[0077] At least one first dedicated processing unit 021 in the first chip 001 is sequentially denoted as DPU_1, DPU_2... DPU_n from left to right. Among them, one or more first dedicated processing units 021 in the at least one first dedicated processing unit 021 are used to obtain and execute data processing tasks based on corresponding computing functions.
[0078] Optionally, the at least one first dedicated processing unit 021 may include one or more of a graphics processing unit (GPU), an image signal processor (ISP), a digital signal processor (DSP), and a neural network processing unit (NPU). For example, the GPU and the ISP can be used to process graphics data in the smart terminal; the DSP can be used to process digital signal data executed in the smart terminal; the NPU can be used to process a large amount of multimedia data such as videos and images in the smart terminal. Therefore, the at least one first dedicated processing unit 021 can be respectively used to execute different data processing tasks through different computing functions, so that the smart terminal can adapt to the increasing data processing requirements of users. The DPU_1 in the first chip 001 of the embodiments of the present application is used to execute data processing tasks.
[0079] Optionally, the inter-chip interconnect line is connected between the one or more first dedicated processing units and the second dedicated processing unit; when the one or more first dedicated processing units 021 in the first chip 001 execute the data processing tasks, the one or more first dedicated processing units 021 are further configured to send start information to the second chip 002 and allocate part or all of the processing tasks of the data processing tasks to the second dedicated processing unit 012 of the second chip 002. So that the second dedicated processing unit 012 responds to the start information, switches from the waiting state to the start state, and processes at least a part of the data processing tasks based on its computing function that is at least partially the same as that of the one or more first dedicated processing units 021. Wherein, the power consumption in the waiting state is lower than that in the start state. Therefore, when the one or more first dedicated processing units 021 do not send start information to the second dedicated processing unit 012, the second dedicated processing unit 012 will always be in the waiting state, and thus the power consumption of the stacked chips can be effectively controlled. It should be noted that the power consumption in the waiting mode is lower than that in the working mode.
[0080] Optionally, the inter-chip interconnecting line is connected between the one or more first dedicated processing units and the second dedicated processing unit; the one or more first dedicated processing units 021 in the first chip 001 are configured to send start-up information to the second dedicated processing unit 012 through the inter-chip interconnecting line when the computing power of the one or more first dedicated processing units 021 does not meet the requirements. When one or more first dedicated processing units 021 in the first chip 001 are executing a data processing task and the computing power of the one or more first dedicated processing units 021 is insufficient, the second chip 002 in the stacked chips can receive the start-up information sent by the one or more first dedicated processing units 021, and convert the second chip 002 from the waiting state to the start-up state according to the start-up information, and assist the one or more first dedicated processing units 021 in the first chip 001 to execute the data processing task, so as to enhance or supplement the computing power of the one or more first dedicated processing units 021, and avoid that the chip cannot complete the data processing task of the target data due to the insufficient computing power of the one or more first dedicated processing units 021, so as to meet the computing requirements of the user.
[0081] The bus 00 in the first chip 001 is used to connect the general-purpose processor 011 and the at least one first dedicated processing unit 021.
[0082] The memory 031 in the first chip 001 is used to store the target data and the data processing task corresponding to the target data. Among them, the target data type may include graphic data, video data, audio data, text data, and so on.
[0083] The analog module 041 in the first chip 001 mainly realizes analog processing functions, such as radio frequency front-end analog, port physical layer (PHY), etc.
[0084] The input / output module 051 in the first chip 001 is a general interface of the SOC chip to external devices and is used for data input and output. Generally, it includes a controller and a port physical layer (PHY), such as a Universal Serial Bus (USB) interface, a Mobile Industry Processor Interface (MIPI), and so on.
[0085] The second dedicated processing unit 012 in the second chip 002 has at least partially the same computing functions as one or more of the at least one first dedicated processing unit 021, and can be used to execute part or all of the data processing tasks assigned by the general-purpose processor 011 or one or more first dedicated processing units 021. For example, when one of the at least one first dedicated processing unit 021 is an NPU, the core of the NPU is a unit of a parallel matrix operation array. Then the second dedicated processing unit 012 is also used to implement the unit of the parallel matrix operation array, and the algorithm type is also the same as the algorithm type used by the NPU, such as Int8, Int6, F16, etc. However, in terms of the number of arithmetic units, the second dedicated processing unit 012 may be different from the NPU, that is. The second dedicated processing unit 012 and the first dedicated processing unit 021 implement the same computing functions, but there are differences in computing power.
[0086] Optionally, the second dedicated processing unit 012 in the second chip 002 has the same computing functions as one or more of the at least one first dedicated processing unit 021. Therefore, the second dedicated processing unit 012 can process the same data processing tasks as one or more first dedicated processing units 021. For example, when one or more first dedicated processing units 021 execute data processing tasks, the second dedicated processing unit 012 can assist the one or more first dedicated processing units 021 to jointly execute the data processing tasks to more efficiently meet the increasingly high computing requirements of users.
[0087] Optionally, the second dedicated processing unit in the second chip 002 is used to respond to the start information, convert from the waiting state to the start state, and process at least a part of the data processing tasks based on the computing functions. When one or more first dedicated processing units 021 in the first chip 001 are executing data processing tasks, the second chip 002 in the stacked chip system can receive the start information sent by the one or more first dedicated processing units 021, and convert the second chip 002 from the waiting mode to the working mode according to the start information to assist the one or more first dedicated processing units 021 of the first chip 001 to execute data processing tasks, so as to enhance or supplement the computing power of the one or more first dedicated processing units 021 and avoid the chip being unable to complete the data processing tasks of the target data due to insufficient computing power of the one or more first dedicated processing units 021. Therefore, when the general-purpose processor does not send start information to the second dedicated processing unit, the second dedicated processing unit will always be in the waiting state, thereby effectively controlling the power consumption of the stacked chips.
[0088] In a possible implementation, the second dedicated processing unit 012 includes corresponding arithmetic units within one or more first dedicated processing units 021 that have at least partially the same computing functions. The arithmetic units corresponding to the one or more first dedicated processing units 021 are respectively used to process target data through arithmetic logic. For example, when the second dedicated processing unit 012 and the neural network processing unit NPU have at least partially the same computing functions, since the arithmetic units of the neural network processing unit NPU include: Matrix Unit, Vector Unit, and Scalar Unit, the second dedicated processing unit 012 within the second chip 002 can also include one or more of Matrix Unit, Vector Unit, and Scalar Unit, so as to perform matrix multiplication, vector operations, scalar operations, etc. on the data for executing the data processing tasks assigned to the second chip 002.
[0089] Optionally, the inter-chip interconnect lines between the first chip 001 and the second chip 002 are TSVs. Since the TSV process technology can achieve a single silicon via aperture as small as the order of ~10um, the number of TSV interconnect lines between the first chip 001 and the second chip 002 can be determined as needed without occupying too much area, and no specific limitation is made in this application comparison.
[0090] In a possible implementation, Figure 2B the first chip 001 in may include multiple first dedicated processing units 021. Among them, each first dedicated processing unit 021 can receive data processing tasks sent by the general-purpose processor 011 and is used to process a corresponding certain type of data processing task. For example, the first chip 001 usually includes a graphics processing unit GPU, an image signal processor ISP, a digital signal processor DSP, and a neural network processing unit NPU, etc. Among them, the graphics processing unit GPU and the image signal processor ISP can be used to process graphics data in the smart terminal; the digital signal processor DSP can be used to process digital signal data executed in the smart terminal; the neural network processing unit NPU can be used to process a large amount of multimedia data such as videos and images in the smart terminal.
[0091] In a possible implementation, the second dedicated processing unit 012 in the second chip 002 may include one or more arithmetic units, and the one or more arithmetic units are used to process data through corresponding arithmetic logics respectively. For example, when the second dedicated processing unit 012 has at least partially the same computing functions as the NPU, since the core computing units of the NPU are the Matrix Unit, the Vector Unit, and the Scalar Unit, the second dedicated processing unit 012 in the second chip 002 may also include one or more of the Matrix Unit, the Vector Unit, and the Scalar Unit, and is used to perform matrix multiplication, vector operation, scalar operation, etc. on the data for the data processing tasks assigned to the second chip 002.
[0092] It should be noted that after the general-purpose processor in the first chip generates a data processing task, since the second dedicated processing unit has at least partially the same computing functions as one or more of the at least one first dedicated processing units, the second dedicated processing unit may process at least partially the same data processing tasks as one or more first dedicated processing units. Therefore, the data processing device can flexibly allocate the data processing tasks to the first dedicated processing unit and the second dedicated processing unit respectively according to the needs of the data tasks to meet the increasing computing requirements of users. For example, when the amount of data processing tasks is small, it can be allocated to the first dedicated processing unit alone; for another example, when the amount of data processing tasks is large, it can be allocated to the first dedicated processing unit and the second dedicated processing unit at the same time, or allocated to the second dedicated processing unit alone, etc. At the same time, stacking the second chip on the first chip can meet the volume requirements of the product without significantly increasing the volume of the chip architecture. Moreover, when only one chip executes the data processing task, the other chip can be in a low-power state, and the power consumption of the stacked chips during data processing can be effectively controlled. Therefore, the data processing device can meet the increasing power consumption and computing requirements of users without increasing the product volume, while improving the product performance and meeting the flexibility of task processing.
[0093] Based on the above Figure 2A provided data processing architecture, and Figure 2B the provided stacked packaging chip architecture, the following is an application scenario of calling the second dedicated processing unit 012 for enhanced computing power calculation when the computing power of one or more first dedicated processing units 021 is insufficient. Please refer to the attached Figure 2C and the attached Figure 2D , the attached Figure 2C is a schematic diagram of a stacked packaging chip architecture in an actual application provided by an embodiment of the present application. The attached Figure 2Dis a kind of provided by the embodiments of the present application as described above Figure 2C Schematic diagram of the interaction between the second dedicated processing unit and the first dedicated processing unit in the stacked package chip shown above. Among them, the above Figure 2C provided stacked package chip architecture is used to support and execute the following method flow steps 1-step 7.
[0094] As Figure 2C shown, in the stacked package chip architecture, the second dedicated processing unit DPU-A included in the second chip 002 is connected to the first dedicated processing unit DPU_1 through TSV interconnections. Among them, the second dedicated processing unit DPU_A may include a memory (Buffer) 0121 and a task scheduler (Task Scheduler) 0122, and an AI algorithm operation module 0123 that directly schedules task sequences and data by the first dedicated processing unit DPU_1 (the AI algorithm operation module 0123 includes: a matrix unit 01 (Matrix Unit), a vector unit 02 (Vector Unit), and a scalar unit 03 (Scalar Unit)). Among them, the memory (Buffer) 0121 is used to store data and corresponding data processing tasks, the task scheduler (Task Scheduler) 0122 is used to receive the assigned data processing tasks, and the AI algorithm operation module 0123 is used to execute data processing tasks. The first dedicated processing unit DPU_1 may also include: a memory (Buffer) 0211 and a task scheduler (Task Scheduler) 0212, where the corresponding functions can refer to the relevant descriptions of each functional module of the first dedicated processing unit DPU_1 above.
[0095] Specifically, since DPU_A is directly scheduled by DPU_1, the signals directly interconnected by the two devices through TSV vias at least include two types: data signals (Data Signals) and control signals (Control Signals). The number of bits of the two signals depends on specific design requirements. The number of bits of Data Signals is the number of bits required for parallel computing data, generally at least a multiple of 64 bits, such as 64b, 128b,... 1028b, etc., while Control signals generally include single-bit signals such as enable signal, start / stop control signal, and interrupt signal. As Figure 2D shown by the dashed line in, first, DPU_1 transports the data processed in the previous link from the memory (Memory) 031 and stores it in the buffer 0211 of the DPU_1 unit, and then decides whether to send it to DPU_A for assisted computing processing according to the computing requirements.
[0096] It should be noted that for the relevant descriptions of other functional units in the data processing device described in the embodiments of the present application, reference can be made to the relevant descriptions of the stacked package chip architecture provided in the above Figure 2B and the relevant descriptions of steps 1 to 7 of the following method flow, which will not be elaborated here.
[0097] Among them, as Figure 2D shown in the interaction schematic diagram of the second dedicated processing unit and the first dedicated processing unit in the stacked package chip, the data processing method may include:
[0098] 1. After DPU_1 receives the data processing task sent by CPU 011 through Task Scheduler0212, it moves the target data from the memory to the temporary cache Buffer0211 inside DPU_1;
[0099] 2. DPU_1 predicts whether the computing power of the data processing task is sufficient through Task Scheduler0212. If the computing and processing ability of DPU_1 cannot meet the requirements of the data processing task, then Task Scheduler0212 sends a start message to DPU_A through the Controlsignals of the TSV;
[0100] 3. After receiving the start message, DPU_A wakes up from the low-power state, enters the start state, and feeds back a waiting signal to DPU_1 through the Control signals of the TSV;
[0101] 4. The Task Scheduler0212 of DPU_1 distributes the allocated data processing task to the TaskScheduler0122 of DPU_A, and DPU_A moves the data that needs to be processed by DPU_A into the Buffer0121 of the DPU_A unit;
[0102] 5. The Task Scheduler0122 of DPU_A starts the AI algorithm operation module 0123 (including Matrix Unit 01, Vector Unit 02, Scalar Unit 03, etc.) according to the data processing task, and the AI algorithm operation module 0123 reads the data in Buffer0121 and starts to execute the data processing task;
[0103] 6. The AI algorithm operation module 0123 stores the computed data into the Buffer0121 of DPU_A;
[0104] 7. The Task Scheduler 0122 of DPU_A sends a processing completion signal back to the Task Scheduler 0212 of DPU_1 and writes the data from Buffer 0121 of DPU_A back to Buffer 0211 inside DPU_1.
[0105] It should be noted that as a dedicated processing unit for AI processing, the core computing part of DPU_1 generally includes parallel Matrix Unit 01, Vector Unit 02, Scalar Unit 03, etc., and also includes an internal temporary data storage buffer 0211 and a task scheduler 0212. As a computing power enhancement unit of DPU_1, the computing core of DPU_A includes Matrix Unit 01, optionally includes Vector Unit 02 and Scalar Unit 03. At the same time, DPU_A can also include Buffer 0121 and Task Scheduler 0122 according to needs. Although the computing core of DPU_A can include Matrix Unit 01, Vector Unit 02 and Scalar Unit 03, the number of operators or MAC numbers inside each unit can be different from that of DPU_1.
[0106] In the embodiment of the present application, DPU_A of the second chip 002 is positioned as a computing power enhancement module and is directly scheduled and controlled by DPU_1 in the first chip 001. In a conventional scenario, the computing power of DPU_1 itself is sufficient to meet the requirements. At this time, DPU_A of the second chip 002 can be in a waiting state to save the overall power consumption of the chip system. When some scenarios such as video recording processing require high computing power for AI processing or assistance in processing, DPU_1 can activate DPU_A through the Control signals of TSV to participate in the operation processing together.
[0107] It can be understood that in the embodiment of the present application, the AI enhanced operation unit is taken as an example. The present application is not limited to this scenario. It can also be that other first dedicated processing units (such as: DPU_1... DPU_n) are connected to DPU_A as the computing power enhancement of other dedicated processing units, such as the computing power enhancement of GPU, ISP, etc. The embodiment of the present application does not make specific limitations on this.
[0108] Based on the above Figure 2A provided data processing architecture, and Figure 2B provided stacked package chip architecture, for the following application scenario where a general-purpose processor 011 calls a second dedicated processing unit 012 for operation, please refer to Appendix Figure 2E and Appendix Figure 2F , Appendix Figure 2EIt is a schematic diagram of a chip architecture of stacked packaging in another practical application provided by an embodiment of the present application. Attached Figure 2F is one provided by an embodiment of the present application above Figure 2E Schematic diagram of the interaction between the second dedicated processing unit and the first dedicated processing unit in the stacked packaging chip shown above. Among them, the above Figure 2E The provided stacked packaging chip architecture is used to support and execute the following method flow steps 1-step 6.
[0109] As Figure 2E shown, the stacked chip system includes a first chip 001 and a second chip 002. In this stacked packaging chip architecture, the second chip 002 is connected to the bus 00 in the first chip 001 through TSV interconnections. That is, the inter-chip interconnection is connected between the second dedicated processing unit 012 and the bus 00. Among them, the second dedicated processing unit DPU_A may include a memory (Buffer) 0121 and a task scheduler (Task Scheduler) 0122, and an AI algorithm operation module 0123 that directly schedules task sequences and data by the first dedicated processing unit DPU_1 (the AI algorithm operation module 0123 includes: Matrix Unit 01, Vector Unit 02, and Scalar Unit 03). Among them, Buffer0121 is used to store data and corresponding data processing tasks, Task Scheduler0122 is used to receive the assigned data processing tasks, and the AI algorithm operation module 0123 is used to execute data processing tasks. The first dedicated processing unit DPU_1 may also include: a memory (Buffer) 0211 and a task scheduler (Task Scheduler) 0212. For the corresponding functions, reference can be made to the relevant descriptions of the respective functional modules of the first dedicated processing unit DPU_1 above.
[0110] Specifically, DPU_A is directly controlled and scheduled by the CPU 011 of the first chip 001. DPU_A is directly connected to the bus 00 of the first chip 001. Therefore, the TSV interconnection signals between DPU_A and the first chip 001 generally use a standard protocol bus (Advanced eXtensible Interface, AXI) and a peripheral bus (Advanced Peripheral Bus, APB). Among them, the AXI bus is used for reading and writing data signals, and the APB bus is used for controlling signal configuration.
[0111] Among them, for the functions of other functional units in the data processing device described in the embodiments of the present application, reference can be made to the relevant descriptions of the stacked packaging chip architecture provided above Figures 2B - 2C and the relevant descriptions of the following method flow steps 1-step 6, which will not be elaborated here.
[0112] Among them, as Figure 2F the interaction schematic diagram of the second dedicated processing unit 012 and the first dedicated processing unit 021 in the stacked packaged chip system shown, the data processing method includes:
[0113] 1. As the general processor of the first chip 001, the CPU 011 initiates the start information of the DPU_A participating in the DPU_1 of the second chip 002 through the bus 00, and configures at least a part of the data processing tasks into the TaskScheduler0122 of the DPU_A unit;
[0114] 2. After receiving the start information, the DPU_A enters the start state from the waiting state, that is, it is awakened from the low-power state; the Task Scheduler0122 of the DPU_A receives the data processing tasks sent by the CPU 011, and transfers data from the memory 031 in the chip to the Buffer0121 inside the DPU_A through the bus 00;
[0115] 3. The Task Scheduler0122 of the DPU_A starts the AI algorithm operation module 0123 (including Matrix Unit 01, Vector Unit 02 and Scalar Unit 03) according to the data processing tasks, and the AI algorithm operation module 0123 reads the data in the buffer0121 and starts to execute the data processing tasks;
[0116] 4. The AI algorithm operation module 0123 stores the processed data into the Buffer0121 of the DPU_A;
[0117] 5. The DPU_A writes the processed data back to the memory 031 in the chip from the Buffer0121 through the bus 00;
[0118] 6. The Task Scheduler0122 of the DPU_A sends a processing completion signal back to the CPU 011 main control of the first chip 001 to complete the operation of this data processing task.
[0119] Optionally, in the embodiment of the present application, the second chip 002 is connected to the bus in the first chip 001 through the TSV interconnection line, that is, the data processing tasks of the target data can be executed by the CPU 011 through the DPU_1 alone, in parallel with the DPU_1 and the DPU_A, or by the DPU_A alone.
[0120] In the embodiment of the present application, the second dedicated processing unit DPU_A of the second chip 002 is a computing power enhancement module, which is directly scheduled and controlled by the CPU 011 in the first chip 001. In a conventional scenario, the computing power of DPU_1 in the first chip 001 is sufficient to meet the requirements. At this time, the DPU_A of the second chip 002 can be in a waiting state; or, only the DPU_A of the second chip 002 is in a startup state. When some scenarios, such as video processing, require high computing power for AI processing, the CPU can activate the DPU_A computing unit through the TSV control line to participate in the high-computing-power processing together.
[0121] Optionally, based on the Figure 2A provided data processing architecture, and Figure 2B the provided stacked packaging chip architecture, the first chip can also be stacked and packaged with multiple second chips into a chip system. Each second chip included in the multiple second chips can respectively match one or more first dedicated processing units in at least one first dedicated processing unit with the same or partially the same computing functions. For another example: the first chip can also be stacked and packaged with the second chip and the third chip respectively, where the third chip includes one or more of a memory, a power transmission circuit module, an input / output circuit module, or an analog module. Please refer to the attached Figure 2G , Figure 2G which is a schematic diagram of another stacked packaging chip architecture provided by the embodiment of the present application. As Figure 2G shown, outside the computing power units with stacked architectures, a memory 003 (Stacked Memory) can also be stacked, further bringing the storage closer to the computing units, improving the data processing bandwidth and efficiency of the overall system, and solving the problem of insufficient storage bandwidth of the stacked chips. Optionally, one or more of a power transmission circuit module, an input / output circuit module, or an analog module can also be stacked to achieve the purpose of separating and decoupling the analog and logical calculations of the chip, and continuing the increasingly high demands of chip evolution and business scenarios for chips.
[0122] It should be noted that Figure 2C and Figure 2E the stacked packaging chips in are only an exemplary implementation manner in the embodiment of the present application. The stacked packaging chip architecture in the embodiment of the present application includes but is not limited to the above-mentioned stacked packaging chip architectures.
[0123] It should also be noted that the stacked packaging chips can be configured in different data processing devices, corresponding to different main control forms in different data processing devices. The embodiment of the present application does not limit the form of the main control. For example, servers, laptops, smartphones, in-vehicle TVs, and so on.
[0124] Please refer to Figure 3 , Figure 3It is a schematic flowchart of a data processing method provided by an embodiment of the present application. This method can be applied to the structure of the data processing device described above Figure 2A and the stacked packaging chip architecture provided in Figure 2B . The data processing device may include the stacked packaging chip architecture provided in the above Figure 2B and is used to support and execute the Figure 3 method flow steps S301 - S302 shown in
[0125] Among them,
[0126] Step S301: Generate a data processing task through a general - purpose processor in the first chip. Specifically, the data processing device generates a data processing task through the general - purpose processor. The first chip includes the general - purpose processor, a bus, and at least one first dedicated processing unit (DPU). The general - purpose processor and the at least one first dedicated processing unit are connected to the bus. The first chip and the second chip are stacked and packaged into a chip architecture.
[0127] Optionally, the general - purpose processor includes a central processing unit (CPU).
[0128] Optionally, each of the one or more first dedicated processing units and the second dedicated processing unit in the second chip package includes at least one of a graphics processing unit (GPU), an image signal processor (ISP), a digital signal processor (DSP), or a neural network processing unit (NPU). For example, the GPU for graphics processing can be used to process graphics data in a smart terminal; the DSP for signal processing can be used to process digital signal data executed in a smart terminal; the neural network processing unit (NPU) can be used to process a large amount of multimedia data such as videos and images in a smart terminal. Therefore, the at least one first dedicated processing unit (DPU) can be respectively used to execute data processing tasks of different data types, so that the smart terminal can adapt to the increasing data processing requirements of users.
[0129] Step S302: Process at least a part of the data processing task through one or more first dedicated processing units in at least one first dedicated processing unit and at least one second dedicated processing unit in the second chip package. Specifically, the data processing device processes at least a part of the data processing task through one or more first dedicated processing units in at least one first dedicated processing unit and at least one second dedicated processing unit in the second chip package. The second dedicated processing unit has the same computing function as one or more first dedicated processing units in the at least one first dedicated processing unit. The first chip and the second chip are stacked and packaged and are interconnected through inter - chip interconnections.
[0130] Optionally, processing at least a part of the data processing task by one or more of the at least one first dedicated processing unit in the at least one first dedicated processing unit and the second dedicated processing unit in the second chip package includes: sending startup information to the second dedicated processing unit through the inter-chip interconnect by the general-purpose processor; in response to the startup information by the second dedicated processing unit, transitioning from a waiting state to a startup state and processing at least a part of the data processing task based on the computing function.
[0131] Optionally, sending startup information to the second dedicated processing unit through the inter-chip interconnect by the general-purpose processor includes: when the computing power of the one or more first dedicated processing units does not meet the requirements, sending startup information to the second dedicated processing unit through the inter-chip interconnect by the general-purpose processor.
[0132] Optionally, processing at least a part of the data processing task by one or more of the at least one first dedicated processing unit in the at least one first dedicated processing unit and the second dedicated processing unit in the second chip package includes: sending startup information to the second dedicated processing unit through the inter-chip interconnect by the one or more first dedicated processing units; the second dedicated processing unit is configured to, in response to the startup information, transition from a waiting state to a startup state and process at least a part of the data processing task based on the computing function.
[0133] Optionally, sending startup information to the second dedicated processing unit through the inter-chip interconnect by the one or more first dedicated processing units includes: when the computing power of the one or more first dedicated processing units does not meet the requirements, sending startup information to the second dedicated processing unit through the inter-chip interconnect by the one or more first dedicated processing units.
[0134] It should be noted that the relevant descriptions of steps S301 - S302 in the embodiments of the present application can also be correspondingly referred to the relevant descriptions of the above Figures 2A - 2G respective embodiments, which will not be elaborated here.
[0135] In the implementation of the embodiments of the present application, the stacked chip architecture in the data processing device can flexibly allocate data processing tasks to the first dedicated processing unit and the second dedicated processing unit respectively. When the first dedicated processing unit processes the data processing task alone, the second processing unit of the stacked second chip will be in a waiting state, and it will only change from the waiting state to the startup state when the second dedicated processing unit receives the startup information. Since the power consumption in the waiting state is lower than that in the startup state, the overall power consumption control of the stacked chip architecture is made more flexible and efficient, improving the energy efficiency of the chip system. Moreover, this stacked chip architecture can also, without increasing the overall volume, have the second dedicated processing unit assist the first dedicated processing unit in processing the data processing task, enhancing the computing power of the first dedicated processing unit, and largely alleviating and solving the demand for improving the chip computing power due to the rapid development of current algorithms, and avoiding the situation where the chip cannot complete the data processing task of the target data due to insufficient computing power of the first dedicated processing unit. Further, this computing power stacked chip architecture can also, in the case of the slowdown of Moore's Law and limited terminal chip area, continue the increasing demand for computing power in chip evolution and business scenarios through vertical computing power stacking (i.e., stacking the first chip and the second chip).
[0136] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0137] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the shown or discussed connections or communication connections to each other can be the communication connections of devices or units connected through some wires, conductors, interfaces, and can also be in electrical or other forms.
[0138] The units described as separate components above may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0139] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, can exist separately as individual physical units, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0140] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A chip architecture, characterized in that, it includes a first chip and a second chip in stacked packaging, and the first chip and the second chip are bare dies; the first chip includes a general-purpose processor, a bus, and at least one first dedicated processing unit (DPU), the general-purpose processor and the at least one first dedicated processing unit are connected to the bus, and the general-purpose processor is used to generate data processing tasks; the second chip includes a second dedicated processing unit, the second dedicated processing unit has at least partially the same computing function as one or more of the at least one first dedicated processing unit, and at least one of the one or more first dedicated processing units and the second dedicated processing unit can process at least a part of the data processing task based on the computing function; wherein, the first chip and the second chip are interconnected through inter-chip interconnect lines.
2. The chip architecture according to claim 1, characterized in that, the second dedicated processing unit has the same computing function as one or more of the at least one first dedicated processing unit.
3. The chip architecture according to claim 1 or 2, characterized in that, the general-purpose processor includes a central processing unit (CPU).
4. The chip architecture according to any one of claims 1-3, characterized in that, each of the one or more first dedicated processing units and the second dedicated processing unit includes at least one of a graphics processing unit (GPU), an image signal processor (ISP), a digital signal processor (DSP), or a neural network processing unit (NPU).
5. The chip architecture according to any one of claims 1-4, characterized in that, the inter-chip interconnect lines include at least one of through-silicon via (TSV) interconnect lines or wire bonding interconnect lines.
6. The chip architecture according to any one of claims 1-5, characterized in that, the chip architecture further includes a third chip, the third chip is stacked and packaged with the first chip and the second chip, and the third chip is connected to at least one of the first chip or the second chip through the inter-chip interconnect lines; the third chip includes at least one of a memory, a power transmission circuit module, an input / output circuit module, or an analog module.
7. The chip architecture according to any one of claims 1-6, characterized in that, the inter-chip interconnect lines are connected between the one or more first dedicated processing units and the second dedicated processing unit; the second dedicated processing unit is used to obtain at least a part of the data processing task from the one or more first dedicated processing units.
8. The chip architecture according to any one of claims 1-6, characterized in that, the inter-chip interconnect lines are connected between the second dedicated processing unit and the bus, and the second dedicated processing unit is used to obtain at least a part of the data processing task from the general-purpose processor through the bus.
9. The chip architecture according to any one of claims 1-8, characterized in that, The general-purpose processor is configured to send startup information to the second dedicated processing unit via the inter-chip interconnect line; The second dedicated processing unit is configured to, in response to the startup information, transition from a waiting state to a startup state and process at least a part of the data processing task based on the computing function.
10. The chip architecture according to claim 9, wherein, the general-purpose processor is configured to send startup information to the second dedicated processing unit via the inter-chip interconnect line when the computing power of the one or more first dedicated processing units does not meet the requirements.
11. The chip architecture according to any one of claims 1-8, wherein, the one or more first dedicated processing units are configured to send startup information to the second dedicated processing unit via the inter-chip interconnect line; the second dedicated processing unit is configured to, in response to the startup information, transition from a waiting state to a startup state and process at least a part of the data processing task based on the computing function.
12. The chip architecture according to claim 11, wherein, the one or more first dedicated processing units are configured to send startup information to the second dedicated processing unit via the inter-chip interconnect line when the computing power of the one or more first dedicated processing units does not meet the requirements.
13. A data processing method, wherein, it includes: generating a data processing task by a general-purpose processor in a first chip of a chip architecture, the first chip including the general-purpose processor, a bus, and at least one first dedicated processing unit (DPU), the general-purpose processor and the at least one first dedicated processing unit being connected to the bus; processing at least a part of the data processing task by at least one of the one or more first dedicated processing units in the at least one first dedicated processing unit and at least one of the second dedicated processing units in a second chip package of the chip architecture, the second dedicated processing unit having at least partially the same computing function as the one or more first dedicated processing units in the at least one first dedicated processing unit, the first chip and the second chip being stacked and packaged and interconnected via an inter-chip interconnect line, and the first chip and the second chip being bare dies.