Semiconductor device

The semiconductor device addresses high heat and power consumption issues by layering memory and arithmetic circuits with OS and Si transistors, reducing area and parasitic capacitance, enabling efficient AI calculations.

JP2025118903AInactive Publication Date: 2025-08-13SEMICON ENERGY LAB CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025082361
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-03-13
Filing Date
2025-05-16
Publication Date
2025-08-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing semiconductor devices face challenges in high heat generation, power consumption, and increased area due to the need for high-frequency data transfer between memory cell arrays and arithmetic circuits, which is exacerbated by the integration of transistors using bonding techniques that increase parasitic capacitance.

Method used

A semiconductor device is designed with memory circuits in a first layer and switching and arithmetic circuits in a second layer, utilizing OS transistors with metal oxides in the first layer and Si transistors in the second layer, stacked vertically to reduce area overhead and parasitic capacitance, and incorporating NOSRAM for non-destructive readout and low leakage current.

Benefits of technology

This configuration results in a miniaturized, low-power consumption device with improved arithmetic processing speed and efficient data transfer, suitable for AI calculations and neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025118903000001_ABST
    Figure 2025118903000001_ABST
Patent Text Reader

Abstract

To provide a semiconductor device with a novel structure.SOLUTION: A plurality of memory circuits, a switching circuit, and a calculation circuit are provided. The respective memory circuits have a function of holding weight data and a function of outputting the weight data to first wires. The switching circuit has a function of switching a conduction state of any one of the first wires and a second wire. The calculation circuit has a function of performing a calculation process using input data and the weight data applied to the second wire. The memory circuit is provided in a first layer including a first transistor. The switching circuit and the calculation circuit are provided in a second layer including a second transistor. The first layer is provided in a layer different from the second layer.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This specification describes semiconductor devices and the like.

[0002] Note that one embodiment of the present invention is not limited to the above technical field. Examples of the technical field of one embodiment of the present invention disclosed in this specification and the like include a semiconductor device, an imaging device, a display device, a light-emitting device, a power storage device, a memory device, a display system, an electronic device, a lighting device, an input device, an input / output device, a driving method thereof, or a manufacturing method thereof. [Background technology]

[0003] Electronic devices having semiconductor devices including a CPU (Central Processing Unit) and the like are becoming widespread. To process large amounts of data at high speed, such electronic devices are subject to active technological development aimed at improving the performance of semiconductor devices. One example of a technology that achieves high performance is the so-called SoC (System on Chip) technology, which tightly couples an accelerator such as a GPU (Graphics Processing Unit) with a CPU. With semiconductor devices that achieve high performance through SoC technology, increased heat generation and power consumption become problems.

[0004] In AI (Artificial Intelligence) technology, the amount of calculations and the number of parameters become enormous, resulting in an increase in the amount of calculations. Since an increase in the amount of calculations leads to an increase in heat generation and power consumption, architectures for reducing the amount of calculations have been actively proposed. Typical architectures include the Binary Neural Network (BNN) and the Ternary Neural Network (TNN), which are particularly effective in reducing circuit scale and power consumption (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0005] [Patent Document 1] International Publication No. 2019 / 078924 Summary of the Invention [Problem to be solved by the invention]

[0006] AI technology calculations require a huge number of repeated product-sum operations using weight data and input data, which requires faster calculation processing. Memory cell arrays must store large amounts of weight data and intermediate data. These memory cell arrays store large amounts of weight data and intermediate data, which are then read out to the calculation circuit via bit lines. As the frequency of reading weight data and intermediate data increases, the bandwidth between the memory cell array and the calculation circuit can become a limiting factor in operating speed.

[0007] Increasing the number of parallel interconnections between the memory cell array and the arithmetic circuit allows the memory cell array and the arithmetic circuit to be connected with a high bandwidth, which is advantageous for speeding up arithmetic processing. However, since the number of interconnections between the arithmetic circuit and the memory cell array increases, there is a risk that the area of the peripheral circuit will increase significantly.

[0008] Furthermore, in AI technology calculations, how to reduce the charge and discharge energy of bit lines is important in achieving low power consumption.

[0009] Shortening the bit lines is an effective way to reduce the charge / discharge energy of the bit lines. However, this requires alternating arrangement of arithmetic circuits and memory cell arrays, which can significantly increase the area of the peripheral circuits. Another technique for shortening the bit lines is to integrate transistors vertically using bonding techniques. However, bonding techniques require large spacing between the electrical connections, which can increase parasitic capacitance and make it difficult to reduce the charge / discharge energy.

[0010] An object of one embodiment of the present invention is to provide a miniaturized semiconductor device.An object of one embodiment of the present invention is to provide a semiconductor device with low power consumption.An object of one embodiment of the present invention is to provide a semiconductor device with improved arithmetic processing speed.An object of one embodiment of the present invention is to provide a semiconductor device with a novel structure.

[0011] Note that one embodiment of the present invention does not necessarily have to solve all of the above problems, but it is sufficient that it can solve at least one of the problems. Furthermore, the description of the above problems does not preclude the existence of other problems. Problems other than these will become apparent from the description in the specification, claims, drawings, etc., and other problems can be extracted from the description in the specification, claims, drawings, etc. [Means for solving the problem]

[0012] One embodiment of the present invention is a semiconductor device including a plurality of memory circuits, a switching circuit, and an arithmetic circuit, each of which has a function of holding weight data, and which switches conduction between one of the memory circuits and the arithmetic circuit, the plurality of memory circuits being provided in a first layer, and the switching circuit and the arithmetic circuit being provided in a second layer, the first layer being a layer different from the second layer.

[0013] One embodiment of the present invention is a semiconductor device including a plurality of memory circuits, a switching circuit, and an arithmetic circuit, each of which has a function of holding weight data and a function of outputting the weight data to a first wiring, and the switching circuit has a function of switching conduction between any one of the plurality of first wirings and the arithmetic circuit, the plurality of memory circuits being provided in a first layer, and the switching circuit and the arithmetic circuit being provided in a second layer, the first layer being a layer different from the second layer.

[0014] One embodiment of the present invention is a semiconductor device including a plurality of memory circuits, a switching circuit, and an arithmetic circuit, each of which has a function of holding weight data and a function of outputting the weight data to a first wiring, a function of switching conduction between one of the plurality of first wirings and a second wiring, and a function of performing arithmetic processing using input data and weight data provided to the second wiring, the plurality of memory circuits being provided in a first layer, and the switching circuit and the arithmetic circuit being provided in a second layer, the first layer being a layer different from the second layer.

[0015] In one aspect of the present invention, the semiconductor device preferably has the second wiring provided approximately parallel to the surface of the substrate.

[0016] In one aspect of the present invention, the semiconductor device preferably has the first wiring provided approximately perpendicular to the surface of the substrate.

[0017] In one embodiment of the present invention, the semiconductor device preferably includes a first transistor in the first layer, and the first transistor includes a semiconductor layer having a metal oxide in a channel formation region.

[0018] In one embodiment of the present invention, the metal oxide preferably contains In, Ga, and Zn.

[0019] In one embodiment of the present invention, the second layer preferably includes a second transistor, and the second transistor preferably includes a semiconductor layer having silicon in a channel formation region.

[0020] In one embodiment of the present invention, the semiconductor device is preferably a circuit that performs a product-sum operation as the arithmetic circuit.

[0021] In one embodiment of the present invention, the semiconductor device is preferably such that the first layer is stacked on the second layer.

[0022] In one embodiment of the present invention, the weight data is data of a first number of bits, and the weight data is data obtained by converting weight data of a second number of bits optimized using learning data, and the first number of bits is smaller than the second number of bits.

[0023] Other aspects of the present invention will be described in the following embodiments and in the drawings. [Effects of the Invention]

[0024] According to one embodiment of the present invention, a miniaturized semiconductor device can be provided. Alternatively, according to one embodiment of the present invention, a semiconductor device with low power consumption can be provided. Alternatively, according to one embodiment of the present invention, a semiconductor device with improved processing speed can be provided. Alternatively, a semiconductor device with a novel structure can be provided.

[0025] The description of multiple effects does not preclude the existence of other effects. Furthermore, one embodiment of the present invention does not necessarily have all of the exemplified effects. Furthermore, problems, effects, and novel features of one embodiment of the present invention other than those described above will become apparent from the description and drawings of this specification. [Brief explanation of the drawings]

[0026] [Figure 1] 1A and 1B are diagrams illustrating an example of the configuration of a semiconductor device. [Figure 2] 2A and 2B are diagrams illustrating an example of the configuration of a semiconductor device. [Figure 3] 3A and 3B are diagrams illustrating an example of the configuration of a semiconductor device. [Figure 4] FIG. 4 is a diagram illustrating an example of the configuration of a semiconductor device. [Figure 5] 5A and 5B are diagrams illustrating an example of the configuration of a semiconductor device. [Figure 6] FIG. 6 is a diagram illustrating an example of the configuration of a semiconductor device. [Figure 7]7A and 7B are diagrams illustrating an example of the configuration of a semiconductor device. [Figure 8] 8A and 8B are diagrams illustrating a configuration example of a semiconductor device. [Figure 9] 9A, 9B, and 9C are diagrams illustrating configuration examples of a semiconductor device. [Figure 10] FIG. 10 is a diagram illustrating an example of the configuration of a semiconductor device. [Figure 11] FIG. 11 is a diagram illustrating an example of the configuration of a semiconductor device. [Figure 12] 12A and 12B are diagrams illustrating a configuration example of a semiconductor device. [Figure 13] 13A and 13B are diagrams illustrating a configuration example of a semiconductor device. [Figure 14] 14A and 14B are diagrams showing configuration examples of an integrated circuit. [Figure 15] FIG. 15 is a diagram illustrating an example of the configuration of a transistor. [Figure 16] FIG. 16 is a diagram illustrating an example of the configuration of a processing system. [Figure 17] FIG. 17 is a diagram illustrating an example of the configuration of a CPU. [Figure 18] 18A and 18B are diagrams illustrating an example of the configuration of a CPU. [Figure 19] FIG. 19 illustrates an example of the configuration of a CPU. [Figure 20] FIG. 20 is a diagram illustrating an example of the configuration of a transistor. [Figure 21] 21A and 21B are diagrams showing examples of the configuration of a transistor. [Figure 22] 22A and 22B are diagrams illustrating an example of the configuration of an integrated circuit. [Figure 23] 23A and 23B are diagrams illustrating an application example of an integrated circuit. [Figure 24] 24A and 24B are diagrams illustrating an application example of an integrated circuit. [Figure 25]25A, 25B and 25C are diagrams illustrating an application example of an integrated circuit. [Figure 26] FIG. 26 is a diagram illustrating an application example of an integrated circuit. [Figure 27] 27A and 27B are diagrams illustrating an application example of an integrated circuit. [Figure 28] 28A and 28B are diagrams illustrating weight data. DETAILED DESCRIPTION OF THE INVENTION

[0027] The following describes an embodiment of the present invention. However, one embodiment of the present invention is not limited to the following description, and it will be readily understood by those skilled in the art that various changes in form and details can be made without departing from the spirit and scope of the present invention. Therefore, one embodiment of the present invention should not be interpreted as being limited to the description of the embodiment shown below.

[0028] In this specification, the ordinal numbers "first," "second," and "third" are used to avoid confusion between components. Therefore, they do not limit the number of components. Furthermore, they do not limit the order of the components. For example, a component referred to as "first" in one embodiment of this specification may be a component referred to as "second" in another embodiment or in the claims. For example, a component referred to as "first" in one embodiment of this specification may be omitted in another embodiment or in the claims.

[0029] In the drawings, the same elements or elements having similar functions, elements made of the same material, or elements formed at the same time may be given the same reference numerals, and repeated description thereof may be omitted.

[0030] In this specification, for example, the power supply potential VDD may be abbreviated to potential VDD, VDD, etc. This also applies to other components (for example, signals, voltages, circuits, elements, electrodes, wiring, etc.).

[0031] Furthermore, when the same symbol is used for multiple elements, and particularly when it is necessary to distinguish between them, the symbol may be accompanied by an identifying symbol such as "_1", "_2", "[n]", or "[m,n]". For example, the second wiring GL is written as wiring GL[2].

[0032] (Embodiment 1) The structure, operation, and the like of a semiconductor device according to one embodiment of the present invention will be described.

[0033] In this specification and the like, a semiconductor device refers to any device that can function by utilizing semiconductor characteristics. Semiconductor elements such as transistors, semiconductor circuits, arithmetic devices, and memory devices are all embodiments of semiconductor devices. Display devices (liquid crystal display devices, light-emitting display devices, etc.), projection devices, lighting devices, electro-optical devices, power storage devices, memory devices, semiconductor circuits, imaging devices, electronic devices, and the like may be considered to include semiconductor devices.

[0034] FIG. 1A is a diagram illustrating a semiconductor device 10 according to one aspect of the present invention.

[0035] The semiconductor device 10 functions as an accelerator that executes a program (also called a kernel or kernel program) called from a host program. The semiconductor device 10 can perform, for example, parallel processing of matrix operations in graphic processing, parallel processing of product-sum operations in neural networks, and parallel processing of floating-point operations in scientific and technological calculations.

[0036] The semiconductor device 10 includes a memory circuit unit 20 (also referred to as a memory cell array), an arithmetic circuit 30, and a switching circuit 40. The arithmetic circuit 30 and the switching circuit 40 are provided in a layer 11 having transistors in the xy plane in the drawing. The memory circuit unit 20 is provided in a layer 12 having transistors in the xy plane in the drawing.

[0037] Layer 11 includes a transistor having silicon in a channel formation region (a Si transistor). Layer 12 includes a transistor having an oxide semiconductor in a channel formation region (an OS transistor). Layer 11 and layer 12 are provided in different layers in a direction substantially perpendicular to the xy plane (the z direction in FIG. 1A ).

[0038] Alternatively, layer 12 may have a Si transistor. In this case, layers 11 and 12 can be provided on different layers in a direction substantially perpendicular to the xy plane (z direction in FIG. 1A) by using a bonding technique or the like. Examples of bonding techniques that can be used include plasma activated bonding and Cu-Cu bonding, which are techniques for bonding semiconductor substrates.

[0039] When the layer 12 is composed of OS transistors, the memory circuit unit 20 can be stacked with the arithmetic circuit 30 and the switching circuit 40, which can be composed of Si transistors. That is, the memory circuit unit 20 is provided on the substrate on which the arithmetic circuit 30 and the switching circuit 40 are provided. Therefore, the memory circuit unit 20 can be arranged without increasing the circuit area. By providing the memory circuit unit 20 on the substrate on which the arithmetic circuit 30 and the switching circuit 40 are provided, the memory capacity required for the arithmetic processing in the semiconductor device 10 functioning as an accelerator can be increased compared to when the memory circuit unit 20, the arithmetic circuit 30, and the switching circuit 40 are arranged on the same layer. The increased memory capacity reduces the number of times data required for the arithmetic processing is transferred from an external storage device to the semiconductor device, thereby reducing power consumption.

[0040] The memory circuit unit 20 is illustrated as an example of a plurality of memory circuit units 20_1 to 20_4. Each memory circuit unit has a plurality of memory circuits 21. In each of the memory circuit units 20_1 to 20_4, the plurality of memory circuits 21 are connected to the switching circuit 40 via wirings LBL_1 to LBL_4 (also referred to as local bit lines or read bit lines) as illustrated in FIG. 1A.

[0041] The memory circuit 21 may have a NOSRAM circuit configuration. "NOSRAM (registered trademark)" is an abbreviation for "Nonvolatile Oxide Semiconductor RAM." NOSRAM refers to a memory in which memory cells are two-transistor (2T) or three-transistor (3T) gain cells and access transistors are OS transistors. The memory circuit 21 is a memory configured with OS transistors. The layer 12 including the memory circuit 21 may be stacked on the layer 11 including the arithmetic circuit 30 and the switching circuit 40. Since the memory circuit unit 20 including the memory circuit 21 is provided on the layer 11 including the arithmetic circuit 30 and the switching circuit 40, it is possible to reduce the area overhead caused by having the memory circuit unit 20.

[0042] Furthermore, OS transistors have an extremely small leakage current, the current that flows between the source and drain when they are off. NOSRAM can be used as nonvolatile memory by utilizing its extremely small leakage current characteristics to retain a charge corresponding to the data within the memory circuit. NOSRAM is particularly suitable for parallel processing of product-sum operations in neural networks, which require a large number of repeated data read operations, because it can read the stored data without destroying it (non-destructive readout).

[0043] The memory circuit 21 is preferably a memory having an OS transistor such as NOSRAM or DOSRAM (hereinafter also referred to as OS memory). Since the band gap of a metal oxide functioning as an oxide semiconductor is 2.5 eV or more, the OS transistor has a very small off-state current. For example, when the voltage between the source and drain is 3.5 V and the temperature is room temperature (25° C.), the off-state current per 1 μm of channel width is 1×10 -20 Less than A, 1 x 10 -22 Less than A or 1 x 10 -24Therefore, the amount of charge leaked from the retention node of the OS memory via the OS transistor is extremely small. Therefore, the OS memory can function as a nonvolatile memory circuit, which enables power gating of the semiconductor device 10.

[0044] Semiconductor devices with highly integrated transistors may generate heat due to circuit operation. This heat increases the temperature of the transistor, which can change the transistor's characteristics, resulting in changes in field-effect mobility and a decrease in operating frequency. OS transistors have higher heat resistance than Si transistors, making them less susceptible to temperature changes in field-effect mobility and a decrease in operating frequency. Furthermore, OS transistors tend to maintain the characteristic that their drain current increases exponentially with respect to the gate-source voltage, even at high temperatures. Therefore, OS transistors enable stable operation in high-temperature environments.

[0045] Metal oxides suitable for OS transistors include Zn oxide, Zn-Sn oxide, Ga-Sn oxide, In-Ga oxide, In-Zn oxide, and In-M-Zn oxide (where M is Ti, Ga, Y, Zr, La, Ce, Nd, Sn, or Hf). Metal oxides using Ga as M are particularly preferred for OS transistors because they can provide transistors with excellent electrical properties, such as field-effect mobility, by adjusting the ratio of elements. Furthermore, the oxide containing indium and zinc may contain one or more elements selected from the group consisting of aluminum, gallium, yttrium, copper, vanadium, beryllium, boron, silicon, titanium, iron, nickel, germanium, zirconium, molybdenum, lanthanum, cerium, neodymium, hafnium, tantalum, tungsten, and magnesium.

[0046] To improve the reliability and electrical characteristics of OS transistors, the metal oxide used in the semiconductor layer is preferably a metal oxide having a crystalline portion, such as CAAC-OS, CAC-OS, or nc-OS. CAAC-OS is an abbreviation for c-axis-aligned crystalline oxide semiconductor. CAC-OS is an abbreviation for cloud-aligned composite oxide semiconductor. nc-OS is an abbreviation for nanocrystalline oxide semiconductor.

[0047] CAAC-OS has a c-axis orientation and a distorted crystal structure in which multiple nanocrystals are connected in the ab-plane direction. The distorted crystal structure refers to the change in the lattice orientation between regions with a uniform lattice arrangement and regions with a different uniform lattice arrangement in the regions where multiple nanocrystals are connected.

[0048] CAC-OS has the function of both allowing electrons (or holes) to flow and preventing electrons from flowing. By separating the electron flow function from the electron blocking function, both functions can be maximized. In other words, using CAC-OS in the channel formation region of an OS transistor can achieve both a high on-state current and an extremely low off-state current.

[0049] Metal oxides have a wide band gap, which makes it difficult for electrons to be excited, and they have a large effective mass for holes. This means that OS transistors are less susceptible to avalanche breakdown and other problems than typical Si transistors. Therefore, for example, hot carrier degradation caused by avalanche breakdown can be suppressed. Suppressing hot carrier degradation allows OS transistors to be driven at a high drain voltage.

[0050] OS transistors are accumulation-type transistors that use electrons as majority carriers. Therefore, they are less susceptible to drain-induced barrier lowering (DIBL), a short-channel effect, compared to inversion-type transistors (typically, Si transistors) with pn junctions. In other words, OS transistors have higher resistance to short-channel effects than Si transistors.

[0051] Because OS transistors have high resistance to short-channel effects, their channel length can be reduced without degrading their reliability, allowing for increased circuit integration. As the channel length decreases, the drain electric field becomes stronger, but as mentioned above, OS transistors are less susceptible to avalanche breakdown than Si transistors.

[0052] Furthermore, because OS transistors have high resistance to short-channel effects, their gate insulating films can be thicker than those of Si transistors. For example, even for miniaturized transistors with channel lengths and widths of 50 nm or less, it may be possible to provide a gate insulating film as thick as about 10 nm. By increasing the gate insulating film thickness, parasitic capacitance can be reduced, thereby improving the operating speed of the circuit. Furthermore, by increasing the gate insulating film thickness, leakage current through the gate insulating film can be reduced, leading to a reduction in static current consumption.

[0053] As described above, the semiconductor device 10 can retain data even when the supply of power supply voltage is stopped by including the memory circuit 21, which is an OS memory. This enables power gating of the semiconductor device 10, thereby enabling a significant reduction in power consumption.

[0054] The data stored in the memory circuit 21 is data (weight data) corresponding to weight parameters used in product-sum operations of a neural network. By using digital weight data, the semiconductor device can be made noise-resistant and capable of high-speed operations. The weight data may also be analog data. Since NOSRAM can hold analog potentials, the data can be appropriately converted to digital data and used. The memory circuit 21, which can hold analog data, can hold weight data with a high number of bits without increasing the number of memory circuits.

[0055] The switching circuits 40_1 to 40_4 shown in the figure as an example of the switching circuit 40 have a function of selecting the potentials of the wirings LBL_1 to LBL_4 extending from the memory circuit units 20_1 to 20_4, respectively, and transmitting the selected potentials to the wiring GBL (also referred to as a global bit line). The wiring GBL is connected to the output terminals of the switching circuits 40_1 to 40_4. The switching circuits 40 need to prevent the output potentials of the selected switching circuit 40 and the unselected switching circuits 40 from being simultaneously supplied with each other, thereby preventing a through current from occurring. The switching circuit 40 can be, for example, a three-state buffer whose output potential state is controlled by a control signal. In this configuration example, the selected switching circuit buffers and outputs the input potential to the wiring GBL, and the outputs of the unselected switching circuits have high impedance, preventing the output potentials from being simultaneously supplied. Note that the switching circuit 40 is preferably configured with a Si transistor. This configuration allows for high-speed switching of the connection state.

[0056] The arithmetic circuits 30_1 to 30_4 shown in the figure as examples of the arithmetic circuit 30 have the function of repeatedly executing the same processing, such as a multiply-and-accumulate operation. The input data and weight data input for the multiply-and-accumulate operation in the arithmetic circuit 30 are preferably digital data. Digital data is less susceptible to noise. Therefore, the arithmetic circuit 30 is suitable for performing arithmetic processing that requires highly accurate calculation results. The arithmetic circuit 30 is preferably composed of Si transistors. This configuration allows the arithmetic circuit 30 to be stacked with OS transistors.

[0057] The arithmetic circuits 30_1 to 30_4 are supplied with weight data held in the memory circuit 21 via wirings LBL_1 to LBL_4 and wiring GBL. Furthermore, the arithmetic circuits 30_1 to 30_4 are supplied with input data (A1, A2, A3, A4) input from the outside. The arithmetic circuits 30_1 to 30_4 perform product-sum arithmetic processing using the weight data held in the memory circuit 21 and the input data input from the outside.

[0058] The weight data provided to the arithmetic circuits 30_1 to 30_4 is weight data selected by the plurality of memory circuit units 20_1 to 20_4, switched by the switching circuits 40_1 to 40_4, and provided via the wiring GBL. That is, the arithmetic circuits 30_1 to 30_4 can perform arithmetic processing using the same weight data, such as a product-sum operation. Therefore, the semiconductor device 10 according to one embodiment of the present invention can efficiently perform processing using the same weight data, such as a convolutional neural network.

[0059] Furthermore, the weight data provided to the arithmetic circuits 30_1 to 30_4 can be provided to the wiring GBL by switching data previously provided to the wirings LBL_1 to LBL_4 using the switching circuits 40_1 to 40_4, so that the weight data provided to the wiring GBL can be switched at a speed corresponding to the electrical characteristics of Si transistors. Therefore, even if it takes a long time to read the weight data from the memory circuit portions 20_1 to 20_4 to the wirings LBL_1 to LBL_4, by reading the weight data to the wirings LBL_1 to LBL_4 in advance, the weight data can be switched at high speed for arithmetic processing.

[0060] The wiring LBL extending from the memory circuit unit 20 to the switching circuit 40 is connected to the weight data W data This is the wiring for transmitting the weight data W from the layer 12 to the wiring LBL. dataIn order to read the data at high speed, it is preferable to shorten the wiring LBL. Also, it is preferable to shorten the wiring LBL in order to reduce the energy consumption associated with charging and discharging. In other words, it is preferable to dispose the switching circuits 40 in a dispersed manner in the xy plane of the layer 11 so as to be close to the wiring LBL (indicated by the arrows extending in the z direction in the drawing) that extend in the z direction.

[0061] The arithmetic circuits 30_1 to 30_4 can be configured to be provided for each of the wirings LBL_1 to LBL_4, which are read bit lines of the memory circuit 21, i.e., for each column (column-parallel calculation). This configuration allows parallel calculations of data for the number of columns of wiring LBL. Compared to multiply-and-accumulate operations using a CPU or GPU, column-parallel calculations are not limited by data bus size (e.g., 32 bits). This significantly increases the parallelism of calculations, thereby improving the efficiency of massive calculations, such as deep neural network learning (AI technology) and scientific and engineering calculations using floating-point arithmetic. Furthermore, because the calculations on data output from the arithmetic circuit 30 can be completed and read, the power consumed by memory access (e.g., data transfer between the arithmetic circuit and memory) can be reduced, suppressing increases in heat generation and power consumption. Furthermore, by shortening the physical distance between the arithmetic circuit 30 and the memory circuit section 20, for example by stacking the layers, the wiring distance can be shortened, thereby reducing the parasitic capacitance generated in the signal lines and enabling lower power consumption.

[0062] Next, referring to FIG. 2A, a block diagram showing the entire processing system 100 including the semiconductor device 10 that functions as an AI accelerator will be described.

[0063] 2A illustrates the semiconductor device 10 described in FIGS. 1A and 1B, as well as a CPU 110 and a bus 120. The CPU 110 has a CPU core 200 and a backup circuit 222. The semiconductor device 10 functioning as an accelerator illustrates a drive circuit 50, memory circuit units 20_1 to 20_N (N is a natural number of 2 or more), a memory circuit 21, a switching circuit 40, and arithmetic circuits 30_1 to 30_N.

[0064] The CPU 110 has the function of performing general-purpose processing, such as running an operating system, controlling data, and executing various calculations and programs. The CPU 110 has a CPU core 200. The CPU core 200 corresponds to one or more CPU cores. The CPU 110 also has a backup circuit 222 that can retain data in the CPU core 200 even if the supply of power voltage is stopped. The supply of power voltage can be controlled by electrically disconnecting it from the power domain using a power switch or the like. The power voltage is sometimes called a drive voltage. An OS memory having an OS transistor, for example, is suitable as the backup circuit 222.

[0065] The backup circuit 222, which is configured with OS transistors, can be stacked on the CPU core 200, which can be configured with Si transistors. Because the area of the backup circuit 222 is smaller than the area of the CPU core 200, the backup circuit 222 can be placed on the CPU core 200 without increasing the circuit area. The backup circuit 222 has a function of retaining data in the registers of the CPU core 200. The backup circuit 222 is also referred to as a data retention circuit. Details of the configuration of the CPU core 200 including the backup circuit 222 with OS transistors will also be described in the fourth embodiment.

[0066] The memory circuit units 20_1 to 20_N respectively store weight data W1 to W2 held in the memory circuit 21. NThe switching circuit 40 outputs the selected weight data to the weight data W via a line LBL (not shown). SEL The driving circuit 50 outputs the input data A1 to A2 to the arithmetic circuits 30_1 to 30_N via the input data lines. N Output.

[0067] The driving circuit 50 has a function of outputting a signal for controlling the writing and reading of weight data in the memory circuit units 20_1 to 20_N. The driving circuit 50 also has a function of providing input data to the arithmetic circuits 30_1 to 30_N to cause them to execute product-sum operations of the neural network, and a function of holding output data obtained by the product-sum operations of the neural network.

[0068] The bus 120 electrically connects the CPU 110 and the semiconductor device 10. That is, the CPU 110 and the semiconductor device 10 can transmit data via the bus 120.

[0069] FIG. 2B is a diagram for explaining the positional relationship of each component in the semiconductor device 10 shown in FIG. 2A when N is 6.

[0070] The memory circuit units 20_1 to 20_6, which are configured with OS transistors, and the arithmetic circuits 30_1 to 30_N are electrically connected via wirings LBL_1 to LBL_6 that extend in a direction substantially perpendicular to the surface of the substrate on which the drive circuit 50, the switching circuit 40, and the arithmetic circuits 30_1 to 30_6 are provided. Note that "substantially perpendicular" refers to a state in which they are arranged at an angle of 85 degrees or more and 95 degrees or less. Note that in this specification, the X, Y, and Z directions shown in Figure 2B and other figures are orthogonal or intersect with each other. Furthermore, the X and Y directions are parallel or substantially parallel to the substrate surface, and the Z direction is perpendicular or substantially perpendicular to the substrate surface.

[0071] Each of the memory circuit units 20_1 to 20_6 includes a memory circuit 21. The memory circuit units 20_1 to 20_6 may be referred to as device memories or shared memories. The memory circuit 21 includes a transistor 22. The semiconductor layer 23 of the transistor 22 may be an oxide semiconductor (metal oxide), thereby forming the memory circuit 21 using the OS transistor described above.

[0072] The memory circuits 21 included in the memory circuit units 20_1 to 20_6 are connected to wirings LBL_1 to LBL_6, respectively. The wirings LBL_1 to LBL_6 are connected to the switching circuit 40 via wirings extending substantially perpendicular to the surface of the substrate on which the Si transistors are provided, i.e., in the z-direction. The switching circuit 40 is configured to amplify the potential of one of the wirings LBL_1 to LBL_6 and transmit the amplified potential to the wiring GBL. The wiring GBL is a wiring extending substantially parallel to the surface of the substrate on which the Si transistors are provided, i.e., in the xy plane. With this configuration, the weight data to be provided to the wiring GBL can be switched at high speed by controlling the switching circuit 40.

[0073] The arithmetic circuits 30_1 to 30_6 receive weight data input via the wiring GBL and input data A given from the driving circuit 50 via the input data line. IN and performs calculations based on the weight data. Since the memory circuit units 20_1 to 20_6 that hold the weight data can be arranged in the upper layer, the calculation circuits 30_1 to 30_6 can be arranged efficiently. Therefore, the input data lines extending from the drive circuit 50 can be shortened, and the semiconductor device 10 can achieve low power consumption and high speed.

[0074] Next, the advantages of the configuration of FIG. 2B will be described. For the sake of explanation, FIG. 3A shows each configuration of FIG. 2B in a block diagram. Note that the following description will be given assuming that weight data W1 to W6 are read out to wirings LBL_1 to LBL_6 from memory circuits 21 in six memory circuit units 20_1 to 20_6. The switching circuit 40 will be described as switching circuits 40_1 to 40_6 connected to wirings LBL_1 to LBL_6. The weight data selected from the weight data W1 to W6 by the switching circuit 40 and provided to wirings GBL will be referred to as weight data W1 to W6. SEL The following description will be given assuming that input data A1 to A6 are given to the arithmetic circuits 30_1 to 30_6, respectively, and output data MAC1 to MAC6 are obtained.

[0075] The wiring LBL_1 to LBL_6 extends in the vertical direction (see FIG. 2B) connecting the upper and lower layers. P are shorter than the wires extending in the horizontal direction. Therefore, the parasitic capacitance of the wires LBL_1 to LBL_6 can be reduced, the charge required for charging and discharging the wires can be reduced, and power consumption can be reduced and calculation efficiency can be improved. In addition, reading from the memory circuit 21 to the wires LBL_1 to LBL_6 can be performed at high speed.

[0076] The arithmetic circuits 30_1 to 30_6 can perform arithmetic processing using the same weight data via the wiring GBL. This configuration is suitable for arithmetic processing of a convolutional neural network that performs arithmetic processing using the same weight data.

[0077] 3B shows an example of a circuit configuration applicable to the switching circuit 40 shown in FIG. 3A. The three-state buffer shown in FIG. 3B has a function of amplifying and transmitting the potential of the line LBL to the line GBL in response to a control signal EN. The switching circuit 40 can be considered as a multiplexer. It has a function of selecting one of multiple input signals.

[0078] In FIG. 3A, the switching circuit 40 selects one wiring from the plurality of wirings LBL and outputs the weight data W SELHowever, other configurations may be used. For example, as shown in FIG. 4, a configuration may be adopted in which a switching circuit 40A and a switching circuit 40B are provided as the switching circuits.

[0079] The switching circuit 40A includes switching circuits 40_1 to 40_12. The configuration of the switching circuit 40A is similar to that of the switching circuit 40. The switching circuits 40_1 to 40_6 and the switching circuits 40_7 to 40_12 may be located apart from each other. The switching circuit 40A selects one of the wirings LBL_1 to LBL_6 and outputs weight data W selected from the weight data W1 to W6. SEL_A The switching circuit 40A selects one of the lines LBL_7 to LBL_12 and supplies the weight data W7 to W8. 12 Weight data W selected from SEL_B is applied to wiring GBL_B.

[0080] The switching circuit 40B includes switching circuits 40X to 40Y. The configuration of the switching circuit 40B is similar to that of the switching circuit 40. The switching circuit 40B selects the wiring GBL_A or the wiring GBL_B and outputs the weight data W SEL_A or weight data W SEL_B Weight data W selected from SEL is provided to the wiring GBL. Through the wiring GBL, the arithmetic circuits 30_1 to 30_6 and the arithmetic circuits 30_7 to 30_12 can perform arithmetic processing using the same weight data. This configuration is suitable for arithmetic processing of a convolutional neural network that performs arithmetic processing using the same weight data.

[0081] 3A has been described as a configuration in which each memory circuit 21 holds one bit of data (i.e., data of '1' or '0') and performs arithmetic processing using that data, but one embodiment of the present invention can also be applied to a configuration in which arithmetic processing is performed using multi-bit data. Similar to FIG. 3A, this configuration is illustrated in FIG. 5A. In the case of multi-bit (e.g., n-bit) data, as illustrated in FIG. 5A, a configuration may be adopted in which multi-bit weight data to be provided to wiring GBL is selected using a switching circuit 40M connected to wirings LBL_1 to LBL_n, the number of which corresponds to the number of bits. Note that when the multi-bit weight data is an analog value, the switching circuit 40M can be configured using an analog switch (transfer gate) or the like.

[0082] When the memory circuit unit 20 and the arithmetic circuit 30 are on separate chips, the bus width is limited by the number of pins on the chip. On the other hand, in a configuration in which the memory circuit unit 20 and the arithmetic circuit 30 are stacked as in one embodiment of the present invention, the number of parallel data required for arithmetic processing can be increased depending on the opening where the wiring LBL is provided, thereby enabling efficient arithmetic processing.

[0083] Fig. 5B is an example of a circuit configuration applicable to the switching circuit 40M shown in Fig. 5A. The three-state buffer shown in Fig. 5B has a function of amplifying and transmitting the potentials of n wirings LBL to n wirings GBL in response to n control signals EN.

[0084] FIG. 6 shows a timing chart for explaining the operation of the configuration described in FIG. 3A. The semiconductor device 10 performs arithmetic processing in response to the toggle operation of the clock signal CLK (for example, times T1 to T7). By configuring the clock signal CLK to have a higher frequency, it is possible to speed up the arithmetic processing. In FIG. 6, W a or W f , W1 to W 17 is the weight data.

[0085] When the input data A1 to A6 are switched at high speed to A1a to A111, A2a to A211, A3a to A311, A4a to A411, A5a to A511, and A6a to A611, respectively, as shown in the figure, in response to the clock signal CLK, the data of the wiring GBL that provides the weight data must be switched at high speed.

[0086] In one embodiment of the present invention, weight data selected from the wiring LBL to the wiring GBL by the switching circuit 40 is read in advance onto the wirings LBL_1 to LBL_6, thereby enabling high-speed switching of data on the wiring GBL to which weight data is applied. For example, weight data W1 can be read onto the wiring LBL_1 at time T1, and the switching circuit 40 can be switched at time T6 to output the weight data W1 from the wiring LBL_1 to the wiring GBL. From time T2 to T7 and after time T7, the weight data can be read onto the wiring LBL at different times from the weight data selected on the wiring GBL, thereby enabling switching of the weight data in response to the clock signal CLK.

[0087] FIG. 7A shows a specific example of the configuration of an arithmetic circuit. FIG. 7A illustrates an example of the configuration of an arithmetic circuit 30 that can perform a product-sum operation between 8-bit weight data and 8-bit input data. FIG. 7A also illustrates a multiplier circuit 24, an adder circuit 25, and a register 26. The 16-bit data multiplied by the multiplier circuit 24 is input to the adder circuit 25. The output of the adder circuit 25 is held in the register 26, and is added to the data multiplied by the multiplier circuit 24 in the adder circuit 25 to perform the product-sum operation. The register is controlled by a clock signal CLK and a reset signal reset_B. Note that the "α" in "17+α" in the figure indicates a carry that occurs when the multiplied data is added. With this configuration, the weight data W SEL and input data A IN It is possible to obtain output data MAC corresponding to the sum-of-products operation with

[0088] 7A illustrates a configuration in which arithmetic processing is performed using 8-bit data, but one embodiment of the present invention can also be applied to a configuration in which 1-bit data is used. This configuration is illustrated in FIG. 7B, similar to FIG. 7A. In the case of 1-bit data, arithmetic processing according to the number of bits can be performed, as illustrated in FIG. 7B.

[0089] 8A is a diagram illustrating an example of a circuit configuration applicable to the memory circuit unit 20 of the semiconductor device 10 of the present invention. Fig. 8A illustrates write word lines WWL_1 to WWL_M, read word lines RWL_1 to RWL_M, write bit lines WBL_1 to WBL_N, and wirings LBL_1 to LBL_N, which are arranged in a matrix of M rows and N columns (M and N are natural numbers of 2 or more). Also illustrated is a memory circuit 21 connected to each word line and bit line.

[0090] 8B is a diagram illustrating an example of a circuit configuration applicable to the memory circuit 21. The memory circuit 21 includes a transistor 61, a transistor 62, a transistor 63, and a capacitor 64 (also referred to as a capacitor).

[0091] One of the source or drain of the transistor 61 is connected to the write bit line WBL. The gate of the transistor 61 is connected to the write word line WWL. The other of the source or drain of the transistor 61 is connected to one electrode of the capacitor 64 and the gate of the transistor 62. One of the source or drain of the transistor 62 and the other electrode of the capacitor 64 are connected to a wiring that applies a fixed potential, such as a ground potential. The other of the source or drain of the transistor 62 is connected to one of the source or drain of the transistor 63. The gate of the transistor 63 is connected to the read word line RWL. The other of the source or drain of the transistor 63 is connected to a wiring LBL. The wiring LBL is connected to a wiring GBL via the switching circuit 40. As described above, the wiring LBL is connected to the switching circuit 40 via a wiring that extends in a direction approximately perpendicular to the surface of the substrate on which the arithmetic circuit 30 is provided.

[0092] The circuit configuration of the memory circuit 21 shown in FIG. 8B corresponds to a three-transistor (3T) gain cell NOSRAM. Transistors 61 to 63 are OS transistors. OS transistors have an extremely small leakage current, i.e., a current flowing between the source and drain in the off state. NOSRAM can be used as a nonvolatile memory by utilizing its extremely small leakage current characteristic to retain charge corresponding to data within the memory circuit. Note that when the transistor 61 shown in FIG. 8B is a Si transistor, it is designed to have an extremely small leakage current, i.e., a current flowing between the source and drain in the off state. For example, the channel length is designed to be sufficiently long relative to the channel width.

[0093] The circuit configuration applicable to the memory circuit 21 of FIG. 8A is not limited to the 3T NOSRAM of FIG. 8B. For example, a circuit equivalent to the DOSRAM illustrated in FIG. 9A may also be used. FIG. 9A illustrates a memory circuit 21A having a transistor 61A and a capacitive element 64A. The transistor 61A is an OS transistor. The memory circuit 21A is illustrated as being connected to a bit line BL, a word line WL, and a back gate line BGL.

[0094] A circuit configuration applicable to the memory circuit 21 of FIG. 8A may be a circuit equivalent to a 2T NOSRAM shown in FIG. 9B. FIG. 9B illustrates the memory circuit 21B including a transistor 61B, a transistor 62B, and a capacitor 64B. The transistors 61B and 62B are OS transistors. The transistors 61B and 62B may be OS transistors whose semiconductor layers are arranged in different layers or may be OS transistors whose semiconductor layers are arranged in the same layer. The memory circuit 21B is illustrated as being connected to a write bit line WBL, a wiring LBL functioning as a read bit line, a write word line WWL, a read word line RWL, a source line SL, and a back gate line BGL.

[0095] A circuit configuration applicable to the memory circuit 21 of FIG. 8A may be a circuit combining 3T-type NOSRAMs as shown in FIG. 9C. FIG. 9C illustrates a memory circuit 21C including a memory circuit 21_P and a memory circuit 21_N, each capable of holding data of different logic levels. FIG. 9C illustrates a memory circuit 21_P including a transistor 61_P, a transistor 62_P, a transistor 63_P, and a capacitor 64_P, and a memory circuit 21_N including a transistor 61_N, a transistor 62_N, a transistor 63_N, and a capacitor 64_N. Each transistor included in the memory circuit 21_P and the memory circuit 21_N is an OS transistor. Each transistor included in the memory circuit 21_P and the memory circuit 21_N may be an OS transistor having semiconductor layers arranged in different layers or may be an OS transistor having semiconductor layers arranged in the same layer. The memory circuit 21C is shown as being connected to a write bit line WBL_P, a wiring LBL_P, a write bit line WBL_N, a wiring LBL_N, a write word line WWL, and a read word line RWL. The memory circuit 21C can hold data of different logics, read the data of different logics to the wiring LBL_P and the wiring LBL_N, and output the data to the wiring GBL via the switching circuit 40, similar to FIG. 3 etc.

[0096] 9C, an exclusive OR circuit (XOR circuit) may be provided so that data corresponding to the multiplication of data held in the memory circuit 21_P and the memory circuit 21_N is output to the wiring LBL. This configuration can omit the operation corresponding to the multiplication in the arithmetic circuit 30, thereby achieving low power consumption.

[0097] FIG. 10 illustrates the flow of computational processing in a convolutional neural network. FIG. 10 illustrates an input layer 90A, an intermediate layer 90B (also referred to as a hidden layer), and an output layer 90C. The input layer 90A illustrates input data input processing 91 (denoted as “Input” in the figure). The intermediate layer 90B illustrates convolutional computations 92, 93, and 95 (denoted as “Conv.” in the figure) and multiple pooling computations 94 and 96 (denoted as “Pool.” in the figure). The output layer 90C illustrates a fully connected computation 97 (denoted as “Full” in the figure). The computational processing flow in the input layer 90A, intermediate layer 90B, and output layer 90C is merely an example; in actual computational processing in a convolutional neural network, other computations such as softmax computation may be performed.

[0098] 10, convolutional neural network performs convolutional calculations 92, 93, and 95 multiple times. In the convolutional calculations, the same weight data is used. Therefore, by applying the configuration of one embodiment of the present invention in which the same weight data is used, it is possible to achieve both high operating speed and low power consumption.

[0099] Next, a detailed block diagram of the semiconductor device 10 is shown in FIG.

[0100] FIG. 11 illustrates an example configuration of the drive circuit 50 illustrated in FIGS. 2A and 2B, in addition to the configurations corresponding to the memory circuit section 20, memory circuit 21, arithmetic circuit 30, switching circuit 40, layer 11, and layer 12 described in FIGS. 1A and 1B, and FIGS. 2A and 2B.

[0101] FIG. 11 illustrates a controller 71, a row decoder 72, a word line driver 73, a column decoder 74, a write driver 75, a precharge circuit 76, an input / output buffer 81, and an arithmetic control circuit 82 as components corresponding to the drive circuit 50 described in FIGS. 2A and 2B.

[0102] Fig. 12A is a diagram illustrating blocks that control the memory circuit unit 20 for each configuration illustrated in Fig. 11. Fig. 12A illustrates a controller 71, a row decoder 72, a word line driver 73, a column decoder 74, a write driver 75, and a precharge circuit 76.

[0103] The controller 71 processes externally input signals and generates control signals for the row decoder 72 and the column decoder 74. The externally input signals are control signals such as a write enable signal and a read enable signal for controlling the memory circuit unit 20. The controller 71 also inputs and outputs data between the CPU 110 and the semiconductor device 10 via a bus 120.

[0104] The row decoder 72 generates signals for driving the word line driver 73. The word line driver 73 generates signals to be provided to the write word lines WWL and the read word lines RWL. The column decoder 74 generates signals for driving the write driver 75. The write driver 75 generates weight data to be provided to the memory circuit 21. The precharge circuit 76 has the function of precharging the wiring LBL and the like. A signal corresponding to the weight data read from the memory circuit 21 of the memory circuit unit 20 is input to the switching circuit 40 via the wiring LBL, as described with reference to Figures 2A and 2B, etc.

[0105] FIG. 12B is a diagram illustrating the blocks that control the arithmetic circuit 30 and the switching circuit 40 in each configuration shown in FIG.

[0106] The controller 71 processes an input signal from the outside and generates a control signal for the arithmetic control circuit 82. The controller 71 also generates various signals such as an address signal and a clock signal for controlling the arithmetic circuit 30. The arithmetic control circuit 82 processes input data A1 to A2 given to the data input lines in accordance with the control of the controller 71 and the output of the input / output buffer 81. NThe arithmetic control circuit 82 outputs a control signal for controlling the switching circuit 40. As described with reference to FIGS. 2A and 2B, the switching circuit 40 supplies one of the weight data supplied from the plurality of wirings LBL to the plurality of arithmetic circuits 30 via the wiring GBL. The arithmetic circuit 30 generates output data MAC according to the product-sum operation by switching the supplied weight data and input data. The generated output data MAC is temporarily stored as intermediate data in a memory such as an SRAM or a register in the arithmetic control circuit 82 via the input / output buffer 81. The stored intermediate data is re-input to the arithmetic circuit 30.

[0107] Note that it is preferable to use a plurality of semiconductor devices 10 in combination to enable parallel calculation with an increased number of parallel operations, in accordance with an embodiment of the present invention. An example of such a configuration will be described with reference to FIGS. 13A and 13B.

[0108] 13A illustrates a configuration corresponding to the semiconductor device 10 described above, in which semiconductor devices 10_1 to 10_n (n is a number equal to or greater than 2) and a controller 71G that controls and inputs / outputs data between the semiconductor devices 10_1 to 10_n. The controller 71G has an internal memory circuit 60 such as an SRAM. The controller 71G stores output data MAC obtained from the plurality of semiconductor devices 10_1 to 10_n in the memory circuit 60. The output data MAC stored in the memory circuit 60 is then used as input data A for the plurality of semiconductor devices 10_1 to 10_n. IN With this configuration, parallel calculations can be performed using a plurality of semiconductor devices with an increased number of parallel operations.

[0109] 13B, which is a different configuration example from that of FIG. 13A, a controller 71G performs another arithmetic process on the output data held in the memory circuit 60, and outputs the resultant input data as input data A to the plurality of semiconductor devices 10_1 to 10_n. IN _1 to A IN_n as the output. In this configuration, for example, the controller 71G is configured to perform arithmetic processing based on an activation function, pooling processing, normalization processing, etc. on the output data stored in the memory circuit 60. With this configuration, in addition to parallel calculation with an increased number of parallel operations using multiple semiconductor devices, arithmetic processing other than convolution processing can be efficiently performed.

[0110] In the semiconductor device 10, output data MAC corresponding to the calculation result of the arithmetic circuit 30 is input as intermediate data to the arithmetic control circuit 82 using the buffer memory in the input / output buffer 81. The arithmetic control circuit 82 can output this intermediate data again as input data to the arithmetic circuit 30. Therefore, arithmetic processing can be performed without reading data during calculation into a main memory or the like external to the semiconductor device 10. Furthermore, in the semiconductor device 10, electrical connection between the memory circuit unit and the arithmetic circuit can be made via wiring in an opening provided in an insulating film or the like, so the number of parallel connections can be increased by increasing the number of wirings. Therefore, the semiconductor device 10 can perform parallel calculations with a number of bits greater than the data bus width of the CPU 110. Furthermore, the number of times a huge amount of weight data needs to be transferred to and from the CPU 110 can be reduced, thereby achieving low power consumption.

[0111] As described above, one embodiment of the present invention can provide a miniaturized semiconductor device that functions as an accelerator. Alternatively, one embodiment of the present invention can provide a semiconductor device that functions as an accelerator and has low power consumption. Alternatively, one embodiment of the present invention can provide a semiconductor device that functions as an accelerator with a novel structure.

[0112] (Embodiment 2) In this embodiment, a description will be given of the configuration of an integrated circuit having Si transistors applicable to the accelerator described as the semiconductor device 10. This configuration increases the degree of freedom in designing the semiconductor device and also increases the degree of integration of the semiconductor device.

[0113] 14A is an example of a schematic cross-sectional view illustrating an integrated circuit 390. In the integrated circuit 390, the semiconductor device 10 described in the above embodiment is provided on a package substrate 400. The package substrate 400 is provided with solder balls 401 for connection to another printed circuit board or the like. The semiconductor device 10 is connected to the package substrate 400 via an interposer or the like. The package substrate 400 can be a ceramic substrate, a plastic substrate, a glass epoxy substrate, or the like.

[0114] The cross-sectional view of integrated circuit 390 shown in Figure 14A illustrates, on the layer 11 side, a semiconductor substrate 402, a plurality of transistors 403 provided on semiconductor substrate 402, wiring 404, and electrodes 405. Also, on the layer 12 side, a semiconductor substrate 412, a plurality of transistors 413 provided on semiconductor substrate 412, wiring 414, and electrodes 415. The configuration of region 420 shown in Figure 14A will be described with reference to Figure 14B.

[0115] Fig. 14B illustrates the semiconductor substrate 402, transistor 403, wiring 404, and electrode 405 illustrated in Fig. 14A. Fig. 14B also illustrates the semiconductor substrate 412, multiple transistors 413 provided on the semiconductor substrate 412, wiring 414, and electrode 415 illustrated in Fig. 14A.

[0116] When the layers 11 and 12 are bonded together, the transistors 403 and 413 provided on the respective semiconductor substrates are connected to the electrodes 405 and 415 via the wirings 404 and 414. The electrodes 405 and 415 are bonded together using bonding techniques such as Cu-Cu bonding or microbumps. Note that Cu-Cu bonding is a technique for achieving electrical conductivity by connecting Cu (copper) pads together. Note that a through-silicon via (TSV) may be formed in the semiconductor substrates 402 and 412 to connect to the electrodes 405 and 415. The thickness of the semiconductor substrates 402 and 412 is 100 μm to 300 μm, but may be thinned to 10 μm to 100 μm by polishing.

[0117] 15 will be used to explain the semiconductor substrate 402, transistor 403, wiring 404, and electrode 405 in layer 11, and the semiconductor substrate 412, transistor 413, wiring 414, and electrode 415 in layer 12. To avoid repetitive explanations, the explanations of the semiconductor substrate 412, transistor 413, wiring 414, and electrode 415 in layer 12, which correspond to the semiconductor substrate 402, transistor 403, wiring 404, and electrode 405 in layer 11, will be simplified.

[0118] The transistor 403 is provided over a semiconductor substrate 402 and includes a conductor 430 functioning as a gate, an insulator 431 functioning as a gate insulator, a semiconductor region 432 formed of part of the semiconductor substrate 402, and low-resistance regions 433a and 433b functioning as source and drain regions. The transistor 403 may be either a p-channel type or an n-channel type.

[0119] The semiconductor substrate 402 having the semiconductor region 432, the low-resistance region 433a, and the low-resistance region 433b preferably includes a semiconductor such as a silicon-based semiconductor, and preferably includes single-crystal silicon. Alternatively, it may be formed of a material including Ge (germanium), SiGe (silicon germanium), GaAs (gallium arsenide), GaAlAs (gallium aluminum arsenide), or the like. It may also be configured using silicon whose effective mass is controlled by applying stress to the crystal lattice and changing the lattice spacing. Alternatively, the transistor 403 may be a high electron mobility transistor (HEMT) by using GaAs and GaAlAs, or the like.

[0120] In addition to the semiconductor material applied to semiconductor region 432, low resistance region 433a, and low resistance region 433b, an element that imparts n-type conductivity, such as arsenic or phosphorus, or an element that imparts p-type conductivity, such as boron, is included.

[0121] The conductor 430 functioning as the gate electrode can be made of a conductive material such as a semiconductor material, metal material, alloy material, or metal oxide material, such as silicon containing an element that imparts n-type conductivity, such as arsenic or phosphorus, or an element that imparts p-type conductivity, such as boron.

[0122] Since the work function is determined by the material of the conductor, the threshold voltage can be adjusted by changing the material of the conductor. Specifically, it is preferable to use materials such as titanium nitride and tantalum nitride for the conductor. Furthermore, in order to achieve both conductivity and embeddability, it is preferable to use metal materials such as tungsten and aluminum as a laminate for the conductor, and tungsten is particularly preferable in terms of heat resistance.

[0123] Note that the transistor 403 shown in FIG. 15 is just an example, and the structure is not limited to this, and an appropriate transistor may be used depending on the circuit configuration and driving method.

[0124] An insulator 440, an insulator 442, an insulator 444, and an insulator 446 are stacked in this order to cover the transistor 403.

[0125] The insulators 440, 442, 444, and 446 can be formed using, for example, silicon oxide, silicon oxynitride, silicon nitride oxide, silicon nitride, aluminum oxide, aluminum oxynitride, aluminum nitride oxide, aluminum nitride, or the like.

[0126] The insulator 442 may function as a planarizing film that flattens a step caused by the transistor 403 or the like provided thereunder. For example, the top surface of the insulator 442 may be planarized by planarization treatment using a chemical mechanical polishing (CMP) method or the like to improve the planarity.

[0127] Note that the insulator 446 preferably has a lower dielectric constant than the insulator 444. For example, the relative dielectric constant of the insulator 446 is preferably less than 4, and more preferably less than 3. Furthermore, for example, the relative dielectric constant of the insulator 446 is preferably 0.7 times or less, and more preferably 0.6 times or less, the relative dielectric constant of the insulator 444. By using a material with a low dielectric constant as the interlayer film, the parasitic capacitance generated between wirings can be reduced.

[0128] A conductor 448 electrically connected to the transistor 403 and a conductor functioning as the wiring 404 are embedded in the insulators 440, 442, 444, and 446. The conductor 448 functions as a plug or a wiring. A plurality of conductors functioning as a plug or a wiring may be collectively denoted by the same reference numeral. In this specification and the like, a wiring and a plug electrically connected to the wiring may be integrated. That is, a part of the conductor may function as a wiring, and a part of the conductor may function as a plug.

[0129] The materials for each plug and wiring (such as the conductor 448 and the wiring 404) can be a conductive material such as a metal material, an alloy material, a metal nitride material, or a metal oxide material, and can be used in a single layer or a laminated layer. High-melting-point materials such as tungsten and molybdenum, which have both heat resistance and conductivity, are preferably used, and tungsten is preferred. Alternatively, they are preferably formed from a low-resistance conductive material such as aluminum or copper. The use of a low-resistance conductive material can reduce the wiring resistance.

[0130] The electrode 405 can be provided over the insulator 446 and the wiring 404. For example, in FIG. 15 , an insulator 450, an insulator 452, and an insulator 454 are stacked in this order. The electrode 405 can be formed by forming openings after the insulators 450, 452, and 454 are formed, providing a conductive layer to fill the openings, and polishing the surface by a CMP method.

[0131] The electrode 405 can be, for example, a metal film containing an element selected from Al, Cr, Cu, Ta, Ti, Mo, or W, or a metal nitride film containing the above elements (titanium nitride film, molybdenum nitride film, tungsten nitride film). By using a conductive bump (hereinafter referred to as a bump) as the electrode 405, a Cu-Cu (copper-copper) direct bond can be achieved. The Cu-Cu direct bond is a technique for achieving electrical conductivity by connecting Cu (copper) pads together. The electrode 405 functions as a plug or wiring. The electrode 405 can be formed using the same material as the conductor 448 and wiring 404.

[0132] This embodiment mode can be combined with the descriptions of other embodiment modes as appropriate.

[0133] (Embodiment 3) In this embodiment, an example of operation will be described in which part of the calculations of a program executed by the CPU 110 described in the above embodiment is executed by an accelerator described as the semiconductor device 10.

[0134] FIG. 16 is a diagram illustrating an example of an operation when part of the calculations of a program executed by a CPU is executed by an accelerator.

[0135] The host program is executed by the CPU (host program execution; step S1).

[0136] When the CPU confirms an instruction to reserve a data area required for calculations using the accelerator in the memory circuit unit (memory reservation instruction; step S2), it reserves the data area in the memory circuit unit (memory reservation; step S3).

[0137] Next, the CPU transmits weight data, which is input data, from the main memory or an external storage device to the memory circuit unit (data transmission; step S4). The memory circuit unit receives the weight data and stores it in the area reserved in step S2 (data reception; step S5).

[0138] When the CPU confirms an instruction to start the kernel program (start kernel program; step S6), the accelerator starts executing the kernel program (start operation; step S7).

[0139] Immediately after the accelerator starts executing the kernel program, the CPU may be switched from a state performing calculations to a PG (power gating) state (PG state transition; step S8). In this case, immediately before the accelerator finishes executing the kernel program, the CPU is switched from the PG state to a state performing calculations (PG state stop step S9). By keeping the CPU in the PG state during the period from step S8 to step S9, it is possible to suppress power consumption and heat generation in the entire calculation processing system.

[0140] When the accelerator finishes executing the kernel program, the output data is stored in a storage unit that holds the calculation results in the accelerator (computation end; step S10).

[0141] After the execution of the kernel program is completed, if the CPU confirms an instruction to transmit the output data stored in the memory unit to the main memory or an external storage device (data transmission request; step S11), the output data is transmitted to the main memory or the external storage device and stored in the main memory or the external storage device (data transmission; step S12).

[0142] By repeating the above operations from step S1 to step S14, the power consumption and heat generation of the CPU and the accelerator can be suppressed, and part of the computations to be performed by the CPU can be performed by the accelerator. The semiconductor device of one embodiment of the present invention has a non-von Neumann architecture, and can perform computations with significantly less power consumption than a von Neumann architecture, which consumes more power as the processing speed increases.

[0143] This embodiment mode can be combined with the descriptions of other embodiment modes as appropriate.

[0144] (Fourth embodiment) In this embodiment, an example of a CPU having a CPU core capable of power gating will be described.

[0145] 17 shows an example of the configuration of the CPU 110. The CPU 110 has a CPU core 200, an L1 (level 1) cache memory device (L1 Cache) 202, an L2 cache memory device (L2 Cache) 203, a bus interface unit (Bus I / F) 205, power switches 210 to 212, and a level shifter (LS) 214. The CPU core 200 has a flip-flop 220.

[0146] The CPU core 200, the L1 cache memory device 202, and the L2 cache memory device 203 are interconnected by a bus interface unit 205.

[0147] The PMU 193 generates a clock signal GCLK1 and various PG (power gating) control signals in response to interrupt signals (Interrupts) input from the outside and signals such as the SLEEP1 signal issued by the CPU 110. The clock signal GCLK1 and the PG control signals are input to the CPU 110. The PG control signals control the power switches 210 to 212 and the flip-flop 220.

[0148] Power switches 210 and 211 control the supply of voltages VDDD and VDD1 to a virtual power line V_VDD (hereinafter referred to as a V_VDD line), respectively. A power switch 212 controls the supply of a voltage VDDH to a level shifter (LS) 214. A voltage VSSS is input to the CPU 110 and PMU 193 without passing through a power switch. A voltage VDDD is input to the PMU 193 without passing through a power switch.

[0149] The voltages VDDD and VDD1 are drive voltages for the CMOS circuit. The voltage VDD1 is lower than the voltage VDDD and is the drive voltage in the sleep state. The voltage VDDH is the drive voltage for the OS transistors and is higher than the voltage VDDD.

[0150] Each of the L1 cache memory device 202, the L2 cache memory device 203, and the bus interface unit 205 has at least one power domain that can be power-gated. Each power domain that can be power-gated has one or more power switches. These power switches are controlled by a PG control signal.

[0151] The flip-flop 220 is used as a register. A backup circuit is provided in the flip-flop 220. The flip-flop 220 will be described below.

[0152] 18 shows an example of the circuit configuration of the flip-flop 220. The flip-flop 220 has a scan flip-flop 221 and a backup circuit 222.

[0153] The scan flip-flop 221 has nodes D1, Q1, SD, SE, RT, CK, and a clock buffer circuit 221A.

[0154] Node D1 is a data input node, node Q1 is a data output node, and node SD is an input node for scan test data. Node SE is an input node for signal SCE. Node CK is an input node for clock signal GCLK1. Clock signal GCLK1 is input to clock buffer circuit 221A. The analog switch of scan flip-flop 221 is connected to nodes CK1 and CKB1 of clock buffer circuit 221A. Node RT is an input node for a reset signal.

[0155] The signal SCE is a scan enable signal and is generated by the PMU 193. The PMU 193 generates signals BK and RC. The level shifter 214 level-shifts the signals BK and RC to generate signals BKH and RCH. The signal BK is a backup signal, and the signal RC is a recovery signal.

[0156] The circuit configuration of the scan flip-flop 221 is not limited to that shown in Fig. 18. Flip-flops available in a standard circuit library can be applied.

[0157] The backup circuit 222 includes nodes SD_IN and SN11, transistors M11 to M13, and a capacitive element C11.

[0158] The node SD_IN is an input node for scan test data and is connected to the node Q1 of the scan flip-flop 221. The node SN11 is a storage node of the backup circuit 222. The capacitive element C11 is a storage capacitor for storing the voltage of the node SN11.

[0159] The transistor M11 controls the conduction state between the node Q1 and the node SN11. The transistor M12 controls the conduction state between the node SN11 and the node SD. The transistor M13 controls the conduction state between the node SD_IN and the node SD. The on / off of the transistors M11 and M13 is controlled by a signal BKH, and the on / off of the transistor M12 is controlled by a signal RCH.

[0160] The transistors M11 to M13 are OS transistors, similar to the transistors 61 to 63 included in the memory circuit 21. The transistors M11 to M13 are illustrated as having back gates. The back gates of the transistors M11 to M13 are connected to a power supply line that supplies a voltage VBG1.

[0161] At least the transistors M11 and M12 are preferably OS transistors. The OS transistors have an extremely small off-state current, which prevents a voltage drop at the node SN11. Furthermore, the backup circuit 222 consumes almost no power to retain data, making it nonvolatile. Because data is rewritten by charging and discharging the capacitive element C11, the backup circuit 222 is theoretically capable of writing and reading data without any restrictions on the number of times it can be rewritten, and with low energy consumption.

[0162] It is highly preferable that all transistors in the backup circuit 222 are OS transistors. As shown in Fig. 18B, the backup circuit 222 can be stacked on a scan flip-flop 221 made up of a silicon CMOS circuit.

[0163] Since the backup circuit 222 has an extremely small number of elements compared to the scan flip-flop 221, stacking the backup circuit 222 does not require changing the circuit configuration and layout of the scan flip-flop 221. In other words, the backup circuit 222 is a highly versatile backup circuit. Furthermore, since the backup circuit 222 can be provided in the region where the scan flip-flop 221 is formed, even if the backup circuit 222 is incorporated, the area overhead of the flip-flop 220 can be reduced to zero. Therefore, providing the backup circuit 222 in the flip-flop 220 enables power gating of the CPU core 200. Because little energy is required for power gating, the CPU core 200 can be power gated with high efficiency.

[0164] By providing the backup circuit 222, a parasitic capacitance due to the transistor M11 is added to the node Q1, but since it is small compared to the parasitic capacitance due to the logic circuit connected to the node Q1, it does not affect the operation of the scan flip-flop 221. In other words, even if the backup circuit 222 is provided, the performance of the flip-flop 220 does not substantially deteriorate.

[0165] For example, a clock gating state, a power gating state, or a sleep state can be set as the low power consumption state of the CPU core 200. The PMU 193 selects the low power consumption mode of the CPU core 200 based on an interrupt signal, a signal SLEEP1, etc. For example, when transitioning from a normal operating state to a clock gating state, the PMU 193 stops generating the clock signal GCLK1.

[0166] For example, when transitioning from a normal operating state to a hibernation state, the PMU 193 performs voltage and / or frequency scaling. For example, when performing voltage scaling, the PMU 193 turns off the power switch 210 and turns on the power switch 211 to input the voltage VDD1 to the CPU core 200. The voltage VDD1 is a voltage that does not cause data to be lost in the scan flip-flop 221. When performing frequency scaling, the PMU 193 reduces the frequency of the clock signal GCLK1.

[0167] When the CPU core 200 is transitioned from the normal operation state to the power gating state, an operation is performed to back up the data of the scan flip-flop 221 to the backup circuit 222. When the CPU core 200 is returned from the power gating state to the normal operation state, an operation is performed to recover the data of the backup circuit 222 to the scan flip-flop 221.

[0168] 19 shows an example of a power gating sequence of the CPU core 200. In FIG. 19, t1 to t7 represent time. Signals PSE0 to PSE2 are control signals for the power switches 210 to 212, and are generated by the PMU 193. When the signal PSE0 is "H" / "L", the power switch 210 is on / off. The same applies to the signals PSE1 and PSE2.

[0169] Before time t1, the state is normal operation. The power switch 210 is on, and the voltage VDDD is input to the CPU core 200. The scan flip-flop 221 performs normal operation. At this time, the level shifter 214 does not need to operate, so the power switch 212 is off, and the signals SCE, BK, and RC are "L". Since the node SE is "L", the scan flip-flop 221 stores the data of the node D1. In the example of FIG. 19, at time t1, the node SN11 of the backup circuit 222 is "L".

[0170] At operation time t1, the PMU 193 stops the clock signal GCLK1 and sets the signals PSE2 and BK to "H." The level shifter 214 becomes active and outputs the signal BKH at "H" to the backup circuit 222.

[0171] The transistor M11 of the backup circuit 222 turns on, and the data at the node Q1 of the scan flip-flop 221 is written to the node SN11 of the backup circuit 222. If the node Q1 of the scan flip-flop 221 is "L", the node SN11 remains "L", and if the node Q1 is "H", the node SN11 becomes "H".

[0172] The PMU 193 sets the signals PSE2 and BK to "L" at time t2, and sets the signal PSE0 to "L" at time t3. At time t3, the state of the CPU core 200 transitions to the power gating state. Note that the signal PSE0 may also fall at the same timing as the signal BK falls.

[0173] The operation during power gating will be described. When the signal PSE0 goes to "L", the voltage of the V_VDD line drops, and the data at node Q1 is lost. Node SN11 continues to hold the data at node Q1 at time t3.

[0174] The operation during recovery will be explained below. At time t4, the PMU 193 sets the signal PSE0 to "H", transitioning from the power gating state to the recovery state. Charging of the V_VDD line begins, and when the voltage on the V_VDD line reaches VDDD (time t5), the PMU 193 sets the signals PSE2, RC, and SCE to "H".

[0175] Transistor M12 turns on, and the charge of capacitive element C11 is distributed between node SN11 and node SD. If node SN11 is "H," the voltage of node SD rises. Since node SE is "H," the data of node SD is written to the input latch circuit of scan flip-flop 221. When clock signal GCLK1 is input to node CK at time t6, the data of the input latch circuit is written to node Q1. In other words, the data of node SN11 has been written to node Q1.

[0176] At time t7, the PMU 193 sets the signals PSE2, SCE, and RC to "L," and the recovery operation ends.

[0177] The backup circuit 222 using OS transistors consumes low dynamic and static power, making it highly suitable for normally-off computing. A CPU 110 including a CPU core 200 with a backup circuit 222 using OS transistors can be called an NoffCPU (registered trademark). The NoffCPU has nonvolatile memory and can stop power supply when operation is not required. Even if the flip-flop 220 is installed, it is possible to minimize the degradation of CPU core 200 performance and the increase in dynamic power.

[0178] The CPU core 200 may have multiple power domains that can be power-gated. Each of the multiple power domains is provided with one or more power switches for controlling the input of voltage. The CPU core 200 may also have one or more power domains in which power gating is not performed. For example, a power domain in which power gating is not performed may be provided with a power gating control circuit for controlling the flip-flop 220 and the power switches 210 to 212.

[0179] The application of the flip-flop 220 is not limited to the CPU 110. In the CPU 110, the flip-flop 220 can be applied to a register provided in a power domain that is capable of power gating.

[0180] This embodiment mode can be combined with the descriptions of other embodiment modes as appropriate.

[0181] (Embodiment 5) In this embodiment, an example of a configuration of a transistor applicable to the CPU 110 described in the above embodiment and the accelerator described as the semiconductor device 10 will be described. As an example, a configuration in which transistors having different electrical characteristics are stacked will be described. This configuration can increase the degree of freedom in designing a semiconductor device. Furthermore, stacking transistors having different electrical characteristics can increase the degree of integration of a semiconductor device.

[0182] FIG. 20 shows a part of a cross-sectional structure of a semiconductor device. The semiconductor device shown in FIG. 20 includes a transistor 550, a transistor 500, and a capacitor 600. FIG. 21A is a cross-sectional view of the transistor 500 in the channel length direction, and FIG. 21B is a cross-sectional view of the transistor 500 in the channel width direction. For example, the transistor 500 corresponds to an OS transistor included in the memory circuit 21 described in the above embodiment, that is, a transistor having an oxide semiconductor in a channel formation region. The transistor 550 corresponds to a Si transistor included in the arithmetic circuit 30 described in the above embodiment, that is, a transistor having silicon in a channel formation region. The capacitor 600 corresponds to a capacitor included in the memory circuit 21.

[0183] The transistor 500 is an OS transistor. An OS transistor has an extremely low off-state current. Therefore, a data voltage or charge written to a storage node through the transistor 500 can be held for a long period of time. That is, the frequency of refresh operations of the storage node can be reduced or no refresh operations are required, thereby reducing the power consumption of the semiconductor device.

[0184] In FIG. 20, the transistor 500 is provided above the transistor 550 , and the capacitor 600 is provided above the transistor 550 and the transistor 500 .

[0185] The transistor 550 is provided on a substrate 311. The substrate 311 is, for example, a p-type silicon substrate. The substrate 311 may also be an n-type silicon substrate. The oxide layer 314 is preferably an insulating layer (also referred to as a BOX layer) formed by buried oxidation (buried oxide) in the substrate 311, such as silicon oxide. The transistor 550 is provided on a single-crystal silicon substrate provided on the substrate 311 with the oxide layer 314 interposed therebetween, a so-called SOI (Silicon On Insulator) substrate.

[0186] A substrate 311 in the SOI substrate is provided with an insulator 313 that functions as an element isolation layer. The substrate 311 also has a well region 312. The well region 312 is a region that is given n-type or p-type conductivity depending on the conductivity type of the transistor 550. The single crystal silicon in the SOI substrate is provided with a semiconductor region 315, and low-resistance regions 316a and 316b that function as source and drain regions. A low-resistance region 316c is also provided on the well region 312.

[0187] The transistor 550 can be provided overlapping a well region 312 to which an impurity element imparting conductivity is added. The well region 312 can function as a bottom gate electrode of the transistor 550 by independently changing the potential through the low-resistance region 316c. This allows the threshold voltage of the transistor 550 to be controlled. In particular, applying a negative potential to the well region 312 can increase the threshold voltage of the transistor 550 and reduce the off-state current. Therefore, applying a negative potential to the well region 312 can reduce the drain current when the potential applied to the gate electrode of the Si transistor is 0 V. As a result, power consumption due to a through current or the like in the arithmetic circuit 30 including the transistor 550 can be reduced, thereby improving arithmetic efficiency.

[0188] The transistor 550 is preferably a so-called fin type transistor in which the top surface of the semiconductor layer and the side surfaces in the channel width direction are covered with a conductor 318 via an insulator 317. By using the fin type transistor 550, the effective channel width can be increased, thereby improving the on-state characteristics of the transistor 550. Furthermore, the contribution of the electric field of the gate electrode can be increased, thereby improving the off-state characteristics of the transistor 550.

[0189] Note that the transistor 550 may be either a p-channel transistor or an n-channel transistor.

[0190] The conductor 318 may function as a first gate (also called a top gate) electrode, and the well region 312 may function as a second gate (also called a bottom gate) electrode. In this case, the potential applied to the well region 312 can be controlled via the low-resistance region 316c.

[0191] The region where the channel of the semiconductor region 315 is formed, the region nearby, the low-resistance region 316a and low-resistance region 316b that serve as the source or drain region, and the low-resistance region 316c connected to an electrode that controls the potential of the well region 312 preferably contain a semiconductor such as a silicon-based semiconductor, and preferably single-crystal silicon. Alternatively, they may be formed of a material containing Ge (germanium), SiGe (silicon germanium), GaAs (gallium arsenide), GaAlAs (gallium aluminum arsenide), or the like. A configuration using silicon in which the effective mass is controlled by applying stress to the crystal lattice and changing the lattice spacing may also be used. Alternatively, the transistor 550 may be a high electron mobility transistor (HEMT) by using GaAs and GaAlAs, or the like.

[0192] Well region 312, low resistance region 316a, low resistance region 316b, and low resistance region 316c contain, in addition to the semiconductor material applied to semiconductor region 315, an element that imparts n-type conductivity, such as arsenic or phosphorus, or an element that imparts p-type conductivity, such as boron.

[0193] The conductor 318 functioning as the gate electrode can be made of a conductive material such as a semiconductor material, metal material, alloy material, or metal oxide material, including an element that imparts n-type conductivity, such as arsenic or phosphorus, or an element that imparts p-type conductivity, such as boron. The conductor 318 may also be made of a silicide, such as nickel silicide.

[0194] Since the work function is determined by the material of the conductor, the threshold voltage of the transistor can be adjusted by selecting the material of the conductor. Specifically, it is preferable to use a material such as titanium nitride or tantalum nitride as the conductor. Furthermore, in order to achieve both conductivity and embeddability, it is preferable to use a metal material such as tungsten or aluminum as the conductor in a laminated state, and tungsten is particularly preferable in terms of heat resistance.

[0195] The low-resistance regions 316a, 316b, and 316c may be formed by stacking another conductor, for example, a silicide such as nickel silicide. This configuration can increase the conductivity of the regions that function as electrodes. In this case, an insulator that functions as a sidewall spacer (also referred to as a sidewall insulating layer) may be provided on the side surface of the conductor 318 that functions as the gate electrode and on the side surface of the insulator that functions as the gate insulating film. This configuration can prevent electrical conduction between the conductor 318 and the low-resistance regions 316a and 316b.

[0196] An insulator 320, an insulator 322, an insulator 324, and an insulator 326 are stacked in this order over the transistor 550.

[0197] The insulators 320, 322, 324, and 326 can be made of, for example, silicon oxide, silicon oxynitride, silicon nitride oxide, silicon nitride, aluminum oxide, aluminum oxynitride, aluminum nitride oxide, aluminum nitride, or the like.

[0198] In this specification, silicon oxynitride refers to a material whose composition contains more oxygen than nitrogen, silicon nitride oxide refers to a material whose composition contains more nitrogen than oxygen, aluminum oxynitride refers to a material whose composition contains more oxygen than nitrogen, and aluminum nitride oxide refers to a material whose composition contains more nitrogen than oxygen.

[0199] The insulator 322 may function as a planarizing film that flattens steps caused by the transistor 550 or the like provided thereunder. For example, the top surface of the insulator 322 may be planarized by planarization treatment using a chemical mechanical polishing (CMP) method or the like to improve the planarity.

[0200] The insulator 324 is preferably a film having a barrier property that prevents hydrogen or impurities from diffusing from the substrate 311, the transistor 550, or the like to a region where the transistor 500 is provided.

[0201] An example of a film having a barrier property against hydrogen is silicon nitride formed by a CVD method. Here, hydrogen diffusion into a semiconductor element having an oxide semiconductor, such as the transistor 500, may degrade the characteristics of the semiconductor element. Therefore, it is preferable to use a film that suppresses hydrogen diffusion between the transistor 500 and the transistor 550. Specifically, the film that suppresses hydrogen diffusion is a film that releases a small amount of hydrogen.

[0202] The amount of desorption of hydrogen can be analyzed using, for example, thermal desorption spectroscopy (TDS). For example, the amount of desorption of hydrogen from the insulator 324 is calculated as 10×10 per area of the insulator 324 when the surface temperature of the film is in the range of 50° C. to 500° C. in TDS analysis. 15 atoms / cm 2 Less than or equal to 5 x 10 15 atoms / cm 2 The following is fine.

[0203] It is preferable that the insulator 326 has a lower dielectric constant than the insulator 324. For example, the relative dielectric constant of the insulator 326 is preferably less than 4, and more preferably less than 3. Furthermore, for example, the relative dielectric constant of the insulator 326 is preferably 0.7 times or less, and more preferably 0.6 times or less, the relative dielectric constant of the insulator 324. By using a material with a low dielectric constant as the interlayer film, the parasitic capacitance that occurs between wirings can be reduced.

[0204] Conductors 328 and 330, which connect to the capacitor 600 or the transistor 500, are embedded in the insulators 320, 322, 324, and 326. The conductors 328 and 330 function as plugs or wiring. A plurality of conductors that function as plugs or wiring may be collectively assigned the same reference numeral. In this specification and the like, a wiring and a plug connected to the wiring may be integrated. That is, a part of a conductor may function as a wiring, and a part of a conductor may function as a plug.

[0205] The materials for each plug and wiring (conductor 328, conductor 330, etc.) can be a conductive material such as a metal material, an alloy material, a metal nitride material, or a metal oxide material, and can be used in a single layer or a laminated layer. High-melting-point materials such as tungsten and molybdenum, which have both heat resistance and conductivity, are preferably used, and tungsten is preferred. Alternatively, they are preferably formed from a low-resistance conductive material such as aluminum or copper. The use of a low-resistance conductive material can reduce the wiring resistance.

[0206] A wiring layer may be provided over the insulator 326 and the conductor 330. For example, in FIG. 20 , the insulator 350, the insulator 352, and the insulator 354 are stacked in this order. The conductor 356 is formed in the insulator 350, the insulator 352, and the insulator 354. The conductor 356 functions as a plug or wiring connected to the transistor 550. Note that the conductor 356 can be formed using a material similar to that of the conductor 328 and the conductor 330.

[0207] Note that, for example, the insulator 350 preferably uses an insulator having a barrier property against hydrogen, similar to the insulator 324. The conductor 356 preferably includes a conductor having a barrier property against hydrogen. In particular, a conductor having a barrier property against hydrogen is formed in an opening of the insulator 350 having a barrier property against hydrogen. With this structure, the transistor 550 and the transistor 500 can be separated by a barrier layer, and diffusion of hydrogen from the transistor 550 to the transistor 500 can be suppressed.

[0208] Note that, for example, tantalum nitride or the like is preferably used as a conductor having a barrier property against hydrogen. Stacking tantalum nitride and highly conductive tungsten can suppress diffusion of hydrogen from the transistor 550 while maintaining the conductivity of the wiring. In this case, it is preferable that the tantalum nitride layer having a barrier property against hydrogen be in contact with the insulator 350 having a barrier property against hydrogen.

[0209] A wiring layer may be provided over the insulator 354 and the conductor 356. For example, in FIG. 20, the insulator 360, the insulator 362, and the insulator 364 are stacked in this order. The conductor 366 is formed in the insulator 360, the insulator 362, and the insulator 364. The conductor 366 functions as a plug or wiring. The conductor 366 can be provided using the same material as the conductor 328 and the conductor 330.

[0210] Note that, for example, the insulator 360 preferably uses an insulator having a barrier property against hydrogen, similar to the insulator 324. The conductor 366 preferably includes a conductor having a barrier property against hydrogen. In particular, a conductor having a barrier property against hydrogen is formed in an opening of the insulator 360 having a barrier property against hydrogen. With this structure, the transistor 550 and the transistor 500 can be separated by a barrier layer, and diffusion of hydrogen from the transistor 550 to the transistor 500 can be suppressed.

[0211] A wiring layer may be provided over the insulator 364 and the conductor 366. For example, in FIG. 20, an insulator 370, an insulator 372, and an insulator 374 are stacked in this order. A conductor 376 is formed in the insulator 370, the insulator 372, and the insulator 374. The conductor 376 functions as a plug or wiring. The conductor 376 can be formed using the same material as the conductors 328 and 330.

[0212] Note that, for example, the insulator 370 preferably uses an insulator having a barrier property against hydrogen, similar to the insulator 324. The conductor 376 preferably includes a conductor having a barrier property against hydrogen. In particular, a conductor having a barrier property against hydrogen is formed in an opening of the insulator 370 having a barrier property against hydrogen. With this structure, the transistor 550 and the transistor 500 can be separated by a barrier layer, and diffusion of hydrogen from the transistor 550 to the transistor 500 can be suppressed.

[0213] A wiring layer may be provided over the insulator 374 and the conductor 376. For example, in FIG. 20, an insulator 380, an insulator 382, and an insulator 384 are stacked in this order. A conductor 386 is formed in the insulator 380, the insulator 382, and the insulator 384. The conductor 386 functions as a plug or wiring. The conductor 386 can be formed using a material similar to that of the conductor 328 and the conductor 330.

[0214] Note that, for example, the insulator 380 preferably uses an insulator having a barrier property against hydrogen, similar to the insulator 324. The conductor 386 preferably includes a conductor having a barrier property against hydrogen. In particular, a conductor having a barrier property against hydrogen is formed in an opening of the insulator 380 having a barrier property against hydrogen. With this structure, the transistor 550 and the transistor 500 can be separated by a barrier layer, and diffusion of hydrogen from the transistor 550 to the transistor 500 can be suppressed.

[0215] Although the above describes a wiring layer including the conductor 356, a wiring layer including the conductor 366, a wiring layer including the conductor 376, and a wiring layer including the conductor 386, the semiconductor device according to this embodiment is not limited to this. There may be three or fewer wiring layers similar to the wiring layer including the conductor 356, or there may be five or more wiring layers similar to the wiring layer including the conductor 356.

[0216] An insulator 510, an insulator 512, an insulator 514, and an insulator 516 are stacked in this order on the insulator 384. Any of the insulator 510, the insulator 512, the insulator 514, and the insulator 516 is preferably made of a substance that has a barrier property against oxygen and hydrogen.

[0217] For example, the insulator 510 and the insulator 514 are preferably formed using a film having a barrier property against hydrogen and impurities in a region from the substrate 311 or a region where the transistor 550 is provided to a region where the transistor 500 is provided. Therefore, a material similar to that of the insulator 324 can be used.

[0218] An example of a film having a barrier property against hydrogen is silicon nitride formed by a CVD method. Here, hydrogen diffusion into a semiconductor element including an oxide semiconductor, such as the transistor 500, may degrade the characteristics of the semiconductor element. Therefore, a film that suppresses hydrogen diffusion is preferably used between the transistor 500 and the transistor 550.

[0219] As a film having a barrier property against hydrogen, for example, the insulators 510 and 514 are preferably made of a metal oxide such as aluminum oxide, hafnium oxide, or tantalum oxide.

[0220] In particular, aluminum oxide has a high blocking effect against both oxygen and impurities such as hydrogen and moisture, which can cause fluctuations in the electrical characteristics of a transistor. Therefore, aluminum oxide can prevent impurities such as hydrogen and moisture from entering the transistor 500 during and after the transistor manufacturing process. Furthermore, aluminum oxide can suppress the release of oxygen from the oxide that constitutes the transistor 500. Therefore, aluminum oxide is suitable for use as a protective film for the transistor 500.

[0221] For example, the insulator 512 and the insulator 516 can be made of a material similar to that of the insulator 320. By using a material with a relatively low dielectric constant for these insulators, the parasitic capacitance generated between wirings can be reduced. For example, the insulators 512 and 516 can be made of a silicon oxide film or a silicon oxynitride film.

[0222] A conductor 518, a conductor constituting the transistor 500 (for example, the conductor 503), and the like are embedded in the insulators 510, 512, 514, and 516. The conductor 518 functions as a plug or wiring connected to the capacitor 600 or the transistor 550. The conductor 518 can be formed using a material similar to that of the conductor 328 and the conductor 330.

[0223] In particular, the conductor 518 in the region in contact with the insulator 510 and the insulator 514 is preferably a conductor having a barrier property against oxygen, hydrogen, and water. With this structure, the transistor 550 and the transistor 500 can be separated by a layer having a barrier property against oxygen, hydrogen, and water, and diffusion of hydrogen from the transistor 550 to the transistor 500 can be suppressed.

[0224] Above the insulator 516 is the transistor 500 .

[0225] As shown in Figures 21A and 21B, transistor 500 has conductor 503 arranged so as to be embedded in insulator 514 and insulator 516, insulator 522 arranged on insulator 516 and conductor 503, insulator 524 arranged on insulator 522, oxide 530a arranged on insulator 524, oxide 530b arranged on oxide 530a, conductors 542a and 542b arranged spaced apart from each other on oxide 530b, insulator 580 arranged on conductors 542a and 542b and having an opening formed therein overlapping with conductors 542a and 542b, insulator 545 arranged on the bottom and side surfaces of the opening, and conductor 560 arranged on the surface on which insulator 545 is formed.

[0226] 21A and 21B, it is preferable that insulator 544 be disposed between oxide 530a, oxide 530b, conductor 542a, and conductor 542b and insulator 580. It is preferable that conductor 560 have conductor 560a disposed inside insulator 545 and conductor 560b disposed so as to be embedded inside conductor 560a. It is preferable that insulator 574 be disposed on insulator 580, conductor 560, and insulator 545, as shown in FIGS.

[0227] In this specification and other documents, oxide 530a and oxide 530b may be collectively referred to as oxide 530.

[0228] Note that although the transistor 500 has a structure in which two layers of the oxide 530a and the oxide 530b are stacked in and around the channel formation region, the present invention is not limited to this. For example, a single layer of the oxide 530b or a stacked structure of three or more layers may be used.

[0229] Although the transistor 500 has a two-layer structure in which the conductor 560 is stacked, the present invention is not limited to this. For example, the conductor 560 may have a single-layer structure or a stacked structure of three or more layers. The transistor 500 shown in FIGS. 20, 21A, and 21B is merely an example and is not limited to this structure. An appropriate transistor may be used depending on the circuit configuration, driving method, and the like.

[0230] Here, the conductor 560 functions as the gate electrode of the transistor, and the conductors 542a and 542b function as the source and drain electrodes, respectively. As described above, the conductor 560 is formed so as to be embedded in the opening of the insulator 580 and in the region sandwiched between the conductors 542a and 542b. The arrangement of the conductors 560, 542a, and 542b is selected in a self-aligned manner with respect to the opening of the insulator 580. That is, in the transistor 500, the gate electrode can be positioned between the source and drain electrodes in a self-aligned manner. Therefore, the conductor 560 can be formed without providing a margin for alignment, thereby reducing the area occupied by the transistor 500. This allows for miniaturization and high integration of semiconductor devices.

[0231] Furthermore, since the conductor 560 is formed in a self-aligned manner in the region between the conductor 542a and the conductor 542b, the conductor 560 does not have a region that overlaps with the conductor 542a or the conductor 542b. This reduces the parasitic capacitance formed between the conductor 560 and the conductor 542a and between the conductor 560 and the conductor 542b. This improves the switching speed of the transistor 500 and provides high frequency characteristics.

[0232] The conductor 560 may function as a first gate (also referred to as a top gate) electrode. The conductor 503 may function as a second gate (also referred to as a bottom gate) electrode. In this case, the threshold voltage of the transistor 500 can be controlled by changing the potential applied to the conductor 503 independently of the potential applied to the conductor 560. In particular, applying a negative potential to the conductor 503 can increase the threshold voltage of the transistor 500 and reduce the off-state current. Therefore, applying a negative potential to the conductor 503 can reduce the drain current when the potential applied to the conductor 560 is 0 V compared to not applying a negative potential to the conductor 503.

[0233] The conductor 503 is arranged to overlap the oxide 530 and the conductor 560. In this way, when a potential is applied to the conductor 560 and the conductor 503, the electric field generated from the conductor 560 and the electric field generated from the conductor 503 are connected, and a channel formation region formed in the oxide 530 can be covered.

[0234] In this specification and the like, a transistor configuration in which a channel formation region is electrically surrounded by the electric field of a pair of gate electrodes (a first gate electrode and a second gate electrode) is called a surrounded channel (S-channel) configuration. The S-channel configuration disclosed in this specification and the like differs from the fin type configuration and the planar type configuration. By adopting the S-channel configuration, the transistor can be made more resistant to the short channel effect, in other words, less susceptible to the short channel effect.

[0235] The conductor 503 has a structure similar to that of the conductor 518, in which the conductor 503a is formed in contact with the inner walls of the openings of the insulators 514 and 516, and the conductor 503b is formed further inward. Note that although the transistor 500 has a structure in which the conductors 503a and 503b are stacked, the present invention is not limited to this. For example, the conductor 503 may have a single layer structure or a stacked structure of three or more layers.

[0236] Here, the conductor 503a is preferably made of a conductive material that has the function of suppressing the diffusion of impurities such as hydrogen atoms, hydrogen molecules, water molecules, and copper atoms (the impurities are less likely to permeate). Alternatively, it is preferably made of a conductive material that has the function of suppressing the diffusion of oxygen (for example, at least one of oxygen atoms, oxygen molecules, etc.) (the oxygen is less likely to permeate). In this specification, the function of suppressing the diffusion of impurities or oxygen refers to the function of suppressing the diffusion of any one or all of the impurities and oxygen.

[0237] For example, the conductor 503a has a function of suppressing the diffusion of oxygen, so that the conductor 503b can be prevented from being oxidized and its conductivity from decreasing.

[0238] Furthermore, when the conductor 503 also functions as a wiring, it is preferable that the conductor 503b be made of a highly conductive material containing tungsten, copper, or aluminum as a main component. Note that, although the conductor 503 is illustrated in this embodiment as a stack of the conductors 503a and 503b, the conductor 503 may have a single-layer structure.

[0239] The insulator 522 and the insulator 524 function as a second gate insulating film.

[0240] Here, the insulator 524 in contact with the oxide 530 preferably contains more oxygen than the oxygen required for the stoichiometric composition. The oxygen is easily released from the film by heating. In this specification and elsewhere, oxygen released by heating may be referred to as "excess oxygen." In other words, the insulator 524 preferably has a region containing excess oxygen (also referred to as an "excess oxygen region"). By providing such an insulator containing excess oxygen in contact with the oxide 530, oxygen vacancies (V O When hydrogen enters the oxygen vacancy in the oxide 530, the defect (hereinafter referred to as V O H.) functions as a donor and may generate electrons as carriers. In addition, some of the hydrogen may bond with oxygen that is bonded to a metal atom to generate electrons as carriers. Therefore, a transistor using an oxide semiconductor containing a large amount of hydrogen is likely to have normally-on characteristics. Furthermore, hydrogen in an oxide semiconductor is easily moved by stress such as heat or an electric field. Therefore, if an oxide semiconductor contains a large amount of hydrogen, the reliability of the transistor may be reduced. In one embodiment of the present invention, V in the oxide 530 O It is preferable to reduce H as much as possible to obtain high-purity intrinsic or substantially high-purity intrinsic V.O To obtain an oxide semiconductor with sufficiently reduced H, it is important to remove impurities such as moisture and hydrogen from the oxide semiconductor (also called "dehydration" or "dehydrogenation treatment") and to supply oxygen to the oxide semiconductor to compensate for oxygen vacancies (also called "oxygenation treatment"). O By using an oxide semiconductor in which H and other elements are sufficiently reduced for a channel formation region of a transistor, stable electrical characteristics can be obtained.

[0241] Specifically, it is preferable to use an oxide material from which a portion of oxygen is released by heating as an insulator having an excess oxygen region. The oxide material from which oxygen is released by heating is an oxide material from which the amount of released oxygen converted to oxygen atoms is 1.0 × 10 in TDS (Thermal Desorption Spectroscopy) analysis. 18 atoms / cm 3 or more, preferably 1.0 × 10 19 atoms / cm 3 More preferably, 2.0 × 10 19 atoms / cm 3 or more, or 3.0 x 10 20 atoms / cm 3 The oxide film is one having the above properties. The surface temperature of the film during the TDS analysis is preferably in the range of 100°C or higher and 700°C or lower, or 100°C or higher and 400°C or lower.

[0242] Alternatively, the oxide 530 may be brought into contact with the insulator having the excess oxygen region and subjected to one or more of heat treatment, microwave treatment, and RF treatment. By performing such treatment, water or hydrogen in the oxide 530 can be removed. For example, a reaction occurs in the oxide 530 that breaks the VOH bond, in other words, "V O The reaction "H → Vo + H" occurs, resulting in dehydrogenation. Some of the generated hydrogen may combine with oxygen to form HO, which may be removed from the oxide 530 or an insulator near the oxide 530. Some of the hydrogen may also be gettered to the conductor 542.

[0243] The microwave treatment is preferably performed using, for example, an apparatus having a power source for generating high-density plasma or an apparatus having a power source for applying RF to the substrate side. For example, high-density oxygen radicals can be generated by using an oxygen-containing gas and high-density plasma, and the oxygen radicals generated by the high-density plasma can be efficiently introduced into the oxide 530 or an insulator near the oxide 530 by applying RF to the substrate side. The microwave treatment is performed at a pressure of 133 Pa or higher, preferably 200 Pa or higher, and more preferably 400 Pa or higher. The gases introduced into the microwave treatment apparatus may be, for example, oxygen and argon, with an oxygen flow ratio (O2 / (O2+Ar)) of 50% or less, preferably 10% to 30%.

[0244] During the manufacturing process of the transistor 500, heat treatment is preferably performed with the surface of the oxide 530 exposed. The heat treatment may be performed, for example, at a temperature of 100° C. to 450° C., more preferably 350° C. to 400° C. Note that the heat treatment is performed in a nitrogen gas or inert gas atmosphere, or an atmosphere containing an oxidizing gas at 10 ppm or more, 1% or more, or 10% or more. For example, the heat treatment is preferably performed in an oxygen atmosphere. This supplies oxygen to the oxide 530, thereby eliminating oxygen vacancies (V O ) can be reduced. The heat treatment may be performed under reduced pressure. Alternatively, the heat treatment may be performed in an atmosphere containing 10 ppm or more, 1% or more, or 10% or more of an oxidizing gas after the heat treatment in a nitrogen gas or inert gas atmosphere to compensate for the desorbed oxygen. Alternatively, the heat treatment may be performed in an atmosphere containing 10 ppm or more, 1% or more, or 10% or more of an oxidizing gas, and then the heat treatment may be performed in a nitrogen gas or inert gas atmosphere.

[0245] By subjecting the oxide 530 to oxygen addition treatment, the oxygen vacancies in the oxide 530 can be repaired by the supplied oxygen, in other words, the reaction "Vo + O → null" can be promoted. Furthermore, the supplied oxygen reacts with the hydrogen remaining in the oxide 530, and the hydrogen can be removed as H2O (dehydration). As a result, the hydrogen remaining in the oxide 530 recombines with the oxygen vacancies to form V O The formation of H can be suppressed.

[0246] When the insulator 524 has an excess oxygen region, the insulator 522 preferably has a function of suppressing the diffusion of oxygen (for example, oxygen atoms, oxygen molecules, etc.) (preferably making the oxygen less permeable).

[0247] The insulator 522 preferably has a function of suppressing diffusion of oxygen and impurities, which prevents oxygen contained in the oxide 530 from diffusing toward the conductor 503. Furthermore, reaction of the conductor 503 with oxygen contained in the insulator 524 or the oxide 530 can be suppressed.

[0248] The insulator 522 is preferably a single-layer or multi-layer insulator containing a high-k material, such as aluminum oxide, hafnium oxide, oxide containing aluminum and hafnium (hafnium aluminate), tantalum oxide, zirconium oxide, lead zirconate titanate (PZT), strontium titanate (SrTiO3), or (Ba,Sr)TiO3 (BST). As transistors become smaller and more highly integrated, thinner gate insulating films can cause problems such as leakage current. Using a high-k material for the insulator that functions as the gate insulating film allows for a reduction in the gate potential during transistor operation while maintaining the physical film thickness.

[0249] In particular, an insulator containing an oxide of one or both of aluminum and hafnium, which is an insulating material that has the function of suppressing the diffusion of impurities and oxygen (i.e., is difficult for oxygen to permeate), is preferably used. As an insulator containing an oxide of one or both of aluminum and hafnium, aluminum oxide, hafnium oxide, or an oxide containing aluminum and hafnium (hafnium aluminate) is preferably used. When the insulator 522 is formed using such a material, the insulator 522 functions as a layer that suppresses oxygen release from the oxide 530 and the intrusion of impurities such as hydrogen into the oxide 530 from the periphery of the transistor 500.

[0250] Alternatively, for example, aluminum oxide, bismuth oxide, germanium oxide, niobium oxide, silicon oxide, titanium oxide, tungsten oxide, yttrium oxide, or zirconium oxide may be added to these insulators. Alternatively, these insulators may be nitrided. Silicon oxide, silicon oxynitride, or silicon nitride may be stacked on the above insulators.

[0251] 21A and 21B illustrate the insulators 522 and 524 as the second gate insulating film having a three-layer stack structure, the second gate insulating film may have a single-layer, two-layer, or four or more-layer stack structure. In this case, the second gate insulating film is not limited to a stack structure made of the same material, and may have a stack structure made of different materials.

[0252] The transistor 500 uses a metal oxide functioning as an oxide semiconductor for the oxide 530 including the channel formation region. For example, the oxide 530 may be a metal oxide such as In-M-Zn oxide (wherein M is one or more elements selected from aluminum, gallium, yttrium, copper, vanadium, beryllium, boron, titanium, iron, nickel, germanium, zirconium, molybdenum, lanthanum, cerium, neodymium, hafnium, tantalum, tungsten, magnesium, or the like).

[0253] The metal oxide functioning as an oxide semiconductor may be formed by a sputtering method or an ALD (Atomic Layer Deposition) method. Note that the metal oxide functioning as an oxide semiconductor will be described in detail in other embodiments.

[0254] The metal oxide that functions as a channel formation region in the oxide 530 preferably has a band gap of 2 eV or more, preferably 2.5 eV or more. By using a metal oxide with a wide band gap, the off-state current of the transistor can be reduced.

[0255] The oxide 530 has the oxide 530a below the oxide 530b, and thus can suppress the diffusion of impurities from components formed below the oxide 530a to the oxide 530b.

[0256] Note that oxide 530 preferably has a stacked structure of multiple oxide layers with different atomic ratios of each metal atom. Specifically, the atomic ratio of element M among the constituent elements in the metal oxide used for oxide 530a is preferably greater than the atomic ratio of element M among the constituent elements in the metal oxide used for oxide 530b. Furthermore, the atomic ratio of element M to In in the metal oxide used for oxide 530a is preferably greater than the atomic ratio of element M to In in the metal oxide used for oxide 530b. Furthermore, the atomic ratio of In to element M in the metal oxide used for oxide 530b is preferably greater than the atomic ratio of In to element M in the metal oxide used for oxide 530a.

[0257] The energy of the conduction band minimum of the oxide 530a is preferably higher than that of the oxide 530b, or in other words, the electron affinity of the oxide 530a is preferably smaller than that of the oxide 530b.

[0258] Here, the energy level of the conduction band minimum changes gradually at the junction between the oxide 530a and the oxide 530b. In other words, the energy level of the conduction band minimum at the junction between the oxide 530a and the oxide 530b changes continuously or forms a continuous junction. To achieve this, it is preferable to reduce the defect level density of the mixed layer formed at the interface between the oxide 530a and the oxide 530b.

[0259] Specifically, when the oxide 530a and the oxide 530b have a common element (main component) other than oxygen, a mixed layer with a low density of defect states can be formed. For example, when the oxide 530b is an In-Ga-Zn oxide, the oxide 530a may be an In-Ga-Zn oxide, a Ga-Zn oxide, a gallium oxide, or the like.

[0260] In this case, the oxide 530b serves as the main carrier path. By configuring the oxide 530a as described above, the defect state density at the interface between the oxide 530a and the oxide 530b can be reduced. As a result, the influence of interface scattering on carrier conduction is reduced, and the transistor 500 can obtain a high on-state current.

[0261] Conductors 542a and 542b, which function as a source electrode and a drain electrode, are provided on oxide 530b. Conductors 542a and 542b are preferably made of a metal element selected from aluminum, chromium, copper, silver, gold, platinum, tantalum, nickel, titanium, molybdenum, tungsten, hafnium, vanadium, niobium, manganese, magnesium, zirconium, beryllium, indium, ruthenium, iridium, strontium, and lanthanum, or an alloy containing any of the above metal elements or an alloy combining any of the above metal elements. For example, tantalum nitride, titanium nitride, tungsten, a nitride containing titanium and aluminum, a nitride containing tantalum and aluminum, ruthenium oxide, ruthenium nitride, an oxide containing strontium and ruthenium, or an oxide containing lanthanum and nickel is preferably used. In addition, tantalum nitride, titanium nitride, nitrides containing titanium and aluminum, nitrides containing tantalum and aluminum, ruthenium oxide, ruthenium nitride, oxides containing strontium and ruthenium, and oxides containing lanthanum and nickel are preferred because they are conductive materials that are resistant to oxidation or materials that maintain conductivity even when absorbing oxygen.Furthermore, metal nitride films such as tantalum nitride are preferred because they have barrier properties against hydrogen or oxygen.

[0262] 21A shows the conductor 542a and the conductor 542b as a single layer, but they may be stacked with two or more layers. For example, a tantalum nitride film and a tungsten film may be stacked. Alternatively, a titanium film and an aluminum film may be stacked. Alternatively, a two-layer structure in which an aluminum film is stacked on a tungsten film, a two-layer structure in which a copper film is stacked on a copper-magnesium-aluminum alloy film, a two-layer structure in which a copper film is stacked on a titanium film, or a two-layer structure in which a copper film is stacked on a tungsten film may be used.

[0263] Other examples include a three-layer structure in which a titanium film or titanium nitride film is laminated on the titanium film or titanium nitride film, an aluminum film or copper film is laminated on the titanium film or titanium nitride film, and a titanium film or titanium nitride film is further formed thereon, and a three-layer structure in which a molybdenum film or molybdenum nitride film is laminated on the molybdenum film or molybdenum nitride film, an aluminum film or copper film is laminated on the molybdenum film or molybdenum nitride film, and a molybdenum film or molybdenum nitride film is further formed thereon. Note that a transparent conductive material containing indium oxide, tin oxide, or zinc oxide may also be used.

[0264] 21A, regions 543a and 543b may be formed as low-resistance regions at and near the interface of the oxide 530 with the conductor 542a (conductor 542b). In this case, the region 543a functions as one of the source region and the drain region, and the region 543b functions as the other of the source region and the drain region. A channel formation region is formed in the region sandwiched between the regions 543a and 543b.

[0265] By providing the conductor 542a (conductor 542b) so as to be in contact with the oxide 530, the oxygen concentration in the region 543a (region 543b) may be reduced. Also, a metal compound layer containing the metal contained in the conductor 542a (conductor 542b) and components of the oxide 530 may be formed in the region 543a (region 543b). In such a case, the carrier density in the region 543a (region 543b) increases, and the region 543a (region 543b) becomes a low-resistance region.

[0266] The insulator 544 is provided to cover the conductors 542a and 542b and suppresses oxidation of the conductors 542a and 542b. In this case, the insulator 544 may be provided to cover the side surface of the oxide 530 and to be in contact with the insulator 524.

[0267] The insulator 544 can be a metal oxide containing one or more elements selected from hafnium, aluminum, gallium, yttrium, zirconium, tungsten, titanium, tantalum, nickel, germanium, neodymium, lanthanum, magnesium, etc. Alternatively, the insulator 544 can be silicon nitride oxide, silicon nitride, or the like.

[0268] In particular, it is preferable to use, as the insulator 544, an insulator containing an oxide of either or both of aluminum and hafnium, such as aluminum oxide, hafnium oxide, or an oxide containing aluminum and hafnium (hafnium aluminate). Hafnium aluminate is particularly preferable because it has higher heat resistance than hafnium oxide film. Therefore, it is less likely to crystallize during heat treatment in a later process. Note that if the conductors 542a and 542b are made of oxidation-resistant materials or if their conductivity does not decrease significantly even when they absorb oxygen, the insulator 544 is not an essential component. It can be designed appropriately depending on the desired transistor characteristics.

[0269] The insulator 544 can prevent impurities such as water and hydrogen contained in the insulator 580 from diffusing into the oxide 530b through the insulator 545. The insulator 580 can also prevent the conductor 560 from being oxidized by excess oxygen.

[0270] The insulator 545 functions as a first gate insulating film. Like the insulator 524, the insulator 545 is preferably formed using an insulator that contains excess oxygen and releases oxygen by heating.

[0271] Specifically, silicon oxide having excess oxygen, silicon oxynitride, silicon nitride oxide, silicon nitride, silicon oxide to which fluorine is added, silicon oxide to which carbon is added, silicon oxide to which carbon and nitrogen are added, and silicon oxide having vacancies can be used. In particular, silicon oxide and silicon oxynitride are preferable because they are stable against heat.

[0272] By providing the insulator 545 as an insulator containing excess oxygen, oxygen can be effectively supplied from the insulator 545 to the channel formation region of the oxide 530b. Similarly to the insulator 524, the concentration of impurities such as water or hydrogen in the insulator 545 is preferably reduced. The thickness of the insulator 545 is preferably 1 nm or more and 20 nm or less. The microwave treatment described above may be performed before and / or after the formation of the insulator 545.

[0273] Furthermore, a metal oxide may be provided between the insulator 545 and the conductor 560 to efficiently supply excess oxygen contained in the insulator 545 to the oxide 530. The metal oxide preferably suppresses oxygen diffusion from the insulator 545 to the conductor 560. By providing a metal oxide that suppresses oxygen diffusion, the diffusion of excess oxygen from the insulator 545 to the conductor 560 is suppressed. That is, a decrease in the amount of excess oxygen supplied to the oxide 530 can be suppressed. Furthermore, oxidation of the conductor 560 due to excess oxygen can be suppressed. As the metal oxide, any material that can be used for the insulator 544 may be used.

[0274] Note that the insulator 545 may have a layered structure, similar to the second gate insulating film. As transistors become smaller and more highly integrated, thinner gate insulating films can cause problems such as leakage current. Therefore, by using a layered structure of a high-k material and a thermally stable material for the insulator that functions as the gate insulating film, it is possible to reduce the gate potential during transistor operation while maintaining the physical film thickness. Furthermore, a layered structure that is thermally stable and has a high dielectric constant can be achieved.

[0275] Although the conductor 560 functioning as the first gate electrode is shown as having a two-layer structure in FIGS. 21A and 21B, it may have a single-layer structure or a laminated structure of three or more layers.

[0276] The conductor 560a is preferably made of a conductive material that suppresses the diffusion of impurities such as hydrogen atoms, hydrogen molecules, water molecules, nitrogen atoms, nitrogen molecules, nitrogen oxide molecules (e.g., NO, NO, and the like), and copper atoms. Alternatively, a conductive material that suppresses the diffusion of oxygen (e.g., at least one of oxygen atoms, oxygen molecules, and the like) is preferably used. The oxygen-suppressing function of the conductor 560a can suppress the oxidation of the conductor 560b due to oxygen contained in the insulator 545, which can reduce the conductivity. Examples of conductive materials that suppress the diffusion of oxygen include tantalum, tantalum nitride, ruthenium, and ruthenium oxide. Alternatively, an oxide semiconductor that can be used for the oxide 530 can be used for the conductor 560a. In this case, the conductor 560b can be formed by sputtering to reduce the electrical resistance of the conductor 560a, thereby making it a conductor. This can be called an OC (Oxide Conductor) electrode.

[0277] The conductor 560b is preferably made of a conductive material containing tungsten, copper, or aluminum as a main component. Because the conductor 560b also functions as wiring, it is preferable to use a conductor with high conductivity. For example, a conductive material containing tungsten, copper, or aluminum as a main component can be used. The conductor 560b may have a layered structure, such as a layered structure of titanium or titanium nitride and the above-mentioned conductive material.

[0278] The insulator 580 is provided over the conductor 542a and the conductor 542b with the insulator 544 interposed therebetween. The insulator 580 preferably has an excess oxygen region. For example, the insulator 580 preferably includes silicon oxide, silicon oxynitride, silicon nitride oxide, silicon nitride, silicon oxide doped with fluorine, silicon oxide doped with carbon, silicon oxide doped with carbon and nitrogen, silicon oxide having voids, or a resin. Silicon oxide and silicon oxynitride are particularly preferred because they are thermally stable. Silicon oxide and silicon oxide having voids are particularly preferred because they allow for easy formation of an excess oxygen region in a later step.

[0279] The insulator 580 preferably has an excess oxygen region. By providing the insulator 580 from which oxygen is released by heating, oxygen in the insulator 580 can be efficiently supplied to the oxide 530. Note that the concentration of impurities such as water or hydrogen in the insulator 580 is preferably reduced.

[0280] The opening of the insulator 580 is formed to overlap the region between the conductor 542a and the conductor 542b, so that the conductor 560 is formed to be embedded in the opening of the insulator 580 and the region sandwiched between the conductor 542a and the conductor 542b.

[0281] When miniaturizing semiconductor devices, it is necessary to shorten the gate length, but it is also necessary to prevent the conductivity of the conductor 560 from decreasing. If the film thickness of the conductor 560 is increased to achieve this, the conductor 560 may have a shape with a high aspect ratio. In this embodiment, the conductor 560 is provided so as to be embedded in the opening of the insulator 580. Therefore, even if the conductor 560 has a shape with a high aspect ratio, the conductor 560 can be formed without collapsing during the process.

[0282] The insulator 574 is preferably provided in contact with the top surface of the insulator 580, the top surface of the conductor 560, and the top surface of the insulator 545. By forming the insulator 574 by a sputtering method, an excess oxygen region can be provided in the insulator 545 and the insulator 580. This allows oxygen to be supplied from the excess oxygen region into the oxide 530.

[0283] For example, the insulator 574 can be a metal oxide containing one or more selected from hafnium, aluminum, gallium, yttrium, zirconium, tungsten, titanium, tantalum, nickel, germanium, magnesium, and the like.

[0284] In particular, aluminum oxide has high barrier properties and can suppress the diffusion of hydrogen and nitrogen even when it is a thin film with a thickness of 0.5 nm to 3.0 nm. Therefore, aluminum oxide formed by sputtering can function as both an oxygen source and a barrier film against impurities such as hydrogen.

[0285] An insulator 581 functioning as an interlayer film is preferably provided over the insulator 574. Like the insulator 524, the insulator 581 preferably has a reduced concentration of impurities such as water or hydrogen.

[0286] Furthermore, conductors 540a and 540b are arranged in openings formed in insulators 581, 574, 580, and 544. Conductor 540a and 540b are arranged opposite each other with conductor 560 interposed therebetween. Conductor 540a and 540b have the same configuration as conductors 546 and 548, which will be described later.

[0287] An insulator 582 is provided over the insulator 581. The insulator 582 is preferably formed using a substance that has a barrier property against oxygen and hydrogen. Therefore, the insulator 582 can be formed using a material similar to that of the insulator 514. For example, the insulator 582 is preferably formed using a metal oxide such as aluminum oxide, hafnium oxide, or tantalum oxide.

[0288] In particular, aluminum oxide has a high blocking effect against both oxygen and impurities such as hydrogen and moisture, which can cause fluctuations in the electrical characteristics of a transistor. Therefore, aluminum oxide can prevent impurities such as hydrogen and moisture from entering the transistor 500 during and after the transistor manufacturing process. Furthermore, aluminum oxide can suppress the release of oxygen from the oxide that constitutes the transistor 500. Therefore, aluminum oxide is suitable for use as a protective film for the transistor 500.

[0289] An insulator 586 is provided over the insulator 582. The insulator 586 can be formed using a material similar to that of the insulator 320. By using a material with a relatively low dielectric constant for these insulators, parasitic capacitance between wirings can be reduced. For example, a silicon oxide film, a silicon oxynitride film, or the like can be used as the insulator 586.

[0290] Furthermore, conductors 546, 548, etc. are embedded in insulators 522, 524, 544, 580, 574, 581, 582, and 586.

[0291] The conductor 546 and the conductor 548 function as plugs or wirings that connect to the capacitor 600, the transistor 500, or the transistor 550. The conductor 546 and the conductor 548 can be formed using a material similar to that of the conductor 328 and the conductor 330.

[0292] After the transistor 500 is formed, an opening may be formed to surround the transistor 500, and an insulator with high barrier properties against hydrogen or water may be formed to cover the opening. By surrounding the transistor 500 with the insulator with high barrier properties, it is possible to prevent moisture and hydrogen from entering from the outside. Alternatively, multiple transistors 500 may be collectively surrounded by an insulator with high barrier properties against hydrogen or water. When forming an opening to surround the transistor 500, for example, it is preferable to form an opening that reaches the insulator 522 or the insulator 514 and form the insulator with high barrier properties in contact with the insulator 522 or the insulator 514, because this can serve as part of the manufacturing process of the transistor 500. For example, the insulator with high barrier properties against hydrogen or water may be made of a material similar to that of the insulator 522 or the insulator 514.

[0293] Subsequently, a capacitor 600 is provided above the transistor 500. The capacitor 600 includes a conductor 610, a conductor 620, and an insulator 630.

[0294] A conductor 612 may be provided over the conductor 546 and the conductor 548. The conductor 612 functions as a plug or a wiring connected to the transistor 500. The conductor 610 functions as an electrode of the capacitor 600. Note that the conductor 612 and the conductor 610 can be formed at the same time.

[0295] A metal film containing an element selected from molybdenum, titanium, tantalum, tungsten, aluminum, copper, chromium, neodymium, and scandium, or a metal nitride film containing any of the above elements (tantalum nitride film, titanium nitride film, molybdenum nitride film, tungsten nitride film), etc. can be used for the conductor 612 and the conductor 610. Alternatively, a conductive material such as indium tin oxide, indium oxide containing tungsten oxide, indium zinc oxide containing tungsten oxide, indium oxide containing titanium oxide, indium tin oxide containing titanium oxide, indium zinc oxide, or indium tin oxide with silicon oxide added can also be used.

[0296] In this embodiment, the conductor 612 and the conductor 610 have a single-layer structure, but the present invention is not limited to this structure and may have a stacked structure of two or more layers. For example, a conductor having a barrier property and a conductor having high adhesion to the conductor having high conductivity may be formed between a conductor having a barrier property and a conductor having high conductivity.

[0297] The conductor 620 is provided so as to overlap with the conductor 610 with the insulator 630 interposed therebetween. Note that the conductor 620 can be formed using a conductive material such as a metal material, an alloy material, or a metal oxide material. It is preferable to use a high-melting-point material such as tungsten or molybdenum that has both heat resistance and conductivity, and tungsten is particularly preferable. Furthermore, when the conductor 620 is formed simultaneously with other components such as a conductor, a low-resistance metal material such as Cu (copper) or Al (aluminum) can be used.

[0298] An insulator 640 is provided over the conductor 620 and the insulator 630. The insulator 640 can be provided using a material similar to that of the insulator 320. The insulator 640 may also function as a planarizing film that covers the uneven shape underneath.

[0299] With this structure, miniaturization or high integration can be achieved in a semiconductor device including a transistor including an oxide semiconductor.

[0300] The configurations, structures, methods, and the like described in this embodiment can be used in appropriate combination with the configurations, structures, methods, and the like described in other embodiment modes and examples.

[0301] (Embodiment 6) In this embodiment, the configuration of an integrated circuit including each component of the arithmetic processing system 100 described in the above embodiment will be described with reference to FIGS. 22A and 22B.

[0302] 22A is an example of a schematic diagram for explaining an integrated circuit including each component of the arithmetic processing system 100. The integrated circuit 390 shown in FIG. 22A can be formed as a single integrated circuit by integrating each circuit of the CPU 110 and the accelerator described as the semiconductor device 10, by configuring some of the circuits of the accelerator using OS transistors.

[0303] As shown in FIG. 22A, the CPU 110 may be configured to include a backup circuit 222 in a layer having OS transistors above the CPU core 200. Also, as shown in FIG. 22A, in the accelerator described as the semiconductor device 10, a memory circuit unit 20 may be provided in a layer having OS transistors above a layer having Si transistors constituting the arithmetic circuit 30 and the switching circuit 40. Alternatively, a driver circuit 50 may be provided in a layer having Si transistors, and an OS memory 300N may be provided in a layer having OS transistors. As the OS memory 300N, in addition to the NOSRAM described in the above embodiment, a DOSRAM may be used. Furthermore, in the OS memory 300N, stacking a layer having OS transistors on a driver circuit provided in a layer having Si transistors can improve memory density.

[0304] As shown in FIG. 22A, in the case of an SoC in which circuits such as the CPU 110, the accelerator described as the semiconductor device 10, and the OS memory 300N are tightly coupled, there is a problem of heat generation. However, OS transistors are preferable because the amount of change in electrical characteristics due to heat is smaller than that of Si transistors. Furthermore, by integrating circuits in a three-dimensional direction as shown in FIG. 22A, parasitic capacitance can be reduced compared to stacked structures using silicon through electrodes (TSVs). The power consumption required for charging and discharging each wiring can be reduced. As a result, the efficiency of computational processing can be improved.

[0305] FIG. 22B shows an example of a semiconductor chip incorporating an integrated circuit 390. The semiconductor chip 391 shown in FIG. 22B has leads 392 and an integrated circuit 390. As described with reference to FIG. 22A, the integrated circuit 390 has the various circuits described in the above embodiments provided on a single die. The integrated circuit 390 has a layered structure and is broadly divided into a layer having Si transistors (Si transistor layer 393), a wiring layer 394, and a layer having OS transistors (OS transistor layer 395). The OS transistor layer 395 can be provided by being layered on the Si transistor layer 393, which facilitates miniaturization of the semiconductor chip 391.

[0306] 22B, a QFP (Quad Flat Package) is used for the package of the semiconductor chip 391, but the package form is not limited to this. Other examples of configurations that can be used as appropriate include an insertion mounting type DIP (Dual In-line Package) and PGA (Pin Grid Array), a surface mounting type SOP (Small Outline Package), SSOP (Shrink Small Outline Package), TSOP (Thin-Small Outline Package), LCC (Leaded Chip Carrier), QFN (Quad Flat Non-leaded package), BGA (Ball Grid Array), FBGA (Fine pitch Ball Grid Array), and a contact mounting type DTP (Dual Tape carrier Package) and QTP (Quad Tape-carrier Package).

[0307] The arithmetic circuit and switching circuit having Si transistors and the memory circuit having OS transistors can all be formed in the Si transistor layer 393, the wiring layer 394, and the OS transistor layer 395. That is, the elements constituting the semiconductor device can be formed using the same manufacturing process. Therefore, even if the number of constituent elements increases, the IC shown in FIG. 22B does not require an increase in the manufacturing process, and the semiconductor device can be incorporated at low cost.

[0308] According to the above-described embodiment of the present invention, a novel semiconductor device and electronic device can be provided. Alternatively, according to one embodiment of the present invention, a semiconductor device and electronic device with low power consumption can be provided. Alternatively, according to one embodiment of the present invention, a semiconductor device and electronic device in which heat generation can be suppressed can be provided.

[0309] This embodiment mode can be combined with the descriptions of other embodiment modes as appropriate.

[0310] (Embodiment 7) In this embodiment mode, electronic devices, mobile objects, and arithmetic systems to which the integrated circuit 390 described in the above embodiment mode can be applied will be described with reference to FIGS.

[0311] Fig. 23A shows an external view of an automobile as an example of a moving body. Fig. 23B is a simplified diagram of data exchange within the automobile. The automobile 590 has a plurality of cameras 591 and the like. The automobile 590 also has various sensors (not shown) such as infrared radar, millimeter-wave radar, and laser radar.

[0312] In an automobile 590, the above-mentioned integrated circuit 390 (or a semiconductor chip 391 incorporating the above-mentioned integrated circuit 390) can be used in a camera 591 or the like. The automobile 590 processes a plurality of images acquired by the camera 591 in a plurality of imaging directions 592 using the integrated circuit 390 described in the above embodiment, and analyzes the plurality of images collectively using a host controller 594 or the like via a bus 593 or the like, thereby determining the surrounding traffic conditions, such as the presence or absence of guardrails or pedestrians, and can perform autonomous driving. The automobile 590 can also be used in systems that provide road guidance, hazard prediction, and the like.

[0313] In the integrated circuit 390, the obtained image data is subjected to arithmetic processing such as neural networks, making it possible to perform processes such as increasing the image resolution, reducing image noise, facial recognition (for security purposes, etc.), object recognition (for autonomous driving purposes, etc.), image compression, image correction (wide dynamic range), image restoration for lensless image sensors, positioning, character recognition, and reducing reflected glare.

[0314] Although an automobile is described above as an example of a moving body, the moving body is not limited to an automobile. For example, moving bodies may include trains, monorails, ships, and flying bodies (helicopters, unmanned aerial vehicles (drones), airplanes, and rockets). A computer according to one embodiment of the present invention may be applied to these moving bodies to provide a system using artificial intelligence.

[0315] Fig. 24A is an external view showing an example of a portable electronic device. Fig. 24B is a simplified diagram showing data exchange within the portable electronic device. Portable electronic device 595 has printed circuit board 596, speaker 597, camera 598, microphone 599, etc.

[0316] In portable electronic device 595, the integrated circuit 390 can be provided on printed circuit board 596. Portable electronic device 595 can improve user convenience by processing and analyzing a plurality of pieces of data obtained by speaker 597, camera 598, microphone 599, etc. using integrated circuit 390 described in the above embodiment. In addition, the portable electronic device 595 can be used in systems that perform voice guidance, image search, etc.

[0317] In the integrated circuit 390, the obtained image data is subjected to arithmetic processing such as neural networks, making it possible to perform processes such as increasing the image resolution, reducing image noise, facial recognition (for security purposes, etc.), object recognition (for autonomous driving purposes, etc.), image compression, image correction (wide dynamic range), image restoration for lensless image sensors, positioning, character recognition, and reducing reflected glare.

[0318] 25A includes a housing 1101, a housing 1102, a housing 1103, a display unit 1104, a connection unit 1105, operation keys 1107, and the like. The housings 1101, 1102, and 1103 are detachable. By attaching the connection unit 1105 provided on the housing 1101 to the housing 1108, a video image displayed on the display unit 1104 can be output to another video device. On the other hand, by attaching the housings 1102 and 1103 to the housing 1109, the housings 1102 and 1103 are integrated and function as an operation unit. The integrated circuit 390 described in the above embodiment can be incorporated into a chip provided on a substrate of the housing 1102 or the housing 1103.

[0319] 25B shows a stick-type electronic device 1120 that is USB-connected. The electronic device 1120 has a housing 1121, a cap 1122, a USB connector 1123, and a board 1124. The board 1124 is housed in the housing 1121. For example, a memory chip 1125 and a controller chip 1126 are attached to the board 1124. The integrated circuit 390 shown in the previous embodiment can be incorporated into the controller chip 1126 of the board 1124.

[0320] 25C shows a humanoid robot 1130. The robot 1130 has sensors 2101 to 2106 and a control circuit 2110. For example, the control circuit 2110 can incorporate the integrated circuit 390 shown in the previous embodiment.

[0321] The integrated circuit 390 described in the above embodiment can be used in a server that communicates with the electronic device instead of being built into the electronic device. In this case, the electronic device and the server constitute a computing system. Figure 26 shows an example of the configuration of a system 3000.

[0322] The system 3000 is configured by an electronic device 3001 and a server 3002. Communication between the electronic device 3001 and the server 3002 can be performed via an internet line 3003.

[0323] Server 3002 has a plurality of racks 3004. A plurality of circuit boards 3005 are provided in the racks, and integrated circuits 390 described in the above embodiment can be mounted on the circuit boards 3005. This forms a neural network in server 3002. Server 3002 can perform neural network calculations using data input from electronic device 3001 via internet line 3003. The results of calculations by server 3002 can be transmitted to electronic device 3001 via internet line 3003 as necessary. This reduces the calculation load on electronic device 3001.

[0324] This embodiment mode can be combined with the descriptions of other embodiment modes as appropriate.

[0325] (Embodiment 8) In this embodiment, an example of the configuration of weight data used in convolution calculation processing in a convolutional neural network (hereinafter referred to as CNN) or the like in an integrated circuit 390 having a semiconductor device 10 will be described with reference to Figures 27 and 28.

[0326] 27A is a conceptual diagram showing how weight data, which are connection parameters of CNN, are generated by inputting learning (training) data. TR , training data D TR 27A shows a computer device 32 to which the learning data D TR For the weight data 34 (W TR ) is used to perform processing such as product-sum operations 33A and processing such as activation functions 33B, and the convolutional data D for learning is obtained. CT is illustrated.

[0327] Training data D TR The weight data 34 (W) corresponds to voice data, image data, text data, etc. Each data is preferably standardized to a data size and format suitable for the content of machine learning so that it can be easily processed in the computer device 32. TR ) is the training data D TR The training data D is generated by calculations such as backpropagation. TR The computer device 32 that processes the learning data D is a stationary type that can be supplied with stable power, and therefore can execute computations that consume a lot of power using a huge memory and a computing device with high computing performance. TR The number of bits of the weight data 34 (W TR) can be optimized. In addition, since the bit precision of the data can affect the convergence of calculations depending on the calculation algorithm, it is preferable to be able to perform calculations with a wide range of bit numbers.

[0328] 27B is a conceptual diagram showing how CNN performs arithmetic processing, in which inference data is input and inferred data is output. In FIG. 27B, voice data uttered by a user to an electronic device 35 or image data acquired by an imaging device mounted on an automobile 36 is input as inference data D IN Inference data D IN is input to an integrated circuit 390 having the semiconductor device 10 described in the above embodiment. In the integrated circuit 390, the inference data D IN is used as input data, and the weight data 37 (W INF ) and other calculation processes such as convolution calculations are performed using the inference data D IN For the weight data 37 (W INF ) is used to perform processing such as multiplication and addition 38A and processing such as activation function 38B, and the convolution data D for inference is obtained. CI The integrated circuit 390 performs calculations including convolution calculations to generate the inferred output data D JD Output.

[0329] Inference data D IN The integrated circuit 390 performs arithmetic processing in an environment with limited processing power. Compared to the computer device 32 in FIG. 27A, only arithmetic processing that requires fewer circuit resources is performed. The integrated circuit 390 is required to perform arithmetic processing at high speed and with low power consumption in an environment with limited processing power. The semiconductor device 10 of one embodiment of the present invention can be a semiconductor device that functions as an accelerator that is small, consumes low power, or is fast. Therefore, it is suitable for use in an environment with limited processing power, such as an edge device.

[0330] In addition, the inference data D IN The number of bits of the training data D TRFor example, the number of bits of the training data D TR When the number of bits is set to a high value such as 8 bits to 64 bits, the inference data D IN is data having a low number of bits (first number of bits) of 16 bits or less, preferably 8 bits or less, preferably 4 bits or less, preferably 2 bits or less. In other words, the number of bits for inference is set to be the same as that of the learning data D TR It is preferable that the number of bits is smaller than the higher number of bits (second number of bits).

[0331] Similarly, the weight data 37 (W INF ) is the weight data 34 (W TR ), it is preferable to use data with a low bit count of 16 bits or less, preferably 8 bits or less, preferably 4 bits or less, and preferably 2 bits or less. This configuration makes it possible to perform calculations with little degradation in accuracy even in an environment with limited circuit resources where only limited memory capacity and calculation performance can be realized in calculation processing. In such a configuration, it is desirable to set the bit count according to the neural network model within conditions that minimize degradation in inference accuracy.

[0332] Weight data 34(W TR ) to weight data 37(W INF ) is converted to weight data 34 (W) by standardizing the process to maintain the relative relationship between the weight data and reducing the number of bits. TR ) to weight data 37(W INF ) can be achieved by reducing the number of bits in the exponent and / or mantissa. For example, the weight data W TR From the weight data W INF In the conversion to the weight data W with the reduced number of bits, the sign part 39A is left as it is and the number of bits of the exponent part 39B and the mantissa part 39C is reduced. INF It states that:

[0333] Also, the weight data WTR From the weight data W INF In the conversion to , the sign part 39A and the exponent part 39B are left as they are, and the number of bits of the mantissa part 39C is significantly reduced to obtain weight data W INF It states that:

[0334] As a configuration other than those shown in FIGS. 28A and 28B, the number of bits can also be reduced by converting a floating-point format such as FP32 into an integer format such as INT8.

[0335] Weight data W with reduced number of bits INF In this case, the reduction in the number of bits can cause rounding errors in the numbers and narrow the range of numbers that can be expressed. On the other hand, even if the number of bits is reduced, the magnitude relationship (relative relationship) between the weight data can be maintained, so the magnitude relationship of the output values from the convolution calculation process is maintained. Therefore, depending on the neural network model, it is possible to perform calculations with little loss in calculation accuracy. Also, in environments with limited processing power such as edge devices, weight data W with a reduced number of bits can be used. INF Inference processing using

[0336] It is also preferable to configure the neural network model so that the bit width is optimized for each layer, or that less important neurons are eliminated. This configuration can reduce the amount of calculation while preventing a decrease in calculation accuracy.

[0337] (Notes regarding the present specification) The above-described embodiment and each configuration in the embodiment will be described below with additional notes.

[0338] The configurations shown in each embodiment can be combined as appropriate with configurations shown in other embodiments or examples to form one aspect of the present invention. Furthermore, when multiple configuration examples are shown in one embodiment, the configuration examples can be combined as appropriate.

[0339] In addition, the content (or even a part of the content) described in one embodiment can be applied to, combined with, or replaced with another content (or even a part of the content) described in that embodiment, and / or with the content (or even a part of the content) described in one or more other embodiments.

[0340] The contents described in the embodiments refer to the contents described in each embodiment using various figures or the contents described using text in the specification.

[0341] Furthermore, a figure (or even a part thereof) described in one embodiment can be combined with another part of that figure, another figure (or even a part thereof) described in that embodiment, and / or a figure (or even a part thereof) described in one or more other embodiments to form even more figures.

[0342] In addition, in the present specification and the like, in the block diagrams, components are classified by function and shown as independent blocks. However, in actual circuits, etc., it is difficult to separate components by function, and there may be cases where one circuit is involved in multiple functions, or where one function is involved across multiple circuits. Therefore, the blocks in the block diagrams are not limited to the components described in the specification, but may be rephrased appropriately depending on the situation.

[0343] In addition, in the drawings, the size, layer thickness, or region is shown at an arbitrary size for convenience of explanation. Therefore, it is not necessarily limited to the scale. Note that the drawings are shown schematically for clarity, and are not limited to the shapes or values shown in the drawings. For example, it is possible to include variations in signal, voltage, or current due to noise, or variations in signal, voltage, or current due to timing deviations.

[0344] Furthermore, the positional relationships of components shown in the drawings are relative. Therefore, when describing components with reference to the drawings, terms such as "above" and "below" indicating the positional relationships may be used for convenience. The positional relationships of components are not limited to the content described in this specification, and can be rephrased appropriately depending on the situation.

[0345] In this specification and the like, when describing the connection relationship of a transistor, the term "one of the source or drain" (or first electrode or first terminal) is used, and the other of the source and drain is referred to as "the other of the source or drain" (or second electrode or second terminal). This is because the source and drain of a transistor vary depending on the structure or operating conditions of the transistor. Note that the names of the source and drain of a transistor can be appropriately changed to source (drain) terminal, source (drain) electrode, etc. depending on the situation.

[0346] Furthermore, the terms "electrode" and "wiring" used in this specification and elsewhere do not limit the functionality of these components. For example, an "electrode" may be used as part of a "wiring," and vice versa. Furthermore, the terms "electrode" and "wiring" also include cases where multiple "electrodes" or "wirings" are integrally formed.

[0347] Furthermore, in this specification and the like, voltage and potential can be interchanged as appropriate. Voltage refers to the potential difference from a reference potential. For example, if the reference potential is a ground voltage (earth voltage), voltage can be interchanged with potential. Ground potential does not necessarily mean 0 V. Note that potential is relative, and the potential applied to wiring, etc. may change depending on the reference potential.

[0348] In this specification and the like, a node can be referred to as a terminal, a wiring, an electrode, a conductive layer, a conductor, an impurity region, etc. depending on the circuit configuration, device structure, etc. Also, a terminal, a wiring, etc. can be referred to as a node.

[0349] In this specification, "A and B are connected" means that A and B are electrically connected. Here, "A and B are electrically connected" means a connection in which an electrical signal can be transmitted between A and B when an object (such as a switch, transistor element, or diode, or a circuit including such an object and wiring) is present between A and B. Note that "A and B are electrically connected" also includes a case in which A and B are directly connected. Here, "A and B are directly connected" means a connection in which an electrical signal can be transmitted between A and B via wiring (or electrodes) or the like, without passing through the object. In other words, a direct connection means a connection that can be regarded as the same circuit diagram when represented by an equivalent circuit.

[0350] In this specification, a switch refers to a device that has the function of controlling whether a current flows by being in a conductive state (on state) or a non-conductive state (off state), or a device that has the function of selecting and switching a path for a current to flow.

[0351] In this specification, the channel length refers to, for example, in a top view of a transistor, a region where a semiconductor (or a portion in the semiconductor through which current flows when the transistor is on) and a gate overlap, or a distance between a source and a drain in a region where a channel is formed.

[0352] In this specification, the channel width refers to, for example, the length of the region where the semiconductor (or the portion in the semiconductor through which current flows when the transistor is on) and the gate electrode overlap, or the length of the portion where the source and drain face each other in the region where the channel is formed.

[0353] In this specification and the like, terms such as "film" and "layer" can be interchangeable depending on the circumstances. For example, the term "conductive layer" can be changed to the term "conductive film." Or, for example, the term "insulating film" can be changed to the term "insulating layer." [Explanation of symbols]

[0354] AIN_1: Input data, AIN: Input data, BGL: Back gate line, BK: Signal, BKH: Signal, BL: Bit line, C11: Capacitor, CK: Node, CLK: Clock signal, DIN: Inference data, DJD: Output data, DTR: Learning data, EN: Control signal, GBL_A: Wiring, GBL_B: Wiring, GBL_N: Wiring, GBL_P: Wiring, GBL: Wiring, GL[2]: Wiring, GL: Wiring, LBL_1: Wiring, LBL_7: Wiring, LBL_N: Wiring, LBL_P: Wiring, LBL: Wiring, LBLP: Wiring, M11: Transistor, M12: Transistor Register, M13: Transistor, MAC: Output Data, RC: Signal, RCH: Signal, RT: Node, RWL_1: Read Word Line, RWL: Read Word Line, SCE: Signal, SD_IN: Node, SD: Node, SE: Node, SL: Source Line, SN11: Node, WBL_N: Write Bit Line, WBL_P: Write Bit Line, WBL: Write Bit Line, Wdata: Weight Data, WINF: Weight Data, WL: Word Line, WSEL_A: Weight Data, WSEL_B: Weight Data, WSEL: Weight Data, WTR: Weight Data, WWL_1: Write word line, WWL: write word line, 10_1: semiconductor device, 10_n: semiconductor device, 10: semiconductor device, 11: layer, 12: layer, 20_1: memory circuit section, 20_4: memory circuit section, 20_6: memory circuit section, 20_N: memory circuit section, 20_N(N: memory circuit section, 20: memory circuit section, 21_N: memory circuit, 21_P: memory circuit, 21A: memory circuit, 21B: memory circuit, 21C: memory circuit, 21: memory circuit, 22: transistor, 23: semiconductor layer, 24: multiplication circuit, 25: addition circuit, 26: register, 30_1: arithmetic circuit, 30_1 2: arithmetic circuit, 30_4: arithmetic circuit, 30_6: arithmetic circuit, 30_7: arithmetic circuit, 30_N: arithmetic circuit, 30: arithmetic circuit, 31: server, 32: computer device, 33A: processing, 33B: processing, 34: weight data, 35: electronic device, 36: automobile, 37: weight data, 38A: processing, 38B: processing, 39A: sign part, 39B: exponent part, 39C: mantissa part, 40_1: switching circuit, 40_12: switching circuit, 40_4: switching circuit, 40_6: switching circuit, 40_7: switching circuit, 40A: switching circuit, 40B: switching circuit, 40M: switching circuit, 40X: switching circuit,40Y: switching circuit, 40: switching circuit, 50: driving circuit, 60: memory circuit, 61_N: transistor, 61_P: transistor, 61A: transistor, 61B: transistor, 61: transistor, 62_N: transistor, 62_P: transistor, 62B: transistor, 62: transistor, 63_N: transistor, 63_P: transistor, 63: transistor, 64_N: capacitance element, 64_P: capacitance element, 64A: capacitance element, 64B: capacitance element, 64: capacitance element, 71G: controller, 71: controller, 72: row decoder, 73: Word line driver, 74: column decoder, 75: write driver, 76: precharge circuit, 81: input / output buffer, 82: arithmetic control circuit, 90A: input layer, 90B: hidden layer, 90C: output layer, 92: convolution operation processing, 93: convolution operation processing, 94: pooling operation processing, 95: convolution operation processing, 96: pooling operation processing, 100: arithmetic processing system, 110: CPU, 120: bus, 193: PMU, 200: CPU core, 202: L1 cache memory device, 203: L2 cache memory device, 205: bus interface unit, 210 : Power switch, 211: Power switch, 212: Power switch, 214: Level shifter, 220: Flip-flop, 221A: Clock buffer circuit, 221: Scan flip-flop, 222: Backup circuit, 300N: OS memory, 311: Substrate, 312: Well region, 313: Insulator, 314: Oxide layer, 315: Semiconductor region, 316a: Low resistance region, 316b: Low resistance region, 316c: Low resistance region, 317: Insulator, 318: Conductor, 320: Insulator, 322: Insulator, 324: Insulator, 326: Insulator, 328: Conductor, 33 0: conductor, 350: insulator, 352: insulator, 354: insulator, 356: conductor, 360: insulator, 362: insulator, 364: insulator, 366: conductor, 370: insulator, 372: insulator, 374: insulator, 376: conductor, 380: insulator, 382: insulator, 384: insulator, 386: conductor, 390: integrated circuit, 391: semiconductor chip, 392: lead, 393: Si transistor layer, 394: wiring layer, 395: OS transistor layer, 400: package substrate, 401: solder ball, 402: semiconductor substrate, 403: transistor, 404: wiring,405: electrode, 412: semiconductor substrate, 413: transistor, 414: wiring, 415: electrode, 420: region, 430: conductor, 431: insulator, 432: semiconductor region, 433a: low resistance region, 433b: low resistance region, 440: insulator, 442: insulator, 444: insulator, 446: insulator, 448: conductor, 450: insulator, 452: insulator, 454: insulator, 500: transistor, 503a: conductor, 503b: conductor, 503: conductor, 510: insulator, 512: insulator, 514: insulator, 516: insulator, 518: conductor, 522: insulator, 524: insulator, 530a: oxide, 530b: oxide, 530: oxide, 540a: conductor, 540b: conductor, 542a: conductor, 542b: conductor, 542: conductor, 543a: region, 543b: region, 544: insulator, 545: insulator, 546: conductor, 548: conductor, 550: transistor, 560a: conductor, 560b: conductor, 560: conductor, 574: insulator, 580: insulator, 581 : insulator, 582: insulator, 586: insulator, 590: automobile, 591: camera, 592: imaging direction, 593: bus, 594: host controller, 595: portable electronic device, 596: printed wiring board, 597: speaker, 598: camera, 599: microphone, 600: capacitive element, 610: conductor, 612: conductor, 620: conductor, 630: insulator, 640: insulator, 1100: portable game console, 1101: housing, 1102: housing, 1103: housing, 1104: surface display unit, 1105: connection unit, 1107: operation keys, 1108: housing, 1109: housing, 1120: electronic device, 1121: housing, 1122: cap, 1123: USB connector, 1124: board, 1125: memory chip, 1126: controller chip, 1130: robot, 2101: sensor, 2106: sensor, 2110: control circuit, 3000: system, 3001: electronic device, 3002: server, 3003: internet line, 3004: rack, 3005: board,

Claims

1. The device includes a plurality of memory circuits, a switching circuit, and an arithmetic circuit, each of the plurality of memory circuits has a function of holding weight data and a function of outputting the weight data to a first wiring; the switching circuit has a function of switching a conduction state between any one of the plurality of first wirings and a second wiring, the arithmetic circuit has a function of performing arithmetic processing using input data and the weight data provided to the second wiring, the weight data is data of a first number of bits, the weight data is data obtained by converting weight data of a second number of bits optimized using learning data, The semiconductor device, wherein the first number of bits is smaller than the second number of bits.

2. In claim 1, The second wiring has wiring that is provided parallel or approximately parallel to the surface of the substrate.

3. In claim 1 or claim 2, The semiconductor device includes a first wiring provided substantially perpendicular to a surface of a substrate.

4. In any one of claims 1 to 3, The semiconductor device, wherein the arithmetic circuit is a circuit that performs a product-sum operation.

Citation Information

Patent Citations

  • Laminate memory device, method for operating the same, and memory system

    JP2019061677A

  • Accelerated deep learning

    US20180314941A1

  • Building a binary neural network architecture

    WO2019078924A1