Semiconductor device

WO2026202648A1PCT designated stage Publication Date: 2026-10-01SEMICON ENERGY LAB CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2026/052557
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-10-03
Filing Date
2026-03-17
Publication Date
2026-10-01

Smart Images

  • Figure IB2026052557_01102026_PF_FP_ABST
    Figure IB2026052557_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a novel semiconductor device. The semiconductor device comprises: a first circuit that has the function of supplying, to a second circuit, a current which has the same value as the total of currents supplied to n input terminals; and a second circuit that has the function of setting the value of the current which has been supplied from the first circuit to 1 / n and supplying the current to an output terminal, wherein the first circuit has a transistor that includes an oxide semiconductor in a semiconductor layer in which a channel is formed, the second circuit has a transistor that includes silicon in a semiconductor layer in which a channel is formed, and the first circuit and the second circuit have mutually overlapping regions.
Need to check novelty before this filing date? Find Prior Art

Description

Semiconductor equipment

[0001] One aspect of the present invention relates to a semiconductor device.

[0002] Furthermore, one aspect of the present invention is not limited to the above-mentioned technical field. The technical field of one aspect of the invention disclosed herein relates to a product, a method, or a method of manufacture. Alternatively, one aspect of the present invention relates to a process, a machine, a manufacture, or a composition of matter. More specifically, examples of the technical fields of one aspect of the present invention disclosed herein include semiconductor devices, display devices, liquid crystal display devices, light-emitting devices, lighting devices, energy storage devices, memory devices, methods for driving them, or methods for manufacturing them.

[0003] Converting analog data to digital data before performing calculations requires a massive amount of computation, making it difficult to reduce the time required for these calculations. Therefore, a method has been proposed that performs multiply-accumulate operations without converting analog data to digital data, similar to the information processing of analog data performed in the brain, which uses neurons as its basic elements.

[0004] For example, Patent Document 1 below discloses an arithmetic circuit that performs nonlinear transformation operations and multiply-accumulate operations using an analog circuit.

[0005] Japanese Patent Publication No. 2004-110421

[0006] For example, in a convolutional neural network used for image recognition, if the results of a multiply-accumulate operation performed on analog data in the convolutional layer are sent to the pooling layer, and the pooling layer does not support analog data, then the analog data needs to be converted to digital data. This leads to problems such as an increase in the circuit size of the semiconductor device required to implement the neural network, and an increase in the time required for computation and power consumption.

[0007] One aspect of the present invention aims to reduce the circuit size of a semiconductor device. Alternatively, it aims to provide a semiconductor device that can reduce the time required for computational processing. Alternatively, it aims to provide a semiconductor device that can reduce the power consumption required for computational processing. Alternatively, it aims to provide a novel semiconductor device.

[0008] Furthermore, the description of the above problems does not preclude the existence of other problems. Other problems can be naturally identified by those skilled in the art from the description in the specification, drawings, and claims, and it is possible to extract other problems from the description in the specification, drawings, and claims. Moreover, one aspect of the present invention does not need to solve all of these problems (the above problems and other problems).

[0009] (1) One aspect of the present invention comprises a first circuit having n (where n is 2 or more) input terminals, an output terminal, and first to fourth transistors, and a second circuit having fifth to eighth transistors, wherein the first terminal of the first transistor is electrically connected to the gate of the first transistor, the gate of the third transistor and the n input terminals, the second terminal of the first transistor is electrically connected to the first terminal of the second transistor, the gate of the second transistor and the gate of the fourth transistor, and the first terminal of the third transistor is electrically connected to the first terminal of the fifth transistor and the fifth transistor This semiconductor device is electrically connected to the gate of the first transistor and the gate of the seventh transistor, the second terminal of the third transistor is electrically connected to the first terminal of the fourth transistor, the second terminal of the fifth transistor is electrically connected to the first terminal of the sixth transistor, the gate of the sixth transistor and the gate of the eighth transistor, the first terminal of the eighth transistor is electrically connected to the second terminal of the seventh transistor, and the first terminal of the seventh transistor is electrically connected to the output terminal, and has the function of supplying a current to the output terminal that is 1 / n of the sum of the currents supplied to the n input terminals.

[0010] Furthermore, in (1), it is preferable that the channel width of the seventh transistor is smaller than the channel width of the fifth transistor, and the channel width of the eighth transistor is smaller than the channel width of the sixth transistor.

[0011] Furthermore, in (1), it is preferable that the first circuit has the function of supplying a current to the second circuit that is k (where k is a number greater than 1) times the current supplied from the n input terminals, and the second circuit has the function of supplying a current to the output terminal that is 1 / (n × k) times the current supplied from the first circuit. In this case, the channel width of the third transistor is larger than the channel width of the first transistor, and the channel width of the fourth transistor is larger than the channel width of the second transistor.

[0012] Furthermore, in (1), it is preferable that the first to fourth transistors include an oxide semiconductor in the semiconductor layer where the channel is formed. It is preferable that the fifth to eighth transistors include silicon in the semiconductor layer where the channel is formed. It is preferable that the first to fourth transistors are n-type transistors and the fifth to eighth transistors are p-type transistors. It is preferable that the first circuit and the second circuit have overlapping regions.

[0013] (2) Another aspect of the present invention is a semiconductor device having a first circuit that has the function of supplying a second circuit with a current equal to the sum of the currents supplied to each of n (where n is a number of 2 or more) input terminals, and a second circuit that has the function of supplying an output terminal with the value of the current supplied from the first circuit divided by n, wherein the first circuit has a transistor in which an oxide semiconductor is included in the semiconductor layer in which a channel is formed, and the second circuit has a transistor in which silicon is included in the semiconductor layer in which a channel is formed, and the first circuit and the second circuit have overlapping regions.

[0014] (3) Another aspect of the present invention is a semiconductor device having a first circuit that has the function of supplying a second circuit with a current k (a number greater than 1) times the sum of the currents supplied to each of n (n is a number of 2 or more) input terminals, and a second circuit that has the function of supplying the current supplied from the first circuit to an output terminal at a multiplier of 1 / (n × k), wherein the first circuit has a transistor in which an oxide semiconductor is included in the semiconductor layer in which the channel is formed, and the second circuit has a transistor in which silicon is included in the semiconductor layer in which the channel is formed, and the first circuit and the second circuit have overlapping regions.

[0015] Furthermore, in (2) and (3), the aforementioned transistor in the first circuit is preferably an n-type transistor, and the aforementioned transistor in the second circuit is preferably a p-type transistor.

[0016] Furthermore, in (1), (2), and (3), it is preferable that the oxide semiconductor contains indium.

[0017] According to one aspect of the present invention, it is possible to reduce the circuit size of a semiconductor device, or to provide a semiconductor device that can reduce the time required for computational processing, or to provide a semiconductor device that can reduce the power consumption required for computational processing.

[0018] Furthermore, the description of the above effects does not preclude the existence of other effects. Other effects can be naturally found by those skilled in the art from the description in the specification, drawings, and claims, and it is possible to extract other effects from the description in the specification, drawings, and claims. Furthermore, one aspect of the present invention does not need to have all of these effects (the above effects and other effects).

[0019] Figures 1A, 1B, and 1C show examples of semiconductor device configurations. Figures 2A and 2B show examples of semiconductor device configurations. Figure 3 shows an example of semiconductor device configuration. Figure 4 is a timing chart illustrating an example of semiconductor device operation. Figure 5 is a diagram illustrating an example of semiconductor device operation. Figure 6 is a diagram illustrating an example of semiconductor device operation. Figure 7 is a diagram illustrating an example of semiconductor device operation. Figure 8 shows an example of semiconductor device configuration. Figure 9 shows an example of a convolutional neural network. Figure 10 is a diagram illustrating an example of convolution processing. Figure 11 is a diagram illustrating an example of convolution processing. Figures 12A1, 12A2, 12B1, and 12B2 illustrate an example of pooling processing. Figures 13A, 13B, 13C, and 13D illustrate examples of transistor configurations. Figures 14A and 14B illustrate examples of transistor configurations. Figures 15A and 15B illustrate examples of transistor configurations. Figures 16A and 16B illustrate examples of transistor configurations. Figures 17A, 17B, 17C, and 17D show examples of electronic components. Figure 18 shows an example of an information processing system. Figure 19 shows an example of space equipment. Figure 20A is a block diagram showing the configuration of an artificial neural network. Figure 20B is a block diagram showing an example of a multiply-accumulate circuit configuration. Figures 21A and 21B are block diagrams showing examples of multiply-accumulate circuit configurations. Figure 22 shows the inference results using an artificial neural network.

[0020] The embodiments will be described below with reference to the drawings. However, it will be readily apparent to those skilled in the art that the embodiments can be implemented in many different ways, and their form and details can be modified in various ways without departing from the spirit and scope thereof. Accordingly, the present invention shall not be construed as being limited to the contents of the following embodiments.

[0021] In this specification, a semiconductor device refers to a device that utilizes semiconductor properties, including circuits containing semiconductor elements (transistors, diodes, photodiodes, etc.), devices having such circuits, etc. It also refers to any device that can function by utilizing semiconductor properties. For example, integrated circuits, chips equipped with integrated circuits, and electronic components with chips housed in packages are examples of semiconductor devices. Furthermore, memory devices, display devices, light-emitting devices, lighting devices, and electronic devices are themselves semiconductor devices, and may also contain semiconductor devices.

[0022] In the drawings relating to this specification, the size, layer thickness, etc., may be exaggerated for clarity. Therefore, the dimensions or aspect ratio are not necessarily limited to those shown. Furthermore, the drawings schematically represent ideal examples, and one aspect of the present invention is not limited to the shapes, values, etc., shown in the drawings.

[0023] In the configuration of the embodiment of the invention, the same reference numerals are used in common across different drawings for the same part or parts having similar functions, and repeated explanations may be omitted. Also, when referring to similar functions, the same hatching pattern may be used, and no reference numerals may be assigned. Furthermore, in order to make the drawings easier to understand, the description of some components may be omitted in perspective views or plan views, etc.

[0024] In this specification, ordinal numbers such as "first," "second," etc., are used to avoid confusion of constituent elements. Therefore, they do not limit the number of constituent elements, nor do they limit the order of the constituent elements. For example, a constituent element referred to as "first" in one embodiment of this specification may be referred to as "second" in another embodiment or in the claims. Also, a constituent element referred to as "first" in one embodiment of this specification may be omitted in another embodiment or in the claims. Furthermore, even if a term does not have an ordinal number in this specification, an ordinal number may be added in the claims to avoid confusion of constituent elements. Also, even if a term has an ordinal number in this specification, a different ordinal number may be added in the claims. Furthermore, even if a term has an ordinal number in this specification, the ordinal number may be omitted in the claims.

[0025] In this specification, phrases indicating arrangement such as "above," "below," "upward," or "downward" are sometimes used for convenience to explain the positional relationship between components with reference to the drawings. Furthermore, the positional relationship between components changes as appropriate depending on the direction in which each component is depicted. Therefore, the phrases described in the specification are not limited to those described and can be appropriately rephrased depending on the situation. For example, the expression "insulator located on the upper surface of the conductor" can be rephrased as "insulator located on the lower surface of the conductor" by rotating the orientation of the drawing shown by 180 degrees.

[0026] Furthermore, the terms "above" and "below" do not limit the positional relationship of the components to being directly above or below each other and in direct contact. For example, the expression "electrode B on insulating layer A" does not require that electrode B be formed in direct contact with insulating layer A, and does not exclude cases where other components are included between insulating layer A and electrode B.

[0027] In this specification, terms such as "overlapping" do not limit the stacking order or other states of the constituent elements. For example, the expression "electrode B overlapping insulating layer A" does not exclude not only the state in which electrode B is formed on top of insulating layer A, but also the state in which electrode B is formed below insulating layer A, or the state in which electrode B is formed to the right (or left) of insulating layer A, etc.

[0028] In this specification, terms such as "film" and "layer" can be interchanged as needed. For example, the term "conductive layer" may be changed to the term "conductive film." Or, for example, the term "insulating film" may be changed to the term "insulating layer." Alternatively, depending on the circumstances, terms such as "film" and "layer" can be omitted and replaced with other terms. For example, the term "conductive layer" or "conductive film" may be changed to the term "conductor." Or, the term "conductor" may be changed to the term "conductive layer" or "conductive film." Alternatively, for example, the term "insulating layer" or "insulating film" may be changed to the term "insulator." Or, the term "insulator" may be changed to the term "insulating layer" or "insulating film."

[0029] In this specification, terms such as "electrode," "wiring," and "terminal" do not functionally limit these components. For example, "electrode" may be used as part of "wiring," and vice versa. Furthermore, the terms "electrode" or "wiring" include cases where multiple "electrodes" or "wiring" are formed as a single unit. Similarly, for example, "terminal" may be used as part of "wiring" or "electrode," and vice versa. Furthermore, the term "terminal" also includes cases where multiple "electrodes," "wiring," or "terminals" are formed as a single unit. Therefore, for example, an "electrode" can be part of "wiring" or a "terminal," and for example, a "terminal" can be part of "wiring" or an "electrode." In addition, terms such as "electrode," "wiring," and "terminal" may be replaced with terms such as "region" or "conductive layer" depending on the context.

[0030] In this specification, terms such as "wiring," "signal line," and "power line" can be interchanged with each other as appropriate or depending on the circumstances. For example, the term "wiring" may be changed to the term "signal line." Similarly, the term "wiring" may be changed to the term "power line," and vice versa. Terms such as "power line" may be changed to the term "wiring." Terms such as "power line" may be changed to the term "signal line," and vice versa. Furthermore, the term "potential" applied to wiring may be changed to the term "signal," and vice versa.

[0031] In this specification, "source" refers to a source region, a source electrode, or a source wiring. A source region refers to one of two regions in a semiconductor layer adjacent to a channel formation region. A source electrode refers to a conductive layer that includes the portion connected to the source region. A source wiring refers to a conductive layer used to electrically connect the source electrode of at least one transistor to another electrode or another wiring.

[0032] In this specification, "drain" refers to the drain region, drain electrode, or drain wiring. The drain region refers to the other of two regions adjacent to the channel formation region within the semiconductor layer. The drain electrode refers to the conductive layer that includes the portion connected to the drain region. The drain wiring refers to the conductive layer used to electrically connect the drain electrode of a transistor to another electrode or another wiring.

[0033] In this specification, "gate" refers to the gate electrode or gate wiring. The gate electrode is an electrode that overlaps with the semiconductor layer of a transistor and has the function of controlling the resistance between the source and drain of the transistor by the supplied voltage. Gate wiring refers to a conductive layer that electrically connects the gate electrode of a transistor to another electrode or another wiring.

[0034] In this specification, one of the sources or drains of a transistor may be referred to as the "first terminal of the transistor." The other of the sources or drains of a transistor may be referred to as the "second terminal of the transistor."

[0035] Furthermore, voltage often refers to the potential difference between a certain potential and a reference potential (e.g., ground potential or source potential). Therefore, voltage and potential are often interchangeable. In this specification, unless otherwise specified, voltage and potential may be considered interchangeable.

[0036] Furthermore, in this specification, the high power supply potential VDD (hereinafter also simply referred to as "VDD") refers to a power supply potential that is higher than the low power supply potential VSS. The low power supply potential VSS (hereinafter also simply referred to as "VSS") refers to a power supply potential that is lower than the high power supply potential VDD. The ground potential GND (hereinafter also simply referred to as "GND") can also be used as VDD or VSS. For example, if VDD is GND, then VSS is at a lower potential than GND, and if VSS is GND, then VDD is at a higher potential than GND.

[0037] In this specification, the "on state" of a transistor means that the source and drain of the transistor are conductive (a state in which current can be conducted). The "off state" of a transistor means that the source and drain of the transistor are non-conductive (a state in which it can be considered electrically blocked).

[0038] Furthermore, in this specification, "on-current" refers to the current that flows between the source and drain when the transistor is in the ON state. "Off-current" refers to the current that flows between the source and drain when the transistor is in the OFF state.

[0039] Furthermore, unless otherwise explicitly stated, the transistors described herein are enhancement-type (normally-off) field-effect transistors. Also, with a reference potential of 0V, the threshold voltage (also called "Vth") of an n-channel field-effect transistor (also called "n-type transistor") is a positive voltage with an absolute value greater than 0V, and the Vth of a p-channel field-effect transistor (also called "p-type transistor") is a negative voltage with an absolute value greater than 0V.

[0040] In this specification, when count values ​​and measured values ​​are referred to as "identical," "same," "equal," or "uniform" (including synonyms thereof), unless otherwise explicitly stated, this shall include an error margin of plus or minus 10%.

[0041] Furthermore, in drawings and other illustrations relating to this specification, arrows indicating the X, Y, and Z directions may be included. In this specification, the "X direction" refers to the direction along the X-axis, and unless explicitly stated, the forward and reverse directions may not be distinguished. The same applies to the "Y direction" and "Z direction." Also, the X, Y, and Z directions are directions that intersect each other. For example, the X, Y, and Z directions are directions that are orthogonal to each other. In this specification, one of the X, Y, or Z directions may be referred to as the "first direction" or "first direction." Another may be referred to as the "second direction" or "second direction." The remaining one may be referred to as the "third direction" or "third direction."

[0042] In explaining circuit operation, rise and fall times may occur when the potential changes, for example, due to loads such as wiring (parasitic capacitance and resistance). Therefore, even if two different operations are shown to occur at the same time, this does not necessarily mean they are strictly at the same time. For example, even if there is a slight time difference due to signal delay in the wiring, they may still be considered to occur at the same time.

[0043] In timing charts used to explain circuit operation, each period is sometimes depicted as having the same length for clarity. However, the length of each period can be changed.

[0044] In this specification, when the same symbol is used for multiple elements, and especially when it is necessary to distinguish them, an identifying symbol such as "A", "b", "_1", "[n]", or "[m,n]" may be added to the symbol.

[0045] In this specification, "connection" includes "electrical connection." The term "electrical connection" is sometimes used to describe the connection relationships of circuit elements as physical objects. Furthermore, "electrical connection" includes both "direct connection" and "indirect connection." "A and B are directly connected" means that A and B are connected without the use of circuit elements (e.g., transistors, switches, etc.; wiring is not considered a circuit element). On the other hand, "A and B are indirectly connected" means that A and B are connected through one or more circuit elements.

[0046] For example, assuming a circuit including A and B is in operation, if there is a timing during the circuit's operation when electrical signals are exchanged or potential interactions occur between A and B, then it can be defined that "A and B are indirectly connected" as physical objects. Furthermore, even if there is a timing during the circuit's operation when no electrical signals are exchanged or potential interactions occur between A and B, if there is a timing during the circuit's operation when electrical signals are exchanged or potential interactions occur between A and B, then it can be defined that "A and B are indirectly connected."

[0047] An example of a case where "A and B are indirectly connected" is when A and B are connected via the source and drain of one or more transistors. On the other hand, an example of a case where "A and B are not indirectly connected" is when an insulator is interposed in the path from A to B. Specifically, this includes cases where a capacitive element is connected between A and B, or where a transistor gate insulating film is interposed between A and B. Therefore, it cannot be said that "the gate (A) of a transistor and the source or drain (B) of a transistor are indirectly connected."

[0048] Another example of a situation where it cannot be said that "A and B are indirectly connected" is when multiple transistors are connected via source and drain in the path from A to B, and a constant potential V is supplied to the nodes between the transistors from a power supply, GND, etc.

[0049] Furthermore, one aspect of the present invention is all or part of the circuit configuration described herein. Therefore, one aspect of the present invention is all or part of the operation described herein.

[0050] (Embodiment 1) In this embodiment, an example of the circuit configuration and operation of a semiconductor device 100 according to one aspect of the present invention will be described.

[0051] <Circuit Configuration Example> Figure 1A shows a circuit configuration example of a semiconductor device 100 according to one aspect of the present invention. The semiconductor device 100 has n input terminals IN (where n is 2 or more), an output terminal OUT, a first circuit 101, and a second circuit 102.

[0052] The first circuit 101 has transistors M11, M12, M13, and M14. The second circuit 102 has transistors M21, M22, M23, and M24.

[0053] In the first circuit 101, either the source or drain of transistor M11 is connected to each of the n input terminals IN. In Figure 1A, the first input terminal IN is shown as input terminal IN[1], the second input terminal IN as input terminal IN[2], the third input terminal IN as input terminal IN[3], and the nth input terminal IN as input terminal IN[n].

[0054] Furthermore, one of the sources or drains of transistor M11 is connected to the gate of transistor M11 and the gate of transistor M13. The other of the sources or drains of transistor M11 is connected to one of the sources or drains of transistor M12, the gate of transistor M12, and the gate of transistor M14. The other of the sources or drains of transistor M12 is connected to wiring 15. Wiring 15 is supplied with, for example, VSS.

[0055] In this embodiment, the region where the source or drain of transistor M11, the source or drain of transistor M12, the gate of transistor M12, and the gate of transistor M14 are connected and are always at the same potential is indicated as node Nd1.

[0056] Either the source or drain of transistor M13 is connected to either the source or drain of transistor M21, the gate of transistor M21, and the gate of transistor M23, all of which are located in the second circuit 102.

[0057] Furthermore, the other source or drain of transistor M13 is connected to one of the source or drains of transistor M14. The other source or drain of transistor M14 is connected to wiring 15.

[0058] In this embodiment, the region connected to the other source or drain of transistor M13 and one source or drain of transistor M14, where they are always at the same potential, is referred to as node Nd2.

[0059] In the second circuit 102, the source or drain of transistor M21 is connected to either the source or drain of transistor M22, the gate of transistor M22, and the gate of transistor M24. The other source or drain of transistor M22 is connected to wiring 25. Wiring 25 is supplied with, for example, VDD.

[0060] In this embodiment, the region where the source or drain of transistor M21, the source or drain of transistor M22, the gate of transistor M22, and the gate of transistor M24 are connected and are always at the same potential is indicated as node Nd3.

[0061] One of the sources or drains of transistor M23 is connected to the output terminal OUT. The other of the sources or drains of transistor M23 is connected to one of the sources or drains of transistor M24. The other of the sources or drains of transistor M24 is connected to wiring 25.

[0062] In this embodiment, the region where the source or drain of transistor M23 and the source or drain of transistor M24 are connected and are always at the same potential is indicated as node Nd4. In addition, in Figure 1 and other diagrams, arrows indicating the direction of current flow may be added.

[0063] Various types of transistors can be used as transistors constituting the semiconductor device 100. For example, various types of transistors can be used, such as top-gate type (e.g., planar type and staggered type), bottom-gate type (e.g., inverse planar type and inverse staggered type), dual-gate type (a structure in which gates are arranged on both sides (e.g., top and bottom) of the channel formation region), FIN type, TRI-GATE type, and GAA type (gate all-around type). In addition, for example, vertical transistors (transistors in which the channel length direction has a component in the vertical direction (also called the height direction or the direction perpendicular to the formed surface)) can be used.

[0064] Furthermore, it is also possible to use transistors with back gates as transistors constituting the semiconductor device 100. Figure 1B shows an example of the circuit configuration of the first circuit 101 when transistors M11 to M14 are transistors with back gates. Figure 1C shows an example of the circuit configuration of the second circuit 102 when transistors M21 to M24 are transistors with back gates.

[0065] Here, let's explain the back gate of a transistor. The gate and back gate of a transistor are positioned so as to sandwich the channel formation region of the semiconductor layer. Both the gate and back gate are formed from a conductive layer or a semiconductor layer with low resistivity. The back gate can function in the same way as the gate. When the gate is used to control the on and off states of the transistor, the potential of the back gate can be set to the same potential as the gate.

[0066] For example, when turning on a transistor with a back gate, supplying the potential required to turn on the transistor to both the gate and the back gate increases the on-current compared to supplying it to only one. Connecting the gate and the back gate makes it possible to keep them at the same potential at all times. Alternatively, by not connecting the gate and the back gate and controlling the back gate's potential independently of the gate's potential, the transistor's Vth can be adjusted. It is also possible to supply a fixed potential, such as GND, to the back gate.

[0067] Since the gate and back gate are formed by conductive layers, the gate and back gate sandwich the channel formation region of the semiconductor layer, making it difficult for electric fields generated outside the transistor to act on the channel formation region (also known as the "field shielding effect"). Therefore, providing a back gate to a transistor stabilizes its operation. In addition, providing a back gate to a transistor reduces the variation in characteristics between multiple transistors. Providing a back gate to a transistor can improve the reliability of the transistor. Therefore, the reliability of the semiconductor device containing the transistor can be improved. Note that the field shielding effect can be obtained even if one or both of the gate and back gate are electrically floating (also known as the "floating state"), but the effect can be enhanced by supplying potential to the gate and back gate.

[0068] Furthermore, when light is shone on the channel formation region of a transistor, the transistor's electrical characteristics may fluctuate. Also, when light is shone on the channel formation region while a voltage is applied to the transistor, the transistor's electrical characteristics may deteriorate. In other words, the reliability of the transistor may decrease. By using light-shielding conductive materials for both the gate and back gate, the deterioration of the transistor's electrical characteristics can be suppressed, and its reliability can be improved.

[0069] Each of the first circuit 101 and the second circuit 102 included in the semiconductor device 100 functions as a current mirror circuit. When VSS is supplied to the wiring 15, it is preferable to use n-type transistors as transistors M11 to M14. However, by adjusting the connection configuration, it is also possible to use p-type transistors as transistors M11 to M14. Furthermore, when VDD is supplied to the wiring 25, it is preferable to use p-type transistors as transistors M21 to M24. However, by adjusting the connection configuration, it is also possible to use n-type transistors as transistors M21 to M24.

[0070] As the semiconductor layer in which the channels of transistors M11 to M14 and transistors M21 to M24 are formed, it is possible to use semiconductors made of single elements such as silicon and germanium, or compound semiconductors such as silicon germanium, silicon carbide, gallium arsenide, oxide semiconductors, and nitride semiconductors.

[0071] When using a p-type transistor as the transistor constituting the semiconductor device 100, for example, a transistor using silicon in the semiconductor layer where the channel is formed (also referred to as a "Si transistor" or "SiFET") can be used. Si transistors are easier to realize as p-type transistors than OS transistors. Furthermore, Si transistors using crystalline silicon in the semiconductor layer have a faster operating speed than OS transistors.

[0072] When an n-type transistor is used as the transistor constituting the semiconductor device 100, it is possible to use a transistor (also referred to as an "OS transistor" or "OSFET") that includes an oxide semiconductor, a type of metal oxide, in the semiconductor layer where the channel is formed. Since oxide semiconductors have a band gap of 2 eV or more, the off-current is extremely small. Specifically, the off-current value of an OS transistor per 1 μm of channel width at room temperature is 1 pA (1 × 10⁻¹⁶). −12 A) Below, 1aA (1×10 −18 A) Below, 1zA (1×10 −21 A) Less than or equal to 1yA (1 × 10 −24 A) The following is possible:

[0073] Furthermore, OS transistors exhibit virtually no change in electrical characteristics even in high-temperature environments. Specifically, the off-current hardly increases even at ambient temperatures between room temperature and 200°C. Also, the on-current does not decrease significantly even in high-temperature environments. Semiconductor devices containing OS transistors operate stably and with high reliability even in high-temperature environments.

[0074] The metal oxide used as the oxide semiconductor preferably contains at least indium (In). By increasing the ratio of indium atoms to the sum of the total number of atoms of all metal elements contained in the metal oxide used in the semiconductor layer, the field-effect mobility of the OS transistor can be increased. Typically, by using single-crystal or polycrystalline indium oxide (also called "indium oxide") in the semiconductor layer, the field-effect mobility of the OS transistor can be significantly increased.

[0075] Furthermore, since OS transistors can be formed as a type of thin-film transistor, it is possible to stack the transistors that constitute the semiconductor device 100. Figure 2A shows a perspective view in which a first circuit 101 formed of an OS transistor is placed above a second circuit 102 formed of a Si transistor. For example, it is possible to form the second circuit 102 on an element layer 20 containing a Si transistor and the first circuit 101 on an element layer 10 containing an OS transistor (see Figure 2B). In this way, by using a combination of transistors with different semiconductor layer compositions, it becomes easy to stack the transistors that constitute the semiconductor device 100.

[0076] Furthermore, by overlapping the first circuit 101 and the second circuit 102 that constitute the semiconductor device 100, the occupied area of ​​the semiconductor device 100 can be reduced. Therefore, the mounting density of the semiconductor device 100 can be increased.

[0077] Furthermore, as mentioned above, the electrical characteristics of OS transistors hardly change even in high-temperature environments. For this reason, the first circuit 101, which is formed of OS transistors and is located above the second circuit 102, which is formed of Si transistors, is less affected by the heat generated by the second circuit 102. Thus, the reliability of the semiconductor device 100 can be improved.

[0078] As described above, the first circuit 101 and the second circuit 102 included in the semiconductor device 100 each function as a current mirror circuit. Therefore, for example, in the first circuit 101, it is preferable that the electrical characteristics of transistor M11 and transistor M13 are equal. It is also preferable that the electrical characteristics of transistor M12 and transistor M14 are equal. In particular, it is preferable that the Vth of transistor M11 and transistor M13 are equal. It is also preferable that the Vth of transistor M12 and transistor M14 are equal.

[0079] For example, if the channel sizes of transistors M12 and M14 are equal, the current I1 flowing through node Nd1 and the current I2 flowing through node Nd2 will be equal. Current I2 will increase or decrease in conjunction with increases or decreases in current I1. In this specification, channel size refers to the ratio of the channel width W to the channel length L of a transistor. Therefore, the channel size increases as the channel width W is larger relative to the channel length L.

[0080] Furthermore, if the channel length L of transistor M12 is L12 and the channel width W is W12, and the channel length L of transistor M14 is L14 and the channel width W is W14, then the relationship between currents I1 and I2 can be expressed by equation (1).

[0081]

[0082] From equation (1), it can be seen that the magnitude of current I2 relative to current I1 can be adjusted by adjusting the channel size. For example, by making the channel size of transistor M14 twice that of transistor M12, current I2 can be made twice that of current I1. In this specification, the right-hand side of equation (1) is sometimes referred to as the "channel size ratio". For example, if the channel size ratio of the first circuit 101 is 3, the amount of current I2 becomes three times that of current I1.

[0083] Furthermore, changing the channel length L often changes the Vth of the transistor. For this reason, it is preferable to adjust the channel size by changing the channel width W without changing the channel length L. For example, if the channel size ratio of the first circuit 101 is set to 2 and the current I2 is set to twice the current I1, it is preferable to make the channel length L12 and channel length L14 equal and the channel width W14 twice the channel width W12.

[0084] In the first circuit 101, one of the source or drain of transistor M12 functions as the drain of transistor M12, and the other of the source or drain of transistor M12 to which VSS is supplied functions as the source of transistor M12. Also, one of the source or drain of transistor M14 functions as the drain of transistor M14, and the other of the source or drain of transistor M14 to which VSS is supplied functions as the source of transistor M14.

[0085] Furthermore, when determining the current I2 according to the ratio of the channel sizes of transistors M12 and M14, the error increases if the drain voltages of transistors M12 and M14 are different. Therefore, it is preferable that the potentials of nodes Nd1 and Nd2 are equal.

[0086] By providing transistors M11 and M13 in the first circuit 101 in the configuration disclosed in this embodiment, the potentials of node Nd1 and node Nd2 can be aligned. Furthermore, in order to align the potentials of node Nd1 and node Nd2, it is preferable that the Vth of transistors M11 and M13 are equal. In addition, it is preferable that the ratio of the channel sizes of transistors M11 and M13 is equal to the ratio of the channel sizes of transistors M12 and M14.

[0087] If the channel length L of transistor M11 is L11 and the channel width W is W11, and the channel length L of transistor M13 is L13 and the channel width W is W13, then the relationship between the ratio of the channel sizes of transistors M11 and M13 and the ratio of the channel sizes of transistors M12 and M14 can be expressed by equation (2).

[0088]

[0089] Furthermore, by making the channel length L of transistors M12 and M14 sufficiently large, errors caused by the difference in drain voltage between transistors M12 and M14 (potential difference between node Nd1 and node Nd2) can be reduced. By making the channel length L of transistors M12 and M14 sufficiently large, the formation of transistors M11 and M13 can be made unnecessary. On the other hand, by providing transistors M11 and M13, the effect of increasing the output resistance due to cascode connection can be obtained, and the increase or decrease of the current value can be performed accurately with a small occupied area.

[0090] The second circuit 102 functions in the same way as the first circuit 101. For example, if the channel sizes of transistors M22 and M24 are equal, the current I2 flowing through node Nd3 and the current I3 flowing through node Nd4 will be equal. Current I3 is supplied to the output terminal OUT. Furthermore, the second circuit 102 can be understood by replacing transistor M11 with transistor M21, transistor M12 with transistor M22, transistor M23 with transistor M13, and transistor M24 with transistor M14 in the explanation of the first circuit 101 described above.

[0091] In the second circuit 102, one of the sources or drains of transistor M22 connected to node Nd3 functions as the drain of transistor M22, and the other of the sources or drains of transistor M22 to which VDD is supplied functions as the source of transistor M22. Also, one of the sources or drains of transistor M24 connected to node Nd4 functions as the drain of transistor M24, and the other of the sources or drains of transistor M24 to which VDD is supplied functions as the source of transistor M24.

[0092] In the second circuit 102, the magnitude of the current I3 can be determined by the ratio of the channel sizes of transistors M22 and M24. As mentioned above, it is preferable to adjust the ratio of the channel sizes of transistors M22 and M24 by changing the channel width W without changing the channel length L. The channel width W can also be achieved by connecting the transistors in parallel.

[0093] Figure 3 shows an example configuration of a semiconductor device 100 in which four transistors M22 are connected in parallel. In Figure 3, the four transistors M22 are referred to as transistors M22[1] to M22[4]. By connecting four transistors with the same channel size in parallel, the same effect as quadrupling the channel width W can be obtained. Therefore, for example, "increasing the channel width W of one transistor by n times" and "connecting n transistors with the same channel size in parallel" are synonymous. Also, for example, "increasing the channel width W of a transistor" and "increasing the number of transistors connected in parallel" can be synonymous. Also, for example, "decreasing the channel width W of a transistor" and "decreasing the number of transistors connected in parallel" can be synonymous.

[0094] Furthermore, in order to satisfy equation (2), when connecting multiple transistors M22 in parallel, it is preferable to connect the same number of transistors M21 in parallel. In the semiconductor device 100 shown in Figure 3, four transistors M21 are also connected in parallel. In Figure 3, the four transistors M21 are shown as transistors M21[1] to M21[4].

[0095] If transistors M22 and M24 have the same configuration, the amount of current I3 can be reduced to one-nth of current I2 by connecting n transistors M22 in parallel to one transistor M24. In the circuit configuration shown in Figure 3, the sum of the currents supplied from the four paths from input terminal IN[1] to input terminal IN[4] corresponds to current I1. Also, current I2, which has the same value as current I1, flows to node Nd3. Then, the second circuit 102 supplies current I3, which has a value of 1 / 4 of current I2, to the output terminal OUT.

[0096] In one embodiment of the present invention, in a semiconductor device 100, for example, by making the number of input terminals IN the same as the number of transistors M22 in parallel, the average value of the current supplied from the input terminals IN can be supplied to the output terminal OUT.

[0097] <Operation Example> Next, an operation example of the semiconductor device 100 will be described. As an example, an operation example of the semiconductor device 100 that outputs the average value of the current input from four input terminals IN (input terminal IN[1] to input terminal IN[4]) will be described.

[0098] Figure 4 is a timing chart illustrating an example of the operation of the semiconductor device 100. Figures 5 to 7 are circuit diagrams illustrating an example of the operation of the semiconductor device 100. In Figures 5 to 7, the standardized channel size is indicated by a number enclosed in an "×". Arrows indicating the direction of current flow are also included.

[0099] In the semiconductor device 100 shown in Figures 5 to 7, the normalized channel sizes of transistors M12 and M14 in the first circuit 101 are both 1, and the channel size ratio between them is 1. In other words, the channel size ratio of the first circuit 101 is 1. Furthermore, in order to satisfy equation (2), the channel size ratio of transistors M12 and M14 is also assumed to be 1.

[0100] Furthermore, the normalized channel size of transistor M22 in the second circuit 102 is assumed to be 4, and the normalized channel size of transistor M24 is assumed to be 1. Therefore, the channel size ratio of the two is 1 / 4. Thus, the channel size ratio of the second circuit 102 is also 1 / 4. For example, by making the channel lengths L of transistors M22 and M24 equal, and making the channel width W of transistor M22 four times the channel width W of transistor M24, the channel size ratio can be made 1 / 4. In addition, in order to satisfy equation (2), the channel size ratio of transistors M21 and M23 is also assumed to be 1 / 4.

[0101] Consider the case where a current of 1 nA is supplied from input terminal IN[1], a current of 2 nA is supplied from input terminal IN[2], a current of 3 nA is supplied from input terminal IN[3], and a current of 4 nA is supplied from input terminal IN[4] (period T11, see Figures 4 and 5). Then, the currents from input terminals IN[1] through IN[4] are added together, and a current of 10 nA I1 flows through node Nd1.

[0102] Since the channel size ratio in the first circuit 101 is 1, the current flowing to node Nd1 is copied, and a current of 10nA I2 flows to node Nd2 and node Nd3 of the second circuit 102 (period T12, see Figures 4 and 6).

[0103] Since the channel size ratio in the second circuit 102 is 1 / 4, a current of 2.5 nA, I3, which is 1 / 4 of the current I2 flowing through node Nd3, flows through node Nd4 (period T13, see Figures 4 and 7). In addition, current I3 is supplied to the output terminal OUT.

[0104] In this way, the average value of the currents input to multiple input terminals IN can be output from the output terminal OUT.

[0105] <Modification> When the channel size ratio of the first circuit 101 is 1, if the current I1, which is the sum of the currents supplied from multiple input terminals IN, is small, the current I2 flowing through node Nd3 of the second circuit 102 will also be small. In this case, the potential rise of node Nd3 will be slower, and the time taken from input to output will be slower. In addition, there is a risk that an accurate current value output may not be obtained. Therefore, there is a risk of a decrease in the operating speed and reliability of the semiconductor device 100.

[0106] Therefore, as shown in Figure 8, it is preferable to make the channel size ratio of the first circuit 101 greater than 1. For example, it is preferable to make the channel width W of transistor M14 greater than the channel width W of transistor M12. In this case, in order to satisfy equation (2), it is preferable to make the channel width W of transistor M11 greater than the channel width W of transistor M13. By making the channel size ratio of the first circuit 101 greater than 1, the amount of current I2 can be made greater than the amount of current I1. By making the amount of current I2 greater, the rate of potential rise at node Nd3 of the second circuit 102 can be increased. Thus, the operating speed of the semiconductor device 100 can be increased. In addition, the reliability of the semiconductor device 100 can be increased.

[0107] Furthermore, in order to obtain an accurate average value, it is necessary to decrease the channel size ratio of the second circuit 102 by the same amount as increasing the channel size ratio of the first circuit 101. If the channel size ratio of the first circuit 101 is k (where k is a number greater than 1), then in order to obtain the average value of the current supplied from n input terminals IN, the channel size ratio of the second circuit 102 must be 1 / (n × k). For example, if the average value of the current supplied from four input terminals IN (n=4) is output to the output terminal OUT, and the channel size ratio of the first circuit 101 is 2 (k=2), then it is possible to output the average value by reducing the channel size ratio of the first circuit 101 to 1 / 8.

[0108] Furthermore, if the channel size ratio of the second circuit 102 is less than 1, increasing the number of parallel transistors M22 too much may increase the parasitic capacitance of node Nd3, potentially reducing the rate of potential rise of node Nd3. This may reduce the operating speed of the semiconductor device 100. It may also increase the power consumption of the semiconductor device 100. Additionally, it may increase the occupied area of ​​the semiconductor device 100. For these reasons, the number of parallel transistors M22 is preferably 10 or less, and more preferably 5 or less.

[0109] For example, when the channel size ratio of the second circuit 102 is set to 1 / 8, it is preferable to set the number of parallel transistors M22 to 4 and the channel size of transistor M24 to 1 / 2 that of one transistor M22. For example, it is preferable to make the channel lengths L of transistor M24 and transistor M22 equal and the channel width W of transistor M24 to 1 / 2 that of one transistor M22. By doing so, the increase in parasitic capacitance of node Nd3 can be suppressed.

[0110] This embodiment can be implemented in appropriate combination with other embodiments described herein.

[0111] (Embodiment 2) A semiconductor device 100 according to one aspect of the present invention is suitable, for example, as a pooling layer of a convolutional neural network (CNN), which is a type of artificial neural network. In this embodiment, a CNN will be described. CNNs are often used for image recognition. A CNN mainly consists of a convolutional layer that extracts features, a pooling layer that compresses the feature quantities obtained in the convolutional layer, and a fully connected layer that performs classification based on the feature quantities obtained through the convolutional layer and the pooling layer.

[0112] As an example of a CNN, AlexNet will be described. AlexNet is an artificial neural network model shown in Figure 9, and AlexNet has an input layer INLY, convolutional layers CNV1 to CNV5, pooling layer PL1, pooling layer PL2, pooling layer PL5, and fully connected layers FC6 to FC8. As shown in Figure 9, AlexNet is composed in the following order: input layer INLY, convolutional layer CNV1, pooling layer PL1, convolutional layer CNV2, pooling layer PL2, convolutional layer CNV3, convolutional layer CNV4, convolutional layer CNV5, pooling layer PL5, fully connected layer FC6, fully connected layer FC7, and fully connected layer FC8.

[0113] [Input Layer INLY] As an input to AlexNet, the input layer INLY contains, for example, a 224x224 pixel image P. inis input. It is assumed that one pixel includes red, green, and blue sub-pixels, and the total number of sub-pixels is 3 colors (red, green, blue) × 224 × 224. Further, image P in has three channels, namely red, green, and blue. Therefore, image P in the number of image data included in the input layer INLY is 3 × 224 × 224.

[0114] Note that in the present embodiment, image P in , the input value at the x-th row and y-th column included in the z-th input channel (where z is an integer of 1 to 3 inclusive) is denoted as p in [x, y, z]. x indicates the row address of image P in and y indicates the column address of image P in That is, in the input layer INLY, x is an integer of 1 to 224 inclusive, and y is an integer of 1 to 224 inclusive.

[0115] [Convolutional Layer CNV1] In the convolutional layer CNV1, a convolution process is performed on image P in Specifically, a product-sum operation is performed between a filter (also referred to as a "kernel") used for the convolution process and image data included in a region A selected from image P in selected from in .

[0116] In the convolutional layer CNV1, the filter size (also referred to as "kernel size" or "window size") is set to 11, the number of output channels (also referred to as "number of kernels") is set to 96, and the stride is set to 4, and a convolution process is performed on a region selected from image P in Further, the number of filter values in one kernel is (filter size) 2 × (number of input channels). Since the number of input channels of image P in is 3, the number of filter values in one kernel in the convolutional layer CNV1 is 11 × 11 × 3 = 363.

[0117] Here, the s-th kernel in the convolutional layer CNV1 (where s is an integer of 1 to 96 inclusive) is denoted as K C1(s) It is stated as follows. Also, kernel K C1 Among the filter values ​​included, select any filter value k C1 (s) This is denoted as [p, q, r]. Here, p represents the row address of the kernel, q represents the column address of the kernel, and r represents the ordinal number of the input channel. In other words, in the convolutional layer CNV1, p is an integer between 1 and 11 (inclusive), q is an integer between 1 and 11 (inclusive), and r is an integer between 1 and 3 (inclusive).

[0118] Figure 10 shows image P in Selected region A in (1) and kernel K C1 (1) Then, a sum-of-products operation is performed, and the result of the operation is the data p. C1 (1) An example of outputting (1) is shown. Note that area A in (x) where x is image P in This shows the ordinal number of the selected region. Also, data p C1 (s) In (x), s indicates the ordinal number of the output channel. Also, data p C1 (s) (x) x is in region A in This corresponds to the ordinal number x of (x).

[0119] Also, since the stride is 4, region A in (1) The region shifted four rows from the original region is region A in (2) For example, Figure 11 shows image P in Selected region A in (2) and kernel K C1 (1) Then, a sum-of-products operation is performed, and the result of the operation is the data p. C1 (1) This shows an example of outputting (2).

[0120] Note: Image P in If the number of pixels is 224 x 224 and the stride is 4, then image P in The number of regions to be selected is 3025 (= 55 2 ) And in this embodiment, each of them is region A in(1) to area A in It is referred to as (3025).

[0121] As described above, depending on the number of strides, Image P in The region selected from is sequentially shifted, and each time the region is shifted, the region and kernel K C1 (1) By performing a sum-of-accumulate operation with this, a 55x55 matrix-like output data is obtained. Furthermore, the kernel included in the convolutional layer CNV1 is kernel K C1 (1) or kernel K C1 (96) Therefore (because the number of kernels in the convolutional layer CNV1 is 96), as a result, the output data P from the convolutional layer CNV1 is 55 × 55 × 96. C1 The following will be output.

[0122] In this way, region A in (1) to area A in (3025) and kernel K C1 (1) ~K C1 (96) Using the filter value and as input values, data P is extracted from the convolutional layer CNV1. C1 You can obtain this.

[0123] [Pooling layer PL1] In pooling layer PL1, the output data P from convolutional layer CNV1 is output. C1 Pooling is performed on the data. Pooling is a process in which predetermined regions are sequentially selected from the data output from a convolutional layer, and features are extracted by performing predetermined processing on each region, and these features are arranged in a matrix.

[0124] As shown in Figure 9, in the pooling layer PL1, with a kernel size of 3, data P C1Pooling processing shall be performed on each region selected from the above. The stride is set to 2, and the pooling processing is average pooling. It is also possible to make the kernel size equal to the stride in the pooling processing. For example, by performing pooling processing with both the kernel size and the stride set to 3, the data amount can be compressed to 1 / 9.

[0125] For example, in FIGS. 12A1 and 12A2, data P C1 region A selected from the first input channel of C1in (1) (1), average pooling processing is performed, and the processing result data p p1 (1) (1) is output as an example. In addition, since the kernel size is 3, region A C1in (1) (1) includes 3×3 pieces of data.

[0126] Note that s in region A C1in (s) (sA) indicates the ordinal number of the input channel of data P C1 and A in region A C1in (s) (sA) indicates the ordinal number of the region selected from data P C1 Furthermore, s in data p p1 (s) (sA) indicates the ordinal number of the output channel, and A in data p p1 (s) (sA) corresponds to the ordinal number A of region A C1in (s) (sA).

[0127] Furthermore, since the stride is 2, a region shifted by 2 in the row direction from region A C1in (1) (1) is defined as region A C1in (1) (2). For example, in FIGS. 12B1 and 12B2, data P C1 region A selected from C1in (1) (2), average pooling processing is performed, and the processing result data p p1 (1)shows an example of outputting (2).

[0128] Note that, for data P C1 , when the number is 55×55×96 and the stride is 2, the number of regions selected from data P C1 is 729 (=27 2 ).

[0129] As described above, by sequentially performing pooling processing on regions selected from data P C1 according to the number of strides, matrix-shaped output data of 27 rows and 27 columns can be obtained. Further, although the first input channel of data P C1 has been described above, the same pooling processing is also performed for the 2nd to 96th input channels, and as a result, data P, which is 27×27×96 output data, is output from the pooling layer PL1. P1 is output.

[0130] The semiconductor device 100 according to one aspect of the present invention is suitable for a pooling layer that performs average pooling processing. For example, when the kernel size of the pooling layer is 3, the average pooling processing can be performed by setting the number of input terminals IN of the semiconductor device 100 to 9.

[0131] The semiconductor device 100 according to one aspect of the present invention can implement average pooling processing without converting input data into digital values. Therefore, especially when convolution processing is performed by analog multiply-accumulate operations, there is no need to convert the multiply-accumulate operation result into a digital value. Further, there is no need to convert the result of pooling processing performed with digital values into analog values.

[0132] In other words, since there is no need to interpose an analog-to-digital converter (ADC) and a digital-to-analog converter (DAC) between the convolutional layer and the pooling layer, the circuit size of the semiconductor device that implements the CNN can be kept small. Furthermore, by using the semiconductor device 100 according to one aspect of the present invention as the pooling layer, the ADC and DAC become unnecessary, thereby increasing the processing speed of the CNN. In addition, by using the semiconductor device 100 according to one aspect of the present invention as the pooling layer, the power consumption of the CNN can be reduced.

[0133] [Convolutional layer CNV2] In convolutional layer CNV2, the data P output by pooling layer PL1 P1 A convolution operation is performed on the data P. Specifically, a kernel used for the convolution operation and the data P are used. P1 A sum-of-products operation is performed on the data contained in the selected region.

[0134] As shown in Figure 9, in the convolutional layer CNV2, with a kernel size of 5 and 256 kernels, data P P1 A convolution operation is performed on the selected region. The stride is set to 1.

[0135] Similar to the explanation for convolutional layer CNV1, convolutional processing is performed in convolutional layer CNV2, resulting in the output data P, which is 27 × 27 × 256 pixels, from convolutional layer CNV2. C2 The following will be output.

[0136] [Pooling layer PL2] In pooling layer PL2, the output data P from convolutional layer CNV2 is processed. C2 Pooling is performed on the data.

[0137] As shown in Figure 9, in the pooling layer PL2, with a kernel size of 3, the data P C2 Pooling is performed on each region selected from the set. The stride is set to 2, and the pooling is performed using average pooling.

[0138] Similar to the explanation for pooling layer PL1, pooling processing is performed in pooling layer PL2, resulting in output data P, which is 13 × 13 × 256 pixels from pooling layer PL2. P2 The following will be output.

[0139] Similar to the explanation for the pooling layer PL1, the semiconductor device 100 according to one aspect of the present invention is also suitable for the pooling layer PL2.

[0140] [Convolutional layer CNV3] In the convolutional layer CNV3, the data P output by the pooling layer PL2 P2 A convolution operation is performed on the data P. Specifically, a kernel used for the convolution operation and the data P are used. P2 A sum-of-products operation is performed on the data contained in the selected region.

[0141] As shown in Figure 9, in the convolutional layer CNV3, the kernel size is 3 and the number of kernels is 384, and the data P P2 A convolution operation is performed on the selected region. The stride is set to 1.

[0142] Similar to the explanation for convolutional layer CNV1, convolutional processing is performed in convolutional layer CNV3, resulting in the output data P, which is 13 × 13 × 384 pixels, being generated from convolutional layer CNV3. C3 The following will be output.

[0143] [Convolutional layer CNV4] In convolutional layer CNV4, the data P output by convolutional layer CNV3 C3 A convolution operation is performed on the data P. Specifically, a kernel used for the convolution operation and the data P are used. C3 A sum-of-products operation is performed on the data contained in the selected region.

[0144] As shown in Figure 9, in the convolutional layer CNV4, the kernel size is 3 and the number of kernels is 384, and the data P C3 A convolution operation is performed on the selected region. The stride is set to 1.

[0145] Similar to the explanation for convolutional layer CNV1, convolutional processing is performed in convolutional layer CNV4, resulting in the output data P, which is 13 × 13 × 384 pixels, being generated from convolutional layer CNV4. C4 The following will be output.

[0146] [Convolutional layer CNV5] In convolutional layer CNV5, the data P output by convolutional layer CNV4 C4 A convolution operation is performed on the data P. Specifically, a kernel used for the convolution operation and the data P are used. C4 A sum-of-products operation is performed on the data contained in the selected region.

[0147] As shown in Figure 9, in the convolutional layer CNV5, with a kernel size of 3 and 256 kernels, data P C4 A convolution operation is performed on the selected region. The stride is set to 1.

[0148] Similar to the explanation for convolutional layer CNV1, convolutional processing is performed in convolutional layer CNV5, resulting in the output data P, which is 13 × 13 × 256 pixels, being generated from convolutional layer CNV5. C5 The following will be output.

[0149] [Pooling layer PL5] In pooling layer PL5, the output data P from convolutional layer CNV5 is processed. C5 Pooling is performed on the data.

[0150] As shown in Figure 9, in the pooling layer PL5, with a kernel size of 3, the data P C5 Pooling is performed on each region selected from the set. The stride is set to 2, and the pooling is performed using average pooling.

[0151] Similar to the explanation for pooling layer PL1, pooling processing is performed in pooling layer PL5, resulting in 6x6x256 output data, data P. P5 The output is as follows. Furthermore, the semiconductor device 100 according to one aspect of the present invention is also suitable for the pooling layer PL5.

[0152] [Fully Connected Layer FC6] In the fully connected layer FC6, the output data P from the pooling layer PL5 is output. P5 For this, the fully connected layer operations are performed.

[0153] As shown in Figure 9, the fully connected layer FC6 has 9126 input channels (6 × 6 × 256) and 4096 output channels. In a fully connected layer, a sum-of-products operation is performed on all the input channel data and corresponding weight coefficients (first data) as one output channel, and the value of the activation function is calculated using the result as the input value. Therefore, the number of weight coefficients (first data) required in the fully connected layer FC6 is 4096 × 9126.

[0154] The data that will become the output channel of the Nth (where N is an integer between 1 and 4096) of the fully connected layer FC6 is z FC6 When (N) is the case, z FC6 (N) can be expressed by the following formula (3.1).

[0155]

[0156] However, f is the activation function in the fully connected layer FC6. Examples of activation functions include the sigmoid function, tanh function, softmax function, ReLU function, or threshold function. Also, u FC6 (N) is as shown in the following formula (3.2).

[0157]

[0158] Note, p p5 (s) (A) is the A-th data of the s-th output channel output in pooling layer PL5. Also, w FC6(N) (s) (A) is the Nth channel of the fully connected layer FC6 and p p5 (s) (A) and the corresponding weight coefficients (first data).

[0159] By using equations (2.1) and (2.2) above, the data of the 1st to 4096th output channels of the fully connected layer FC6, z FC6 (1) to zFC6 (4096) can be found.

[0160] [Fully Connected Layer FC7] In the fully connected layer FC7, the data from the output channel of the fully connected layer FC6 is z FC6 (1) to z FC6 The fully connected layer performs calculations on (4096).

[0161] As shown in Figure 9, the fully connected layer FC7 has 4096 input channels and 4096 output channels. Similar to the fully connected layer FC6, the fully connected layer FC7 performs a sum-of-products operation on all the input channel data and corresponding weight coefficients (first data) as a single output channel, and calculates the value of the activation function using the result as the input value. Therefore, the number of weight coefficients (first data) required in the fully connected layer FC7 is 4096 × 4096.

[0162] For details on sum-of-products operations and activation function calculations in the fully connected layer FC7, please refer to the explanation for the fully connected layer FC6.

[0163] In the fully connected layer FC7, the data from the output channel of the fully connected layer FC6 is z FC6 (1) to z FC6 When (4096) is input, the data of the output channels from the 1st to the 4096th of the fully connected layer FC7, z FC7 (1) to z FC7 Output (4096).

[0164] [Fully Connected Layer FC8] In the fully connected layer FC8, the data from the output channel of the fully connected layer FC7 is z FC7 (1) to z FC7 The fully connected layer performs calculations on (4096).

[0165] As shown in Figure 9, the fully connected layer FC8 has 4096 input channels and 1000 output channels. Similar to the fully connected layer FC6, the fully connected layer FC8 performs a sum-of-products operation on all the input channel data and corresponding weight coefficients (first data) as one output channel, and calculates the value of the activation function using the result as the input value. Therefore, the number of weight coefficients (first data) required in the fully connected layer FC8 is 1000 × 4096.

[0166] For details on sum-of-products operations and activation function calculations in the fully connected layer FC8, please refer to the explanation for the fully connected layer FC6.

[0167] In the fully connected layer FC8, the data from the output channel of the fully connected layer FC7 is z FC7 (1) to z FC7 When (4096) is input, the data of the first to 1000th output channels of the fully connected layer FC8, z FC8 (1) to z FC8 Output (1000).

[0168] Furthermore, for the fully coupled layer FC8, you can refer to the explanation of the operation of the fully coupled layer FC7 described above.

[0169] In this embodiment, for example, in the input layer INLY, a 224 x 224 pixel image P in While this method uses [specific method / framework], the image size is not limited to this. Furthermore, the number of kernels used in convolutional layers CNV1 through CNV5, and the filter values ​​included in them, can be arbitrarily determined. Note that pooling is not limited to average pooling; maximum pooling, minimum pooling, Lp pooling, etc., can also be used.

[0170] As described above, the semiconductor device according to one aspect of the present invention can be used in AlexNet shown in Figure 9. Furthermore, the computational model to which the semiconductor device according to one aspect of the present invention can be applied is not limited to AlexNet. The semiconductor device according to one aspect of the present invention can also be used in convolutional neural networks composed of computational models other than AlexNet.

[0171] This embodiment can be implemented in appropriate combination with other embodiments described herein.

[0172] (Embodiment 3) This embodiment describes an example of a transistor configuration applicable to a semiconductor device according to one aspect of the present invention.

[0173] <Transistor Configuration Example 1> Figures 13A to 13D are plan views and cross-sectional views of transistor 200A. Figure 13A is a plan view of transistor 200A, and Figures 13B to 13D are schematic cross-sectional views corresponding to the cutting lines A1-A2, A3-A4, and A5-A6 in Figure 13A, respectively. Figure 13B corresponds to the cross-section of transistor 200A in the channel length direction, and Figures 13C and 13D correspond to the cross-section in the channel width direction, respectively. Figures 14A and 14B are cross-sectional views of transistor 200A, corresponding to enlarged views of Figure 13B. As mentioned above, in order to make the drawings easier to understand, some components may be omitted from the plan views, cross-sectional views, etc.

[0174] The transistor 200A includes a semiconductor layer 230 provided on an insulating layer 201 provided on a substrate (not shown), a conductive layer 242 (conductive layer 242a and conductive layer 242b) on the semiconductor layer 230, an insulating layer 250 on the semiconductor layer 230, and a conductive layer 260 on the insulating layer 250. An insulating layer 275 is provided covering the semiconductor layer 230 and the conductive layer 242, and an insulating layer 280 is provided on the insulating layer 275. The insulating layer 280 and the insulating layer 275 are provided with grooves that reach the semiconductor layer 230, and the conductive layer 242a and the conductive layer 242b are separated by these grooves. The insulating layer 250 is provided inside the grooves along the surfaces of the insulating layer 280, the insulating layer 275, the conductive layer 242a, the conductive layer 242b, and the semiconductor layer 230. The conductive layer 260 is provided on the insulating layer 250 so as to fill the grooves. Furthermore, insulating layers 282 and 285 are provided in order, covering insulating layer 280, insulating layer 250, and conductive layer 260.

[0175] The semiconductor layer 230 includes a region that functions as a channel formation region of the transistor 200A. The conductive layer 260 functions as a gate electrode of the transistor 200A. The insulating layer 250 functions as a gate insulating layer of the transistor 200A. The conductive layer 242a functions as either the source electrode or the drain electrode of the transistor 200A, and the conductive layer 242b functions as the other. The region of the semiconductor layer 230 that overlaps with the conductive layer 260 via the insulating layer 250, located between the region in contact with conductive layer 242a and the region in contact with conductive layer 242b, functions as a channel formation region. Furthermore, the region of the semiconductor layer 230 that overlaps with conductive layer 242a functions as either the source region or the drain region, and the region of the semiconductor layer 230 that overlaps with conductive layer 242b functions as the other source region or drain region.

[0176] The channel length L of transistor 200A can be expressed as the length in the X direction of the conductive layer 260 in the region overlapping with the semiconductor layer 230 (see Figure 13B). Furthermore, the channel of transistor 200A is formed between the source region and the drain region of the semiconductor layer 230. Therefore, the channel length L of transistor 200A can be expressed as the distance from the edge of opposite conductive layer 242a to the edge of conductive layer 242b.

[0177] Furthermore, the channel width W of transistor 200A can be expressed as the length in the Y direction of the semiconductor layer 230 in the region where it overlaps with the conductive layer 260 (see Figure 13C).

[0178] It is preferable that the conductive layers 242a and 242b have a laminated structure. It is preferable to use a conductor that is resistant to oxidation, such as a metal nitride, on the side in contact with the semiconductor layer 230. This prevents the conductive layers 242a and 242b from being excessively oxidized by the oxygen contained in the semiconductor layer 230. It is also preferable to use a metal or alloy with higher conductivity than the layer in contact with the semiconductor layer 230 on the side that does not contact the semiconductor layer 230. This allows the conductive layers 242a and 242b to function as highly conductive wiring or electrodes.

[0179] In conductive layers 242a and 242b, it is preferable to use a metal nitride on the side in contact with the semiconductor layer 230. For example, it is preferable to use nitrides containing tantalum, nitrides containing titanium, nitrides containing molybdenum, nitrides containing tungsten, nitrides containing ruthenium, nitrides containing tantalum and aluminum, or nitrides containing titanium and aluminum. In addition, for example, ruthenium, oxides containing ruthenium, oxides containing strontium and ruthenium, or oxides containing lanthanum and nickel can be used. These materials are preferred because they are conductive materials that are resistant to oxidation or materials that maintain conductivity even when absorbing oxygen.

[0180] The insulating layer 201 is a film that is in contact with the semiconductor layer 230, and it is preferable to use an oxide insulating film. For example, it is preferable to use silicon oxide or silicon oxynitride as the insulating layer 201.

[0181] It is also possible to provide an insulating layer that functions as a barrier layer (also called a "barrier insulating layer") between the insulating layer 201 and the semiconductor layer 230. Preferably, this insulating layer has barrier properties against hydrogen. Examples of hydrogen barrier insulating layers include oxides such as aluminum oxide, hafnium oxide, and tantalum oxide, and nitrides such as silicon nitride. In particular, when silicon oxide is used for the insulating layer 201, it is especially preferable to provide an insulating layer that has barrier properties against hydrogen because it has the characteristic of easily diffusing hydrogen. This makes it possible to keep the hydrogen concentration in the semiconductor layer 230 low, thereby improving the reliability of the transistor 200A.

[0182] To stabilize the electrical characteristics of transistor 200A, it is effective to reduce the impurity concentration in the semiconductor layer 230. Furthermore, in order to reduce the impurity concentration in the semiconductor layer 230, it is preferable to also reduce the impurity concentration in adjacent films. For example, when an oxide semiconductor is used as the semiconductor layer 230, impurities include hydrogen, nitrogen, alkali metals, alkaline earth metals, iron, nickel, silicon, etc. Note that impurities in the semiconductor layer 230 refer to elements other than the main components that make up the semiconductor layer 230. For example, elements with a concentration of less than 0.1 atomic percent can be considered impurities.

[0183] It is preferable to use an oxide semiconductor for the semiconductor layer 230. It is preferable to use indium oxide for the semiconductor layer 230. In particular, it is preferable to use a single-crystal indium oxide film. Furthermore, it is preferable to use a crystalline film for the semiconductor layer 230, and particularly preferable to use indium oxide with a single-crystal structure. By using indium oxide with a single-crystal structure, carrier scattering at grain boundaries can be suppressed, enabling the realization of a transistor with high field-effect mobility. Additionally, it is possible to realize a transistor with less variation in characteristics between multiple transistors. Furthermore, it is possible to realize a highly reliable transistor.

[0184] Furthermore, indium oxide having a polycrystalline or microcrystalline structure can also be used for the semiconductor layer 230. When using indium oxide having a polycrystalline structure, it is preferable that no grain boundaries are observed, at least in the channel formation region (the region superimposed with the conductive layer 260). This allows the same effects as when using indium oxide having a single-crystal structure to be achieved, even when using indium oxide having a polycrystalline structure.

[0185] When an oxide semiconductor is used for the semiconductor layer 230, the film thickness of the semiconductor layer 230 is more preferably 2 nm to 50 nm, more preferably 2.5 nm to 30 nm, more preferably 2.5 nm to 20 nm, more preferably 5 nm to 20 nm, and even more preferably 5 nm to 10 nm. By setting the film thickness of the semiconductor layer 230 within the above range, the crystallinity of the semiconductor layer 230 can be improved.

[0186] When an oxide semiconductor is used for the semiconductor layer 230, it is preferable that the insulating layer 250, which functions as a gate insulating layer, has the function of capturing and fixing hydrogen. This reduces the hydrogen concentration in the channel formation region of the semiconductor layer 230. This makes it possible to make the channel formation region i-type or substantially i-type.

[0187] In this case, the insulating layer 250 is preferably a laminated structure comprising a first layer in contact with the semiconductor layer 230, a second layer on the first layer, and a third layer on the second layer. In this case, it is preferable that the first layer has the function of capturing hydrogen and fixing hydrogen.

[0188] Examples of insulators having the function of capturing and fixing hydrogen include metal oxides having an amorphous structure. For the first layer, it is preferable to use a metal oxide such as magnesium oxide, or an oxide containing one or both of aluminum and hafnium. In such amorphous metal oxides, oxygen atoms have dangling bonds, and these dangling bonds may have the property of capturing or fixing hydrogen. In other words, amorphous metal oxides have a high ability to capture or fix hydrogen.

[0189] Furthermore, it is preferable to use a high-k material for the first layer. An example of a high-k material is an oxide containing either or both aluminum and hafnium. By using a high-k material as the first layer, it becomes possible to reduce the gate potential applied during transistor operation while maintaining the physical thickness of the gate insulating layer. Additionally, it becomes possible to thin the EOT (End-of-Touch) of the insulator that functions as the gate insulating layer.

[0190] It is preferable to use an oxide containing one or both of aluminum and hafnium as the first layer, and it is more preferable to use an oxide having an amorphous structure that contains one or both of aluminum and hafnium, and it is even more preferable to use aluminum oxide having an amorphous structure.

[0191] Next, it is preferable to use an insulator with a thermally stable structure, such as silicon oxide or silicon oxide-nitride, for the second layer.

[0192] Furthermore, a structure can be provided in which a fourth layer is placed on top of the second layer. In this case, the fourth layer can be an insulator that can be used for the first layer. For example, hafnium oxide can be used as the fourth layer. By providing the fourth layer between the third layer and the second layer, hydrogen contained in the second layer and the like can be captured and fixed more effectively.

[0193] The third layer preferably has oxygen barrier properties. The third layer is provided between the channel-forming region of the semiconductor layer 230 and the conductive layer 260, and between the insulating layer 280 and the conductive layer 260. This configuration suppresses the diffusion of oxygen contained in the channel-forming region of the semiconductor layer 230 into the conductive layer 260, thereby preventing the formation of oxygen vacancies in the channel-forming region of the semiconductor layer 230. Furthermore, it suppresses the diffusion of oxygen contained in the semiconductor layer 230 and the oxygen contained in the insulating layer 280 into the conductive layer 260, thereby preventing the oxidation of the conductive layer 260. The third layer preferably has lower oxygen permeability than at least the insulating layer 280. For example, it is preferable to use a silicon nitride film as the third layer. In this case, the third layer is an insulator containing at least nitrogen and silicon.

[0194] Furthermore, it is preferable that the third layer has hydrogen barrier properties. This prevents impurities such as hydrogen contained in the conductive layer 260 from diffusing into the semiconductor layer 230.

[0195] The insulating layer 275 preferably has barrier properties against oxygen. The insulating layer 275 is provided between the insulating layer 280 and the conductive layer 242a, and between the insulating layer 280 and the conductive layer 242b. This configuration suppresses the diffusion of oxygen contained in the insulating layer 280 into the conductive layers 242a and 242b. Therefore, it is possible to suppress the oxidation of the conductive layers 242a and 242b by the oxygen contained in the insulating layer 280, which increases their resistivity and reduces the on-current.

[0196] The insulating layer 275 is preferably less permeable to oxygen than the insulating layer 280. In addition, it is preferable that it is less permeable to hydrogen. For example, silicon nitride is preferably used as the insulating layer 275. In this case, the insulating layer 275 is an insulator having at least nitrogen and silicon.

[0197] Furthermore, in this embodiment, it is preferable to configure the device to suppress the diffusion of hydrogen from the outside into the transistor 200A and the like. For example, it is preferable to provide an insulator having the function of suppressing hydrogen diffusion so as to cover the transistor 200A. In the semiconductor device described in this embodiment, the insulator is, for example, an insulating layer 282. It is also possible to provide a similar film below the transistor 200A, as shown in Figure 14B. Figure 14B shows an example in which an insulating layer 283 is provided below the insulating layer 201.

[0198] The insulating layer 282 and insulating layer 283 preferably function as barrier insulating layers that suppress the diffusion of impurities such as water and hydrogen from the outside into the transistor 200A. Therefore, the insulating layer 282 and insulating layer 283 preferably function as barrier insulating layers that suppress the diffusion of impurities such as hydrogen atoms, hydrogen molecules, water molecules, nitrogen atoms, nitrogen molecules, nitrogen oxide molecules (N 2 O, NO, NO 2 It is preferable that the material be made of an insulating material that has the function of suppressing the diffusion of impurities such as copper atoms (i.e., the above-mentioned impurities do not easily permeate). Alternatively, it is preferable that the material be made of an insulating material that has the function of suppressing the diffusion of oxygen (i.e., at least one such as oxygen atoms and oxygen molecules) (i.e., the above-mentioned oxygen does not easily permeate).

[0199] The insulating layer 282 and insulating layer 283 preferably have an insulator that has the function of suppressing the diffusion of impurities such as water and hydrogen, and oxygen. For example, aluminum oxide, magnesium oxide, hafnium oxide, gallium oxide, silicon nitride, or silicon nitride oxide can be used. For example, it is preferable to use silicon nitride, which has higher hydrogen barrier properties, as the insulating layer 283. Also, for example, it is preferable that the insulating layer 282 has aluminum oxide or magnesium oxide, which have high hydrogen capture and hydrogen fixation functions. This makes it possible to suppress the diffusion of impurities such as water and hydrogen from the interlayer insulating film located outside the insulating layer 283 to the transistor 200A, etc. Also, it is possible to suppress the diffusion of oxygen contained in the insulating layer 280, etc., upward from the transistor 200A, etc. via the insulating layer 282, etc. Furthermore, by providing a film similar to one or both of the insulating layers 282 and 283 below the transistor 200A, it is possible to suppress the diffusion of impurities such as water and hydrogen from the substrate side to the transistor 200A, etc.

[0200] Furthermore, as shown in Figures 14A and 14B, an insulating layer 271a can be provided between the conductive layer 242a on the conductive layer 242a and the insulating layer 275, and an insulating layer 271b can be provided between the conductive layer 242b on the conductive layer 242b and the insulating layer 275. The insulating layer 271a functions as an etching stopper when the conductive layer 242a is processed. The insulating layer 271b functions as an etching stopper when the conductive layer 242b is processed. In other words, the insulating layer 271 (insulating layer 271a and insulating layer 271b) has the function of protecting the conductive layer 242. Also, since the insulating layer 271a is in contact with the conductive layer 242a and the insulating layer 271b is in contact with the conductive layer 242b, it is preferable that the insulating layer 271 is an inorganic insulator that does not easily oxidize the conductive layer 242. For example, the insulating layer 271 can be made into a laminated structure, with silicon nitride used on the side in contact with the conductive layer 242 and silicon oxide used on the other sides.

[0201] In Figures 14A and 14B, openings reaching the conductive layer 242a are formed in insulating layers 285, 283, 282, 280, 275, and 271a, and the conductive layer 240a and insulating layer 241a are provided within these openings. The insulating layer 241a is provided adjacent to the side wall of the opening, and the conductive layer 240a is provided inside the insulating layer 241a. In addition, openings reaching the conductive layer 242b are formed in insulating layers 285, 283, 282, 280, 275, and 271b, and the conductive layer 240b and insulating layer 241b are provided within these openings. The insulating layer 241b is provided adjacent to the side wall of the opening, and the conductive layer 240b is provided inside the insulating layer 241b. The conductive layer 240 (conductive layer 240a and conductive layer 240b) functions as a via connecting the conductive layer provided on the transistor 200A to the source or drain of the transistor 200A.

[0202] The conductive layers 240a and 240b are preferably made of conductive materials mainly composed of, for example, tungsten, copper, or aluminum. Furthermore, the conductive layers 240a and 240b can be arranged in a laminated structure.

[0203] For example, as shown in Figures 14A and 14B, it is also possible to have a two-layer laminated structure of conductive layers 240a and 240b. Conductive layer 240a has conductive layer 240a1 formed along the opening and conductive layer 240a2 formed inside conductive layer 240a1. Conductive layer 240b has conductive layer 240b1 formed along the opening and conductive layer 240b2 formed inside conductive layer 240b1.

[0204] It is preferable to use conductive materials such as tantalum, tantalum nitride, titanium, titanium nitride, ruthenium, or ruthenium oxide for conductive layers 240a1 and 240b1, which have the function of suppressing the permeation of impurities such as water and hydrogen. Furthermore, the conductive material having the function of suppressing the permeation of impurities such as water and hydrogen can be used as a single layer or in a laminated form. By providing conductive layers 240a1 and 240b1, it is possible to suppress the diffusion of impurities such as water and hydrogen into the semiconductor layer 230 through conductive layers 240a2 and 240b2. The conductive layers 240a2 and 240b2 may be made of conductive materials that can be used for conductive layers 240a and 240b described above.

[0205] Furthermore, as shown in Figure 13B, the upper surfaces of the conductive layers 240a and 240b can be formed to coincide with or substantially coincide with the upper surface of the insulating layer 285. Also, as shown in Figures 14A and 14B, the lower part of the conductive layer 240a may be formed to be embedded in the conductive layer 242a. Similarly, the lower part of the conductive layer 240b may be formed to be embedded in the conductive layer 242b.

[0206] As the insulating layer 241 (insulating layer 241a and insulating layer 241b), a barrier insulating layer that can be used for insulating layer 275, etc., may be used. For example, silicon nitride may be used as the insulating layer 241. The insulating layer 241 is provided in contact with insulating layers 285, 283, 282, 275, 271a, and 271b. This suppresses the diffusion of impurities such as water and hydrogen contained in the insulating layer 280, etc., into the semiconductor layer 230 through the conductive layer 240a and conductive layer 240b. Silicon nitride is particularly suitable because it has high barrier properties against hydrogen. In addition, it is possible to prevent oxygen contained in the insulating layer 280 from being absorbed by the conductive layer 240a and conductive layer 240b.

[0207] The conductive layer 260 functions as the gate electrode of the transistor 200A. Here, it is preferable that the conductive layer 260 extends in the channel width direction, as shown in Figures 13A and 13C. With this configuration, when multiple transistors are provided, the conductive layer 260 functions as wiring.

[0208] The conductive layer 260 may also have a laminated structure. Figures 14A and 14B show an example in which the conductive layer 260 has a conductive layer 260a located on the side in contact with the insulating layer 250 and a conductive layer 260b above it. In this case, it is preferable to use a conductive material that is resistant to oxidation, such as titanium, titanium nitride, tantalum, tantalum nitride, ruthenium, or ruthenium oxide, or a conductive material that has the function of suppressing oxygen diffusion, for the conductive layer 260a. It is also preferable to use a low-resistance conductive material such as tungsten, copper, or aluminum for the conductive layer 260b.

[0209] The insulating layer 280 and insulating layer 285 preferably have a low dielectric constant. By using a material with a low dielectric constant as the interlayer film, parasitic capacitance between wirings can be reduced. For example, the insulating layer 280 and insulating layer 285 preferably contain one or more of the following: silicon oxide, silicon oxynitride, silicon oxide with added fluorine, silicon oxide with added carbon, silicon oxide with added carbon and nitrogen, and silicon oxide with vacancies. Silicon oxide and silicon oxynitride are preferred because they are thermally stable. In particular, materials such as silicon oxide, silicon oxynitride, and silicon oxide with vacancies are preferred because they can easily form regions containing oxygen that is desorbed by heating.

[0210] <Transistor Configuration Example 2> Below, we will describe a configuration example of transistor 200B, which differs in some aspects from transistor 200A. Note that the following mainly describes the differences from the above. Therefore, explanations of parts that overlap with the above may be omitted.

[0211] Transistor 200B is a modified version of transistor 200A, and Figures 15A and 15B are enlarged cross-sectional views of transistor 200B. Transistor 200B has a conductive layer 205 that functions as a back gate. Transistor 200B shown in Figure 15A has a conductive layer 205 and an insulating layer 202 beneath an insulating layer 201.

[0212] In the transistor 200B shown in Figure 15A, the conductive layer 205 is provided so as to be embedded in the insulating layer 202. The insulating layer 201 is provided so as to cover the insulating layer 202 and the conductive layer 205.

[0213] The conductive layer 260 functions as the first gate of transistor 200B, and the conductive layer 205 functions as the second gate (back gate) of transistor 200B. The conductive layer 260 and the conductive layer 205 have overlapping regions via the semiconductor layer 230. The conductive layer 205 can be made of the same material as the conductive layer 260. The conductive layer 205 can also have a layered structure.

[0214] Furthermore, the insulating layer 250 functions as a first gate insulating layer, and the insulating layer 201 functions as a second gate insulating layer. In this case, the insulating layer 201 can be made into a laminated structure, and a high dielectric constant material such as hafnium oxide, aluminum oxide, or hafnium aluminate can be used in part thereof. For example, silicon oxide can be used as the insulating layer 202.

[0215] Furthermore, it is preferable to provide an insulating film, such as silicon nitride or aluminum oxide, which has oxygen barrier properties, between the insulating layer 202 and the conductive layer 205, as this can suppress oxidation of the conductive layer 205.

[0216] Furthermore, as shown in Figure 15B, it is also possible to provide an insulating layer 283 that functions as a barrier insulating layer between the insulating layer 201 and the conductive layer 205, as in the transistor 200B.

[0217] <Transistor Configuration Example 3> Figure 16A is a plan view of transistor 200C, which has a different configuration from transistors 200A and 200B. Figure 16B is a cross-sectional view corresponding to the cutting line shown A1-A2 in Figure 16A. In the following, we will mainly explain the parts that differ from the above. Therefore, explanations of parts that overlap with the above may be omitted.

[0218] Transistor 200C has a conductive layer 255 on top of an insulating layer 201. It also has an insulating layer 257 on top of the conductive layer 255, an insulating layer 258 on top of the insulating layer 257, and an insulating layer 259 on top of the insulating layer 258. In this specification, insulating layers 257, 258, and 259 may be collectively referred to as an insulating layer 256 or a spacer layer. It also has a conductive layer 261 on top of the insulating layer 259.

[0219] Furthermore, an opening 262 is provided in a region that overlaps with a part of the conductive layer 255, penetrating the conductive layer 261, the insulating layer 259, the insulating layer 258, and the insulating layer 257. A semiconductor layer 230 is also provided covering the inner wall of the opening 262.

[0220] The semiconductor layer 230 has a region that overlaps with the bottom of the opening 262 and a region that overlaps with the side of the opening 262. That is, the semiconductor layer 230 has a region that is in contact with the insulating layer 256 inside the opening 262. The semiconductor layer 230 also has a region that is in contact with the conductive layer 255 and a region that is in contact with the conductive layer 261 inside the opening 262.

[0221] Furthermore, an insulating layer 250 is provided on top of the insulating layer 259, the conductive layer 261, and the semiconductor layer 230. A conductive layer 265 is also provided on top of the insulating layer 250. The conductive layer 265 has a region that overlaps with the semiconductor layer 230. The conductive layer 265 has a region that overlaps with the semiconductor layer 230 via the insulating layer 250. The conductive layer 265 functions as a gate electrode. Therefore, the conductive layer 265 corresponds to the conductive layer 260 in transistors 200A and 200B.

[0222] Furthermore, each of the insulating layer 250 and the conductive layer 265 has a region that overlaps with the opening 262. Also, each of the insulating layer 250 and the conductive layer 265 has a region that overlaps with the inside of the opening 262. Inside the opening 262, the semiconductor layer 230 has a region that overlaps with the conductive layer 265 via the insulating layer 250, and a region that overlaps with the side surface of the opening 262 (the side surface of the insulating layer 256).

[0223] Furthermore, an insulating layer 285 is provided on top of the insulating layer 250. It is preferable that the upper surface of the insulating layer 285 is flat. Alternatively, it is preferable that the heights (positions in the Z-direction) of the upper surfaces of the insulating layer 285 and the conductive layer 265 coincide or substantially coincide. For example, the flatness of the upper surface of the insulating layer 285 can be improved by performing chemical mechanical polishing (CMP) treatment. Also, by performing CMP treatment, the positions of the upper surfaces of the insulating layer 285 and the conductive layer 265 can be made to coincide or substantially coincide. By performing CMP treatment, surface irregularities of the sample can be reduced, thereby improving the coverage of the insulating layer and conductive layer formed thereafter.

[0224] Furthermore, when an oxide semiconductor is used for the semiconductor layer 230, it is preferable that the conductive layer 255 and the conductive layer 261 in contact with the semiconductor layer 230 use a conductive material that converts the oxide semiconductor to n-type. For example, a conductive material containing nitrogen may be used. For example, a conductive material containing titanium or tantalum and nitrogen may be used. It is also possible to provide other conductive materials on top of the conductive material containing nitrogen.

[0225] Furthermore, when an oxide semiconductor is used for the semiconductor layer 230, it is preferable to use a material with reduced hydrogen and containing oxygen for the insulating layer 258. For example, a material containing silicon and oxygen may be used. Specifically, silicon oxide or silicon oxynitride may be used. Since hydrogen is an impurity element in oxide semiconductors, the contact between the oxide semiconductor semiconductor layer 230 and the hydrogen-reduced insulating layer 258 makes it less likely for the semiconductor layer 230 to become n-type. In addition, the contact between the oxide semiconductor semiconductor layer 230 and the oxygen-containing insulating layer 258 reduces oxygen vacancies in the semiconductor layer 230, stabilizing the transistor's characteristics and improving reliability.

[0226] Furthermore, when an oxide semiconductor is used for the semiconductor layer 230, the insulating layer 258 may contain excess oxygen. In this specification, excess oxygen refers to oxygen that is desorbed by heating. A material that desorbs oxygen by heating is one in which the amount of oxygen desorbed, converted to oxygen atoms, is 1.0 × 10¹⁶ by TDS (Thermal Desorption Spectroscopy) analysis.18 atoms / cm 3 Preferably 1.0 × 10 19 atoms / cm 3 More preferably 2.0 × 10 19 atoms / cm 3 The above or 3.0 x 10 20 atoms / cm 3 The material is as described above. Furthermore, the surface temperature of the film during the TDS analysis is preferably in the range of 100°C to 700°C or 100°C to 400°C.

[0227] Furthermore, when using a material containing excess oxygen for the insulating layer 258, it is preferable to use materials that are impermeable to oxygen for the insulating layers 257 and 259. Examples of materials that are impermeable to oxygen include oxides containing one or both of aluminum and hafnium, and silicon nitrides. By using materials that are impermeable to oxygen for the insulating layers 257 and 259, excess oxygen contained in the insulating layer 258 is less likely to desorb to the lower or upper layer. Therefore, sufficient oxygen can be supplied to the oxide semiconductor. For example, a configuration having an insulating layer (insulating layer 258) containing silicon and oxygen between two insulating layers (insulating layer 257 and insulating layer 259) containing silicon and nitrogen is preferred.

[0228] Furthermore, when an oxide semiconductor is used for the semiconductor layer 230, by using hydrogen-containing materials for the insulating layer 257 and the insulating layer 259, hydrogen is supplied to the region of the semiconductor layer 230 in contact with the insulating layer 257 and the region of the semiconductor layer 230 in contact with the insulating layer 259, and depending on the composition of the oxide semiconductor used for the semiconductor layer 230, each region becomes n-type. Therefore, the region of the semiconductor layer 230 in contact with the conductive layer 261 and the region of the semiconductor layer 230 in contact with the insulating layer 259 function as either a source region or a drain region. Also, the region of the semiconductor layer 230 in contact with the conductive layer 255 and the region of the semiconductor layer 230 in contact with the insulating layer 257 function as either a source region or a drain region.

[0229] The conductive layer 261 functions as either the source electrode or the drain electrode of transistor 200C. The conductive layer 255 functions as the other source electrode or drain electrode of transistor 200C. Therefore, the conductive layer 261 functions as either the conductive layer 242a or the conductive layer 242b in transistors 200A and 200B. Also, the conductive layer 255 functions as the other conductive layer 242a or the conductive layer 242b in transistors 200A and 200B.

[0230] Transistor 200C is a transistor in which the source electrode and drain electrode are arranged in the Z direction. That is, the source and drain of transistor 200C are positioned at different heights. In other words, the source and drain of transistor 200C are positioned at different locations in the Z direction. Such a transistor is also called a "vertical channel transistor," "vertical transistor," or "VFET (Vertical Field Effect Transistor)."

[0231] In the above configuration, for the VFET transistor 200C, the length of the side surface of the insulating layer 258 viewed from the X or Y direction becomes the channel length L (channel length L1) (see Figure 16B). Therefore, the channel length L of the transistor 200C is determined according to the thickness t1 of the insulating layer 258.

[0232] Furthermore, it is possible to use materials that do not contain hydrogen or contain very little hydrogen for the insulating layer 257 and insulating layer 259. When silicon nitride or silicon nitride oxide with very little hydrogen is used for the insulating layer 257 and insulating layer 259, the regions of the semiconductor layer 230 that are in contact with the insulating layer 257 and the regions of the semiconductor layer 230 that are in contact with the insulating layer 259 are not n-type. Therefore, the region of the semiconductor layer 230 that is in contact with the conductive layer 261 functions as either a source region or a drain region. Also, the region of the semiconductor layer 230 that is in contact with the conductive layer 255 functions as either a source region or a drain region. Furthermore, the region of the semiconductor layer 230 that is in contact with the insulating layer 258 functions as a channel-forming region.

[0233] In this case, the sum of the lengths of the sides of insulating layers 257, 258, and 259 as viewed from the X or Y direction becomes the channel length L (channel length L2). Therefore, the channel length L of transistor 200C is determined according to the thickness t2 obtained by adding the thicknesses of insulating layers 257, 258, and 259. Thus, transistor 200C has a channel formation region that is aligned with the side surface of insulating layer 256.

[0234] Furthermore, since the semiconductor layer 230 is provided in the opening 262, the length of the perimeter of the opening 262 when viewed from the Z direction becomes the channel width W of the transistor 200C (see Figure 16A). The length of the perimeter can be determined, for example, at a point where the thickness t1 of the insulating layer 258 is halfway, or at a point where the thickness t2 is halfway. If necessary, the length of the perimeter at any position of the opening 262 can be used as the channel width W. For example, the length of the perimeter at the bottom of the opening 262 can be used as the channel width W, or the length of the perimeter at the top of the opening 262 can be used as the channel width W. Also, although the contour (planar shape) of the opening 262 when viewed from the Z direction is shown as a circle in Figure 16A, it is not limited to this. For example, the contour of the opening 262 when viewed from the Z direction can be an ellipse, a rectangle, etc.

[0235] Furthermore, in order to improve the coverage of the semiconductor layer 230, insulating layer 250, and conductive layer 265 formed inside the opening 262, it is preferable that the taper angle θ of the side surface of the opening 262, that is, the taper angle θ of the side surfaces of the insulating layer 257, insulating layer 258, and insulating layer 259, be 45° or more and less than 90°, preferably 50° or more and 75° or less. Note that the taper angle θ of the side surface of a layer (insulating layer, conductive layer, or semiconductor layer) refers to the angle between the bottom surface and the side surface of the layer (see Figure 16B).

[0236] Vertical transistors can reduce the area occupied by a transistor (also called a "horizontal transistor") in which the channel formation region, source region, and drain region are separately located on the XY plane. Therefore, by using vertical channel transistors in semiconductor devices, the area occupied by the semiconductor device can be reduced. Furthermore, by using vertical channel transistors in semiconductor devices, high integration of the semiconductor device can be achieved.

[0237] Furthermore, in lateral transistors, the channel length was limited by the exposure limit of photolithography. In one aspect of the present invention, the channel length can be set by the thickness of the insulating layer 256 or insulating layer 258. Therefore, the channel length of the transistor can be made into an extremely fine structure below the exposure limit of photolithography (for example, 60 nm or less, 50 nm or less, 40 nm or less, 30 nm or less, 20 nm or less, or 10 nm or less, and 1 nm or more or 5 nm or more). This increases the on-current of transistor 200C, and improves the frequency characteristics. By using a vertical channel transistor, a semiconductor device with a high operating speed can be provided.

[0238] [Insulating Layer] Unless otherwise specified, various inorganic insulators can be used for the insulating layer (insulating layer 283, insulating layer 201, insulating layer 202, insulating layer 271, insulating layer 275, insulating layer 280, insulating layer 282, insulating layer 285, insulating layer 250, etc.) in a semiconductor device according to one embodiment of the present invention. For example, oxide insulators, nitride insulators, oxidized nitride insulators and nitrided oxide insulators can be used.

[0239] Examples of oxide insulators include silicon oxide, aluminum oxide, magnesium oxide, gallium oxide, germanium oxide, yttrium oxide, zirconium oxide, lanthanum oxide, neodymium oxide, hafnium oxide, tantalum oxide, cerium oxide, zinc gallium oxide, and hafnium aluminate. Examples of nitride insulators include silicon nitride and aluminum nitride. Examples of oxidative nitride insulators include silicon oxidative nitride, aluminum oxidative nitride, gallium oxidative nitride, yttrium oxidative nitride, and hafnium oxidative nitride. Examples of nitride oxide insulators include silicon nitride and aluminum nitride. Furthermore, organic insulators can also be used in the insulating layer of a semiconductor device according to one embodiment of the present invention.

[0240] In this specification, "oxide-nitride" refers to a material in which the oxygen content is greater than the nitrogen content, and "nitride oxide" refers to a material in which the nitrogen content is greater than the oxygen content. For example, when "silicon oxynitride" is written, it refers to a material in which the oxygen content is greater than the nitrogen content, and when "silicon nitride oxide" is written, it refers to a material in which the nitrogen content is greater than the oxygen content. The content of each element can be measured using methods such as Rutherford backscattering (RBS).

[0241] For example, as transistors become smaller and more integrated, thinning of the gate insulating layer can lead to problems such as increased leakage current. By using a high-k material for the insulating layer that functions as the gate insulating layer, it becomes possible to lower the voltage during transistor operation while maintaining the physical film thickness. It also becomes possible to thin the equivalent oxide film thickness (EOT) of the gate insulating layer. On the other hand, by using a material with a low relative permittivity for the insulating layer that functions as an interlayer insulating film (e.g., insulating layer 280, insulating layer 285, etc.), parasitic capacitance that occurs between conductive layers such as wiring can be reduced. Therefore, it is important to select the material according to the function of the insulating layer. It should be noted that materials with a low relative permittivity also have high dielectric strength.

[0242] Examples of materials with a high dielectric constant (high-k) include aluminum oxide, gallium oxide, hafnium oxide, tantalum oxide, zirconium oxide, hafnium-zirconium oxide, oxides containing aluminum and hafnium, oxides containing aluminum and hafnium, oxides containing silicon and hafnium, oxides containing silicon and hafnium, and nitrides containing silicon and hafnium.

[0243] Examples of materials with low dielectric constant include inorganic insulating materials such as silicon oxide, silicon oxide-nitride, and silicon nitride-oxide, as well as resins such as polyester, polyolefin, polyamide (nylon, aramid, etc.), polyimide, polycarbonate, and acrylic resin. Other inorganic insulating materials with low dielectric constant include, for example, silicon oxide with added fluorine, silicon oxide with added carbon, and silicon oxide with added carbon and nitrogen. Also, for example, silicon oxide with voids is another example. These silicon oxides may contain nitrogen.

[0244] [Conductive Layer] Unless otherwise specified, the conductive layers (conductive layer 205, conductive layer 242, conductive layer 240, conductive layer 255, conductive layer 260, conductive layer 261, conductive layer 265, etc.) in a semiconductor device according to one aspect of the present invention may be made of a metal element selected from aluminum, chromium, copper, silver, gold, platinum, zinc, tantalum, nickel, titanium, iron, cobalt, molybdenum, tungsten, hafnium, vanadium, niobium, manganese, magnesium, zirconium, beryllium, indium, ruthenium, iridium, strontium, lanthanum, etc., or an alloy containing the aforementioned metal elements, or an alloy combining the aforementioned metal elements.

[0245] As alloys composed of the aforementioned metal elements, nitrides or oxides of the alloys can also be used. For example, tantalum nitride, titanium nitride, nitrides containing titanium and aluminum, nitrides containing tantalum and aluminum, ruthenium oxide, ruthenium nitride, oxides containing strontium and ruthenium, oxides containing lanthanum and nickel, etc. can be used. In addition, semiconductors with high electrical conductivity, such as polycrystalline silicon containing impurity elements such as phosphorus, and silicides such as nickel silicide can also be used.

[0246] Furthermore, conductive materials containing nitrogen, such as nitrides containing tantalum, nitrides containing titanium, nitrides containing molybdenum, nitrides containing tungsten, nitrides containing ruthenium, nitrides containing tantalum and aluminum, nitrides containing titanium and aluminum, conductive materials containing oxygen, such as oxides containing ruthenium oxide, strontium and ruthenium, oxides containing lanthanum and nickel, and materials containing metallic elements such as titanium, tantalum, and ruthenium, are preferred because they are conductive materials that are resistant to oxidation, conductive materials that have the function of suppressing oxygen diffusion, or materials that maintain conductivity even when absorbing oxygen. Examples of conductive materials containing oxygen include oxides containing tungsten and indium, oxides containing titanium and indium, oxides containing indium and tin (ITO: Indium Tin Oxide), oxides containing titanium, indium and tin, silicon-added indium and tin oxide (also called ITSO), oxides containing indium and zinc (also called IZO®), and oxides containing tungsten, indium and zinc.

[0247] Furthermore, it is possible to use multiple conductive layers formed from the above materials in a laminated structure. For example, a laminated structure can be formed by combining the aforementioned metal element material with an oxygen-containing conductive material. Alternatively, a laminated structure can be formed by combining the aforementioned metal element material with a nitrogen-containing conductive material. Furthermore, a laminated structure can be formed by combining the aforementioned metal element material with an oxygen-containing conductive material and a nitrogen-containing conductive material.

[0248] When using an oxide semiconductor, which is a type of metal oxide, as the semiconductor layer 230, it is preferable to use a conductive material that is resistant to oxidation, a conductive material that maintains low electrical resistance even when oxidized, a conductive metal oxide (also called an "oxide conductive layer"), or a conductive material that has the function of suppressing oxygen diffusion as the conductive layer in contact with the semiconductor layer 230. Examples of such conductive materials include conductive materials containing nitrogen and conductive materials containing oxygen. This makes it possible to suppress a decrease in the conductivity of the conductive layer.

[0249] Furthermore, by using an oxide conductive layer as the conductive layer, the conductive layer can maintain its conductivity even if it absorbs oxygen. For example, even when an insulating layer containing oxygen that is desorbed by heating (also called "excess oxygen") is used as the insulating layer in contact with the conductive layer, the conductive layer can maintain its conductivity, making it suitable. For example, ITO, ITSO, IZO (registered trademarks), etc., can be used as the conductive layer.

[0250] [Semiconductor Layer] As the semiconductor layer, single-crystal semiconductors, polycrystalline semiconductors, microcrystalline semiconductors, amorphous semiconductors, etc., can be used individually or in combination. Using a single-crystal semiconductor or a crystalline semiconductor in the semiconductor layer where the channel is formed is preferable because it can suppress the degradation of transistor characteristics.

[0251] As semiconductor materials, for example, semiconductors composed of elemental elements such as silicon and germanium can be used. Alternatively, compound semiconductors such as silicon germanium, silicon carbide, gallium arsenide, and nitride semiconductors can be used. As compound semiconductors, organic materials with semiconductor properties (also called "organic semiconductors"), metal nitrides with semiconductor properties (also called "nitride semiconductors"), or metal oxides with semiconductor properties (also called "oxide semiconductors") can be used. These semiconductor materials may contain impurities as dopants.

[0252] When silicon is used as a semiconductor layer, examples of silicon that can be used for the semiconductor layer include single-crystal silicon, polycrystalline silicon, microcrystalline silicon, and amorphous silicon. An example of polycrystalline silicon is low-temperature polysilicon.

[0253] It is also possible to use a two-dimensional material that functions as a semiconductor as the semiconductor layer of a transistor. Two-dimensional materials, also called layered materials, are a general term for a group of materials that have a layered crystalline structure. Layered materials have high electrical conductivity within a unit layer, that is, high two-dimensional electrical conductivity. By using a material that functions as a semiconductor and has high two-dimensional electrical conductivity as the semiconductor layer, it is possible to provide a transistor with a large on-current.

[0254] Examples of the above-mentioned layered materials include graphene, silicene, and chalcogenides. Chalcogenides are compounds containing chalcogens (elements belonging to Group 16). Examples of chalcogenides include transition metal chalcogenides and Group 13 chalcogenides. Specifically, a transition metal chalcogenide applicable as a semiconductor layer in transistors is molybdenum sulfide (typically MoS 2 ), molybdenum selenide (typically MoSe 2 ), molybdenum tellurium (typically MoTe 2 ), tungsten sulfide (typically WS 2 ), tungsten selenide (typically WSe 2 ), tungsten tellurium (typically WTe 2 ), hafnium sulfide (typically HfS 2 ), hafnium selenide (typically HfSe 2 ), zirconium sulfide (typically ZrS 2 ), zirconium selenide (typically ZrSe 2 ) are some examples.

[0255] [Metal Oxide Layer] In one aspect of the present invention, it is preferable that the semiconductor layer 230 including the channel formation region has an oxide semiconductor, which is a type of metal oxide. That is, it is preferable to use an OS transistor as the transistor 200 (transistor 200A to transistor 200C).

[0256] OS transistors have oxygen vacancies (V) in the channel-forming region of a metal oxide that functions as a semiconductor. O ) and impurities can cause electrical properties to fluctuate easily, potentially leading to poor reliability. In addition, defects in which hydrogen enters the oxygen vacancy (hereinafter referred to as V) OThis can form an oxygen vacancy (sometimes called H) and generate electrons that act as carriers. Therefore, if the channel formation region in the metal oxide contains oxygen vacancies, the OS transistor is likely to exhibit normally-on characteristics. Consequently, it is preferable that oxygen vacancies and impurities are reduced as much as possible in the channel formation region of the metal oxide. In other words, it is preferable that the carrier concentration in the channel formation region of the metal oxide is reduced and that it is i-type (intrinsic) or substantially i-type.

[0257] On the other hand, the source and drain regions in the metal oxide that function as the semiconductor of the OS transistor have more oxygen vacancies than the channel formation region. O It is preferable that the region has a high concentration of H or high concentrations of impurities such as hydrogen, nitrogen, and metallic elements, which increases the carrier concentration and lowers the resistance. In other words, it is preferable that the source region and drain region of an OS transistor are n-type regions with a higher carrier concentration and lower resistance compared to the channel formation region.

[0258] The band gap of the metal oxide that functions as a semiconductor is preferably 2.0 eV or higher, and more preferably 2.5 eV or higher. By using a metal oxide that functions as a semiconductor and has a larger band gap than silicon in the semiconductor layer 230, the off-current of the transistor 200 can be reduced. Because the OS transistor has a small off-current, the power consumption of the semiconductor device can be significantly reduced. In addition, because the OS transistor has high frequency characteristics, the semiconductor device can be operated at high speed.

[0259] The metal oxide that can be used in the semiconductor layer of an OS transistor preferably contains at least indium (In). Furthermore, it is preferable that the metal oxide contains at least one of indium (In) or zinc (Zn). Moreover, it is preferable that the metal oxide contains two or three elements selected from indium, element M, and zinc. Element M is a metal or metalloid element with a high bond energy with oxygen, for example, a metal or metalloid element with a higher bond energy with oxygen than indium.

[0260] Specific examples of element M include aluminum, gallium, tin, yttrium, titanium, vanadium, chromium, manganese, iron, cobalt, nickel, zirconium, molybdenum, hafnium, tantalum, tungsten, lanthanum, cerium, neodymium, magnesium, calcium, strontium, barium, boron, silicon, germanium, and antimony. The element M contained in the metal oxide is preferably one or more of the above elements, more preferably one or more selected from aluminum, gallium, tin, and yttrium, and even more preferably gallium.

[0261] For example, indium oxide (In oxide, indium oxide) can be used as a metal oxide for the semiconductor layer of an OS transistor. Other metal oxides include zinc oxide (Zn oxide, zinc oxide), indium zinc oxide (In-Zn oxide), indium tin oxide (In-Sn oxide), indium titanium oxide (In-Ti oxide), indium gallium oxide (In-Ga oxide), indium gallium aluminum oxide (In-Ga-Al oxide), indium gallium tin oxide (In-Ga-Sn oxide), gallium zinc oxide (Ga-Zn oxide, also written as "GZO"), aluminum zinc oxide (Al-Zn oxide, also written as "AZO"), and indium Aluminum zinc oxide (In-Al-Zn oxide, also written as "IAZO"), indium tin zinc oxide (In-Sn-Zn oxide), indium titanium zinc oxide (In-Ti-Zn oxide), indium gallium zinc oxide (In-Ga-Zn oxide, also written as "IGZO"), indium gallium tin zinc oxide (In-Ga-Sn-Zn oxide, also written as "IGZTO"), indium gallium aluminum zinc oxide (In-Ga-Al-Zn oxide, also written as "IGAZO" or "IAGZO"), etc., can be used. Alternatively, silicon-containing indium tin oxide, gallium tin oxide (Ga-Sn oxide), aluminum tin oxide (Al-Sn oxide), etc., can be used.

[0262] Examples of crystal structures for metal oxides that function as semiconductors include amorphous (including completely amorphous), CAAC (c-axis-aligned crystalline), nc (nanocrystalline), CAC (cloud-aligned composite), single crystal, and polycrystalline.

[0263] Furthermore, by increasing the ratio of zinc atoms to the sum of the atoms of other metal elements in a metal oxide that functions as a semiconductor, a highly crystalline metal oxide can be obtained, suppressing the diffusion of impurities within the metal oxide. Consequently, fluctuations in the electrical properties of the transistor can be suppressed, improving reliability.

[0264] Furthermore, by increasing the ratio of element M atoms to the sum of the atoms of metal elements among the main constituent elements contained in the metal oxide, the formation of oxygen vacancies in the metal oxide can be suppressed. Therefore, carrier generation caused by oxygen vacancies is suppressed, resulting in a transistor with low off-current. In addition, fluctuations in the electrical characteristics of the transistor are suppressed, and reliability can be improved.

[0265] By increasing the ratio of indium atoms to the sum of all metal element atoms in a metal oxide, the field-effect mobility of a transistor can be improved. Typically, using single-crystal or polycrystalline indium oxide in the semiconductor layer significantly increases the field-effect mobility of a transistor. Furthermore, transistors using single-crystal or polycrystalline indium oxide in the semiconductor layer can achieve excellent frequency characteristics.

[0266] This embodiment can be implemented in appropriate combination with other embodiments described herein.

[0267] (Embodiment 4) This embodiment describes an indium oxide film that can be used in the semiconductor layer of a transistor in a semiconductor device according to one aspect of the present invention.

[0268] In this specification, indium oxide having at least a crystalline portion or crystalline region in the film is referred to as crystalline indium oxide (crystal IO) or crystalline indium oxide (crystalline IO). Examples of crystal IO or crystalline IO include single-crystal indium oxide, polycrystalline indium oxide, and microcrystalline indium oxide.

[0269] Indium oxide is a semiconductor material with completely different physical properties from oxide semiconductors such as In-Ga-Zn oxide (hereinafter also referred to as IGZO) and zinc oxide.

[0270] This paper describes the carrier concentration dependence of the hole mobility of indium oxide, silicon, and IGZO.

[0271] IGZO tends to exhibit higher hole mobility as the carrier concentration increases. On the other hand, single-crystal indium oxide tends to exhibit higher hole mobility as the carrier concentration decreases. This trend is similar to that of silicon, where lower dopant (impurity) concentrations in the material reduce impurity scattering and increase hole mobility. In other words, the higher the purity and intrinsic nature of single-crystal indium oxide, the higher its hole mobility. From these results, it can be said that single-crystal indium oxide, unlike IGZO, is a material with physical properties similar to silicon. Note that when indium oxide is not single-crystal (e.g., polycrystalline), the trend may differ from that of single crystals.

[0272] The range of carrier concentrations suitable for the channel formation region of a transistor is 1 × 10⁻⁶. 15 cm −3 This range includes, for example, 1 × 10 14 cm −3 The above is 1 x 10 18 cm −3 The range is as follows: By sufficiently reducing the carrier concentration, the hole mobility value can be increased to 270 cm⁻¹. 2 It can be expected to be raised to the level of / (V・s).

[0273] Indium oxide can contain elements that lower the carrier concentration. Examples of elements that lower the carrier concentration include magnesium, calcium, zinc, cadmium, and copper. These elements can lower the carrier concentration by substituting for indium. Other examples include nitrogen, phosphorus, arsenic, and antimony. These elements can lower the carrier concentration by substituting for oxygen.

[0274] On the other hand, electrical resistance can be reduced by increasing the carrier concentration. For example, the suitable carrier concentration range for the source and drain regions of a transistor, or for a resistor or transparent conductive film, is when the carrier concentration value is 1 × 10⁻⁶ 20 cm −3 This range includes, for example, 1 × 10 19 cm −3 The above is 1 x 10 22 cm −3 The range is as follows: By making the carrier concentration sufficiently high, the resistivity can be increased to 1 × 10⁻⁶. −4 It is expected that the level can be reduced to below Ω·cm.

[0275] Indium oxide may contain elements that increase the carrier concentration. For example, it is preferable to include elements common to the source and drain electrodes of the transistor. Examples of elements that increase the carrier concentration include titanium, zirconium, hafnium, tantalum, tungsten, molybdenum, tin, silicon, and boron. In particular, it is more preferable to use elements in which the oxide is conductive or semiconducting.

[0276] Because indium oxide is an oxide whose valence electrons can be controlled, the region with a low carrier concentration can be used for the channel formation region of the transistor, and the region with a high carrier concentration can be used for the source and drain regions of the transistor. This makes it possible to create a so-called n-i-n junction (a junction between an n-type region, an i-type region, and an n-type region). Valence electron control in transistors using silicon is generally known. On the other hand, valence electron control in transistors using indium oxide is a novel technological concept that would not normally be conceived. By using this technological concept, it is possible to realize a transistor with high mobility, low off-current, normally-off capability, and high reliability.

[0277] The indium oxide film is preferably crystalline. In particular, the indium oxide film is preferably polycrystalline, and more preferably single-crystal. A single-crystal film does not have grain boundaries. By using a single-crystal film, carrier scattering at grain boundaries can be suppressed, enabling the realization of transistors that exhibit high field-effect mobility. Furthermore, it has the excellent effect of suppressing variations in transistor characteristics caused by these grain boundaries.

[0278] Furthermore, polycrystalline films are preferable because they can reduce carrier scattering and exhibit high field-effect mobility compared to microcrystalline or amorphous films. When using polycrystalline films, it is preferable to use films with the largest possible grain size and few grain boundaries. In a transistor to which a polycrystalline film is applied, if there are no grain boundaries in the channel formation region, or if no grain boundaries are observed, the channel formation region is located within the single-crystal region contained in the polycrystalline film, and therefore it can be considered a transistor to which a single-crystal film is applied.

[0279] The crystallinity of indium oxide can be analyzed, for example, by X-ray diffraction (XRD), transmission electron microscopy (TEM), or electron diffraction (ED). Alternatively, a combination of these methods may be used for analysis.

[0280] Furthermore, in this specification, a semiconductor layer in which no grain boundaries are observed in the channel formation region, a semiconductor layer in which the channel formation region is contained within a single crystal grain, or a semiconductor layer in which the crystal axis directions are the same in at least two regions within the channel formation region can be considered as a single crystal film.

[0281] The channel formation region refers to the region of the semiconductor layer that overlaps with (or faces) the gate electrode via the gate insulating layer, and is located between the region in contact with the source electrode and the region in contact with the drain electrode. The crystal grains, grain boundaries, crystal axes, and crystal orientation in the channel formation region can be confirmed by cross-sectional observation including the semiconductor layer, source electrode, and drain electrode.

[0282] Impurities in the indium oxide film can act as a source of carrier scattering, thus potentially causing a decrease in field-effect mobility and inhibiting crystal growth. Examples of impurities in the indium oxide film include boron and silicon. In the channel-forming region of the indium oxide film, lower concentrations of these impurities are preferable. For example, the concentration of each of the above impurity elements should be 0.1% or less, more preferably 0.01% (100 ppm) or less. Note that elements such as carbon and hydrogen may be present in the deposition gas or precursor during film formation, and may remain in the indium oxide film in higher concentrations than the above impurities.

[0283] Furthermore, indium oxide films can also contain elements that can become trivalent cations like indium, as long as their crystals maintain a cubic crystal structure (Bixbite type). Examples include group 13 elements of the periodic table such as gallium and aluminum, and group 3 elements of the periodic table. Since these elements mainly exist as trivalent cations in oxides, the carrier concentration of indium oxide can be kept low.

[0284] By using such an indium oxide film in a transistor, the field-effect mobility of the transistor can be increased to 50 cm². 2 / (V·s) or more, preferably 100 cm 2 / (V·s) or more, more preferably 150 cm 2 / (V·s) or more, more preferably 200 cm 2 / (V·s) or more, more preferably 250 cm 2 It can be set to (V・s) or more.

[0285] One of the characteristics of indium oxide films is their higher oxygen permeability (diffusivity) compared to IGZO films. For example, oxygen diffusing into an indium oxide film permeates the film and is released as oxygen molecules. In some cases, it may also be released as water molecules by reacting with hydrogen contained in the film. Furthermore, if there is an oxygen deficiency in the film, diffusing oxygen atoms will fill the deficiency. Because oxygen diffuses easily through indium oxide films, it can be said that oxygen deficiencies are more easily filled in compared to IGZO films.

[0286] Thus, because indium oxide films are more likely to reduce oxygen vacancies in the film compared to IGZO films, applying such indium oxide films to transistors makes it possible to realize transistors with extremely high reliability.

[0287] Furthermore, the indium oxide film diffuses hydrogen. Hydrogen diffusing into the indium oxide film from the outside permeates the film and is released as hydrogen molecules. Alternatively, it reacts with oxygen contained in the film and is released as water molecules.

[0288] Indium oxide is characterized by a small effective electron mass and a large effective hole mass. Furthermore, the effective electron mass of indium oxide is largely independent of the crystal orientation. Therefore, using crystalline indium oxide in transistors allows for the realization of transistors with high field-effect mobility and high frequency characteristics (also known as f-response). Moreover, due to the large effective hole mass, transistors with extremely low off-currents can be realized. For example, by applying an indium oxide film to a transistor, the off-current per 1 μm of channel width is 1 fA (1 × 10⁻¹⁶) at 125°C. −15 A) Less than or equal to, or 1aA (1 × 10 −18 A) Less than or equal to 1aA (1 × 10) in a room temperature (25°C) environment. −18 A) Less than or equal to, or 1zA (1 × 10⁻¹⁰−21 A) The following is possible. Furthermore, because indium oxide has a smaller effective electron mass and a larger effective hole mass than silicon, it may be possible to realize transistors with higher field-effect mobility and lower off-current than Si transistors.

[0289] It is preferable to provide a seed layer so as to be in contact with at least a portion of the crystalline indium oxide film. It is preferable to use a material containing crystals with a small difference in lattice constant (also called lattice mismatch) with the indium oxide for the seed layer. This improves the crystallinity of the indium oxide film. A substrate (e.g., a single-crystal substrate) may be used as one of the layers in contact with at least a portion of the crystalline indium oxide film.

[0290] One method for evaluating the degree of lattice mismatch is to use the following lattice mismatch value. The lattice mismatch Δa [%] of the crystals in the formed film (in this case, the indium oxide film) relative to the crystals in the seed layer is given by Δa = ((L 1 -L 2 ) / L 2 It is calculated as ) × 100. Here L 1 L is the length of the unit cell vector of the crystals in the formed film, or the lattice constant. 2 This is the length of the unit cell vector of the crystal in the seed layer, or the lattice constant.

[0291] The lattice mismatch Δa between the seed layer and the indium oxide film is preferably small in absolute value, and most preferably zero. For example, Δa can be -5% or more and 5% or less, preferably -4% or more and 4% or less, more preferably -3% or more and 3% or less, and even more preferably -2% or more and 2% or less.

[0292] Here, the indium oxide crystal has a cubic structure (bixbite type). For example, yttria-stabilized zirconia (YSZ) crystals can have a cubic structure (fluorite type). The lattice mismatch of the indium oxide crystal with respect to the cubic YSZ crystal is in the range of -2% to 2%, and a single crystal film of indium oxide can be epitaxially grown on a YSZ substrate.

[0293] Furthermore, the crystal structure of the seed layer and the crystal structure of the indium oxide film do not necessarily have to be the same in terms of crystal system or crystal orientation. For example, a film with a hexagonal or trigonal crystal structure can be used beneath an indium oxide film with a cubic crystal structure. For example, by setting the crystal orientation of the surface of the seed layer to

[001] and the crystal orientation of the underside of the indium oxide film to

[111] , the requirements related to crystal orientation necessary for epitaxial growth can be met. Examples of hexagonal or trigonal crystals include wurtzite-type structures and YbFe. 2 O 4 Type structure, Yb 2 Fe 3 O 7 These include type structures and their modified type structures. YbFe 2 O 4 Type structure or Yb 2 Fe 3 O 7 An example of a crystal having a type structure is IGZO.

[0294] This embodiment can be implemented in appropriate combination with other embodiments described herein.

[0295] (Embodiment 5) This embodiment describes an electronic component that can use the semiconductor device described in the above embodiment. An electronic component using a semiconductor device according to one aspect of the present invention is effective in improving performance, such as reducing power consumption.

[0296] [Electronic Components] A perspective view of electronic component 1700A is shown in Figure 17A. The electronic component 1700A shown in Figure 17A comprises a substrate 1701, a semiconductor device 1710 on the substrate 1701, and a mold 1711. In particular, the semiconductor device 1710 is sealed by the mold 1711. Note that in Figure 17A, some details have been omitted in order to show the inside of the electronic component 1700A.

[0297] For example, the substrate 1701 can be a ceramic substrate, a plastic substrate, or a glass epoxy substrate.

[0298] The electronic component 1700A is provided with, for example, a lead frame 1712. A portion of the lead frame 1712 located on the substrate 1701 is covered by a mold 1711, while another portion of the lead frame 1712 is exposed to the outside of the mold 1711. In particular, the lead frame 1712 exposed to the outside of the mold 1711 functions, for example, as a terminal for mounting the electronic component 1700A onto a printed circuit board.

[0299] Within the mold 1711, electrode pads 1713 are provided on the lead frame 1712, and the electrode pads 1713 are connected to the semiconductor device 1710 via wires 1714. The electronic component 1700A is mounted on the printed circuit board, for example, by bringing the lead frame 1712 into contact with the wiring on the printed circuit board side. In this way, multiple electronic components are combined and connected on the printed circuit board to complete the mounted circuit board.

[0300] Next, the semiconductor device 1710 will be described. For example, as shown in Figure 17B, the semiconductor device 1710 has a drive circuit layer 1715 and a storage layer 1716. The storage layer 1716 can be configured by stacking multiple cell arrays. The cell array can include the arithmetic cells, drive cells, and storage cells described in the above embodiment. The configuration in which the drive circuit layer 1715 and the storage layer 1716 are stacked can be a monolithic stack configuration. In a monolithic stack configuration, the layers can be connected without using through-electrode technology (for example, TSV (Through Silicon Via)) and bonding technology such as Cu-Cu direct bonding. By configuring the drive circuit layer 1715 and the storage layer 1716 in a monolithic stack configuration, for example, a so-called on-chip memory configuration can be achieved in which memory is directly formed on the processor. By using an on-chip memory configuration, it is possible to speed up the operation of the interface portion between the processor and the memory. For example, by using the arithmetic unit described in the above embodiment as the processor, the transmission speed of data (e.g., weight coefficients) from memory to the arithmetic unit can be increased.

[0301] Furthermore, by using an on-chip memory configuration, it is possible to reduce the size of connection wiring and other components compared to technologies that use through-hole electrodes such as TSVs, thus increasing the number of connection pins. Increasing the number of connection pins enables parallel operation, which in turn improves the memory bandwidth (also called memory bandwidth).

[0302] Furthermore, it is preferable to form the multiple memory cell arrays of the memory layer 1716 using OS transistors and to stack these multiple memory cell arrays monolithically. By configuring the multiple memory cell arrays in a monolithic stack, it is possible to improve either or both of the memory bandwidth and / or memory access latency. Bandwidth is the amount of data transferred per unit time, and access latency is the time from access to the start of data exchange. In the case of a configuration using Si transistors in the memory layer 1716, it is difficult to create a monolithic stack configuration compared to OS transistors. Therefore, in a monolithic stack configuration, OS transistors can be said to have a superior structure compared to Si transistors.

[0303] Furthermore, the semiconductor device 1710 may also be referred to as a die. In this specification, a die refers to a chip piece obtained in the semiconductor chip manufacturing process by forming a circuit pattern on, for example, a disc-shaped substrate (also called a wafer) and cutting it into cubes. Examples of semiconductor materials that can be used for dies include silicon (Si), silicon carbide (SiC), and gallium nitride (GaN). For example, a die obtained from a silicon substrate (also called a silicon wafer) is sometimes called a silicon die.

[0304] Next, Figure 17C shows electronic component 1700B, which is a modified version of electronic component 1700A. Unlike electronic component 1700A, electronic component 1700B shown in Figure 17C does not use a lead frame 1712, and instead has electrodes 1733 provided at the bottom of the substrate 1701. The electrodes 1733 function as connection terminals for mounting electronic component 1700B onto a printed circuit board.

[0305] Figure 17C shows an example in which the electrode 1733 is formed with solder balls. By arranging solder balls in a matrix at the bottom of the substrate 1701, BGA (Ball Grid Array) mounting can be realized. For this purpose, the substrate 1701 is provided with through-hole vias, and a conductive layer 1732 that functions as wiring is provided on these vias. On the substrate 1701, the electrode pad 1713 is provided in contact with the conductive layer 1732 above it, and on the substrate 1701, the electrode 1733 is provided in contact with the conductive layer 1732 below it.

[0306] Furthermore, the electrodes 1733 can be formed with conductive pins instead of solder balls. By arranging conductive pins in a matrix at the bottom of the substrate 1701, PGA (Pin Grid Array) mounting can be realized.

[0307] Furthermore, the electronic component 1700B can be mounted on other boards using various mounting methods, not limited to BGA and PGA. Examples of mounting methods include SPGA (Staggered Pin Grid Array), LGA (Land Grid Array), QFP (Quad Flat Package), QFJ (Quad Flat J-leaded package), and QFN (Quad Flat Non-leaded package).

[0308] Furthermore, an electronic component according to one aspect of the present invention can also be in the form of SiP (System in Package) or MCM (Multi-Chip Module). For example, the electronic component 1700C shown in Figure 17D has an interposer 1731 provided on a package substrate 1734 (printed circuit board), and a semiconductor device 1735 and a plurality of semiconductor devices 1710 are provided on the interposer 1731.

[0309] In Figure 17D, the electronic component 1700C shows, as an example, an example in which the semiconductor device 1710 is used as a high-bandwidth memory (HBM). For example, the semiconductor device 1735 can be used as an arithmetic circuit in an integrated circuit such as a CPU, GPU, or FPGA (Field Programmable Gate Array). Furthermore, a semiconductor device according to one aspect of the present invention can be used as the semiconductor device 1735.

[0310] The package substrate 1734, like the substrate 1701, can be made of, for example, a ceramic substrate, a plastic substrate, a glass epoxy substrate, etc. The interposer 1731 can be made of, for example, a silicon interposer or a resin interposer.

[0311] The interposer 1731 has multiple wirings and functions to connect multiple integrated circuits with different terminal pitches. The multiple wirings are provided in a single layer or multiple layers. The interposer 1731 also has the function of connecting integrated circuits provided on the interposer 1731 to electrodes provided on the package substrate 1734. For these reasons, the interposer is sometimes called a "redistribution substrate" or "intermediate substrate". In addition, through electrodes may be provided on the interposer 1731, and these through electrodes may be used to connect the integrated circuits and the package substrate 1734. Furthermore, in silicon interposers, TSVs can also be used as through electrodes.

[0312] In HBMs, many connections are necessary to achieve a wide memory bandwidth. Therefore, the interposer on which the HBM is mounted requires fine and high-density wiring. For this reason, it is preferable to use a silicon interposer for mounting the HBM.

[0313] Furthermore, in SiP and MCM using silicon interposers, reliability degradation due to differences in expansion coefficients between the integrated circuit and the interposer is less likely to occur. In addition, because silicon interposers have high surface flatness, connection failures between the integrated circuit placed on the silicon interposer and the silicon interposer are less likely to occur. In particular, in 2.5D packages (2.5-dimensional packaging) where multiple integrated circuits are arranged side by side on the interposer, it is preferable to use a silicon interposer.

[0314] On the other hand, when connecting multiple integrated circuits with different terminal pitches using silicon interposers and TSVs, space is required, such as the width of the terminal pitch. Therefore, when trying to reduce the size of the electronic component 1700C, the width of the terminal pitch becomes a problem, and it may become difficult to provide the many wires necessary to achieve a wide memory bandwidth. For this reason, as mentioned above, a monolithic stacked configuration using OS transistors is preferable. Also, for example, a memory cell array stacked using TSVs and a monolithic stacked memory cell array can be combined. Furthermore, a structure that combines a memory cell array stacked using TSVs and a monolithic stacked memory cell array is sometimes called a composite structure.

[0315] Furthermore, if the temperature of the electronic component 1700C rises due to current-induced heat or other factors, the characteristics of the circuit elements (such as transistors) provided in the electronic component 1700C may deteriorate. Therefore, it is preferable to provide a heat sink (heat dissipation plate) on top of the electronic component 1700C. When a heat sink is provided, it is preferable to align the heights of the integrated circuits provided on the interposer 1731. For example, in the electronic component 1700C shown in this embodiment, it is preferable to align the heights of the semiconductor device 1710 and the semiconductor device 1735.

[0316] This embodiment can be implemented in appropriate combination with other embodiments described herein.

[0317] (Embodiment 6) This embodiment describes an electronic device using the electronic components described in the above embodiment, and an information processing system using the electronic device.

[0318] Figure 18 shows an example of the configuration of an information processing system. The information processing system 8000 shown in Figure 18 includes an example of various electronic devices and a server located within the network.

[0319] Figure 18 shows, as examples of such electronic devices, a portable information terminal 8200, a wearable information terminal 8300, a notebook personal computer 8400, an automobile 8500, an industrial robot 8600, and a camera 8700. Figure 18 also shows a network 8100 and a large computer 8110 located within the network 8100.

[0320] The term "large-scale computer 8110" can sometimes refer to multiple computers installed in a server room or similar location. For example, a rack-mount type large-scale computer 8110 is one in which multiple computers are housed in a rack. The large-scale computer 8110 is sometimes referred to as a supercomputer. Furthermore, in the information processing system 8000, the large-scale computer 8110 may also be referred to as a server or cloud server.

[0321] Each of the multiple computers in the large computer 8110 has a motherboard, which is provided with multiple slots, multiple connection terminals, etc. One or more PC cards can be inserted into the slots, for example. A PC card is an example of a processing board equipped with processing units such as a CPU and a GPU. For example, one or more of the electronic components 1700A, 1700B, and 1700C can be used as the processing units.

[0322] The mainframe computer 8110 can also function as a parallel computer. By using the mainframe computer 8110 as a parallel computer, it is possible to perform large-scale calculations necessary for, for example, the training and inference of artificial intelligence.

[0323] When performing wired communication as Network 8100, specifications standardized by IEEE, such as Ethernet (registered trademark), can be used. Furthermore, types of communication include electrical communication using wires such as twisted-pair cables, and optical communication using optical fibers.

[0324] On the other hand, when performing wireless communication as network 8100, communication protocols or communication technologies that can be used include communication standards such as the fourth-generation mobile communication system (4G), the fifth-generation mobile communication system (5G), the sixth-generation mobile communication system (6G), or specifications standardized by IEEE such as Wi-Fi® and Bluetooth®.

[0325] Network 8100 can be, for example, a PAN (Personal Area Network), a LAN (Local Area Network), a CAN (Campus Area Network), a MAN (Metropolitan Area Network), a WAN (Wide Area Network), or a GAN (Global Area Network). For example, by using a GAN in network 8100, it is possible to use the Internet, which is the foundation of the World Wide Web (WWW).

[0326] Furthermore, if the information processing system 8000 is built on a LAN as a network 8100, the possibility of confidential information leakage can be reduced compared to using the Internet.

[0327] Furthermore, a company or individual managing the mainframe computer 8110 can, for example, use the network 8100 to provide services using the information processing system 8000 to users of each electronic device. One example of such services is a usage model called cloud computing. Through this cloud computing, users of the aforementioned electronic devices can utilize the mainframe computer 8110's functions for storing large amounts of data, performing large-scale calculations, and other applications.

[0328] In particular, by using a semiconductor device according to one aspect of the present invention in the aforementioned electronic devices and the large computer 8110, large-scale calculations such as artificial neural networks can be performed with low power consumption. Furthermore, this enables the information processing system 8000 to provide services to users in a usage form known as cloud AI or edge AI.

[0329] Cloud AI generally refers to a service where a large-scale computer (8110) performs the training and inference of an artificial neural network. The large-scale computer 8110 is pre-trained on collected data, and each electronic device transmits input data to the large-scale computer 8110 for the artificial neural network, where the large-scale computer 8110 performs inference on that input data. The large-scale computer 8110 also transmits the results of this inference to each electronic device, which can then use. Because training and inference are performed by the large-scale computer 8110, cloud AI is suitable for processing large amounts of data and for handling complex calculations.

[0330] On the other hand, edge AI generally refers to a service where each electronic device performs the learning and inference of an artificial neural network. In this case, the mainframe computer 8110 provides each electronic device with the artificial neural network model, weight coefficients (sometimes called weight data, connection coefficients, etc.), etc. The results of the learning and inference performed on each electronic device are also transmitted to the mainframe computer 8110. Furthermore, a usage model in which the mainframe computer 8110 learns the artificial neural network and each electronic device performs inference using the learned neural network is also sometimes referred to as edge AI.

[0331] Edge AI performs artificial neural network inference on each individual electronic device, thus reducing the communication time required compared to cloud AI. In other words, edge AI is well-suited for real-time analysis of input data. Furthermore, the amount of data transmitted between each electronic device and the mainframe computer 8110 is reduced, lowering data communication costs and power consumption. The reduced data transmission also minimizes security risks such as information leaks. For these reasons, edge AI is suitable for building small-scale systems, for example.

[0332] Furthermore, a semiconductor device according to one aspect of the present invention is suitable for edge AI because it consumes little power. A specific example of an edge AI system is described below.

[0333] [Personal Information Terminal] The personal information terminal 8200 shown in Figure 18 is an electronic device that integrates a display device and a touch panel. The personal information terminal 8200 can also be equipped with electronic component 8201 as one of the electronic components 1700 (electronic component 1700A, electronic component 1700B, or electronic component 1700C) mentioned above, thereby enabling the personal information terminal 8200 to perform large-scale calculations such as artificial neural networks. The personal information terminal 8200 can also be equipped with a camera.

[0334] By equipping the personal digital assistant (PDA) 8200 with a camera, image recognition using edge AI can be performed on images captured by the PDA 8200. The objects that can be recognized include humans, animals, plants, characters, and pictograms. In particular, by performing image recognition on images of human faces, fingerprints, palm prints, irises, and veins, it can be used for biometric authentication.

[0335] [Wearable Information Terminal] The wearable information terminal 8300 shown in Figure 18 is an electronic device that can be worn on a person's head. The wearable information terminal 8300 in Figure 18 has an eye cover, a display device, temples (arms) that hook onto the ears, and earphones, but other examples include HMDs (head-mounted displays) and glasses-type XR devices. The wearable information terminal 8300 can also be equipped with a camera, similar to the portable information terminal 8200.

[0336] Furthermore, the wearable information terminal 8300 can be equipped with the electronic component 8301 as the aforementioned electronic component 1700, thereby enabling large-scale computations such as artificial neural networks to be performed in the wearable information terminal 8300.

[0337] By equipping the wearable information terminal 8300 with a camera, images captured by the wearable information terminal 8300 can be displayed on a display device in real time. Furthermore, by performing image recognition using edge AI, information about objects included in the image displayed on the display device can be added to the display device. In addition, by performing image recognition on moving objects such as pedestrians, bicycles, cars, and trains displayed on the display device, it is possible to perform risk prediction to determine whether or not there is a risk of collision.

[0338] [Notebook Personal Computer] The notebook personal computer 8400 shown in Figure 18 is an electronic device primarily used on a desktop. The notebook personal computer 8400 can also be equipped with electronic component 8401 as the electronic component 1700 mentioned above, thereby enabling the notebook personal computer 8400 to perform large-scale calculations such as artificial neural networks.

[0339] The 8400 notebook personal computer can, for example, use edge AI as part of its computational processing when using applications. Examples of its applications include upconversion, which increases the screen resolution of images (including still images and videos) displayed on a display device in real time; translation, which converts text into another language; and editing tasks for text or images.

[0340] [Automobile] The automobile 8500 shown in Figure 18 is an example of a mobile device. The automobile 8500 can also be equipped with the electronic component 8501 as the electronic component 1700 described above, thereby enabling the automobile 8500 to be used as an electronic device for edge AI.

[0341] Edge AI in the Automobile 8500 can be used for applications such as autonomous driving, hazard prediction in autonomous driving, and in-car air conditioning management.

[0342] In this specification, automobiles are used as an example of a mobile device, but other examples of mobile devices include trains, monorails, ships, and aircraft (e.g., helicopters, unmanned aerial vehicles (drones), airplanes, and rockets). The aforementioned mobile devices can also be used as electronic devices for edge AI.

[0343] [Industrial Robot] The industrial robot 8600 shown in Figure 18 can be deployed, for example, in a production plant. The industrial robot 8600 preferably has multiple drive axes to finely control the drive range. The industrial robot 8600 may also have one or more functions such as grasping, cutting, welding, coating, and attaching objects. In addition, the industrial robot 8600 is preferably equipped with sensors such as an image detection module or a camera to detect the object. The industrial robot 8600 is also preferably equipped with a sensor that detects minute currents to determine whether or not it has grasped an object.

[0344] In addition, the industrial robot 8600 can be provided with the electronic component 8601 as the above-mentioned electronic component 1700, whereby the industrial robot 8600 can be used as an edge AI electronic device. The edge AI in the industrial robot 8600 can be used for applications such as performing image recognition on objects, classification by type, classification by size, and inspection for determining non-defective and defective products, for example.

[0345] [Camera] The camera 8700 shown in FIG. 18 can be used for, for example, surveillance cameras, security cameras, pet cameras, and the like. In addition, the housing of the camera 8700 is not limited to the shape for ceiling mounting as shown in FIG. 18, and there are various types such as a desk-mounted type and a wall-mounted type.

[0346] Note that the terms "surveillance camera", "security camera", and "pet camera" are conventional names, and do not limit the application according to the names. For example, a pet camera may be used for the purpose of a surveillance camera or a security camera, and vice versa. In addition, the camera 8700 may also be referred to as a video camera in some cases.

[0347] In addition, the camera 8700 can be provided with the electronic component 8701 as the above-mentioned electronic component 1700, whereby the camera 8700 can be used as an edge AI electronic device. The edge AI in the camera 8700 can be used, for example, for security applications, for moving object detection on objects displayed in images (still images and moving images) captured by the camera 8700. It can also be used for disaster prevention applications for detection of river flooding, tsunamis, and the like.

[0348] This embodiment can be implemented in appropriate combination with other embodiments described in this specification.

[0349] (Embodiment 7) In this embodiment, space equipment and a data center (also referred to as Data Center: DC) that can use the semiconductor device according to one aspect of the present invention will be described. The semiconductor device according to one aspect of the present invention is effective in power saving and is suitable for space equipment and data centers.

[0350] [Space Equipment] A semiconductor device according to one aspect of the present invention is suitable for space equipment (for example, equipment having the function of processing and storing information).

[0351] A semiconductor device according to one aspect of the present invention preferably includes an OS transistor. The OS transistor exhibits minimal fluctuations in electrical characteristics due to radiation exposure. In other words, it has high resistance to radiation and is therefore suitable for environments where radiation may be incident. For example, an OS transistor is suitable for use in outer space.

[0352] Figure 19 shows an example of space equipment, specifically a satellite 6800. The satellite 6800 comprises a body 6801, a solar panel 6802, an antenna 6803, a secondary battery 6805, and a control device 6807. In Figure 19, a planet 6804 is shown as an example in outer space. Outer space refers to, for example, an altitude of 100 km or more, but as described herein, outer space includes the thermosphere, mesosphere, and stratosphere.

[0353] Furthermore, although not shown in Figure 19, a battery management system (also known as a BMS) or battery control circuit can be provided in the secondary battery 6805. Using an OS transistor in the above-mentioned battery management system or battery control circuit is preferable because it consumes little power and has high reliability even in outer space.

[0354] Furthermore, outer space is an environment with radiation levels more than 100 times higher than those on Earth. Examples of radiation include electromagnetic waves (electromagnetic radiation) such as X-rays and gamma rays, as well as particle radiation such as alpha rays, beta rays, neutron rays, proton rays, heavy ion rays, and meson rays.

[0355] When sunlight shines on the solar panel 6802, the power necessary for the satellite 6800 to operate is generated. However, if, for example, the solar panel is not exposed to sunlight, or if the amount of sunlight hitting the solar panel is low, the amount of power generated will decrease. Therefore, there is a possibility that the power necessary for the satellite 6800 to operate may not be generated. To operate the satellite 6800 even under conditions of low power generation, it is advisable to equip the satellite 6800 with a secondary battery 6805. Note that solar panels are sometimes called solar cell modules.

[0356] The satellite 6800 can generate a signal. This signal is transmitted via antenna 6803, and can be received, for example, by a receiver on the ground or another satellite. By receiving the signal transmitted by satellite 6800, the position of the receiver that received the signal can be measured. Thus, satellite 6800 can constitute a satellite positioning system.

[0357] Furthermore, the control device 6807 has the function of controlling the artificial satellite 6800. The control device 6807 is composed of one or more selected from, for example, a CPU, a GPU, and a memory circuit. In addition, a semiconductor device according to one aspect of the present invention can be used for the control device 6807. Compared to Si transistors, OS transistors exhibit less fluctuation in electrical characteristics due to radiation irradiation. In other words, they operate stably and are highly reliable even in environments where radiation may be incident.

[0358] Furthermore, the satellite 6800 can be configured to include sensors. For example, by configuring it to include a visible light sensor, the satellite 6800 can have the function of detecting sunlight reflected from an object on the ground. Alternatively, by configuring it to include a thermal infrared sensor, the satellite 6800 can have the function of detecting thermal infrared radiation emitted from the Earth's surface. Thus, the satellite 6800 can function, for example, as an Earth observation satellite.

[0359] In this embodiment, an artificial satellite is given as an example of space equipment, but the invention is not limited to this. For example, a semiconductor device according to one aspect of the present invention can also be used in space equipment such as spacecraft, space capsules, and space probes.

[0360] This embodiment can be implemented in appropriate combination with other embodiments described herein.

[0361] In inference using artificial neural networks, if a malfunction such as damage occurs in a part of the multiply-accumulate unit that constitutes the artificial neural network, the inference results may be extremely biased. In this embodiment, we will explain the inference results when a malfunction occurs in a part of the multiply-accumulate unit in the inference of handwritten digits from 0 to 9, and the inference results when the malfunction is resolved.

[0362] Figure 20A is a block diagram showing the configuration of the artificial neural network NN used for inference. The artificial neural network NN is a fully connected artificial neural network consisting of an input layer IL, a hidden layer HL, and an output layer OL. The hidden layer HL and the output layer OL each function as a multiply-accumulate unit.

[0363] The input image size was set to 23 pixels vertically and 22 pixels horizontally. Therefore, the number of neurons NE constituting the input layer (IL) was set to 506 (23 x 22) to match the size of the input image. The number of neurons NE constituting the hidden layer (HL) was set to 256. In addition, to perform inference from 0 to 9, the number of neurons NE constituting the output layer (OL) was set to 10.

[0364] Each neuron NE in the input layer IL receives data corresponding to each dot of the input image. Each neuron NE in the hidden layer HL receives the output of the neuron NE in the input layer IL as the input signal Xin. Each neuron NE in the output layer OL receives the output signal Yout from the hidden layer HL as the input signal.

[0365] As an example, Figure 20B shows a block diagram of a Multiply Accumulate Circuit (MAC) applicable to the intermediate layer HL and the output layer OL, respectively. The Multiply Accumulate Circuit (MAC) has multiple Multiplication Circuits (MC) arranged in a matrix. Figure 20B illustrates a case where the Multiply Accumulate Circuit (MAC) has Multiplication Circuits MC arranged in a 5x5 matrix.

[0366] Furthermore, the multiply-accumulate circuit MAC has multiple wires 91 extending in the row direction, wires 92 extending in the column direction, and wires 93 extending in the column direction. In Figure 20B, five wires each of wires 91, 92, and 93 are shown.

[0367] The input signal Xin[1] is supplied to the first row of wiring 91, and the input signal Xin[2] is supplied to the second row of wiring 91. Similarly, the input signals Xin[3] to Xin[5] are supplied to each of the third to fifth rows of wiring 91, respectively. In addition, the weight coefficient W[1] is supplied to the first column of wiring 92, and the weight coefficient W[2] is supplied to the second column of wiring 92. Similarly, the weight coefficients W[3] to W[5] are supplied to each of the third to fifth columns of wiring 92, respectively. In addition, the output signal Yout[1] is supplied to the first column of wiring 93, and the output signal Yout[2] is supplied to the second column of wiring 93. Similarly, the output signals Yout[3] to Yout[5] are supplied to each of the third to fifth columns of wiring 93, respectively.

[0368] Each multiplier circuit MC located in a row is connected to a wiring 91 for that row. That is, the multiplier circuit MC located in the first row is connected to the wiring 91 of the first row. Each multiplier circuit MC located in a column is connected to wirings 92 and 93 for that column. That is, the multiplier circuit MC located in the first column is connected to wirings 92 and 93 of the first column.

[0369] The multiplier circuit MC outputs the product of the input signal Xin supplied via wiring 91 and the weight coefficient W to wiring 93. For example, the sum of the calculation results of all multiplier circuits MC located in the first column is output as the output signal Yout[1] from wiring 93 in the first column. A bias value can be added to the output signal Yout as needed. In this way, a set of multiple multiplier circuits MC in each column functions as a single neuron NE. Figure 20B shows an example in which the multiply-accumulate circuit MAC has five neurons NE. However, the number of neurons NE that the multiply-accumulate circuit MAC has is not limited to this.

[0370] Prior to inference, a learning model with an accuracy of 97.9% was generated using the MNIST database. Using this learning model, inference was performed in the case where a defect occurred in a part of the sum-of-accumulate unit, and then in the case where the defect was resolved.

[0371] Inference in the event of a malfunction in part of the multiply-accumulate unit was performed assuming an artificial neural network NN in which a multiplication circuit MCe that outputs an abnormal value is included in some of the neurons NE that make up the hidden layer HL. As an example, Figure 21A shows a block diagram of the multiply-accumulate unit MAC including the multiplication circuit MCe. In Figure 21A, four neurons NE (neuron NE[1] to neuron NE[4]) are shown out of the 256 neurons NE that make up the hidden layer HL. In Figure 21A, neuron NE[3] includes a multiplication circuit MCe that outputs an abnormal value.

[0372] To resolve the above-mentioned problem, the neuron NE containing the faulty multiplication circuit MCe is not used, and instead, a neuron NE without the faulty component is used as neuron NE[3]. Figure 21B shows an example where a neuron NE without the multiplication circuit MCe is used as neuron NE[3]. After implementing the above countermeasures and resolving the problem, inference was performed.

[0373] As mentioned above, the input image used for inference is 23 pixels high and 22 pixels wide. Furthermore, this input image is a monochrome grayscale image and contains noise components in 40% of the entire image.

[0374] Figure 22 shows the input image, the inference result before the countermeasure, and the inference result after the countermeasure. In the inference using the artificial neural network (NN) before the countermeasure, which included the faulty area, the inference result for all input images was "4," indicating a significant decrease in inference accuracy. In the inference using the artificial neural network (NN) after the countermeasure, which did not include the faulty area, correct results were obtained for all input images.

[0375] It was confirmed that if a malfunction occurs in a neuron (NE) that makes up an artificial neural network (NN), the decrease in inference accuracy can be suppressed by replacing the malfunctioning neuron (NE) with a neuron (NE) that is not malfunctioning.

[0376] 10: Element layer, 15: Wiring, 20: Element layer, 25: Wiring, 100: Semiconductor device, 101: First circuit, 102: Second circuit, 200: Transistor, 200A: Transistor, 200B: Transistor, 200C: Transistor, 201: Insulating layer, 202: Insulating layer, 205: Conductive layer, 230: Semiconductor layer, 240: Conductive layer, 240a: Conductive layer, 240b: Conductive layer, 241: Insulating layer, 241a: Insulating layer, 241b: Insulating layer, 242: Conductive layer, 242a: Conductive layer, 242b: Conductive layer, 250: Insulating layer, 255: Conductive layer, 256: Insulating layer, 257: Insulating layer, 258: 259: Insulating layer, 260: Conductive layer, 260a: Conductive layer, 260b: Conductive layer, 261: Conductive layer, 262: Opening, 265: Conductive layer, 271: Insulating layer, 271a: Insulating layer, 271b: Insulating layer, 275: Insulating layer, 280: Insulating layer, 282: Insulating layer, 283: Insulating layer, 285: Insulating layer, 1700: Electronic component, 1700A: Electronic component, 1700B: Electronic component, 1700C: Electronic component, 1701: Substrate, 1710: Semiconductor device, 1711: Mold, 1712: Lead frame, 1713: Electrode pad, 1714: Wire, 1715: Drive circuit layer, 1716: Note 1731: Interposer, 1732: Conductive layer, 1733: Electrode, 1734: Package substrate, 1735: Semiconductor device, 6800: Artificial satellite, 6801: Aircraft body, 6802: Solar panel, 6803: Antenna, 6804: Planet, 6805: Secondary battery, 6807: Control device, 8000: Information processing system, 8100: Network, 8110: Large-scale computer, 8200: Portable information terminal, 8201: Electronic component, 8300: Wearable information terminal, 8301: Electronic component, 8400: Notebook personal computer, 8401: Electronic component, 850 0: Automobile, 8501: Electronic component, 8600: Industrial robot, 8601: Electronic component, 8700: Camera, 8701: Electronic component, Ain: Region, IN: Input terminal, IN[1]: Input terminal, IN[2]: Input terminal, IN[3]: Input terminal, IN[4]: Input terminal, IN[n]: Input terminal, INLY: Input layer, KC1: Kernel, L12: Channel length, L14: Channel length, M11: Transistor, M12: Transistor, M13: Transistor, M14: Transistor, M21: Transistor, M21[1]: Transistor, M21[4]: Transistor,M22: Transistor, M22[1]: Transistor, M22[4]: Transistor, M23: Transistor, M24: Transistor, OUT: Output terminal, Pin: Image, RL: Function circuit, T11: Period, T12: Period, T13: Period, VDD: High power supply potential, VSS: Low power supply potential, W12: Channel width, W14: Channel width,

Claims

1. A first circuit having n (where n is a number of 2 or more) input terminals, an output terminal, and first to fourth transistors, and a second circuit having fifth to eighth transistors, wherein the first terminal of the first transistor is electrically connected to the gate of the first transistor, the gate of the third transistor, and the n input terminals, the second terminal of the first transistor is electrically connected to the first terminal of the second transistor, the gate of the second transistor, and the gate of the fourth transistor, the first terminal of the third transistor is electrically connected to the first terminal of the fifth transistor, the gate of the fifth transistor, and the gate of the seventh transistor, the second terminal of the third transistor is electrically connected to the first terminal of the fourth transistor, the second terminal of the fifth transistor is electrically connected to the first terminal of the sixth transistor, the gate of the sixth transistor, and the gate of the eighth transistor, the first terminal of the eighth transistor is electrically connected to the second terminal of the seventh transistor, and the first terminal of the seventh transistor is electrically connected to the output terminal. A semiconductor device having the function of supplying a current to the output terminal that is 1 / n of the total current supplied to the n input terminals.

2. The semiconductor device according to claim 1, wherein the channel width of the seventh transistor is smaller than the channel width of the fifth transistor, and the channel width of the eighth transistor is smaller than the channel width of the sixth transistor.

3. The semiconductor device according to claim 1, wherein the first circuit has the function of supplying a current to the second circuit that is k (where k is a number greater than 1) times the current supplied from the n input terminals, and the second circuit has the function of supplying a current to the output terminal that is 1 / (n × k) times the current supplied from the first circuit.

4. The semiconductor device according to claim 3, wherein the channel width of the third transistor is greater than the channel width of the first transistor, and the channel width of the fourth transistor is greater than the channel width of the second transistor.

5. The semiconductor device according to any one of claims 1 to 4, wherein the first to fourth transistors include an oxide semiconductor in the semiconductor layer on which the channel is formed.

6. The semiconductor device according to claim 5, wherein the oxide semiconductor comprises indium.

7. The semiconductor device according to any one of claims 1 to 4, wherein the fifth to eighth transistors include silicon in the semiconductor layer on which the channel is formed.

8. A semiconductor device according to any one of claims 1 to 4, wherein the first to fourth transistors are n-type transistors and the fifth to eighth transistors are p-type transistors.

9. The semiconductor device according to any one of claims 1 to 4, wherein the first circuit and the second circuit have overlapping regions.

10. A semiconductor device comprising: a first circuit having the function of supplying a second circuit with a current equal to the sum of the currents supplied to each of n (where n is a number of 2 or more) input terminals; and a second circuit having the function of supplying an output terminal with the current supplied from the first circuit divided by n, wherein the first circuit has a transistor containing an oxide semiconductor in the semiconductor layer in which the channel is formed, and the second circuit has a transistor containing silicon in the semiconductor layer in which the channel is formed, and the first circuit and the second circuit have overlapping regions.

11. A semiconductor device comprising: a first circuit having the function of supplying a second circuit with a current k (a number greater than 1) times the sum of the currents supplied to each of n (n is a number greater than or equal to 2) input terminals; and a second circuit having the function of supplying the current supplied from the first circuit to an output terminal at a multiplier of 1 / (n × k), wherein the first circuit has a transistor containing an oxide semiconductor in the semiconductor layer in which the channel is formed, and the second circuit has a transistor containing silicon in the semiconductor layer in which the channel is formed, and the first circuit and the second circuit have overlapping regions.

12. The semiconductor device according to claim 10 or claim 11, wherein the oxide semiconductor comprises indium.

13. A semiconductor device according to claim 10 or claim 11, wherein the transistor in the first circuit is an n-type transistor and the transistor in the second circuit is a p-type transistor.