Artificial visual nervous system based on bimodal nerve device
By designing an artificial visual nervous system based on bimodal neural devices, using physical convolutional kernel arrays and synaptic computing arrays to simulate the biological retina and central nervous system, the problems of low computing efficiency and high energy consumption in the existing technology are solved, and the integration of image perception and computing with low power consumption and high integration are achieved.
Patent Information
- Application Number
- CN202510559698.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing artificial retinal devices can only simulate the functions of photoreceptor cells, making it difficult to realize the convolution operation of optical signal and multimodal biological neurobionic functions, resulting in low computational efficiency and high energy consumption.
Design an artificial visual nervous system based on bimodal neural devices, including physical convolutional kernel arrays and synaptic computing arrays, and use bimodal transistors to realize optical signal perception and synaptic computing, simulating the functions of biological retina and central nervous system.
It realizes the integration of image perception and computing with low power consumption and high integration, has the ability to efficiently convert optical signal-electric signal, supports online adjustment of synaptic weights, and completes advanced calculation and recognition of visual information.
Smart Images

Figure CN120471114A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of semiconductor technology, and in particular to an artificial visual neural system based on a dual-modal neural device. Background Art
[0002] With the rise of artificial intelligence and artificial neural networks, machine vision has been widely used in fields such as autonomous driving, robot navigation, medical image processing, and security monitoring. The core goal of machine vision is to collect visual data of the surrounding environment through cameras or other sensors, analyze and process this data, and complete specific tasks. Convolution operations are often used during the data analysis and processing stages to extract different image features. However, such systems based on traditional hardware often suffer from low information computation efficiency and high computational energy consumption due to the separate sensing, digital-to-analog conversion, computing, and storage modules. Inspired by the biological nervous system, researchers have recently begun to physically realize the biological visual nervous system. The biological visual nervous system is divided into the peripheral visual nervous system and the central visual nervous system.
[0003] The peripheral visual nervous system is primarily responsible for the perception and feature extraction (information preprocessing) of visual signals, and its primary component is the retina. These two primary functions of the retina are performed by the rods and cones in the upper layer, and the bipolar cells in the lower layer. Within the photoreceptor layer, rods are primarily responsible for visual perception in low-light environments, sensing light intensity but not color. Cones are primarily responsible for color vision and detail recognition in bright environments. Cones come in three types, sensitive to red, green, and blue light. Bipolar cells, on the other hand, are located between photoreceptors and ganglion cells. They are responsible for transmitting light signals from photoreceptors to ganglion cells, and to a certain extent, perform convolution-like processing and integration of the signals. The central visual nervous system, located in the brain, integrates and recognizes preprocessed feature information. This function primarily relies on the biological synaptic structure within complex neural networks. Researchers hope to build a new neuromorphic vision system by preparing electronic units with neural simulation behavior, realizing a low-power, highly integrated miniaturized image sensing, storage and computing architecture.
[0004] Currently, teams have developed artificial retinas for light perception and artificial synapses for neural network computation. However, most current artificial retinas can only mimic the photoreceptors found in biological retinas, capable only of generating basic non-volatile photocurrents and integrating low-level signals. They are unable to simultaneously generate positive and negative photocurrents, making convolution operations difficult. Devices capable of performing optical signal convolution operations are limited to processing the light signal, making it difficult to integrate multimodal biological neuromimetic functions into a single component.
[0005] Therefore, designing a new visual convolution system that can simultaneously realize optical signal convolution operation and salient calculation through only a single component is a technical problem that needs to be solved urgently. Summary of the Invention
[0006] The purpose of the present invention is to solve the problem in the prior art that multiple types of semiconductor devices are needed to simulate the biological visual nervous system.
[0007] The present application provides an artificial visual neural system based on a dual-modal neural device, including a dual-modal visual information computing array composed of a physical convolution kernel array and a synaptic computing array. The physical convolution kernel array is used to receive ambient light signals and convert them into electrical signals for output. The synaptic computing array is used to receive the electrical signals and convert them into visual signals. The physical convolution kernel array and the synaptic computing array are both composed of a number of dual-modal transistors, and the dual-modal transistors have light signal perception and synaptic computing capabilities.
[0008] As a further improvement of the present application, the bimodal transistor includes, from bottom to top, a substrate, a dielectric layer, a semiconductor layer, a source and a drain respectively arranged on both sides of the semiconductor layer away from the surface of the dielectric layer, and a light absorption layer arranged on the semiconductor layer away from the surface of the dielectric layer and located between the source and the drain.
[0009] As a further improvement of the present application, the substrate is a P-type silicon substrate.
[0010] As a further improvement of the present application, the dielectric layer material is selected from Al2O3, ZrO x , HfO x Any one of .
[0011] As a further improvement of the present application, the semiconductor layer is made of n-type semiconductor material.
[0012] As a further improvement of the present application, the semiconductor layer material is selected from InO x , IGZO, or IZO.
[0013] As a further improvement of the present application, the light absorption layer material is selected from any one of PbS, perovskite, MXene, and MoFs.
[0014] As a further improvement of the present application, the preparation of the bimodal transistor includes the following steps:
[0015] S1, cleaning the substrate;
[0016] S2, growing a dielectric layer on the upper surface of the substrate;
[0017] S3, growing a semiconductor layer on the surface of the dielectric layer away from the substrate;
[0018] S4, preparing a source electrode and a drain electrode on a surface of the semiconductor layer away from the dielectric layer;
[0019] S5. Growing a light absorbing layer on a surface of the semiconductor layer away from the dielectric layer and between the source electrode and the drain electrode.
[0020] As a further improvement of the present application, the bimodal transistor adopts a solution method to prepare the dielectric layer, the semiconductor layer and the light absorption layer.
[0021] As a further improvement of the present application, a substrate hydrophilic treatment step is further included between step S1 and step S2.
[0022] Beneficial effects of this application:
[0023] 1. The bimodal transistor designed in this application has the characteristics of photoelectric fusion and non-volatile storage, and has the ability to sense optical signals. Through the coordinated design of the light absorption layer (such as PbS quantum dots) and the n-type semiconductor layer (InO x, etc.), it realizes the efficient conversion and gain adjustment of optical signals to electrical signals. x The ion migration effect of the dielectric layer gives the device long-term potentiation / depression characteristics, supporting online adjustment of synaptic weights.
[0024] 2. The physical convolution kernel array constructed in this application can simulate the collaborative visual signal processing behavior of photoreceptor cells and bipolar cells in the retina, and has the detection, processing and recognition functions of visual information. Through different gate voltage adjustment strategies, a variety of dynamic configurations of convolution kernel arrays (such as sharpening, edge detection, blurring, etc.) can be obtained, with flexible image processing capabilities. This application can be used to implement visual information processing and recognition operations, and can be used in future neuromorphic image feature enhancement systems in machine vision.
[0025] 3. This application has excellent dual-modal collaborative computing capabilities. It simulates the receptive field mechanism (excitatory / inhibitory response) of retinal bipolar cells through a physical convolution kernel array to achieve image feature extraction and preprocessing. It combines the synaptic plasticity (LTP / LTD mechanism) of the bionic central nervous system with the synaptic computing array to complete advanced calculations of visual signals, forming a complete optic nerve bionic closed loop with the help of a single device. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a schematic diagram of the visual convolution system in this application;
[0027] Figure 2 is a schematic structural diagram of a bimodal transistor in this application;
[0028] Figure 3This is a schematic diagram of the structure of the biological retina;
[0029] Figure 4 This is a schematic diagram of the photoreceptor differences between different types of bipolar cells in this application;
[0030] Figure 5 is the excitatory photocurrent of the bimodal transistor in this application;
[0031] Figure 6 is the inhibitory photocurrent of the bimodal transistor in this application;
[0032] Figure 7 is the change in the current response intensity of the bimodal transistor in this application;
[0033] Figure 8 It is the equivalent circuit diagram of the physical convolution kernel matrix in this application;
[0034] Figure 9 It is the different gate voltage adjustment strategies of the convolution kernel matrix in this application;
[0035] Figure 10 This is a schematic diagram of the convolution kernel matrix processing image in this application;
[0036] Figure 11 is the ID-VG scanning curve of the bimodal transistor in this application;
[0037] Figure 12 is the non-volatile current curve of the bimodal transistor in this application;
[0038] Figure 13 are the LTP and LTD conductance values of the bimodal transistor in this application. DETAILED DESCRIPTION
[0039] To make the purpose, technical solutions, and advantages of this application more clear, the specific embodiments of the present invention are described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0040] In the following description of this application, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or components referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0041] In the following description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be internal communication between two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0042] like Figure 1 As shown, the present application provides a visual convolution system based on a dual-modal neural device. A dual-modal visual information calculation matrix is constructed using a dual-modal neural device (dual-modal transistor) with photoelectric response capability, and the calculation matrix includes a physical convolution kernel array and a synaptic calculation array. The physical convolution kernel array can process the input ambient light signal, extract image feature information and ultimately output an electrical signal. The synaptic calculation array can receive electrical signals and ultimately convert them into visual electrical signals for output. The visual convolution system designed in this application can simulate the peripheral visual nervous system (visual information perception layer) and the central visual nervous system (visual information calculation layer) in the biological visual nervous system.
[0043] like Figure 2 As shown in FIG, the bimodal transistor includes, from bottom to top, a substrate (gate), a dielectric layer, a semiconductor layer, a source electrode and a drain electrode respectively arranged on both sides of the semiconductor layer away from the surface of the dielectric layer, and a light absorption layer arranged on the semiconductor layer away from the surface of the dielectric layer and located between the source electrode and the drain electrode. The substrate is a heavily doped Si (p++) substrate; the dielectric layer material is a metal oxide semiconductor material with a large bandgap width, such as Al2O3, ZrO x , HfO x The semiconductor layer is selected from n-type metal oxide semiconductor or n-type van der Waals semiconductor materials, such as InO xThe light absorbing layer material may be a photosensitive material having effective absorption in the visible light band, such as any one of PbS, perovskite, MXene, and metal organic frameworks (MOFs).
[0044] This application also provides a method for preparing a bimodal transistor. Taking a substrate (gate) of highly doped Si (p++), a dielectric layer of AlOx, a semiconductor layer of InOx, a light absorbing layer of Pbs, and a source / drain of aluminum as an example, the specific preparation method is as follows:
[0045] S1: Substrate cleaning: Ultrasonic cleaning of heavily doped Si(p++) substrates with deionized water for 15 minutes, followed by drying under N2 flow. Hydrophilic treatment: Plasma treatment of the cleaned substrate for 15 minutes to hydrophilize the membrane surface.
[0046] S2: Prepare the dielectric layer and spin coat the Al2O3-Li+ precursor solution on the substrate at 3000-5000 rpm to prepare AlO x The film is then annealed in air or nitrogen at 300-450°C for 30-60 minutes. During the specific preparation process, the appropriate spin-coating speed and time are selected based on the preset dielectric layer thickness. The dielectric layer annealing process should be performed at a gradually increasing temperature to avoid cracking of the dielectric film.
[0047] S3: Prepare semiconductor layer, wait for AlO x After the film is cooled, the InO x The precursor solution is spin-coated onto the dielectric layer at 2000-4000 rpm for 20-40 seconds, followed by annealing at 300-500°C for 40-100 minutes. During the fabrication process, the appropriate spin-coating speed and time are selected based on the desired semiconductor layer thickness. The annealing temperature should be gradually increased to prevent cracking of the semiconductor film.
[0048] S4: Prepare source and drain electrodes, and use thermal evaporation to prepare 30nm thick aluminum metal source / drain electrodes.
[0049] S5: preparing a light absorption layer, spin-coating the prepared PbS quantum dot dispersion on the semiconductor layer at 1000-2000 rpm for 20-40 seconds, and then annealing at 100-300° C. for 30 minutes to obtain a bimodal transistor.
[0050] This application applies the bionic principle of the peripheral optic nerve:
[0051] The peripheral visual nervous system is mainly responsible for the perception and feature extraction of visual signals (information preprocessing), and its main component is the retina. Figure 3As shown in the figure, the retina is a key component of the peripheral visual nervous system, responsible for sensing and processing external light signals, converting them into neural signals, and then transmitting them to the brain for further analysis and understanding. The function of the retina is not only to receive light, it also undertakes preliminary image processing, including image enhancement, contrast adjustment, motion detection, etc. In this process, the retina performs complex perception and processing functions to ensure that the brain receives efficient and useful visual information. The photoreceptor cells of the retina (rods and cones) are the main photoreceptors. Rods are sensitive in low-light environments and can perceive the brightness of light, while cones can perceive color and are responsible for capturing details. Photoreceptor cells convert light into electrical signals through photoreceptor substances (such as rhodopsin).
[0052] Bipolar cells are located in the middle layer of the retina, between the photoreceptor cells and the ganglion cells. Bipolar cells are mainly divided into two categories according to the different ways they transmit signals: Excitatory bipolar cells (ON bipolar cells): This type of bipolar cell is mainly associated with the ON (bright) response of the photoreceptor cells. They respond positively to the increase in light. When the photoreceptor cells sense the increase in light, the excitatory bipolar cells will produce depolarization (the potential rises) and transmit the signal to the downstream ganglion cells. Inhibitory bipolar cells (OFF bipolar cells): This type of bipolar cell is associated with the OFF (dark) response of the photoreceptor cells. They are activated when the light intensity of the photoreceptor cells decreases, produce a depolarization response, and transmit the signal to the downstream ganglion cells. Simply put, the stronger the light stimulus, the more excited the excitatory bipolar cells are, and conversely, the more inhibited the inhibitory bipolar cells are.
[0053] The ability to have two cells respond in opposite directions to the same light stimulus is one of the key settings for the biological visual system to have excellent encoding performance. This setting plays a crucial role in image enhancement tasks. The spatial area formed by all stimulus sites that can trigger a cell response is called the receptive field of the cell. Figure 4 As shown in the figure, the receptive field of a bipolar cell is generally a concentric circle structure. It is divided into a central circle and a peripheral ring. The former is called the receptive field center, and the latter is called the receptive field periphery. Depending on the type of bipolar cells in the center and periphery, they can be divided into two types: light-on type and light-off type. They will produce different or even opposite responses to the same stimulus applied to the center and periphery of the receptive field. For example, some bipolar cells become more activated when the light in the center of the receptive field increases, but their activation decreases when the light in the periphery increases. Other bipolar cells become more activated when the light in the center of the receptive field decreases, but their activation decreases when the light in the periphery decreases. In this way, the receptive field formed by bipolar cells completes the encoding and integration of visual signals in a local area.
[0054] This application uses bimodal transistors to construct a physical convolution kernel array and uses convolution calculation to simulate the behavior of bipolar cells that react differently to light stimulation. The bimodal transistor designed in this application can use light signals as input and achieve excitability (such as Figure 5 shown) and inhibitory (as Figure 6 (As shown) There are two photocurrents. When a positive gate voltage is applied, the top light absorbing layer generates photogenerated carriers after receiving light stimulation, and the electrons drift downward under the action of the electric field force and enter the semiconductor layer. Since electrons are the main carriers in n-type semiconductors, the carrier concentration increases after receiving electrons from the light absorbing layer, and the current is gained, showing excitability. When a negative gate voltage is applied, the light absorbing layer generates photogenerated carriers after receiving light stimulation, and the electrons drift upward under the action of the electric field force, leaving holes to accumulate at the interface between the light absorbing layer and the semiconductor layer. At this time, the electrons in the semiconductor layer also drift upward under the longitudinal electric field force, and recombine with the holes at the interface. This process reduces the electron concentration in the semiconductor layer, reduces the current, and shows inhibition. As shown Figure 7 As shown, the present application can also change the intensity of the current response by changing the magnitude of the forward gate voltage. The greater the forward gate voltage, the greater the current response intensity.
[0055] In some specific embodiments of the present application, a bimodal transistor is prepared by the following method:
[0056] The heavily doped Si(p++) substrate was cleaned and ultrasonically cleaned with deionized water for 15 minutes, then dried under a stream of N2. The cleaned substrate was then subjected to a hydrophilic treatment and plasma treatment for 15 minutes to hydrophilize the film surface. An Al2O3-Li+ precursor solution was spin-coated on the treated substrate at 4500 rpm for 20 seconds, followed by annealing at 320°C in air for 60 minutes. After the AlOx film cooled, an InOx precursor solution was spin-coated on the dielectric layer at 3500 rpm for 30 seconds, followed by annealing at 270°C for 60 minutes. 30nm thick aluminum source / drain electrodes were deposited on the semiconductor layer using thermal evaporation. A PbS quantum dot dispersion was then spin-coated on the semiconductor layer at 2000 rpm for 20 seconds, followed by annealing at 200°C for 30 minutes, resulting in a bimodal transistor.
[0057] By wiring and packaging the 9 prepared bimodal transistors, a 3×3 physical convolution kernel array can be obtained, and its equivalent circuit diagram is shown in the following figure. Figure 8 By applying different gate voltage strategies to each transistor, several common types of convolution kernels for image enhancement and feature extraction can be realized. Figure 9 As shown, for T 1,1 The cell applies a forward gate voltage, T 0,1 、T1,0 、T 1,2 、T 2,1 Applying a negative gate voltage can obtain a convolution kernel for image sharpening. 1,1 Applying a positive gate voltage to the unit and a negative gate voltage to the other eight units can obtain a convolution kernel for image edge detection. Applying a small positive gate voltage to each unit can obtain a convolution kernel for image blurring. Figure 10 As shown, the physical convolution kernel is used to perform sharpening, edge detection and image blur simulation processing on the frog image respectively. The physical convolution kernel array designed in this application shows excellent image feature extraction function.
[0058] This application applies the bionic principle of the central optic nerve:
[0059] Biological synapses are the core units of information and computation in the central nervous system. They dynamically regulate the strength of connections between neurons through synaptic plasticity, forming the physiological basis for learning and memory. Among them, long-term potentiation (LTP) and long-term depression (LTD) are key mechanisms for non-volatile weight regulation: high-frequency neural activity triggers LTP, which promotes an increase in the number of postsynaptic membrane receptors through calcium influx, permanently enhancing signal transmission efficiency; low-frequency activity reduces receptor density through LTD, weakening connection strength over the long term. This bidirectional regulation of synaptic weights based on activity patterns enables biological neural networks to adapt to their environment.
[0060] Artificial synapses, inspired by biological mechanisms, should be able to mimic nonvolatile weight regulation. By controlling the device's conductance (weight) through the amplitude, frequency, or timing of electrical pulses, they achieve "write-hold-erase" functionality equivalent to that of biological synapses. This biomimetic design enables neuromorphic chips to achieve parallel computing and online learning with extremely low energy consumption, providing the hardware foundation for overcoming the bottlenecks of the von Neumann architecture.
[0061] The bimodal transistor designed in this application has excellent synaptic computing capabilities. Figure 11 As shown in Figure 2, when the gate voltage of the device is swept bidirectionally, the channel current exhibits obvious hysteresis behavior. The observed memory window of approximately 2V is the key to the device's non-volatile current. When a positive voltage pulse is applied to the gate, lithium ions are ejected from the AlO x Migrate to the channel, leading to the improvement of channel current and the formation of excitatory postsynaptic current. Figure 12 As shown, when the pulse stimulation ends, the ions near the channel electrolyte interface will be + The concentration gradient diffuses away from the interface, and the current slowly decreases, showing long-term non-volatility, corresponding to the long-term potentiation in biological synapses. Thereafter, when the presynaptic stimulus is negative, the downward trend of the channel helps to inhibit the postsynaptic current, thereby generating a long-term inhibitory current. Figure 13As shown, the LTP (Long-Term Potentiation) and LTD (Long-Term Depression) conductance values of the device are normalized, which can be defined as the weight modulation rules in the artificial neural network and participate in the weight deployment in the subsequent array network.
[0062] The embodiments of the present invention are described in detail above, but the present invention is not limited to the described embodiments. It is apparent to those skilled in the art that various changes, modifications, substitutions, and variations of these embodiments may be made without departing from the principles and spirit of the present invention, and the changes still fall within the scope of protection of the present invention.
Claims
1. An artificial visual neural system based on a dual-modal neural device, characterized in that: It includes a dual-modal visual information computing array composed of a physical convolution kernel array and a synaptic computing array. The physical convolution kernel array is used to receive ambient light signals and convert them into electrical signals for output. The synaptic computing array is used to receive the electrical signals and convert them into visual signals. The physical convolution kernel array and the synaptic computing array are both composed of a number of dual-modal transistors, and the dual-modal transistors have light signal perception and synaptic computing capabilities.
2. The artificial visual nervous system according to claim 1, characterized in that The bimodal transistor includes, from bottom to top, a substrate, a dielectric layer, a semiconductor layer, a source electrode and a drain electrode respectively arranged on both sides of the semiconductor layer away from the surface of the dielectric layer, and a light absorption layer arranged on the semiconductor layer away from the surface of the dielectric layer and located between the source electrode and the drain electrode.
3. The artificial visual nervous system according to claim 2, characterized in that The substrate is a P-type silicon substrate.
4. The artificial visual nervous system according to claim 2, characterized in that The dielectric layer material is selected from Al2O3, ZrO x , HfO x Any one of .
5. The artificial visual nervous system according to claim 2, characterized in that The semiconductor layer is made of n-type semiconductor material.
6. The artificial visual nervous system according to claim 5, characterized in that The semiconductor layer material is selected from InO x , IGZO, or IZO.
7. The artificial visual nervous system according to claim 2, characterized in that The light absorption layer material is selected from any one of PbS, perovskite, MXene, and MoFs.
8. The artificial visual nervous system according to claim 2, characterized in that The preparation of the bimodal transistor comprises the following steps: S1, cleaning the substrate; S2, growing a dielectric layer on the upper surface of the substrate; S3, growing a semiconductor layer on the surface of the dielectric layer away from the substrate; S4, preparing a source electrode and a drain electrode on a surface of the semiconductor layer away from the dielectric layer; S5. Growing or transferring a light absorbing layer on a surface of the semiconductor layer away from the dielectric layer and between the source electrode and the drain electrode.
9. The artificial visual nervous system according to claim 8, characterized in that The bimodal transistor adopts a solution method to prepare the dielectric layer, the semiconductor layer and the light absorption layer.
10. The artificial visual nervous system according to claim 9, characterized in that A substrate hydrophilic treatment step is also included between step S1 and step S2.
Citation Information
Patent Citations
Artificial nerve synapse transistor based on graphene / carbon nanotube composite absorbing layer
CN106653850A
Bionic synaptic device, manufacturing method and application thereof
CN110739393A
Retina form photoelectric sensor array and picture convolution processing method thereof
CN111370526A
Novel brain-like vision system
CN111950720A
Multi-sensing ferroelectric semiconductor synaptic device for neuromorphic calculation as well as preparation method and application of multi-sensing ferroelectric semiconductor synaptic device
CN118139518A
Cited By
Direction-selective sensing-storage-calculation integrated chip based on heterogeneous integration and preparation method thereof
CN121705237A