Imaging method
The imaging system with a colorless sensor and AI-based colorization addresses light attenuation issues, providing high-sensitivity and accurate color reproduction in low-light environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SEMICON ENERGY LAB CO LTD
- Filing Date
- 2025-02-25
- Publication Date
- 2026-05-19
AI Technical Summary
Conventional image sensors use color filters that attenuate light, leading to reduced light reception and poor image quality in low-light conditions, especially in security cameras and night vision systems, where accurate color reproduction is essential.
An imaging system using a solid-state image sensor without a color filter, combined with an AI system that colorizes black and white images using a neural network, enabling high-sensitivity imaging and faithful color reproduction.
The system achieves clear and vivid color images even in low-light conditions, allowing for accurate identification of individuals and objects, enhancing security and surveillance capabilities.
Smart Images

Figure 0007862129000001 
Figure 0007862129000002 
Figure 0007862129000003
Abstract
Description
[Technical Field]
[0001] One aspect of the present invention relates to a neural network and an imaging system using the same. Another aspect of the present invention relates to an electronic device using a neural network. Another aspect of the present invention relates to a vehicle using a neural network. The present invention relates to an imaging system that obtains a color image from a grayscale image obtained with a solid-state image sensor using image processing technology. The present invention also relates to a video surveillance system, security system, or safety information provision system using the imaging system.
[0002] Furthermore, one aspect of the present invention is not limited to the above-mentioned technical field. One aspect of the invention disclosed herein relates to a product, a method, or a method of manufacture. One aspect of the present invention relates to a process, a machine, a manufacture, or a composition of matter. More specifically, examples of technical fields of one aspect of the present invention disclosed herein include semiconductor devices, display devices, light-emitting devices, energy storage devices, memory devices, electronic devices, lighting devices, input devices, input / output devices, methods for driving them, or methods for manufacturing them.
[0003] In this specification, the term "semiconductor device" refers to all devices that can function by utilizing semiconductor properties, and electro-optical devices, semiconductor circuits, and electronic devices are all considered semiconductor devices. [Background technology]
[0004] Traditionally, images captured using an image sensor have been colorized using a color filter. Image sensors are widely used as components for imaging in digital cameras and video cameras. They are also used as part of security equipment such as security cameras, and in such equipment, accurate imaging is necessary not only in bright daylight but also at night or in dark places with little or no lighting, requiring an image sensor with a wide dynamic range.
[0005] Furthermore, advancements in AI (Artificial Intelligence) technology are remarkable, and for example, there is active development of automatic colorization technology that uses AI to colorize old black and white photographs taken with photographic film. Known methods for AI-based colorization include training with a large amount of image data to generate a model, and then achieving colorization through inference using the resulting generative model. Machine learning is a part of AI.
[0006] A technology for constructing transistors using oxide semiconductor thin films formed on a substrate is attracting attention. For example, Patent Document 1 discloses an imaging device in which an oxide semiconductor transistor with an extremely low off-current is used in the pixel circuit.
[0007] Furthermore, a technology for adding computational functions to an imaging device is disclosed in Patent Document 2. In addition, a technology related to super-resolution processing is disclosed in Patent Document 3. [Prior art documents] [Patent Documents]
[0008] [Patent Document 1] Japanese Patent Publication No. 2011-119711 [Patent Document 2] Japanese Patent Publication No. 2016-123087 [Patent Document 3] Japanese Patent Publication No. 2010-262276 [Overview of the project] [Problems that the invention aims to solve]
[0009] Conventional image sensors and other imaging devices use color filters to produce color images. Image sensors are sold with color filters already attached, and these image sensors are combined with lenses and other components as needed to be mounted in electronic devices. Simply placing a color filter over the light-receiving area of the image sensor reduces the amount of light reaching that area. Therefore, attenuation of the received light due to the color filter is unavoidable. Various color filters are used in combination with image sensors by different manufacturers, and optimization efforts are made, but they do not capture the wavelength of light to be passed through with pinpoint accuracy, but rather receive broad light across a certain wavelength range.
[0010] For example, in security cameras used in surveillance systems, there is a challenge in shooting in dim light conditions: insufficient light can result in images clear enough to recognize faces. For those installing security cameras, color images are preferred over black and white images from infrared cameras.
[0011] For example, when photographing underwater, there are many places with insufficient light, and while a light source is necessary in deep water, if you want to photograph fish, the light source may scare the fish away. Also, because light does not penetrate easily underwater, it is difficult to photograph distant fish even with a light source.
[0012] One of the challenges is to provide an imaging method and system that obtains images with high visibility and faithful reproduction of actual colors without using color filters.
[0013] In particular, when imaging in dark places with little lighting, such as in the evening or at night, there is a challenge in achieving high-sensitivity imaging due to the low amount of light received. As a result, images taken in dark places are inferior in terms of visibility compared to images taken in bright places.
[0014] Since a security camera placed in a light environment with a narrow wavelength range needs to accurately grasp the situation reflected in the captured image as a clue leading to an incident or accident, etc., it is necessary to accurately capture the characteristics of the objects reflected in real time. Therefore, for example, in the case of a night vision camera, it is important to focus on the object and capture an image with good visibility of the object even in a dark place.
[0015] In some night vision cameras, a color video is obtained by using an infrared light source and separating colors with a special color filter. However, since the reflected light of infrared rays is used, there is a problem that the color is not reflected or is expressed as a different color depending on the subject. This problem often occurs when the subject is a material that absorbs infrared rays. For example, human skin is likely to be imaged whiter than it actually is, and warm colors such as yellow may be imaged blue.
[0016] One of the problems is to provide an imaging method and an imaging system that have high visibility even in imaging at night or in a dark place and can obtain an image faithful to the same color as when there is external light.
Means for Solving the Problems
[0017] The imaging system of the present invention includes a solid-state imaging device that does not have a color filter, a storage device, and a learning device. Since a color filter is not used, light attenuation can be avoided, and imaging with high sensitivity can be performed even with a small amount of light.
[0018] Since the camera does not have a color filter, the obtained black and white image data (analog data) is colorized using an AI system. The AI system, that is, a learning device that uses training data stored in memory, uses extracted features to perform inference and adjust the focus, enabling the acquisition of highly visible color images (colorized image data (digital data)) even when imaging at night or in dark places. The learning device includes at least a neural network section and is capable of not only learning but also inference and data output. In addition, the learning device may perform inference using pre-trained features. In that case, the pre-trained features are stored in memory and calculations are performed to achieve data output at a level comparable to when pre-trained features are not used.
[0019] Furthermore, if the subject's outline becomes unclear due to insufficient light, the boundary may not be distinguishable, potentially resulting in incomplete coloring of that area.
[0020] Therefore, it is preferable to use an image sensor without a color filter to acquire black and white image data with a wide dynamic range, repeat the super-resolution processing multiple times, then identify the color boundaries and colorize the image. Alternatively, the training data may be subjected to super-resolution processing at least once. By mixing color illustrations (animations) in addition to color photographic images as training data when creating a learning model for identifying color boundaries, it is possible to obtain colorized image data with clearly defined color boundaries. Super-resolution processing refers to image processing that generates high-resolution images from low-resolution images.
[0021] Furthermore, by preparing a pre-trained model, it is possible to acquire relatively bright monochrome image data through imaging without using a flash light source in situations with insufficient light levels, and then obtain vividly colorized image data by colorizing that monochrome image data.
[0022] A video surveillance system, security system, or safety information provision system using the above imaging system can clearly capture images even in relatively dark places.
[0023] Specifically, it is a surveillance system equipped with a security camera, the security camera having a solid-state image sensor without a color filter, a learning device, and a storage device, and while the security camera detects a person, the solid-state image sensor takes an image and runs a software program that creates colorized image data through inference by the learning device using training data in the storage device. [Effects of the Invention]
[0024] The imaging system disclosed herein allows for the acquisition of clear color images even in low-light and dimly lit conditions. Therefore, it is relatively easy to identify individuals (such as faces) or distinguish clothing features based on the obtained color images. By applying this imaging system to security cameras, it is also possible to estimate a person's face in the color image and display it on a display device.
[0025] In particular, when acquiring 8K-size input image data with a solid-state image sensor, the light-receiving area of each individual pixel of the solid-state image sensor becomes smaller, resulting in a reduced amount of light being obtained. However, the imaging system disclosed herein does not use a color filter with the solid-state image sensor, so there is no reduction in light intensity due to the color filter. As a result, 8K-size image data can be captured with high sensitivity. [Brief explanation of the drawing]
[0026] [Figure 1] Figure 1 is a block diagram showing one embodiment of the present invention. [Figure 2] Figure 2 shows an example of source code illustrating one aspect of the present invention. [Figure 3] Figure 3 shows an example of source code illustrating one aspect of the present invention. [Figure 4] Figure 4 is an example of a flowchart illustrating one aspect of the present invention. [Figure 5] Figure 5 is a block diagram illustrating the imaging device. [Figure 6] Figure 6 is a diagram illustrating the pixel block 200 and circuit 201. [Figure 7] Figure 7 is a diagram illustrating pixel 100. [Figure 8] Figures 8A, 8B, and 8C illustrate the filter. [Figure 9] Figure 9 is an example of a flowchart illustrating one aspect of the present invention. [Figure 10] Figure 10 shows an example of source code illustrating one aspect of the present invention. [Figure 11] Figures 11A, 11B, and 11C show examples of application products illustrating one aspect of the present invention. [Modes for carrying out the invention]
[0027] Embodiments of the present invention will be described in detail below with reference to the drawings. However, it will be readily apparent to those skilled in the art that the present invention is not limited to the following description, and its form and details can be modified in various ways. Furthermore, the present invention is not to be interpreted as being limited to the embodiments described below.
[0028] (Embodiment 1) An example of the configuration of an imaging system 21 used in a video surveillance system or security system will be explained with reference to the block diagram shown in Figure 1.
[0029] The data acquisition device 10 is a semiconductor chip including a solid-state image sensor 11 and an analog arithmetic circuit 12, and does not have a color filter. The data acquisition device 10 has an optical system such as a lens. The optical system can have any configuration as long as its imaging characteristics are known, and is not particularly limited.
[0030] The A / D circuit 13 (also called an A / D converter) is an analog-to-digital conversion circuit that converts the analog data output from the data acquisition device 10 into digital data. If necessary, an amplification circuit may be provided between the data acquisition device 10 and the A / D circuit 13 to amplify the analog signal before conversion to digital data.
[0031] The memory unit 14 is a circuit that stores the converted digital data, and is configured to store the data before inputting it to the neural network unit 16, but is not limited to this configuration. Depending on the amount of data output from the data acquisition device or the data processing capacity of the image processing device, for small amounts of data, the output from the A / D circuit 13 may be input directly to the neural network unit 16 without storing it in the memory unit 14. Alternatively, the output from the A / D circuit 13 may be input to the neural network unit 16 located remotely using internet communication. For example, the neural network unit 16 may be built on a server capable of bidirectional communication.
[0032] The image processing device 20 is a device for estimating contours or colors corresponding to the grayscale images obtained by the data acquisition device 10. The image processing device 20 is executed in two distinct stages: a first stage of learning and a second stage of estimation. In this embodiment, the data acquisition device 10 and the image processing device 20 are configured as separate devices, but it is also possible to configure the data acquisition device 10 and the image processing device 20 as an integrated unit. When configured as an integrated unit, it is also possible to update the feature quantities obtained by the neural network unit in real time.
[0033] The neural network unit 16 is implemented by software computation using a microcontroller. A microcontroller is a computer system integrated into a single integrated circuit (IC). If the computational scale or the amount of data to be handled is large, multiple ICs may be combined to configure the neural network unit 16. These multiple ICs include at least one learning device. Furthermore, a microcontroller equipped with Linux® is preferable because it allows the use of free software, thereby reducing the total cost of configuring the neural network unit 16. However, it is not limited to Linux®, and other operating systems (OS) may also be used.
[0034] The learning process for the neural network unit 16 shown in Figure 1 is described below. Training data is pre-stored in the memory unit 18. The training data can be a training set or similar, and the image types include landscape photographs, portraits, and illustrations. The learning device performs inference using this training data. The learning device can be configured in any way as long as it can perform inference using the training data in the memory unit 18 based on black and white analog data and output colorized image data. Furthermore, if pre-trained features are used, the learning device can be configured in any way as long as it can perform calculations using the feature data in the memory unit 18 based on black and white analog data and output colorized image data. Using pre-trained features reduces the amount of data and computation, offering the advantage of a smaller configuration, for example, a learning device consisting of one or two ICs.
[0035] The program will be created using Python under a Linux (registered trademark) operating environment. In this embodiment, Keras dataframes will be used. Keras is a library that provides convenient functions for deep learning. In particular, it makes it easy to read and write data to the hidden layers of the neural network and to change the weight coefficients of the neural network. In addition, Numpy will be loaded in order to handle the Numpy library in Python. Furthermore, OpenCV will be loaded for image editing.
[0036] The source code corresponding to the above process is shown sequentially below the line labeled In[1] in Figure 2.
[0037] Next, the data is loaded. Color images are used as the data. It is preferable to prepare thousands or tens of thousands of color images. If the number of color image files is enormous, the loading process will be computationally burdensome, so it may be better to convert them to a format such as h5py before loading.
[0038] Colorization can be achieved using the following convolutional neural network (also called an artificial neural network). For example, in Python, it can be written using Keras as shown below the row labeled In[X] in Figure 3 (where X is any number and can be changed according to the program). Such a neural network is also called a U-net. The U-net has a data interpolation function called skip connections, which moves data from the convolutional layer to the deconvolutional layer, preventing the vanishing of gradients and enabling the construction of a good learning model.
[0039] During training, the grayscale versions of the color images used as training data are input to the neural network shown in In[X] in Figure 3. For example, OpenCV can be used to convert the images to grayscale.
[0040] You may also use GaN (Generative Adversarial Networks) with the above neural network as the Generator.
[0041] The output data from the neural network unit 16 and the time data from the time information acquisition device 17 are linked and stored in the large-scale storage device 15. The large-scale storage device 15 stores the data obtained from the start of imaging onwards.
[0042] The display unit 19 (including time display and video display) may be equipped with an operation input unit such as a touch panel, allowing the user to select and observe data stored in the large-scale storage device 15 as needed. The display unit 19 may also be able to access the large-scale storage device 15 remotely via internet communication, and the large-scale storage device 15 may be equipped with a transmitting antenna or a receiving antenna. The imaging system 21 can be used in a video surveillance system or a security system.
[0043] Furthermore, the display unit 19 can also be the display unit 19 of the user's mobile device (such as a smartphone). By accessing the large-scale storage device 15 from the display unit of the mobile device, the user can be monitored regardless of their location.
[0044] The imaging system 21 is not limited to being installed on the walls of a room; by mounting all or part of the imaging system 21 on an unmanned aerial vehicle (also called a drone) equipped with rotor blades, it is also possible to perform aerial video surveillance. This is particularly useful for imaging in low-light environments such as evening or nighttime when streetlights are not on.
[0045] Furthermore, although this embodiment has described a video surveillance system or security system, it is not particularly limited, and can also be applied to vehicles capable of semi-autonomous driving or fully autonomous driving by combining a camera or radar that captures images of the area around the vehicle with an ECU (Electronic Control Unit) that performs image processing, etc. Vehicles using electric motors have multiple ECUs, which perform engine control, etc. The ECU includes a microcomputer. The ECU is connected to a CAN (Controller Area Network) installed in the electric vehicle. CAN is one of the serial communication standards used as an in-vehicle LAN. The ECU uses a CPU or GPU. For example, one of the multiple cameras mounted on the electric vehicle (such as a drive recorder camera or rear camera) may be a solid-state image sensor without a color filter, and the obtained monochrome image may be inferred by the ECU via CAN to create a colorized image, which can then be displayed on an in-vehicle display device or a portable information terminal display.
[0046] (Embodiment 2) In this embodiment, Figure 4 shows an example of a flow for colorizing a black and white image obtained by the solid-state image sensor 11 using the block diagram and program shown in Embodiment 1.
[0047] Install the imaging system 21 shown in Embodiment 1 in the location to be monitored (room, parking lot, entrance, etc.), activate it, and start continuous shooting.
[0048] First, we begin preparing to retrieve the data (S1).
[0049] Black and white image data is acquired using a solid-state image sensor without a color filter (S2). Note that a configuration in which multiple solid-state image sensors are arranged in a matrix direction is sometimes called a pixel array.
[0050] Next, the obtained analog data is filtered using a multiply-accumulate circuit (S3).
[0051] Steps S2 and S3 are performed using the imaging device shown in Figure 5. The imaging device will be described below.
[0052] Figure 5 is a block diagram illustrating the imaging device. The imaging device includes a pixel array 300, and circuits 201, 301, 302, 303, 304, 305, and 306. Note that each of circuits 201 and 301 through 306 is not limited to a single circuit configuration, but may be composed of a combination of multiple circuits. Alternatively, any multiple of the above circuits may be integrated. In addition, other circuits may be connected.
[0053] The pixel array 300 has imaging and calculation functions. Circuits 201 and 301 have calculation functions. Circuit 302 has either calculation or data conversion functions. Circuits 303, 304, and 306 have selection functions. Circuit 303 is electrically connected to the pixel block 200 via wiring 124. Circuit 304 is electrically connected to the pixel block 200 via wiring 123. Circuit 305 has the function of supplying potential for multiply-accumulate calculations to the pixels. A shift register or decoder can be used for the circuit with the selection function. Circuit 306 is electrically connected to the pixel block 200 via wiring 113. Note that circuits 301 and 302 may be provided externally.
[0054] The pixel array 300 has a plurality of pixel blocks 200. As shown in Figure 6, each pixel block 200 has a plurality of pixels 100 arranged in a matrix, and each pixel 100 is electrically connected to a circuit 201 via wiring 112. The circuit 201 can also be provided within the pixel block 200.
[0055] Furthermore, each pixel 100 is electrically connected to an adjacent pixel 100 via transistors 150 (transistors 150g to 150j). The function of transistor 150 will be described later.
[0056] At pixel 100, image data can be acquired and data can be generated by adding the image data to a weight coefficient. In Figure 6, the number of pixels in pixel block 200 is shown as 3x3 as an example, but this is not limited to this. For example, it can be 2x2, 4x4, etc. Alternatively, the number of pixels in the horizontal and vertical directions may differ. Furthermore, some pixels may be shared between adjacent pixel blocks.
[0057] The pixel block 200 and circuit 201 can be operated as a multiply-accumulate circuit.
[0058] As shown in Figure 7, the pixel 100 may have a photoelectric conversion device 101, a transistor 102, a transistor 103, a transistor 104, a transistor 105, a transistor 106, and a capacitor 107.
[0059] One electrode of the photoelectric conversion device 101 is electrically connected to either the source or drain of transistor 102. The other electrode of the source or drain of transistor 102 is electrically connected to either the source or drain of transistor 103, the gate of transistor 104, and one electrode of capacitor 107. One electrode of the source or drain of transistor 104 is electrically connected to either the source or drain of transistor 105. The other electrode of capacitor 107 is electrically connected to either the source or drain of transistor 106.
[0060] The other electrode of the photoelectric conversion device 101 is electrically connected to the wiring 114. The other source or drain of transistor 103 is electrically connected to the wiring 115. The other source or drain of transistor 105 is electrically connected to the wiring 112. The other source or drain of transistor 104 is electrically connected to the GND wiring, etc. The other source or drain of transistor 106 is electrically connected to the wiring 111. The other electrode of capacitor 107 is electrically connected to the wiring 117.
[0061] The gate of transistor 102 is electrically connected to wiring 121. The gate of transistor 103 is electrically connected to wiring 122. The gate of transistor 105 is electrically connected to wiring 123. The gate of transistor 106 is electrically connected to wiring 124.
[0062] Here, node FD is defined as the electrical connection point between the other source or drain of transistor 102, the one source or drain of transistor 103, one electrode of capacitor 107, and the gate of transistor 104. Also, node FDW is defined as the electrical connection point between the other electrode of capacitor 107 and one source or drain of transistor 106.
[0063] Wires 114 and 115 can function as power lines. For example, wire 114 can function as a high-potential power line, and wire 115 can function as a low-potential power line. Wires 121, 122, 123, and 124 can function as signal lines that control the conduction of each transistor. Wire 111 can function as a wire that supplies a potential corresponding to a weighting coefficient to pixel 100. Wire 112 can function as a wire that electrically connects pixel 100 and circuit 201. Wire 117 can function as a wire that electrically connects the other electrode of capacitor 107 of one pixel to the other electrode of capacitor 107 of another pixel via transistor 150 (see Figure 6).
[0064] An amplification circuit or a gain adjustment circuit may also be electrically connected to wiring 112.
[0065] A photodiode can be used as the photoelectric conversion device 101. Any type of photodiode is acceptable, including Si photodiodes with silicon as the photoelectric conversion layer, and organic photodiodes with an organic photoconductive film as the photoelectric conversion layer. Furthermore, if you wish to improve light detection sensitivity at low light levels, it is preferable to use an avalanche photodiode.
[0066] Transistor 102 may have the function of controlling the potential of node FD. Transistor 103 may have the function of initializing the potential of node FD. Transistor 104 may have the function of controlling the current that circuit 201 flows according to the potential of node FD. Transistor 105 may have the function of selecting pixels. Transistor 106 may have the function of supplying a potential corresponding to a weighting coefficient to node FDW.
[0067] When an avalanche photodiode is used in the photoelectric conversion device 101, a high voltage may be applied, and it is preferable to use a high-voltage transistor for the transistor connected to the photoelectric conversion device 101. For example, a transistor using a metal oxide in the channel formation region (hereinafter referred to as an OS transistor) can be used as the high-voltage transistor. Specifically, it is preferable to apply an OS transistor to transistor 102.
[0068] Furthermore, OS transistors also possess the characteristic of extremely low off-current. By using OS transistors for transistors 102, 103, and 106, the period during which charge can be held at nodes FD and FDW can be made extremely long. Therefore, a global shutter method that performs charge accumulation operation simultaneously at all pixels can be applied without complicating the circuit configuration or operating method. In addition, it is possible to hold image data at node FD and perform multiple calculations using that image data.
[0069] On the other hand, it is sometimes desirable for transistor 104 to have excellent amplification characteristics. Also, it is sometimes preferable to use a transistor with high mobility that can operate at high speed for transistor 106. Therefore, transistors 104 and 106 may be transistors that use silicon in the channel formation region (hereinafter referred to as Si transistors).
[0070] Furthermore, OS transistors and Si transistors may be applied in any combination, not limited to the above. Alternatively, all transistors may be OS transistors, or all transistors may be Si transistors. Examples of Si transistors include those made of amorphous silicon, and those made of crystalline silicon (microcrystalline silicon, low-temperature polysilicon, single-crystal silicon).
[0071] The potential of node FD at pixel 100 is determined by the sum of the reset potential supplied from wiring 115 and the potential (image data) generated by photoelectric conversion by photoelectric conversion device 101. Alternatively, it is determined by capacitive coupling with a potential corresponding to a weighting coefficient supplied from wiring 111. Therefore, transistor 105 can supply a current corresponding to data obtained by adding an arbitrary weighting coefficient to the image data.
[0072] Note that the above is just one example of a circuit configuration for pixel 100, and the photoelectric conversion operation can also be performed with other circuit configurations.
[0073] As shown in Figure 6, each pixel 100 is electrically connected to one another by wiring 112. Circuit 201 can perform calculations using the sum of the currents flowing through the transistors 104 of each pixel 100.
[0074] Circuit 201 includes a capacitor 202, transistors 203, 204, 205, 206, and a resistor 207.
[0075] One electrode of capacitor 202 is electrically connected to either the source or drain of transistor 203. One of the sources or drains of transistor 203 is electrically connected to the gate of transistor 204. One of the sources or drains of transistor 204 is electrically connected to either the source or drain of transistor 205. One of the sources or drains of transistor 205 is electrically connected to either the source or drain of transistor 206. One electrode of resistor 207 is electrically connected to the other electrode of capacitor 202.
[0076] The other electrode of capacitor 202 is electrically connected to wiring 112. The other source or drain of transistor 203 is electrically connected to wiring 218. The other source or drain of transistor 204 is electrically connected to wiring 219. The other source or drain of transistor 205 is electrically connected to a reference power line such as the GND wiring. The other source or drain of transistor 206 is electrically connected to wiring 212. The other electrode of resistor 207 is electrically connected to wiring 217.
[0077] Wires 217, 218, and 219 can function as power lines. For example, wire 218 can function as a wire supplying a dedicated potential for reading. Wires 217 and 219 can function as high-potential power lines. Wires 213, 215, and 216 can function as signal lines controlling the conduction of each transistor. Wire 212 is an output line and can be electrically connected to, for example, the circuit 301 shown in Figure 5.
[0078] Transistor 203 can have a function of resetting the potential of wiring 211 to the potential of wiring 218. Wiring 211 is a wiring connected to one electrode of capacitor 202, one of the source or drain of transisitor 203, and the gate of transisitor 204. Transistors 204 and 205 can have a function as a source follower circuit. Transistor 206 can have a function of controlling reading. Note that circuit 201 has a function as a correlated double sampling circuit (CDS circuit), and can be replaced with a circuit of another configuration having the function.
[0079] In one aspect of the present invention, an offset component other than the product of image data (X) and weight coefficient (W) is removed, and target WX is extracted. WX can be calculated using data with and without imaging for the same pixel, and data when a weight is added to each of them.
[0080] The total current (I p ) flowing through pixel 100 when imaging is kΣ(X - V th ) 2 , and the total current (I p ) flowing through pixel 100 when a weight is added is kΣ(W + X - V th ) 2 . Also, the total current (I ref ) flowing through pixel 100 when not imaging is kΣ(0 - V th ) 2 , and the total current (I ref ) flowing through pixel 100 when a weight is added is kΣ(W - V th ) 2 . Here, k is a constant, and V th is the threshold voltage of transisitor 105.
[0081] First, the difference (data A) between the data with imaging and the data with a weight added to the data is calculated. kΣ((X - V th ) 2 - (W + X - V th ) 2 ) = kΣ(-W 2 - 2W·X + 2W·V th)
[0082] Next, we calculate the difference (data B) between the data without imaging and the data with weights added to it. kΣ((0-V th ) 2 -(WV th ) 2 )=kΣ(-W 2 +2W·V th )
[0083] Then, we take the difference between data A and data B. kΣ(-W 2 -2W·X+2W·V th -(-W 2 +2W·V th )) = kΣ(-2W·X). In other words, offset components other than the product of the image data (X) and the weight coefficient (W) can be removed.
[0084] Circuit 201 can read out data A and data B. The difference calculation between data A and data B can be performed, for example, by circuit 301.
[0085] Here, the weights supplied to the entire pixel block 200 function as a filter. As this filter, for example, a convolutional filter of a convolutional neural network (CNN) can be used. Alternatively, an image processing filter such as an edge detection filter can be used. Examples of edge detection filters include the Laplacian filter shown in Figure 8A, the Prewitt filter shown in Figure 8B, and the Sobel filter shown in Figure 8C.
[0086] If the pixel block 200 has 3x3 pixels 100, the elements of the edge extraction filter can be assigned as weights to each pixel 100 and supplied. As mentioned above, to calculate data A and data B, data with and without imaging, and the data with weights added to each of them can be used. Here, the data with and without imaging is data without weights, which can also be rephrased as data with a weight of 0 added to all pixels 100.
[0087] The edge extraction filters illustrated in Figures 8A to 8C are filters in which the sum of the filter elements (weights: ΔW) (ΣΔW / N, where N is the number of elements) is 0. Therefore, without having to supply ΔW=0 from another circuit, by performing the operation to acquire ΣΔW / N, it is possible to obtain data in which ΔW=0 is added to all 100 pixels.
[0088] This operation is equivalent to making the transistors 150 (transistors 150a to 150f) installed between the pixels 100 conduct (see Figure 6). By making the transistors 150 conduct, all node FDWs of each pixel 100 are short-circuited via the wiring 117. At this time, the charge accumulated in the node FDWs of each pixel 100 is redistributed, and when using the edge extraction filters exemplified in Figures 8A to 8C, the potential (ΔW) of the node FDW becomes 0 or approximately 0. Therefore, data equivalent to ΔW=0 can be obtained.
[0089] Furthermore, when supplying charge from a circuit outside the pixel array 300 to rewrite the weight (ΔW), the rewriting process takes time due to factors such as the capacitance of the long wiring 111. On the other hand, the pixel block 200 is a tiny area, and the wiring 117 is short and has small capacitance. Therefore, the operation of redistributing the charge accumulated in node FDW within the pixel block 200 allows for high-speed rewriting of the weight (ΔW).
[0090] In the pixel block 200 shown in Figure 6, transistors 150a to 150f are electrically connected to different gate lines (wirings 113a to 113f). In this configuration, the conduction of transistors 150a to 150f can be controlled independently, and the operation to acquire ΣΔW / N can be selectively performed. Also in Figure 6, transistors 150g to 150j are electrically connected to different gate lines (113g to 113i).
[0091] For example, when using the filters shown in Figures 8B and 8C, there are pixels that are initially supplied with ΔW=0. Assuming that ΣΔW / N=0, pixels supplied with ΔW=0 may be excluded from the pixels subject to summation. By excluding these pixels, it becomes unnecessary to supply potential to operate some of the transistors 150a to 150f, thus reducing power consumption. In Figures 6 and 8, an example is shown in which nine transistors 150 (transistors 150a to 150f) are placed between pixels 100, but the number of transistors 150 may be increased further. Also, in transistors 150g to 150j, some transistors may be omitted to eliminate parallel paths.
[0092] The data resulting from the sum-of-accumulate operation output from circuit 201 is sequentially input to circuit 301. In addition to the function of calculating the difference between data A and data B as described above, circuit 301 may have various other calculation functions. For example, circuit 301 can have the same configuration as circuit 201. Alternatively, the functions of circuit 301 may be replaced by software processing.
[0093] Furthermore, circuit 301 may have a circuit that performs calculations on the activation function. For example, a comparator circuit can be used for this circuit. In a comparator circuit, the result of comparing the input data with a set threshold value is output as binary data. That is, the pixel block 200 and circuit 301 can act as elements of a neural network.
[0094] The data output from circuit 301 is sequentially input to circuit 302. Circuit 302 can be configured to include, for example, a latch circuit and a shift register. This configuration allows for parallel-to-serial conversion, and the data input in parallel can be output as serial data to wiring 311.
[0095] The destination of wiring 311 is not limited. For example, it can be connected to the A / D circuit 13 or the neural network section 16 shown in Figure 1. Alternatively, wiring 311 may be connected to an FPGA (field-programmable gate array).
[0096] Next, the A / D circuit 13 converts the filtered analog data into digital data (S4).
[0097] Next, the converted digital data is stored in the memory unit 14 (digital memory unit) (S5).
[0098] Next, the digital data is converted into a signal format required by the subsequent inference program (such as JPEG®) (S6).
[0099] Next, the converted digital data is subjected to convolution processing using a CPU or the like to infer contours, colors, etc., and then colorized (S7). Instead of a CPU, a single IC chip integrated with a GPU (Graphics Processing Unit) and a PMU (Power Management Unit) may be used. The colorized image data is then output (S8). The colorized image data is then saved along with time data such as the date and time (S9). The data is stored in a large-capacity storage device 15, a so-called high-capacity storage device (such as a hard disk), or a database.
[0100] The above colorized image data is acquired repeatedly (during operation). By repeating this process, real-time colorization can also be performed.
[0101] The resulting colorized image data is based on a wide dynamic range black and white image obtained using an image sensor without a color filter. Therefore, even in cases where conventional image sensors with color filters would be unable to distinguish due to insufficient light, distinguishable colorized image data can be obtained. The imaging system shown in this embodiment allows each of the above steps (S4 to S9) to be implemented by one or more computers.
[0102] Furthermore, as a means of real-time colorization, latent variables (features) are monitored using cosine similarity, and the optical system is used to adjust the focus to minimize fluctuations. This makes it possible to adjust the focus so that the subject remains in focus even if it moves during shooting.
[0103] Furthermore, inference may be performed using the extracted features. Focus adjustment may be performed to minimize variability in the inference results. For example, when a person is inferred, focus adjustment may be performed so that its likelihood remains constant or increases.
[0104] By performing inference, it is possible to instantly identify what the imaged object is. For example, in security applications, if an object is determined to be dangerous, it is possible to immediately notify the necessary contacts via a smartphone or other mobile device. Furthermore, even if the image is out of focus, it is possible to remove blur and sharpen the image before performing inference.
[0105] (Embodiment 3) This embodiment demonstrates an example that enables smoother image processing or finer colorization compared to the colorized image data obtained in Embodiment 2.
[0106] Figure 9 shows a flowchart. Note that the same reference numerals are used for the same steps as in the flowchart shown in Figure 4 of Embodiment 2. Since steps S1-S6 and S8-S9 in Figure 4 are identical, a detailed explanation is omitted here.
[0107] As shown in Figure 9, after step S6, the converted digital data is subjected to super-resolution processing multiple times using the first learning model to infer contours (S7a).
[0108] Then, the digital data after super-resolution processing is colorized using a second learning model to infer color and other properties (S7b). The subsequent steps are the same as in Embodiment 2.
[0109] The training data for the second learning model is pre-processed by applying super-resolution processing multiple times, mixing animated images with photographic images, or enhancing contours using functions such as OPENCV's drawcontours function. The ratio of animated images to photographic images is set so that if the number of photographic images is 2, the number of animated images is 1. Animated images are a type of illustration, but they contain many edge components or color components. In the process of colorizing a black and white image, edges are extracted as features in the convolutional layer, and the color of each region of the image is inferred based on these features. Therefore, using images with many edge components as training data is effective in improving the efficiency of machine learning. By using animated images as training data, the amount of training data required to obtain a color image that can reach a certain standard can be reduced, shortening the time required for machine learning and simplifying the configuration of the neural network. The neural network is a part of machine learning. Deep learning is also a part of the neural network.
[0110] The training of the neural network portion of this embodiment is described below.
[0111] The program will be created using Python under a Linux (registered trademark) operating environment. In this embodiment, Keras dataframes will be used. Keras is a library that provides convenient functions for deep learning. In particular, it makes it easy to read and write data to the hidden layers of a neural network and to change the weight coefficients of the neural network. In addition, Numpy will be loaded in order to handle the Numpy library in Python. Furthermore, scipy will be used for image processing.
[0112] The source code corresponding to the above-mentioned process is shown in order below the line In[Y1] in Figure 10 (where Y1 is any number and can be changed to suit the program).
[0113] Next, the data is loaded. In this embodiment, the training data is a set of high-resolution images. Training data is a set of data used in supervised learning or data that has been classified. The above images can be color images or black and white images. It is preferable to prepare several thousand or tens of thousands of images. If the number of image files is enormous, the loading process will be computationally burdensome, so it may be better to convert them to, for example, h5py format before loading.
[0114] The super-resolution processing consists of a three-layer convolutional neural network. For example, in Python, the Keras Sequential model can be used and written as shown below the row In[Y2] in Figure 10 (where Y2 is any number and can be changed to suit the program as long as it is after Y1). Note that In[Y2] shows an example of inputting a 33x33 pixel image. In this embodiment, an example of a 33x33 pixel image is shown, but it is not particularly limited and a 2K size (1920x1080 pixel) image or a 4K size (3840x2160 pixel) image may also be used. Note that the size is the resolution, and the neural network should be designed to match the input image data or output image data. Also, the input image data and output image data may be different; for example, input image data may be obtained in QHD size (960x540 pixels) and then output as a 2K size image. When obtaining 8K-size input image data with a solid-state image sensor, the amount of light obtained by each individual sensor is small. Therefore, it is particularly preferable not to use a color filter, as this allows for a greater amount of light to be obtained.
[0115] During training, the neural network shown in In[Y3] is input with images obtained by reducing the resolution of the training data images. This resolution reduction process can be written, for example, using scipy, as shown below the row for In[Y3] in Figure 10 (where Y3 is any number and can be changed according to the program as long as it is after Y2). Here, we show the process of reducing the resolution to 1 / 3.
[0116] Using such training data and code, a model capable of outputting images can be created. During inference, this model can be used to input low-resolution images and output high-resolution images. The imaging system shown in this embodiment can implement each of the above steps (S4 to S9) on one or more computers.
[0117] A GaN (GaN) with the above neural network as the Generator may also be used.
[0118] The colorized image data obtained in this embodiment has smoother contours than the image data in Embodiment 2, and optimal colorization is achieved.
[0119] Furthermore, when preparing training images independently, for example, when you want to colorize a rare fish, even if you only have training images of similar fish with blurred outlines, you can still effectively train a colorization model using these training images.
[0120] (Embodiment 4) In this embodiment, electronic devices that can use the imaging device used in one aspect of the present invention include display devices, personal computers, image storage devices or image playback devices equipped with recording media, mobile phones, game consoles including portable ones, portable data terminals, e-book readers, cameras such as video cameras and digital still cameras, goggle-type displays (head-mounted displays), navigation systems, sound playback devices (car audio, digital audio players, etc.), photocopiers, facsimile machines, printers, printer-multifunction devices, automated teller machines (ATMs), and vending machines. Specific examples of these electronic devices are shown in Figure 11.
[0121] Figure 11A shows a surveillance camera, which has a housing 951, a lens 952, a support 953, etc. To acquire images from this surveillance camera, a shooting system according to one embodiment of the present invention can be provided. The housing 951 contains a neural network unit. Note that "surveillance camera" is a conventional name and does not limit its use. For example, a device that functions as a surveillance camera is also called a camera or video camera. The surveillance camera uses an image sensor that does not use a color filter. Furthermore, by incorporating the program shown in Embodiment 2 or Embodiment 3 as a software program and executing it in the neural network unit, colorized image data can be created. When multiple surveillance cameras are used, if at least one of them is the surveillance camera of this embodiment, color images can be acquired in dimly lit environments, which are difficult to acquire with conventional surveillance cameras, thus enhancing the surveillance system when combined with conventional surveillance cameras.
[0122] Figure 11B also shows a surveillance camera, which includes a support base 954, a camera unit 955, a protective cover 956, etc. The camera unit 955 is equipped with a rotation mechanism, and when installed on the ceiling, it is possible to capture images of the entire surroundings. The camera unit 955 can be used as an imaging device in a surveillance system according to one embodiment of the present invention. Furthermore, based on the data obtained by the camera unit 955, the neural network section of the camera unit 955 makes estimations, making it possible to identify suspicious persons from the information captured through colorization or super-resolution.
[0123] Figure 11C shows an example of an aircraft. The aircraft 6500 shown in Figure 11C has a propeller 6501, a camera 6502, a battery 6503, etc., and is capable of autonomous flight.
[0124] For example, image data captured by camera 6502 is stored in electronic component 6504. Electronic component 6504 can analyze the image data and detect the presence or absence of obstacles during movement. Multiple types of imaging devices may be used as camera 6502. Camera 6502 can be used as the imaging device in a surveillance system according to one embodiment of the present invention. Furthermore, by having the neural network unit make estimations based on the data obtained by camera 6502, it is possible to identify suspicious persons from the information captured through colorization or super-resolution.
[0125] The configurations, structures, and methods shown in this embodiment can be used in appropriate combination with the configurations, structures, and methods shown in other embodiments. [Explanation of symbols]
[0126] 10: Data acquisition device, 11: Solid-state image sensor, 12: Analog arithmetic circuit, 13: A / D circuit, 14: Memory section, 15: Large-scale memory device, 16: Neural network section, 17: Time information acquisition device, 18: Storage section, 19: Display section, 20: Image processing device, 21: Imaging system, 100: Pixel, 101: Photoelectric conversion device, 102: Transistor, 103: Transistor, 104: Transistor, 105: Transistor, 106: Transistor, 107: Capacitor, 111: Wiring, 112: Wiring, 113a: Wiring, 113f: Wiring, 114: Wiring, 115: Wiring, 117: Wiring, 121: Wiring, 122: Wiring, 123: Wiring, 124: Wiring, 150: Transistor, 150g: Transistor, 150h: Transistor 150i: Transistor, 150j: Transistor, 200: Pixel block, 201: Circuit, 202: Capacitor, 203: Transistor, 204: Transistor, 205: Transistor, 206: Transistor, 207: Resistor, 211: Wiring, 212: Wiring, 213: Wiring, 215: Wiring, 216: Wiring, 217: Wiring, 218: Wiring, 219: Wiring, 300: Pixel array, 301: Circuit, 302: Circuit, 303: Circuit, 304: Circuit, 305: Circuit, 306: Circuit, 311: Wiring, 951: Housing, 952: Lens, 953: Support, 954: Support base, 955: Camera unit, 956: Protective cover, 6500: Aircraft, 6501: Propeller, 6502: Camera, 6503: Battery, 6504: Electronic component
Claims
1. A solid-state image sensor without a color filter, A sum-of-accumulate circuit, A / D circuit and The memory section, Memory unit and, An imaging method using an imaging system having a memory device, The first step is to acquire analog data, which is monochrome image data, using the solid-state image sensor, A second step involves filtering the analog data using the sum-of-accumulate circuit, In the A / D circuit, a third step is to convert the filtered analog data into digital data, A fourth step is to store the converted digital data in the memory unit, A fifth step involves converting the digital data output from the memory unit into a signal format required by the inference program. A sixth step involves performing super-resolution processing multiple times on the digital data after conversion to a signal format using the first learning model stored in the memory unit to infer contours, A seventh step involves using a second learning model stored in the memory unit to infer color from the digital data after super-resolution processing and then colorizing it. An eighth step is to output the colorized digital data as image data, The ninth step is to store the colorized digital data together with time data in the storage device, Steps 1 through 9 above are repeated, The training data for the second learning model comprises color photographic images and color animated images. An imaging method wherein the proportion of color animation images in the training data for the second learning model is smaller than the proportion of color photographic images in the training data for the second learning model.
2. In claim 1, The first and second steps described above are performed using an imaging device. The imaging device comprises a pixel array having a plurality of pixel blocks, a first circuit, and a first wiring. Each of the plurality of pixel blocks comprises a first to third pixel and a first to third transistor, The first circuit comprises a fourth to seventh transistor, a capacitor, and a resistor. The first pixel is connected to the second pixel via the first transistor. The second pixel is connected to the third pixel via the second transistor and the third transistor, Each of the first pixel and the third pixel is connected via the first wiring to one electrode of the capacitor and one electrode of the resistor. The other electrode of the capacitor is connected to one of the source and drain of the fourth transistor and to the gate of the fifth transistor. An imaging method in which one of the source and drain of the fifth transistor is connected to one of the source and drain of the sixth transistor and one of the source and drain of the seventh transistor.
3. In claim 2, The other electrode of the resistor is connected to a second wire that functions as a power line. The source and the other drain of the fourth transistor are connected to a third wire that functions as a power line. The source and drain of the fifth transistor are connected to a fourth wire that functions as a power line. An imaging method in which the source and the other drain of the seventh transistor are connected to a fifth wiring that functions as an output line.