Imaging system

The imaging system addresses the challenge of capturing clear, colorized images in low-light conditions by using a solid-state imaging device without a color filter and AI-driven neural networks for colorization, resulting in enhanced visibility and color fidelity.

JP2025084825AActive Publication Date: 2025-06-03SEMICON ENERGY LAB CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2025027912
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-11
Filing Date
2025-02-25
Publication Date
2025-06-03
Estimated Expiration
2040-09-29

AI Technical Summary

Technical Problem

Conventional imaging systems using color filters face challenges in capturing clear, colorized images in low-light conditions due to light attenuation and inaccurate wavelength transmission, leading to inferior image visibility and color fidelity.

Method used

An imaging system that employs a solid-state imaging device without a color filter, utilizing AI-driven neural networks for colorization of black-and-white image data, thereby avoiding light attenuation and enhancing sensitivity and color accuracy even in low-light environments.

Benefits of technology

The system achieves high visibility and faithful color representation in low-light conditions, enabling clear identification of objects and characteristics, such as facial features, without the need for external light sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025084825000001_ABST
    Figure 2025084825000001_ABST
Patent Text Reader

Abstract

To solve a problem in which color images captured by conventional imaging devices such as image sensors use color filters, and image sensors are sold with color filters attached, and the image sensors are appropriately combined with lenses and installed in electronic devices, and simply placing a color filter over the light receiving area of an image sensor reduces the amount of light that reaches the light receiving area.SOLUTION: An imaging system according to the present invention includes a solid-state imaging element that does not have a color filter, a storage device, and a learning device. Since the system does not have a color filter, the obtained black and white image data (analog data) is colorized, but the coloring is performed using an AI system.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present invention relates to a neural network and an imaging system using the same. Another aspect of the present invention relates to an electronic device using a neural network. Another aspect of the present invention relates to a vehicle using a neural network. The present invention relates to an imaging system that obtains a color image from a black-and-white image obtained by a solid-state imaging device using image processing technology. The present invention relates to a video surveillance system, a security system, or a safety information providing system using such an imaging system.

[0002] Note that one aspect of the present invention is not limited to the above technical field. One aspect of the invention disclosed in this specification or the like relates to an article, a method, or a manufacturing method. One aspect of the present invention relates to a process, a machine, a manufacture, or a composition of matter. Therefore, as a more specific technical field of one aspect of the present invention disclosed in this specification or the like, semiconductor devices, display devices, light-emitting devices, power storage devices, storage devices, electronic devices, lighting devices, input devices, input / output devices, their driving methods, or their manufacturing methods can be cited as an example.

[0003] Note that in this specification, the semiconductor device refers to all devices that can function by utilizing semiconductor characteristics, and electro-optical devices, semiconductor circuits, and electronic devices are all semiconductor devices.

Background Art

[0004] Conventionally, a technique for colorizing an image captured using an image sensor using a color filter is known. Image sensors are widely used as components for imaging such as digital cameras or video cameras. In addition, since it is also used as part of security equipment such as security cameras, in such equipment, it is necessary to perform accurate imaging not only in bright places during the day but also at night or in dark places with little lighting and poor light, and an image sensor with a wide dynamic range is required.

[0005] Moreover, the progress of technologies using AI (Artificial Intelligence) is remarkable. For example, the development of automatic coloring technology that colors black-and-white photos using old photographic films with AI is also actively underway. For colorization by AI, there are known methods such as learning with a large amount of image data to generate a model and realizing colorization by inference using the obtained generation model. Note that machine learning is a part of AI.

[0006] Techniques for constructing transistors using oxide semiconductor thin films formed on a substrate have attracted attention. For example, an imaging device having a configuration in which a transistor having an extremely low off-current with an oxide semiconductor is used in a pixel circuit is disclosed in Patent Document 1.

[0007] Moreover, a technique for adding an arithmetic function to an imaging device is disclosed in Patent Document 2. Also, a technique related to super-resolution processing is disclosed in Patent Document 3.

Prior Art Documents

Patent Documents

[0008]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0009] Color images captured by imaging devices such as conventional image sensors use color filters. The imaging device is sold with a color filter attached, and the imaging device and lenses are appropriately combined and installed in an electronic device. The color filter is simply placed over the light-receiving area of the image sensor, reducing the amount of light reaching the light-receiving area. Therefore, attenuation of the received light amount due to the color filter is inevitable. Color filters combined with imaging devices vary from company to company and are being optimized, but they do not precisely acquire the wavelength of the light to be transmitted and receive broad light in a certain wavelength range.

[0010] For example, in a security camera used in a surveillance system or the like, there is also a problem that in shooting in dim light, the amount of light is insufficient, so an image clear enough to recognize a face cannot be obtained. For the installer of a security camera, a color image is preferred over a black-and-white image obtained by an infrared camera.

[0011] For example, in underwater photography, there are many places where the amount of light is insufficient, and a light source is required in deep water areas. However, when trying to photograph fish, the fish may escape due to the light source. Also, since light hardly reaches underwater, it is difficult to image fish in the distance even with a light source.

[0012] One of the problems is to provide an imaging method and an imaging system that can obtain an image with high visibility and faithful to the actual color without using a color filter.

[0013] In particular, in imaging in dark places with little lighting such as in the evening or at night, there is a problem that it is difficult to perform imaging with good sensitivity because the amount of received light is small. Therefore, images taken in dark places are inferior in terms of visibility compared to those taken in bright places.

[0014] Since a security camera placed in a light environment with a narrow wavelength range needs to accurately grasp the situation shown in the captured image as a clue leading to an incident or accident, etc., it is necessary to accurately capture the characteristics of the objects shown in real time. Therefore, for example, in the case of a night vision camera, it is important to focus on the object and capture an image with good visibility of the object even in a dark place.

[0015] In some night vision cameras, a color video is obtained by using an infrared light source and separating colors with a special color filter. However, since the reflected light of infrared rays is used, there is a problem that the color is not reflected or is expressed as a different color depending on the subject. This problem often occurs when the subject is a material that absorbs infrared rays. For example, human skin is likely to be imaged whiter than it actually is, and warm colors such as yellow may be imaged blue.

[0016] One of the problems is to provide an imaging method and an imaging system that have high visibility even in imaging at night or in a dark place and can obtain an image faithful to the same color as when there is external light.

Means for Solving the Problem

[0017] The imaging system of the present invention includes a solid-state imaging device that does not have a color filter, a storage device, and a learning device. Since a color filter is not used, light attenuation can be avoided, and imaging with high sensitivity can be performed even with a small amount of light.

[0018] Since it does not have a color filter, colorization is performed on the obtained black-and-white image data (analog data). Color is applied using an AI system. An inference is made using the extracted feature amounts by the AI system, that is, a learning device that uses teacher data stored in a storage device, and focusing is adjusted to obtain a highly visible color image (colorized image data (digital data)) even in imaging at night or in a dark place. Note that the learning device includes at least a neural network unit, performs not only learning but also inference, and can output data. Also, the learning device may perform inference using pre-learned feature amounts. In that case, the pre-learned feature amounts are stored in the storage device, and by performing calculations, data output can be performed at a level comparable to the case where pre-learned feature amounts are not used.

[0019] Also, when a part of the outline of the subject becomes unclear due to insufficient light amount to the subject, the boundary cannot be discriminated, and there is a risk that the color application to that part becomes incomplete.

[0020] Therefore, it is preferable to use an image sensor that does not use a color filter, acquire black-and-white image data with a wide dynamic range, repeatedly perform super-resolution processing a plurality of times, then discriminate color boundaries and apply color. Also, super-resolution processing may be performed on the teacher data at least once. By mixing not only color photographic images but also color illustration (animation) images as teacher data for creating a learning model for discriminating color boundaries, colorized image data with clear color boundaries can be obtained. Note that super-resolution processing refers to image processing for generating a high-resolution image from a low-resolution image.

[0021] Also, if a learning model is prepared in advance, relatively bright black-and-white image data can be acquired by imaging without using a flash light source in a situation where the light amount is insufficient, and by colorizing based on the black-and-white image data, vividly colorized image data can be obtained.

[0022] The video surveillance system, security system, or security information providing system using the above imaging system can clearly achieve imaging in a relatively dark place.

[0023] Specifically, it is a surveillance system equipped with a security camera. The security camera has a solid-state imaging device without a color filter, a learning device, and a storage device. While the security camera is detecting a person, the solid-state imaging device captures an image, and a software program is executed to create colorized image data by inference of the learning device using the teacher data in the storage device.

Advantages of the Invention

[0024] With the imaging system disclosed in this specification, even in shooting situations with low light levels and dim conditions, clear colorized images can be obtained. Therefore, based on the obtained colorized images, it is relatively easy to identify a person (such as the face) or the characteristics of clothing. By applying this imaging system to a security camera, the face of a person can also be estimated in color video and displayed on a display device.

[0025] In particular, when obtaining input image data of 8K size with a solid-state imaging device, the light-receiving area of the solid-state imaging device arranged for each pixel becomes narrow, resulting in a decrease in the amount of light obtained. However, in the imaging system disclosed in this specification, since a color filter is not used in the solid-state imaging device, there is no reduction in the amount of light due to the color filter. As a result, 8K-sized image data can be captured with good sensitivity.

Brief Description of the Drawings

[0026]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Embodiments for Carrying Out the Invention

[0027] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the present invention is not limited to the following description, and it will be easily understood by those skilled in the art that its forms and details can be variously changed. Also, the present invention is not to be construed as being limited to the description of the embodiments shown below.

[0028] (Embodiment 1) An example of the configuration of an imaging system 21 used in a video surveillance system or a security system will be described with reference to the block diagram shown in FIG. 1.

[0029] The data acquisition device 10 is a semiconductor chip including a solid-state imaging device 11 and an analog arithmetic circuit 12, and does not have a color filter. The data acquisition device 10 has an optical system such as a lens. Note that the optical system may have any configuration as long as its imaging characteristics are known, and is not particularly limited.

[0030] The A / D circuit 13 (also called an A / D converter) represents an analog-to-digital conversion circuit and converts the analog data output from the data acquisition device 10 into digital data. If necessary, an amplification circuit may be provided between the data acquisition device 10 and the A / D circuit 13 to amplify the analog signal before converting it into digital data.

[0031] The memory unit 14 is a circuit that stores the converted digital data and is configured to store the data before inputting it to the neural network unit 16, but is not particularly limited to this configuration. Depending on the data amount output from the data acquisition device or the data processing capacity of the image processing device, for small-scale data, the output from the A / D circuit 13 may be directly input to the neural network unit 16 without being stored in the memory unit 14. Also, the output from the A / D circuit 13 may be configured to be input to the neural network unit 16 located remotely using Internet communication. For example, the neural network unit 16 may be constructed on a server capable of two-way communication.

[0032] The image processing device 20 is a device for estimating the contour or color corresponding to the black-and-white image obtained by the data acquisition device 10. The image processing device 20 is executed separately in a first stage for learning and a second stage for estimation. In this embodiment, the data acquisition device 10 and the image processing device 20 are configured as separate devices, but it is also possible to integrally configure the data acquisition device 10 and the image processing device 20. In the case of an integrated configuration, it is also possible to update the feature amount obtained by the neural network unit in real time.

[0033] The neural network unit 16 is realized by software operations performed by a microcontroller. A microcontroller is a computer system integrated into a single integrated circuit (IC). When the scale of operations or the data to be processed is large, the neural network unit 16 may be configured by combining multiple ICs. The learning device at least includes these multiple ICs. Also, since free software can be used if the microcontroller is equipped with Linux (registered trademark), it is preferable because the total cost for configuring the neural network unit 16 can be reduced. Also, it is not limited to Linux (registered trademark), and other operating systems (OS) may be used.

[0034] The learning of the neural network unit 16 shown in FIG. 1 is described below. In advance, the training teacher data is stored in the storage unit 18. The training teacher data may be obtained using a training set or the like, and the types of images include landscape photos, portrait photos, illustrations, and the like. Using this training teacher data, the learning device makes inferences. The learning device may have any configuration as long as it can output colorized image data by making inferences using the teacher data in the storage unit 18 based on black-and-white analog data. Also, when using pre-trained feature amounts, the learning device may have any configuration as long as it can output colorized image data by performing operations using the feature amount data in the storage unit 18 based on black-and-white analog data. When using pre-trained feature amounts, there is an advantage that a small-scale configuration, for example, a learning device can be configured with one or two ICs, because the amount of data and the amount of operations are reduced.

[0035] Create a program using Python in the operating environment of Linux (registered trademark). In this embodiment, the data frame of Keras is used. Keras is a library that provides functions convenient for deep learning. In particular, it is easy to read and write data to the intermediate layer of the neural network part and change the weight coefficients of the neural network part. Also, to handle the Numpy library on Python, load Numpy. Also, load openCV for image editing.

[0036] The source code corresponding to the above-described processing is shown in order below the line of In[1] in Figure 2.

[0037] Next, read the data. As the data, use color images. It is preferable to prepare thousands or tens of thousands of color images. When the number of files of the color images is huge, since a computational load is imposed in the reading process, for example, it may be read after converting to the h5py format.

[0038] For colorization, the following convolutional neural network (also called an artificial neural network) can be used. For example, in Python, using keras, it can be described as follows under the line of In[X] in Figure 3 (where X is an arbitrary number and can be changed according to the program). Such a neural network is also called U-net. U-net has a data complementation function from the convolutional layer to the deconvolutional layer called skip connection, can prevent the vanishing gradient, and can construct a good learning model.

[0039] During learning, for the color image of the teacher data, input the image grayscaled image into the neural network shown in In[X] of Figure 3. For the process of grayscaling the image, for example, openCV can be used.

[0040] You may use GaN (Generative Adversarial Networks) with the above neural network as the Generator.

[0041] The output data of the neural network unit 16 and the time data of the time information acquisition device 17 are associated and stored in the large-scale storage device 15. In the large-scale storage device 15, the data obtained since the start of imaging is accumulated and stored.

[0042] The display unit 19 (including video display with time display, etc.) may be provided with an operation input unit such as a touch panel, and the user can select from the data stored in the large-scale storage device 15 and observe it as appropriate. Also, the display unit 19 may be enabled to access the large-scale storage device 15 by remote operation via Internet communication, and the large-scale storage device 15 may be provided with a transmission antenna or a reception antenna. The imaging system 21 can be used in a video surveillance system or a security system.

[0043] Also, the display unit of the user's portable information terminal (such as a smartphone) can be used as the display unit 19. By accessing the large-scale storage device 15 from the display unit of the portable information terminal, it is possible to monitor regardless of the user's location.

[0044] The installation of the imaging system 21 is not limited to the wall of the room. By mounting all or part of the configuration of the imaging system 21 on an unmanned aircraft (also called a drone) equipped with a rotary wing, it is also possible to perform video surveillance from the air. In particular, imaging can be performed in an environment with a small amount of light, such as in the evening or at night when streetlights are not lit.

[0045] In addition, in this embodiment, although the video surveillance system or the security system has been described, it is not particularly limited. By combining a camera or radar that images the periphery of the vehicle with an ECU (Electronic Control Unit) that performs image processing or the like, it can also be applied to a vehicle capable of semi-automatic driving or a vehicle capable of fully automatic driving. A vehicle using an electric motor has a plurality of ECUs, and the ECU performs engine control and the like. The ECU includes a microcomputer. The ECU is connected to a CAN (Controller Area Network) provided in the electric vehicle. CAN is one of the serial communication standards used as an in-vehicle LAN. The ECU uses a CPU or a GPU. For example, as one of a plurality of cameras (such as a camera for a drive recorder and a rear camera) mounted on an electric vehicle, a solid-state imaging device without a color filter is used, and the obtained black-and-white image is inferred by the ECU via the CAN, and a colorized image is created and configured to be displayed on a display device in the vehicle or a display unit of a portable information terminal.

[0046] (Embodiment 2) In this embodiment, an example of a flow for colorizing a black-and-white video obtained by the solid-state imaging device 11 is shown in FIG. 4 using the block diagram and program shown in Embodiment 1.

[0047] Install the imaging system 21 shown in Embodiment 1 at a location to be monitored (such as a room, a parking lot, or an entrance), start it, and start continuous shooting.

[0048] First, prepare to acquire data (S1).

[0049] Acquire black-and-white image data using a solid-state imaging device without a color filter (S2). Note that a plurality of solid-state imaging devices arranged in a matrix direction may also be called a pixel array.

[0050] Next, filter the obtained analog data using a sum-of-products operation circuit (S3).

[0051] Steps S2 and S3 are performed by the imaging device shown in FIG. 5. The imaging device will be described below.

[0052] FIG. 5 is a block diagram for explaining the imaging device. The imaging device includes a pixel array 300, a circuit 201, a circuit 301, a circuit 302, a circuit 303, a circuit 304, a circuit 305, and a circuit 306. Note that each of the circuit 201 and the circuits 301 to 306 is not limited to a single circuit configuration, and may be configured by a combination of a plurality of circuits. Alternatively, any one or more of the above circuits may be integrated. Further, circuits other than the above may be connected.

[0053] The pixel array 300 has an imaging function and an arithmetic function. The circuits 201 and 301 have an arithmetic function. The circuit 302 has an arithmetic function or a data conversion function. The circuits 303, 304, and 306 have a selection function. The circuit 303 is electrically connected to the pixel block 200 via the wiring 124. The circuit 304 is electrically connected to the pixel block 200 via the wiring 123. The circuit 305 has a function of supplying a potential for the multiplication and accumulation operation to the pixels. For the circuits having a selection function, a shift register or a decoder or the like can be used. The circuit 306 is electrically connected to the pixel block 200 via the wiring 113. Note that the circuits 301 and 302 may be provided outside.

[0054] The pixel array 300 includes a plurality of pixel blocks 200. As shown in FIG. 6, each pixel block 200 includes a plurality of pixels 100 arranged in a matrix, and each pixel 100 is electrically connected to the circuit 201 via the wiring 112. Note that the circuit 201 can also be provided inside the pixel block 200.

[0055] Further, the pixel 100 is electrically connected to an adjacent pixel 100 via a transistor 150 (transistors 150g to 150j). The function of the transistor 150 will be described later.

[0056] In pixel 100, it is possible to acquire image data and generate data obtained by adding the image data and a weighting coefficient. Note that in FIG. 6, as an example, the number of pixels included in pixel block 200 is 3×3, but it is not limited to this. For example, it can be 2×2, 4×4, etc. Alternatively, the number of pixels in the horizontal direction and the vertical direction may be different. Also, some pixels may be shared by adjacent pixel blocks.

[0057] Pixel block 200 and circuit 201 can be operated as a multiplication and accumulation circuit.

[0058] As shown in FIG. 7, pixel 100 can include a photoelectric conversion device 101, a transistor 102, a transistor 103, a transistor 104, a transistor 105, a transistor 106, and a capacitor 107.

[0059] One electrode of the photoelectric conversion device 101 is electrically connected to one of the source or drain of the transistor 102. The other of the source or drain of the transistor 102 is electrically connected to one of the source or drain of the transistor 103, the gate of the transistor 104, and one electrode of the capacitor 107. One of the source or drain of the transistor 104 is electrically connected to one of the source or drain of the transistor 105. The other electrode of the capacitor 107 is electrically connected to one of the source or drain of the transistor 106.

[0060] The other electrode of the photoelectric conversion device 101 is electrically connected to the wiring 114. The other of the source or drain of the transistor 103 is electrically connected to the wiring 115. The other of the source or drain of the transistor 105 is electrically connected to the wiring 112. The other of the source or drain of the transistor 104 is electrically connected to a GND wiring or the like. The other of the source or drain of the transistor 106 is electrically connected to the wiring 111. The other electrode of the capacitor 107 is electrically connected to the wiring 117.

[0061] The gate of transistor 102 is electrically connected to wiring 121. The gate of transistor 103 is electrically connected to wiring 122. The gate of transistor 105 is electrically connected to wiring 123. The gate of transistor 106 is electrically connected to wiring 124.

[0062] Here, an electrical connection point between the other of the source or drain of transistor 102, one of the source or drain of transistor 103, one electrode of capacitor 107, and the gate of transistor 104 is defined as node FD. Also, an electrical connection point between the other electrode of capacitor 107 and one of the source or drain of transistor 106 is defined as node FDW.

[0063] Wiring 114 and 115 can function as power supply lines. For example, wiring 114 can function as a high-potential power supply line and wiring 115 can function as a low-potential power supply line. Wiring 121, 122, 123, and 124 can function as signal lines for controlling the conduction of each transistor. Wiring 111 can function as a wiring for supplying a potential corresponding to a weighting factor to pixel 100. Wiring 112 can function as a wiring for electrically connecting pixel 100 and circuit 201. Wiring 117 can function as a wiring for electrically connecting the other electrode of capacitor 107 of the pixel and the other electrode of capacitor 107 of another pixel via transistor 150 (see FIG. 6).

[0064] Note that an amplification circuit or a gain adjustment circuit may be electrically connected to wiring 112.

[0065] As the photoelectric conversion device 101, a photodiode can be used. Regardless of the type of photodiode, an Si photodiode having silicon in the photoelectric conversion layer, an organic photodiode having an organic photoconductive film in the photoelectric conversion layer, etc. can be used. When it is desired to enhance the light detection sensitivity at low illuminance levels, it is preferable to use an avalanche photodiode.

[0066] Transistor 102 can have a function of controlling the potential of node FD. Transistor 103 can have a function of initializing the potential of node FD. Transistor 104 can have a function of controlling the current flowing through circuit 201 according to the potential of node FD. Transistor 105 can have a function of selecting a pixel. Transistor 106 can have a function of supplying a potential corresponding to a weighting factor to node FDW.

[0067] When an avalanche photodiode is used for the photoelectric conversion device 101, a high voltage may be applied, and it is preferable to use a high-voltage withstand transistor for the transistor connected to the photoelectric conversion device 101. As the high-voltage withstand transistor, for example, a transistor using a metal oxide in the channel formation region (hereinafter, OS transistor) can be used. Specifically, it is preferable to apply an OS transistor to transistor 102.

[0068] Also, the OS transistor also has a characteristic of extremely low off-current. By using the OS transistor for transistors 102, 103, and 106, the period during which charge can be held at node FD and node FDW can be made extremely long. Therefore, a global shutter method that performs a charge accumulation operation simultaneously for all pixels can be applied without complicating the circuit configuration or the operation method. Also, while holding image data at node FD, a plurality of operations using the image data can be performed.

[0069] On the other hand, in some cases, it may be desired that transistor 104 has excellent amplification characteristics. Also, in some cases, it may be preferable to use a transistor with high mobility that enables high-speed operation for transistor 106. Therefore, a transistor using silicon in the channel formation region (hereinafter, Si transistor) may be applied to transistors 104 and 106.

[0070] Note that, not limited to the above, the OS transistors and Si transistors may be arbitrarily combined and applied. Also, all the transistors may be OS transistors. Or, all the transistors may be Si transistors. Examples of Si transistors include transistors having amorphous silicon, transistors having crystalline silicon (microcrystalline silicon, low-temperature polysilicon, single-crystalline silicon), and the like.

[0071] The potential of the node FD in the pixel 100 is determined by the potential obtained by adding the reset potential supplied from the wiring 115 and the potential (image data) generated by the photoelectric conversion by the photoelectric conversion device 101. Or, further, the potential corresponding to the weight coefficient supplied from the wiring 111 is capacitively coupled and determined. Therefore, the transistor 105 can pass a current corresponding to the data obtained by adding an arbitrary weight coefficient to the image data.

[0072] Note that the above is an example of the circuit configuration of the pixel 100, and the photoelectric conversion operation can also be performed with other circuit configurations.

[0073] As shown in FIG. 6, the pixels 100 are electrically connected to each other by the wiring 112. The circuit 201 can perform an operation using the sum of the currents flowing through the transistors 104 of the pixels 100.

[0074] The circuit 201 includes a capacitor 202, a transistor 203, a transistor 204, a transistor 205, a transistor 206, and a resistor 207.

[0075] One electrode of the capacitor 202 is electrically connected to one of the source or drain of the transistor 203. One of the source or drain of the transistor 203 is electrically connected to the gate of the transistor 204. One of the source or drain of the transistor 204 is electrically connected to one of the source or drain of the transistor 205. One of the source or drain of the transistor 205 is electrically connected to one of the source or drain of the transistor 206. One electrode of the resistor 207 is electrically connected to the other electrode of the capacitor 202.

[0076] The other electrode of the capacitor 202 is electrically connected to the wiring 112. The other of the source or drain of the transistor 203 is electrically connected to the wiring 218. The other of the source or drain of the transistor 204 is electrically connected to the wiring 219. The other of the source or drain of the transistor 205 is electrically connected to a reference power line such as a GND wiring. The other of the source or drain of the transistor 206 is electrically connected to the wiring 212. The other electrode of the resistor 207 is electrically connected to the wiring 217.

[0077] The wirings 217, 218, and 219 can have the function as a power line. For example, the wiring 218 can have the function as a wiring for supplying a dedicated potential for reading. The wirings 217 and 219 can function as high-potential power lines. The wirings 213, 215, and 216 can function as signal lines for controlling the conduction of each transistor. The wiring 212 is an output line and can be electrically connected to, for example, the circuit 301 shown in FIG. 5.

[0078] Transistor 203 can have a function of resetting the potential of wiring 211 to the potential of wiring 218. Wiring 211 is a wiring connected to one electrode of capacitor 202, one of the source or drain of transistor 203, and the gate of transistor 204. Transistors 204 and 205 can have a function as a source follower circuit. Transistor 206 can have a function of controlling reading. Note that circuit 201 has a function as a correlated double sampling circuit (CDS circuit), and can also be replaced with a circuit of another configuration having the same function.

[0079] In one aspect of the present invention, an offset component other than the product of image data (X) and weight coefficient (W) is removed, and the target WX is extracted. WX can be calculated using data with and without imaging for the same pixel, and data when weights are added to each of them.

[0080] The total current (I p ) flowing through pixel 100 when imaging is performed is kΣ(X - V th ) 2 , and the total current (I p ) flowing through pixel 100 when weights are added is kΣ(W + X - V th ) 2 . Also, the total current (I ref ) flowing through pixel 100 when no imaging is performed is kΣ(0 - V th ) 2 , and the total current (I ref ) flowing through pixel 100 when weights are added is kΣ(W - V th ) 2 . Here, k is a constant, and V th is the threshold voltage of transistor 105.

[0081] First, the difference (data A) between the data with imaging and the data with weights added to the data is calculated. kΣ((X - V th ) 2 - (W + X - V th ) 2 ) = kΣ(-W 2 - 2W·X + 2W·V th) is obtained.

[0082] Next, the difference (data B) between the data without imaging and the data with weights added to the data is calculated. kΣ((0 - V th ) 2 - (W - V th ) 2 ) = kΣ(-W 2 + 2W·V th ) is obtained.

[0083] Then, the difference between data A and data B is taken. kΣ(-W 2 - 2W·X + 2W·V th - (-W 2 + 2W·V th )) = kΣ(-2W·X) is obtained. That is, the offset components other than the product of the image data (X) and the weight coefficient (W) can be removed.

[0084] In circuit 201, data A and data B can be read out. Note that the difference operation between data A and data B can be performed, for example, in circuit 301.

[0085] Here, the weight supplied to the entire pixel block 200 functions as a filter. As such a filter, for example, a convolution filter of a convolutional neural network (CNN) can be used. Alternatively, an image processing filter such as an edge extraction filter can be used. Examples of the edge extraction filter include, for example, the Laplacian filter shown in FIG. 8A, the Prewitt filter shown in FIG. 8B, the Sobel filter shown in FIG. 8C, etc.

[0086] When the number of pixels 100 included in the pixel block 200 is 3×3, the elements of the edge extraction filter can be supplied to each pixel 100 as weights. As described above, in order to calculate the data A and the data B, it is possible to calculate using the data with and without imaging and the data when weights are added to each of them. Here, the data with and without imaging is data without adding weights, and can also be paraphrased as data obtained by adding a weight of 0 to all the pixels 100.

[0087] The edge extraction filter illustrated in FIGS. 8A to 8C is a filter in which the sum of the elements (weights: ΔW) of the filter (ΣΔW / N, where N is the number of elements) is 0. Therefore, even if an operation of supplying ΔW = 0 from another circuit is not newly performed, by performing an operation of obtaining ΣΔW / N, it is possible to obtain data in which an equivalent of ΔW = 0 is added to all the pixels 100.

[0088] This operation corresponds to turning on the transistors 150 (transistors 150a to 150f) provided between the pixels 100 (see FIG. 6). By turning on the transistors 150, the nodes FDW of each pixel 100 are all short-circuited via the wiring 117. At this time, the charges accumulated in the nodes FDW of each pixel 100 are redistributed, and when the edge extraction filter illustrated in FIGS. 8A to 8C is used, the potential (ΔW) of the node FDW becomes 0 or substantially 0. Therefore, it is possible to obtain data in which an equivalent of ΔW = 0 is added.

[0089] When rewriting the weight (ΔW) by supplying charges from a circuit outside the pixel array 300, it takes time until the rewriting is completed due to the capacitance of the long wiring 111 and the like. On the other hand, the pixel block 200 is a minute area, and the distance of the wiring 117 is short and the capacitance is small. Therefore, in the operation of redistributing the charges accumulated in the nodes FDW in the pixel block 200, the weight (ΔW) can be rewritten at high speed.

[0090] In the pixel block 200 shown in FIG. 6, the transistors 150a to 150f are each shown to be electrically connected to different gate lines (wiring 113a to 113f). In this configuration, the conduction of the transistors 150a to 150f can be independently controlled, and the operation of obtaining ΣΔW / N can be selectively performed. Also, FIG. 6 shows a configuration in which the transistors 150g to 150j are each electrically connected to different gate lines (113g to 113i).

[0091] For example, when using the filters shown in FIGS. 8B and 8C, there are pixels where ΔW = 0 is initially supplied. Assuming that ΣΔW / N = 0, the pixels where ΔW = 0 may be excluded from the pixels to be summed. By excluding such pixels, the supply of potential for operating a part of the transistors 150a to 150f becomes unnecessary, so that the power consumption can be suppressed. Note that in FIGS. 6 and 8, an example is shown in which nine transistors 150 (transistors 150a to 150f) are provided between the pixels 100, but the number of transistors 150 may be further increased. Also, in the transistors 150g to 150j, some transistors may be omitted so as to eliminate parallel paths.

[0092] The data of the sum-of-products operation result output from the circuit 201 is sequentially input to the circuit 301. The circuit 301 may have various arithmetic functions in addition to the function of calculating the difference between the aforementioned data A and data B. For example, the circuit 301 can have the same configuration as the circuit 201. Or, the function of the circuit 301 may be replaced by software processing.

[0093] Also, the circuit 301 may have a circuit that performs the operation of an activation function. For example, a comparator circuit can be used for such a circuit. In the comparator circuit, the result of comparing the input data with a set threshold value is output as binary data. That is, the pixel block 200 and the circuit 301 can act as part of the elements of a neural network.

[0094] The data output from circuit 301 is sequentially input to circuit 302. Circuit 302 can be configured to have, for example, a latch circuit and a shift register. With this configuration, parallel-to-serial conversion can be performed, and the data input in parallel can be output as serial data to wiring 311.

[0095] The connection destination of wiring 311 is not limited. For example, it can be connected to the A / D circuit 13 or the neural network unit 16 shown in FIG. 1, etc. Also, the connection destination of wiring 311 may be an FPGA (field-programmable gate array).

[0096] Next, the analog data after filtering is converted into digital data by the A / D circuit 13 (S4).

[0097] Next, the converted digital data is stored in the memory unit 14 (digital memory unit) (S5).

[0098] Next, the digital data is converted into a signal format (such as JPEG (registered trademark)) required by the subsequent inference program (S6).

[0099] Next, the converted digital data is subjected to convolution processing using a CPU or the like, and inferences such as contours and colors are made and colorized (S7). Instead of the CPU, one IC chip integrated with a GPU (Graphics Processing Unit), a PMU (Power Management Unit), etc. may be used. And colorized image data is output (S8). And the colorized image data is stored together with time data such as the date and time (S9). The storage is accumulated in the large-scale storage device 15, a so-called large-capacity storage device (hard disk, etc.) or a database.

[0100] The acquisition of the above colorized image data is repeated (during operation). By repeating, colorization can also be performed in real time.

[0101] The colorized image data thus obtained is based on a black-and-white image with a wide dynamic range using an image sensor without a color filter. Therefore, even when the amount of light is small and indistinguishable with a conventional image sensor with a color filter, distinguishable colorized image data can be obtained. The imaging system shown in this embodiment can be realized on one or more computers for each of the above-described steps (S4 to S9).

[0102] Also, as a means for real-time colorization, the latent variable (feature amount) is monitored by cosine similarity, and focus adjustment is performed by an optical system so that fluctuations are reduced. Even if the object moves during shooting, it is possible to adjust the focus so that the object is in focus.

[0103] Also, inference may be performed using the extracted feature amounts. In the result of the inference, focus adjustment may be performed so that fluctuations are reduced. For example, when a person is inferred, focus adjustment may be performed so that the likelihood is constant or large.

[0104] By performing inference, it is possible to immediately determine what the captured object is. Therefore, for example, in a security application, if the object is determined to be dangerous, it is also possible to notify a mobile information terminal such as a smartphone of the object to the necessary contact at that time. Also, even if the focus is out of focus, it is possible to infer an image that has been sharpened by removing image blurring.

[0105] (Embodiment 3) In this embodiment, an example is shown that enables smoother image processing or finer coloring processing compared to the colorized image data obtained in Embodiment 2.

[0106] A flowchart is shown in FIG. 9. The same reference numerals are used for the same steps as the flowchart shown in FIG. 4 of Embodiment 2. Since S1 to S6 and S8 to S9 in FIG. 4 are the same, detailed description thereof will be omitted here.

[0107] As shown in FIG. 9, after step S6, super-resolution processing is performed multiple times on the converted digital data using a first learning model to infer the contour (S7a).

[0108] Then, the digital data after the super-resolution processing is used to infer colors and the like using a second learning model, and colorization is performed (S7b). The subsequent steps are the same as those in Embodiment 2.

[0109] In the teacher data of the second learning model, multiple super-resolution processes are performed in advance, or an animation image is mixed with a photographic image, or the contour is emphasized using the OPENCV drawcontours function or the like. The ratio of mixing the animation image with the photographic image is such that when the photographic image is 2, the animation image is 1. The animation image is a type of illustration but contains many edge components or color components. In the process of colorizing a black-and-white image, edges are extracted as feature amounts in the convolutional layer, and the colors of each region of the image are inferred based on the feature amounts. Therefore, using an image containing many edge components as teacher data is effective for improving the efficiency of machine learning. By using the animation image as teacher data, the number of teacher data required to obtain a color image that can reach a certain standard can be reduced, the time required for machine learning can be shortened, and the configuration of the neural network part can be simplified. Note that the neural network part is a part of machine learning. Also, deep learning is a part of the neural network part.

[0110] The learning of the neural network part of this embodiment is shown below.

[0111] Create a program using Python in the operating environment of Linux (registered trademark). In this embodiment, the data frame of Keras is used. Keras is a library that provides functions convenient for deep learning. In particular, it is easy to read and write data to the intermediate layer of the neural network and to change the weight coefficients of the neural network. Also, to handle the Numpy library on Python, Numpy is loaded. Also, as image processing, scipy is used.

[0112] The source code corresponding to the above-described processing is shown in order below the line In[Y1] in FIG. 10 (where Y1 is an arbitrary number and can be changed according to the program).

[0113] Next, read the data. In the case of this embodiment, the teacher data is a high-resolution image. Teacher data is a set of data used in supervised learning or classified data. As the above image, a color image or a black-and-white image can be used. It is preferable to prepare thousands or tens of thousands of images. If the number of files of the images is huge, since a computational load is imposed in the read process, for example, it may be read after converting to the h5py format.

[0114] The super-resolution process consists of a three-layer convolutional neural network. For example, in Python, using the Sequential model of Keras, it can be described as follows under the line of In[Y2] in Figure 10 (where Y2 is an arbitrary number and can be changed according to the program as long as it is after Y1). Note that an example of inputting an image of 33×33 pixels is shown in In[Y2]. In this embodiment, an example of an image of 33×33 pixels is shown, but it is not particularly limited, and it may be an image of 2K size (1920×1080 pixels) or an image of 4K size (3840×2160 pixels). The size refers to the resolution, and the neural network can be designed according to the input image data or output image data. Also, the input image data and output image data may be different. For example, after obtaining input image data of QHD size (960×540 pixels), it may be used as an output image of 2K size. When obtaining input image data of 8K size with a solid-state imaging device, since the amount of light obtained by each solid-state imaging device is reduced, it is particularly preferable not to use a color filter in order to obtain more light.

[0115] During training, for the image of the teacher data, an image with a reduced resolution of the image is input into the neural network shown in In[Y3]. The process of reducing the resolution can be described, for example, using scipy, as follows under the line of In[Y3] in Figure 10 (where Y3 is an arbitrary number and can be changed according to the program as long as it is after Y2). Here, a process of reducing the resolution to 1 / 3 is shown.

[0116] Using such teacher data and code, a model capable of outputting an image can be created. During inference, using this model, an image with a low resolution can be input and an image with a high resolution can be output. The imaging system shown in this embodiment can be realized on one or more computers for each of the above steps (S4~S9).

[0117] GaN using the above neural network as a Generator may also be used.

[0118] The colorized image data obtained in this embodiment has smoother contours than the image data of Embodiment 2, and optimal colorization is performed.

[0119] Also, when preparing a teacher image independently, for example, when wanting to colorize a rare fish, or when using a teacher image of a similar fish and only having a teacher image with blurred contours, this teacher image can effectively train the colorization model.

[0120] (Embodiment 4) In this embodiment, as an electronic device that can use the imaging device used in the imaging system of one aspect of the present invention, a display device, a personal computer, an image storage device or an image playback device equipped with a recording medium, a mobile phone, a game machine including a portable type, a portable data terminal, an e-book terminal, a video camera, a camera such as a digital still camera, a goggle-type display (head-mounted display), a navigation system, an audio playback device (car audio, digital audio player, etc.), a copying machine, a facsimile machine, a printer, a printer multifunction machine, a cash dispenser (ATM), a vending machine, etc. can be mentioned. Specific examples of these electronic devices are shown in FIG. 11.

[0121] FIG. 11A is a surveillance camera, which has a housing 951, a lens 952, a support portion 953, etc. In order to acquire an image in the surveillance camera, a photographing system according to an aspect of the present invention can be provided. A neural network portion is provided in the housing 951. Note that the surveillance camera is a common name and does not limit the use. For example, a device having a function as a surveillance camera is also called a camera or a video camera. The surveillance camera uses an image sensor that does not use a color filter. Further, by incorporating the program shown in Embodiment 2 or Embodiment 3 as a software program and executing it in the neural network portion, colorized image data can be created. When a plurality of surveillance cameras are used, if at least one of them is the surveillance camera of the present embodiment, a color image in a dim environment, which is difficult to acquire with a conventional surveillance camera, can be acquired. Thus, the surveillance system can be enhanced by combining it with a conventional surveillance camera.

[0122] FIG. 11B is also a surveillance camera, which has a support base 954, a camera unit 955, a protective cover 956, etc. A rotation mechanism or the like is provided in the camera unit 955, and by installing it on the ceiling, imaging of the entire circumference becomes possible. The camera unit 955 can be used as an imaging device included in a surveillance system according to an aspect of the present invention. Further, by estimating with the neural network portion of the camera unit 955 based on the data obtained by the camera unit 955, a suspicious person can be identified from the information imaged by colorization or super-resolution.

[0123] FIG. 11C shows an example of an aircraft. The aircraft 6500 shown in FIG. 11C has a propeller 6501, a camera 6502, a battery 6503, etc., and has a function of flying autonomously.

[0124] For example, the image data captured by the camera 6502 is stored in the electronic component 6504. The electronic component 6504 can analyze the image data and detect the presence or absence of obstacles when moving. As the camera 6502, imaging devices of multiple types of systems may be used. As the imaging device included in the monitoring system according to one aspect of the present invention, the camera 6502 can be used. Further, by estimating with the neural network unit based on the data obtained by the camera 6502, a suspicious person can be identified from the information captured by colorization or super-resolution.

[0125] The configurations, structures, methods, etc. shown in the present embodiment can be used in appropriate combination with the configurations, structures, methods, etc. shown in other embodiments, etc.

Explanation of Signs

[0126] 10: Data acquisition device, 11: Solid-state imaging device, 12: Analog arithmetic circuit, 13: A / D circuit, 14: Memory section, 15: Mass storage device, 16: Neural network section, 17: Time information acquisition device, 18: Memory section, 19: Display section, 20: Image processing device, 21: Imaging system, 100: Pixel, 101: Photoelectric conversion device, 102: Transistor, 103: Transistor, 104: Transistor, 105: Transistor, 106: Transistor, 107: Capacitor, 111: Wiring, 112: Wiring, 113a: Wiring, 113f: Wiring, 114: Wiring, 115: Wiring, 117: Wiring, 121: Wiring, 122: Wiring, 123: Wiring, 124: Wiring, 150: Transistor, 150g: Transistor, 150h: Transistor, 150i: Transistor, 150j: Transistor, 200: Pixel block, 201: Circuit, 202: Capacitor, 203: Transistor, 204: Transistor, 205: Transistor, 206: Transistor, 207: Resistor, 211: Wiring, 212: Wiring, 213: Wiring, 215: Wiring, 216: Wiring, 217: Wiring, 218: Wiring, 219: Wiring, 300: Pixel array, 301: Circuit, 302: Circuit, 303: Circuit, 304: Circuit, 305: Circuit, 306: Circuit, 311: Wiring, 951: Housing, 952: Lens, 953: Support section, 954: Support base, 955: Camera unit, 956: Protective cover, 6500: Aircraft, 6501: Propeller, 6502: Camera, 6503: Battery, 6504: Electronic component

Claims

[Claim 1] An imaging system including a solid-state imaging element having no color filter, a storage device, and a learning device, The solid-state imaging device acquires black and white image data; An imaging system in which the learning device colorizes the black-and-white image data using teacher data stored in the storage device, thereby creating colorized image data.

Citation Information

Patent Citations

  • Human face area detection device

    JP1995282227A

  • Intruder detecting device

    JP1998083487A

  • Monitoring recording device

    JP2006178516A

  • Image processing apparatus,monitoring center, monitoring system, image processing method, and image processing program

    JP2007208481A

  • Color information estimation model generating device, moving image colorization device, and programs for the same

    JP2019117559A