Solid-state imaging device, electronic device, and imaging system

The solid-state imaging device addresses the challenge of improving neural network processing accuracy by optimizing image acquisition and neural network model training within the device, enhancing recognition rates and reducing costs while ensuring privacy and security.

JP7680969B2Active Publication Date: 2025-05-21SONY SEMICON SOLUTIONS CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021574438
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-30
Filing Date
2020-05-01
Publication Date
2025-05-21
Estimated Expiration
2040-05-01

AI Technical Summary

Technical Problem

Existing solid-state imaging devices face challenges in improving the accuracy of neural network processing for image data captured by image sensors, particularly in maintaining high recognition rates under varying environmental conditions.

Method used

A solid-state imaging device is designed with a pixel array, a converter, an image processing unit, a digital signal processing unit, and a control unit that optimizes the acquisition processes of analog pixel signals, digital image data, and recognition processing results based on feedback from the recognition process. This includes controlling exposure time, optimizing image processing parameters, and retraining neural network models.

Benefits of technology

The solution enhances the accuracy of neural network processing, improves recognition rates, and reduces system costs by optimizing image acquisition conditions and neural network models within the imaging device, while also ensuring privacy and security by processing data internally.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007680969000001
    Figure 0007680969000001
  • Figure 0007680969000002
    Figure 0007680969000002
  • Figure 0007680969000003
    Figure 0007680969000003
Patent Text Reader

Abstract

The present invention improves accuracy of recognition processing used in an image sensor. This solid-state imaging device is provided with a pixel array, a converter, an image processing unit, a digital signal processing unit, and a control unit. The pixel array has a plurality of pixels for performing photoelectric conversion. The converter converts an analog pixel signal outputted from the pixel array to digital image data. The image processing unit performs image processing of the digital image data. The digital signal processing unit performs recognition processing on the digital image data outputted by the image processing unit. The control unit performs, on the basis of the result of the recognition processing, optimization related to at least one acquisition processing among acquisition of the analog pixel signal, acquisition of the digital image data, and acquisition of the result of the recognition processing.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application is a related application to U.S. Provisional Patent Application No. 62 / 967,869, filed January 30, 2020, and claims priority to that provisional application, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates to a solid-state imaging device, an electronic device, and an imaging system. [Background technology]

[0003] In recent years, there has been a demand for high-speed signal processing of image data captured by an image sensor. In addition, with the advancement of semiconductor process technology, a semiconductor device has been proposed in which a plurality of chips, such as an image sensor chip, a memory chip, and a signal processing chip, are connected to each other via bumps and packaged, and a die on which an image sensor is arranged and a die on which a memory, a signal processing circuit, and the like are arranged are stacked and packaged.

[0004] When a semiconductor device (hereinafter referred to as an imaging device) incorporating an image sensor and a signal processing circuit is mounted on an electronic device such as a smartphone, the signal processing circuit in the imaging device often performs various signal processing according to instructions from an application processor mounted on the electronic device. For example, by executing neural network processing in the signal processing circuit and outputting the results, it is possible to realize the latency of transmitting a captured image to an external device before processing it, and the privacy and security improvement by transmitting a captured image to an external device. For this reason, a technology is desired that improves the recognition rate of a neural network used when performing estimation processing, etc. in a built-in signal processing circuit. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] International Publication WO2018 / 051809A1 Summary of the Invention [Problem to be solved by the invention]

[0006] Therefore, the present disclosure provides a solid-state imaging device, an electronic device, and an imaging system that realizes improved accuracy of processing by a neural network used in an image sensor. [Means for solving the problem]

[0007] According to one embodiment, a solid-state imaging device includes a pixel array having a plurality of pixels that perform photoelectric conversion, a converter that converts an analog pixel signal output from the pixel array into digital image data, an image processing unit that performs image processing of the digital image data, a digital signal processing unit that performs recognition processing on the digital image data output by the image processing unit, and a control unit that optimizes at least one of the acquisition processes of the analog pixel signals, the digital image data, or the result of the recognition processing based on a result of the recognition processing.

[0008] The control unit may perform optimization by feeding back a result of the recognition process.

[0009] The control unit may optimize at least one of the acquisition processes of the analog pixel signal, the digital image data, or the result of the recognition process when the recognition rate of the recognition process decreases.

[0010] The controller may control an exposure time of the pixels in the pixel array.

[0011] The control unit may optimize parameters relating to image processing by the image processing unit.

[0012] The digital signal processor may perform the recognition process using a trained neural network model.

[0013] The control unit may retrain the neural network model.

[0014] The solid-state imaging device may be configured to include a semiconductor device having a first substrate on which the pixel array is arranged, and a second substrate stacked on the first substrate and on which the converter, the image processing unit, the digital signal processing unit, and the control unit are arranged.

[0015] The first substrate and the second substrate may be bonded together by any one of a CoC (Chip on Chip) method, a CoW (Chip on Wafer) method, and a WoW (Wafer on Wafer) method.

[0016] The solid-state imaging device may further include a selector that selects an output from at least one of the converter, the image processing unit, and the digital signal processing unit, and outputs the data to an external device.

[0017] When outputting image information, only information within the region of interest may be selected and output.

[0018] When image information is output, different compression methods may be used for the region of interest and other regions.

[0019] The region of interest and other regions may be compressed at different compression rates and output.

[0020] When the recognition rate drops in the recognition process, the image information and the recognition information may be output together.

[0021] According to one embodiment, an electronic device includes the above-described solid-state imaging device, and an application processor on a substrate separate from the semiconductor substrate, and the application processor inputs and outputs data via the selector.

[0022] According to one embodiment, an imaging system includes the solid-state imaging device and a server connected to the solid-state imaging device via a network, and the solid-state imaging device links at least one of the outputs of the converter or the image processing unit to the output of the digital signal processing unit and outputs them to the server, and the server optimizes at least one of the acquisition processes of the analog pixel signal, the digital image data, or the result of the recognition processing based on the received information, and deploys it to the solid-state imaging device. [Brief description of the drawings]

[0023] [Figure 1] FIG. 1 is a block diagram showing an outline of an electronic device including an imaging device according to an embodiment. [Diagram 2] FIG. 2 is a block diagram showing processing of an imaging apparatus according to an embodiment. [Diagram 3] FIG. 2 is a block diagram showing processing of an imaging apparatus according to an embodiment. [Figure 4] FIG. 1 is a block diagram showing an outline of an imaging system according to an embodiment. [Diagram 5] FIG. 4 is a diagram showing an example of an acquired image according to an embodiment. [Figure 6] FIG. 2 is a diagram showing an example of the configuration of a semiconductor substrate of an imaging device according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0024] Hereinafter, an embodiment of an imaging device and an electronic device will be described with reference to the drawings. The following description will focus on the main components of the imaging device and the electronic device, but the imaging device and the electronic device may include components and functions that are not shown or described. The following description does not exclude components and functions that are not shown or described.

[0025] (First embodiment) 1 is a block diagram showing a schematic configuration of an electronic device 2 including an imaging device 1 according to an embodiment. The electronic device 2 includes the imaging device 1 and an application processor (hereinafter, referred to as AP 20). The electronic device 2 is a smartphone, a mobile phone, a tablet terminal, a PC, a digital camera, a digital video camera, or the like, which has an imaging function, and the specific form of the device is not limited.

[0026] The imaging device 1 can be realized as one semiconductor device. This semiconductor device may also be called an image sensor or a solid-state imaging device. The imaging device 1 includes at least a CMOS image sensor (hereinafter, referred to as CIS 10), an image processing unit (hereinafter, referred to as ISP 11), a digital signal processing unit (hereinafter, referred to as DSP 12), a control unit (hereinafter, referred to as CPU 13), a memory unit 14, a shutter 15, and a selector (hereinafter, referred to as SEL 16).

[0027] The CIS 10 is an image sensor having an optical system 100, an imaging section including a pixel array 102, and an analog-to-digital conversion circuit (hereinafter, referred to as ADC 104).

[0028] The optical system 100 includes, for example, a zoom lens, a single focal length lens, an aperture, etc. The optical system 100 guides incident light to a pixel array 102.

[0029] The pixel array 102 has a plurality of pixels arranged in a two-dimensional direction. Each pixel may be composed of a plurality of unit pixels of a plurality of colors such as RGB. Each unit pixel has a light receiving element such as a photodiode. The light receiving element photoelectrically converts the incident light and outputs an analog pixel signal. The light incident on the imaging unit is imaged via the optical system 100 on a light receiving surface on which a plurality of light receiving elements are arranged, and each light receiving element accumulates electric charge according to the intensity of the incident light and outputs an analog pixel signal according to the amount of accumulated electric charge.

[0030] The ADC 104 converts the analog pixel signal output from the pixel array 102 into digital image data. Since the ADC 104 performs A / D conversion, the ISP 11, DSP 12, CPU 13, storage unit 14, shutter 15, and SEL 16, which are downstream of the ADC 104, handle digital image data. A voltage generation circuit that generates a drive voltage for driving the imaging unit from a power supply voltage or the like supplied to the imaging device 1 may be provided inside the ADC 104 or separately from the ADC 104. A digital-to-analog conversion circuit (DAC) required to realize the A / D conversion of the ADC 104 may also be provided separately.

[0031] The ISP 11 performs various image processing on the digital image data. The ISP 11 may perform image processing on the digital image data output from the CIS 10, or may perform image processing on the digital image data output from the CIS 10 and stored in the storage unit 14. The ISP 11 executes image processing according to an external instruction. For example, the ISP 11 converts the digital image data so that it becomes data suitable for signal processing by the DSP 12. Hereinafter, the digital image data will be simply referred to as image, image data, etc.

[0032] The DSP 12 has a function of an information processing unit that executes at least one process, such as a predetermined recognition process and a detection process, based on the data image-processed by the ISP 11. That is, the DSP 12 executes a recognition process, etc., on the data processed by the ISP 11. The DSP 12 may also execute a predetermined process based on the result of the recognition process, etc. For example, the DSP 12 may input the data image-processed by the ISP 11 to a pre-trained, learned neural network model, and execute object recognition, etc. Then, the DSP 12 may execute a process to improve the accuracy of the neural network model based on the result of the recognition, etc. The neural network model is, for example, one trained by deep learning, but is not limited to this, and may be another model that can be retrained. When pre-trained data is used, for example, trained parameters may be stored in the storage unit 14, and a neural network model may be formed based on the parameters.

[0033] The CPU 13 controls each part in the imaging device 1 according to an instruction from the AP 20 or the like. The CPU 13 may also execute a predetermined process based on a program stored in the storage unit 14. The CPU 13 may be integrated with the DSP 12. That is, the CPU 13 may be configured to realize the process of the DSP 12. The CPU 13 may set the imaging conditions of the CIS 10 based on the result of the recognition of the DSP 12, for example. In the present disclosure, the exposure time is controlled as an example, but as described later, the process is not limited to this process, and various processes that can be processed in the ISP 11 may be controlled.

[0034] The DSP 12 executes a program stored in the storage unit 14, for example, to execute a computation process using a computation model learned by machine learning. The ISP 11 and the CPU 13 may also execute a program stored in the storage unit 14 to execute various computations. The storage unit 14 stores in advance various information related to the learned computation model, various information used for various processes, and programs. The ISP 11, the DSP 12, and the CPU 13 read out necessary information from the storage unit 14 and execute computation processes.

[0035] The processing related to the machine learning model by the DSP 12 is, for example, a computation model by a deep neural network (hereinafter, referred to as DNN) trained by deep learning as described above. This computation model can be designed based on parameters generated by inputting, for example, the output data of the ISP 11 as an input city and inputting learning data associated with a label for this input into a trained computation model.

[0036] The DSP 12 can execute, for example, a predetermined recognition process by a calculation process using the DNN. Here, the recognition process is a process of automatically recognizing whether or not image data that is output data of the ISP 11 contains characteristic image information. More specifically, the recognition process is a process in which input data is given to a calculation model formed by parameters generated by machine learning and a calculation is performed, and the input data is the output data of the ISP 11.

[0037] The DSP 12 may perform a multiplication and accumulation operation between the dictionary coefficients stored in the storage unit 14 and image data in the process of performing calculation processing based on the learned calculation model stored in the storage unit 14. The calculation result by the DSP 12 is stored in the storage unit 14 and output to the SEL 16. The result of the calculation processing by the DSP 12 using the calculation model may be image data or various information (metadata) obtained from the image data. The DSP 12 or the above-mentioned CPU 13 may have a function of a memory controller that controls writing and reading to the storage unit 14, or a memory controller may be provided separately from the DSP 12 and the CPU 13. The DSP 12 may also perform detection processing such as a motion detection process and a face detection error. The detection process may be performed by the ISP 11 instead of the DSP 12. Alternatively, the ISP 11 and the DSP 12 may cooperate to perform the detection process.

[0038] The storage unit 14 stores digital pixel data output from the CIS 10, programs executed by the ISP 11, the DSP 12, and the CPU 13, and various information related to the learned calculation model used by the DSP 12 for calculation processing. The storage unit 14 may also store data related to the progress and results of the calculation processing of each of the above-mentioned units. The storage unit 14 is, for example, a readable and writable RAM (Random Access Memory). By replacing information related to the calculation model in the storage unit 14, each of the above-mentioned units can execute various calculations, making it possible to realize highly versatile and wide-ranging processing. When each of the above-mentioned units executes calculation processing using a calculation model for a specific purpose, the storage unit 14 may include a ROM (Read Only Memory) as a part thereof.

[0039] The shutter 15 controls the exposure time in the CIS 10 under control of the CPU 13. For example, if the CPU 13 determines that it is too bright for the DSP 12 to perform recognition, the shutter 15 shortens the exposure time. Conversely, if it is too dark, the shutter 15 lengthens the exposure time. This shutter 15 does not have to be provided outside the CIS 10, but may be provided inside. Also, this shutter 15 may be an analog shutter or a digital shutter.

[0040] The SEL 16 selects and outputs the output data processed by the DSP 12 based on a selection control signal from the CPU 13. The data to be output may be the output of the DSP 12 as is, or may be data stored in the storage unit 14 by the DSP 12. In some cases, the SEL 16 may select and output the digital pixel data output by the ADC 104 based on a control signal from the CPU 13. The SEL 16 outputs necessary data to the AP 20 via an interface such as a Mobile Industry Processor Interface (MIPI) or an Inter-Integrated Circuit (I2C).

[0041] The AP 20 is a semiconductor device separate from the imaging device 1, and is mounted on the same or a different base substrate as the imaging device 1. The AP 20 has an internal CPU separate from the CPU 13 of the imaging device 1, and executes programs such as an operating system and various application software. The AP 20 may have a DSP separate from the DSP 12 of the imaging device 1, and various signal processing may be executed by the DSP. The DSP in the AP 20 may be capable of executing more advanced signal processing at higher speeds than the ISP 11, DSP 12, etc. in the imaging device 1.

[0042] In addition, the AP 20 may be equipped with a function for executing image processing, signal processing, etc., such as a GPU (Graphics Processing Unit), a baseband processor, etc. The AP 20 may execute various processes on the image data and calculation results from the imaging device 1 as necessary, may control the display of an image together with the identification result on the display unit of the electronic device 2, and may transmit the processing result or data related to the recognition result to an external cloud server via a predetermined wired or wireless network.

[0043] The predetermined network may be, for example, various communication networks such as the Internet, a wired LAN (Local Area Network), a wireless LAN, a mobile communication network, a close-proximity wireless communication such as Bluetooth (registered trademark), etc. Furthermore, the destination of the image data and the calculation result is not limited to a cloud server, but may be various information processing devices having a communication function, such as a stand-alone server, a file server, or a communication terminal such as a mobile phone.

[0044] 1 shows an example in which the AP 20 transmits an instruction to the imaging device 1. In the following, an example in which the AP 20 transmits an instruction to the imaging device 1 will be described, but the actual implementation is not limited to this form, and can be interpreted as including a case in which a processor other than the AP 20 transmits an instruction to the imaging device 1.

[0045] 2 is a flowchart showing an image capturing process according to an embodiment. The image capturing device 1 executes the image capturing process based on this flowchart. The image capturing device 1 is installed so as to capture an image within a predetermined range of a factory line, for example. This is shown as an example, and the image capturing device 1 can be applied to any device that performs recognition processing without being limited to this example. Also, other inference processing may be performed instead of recognition processing.

[0046] First, the imaging device 1 executes pre-processing (S100). The pre-processing is, for example, processing for determining whether or not an object to be recognized has been properly captured as image data by imaging, and setting imaging conditions, etc.

[0047] FIG. 3 is a flow chart illustrating this pre-processing according to one embodiment.

[0048] First, as pre-processing, the CIS 10 performs shooting and acquires an image (S1000).

[0049] Next, the ISP 11 performs appropriate image processing on the acquired image (S1002). This image processing includes, for example, various filter processes such as sharpening, resolution conversion, gain adjustment, dynamic range conversion, area cropping, color correction, color conversion, and parameters related to normalization. In addition, the ISP 11 may perform processes such as enlarging, reducing, rotating, and correcting distortion of the image in order to convert the image into an image suitable for recognition processing, which will be described later. In addition, if the image to be acquired is a moving image, the ISP 11 may perform optical flow acquisition, etc.

[0050] Next, the DSP 12 executes a recognition process (S1004) using the image that has been subjected to a predetermined image processing by the ISP 11. This recognition process is a process for recognizing whether or not it is possible to recognize, for example, whether a worker is present at an appropriate location on the line, whether an object is passing through an appropriate location on the line, and so on.

[0051] Next, the CPU 13 judges the result of the recognition by the DSP 12 and judges whether the recognition rate is sufficiently high or not (S1006). For example, the CPU 13 compares the recognition rate recognized by the DSP 12 with a predetermined value and judges whether the recognition rate is sufficiently high or not. This recognition rate may be calculated based on the number of images that have been successfully recognized out of a predetermined number or more of captured images. For example, the recognition rate may be calculated using a predetermined number or more of frame images that have been continuously captured.

[0052] If the recognition rate is not high enough (S1006: NO), the CPU 13 changes the image acquisition conditions (S1008). The image acquisition conditions may be image shooting conditions, such as conditions related to the exposure time of the shutter 15, or conditions related to image processing such as a filter coefficient in the ISP 11, such as adjustment of the filter coefficient and gain. By changing the image acquisition conditions in this way, the imaging device 1 executes setting of parameters that provide a high recognition rate as preprocessing.

[0053] The image acquisition conditions may be a combination of a plurality of the above conditions, for example, a combination of parameters related to the CIS 10, parameters related to the ISP 11, and parameters related to the DSP 12.

[0054] If the recognition rate is sufficiently high (S1008: YES), the pre-processing may be ended, or a determination as to whether or not to end the pre-processing may be made.

[0055] Next, the CPU 13 judges whether or not to end the pre-processing (S1010). This pre-processing judgment may be made if a recognition rate sufficiently higher than a predetermined recognition rate can be obtained. Another condition may be, for example, that the exposure time by the shutter 15 is changed from a predetermined shortest exposure time to a longest exposure time and then the recognition rate is obtained. In this case, the CPU 13 may set the parameter that results in the highest recognition rate as the image acquisition condition. Although the exposure time has been described, it is of course possible to optimize other image acquisition conditions in this manner as described above.

[0056] When it is determined that the pre-processing is to be ended (S1010: YES), the CPU 13 ends the pre-processing.

[0057] On the other hand, if it is determined not to end the pre-processing (S1010: NO), the CPU 13 may repeat the processing from the photographing (S1000) in the CIS 10. As another example, when adjusting a filter or the like of an already acquired image, the processing from the image processing (S1002) may be repeated, and when changing the conditions used in the recognition processing in S1008, the processing from the recognition processing (S1004) may be repeated. When a combination of multiple conditions is used, all of the combinations of the multiple conditions or a combination of several conditions appropriately selected from all the combinations may be used.

[0058] Returning to FIG. 2, the processing after the preprocessing will be described.

[0059] After the pre-processing is completed, the CIS 10 starts capturing images for performing the actual recognition process (S102). The captured images may be still images or videos, for example. The captured images are captured using the parameters related to the capture that were set in the pre-processing.

[0060] Next, the ISP 11 executes a predetermined image processing on the image captured by the CIS 10 (S104). The ISP 11 executes the above-mentioned filter processing, gain adjustment, etc. on the image. If parameters have been changed by pre-processing, the ISP 11 executes image processing based on the changed parameters.

[0061] Next, the DSP 12 executes a recognition process on the image processed by the ISP 11 (S106). The DSP 12 executes a recognition process to determine, for example, whether an object passes through a predetermined position on a factory line, whether the object is in an appropriate state, or whether a worker is working appropriately. The DSP 12 may also execute face authentication of a worker. For example, the imaging device 1 is fixedly disposed as described above, and appropriately recognizes the situation within a predetermined area of ​​an image acquired from this arrangement.

[0062] The processes up to this point can all be realized within the imaging device 1. In other words, it is possible to acquire conditions for realizing appropriate recognition processing without outputting from the imaging device 1 to an external device such as the AP 20 via an interface. These conditions can be set based on the environment by executing the pre-processing as described above. Note that the recognition rate of the recognition process in S106 may decrease due to changes in the environment. In such a case, the processes from the pre-processing in S100 may be repeated.

[0063] The pre-processing may be repeated, for example, by monitoring the recognition rate and executing the pre-processing again when the recognition rate falls below a predetermined value. As another example, the pre-processing may be repeated at predetermined time intervals or after a predetermined number of images are captured. In this manner, the pre-processing may be repeated periodically.

[0064] As described above, according to this embodiment, it is possible to set parameters and the like that realize recognition processing appropriate to the environment while the imaging device 1 is closed. All the components of the imaging device 1 can be implemented in one chip or one stacked semiconductor device, as described later. If the CIS 10, ISP 11, DSP 12, CPU 13, storage unit 14, and SEL 16 are implemented in such one semiconductor device, it is possible to set the above parameters without outputting information from the semiconductor device to the outside. Therefore, it is possible to set the parameters at a higher speed than when the information is output to the outside, and to protect privacy and security, and further, it is possible to improve the recognition rate while reducing the cost of the system.

[0065] In a factory line, for example, when sunlight is incident on the factory, the recognition rate may decrease under the same conditions depending on the sunlight. Even in such a case, according to the present embodiment, it is possible to suppress the decrease in the recognition rate. Furthermore, it is possible to avoid the decrease in the recognition rate depending on changes in various environments, not limited to sunlight, while maintaining the above-mentioned advantages, and it is possible to realize a system that is robust against the environment, etc.

[0066] According to the example of FIG. 2, processing continues further.

[0067] After the recognition process, the CPU 13 may further execute post-processing (S108). This post-processing does not need to be performed for each recognition process, and may be executed at an appropriate interval. For example, it may be executed at a predetermined time interval or after a predetermined number of shots. In another example, this post-processing may be executed sequentially. The post-processing is a process for improving the accuracy of the recognition process in the DSP 12. The CPU 13 may retrain the neural network model used by the DSP 12, for example, by feeding back the recognition result by the DSP 12. By using the retrained neural network model, it is possible to further improve the recognition rate. This re-training can also be executed in the imaging device 1, i.e., one semiconductor device, like the above pre-processing.

[0068] Next, the SEL 16 outputs appropriate data (S110). The appropriate data is, for example, data captured by the CIS 10, data processed by the ISP 11, or data related to the recognition result of the DSP 12. For example, an image may not be output under normal circumstances, and an image may be output when an abnormality occurs in the recognition. The recognition rate may be output all the time, or may be output only when an abnormality occurs. In this way, appropriate data is selected and output from the SEL 16 to the AP 20. The data to be output may be temporarily stored in the storage unit 14. For example, the data stored in the storage unit 14 may be output to some extent in a lump in response to a request from the AP 20. Also, the SEL 16 may receive data, requests, etc. from the AP 20. The input data, requests, etc. may be stored in the storage unit 14, for example, or output to the CPU 13.

[0069] As described above, according to this embodiment, the imaging device 1 can retrain the neural network model to improve the recognition processing without outputting data to the outside. As with the pre-processing, this processing can be realized within one semiconductor device. Therefore, as with the above, it is possible to protect privacy and security at a high speed compared to the case of outputting data to the outside, and furthermore, it is possible to reduce the recognition accuracy while reducing the system cost.

[0070] For example, when replacing an existing photographing device with a new one, learning costs are generally incurred. Also, the cost of tuning for each installation location of the photographing device is high. Even in such a case, it is possible to reduce costs by optimizing the image acquisition conditions and the neural network model in an environment closed to the imaging device 1 as in this embodiment.

[0071] Second embodiment 4 is a schematic diagram showing an example of an imaging system according to the second embodiment. The imaging system 3 includes an electronic device 2 and a cloud 30. The cloud 30 is a broad concept that may be a general cloud on the Internet, or may be, for example, a server or a group of servers closed within an intranet. The electronic device 2 is connected to the cloud 30 via, for example, a wired or wireless network.

[0072] In the first embodiment described above, the pre-processing and the re-training of the model are all closed processes within one semiconductor device, but the re-training of the model may of course be realized externally. In this case, data may be transmitted from the electronic device 2 to a server or the like in the cloud 30 at a predetermined timing. In this case, the image data to be transmitted may be encrypted in order to protect privacy and improve security.

[0073] In the cloud 30 to which the data is transmitted, retraining can be performed using a CPU, GPU, or the like with higher accuracy than the CPU 13, etc., provided in the imaging device 1. For this reason, it is possible to retrain the model in the imaging device 1 while executing the recognition process, and to perform training in parallel in a server, etc. with higher accuracy.

[0074] In this embodiment, for example, when data is transmitted from the electronic device 2 to the cloud 30, parameters of the neural network model used in the DSP 12 may be transmitted together. Then, the server or the like in the cloud 30 may retrain the received parameters using the received information. After that, the server or the like can feed back the optimized parameters to the electronic device 2 that transmitted the data, thereby optimizing image acquisition in the electronic device 2 and improving the recognition accuracy by the neural network model. In this way, various parameters optimized based on the output from the imaging device 1 may be deployed from the server side to the imaging device 1.

[0075] For example, when the recognition rate drops or an abnormal value occurs in the recognition, the electronic device 2 may transmit only the image data that is the cause of this. Then, the data that caused the drop in the recognition rate may be analyzed in a server or the like. This data may be used to optimize parameters of the image acquisition conditions or a model used in the recognition. Furthermore, instead of transmitting only the image data that is the cause, image data relating to the previous and next frames may also be transmitted. Here, the previous and next frames are not limited to one previous and next frame, but may include multiple frames.

[0076] A server or the like on the cloud 30 may accept data from a plurality of electronic devices 2. In such a case, an identifier (ID) or the like uniquely assigned to the electronic device 2 may be linked and transmitted together with the recognition rate and image data. In this case, the model number or the like of the electronic device 2 may also be transmitted. By transmitting the ID in this manner, it becomes possible to process information from a plurality of electronic devices 2 by one server or the like. Also, by transmitting the model number or the like, it is possible to transmit parameters or the like obtained by performing similar processing on electronic devices having the same model number or devices having similar imaging systems.

[0077] The information to be transmitted may not be the image data itself as described above, but may be appropriate information related to the image. For example, the imaging device 1 may include a motion detection circuit in addition to the configuration of FIG. 1 and the like. This motion detection circuit is a circuit that acquires how much movement or how much brightness change has occurred in the acquired image. This motion detection circuit can also be included in one semiconductor device, like other components. The electronic device 2 may transmit the information output by the motion detection circuit output from the imaging device 1 to a server or the like. The server or the like may identify the cause of the decrease in the recognition rate by analyzing the motion detection result. The server or the like may optimize the image acquisition conditions or the neural network model based on the analysis result.

[0078] The imaging device 1 may transmit information on the detection value used for adjusting the exposure to a server, etc. The detection value is, for example, information on a change in brightness.

[0079] Furthermore, the image capturing device 1 may transmit to a server, etc., image data of frames before and after an image in which the recognition rate has decreased, as well as the recognition results of the frames before and after and the change in the recognition rate. In this case, parameters related to the image quality of the frames before and after may also be transmitted. By transmitting these, the server, etc., can analyze which parameter caused the recognition rate to decrease, making it possible to further optimize the model.

[0080] Furthermore, the imaging device 1 may transmit not only the recognition result but also data in the intermediate layer of the neural network model. The intermediate layer data may be, for example, data representing feature quantities that have been dimension-compressed, and it is possible to obtain optimal parameters by analyzing such data in a server or the like. As another example, an error backpropagation may be performed from the intermediate layer data to optimize the encoder layer from the input layer to the intermediate layer in the neural network, or conversely, a layer that realizes recognition from the feature quantities from the intermediate layer to the output layer may be optimized.

[0081] In the above, parameters related to the internal state of the imaging device 1 are output, but this is not limited to this. For example, the imaging device 1, i.e., data indicating the location, time, temperature, humidity, and other external environment of the electronic device 2 may be transmitted to a server, etc. By transmitting data related to the external environment in this manner, optimization based on the external environment can also be realized.

[0082] Moreover, by transmitting the position information, it is possible to perform optimization with respect to the conditions for acquiring an image. The position information may be, for example, information from a Global Positioning System (GPS). Furthermore, information such as the shooting angle of the imaging device 1 with respect to an object may be output. This angle may be, for example, information acquired by a gyro sensor, an acceleration sensor, or the like further provided in the imaging device 1.

[0083] Also, by transmitting time information, it is possible to optimize the change in the recognition rate depending on the time. In this case, the imaging device 1 may be configured to control parameters depending on the time even in normal shooting based on a program deployed from a server or the like. Of course, this parameter control may be performed based on other external environments such as temperature and humidity.

[0084] In this manner, the imaging system 3 may perform processes on the cloud 30 if the processes are too costly to be realized in the imaging device 1 .

[0085] As described above, according to this embodiment, it is possible to execute advanced retraining using a CPU or the like with higher performance than that of the imaging device 1. Furthermore, in this case, the server or the like can analyze the received inference result to further improve the accuracy. For example, by using the data received by the server or the like as training data, it is possible to generate a neural network model with higher accuracy, and by optimizing the image acquisition conditions, it is possible to acquire an image that is more suitable for recognition.

[0086] Third embodiment 5 is a diagram showing an example of an image captured by the imaging device 1. The image Im is an image captured by the imaging device 1 installed in, for example, a factory that produces bottle-shaped objects. In this factory, for example, a worker processes or visually inspects bottles that flow through a line. An electronic device 2 is installed so as to capture an image of the factory line in order to capture an image of the worker and the bottle.

[0087] As described above, the image Im contains images of a worker and bottles flowing through the line. As described in the above embodiment, for example, the imaging device 1 may output image information together with the recognition result in the DSP 12. However, if all images are output as low-compression, high-resolution images, the bandwidth for transmission and the memory area for storing image information may become strained.

[0088] Therefore, in this embodiment, for example, when image information is transmitted for retraining, a method for reducing the amount of data while maintaining accuracy will be described.

[0089] In the imaging device 1, for example, the SEL 16 may output a part of the image information to the outside. The image data output to the outside, for example, the AP 20, is transmitted to a server on the cloud 30. The imaging device 1 may crop and output only an area required for recognition.

[0090] The imaging device 1 may output only information within a region of interest (hereinafter, referred to as ROI R1, etc.) where, for example, a worker is likely to be captured. In this case, information of regions other than ROI R1 may be deleted, and only ROI R1 may be cropped and transmitted after appropriate image processing. As another example, only ROI R1 may be compressed with high accuracy. For example, the ISP 11 may change the data compression method for the ROI R1 and the data compression method for the other regions. The ROI R1 may be compressed with a method that allows high accuracy recovery but does not allow much data compression, and the other regions may be compressed with a method that does not allow high accuracy recovery but has a higher data compression rate than the ROI R1.

[0091] Also, the same data compression method may be used, with compression parameters enabling high-resolution decompression within ROI R1, and compression parameters enabling less resolution but greater data reduction for other regions.

[0092] As another example, raw data may be output within ROI R1, while data compressed at a high compression rate may be transmitted to other regions.

[0093] The number of ROIs is not limited to one, and may be multiple. For example, as shown in FIG. 5, ROI R1 and ROI R2 may exist, and the data within these ROIs may be compressed in a manner that can perform high-precision image restoration different from that of the other regions. In addition, it is not necessary to maintain the same precision between the ROIs. That is, ROI R1 and ROI R2 may be compressed with different compression rates or different compression methods. For example, the compression method used for ROI R2 may be one that can perform high-precision image restoration compared to the compression method used for ROI R1. Of course, in this case, the compression method that can reduce the amount of data may be used for the regions other than ROI R1 and R2. In addition, it is not necessary to transmit images in the regions other than ROI R1 and R2.

[0094] Furthermore, when transmitting consecutive frames of information, information within the ROI may be transmitted in every frame, while information outside the ROI may be transmitted by thinning out some frames.

[0095] Then, in a server in the cloud 30, the highly accurate restored images, e.g., high resolution images, or raw data may be used to optimize image acquisition conditions or retrain the neural network model.

[0096] As described above, according to this embodiment, it is possible to acquire images under appropriate image acquisition conditions within the imaging device 1, and to realize optimization of models and the like with further improved performance outside the imaging device 1. Furthermore, it is possible to transmit and receive data without straining the bandwidth and the like while continuing recognition.

[0097] (Chip structure of imaging device 1) Next, the chip structure of the imaging device 1 in FIG. 1 will be described. FIG. 6 is a diagram showing an example of the chip structure of the imaging device 1 in FIG. 1. The imaging device 1 in FIG. 6 is a stacked pair in which a first substrate 40 and a second substrate 41 are stacked. The first substrate 40 and the second substrate 41 are sometimes called dies. In the example of FIG. 6, an example in which the first substrate 40 and the second substrate 41 are rectangular is shown, but the specific shapes and sizes of the first substrate 40 and the second substrate 41 are arbitrary. In addition, the first substrate 40 and the second substrate 41 may be the same size or different sizes.

[0098] 1 is disposed on the first substrate 40. Also, at least a part of the optical system 100 of the CIS 10 may be mounted on-chip on the first substrate 40. Also, although not shown, a shutter 15 may be mounted on the first substrate 40. In the case of an optical shutter, for example, the shutter may be provided so as to cover the light receiving surface of the light receiving element of the first substrate 40, and in the case of a digital shutter, the shutter may be provided so as to control the light receiving element of the pixel array 102.

[0099] 1 are arranged on the second board 41. In addition, the second board 41 may also include components necessary for controlling the imaging device 1, such as an input / output interface, a power supply circuit, etc. (not shown).

[0100] The first substrate 40 and the second substrate 41 are laminated by a predetermined method to form a single semiconductor device. As a specific example of lamination, the so-called CoC (Chip on Chip) method may be adopted in which the first substrate 40 and the second substrate 41 are cut out from a wafer, diced, and then laminated together. Alternatively, the so-called CoW (Chip on Wafer) method may be adopted in which one of the first substrate 40 and the second substrate 41 (for example, the first substrate 40) is taken out from the wafer and diced, and then the diced first substrate 40 is laminated to the second substrate 41 before dicing. Alternatively, the so-called WoW (Wafer on Wafer) method may be adopted in which the first substrate 40 and the second substrate 41 are laminated together in the wafer state.

[0101] The first substrate 40 and the second substrate 41 can be bonded together by, for example, via holes, microbumps, micropads, plasma bonding, etc. However, various other bonding methods may also be used.

[0102] 6 is given as an example, and the arrangement of each component on the first substrate 40 and the second substrate 41 is not limited thereto. For example, at least one component arranged on the second substrate 41 shown in FIG. 6 may be provided on the first substrate 40. In addition, although a stacked type is shown, this is not limiting, and the above components may be arranged on a single semiconductor substrate.

[0103] (Applicability to other sensors) In the above-described embodiment, the technology according to the present disclosure is applied to an imaging device 1 (image sensor) that acquires a two-dimensional image, but the application of the technology according to the present disclosure is not limited to imaging devices. For example, the technology according to the present disclosure can be applied to various light receiving sensors such as a ToF (Time of Flight) sensor, an infrared (IR) sensor, and a DVS (Dynamic Vision Sensor). In other words, by making the chip structure of the light receiving sensor a stacked type, it is possible to reduce noise contained in the sensor result and reduce the size of the sensor chip.

[0104] In addition, by setting image acquisition conditions and optimizing the neural network model using the above-mentioned method, feedback processing can be performed in a closed state within the imaging device 1, i.e., while ensuring privacy and security, and depending on the situation, more cost-effective and accurate optimization can be achieved on a server on the cloud 30.

[0105] The above-described embodiment may be modified as follows.

[0106] (1) A pixel array having a plurality of pixels that perform photoelectric conversion; a converter for converting analog pixel signals output from the pixel array into digital image data; an image processing unit that processes the digital image data; a digital signal processing unit that performs recognition processing on the digital image data output by the image processing unit; a control unit that optimizes at least one of the acquisition processes of the analog pixel signal, the digital image data, or the result of the recognition process based on a result of the recognition process; A solid-state imaging device comprising:

[0107] (2) The control unit performs optimization by feeding back a result of the recognition process. A solid-state imaging device according to (1).

[0108] (3) When a recognition rate of the recognition process is decreased, the control unit optimizes at least one of the acquisition processes of the analog pixel signal, the digital image data, and the result of the recognition process. A solid-state imaging device according to (2).

[0109] (4) The control unit controls an exposure time of the pixels of the pixel array. A solid-state imaging device according to (2) or (3).

[0110] (5) The control unit optimizes parameters related to image processing of the image processing unit. A solid-state imaging device according to any one of (2) to (4).

[0111] (6) The digital signal processing unit performs recognition processing using a trained neural network model. A solid-state imaging device according to any one of (2) to (5).

[0112] (7) The control unit retrains the neural network model. A solid-state imaging device according to (6).

[0113] (8) a first substrate on which the pixel array is disposed; a second substrate laminated on the first substrate, on which the converter, the image processing unit, the digital signal processing unit, and the control unit are disposed; A semiconductor device having The solid-state imaging device according to any one of (1) to (7), comprising:

[0114] (9) The first substrate and the second substrate are bonded together by any one of a CoC (Chip on Chip) method, a CoW (Chip on Wafer) method, and a WoW (Wafer on Wafer) method. A solid-state imaging device according to (8).

[0115] (10) a selector that selects an output from at least one of the converter, the image processing unit, or the digital signal processing unit and outputs the data to an outside; The solid-state imaging device according to (8) or (9), further comprising:

[0116] (11) When outputting image information, only information within the region of interest is selected and output. A solid-state imaging device according to (10).

[0117] (12) When outputting image information, different compression methods are used for the region of interest and other regions. A solid-state imaging device according to (10).

[0118] (13) The region of interest and other regions are compressed at different compression rates and output. A solid-state imaging device according to (12).

[0119] (14) When a recognition rate is decreased in the recognition process, the image information and the recognition information are output together. A solid-state imaging device according to any one of (11) to (13).

[0120] (15) A solid-state imaging device according to any one of (9) to (14), an application processor on a substrate separate from the semiconductor substrate; The application processor inputs and outputs data via the selector. electronic equipment.

[0121] (16) A solid-state imaging device according to any one of (1) to (13), a server connected to the solid-state imaging device via a network; Equipped with the solid-state imaging device links at least one of the output of the converter or the output of the image processing unit with the output of the digital signal processing unit and outputs the linked output to the server; The server optimizes at least one of the acquisition processes of the analog pixel signal, the digital image data, and the result of the recognition process based on the received information, and deploys the optimization results to the solid-state imaging device. Imaging system.

[0122] The aspects of the present disclosure are not limited to the above-mentioned individual embodiments, but include various modifications that may be conceived by a person skilled in the art, and the effects of the present disclosure are not limited to the above-mentioned contents. In other words, various additions, modifications, and partial deletions are possible within the scope of the conceptual idea and intent of the present disclosure derived from the contents defined in the claims and their equivalents.

[0123] In addition to the above-mentioned applications to moving objects and medical care, the present invention can also be applied to devices that detect movement and perform recognition processing within the imaging device 1, such as surveillance cameras. [Explanation of symbols]

[0124] 1: Imaging device, 10: CIS, 100:Optical system, 102: pixel array, 104: ADC, 11: ISP, 12: DSP, 13: CPU, 14: Storage part, 15: Shutter, 16:SEL, 2:Electronic equipment, 20: AP, 3: Imaging system, 30: Cloud, 40: first substrate, 41: Second board

Claims

1. A pixel array having a plurality of pixels that perform photoelectric conversion; a converter for converting analog pixel signals output from the pixel array into digital image data; an image processing unit that processes the digital image data; a digital signal processing unit that performs recognition processing on the digital image data output by the image processing unit; a control unit that, when a recognition rate of the result of the recognition process is reduced, executes a pre-processing for optimizing at least one of the acquisition processes of the analog pixel signal, the digital image data, and the result of the recognition process, and further repeats the pre-processing as necessary based on the result of the optimization, and controls the actual shooting using the result of the optimization; Equipped with The recognition rate is determined based on the number of images that have been successfully recognized among the plurality of digital image data captured continuously. Solid-state imaging device.

2. The control unit determines the recognition rate based on a predetermined number or more of images captured in succession.

2. A solid-state imaging device according to claim 1.

3. The control unit controls exposure times of the pixels of the pixel array as an optimization.

3. The solid-state imaging device according to claim 1 or 2.

4. The control unit optimizes parameters related to image processing of the image processing unit.

4. The solid-state imaging device according to claim 1.

5. The digital signal processing unit performs recognition processing using a trained neural network model.

5. The solid-state imaging device according to claim 1.

6. The control unit retrains the neural network model.

6. A solid-state imaging device according to claim 5.

7. a first substrate on which the pixel array is disposed; a second substrate laminated on the first substrate, on which the converter, the image processing unit, the digital signal processing unit, and the control unit are disposed; A semiconductor device having 7. The solid-state imaging device according to claim 1, comprising:

8. The first substrate and the second substrate are bonded together by any one of a CoC (Chip on Chip) method, a CoW (Chip on Wafer) method, and a WoW (Wafer on Wafer) method.

8. A solid-state imaging device according to claim 7.

9. a selector that selects an output from at least one of the converter, the image processing unit, or the digital signal processing unit and outputs the data to an outside; 9. The solid-state imaging device according to claim 7, further comprising:

10. When outputting image information, only information within the region of interest is selected and output.

10. The solid-state imaging device according to claim 9.

11. When outputting image information, different compression methods are used for the region of interest and other regions.

10. The solid-state imaging device according to claim 9.

12. The region of interest and other regions are compressed at different compression rates and output.

12. A solid-state imaging device according to claim 11.

13. When the recognition rate is decreased in the recognition processing, the image information and a result of the recognition processing are output together.

13. A solid-state imaging device according to claim 10.

14. A solid-state imaging device according to any one of claims 9 to 13, an application processor is provided on a substrate independent of the semiconductor device; The application processor inputs and outputs data via the selector. electronic equipment.

15. A solid-state imaging device according to any one of claims 1 to 12, a server connected to the solid-state imaging device via a network; Equipped with the solid-state imaging device links at least one of the output of the converter or the output of the image processing unit with the output of the digital signal processing unit and outputs the linked output to the server; The server optimizes at least one of the acquisition processes of the analog pixel signal, the digital image data, or the result of the recognition process based on the received information, and deploys the optimization results to the solid-state imaging device. Imaging system.

Citation Information

Patent Citations

  • Target detection method and device and fuzzy processing method and device

    CN108460395A

  • Image recognition camera

    JP2008017259A

  • On-vehicle camera system and camera lens abnormality detecting method

    JP2014115814A

  • Semiconductor device and manufacturing method

    JP2017079281A

  • Stacked image sensor package and stacked image sensor module including the same

    US20180040584A1