Imaging apparatus and method for controlling imaging apparatus
By designing machine learning models containing multiple computing layers and machine learning computing components of quantization computing layers in the camera device, the problem of performance compromise between convolutional neural network accelerator in IoT devices is solved, and efficient, high-speed and high-versatility image processing operations are achieved.
Patent Information
- Application Number
- CN202380074658.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-31
- Filing Date
- 2023-10-05
- Publication Date
- 2025-06-03
AI Technical Summary
In embedded devices such as IoT devices, in order to realize image recognition and image quality improvement processing of convolutional neural networks, efficient, high-speed and high-versatility accelerators are needed, but due to the compromise of performance, the computing efficiency is reduced.
An imaging device is designed, including an image sensor, a first machine learning computing component and a second machine learning computing component. The first machine learning computing component uses a machine learning model containing multiple computing layers to process the image, and decides whether to perform processing through the decision component; the second machine learning computing component uses different models to further process the processing results, including quantizing the computing layer to improve computing efficiency.
It realizes efficient and high-speed operation related to machine learning in IoT devices, improves the computing efficiency and versatility of the circuit, and reduces power consumption.
Smart Images

Figure CN120092457A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an imaging device and a control method for the imaging device.
[0002] This application claims priority based on Japanese Patent Application No. 2022-174141 filed in Japan on October 31, 2022, and incorporates by reference all the descriptions recorded in this application. Background Art
[0003] In recent years, a convolutional neural network (CNN) has been used as a model for image recognition and the like. A convolutional neural network has a multi-layer structure including a convolutional layer and a pooling layer, and requires multiple operations to be performed in parallel. Various operation methods for accelerating operations based on a convolutional neural network have been proposed (Patent Document 1, etc.).
[0004] Prior Art Documents
[0005] Patent Documents
[0006] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2018-077829 Summary of the Invention
[0007] Problems to be Solved by the Invention
[0008] On the other hand, in embedded devices such as IoT devices, image recognition, image quality improvement processing, etc. using a convolutional neural network are also performed. In order to implement such a convolutional neural network, an accelerator with high versatility in addition to high speed and power saving is required. However, each performance is a trade-off relationship, and in particular, a certain redundancy is required to improve versatility, resulting in a decrease in operation efficiency with respect to the circuit scale or power consumption. Therefore, for an accelerator that processes CNN, versatility is also desired in addition to high speed and power saving.
[0009] In view of the above circumstances, an object of the present invention is to provide an imaging device and a control method for the imaging device, the imaging device being an imaging device that can acquire an image as an IoT device, efficiently and rapidly performing operations related to machine learning on the acquired image, and the control method for the imaging device being for efficiently and rapidly operating a circuit and a model that perform operations related to machine learning.
[0010] Means for Solving the Problems
[0011] In order to solve the above problems, the invention proposes the following means.
[0012] The imaging device according to an embodiment of the present invention includes an image sensor having a plurality of pixels for converting a subject image into an electrical signal. The imaging device is characterized by including: a first machine learning operation component for processing a signal obtained from the plurality of pixels using a first machine learning model including a plurality of operation layers; a decision component for determining whether to perform processing by the first machine learning operation component; and a second machine learning operation component for processing based on the signal or a processing result obtained by the first machine learning operation component using a second machine learning model different from the first machine learning model. The first machine learning model includes at least a quantization operation layer for performing a quantization operation on an operation result of a convolution operation layer that performs a convolution operation.
[0013] Effect of the Invention
[0014] The imaging device and the control method of the imaging device according to the present invention can provide the following imaging device and the control method of the imaging device: The imaging device is an imaging device capable of acquiring an image as an IoT device, and performs operations related to machine learning on the acquired image efficiently and at high speed. The control method of the imaging device is for causing a circuit and a model that perform operations related to machine learning to operate efficiently and at high speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a functional block diagram of the imaging device according to the first embodiment.
[0016] Figure 2 It is a functional block diagram of the sensor according to the first embodiment.
[0017] Figure 3 It is a functional block diagram of the first machine learning operation unit according to the first embodiment.
[0018] Figure 4 It is a timing chart for explaining the operation of the buffer according to the first embodiment.
[0019] Figure 5 It is a diagram showing the network structure of the machine learning model according to the first embodiment.
[0020] Figure 6 It is a functional block diagram of the second machine learning operation unit according to the first embodiment.
[0021] Figure 7 It is a diagram showing each operation layer in the feature extraction unit according to the first embodiment.
[0022] Figure 8 It is a flowchart for explaining the control method of the imaging device according to the first embodiment.
[0023] Figure 9It is a functional block diagram of the imaging device according to the second embodiment.
[0024] Figure 10 It is a timing chart for explaining the operations of the buffer and the ISP according to the second embodiment.
[0025] Figure 11 It is a functional block diagram of the imaging device according to the third embodiment.
[0026] Figure 12 It is a functional block diagram of the imaging device according to the fourth embodiment. Detailed Embodiments
[0027] (First Embodiment)
[0028] Refer to Figures 1 to 8 to describe the embodiments of the present invention. Figure 1 It is a diagram showing the imaging device 1000 of the present embodiment.
[0029] [Imaging Device 1000]
[0030] Figure 1 It is a functional block diagram of the imaging device of the present embodiment. Referring to this figure, the imaging device 1000 of the present embodiment will be described. The imaging device 1000 is a device for acquiring a subject image generated by a predetermined condensing device such as an optical member such as a lens. As an example, it is a digital camera, a surveillance camera, a vehicle-mounted camera, etc. However, as long as it is a device equipped with an imaging component such as a smartphone or a robot, the present invention can be applied. It should be noted that the invention of the present embodiment is preferably applied to products such as embedded devices with limited power consumption such as battery drive.
[0031] The imaging device 1000 of the present embodiment includes a sensor 100, a first machine learning operation unit 200, a sensor I / F 300, an ISP 400, an input / output unit 500, a display unit 600, a CPU 700, a memory 800, and a second machine learning operation unit 900.
[0032] The sensor 100 is an individual imaging element that converts a subject image imaged by an optical component (not shown) into an electrical signal through photoelectric conversion. As an example, it is a CMOS image sensor. The sensor 100 of the present embodiment is as Figure 2As shown, there are a plurality of pixels 110 having at least more pixel numbers than 2000×1500 pixels. Each pixel 110 has a predetermined color filter, and the sensor 100 of the present embodiment has a color filter with a so-called Bayer arrangement. In addition, the sensor 100 includes an analog-to-digital conversion circuit (ADC) 120 that converts an analog electrical signal obtained by photoelectric conversion into a digital value based on the timing of a synchronization signal generated by a sensor control unit (not shown) controlled by the CPU 700 described later. Furthermore, it includes a multi-channel high-speed I / F 130 capable of high-speed output of multi-bit digital signals converted by the analog-digital conversion circuit 120.
[0033] Here, in the present embodiment, the analog-digital conversion circuit 120 included in the sensor 100 has a resolution capable of converting each pixel value into a digital value of 12 bits or more, and operates in a plurality of drive modes under the control of a sensor control unit (not shown). As an example, it may have a mode of reading signals from all the pixels 110 included in the sensor 100 through a rolling shutter operation and outputting a digital value of 12 bits, a mode of reading by partially adding or thinning signals from a part of the pixels 110 included in the sensor 100 and outputting a digital value of 10 bits, etc. In addition, as a high-definition dynamic image mode, it may also have a mode of outputting 30 frames or 60 frames per second for the pixel numbers in 4K or 8K format. It should be noted that the signals of each pixel 110 output by the sensor 100 are so-called RAW image data, including information with a bit accuracy of 12 bits or 14 bits.
[0034] The first machine learning operation unit 200 is an operation unit for taking as input the multi-bit digital value, i.e., RAW image data, which is the output of the sensor 100, and performing an operation based on a predetermined machine learning model on the input. Figure 3 It is a functional block diagram of the first machine learning operation unit 200. The first machine learning operation unit 200 includes a buffer 210, a preprocessing unit 220, a first inference unit 230, and a post-processing unit 240.
[0035] The buffer 210 is a buffer that receives the output of the sensor 100 and temporarily holds it. The sensor 100 in this embodiment repeatedly outputs the pixel values of a predetermined unit pixel at the period of the horizontal synchronization signal (HD). As an example, the sensor 100 outputs the pixel values of one line in sequence for one period of the horizontal synchronization signal. That is, when the sensor 100 has 1500 rows of pixels 110, 1500 pixel values of 12 bits or 14 bits, which are more than 8 bits, are output in one horizontal synchronization. Moreover, the pixel values of one frame are output during 1500 periods of the horizontal synchronization signal period. In particular, when processing an image by a machine learning model, convolution operation is used, so it is necessary to hold the pixel values of multiple rows. Therefore, the buffer 210 has a capacity to hold the pixel values of multiple rows of 3 rows or more.
[0036] Figure 4 is a timing diagram for explaining the operation of the buffer 210. In this embodiment, for simplicity of explanation, an example of holding the pixel values of one line in the buffer 210 is shown. The sensor 100 outputs the pixel values of one screen at the period of the vertical synchronization signal VD. Then, the period of the vertical synchronization signal VD is divided into multiple periods of the horizontal synchronization signal HD, and the sensor 100 outputs the pixel values of a predetermined unit (for example, one line) based on the period of the horizontal synchronization signal HD. In Figure 4 A, the data output timing of the pixel values output from the sensor 100 is shown. The sensor 100 synchronizes with the timing of the horizontal synchronization signal HD and sequentially outputs pixel value data during the period Ta, such as pixel A, pixel B, and pixel C. After the period Ta when the sensor 100 outputs the pixel value data, power saving operations such as cutting off the power of each block are performed. Therefore, by using multiple output CHs etc. for high-speed data transmission, the shorter the period Ta, the more power reduction in the sensor 100 can be achieved. In Figure 4 B, the data output timing of the pixel values output from the buffer 210 is shown. Synchronized with the timing of the horizontal synchronization signal HD, the pixel value data is sequentially read out during the period Tb. It should be noted that the slower the peak value of the data rate of the processed data at the subsequent stage of the buffer 210, the more power reduction can be achieved. Therefore, it is preferable to perform data rate conversion when reading from the buffer 210. That is, in the buffer 210, by making the data rate at the time of reading slower than the peak data rate at the time of input during the period of the horizontal synchronization signal HD, the effect of improving the processing efficiency can be obtained.
[0037] In Figure 4In this case, an example in which the buffer 210 holds pixel values of one line is described, but it is not limited thereto. When multiple lines of data are required in the first inference unit 230 described later, multiple lines may also be held. For example, in the first inference unit 230, when a 3×3 weight is used for calculation, at least 3 lines may be held. It should be noted that the unit held in the buffer 210 may not be a line unit. When the sensor 100 outputs pixel values in units of a predetermined area, the unit of that area may be set to be held.
[0038] In Figure 3 this case, the preprocessing unit 220 uses the output of the buffer 210 as input and performs predetermined processing. Examples of the processing performed by the preprocessing unit 220 include sorting of pixel values held in the buffer 210, correction of defective pixels, black level correction, fixed pattern noise correction, and the like. By performing predetermined processing at the stage before the first inference unit 230 described later, the processing accuracy in the first inference unit 230 can be improved. The predetermined processing in the preprocessing unit 220 is rule-based processing. The correction of defective pixels is performed based on comparison with surrounding pixels, and the correction of fixed pattern noise is performed based on a correction value calculated from the black level.
[0039] The first inference unit 230 includes a processing circuit that corrects pixel values through an inference process using a machine learning model. In the present embodiment, the machine learning model included in the first inference unit 230 is assembled in a state where learning has been completed based on a previously captured image. For example, in the case of learning to reduce noise in an image by a machine learning model, a low-sensitivity image with less noise and a high-sensitivity image with more noise of the same subject and at the same exposure are prepared, and the image with more noise is inferred using the image with less noise as training data, thereby enabling learning. It should be noted that in addition to or on the basis of learning to reduce noise, the machine learning model may also learn to improve the quality of an image in reducing jitter in the image, reducing optical aberration, and the like. As an example, the machine learning model included in the first inference unit 230 may be a neural network having a network structure such as a U-NET structure.
[0040] Here, the machine learning model included in the first inference unit 230 has a multi-layer structure including a plurality of operation layers, and convolution operations using weights are performed in each layer. Figure 5It is a diagram showing the network structure of the machine learning model included in the first inference unit 230. The machine learning model included in the first inference unit 230 includes an input layer 231, a convolutional operation layer 232, a quantization operation layer 233, and an output layer 234. The convolutional operation layer 232 and the quantization operation layer 233 include multiple layers (n layers), and each layer is alternately connected, but it is also possible to skip some layers and connect. In addition, the machine learning model may also have layers with other functions such as a fully connected layer and a pooling layer. It should be noted that the machine learning model included in the first machine learning operation unit 200 corresponds to the first machine learning model.
[0041] An input signal IN is input to the input layer 231. The input signal IN is a signal corresponding to multiple pixel values generated based on the output of the preprocessing unit 220, and in this embodiment, it is a signal with a bit precision of 8 bits or more. The input layer converts each input signal IN into a vector with multiple elements. The converted vector is used as the input to the first layer convolutional operation layer 232-1.
[0042] The convolutional operation layer 232 performs a convolutional operation using the weight W on the input vector or a tensor formed by combining multiple vectors (hereinafter referred to as activation). In particular, in the convolutional operation layer 232 of this embodiment, the activation or weight W used for the operation is quantized to 8 bits or less. As an example, the activation is quantized to an 8-bit value, and the weight W is quantized to a 1-bit value. In this way, by performing the operation with the quantized low bits, effects such as reducing the storage capacity of the memory for holding the parameters themselves, saving space in the operation circuit, and increasing the operation speed can be obtained. It should be noted that for the activation, it can also be quantized to 2 bits for the purpose of reducing the operation load, etc.
[0043] The quantization operation layer 233 takes the convolution operation result in the convolution operation layer 232 as input and performs an operation of quantization using a predetermined function. The quantized convolution operation result becomes the input to the next convolution operation layer 232. In this embodiment, each element of the matrix of the convolution operation result output from the convolution operation layer 232 is a 16-bit integer, and its quantization result is a lower bit than the input signal IN, and as an example, it is an 8-bit integer. In this case, the function shown in Equation 1 below is used for quantization. It should be noted that as a quantization method, it is also possible to use multiple thresholds or tables instead of a function. In the case of quantization to 2 bits, it can be achieved by comparing with three thresholds.
[0044] [Mathematical formula 1]
[0045]
[0046] As Figure 5As shown, operations are repeatedly performed using multiple convolutional operation layers 232 and quantization operation layers 233, and the result of the n-th convolutional operation layer 232-n is input to the output layer 234. The output layer 234 outputs the operation result in the machine learning model.
[0047] In Figure 3 the post-processing unit 240 takes the operation result of the first inference unit 230 as input, performs predetermined processing, and outputs the data of the operation result in synchronization with the horizontal synchronization signal. In the present embodiment, examples of the processing performed by the post-processing unit 240 include sorting of pixel values, addition of pixel values, thinning, serial signal conversion, addition of header information and synchronization signals, etc. Through the processing of the post-processing unit 240, high-speed data transmission to the subsequent block can be achieved.
[0048] It should be noted that all or part of the functions of the first machine learning operation unit 200 can also be implemented using hardware such as ASIC (Application Specific Integrated Circuit), PLD (Programmable Logic Device), or FPGA (Field-Programmable Gate Array). The first machine learning operation unit 200 can reduce the operation resources by quantifying its elements in the convolutional operation that requires a large amount of operation resources. The connection between the sensor 100 and the ISP 400 described later is performed through multi-channel high-speed communication. Therefore, configuring a large-scale operation circuit may cause a delay in this communication. However, by using quantization technology, miniaturization of this operation circuit can be achieved, and operations in a machine learning model with multiple layers can be performed.
[0049] For example, in order to constitute the functions of the first machine learning operation unit 200, a processor that executes program processing and an accelerator that executes operations related to neural networks can also be combined. Specifically, a neural network operation accelerator for repeatedly performing convolutional operations and quantization operations can be used in combination with a processor.
[0050] Return Figure 1 to the description of
[0051] The sensor I / F 300 receives the output of the first machine learning operation unit 200 and outputs data to the internal bus IB. As an example, the sensor I / F 300 outputs the data received from the first machine learning operation unit 200 to the memory 800 for subsequent processing. Additionally, as another example, the sensor I / F 300 outputs to the ISP 400 described later in order to perform predetermined image processing on the data received from the first machine learning operation unit 200. Note that, in order to perform data format conversion, etc. in the sensor I / F 300, a buffer for temporarily holding data may be provided.
[0052] The ISP 400 is an image processing unit that selectively performs predetermined image processing on data based on pixel values acquired by the sensor 100 (hereinafter referred to as image data). As an example, it performs demosaicing processing, encoding compression processing, color adjustment processing, gamma correction processing, etc. Each processing is pipelined, and the input image data is processed continuously, and the processing result is output. The processing result in the ISP 400 can be output to the outside via the input / output unit 500, or can be displayed on the display unit 600.
[0053] The input / output unit 500 communicates image data, etc. between the imaging device 1000 and an external device (not shown). As a communication method, it can be a wired means using a cable, etc., or wireless communication without using a cable, etc. Additionally, the input / output unit 500 can receive, in addition to image data, commands including operation instructions, etc., a machine learning model operating in the imaging device 1000, various parameters, and other programs from an external device.
[0054] The display unit 600 includes a display for displaying image data, etc. captured by the imaging device 1000, and also displays a predetermined UI / UX, notifications, etc. in addition to the image data. Additionally, it can be used as an operation unit by providing a touch panel on the display of the display unit 600.
[0055] The CPU 700 is a control unit including a processor that overall controls each block of the imaging device 1000. The CPU 700 realizes various functions by executing a program previously stored in the memory 800. As an example, based on a user instruction from an operation unit (not shown), it controls the switching of the operation mode of the imaging device 1000. The operation modes include a still image mode, a moving image mode, a night scene mode, etc. Additionally, the CPU 700 sets a machine learning model previously stored in the memory 800 in the first machine learning operation unit 200 or the second machine learning operation unit 900 according to the operation mode, thereby controlling machine learning operations. Additionally, the CPU 700 can also be configured to include a sensor control unit as a control unit, and this sensor control unit generates and supplies a clock and a synchronization signal for controlling the imaging device 1000.
[0056] The memory 800 is composed of DRAM or the like, and holds firmware for controlling the entire imaging device 1000, UI data, data related to the operation mode, data related to the machine learning model, etc. in a plurality of storage areas. In the present embodiment, the data related to the machine learning model includes network information, weights, quantization parameters, etc. In addition, the memory 800 includes an area for holding image data, a buffer area during operation, a storage area for holding still images and moving images captured, etc. In the present embodiment, the memory 800 corresponds to a holding component that holds various data and programs including image data.
[0057] The second machine learning operation unit 900 is an operation unit that uses the image data processed by the ISP 400 as input and performs an operation based on a predetermined machine learning model on the input. Figure 6 is a functional block diagram of the second machine learning operation unit 900. The second machine learning operation unit 900 includes a feature extraction unit 910, a second inference unit 920, and an output processing unit 930. Each block performs processing based on a clock signal that is the same as or doubled that of the CPU 700. The machine learning model included in the second machine learning operation unit 900 may also have a network structure different from the U-NET structure, such as a Transformer structure, a recurrent neural network structure, etc.
[0058] The feature extraction unit 910 includes a processing circuit that performs feature extraction processing using a machine learning model. In the present embodiment, it is assembled in a state where learning has been completed based on a previously captured image. For example, in the case of learning to detect an intended object using a machine learning model for object detection in an image, a plurality of labeled images are prepared, and the labeling results are used as training data to detect the object, thereby enabling learning. It should be noted that the machine learning model can also perform learning for pose detection, object recognition, object tracking, reducing jitter in the image or reducing optical aberration to improve the quality of the image, in addition to object detection. It has a plurality of convolutional operation layers L, and operations are sequentially performed in each of them. The result obtained by performing the operation is output as a feature map.
[0059] Figure 7This is a diagram showing each operation layer L in the feature extraction unit 910 of the present embodiment. In the feature extraction unit 910, operations are repeatedly performed using a layer that performs a convolution operation on the input image data and a layer that performs a pooling operation. In the feature quantity extraction unit 910 of the present embodiment, as the operations are performed, the size corresponding to the horizontal and vertical directions of the original image data decreases. On the other hand, the size in the depth direction or the channel direction increases. In the case of performing such operations, in order to appropriately extract the feature quantity, several lines of image data are not sufficient, and the image data of the entire screen is required. Therefore, it is preferable that the second machine learning operation unit 900 uses the image data stored in the memory 800 as input.
[0060] Based on the feature map generated by the feature extraction unit 910, the second inference unit 920 performs an inference operation using a machine learning model to detect whether a predetermined subject is captured in the image data. Specifically, the likelihood of the class that has been learned in advance for the intended detection object is output. As an example of the class that is the detection object, there are a person, a vehicle, etc. A specific object can be used as the detection object, or multiple types can be detected simultaneously. In addition, in addition to the class, the coordinates of the region where the detection object exists can also be output as a bounding box.
[0061] The output processing unit 930 outputs the final detection result based on the likelihood of each class output by the second inference unit 920. Specifically, the class with the highest likelihood among the likelihoods for multiple classes is selected, and this class is used as the final detection result. In addition, when the likelihoods for all classes are lower than a certain value, it is assumed that the detection object is not included in the image data.
[0062] It should be noted that the operations performed by the second machine learning operation unit 900 have a bit precision of 8 bits or more. As an example, they are operations based on 16-bit floating-point numbers. Therefore, it is possible to easily install a machine learning model that can be used in a general environment such as a GPU, and high versatility can be achieved.
[0063] It should be noted that all or part of the functions of the second machine learning operation unit 900 can also be implemented using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field-Programmable Gate Array).
[0064] For example, to constitute the functions of the second machine learning operation unit 900, a processor that executes program processing and an accelerator that executes operations related to a neural network may also be combined. Specifically, a neural network operation accelerator for repeatedly executing convolution operations or the like may be used in combination with a processor.
[0065] Figure 8 FIG. 4 is a flowchart for explaining the imaging operation of the imaging device 1000. Each process in this flowchart is implemented by a processor included in the CPU 700 executing a predetermined program stored in advance in the memory 800 and controlling each block of the imaging device 1000. The operation of this flowchart starts by pressing a power button or starting a predetermined camera application program in the imaging device 1000.
[0066] At the start of the process, in step S800, the CPU 700 determines whether to start the imaging process. Specifically, it waits until a user instruction for imaging is received based on whether there is a transition instruction to the imaging mode or a shooting start instruction such as a shutter button included in an operation unit (not shown). Then, when an instruction to perform imaging is received, the process proceeds to the next step.
[0067] In step S810, the CPU 700 controls the sensor 100 via a sensor control unit (not shown) to start exposure. Specifically, power supply and clock signal supply to the sensor 100, supply of a vertical synchronization signal or a horizontal synchronization signal, and writing of parameters to a control register in the sensor 100 are performed. Here, the control register includes registers related to exposure such as exposure time and gain; and registers related to readout control such as pixel addition and thinning readout. Then, based on the parameters written to the control register, a reset operation of the charge generated in the pixel 110 of the sensor 100 and a light charge accumulation operation are performed, and signals are sequentially read out from each pixel. After the exposure control in the pixel 110 is completed, the process proceeds to the next step.
[0068] The reset operation, light charge accumulation operation, and readout operation of the charge generated in the pixel 110 of the sensor 100 will be further described. In the present embodiment, for the sake of explanation, the signal readout from the sensor 100 is performed by a rolling shutter operation, and signals are read out from the pixel 110 row by row. It should be noted that the readout method and the readout unit are merely examples. As different examples, as the readout method, exposure and readout may be performed by a global shutter operation, and as the readout unit, multiple rows or a predetermined block area may be used as a unit.
[0069] In step S820, the CPU 700 determines the exposure conditions based on the parameters set for the sensor 100 in step S800 and the previously acquired exposure information. More specifically, it determines whether the ISO sensitivity used in imaging is equal to or higher than a predetermined value. In the present embodiment, when the ISO sensitivity used in imaging is greater than ISO 3200, the process proceeds to step S830, and when it is equal to or lower than ISO 3200, the process proceeds to step S840. Note that, as the determination target in this step, instead of the ISO sensitivity, the analog gain value or digital gain value set for the sensor 100, the exposure time, or a combination thereof may be used. Each determination condition is a condition for determining whether the amount of noise included in the signal obtained from each pixel 110 is equal to or greater than a certain level. As an example, conditions such as temperature and the type of the sensor 100 that may cause an increase in the amount of noise may be added for determination.
[0070] In step S830, the CPU 700 controls the first machine learning operation unit 200 to perform arithmetic processing on the signals obtained from each pixel 110 using a machine learning model. The machine learning model in the present embodiment is previously learned to reduce the amount of noise in the signal, and the CPU 700 controls the operation by setting the parameters of the machine learning model in registers and the like included in the first machine learning operation unit 200. Then, after the operation in the first machine learning operation unit 200 ends, the process proceeds to the next step. Note that, in the present embodiment, the parameters of the machine learning model include weight parameters for convolution operation, quantization parameters for quantizing the convolution operation result, commands for controlling each block included in the first machine learning operation unit 200, and the like.
[0071] In the present embodiment, in cases such as exposure conditions where the signals obtained from each pixel 110 contain a large amount of noise, by controlling the first machine learning operation unit 200 to improve the signal quality of each signal, more appropriate image data can be obtained. On the other hand, in cases such as exposure conditions where the noise component included in the signal is small and there is no need to improve the signal quality in the first machine learning operation unit 200, the process subsequent to the first machine learning operation unit 200 is performed without performing the processing in the first machine learning operation unit 200, thereby enabling an improvement in response speed and power saving.
[0072] Note that, in the present embodiment, for the purpose of explanation, an example is shown in which the processing in the first machine learning operation unit 200 is controlled based on whether the ISO sensitivity is a certain value or more, but it is not limited thereto. As another example, the processing in the first machine learning operation unit 200 may be controlled by switching multiple machine learning models learned according to the noise amount or each ISO sensitivity according to the exposure conditions. More specifically, the control may be performed in the following manner: when the ISO sensitivity is between ISO 800 and ISO 3200, a machine learning model learned based on an image with noise equivalent to ISO 1600 superimposed thereon is used, and when the ISO sensitivity is ISO 3200 or more, a machine learning model learned based on an image with noise equivalent to ISO 3200 superimposed thereon is used. Note that three or more machine learning models may be switched, or only a part of parameters such as learning parameters may be switched without switching the machine learning model.
[0073] In addition, when learning has been performed in the machine learning model in a manner of reducing shake, control may also be performed according to the exposure time during which shake may occur. As an example, when the exposure time is longer than 1 / 15 second, control may be performed in a manner of performing processing using the machine learning model.
[0074] In step S840, the CPU 700 controls the ISP 400 to perform predetermined image processing on the image data. In the present embodiment, at least demosaicing processing and encoding compression processing are performed. Through this processing, the ISP 400 processes the RAW image data processed by the first machine learning operation unit 200 to generate compressed and encoded data for data storage or display, and causes the processing to proceed to the next step. As an example of the compressed and encoded data, not only data formats for still images such as the JPEG format and the BMP format, but also data formats for moving images such as the MPEG format, the H.264 format, or the H.265 format may be used. Note that, as the processing of the ISP 400, in order to easily perform the arithmetic processing in the second machine learning operation unit 900, processing such as cropping, size change, deformation, and synthesis may also be performed on the compressed and encoded data.
[0075] In step S850, the CPU 700 controls the second machine learning operation unit 900 to perform operation processing on the compression-encoded data using a machine learning model. The machine learning model in the present embodiment performs an operation for detecting whether a predetermined detection object exists in the image. The CPU 700 controls the operation by setting the parameters of the machine learning model in registers and the like included in the second machine learning operation unit 900. Then, after the operation in the second machine learning operation unit 900 ends, the process proceeds to the next step. It should be noted that, in the present embodiment, the parameters of the machine learning model include weight parameters for convolution operation, commands for controlling each block included in the second machine learning operation unit 900, and the like. In addition, as shown in the present embodiment, when the second machine learning operation unit 900 processes the compression-encoded data, since the amount of image data itself as the processing object is reduced, the storage capacity required in the memory 800 can be suppressed.
[0076] In the present embodiment, when controlling each block of the imaging device 1000 such as the ISP 400 according to the object included in the image data, the second machine learning operation unit 900 can be controlled to appropriately detect the object, and more appropriate control can be achieved. In addition, when controlling each block of the imaging device 1000 according to the detection result, it is also possible to perform control by switching multiple machine learning models. As an example, when controlling the focus position of an optical unit (not shown) according to the detection result, a machine learning model for detecting the distance of the detection object can also be used. In addition, when switching the image processing in the ISP 400 according to the detection object, a machine learning model for detecting the area occupied by the detection object in the image data can also be used. In addition, when using the posture of the detection object as a user interface to control the imaging device 1000, a machine learning model for detecting the posture of the detection object can also be used. In addition, when performing authentication of a person or the like, a machine learning model for detecting at least a part of the human body can also be used.
[0077] In the present embodiment, since the machine learning model used in the second machine learning operation unit 900 needs to implement various functions, in addition to the operation accuracy and operation speed, high versatility can also be cited as the capabilities required for the second machine learning operation unit 900. Therefore, in the second machine learning operation unit 900, circuit redundancy is also required.
[0078] In step S860, the CPU 700 determines whether a detection object is detected as a result of the operation using the machine learning model in the second machine learning operation unit 900. If a detection object is detected, the process proceeds to step S870, and the detection result is displayed on the display unit 600. On the other hand, if a detection object is not detected, the process proceeds to step S880. In the present embodiment, an example is shown in which the operation result of the machine learning model in the second machine learning operation unit 900 is displayed on the display unit 600, but it is not limited thereto. Depending on which module of the imaging device 1000 the operation result of the machine learning model in the second machine learning operation unit 900 is used to control, the processes in step S860 and step S870 can be replaced. It should be noted that in the present embodiment, an example is shown in which the processes of step S850 to step S870 are performed once, but the control can also be performed in a manner of repeating a predetermined number of times.
[0079] In step S880, the CPU 700 determines whether to end the imaging operation. More specifically, the CPU 700 determines whether to end the processing of this flowchart based on the user's imaging end instruction and the end instruction of the application program, and repeatedly executes the processing of this flowchart until an end determination is made.
[0080] As described above, as the imaging device 1000 and its control method of the present embodiment are described with reference to the respective drawings, by including the first machine learning operation unit 200 and the second machine learning operation unit 900 each having different operation components as special features, it is possible to achieve both high processing speed and versatility. Ordinary operations related to machine learning require a large number of multi-bit product-sum operations to be executed in parallel, and even require a large-scale processing device such as a server. As a method for reducing the amount of calculation, there is a method of performing quantization processing. However, if the bit accuracy decreases due to quantization, a new problem of a decrease in calculation accuracy will occur accordingly. In addition, in the operation of machine learning, sometimes the operation accuracy reduction caused by quantization is suppressed by limiting the executed tasks to specific contents and ranges. In other words, these indicate that in edge devices such as embedded devices where there are limitations on power consumption and computing resources, it is a very difficult problem to simultaneously meet the requirements of versatility for accurately executing various operations related to machine learning and the requirements of high computing efficiency for suppressing circuit scale and power consumption.
[0081] Regarding the issue of balancing generality and operation efficiency, in the imaging device 1000 of the present embodiment, a first machine learning operation unit 200 dedicated to processing signals obtained from each pixel on a per-pixel basis is pipelined between the sensor 100 and the ISP 400 to improve the operation efficiency. Moreover, a second machine learning operation unit with higher generality is arranged at the subsequent stage of the ISP 400. In other words, the first machine learning operation unit 200 performs processing based on a machine learning model including quantization operations in units output from the sensor 100 based on a synchronization signal, thereby suppressing the memory consumption and performing low-latency and high-efficiency operations. In addition, by limiting the tasks performed by the machine learning model to per-pixel processing such as noise reduction, the reduction in operation accuracy caused by the quantization operation can be suppressed. Further, since the signal output from the sensor 100 contains multi-bit information, it is suitable for performing image processing related to image quality improvement. Also, by further including the second machine learning operation unit 900, the generality for handling tasks in various machine learning models can be maintained as a whole.
[0082] (Second Embodiment)
[0083] In the first embodiment, an example in which the first machine learning operation unit 200 and the internal bus IB are connected via the sensor I / F 300 is shown. Figure 9 FIG. is a functional block diagram of the imaging device 1100 according to the second embodiment. The same reference numerals are used to denote the same structures as those in the imaging device 1000 in the first embodiment, and the description thereof may be omitted sometimes.
[0084] In the imaging device 1100, the difference from the imaging device 1000 in the first embodiment lies in the connection manner between the first machine learning operation unit 200 and the internal bus IB. The sensor 100 and the first machine learning operation unit 200 are connected in the same manner as in the first embodiment by a high-speed multi-channel communication method via the high-speed I / F 130 in the sensor 100. On the other hand, the first machine learning operation unit 200 is connected to each functional block of the imaging device 1100 through the internal bus IB capable of high-speed communication. In other words, in the present embodiment, the first machine learning operation unit 200, the ISP 400, the input / output unit 500, the display unit 600, the CPU 700, the memory 800, and the second machine learning operation unit 900 are formed on the same silicon chip and are respectively connected to the high-speed internal bus IB.
[0085] Figure 10 FIG. is a timing chart for explaining the operations of the buffer 210 and the ISP 400 in the present embodiment. An example of buffering pixel values of one row is shown in the present embodiment. In Figure 10In A, it represents the data output timing of the pixel values output from the sensor 100. Synchronized with the timing of the horizontal synchronization signal, the pixel value data is sequentially output during period Ta. Then, in Figure 10 In B, it represents the data output timing of the pixel values output from the buffer 210. Synchronized with the timing of the horizontal synchronization signal, the pixel value data is sequentially output during period Tb. At the subsequent stage of the buffer 210, the slower the data rate of the processed data, the more power reduction can be achieved. Therefore, the data rate of the data read from the buffer 210 is slower than the data rate at the time of input. In Figure 10 In C, it represents the execution timing of the image processing for the image data in the ISP 400. The result processed in the first machine learning operation unit 200 is input to the ISP 400 with a delay compared to the timing shown in Figure 10 B (period Tc1). Then, the image data input during period Tc2 is sequentially processed in a pipeline manner.
[0086] In Figure 10 , an example in which the buffer 210 holds one line of pixel values is described, but it is not limited thereto. When multiple lines of data are required in the first inference unit 230 or the ISP 400, multiple lines can also be held. For example, in the first inference unit 230, when a 3×3 weight is used for the operation, etc., at least 3 lines can also be held. Additionally, when 7 lines of image data are required in the ISP 400, etc., at least 7 lines can also be held.
[0087] As Figure 10 shown, by directly connecting the first machine learning operation unit 200 to the internal bus IB, the image data can be processed in a pipeline manner. Thereby, the overall processing rate in the imaging device 1100 can be improved.
[0088] The second embodiment of the present invention has been described in detail with reference to the accompanying drawings, but the specific structure is not limited to this embodiment, and also includes design changes and the like within the scope not departing from the gist of the present invention. Additionally, the components shown in the above embodiments and modification examples can be appropriately combined to form.
[0089] (Third Embodiment)
[0090] Figure 11 FIG. is a functional block diagram showing the imaging device 1200 of the third embodiment. In the description of the imaging device 1200, the same structural elements as those of the imaging device 1000 or the imaging device 1100 are denoted by the same reference numerals, and the description thereof may be omitted sometimes. It should be noted that the imaging device 1200 of the third embodiment can be applied to, for example, an imaging device having a stacked CMOS image sensor or the like.
[0091] The imaging device 1200 differs from the imaging devices 1000 and 1100 in that the sensor 100 and the first machine learning operation unit 210 are respectively and independently connected to the internal bus IB. The first machine learning operation unit 210 is a modified example of the first machine learning operation unit 200. In addition, the imaging device 1200 differs from the imaging devices 1000 and 1100 in that it is provided with a sensor I / F 310 and an I / F 320 instead of the sensor I / F 300.
[0092] The sensor I / F 310 stores the information (RAW image data) output by the sensor 100 into a predetermined storage unit via the internal bus IB. The predetermined storage unit may be, for example, the memory 800.
[0093] The first machine learning operation unit 210 takes the RAW image data stored in a predetermined storage unit, such as the memory 800, as an input, and performs an operation based on a predetermined machine learning model on this input. In addition, the first machine learning operation unit 210 may directly take the RAW image data as an input via the sensor I / F 310 without using the memory 800 or the like, and perform an operation based on a predetermined machine learning model on this input. The internal structure of the first machine learning operation unit 210 may be the same as that of the first machine learning operation unit 200.
[0094] The I / F 320 outputs the output of the first machine learning operation unit 210 to the internal bus IB. The data output to the internal bus IB is stored in the memory 800, for example. The I / F 320 may also output to the ISP 400 in order to perform predetermined image processing on the data obtained from the first machine learning operation unit 210. It should be noted that the I / F 320 may also be provided with a buffer for temporarily holding data in order to perform data format conversion or the like. It should be noted that the I / F 320 is configured to be connected to the internal bus IB, but as long as a sufficiently high-speed and low-latency communication method such as a high-speed serial communication standard can be used, it may also be configured to be connected to the internal bus IB via a predetermined communication I / F.
[0095] Here, as a specific implementation means of the structure of the present embodiment, it may also be installed as three ICs (Integrated Circuit) as shown in the figure. The integrated circuit IC1 has the functions of the sensor 100 and the sensor I / F 310. The integrated circuit IC2 has the functions of the first machine learning operation unit 200 and the I / F 320. The integrated circuit IC3 has the functions of the ISP 400, the input / output unit 500, the display unit 600, the CPU 700, the memory 800, and the second machine learning operation unit 900. The integrated circuit IC1, the integrated circuit IC2, and the integrated circuit IC3 may also be electrically connected to each other via the internal bus IB so as to be able to perform data communication.
[0096] (Fourth Embodiment)
[0097] Figure 12 FIG. is a functional block diagram showing a camera device 1300 according to the fourth embodiment. In the description of the camera device 1300, the same reference numerals are used for the same structures as those of the camera device 1000, and the description thereof may be omitted. The difference between the camera device 1300 and the camera device 1000 is that it includes a sensor 110, a third machine learning operation unit 210, a sensor I / F 320, and a sensor control unit 330. It should be noted that the camera device 1300 according to the fourth embodiment can be applied to, for example, a control device for controlling a vehicle. In a control device for controlling a vehicle or the like, there is a need to minimize the occupation time of the internal bus IB.
[0098] The sensor 110 is a sensor that captures images, like the sensor 100, but is a different sensor from the sensor 100. That is, the camera device 1300 has multiple sensors. When the camera device 1300 is applied to a control device for controlling a vehicle, the sensor 100 is a sensor that captures the front of the vehicle, and the sensor 110 can be either a sensor that captures the rear of the vehicle or can overlap with the field of view angle part captured by the sensor 100. It should be noted that the camera device 1300 may also have multiple other sensors not shown. The output data of the sensor 110 is not output to the internal bus IB, but is output to the third machine learning operation unit 220.
[0099] The third machine learning operation unit 220 takes the information (RAW image data) output by the sensor 110 as input and performs operations based on a predetermined machine learning model. The third machine learning operation unit 220 can, for example, have the same structure as the second machine learning operation unit 900, and can also perform object detection, pose detection, object recognition, object tracking, reduction of jitter in an image, reduction processing of optical aberration, etc. based on feature amounts and the like.
[0100] The sensor I / F 320 outputs the output result of the third machine learning operation unit 220 to the sensor control unit 330. The sensor control unit 330 outputs the result obtained by the operation of the third machine learning operation unit 220 to the internal bus IB. Here, for example, when the third machine learning operation unit 220 performs object detection, the result obtained by the operation of the third machine learning operation unit 220 may also be information such as a class and a likelihood. That is, according to the present embodiment, only information such as a class and a likelihood obtained as a result of performing machine learning is output, and thus image data is not output to the internal bus IB. Therefore, according to the present embodiment, the capacity of the data output to the internal bus IB can be reduced, and the occupation time of the internal bus IB can be shortened. Therefore, according to the present embodiment, even when a plurality of sensors are provided, it is possible to suppress a situation where the internal bus IB becomes congested and a delay occurs.
[0101] (Modification Example 1)
[0102] For example, in the first machine learning operation unit 200 and the second machine learning operation unit 900 described in the above embodiment, the data to be operated on is not limited to a single form, and can be composed of still images, moving images, sounds, texts, numerical values, and combinations thereof. It should be noted that the data input to the first machine learning operation unit 200 and the second machine learning operation unit 900 may also be combined with measurement results of physical quantity measuring devices such as optical sensors, thermometers, global positioning system (GPS) measuring devices, angular velocity measuring devices, and anemometers. Different information such as base station information received from peripheral devices via wired or wireless communication, information of vehicles / ships, weather information, information related to congestion status, financial information, and personal information may also be combined.
[0103] (Modification Example 2)
[0104] As the imaging device 1000 or the imaging device 1100, it is assumed to be a communication device such as a mobile phone driven by a battery or the like, an intelligent device such as a personal computer, a digital video camera, a game device, a mobile device such as a robot product, but is not limited thereto. By using the peak power supply limit in Power over Ethernet (PoE) or the like, reducing product heat generation, or a product with a high requirement for long-term driving, effects not obtained in other existing examples can also be obtained.
[0105] For example, by applying in-vehicle cameras mounted on vehicles, ships, etc., surveillance cameras installed in public facilities, on roads, etc., not only can long-term shooting be achieved, but it also helps with weight reduction and high durability. In addition, by applying to display devices such as TVs and monitors, medical devices such as medical cameras or surgical robots, and work robots used at manufacturing sites or construction sites, the same effects can also be achieved.
[0106] (Variant Example 3)
[0107] One or more processors can also be used to implement part or all of the first machine learning operation unit 200 and the second machine learning operation unit 900. For example, the first machine learning operation unit 200 and the second machine learning operation unit 900 can also implement part or all of the input layer or the output layer through software processing based on a processor. Part of the input layer or output layer implemented through software processing is, for example, the normalization and conversion of data. Thus, various input forms or output forms can be handled. It should be noted that the software executed by the processor can also be configured to be rewritten using a communication component or an external medium.
[0108] (Variant Example 4)
[0109] Part of the processing in the second machine learning operation unit 900 can also be implemented by combining graphics processing units (GPUs) on the cloud, etc. The second machine learning operation unit 900 further performs processing on the cloud in addition to the processing performed by the imaging device 1000 or the imaging device 1100, or performs processing in addition to the processing on the cloud, thereby enabling more complex processing to be achieved with fewer resources.
[0110] (Variant Example 5)
[0111] There are differences between the first machine learning operation unit 200 and the second machine learning operation unit 900 in terms of whether quantization operations are included. Therefore, for the machine learning models operating in each of them, different learning methods can also be used. As an example, a machine learning model operating in the first machine learning operation unit 200 preferably adopts a method (hereinafter referred to as the QAT method) including the following learning steps. The learning steps are to perform learning in a form including quantization operations after generating a network including quantization operations. By performing learning in the QAT method in this way, the reduction in operation accuracy caused by quantization can be reduced. On the other hand, the QAT method requires the design of learning methods, learning parameters, etc., so there is a case where the generality is reduced. Therefore, for the machine learning model operating in the second machine learning operation unit 900, it is preferable not to use the QAT method for learning. In this way, it is preferable to determine the learning method of the machine learning model according to whether it is used in one of the first machine learning operation unit 200 and the second machine learning operation unit 900.
[0112] In addition, the effects described in this specification are merely illustrative or exemplary effects and are not limiting effects. That is, the technology related to the present disclosure can achieve other effects that are clear to those skilled in the art from the description of this specification in addition to or instead of the above effects.
[0113] Description of Reference Numerals
[0114] 100: Image sensor; 200: First machine learning operation unit; 300: Sensor I / F; 400: ISP; 500: Input / output unit; 600: Display unit; 700: CPU; 800: Memory; 900: Second machine learning operation unit; 1000: Imaging device of the first embodiment; 1100: Imaging device of the second embodiment.
Claims
1. An imaging device includes an image sensor having a plurality of pixels for converting a subject image into an electrical signal. The imaging device is characterized by comprising: a first machine learning arithmetic component configured to process signals obtained from the plurality of pixels using a first machine learning model including a plurality of arithmetic layers; a determination component configured to determine whether to perform processing by the first machine learning arithmetic component; and a second machine learning arithmetic component configured to process based on the signals or processing results obtained by the first machine learning arithmetic component using a second machine learning model different from the first machine learning model, wherein the first machine learning model at least includes a quantization arithmetic layer configured to perform a quantization operation on an operation result of a convolution arithmetic layer that performs a convolution operation.
2. The imaging device according to claim 1, wherein, the second machine learning model has a network structure different from that of the first machine learning model and at least includes a convolution arithmetic layer that performs a convolution operation and a pooling layer that performs a pooling operation on a convolution operation result.
3. The imaging device according to claim 1, wherein, the imaging device further includes a control component configured to generate a synchronization signal for controlling the image sensor, and perform processing in the first machine learning arithmetic component synchronously with the synchronization signal.
4. The imaging device according to claim 1, wherein, the quantization arithmetic layer included in the first machine learning model quantizes the operation result of the convolution arithmetic layer into a value of 8 bits or less.
5. The imaging device according to claim 1, wherein, the image sensor further includes a conversion component configured to convert an analog signal in the plurality of pixels into a digital signal, and the quantization arithmetic layer included in the first machine learning model quantizes the operation result of the convolution arithmetic layer into a value not exceeding the resolution of the conversion component.
6. The imaging device according to claim 1, wherein, the imaging device further includes an image processing component configured to perform predetermined image processing on the signals processed by the first machine learning arithmetic component, and the predetermined image processing performed in the image processing component at least includes a demosaicing process and an encoding and compression process.
7. The imaging device according to claim 6, wherein, the first machine learning model performs an inference operation for reducing noise included in the signals of the plurality of pixels, and the second machine learning model performs processing for detecting a predetermined detection object in the image data that is a result of the image processing component.
8. The imaging device according to claim 7, wherein, the imaging device further includes a display component configured to display the image data that is a result of the image processing component, and the display component displays a detection result of the detection object in the second machine learning model corresponding to the image data displayed on the display component.
9. The imaging device according to claim 6, wherein, The imaging device further includes a holding component for holding image data that is the result of the image processing component. The first machine learning operation component performs an inference operation on signals of the plurality of pixels in a predetermined unit output by the image sensor. The second machine learning operation component performs processing for detecting a predetermined detection target in units of the image data held in the holding component.
10. The imaging device according to claim 9, wherein, in the predetermined unit, there are 1500 or more pixel values of 8 bits or more, the first machine learning operation component and the image processing component perform processing sequentially in a pipeline based on the unit.
11. The imaging device according to claim 9 or 10, wherein, the first machine learning operation component further includes a buffer component that temporarily holds signals of the plurality of pixels in a predetermined unit output by the image sensor, the data rate at which the first machine learning operation component reads out the signals of the plurality of pixels temporarily held in the buffer component for processing using the first machine learning model is slower than the data rate of the signals of the plurality of pixels when input to the buffer component.
12. The imaging device according to claim 1, wherein, the first machine learning operation component further includes a switching component for switching between a plurality of machine learning models, the first machine learning operation component switches the machine learning model based on the exposure condition in the image sensor.
13. The imaging device according to claim 12, wherein, the second machine learning operation component further includes a switching component for switching between a plurality of machine learning models, the plurality of machine learning models switched in the second machine learning operation component include a machine learning model for detecting at least a part of a human body.
14. The imaging device according to claim 1, wherein, the first machine learning model in the first machine learning operation component is a learned machine learning model that has been pre-learned in an external device using a learning method different from the second machine learning model, the learning method includes a learning step performed in a form including quantization operations in the first machine learning model.
15. The imaging device according to claim 1, wherein, the imaging device further includes: a sensor for acquiring an image; and a third machine learning operation component for performing predetermined detection processing based on the image acquired by the sensor, the third machine learning operation component outputs a detection result.
16. A control method for an imaging device, the imaging device including an image sensor having a plurality of pixels for converting a subject image into an electrical signal, the control method for the imaging device is characterized in that it includes: a first machine learning operation step for processing signals of the plurality of pixels using a first machine learning model including a plurality of operation layers; a determination step for determining whether to perform processing through the first machine learning operation step; and A second machine learning operation step, which is used to perform processing based on the signal or the processing result obtained through the first machine learning operation step by using a second machine learning model different from the first machine learning model. The first machine learning model at least includes a quantization operation step for quantizing the operation result of a convolution operation layer that performs a convolution operation.
Citation Information
Patent Citations
Information processing method, information processing device and program
JP2018077829A
Drug delivery device having a sensing system
JP2022174141A