Information processing device, solid-state imaging device, and information processing method
The information processing device optimizes image recognition by using a learning model to sequentially process pixel signals and terminate processing based on conditions, addressing the inefficiencies in conventional methods.
Patent Information
- Application Number
- JP2024006216
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-08-31
- Filing Date
- 2024-01-18
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2039-08-30
AI Technical Summary
Conventional image recognition in imaging devices requires processing multiple frames of image data, leading to increased processing time and power consumption.
An information processing device that utilizes a recognition unit with a learning model to process pixel signals sequentially, terminating the recognition process if a predetermined condition is met, thereby reducing processing time and power consumption.
The solution significantly reduces processing time and power consumption by optimizing the image recognition process in imaging devices.
Smart Images

Figure 0007740396000001 
Figure 0007740396000002 
Figure 0007740396000003
Abstract
Description
[Technical Field]
[0001] The present disclosure provides: information processing device, solid-state image sensor Child Call information Regarding the processing method. [Background technology]
[0002] In recent years, as imaging devices such as digital still cameras, digital video cameras, and compact cameras installed in multi-function mobile phones (smartphones) have become more powerful, imaging devices equipped with image recognition functions that can recognize specific objects included in captured images have been developed. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-112409 Summary of the Invention [Problem to be solved by the invention]
[0004] However, conventionally, image processing of one to several frames of image data is required to execute the image recognition function, which poses a problem of increased processing time and power consumption to realize the function.
[0005] The present disclosure provides a method for reducing the processing time and power consumption required to realize a function. information processing device, solid-state image sensor Child Call information The present invention aims to provide a processing method. [Means for solving the problem]
[0006] An information processing device according to the present disclosure is an information processing device that processes pixel signals read out from pixels included in a pixel region in which a plurality of pixels are arranged, and includes a recognition unit that performs recognition processing on the pixel signals corresponding to the readout unit set as part of the pixel region using a learning model that is trained based on teacher data for the readout unit, and outputs a recognition result of the recognition processing. The recognition unit performs the recognition process on the pixel signals sequentially read out from the pixels included in the pixel area for each read unit, and if the recognition result does not satisfy a predetermined condition, continues to perform the recognition process on the pixel signals of the read unit, and if the recognition result satisfies the predetermined condition, ends the recognition process. . [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 is a block diagram illustrating a configuration of an example of an imaging device applicable to each embodiment of the present disclosure. [Figure 2A] FIG. 2 is a schematic diagram illustrating an example of a hardware configuration of an imaging device according to each embodiment. [Figure 2B] FIG. 2 is a schematic diagram illustrating an example of a hardware configuration of an imaging device according to each embodiment. [Figure 3A] 10 is a diagram showing an example in which the imaging device according to each embodiment is formed using a stacked CIS with a two-layer structure. FIG. [Figure 3B] FIG. 10 is a diagram showing an example in which the imaging device according to each embodiment is formed using a stacked CIS with a three-layer structure. [Figure 4] FIG. 2 is a block diagram showing a configuration of an example of a sensor unit applicable to each embodiment. [Figure 5A] FIG. 1 is a schematic diagram illustrating a rolling shutter system. [Figure 5B] FIG. 1 is a schematic diagram illustrating a rolling shutter system. [Figure 5C] FIG. 1 is a schematic diagram illustrating a rolling shutter system. [Figure 6A] FIG. 10 is a schematic diagram for explaining line thinning in the rolling shutter method. [Figure 6B] FIG. 10 is a schematic diagram for explaining line thinning in the rolling shutter method. [Figure 6C] FIG. 10 is a schematic diagram for explaining line thinning in the rolling shutter method. [Figure 7A]10A and 10B are diagrams illustrating an example of another imaging method using the rolling shutter system. [Figure 7B] 10A and 10B are diagrams illustrating an example of another imaging method using the rolling shutter system. [Figure 8A] FIG. 1 is a schematic diagram for explaining a global shutter system. [Figure 8B] FIG. 1 is a schematic diagram for explaining a global shutter system. [Figure 8C] FIG. 1 is a schematic diagram for explaining a global shutter system. [Figure 9A] 1A and 1B are diagrams illustrating examples of sampling patterns that can be realized in a global shutter system. [Figure 9B] 1A and 1B are diagrams illustrating examples of sampling patterns that can be realized in a global shutter system. [Figure 10] FIG. 1 is a diagram for explaining an outline of image recognition processing by CNN. [Figure 11] FIG. 1 is a diagram for explaining an outline of an image recognition process for obtaining a recognition result from a part of an image to be recognized. [Figure 12A] FIG. 10 is a diagram illustrating an example of a classification process using a DNN when time-series information is not used. [Figure 12B] FIG. 10 is a diagram illustrating an example of a classification process using a DNN when time-series information is not used. [Figure 13A] FIG. 1 is a diagram schematically illustrating a first example of a classification process using a DNN when time-series information is used. [Figure 13B] FIG. 1 is a diagram schematically illustrating a first example of a classification process using a DNN when time-series information is used. [Figure 14A] FIG. 10 is a diagram schematically illustrating a second example of a classification process using a DNN when time-series information is used. [Figure 14B] FIG. 10 is a diagram schematically illustrating a second example of a classification process using a DNN when time-series information is used. [Figure 15A] 10A and 10B are diagrams for explaining the relationship between the frame drive speed and the readout amount of pixel signals. [Figure 15B] 10A and 10B are diagrams for explaining the relationship between the frame drive speed and the readout amount of pixel signals. [Figure 16] FIG. 2 is a schematic diagram for explaining a recognition process according to each embodiment of the present disclosure. [Figure 17] 4 is a flowchart illustrating an example of a recognition process performed by a recognition processing unit according to the first embodiment. [Figure 18] FIG. 2 is a diagram showing an example of image data for one frame. [Figure 19] FIG. 2 is a diagram illustrating the flow of machine learning processing executed by a recognition processing unit according to the first embodiment. [Figure 20A] FIG. 4 is a schematic diagram for explaining an application example of the first embodiment. [Figure 20B] FIG. 4 is a schematic diagram for explaining an application example of the first embodiment. [Figure 21] FIG. 10 is a functional block diagram illustrating an example of functions of an imaging device according to a second embodiment. [Figure 22] 10 is a schematic diagram illustrating in more detail an example of processing in a recognition processing unit according to the second embodiment. FIG. [Figure 23] FIG. 10 is a functional block diagram illustrating an example of functions according to a second embodiment. [Figure 24] FIG. 10 is a schematic diagram for explaining a frame readout process according to the second embodiment. [Figure 25] FIG. 10 is a schematic diagram illustrating a recognition process according to a second embodiment. [Figure 26] FIG. 10 is a diagram illustrating an example in which recognition processing is terminated midway through frame reading. [Figure 27] FIG. 10 is a diagram illustrating an example in which recognition processing is terminated midway through frame reading. [Figure 28] 10 is a flowchart illustrating an example of a recognition process according to a second embodiment. [Figure 29A] 10 is a time chart illustrating an example of control of a readout and recognition process according to the second embodiment. [Figure 29B]10 is a time chart illustrating an example of control of a readout and recognition process according to the second embodiment. [Figure 30] 10 is a time chart illustrating another example of control of the readout and recognition process according to the second embodiment. [Figure 31] FIG. 10 is a schematic diagram for explaining a frame readout process according to a first modified example of the second embodiment. [Figure 32] FIG. 10 is a schematic diagram for explaining a frame readout process according to a second modified example of the second embodiment. [Figure 33] FIG. 10 is a schematic diagram for explaining a frame readout process according to a third modified example of the second embodiment. [Figure 34] FIG. 10 is a schematic diagram illustrating a recognition process according to a third modified example of the second embodiment. [Figure 35] 10A and 10B are diagrams for explaining an example in which recognition processing is terminated midway through frame reading when the reading unit is an area. [Figure 36] 10A and 10B are diagrams for explaining an example in which recognition processing is terminated midway through frame reading when the reading unit is an area. [Figure 37] FIG. 10 is a schematic diagram for explaining a frame readout process according to a fourth modified example of the second embodiment. [Figure 38] FIG. 10 is a schematic diagram illustrating a recognition process applicable to a fourth modified example of the second embodiment. [Figure 39] 10 is a time chart illustrating an example of readout and control according to a fourth modified example of the second embodiment. [Figure 40] FIG. 10 is a diagram for more specifically explaining the frame readout process according to the fourth modified example of the second embodiment. [Figure 41] FIG. 10 is a diagram for more specifically explaining the frame readout process according to the fourth modified example of the second embodiment. [Figure 42] FIG. 11 is a schematic diagram for explaining a frame readout process according to a fifth modified example of the second embodiment. [Figure 43]FIG. 13 is a schematic diagram for explaining a frame readout process according to a sixth modified example of the second embodiment. [Figure 44] FIG. 13 is a diagram showing an example of a pattern for performing a readout and recognition process according to a sixth modified example of the second embodiment. [Figure 45] FIG. 13 is a schematic diagram for explaining a first example of a frame readout process according to a seventh modified example of the second embodiment. [Figure 46] FIG. 13 is a schematic diagram for explaining a frame readout process according to a first other example of the seventh modified example of the second embodiment. [Figure 47] FIG. 13 is a schematic diagram for explaining a frame readout process according to a second alternative example of the seventh modified example of the second embodiment. [Figure 48] FIG. 13 is a schematic diagram for explaining a frame readout process according to a third alternative example of the seventh modified example of the second embodiment. [Figure 49] FIG. 13 is a functional block diagram illustrating an example of functions according to an eighth modified example of the second embodiment. [Figure 50] 13 is a flowchart illustrating an example of a recognition process according to an eighth modified example of the second embodiment. [Figure 51A] FIG. 13 is a diagram for explaining a first example of a readout and recognition process according to an eighth modified example of the second embodiment. [Figure 51B] FIG. 13 is a diagram for explaining a first example of a readout and recognition process according to an eighth modified example of the second embodiment. [Figure 52] FIG. 13 is a diagram showing an example of an exposure pattern according to a ninth modified example of the second embodiment. [Figure 53] FIG. 20 is a functional block diagram illustrating an example of functions according to a tenth modification of the second embodiment. [Figure 54] 13 is a flowchart illustrating an example of processing according to a tenth modified example of the second embodiment. [Figure 55] FIG. 20 is a schematic diagram for explaining a first process according to a tenth modification of the second embodiment. [Figure 56]FIG. 20 is a schematic diagram for explaining a second process according to a tenth modification of the second embodiment. [Figure 57A] FIG. 20 is a schematic diagram for explaining a third process according to a tenth modification of the second embodiment. [Figure 57B] FIG. 20 is a schematic diagram for explaining a third process according to a tenth modification of the second embodiment. [Figure 58A] FIG. 20 is a schematic diagram for explaining a third process according to a tenth modification of the second embodiment. [Figure 58B] FIG. 20 is a schematic diagram for explaining a third process according to a tenth modification of the second embodiment. [Figure 59A] FIG. 20 is a schematic diagram for explaining a third process according to a tenth modification of the second embodiment. [Figure 59B] FIG. 20 is a schematic diagram for explaining a third process according to a tenth modification of the second embodiment. [Figure 60] FIG. 20 is a schematic diagram illustrating in more detail an example of processing in a recognition processing unit according to a tenth modification of the second embodiment. [Figure 61] FIG. 11 is a functional block diagram illustrating an example of functions according to a third embodiment. [Figure 62] FIG. 11 is a schematic diagram showing an example of a read unit pattern applicable to the third embodiment. [Figure 63] FIG. 11 is a schematic diagram showing an example of a reading order pattern applicable to the third embodiment. [Figure 64] FIG. 13 is a functional block diagram illustrating an example of functions according to a first modified example of the third embodiment. [Figure 65] FIG. 11 is a schematic diagram for explaining a first setting method of a first modified example of the third embodiment. [Figure 66] FIG. 10 is a schematic diagram for explaining a second setting method of the first modified example of the third embodiment. [Figure 67] FIG. 10 is a schematic diagram for explaining a third setting method of the first modified example of the third embodiment. [Figure 68]FIG. 13 is a functional block diagram illustrating an example of a function of an imaging device according to a third modified example of the third embodiment. [Figure 69] FIG. 10 is a functional block diagram illustrating an example of functions of an imaging device according to a fourth embodiment. [Figure 70] FIG. 10 is a schematic diagram for explaining image processing according to a fourth embodiment. [Figure 71] FIG. 13 is a diagram illustrating an example of a read process according to the fourth embodiment. [Figure 72] 13 is a flowchart illustrating an example of processing according to the fourth embodiment. [Figure 73] FIG. 13 is a diagram illustrating a third example of control of an image data storage unit according to the fourth embodiment. [Figure 74] FIG. 10 is a diagram for explaining a first modified example of the fourth embodiment. [Figure 75] FIG. 10 is a diagram for explaining a second modified example of the fourth embodiment. [Figure 76] FIG. 10 is a diagram for explaining a first example of a third modified example of the fourth embodiment. [Figure 77] 1A and 1B are diagrams illustrating an example of use of an imaging device to which the technology of the present disclosure is applied. [Figure 78] 1 is a block diagram showing an example of a schematic configuration of a vehicle control system; [Figure 79] FIG. 2 is an explanatory diagram showing an example of the installation positions of an outside-vehicle information detection unit and an imaging unit. DETAILED DESCRIPTION OF THE INVENTION
[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are denoted by the same reference numerals, and redundant description will be omitted.
[0009] Hereinafter, embodiments of the present disclosure will be described in the following order. 1. Configuration Examples According to Each Embodiment of the Present Disclosure 2. Examples of existing technologies applicable to this disclosure 2-1. Overview of rolling shutter 2-2.Global Shutter Overview 2-3.About DNN (Deep Neural Network) 2-3-1. Overview of CNN (Convolutional Neural Network) 2-3-2. Overview of RNN (Recurrent Neural Network) 2-4. Drive speed 3. Overview of this Disclosure 4. First Embodiment 4-1. Example of operation by the recognition processing unit 4-2. Specific examples of operations performed by the recognition processing unit 4-3. Application example of the first embodiment 5. Second embodiment 5-0-1. Configuration example according to the second embodiment 5-0-2. Example of processing in the recognition processing unit according to the second embodiment 5-0-3. Details of Recognition Processing According to the Second Embodiment 5-0-4. Example of control of readout and recognition processing according to the second embodiment 5-1. First Modification of the Second Embodiment 5-2. Second Modification of the Second Embodiment 5-3. Third Modification of the Second Embodiment 5-4. Fourth Modification of the Second Embodiment 5-5. Fifth Modification of the Second Embodiment 5-6. Sixth Modification of the Second Embodiment 5-7. Seventh Modification of the Second Embodiment 5-8. Eighth Modification of the Second Embodiment 5-9. Ninth Modification of the Second Embodiment 5-10. Tenth Modification of the Second Embodiment 6. Third Embodiment 6-0-1. How to determine the read unit pattern and read order pattern 6-0-1-1. Examples of read unit patterns and read order patterns 6-0-1-2.Example of how to set the priority of read unit patterns 6-0-1-3.Example of how to set the priority of the reading order pattern 6-1. First Modification of the Third Embodiment 6-2. Second Modification of the Third Embodiment 6-3. Third Modification of the Third Embodiment 7. Fourth Embodiment 7-1. First Modification of the Fourth Embodiment 7-2. Second Modification of the Fourth Embodiment 7-3. Third Modification of the Fourth Embodiment 8. Fifth Embodiment
[0010] 1. Configuration Examples According to Each Embodiment of the Present Disclosure The configuration of an imaging device according to the present disclosure will be described briefly. FIG. 1 is a block diagram showing the configuration of an example of an imaging device applicable to each embodiment of the present disclosure. In FIG. 1, the imaging device 1 includes a sensor unit 10, a sensor control unit 11, a recognition processing unit 12, a memory 13, a visual recognition processing unit 14, and an output control unit 15, and is a CMOS image sensor (CIS) in which each of these units is integrally formed using a CMOS (Complementary Metal Oxide Semiconductor). Note that the imaging device 1 is not limited to this example, and may be another type of optical sensor, such as an infrared light sensor that captures images using infrared light.
[0011] The sensor unit 10 outputs pixel signals corresponding to light irradiated onto the light-receiving surface via the optical unit 30. More specifically, the sensor unit 10 has a pixel array in which pixels, each including at least one photoelectric conversion element, are arranged in a matrix. The pixels arranged in a matrix in the pixel array form a light-receiving surface. The sensor unit 10 further includes a drive circuit for driving each pixel included in the pixel array, and a signal processing circuit for performing predetermined signal processing on signals read from each pixel and outputting the processed signals as pixel signals for each pixel. The sensor unit 10 outputs the pixel signals of each pixel included in the pixel area as digital image data.
[0012] Hereinafter, the region in the pixel array of the sensor unit 10 where pixels effective for generating pixel signals are arranged will be referred to as a frame. Frame image data is formed from pixel data based on each pixel signal output from each pixel included in the frame. Each row in the pixel array of the sensor unit 10 will be referred to as a line, and line image data will be formed from pixel data based on the pixel signals output from each pixel included in the line. Furthermore, the operation of the sensor unit 10 to output pixel signals in response to light irradiated onto the light-receiving surface will be referred to as imaging. The sensor unit 10 controls the exposure during imaging and the gain (analog gain) for the pixel signals according to an imaging control signal supplied from the sensor control unit 11, which will be described later.
[0013] The sensor control unit 11 is configured by, for example, a microprocessor, controls the reading of pixel data from the sensor unit 10, and outputs pixel data based on each pixel signal read from each pixel included in a frame. The pixel data output from the sensor control unit 11 is passed to the recognition processing unit 12 and the visual recognition processing unit 14.
[0014] The sensor control unit 11 also generates an imaging control signal for controlling imaging in the sensor unit 10. The sensor control unit 11 generates the imaging control signal, for example, in accordance with instructions from the recognition processing unit 12 and the visual recognition processing unit 14, which will be described later. The imaging control signal includes information indicating the exposure and analog gain used when imaging in the sensor unit 10, as described above. The imaging control signal further includes control signals (such as a vertical synchronization signal and a horizontal synchronization signal) used by the sensor unit 10 to perform imaging operations. The sensor control unit 11 supplies the generated imaging control signal to the sensor unit 10.
[0015] The optical unit 30 is used to irradiate the light receiving surface of the sensor unit 10 with light from the subject, and is disposed, for example, at a position corresponding to the sensor unit 10. The optical unit 30 includes, for example, a plurality of lenses, an aperture mechanism for adjusting the size of an aperture for incident light, and a focus mechanism for adjusting the focus of light irradiated onto the light receiving surface. The optical unit 30 may further include a shutter mechanism (mechanical shutter) for adjusting the time for which light is irradiated onto the light receiving surface. The aperture mechanism, focus mechanism, and shutter mechanism of the optical unit 30 can be controlled, for example, by the sensor control unit 11. Alternatively, the aperture and focus of the optical unit 30 can be controlled from outside the imaging device 1. The optical unit 30 can also be configured integrally with the imaging device 1.
[0016] The recognition processing unit 12 performs a recognition process of an object included in an image using pixel data based on the pixel data passed from the sensor control unit 11. In the present disclosure, for example, a DSP (Digital Signal Processor) reads and executes a program that has been learned in advance using teacher data and stored as a learning model in the memory 13, thereby configuring the recognition processing unit 12 as a machine learning unit that performs recognition processing using a DNN (Deep Neural Network). The recognition processing unit 12 can instruct the sensor control unit 11 to read pixel data required for the recognition processing from the sensor unit 10. The recognition result by the recognition processing unit 12 is passed to the output control unit 15.
[0017] The visual recognition processing unit 14 processes the pixel data passed from the sensor control unit 11 to obtain an image suitable for human visual recognition, and outputs image data consisting of, for example, a group of pixel data. For example, the visual recognition processing unit 14 is configured by an ISP (Image Signal Processor) reading and executing a program stored in advance in a memory (not shown).
[0018] For example, when a color filter is provided for each pixel included in the sensor unit 10 and the pixel data has color information of R (red), G (green), and B (blue), the visual recognition processing unit 14 can perform demosaic processing, white balance processing, etc. Furthermore, the visual recognition processing unit 14 can instruct the sensor control unit 11 to read out pixel data required for the visual recognition processing from the sensor unit 10. The image data obtained by image processing the pixel data by the visual recognition processing unit 14 is passed to the output control unit 15.
[0019] The output control unit 15 is configured by, for example, a microprocessor, and outputs one or both of the recognition result passed from the recognition processing unit 12 and the image data passed from the visual recognition processing unit 14 as a visual recognition processing result to the outside of the imaging device 1. The output control unit 15 can output the image data to, for example, a display unit 31 having a display device. This allows the user to visually recognize the image data displayed by the display unit 31. The display unit 31 may be built into the imaging device 1 or may be configured external to the imaging device 1.
[0020] 2A and 2B are schematic diagrams showing examples of the hardware configuration of the imaging device 1 according to each embodiment. Fig. 2A shows an example in which a single chip 2 is equipped with the sensor unit 10, sensor control unit 11, recognition processing unit 12, memory 13, visual recognition processing unit 14, and output control unit 15 of the configuration shown in Fig. 1. Note that in Fig. 2A, the memory 13 and output control unit 15 are omitted to avoid complexity.
[0021] 2A, the recognition result by the recognition processing unit 12 is output to the outside of the chip 2 via an output control unit 15 (not shown). Also, in the configuration of FIG. 2A, the recognition processing unit 12 can obtain pixel data to be used for recognition from the sensor control unit 11 via an interface inside the chip 2.
[0022] 2B shows an example in which the sensor unit 10, sensor control unit 11, visual recognition processing unit 14, and output control unit 15 of the configuration shown in Fig. 1 are mounted on one chip 2, and the recognition processing unit 12 and memory 13 (not shown) are placed outside the chip 2. In Fig. 2B, as in Fig. 2A described above, the memory 13 and output control unit 15 are omitted to avoid complexity.
[0023] In the configuration of Fig. 2B, the recognition processing unit 12 acquires pixel data to be used for recognition via an interface for communication between chips. Also, in Fig. 2B, the recognition result by the recognition processing unit 12 is shown as being directly output from the recognition processing unit 12 to the outside, but this is not limited to this example. That is, in the configuration of Fig. 2B, the recognition processing unit 12 may return the recognition result to the chip 2 and output it from an output control unit 15 (not shown) mounted on the chip 2.
[0024] In the configuration shown in FIG. 2A, the recognition processing unit 12 is mounted on the chip 2 together with the sensor control unit 11, and communication between the recognition processing unit 12 and the sensor control unit 11 can be performed at high speed via an interface inside the chip 2. On the other hand, in the configuration shown in FIG. 2A, the recognition processing unit 12 cannot be replaced, making it difficult to change the recognition processing. In contrast, in the configuration shown in FIG. 2B, the recognition processing unit 12 is provided outside the chip 2, so communication between the recognition processing unit 12 and the sensor control unit 11 must be performed via an interface between the chips. Therefore, communication between the recognition processing unit 12 and the sensor control unit 11 is slower than in the configuration of FIG. 2A, and there is a possibility of delays in control. On the other hand, the recognition processing unit 12 can be easily replaced, making it possible to realize a variety of recognition processes.
[0025] Unless otherwise specified, the imaging device 1 will be assumed to have the configuration shown in Figure 2A, in which a sensor unit 10, a sensor control unit 11, a recognition processing unit 12, a memory 13, a visual recognition processing unit 14, and an output control unit 15 are mounted on one chip 2.
[0026] 2A, the imaging device 1 can be formed on a single substrate. However, the imaging device 1 may be a stacked CIS in which multiple semiconductor chips are stacked and integrally formed.
[0027] As an example, the imaging device 1 can be formed with a two-layer structure in which semiconductor chips are stacked in two layers. FIG. 3A is a diagram showing an example in which the imaging device 1 according to each embodiment is formed using a two-layer stacked CIS structure. In the structure of FIG. 3A, a pixel section 20a is formed on a first-layer semiconductor chip, and a memory and logic section 20b is formed on a second-layer semiconductor chip. The pixel section 20a includes at least a pixel array in the sensor section 10. The memory and logic section 20b includes, for example, a sensor control section 11, a recognition processing section 12, a memory 13, a visual recognition processing section 14, and an output control section 15, as well as an interface for communicating between the imaging device 1 and the outside. The memory and logic section 20b further includes a part or all of a drive circuit for driving the pixel array in the sensor section 10. Although not shown, the memory and logic section 20b can further include, for example, a memory used by the visual recognition processing section 14 to process image data.
[0028] As shown on the right side of FIG. 3A, the imaging device 1 is configured as a single solid-state imaging element by bonding the semiconductor chip of the first layer and the semiconductor chip of the second layer together while making electrical contact with each other.
[0029] As another example, the imaging device 1 can be formed with a three-layer structure in which semiconductor chips are stacked in three layers. FIG. 3B is a diagram showing an example in which the imaging device 1 according to each embodiment is formed using a three-layer stacked CIS. In the structure of FIG. 3B, the pixel unit 20a is formed in the first semiconductor chip, the memory unit 20c is formed in the second semiconductor chip, and the logic unit 20b' is formed in the third semiconductor chip. In this case, the logic unit 20b' includes, for example, the sensor control unit 11, the recognition processing unit 12, the visual recognition processing unit 14, and the output control unit 15, as well as an interface for communicating between the imaging device 1 and the outside. The memory unit 20c can also include the memory 13 and a memory used by the visual recognition processing unit 14 to process image data. The memory 13 may be included in the logic unit 20b'.
[0030] As shown on the right side of Figure 3B, the imaging device 1 is constructed as a single solid-state imaging element by bonding the first layer semiconductor chip, the second layer semiconductor chip, and the third layer semiconductor chip together while maintaining electrical contact.
[0031] Fig. 4 is a block diagram showing an example of the configuration of a sensor unit 10 applicable to each embodiment. In Fig. 4, the sensor unit 10 includes a pixel array unit 101, a vertical scanning unit 102, an AD (Analog to Digital) conversion unit 103, pixel signal lines 106, vertical signal lines VSL, a control unit 1100, and a signal processing unit 1101. Note that in Fig. 4, the control unit 1100 and the signal processing unit 1101 may also be included in, for example, the sensor control unit 11 shown in Fig. 1.
[0032] The pixel array unit 101 includes a plurality of pixel circuits 100, each of which includes a photoelectric conversion element, such as a photodiode, that performs photoelectric conversion on received light, and a circuit that reads out electric charges from the photoelectric conversion element. In the pixel array unit 101, the plurality of pixel circuits 100 are arranged in a matrix array in the horizontal direction (row direction) and the vertical direction (column direction). In the pixel array unit 101, the row direction arrangement of the pixel circuits 100 is called a line. For example, if one frame of image is formed with 1920 pixels x 1080 lines, the pixel array unit 101 includes at least 1080 lines, each including at least 1920 pixel circuits 100. One frame of image (image data) is formed by pixel signals read out from the pixel circuits 100 included in the frame.
[0033] Hereinafter, the operation of reading pixel signals from each pixel circuit 100 included in a frame in the sensor unit 10 will be appropriately described as "reading pixels from a frame," etc. Also, the operation of reading pixel signals from each pixel circuit 100 of a line included in a frame will be appropriately described as "reading a line," etc.
[0034] Furthermore, pixel signal lines 106 are connected to the pixel array unit 101 for each row and column of each pixel circuit 100, and vertical signal lines VSL are connected to each column. The ends of the pixel signal lines 106 that are not connected to the pixel array unit 101 are connected to the vertical scanning unit 102. Under the control of a control unit 1100 (described later), the vertical scanning unit 102 transmits control signals such as drive pulses used to read pixel signals from pixels to the pixel array unit 101 via the pixel signal lines 106. The ends of the vertical signal lines VSL that are not connected to the pixel array unit 101 are connected to an AD conversion unit 103. The pixel signals read from the pixels are transmitted to the AD conversion unit 103 via the vertical signal lines VSL.
[0035] The following provides an overview of the control of reading out pixel signals from the pixel circuit 100. Reading out pixel signals from the pixel circuit 100 is performed by transferring charges accumulated in a photoelectric conversion element upon exposure to a floating diffusion layer (FD) and converting the transferred charges into a voltage in the floating diffusion layer. The voltage into which the charges are converted in the floating diffusion layer is output to a vertical signal line VSL via an amplifier.
[0036] More specifically, in the pixel circuit 100, during exposure, the connection between the photoelectric conversion element and the floating diffusion layer is turned off (open), and charges generated in response to incident light through photoelectric conversion are accumulated in the photoelectric conversion element. After exposure is completed, the floating diffusion layer is connected to the vertical signal line VSL in response to a selection signal supplied via the pixel signal line 106. Furthermore, the floating diffusion layer is connected to a power supply voltage VDD or a black level voltage supply line for a short period in response to a reset pulse supplied via the pixel signal line 106, thereby resetting the floating diffusion layer. A voltage (referred to as voltage A) at the reset level of the floating diffusion layer is output to the vertical signal line VSL. Thereafter, a transfer pulse supplied via the pixel signal line 106 turns the connection between the photoelectric conversion element and the floating diffusion layer on (closed), and the charges accumulated in the photoelectric conversion element are transferred to the floating diffusion layer. A voltage (referred to as voltage B) corresponding to the amount of charge in the floating diffusion layer is output to the vertical signal line VSL.
[0037] The AD conversion unit 103 includes an AD converter 107 provided for each vertical signal line VSL, a reference signal generation unit 104, and a horizontal scanning unit 105. The AD converter 107 is a column AD converter that performs AD conversion processing on each column of the pixel array unit 101. The AD converter 107 performs AD conversion processing on pixel signals supplied from the pixel circuits 100 via the vertical signal lines VSL, and generates two digital values (values corresponding to voltage A and voltage B, respectively) for correlated double sampling (CDS) processing that reduces noise.
[0038] The AD converter 107 supplies the two generated digital values to the signal processing unit 1101. The signal processing unit 1101 performs CDS processing based on the two digital values supplied from the AD converter 107, and generates a pixel signal (pixel data) as a digital signal. The pixel data generated by the signal processing unit 1101 is output to the outside of the sensor unit 10.
[0039] The reference signal generating unit 104 generates, as a reference signal, a ramp signal used by each AD converter 107 to convert a pixel signal into two digital values, based on a control signal input from the control unit 1100. A ramp signal is a signal whose level (voltage value) decreases at a constant slope over time, or a signal whose level decreases in a step-like manner. The reference signal generating unit 104 supplies the generated ramp signal to each AD converter 107. The reference signal generating unit 104 is configured using, for example, a DAC (Digital to Analog Converter) or the like.
[0040] When a ramp signal, whose voltage drops stepwise according to a predetermined slope, is supplied from the reference signal generator 104, the counter starts counting in accordance with the clock signal. The comparator compares the voltage of the pixel signal supplied from the vertical signal line VSL with the voltage of the ramp signal, and stops counting by the counter when the voltage of the ramp signal crosses the voltage of the pixel signal. The AD converter 107 converts the analog pixel signal into a digital value by outputting a value corresponding to the count value at the time when the counting was stopped.
[0041] The AD converter 107 supplies the two generated digital values to the signal processing unit 1101. The signal processing unit 1101 performs CDS processing based on the two digital values supplied from the AD converter 107, and generates a pixel signal (pixel data) based on a digital signal. The pixel signal based on a digital signal generated by the signal processing unit 1101 is output to the outside of the sensor unit 10.
[0042] Under the control of the control unit 1100, the horizontal scanning unit 105 performs selective scanning to select each AD converter 107 in a predetermined order, thereby causing each AD converter 107 to sequentially output each digital value temporarily held therein to the signal processing unit 1101. The horizontal scanning unit 105 is configured using, for example, a shift register, an address decoder, etc.
[0043] The control unit 1100 controls the driving of the vertical scanning unit 102, the AD conversion unit 103, the reference signal generation unit 104, the horizontal scanning unit 105, etc. in accordance with the imaging control signal supplied from the sensor control unit 11. The control unit 1100 generates various driving signals that serve as references for the operations of the vertical scanning unit 102, the AD conversion unit 103, the reference signal generation unit 104, and the horizontal scanning unit 105. The control unit 1100 generates control signals that the vertical scanning unit 102 supplies to each pixel circuit 100 via the pixel signal line 106, based on, for example, a vertical synchronization signal or an external trigger signal included in the imaging control signal and a horizontal synchronization signal. The control unit 1100 supplies the generated control signals to the vertical scanning unit 102.
[0044] Furthermore, the control unit 1100 passes, for example, information indicating an analog gain included in an imaging control signal supplied from the sensor control unit 11 to the AD conversion unit 103. The AD conversion unit 103 controls the gain of a pixel signal input to each AD converter 107 included in the AD conversion unit 103 via a vertical signal line VSL in accordance with the information indicating the analog gain.
[0045] Based on a control signal supplied from the control unit 1100, the vertical scanning unit 102 supplies various signals including drive pulses to the pixel signal lines 106 of a selected pixel row of the pixel array unit 101, to each pixel circuit 100 for each line, and causes each pixel circuit 100 to output a pixel signal to a vertical signal line VSL. The vertical scanning unit 102 is configured using, for example, a shift register, an address decoder, etc. Furthermore, the vertical scanning unit 102 controls the exposure of each pixel circuit 100 in accordance with information indicating exposure supplied from the control unit 1100.
[0046] The sensor unit 10 configured in this manner is a column AD type CMOS (Complementary Metal Oxide Semiconductor) image sensor in which AD converters 107 are arranged for each column.
[0047] [2. Examples of existing technologies applicable to the present disclosure] Prior to describing each embodiment of the present disclosure, a brief description of existing technologies applicable to the present disclosure will be given to facilitate understanding.
[0048] (2-1. Overview of Rolling Shutter) Known imaging methods for capturing images using the pixel array unit 101 include the rolling shutter (RS) method and the global shutter (GS) method. First, the rolling shutter method will be briefly described. Figures 5A, 5B, and 5C are schematic diagrams for explaining the rolling shutter method. In the rolling shutter method, as shown in Figure 5A, images are captured line by line in sequence, starting from, for example, line 201 at the top of a frame 200.
[0049] In the above description, "imaging" refers to the operation of the sensor unit 10 to output a pixel signal in response to light irradiated onto the light-receiving surface. More specifically, "imaging" refers to a series of operations from exposing a pixel to transferring a pixel signal based on charges accumulated in a photoelectric conversion element included in the pixel by the exposure to the sensor control unit 11. Also, as described above, a frame refers to an area in the pixel array unit 101 where pixel circuits 100 effective for generating pixel signals are arranged.
[0050] For example, in the configuration of Fig. 4, exposure is performed simultaneously in each pixel circuit 100 included in one line. After exposure is completed, pixel signals based on charges accumulated by exposure are simultaneously transferred in each pixel circuit 100 included in that line via each vertical signal line VSL corresponding to each pixel circuit 100. By performing this operation sequentially line by line, imaging using a rolling shutter can be achieved.
[0051] FIG. 5B schematically illustrates an example of the relationship between imaging and time in the rolling shutter method. In FIG. 5B, the vertical axis represents line position, and the horizontal axis represents time. In the rolling shutter method, exposure for each line is performed line-by-line sequentially, so as shown in FIG. 5B, the exposure timing for each line is shifted sequentially according to the line position. Therefore, for example, if the horizontal positional relationship between the imaging device 1 and the subject changes rapidly, distortion occurs in the image of the captured frame 200, as illustrated in FIG. 5C. In the example of FIG. 5C, image 202 corresponding to frame 200 is tilted at an angle corresponding to the speed and direction of change in the horizontal positional relationship between the imaging device 1 and the subject.
[0052] In the rolling shutter method, it is also possible to thin out lines when capturing an image. Figures 6A, 6B, and 6C are schematic diagrams for explaining line thinning in the rolling shutter method. As shown in Figure 6A, similar to the example of Figure 5A described above, capturing is performed line by line from line 201 at the top of frame 200 toward the bottom of frame 200. At this time, capturing is performed while skipping every predetermined number of lines.
[0053] For the sake of explanation, let us assume that every other line is imaged by thinning out one line. That is, after imaging the nth line, imaging of the (n+2)th line is performed. In this case, the time from imaging the nth line to imaging the (n+2)th line is assumed to be equal to the time from imaging the nth line to imaging the (n+1)th line when thinning out is not performed.
[0054] FIG. 6B schematically illustrates an example of the relationship between imaging and time when one line is thinned out using the rolling shutter method. In FIG. 6B, the vertical axis represents line position, and the horizontal axis represents time. In FIG. 6B, exposure A corresponds to the exposure in FIG. 5B without thinning out, and exposure B represents the exposure when one line is thinned out. As shown in exposure B, by thinning out the lines, the difference in exposure timing at the same line position can be reduced compared to when line thinning is not performed. Therefore, as illustrated by image 203 in FIG. 6C, the tilt distortion generated in the image of captured frame 200 is smaller than when line thinning is not performed as shown in FIG. 5C. On the other hand, when line thinning is performed, the image resolution is lower than when line thinning is not performed.
[0055] Although the above description has been given of an example in which imaging is performed line-sequentially from the top to the bottom of the frame 200 using the rolling shutter method, this is not limiting. Figures 7A and 7B are diagrams schematically illustrating examples of other imaging methods using the rolling shutter method. For example, as shown in Figure 7A, imaging can be performed line-sequentially from the bottom to the top of the frame 200 using the rolling shutter method. In this case, the horizontal direction of distortion in the image 202 is reversed compared to when imaging is performed line-sequentially from the top to the bottom of the frame 200.
[0056] Also, for example, by setting the range of the vertical signal line VSL that transfers pixel signals, it is possible to selectively read out a portion of the line. Furthermore, by respectively setting the line where imaging is performed and the vertical signal line VSL that transfers pixel signals, it is possible to set the lines where imaging starts and ends other than the top and bottom ends of the frame 200. Fig. 7B schematically shows an example in which the imaging range is a rectangular region 205 whose width and height are less than the width and height of the frame 200. In the example of Fig. 7B, imaging is performed line by line from line 204 at the top end of the region 205 toward the bottom end of the region 205.
[0057] (2-2. Global Shutter Overview) Next, a global shutter (GS) system will be briefly described as an imaging system used when capturing an image using the pixel array unit 101. Figures 8A, 8B, and 8C are schematic diagrams for explaining the global shutter system. In the global shutter system, as shown in Figure 8A, all pixel circuits 100 included in a frame 200 are exposed simultaneously.
[0058] 4, one possible configuration is to further provide a capacitor between the photoelectric conversion element and the FD in each pixel circuit 100. A first switch is provided between the photoelectric conversion element and the capacitor, and a second switch is provided between the capacitor and the floating diffusion layer, and the opening and closing of each of the first and second switches is controlled by a pulse supplied via the pixel signal line 106.
[0059] With this configuration, during the exposure period, the first and second switches are opened in all pixel circuits 100 included in frame 200, and when exposure ends, the first switch is switched from open to closed to transfer charge from the photoelectric conversion element to the capacitor. Thereafter, the capacitor is regarded as a photoelectric conversion element, and charge is read out from the capacitor in a sequence similar to the readout operation described for the rolling shutter method. This enables simultaneous exposure in all pixel circuits 100 included in frame 200.
[0060] FIG. 8B schematically shows an example of the relationship between imaging and time in the global shutter system. In FIG. 8B, the vertical axis represents line position, and the horizontal axis represents time. In the global shutter system, exposure is performed simultaneously in all pixel circuits 100 included in frame 200, so the exposure timing for each line can be made the same, as shown in FIG. 8B. Therefore, even if, for example, the horizontal positional relationship between imaging device 1 and a subject changes rapidly, no distortion corresponding to the change occurs in image 206 of captured frame 200, as shown in FIG. 8C.
[0061] The global shutter method ensures simultaneous exposure timing for all pixel circuits 100 included in the frame 200. Therefore, by controlling the timing of each pulse supplied by the pixel signal line 106 of each line and the timing of transfer by each vertical signal line VSL, sampling (reading of pixel signals) in various patterns can be realized.
[0062] 9A and 9B are diagrams schematically showing examples of sampling patterns that can be achieved with the global shutter system. Fig. 9A shows an example in which samples 208 for reading out pixel signals are extracted in a checkerboard pattern from each pixel circuit 100 arranged in a matrix and included in a frame 200. Fig. 9B shows an example in which samples 208 for reading out pixel signals are extracted in a grid pattern from each pixel circuit 100. Similarly to the rolling shutter system described above, the global shutter system also allows for line-sequential imaging.
[0063] (2-3. About DNN) Next, a recognition process using a DNN (Deep Neural Network) applicable to each embodiment will be briefly described. In each embodiment, a recognition process for image data is performed using a CNN (Convolutional Neural Network) and an RNN (Recurrent Neural Network) among DNNs. Hereinafter, "recognition process for image data" will be referred to as "image recognition process" or the like as appropriate.
[0064] (2-3-1. Overview of CNN) First, a brief description of CNN will be given. Image recognition processing using CNN generally involves performing image recognition processing based on image information consisting of pixels arranged in a matrix, for example. FIG. 10 is a diagram for explaining an outline of image recognition processing using CNN. Processing is performed by a CNN 52 that has been trained in a predetermined manner on all pixel information 51 of an image 50 depicting the number "8," which is an object to be recognized. As a result, the number "8" is recognized as a recognition result 53.
[0065] Alternatively, it is possible to perform CNN processing based on an image for each line, and obtain a recognition result from a portion of the image to be recognized. FIG. 11 is a diagram for explaining an outline of image recognition processing that obtains a recognition result from a portion of the image to be recognized. In FIG. 11, an image 50' is a partial line-by-line acquisition of the number "8," which is the object to be recognized. For example, pixel information 54a, 54b, and 54c for each line that forms pixel information 51' of this image 50' is sequentially processed by a CNN 52' that has been trained in a predetermined manner.
[0066] For example, the recognition result 53a obtained by the CNN 52' in the recognition process for pixel information 54a on the first line is not a valid recognition result. Here, a valid recognition result refers to, for example, a recognition result with a score indicating the reliability of the recognition result equal to or greater than a predetermined value. The CNN 52' updates its internal state 55 based on the recognition result 53a. Next, the CNN 52', whose internal state has been updated 55 based on the previous recognition result 53a, performs recognition process on pixel information 54b on the second line. In FIG. 11, the recognition result 53b indicates that the digit to be recognized is either "8" or "9." Furthermore, the CNN 52' updates its internal information 55 based on the recognition result 53b. Next, the CNN 52', whose internal state has been updated 55 based on the previous recognition result 53b, performs recognition process on pixel information 54c on the third line. In FIG. 11, the digit to be recognized is narrowed down to "8" from among "8" and "9."
[0067] Here, the recognition process shown in Fig. 11 updates the internal state of the CNN using the result of the previous recognition process, and the CNN with this updated internal state performs recognition processing using pixel information of lines adjacent to the line on which the previous recognition process was performed. In other words, the recognition process shown in Fig. 11 is performed line by line on the image while updating the internal state of the CNN based on the previous recognition result. Therefore, the recognition process shown in Fig. 11 is a process that is performed recursively line by line, and can be considered to have a structure equivalent to an RNN.
[0068] (2-3-2. Overview of RNN) Next, an outline of the RNN will be explained. Figures 12A and 12B are diagrams that schematically show an example of classification processing (recognition processing) by the DNN when time-series information is not used. In this case, as shown in Figure 12A, one image is input to the DNN. In the DNN, classification processing is performed on the input image, and a classification result is output.
[0069] Fig. 12B is a diagram for explaining the processing of Fig. 12A in more detail. As shown in Fig. 12B, the DNN executes a feature extraction process and a classification process. In the DNN, feature amounts are extracted from an input image by the feature extraction process. In addition, in the DNN, classification process is executed on the extracted feature amounts to obtain a classification result.
[0070] 13A and 13B are diagrams schematically illustrating a first example of a classification process using a DNN when time-series information is used. In the examples of FIGS. 13A and 13B, a fixed number of past pieces of time-series information are used to perform the classification process using a DNN. In the example of FIG. 13A, an image [T] at time T, an image [T-1] at time T-1 before time T, and an image [T-2] at time T-2 before time T-1 are input to the DNN. The DNN performs classification processing on the input images [T], [T-1], and [T-2], and obtains a classification result [T] at time T.
[0071] 13B is a diagram for explaining the processing of FIG. 13A in more detail. As shown in FIG. 13B, in the DNN, the feature extraction processing described above using FIG. 12B is performed one-to-one for each of the input images [T], [T-1], and [T-2], and feature quantities corresponding to the images [T], [T-1], and [T-2] are extracted. In the DNN, the feature quantities obtained based on these images [T], [T-1], and [T-2] are integrated, and a classification process is performed on the integrated feature quantities to obtain a classification result [T] at time T.
[0072] The method of Figures 13A and 13B requires multiple configurations for feature extraction, and the number of configurations required for feature extraction depends on the number of past images available, which could result in the DNN configuration becoming large-scale.
[0073] 14A and 14B are diagrams schematically illustrating a second example of a classification process using a DNN when time-series information is used. In the example of Fig. 14A, an image [T] at time T is input to a DNN whose internal state has been updated to the state at time T-1, and a classification result [T] at time T is obtained.
[0074] FIG. 14B is a diagram for explaining the processing of FIG. 14A in more detail. As shown in FIG. 14B, in the DNN, the feature extraction processing described above with reference to FIG. 12B is performed on the input image [T] at time T, and features corresponding to the image [T] are extracted. In the DNN, the internal state is updated by the image before time T, and features related to the updated internal state are stored. The features related to this stored internal information are integrated with the features in the image [T], and a classification process is performed on the integrated features.
[0075] 14A and 14B is a recursive process that is executed using a DNN whose internal state has been updated using, for example, the previous classification result. A DNN that performs recursive processing in this way is called an RNN (Recurrent Neural Network). Classification processing using an RNN is generally used for video image recognition, and it is possible to improve classification accuracy by, for example, sequentially updating the internal state of the DNN using frame images that are updated in a time series.
[0076] In the present disclosure, an RNN is applied to a rolling shutter structure. That is, in the rolling shutter method, pixel signals are read out line-sequentially. Therefore, the pixel signals read out line-sequentially are applied to an RNN as time-series information. This makes it possible to perform classification processing based on multiple lines with a smaller configuration than when a CNN is used (see FIG. 13B). Not limited to this, an RNN can also be applied to a global shutter structure. In this case, for example, adjacent lines can be considered as time-series information.
[0077] (2-4. Drive speed) Next, the relationship between the frame drive speed and the amount of pixel signal readout will be explained using Figures 15A and 15B. Figure 15A is a diagram showing an example of reading out all lines in an image. Here, it is assumed that the resolution of the image to be subjected to recognition processing is 640 pixels horizontally by 480 pixels vertically (480 lines). In this case, driving at a drive speed of 14,400 lines / second enables output at 30 frames per second (fps).
[0078] Next, consider imaging with line thinning. For example, as shown in FIG. 15B, imaging is performed using 1 / 2 thinning readout, in which imaging is performed by skipping one line at a time. As a first example of 1 / 2 thinning, when the drive speed is 14,400 lines / second as described above, the number of lines read from the image is halved, resulting in a decrease in resolution. However, output at 60 fps, double the speed when no thinning is performed, is possible, thereby improving the frame rate. As a second example of 1 / 2 thinning, when the drive speed is set to 7,200 fps, half that of the first example, the frame rate is 30 fps, the same as when no thinning is performed, but power consumption can be reduced.
[0079] When reading out lines of an image, whether to perform no thinning, to perform thinning and increase the drive speed, or to perform thinning and keep the drive speed the same as when no thinning is performed can be selected depending on, for example, the purpose of the recognition processing based on the read-out pixel signals.
[0080] 3. Overview of this Disclosure Each embodiment of the present disclosure will be described in more detail below. First, a summary of the processing according to each embodiment of the present disclosure will be given. Fig. 16 is a schematic diagram for explaining the recognition processing according to each embodiment of the present disclosure. In Fig. 16, in step S1, the imaging device 1 (see Fig. 1) according to each embodiment starts capturing an image to be recognized.
[0081] It is assumed that the target image is, for example, an image of the number "8" drawn by hand. Furthermore, a learning model that has been trained to be able to identify numbers using predetermined training data is stored in advance in the memory 13 as a program, and the recognition processing unit 12 is able to identify numbers included in the image by reading and executing this program from the memory 13. Furthermore, it is assumed that the imaging device 1 captures images using a rolling shutter method. It is noted that even when the imaging device 1 captures images using a global shutter method, the following processing can be applied in the same way as in the case of the rolling shutter method.
[0082] When imaging starts, in step S2, the imaging device 1 sequentially reads out the frame line by line from the top end to the bottom end of the frame.
[0083] When the lines are read up to a certain position, the recognition processing unit 12 identifies the number "8" or "9" from the image of the read lines (step S3). For example, the numbers "8" and "9" contain a common feature in the upper half, so when the lines are read from the top and the feature is recognized, the recognized object can be identified as either the number "8" or "9."
[0084] Here, as shown in step S4a, by reading up to the bottom line of the frame or a line near the bottom, the entire recognized object appears, and the object identified in step S2 as either the number "8" or "9" is confirmed to be the number "8".
[0085] On the other hand, steps S4b and S4c are processes related to the present disclosure.
[0086] As shown in step S4b, by reading further from the line position read in step S3, it is possible to identify the recognized object as the number "8" even when the bottom end of the number "8" is reached. For example, the bottom half of the number "8" and the bottom half of the number "9" each have different characteristics. By reading the line up to the point where the difference in these characteristics becomes clear, it becomes possible to identify whether the object recognized in step S3 is the number "8" or "9." In the example of FIG. 16, the object is determined to be the number "8" in step S4b.
[0087] Furthermore, as shown in step S4c, it is also possible to jump from the line position of step S3 to a line position where it is possible to determine whether the object identified in step S3 is the number "8" or "9" by further reading in the state of step S3. By reading this jump destination line, it is possible to determine whether the object identified in step S3 is the number "8" or "9." The jump destination line position can be determined based on a learning model that has been trained in advance based on predetermined training data.
[0088] Here, when the object is confirmed in step S4b or step S4c described above, the image capture device 1 can end the recognition process, thereby realizing a reduction in the time and power consumption of the recognition process in the image capture device 1.
[0089] The training data is data that holds multiple combinations of input signals and output signals for each reading unit. As an example, in the above-mentioned task of identifying numbers, data for each reading unit (line data, subsampled data, etc.) can be applied as the input signal, and data indicating the "correct number" can be applied as the output signal. As another example, in a task of detecting an object, data for each reading unit (line data, subsampled data, etc.) can be applied as the input signal, and object class (human body / vehicle / non-object) or object coordinates (x, y, h, w) can be applied as the output signal. Furthermore, the output signal may be generated from only the input signal using self-supervised learning.
[0090] 4. First Embodiment Next, a first embodiment of the present disclosure will be described.
[0091] (4-1. Example of operation by the recognition processing unit) In the imaging device 1 according to the first embodiment, as described above, the recognition processing unit 12 functions as a recognizer using DNN by reading and executing a program stored in the memory 13 as a learning model that has been pre-trained based on predetermined teacher data.
[0092] 17 is a flowchart illustrating an example of recognition processing by the recognition processing unit 12 according to the first embodiment. In step S121 of FIG. 17, the DSP constituting the recognition processing unit 12 in the imaging device 1 reads and executes a learning model from the memory 13. This causes the DSP to function as the recognition processing unit 12.
[0093] Next, in step S122, the recognition processing unit 12 in the imaging device 1 instructs the sensor control unit 11 to start reading frames from the sensor unit 10. In this frame reading, for example, image data for one frame is read out line by line (also referred to as row by row). The recognition processing unit 12 determines whether image data for a predetermined number of lines in one frame has been read out.
[0094] When the recognition processing unit 12 determines that image data for a predetermined number of lines in one frame has been read out (step S123, "YES"), it transitions the process to step S124. In step S124, the recognition processing unit 12 performs recognition processing as machine learning processing using CNN on the read image data for the predetermined number of lines. That is, the recognition processing unit 12 performs machine learning processing using a learning model with the image data for the predetermined number of lines as a unit area. Furthermore, in the machine learning processing using CNN, recognition processing and detection processing such as face detection, face authentication, gaze detection, facial expression recognition, face direction detection, object detection, object recognition, movement (moving object) detection, pet detection, scene recognition, state detection, and avoidance target recognition are performed.
[0095] Here, face detection is a process of detecting the face of a person included in image data. Face authentication is a type of biometric authentication, and is a process of authenticating whether or not the face of a person included in image data matches the face of a person registered in advance. Gaze detection is a process of detecting the gaze direction of a person included in image data. Facial expression recognition is a process of recognizing the facial expression of a person included in image data. Facial direction detection is a process of detecting the up-down direction of the face of a person included in image data. Object detection is a process of detecting an object included in image data. Object recognition is a process of recognizing what an object included in image data is. Motion (animal object) detection is a process of detecting an animal object included in image data. Pet detection is a process of detecting a pet such as a dog or cat included in image data. Scene recognition is a process of recognizing the scene being photographed (the sea, mountains, etc.). State detection is a process of detecting the state of a person, etc. included in image data (whether they are in a normal state, an abnormal state, etc.). Avoidable object recognition is a process of recognizing an object to be avoided that exists ahead in the traveling direction of the subject when the subject moves. The machine learning processes executed by the recognition processing unit 12 are not limited to these examples.
[0096] In step S125, the recognition processing unit 12 determines whether the machine learning process using CNN in step S124 was successful. If the recognition processing unit 12 determines that the machine learning process using CNN in step S124 was successful (step S125, "YES"), it shifts the process to step S129. On the other hand, if the recognition processing unit 12 determines that the machine learning process using CNN in step S124 was unsuccessful (step S125, "NO"), it shifts the process to step S126. In step S126, the recognition processing unit 12 waits for the next predetermined number of lines of image data to be read from the sensor control unit 11 (step S126, "NO").
[0097] In this description, a successful machine learning process means that a certain detection result, recognition result, or authentication is obtained, for example, in the face detection, face authentication, etc. exemplified above. On the other hand, a failure in the machine learning process means that a sufficient detection result, recognition result, or authentication is not obtained, for example, in the face detection, face authentication, etc. exemplified above.
[0098] Next, in step S126, when the next predetermined number of lines of image data (unit area) is read out (step S126, "YES"), the recognition processing unit 12 performs machine learning processing using an RNN on the read image data of the predetermined number of lines in step S127. In the machine learning processing using an RNN, for example, the results of machine learning processing using a CNN or an RNN that has been previously performed on image data of the same frame are also used.
[0099] In step S128, if the recognition processing unit 12 determines that the machine learning process using the RNN in step S127 has been successful (step S128, "YES"), it shifts the process to step S129.
[0100] In step S129, the recognition processing unit 12 supplies the successful machine learning result in step S124 or step S127 to, for example, the output control unit 15. The machine learning result output in step S129 is, for example, a valid recognition result by the recognition processing unit 12. The recognition processing unit 12 may store the machine learning result in the memory 13.
[0101] Furthermore, if the recognition processing unit 12 determines in step S128 that the machine learning process using the RNN in step S127 has failed (step S128, "NO"), it shifts the process to step S130. In step S130, the recognition processing unit 12 determines whether or not the reading of one frame's worth of image data has been completed. If the recognition processing unit 12 determines that the reading of one frame's worth of image data has not been completed (step S130, "NO"), it returns the process to step S126, and processing is performed on the next predetermined number of lines of image data.
[0102] On the other hand, if the recognition processing unit 12 determines in step S130 that reading of image data for one frame has been completed (step S130, "YES"), then, for example, the recognition processing unit 12 determines in step S131 whether or not to end the series of processes according to the flowchart in Fig. 17. If the recognition processing unit 12 determines not to end the series of processes (step S131, "NO"), it returns the process to step S122 and executes the same operation for the next frame. If the recognition processing unit 12 determines to end the series of processes according to the flowchart in Fig. 17 (step S131, "YES"), then the recognition processing unit 12 ends the series of processes according to the flowchart in Fig. 17.
[0103] The determination of whether to proceed to the next frame in step S131 may be made, for example, based on whether an instruction to end has been input from outside the imaging device 1, or based on whether a series of processes for a predetermined number of frames of image data has been completed.
[0104] Furthermore, when machine learning processes such as face detection, face authentication, gaze detection, facial expression recognition, face direction detection, object detection, object recognition, movement (moving object) detection, scene recognition, and state detection are performed consecutively, the next machine learning process may be skipped if the immediately preceding machine learning process fails. For example, when face authentication is performed following face detection, the next face authentication may be skipped if face detection fails.
[0105] (4-2. Specific examples of operations by the recognition processing unit) Next, a specific example will be used to explain the operation of the machine learning unit described with reference to Fig. 17. Note that the following will exemplify a case where face detection is performed using DNN.
[0106] Fig. 18 is a diagram showing an example of image data for one frame. Fig. 19 is a diagram for explaining the flow of machine learning processing executed by the recognition processing unit 12 according to the first embodiment.
[0107] When performing face detection using machine learning on image data such as that shown in FIG. 18, as shown in section (a) of FIG. 19, image data for a predetermined number of lines is first input to the recognition processing unit 12 (corresponding to step S123 in FIG. 17). The recognition processing unit 12 performs machine learning processing using CNN on the input image data for the predetermined number of lines, thereby performing face detection (corresponding to step S124 in FIG. 17). However, at the stage of section (a) of FIG. 19, image data of the entire face has not yet been input, so the recognition processing unit 12 fails to detect the face (corresponding to "NO" in step S125 in FIG. 17).
[0108] 19, image data for the next predetermined number of lines is input to the recognition processing unit 12 (corresponding to step S126 in FIG. 17). The recognition processing unit 12 performs face detection by performing machine learning processing using an RNN on newly input image data for the predetermined number of lines, while using the result of the machine learning processing using a CNN that was performed on the image data for the predetermined number of lines that was input in section (a) in FIG. 19 (corresponding to step S127 in FIG. 17).
[0109] At the stage of section (b) in Fig. 19, image data of the entire face is input together with the pixel data for the predetermined number of lines input at the stage of section (a) in Fig. 19. Therefore, at the stage of section (b) in Fig. 19, the recognition processing unit 12 succeeds in face detection (corresponding to "YES" in step S128 in Fig. 17). Then, in this operation, the result of face detection is output (corresponding to step S129 in Fig. 17) without reading out the next and subsequent image data (image data in sections (c) to (f) in Fig. 19).
[0110] In this way, by performing machine learning processing using DNN on image data for a predetermined number of lines at a time, it is possible to omit reading out image data and performing machine learning processing on it after face detection is successful. This makes it possible to complete detection, recognition, authentication, and other processes in a short time, thereby reducing processing time and power consumption.
[0111] The predetermined number of lines is determined by the size of the filter required by the learning model algorithm, and the minimum number is one line.
[0112] Furthermore, the image data read from the sensor unit 10 by the sensor control unit 11 may be image data thinned in at least one of the column direction and the row direction. In this case, for example, when image data is read every other row in the column direction, image data on the 2(N-1)th line (N is an integer equal to or greater than 1) is read.
[0113] Furthermore, if the filter required by the learning model algorithm is not in units of lines, but is a rectangular area in units of pixels, such as 1x1 pixels or 5x5 pixels, image data of a rectangular area corresponding to the shape and size of the filter may be input to the recognition processing unit 12 as image data of a unit area on which the recognition processing unit 12 performs machine learning processing, instead of image data of a predetermined number of lines.
[0114] Furthermore, in the above description, CNN and RNN are given as examples of DNNs, but the present invention is not limited to these, and other learning models can also be used.
[0115] (4-3. Application example of the first embodiment) Next, an application example of the first embodiment will be described. Here, as an application example of the first embodiment, an example will be described in which exposure for a predetermined number of lines to be read next is controlled based on the result of machine learning processing by CNN in step S124 of the flowchart in Fig. 17 or the result of machine learning processing by RNN in step S127. Fig. 20A and Fig. 20B are schematic diagrams for explaining an application example of the first embodiment.
[0116] Section (a) of Figure 20A is a schematic diagram showing an example of an overexposed image 60a. Because the image 60a is overexposed, the image 60a appears whitish overall, and for example, a monitor 62, which is an object included in the image 60a, has a so-called blown-out highlight within the screen, making it difficult for the human eye to distinguish details. On the other hand, a person 61, which is an object included in the image 60a, appears slightly whitish due to overexposure, but is easily distinguishable to the human eye compared to the monitor 62.
[0117] Section (b) of Figure 20A is a schematic diagram showing an example of an underexposed image 60b. Because image 60b is underexposed, image 60b appears dark overall, making it difficult for the human eye to distinguish, for example, a person 61 that is visible in image 60a. On the other hand, a monitor 62 included in image 60b can be distinguished in finer detail by the human eye compared to image 60a.
[0118] Fig. 20B is a schematic diagram illustrating a readout method according to an application example of the first embodiment. Sections (a) and (b) of Fig. 20B show a case where frame readout is started in an underexposed state in step S122 of the flowchart of Fig. 17 described above.
[0119] Section (a) of FIG. 20B illustrates a first example of a readout method in an application example of the first embodiment. In image 60c in section (a) of FIG. 20B, for example, it is assumed that the recognition process using CNN in step S124 for the first line L#1 of the frame failed, or the score indicating the reliability of the recognition result was below a predetermined level. In this case, the recognition processing unit 12 instructs the sensor control unit 11 to set the exposure of line L#2 to be read out in step S126 to an exposure suitable for the recognition process (in this case, to set a large amount of exposure). Note that in FIG. 20B, each of lines L#1, L#2, ... may be a single line, or multiple adjacent lines.
[0120] In the example of section (a) in FIG. 20B, the exposure amount of line L#2 is set to be greater than the exposure amount of line L#1. As a result, line L#2 becomes overexposed, and for example, the recognition process using RNN in step S127 fails or the score is below a predetermined value. The recognition processing unit 12 returns the process from step S130 to step S126, and instructs the sensor control unit 11 to set the exposure amount of line L#3, which is read out, to be less than the exposure amount of line L#2. Similarly, for lines L#4, ..., L#m, ..., the exposure amount of the next line is sequentially set according to the result of the recognition process.
[0121] In this way, by adjusting the exposure amount of the next line to be read based on the recognition result of a certain line, it is possible to perform the recognition process with higher accuracy.
[0122] As a further application of the above application example, as shown in section (b) of Figure 20B, a method can be considered in which the exposure is reset when a predetermined line has been read out, and reading is resumed from the first line of the frame. As shown in section (b) of Figure 20B, the recognition processing unit 12 reads out the first line of the frame, for example, L#1 to line L#m, in the same manner as in section (a) above, and resets the exposure based on the recognition results. The recognition processing unit 12 then reads out each line of the frame, L#1, L#2, ..., based on the reset exposure (second).
[0123] In this way, by resetting the exposure based on the results of reading a predetermined number of lines, and then re-reading lines L#1, L#2, ... from the beginning of the frame based on the reset exposure, it is possible to perform recognition processing with even higher accuracy.
[0124] 5. Second Embodiment (5-0-1. Configuration Example According to Second Embodiment) Next, a second embodiment of the present disclosure will be described. The second embodiment is an extension of the recognition processing according to the first embodiment described above. FIG. 21 is a functional block diagram of an example for explaining the function of an imaging device according to the second embodiment. Note that in FIG. 21, the optical unit 30, sensor unit 10, memory 13, and display unit 31 shown in FIG. 1 are omitted. Also, in FIG. 21, a trigger generation unit 16 is added to the configuration of FIG. 1.
[0125] 21 , the sensor control unit 11 includes a readout unit 110 and a readout control unit 111. The recognition processing unit 12 includes a feature amount calculation unit 120, a feature amount storage control unit 121, a readout determination unit 123, and a recognition processing execution unit 124, and the feature amount storage control unit 121 includes a feature amount storage unit 122. The visual recognition processing unit 14 includes an image data storage control unit 140, a readout determination unit 142, and an image processing unit 143, and the image data storage control unit 140 includes an image data storage unit 141.
[0126] In the sensor control unit 11, the read control unit 111 receives read area information indicating a read area to be read in the recognition processing unit 12 from the read determination unit 123 included in the recognition processing unit 12. The read area information is, for example, the line numbers of one or more lines. However, the read area information may be information indicating pixel positions within one line. Furthermore, by combining one or more line numbers with information indicating pixel positions of one or more pixels within a line as the read area information, it is possible to specify read areas of various patterns. Note that the read area is equivalent to the read unit. However, the read area and the read unit may be different.
[0127] Similarly, the read control unit 111 receives read area information indicating the read area from which the visual recognition processing unit 14 performs reading, from a read determination unit 142 included in the visual recognition processing unit 14 .
[0128] The read control unit 111 passes read area information indicating the read area from which reading will actually be performed to the read unit 110 based on the read determination units 123 and 142. For example, if a conflict occurs between the read area information received from the read determination unit 123 and the read area information received from the read determination unit 142, the read control unit 111 can arbitrate and adjust the read area information to be passed to the read unit 110.
[0129] Furthermore, the read control unit 111 can receive information indicating exposure and analog gain from the read determination unit 123 or the read determination unit 142. The read control unit 111 passes the received information indicating exposure and analog gain to the read unit 110.
[0130] The readout unit 110 reads pixel data from the sensor unit 10 in accordance with the readout area information passed from the readout control unit 111. For example, the readout unit 110 obtains a line number indicating the line to be read out and pixel position information indicating the position of the pixel to be read out on that line based on the readout area information, and passes the obtained line number and pixel position information to the sensor unit 10. The readout unit 110 passes each pixel data acquired from the sensor unit 10, together with the readout area information, to the recognition processing unit 12 and the visual recognition processing unit 14.
[0131] The readout unit 110 also sets exposure and analog gain (AG) for the sensor unit 10 in accordance with information indicating the exposure and analog gain received from the readout control unit 111. Furthermore, the readout unit 110 can generate a vertical synchronization signal and a horizontal synchronization signal and supply them to the sensor unit 10.
[0132] In the recognition processing unit 12, the read determination unit 123 receives read information indicating the read area to be read next from the feature accumulation control unit 121. The read determination unit 123 generates read area information based on the received read information and passes it to the read control unit 111.
[0133] Here, the read determination unit 123 can use, for example, information in which a predetermined read unit is added with read position information for reading pixel data of the read unit as the read area indicated in the read area information. The read unit is a set of one or more pixels and serves as a unit of processing by the recognition processing unit 12 and the visual recognition processing unit 14. As an example, if the read unit is a line, a line number [L#x] indicating the position of the line is added as the read position information. Furthermore, if the read unit is a rectangular area including multiple pixels, information indicating the position of the rectangular area in the pixel array unit 101, for example, information indicating the position of the pixel in the upper left corner, is added as the read position information. The read determination unit 123 is assigned a read unit to be applied in advance. However, the read determination unit 123 can also determine the read unit in response to, for example, an instruction from outside the read determination unit 123. Therefore, the read determination unit 123 functions as a read unit control unit that controls the read unit.
[0134] The read determination unit 123 can also determine the next read area to be read based on recognition information passed from the recognition process execution unit 124, which will be described later, and generate read area information indicating the determined read area.
[0135] Similarly, in the visual recognition processing unit 14, the read determination unit 142 receives read information indicating the read area to be read next, for example, from the image data storage control unit 140. The read determination unit 142 generates read area information based on the received read information and passes it to the read control unit 111.
[0136] In the recognition processing unit 12, the feature calculation unit 120 calculates the feature in the area indicated by the read-out area information based on the pixel data and the read-out area information passed from the read-out unit 110. The feature calculation unit 120 passes the calculated feature to the feature accumulation control unit 121.
[0137] As will be described later, the feature amount calculation unit 120 may calculate the feature amount based on the pixel data passed from the readout unit 110 and the past feature amount passed from the feature amount accumulation control unit 121. However, the feature amount calculation unit 120 may also obtain, for example, information for setting exposure and analog gain from the readout unit 110, and calculate the feature amount by further using the obtained information.
[0138] In the recognition processing unit 12, the feature accumulation control unit 121 accumulates the feature passed from the feature calculation unit 120 in the feature accumulation unit 122. Furthermore, when the feature is passed from the feature calculation unit 120, the feature accumulation control unit 121 generates read information indicating the read area from which the next readout will be performed, and passes the read information to the read determination unit 123.
[0139] Here, the feature accumulation control unit 121 can integrate and accumulate already accumulated feature amounts and newly passed feature amounts. The feature accumulation control unit 121 can also delete feature amounts that are no longer needed from the feature amounts accumulated in the feature accumulation unit 122. Possible examples of unnecessary feature amounts include feature amounts related to the previous frame and feature amounts that have been calculated and already accumulated based on frame images of scenes different from the frame images for which new feature amounts have been calculated. The feature accumulation control unit 121 can also initialize the feature accumulation unit 122 by deleting all feature amounts accumulated therein as necessary.
[0140] Furthermore, the feature accumulation control unit 121 generates a feature to be used in the recognition process by the recognition process execution unit 124, based on the feature passed from the feature calculation unit 120 and the feature accumulated in the feature accumulation unit 122. The feature accumulation control unit 121 passes the generated feature to the recognition process execution unit 124.
[0141] The recognition process execution unit 124 executes the recognition process based on the feature amounts passed from the feature amount accumulation control unit 121. The recognition process execution unit 124 performs object detection, face detection, and the like through the recognition process. The recognition process execution unit 124 passes the recognition result obtained through the recognition process to the output control unit 15. The recognition process execution unit 124 can also pass recognition information including the recognition result generated through the recognition process to the readout determination unit 123. Note that the recognition process execution unit 124 can receive the feature amounts from the feature amount accumulation control unit 121 and execute the recognition process, for example, in response to a trigger generated by the trigger generation unit 16.
[0142] Meanwhile, in the visual recognition processing unit 14, the image data storage control unit 140 receives pixel data read from the read area and read area information corresponding to the image data from the read unit 110. The image data storage control unit 140 associates the pixel data and the read area information and stores them in the image data storage unit 141.
[0143] The image data storage control unit 140 generates image data for image processing by the image processing unit 143 based on the pixel data passed from the reading unit 110 and the image data stored in the image data storage unit 141. The image data storage control unit 140 passes the generated image data to the image processing unit 143. However, the image data storage control unit 140 can also pass the pixel data passed from the reading unit 110 to the image processing unit 143 as is.
[0144] Furthermore, based on the read area information passed from the reading unit 110, the image data storage control unit 140 generates read information indicating the read area from which the next read will be performed, and passes this information to the read determination unit 142.
[0145] Here, the image data storage control unit 140 can integrate and store already stored image data and newly passed pixel data, for example, by averaging. The image data storage control unit 140 can also delete image data stored in the image data storage unit 141 that is no longer needed. Examples of image data that is no longer needed include image data relating to the previous frame and image data that has been calculated and already stored based on a frame image of a different scene from the frame image for which the new image data was calculated. The image data storage control unit 140 can also delete and initialize all image data stored in the image data storage unit 141 as necessary.
[0146] Furthermore, the image data storage control unit 140 can acquire information for setting exposure and analog gain from the reading unit 110, and store image data corrected using the acquired information in the image data storage unit 141.
[0147] The image processing unit 143 performs predetermined image processing on the image data passed from the image data storage control unit 140. For example, the image processing unit 143 can perform predetermined image quality improvement processing on the image data. Furthermore, if the passed image data is image data in which data has been spatially reduced by line thinning or the like, it is also possible to fill in the thinned-out portions with image information by interpolation processing. The image processing unit 143 passes the image data that has been subjected to image processing to the output control unit 15.
[0148] The image processing unit 143 can receive image data from the image data storage control unit 140 and execute image processing, for example, in response to a trigger generated by the trigger generating unit 16.
[0149] The output control unit 15 outputs either or both of the recognition result passed from the recognition processing execution unit 124 and the image data passed from the image processing unit 143. The output control unit 15 outputs either or both of the recognition result and the image data in response to a trigger generated by the trigger generation unit 16, for example.
[0150] Based on the information related to the recognition processing passed from the recognition processing unit 12 and the information related to the image processing passed from the visual recognition processing unit 14, the trigger generation unit 16 generates a trigger to be passed to the recognition processing execution unit 124, a trigger to be passed to the image processing unit 143, and a trigger to be passed to the output control unit 15. The trigger generation unit 16 passes each of the generated triggers to the recognition processing execution unit 124, the image processing unit 143, and the output control unit 15 at a predetermined timing, respectively.
[0151] (5-0-2. Example of processing in the recognition processing unit according to the second embodiment) 22 is a schematic diagram showing in more detail an example of processing in the recognition processing unit 12 according to the second embodiment. Here, the readout area is a line, and the readout unit 110 reads pixel data line by line from the top to the bottom of the frame of the image 60. The line image data (line data) of the line L#x read out line by line by the readout unit 110 is input to the feature calculation unit 120.
[0152] The feature calculation unit 120 executes a feature extraction process 1200 and an integration process 1202. The feature calculation unit 120 performs the feature extraction process 1200 on the input line data to extract a feature 1201 from the line data. Here, the feature extraction process 1200 extracts the feature 1201 from the line data based on parameters obtained in advance by learning. The feature 1201 extracted by the feature extraction process 1200 is integrated with a feature 1212 processed by the feature accumulation control unit 121 by the integration process 1202. The integrated feature 1210 is passed to the feature accumulation control unit 121.
[0153] The feature accumulation control unit 121 executes internal state update processing 1211. The feature 1210 passed to the feature accumulation control unit 121 is passed to the recognition processing execution unit 124 and also subjected to internal state update processing 1211. The internal state update processing 1211 reduces the feature 1210 based on pre-learned parameters to update the internal state of the DNN and generates a feature 1212 related to the updated internal state. This feature 1212 is integrated with the feature 1201 by integration processing 1202. This processing by the feature accumulation control unit 121 corresponds to processing using an RNN.
[0154] The recognition processing execution unit 124 executes recognition processing 1240 on the feature 1210 passed from the feature accumulation control unit 121 based on parameters previously learned using, for example, predetermined training data, and outputs the recognition result.
[0155] As described above, in the recognition processing unit 12 according to the second embodiment, the feature extraction processing 1200, the integration processing 1202, the internal state update processing 1211, and the recognition processing 1240 are executed based on pre-trained parameters. The parameters are learned using training data based on an expected recognition target, for example.
[0156] The functions of the feature amount calculation unit 120, feature amount storage control unit 121, read determination unit 123, and recognition processing execution unit 124 described above are realized, for example, by a DSP included in the imaging device 1 reading and executing a program stored in the memory 13 or the like. Similarly, the functions of the image data storage control unit 140, read determination unit 142, and image processing unit 143 described above are realized, for example, by an ISP included in the imaging device 1 reading and executing a program stored in the memory 13 or the like. These programs may be stored in the memory 13 in advance, or may be supplied to the imaging device 1 from outside and written into the memory 13.
[0157] (5-0-3. Details of Recognition Processing According to Second Embodiment) Next, the second embodiment will be described in more detail. Fig. 23 is an example functional block diagram for explaining the functions according to the second embodiment. In the second embodiment, the recognition processing by the recognition processing unit 12 is the main focus, so in Fig. 23, the visual recognition processing unit 14, the output control unit 15, and the trigger generation unit 16 are omitted from the configuration of Fig. 21 described above. Also, in Fig. 23, the readout control unit 111 is omitted from the sensor control unit 11.
[0158] 24 is a schematic diagram for explaining the frame readout process according to the second embodiment. In the second embodiment, the readout unit is a line, and pixel data is read out line-sequentially for frame Fr(x). In the example of FIG. 24, in the mth frame Fr(m), lines are read out line-sequentially starting from line L#1 at the top of frame Fr(m), with lines L#2, L#3, .... After line readout in frame Fr(m) is completed, lines are similarly read out line-sequentially starting from line L#1 at the top of the next (m+1)th frame Fr(m+1).
[0159] 25 is a schematic diagram illustrating the recognition process according to the second embodiment. As shown in FIG. 25, the recognition process is performed by sequentially executing processing by the CNN 52' and updating 55 of the internal information on each of the lines L#1, L#2, L#3, etc. As a result, it is only necessary to input pixel information 54 for one line to the CNN 52', and the recognizer 56 can be configured on an extremely small scale. Note that the recognizer 56 has a configuration as an RNN, since it executes processing by the CNN 52' on the sequentially input information and updates 55 of the internal information.
[0160] By performing line-sequential recognition processing using an RNN, it may be possible to obtain a valid recognition result without reading all lines included in a frame. In this case, the recognition processing unit 12 can terminate the recognition processing when a valid recognition result is obtained. An example of terminating the recognition processing during frame reading will be described with reference to Figures 26 and 27.
[0161] 26 is a diagram showing an example in which the recognition target is the number "8." In the example of FIG. 26, the number "8" is recognized when a range 71 of about 3 / 4 of the vertical direction in frame 70 is read out. Therefore, the recognition processing unit 12 can output a valid recognition result indicating that the number "8" has been recognized when this range 71 is read out, and can end line reading and recognition processing for frame 70.
[0162] Fig. 27 is a diagram showing an example in which the recognition target is a person. In the example of Fig. 27, a person 74 is recognized when a range 73 of about half the vertical range in frame 72 is read out. Therefore, the recognition processing unit 12 can output a valid recognition result indicating that a person 74 has been recognized when this range 73 is read out, and can end line reading and recognition processing for frame 72.
[0163] In this way, in the second embodiment, if a valid recognition result is obtained during line readout for a frame, line readout and recognition processing can be terminated, which makes it possible to save power in the recognition processing and shorten the time required for the recognition processing.
[0164] In the above description, line readout is performed from the top to the bottom of the frame, but this is not limited to this example. For example, line readout may be performed from the bottom to the top of the frame. That is, an object that is far from the image capture device 1 can generally be recognized earlier by performing line readout from the top to the bottom of the frame. On the other hand, an object that is closer to the image capture device 1 can generally be recognized earlier by performing line readout from the bottom to the top of the frame.
[0165] For example, consider a case where the imaging device 1 is installed in a vehicle so as to capture an image of the front. In this case, since a near object (e.g., a vehicle or pedestrian ahead of the vehicle) is present in the lower part of the captured image, it is more effective to perform line readout from the bottom end of the frame to the top end. Furthermore, when an immediate stop is required in an ADAS (Advanced Driver-Assistance Systems), it is sufficient that at least one relevant object is recognized, and it is considered more effective to perform line readout again from the bottom end of the frame once one object is recognized. Furthermore, for example, on a highway, a distant object may be given priority. In this case, it is preferable to perform line readout from the top end of the frame to the bottom end.
[0166] Furthermore, the readout unit may be the column direction in the row and column directions in the pixel array unit 101. For example, it is possible to use a plurality of pixels arranged in one column in the pixel array unit 101 as the readout unit. By applying a global shutter method as the imaging method, column readout, in which the column is the readout unit, is possible. The global shutter method makes it possible to switch between column readout and line readout. When readout is fixed to column readout, it is possible to rotate the pixel array unit 101 by 90° and use a rolling shutter method, for example.
[0167] For example, an object on the left side of the image capture device 1 can be recognized earlier by sequentially reading out the frame from the left edge using column readout. Similarly, an object on the right side of the image capture device 1 can be recognized earlier by sequentially reading out the frame from the right edge using column readout.
[0168] In an example where the imaging device 1 is mounted on a vehicle, for example, when the vehicle is turning, priority may be given to an object on the turning side. In such a case, it is preferable to perform column readout from the end of the turning side. The turning direction can be obtained, for example, based on steering information of the vehicle. However, it is also possible to provide the imaging device 1 with a sensor capable of detecting angular velocity in three directions and obtain the turning direction based on the detection results of this sensor.
[0169] Fig. 28 is a flowchart illustrating an example of recognition processing according to the second embodiment. The processing according to the flowchart in Fig. 28 corresponds to, for example, reading pixel data in a read unit (for example, one line) from a frame. Here, the description will be given assuming that the read unit is a line. For example, the read area information can use a line number indicating the line to be read.
[0170] In step S100, the recognition processing unit 12 reads line data from a line indicated by the read line of the frame. More specifically, the read determination unit 123 in the recognition processing unit 12 passes the line number of the line to be read next to the sensor control unit 11. In the sensor control unit 11, the read unit 110 reads pixel data of the line indicated by the line number from the sensor unit 10 as line data in accordance with the passed line number. The read unit 110 passes the line data read from the sensor unit 10 to the feature calculation unit 120. The read unit 110 also passes read area information (e.g., line number) indicating the area from which pixel data has been read to the feature calculation unit 120.
[0171] In the next step S101, the feature amount calculation unit 120 calculates feature amounts based on the line data, based on the pixel data passed from the reading unit 110, and calculates feature amounts for the lines. In the next step S102, the feature amount calculation unit 120 acquires feature amounts accumulated in the feature amount accumulation unit 122 from the feature amount accumulation control unit 121. In the next step S103, the feature amount calculation unit 120 integrates the feature amount calculated in step S101 with the feature amount acquired from the feature amount accumulation control unit 121 in step S102. The integrated feature amount is passed to the feature amount accumulation control unit 121. The feature amount accumulation control unit 121 accumulates the integrated feature amount passed from the feature amount calculation unit 120 in the feature amount accumulation unit 122 (step S104).
[0172] If the series of processes from step S100 are for the first line of a frame and the feature amount storage unit 122 has been initialized, for example, the processes in steps S102 and S103 can be omitted. In this case, the process in step S104 is to store the line feature amount calculated based on the first line in the feature amount storage unit 122.
[0173] The feature amount accumulation control unit 121 also passes the integrated feature amounts passed from the feature amount calculation unit 120 to the recognition process execution unit 124. In step S105, the recognition process execution unit 124 executes recognition processing using the integrated feature amounts passed from the feature amount accumulation control unit 121. In the next step S106, the recognition process execution unit 124 outputs the recognition result obtained by the recognition process in step S105.
[0174] In step S107, the read determination unit 123 in the recognition processing unit 12 determines the read line to be read next according to the read information passed from the feature accumulation control unit 121. For example, the feature accumulation control unit 121 receives read area information together with the feature from the feature calculation unit 120. Based on this read area information, the feature accumulation control unit 121 determines the read line to be read next according to, for example, a pre-specified read pattern (line sequential in this example). The processing from step S100 is executed again for this determined read line.
[0175] (5-0-4. Example of control of readout and recognition processing according to the second embodiment) Next, an example of control of readout and recognition processing according to the second embodiment will be described. Figures 29A and 29B are time charts showing an example of control of readout and recognition processing according to the second embodiment. The example of Figures 29A and 29B is an example in which a blank period blk in which no imaging operation is performed is provided within one imaging cycle (one frame cycle). In Figures 29A and 29B, time progresses to the right.
[0176] FIG. 29A shows an example in which, for example, 1 / 2 of the imaging cycle is consecutively allocated to blank periods blk. In FIG. 29A, the imaging cycle is a frame cycle, for example, 1 / 30 [sec]. Frames are read out from the sensor unit 10 at this frame cycle. The imaging time is the time required to capture images of all lines included in a frame. In the example of FIG. 29A, a frame includes n lines, and imaging of n lines L#1 to L#n is completed in 1 / 60 [sec], which is 1 / 2 of the frame cycle of 1 / 30 [sec]. The time allocated to capturing images of one line is 1 / (60×n) [sec]. The blank period blk is the 1 / 30 [sec] period from the timing at which the last line L#n in a frame is captured to the timing at which the first line L#1 of the next frame is captured.
[0177] For example, when imaging of line L#1 is completed, imaging of the next line L#2 is started, and the recognition processing unit 12 executes line recognition processing for that line L#1, i.e., recognition processing for the pixel data included in line L#1. The recognition processing unit 12 completes the line recognition processing for line L#1 before imaging of the next line L#2 is started. When the line recognition processing for line L#1 is completed, the recognition processing unit 12 outputs the recognition result of the recognition processing.
[0178] Similarly, for the next line L#2, imaging of the next line L#3 is started at the timing when imaging of the line L#2 is completed, and the recognition processing unit 12 executes line recognition processing for the line L#2, and the executed line recognition processing is terminated before imaging of the next line L#3 is started. In the example of Fig. 29A, imaging of the lines L#1, L#2, L#3, ..., L#m, ... is executed sequentially in this manner. Then, for each of the lines L#1, L#2, L#3, ..., L#m, ..., imaging of the line following the line whose imaging has been completed is started at the timing when imaging is completed, and line recognition processing is executed for the line whose imaging has been completed.
[0179] In this way, by sequentially executing the recognition process for each read unit (in this example, each line), it is possible to sequentially obtain recognition results without inputting all of the image data of the frame into the recognizer (recognition processing unit 12), thereby reducing the delay until the recognition result is obtained. Also, when a valid recognition result is obtained for a certain line, the recognition process can be terminated at that point, thereby shortening the recognition process time and saving power. Also, by propagating and integrating information on the time axis for the recognition results of each line, it is possible to gradually improve the recognition accuracy.
[0180] In the example of FIG. 29A, other processing that should be executed within a frame period (for example, image processing in the visual recognition processing unit 14 using the recognition result) can be executed during the blank period blk within the frame period.
[0181] Fig. 29B shows an example in which a blank period blk is provided for each line imaged. In the example of Fig. 29B, the frame period (image capture period) is 1 / 30 [sec], similar to the example of Fig. 29A. Meanwhile, the image capture time is 1 / 30 [sec], the same as the image capture period. Also, in the example of Fig. 29B, in one frame period, images of n lines, line L#1 to line L#n, are captured at time intervals of 1 / (30×n) [sec], and the image capture time for one line is 1 / (60×n) [sec].
[0182] In this case, a blank period blk of 1 / (60×n) [sec] can be provided for each imaging of each line L#1 to L#n. During each blank period blk of each line L#1 to L#n, other processing to be performed on the imaging image of the corresponding line (e.g., image processing in the visual recognition processing unit 14 using the recognition result) can be performed. At this time, the time (approximately 1 / (30×n) [sec] in this example) until just before imaging of the line next to the target line is completed can be allocated to this other processing. In the example of FIG. 29B, the processing results of this other processing can be output for each line, making it possible to obtain the processing results of this other processing more quickly.
[0183] Fig. 30 is a time chart showing another example of control of the readout and recognition processes according to the second embodiment. In the example of Fig. 29 described above, imaging of all lines L#1 to L#n included in a frame is completed in half the frame period, and the remaining half of the frame period is a blank period. In contrast, in the example shown in Fig. 30, no blank period is provided within the frame period, and imaging of all lines L#1 to L#n included in a frame is performed using the entire frame period.
[0184] Here, if the imaging time for one line is 1 / (60×n) [sec], the same as in Figures 29A and 29B, and the number of lines included in a frame is n, the same as in Figures 29A and 29B, the frame period, i.e., imaging period, is 1 / 60 [sec]. Therefore, in the example shown in Figure 30 where no blank period blk is provided, the frame rate can be increased compared to the examples in Figures 29A and 29B described above.
[0185] [5-1. First Modification of the Second Embodiment] Next, a first modified example of the second embodiment will be described. The first modified example of the second embodiment is an example in which the read unit is a plurality of sequentially adjacent lines. Note that the configuration described using FIG. 23 can be applied as is to the first modified example of the second embodiment and the second to seventh modified examples of the second embodiment described later, so detailed description of the configuration will be omitted.
[0186] Fig. 31 is a schematic diagram for explaining a frame readout process according to a first modified example of the second embodiment. As shown in Fig. 31, in the first modified example of the second embodiment, a line group including a plurality of lines adjacent to each other is used as a readout unit, and pixel data is read out for the frame Fr(m) in line group order. In the recognition processing unit 12, the readout determination unit 123 determines, for example, a line group Ls#x including a predetermined number of lines as a readout unit.
[0187] The read determination unit 123 passes read area information, which is information indicating the read unit determined as the line group Ls#x and read position information for reading out pixel data of the read unit, to the read control unit 111. The read control unit 111 passes the read area information passed from the read determination unit 123 to the read unit 110. The read unit 110 reads out pixel data from the sensor unit 10 in accordance with the read area information passed from the read control unit 111.
[0188] 31, in the mth frame Fr(m), line groups Ls#2, Ls#3, ..., Ls#p, ... and line group Ls#x are read out line by line starting from line group Ls#1 at the top of frame Fr(m). After reading of line group Ls#x in frame Fr(m) is completed, in the next (m+1)th frame Fr(m+1), line groups Ls#2, Ls#3, ... are read out line by line starting from line group Ls#1 at the top.
[0189] In this way, by reading pixel data using a line group Ls#x including multiple lines as a readout unit, it is possible to read one frame's worth of pixel data faster than when reading is performed line-sequentially. Furthermore, the recognition processing unit 12 can use more pixel data in one recognition process, thereby improving the recognition response speed. Furthermore, since the number of readouts per frame is reduced compared to line-sequential reading, it is possible to suppress distortion of the captured frame image when the imaging method of the sensor unit 10 is the rolling shutter method.
[0190] In the first modification of the second embodiment, similarly to the second embodiment described above, the reading of the line group Ls#x may be performed from the bottom to the top of the frame. Furthermore, if a valid recognition result is obtained during the reading of the line group Ls#x for a frame, the reading of the line group and the recognition process can be terminated. This allows for power saving in the recognition process and shortens the time required for the recognition process.
[0191] [5-2. Second Modification of the Second Embodiment] Next, a second modified example of the second embodiment will be described. The second modified example of the second embodiment is an example in which the read unit is a part of one line.
[0192] Fig. 32 is a schematic diagram for explaining a frame readout process according to a second modification of the second embodiment. As shown in Fig. 32, in the second modification of the second embodiment, in line readout performed line-sequentially, a part of each line (referred to as a partial line) is used as a readout unit, and pixel data is read out line-sequentially for the frame Fr(m) for partial lines Lp#x in each line. In the recognition processing unit 12, the readout determination unit 123 determines, for example, from among the pixels included in a line, a plurality of pixels that are adjacent to each other in sequence and whose number is less than the total number of pixels included in the line as a readout unit.
[0193] The read determination unit 123 passes read area information, which is information indicating the read unit determined as the partial line Lp#x and read position information for reading out pixel data of the partial line Lp#x, to the read control unit 111. Here, the information indicating the read unit may be composed of, for example, the position of the partial line Lp#x within one line and the number of pixels included in the partial line Lp#x. Furthermore, the read position information may use the number of the line including the partial line Lp#x to be read. The read control unit 111 passes the read area information passed from the read determination unit 123 to the read unit 110. The read unit 110 reads pixel data from the sensor unit 10 in accordance with the read area information passed from the read control unit 111.
[0194] 32, in the mth frame Fr(m), partial lines Lp#1 included in the top line of frame Fr(m) are read out line by line, starting with partial line Lp#1, followed by partial lines Lp#2, Lp#3, ... and each partial line Lp#x included in each line. After reading of the line group in frame Fr(m) is completed, in the next (m+1)th frame Fr(m+1), partial lines Lp#x are similarly read out line by line, starting with partial line Lp#1 included in the top line.
[0195] In this way, in line readout, by limiting the pixels to be read out to those included in a portion of the line, it is possible to transfer pixel data in a narrower bandwidth than when pixel data is read out from the entire line. In the readout method according to the second modification of the second embodiment, the amount of pixel data transferred is smaller than when pixel data is read out from the entire line, thereby enabling power savings.
[0196] In the second modification of the second embodiment, as in the second embodiment described above, partial lines may be read from the bottom to the top of the frame. Furthermore, if a valid recognition result is obtained during the reading of partial lines for a frame, the reading of partial line Lp#x and the recognition process can be terminated. This allows for power saving in the recognition process and shortens the time required for the recognition process.
[0197] [5-3. Third Modification of the Second Embodiment] Next, a third modified example of the second embodiment will be described. The third modified example of the second embodiment is an example in which the readout unit is an area of a predetermined size within a frame.
[0198] Fig. 33 is a schematic diagram for explaining a frame readout process according to a third modified example of the second embodiment. As shown in Fig. 33, in the third modified example of the second embodiment, an area Ar#xy of a predetermined size including a plurality of pixels adjacent to each other in both the line direction and the vertical direction within a frame is used as a readout unit, and for a frame Fr(m), the area Ar#xy is read out sequentially in the line direction, for example, and further, the sequential readout of the area Ar#xy in the line direction is repeated sequentially in the vertical direction. In the recognition processing unit 12, the readout determination unit 123 determines the area Ar#xy defined by, for example, the size in the line direction (number of pixels) and the size in the vertical direction (number of lines) as the readout unit.
[0199] The read determination unit 123 passes read area information, which is information indicating the read unit determined as the area Ar#xy and read position information for reading pixel data of the area Ar#xy, to the read control unit 111. Here, the information indicating the read unit may be composed of, for example, the above-mentioned line size (number of pixels) and vertical size (number of lines). Furthermore, the read position information may be the position of a predetermined pixel included in the area Ar#xy to be read, for example, the pixel position of the pixel in the upper left corner of the area Ar#xy. The read control unit 111 passes the read area information passed from the read determination unit 123 to the read unit 110. The read unit 110 reads pixel data from the sensor unit 10 in accordance with the read area information passed from the read control unit 111.
[0200] 33, in the mth frame Fr(m), each area Ar#xy is read out sequentially in the line direction, starting from area Ar#1-1 located in the upper left corner of frame Fr(m), then areas Ar#2-1, Ar#3-1, ... In frame Fr(m), when reading out has been completed up to the right end in the line direction, the vertical read position is moved, and again each area Ar#xy is read out sequentially in the line direction, starting from the left end of frame Fr(m), then areas Ar#1-2, Ar#2-2, Ar#3-2, ...
[0201] FIG. 34 is a schematic diagram illustrating a recognition process according to a third modified example of the second embodiment. As shown in FIG. 34, the recognition process is performed by sequentially executing processing by the CNN 52′ and updating 55 of the internal information for each of the pixel information 54 of each of the areas Ar#1-1, Ar#2-1, Ar#3-1, .... Therefore, it is sufficient to input pixel information 54 for one area to the CNN 52′, and the recognizer 56 can be configured on an extremely small scale. Note that the recognizer 56 has a configuration as an RNN, since it executes processing by the CNN 52′ on the sequentially input information and updates 55 of the internal information.
[0202] By performing line-sequential recognition processing using an RNN, it may be possible to obtain a valid recognition result without reading all lines included in a frame. In this case, the recognition processing unit 12 can terminate the recognition processing when a valid recognition result is obtained. An example of terminating the recognition processing midway through frame reading when the reading unit is area Ar#xy will be described with reference to Figures 35 and 36.
[0203] 35 is a diagram showing an example in which the recognition target is the number "8." In the example of FIG. 35, the number "8" is recognized at position P1 when a range 81, which is about two-thirds of the entire frame 80, is read out. Therefore, the recognition processing unit 12 can output a valid recognition result indicating that the number "8" has been recognized when this range 81 is read out, and can end line reading and recognition processing for frame 80.
[0204] Fig. 36 is a diagram showing an example in which the recognition target is a person. In the example of Fig. 36, a person 84 is recognized at position P2 at the time when a range 83 of about half the vertical direction in frame 82 is read out. Therefore, the recognition processing unit 12 can output a valid recognition result indicating that a person 84 has been recognized at the time when this range 83 is read out, and can end line reading and recognition processing for frame 82.
[0205] As described above, in the third modification of the second embodiment, if a valid recognition result is obtained during area readout for a frame, the area readout and recognition process can be terminated. This enables power saving in the recognition process and shortens the time required for the recognition process. Furthermore, in the third modification of the second embodiment, redundant readout is reduced compared to the example in which readout is performed across the entire width in the line direction, as in the second embodiment and the first modification of the second embodiment described above, and the time required for the recognition process can be shortened.
[0206] In the above description, the area Ar#xy is read from the left edge to the right edge in the line direction and from the top edge to the bottom edge of the frame in the vertical direction, but this is not limited to this example. For example, the line direction reading may be performed from the right edge to the left edge, and the vertical direction reading may be performed from the bottom edge to the top edge of the frame.
[0207] 5-4. Fourth Modification of the Second Embodiment Next, a fourth modified example of the second embodiment will be described. The fourth modified example of the second embodiment is an example in which the readout unit is a pattern made up of a plurality of pixels including non-adjacent pixels.
[0208] Fig. 37 is a schematic diagram for explaining a frame readout process according to a fourth modified example of the second embodiment. As shown in Fig. 37, in the fourth modified example of the second embodiment, a pattern Pφ#xy consisting of a plurality of pixels arranged discretely and periodically in each of the line direction and the vertical direction is used as a readout unit. In the example of Fig. 37, the pattern Pφ#xy is composed of six periodically arranged pixels: three pixels arranged at a predetermined interval in the line direction and three pixels arranged at a predetermined interval in the vertical direction so that their positions in the line direction correspond to those of the three pixels. In the recognition processing unit 12, the readout determination unit 123 determines the plurality of pixels arranged according to this pattern Pφ#xy as a readout unit.
[0209] In the above description, the pattern Pφ#xy is described as being composed of a plurality of discrete pixels, but this is not limited to this example. For example, the pattern Pφ#xy may be composed of a plurality of pixel groups each including a plurality of adjacent pixels arranged discretely. For example, the pattern Pφ#xy may be composed of a plurality of pixel groups each consisting of four adjacent pixels (2 pixels x 2 pixels) arranged discretely and periodically as shown in FIG. 37.
[0210] The read determination unit 123 passes read area information, which is information indicating the read unit determined as the pattern Pφ#xy and read position information for reading out the pattern Pφ#xy, to the read control unit 111. Here, the information indicating the read unit may be composed of, for example, information indicating the positional relationship between a predetermined pixel among the pixels constituting the pattern Pφ#xy (for example, the pixel in the upper left corner among the pixels constituting the pattern Pφ#xy) and each of the other pixels constituting the pattern Pφ#xy. Furthermore, the read position information may be information indicating the position of a predetermined pixel included in the pattern Pφ#xy to be read (information indicating the position within the line and the line number). The read control unit 111 passes the read area information passed from the read determination unit 123 to the read unit 110. The read unit 110 reads pixel data from the sensor unit 10 in accordance with the read area information passed from the read control unit 111.
[0211] 37, in the mth frame Fr(m), for example, patterns Pφ#1-1, whose upper left corner pixel is located at the upper left corner of frame Fr(m), are read out sequentially in the line direction, for example, by one pixel at a time, starting with pattern Pφ#1-1, and patterns Pφ#2-1, Pφ#3-1, .... For example, when the right end of pattern Pφ#xy reaches the right end of frame Fr(m), the position is shifted vertically by one pixel (one line) from the left end of frame Fr(m), and patterns Pφ#1-2, Pφ#2-2, Pφ#3-2, ... are read out in the same manner.
[0212] Because the pattern Pφ#xy is configured by periodically arranging pixels, the operation of shifting the pattern Pφ#xy by one pixel can be considered an operation of shifting the phase of the pattern Pφ#xy. That is, in the fourth modification of the second embodiment, each pattern P#xy is read out while shifting the phase of the pattern Pφ#xy in the line direction by Δφ. The pattern Pφ#xy is moved in the vertical direction by shifting the phase Δφ' in the vertical direction relative to the position of the first pattern Pφ#1-y in the line direction, for example.
[0213] FIG. 38 is a schematic diagram illustrating a recognition process applicable to a fourth modification of the second embodiment. FIG. 38 illustrates an example in which a pattern Pφ#z is configured by four pixels spaced one pixel apart in the horizontal (line) and vertical directions. As shown in sections (a), (b), (c), and (d) of FIG. 38, patterns Pφ#1, Pφ#2, Pφ#3, and Pφ#4, each consisting of four pixels shifted in phase by one pixel in the horizontal and vertical directions, enable all 16 pixels included in a 4×4 pixel area to be read out without overlapping. The four pixels read out according to patterns Pφ#1, Pφ#2, Pφ#3, and Pφ#4, respectively, are subsamples Sub#1, Sub#2, Sub#3, and Sub#4, extracted from the 16 pixels included in the 4×4 pixel sample area without overlapping each other.
[0214] In the example of sections (a) to (d) in Fig. 38, the recognition process is performed by executing the process by the CNN 52' and the internal information update 55 for each of the subsamples Sub#1, Sub#2, Sub#3, and Sub#4. Therefore, only four pixel data items need to be input to the CNN 52' for one recognition process, and the recognizer 56 can be configured on an extremely small scale.
[0215] Fig. 39 is a time chart showing an example of readout and control according to the fourth modified example of the second embodiment. In Fig. 39, the imaging period is a frame period, for example, 1 / 30 [sec]. Frames are read out from the sensor unit 10 at this frame period. The imaging time is the time required to capture images of all subsamples Sub#1 to Sub#4 included in a frame. In the example of Fig. 39, the imaging time is 1 / 30 [sec], the same as the imaging period. Note that imaging of one subsample Sub#x is called subsample imaging.
[0216] In the fourth modification of the second embodiment, the imaging time is divided into four periods, and sub-sample imaging of each of the sub-samples Sub#1, Sub#2, Sub#3, and Sub#4 is performed in each period.
[0217] More specifically, the sensor control unit 11 performs subsample imaging using subsample Sub#1 over the entire frame during the first period of the first to fourth periods into which the imaging time is divided. The sensor control unit 11 extracts subsample Sub#1, for example, by moving a 4-pixel by 4-pixel sample area in the line direction so that the sample areas do not overlap. The sensor control unit 11 repeatedly performs the operation of extracting subsample Sub#1 while moving the sample area in the line direction in the vertical direction.
[0218] When extraction of subsamples Sub#1 for one frame is completed, the recognition processing unit 12 inputs the extracted subsamples Sub#1 for one frame to the recognizer 56 for each subsample Sub#1, for example, and executes recognition processing. The recognition processing unit 12 outputs the recognition result after completing recognition processing for one frame. However, the recognition processing unit 12 may output the recognition result if a valid recognition result is obtained during recognition processing for one frame, and terminate recognition processing for that subsample Sub#1.
[0219] Thereafter, in the second, third and fourth periods, sub-sample imaging is similarly performed over the entire frame using sub-samples Sub#2, Sub#3 and Sub#4, respectively.
[0220] The frame readout process according to the fourth modified example of the second embodiment will be described in more detail with reference to Figures 40 and 41. Figure 40 is a diagram showing an example in which the recognition target is the number "8" and three numbers "8" of different sizes are included in one frame. In sections (a), (b), and (c) of Figure 40, frame 90 includes three objects 93L, 93M, and 93S of different sizes, each showing the number "8." Of these, object 93L is the largest and object 93S is the smallest.
[0221] In section (a) of Figure 40, subsample Sub#1 is extracted from sample area 92. By extracting subsample Sub#1 from each sample area 92 included in frame 90, pixels are read out in a grid pattern from frame 90 at intervals of every other pixel in both the horizontal and vertical directions, as shown in section (a) of Figure 40. In the example of section (a) of Figure 40, recognition processing unit 12 recognizes object 93L, which is the largest of objects 93L, 93M, and 93S, based on the pixel data of the pixels read out in this grid pattern.
[0222] After extraction of subsample Sub#1 from frame 90, extraction of subsample Sub#2 from sample area 92 is performed. Subsample Sub#2 is composed of pixels that are shifted by one pixel horizontally and one pixel vertically within sample area 92 relative to subsample Sub#1. Because the internal state of recognizer 56 is updated based on the recognition result of subsample Sub#1, the recognition result for subsample Sub#2 is affected by the recognition process for subsample Sub#1. Therefore, the recognition process for subsample Sub#2 can be considered to be performed based on the pixel data of pixels read in a checkerboard pattern, as shown in section (b) of Figure 40. Therefore, the state in which subsample Sub#2 is further extracted, as shown in section (b) of Figure 40, has improved pixel data-based resolution compared to the state in which only subsample Sub#1 is extracted, as shown in section (a) of Figure 40, enabling more accurate recognition processing. In the example of section (b) of Figure 40, object 93M, which is the next largest object after object 93L, is also recognized.
[0223] Section (c) of Figure 40 shows the state after extraction of all subsamples Sub#1 to Sub#4 included in sample area 92 has been completed for frame 90. In section (c) of Figure 40, all pixels included in frame 90 have been read out, and in addition to objects 93L and 93M recognized in the extraction of subsamples Sub#1 and Sub#2, the smallest object 93S has also been recognized.
[0224] Fig. 41 is a diagram showing an example in which the recognition targets are people, and images of three people at different distances from the imaging device 1 are included in one frame. In sections (a), (b), and (c) of Fig. 41, three objects 96L, 96M, and 96S, each of which is an image of a person and is of different sizes, are included in frame 95. Of these, object 96L is the largest, and of the three people included in frame 95, the person corresponding to object 96L is the person closest to the imaging device 1. Furthermore, object 96S, the smallest of objects 96L, 96M, and 96S, represents the person corresponding to object 96S who is the farthest from the imaging device 1 among the three people included in frame 95.
[0225] In Figure 41, section (a) corresponds to section (a) of Figure 40, and is an example in which the above-mentioned subsample Sub#1 is extracted and the recognition process is performed, resulting in the recognition of the largest object 96L. Section (b) of Figure 41 corresponds to section (b) of Figure 40, and is an example in which subsample Sub#2 is further extracted from the state of section (a) of Figure 41 and the recognition process is performed, resulting in the recognition of the next largest object 96M. Section (c) of Figure 41 corresponds to section (c) of Figure 40, and is an example in which subsamples Sub#3 and Sub#4 are further extracted from the state of section (b) of Figure 41 and the recognition process is performed based on the pixel data of all pixels included in frame 95. Section (c) of Figure 41 shows the recognition of objects 96L and 96M, as well as the smallest object 96S.
[0226] In this way, by extracting subsamples Sub#1, Sub#2, ... and repeating the recognition process, it becomes possible to recognize people who are successively farther away.
[0227] 40 and 41, frame readout and recognition processing can be controlled depending on the time allocable to the recognition processing. As an example, if the time allocable to the recognition processing is short, frame readout and recognition processing can be terminated when extraction of subsample Sub#1 in frame 90 is completed and object 93L is recognized. On the other hand, if the time allocable to the recognition processing is long, frame readout and recognition processing can be continued until extraction of all subsamples Sub#1 to Sub#4 is completed.
[0228] Alternatively, the recognition processing unit 12 may control frame readout and recognition processing according to the reliability (score) of the recognition result. For example, in section (b) of Fig. 41, if the score for the recognition result based on the extraction and recognition processing of subsample Sub#2 is equal to or greater than a predetermined value, the recognition processing unit 12 may terminate the recognition processing and not extract the next subsample Sub#3.
[0229] In this way, in the fourth variant of the second embodiment, the recognition process can be terminated when a predetermined recognition result is obtained, which enables a reduction in the amount of processing in the recognition processing unit 12 and power saving.
[0230] Furthermore, in the fourth modified example of the second embodiment, the recognition response speed for large objects within a frame can be increased, making it possible to increase the frame rate.
[0231] 5-5. Fifth Modification of the Second Embodiment Next, a fifth modified example of the second embodiment will be described. The fifth modified example of the second embodiment is an example in which the readout unit is a pattern in which a plurality of pixels, including non-adjacent pixels, are randomly arranged.
[0232] Fig. 42 is a schematic diagram for explaining a frame readout process according to a fifth modified example of the second embodiment. As shown in Fig. 42, in the fifth modified example of the second embodiment, for example, a pattern Rd#m_x composed of a plurality of pixels arranged discretely and non-periodically within a frame Fr(m) is used as a readout unit. That is, the readout unit according to the fifth modified example of the second embodiment is the entire frame.
[0233] In a fifth modification of the second embodiment, referring to the above-mentioned Fig. 39, one frame period is divided into multiple periods, and a pattern is switched for each period. In the example of Fig. 42, in the first period into which the frame period of the mth frame Fr(m) is divided, the recognition processing unit 12 reads pixels according to a pattern Rd#m_1 consisting of multiple pixels that are discretely and non-periodically arranged in the frame Fr(m) and performs recognition processing. As an example, when the total number of pixels included in the frame Fr(m) is s and the number of divisions of the frame period is D, the recognition processing unit 12 selects (s / D) pixels that are discretely and non-periodically arranged in the frame Fr(m) to form the pattern Rd#m_1.
[0234] In the next period into which the frame cycle is divided, the recognition processing unit 12 reads out pixels according to pattern Rd#m_2 in which pixels different from those in pattern Rd#m_1 are selected in frame Fr(m), and performs recognition processing.
[0235] Similarly, in the next (m+1)th frame Fr(m+1), in the first period into which the frame period of frame Fr(m+1) is divided, the recognition processing unit 12 reads out pixels according to pattern Rd#(m+1)_1, which is made up of multiple pixels that are discretely and non-periodically arranged in frame Fr(m+1), and performs recognition processing. In the next period, the recognition processing unit 12 reads out pixels according to pattern Rd#(m+1)_2, in which pixels different from those in pattern Rd#(m+1)_1 are selected, and performs recognition processing.
[0236] In the recognition processing unit 12, the read determination unit 123, for example, in the first period into which the frame period of frame Fr(m) is divided, selects a predetermined number of pixels from all pixels included in frame Fr(m) based on pseudo-random numbers to determine pattern Rd#m_1 as a readout unit. In the next period, for example, the read determination unit 123 selects a predetermined number of pixels from all pixels included in frame Fr(m) excluding the pixels selected in pattern Rd#m_1 based on pseudo-random numbers to determine pattern Rd#m_2 as a readout unit. However, the recognition processing unit 12 may again select a predetermined number of pixels from all pixels included in frame Fr(m) based on pseudo-random numbers to determine pattern Rd#m_2 as a readout unit.
[0237] The read determination unit 123 passes read area information, which is information indicating the read unit determined as pattern Rd#m_x and read position information for reading out pixel data of the pattern Rd#m_x, to the read control unit 111. The read control unit 111 passes the read area information passed from the read determination unit 123 to the read unit 110. The read unit 110 reads pixel data from the sensor unit 10 in accordance with the read area information passed from the read control unit 111.
[0238] Here, the information indicating the readout unit may be configured, for example, by position information (e.g., information indicating the line number and pixel position within the line) of each pixel included in the pattern Rd#m_1 in the frame Fr(m). In addition, since the readout unit in this case targets the entire frame Fr(m), the readout position information can be omitted. Information indicating the position of a specific pixel within the frame Fr(m) may also be used as the readout position information.
[0239] In this manner, in the fifth modified example of the second embodiment, the frame readout process is performed using a pattern Rd#m_x consisting of a plurality of pixels that are discretely and non-periodically arranged from all pixels of the frame Fr(m). Therefore, compared to using a periodic pattern, it is possible to reduce sampling artifacts. For example, the frame readout process according to the fifth modified example of the second embodiment can suppress false detection or non-detection of temporal periodic patterns (e.g., flicker) in the recognition process. Furthermore, the frame readout process can also suppress false detection or non-detection of spatial periodic patterns (e.g., fences or mesh-like structures) in the recognition process.
[0240] Furthermore, with this frame readout process, the amount of pixel data available for recognition processing increases over time, which makes it possible to speed up the recognition response speed for large objects within the frame Fr(m), for example, and increase the frame rate.
[0241] In the above description, the recognition processing unit 12 generates each pattern Rd#m_x each time, but this is not limited to this example. For example, each pattern Rd#m_x may be generated in advance and stored in a memory, and the readout determination unit 123 may read out and use each stored pattern Rd#m_x from the memory.
[0242] [5-6. Sixth Modification of the Second Embodiment] Next, a sixth modified example of the second embodiment will be described. The sixth modified example of the second embodiment is an example in which the configuration of the readout unit is changed depending on the result of the recognition process. Fig. 43 is a schematic diagram for explaining the frame readout process according to the sixth modified example of the second embodiment. Here, the explanation will be given taking as an example the readout unit by a pattern consisting of a plurality of pixels including non-adjacent pixels, which was explained using Fig. 37.
[0243] 43, in the mth frame Fr(m), similar to the pattern Pφ#xy described with reference to FIG. 37, the read determination unit 123 generates a pattern Pt#xy consisting of a plurality of pixels arranged discretely and periodically in both the line direction and the vertical direction, and sets this as an initial read unit. The read determination unit 123 passes read area information, in which read position information for reading out the pattern Pt#xy is added to information indicating the read unit determined as the pattern Pt#xy, to the read control unit 111. The read control unit 111 passes the read area information passed from the read determination unit 123 to the read unit 110. The read unit 110 reads pixel data from the sensor unit 10 in accordance with the read area information passed from the read control unit 111.
[0244] As shown in Figure 43, in frame Fr(m), the recognition processing unit 12 reads out and performs recognition processing on patterns Pt#1-1, Pt#2-1, Pt#3-1 while shifting their position horizontally from the left end by a phase Δφ, and when the right end of pattern Pt#xy reaches the right end of frame Fr(m), it again reads out and performs recognition processing on patterns Pt#1-2, ... while shifting their position vertically by a phase Δφ' and horizontally by a phase Δφ from the right end of frame Fr(m).
[0245] In response to the recognition result for frame Fr(m), the recognition processing unit 12 generates a new pattern Pt'#xy. As an example, it is assumed that the recognition processing unit 12 recognizes a target object (e.g., a person) in the center of frame Fr(m) during the recognition process for frame Fr(m). In the recognition processing unit 12, in response to this recognition result, the read determination unit 123 generates a pattern Pt'#xy that concentrates on reading pixels in the center of frame Fr(m) as a new read unit.
[0246] The read determination unit 123 can generate the pattern Pt'#x-1 using fewer pixels than the pattern Pt#xy. Also, the read determination unit 123 can make the pixel arrangement of the pattern Pt'#xy denser than the pixel arrangement in the pattern Pt#xy.
[0247] The read determination unit 123 passes read area information, which is information indicating the read unit determined as pattern Pt'#xy and read position information for reading out the pattern Pt'#xy, to the read control unit 111. The read determination unit 123 then applies this pattern Pt'#xy to the next frame Fr(m+1). The read control unit 111 passes the read area information passed from the read determination unit 123 to the read unit 110. The read unit 110 reads pixel data from the sensor unit 10 in accordance with the read area information passed from the read control unit 111.
[0248] 43, in frame Fr(m+1), the recognition processor 12 first reads and recognizes pattern Pt'#1-1 in the center of frame Fr(m+1), then reads and recognizes pattern Pt'#2-1 while shifting its position horizontally by, for example, a phase Δφ. Furthermore, the recognition processor 12 reads patterns Pt'#1-2 and Pt'#2-2 while shifting their positions vertically from the position of pattern Pt'#1-1 by a phase Δφ' and further shifting their positions horizontally by a phase Δφ.
[0249] In this way, in the sixth modification of the second embodiment, a pattern Pt'#xy to be used for reading pixels in the next frame Fr(m+1) is generated based on the recognition result in frame Fr(m) based on the pattern Pt#xy, which is the initial pattern. This enables more accurate recognition processing. Furthermore, by using a new pattern Pt'#xy generated according to the result of the recognition processing, the recognition processing is narrowed down to the part where the object is recognized, thereby reducing the amount of processing in the recognition processing unit 12, saving power, and improving the frame rate.
[0250] (Another example of the sixth modified example) Next, another example of the sixth modified example of the second embodiment will be described. Fig. 44 is a diagram showing an example of a pattern for performing a readout and recognition process according to the sixth modified example of the second embodiment. The pattern Cc#x shown in Fig. 44 is an annular shape, and is a readout unit in which the radius of the annulus changes over time. In the example of Fig. 44, in the first period into which the frame period of the frame Fr(m) is divided, a pattern Cc#1 with a small radius is used, in the next period a pattern Cc#2 with a radius larger than that of the pattern Cc#1 is used, and in the period after that a pattern Cc#3 with a radius even larger than that of the pattern Cc#2 is used.
[0251] For example, as shown in the above-mentioned Figure 43, in frame Fr(m), the recognition processing unit 12 reads out and performs recognition processing on patterns Pt#1-1, Pt#2-1, Pt#3-1 while shifting their position horizontally from the left end side by a phase Δφ, and when the right end of pattern Pt#xy reaches the right end of frame Fr(m), it again reads out and performs recognition processing on patterns Pt#1-2, ... while shifting their position vertically by a phase Δφ' and horizontally by a phase Δφ from the right end side of frame Fr(m).
[0252] In response to the recognition result for frame Fr(m), the recognition processing unit 12 generates a new pattern Cc#1 having a ring shape. As an example, it is assumed that the recognition processing unit 12 recognizes a target object (e.g., a person) in the center of frame Fr(m) during the recognition process for frame Fr(m). In the recognition processing unit 12, the read determination unit 123 generates patterns Cc#1, Cc#2, ... in response to this recognition result, and performs readout and recognition processes based on these patterns Cc#1, Cc#2, ....
[0253] In FIG. 44, the radius of the pattern Cc#m is increased over time, but this is not limited to this example, and the radius of the pattern Cc#m may be decreased over time.
[0254] As another example of the sixth modified example of the second embodiment, the density of pixels in the readout pattern may be changed. In addition, in the pattern Cc#m shown in Fig. 44, the size of the pixels changes from the center to the periphery of the annular shape or from the periphery to the center, but this is not limited to this example.
[0255] [5-7. Seventh Modification of the Second Embodiment] Next, a seventh modified example of the second embodiment will be described. In the second embodiment and the first to fourth modified examples of the second embodiment described above, the lines, areas, and patterns for reading out pixels are moved sequentially according to the order of coordinates within a frame (line numbers, order of pixels within a line, etc.). In contrast, in the seventh modified example of the second embodiment, the lines, areas, and patterns for reading out pixels are set so that the pixels within a frame can be read out more uniformly in a short period of time.
[0256] Fig. 45 is a schematic diagram for explaining a first example of a frame readout process according to the seventh modification of the second embodiment. For the sake of explanation, in Fig. 45, it is assumed that the frame Fr(m) includes eight lines: lines L#1, L#2, L#3, L#4, L#5, L#6, L#7, and L#8.
[0257] The readout process shown in section (a) of Figure 45 corresponds to the readout process described in the second embodiment using Figure 24, and shows an example in which the readout unit is a line, and pixel data is read out line-sequentially for frame Fr(m), in the order of lines L#1, L#2, ..., L#8. In the example of section (a), there is a large delay from when readout of frame Fr(m) starts until the pixel data for the bottom part of frame Fr(m) is obtained.
[0258] Section (b) of Figure 45 shows an example of a readout process according to a seventh modified example of the second embodiment. In this section (b), as in section (a) of Figure 45 described above, the readout unit is a line. In the example of section (b), in frame Fr(m), for each odd-numbered line and each even-numbered line, two lines are paired together, with the inter-line distance being half the number of lines in frame Fr(m). Of these pairs, the pairs with odd line numbers are read out sequentially, and then the pairs with even line numbers are read out sequentially.
[0259] 45, in the first half of the frame period, for example, among the lines L#x included in frame Fr(m), the read order of odd-numbered lines L#1, L#3, L#5, and L#7 is reversed, and lines L#1, L#5, L#3, L#7 are read in that order. Similarly, in the second half of the frame period, among the lines L#x included in frame Fr(m), the read order of even-numbered lines L#2, L#4, L#6, and L#8 is reversed, and lines L#2, L#6, L#4, L#8 are read in that order.
[0260] Such control of the read order of each line L#x can be realized by the read determination unit 123 sequentially setting the read position information according to the read order as shown in section (b) of FIG.
[0261] By determining the readout order in this way, it is possible to reduce the delay from when readout of frame Fr(m) starts until the pixel data at the bottom of frame Fr(m) is obtained, compared to the example in section (a). Furthermore, in the seventh modification of the second embodiment, it is possible to increase the recognition response speed for large objects within a frame, thereby enabling an increase in the frame rate.
[0262] Note that the read order of each line L#x described using section (b) of Figure 45 is just an example, and the read area can be set so as to facilitate recognition of an intended object. For example, the read determination unit 123 can set an area in the frame where recognition processing is to be preferentially performed based on external information provided from outside the imaging device 1, and determine read position information so as to preferentially perform readout of the read area for this area. Furthermore, the read determination unit 123 can also set an area in the frame where recognition processing is to be preferentially performed according to a scene captured in the frame.
[0263] Furthermore, in the seventh modification of the second embodiment, similarly to the second embodiment described above, if a valid recognition result is obtained during the reading of each line for a frame, the line reading and recognition process can be terminated, thereby enabling power saving in the recognition process and shortening the time required for the recognition process.
[0264] (First alternative example of the seventh modified example of the second embodiment) Next, a first alternative example of the seventh modified example of the second embodiment will be described. In the seventh modified example of the second embodiment described above, one line is used as a readout unit, but this is not limited to this example. In this first alternative example, two non-adjacent lines are used as a readout unit.
[0265] Fig. 46 is a schematic diagram for explaining a frame readout process according to a first alternative example of the seventh modified example of the second embodiment. In the example of Fig. 46, in the frame Fr(m) described in section (b) of Fig. 45, each line having an odd line number is paired with each line having an even line number, and the pairing is performed as a readout unit, with the distance between the lines being 1 / 2 the number of lines. More specifically, the pair of lines L#1 and L#5, the pair of lines L#3 and L#7, the pair of lines L#2 and L#6, and the pair of lines L#4 and L#8 are each used as a readout unit. Among these pairs, the pairs with odd line numbers are sequentially read out, followed by the pairs with even line numbers.
[0266] In this first alternative example, since the reading unit includes two lines, it is possible to further reduce the time required for recognition processing compared to the seventh modified example of the second embodiment described above.
[0267] (Second alternative example of the seventh modified example of the second embodiment) Next, a second example of the seventh modified example of the second embodiment will be described. This second example is an application of an example in which a readout area for reading out the readout unit is set so that the pixels in the frame can be read out more uniformly in a short time to the example in which the readout unit is an area of a predetermined size in the frame according to the third modified example of the second embodiment described with reference to Fig. 33.
[0268] Fig. 47 is a schematic diagram for explaining a frame readout process according to a second alternative example of the seventh modified example of the second embodiment. In the example of Fig. 47, the positions of the areas Ar#xy described with reference to Fig. 33 are discretely specified in the frame Fr(m), and the frame Fr(m) is readout. As an example, after reading and recognition processing of the area Ar#1-1 in the upper left corner of the frame Fr(m), reading and recognition processing of the area Ar#3-1, which includes the same line as the area Ar#1-1 in the frame Fr(m) and is located in the center of the frame Fr(m) in the line direction, is performed. Next, reading and recognition processing of the area Ar#1-3 in the upper left corner of the lower half of the frame Fr(m) is performed, and then reading and recognition processing of the area Ar#3-3, which includes the same line as the area Ar#1-3 in the frame Fr(m) and is located in the center of the frame Fr(m) in the line direction, is performed.
[0269] Similarly, the areas Ar#2-2 and Ar#4-2, and the areas Ar#2-4 and Ar#4-4 are subjected to the readout and recognition processes.
[0270] By determining the readout order in this way, it is possible to reduce the delay from when readout of frame Fr(m) starts from the left edge of frame Fr(m) until pixel data at the bottom and right edge of frame Fr(m) is obtained, compared to the example in Fig. 33. Furthermore, in this second other example, it is possible to speed up the recognition response speed for large objects within a frame, and increase the frame rate.
[0271] In this second example, similarly to the second embodiment, if a valid recognition result is obtained during the reading of each area Ar#xy for a frame, the reading of the area Ar#xy and the recognition process can be terminated, thereby enabling power saving in the recognition process and shortening the time required for the recognition process.
[0272] (Third alternative example of the seventh modified example of the second embodiment) Next, a third example of the seventh modified example of the second embodiment will be described. This third example is an application of an example in which a readout area for reading out the readout unit is set so that the pixels in the frame can be read out more uniformly in a short time to the example in which the readout unit is a pattern made up of a plurality of pixels that are discretely and periodically arranged in both the line direction and the vertical direction according to the third modified example of the second embodiment described with reference to Fig. 37.
[0273] Fig. 48 is a schematic diagram for explaining a frame readout process according to a third other example of the seventh modified example of the second embodiment. In the example of Fig. 48, the pattern Pφ#z has the same configuration as the pattern Pφ#xy explained using Fig. 37, and the position of the pattern Pφ#z is discretely specified in the frame Fr(m) to read out the frame Fr(m).
[0274] As an example, the recognition processing unit 12 starts from the upper left corner of frame Fr(m) and reads and recognizes pattern Pφ#1 located at the upper left corner. Next, it reads and recognizes pattern Pφ#2, which is shifted by half the spacing between pixels in pattern Pφ#1 in both the line and vertical directions. Next, it reads and recognizes pattern Pφ#3, which is shifted by half the spacing in the line direction from the position of pattern Pφ#1, and then it reads and recognizes pattern Pφ#4, which is shifted by half the spacing in the vertical direction from the position of pattern Pφ#1. The readout and recognition processing of these patterns Pφ#1 to Pφ#4 is repeatedly performed while shifting the position of pattern Pφ#1, for example, by one pixel in the line direction, and then again by one pixel in the vertical direction.
[0275] By determining the readout order in this way, it is possible to reduce the delay from when readout of frame Fr(m) starts from the left edge of frame Fr(m) until pixel data at the bottom and right edge of frame Fr(m) is obtained, compared to the example in Fig. 37. Furthermore, in this third alternative example, it is possible to speed up the recognition response speed for large objects within a frame, and increase the frame rate.
[0276] In this third example, similarly to the second embodiment, if a valid recognition result is obtained during the reading of each pattern Pφ#z for a frame, the reading of the pattern Pφ#z and the recognition process can be terminated, thereby enabling power saving in the recognition process and shortening the time required for the recognition process.
[0277] [5-8. Eighth Modification of the Second Embodiment] Next, an eighth modified example of the second embodiment will be described. The eighth modified example of the second embodiment is an example in which at least one of the exposure and the analog gain of the sensor unit 10 is controlled according to a predetermined pattern. In the following description, analog gain will be abbreviated as "AG" where appropriate.
[0278] Fig. 49 is a functional block diagram illustrating an example of functions according to an eighth modification of the second embodiment. In Fig. 49, the readout determination unit 123a generates readout region information and information indicating an exposure time and an analog gain based on readout information passed from the feature amount accumulation control unit 121. Without being limited to this, the readout determination unit 123a may generate at least one of the exposure time and the analog gain. The readout determination unit 123a passes the generated readout region information, exposure time, and analog gain to the readout unit 110.
[0279] Fig. 50 is a flowchart illustrating an example of recognition processing according to an eighth modified example of the second embodiment. The processing according to the flowchart in Fig. 50 corresponds to, for example, reading pixel data in a read unit (for example, one line) from a frame. Note that, in this description, the read unit is assumed to be a line, and the read area information is assumed to be a line number indicating the line to be read.
[0280] In the flowchart of FIG. 50, the processing from step S100 to step S107 is the same as the processing from step S100 to step S107 in the flowchart of FIG. 28 described in the second embodiment.
[0281] That is, in step S100, the recognition processing unit 12 reads line data from the line indicated by the read line of the frame. In the next step S101, the feature amount calculation unit 120 calculates line feature amounts based on the line data passed from the reading unit 110. In the next step S102, the feature amount calculation unit 120 acquires feature amounts accumulated in the feature amount accumulation unit 122 from the feature amount accumulation control unit 121. In the next step S103, the feature amount calculation unit 120 integrates the feature amount calculated in step S101 and the feature amount acquired from the feature amount accumulation control unit 121 in step S102, and passes the integrated result to the feature amount accumulation control unit 121.
[0282] In the next step S104, the feature accumulation control unit 121 accumulates the integrated feature in the feature accumulation unit 122 (step S104). In the next step S105, the recognition process execution unit 124 executes recognition processing using the integrated feature. In the next step S106, the recognition process execution unit 124 outputs the recognition result of the recognition process in step S105. In step S107, the read determination unit 123 in the recognition processing unit 12 determines the read line to be read next in accordance with the read information passed from the feature accumulation control unit 121.
[0283] After the process of step S107, the process proceeds to step S1070. In step S1070, the readout determination unit 123 determines the exposure time for the next readout according to a predetermined exposure pattern. In the next step S1071, the readout determination unit 123 determines the analog gain for the next readout according to a predetermined AG pattern.
[0284] The exposure time and analog gain determined in steps S1070 and S1071, respectively, are applied to the readout line determined in step S107, and the process from step S100 is executed again.
[0285] The order of steps S107 to S1071 is not limited to the above order. Also, one of the processes in steps S1070 and S1071 can be omitted.
[0286] (First example of reading and recognition process) Next, a first example of the readout and recognition process according to the eighth modified example of the second embodiment will be described with reference to FIGS. 51A and 51B. FIG. 51A is a schematic diagram showing an example of an exposure pattern applicable to the eighth modified example of the second embodiment. In FIG. 51A, the vertical axis represents lines, and the horizontal axis represents time. In the example of FIG. 51A, the exposure pattern alternates between a first exposure time and a second exposure time shorter than the first exposure time for each line L#1, L#2, .... Therefore, exposure is performed using the first exposure time for the odd-numbered lines L#1, L#3, ..., L#(2n-1), and exposure is performed using the second exposure time shorter than the first exposure time for the even-numbered lines L#2, L#4, ..., L#2n. Note that the transfer time of line data is constant regardless of the exposure time.
[0287] Fig. 51B is a diagram showing an example of a captured image when capturing an image according to the exposure pattern of Fig. 51A. In this example, the first exposure time is set so that the odd-numbered lines L#1, L#3, ..., L#(2n-1) are overexposed, and the even-numbered lines L#2, L#4, ..., L#2n are underexposed.
[0288] 22, the feature extraction process 1200, the integration process 1202, the internal state update process 1211, and the recognition process 1240 of the recognition processing unit 12 are each executed using parameters obtained by learning in advance based on an image in which light and dark change for each line, as shown in Fig. 51B. However, the recognition processing unit 12 may also learn using a general image.
[0289] In this way, by setting different exposures alternately for each line within one frame, the recognition processing unit 12 can recognize both bright objects and dark objects in the recognition processing for one frame.
[0290] 51A, the exposure time is alternately switched for each line according to a predetermined exposure pattern, but this is not limited to this example. That is, the analog gain may be alternately switched for each line according to a predetermined AG pattern, or the exposure time and analog gain may be switched according to the predetermined exposure pattern and AG pattern, respectively.
[0291] 51A, the readout unit is a line, and the first exposure time and the second exposure time are set alternately for each line, but this is not limited to this example. For example, it is also possible to apply the setting of the exposure time and analog gain for each readout unit according to the eighth modified example of the second embodiment to each readout unit according to the first to seventh modified examples of the second embodiment described above.
[0292] [5-9. Ninth Modification of the Second Embodiment] Next, a ninth modified example of the second embodiment will be described. The ninth modified example of the second embodiment is an example in which the length of the exposure time is set by spatially varying the density in a frame. Note that, in the ninth modified example of the second embodiment, the configuration described using Fig. 49 can be applied as is, so detailed description of the configuration will be omitted.
[0293] FIG. 52 is a diagram showing an example of an exposure pattern according to a ninth modification of the second embodiment. In sections (a) to (c) of FIG. 52, each square is, for example, an area including multiple pixels. In each of sections (a) to (c) of FIG. 52, the open squares are areas where no exposure or no readout is performed. Also, in FIG. 52, the exposure time is expressed in three stages by the shade of the squares, with squares with shorter exposure times being filled in more darkly. Hereinafter, the open squares will be considered to indicate areas where no exposure is performed. Also, the filled squares will be described as areas where exposure is performed for short, medium, or long exposure times, depending on the shade of the fill. Note that the medium time is longer than the short time and shorter than the long time. In each of sections (a) to (c) of FIG. 52, for example, a set of filled squares constitutes a readout unit.
[0294] Specifically, frame 210a shown in section (a) is an example (referred to as a first pattern) in which areas where the exposure time is set to a short time are densely located in the center and gradually become sparser toward the periphery. Frame 210b shown in section (b) is an example (referred to as a second pattern) in which areas where the exposure time is set to a medium time are arranged approximately uniformly throughout the frame 210b. Furthermore, frame 210c shown in section (c) is an example (referred to as a third pattern) in which areas where the exposure time is set to a long time are densely located in the periphery and gradually become sparser toward the center.
[0295] In the recognition processing unit 12, the readout determination unit 123 selects a pattern from the above-mentioned first to third patterns according to, for example, readout information passed from the feature accumulation control unit 121 or recognition information (not shown) passed from the recognition processing execution unit 124. However, the readout determination unit 123 may select one pattern from the first to third patterns according to control information from the outside, such as a user operation.
[0296] The readout determination unit 123 generates readout area information including the exposure time of the selected pattern and the position of each area, and passes the generated readout area information to the sensor control unit 11. In the sensor control unit 11, the readout unit 110 reads pixel data from the sensor unit 10 in accordance with the passed readout area information. Note that it is preferable that the readout using the above-mentioned first to third patterns is applied to a global shutter system.
[0297] Now, consider a case where the imaging device 1 according to the ninth modification of the second embodiment is installed in a vehicle so as to capture an image of the front. In this case, headlights will cause spatial differences in brightness in the captured image at night. Specifically, it is considered that the brightness value of the captured image will be high in the center and decrease toward the periphery. Therefore, when the headlights are turned on in a dark place, the first pattern (section (a) in FIG. 52) will be selected.
[0298] Furthermore, areas with different exposure times can be arranged in a pattern different from the above-mentioned patterns 1 to 3. For example, consider a fourth pattern in which areas with long exposure times are densely located in the center and gradually become sparser toward the periphery, and a fifth pattern in which areas with short exposure times are densely located in the periphery and gradually become sparser toward the center.
[0299] In the above-described in-vehicle example, when the vehicle equipped with the imaging device 1 moves forward, blurring occurs in the captured image due to the movement of the vehicle. This blurring is small in the center of the image and increases toward the periphery. Therefore, if you want to perform recognition processing on the center of the image, you can select the fourth pattern, and if you want to perform recognition processing on the periphery of the image, you can select the fifth pattern.
[0300] In this way, by setting the length of the exposure time to vary spatially in density within a frame, the recognition processing unit 12 can recognize bright objects and dark objects in the recognition processing for one frame. Also, because the density of the exposed area is changed depending on the position within the frame, it is possible to detect the spatial appearance frequency of bright objects and dark objects.
[0301] [5-10. Tenth Modification of the Second Embodiment] Next, a tenth modification of the second embodiment will be described. The tenth modification of the second embodiment is an example in which a reading area from which the next reading unit is to be read is determined based on a feature generated in the feature accumulation control unit 121.
[0302] FIG. 53 is a functional block diagram illustrating an example of functions according to a tenth modification of the second embodiment. In section (a) of FIG. 53, the feature accumulation control unit 121a integrates the feature transferred from the feature calculation unit 120 with the feature accumulated in the feature accumulation unit 122, and transfers the integrated feature together with read information to the read determination unit 123b. The read determination unit 123b generates read area information and information indicating an exposure time and an analog gain based on the feature transferred from the feature accumulation control unit 121a and the read information. Alternatively, the read determination unit 123b may generate at least one of the exposure time and the analog gain. The read determination unit 123a / 123b transfers the generated read area information, exposure time, and analog gain to the read unit 110.
[0303] Section (b) of Fig. 53 is an example functional block diagram for explaining in more detail the function of the read determination unit 123b according to the tenth modification of the second embodiment. In section (b) of Fig. 53, the read determination unit 123b includes a read area determination unit 1230, an exposure time determination unit 1231, and an AG amount determination unit 1232.
[0304] The readout information and feature amounts passed from the feature amount accumulation control unit 121 to the readout determination unit 123b are input to the readout area determination unit 1230, exposure time determination unit 1231, and AG amount determination unit 1232, respectively. Based on the input readout information and feature amounts, the readout area determination unit 1230 generates and outputs readout area information (e.g., a line number) indicating the readout area to be read out next. Based on the input feature amounts, the exposure time determination unit 1231 generates and outputs information indicating the exposure time for the next imaging. Furthermore, the AG amount determination unit 1232 generates and outputs information indicating the analog gain for the next imaging based on the input feature amounts.
[0305] Fig. 54 is a flowchart illustrating an example of processing according to a tenth modification of the second embodiment. The processing according to the flowchart in Fig. 54 corresponds to, for example, reading pixel data in a read unit (for example, one line) from a frame. Note that, in this description, the read unit is assumed to be a line, and the read area information is assumed to be a line number indicating the line to be read.
[0306] In the flowchart of FIG. 54, the processing from step S100 to step S106 is equivalent to the processing from step S100 to step S107 in the flowchart of FIG. 28 described in the second embodiment.
[0307] That is, in step S100, the recognition processing unit 12 reads line data from the line indicated by the read line of the frame. In the next step S101, the feature amount calculation unit 120 calculates line feature amounts based on the line data passed from the reading unit 110. In the next step S102, the feature amount calculation unit 120 acquires feature amounts accumulated in the feature amount accumulation unit 122 from the feature amount accumulation control unit 121. In the next step S103, the feature amount calculation unit 120 integrates the feature amount calculated in step S101 and the feature amount acquired from the feature amount accumulation control unit 121 in step S102, and passes the integrated result to the feature amount accumulation control unit 121.
[0308] In the next step S104, the feature amount accumulation control unit 121 accumulates the integrated feature amounts in the feature amount accumulation unit 122. In the next step S105, the recognition process execution unit 124 executes recognition process using the integrated feature amounts. In the next step S106, the recognition process execution unit 124 outputs the recognition result of the recognition process in step S105.
[0309] In the next step S1080, the readout region determination unit 1230 in the readout determination unit 123b determines the readout line to be read out next, using the readout information passed from the feature amount accumulation control unit 121 and a feature amount obtained by integrating the feature amount calculated in step S101 and the feature amount acquired from the feature amount accumulation control unit 121 in step S102. In the next step S1081, the exposure time determination unit 1231 in the readout determination unit 123b determines the exposure time for the next readout, based on the integrated feature amount passed from the feature amount accumulation control unit 121. In the next step S1082, the AG amount determination unit 1232 in the readout determination unit 123b determines the analog gain for the next readout, based on the integrated feature amount passed from the feature amount accumulation control unit 121.
[0310] The exposure time and analog gain determined in steps S1081 and S1082, respectively, are applied to the readout line determined in step S1080, and the process from step S100 is executed again.
[0311] The order of steps S1080 to S1082 is not limited to the above order. Also, one of the processes in steps S1081 and S1082 can be omitted.
[0312] (First process) First, a first process according to the tenth modified example of the second embodiment will be described. Fig. 55 is a schematic diagram for explaining the first process according to the tenth modified example of the second embodiment. The process shown in Fig. 55 corresponds to the processes of steps S1 to S4c in Fig. 16 described above. Here, the readout unit is a line, and imaging is performed by the rolling shutter method.
[0313] 55, in step S1, the imaging device 1 (see FIG. 1) starts capturing an image of a target image (handwritten number "8") to be recognized. Here, a learning model that has been trained to be able to identify numbers using predetermined training data is stored in advance in the memory 13 as a program, and the recognition processing unit 12 is capable of identifying numbers included in the image by reading and executing this program from the memory 13.
[0314] When imaging starts, in step S2, the sensor control unit 11 sequentially reads out the frame line by line from the top end to the bottom end of the frame in accordance with the readout area information passed from the recognition processing unit 12.
[0315] When the lines are read out up to a certain position, the recognition processing unit 12 identifies the number "8" or "9" from the image formed by the read lines (step S3). In the recognition processing unit 12, the read determination unit 123b generates read area information specifying a line L#m that is predicted to enable identification of the object identified in step S3 as either the number "8" or "9" based on the integrated feature amount passed from the feature amount accumulation control unit 121, and passes this information to the reading unit 110. Then, the recognition processing unit 12 executes recognition processing based on the pixel data of the line L#m read out by the reading unit 110 (step S4c).
[0316] Here, if the object is confirmed in step S4c described above, the recognition processing unit 12 can end the recognition processing, thereby realizing a reduction in the time required for the recognition processing and power saving.
[0317] (Second process) Next, a second process according to a tenth modified example of the second embodiment will be described. Fig. 56 is a schematic diagram for explaining the second process according to the tenth modified example of the second embodiment. The process shown in Fig. 56 corresponds to the process shown in Fig. 55 described above. In this second process, reading of a frame line by line is performed while thinning out the lines.
[0318] 56, in step S10, the imaging device 1 starts capturing an image of a target image to be recognized (handwritten number "8"). As described above, the memory 13 is previously stored with a learning model as a program that has been trained to be able to identify numbers when read out line by line using predetermined training data, and the recognition processing unit 12 is capable of identifying numbers included in the image by reading and executing this program from the memory 13.
[0319] When imaging begins, in step S11, the sensor control unit 11 reads out the frame line by line while thinning out the lines from the top to the bottom of the frame in accordance with the readout area information passed from the recognition processing unit 12. In the example of Fig. 56, the sensor control unit 11 first reads out line L#1 at the top of the frame in accordance with the readout area information, and then reads out line L#p with a predetermined number of lines thinned out. The recognition processing unit 12 performs recognition processing on each line data of lines L#1 and L#p each time it is read out.
[0320] Here, it is assumed that reading is further performed line by line by thinning, and that the recognition processing unit 12 performs a recognition process on the line data read from line L#q, resulting in the recognition of the number "8" or "0" (step S12). Based on the integrated features passed from the feature accumulation control unit 121, the read determination unit 123b generates read area information specifying line L#r that is predicted to enable identification of the object identified in step S12 as either the number "8" or "0", and passes this information to the read unit 110. At this time, the position of line L#r may be on the upper or lower end side of the frame with respect to line L#q.
[0321] The recognition processing unit 12 executes recognition processing based on the pixel data read out from the line L#r by the reading unit 110 (step S13).
[0322] In this second process, the lines of the frame are read out while thinning out the lines, which makes it possible to further reduce the time required for the recognition process and to further reduce power consumption.
[0323] (Third Processing) Next, a third process according to a tenth modified example of the second embodiment will be described. The third process according to the tenth modified example of the second embodiment is an example in which the exposure time and analog gain are adaptively set while reading out in units of readout.
[0324] 57A and 57B, 58A and 58B, and 59A and 59B are schematic diagrams for explaining a third process according to a tenth modification of the second embodiment. In the following, a case where the exposure time is adaptively set among the exposure time and the analog gain will be described.
[0325] Figures 57A and 57B correspond to Figures 51A and 51B described above and show an example in which the exposure time is not adaptively set. The exposure pattern shown in Figure 57A has a larger difference between the first exposure time and the second exposure time, which is shorter than the first exposure time, compared to the exposure pattern shown in Figure 51A described above. The exposure pattern, like the example in Figure 51A, is a pattern in which the first exposure time and the second exposure time are applied alternately for each line L#1, L#2, ...
[0326] Figure 57B is a diagram showing an example of a captured image 200a when capturing an image according to the exposure pattern of Figure 57A. In this example, the first exposure time is set so that odd-numbered lines L#1, L#3, ..., L#(2n-1) are overexposed, and the second exposure time is set so that even-numbered lines L#2, L#4, ..., L#2n are underexposed.
[0327] Figures 58A and 58B show an example in which imaging is performed according to the exposure pattern of Figure 57A, and the exposure time for each line L#2, L#3, ..., L#2n is set based on the integrated features corresponding to the pixel data read out from each line L#1, L#2, ..., L#(2n-1).
[0328] When reading from line L#1 is performed, the read determination unit 123b sets the exposure time (hereinafter referred to as exposure time L#3) of the next odd-numbered line L#3 based on the feature amount calculated by the feature amount calculation unit 120 based on the pixel data read from this line L#1 and the feature amount accumulated in the feature amount accumulation unit 122, which are integrated by the feature amount accumulation control unit 121. In the example of FIG. 58A, the exposure time (L#3) is set shorter than the exposure time of line L#1. The read determination unit 123b holds this exposure time (L#3). Note that the process of setting the exposure time by the read determination unit 123b is executed, for example, during the transfer period of the next line (in this case, line L#2).
[0329] Next, reading is performed from line L#2, which is an even-numbered line number. The read determination unit 123b sets and holds the exposure time (L#4) for the next line L#4, which is an even-numbered line number, based on the feature amount calculated by the feature amount calculation unit 120 based on the pixel data read from line L#2 and the feature amount accumulated in the feature amount accumulation unit 122, which are integrated by the feature amount accumulation control unit 121. In the example of Figure 58A, the exposure time (L#4) is set shorter than the exposure time for line L#2.
[0330] When the reading and recognition process of line L#2 is completed, the reading determination unit 123b passes to the reading unit 110 the exposure time (L#3) that it has stored and the reading region information for instructing the reading of the next line L#3.
[0331] Similarly, for reading out line L#3, the read determination unit 123b sets the exposure time (L#5) of the next odd-numbered line L#5 based on the pixel data read out from line L#3. In the example of FIG. 58A, the exposure time (L#5) is set shorter than the exposure time of line L#3. The read determination unit 123b holds this exposure time (L#5). Next, reading out is performed from line L#4, an even-numbered line, and the read determination unit 123b sets the exposure time (L#6) of the next even-numbered line L#6 based on the pixel data read out from line L#4. In the example of FIG. 58A, the exposure time (L#6) is set shorter than the exposure time of line L#4. The read determination unit 123b holds this exposure time (L#6).
[0332] In this way, line readout and exposure setting are repeated alternately for odd-numbered lines and even-numbered lines, so that, as shown in Figure 58B, the exposure times are set appropriately for the odd-numbered lines L#1, L#3, ..., L#(2n-1) that are set to be overexposed, and the even-numbered lines L#2, L#4, ..., L#2n that are set to be underexposed.
[0333] Here, the read determination unit 123b sets the exposure time using the feature amount obtained based on the pixel information read from the line, so that the read determination unit 123b can set an exposure time suitable for the recognition process execution unit 124 to execute the recognition process.
[0334] Comparing Figure 57B and Figure 58B, in the example of Figure 58B, the brightness is reduced in the odd-numbered lines L#1, L#3, ..., which are overexposed, and the brightness is increased in the even-numbered lines L#2, L#4, ..., which are underexposed. Therefore, in the example of Figure 58B, it is expected that the recognition process will be performed with higher accuracy than in the example of Figure 57B for both the odd-numbered lines and the even-numbered lines.
[0335] In addition, the exposure times set for the odd-numbered lines and even-numbered lines at the bottom of the frame can be used to set the exposure times for the odd-numbered lines and even-numbered lines in the next frame.
[0336] 58A and 58B, all lines of the frame are read out, but this is not limited to this example. For example, it is possible to read out lines up to the middle of the frame and then return the read position depending on the results.
[0337] This will be explained using Figures 59A and 59B. Figures 59A and 59B show an example in which, like the example in Figures 58A and 58B described above, imaging is performed according to the exposure pattern in Figure 57A, and the exposure times for each of lines L#2, L#3, ..., L#2n are set based on the integrated feature amounts corresponding to the pixel data read out from each of lines L#1, L#2, ..., L#(2n-1).
[0338] In the example of FIG. 59A, it is assumed that a predetermined object (a person) is recognized by the recognition process performed by the recognition process execution unit 124 based on pixel data read from, for example, line L#2. For example, the recognition process execution unit 124 passes recognition information indicating that a person has been recognized to the read determination unit 123b. In accordance with this recognition information, the read determination unit 123b sets the exposure time of line L#2' corresponding to line L#2 in the next frame. In the example of FIG. 59A, the read determination unit 123b sets the exposure time of line L#2' to be longer than the exposure time of line L#2.
[0339] After setting the exposure time for line L#2', the read determination unit 123b passes read area information indicating line L#2' and the exposure time set for line L#2' to the read unit 110. In accordance with this read area information and exposure time, the read unit 110 starts exposing line L#2', for example, after the transfer process for line L#3 is completed.
[0340] For lines L#3 and onward, the read determination unit 123b sets the exposure time for each of the odd-numbered lines L#3, L#5, ... and the even-numbered lines L#4, L#6 in the same manner as in Figure 58A. Here, in the example of Figure 59A, immediately after the transfer process for line L#7 is completed, the exposure of line L#2' is completed and the transfer of pixel data read from line L#2' is started.
[0341] As shown in FIG. 59B, the line L#2' has a higher luminance than the line L#2 shown in FIG. 58B, and it is expected that recognition with higher accuracy will be possible.
[0342] In addition, the read determination unit 123b sets exposure times for each of the lines L#3', L#4', ... subsequent to line L#2' that are included in the same frame as line L#2', and reads them out in the same manner as the processing described using Figures 58A and 58B.
[0343] The exposure time for line L#2' is set in accordance with the recognition information based on the pixel data read from line L#2 of the previous frame. This allows the recognition processing execution unit 124 to perform the recognition processing based on line L#2' with higher accuracy than the recognition processing based on line L#2.
[0344] In the third process of the tenth variant of the second embodiment, the exposure time of the line is set based on the readout result of the line, but this is not limited to this example, and the analog gain of the line may also be set.
[0345] Increasing the exposure time allows for capturing images with higher brightness values and reduces noise. However, increasing the exposure time may increase blur, for example, in dynamic scenes. On the other hand, adjusting the analog gain does not change the blur, but increasing the analog gain may increase noise.
[0346] Therefore, it is preferable to select whether to set the exposure time or the analog gain depending on the purpose of the recognition process and the subject of the image capture. For example, if the subject of the image capture is a dynamic scene, the analog gain can be increased to suppress blur, and if the subject of the image capture is a static scene, the exposure time can be increased to capture a brighter image and suppress noise.
[0347] Fig. 60 is a schematic diagram showing in more detail an example of processing in the recognition processing unit 12 according to the tenth modified example of the second embodiment. In the configuration shown in Fig. 60, a read determination unit 123b is added to the configuration shown in Fig. 22. The read determination unit 123b receives an input of a feature 1212 related to the internal state updated by an internal state update process 1211.
[0348] The processing in the read determination unit 123b includes next line determination processing 1233, next exposure determination processing 1234, and next AG determination processing 1235. These next line determination processing 1233, next exposure determination processing 1234, and next AG determination processing 1235 are processing executed in the read area determination unit 1230, exposure time determination unit 1231, and AG amount determination unit 1232, respectively, which were described using section (b) of Fig. 53 .
[0349] The next line determination process 1233, the next exposure determination process 1234, and the next AG determination process 1235 are each executed based on pre-trained parameters. The parameters are learned using training data based on, for example, an expected readout pattern or a recognition target. In the above example, the expected readout pattern can be, for example, a pattern in which overexposed lines and underexposed lines are alternately read out.
[0350] 57A and 57B, 58A and 58B, and 59A and 59B, the reading unit is described as a line, but this is not limited to this example. That is, other reading units, such as multiple lines adjacent to each other, may be applied to the tenth modification of the second embodiment.
[0351] In the imaging device 1 according to the second embodiment and its modifications described above, the recognition processing unit 12 performs recognition processing for each readout unit, but this is not limited to this example. For example, it may be possible to switch between recognition processing for each readout unit and normal recognition processing, i.e., recognition processing based on pixel data of pixels read from the entire frame after reading out the entire frame. In other words, normal recognition processing is performed based on pixels of the entire frame, and therefore it is possible to obtain recognition results with higher accuracy than recognition processing for each readout unit.
[0352] For example, normal recognition processing can be performed at regular intervals for recognition processing per read unit. Furthermore, normal recognition processing may be performed in an emergency for recognition processing per read unit to improve the stability of recognition. Furthermore, when switching from recognition processing per read unit to normal recognition processing, the normal recognition processing is less rapid than recognition processing per read unit, so the clock for operating the device in normal recognition processing may be switched to a faster clock. Furthermore, when the reliability of recognition processing per read unit is low, the device may be switched to normal recognition processing, and when the reliability of the recognition processing increases, the device may be switched back to recognition processing per read unit.
[0353] 6. Third Embodiment Next, a third embodiment of the present disclosure will be described. The third embodiment is an example in which parameters for frame readout are adaptively set. Possible parameters for frame readout include a readout unit, a readout order within a frame based on the readout unit, a readout area, an exposure time, and an analog gain. Note that the imaging device 1 according to the third embodiment can apply the functions described with reference to FIG. 21 to its overall function, and therefore a description of the overall configuration will be omitted.
[0354] Fig. 61 is a functional block diagram illustrating an example of functions according to the third embodiment. In the third embodiment, the recognition processing by the recognition processing unit 12 is the main focus, and therefore the configuration shown in section (a) of Fig. 61 omits the visual recognition processing unit 14, the output control unit 15, the trigger generation unit 16, and the readout control unit 111 in the sensor control unit 11 from the configuration of Fig. 21 described above.
[0355] On the other hand, in the configuration shown in section (a) of Fig. 61, an external information acquisition unit 17 is added to the configuration of Fig. 21, and pixel data is passed to a read determination unit 123c from a read unit 110. In addition, recognition information is passed to the read determination unit 123c from a recognition processing execution unit 124.
[0356] Furthermore, the external information acquisition unit 17 acquires external information generated outside the imaging device 1, and passes the acquired external information to the readout determination unit 123c.
[0357] Specifically, the external information acquisition unit 17 can use an interface that transmits and receives signals in a predetermined format. For example, if the imaging device 1 is mounted on a vehicle, vehicle information can be applied as the external information. The vehicle information is information acquired from a vehicle system, such as steering information and speed information. The environmental information is information indicating the environment around the imaging device 1, such as ambient brightness. In the following, unless otherwise specified, the imaging device 1 is assumed to be mounted on a vehicle, and the external information is assumed to be vehicle information acquired from the vehicle in which the imaging device 1 is mounted.
[0358] Section (b) of Figure 61 is an example functional block diagram for explaining in more detail the function of the read determination unit 123c according to the third embodiment. In section (b) of Figure 61, the read determination unit 123c includes a read unit pattern selection unit 300, a read order pattern selection unit 302, and a read determination processing unit 304. The read unit pattern selection unit 300 includes a read unit pattern DB (database) 301 in which a plurality of different read patterns are stored in advance. In addition, the read order pattern selection unit 302 includes a read order pattern DB 303 in which a plurality of different read order patterns are stored in advance.
[0359] The read determination unit 123c sets priorities for each read unit pattern stored in the read unit pattern DB301 and each read order pattern stored in the read order pattern DB303 based on one or more of the recognition information, pixel data, vehicle information, and environmental information passed to it.
[0360] The read unit pattern selection unit 300 selects the read unit pattern set with the highest priority from among the read unit patterns stored in the read unit pattern DB 301. The read unit pattern selection unit 300 passes the read unit pattern selected from the read unit pattern DB 301 to the read determination processing unit 304. Similarly, the read order pattern selection unit 302 selects the read order pattern set with the highest priority from among the read order patterns stored in the read order pattern DB 303. The read order pattern selection unit 302 passes the read order pattern selected from the read order pattern DB 303 to the read determination processing unit 304.
[0361] The read determination processing unit 304 determines the read area to be read next from the frame in accordance with the read information passed from the feature accumulation control unit 121, the read unit pattern passed from the read unit pattern selection unit 300, and the read order pattern passed from the read order pattern selection unit 302, and passes read area information indicating the determined read area to the read unit 110. In addition, the read determination processing unit 304 generates information indicating the exposure time and analog gain when the next read is to be performed from the frame, based on one or more of the pixel data, recognition information, vehicle information, and environmental information passed to the read determination unit 123c, and passes each of the generated information to the read unit 110.
[0362] (6-0-1. How to set the read unit pattern and read order pattern) An example of a method for setting a read unit pattern and a read order pattern according to the third embodiment using the configuration shown in sections (a) and (b) of FIG. 61 will be described.
[0363] (6-0-1-1. Example of read unit pattern and read order pattern) Here, examples of read unit patterns and read order patterns applicable to the third embodiment will be described. Figure 62 is a schematic diagram showing examples of read unit patterns applicable to the third embodiment. In the example of Figure 62, five read unit patterns, 410, 411, 412, 413, and 414, are shown.
[0364] The read unit pattern 410 corresponds to the above-mentioned Fig. 24, and is a read unit pattern in which a line is used as a read unit and reading is performed line by line in the frame 400. The read unit pattern 411 corresponds to the above-mentioned Fig. 33, and is a read pattern in which an area of a predetermined size is used as a read unit in the frame 400 and reading is performed area by area in the frame 400.
[0365] The readout unit pattern 412 corresponds to the above-mentioned Fig. 37 and is a readout unit pattern in which a pixel set made up of a plurality of periodically arranged pixels, which include non-adjacent pixels, is used as a readout unit, and readout is performed for each of these plurality of pixels in the frame 400. The readout unit pattern 413 corresponds to the above-mentioned Fig. 42 and is a readout unit pattern in which a plurality of discretely and non-periodically arranged pixels (random pattern) is used as a readout unit, and readout is performed while updating the random pattern in the frame 400. These readout unit patterns 412 and 413 enable more uniform sampling of pixels from the frame 400.
[0366] The read unit pattern 414 corresponds to the above-mentioned FIG. 43, and is a read unit pattern that is adaptively generated based on the recognition information.
[0367] The read unit patterns 410 to 414 described with reference to FIG. 62 are stored in advance in the read unit pattern DB 301 shown in section (b) of FIG.
[0368] Note that the read unit applicable as the read unit pattern according to the third embodiment is not limited to the example shown in Fig. 62. For example, each read unit described in the second embodiment and each modified example thereof can be applied as the read unit pattern according to the third embodiment.
[0369] Figure 63 is a schematic diagram showing an example of a readout order pattern applicable to the third embodiment. Section (a) of Figure 63 shows an example of a readout order pattern when the readout unit is a line. Section (b) shows an example of a readout order pattern when the readout unit is an area. Section (c) shows an example of a readout order pattern when the readout unit is the above-mentioned pixel set.
[0370] In sections (a), (b), and (c) of FIG. 63, the readout order patterns 420a, 430a, and 440a on the left side respectively show examples of readout order patterns in which readout is performed sequentially in line order or pixel order.
[0371] Here, the readout order pattern 420a of section (a) is an example of line-sequential readout from the top to the bottom of the frame 400. The readout order patterns 430a and 440a of sections (b) and (c) are examples of sequential readout for each area or pixel set along the line direction from the upper left corner of the frame 400, respectively, and repeating this line-direction readout in the vertical direction of the frame 400. These readout order patterns 420a, 430a, and 440a are called forward readout order patterns.
[0372] On the other hand, readout order pattern 420b in section (a) is an example in which readout is performed line-sequentially from the bottom edge toward the top edge of frame 400. Readout order patterns 430b and 440b in sections (b) and (c) are examples in which readout is performed area by area or pixel set by pixel set along the line direction from the bottom right corner of frame 400, and this line-direction readout is repeated in the vertical direction of frame 400. These readout order patterns 420b, 430b, and 440b are called reverse readout order patterns.
[0373] Furthermore, readout order pattern 420c in section (a) is an example in which reading is performed from the top to the bottom of frame 400 while thinning out lines. Readout order patterns 430c and 440c in sections (b) and (c) are examples in which areas within frame 400 are read out at discrete positions and in discrete readout orders, respectively. In readout order pattern 430c, for example, when a readout unit is made up of four pixels, each pixel is read out in the order indicated by the arrows in the figure. In readout order pattern 440c, as shown in area 442 in the figure, for example, each pixel is read out while the pixel serving as the reference for the pattern is moved to a discrete position in an order different from the order of pixel positions in the line and column directions.
[0374] The reading order patterns 420a to 420c, the reading order patterns 430a to 430c, and the reading order patterns 420a to 420c described with reference to FIG. 63 are stored in advance in the reading order pattern DB 303 shown in section (b) of FIG.
[0375] (6-0-1-2. Example of how to set the read unit pattern) An example of a method for setting a read unit pattern according to the third embodiment will be described more specifically with reference to the above-mentioned FIGS.
[0376] First, a method for setting a read unit pattern based on image information (pixel data) will be described. The read determination unit 123c detects noise contained in the pixel data passed from the read unit 110. Here, a group of pixels arranged closely together has higher resistance to noise than a group of individual pixels arranged scatteredly. Therefore, when the pixel data passed from the read unit 110 contains a predetermined level of noise or more, the read determination unit 123c sets the priority of the read unit pattern 410 or 411, of the read unit patterns 410 to 414 stored in the read unit pattern DB 301, to be higher than the priority of the other read unit patterns.
[0377] Next, a method for setting a readout unit pattern based on the recognition information will be described. The first setting method is an example in which many objects larger than a predetermined size are recognized in the frame 400 based on the recognition information passed from the recognition process execution unit 124. In this case, the read determination unit 123c sets the priority of the readout unit pattern 412 or 413, among the readout unit patterns 410 to 414 stored in the readout unit pattern DB 301, to be higher than the priority of the other readout unit patterns. This is because uniform sampling of the entire frame 400 makes it possible to further improve the speed of reporting.
[0378] The second setting method is an example in which flicker is detected in an image captured based on pixel data, for example. In this case, the read determination unit 123c sets the priority of the read unit pattern 413 among the read unit patterns 410 to 414 stored in the read unit pattern DB 301 to be higher than the priorities of the other read unit patterns. This is because, for flicker, artifacts caused by flicker can be suppressed by sampling the entire frame 400 using a random pattern.
[0379] The third setting method is an example of a case where a configuration that is considered to be capable of executing the recognition process more efficiently is generated when adaptively changing the configuration of the reading unit based on the recognition information. In this case, the reading determination unit 123c sets the priority of the reading unit pattern 414 among the reading unit patterns 410 to 414 stored in the reading unit pattern DB301 to be higher than the priority of the other reading unit patterns.
[0380] Next, a description will be given of a method for setting a reading unit pattern based on external information acquired by the external information acquisition unit 17. The first setting method is an example in which the vehicle on which the imaging device 1 is mounted turns to the left or right based on the external information. In this case, the reading determination unit 123c sets the priority of the reading unit pattern 410 or 411, among the reading unit patterns 410 to 414 stored in the reading unit pattern DB301, higher than the priority of the other reading unit patterns.
[0381] Here, in this first setting method, the read determination unit 123c sets the read unit pattern 410 to a column direction among the row and column directions in the pixel array unit 101 as a read unit, and sets the readout to be performed column by column in the line direction of the frame 400. Also, the readout unit pattern 411 sets the area to be read out along the column direction, and this is repeated in the line direction.
[0382] When the vehicle turns left, the read determination unit 123c sets the read determination processing unit 304 to start column-sequential reading or reading along the column direction of the area from the left end side of the frame 400. On the other hand, when the vehicle turns right, the read determination processing unit 123c sets the read determination processing unit 304 to start column-sequential reading or reading along the column direction of the area from the right end side of the frame 400.
[0383] When the vehicle is moving straight, the read determination unit 123c reads the data in the normal line order or along the line direction of the area, and when the vehicle turns left or right, it can initialize the features stored in the feature storage unit 122, for example, and resume the read process by reading the features in the above-mentioned column order or by reading the area along the column direction.
[0384] The second method of setting the read unit pattern based on external information is an example of a case where the vehicle equipped with the imaging device 1 is traveling on a highway based on external information. In this case, the read determination unit 123c sets the priority of the read unit pattern 410 or 411, among the read unit patterns 410 to 414 stored in the read unit pattern DB 301, to be higher than the priority of the other read unit patterns. When traveling on a highway, it is considered important to recognize objects that are distant small objects. Therefore, by sequentially reading frames 400 from the top end, it is possible to further improve the speed of detecting objects that are distant small objects.
[0385] (6-0-1-3. Example of how to set the reading order pattern) An example of a method for setting a reading order pattern according to the third embodiment will be described more specifically with reference to the above-mentioned FIGS.
[0386] First, a method for setting a readout order pattern based on image information (pixel data) will be described. The readout determination unit 123c detects noise contained in the pixel data transferred from the readout unit 110. Here, the smaller the change in the area to be recognized, the less the recognition process is affected by noise, making the recognition process easier. Therefore, when the pixel data transferred from the readout unit 110 contains a predetermined amount of noise or more, the readout determination unit 123c sets the priority of one of the readout order patterns 420a, 430a, and 440a among the readout order patterns 420a to 420c, 430a to 430c, and 440a to 440c stored in the readout order pattern DB 303 to be higher than the priority of the other readout order patterns. However, the priority of one of the readout order patterns 420b, 430b, and 440b may be set to be higher than the priority of the other readout order patterns.
[0387] The priority of which of the read order patterns 420a, 430a and 440a, and the read order patterns 420b, 430b and 440b is set higher can be determined based on, for example, which of the read unit patterns 410 to 414 is set higher in the read unit pattern selection section 300, and whether reading is performed from the top or bottom end of the frame 400.
[0388] Next, a method for setting a reading order pattern based on the recognition information will be described. When a large number of objects larger than a predetermined size are recognized in the frame 400 based on the recognition information passed from the recognition processing execution unit 124, the reading determination unit 123c sets the priority of one of the reading order patterns 420c, 430c, and 440c among the reading order patterns 420a to 420c, 430a to 430c, and 440a to 440c stored in the reading order pattern DB 303 to be higher than the priority of the other reading order patterns. This is because uniform sampling can improve the speed of reporting rather than sequentially reading out the entire frame 400.
[0389] Next, a method for setting a reading order pattern based on external information will be described. The first setting method is an example in which the vehicle on which the imaging device 1 is mounted turns to the left or right direction based on external information. In this case, the reading determination unit 123c sets the priority of one of the reading order patterns 420a, 430a and 440a among the reading order patterns 420a to 420c, 430a to 430c and 440a to 440c stored in the reading order pattern DB303 higher than the priority of the other reading unit patterns.
[0390] Here, in this first setting method, the readout determination unit 123c sets the readout order pattern 420a so that the column direction of the row and column directions in the pixel array unit 101 is the readout unit, and readout is performed column by column in the line direction of the frame 400. Also, the readout order pattern 430a is set so that an area is read out along the column direction, and this is repeated in the line direction. Furthermore, the readout order pattern 440a is set so that a pixel set is read out along the column direction, and this is repeated in the line direction.
[0391] When the vehicle turns left, the read determination unit 123c sets the read determination processing unit 304 to start column-sequential reading or reading along the column direction of the area from the left end side of the frame 400. On the other hand, when the vehicle turns right, the read determination processing unit 123c sets the read determination processing unit 304 to start column-sequential reading or reading along the column direction of the area from the right end side of the frame 400.
[0392] When the vehicle is moving straight, the read determination unit 123c reads the data in the normal line order or along the line direction of the area, and when the vehicle turns left or right, it can initialize the features stored in the feature storage unit 122, for example, and resume the read process by reading the features in the above-mentioned column order or by reading the area along the column direction.
[0393] The second setting method of the reading order pattern based on external information is an example in which the vehicle equipped with the imaging device 1 is traveling on a highway based on external information. In this case, the reading determination unit 123c sets the priority of the reading order patterns 420a, 430a, and 440a among the reading order patterns 420a to 420c, 430a to 430c, and 440a to 440c stored in the reading order pattern DB 303 to be higher than the priority of the other reading order patterns. When traveling on a highway, it is considered important to recognize objects by small objects in the distance. Therefore, by sequentially reading frames 400 from the top, it is possible to improve the speed of detecting objects by small objects in the distance.
[0394] Here, as described above, when the priority of the reading unit pattern or the reading order pattern is set based on a plurality of different information (image information, recognition information, external information), there is a possibility that different reading unit patterns or different reading order patterns may collide with each other. In order to avoid this collision, for example, it is conceivable to make each priority set based on each information different in advance.
[0395] [6-1. First Modification of the Third Embodiment] Next, a first modified example of the third embodiment will be described. The first modified example of the third embodiment is an example in which a readout area is adaptively set when frame readout is performed.
[0396] Fig. 64 is a functional block diagram of an example for explaining functions according to a first modified example of the third embodiment. In the first modified example of the third embodiment, the recognition processing by the recognition processing unit 12 is the main focus, and therefore the configuration shown in section (a) of Fig. 64 omits the visual recognition processing unit 14, the output control unit 15, the trigger generation unit 16, and the read control unit 111 in the sensor control unit 11 from the configuration of Fig. 21 described above. Also, the read determination unit 123d shown in sections (a) and (b) of Fig. 64 differs in function from the read determination unit 123c according to the third embodiment shown in Fig. 61.
[0397] In section (b) of Figure 64, the read determination unit 123b has a configuration corresponding to the read determination unit 123b shown in section (b) of Figure 53 described above, and includes a read area determination unit 1230', an exposure time determination unit 1231', and an AG amount determination unit 1232'.
[0398] The read information and feature amounts passed from the feature amount accumulation control unit 121 to the read determination unit 123d are input to the read area determination unit 1230′, exposure time determination unit 1231′, and AG amount determination unit 1232′, respectively. In addition, the pixel data passed from the read unit 110, the recognition information passed from the recognition processing execution unit 124, and the vehicle information and environmental information passed from the external information acquisition unit 17 are input to the read area determination unit 1230′, exposure time determination unit 1231′, and AG amount determination unit 1232′, respectively.
[0399] The readout area determination unit 1230' generates and outputs readout area information (e.g., a line number) indicating the readout area to be read next, based on at least one of the input readout information, feature amount, pixel data, recognition information, vehicle information, and environmental information. The exposure time determination unit 1231' generates and outputs information indicating the exposure time for the next imaging, based on at least one of the input readout information, feature amount, pixel data, recognition information, vehicle information, and environmental information. Furthermore, the AG amount determination unit 1232' generates and outputs information indicating the analog gain for the next imaging, based on at least one of the input readout information, feature amount, pixel data, recognition information, vehicle information, and environmental information.
[0400] (Adaptive setting method for read area) Next, a more specific description will be given of a method for adaptively setting a readout area according to a first modified example of the third embodiment. Note that the description will be given assuming that the imaging device 1 is used as an in-vehicle device.
[0401] (Example of setting the read area based on recognition information) First, a first setting method for adaptively setting a readout area based on recognition information will be described. In the first setting method, the readout determination unit 123d adaptively sets an area within a frame using an area or class detected by the recognition process of the recognition process execution unit 124, and limits the readout area to be read next. This first setting method will be described in more detail with reference to FIG. 65.
[0402] Fig. 65 is a schematic diagram for explaining a first setting method of a first modified example of the third embodiment. In section (a) of Fig. 65, line readout is performed on frame 500a in line order with line thinning as the readout unit. In the example of section (a) of Fig. 65, the recognition process execution unit 124 executes recognition process on the entire frame 500a based on pixel data read out by line readout. As a result, the recognition process execution unit 124 detects a specific object (a person in this example) in area 501 within frame 500a. The recognition process execution unit 124 passes recognition information indicating this recognition result to the read determination unit 123d.
[0403] In the read determination unit 123d, the read area determination unit 1230′ determines the read area to be read next based on the recognition information passed from the recognition process execution unit 124. For example, the read area determination unit 1230′ determines the read area to be read next to be an area including the recognized area 501 and the periphery of the area 501. The read area determination unit 1230′ passes read area information indicating the read area based on the area 502 to the reading unit 110.
[0404] The reading unit 110 performs frame reading without line thinning, for example, in accordance with the read area information passed from the read area determination unit 1230′, and passes the read pixel data to the recognition processing unit 12. Section (b) of FIG. 65 shows an example of an image read in accordance with the read area. In this example, in frame 500b, which is the frame next to frame 500a, for example, pixel data of area 502 indicated in the read area information is acquired, and the area outside area 502 is ignored. In the recognition processing unit 12, the recognition processing execution unit 124 performs recognition processing on area 502. As a result, the person detected in area 501 is recognized as a pedestrian.
[0405] In this way, by limiting the read area to be read next based on the area detected by the recognition processing execution unit 124, it becomes possible to execute high-precision recognition processing at higher speed.
[0406] Alternatively, in this first determination method, the position of the object in the next frame 500b may be predicted depending on whether the object recognized in frame 500a is a stationary object or a moving object, and the readout area to be read next may be limited based on the predicted position. In addition, if the recognized object is a moving object, the speed may be further predicted, thereby making it possible to more accurately limit the readout area to be read next.
[0407] In addition, in this first determination method, it is also possible to limit the readout area to be read next depending on the type of recognized object. For example, if the object recognized in frame 500a is a traffic light, the readout determination unit 123d can limit the readout area to be read in the next frame 500b to the lamp portion of the traffic light. Furthermore, if the object recognized in frame 500a is a traffic light, the readout determination unit 123d can change the frame readout method to a readout method that reduces the influence of flicker, and perform readout in the next frame 500b. As an example of a readout method that reduces the influence of flicker, the pattern Rd#m_x according to the fifth modified example of the second embodiment described above can be applied.
[0408] Next, a second setting method for adaptively setting a readout area based on the recognition information will be described. In the second setting method, the readout determination unit 123d limits the readout area to be read next by using the recognition information during the recognition process in the recognition process execution unit 124. This second setting method will be described in more detail with reference to FIG.
[0409] Fig. 66 is a schematic diagram for explaining the second setting method of the first modified example of the third embodiment. In this example, the object to be recognized is a vehicle license plate. Section (a) of Fig. 66 is a diagram showing an example in which an object indicating a bus vehicle is recognized in area 503 during recognition processing in response to frame readout of frame 500c. Section (a) of Fig. 66 can correspond to, for example, the example shown in section (b) of Fig. 65 above in which the readout area is limited to area 502 and readout and recognition processing are performed.
[0410] Here, when the object is recognized to be a bus vehicle in area 503 during the recognition process, the read area determination unit 1230′ can predict the position of the license plate of the bus vehicle based on the content recognized from area 503. The read area determination unit 1230′ determines the read area to be read next based on the predicted position of the license plate, and passes read area information indicating the determined read area to the reading unit 110.
[0411] The reading unit 110 reads, for example, frame 500d, which is the next frame after frame 500c, in accordance with the read area information passed from the read area determination unit 1230', and passes the read pixel data to the recognition processing unit 12. Section (b) of Figure 66 shows an example of an image read in accordance with the read area. In this example, pixel data is acquired for frame 500d in area 504, which includes the predicted position of the license plate indicated in the read area information. In the recognition processing unit 12, the recognition processing execution unit 124 performs recognition processing on area 504. As a result, recognition processing is performed on the license plate as an object included in area 504, and it is possible to acquire, for example, the vehicle number of a bus vehicle detected in the recognition processing on area 503.
[0412] In this second setting method, the read area of the next frame 500d is determined in the middle of the recognition processing of the entire target object in response to the reading of frame 500c by the recognition processing execution unit 124, thereby enabling high-precision recognition processing to be performed more quickly.
[0413] In the recognition process for frame 500c shown in section (a) of Fig. 66, if the reliability indicated in the recognition information passed from the recognition process execution unit 124 during recognition is equal to or higher than a predetermined level, the read determination unit 123d determines area 504 as the read area to be read next, and performs the reading shown in section (b) of Fig. 66. In this case, if the reliability indicated in the recognition information is lower than a predetermined level, the recognition process is performed on the entire object in frame 500c.
[0414] Next, a third setting method will be described, in which the read area is adaptively set based on the recognition information. In the third setting method, the read determination unit 123d limits the read area to be read next, using the reliability of the recognition process in the recognition process execution unit 124. This third setting method will be described in more detail with reference to FIG.
[0415] In section (a) of Fig. 67, frame 500e is read out line by line, sequentially and with line thinning, with the readout unit being a line. In the example of section (a) of Fig. 67, the recognition process execution unit 124 executes recognition process on the entire frame 500a based on the pixel data read out by line reading, and detects a specific object (a person in this example) in area 505a within frame 500e. The recognition process execution unit 124 passes recognition information indicating this recognition result to the read determination unit 123d.
[0416] In the read determination unit 123d, the read region determination unit 1230′ generates read region information indicating that reading of the frame next to frame 500e is not to be performed, for example, if the reliability indicated in the recognition information passed from the recognition processing execution unit 124 is equal to or higher than a predetermined level. The read region determination unit 1230′ passes the generated read region information to the read unit 110.
[0417] On the other hand, if the reliability indicated in the recognition information passed from the recognition processing execution unit 124 is less than a predetermined value, the readout region determination unit 1230' generates readout region information so as to perform reading of the frame next to frame 500e. For example, the readout region determination unit 1230' generates readout region information that specifies, as a readout region, an area corresponding to area 505a in which a specific object (person) is detected in frame 500e. The readout region determination unit 1230' passes the generated readout region information to the reading unit 110.
[0418] The readout unit 110 reads the frame next to frame 500e in accordance with the readout region information passed from the readout region determination unit 1230'. Here, the readout region determination unit 1230' can add to the readout region information an instruction to read the region corresponding to region 505a in the frame next to frame 500e without thinning out the pixels. The readout unit 110 reads the frame next to frame 500e in accordance with this readout region information and passes the read pixel data to the recognition processing unit 12.
[0419] Section (b) of FIG. 67 shows an example of an image read out in accordance with the read-out area information. In this example, in frame 500f, which is the frame next to frame 500e, pixel data of area 505b corresponding to area 505a indicated in the read-out area information is acquired. For example, the pixel data of frame 500e in the portion of frame 500f other than area 505b may be used as is without reading. In the recognition processing unit 12, the recognition processing execution unit 124 performs recognition processing on area 505b. This makes it possible to recognize with a higher degree of reliability that a person detected in area 501 is a pedestrian.
[0420] (Example of adaptively setting the readout area based on external information) Next, a first setting method will be described, in which the read area is adaptively set based on external information. In the first setting method, the read determination unit 123d adaptively sets an area within the frame based on vehicle information passed from the external information acquisition unit 17, and limits the read area to be read next. This makes it possible to perform recognition processing suited to the running of the vehicle.
[0421] For example, in the read determination unit 123d, the read area determination unit 1230' acquires the inclination of the vehicle based on the vehicle information and determines the read area according to the acquired inclination. As an example, when the read area determination unit 1230' acquires, based on the vehicle information, that the vehicle has run up on a bump or the like and the front side is raised, it corrects the read area toward the upper end of the frame. Furthermore, when the read area determination unit 1230' acquires, based on the vehicle information, that the vehicle is turning, it determines an unobserved area in the turning direction (for example, the area on the left end if the vehicle is turning left) as the read area.
[0422] Next, a second setting method will be described, in which the readout area is adaptively set based on external information. In the second setting method, map information capable of sequentially reflecting the current location is used as the external information. In this case, the readout area determination unit 1230' generates readout area information that instructs, for example, to increase the frequency of frame readout when the current location is in an area where caution is required for vehicle travel (for example, around a school or nursery school). This makes it possible to prevent accidents caused by children running out into the road.
[0423] Next, a third setting method will be described, in which the readout area is adaptively set based on external information. In the third setting method, detection information from another sensor is used as the external information. An example of the other sensor is a LiDAR (Laser Imaging Detection and Ranging) sensor. The readout area determination unit 1230' generates readout area information that skips reading of an area where the reliability of the detection information from the other sensor is equal to or higher than a predetermined level. This enables power saving and speeding up of frame readout and recognition processing.
[0424] [6-2. Second Modification of the Third Embodiment] Next, a second modified example of the third embodiment will be described. The second modified example of the third embodiment is an example in which at least one of the exposure time and the analog gain is adaptively set when frame readout is performed.
[0425] (Example of adaptively setting exposure time and analog gain based on image information) First, a method for adaptively setting the exposure time and analog gain based on image information (pixel data) will be described. The readout determination unit 123d detects noise contained in the pixel data passed from the readout unit 110. If the pixel data passed from the readout unit 110 contains a predetermined amount of noise or more, the exposure time determination unit 1231' sets a longer exposure time, and the AG amount determination unit 1232' sets a higher analog gain.
[0426] (Example of adaptively setting exposure time and analog gain based on recognition information) Next, a method for adaptively setting the exposure time and analog gain based on the recognition information will be described. If the reliability indicated in the recognition information passed from the recognition processing execution unit 124 is less than a predetermined value, the exposure time determination unit 1231' and the AG amount determination unit 1232' in the read determination unit 123d adjust the exposure time and analog gain, respectively. The read unit 110 reads, for example, the next frame using the adjusted exposure time and analog gain.
[0427] (Example of adaptively setting exposure time and analog gain based on external information) Next, a method for adaptively setting the exposure time and analog gain based on external information will be described. Here, vehicle information is used as the external information.
[0428] As a first example, in the read determination unit 123d, when the vehicle information passed from the external information acquisition unit 17 indicates that the headlights are on, the exposure time determination unit 1231' and the AG amount determination unit 1232' make the exposure time and analog gain different between the central and peripheral parts of the frame. That is, when the headlights are on in the vehicle, the brightness value is high in the central part of the frame and low in the peripheral part. Therefore, when the headlights are on, the exposure time determination unit 1231' shortens the exposure time and the AG amount determination unit 1232' increases the analog gain for the central part of the frame. On the other hand, when the headlights are on, the exposure time determination unit 1231' lengthens the exposure time and the AG amount determination unit 1232' decreases the analog gain for the peripheral part of the frame.
[0429] As a second example, in the read determination unit 123d, the exposure time determination unit 1231' and the AG amount determination unit 1232' adaptively set the exposure time and analog gain based on the vehicle speed indicated in the vehicle information passed from the external information acquisition unit 17. For example, blurring due to vehicle movement is less in the center of the frame. Therefore, in the center of the frame, the exposure time determination unit 1231' sets the exposure time longer, and the AG amount determination unit 1232' sets the analog gain lower. On the other hand, blurring due to vehicle movement is greater in the periphery of the frame. Therefore, in the periphery of the frame, the exposure time determination unit 1231' sets the exposure time shorter, and the AG amount determination unit 1232' sets the analog gain higher.
[0430] Here, if the vehicle speed changes to a higher speed based on the vehicle information, in the center of the frame, the exposure time determination unit 1231' changes the exposure time to a shorter time, and the AG amount determination unit 1232' changes the analog gain to a higher gain.
[0431] In this way, by adaptively setting the exposure time and analog gain, it is possible to suppress the influence of changes in the imaging environment on the recognition process.
[0432] [6-3. Third Modification of the Third Embodiment] Next, a third modified example of the third embodiment will be described. The third modified example of the third embodiment is an example in which the readout area, exposure time, analog gain, and drive speed are set according to a predetermined priority mode. The drive speed is the speed at which the sensor unit 10 is driven, and by increasing the drive speed within the range allowed for the sensor unit 10, it is possible to increase the frame readout speed, for example.
[0433] Fig. 68 is a functional block diagram of an example for explaining functions of an imaging device according to a third modified example of the third embodiment. The configuration shown in section (a) of Fig. 68 is obtained by adding a priority mode instructing unit 2020 to the configuration shown in section (a) of Fig. 64. Furthermore, the read determining unit 123e includes functions different from those of the read determining unit 123d shown in Fig. 64.
[0434] Section (b) of Fig. 68 is an example functional block diagram for explaining in more detail the function of the read determination unit 123e according to a third modified example of embodiment 3. In section (b) of Fig. 68, the read determination unit 123e has a resource adjustment unit 2000, which includes a read area determination unit 2010, an exposure time determination unit 2011, an AG amount determination unit 2012, and a drive speed determination unit 2013.
[0435] The read information and feature amounts passed from the feature amount accumulation control unit 121 to the read determination unit 123e are input to a read area determination unit 2010, an exposure time determination unit 2011, an AG amount determination unit 2012, and a drive speed determination unit 2013. In addition, the vehicle information and environmental information passed from the external information acquisition unit 17 and the pixel data passed from the read unit 110 are input to the read area determination unit 2010, an exposure time determination unit 2011, an AG amount determination unit 2012, and a drive speed determination unit 2013, respectively.
[0436] The readout area determination unit 2010 generates and outputs readout area information (e.g., line number) indicating the readout area to be read next based on the input information. The exposure time determination unit 2011 generates and outputs information indicating the exposure time for the next imaging based on the input information. The AG amount determination unit 2012 generates and outputs information indicating the analog gain for the next imaging based on the input information.
[0437] The drive speed determination unit 2013 also generates and outputs drive speed information for adjusting the drive speed of the sensor unit 10 based on the input information. One method for adjusting the drive speed by the drive speed determination unit 2013 is to change the frequency of the clock signal of the sensor unit 10. With this method, the power consumption of the sensor unit 10 changes depending on the drive speed. Another method for adjusting the drive speed is to adjust the drive speed while keeping the power consumption of the sensor unit 10 fixed. For example, the drive speed determination unit 2013 can adjust the drive speed by changing the number of bits of pixel data read from the sensor unit 10. The drive speed determination unit 2013 passes the generated drive speed information to the readout unit 110. The readout unit 110 passes the drive speed information to the sensor unit 10 by including it in imaging control information.
[0438] The priority mode instruction unit 2020 outputs priority mode setting information for setting a priority mode in response to an instruction according to a user operation or an instruction from a higher-level system. The priority mode setting information is input to the resource adjustment unit 2000. The resource adjustment unit 2000 adjusts the generation of readout area information by the readout area determination unit 2010, the generation of exposure time by the exposure time determination unit 2011, the generation of analog gain by the AG amount determination unit 2012, and the generation of drive speed information by the drive speed determination unit 2013, in accordance with the input priority mode setting information.
[0439] The priority mode instruction unit 2020 can instruct various priority modes. Examples of priority modes include an accuracy priority mode that prioritizes recognition accuracy, a power saving mode that prioritizes power consumption, a speed priority mode that prioritizes speed in reporting recognition results, a wide area priority mode that prioritizes recognition processing for a wide area, a dark place mode that prioritizes recognition processing for images captured in a dark environment, a small object priority mode that prioritizes recognition processing for small objects, and a high-speed object priority mode that prioritizes recognition processing for objects moving at high speed. The priority mode instruction unit 2020 may instruct one priority mode from each of these priority modes, or may instruct multiple priority modes from each priority mode.
[0440] (Example of priority mode operation) An example of the operation of the priority mode will be described. As a first example, an example will be described in which the resource adjustment unit 2000 adopts a setting corresponding to the priority mode instructed by the priority mode instructing unit 2020. For example, the first setting is a setting assuming a case where the imaging environment is a dark environment and the captured pixel data contains a predetermined level of noise, and the exposure time is 10 [msec] and the analog gain is 1. Furthermore, the second setting is a setting assuming a case where the imaging device 1 moves at high speed, such as when a vehicle equipped with the imaging device 1 according to the third modification of the third embodiment travels at a predetermined speed or higher, and the captured image is blurred, and the exposure time is 1 [msec] and the analog gain is 10.
[0441] In this case, when the priority mode instructing unit 2020 instructs the resource adjusting unit 2000 to use the dark place priority mode, the resource adjusting unit 2000 adopts the first setting. The resource adjusting unit 2000 instructs the exposure time determining unit 2011 and the AG amount determining unit 2012 to set the exposure time to 10 msec and the analog gain to 1, in accordance with the adopted first setting. The exposure time determining unit 2011 and the AG amount determining unit 2012 pass the instructed exposure time and analog gain to the reading unit 110, respectively. The reading unit 110 sets the exposure time and analog gain passed from the exposure time determining unit 2011 and the AG amount determining unit 2012 to the sensor unit 10.
[0442] As a second example, an example will be described in which the resource adjustment unit 2000 determines, by weighting, the settings to be adopted for the priority mode instructed by the priority mode instruction unit 2020. Taking the above-mentioned first setting and second setting as an example, the resource adjustment unit 2000 weights each of the first setting and the second setting according to the priority mode instructed by the priority mode instruction unit 2020. The resource adjustment unit 2000 determines the settings for the instructed priority mode according to the weighted first setting and second setting. Note that, for example, the setting targets (exposure time, analog gain, etc.) and the weighting values according to the targets can be set and stored in advance for each priority mode that can be instructed by the priority mode instruction unit 2020.
[0443] As a third example, an example will be described in which the resource adjustment unit 2000 weights the frequency of the setting to be adopted for the priority mode instructed by the priority mode instructing unit 2020 and determines the setting for that priority mode. As an example, consider a third setting in which uniform readout is performed to read out the entire frame approximately uniformly, and a fourth setting in which peripheral readout is performed to read out the peripheral portions of the frame with a focus. Here, the third setting is set as the normal readout setting. Furthermore, the fourth setting is set as the readout setting when an object with a reliability lower than a predetermined level is recognized.
[0444] In this case, for example, when the priority mode instructing unit 2020 instructs the resource adjustment unit 2000 to use the wide area priority mode, the resource adjustment unit 2000 can increase the frequency of readout and recognition processing using the third setting relative to the frequency of readout and recognition processing using the fourth setting. As a specific example, the resource adjustment unit 2000 specifies, for example, in a frame-by-frame time series, "third setting," "third setting," "fourth setting," "third setting," "third setting," "fourth setting," ... and increases the frequency of operation using the third setting relative to the frequency of operation using the fourth setting.
[0445] As another example, when the priority mode instructing unit 2020 instructs the resource adjustment unit 2000 to use the accuracy priority mode, the resource adjustment unit 2000 can increase the frequency of readout and recognition processing using the fourth setting relative to the frequency of readout and recognition processing using the third setting. As a specific example, the resource adjustment unit 2000 executes the following in a frame-by-frame time series, for example: "fourth setting," "fourth setting," "third setting," "fourth setting," "fourth setting," "third setting," ..., thereby increasing the frequency of operation using the fourth setting relative to the frequency of operation using the third setting.
[0446] By determining the operation in the priority mode in this manner, it becomes possible to perform appropriate frame reading and recognition processing in various situations.
[0447] (Example of driving speed adjustment) The drive speed determination unit 2013 can adjust the drive speed at which the sensor unit 10 is driven based on each piece of information passed to the read determination unit 123d. For example, when the vehicle information indicates an emergency, the drive speed determination unit 2013 can increase the drive speed to improve the accuracy and responsiveness of the recognition process. When the recognition information indicates that the reliability of the recognition process is below a predetermined level and that rereading of the frame is necessary, the drive speed determination unit 2013 can increase the drive speed. Furthermore, when the vehicle information indicates that the vehicle is turning and the recognition information indicates that an unobserved area has appeared in the frame, the drive speed determination unit 2013 can increase the drive speed until reading of the unobserved area is completed. Furthermore, when the current location is in an area requiring caution when driving a vehicle, based on map information that can sequentially reflect the current location, the drive speed determination unit 2013 can increase the drive speed.
[0448] On the other hand, when the power saving mode is instructed by the priority mode instructing unit 2020, the drive speed can be prevented from being increased except when the above-mentioned vehicle information indicates an emergency. Similarly, when the power saving mode is instructed, the readout area determining unit 2010 determines an area in the frame where an object is expected to be recognized as the readout area, and the drive speed determining unit 2013 can reduce the drive speed.
[0449] 7. Fourth Embodiment Next, a fourth embodiment of the present disclosure will be described. In the above-mentioned first to third embodiments and their respective modifications, various forms of recognition processing according to the present disclosure have been described. Here, for example, images processed for image recognition processing using machine learning are often not suitable for human viewing. In this fourth embodiment, recognition processing is performed on images that have undergone frame reading, and an image that can withstand human viewing can be output.
[0450] Fig. 69 is a functional block diagram illustrating an example of the function of the imaging device according to the fourth embodiment. The imaging device shown in Fig. 69 differs from the imaging device shown in Fig. 21 in that recognition information is supplied from the recognition processing execution unit 124 to the read determination unit 142a in the visual recognition processing unit 14.
[0451] FIG. 70 is a schematic diagram for outlining image processing according to the fourth embodiment. Here, frame reading is performed by sequentially reading out the frame horizontally and vertically, with the area Ar#z of a predetermined size described with reference to FIG. 33 as the reading unit. Section (a) of FIG. 70 schematically shows how pixel data for each of areas Ar#10, Ar#11, ..., Ar#15 is sequentially read out by the reading unit 110. The recognition processing unit 12 according to the present disclosure can perform recognition processing based on the pixel data for each of areas Ar#10, Ar#11, ..., Ar#15 read out in this order.
[0452] The visual recognition processing unit 14 sequentially updates, for example, the image of a frame using the pixel data of each of the areas Ar#10, Ar#11, ..., Ar#15 that are sequentially read out, as shown in section (b) of Fig. 70. This allows for the generation of an image suitable for visual recognition.
[0453] More specifically, the visual recognition processing unit 14 accumulates the pixel data of each of the areas Ar#10, Ar#11, ..., Ar#15 read out in order by the reading unit 110 in the image data accumulation unit 141 of the image data accumulation control unit 140. At this time, the image data accumulation control unit 140 accumulates the pixel data of each of the areas Ar#10, Ar#11, ..., Ar#15 read out from the same frame in the image data accumulation unit 141 while maintaining their positional relationship within the frame. In other words, the image data accumulation control unit 140 accumulates each pixel data in the image data accumulation unit 141 as image data in which the pixel data is mapped to each position within the frame.
[0454] For example, in response to a request from the image processing unit 143, the image data storage control unit 140 reads out the pixel data of each area Ar#10, Ar#11, ..., Ar#15 of the same frame stored in the image data storage unit 141 from the image data storage unit 141 as image data for that frame.
[0455] Here, in the present disclosure, if the recognition processing unit 12 obtains, for example, a desired recognition result during frame readout, it can end frame readout at that point (see the second embodiment, FIGS. 26 and 27, etc.). Also, if the recognition processing unit 12 obtains a predetermined recognition result during frame readout, it can jump the frame readout position to a position where the desired recognition result is predicted to be obtained based on the recognition result (see the tenth modification of the second embodiment, FIGS. 55 and 56, etc.). In these cases, frame readout ends when the recognition process ends, which may result in missing portions of the frame image.
[0456] Therefore, in the fourth embodiment, if there is an unprocessed area in the frame that has not been read out at the time when the recognition processing by the recognition processing unit 12 is completed, the visual recognition processing unit 14 reads out this unprocessed area after the recognition processing is completed, and fills in the missing part of the frame image.
[0457] FIG. 71 is a diagram showing an example of a readout process according to the fourth embodiment. The readout process according to the fourth embodiment will be described with reference to the example of FIG. 55 described above. Step S20 in FIG. 71 corresponds to the process of step S4c in FIG. 55. That is, with reference to FIG. 55, in step S1, the imaging device 1 starts capturing an image of a target image (handwritten number "8") to be recognized. In step S2, the sensor control unit 11 sequentially reads out the frame line by line from the top to the bottom of the frame in accordance with the readout area information passed from the recognition processing unit 12. When the lines have been read out up to a certain position, the recognition processing unit 12 identifies the number "8" or "9" from the image formed by the readout lines (step S3).
[0458] The pixel data read by the reading unit 110 in step S2 is passed to the recognition processing unit 12 and also to the visual recognition processing unit 14. In the visual recognition processing unit 14, the image data storage control unit 140 sequentially stores the pixel data passed from the reading unit 110 in the image data storage unit 141.
[0459] In the recognition processing unit 12, the read determination unit 123f generates read area information specifying a predicted line along which it is predicted that the object identified in step S3 can be identified as either the number "8" or "9" based on the results of the recognition processing up to step S3, and passes this information to the read unit 110. The read unit 110 passes the predicted line read out in accordance with this read area information to the recognition processing unit 12 and also passes it to the visual recognition processing unit 14. In the visual recognition processing unit 14, the image data accumulation control unit 140 accumulates the pixel data of the predicted line passed from the read unit 110 in the image data accumulation unit 141.
[0460] The recognition processing unit 12 executes recognition processing (step S20) based on the pixel data of the predicted line passed by the reading unit 110. When the object is determined in step S20, the recognition processing unit 12 outputs the recognition result (step S21).
[0461] In the recognition processing unit 12, when the recognition result is output in step S21, the recognition processing execution unit 124 passes recognition information indicating the end of the recognition processing to the visual recognition processing unit 14. In response to the recognition information passed from the recognition processing execution unit 124, the visual recognition processing unit 14 reads out the unprocessed area at the time of step S20 (step S22).
[0462] More specifically, the visual recognition processing unit 14 passes the recognition information indicating the end of the recognition processing, which has been passed to the read determination unit 142a. The read determination unit 142a sets a read area for reading out the unprocessed area according to the passed recognition information. The read determination unit 142a passes read area information indicating the set read area to the read unit 110. The read unit 110 reads out the unprocessed area in the frame according to the passed read area information, and passes the read pixel data to the visual recognition processing unit 14. In the visual recognition processing unit 14, the image data accumulation control unit 140 accumulates the pixel data of the unprocessed area passed from the read unit 110 in the image data accumulation unit 141.
[0463] When the reading unit 110 finishes reading the unprocessed area of the frame, reading of the entire frame is completed. The image data storage unit 141 stores pixel data read for the recognition process and pixel data read from the unprocessed area after the recognition process is completed. Therefore, for example, the image data storage control unit 140 can output image data of the entire frame image by reading pixel data of the same frame from the image data storage unit 141 (step S23).
[0464] It is preferable that the series of processes from step S1 to step S23 described in FIG. 71 be executed within one frame period, since this allows the visibility processing unit 14 to output a moving image in approximately real time relative to the timing of imaging.
[0465] Note that the readout for the recognition process up to step S20 and the readout for the visual confirmation process in step S22 are performed in a different order from the sequential readout of the lines of the frame. Therefore, for example, if the imaging method of the sensor unit 10 is a rolling shutter method, a mismatch between the readout order and time (called a time lag) occurs in each line group where the readout order is divided. This time lag can be corrected by image processing based on, for example, the line number of each line group and the frame period.
[0466] Furthermore, this time difference becomes more pronounced when the image capture device 1 is moving relative to the subject. In this case, if the image capture device 1 is equipped with a gyro capable of detecting angular velocity in three directions, the movement direction and speed of the image capture device 1 can be determined based on the detection output of the gyro, and the time difference can be corrected by further using this movement direction and speed.
[0467] Fig. 72 is a flowchart showing an example of processing according to the fourth embodiment. In the flowchart of Fig. 72, the processing from step S200 to step S206 is equivalent to the processing from step S100 to step S1080 in the flowchart of Fig. 54 described above.
[0468] That is, in step S200, the reading unit 110 reads line data from the line indicated by the read line of the target frame. The reading unit 110 passes the line data consisting of the pixel data of the read line to the recognition processing unit 12 and the visibility processing unit 14.
[0469] When the process of step S200 ends, the process proceeds to steps S201 and S211. The processes of steps S201 to S208 are processes in the recognition processing unit 12. On the other hand, the processes of steps S211 to S214 are processes in the visual recognition processing unit 14. The processes in the recognition processing unit 12 and the processes in the visual recognition processing unit 14 can be executed in parallel.
[0470] First, the processing by the recognition processing unit 12 from step S201 will be described. In step S201, the recognition processing unit 12 determines whether or not the recognition processing for the target frame has been completed. If the recognition processing unit 12 determines that the processing has been completed (step S201, "Yes"), it does not execute the processing from step S202 onwards. On the other hand, if the recognition processing unit 12 determines that the processing has not been completed (step S201, "No"), it transitions the processing to step S202.
[0471] The processing in steps S202 to S208 is equivalent to the processing in steps S101 to S1080 in Fig. 54. That is, in step S202, the feature amount calculation unit 120 in the recognition processing unit 12 calculates line feature amounts based on the line data passed from the reading unit 110. In the next step S203, the feature amount calculation unit 120 acquires the feature amounts accumulated in the feature amount accumulation unit 122 from the feature amount accumulation control unit 121. In the next step S204, the feature amount calculation unit 120 integrates the feature amount calculated in step S202 and the feature amount acquired from the feature amount accumulation control unit 121 in step S203, and passes the integrated result to the feature amount accumulation control unit 121.
[0472] In the next step S205, the feature accumulation control unit 121 accumulates the integrated feature in the feature accumulation unit 122. In the next step S206, the recognition process execution unit 124 executes recognition process using the integrated feature. In the next step S207, the recognition process execution unit 124 outputs the recognition result of the recognition process of step S206. Here, the recognition process execution unit 124 passes recognition information including the recognition result to the readout determination unit 142a of the visual recognition processing unit 14.
[0473] In the next step S208, the readout area determination unit 1230 in the readout determination unit 123f determines the readout line to be read out next using the readout information passed from the feature amount accumulation control unit 121 and a feature obtained by integrating the feature amount calculated in step S202 and the feature amount acquired from the feature amount accumulation control unit 121 in step S203. The readout determination unit 123f passes information indicating the determined readout line (readout area information) to the readout control unit 111 in the sensor control unit 11. When the processing of step S208 ends, the processing proceeds to step S220.
[0474] Next, the processing by the visual recognition processing unit 14 from step S211 will be described. In step S211, the image data storage control unit 140 in the visual recognition processing unit 14 stores the line data passed from the reading unit 110 in the image data storage unit 141. In the next step S212, the image processing unit 143 in the visual recognition processing unit 14 performs image processing for visual recognition on image data based on the line data stored in the image data storage unit 141, for example. In the next step S213, the image processing unit 143 outputs the image data that has been subjected to image processing for visual recognition.
[0475] Alternatively, in step S213, the image processing unit 143 may store the image data that has been subjected to image processing for viewing again in the image data storage unit 141. Furthermore, when the image data of the entire target frame has been stored in the image data storage unit 141, the image processing unit 143 may perform the image processing in step S212 on the image data.
[0476] In the next step S214, the read determination unit 142a in the visual recognition processing unit 14 determines the read line to be read next based on the line information indicating the line data read in step S200 and the recognition information passed from the recognition processing execution unit 124 in step S207. The read determination unit 142a passes information indicating the determined read line (read area information) to the read control unit 111. When the processing of step S214 ends, the process proceeds to step S220.
[0477] In step S220, the read control unit 111 passes read area information indicating either the read line passed from the recognition processing unit 12 in step S208 or the read line passed from the visual processing unit 14 in step S214 to the read unit 110. Here, if the recognition processing unit 12 is performing the recognition processing (step S201, “No”), the read line passed from the recognition processing unit 12 in step S208 matches the read line passed from the visual processing unit 14 in step S214. Therefore, the read control unit 111 may pass read area information indicating either the read line passed from the recognition processing unit 12 or the read line passed from the visual processing unit 14 to the read unit 110. On the other hand, if the recognition processing unit 12 is not performing the recognition processing (step S201, “Yes”), the read control unit 111 passes the read area information passed from the visual processing unit 14 to the read unit 110.
[0478] In this way, in the fourth embodiment, the unprocessed area of the frame is read out after the recognition process is completed, so that the image of the entire frame can be obtained even if the recognition process is terminated midway or if a jump in the read position occurs during the recognition process.
[0479] In the above description, the visual recognition processing unit 14 is described as sequentially updating the frame image using pixel data read out during frame readout, but this is not limited to this example. For example, the visual recognition processing unit 14 may store the pixel data read out during frame readout in, for example, the image data storage unit 141, and when the amount of stored pixel data for the same frame exceeds a threshold, read all of the pixel data for the same frame from the image data storage unit 141. Furthermore, for example, when frame readout is performed by line thinning, the thinned-out portion may be interpolated using surrounding pixel data.
[0480] (Example of trigger for image data output) When image data is output in frame units, the visual recognition processing unit 14 outputs the image data after image data for a frame has been accumulated in the image data accumulation unit 141. On the other hand, when image data is not output in frame units, the visual recognition processing unit 14 can sequentially output line data passed from the reading unit 110, for example.
[0481] (Image data storage control) Next, an example of control of the image data storage unit 141 that can be applied to the fourth embodiment will be described. As a first example of control of the image data storage unit 141, the image data storage control unit 140 stores line data passed from the reading unit 110 in the image data storage unit 141 when the image data stored in the image data storage unit 141 is insufficient.
[0482] As an example, when image data storage unit 141 does not store image data for a unit of image processing by image processing unit 143, line data passed from read-out unit 110 is stored in image data storage unit 141. As a more specific example, when image processing unit 143 performs image processing on a frame-by-frame basis, if the image data of a target frame stored in image data storage unit 141 is less than one frame, pixel data passed from read-out unit 110 may be stored in image data storage unit 141.
[0483] As a second example of control of the image data storage unit 141, when a scene of the imaged object changes, image data stored in the image data storage unit 141 is discarded. A change in the scene of the imaged object occurs, for example, due to a sudden change in the brightness, movement, or screen composition of the imaged object. When the imaging device 1 is mounted on a vehicle and used, when the vehicle enters or exits a tunnel, the brightness of the imaged object suddenly changes, causing a scene change. Furthermore, when the vehicle suddenly accelerates, stops, or makes a sharp turn, the movement of the imaged object suddenly changes, causing a scene change. Furthermore, when the vehicle suddenly exits a crowded area into an open area, the screen composition of the imaged object suddenly changes, causing a scene change. A change in the scene of the imaged object can be determined based on pixel data passed from the readout unit 110. Alternatively, a change in the scene of the imaged object can also be determined based on recognition information passed from the recognition processing execution unit 124 to the visual recognition processing unit 14.
[0484] As a third example of control of the image data storage unit 141, line data passed from the reading unit 110 is not stored in the image data storage unit 141. The third example of control of the image data storage unit 141 according to the fourth embodiment will be described with reference to Fig. 73. In Fig. 73, step S30 shows a state in which all pixel data included in frame 510, which includes area 511 in which a person has been recognized, has been read out.
[0485] In the next step S31, the recognition processing unit 12 reads frames using the readout unit 110 after a certain time has elapsed since step S30, and determines whether or not there has been a change in the recognition result in the area 511 recognized in step S30. In this example, the recognition processing unit 12 makes this determination by executing recognition processing on a portion of the area 511 in which a person has been recognized (in this example, line L#tgt). For example, the recognition processing unit 12 can determine that there has been no change in the recognition result if the amount of change in the recognition score of the portion on which recognition processing was executed in step S31 relative to the state in step S30 is equal to or less than a threshold. If the recognition processing unit 12 determines that there has been no change in the recognition result, it does not store the line data read out in step S31 in the image data storage unit 141.
[0486] For example, if only the movement of the hair of a person included in area 511 is recognized and there is no change in the position of the person, the recognition score may be lower than the threshold, and the recognition result may be determined to be no change. In this case, consistency of the image data can be maintained from the perspective of visibility even if the pixel data read in step S31 is not stored. In this way, by not storing the pixel data read when there is a change in the image but no change in the recognition result, it is possible to save the capacity of image data storage unit 141.
[0487] [7-1. First Modification of the Fourth Embodiment] Next, a first modified example of the fourth embodiment will be described. The first modified example of the fourth embodiment is an example in which, when outputting an image obtained by frame readout, an area in which a specific object has been recognized or is predicted to be recognized is masked.
[0488] A first modified example of the fourth embodiment will be described with reference to Fig. 74. In Fig. 74, the recognition processing unit 12 recognizes a part of a specific object (a person in this example) in the area 521a at the time when the recognition processing unit 12 reads out line data from the top end of the frame 520 to the position of the line L#m in step S41. The visual recognition processing unit 14 outputs an image based on the line data up to this line L#m. Alternatively, the visual recognition processing unit 14 stores the line data up to this line L#m in the image data storage unit 141.
[0489] Here, the recognition processing unit 12 can predict the entire specific object (shown as area 521b in step S42) when it recognizes a part of the specific object. The recognition processing unit 12 passes recognition information including information on area 521b where the specific object has been recognized and predicted to the visibility processing unit 14. The recognition processing unit 12 also ends the recognition processing at the position of line L#m where the specific object has been recognized.
[0490] The visual recognition processing unit 14 continues to read line data from the frame 520 after line L#m, and outputs an image based on the read line data, or stores the read line data in the image data storage unit 141. At this time, the visual recognition processing unit 14 masks the portion read after line L#m in the region 521b predicted to include a specific object (step S42). For example, in the visual recognition processing unit 14, the image processing unit 143 masks a portion of this region 521b and outputs it. Alternatively, the image processing unit 143 may store a frame image in which a portion of this region 521b is masked in the image data storage unit 141. Furthermore, the visual recognition processing unit 14 may not read pixel data after line L#m in this region 521b.
[0491] Furthermore, the visual recognition processing unit 14 may mask the entire region 521b predicted to include a specific object, as shown in step S43. In this case, the masking target is pixel data to be output by the visual recognition processing unit 14, and pixel data to be used by the recognition processing unit 12 for recognition processing, for example, is not masked.
[0492] In the above description, in frame 520, region 521b in which a specific object is recognized is masked, and an image of the other portion is output, for example. This is not limited to this example, and it is also possible to mask the portion other than region 521b, and output an image of region 521b, for example.
[0493] In this way, by masking the area 521b in which a specific object is recognized in the visual image, it is possible to realize privacy protection. For example, when the imaging device 1 according to the first modification of the fourth embodiment is applied to a road monitoring camera, a drive recorder, a drone-mounted camera, or the like, it is possible to erase only personal information from the captured image data (for example, the image data) and make it into a format that is easy to handle. Examples of specific objects to be masked in such applications include people, faces, vehicles, and vehicle license plates.
[0494] [7-2. Second Modification of the Fourth Embodiment] Next, a second modified example of the fourth embodiment will be described. The second modified example of the fourth embodiment is an example in which, when an output of a sensor that performs object detection using another method and an output image of the imaging device 1 are integrated and displayed, an area in the frame that is suitable for display is preferentially read out.
[0495] A second modified example of the fourth embodiment will be described with reference to Fig. 75. Here, a LiDAR sensor (hereinafter referred to as a LiDAR sensor) is applied as the sensor of the other type. Not limited to this, a radar or the like can also be applied as the sensor of the other type. Although not shown in the drawings, the imaging device 1 according to the second modified example of the fourth embodiment inputs the detection result by the sensor of the other type to the readout determination unit 142a instead of, for example, the recognition information from the recognition processing execution unit 124.
[0496] 75, section (a) shows an example of an image 530 acquired by a LiDAR sensor. In this section (a), it is assumed that an area 531 is an image acquired within a range within a predetermined distance (for example, 10 [m] to several tens [m]) from the LiDAR sensor.
[0497] 75, section (b) shows an example of a frame 540 captured by an imaging device 1 according to a second modified example of the fourth embodiment. In this frame 540, a shaded area 541 corresponds to area 531 in section (a) and includes objects such as objects within a predetermined distance from the imaging device 1. On the other hand, area 542 is an area captured beyond the predetermined distance from the imaging device 1, and includes, for example, the sky and distant scenery.
[0498] In section (a) of FIG. 75, an area 531 corresponding to a range within a predetermined distance from the LiDAR sensor and the imaging device 1 is considered to be an area where objects are densely packed and suitable for display. On the other hand, areas other than the area 531 are considered to be areas where objects are sparse and there is little need for prioritized display. Therefore, in the visibility processing unit 14, the read determination unit 142a reads the frame 540 preferentially from the area 541 corresponding to the area 531 out of the areas 541 and 542. For example, the read determination unit 142a reads the area 541 at high resolution without thinning. On the other hand, the read determination unit 142a reads the area 542 at low resolution, for example, by thinning readout, or does not read it at all.
[0499] In this way, the imaging device 1 according to the second variant of the fourth embodiment can optimize the reading of frame 540 by setting the frame reading resolution according to the detection results of a sensor of another type.
[0500] [7-3. Third Modification of the Fourth Embodiment] Next, a third modified example of the fourth embodiment will be described. The third modified example of the fourth embodiment is an example in which frame readout is adaptively performed by the recognition processing unit 12 and the visibility processing unit 14. In a first example of the third modified example of the fourth embodiment, a region in which a specific object is recognized is read out first, and then an unprocessed region is read out. In this case, the unprocessed region is read out using a low-resolution readout method such as thinning readout.
[0501] A first example of the third modified example of the fourth embodiment will be described with reference to Fig. 76. In Fig. 76, the recognition processing unit 12 recognizes a specific object (a person in this example) in the area 541 at the time when the recognition processing unit 12 reads out line data from the top end of the frame 540 to the position of the line L#m in step S50. The visual recognition processing unit 14 outputs an image based on the line data up to this line L#m. Alternatively, the visual recognition processing unit 14 stores the line data up to this line L#m in the image data storage unit 141.
[0502] In step S50 of Figure 76, similar to step S40 of Figure 74 described above, the recognition processing unit 12 predicts the entire area 541 of the specific object at the time when it recognizes a part of the specific object (the area up to line L#m).
[0503] In the next step S51, the recognition processing unit 12 preferentially reads out the region 541 recognized in step S50 in the portion below the line L#m of the frame 540. The recognition processing unit 12 can perform more detailed recognition processing based on the pixel data read out from the region 541. When the recognition processing unit 12 recognizes a specific object in the region 541, it ends the recognition processing. Furthermore, the visual recognition processing unit 14 outputs an image based on the pixel data of this region 541. Alternatively, the visual recognition processing unit 14 stores the pixel data of this region 541 in the image data storage unit 141.
[0504] In the next step S52, the visual recognition processing unit 14 reads out the area from line L#m onwards in frame 540. At this time, the visual recognition processing unit 14 can read out the area from line L#m onwards at a lower resolution than the reading out of the line before L#m. In the example of Fig. 76, the visual recognition processing unit 14 reads out the area from line L#m onwards by thinning out reading. Note that the visual recognition processing unit 14 can perform this reading out, excluding, for example, the area 541.
[0505] The visual recognition processing unit 14 outputs an image based on the line data from line L#m onwards of frame 540, which has been read out by thinning-out reading. In this case, the image processing unit 143 in the visual recognition processing unit 14 can output the thinned-out line data by interpolating between lines. Alternatively, the visual recognition processing unit 14 stores the line data from line L#m onwards in the image data storage unit 141.
[0506] As described above, in the first example of the third modified example of the fourth embodiment, after a specific object is recognized by the recognition processing unit 12, the unprocessed region is read out at a low resolution by the visibility processing unit 14. This enables more accurate recognition processing of the specific object and enables the image of the entire frame to be output at a higher speed.
[0507] Next, a second example of the third modified example of the fourth embodiment will be described. In this second example, the readout conditions are changed between the recognition process and the visual confirmation process. As an example, at least one of the exposure time and the analog gain is made different between the recognition process and the visual confirmation process. As a specific example, the recognition processing unit 12 performs frame readout by capturing an image with the analog gain maximized. On the other hand, the visual confirmation processing unit 14 performs frame readout by capturing an image with the exposure time appropriately set.
[0508] At this time, frame readout by the recognition processing unit 12 and frame readout by the visual recognition processing unit 14 can be executed alternately for each frame. However, the present invention is not limited to this, and frame readout by the recognition processing unit 12 and frame readout by the visual recognition processing unit 14 can also be executed by switching between them for each readout unit (for example, line). This makes it possible to execute the recognition processing and the visual recognition processing under appropriate conditions.
[0509] 8. Fifth Embodiment Next, as a fifth embodiment, an application example of the imaging device 1 according to the first to fourth embodiments and each modified example according to the present disclosure will be described. Fig. 77 is a diagram showing an example of using the imaging device 1 according to the above-described first to fourth embodiments and each modified example.
[0510] The imaging device 1 described above can be used in various cases for sensing light such as visible light, infrared light, ultraviolet light, and X-rays, for example, as follows.
[0511] ·Devices that take images for viewing purposes, such as digital cameras and mobile devices with camera functions. - Devices used for traffic purposes, such as in-vehicle sensors that take pictures of the front, rear, surroundings, and interior of a vehicle for safe driving such as automatic stopping, and for recognizing the driver's condition, surveillance cameras that monitor moving vehicles and roads, and distance measuring sensors that measure distances between vehicles. A device used in home appliances such as TVs, refrigerators, and air conditioners to capture user gestures and operate the appliances in accordance with those gestures. -Devices used for medical and healthcare purposes, such as endoscopes and devices that take blood vessel images by receiving infrared light. -Devices used for security purposes, such as surveillance cameras for crime prevention and cameras for person authentication. - Cosmetic devices such as skin measuring devices that take pictures of the skin and microscopes that take pictures of the scalp. - Devices used for sports, such as action cameras and wearable cameras for sports. Agricultural equipment, such as cameras for monitoring the condition of fields and crops.
[0512] [Further application examples of the technology disclosed herein] The technology according to the present disclosure (the present technology) can be applied to various products. For example, the technology according to the present disclosure may be applied to devices mounted on various moving bodies such as automobiles, electric vehicles, hybrid electric vehicles, motorcycles, bicycles, personal mobility devices, airplanes, drones, ships, and robots.
[0513] FIG. 78 is a block diagram showing a schematic configuration example of a vehicle control system, which is an example of a mobile object control system to which the technology of the present disclosure can be applied.
[0514] The vehicle control system 12000 includes a plurality of electronic control units connected via a communication network 12001. In the example shown in Fig. 78, the vehicle control system 12000 includes a drive system control unit 12010, a body system control unit 12020, an outside-vehicle information detection unit 12030, an inside-vehicle information detection unit 12040, and an integrated control unit 12050. Also shown as functional components of the integrated control unit 12050 are a microcomputer 12051, an audio / video output unit 12052, and an in-vehicle network I / F (interface) 12053.
[0515] The drivetrain control unit 12010 controls the operation of devices related to the drivetrain of the vehicle in accordance with various programs. For example, the drivetrain control unit 12010 functions as a control device for a drive force generating device for generating a drive force for the vehicle, such as an internal combustion engine or a drive motor, a drive force transmission mechanism for transmitting the drive force to the wheels, a steering mechanism for adjusting the steering angle of the vehicle, a braking device for generating a braking force for the vehicle, etc.
[0516] The body system control unit 12020 controls the operation of various devices equipped in the vehicle body according to various programs. For example, the body system control unit 12020 functions as a control device for a keyless entry system, a smart key system, a power window device, or various lamps such as headlamps, backup lamps, brake lamps, turn signals, and fog lamps. In this case, radio waves transmitted from a portable device that serves as a key or signals from various switches may be input to the body system control unit 12020. The body system control unit 12020 receives these radio waves or signals and controls the vehicle's door lock device, power window device, lamps, etc.
[0517] The outside-vehicle information detection unit 12030 detects information outside the vehicle equipped with the vehicle control system 12000. For example, an imaging unit 12031 is connected to the outside-vehicle information detection unit 12030. The outside-vehicle information detection unit 12030 causes the imaging unit 12031 to capture images outside the vehicle and receives the captured images. The outside-vehicle information detection unit 12030 may perform object detection processing or distance detection processing for people, cars, obstacles, signs, characters on the road surface, etc. based on the received images. The outside-vehicle information detection unit 12030, for example, performs image processing on the received images and performs object detection processing or distance detection processing based on the results of the image processing.
[0518] The imaging unit 12031 is an optical sensor that receives light and outputs an electrical signal according to the amount of light received. The imaging unit 12031 can output the electrical signal as an image, or as distance measurement information. The light received by the imaging unit 12031 may be visible light or invisible light such as infrared light.
[0519] The in-vehicle information detection unit 12040 detects information inside the vehicle. For example, a driver state detection unit 12041 that detects the state of the driver is connected to the in-vehicle information detection unit 12040. The driver state detection unit 12041 includes, for example, a camera that captures an image of the driver, and the in-vehicle information detection unit 12040 may calculate the degree of fatigue or concentration of the driver based on the detection information input from the driver state detection unit 12041, or may determine whether the driver is dozing off.
[0520] The microcomputer 12051 can calculate control target values for the driving force generating device, steering mechanism, or braking device based on the information inside and outside the vehicle acquired by the outside-vehicle information detection unit 12030 or the inside-vehicle information detection unit 12040, and output control commands to the drivetrain control unit 12010. For example, the microcomputer 12051 can perform cooperative control aimed at realizing the functions of an ADAS (Advanced Driver Assistance System), including avoiding or mitigating collisions between vehicles, following based on the distance between vehicles, maintaining vehicle speed, warning of vehicle collisions, or warning of vehicle lane departure.
[0521] In addition, the microcomputer 12051 can perform cooperative control for the purpose of autonomous driving, which allows the vehicle to travel autonomously without relying on driver operation, by controlling the driving force generating device, steering mechanism, braking device, etc. based on information about the surroundings of the vehicle obtained by the outside vehicle information detection unit 12030 or the inside vehicle information detection unit 12040.
[0522] Furthermore, the microcomputer 12051 can output a control command to the body system control unit 12020 based on the information about the outside of the vehicle acquired by the outside information detection unit 12030. For example, the microcomputer 12051 can control the headlamps according to the position of a preceding vehicle or an oncoming vehicle detected by the outside information detection unit 12030, and perform cooperative control for the purpose of preventing glare, such as switching from high beams to low beams.
[0523] The audio / video output unit 12052 transmits at least one of audio and video output signals to an output device capable of visually or audibly notifying information to vehicle occupants or the outside of the vehicle. In the example of Fig. 78, an audio speaker 12061, a display unit 12062, and an instrument panel 12063 are exemplified as output devices. The display unit 12062 may include, for example, at least one of an on-board display and a head-up display.
[0524] Fig. 79 is a diagram showing an example of the installation position of the image capturing unit 12031. In Fig. 79, the vehicle 12100 has image capturing units 12101, 12102, 12103, 12104, and 12105 as the image capturing unit 12031.
[0525] The imaging units 12101, 12102, 12103, 12104, and 12105 are provided, for example, at positions such as the front nose, side mirrors, rear bumper, back door, and the top of the windshield inside the vehicle cabin of the vehicle 12100. The imaging unit 12101 provided at the front nose and the imaging unit 12105 provided at the top of the windshield inside the vehicle cabin mainly acquire images of the front of the vehicle 12100. The imaging units 12102 and 12103 provided at the side mirrors mainly acquire images of the sides of the vehicle 12100. The imaging unit 12104 provided at the rear bumper or back door mainly acquires images of the rear of the vehicle 12100. The images of the front acquired by the imaging units 12101 and 12105 are mainly used to detect preceding vehicles, pedestrians, obstacles, traffic lights, traffic signs, lanes, etc.
[0526] 79 shows an example of the imaging ranges of the imaging units 12101 to 12104. Imaging range 12111 indicates the imaging range of the imaging unit 12101 provided on the front nose, imaging ranges 12112 and 12113 indicate the imaging ranges of the imaging units 12102 and 12103 provided on the side mirrors, respectively, and imaging range 12114 indicates the imaging range of the imaging unit 12104 provided on the rear bumper or back door. For example, by overlaying the image data captured by the imaging units 12101 to 12104, a bird's-eye view image of the vehicle 12100 viewed from above can be obtained.
[0527] At least one of the imaging units 12101 to 12104 may have a function of acquiring distance information. For example, at least one of the imaging units 12101 to 12104 may be a stereo camera made up of multiple imaging elements, or an imaging element having pixels for detecting a phase difference.
[0528] For example, the microcomputer 12051 can extract, as a preceding vehicle, the three-dimensional object that is the closest three-dimensional object on the path of the vehicle 12100 and traveling in approximately the same direction as the vehicle 12100 at a predetermined speed (for example, 0 km / h or higher) by calculating the distance to each three-dimensional object within the imaging ranges 12111-12114 and the change in this distance over time (relative speed with respect to the vehicle 12100) based on the distance information obtained from the imaging units 12101-12104. Furthermore, the microcomputer 12051 can set a vehicle-to-vehicle distance to be maintained in advance in front of the preceding vehicle, and perform automatic braking control (including follow-up stop control), automatic acceleration control (including follow-up start control), etc. In this way, cooperative control can be ...
Claims
1. An information processing device that processes pixel signals read out from pixels included in a pixel area in which a plurality of pixels are arranged, a recognition unit that performs a recognition process on the pixel signal corresponding to the readout unit set as part of the pixel region using a learning model that has been learned based on teacher data of the readout unit, and outputs a recognition result of the recognition process; Equipped with The recognition unit performing the recognition processing on the pixel signals sequentially read out for each read unit from the pixels included in the pixel region; If the recognition result does not satisfy a predetermined condition, the recognition process is continued for the pixel signals of the readout unit; If the recognition result satisfies a predetermined condition, the recognition process is terminated. Information processing device.
2. The readout unit is composed of a plurality of the pixels corresponding to a predetermined number of rows or columns of the array. The information processing device according to claim 1 .
3. The readout unit is composed of a plurality of the pixels aligned in one row or column of the array. The information processing device according to claim 1 .
4. The readout unit is composed of a plurality of the pixels corresponding to a plurality of rows or columns of the array. The information processing device according to claim 1 .
5. The readout unit is a pattern consisting of a plurality of the pixels including the pixels that are not adjacent to each other. The information processing device according to claim 1 .
6. The pattern is made up of a plurality of the pixels arranged according to a predetermined rule. The information processing device according to claim 5 .
7. The recognition process includes a face detection process. The information processing device according to claim 1 .
8. an information processing device including: a recognition unit that performs recognition processing on pixel signals read from pixels included in a pixel area, the pixel area corresponding to the readout unit set as a part of a pixel area in which a plurality of pixels are arranged, using a learning model that has been learned based on teacher data of the readout unit, and outputs a recognition result of the recognition processing; an imaging unit having the pixel area; a readout unit that reads out the readout unit set as a part of the pixel area; Equipped with The recognition unit performing the recognition processing on the pixel signals sequentially read out for each read unit from the pixels included in the pixel region; If the recognition result does not satisfy a predetermined condition, the recognition process is continued for the pixel signals of the readout unit; If the recognition result satisfies a predetermined condition, the recognition process is terminated. Solid-state imaging element.
9. The reading unit reading the read unit at the read position determined using a learning model learned based on teacher data; The solid-state imaging device according to claim 8 .
10. The reading unit reading the reading unit at the reading position determined based on the recognition result; The solid-state imaging device according to claim 8 .
11. The reading unit When the recognition unit acquires a candidate for the recognition result that satisfies a predetermined condition, the reading unit reads out the reading unit at a position where the recognition result that satisfies the predetermined condition is expected to be acquired. The solid-state imaging device according to claim 8 .
12. The reading unit thinning out the pixels included in the pixel region in units of readout to read out the pixel signals; The solid-state imaging device according to claim 8 .
13. The reading unit based on the recognition result, controlling at least one of exposure of the pixels included in the pixel area and a gain for the pixel signals read out from the pixels included in the pixel area, and reading out the readout unit. The solid-state imaging device according to claim 8 .
14. The reading unit reading out the readout unit determined based on at least one of pixel information based on the pixel signal, recognition information generated by the recognition processing, and external information acquired from the outside; The solid-state imaging device according to claim 8 .
15. The reading unit reading the read unit determined based on an area within the pixel area indicated by the recognition result; The solid-state imaging device according to claim 8 .
16. The reading unit and reading out the readout unit by controlling at least one of exposure at the pixels included in the pixel area and a gain for the pixel signals read out from the pixels included in the pixel area based on at least one of pixel information based on the pixel signals, recognition information generated by the recognition processing, and external information acquired from the outside. The solid-state imaging device according to claim 8 .
17. The reading unit and controlling at least one of the readout unit, exposure at the pixels included in the pixel area, and gain for the pixel signals read out from the pixels included in the pixel area in accordance with an operation mode instructed from an external device, thereby reading out pixel signals from each of the pixels included in the pixel area. The solid-state imaging device according to claim 8 .
18. the imaging unit, the readout unit, and the recognition unit are arranged in a single chip; The solid-state imaging device according to claim 8 .
19. the single chip has a laminated structure in which a first chip and a second chip are bonded together; the imaging unit is disposed on the first chip, the readout unit and the recognition unit are disposed on the second chip. The solid-state imaging device according to claim 18.
20. Executed by a processor, An information processing method for processing pixel signals read out from pixels included in a pixel area in which a plurality of pixels are arranged, comprising: a recognition step of performing a recognition process on the pixel signal corresponding to the readout unit set as part of the pixel area using a learning model learned based on teacher data of the readout unit, and outputting a recognition result of the recognition process; Including, The recognition step includes: performing the recognition processing on the pixel signals sequentially read out for each read unit from the pixels included in the pixel region; If the recognition result does not satisfy a predetermined condition, the recognition process is continued for the pixel signals of the readout unit; If the recognition result satisfies a predetermined condition, the recognition process is terminated. Information processing methods.
Citation Information
Patent Citations
Character recognizing device
JP1989093875A
Character recognition translation system
JP1997138802A
Image recognition apparatus and image recognition method
JP2015177300A
Imaging apparatus and method
JP2017112409A
Imaging device, method for controlling the same, and imaging element
JP2017183952A