Information processing device, information processing system, information processing method, and information processing program

By controlling pixel signal readout and generating reliability maps based on area, count, and exposure, the solution enhances image recognition reliability and accuracy in partial region processing.

JP7769915B2Active Publication Date: 2025-11-14SONY GROUP CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022538657
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-20
Filing Date
2021-06-25
Publication Date
2025-11-14
Estimated Expiration
2041-06-25

AI Technical Summary

Technical Problem

Conventional image recognition methods using partial regions of image data face reliability issues due to varying line numbers and line widths, leading to decreased accuracy.

Method used

A readout unit controls pixel signal readout in a two-dimensional pixel array, calculating reliability based on area, readout count, dynamic range, and exposure information, generating reliability maps, and correcting reliability values to enhance recognition processing.

Benefits of technology

The solution stabilizes reliability in image recognition by adjusting for variations in line numbers and widths, improving accuracy and consistency in image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007769915000003
    Figure 0007769915000003
  • Figure 0007769915000004
    Figure 0007769915000004
  • Figure 0007769915000005
    Figure 0007769915000005
Patent Text Reader

Abstract

[Problem] To provide an imaging device, an imaging system, an imaging method, and an imaging program that are capable of suppressing a decrease in the precision of reliability, even for a case in which a recognition process is carried out using a partial region of image data. [Solution] Provided is an information processing device comprising: a reading section that sets, as a reading unit, a portion of a pixel region in which a plurality of pixels are arrayed in a two-dimensional array pattern, and controls the reading out of pixel signals from pixels included in the pixel region; and a reliability calculation unit that calculates the reliability of a prescribed region within the pixel region on the basis of at least one of the surface area, number of read-out times, dynamic range, and exposure information, of a region of a captured image that was set as the reading unit and read out.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, an information processing system, an information processing method, and an information processing program. [Background technology]

[0002] In recent years, with the increasing performance of imaging devices such as digital still cameras, digital video cameras, and compact cameras installed in multi-function mobile phones (smartphones), imaging devices equipped with an image recognition function that recognizes predetermined objects contained in captured images have been developed. Furthermore, efforts are being made to speed up recognition processing by using partial regions of image data within one frame. Furthermore, in recognition processing, reliability is generally assigned as an evaluation value of recognition accuracy.

[0003] However, in new recognition methods using partial regions, such as line image data, the number of lines and line width may change depending on the recognition target, which may result in a decrease in accuracy when using conventional reliability. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-112409 Summary of the Invention [Problem to be solved by the invention]

[0005] One aspect of the present disclosure provides an information processing device, an information processing system, an information processing method, and an information processing program that are capable of suppressing a decrease in reliability even when performing recognition processing using a partial region of image data. [Means for solving the problem]

[0006] In order to solve the above problems, the present disclosure provides a readout unit that sets a readout unit as a part of a pixel area in which a plurality of pixels are arranged in a two-dimensional array, and controls readout of pixel signals from pixels included in the pixel area; a reliability calculation unit that calculates the reliability of a predetermined region within the pixel region based on at least one of an area, a readout count, a dynamic range, and exposure information of the region of the captured image that is set as the readout unit and read out; An information processing device is provided.

[0007] the reliability calculation unit calculates a correction value of the reliability for each of the plurality of pixels based on at least one of an area of ​​a region of a captured image, a number of times read out, a dynamic range, and exposure information, and generates a reliability map in which the correction values ​​are arranged in a two-dimensional array; It may further include:

[0008] The reliability calculation unit includes a correction unit that corrects the reliability based on a correction value of the reliability, It may further include:

[0009] The correction unit may correct the reliability in accordance with a representative value of the correction values ​​based on the predetermined region.

[0010] The reading section may read out pixels included in the pixel region as line-shaped image data.

[0011] The reading section may read out pixels included in the pixel region as sampled image data in a grid or checkerboard pattern.

[0012] a recognition processing execution unit that recognizes an object within the predetermined area, Further, it may be provided.

[0013] The correction unit may calculate a representative value of the correction value based on a receptive field in which the feature amount within the predetermined region is calculated.

[0014] the reliability map generating unit generates at least two types of reliability maps based on at least two pieces of information selected from the area, the number of readouts, the dynamic range, and the exposure information; a synthesis unit that synthesizes the at least two types of reliability maps, Further, it may be provided.

[0015] The predetermined region within the pixel region may be a region based on at least one of a label and a category associated with each pixel by semantic segmentation.

[0016] In order to solve the above problems, one aspect of the present disclosure provides a sensor unit including a plurality of pixels arranged in a two-dimensional array; An information processing system comprising: The recognition processing unit a readout unit that sets a readout unit as a part of a pixel area of ​​the sensor unit and controls readout of pixel signals from pixels included in the pixel area; a recognition processing unit having a reliability calculation unit that calculates the reliability of a predetermined region within the pixel region based on at least one of an area, a readout count, a dynamic range, and exposure information of the region of the captured image that is set as the readout unit and read out; An information processing system is provided having:

[0017] In order to solve the above problem, one aspect of the present disclosure provides a readout process that sets a readout unit as a part of a pixel area in which a plurality of pixels are arranged in a two-dimensional array, and controls readout of pixel signals from pixels included in the pixel area; a reliability calculation step of calculating the reliability of a predetermined region within the pixel region based on at least one of an area, a readout count, a dynamic range, and exposure information of the region of the captured image set as the readout unit and read out; An information processing method is provided, comprising:

[0018] In order to solve the above problem, one aspect of the present disclosure is a method for detecting a noise generated by a recognition processing unit, the method comprising: a readout step of setting a readout unit as a part of a pixel area in which a plurality of pixels are arranged in a two-dimensional array and controlling the readout of pixel signals from pixels included in the pixel area; a reliability calculation step of calculating the reliability of a predetermined region within the pixel region based on at least one of an area, a readout count, a dynamic range, and exposure information of the region of the captured image set as the readout unit and read out; A program for causing a computer to execute the above is provided. [Brief explanation of the drawings]

[0019] [Figure 1] FIG. 1 is a block diagram showing a configuration of an example of an imaging device applicable to each embodiment of the present disclosure. [Figure 2A] FIG. 2 is a schematic diagram showing an example of the hardware configuration of an imaging apparatus according to each embodiment. [Figure 2B] FIG. 2 is a schematic diagram showing an example of the hardware configuration of an imaging apparatus according to each embodiment. [Figure 3A] FIG. 10 is a diagram showing an example in which the imaging device according to each embodiment is formed using a stacked CIS with a two-layer structure. [Figure 3B] FIG. 10 is a diagram showing an example in which the imaging device according to each embodiment is formed using a stacked CIS with a three-layer structure. [Figure 4] FIG. 2 is a block diagram showing a configuration of an example of a sensor unit applicable to each embodiment. [Figure 5A] FIG. 1 is a schematic diagram illustrating a rolling shutter system. [Figure 5B] FIG. 1 is a schematic diagram illustrating a rolling shutter system. [Figure 5C] FIG. 1 is a schematic diagram illustrating a rolling shutter system. [Figure 6A] FIG. 10 is a schematic diagram for explaining line thinning in the rolling shutter method. [Figure 6B] FIG. 10 is a schematic diagram for explaining line thinning in the rolling shutter method. [Figure 6C]FIG. 10 is a schematic diagram for explaining line thinning in the rolling shutter method. [Figure 7A] FIG. 10 is a diagram schematically illustrating an example of another imaging method in the rolling shutter system. [Figure 7B] FIG. 10 is a diagram schematically illustrating an example of another imaging method in the rolling shutter system. [Figure 8A] FIG. 1 is a schematic diagram for explaining a global shutter system. [Figure 8B] FIG. 1 is a schematic diagram for explaining a global shutter system. [Figure 8C] FIG. 1 is a schematic diagram for explaining a global shutter system. [Figure 9A] FIG. 10 is a diagram schematically showing an example of a sampling pattern that can be realized in the global shutter system. [Figure 9B] FIG. 10 is a diagram schematically showing an example of a sampling pattern that can be realized in the global shutter system. [Figure 10] FIG. 1 is a diagram for explaining the outline of image recognition processing using CNN. [Figure 11] FIG. 2 is a diagram for explaining an outline of image recognition processing for obtaining a recognition result from a part of an image to be recognized. [Figure 12A] FIG. 1 is a diagram illustrating an example of a classification process using a DNN when time-series information is not used. [Figure 12B] FIG. 1 is a diagram illustrating an example of a classification process using a DNN when time-series information is not used. [Figure 13A] FIG. 1 is a diagram schematically illustrating a first example of a classification process using a DNN when time-series information is used. [Figure 13B] FIG. 1 is a diagram schematically illustrating a first example of a classification process using a DNN when time-series information is used. [Figure 14A] FIG. 10 is a diagram schematically illustrating a second example of a classification process using a DNN when time-series information is used. [Figure 14B] FIG. 10 is a diagram schematically illustrating a second example of a classification process using a DNN when time-series information is used. [Figure 15A] 5A and 5B are diagrams for explaining the relationship between the frame drive speed and the readout amount of pixel signals. [Figure 15B] 5A and 5B are diagrams for explaining the relationship between the frame drive speed and the readout amount of pixel signals. [Figure 16] 1A and 1B are schematic diagrams for explaining a recognition process according to each embodiment of the present disclosure. [Figure 17] FIG. 2 is a functional block diagram illustrating an example of functions of a control unit and a recognition processing unit. [Figure 18A] FIG. 4 is a block diagram showing the configuration of a reliability map generation unit. [Figure 18B] FIG. 10 is a diagram showing a schematic diagram illustrating that the number of times line data is read out varies depending on the section (time) for integration. [Figure 18C] 10A and 10B are diagrams showing examples in which the read position of line data is adaptively changed in accordance with the recognition result of the recognition processing execution unit. [Figure 19] FIG. 4 is a schematic diagram showing in more detail an example of processing in a recognition processing unit. [Figure 20] FIG. 4 is a schematic diagram for explaining a readout process of a readout unit. [Figure 21] FIG. 10 is a diagram showing an area that has been read out line by line and an area that has not been read out. [Figure 22] 10A and 10B are diagrams showing an area that has been read out line by line from the left end to the right end and an area that has not been read out; [Figure 23] FIG. 10 is a diagram schematically illustrating an example of reading out line by line from the left end side to the right end side. [Figure 24] FIG. 10 is a diagram schematically showing reliability map values ​​when the readout area changes within the recognition region. [Figure 25] FIG. 10 is a diagram schematically showing an example in which the read range of line data is limited. [Figure 26] FIG. 1 is a diagram showing an example of discrimination processing (recognition processing) using a DNN when time-series information is not used. [Figure 27A] FIG. 10 is a diagram showing an example of subsampling one image into a grid pattern. [Figure 27B] A diagram showing an example of checkerboard subsampling of an image. [Figure 28]FIG. 10 is a diagram illustrating a case where a reliability map is used in a transportation system. [Figure 29] 10 is a flowchart showing the flow of processing by a reliability calculation unit. [Figure 30] Schematic diagram showing the relationship between features and receptive fields. [Figure 31] A schematic diagram of the recognition area and receptive field. [Figure 32] FIG. 10 is a diagram schematically showing the contribution to the feature amount within the recognition region. [Figure 33] A schematic diagram of an image undergoing recognition processing using general semantic segmentation. [Figure 34] FIG. 10 is a block diagram of a reliability map generation unit according to the second embodiment. [Figure 35] FIG. 4 is a diagram schematically showing the relationship between a recognition area and line data. [Figure 36] FIG. 11 is a block diagram of a reliability map generation unit according to the third embodiment. [Figure 37] FIG. 10 is a diagram schematically showing the relationship between line data and exposure frequency. [Figure 38] FIG. 13 is a block diagram of a reliability map generation unit according to the fourth embodiment. [Figure 39] FIG. 10 is a diagram schematically showing the relationship between line data and the dynamic range. [Figure 40] FIG. 13 is a block diagram of a reliability map generation unit according to the fifth embodiment. [Figure 41] 1A to 1C are diagrams illustrating examples of using an information processing device according to a first embodiment and each of its modifications to a fifth embodiment. [Figure 42] 1 is a block diagram showing an example of a schematic configuration of a vehicle control system; [Figure 43] FIG. 2 is an explanatory diagram showing an example of the installation positions of an outside-vehicle information detection unit and an imaging unit. DETAILED DESCRIPTION OF THE INVENTION

[0020] Hereinafter, embodiments of an information processing device, an information processing system, an information processing method, and an information processing program will be described with reference to the drawings. The following description will focus on the main components of the information processing device, the information processing system, the information processing method, and the information processing program, but the information processing device, the information processing system, the information processing method, and the information processing program may include components and functions that are not shown or described. The following description does not exclude components and functions that are not shown or described.

[0021] 1. Configuration Examples According to Each Embodiment of the Present Disclosure An example of the overall configuration of an information processing system according to each embodiment will be described briefly. FIG. 1 is a block diagram showing an example of the configuration of an information processing system 1. In FIG. 1, the information processing system 1 includes a sensor unit 10, a sensor control unit 11, a recognition processing unit 12, a memory 13, a visual recognition processing unit 14, and an output control unit 15. These units are, for example, a CMOS image sensor (CIS) integrally formed using a CMOS (Complementary Metal Oxide Semiconductor). Note that the information processing system 1 is not limited to this example and may be another type of optical sensor, such as an infrared light sensor that captures images using infrared light. The sensor control unit 11, the recognition processing unit 12, the memory 13, the visual recognition processing unit 14, and the output control unit 15 constitute an information processing device 2.

[0022] The sensor unit 10 outputs pixel signals corresponding to light irradiated onto the light-receiving surface via the optical unit 30. More specifically, the sensor unit 10 has a pixel array in which pixels, each including at least one photoelectric conversion element, are arranged in a matrix. The pixels arranged in a matrix in the pixel array form a light-receiving surface. The sensor unit 10 further includes a drive circuit for driving each pixel included in the pixel array, and a signal processing circuit for performing predetermined signal processing on signals read from each pixel and outputting the processed signals as pixel signals for each pixel. The sensor unit 10 outputs the pixel signals of each pixel included in the pixel area as digital image data.

[0023] Hereinafter, the region in the pixel array of the sensor unit 10 where pixels effective for generating pixel signals are arranged will be referred to as a frame. Frame image data is formed from pixel data based on pixel signals output from each pixel included in the frame. Each row in the pixel array of the sensor unit 10 will be referred to as a line, and line image data will be formed from pixel data based on pixel signals output from each pixel included in the line. Furthermore, the operation of the sensor unit 10 to output pixel signals in response to light irradiated onto the light-receiving surface will be referred to as imaging. The sensor unit 10 controls the exposure during imaging and the gain (analog gain) for the pixel signals according to imaging control signals supplied from the sensor control unit 11, which will be described later.

[0024] The sensor control unit 11 is configured by, for example, a microprocessor, controls the reading of pixel data from the sensor unit 10, and outputs pixel data based on each pixel signal read from each pixel included in a frame. The pixel data output from the sensor control unit 11 is supplied to the recognition processing unit 12 and the visual recognition processing unit 14.

[0025] The sensor control unit 11 also generates an imaging control signal for controlling imaging in the sensor unit 10. The sensor control unit 11 generates the imaging control signal, for example, in accordance with instructions from the recognition processing unit 12 and the visual recognition processing unit 14, which will be described later. The imaging control signal includes information indicating the exposure and analog gain used when imaging in the sensor unit 10, as described above. The imaging control signal further includes control signals (such as a vertical synchronization signal and a horizontal synchronization signal) used by the sensor unit 10 to perform imaging operations. The sensor control unit 11 supplies the generated imaging control signal to the sensor unit 10.

[0026] The optical unit 30 is used to irradiate the light receiving surface of the sensor unit 10 with light from the subject, and is disposed, for example, at a position corresponding to the sensor unit 10. The optical unit 30 includes, for example, a plurality of lenses, an aperture mechanism for adjusting the size of an aperture for incident light, and a focus mechanism for adjusting the focus of light irradiated onto the light receiving surface. The optical unit 30 may further include a shutter mechanism (mechanical shutter) for adjusting the time for which light is irradiated onto the light receiving surface. The aperture mechanism, focus mechanism, and shutter mechanism of the optical unit 30 can be controlled, for example, by the sensor control unit 11. Alternatively, the aperture and focus of the optical unit 30 can be controlled from outside the information processing system 1. The optical unit 30 can also be configured integrally with the information processing system 1.

[0027] The recognition processing unit 12 performs a recognition process of an object included in an image using pixel data based on the pixel data supplied from the sensor control unit 11. In the present disclosure, for example, a DSP (Digital Signal Processor) reads and executes a program that has been learned in advance using teacher data and stored as a learning model in the memory 13, thereby configuring the recognition processing unit 12 as a machine learning unit that performs recognition processing using a DNN (Deep Neural Network). The recognition processing unit 12 can instruct the sensor control unit 11 to read pixel data required for the recognition processing from the sensor unit 10. The recognition result by the recognition processing unit 12 is supplied to the output control unit 15.

[0028] The visual recognition processing unit 14 processes the pixel data supplied from the sensor control unit 11 to obtain an image suitable for human visual recognition, and outputs image data consisting of, for example, a group of pixel data. For example, the visual recognition processing unit 14 is configured by an ISP (Image Signal Processor) reading and executing a program stored in advance in a memory (not shown).

[0029] For example, when a color filter is provided for each pixel included in the sensor unit 10 and the pixel data has color information of R (red), G (green), and B (blue), the visual recognition processing unit 14 can perform demosaic processing, white balance processing, etc. Furthermore, the visual recognition processing unit 14 can instruct the sensor control unit 11 to read out pixel data required for the visual recognition processing from the sensor unit 10. The image data obtained by image processing the pixel data by the visual recognition processing unit 14 is supplied to the output control unit 15.

[0030] The output control unit 15 is configured by, for example, a microprocessor, and outputs one or both of the recognition result supplied from the recognition processing unit 12 and the image data supplied from the visual recognition processing unit 14 as a visual recognition processing result to the outside of the information processing system 1. The output control unit 15 can output the image data to, for example, a display unit 31 having a display device. This allows the user to visually recognize the image data displayed by the display unit 31. The display unit 31 may be built into the information processing system 1 or may be configured external to the information processing system 1.

[0031] 2A and 2B are schematic diagrams showing examples of the hardware configuration of an information processing system 1 according to each embodiment. Fig. 2A shows an example in which a single chip 2 is equipped with a sensor unit 10, a sensor control unit 11, a recognition processing unit 12, a memory 13, a visual recognition processing unit 14, and an output control unit 15 from the configuration shown in Fig. 1. Note that in Fig. 2A, the memory 13 and the output control unit 15 are omitted to avoid complexity.

[0032] 2A, the recognition result by the recognition processing unit 12 is output to the outside of the chip 2 via an output control unit 15 (not shown). Also, in the configuration of FIG. 2A, the recognition processing unit 12 can obtain pixel data to be used for recognition from the sensor control unit 11 via an interface inside the chip 2.

[0033] 2B shows an example in which the sensor unit 10, sensor control unit 11, visual recognition processing unit 14, and output control unit 15 of the configuration shown in Fig. 1 are mounted on one chip 2, and the recognition processing unit 12 and memory 13 (not shown) are placed outside the chip 2. In Fig. 2B, as in Fig. 2A described above, the memory 13 and output control unit 15 are omitted to avoid complexity.

[0034] In the configuration of Fig. 2B, the recognition processing unit 12 acquires pixel data to be used for recognition via an interface for communication between chips. Also, in Fig. 2B, the recognition result by the recognition processing unit 12 is shown as being directly output from the recognition processing unit 12 to the outside, but this is not limited to this example. That is, in the configuration of Fig. 2B, the recognition processing unit 12 may return the recognition result to the chip 2 and output it from an output control unit 15 (not shown) mounted on the chip 2.

[0035] In the configuration shown in FIG. 2A, the recognition processing unit 12 is mounted on the chip 2 together with the sensor control unit 11, and communication between the recognition processing unit 12 and the sensor control unit 11 can be performed at high speed via an interface inside the chip 2. On the other hand, in the configuration shown in FIG. 2A, the recognition processing unit 12 cannot be replaced, making it difficult to change the recognition processing. In contrast, in the configuration shown in FIG. 2B, the recognition processing unit 12 is provided outside the chip 2, so communication between the recognition processing unit 12 and the sensor control unit 11 must be performed via an interface between the chips. Therefore, communication between the recognition processing unit 12 and the sensor control unit 11 is slower than in the configuration of FIG. 2A, and there is a possibility of delays in control. On the other hand, the recognition processing unit 12 can be easily replaced, making it possible to realize a variety of recognition processes.

[0036] Unless otherwise specified, the information processing system 1 will be assumed to have the configuration shown in Figure 2A, in which a sensor unit 10, a sensor control unit 11, a recognition processing unit 12, a memory 13, a visual recognition processing unit 14, and an output control unit 15 are mounted on one chip 2.

[0037] 2A, the information processing system 1 can be formed on a single substrate. However, the information processing system 1 may be a stacked CIS in which multiple semiconductor chips are stacked and integrally formed.

[0038] As an example, the information processing system 1 may be formed with a two-layer structure in which semiconductor chips are stacked in two layers. FIG. 3A illustrates an example in which the information processing system 1 according to each embodiment is formed using a two-layer stacked CIS. In the structure of FIG. 3A, a pixel unit 20a is formed on a first-layer semiconductor chip, and a memory and logic unit 20b is formed on a second-layer semiconductor chip. The pixel unit 20a includes at least a pixel array in the sensor unit 10. The memory and logic unit 20b includes, for example, a sensor control unit 11, a recognition processing unit 12, a memory 13, a visual recognition processing unit 14, and an output control unit 15, as well as an interface for communicating between the information processing system 1 and the outside. The memory and logic unit 20b further includes a part or all of a drive circuit for driving the pixel array in the sensor unit 10. Although not illustrated, the memory and logic unit 20b may further include, for example, a memory used by the visual recognition processing unit 14 for processing image data.

[0039] As shown on the right side of FIG. 3A, the information processing system 1 is configured as one solid-state imaging device by bonding the semiconductor chip of the first layer and the semiconductor chip of the second layer together while making electrical contact with each other.

[0040] As another example, the information processing system 1 can be formed with a three-layer structure in which semiconductor chips are stacked in three layers. FIG. 3B is a diagram showing an example in which the information processing system 1 according to each embodiment is formed using a three-layer stacked CIS. In the structure of FIG. 3B, a pixel unit 20a is formed in a first-layer semiconductor chip, a memory unit 20c is formed in a second-layer semiconductor chip, and a logic unit 20b is formed in a third-layer semiconductor chip. In this case, the logic unit 20b includes, for example, a sensor control unit 11, a recognition processing unit 12, a visual recognition processing unit 14, an output control unit 15, and an interface for communicating between the information processing system 1 and the outside. The memory unit 20c can also include a memory 13 and a memory used by the visual recognition processing unit 14 to process image data. The memory 13 may be included in the logic unit 20b.

[0041] As shown on the right side of Figure 3B, the information processing system 1 is configured as a single solid-state imaging element by bonding the first layer semiconductor chip, the second layer semiconductor chip, and the third layer semiconductor chip together while maintaining electrical contact.

[0042] Fig. 4 is a block diagram showing an example of the configuration of a sensor unit 10 applicable to each embodiment. In Fig. 4, the sensor unit 10 includes a pixel array unit 101, a vertical scanning unit 102, an AD (Analog to Digital) conversion unit 103, pixel signal lines 106, vertical signal lines VSL, a control unit 1100, and a signal processing unit 1101. Note that in Fig. 4, the control unit 1100 and the signal processing unit 1101 may also be included in, for example, the sensor control unit 11 shown in Fig. 1.

[0043] The pixel array unit 101 includes a plurality of pixel circuits 100, each of which includes a photoelectric conversion element, such as a photodiode, that performs photoelectric conversion on received light, and a circuit that reads out electric charges from the photoelectric conversion element. In the pixel array unit 101, the plurality of pixel circuits 100 are arranged in a matrix array in the horizontal direction (row direction) and the vertical direction (column direction). In the pixel array unit 101, the row direction arrangement of the pixel circuits 100 is called a line. For example, if one frame of image is formed with 1920 pixels x 1080 lines, the pixel array unit 101 includes at least 1080 lines, each including at least 1920 pixel circuits 100. One frame of image (image data) is formed by pixel signals read out from the pixel circuits 100 included in the frame.

[0044] Hereinafter, the operation of reading pixel signals from each pixel circuit 100 included in a frame in the sensor unit 10 will be appropriately described as "reading pixels from a frame," etc. Also, the operation of reading pixel signals from each pixel circuit 100 of a line included in a frame will be appropriately described as "reading a line," etc.

[0045] Furthermore, pixel signal lines 106 are connected to the pixel array unit 101 for each row and column of each pixel circuit 100, and vertical signal lines VSL are connected to each column. The ends of the pixel signal lines 106 that are not connected to the pixel array unit 101 are connected to the vertical scanning unit 102. Under the control of a control unit 1100 (described later), the vertical scanning unit 102 transmits control signals such as drive pulses used to read pixel signals from pixels to the pixel array unit 101 via the pixel signal lines 106. The ends of the vertical signal lines VSL that are not connected to the pixel array unit 101 are connected to an AD conversion unit 103. The pixel signals read from the pixels are transmitted to the AD conversion unit 103 via the vertical signal lines VSL.

[0046] The following provides an overview of the control of reading out pixel signals from the pixel circuit 100. Reading out pixel signals from the pixel circuit 100 is performed by transferring charges accumulated in a photoelectric conversion element upon exposure to a floating diffusion layer (FD) and converting the transferred charges into a voltage in the floating diffusion layer. The voltage into which the charges are converted in the floating diffusion layer is output to a vertical signal line VSL via an amplifier.

[0047] More specifically, in the pixel circuit 100, during exposure, the connection between the photoelectric conversion element and the floating diffusion layer is turned off (open), and charges generated in response to incident light through photoelectric conversion are accumulated in the photoelectric conversion element. After exposure is completed, the floating diffusion layer is connected to the vertical signal line VSL in response to a selection signal supplied via the pixel signal line 106. Furthermore, the floating diffusion layer is connected to a power supply voltage VDD or a black level voltage supply line for a short period in response to a reset pulse supplied via the pixel signal line 106, thereby resetting the floating diffusion layer. A voltage (referred to as voltage A) at the reset level of the floating diffusion layer is output to the vertical signal line VSL. Thereafter, a transfer pulse supplied via the pixel signal line 106 turns the connection between the photoelectric conversion element and the floating diffusion layer on (closed), and the charges accumulated in the photoelectric conversion element are transferred to the floating diffusion layer. A voltage (referred to as voltage B) corresponding to the amount of charge in the floating diffusion layer is output to the vertical signal line VSL.

[0048] The AD conversion unit 103 includes an AD converter 107 provided for each vertical signal line VSL, a reference signal generation unit 104, and a horizontal scanning unit 105. The AD converter 107 is a column AD converter that performs AD conversion processing on each column of the pixel array unit 101. The AD converter 107 performs AD conversion processing on pixel signals supplied from the pixel circuits 100 via the vertical signal lines VSL, and generates two digital values ​​(values ​​corresponding to voltage A and voltage B, respectively) for correlated double sampling (CDS) processing that reduces noise.

[0049] The AD converter 107 supplies the two generated digital values ​​to the signal processing unit 1101. The signal processing unit 1101 performs CDS processing based on the two digital values ​​supplied from the AD converter 107, and generates a pixel signal (pixel data) as a digital signal. The pixel data generated by the signal processing unit 1101 is output to the outside of the sensor unit 10.

[0050] The reference signal generation unit 104 generates, as a reference signal, a ramp signal used by each AD converter 107 to convert a pixel signal into two digital values, based on a control signal input from the control unit 1100. A ramp signal is a signal whose level (voltage value) decreases at a constant slope over time, or a signal whose level decreases in a step-like manner. The reference signal generation unit 104 supplies the generated ramp signal to each AD converter 107. The reference signal generation unit 104 is configured using, for example, a DAC (Digital to Analog Converter) or the like.

[0051] When a ramp signal, whose voltage drops stepwise according to a predetermined slope, is supplied from the reference signal generator 104, the counter starts counting in accordance with the clock signal. The comparator compares the voltage of the pixel signal supplied from the vertical signal line VSL with the voltage of the ramp signal, and stops counting by the counter when the voltage of the ramp signal crosses the voltage of the pixel signal. The AD converter 107 converts the analog pixel signal into a digital value by outputting a value corresponding to the count value at the time when the counting was stopped.

[0052] The AD converter 107 supplies the two generated digital values ​​to the signal processing unit 1101. The signal processing unit 1101 performs CDS processing based on the two digital values ​​supplied from the AD converter 107, and generates a pixel signal (pixel data) based on the digital signal. The pixel signal based on the digital signal generated by the signal processing unit 1101 is output to the outside of the sensor unit 10.

[0053] Under the control of the control unit 1100, the horizontal scanning unit 105 performs selective scanning to select each AD converter 107 in a predetermined order, thereby causing each AD converter 107 to sequentially output each digital value temporarily held therein to the signal processing unit 1101. The horizontal scanning unit 105 is configured using, for example, a shift register, an address decoder, etc.

[0054] The control unit 1100 controls the driving of the vertical scanning unit 102, the AD conversion unit 103, the reference signal generation unit 104, the horizontal scanning unit 105, etc. in accordance with the imaging control signal supplied from the sensor control unit 11. The control unit 1100 generates various driving signals that serve as references for the operations of the vertical scanning unit 102, the AD conversion unit 103, the reference signal generation unit 104, and the horizontal scanning unit 105. The control unit 1100 generates control signals that the vertical scanning unit 102 supplies to each pixel circuit 100 via the pixel signal line 106, based on, for example, a vertical synchronization signal or an external trigger signal included in the imaging control signal and a horizontal synchronization signal. The control unit 1100 supplies the generated control signals to the vertical scanning unit 102.

[0055] Furthermore, the control unit 1100 outputs, for example, information indicating an analog gain, which is included in an imaging control signal supplied from the sensor control unit 11, to the AD conversion unit 103. The AD conversion unit 103 controls the gain of a pixel signal input to each AD converter 107 included in the AD conversion unit 103 via a vertical signal line VSL, according to the information indicating the analog gain.

[0056] Based on a control signal supplied from the control unit 1100, the vertical scanning unit 102 supplies various signals including drive pulses to the pixel signal lines 106 of a selected pixel row of the pixel array unit 101, to each pixel circuit 100 for each line, and causes each pixel circuit 100 to output a pixel signal to a vertical signal line VSL. The vertical scanning unit 102 is configured using, for example, a shift register, an address decoder, etc. Furthermore, the vertical scanning unit 102 controls the exposure of each pixel circuit 100 in accordance with information indicating exposure supplied from the control unit 1100.

[0057] The sensor unit 10 configured in this manner is a column AD type CMOS (Complementary Metal Oxide Semiconductor) image sensor in which AD converters 107 are arranged for each column.

[0058] [2. Examples of existing technologies applicable to the present disclosure] Prior to describing each embodiment of the present disclosure, a brief description of existing technologies applicable to the present disclosure will be given to facilitate understanding.

[0059] (2-1. Overview of Rolling Shutter) Known imaging methods for capturing images using the pixel array unit 101 include the rolling shutter (RS) method and the global shutter (GS) method. First, the rolling shutter method will be briefly described. Figures 5A, 5B, and 5C are schematic diagrams for explaining the rolling shutter method. In the rolling shutter method, as shown in Figure 5A, images are captured line by line in sequence, starting from, for example, line 201 at the top of a frame 200.

[0060] In the above description, "imaging" refers to the operation of the sensor unit 10 to output a pixel signal in response to light irradiated onto the light-receiving surface. More specifically, "imaging" refers to a series of operations from exposing a pixel to transferring a pixel signal based on charges accumulated in a photoelectric conversion element included in the pixel by the exposure to the sensor control unit 11. Also, as described above, a frame refers to an area in the pixel array unit 101 where pixel circuits 100 effective for generating pixel signals are arranged.

[0061] For example, in the configuration of Fig. 4, exposure is performed simultaneously in each pixel circuit 100 included in one line. After exposure is completed, pixel signals based on charges accumulated by exposure are simultaneously transferred in each pixel circuit 100 included in that line via each vertical signal line VSL corresponding to each pixel circuit 100. By performing this operation sequentially line by line, imaging using a rolling shutter can be achieved.

[0062] FIG. 5B schematically illustrates an example of the relationship between imaging and time in the rolling shutter method. In FIG. 5B, the vertical axis represents line position, and the horizontal axis represents time. In the rolling shutter method, exposure for each line is performed line-by-line sequentially, so as shown in FIG. 5B, the exposure timing for each line is shifted sequentially according to the line position. Therefore, for example, if the horizontal positional relationship between the information processing system 1 and the subject changes rapidly, distortion occurs in the captured image of frame 200, as illustrated in FIG. 5C. In the example of FIG. 5C, image 202 corresponding to frame 200 is tilted at an angle corresponding to the speed and direction of change in the horizontal positional relationship between the information processing system 1 and the subject.

[0063] In the rolling shutter method, it is also possible to thin out lines when capturing an image. Figures 6A, 6B, and 6C are schematic diagrams for explaining line thinning in the rolling shutter method. As shown in Figure 6A, similar to the example of Figure 5A described above, capturing is performed line by line from line 201 at the top of frame 200 toward the bottom of frame 200. At this time, capturing is performed while skipping every predetermined number of lines.

[0064] For the sake of explanation, let us assume that every other line is imaged by thinning out one line. That is, after imaging the nth line, imaging of the (n+2)th line is performed. In this case, the time from imaging the nth line to imaging the (n+2)th line is assumed to be equal to the time from imaging the nth line to imaging the (n+1)th line when thinning out is not performed.

[0065] FIG. 6B schematically illustrates an example of the relationship between imaging and time when one line is thinned out using the rolling shutter method. In FIG. 6B, the vertical axis represents line position, and the horizontal axis represents time. In FIG. 6B, exposure A corresponds to the exposure in FIG. 5B without thinning out, and exposure B represents the exposure when one line is thinned out. As shown in exposure B, by thinning out the lines, the difference in exposure timing at the same line position can be reduced compared to when line thinning is not performed. Therefore, as illustrated by image 203 in FIG. 6C, the tilt distortion generated in the image of captured frame 200 is smaller than when line thinning is not performed as shown in FIG. 5C. On the other hand, when line thinning is performed, the image resolution is lower than when line thinning is not performed.

[0066] Although the above description has been given of an example in which imaging is performed line-sequentially from the top to the bottom of the frame 200 using the rolling shutter method, this is not limiting. Figures 7A and 7B are diagrams schematically illustrating examples of other imaging methods using the rolling shutter method. For example, as shown in Figure 7A, imaging can be performed line-sequentially from the bottom to the top of the frame 200 using the rolling shutter method. In this case, the horizontal direction of distortion in the image 202 is reversed compared to when imaging is performed line-sequentially from the top to the bottom of the frame 200.

[0067] Also, for example, by setting the range of the vertical signal line VSL that transfers pixel signals, it is possible to selectively read out a portion of the line. Furthermore, by respectively setting the line where imaging is performed and the vertical signal line VSL that transfers pixel signals, it is possible to set the lines where imaging starts and ends other than the top and bottom ends of the frame 200. Fig. 7B schematically shows an example in which the imaging range is a rectangular region 205 whose width and height are less than the width and height of the frame 200. In the example of Fig. 7B, imaging is performed line by line from line 204 at the top end of the region 205 toward the bottom end of the region 205.

[0068] (2-2. Global Shutter Overview) Next, a global shutter (GS) system will be briefly described as an imaging system used when capturing an image using the pixel array unit 101. Figures 8A, 8B, and 8C are schematic diagrams for explaining the global shutter system. In the global shutter system, as shown in Figure 8A, all pixel circuits 100 included in a frame 200 are exposed simultaneously.

[0069] 4, one possible configuration is to further provide a capacitor between the photoelectric conversion element and the FD in each pixel circuit 100. A first switch is provided between the photoelectric conversion element and the capacitor, and a second switch is provided between the capacitor and the floating diffusion layer, and the opening and closing of each of the first and second switches is controlled by a pulse supplied via the pixel signal line 106.

[0070] With this configuration, during the exposure period, the first and second switches are opened in all pixel circuits 100 included in frame 200, and when exposure ends, the first switch is switched from open to closed to transfer charge from the photoelectric conversion element to the capacitor. Thereafter, the capacitor is regarded as a photoelectric conversion element, and charge is read out from the capacitor in a sequence similar to the readout operation described for the rolling shutter method. This enables simultaneous exposure in all pixel circuits 100 included in frame 200.

[0071] FIG. 8B schematically shows an example of the relationship between imaging and time in the global shutter system. In FIG. 8B, the vertical axis represents line position, and the horizontal axis represents time. In the global shutter system, exposure is performed simultaneously in all pixel circuits 100 included in frame 200, so the exposure timing for each line can be made the same, as shown in FIG. 8B. Therefore, even if, for example, the horizontal positional relationship between information processing system 1 and a subject changes rapidly, no distortion corresponding to the change occurs in image 206 of captured frame 200, as exemplified in FIG. 8C.

[0072] The global shutter method ensures simultaneous exposure timing for all pixel circuits 100 included in the frame 200. Therefore, by controlling the timing of each pulse supplied by the pixel signal line 106 of each line and the timing of transfer by each vertical signal line VSL, sampling (reading of pixel signals) in various patterns can be realized.

[0073] 9A and 9B are diagrams schematically showing examples of sampling patterns that can be achieved with the global shutter system. Fig. 9A shows an example in which samples 208 for reading out pixel signals are extracted in a checkerboard pattern from each pixel circuit 100 arranged in a matrix and included in a frame 200. Fig. 9B shows an example in which samples 208 for reading out pixel signals are extracted in a grid pattern from each pixel circuit 100. Similarly to the rolling shutter system described above, the global shutter system also allows for line-sequential imaging.

[0074] (2-3. About DNN) Next, a recognition process using a DNN (Deep Neural Network) applicable to each embodiment will be briefly described. In each embodiment, a recognition process for image data is performed using a CNN (Convolutional Neural Network) and an RNN (Recurrent Neural Network) among DNNs. Hereinafter, "recognition process for image data" will be referred to as "image recognition process" or the like as appropriate.

[0075] (2-3-1. Overview of CNN) First, a brief description of CNN will be given. Image recognition processing using CNN generally involves performing image recognition processing based on image information consisting of pixels arranged in a matrix, for example. FIG. 10 is a diagram for explaining an outline of image recognition processing using CNN. Processing is performed by a CNN 52 that has been trained in a predetermined manner on the entire pixel information 51 of an image 50' depicting the number "8," which is an object to be recognized. As a result, the number "8" is recognized as a recognition result 53.

[0076] Alternatively, it is possible to perform CNN processing based on an image for each line, and obtain a recognition result from a portion of the image to be recognized. FIG. 11 is a diagram for explaining an outline of image recognition processing that obtains a recognition result from a portion of the image to be recognized. In FIG. 11, an image 50' is a partial line-by-line acquisition of the number "8," which is the object to be recognized. For example, pixel information 54a, 54b, and 54c for each line that forms pixel information 51' of this image 50' is sequentially processed by a CNN 52' that has been trained in a predetermined manner.

[0077] For example, the recognition result 53a obtained by the recognition process of the pixel information 54a of the first line by the CNN 52' is not a valid recognition result. Here, a valid recognition result refers to, for example, a recognition result having a score indicating the reliability of the recognition result equal to or greater than a predetermined value. Note that the reliability in this embodiment refers to an evaluation value that indicates how much the recognition result [T] output by the DNN can be trusted. For example, the reliability ranges from 0.0 to 1.0, and the closer the value is to 1.0, the fewer other competing candidates with scores similar to the recognition result [T]. On the other hand, the closer the value is to 0, the more other competing candidates with scores similar to the recognition result [T] have appeared.

[0078] The CNN 52' updates 55 its internal state based on the recognition result 53a. Next, the CNN 52', whose internal state has been updated 55 based on the previous recognition result 53a, performs recognition processing on the pixel information 54b of the second line. In FIG. 11, as a result, a recognition result 53b is obtained indicating that the digit to be recognized is either "8" or "9." Furthermore, the CNN 52' updates 55 its internal information based on the recognition result 53b. Next, the CNN 52', whose internal state has been updated 55 based on the previous recognition result 53b, performs recognition processing on the pixel information 54c of the third line. In FIG. 11, as a result, the digit to be recognized is narrowed down to "8" out of "8" and "9."

[0079] Here, the recognition process shown in Fig. 11 updates the internal state of the CNN using the result of the previous recognition process, and the CNN with this updated internal state performs recognition processing using pixel information of lines adjacent to the line on which the previous recognition process was performed. In other words, the recognition process shown in Fig. 11 is performed line by line on the image while updating the internal state of the CNN based on the previous recognition result. Therefore, the recognition process shown in Fig. 11 is a process that is performed recursively line by line, and can be considered to have a structure equivalent to an RNN.

[0080] (2-3-2. Overview of RNN) Next, an outline of the RNN will be explained. Figures 12A and 12B are diagrams that schematically show an example of classification processing (recognition processing) by the DNN when time-series information is not used. In this case, as shown in Figure 12A, one image is input to the DNN. In the DNN, classification processing is performed on the input image, and a classification result is output.

[0081] Fig. 12B is a diagram for explaining the processing of Fig. 12A in more detail. As shown in Fig. 12B, the DNN executes a feature extraction process and a classification process. In the DNN, feature amounts are extracted from an input image by the feature extraction process. In addition, in the DNN, classification process is executed on the extracted feature amounts to obtain a classification result.

[0082] 13A and 13B are diagrams schematically illustrating a first example of a classification process using a DNN when time-series information is used. In the examples of FIGS. 13A and 13B, a fixed number of past pieces of time-series information are used to perform the classification process using a DNN. In the example of FIG. 13A, an image [T] at time T, an image [T-1] at time T-1 before time T, and an image [T-2] at time T-2 before time T-1 are input to the DNN. The DNN performs classification processing on the input images [T], [T-1], and [T-2], and obtains a classification result [T] at time T. A reliability is assigned to the classification result [T].

[0083] FIG. 13B is a diagram for explaining the processing of FIG. 13A in more detail. As shown in FIG. 13B, in the DNN, the feature extraction processing described above using FIG. 12B is performed one-to-one for each of the input images [T], [T-1], and [T-2], and feature quantities corresponding to the images [T], [T-1], and [T-2] are extracted. In the DNN, the feature quantities obtained based on these images [T], [T-1], and [T-2] are integrated, and a classification process is performed on the integrated feature quantities to obtain a classification result [T] at time T. A reliability is assigned to the classification result [T].

[0084] The method of Figures 13A and 13B requires multiple configurations for feature extraction, and the number of configurations required for feature extraction depends on the number of past images available, which could result in the DNN configuration becoming large-scale.

[0085] 14A and 14B are diagrams schematically illustrating a second example of a classification process using a DNN when time-series information is used. In the example of Fig. 14A, an image [T] at time T is input to a DNN whose internal state has been updated to the state at time T-1, and a classification result [T] at time T is obtained. A reliability is assigned to the classification result [T].

[0086] FIG. 14B is a diagram for explaining the processing of FIG. 14A in more detail. As shown in FIG. 14B, in the DNN, the feature extraction processing described above with reference to FIG. 12B is performed on the input image [T] at time T, and features corresponding to the image [T] are extracted. In the DNN, the internal state is updated by the image before time T, and features related to the updated internal state are stored. The features related to this stored internal information are integrated with the features in the image [T], and a classification process is performed on the integrated features.

[0087] 14A and 14B is a recursive process that is executed using a DNN whose internal state has been updated using, for example, the previous classification result. A DNN that performs recursive processing in this way is called an RNN (Recurrent Neural Network). Classification processing using an RNN is generally used for video image recognition, and it is possible to improve classification accuracy by, for example, sequentially updating the internal state of the DNN using frame images that are updated in a time series.

[0088] In the present disclosure, an RNN is applied to a rolling shutter structure. That is, in the rolling shutter method, pixel signals are read out line-sequentially. Therefore, the pixel signals read out line-sequentially are applied to an RNN as time-series information. This makes it possible to perform classification processing based on multiple lines with a smaller configuration than when a CNN is used (see FIG. 13B). Not limited to this, an RNN can also be applied to a global shutter structure. In this case, for example, adjacent lines can be considered as time-series information.

[0089] (2-4. Drive speed) Next, the relationship between the frame drive speed and the amount of pixel signal readout will be explained using Figures 15A and 15B. Figure 15A is a diagram showing an example of reading out all lines in an image. Here, it is assumed that the resolution of the image to be subjected to recognition processing is 640 pixels horizontally by 480 pixels vertically (480 lines). In this case, driving at a drive speed of 14,400 lines / second enables output at 30 frames per second (fps).

[0090] Next, consider imaging with line thinning. For example, as shown in FIG. 15B, imaging is performed using 1 / 2 thinning readout, in which imaging is performed by skipping one line at a time. As a first example of 1 / 2 thinning, when the drive speed is 14,400 lines / second as described above, the number of lines read from the image is halved, resulting in a decrease in resolution. However, output at 60 fps, double the speed when no thinning is performed, is possible, thereby improving the frame rate. As a second example of 1 / 2 thinning, when the drive speed is set to 7,200 fps, half that of the first example, the frame rate is 30 fps, the same as when no thinning is performed, but power consumption can be reduced.

[0091] When reading out lines of an image, whether to not perform thinning, to perform thinning and increase the drive speed, or to make the drive speed the same as when thinning is not performed can be selected depending on, for example, the purpose of the recognition processing based on the read pixel signals.

[0092] (First embodiment) 16 is a schematic diagram for outlining the recognition process according to this embodiment of the present disclosure. In FIG. 16, in step S1, the information processing system 1 according to this embodiment (see FIG. 1) starts capturing an image of a target to be recognized.

[0093] It is assumed that the target image is, for example, an image of the number "8" drawn by hand. A learning model that has been trained to be able to identify numbers using predetermined training data is pre-stored in the memory 13 as a program, and the recognition processing unit 12 is capable of identifying numbers included in the image by reading and executing this program from the memory 13. Furthermore, it is assumed that the information processing system 1 captures images using a rolling shutter method. Even when the information processing system 1 captures images using a global shutter method, the following process is applicable in the same way as in the case of the rolling shutter method.

[0094] When imaging starts, the information processing system 1 sequentially reads out the frame line by line from the top end side to the bottom end side of the frame in step S2.

[0095] When the lines are read up to a certain position, the recognition processing unit 12 identifies the number "8" or "9" from the image of the read lines (step S3). For example, the numbers "8" and "9" contain a common feature in the upper half, so when the lines are read from the top and the feature is recognized, the recognized object can be identified as either the number "8" or "9."

[0096] Here, as shown in step S4a, by reading up to the bottom line of the frame or a line near the bottom, the entire recognized object appears, and the object identified in step S2 as either the number "8" or "9" is confirmed to be the number "8".

[0097] On the other hand, steps S4b and S4c are processes related to the present disclosure.

[0098] As shown in step S4b, by reading further from the line position read in step S3, it is possible to identify the recognized object as the number "8" even when the bottom end of the number "8" is reached. For example, the bottom half of the number "8" and the bottom half of the number "9" each have different characteristics. By reading the line up to the point where the difference in these characteristics becomes clear, it becomes possible to identify whether the object recognized in step S3 is the number "8" or "9." In the example of FIG. 16, the object is determined to be the number "8" in step S4b.

[0099] Furthermore, as shown in step S4c, it is also possible to jump from the line position of step S3 to a line position where it is possible to determine whether the object identified in step S3 is the number "8" or "9" by further reading in the state of step S3. By reading this jump destination line, it is possible to determine whether the object identified in step S3 is the number "8" or "9." The jump destination line position can be determined based on a learning model that has been trained in advance based on predetermined training data.

[0100] Here, when the object is confirmed in step S4b or step S4c described above, the information processing system 1 can end the recognition process, thereby realizing a reduction in the time and power consumption of the recognition process in the information processing system 1.

[0101] The training data is data that holds multiple combinations of input signals and output signals for each reading unit. As an example, in the above-mentioned task of identifying numbers, data for each reading unit (line data, subsampled data, etc.) can be applied as the input signal, and data indicating the "correct number" can be applied as the output signal. As another example, in a task of detecting an object, data for each reading unit (line data, subsampled data, etc.) can be applied as the input signal, and object class (human body / vehicle / non-object) or object coordinates (x, y, h, w) can be applied as the output signal. Furthermore, the output signal may be generated from only the input signal using self-supervised learning.

[0102] FIG. 17 is a functional block diagram illustrating an example of the functions of the sensor control unit 11 and the recognition processing unit 12 according to this embodiment. 17, the sensor control unit 11 has a readout unit 110. The recognition processing unit 12 has a feature calculation unit 120, a feature accumulation control unit 121, a readout area determination unit 123, a recognition processing execution unit 124, and a reliability calculation unit 125. In addition, the reliability calculation unit 125 has a reliability map generation unit 126 and a score correction unit 127.

[0103] In the sensor control unit 11, the readout unit 110 sets a readout pixel as part of a pixel array unit 101 (see FIG. 4 ) in which a plurality of pixels are arranged in a two-dimensional array, and controls the readout of pixel signals from the pixels included in the pixel area. More specifically, the readout unit 110 receives readout area information indicating a readout area from the readout area determination unit 123 of the recognition processing unit 12, the readout area information indicating a readout area in which the recognition processing unit 12 performs readout. The readout area information is, for example, the line number of one or more lines. Alternatively, the readout area information may be information indicating pixel positions within a line. Furthermore, various patterns of readout areas can be specified by combining one or more line numbers with information indicating pixel positions of one or more pixels within a line as the readout area information. The readout area is equivalent to a readout unit. Alternatively, the readout area and the readout unit may be different.

[0104] The reading unit 110 can also receive information indicating exposure and analog gain from the recognition processing unit 12 or the visual field processing unit 14 (see FIG. 1 ). The reading unit 110 outputs the input information indicating exposure and analog gain, read area information, and the like to the reliability calculation unit 125.

[0105] The readout unit 110 reads pixel data from the sensor unit 10 in accordance with the readout region information input from the recognition processing unit 12. For example, the readout unit 110 obtains a line number indicating the line to be read out and pixel position information indicating the position of the pixel to be read out on that line based on the readout region information, and outputs the obtained line number and pixel position information to the sensor unit 10. The readout unit 110 outputs each pixel data acquired from the sensor unit 10 to the reliability calculation unit 125 together with the readout region information.

[0106] In addition, the readout unit 110 sets exposure and analog gain (AG) for the sensor unit 10 according to the supplied information indicating exposure and analog gain. Furthermore, the readout unit 110 can generate a vertical synchronization signal and a horizontal synchronization signal and supply them to the sensor unit 10.

[0107] In the recognition processing unit 12, the readout area determination unit 123 receives readout information indicating the readout area to be read next from the feature accumulation control unit 121. The readout area determination unit 123 generates readout area information based on the received readout information and outputs it to the reading unit 110.

[0108] Here, the readout region determination unit 123 can use, for example, information in which readout position information for reading pixel data of a predetermined readout unit is added to a predetermined readout unit as the readout region indicated in the readout region information. A readout unit is a set of one or more pixels and serves as a unit of processing by the recognition processing unit 12 and the visual recognition processing unit 14. As an example, if the readout unit is a line, a line number [L#x] indicating the position of the line is added as the readout position information. Also, if the readout unit is a rectangular region including multiple pixels, information indicating the position of the rectangular region in the pixel array unit 101, for example, information indicating the position of the pixel in the upper left corner, is added as the readout position information. The readout region determination unit 123 is assigned a readout unit to be applied in advance. Also, when reading subpixels in a global shutter system, the readout region determination unit 123 can include position information of the subpixels in the readout region. Alternatively, the readout region determination unit 123 can determine the readout unit in response to, for example, an instruction from outside the readout region determination unit 123. Therefore, the reading area determination unit 123 functions as a reading unit control unit that controls the reading unit.

[0109] The readout area determination unit 123 can also determine the readout area to be read next based on recognition information supplied from the recognition processing execution unit 124 described later, and generate readout area information indicating the determined readout area.

[0110] In the recognition processing unit 12, the feature calculation unit 120 calculates the feature in the area indicated by the read-out area information based on the pixel data and read-out area information supplied from the read-out unit 110. The feature calculation unit 120 outputs the calculated feature to the feature accumulation control unit 121.

[0111] The feature amount calculation unit 120 may calculate the feature amount based on the pixel data supplied from the readout unit 110 and the past feature amount supplied from the feature amount accumulation control unit 121. However, the feature amount calculation unit 120 may also obtain, for example, information for setting exposure and analog gain from the readout unit 110, and calculate the feature amount by further using the obtained information.

[0112] In the recognition processing unit 12, the feature accumulation control unit 121 accumulates the feature supplied from the feature calculation unit 120 in the feature accumulation unit 122. When the feature is supplied from the feature calculation unit 120, the feature accumulation control unit 121 generates read information indicating the read area from which the next readout will be performed, and outputs the read information to the read area determination unit 123.

[0113] Here, the feature accumulation control unit 121 can integrate and accumulate already accumulated feature amounts and newly supplied feature amounts. Furthermore, the feature accumulation control unit 121 can delete feature amounts that are no longer needed from the feature amounts accumulated in the feature accumulation unit 122. Possible examples of unnecessary feature amounts include feature amounts related to the previous frame and feature amounts that have been calculated and already accumulated based on frame images of scenes different from the frame images for which new feature amounts have been calculated. Furthermore, the feature accumulation control unit 121 can also initialize the feature accumulation unit 122 by deleting all feature amounts accumulated therein as necessary.

[0114] Furthermore, the feature accumulation control unit 121 generates a feature to be used in the recognition process by the recognition process execution unit 124, based on the feature supplied from the feature calculation unit 120 and the feature accumulated in the feature accumulation unit 122. The feature accumulation control unit 121 outputs the generated feature to the recognition process execution unit 124.

[0115] The recognition process execution unit 124 executes recognition processing based on the feature amounts supplied from the feature amount accumulation control unit 121. The recognition process execution unit 124 performs object detection, face detection, and the like through the recognition process. The recognition process execution unit 124 outputs the recognition result obtained through the recognition process to the output control unit 15 and the reliability calculation unit 125. The recognition result includes information on the detection score. Note that the detection score according to this embodiment corresponds to the reliability.

[0116] The recognition processing execution unit 124 can also output recognition information including the recognition result generated by the recognition processing to the readout area determination unit 123. Note that the recognition processing execution unit 124 can receive features from the feature accumulation control unit 121 and execute the recognition processing based on a trigger generated by, for example, a trigger generation unit (not shown).

[0117] 18A is a block diagram showing the configuration of the reliability map generation unit 126. The reliability map generation unit 126 generates a reliability correction value for each pixel. The reliability map generation unit 126 includes a read count accumulation unit 126a, a read count acquisition unit 126b, an integration time setting unit 126c, and a read area map generation unit 126e. In this embodiment, a two-dimensional layout diagram of the reliability correction values ​​for each pixel is referred to as a reliability map. Furthermore, for example, the final reliability is determined by multiplying the representative value of the correction values ​​within a recognition rectangle by the reliability in that recognition rectangle.

[0118] The read count accumulation unit 126a accumulates the number of reads for each pixel together with the read time in the accumulation unit 126b. This read count accumulation unit 126a can integrate the number of reads for each pixel that has already been accumulated in the accumulation unit 126b and the newly supplied number of reads for each pixel to obtain the number of reads for each pixel.

[0119] FIG. 18B is a diagram showing that the number of line data readouts varies depending on the integration interval (time). The horizontal axis represents time, and the diagram shows an example of line readout in a 1 / 4 cycle interval (time). The line data in one cycle interval (time) covers the entire image data range. On the other hand, considering periodic readout, the number of line data in a 1 / 4 cycle is one-fourth of one cycle. Thus, if the integration time is one-fourth of one cycle, the number of line data is, for example, two lines in FIG. 18B. On the other hand, if the integration time is two-fourths of one cycle, the number of line data is, for example, four lines in FIG. 18B. If the integration time is three-fourths of one cycle, the number of line data is, for example, six lines in FIG. 18B. If the integration time is one cycle, the number of line data is, for example, eight lines, i.e., all pixels, in FIG. 18B. Therefore, the integration time setting unit 126c supplies a signal including information about the integration interval (time) to the read count acquisition unit 126d.

[0120] FIG. 18C is a diagram showing an example in which the line data read position is adaptively changed according to the recognition result of the recognition processing execution unit 124 shown in FIG. 16. In such a case, as shown in the left diagram, line data is read sequentially while thinning out. Next, as shown in the middle diagram, if an "8" or a "0" is identified along the way, as shown in the right diagram, only the part where it is likely to be possible to distinguish between an "8" and a "0" is read back. In such a case, the concept of a period does not exist. Even if such a period does not exist, the number of times the line data is read varies depending on the interval (time) over which the data is accumulated. For this reason, the accumulation time setting unit 126c supplies a signal including information about the interval (time) over which the data is accumulated to the read count acquisition unit 126d.

[0121] The read count acquisition unit 126d acquires the number of reads for each pixel for each acquisition section from the read count accumulation unit 126a. The read count acquisition unit 126d supplies the accumulated time (section to be accumulated) supplied from the accumulated time setting unit 126c and the number of reads for each pixel for each acquisition section to the read area map generation unit 126e. For example, in response to a trigger generated by a trigger generation unit (not shown), the read count acquisition unit 126d can read the number of reads for each pixel from the read count accumulation unit 126a, read them together with the accumulated time, and supply them to the read area map generation unit 126e.

[0122] The readout area map generating unit 126e generates a correction value of the reliability for each pixel based on the number of readouts for each pixel for each acquisition section and the integration time. The readout area map generating unit 126e will be described in detail later.

[0123] 17, the score correction unit 127 calculates the final reliability by multiplying, for example, the representative value of the correction values ​​within the recognition rectangle by the reliability in the recognition rectangle. In this embodiment, the two-dimensional arrangement diagram of the correction values ​​of the reliability for each pixel is referred to as a reliability map. The score correction unit 127 outputs the corrected reliability to the output control unit 15 (see FIG. 1).

[0124] 19 is a schematic diagram showing in more detail an example of processing in the recognition processing unit 12 according to this embodiment. Here, the readout area is a line, and the readout unit 110 reads pixel data line by line from the top to the bottom of the frame of the image 60.

[0125] 20 is a schematic diagram for explaining the readout process of the readout unit 110. For example, the readout unit is a line, and pixel data is read out line by line for a frame Fr(x). In the example of FIG. 20, in the mth frame Fr(m), lines are read out line by line starting from the top line L#1 of the frame Fr(m), with lines L#2, L#3, ... being read out line by line. When line reading in frame Fr(m) is completed, lines are similarly read out line by line starting from the top line L#1 in the next (m+1)th frame Fr(m+1).

[0126] 21(a), which will be described later, in the readout process of the readout unit 110, line data may be read out every third line, such as line L#1 being the first line from the top, line L#2 being the fourth line from the top, and line L#3 being the eighth line from the top. Similarly, line data may be read out every third line, such as line L#1 being the first line from the top, line L#2 being the fourth line from the top, and line L#3 being the eighth line from the top.

[0127] Similarly, as shown in Figure 21(b) described below, in the reading process of the reading unit 110, line data may be read every other line, such as line L#1 being the first line from the top, line L#2 being the third line from the top, and line L#3 being the fifth line from the top.

[0128] The line image data (line data) of the line L#x read by the reading unit 110 on a line-by-line basis is input to the feature calculation unit 120. In addition, information on the line L#x read by the line-by-line basis, i.e., read area information, is supplied to the reliability map generation unit 126.

[0129] The feature calculation unit 120 executes a feature extraction process 1200 and an integration process 1202. The feature calculation unit 120 performs the feature extraction process 1200 on the input line data to extract a feature 1201 from the line data. Here, the feature extraction process 1200 extracts the feature 1201 from the line data based on parameters obtained in advance by learning. The feature 1201 extracted by the feature extraction process 1200 is integrated with a feature 1212 processed by the feature accumulation control unit 121 by the integration process 1202. The integrated feature 1210 is passed to the feature accumulation control unit 121.

[0130] The feature accumulation control unit 121 executes internal state update processing 1211. The feature 1210 passed to the feature accumulation control unit 121 is passed to the recognition processing execution unit 124 and also subjected to internal state update processing 1211. The internal state update processing 1211 reduces the feature 1210 based on pre-learned parameters to update the internal state of the DNN and generates a feature 1212 related to the updated internal state. This feature 1212 is integrated with the feature 1201 by integration processing 1202. This processing by the feature accumulation control unit 121 corresponds to processing using an RNN.

[0131] The recognition processing execution unit 124 executes a recognition processing 1240 on the feature 1210 passed from the feature accumulation control unit 121 based on parameters previously learned using, for example, predetermined training data, and outputs a recognition result including information on the recognition area and reliability.

[0132] As described above, in the recognition processing unit 12 according to this embodiment, the feature extraction process 1200, the integration process 1202, the internal state update process 1211, and the recognition process 1240 are executed based on pre-trained parameters. The parameters are learned using training data based on an expected recognition target, for example.

[0133] The reliability map generation unit 126 of the reliability calculation unit 125 calculates a correction value of the reliability for each pixel based on the readout region information and the integrated time information, for example, using information on the line L#x read out line by line. 21 is a diagram showing areas L20a, L20b (valid areas) that have been read out line by line, and areas L22a, L22b (invalid areas) that have not been read out. In this embodiment, the area from which image information has been read out is referred to as the valid area, and the area from which image information has not been read out is referred to as the invalid area.

[0134] The read area map generating section 126e of the reliability map generating section 126 generates the ratio of the valid area to the entire image area as a screen average. Fig. 21(a) shows a case where the area of ​​region L20a read out line by line at a quarter cycle is one-fourth of the entire image, while Fig. 21(b) shows a case where the area of ​​region L20b read out line by line at a quarter cycle is one-half of the entire image.

[0135] In this case, the area map generator 126e generates a screen average of 1 / 4, which is the ratio of the effective area to the entire image area, for Figure 21(a). Similarly, the readout area map generator 126e generates a screen average of 1 / 2, which is the ratio of the effective area to the entire image area, for Figure 21(b). In this way, the readout area map generator 126e can calculate the screen average using information on the effective area and information on the invalid area.

[0136] The read area map generating unit 126e can also calculate the screen average by filtering. For example, the pixel values ​​in region L20a are set to 1 and the pixel values ​​in region L22a to 0, and a smoothing calculation is performed on the pixel values ​​in the entire image. For example, this smoothing calculation is a filtering process that reduces high-frequency components. In this case, the vertical size of the filter is set to the vertical length of the valid region plus the vertical length of the invalid region. In FIG. 21(a), for example, the vertical length of the invalid region is 12 pixels, and the vertical length of the valid region is 3 pixels. In this case, the vertical size of the filter is equivalent to 16 pixels. With this vertical size of the filter, the result of the filtering process is calculated as one-fourth of the screen average, regardless of the horizontal size.

[0137] Similarly, in Figure 21(b), for example, assume that the vertical length of the valid area is three pixels, and the vertical length of the invalid area is three pixels. In this case, for example, the vertical size of the filter is equivalent to a length of six pixels. With this vertical size of the filter, the result of the filtering process is calculated as half the screen average, regardless of the horizontal size.

[0138] For the recognition area A20a, the score correction unit 127 corrects the reliability corresponding to the recognition area A20a based on a representative value of the correction values ​​within the recognition area A20a. For example, the representative value can be a statistical value such as the average value, median value, or mode of the correction values ​​within the recognition area A20a. For example, the representative value can be set to one-fourth of the average value of the correction values ​​within the recognition area A20a. In this way, the score correction unit 127 can use the screen average of the read screen to calculate the reliability.

[0139] On the other hand, for the recognition area A20b, the score correction unit 127 corrects the reliability corresponding to the recognition area A20b based on a representative value of the correction values ​​within the recognition area A20b. For example, the reliability is corrected to half the average value of the correction values ​​within the recognition area A20b. As a result, the reliability corresponding to the recognition area A20a is corrected based on one-fourth, and the reliability corresponding to the recognition area A20a is corrected based on half. In this embodiment, the final reliability is obtained by multiplying the reliability corresponding to A20b by the representative value of the correction values ​​within the recognition area A20b. Note that a function having a nonlinear input / output relationship may be used to multiply the reliability by the output value obtained after function calculation using the representative value as an input.

[0140] As described above, sensor control results in read-out regions L20a, L20b and unread-out regions L22a, L22b. This differs from general recognition processing in which pixels of the entire region are read. Therefore, if a general reliability is used in a case where read-out regions L20a, L20b and unread regions L22a, L22b occur, the accuracy of the reliability may be reduced. In contrast, in this embodiment, the reliability map generation unit 126 calculates a correction value for each pixel according to the read-out regions L20a, L20b / (read-out regions L20a, L20b+unread regions L22a, L22b) as the screen average. The score correction unit 127 then corrects the reliability based on the correction value, enabling calculation of a more accurate reliability.

[0141] The functions of the above-mentioned feature calculation unit 120, feature accumulation control unit 121, readout area determination unit 123, recognition processing execution unit 124, and reliability calculation unit 125 are realized, for example, by reading and executing a program stored in memory 13 or the like provided in the information processing system 1.

[0142] In the above description, line readout is performed from the top edge of the frame to the bottom edge, but this is not limited to this example. For example, line readout may be performed from the left edge to the right edge, or from the right edge to the left edge.

[0143] Fig. 22 shows areas L21a and L21b that have been read line by line from the left end to the right end, and areas L23a and L23b that have not been read. Fig. 22(a) shows the case where area L21a that has been read line by line is one-fourth of the entire image. On the other hand, Fig. 22(b) shows the case where area L21b that has been read line by line is one-half of the entire image.

[0144] In this case, the read area map generator 126e of the reliability map generator 126 generates a screen average of 1 / 4, which is the ratio of the valid area to the entire image area, for Fig. 22(a). Similarly, the area map generator 126e generates a screen average of 1 / 2, which is the ratio of the valid area to the entire image area, for Fig. 21(b).

[0145] For the recognition area A21a, the score correction unit 127 corrects the reliability corresponding to the recognition area A21a based on a representative value of the correction values ​​in the recognition area A21a, for example, to one-fourth of the average value of the correction values ​​in the recognition area A21a.

[0146] On the other hand, for the recognition area A21b, the score correction unit 127 corrects the reliability corresponding to the recognition area A21b based on a representative value of the correction values ​​in the recognition area A21b. For example, the representative value is set to half the average value of the correction values ​​in the recognition area A21b.

[0147] 23 is a diagram schematically illustrating an example of reading line-by-line from the left end to the right end. The upper diagram shows the read area and the unread area. In the area where recognition area A23a exists, the area ratio where line data exists is one-quarter, and in the area where recognition area A23b exists, the area ratio where line data exists is one-half. In other words, this is an example in which the recognition processing execution unit 124 adaptively changes the area where line data is read.

[0148] The lower diagram shows a reliability map generated by the read area map generator 126e. This diagram illustrates a two-dimensional distribution in the read area map. As described above, the read area map illustrates a two-dimensional distribution of reliability correction values ​​based on the area of ​​the read data. The correction values ​​are represented by grayscale values. For example, the read area map generator 126e assigns 1 to valid areas and 0 to invalid image areas, as described above. The read area map generator 126e then generates an area map by performing a smoothing calculation on the entire image, for example, for each rectangular area centered on a pixel. For example, the rectangular area may be a 5x5 pixel area. As a result of this processing, in Figure 23, although there is variation depending on the pixel position, in an area where the area ratio is one-quarter, the correction value of each pixel is approximately one-quarter. On the other hand, in an area where the area ratio is one-half, there is variation depending on the pixel position, but the correction value of each pixel is approximately one-half. Note that the specified area is not limited to a rectangle and may be, for example, an ellipse or a circle. In this embodiment, a predetermined value is assigned to the valid area and the invalid area, and the image obtained by the smoothing calculation process is called an area map.

[0149] For the recognition area A23a, the score correction unit 127 corrects the reliability corresponding to the recognition area A21b based on a representative value of the correction values ​​within the recognition area A21b. For example, the representative value is set to one-fourth the average value of the correction values ​​within the recognition area A23ab. On the other hand, for the recognition area A23b, the score correction unit 127 corrects the reliability corresponding to the recognition area A21b based on a representative value of the correction values ​​within the recognition area A23b. For example, the representative value is set to one-half the average value of the correction values ​​within the recognition area A23b. In this way, by displaying the reliability map, it is possible to grasp the overall reliability of the recognition areas within the image area in a short time.

[0150] 24 is a diagram schematically showing the values ​​of the reliability map when the readout area changes within the recognition region A24. As shown in FIG. 24, when the readout area changes within the recognition region A24, the values ​​of the reliability map also change within the recognition region A24. In this case, the score correction unit 127 may use, as the representative value within the recognition region A24, the value of the most frequent value within the recognition region A24, the value at the center of the recognition region A24, or a weighted integrated value in which the distance from the center of the recognition region A24 is used as a weight.

[0151] Fig. 25 is a diagram schematically showing an example in which the read range of line data is limited. As shown in Fig. 25, the read range of line data may be changed for each read timing. In this case, too, the read area map generating unit 126e can generate a reliability map using the same method as described above.

[0152] Fig. 26 is a diagram schematically illustrating an example of classification processing (recognition processing) by a DNN when time-series information is not used. In this case, as shown in Fig. 26, one image is subsampled and input to the DNN. In the DNN, classification processing is performed on the input image, and the classification result is output.

[0153] 27A is a diagram showing an example in which one image is subsampled in a grid pattern. Even when the entire image is subsampled in this way, the readout area map generation unit 126e can generate a reliability map by using the ratio between the number of sampled pixels and the total number of pixels. In this case, for the recognition area A26, the score correction unit 127 corrects the reliability corresponding to the recognition area A26 based on a representative value of the correction values ​​within the recognition area A26.

[0154] 27B is a diagram showing an example in which one image is subsampled in a checkerboard pattern. Even when the entire image is subsampled in this way, the readout area map generation unit 126e can generate a reliability map by using the ratio between the number of sampled pixels and the total number of pixels. In this case, for the recognition area A27, the score correction unit 127 corrects the reliability corresponding to the recognition area A27 based on a representative value of the correction values ​​within the recognition area A27.

[0155] Figure 28 is a diagram showing a schematic diagram of a case where a reliability map is used in a transportation system, for example, a moving body. Figure (a) shows the average value of the read area in shades of gray. The shade indicated by "0" indicates that the average value of the read area is 0, and the shade indicated by "1 / 2" indicates that the average value of the read area is 1 / 2.

[0156] Figures (b) and (c) show examples in which a readout area map is used as a reliability map. The correction value for the right region in Figure (b) is lower than the correction value for the right region in Figure (c). As a result, for example, in a situation like Figure (b), if the reliability map is not used, the robot will change course to the right of the camera, even though there is a possibility that an object is there. On the other hand, if the reliability map is used, the correction value for the region to the right of the camera is low, resulting in a low reliability, so the robot can stop on the spot without changing course to the right of the camera, taking into account the possibility that an object may be there.

[0157] On the other hand, as shown in Figure (c), when the correction value for the area to the right of the camera becomes high, the reliability increases, so it is possible to determine that there is no object to the right of the camera and change course to the right of the camera.

[0158] For example, even if the detection score is high, if the reliability is low (if the correction value based on the read area is low), it is necessary to consider the possibility that there is no object. As an example of updating the reliability, as described above, it can be calculated as reliability = detection score (original reliability) x correction value based on the read area. When the urgency is low (e.g., there is no possibility of an imminent collision), even if the detection score is high, it is possible to determine that there is no object there if the reliability (value after correction based on the read area) is low. When the urgency is high (e.g., there is a possibility of an imminent collision), it is possible to determine that there is an object there if the detection score is high, even if the reliability (value after correction based on the read area) is low. In this way, using a reliability map enables safer control of moving objects such as cars.

[0159] 29 is a flowchart showing the flow of processing by the reliability calculation unit 125. Here, an example of processing in the case of line data will be described.

[0160] First, the read count accumulation unit 126a acquires read area information including read line number information from the read unit 110 (step S100), and accumulates information on the read pixels and times in the accumulation unit 126b as information on the number of reads per pixel (step S102).

[0161] Next, the read count acquisition unit 126d determines whether or not a trigger signal for map generation has been input (step S104). If not (No in step S104), the process repeats from step S100. On the other hand, if the trigger signal has been input (Yes in step S104), the read count acquisition unit 126d acquires the number of times each pixel has been read within an integrated time, for example, a time corresponding to a quarter cycle, from the read count accumulation unit 126a (step S106). Here, the number of times each pixel has been read within a time corresponding to a quarter cycle is set to one. For example, there are cases where a pixel is read out several times within a time corresponding to a quarter cycle, and this case will be described later.

[0162] Next, the readout area map generating unit 126e generates a correction value indicating the proportion of the readout area for each pixel (step S108). Subsequently, the readout area map generating unit 126e outputs the two-dimensional arrangement data of the correction values ​​to the output control unit 15 as a reliability map.

[0163] Next, the score correcting unit 127 acquires the detection score for the rectangular area (for example, the recognition area A20a in FIG. 21), that is, the reliability, from the recognition processing executing unit 124 (step S110).

[0164] Next, the score correction unit 127 acquires a representative value of the correction values ​​in a rectangular area (for example, the recognition area A20a of FIG. 21) (step S112). For example, the representative value can be a statistical value such as the average value, median value, or mode value of the correction values ​​in the recognition area A20a.

[0165] Then, the score corrector 127 updates the detection score based on the detection score and the representative value (step S114), outputs it as the final reliability, and ends the overall processing.

[0166] As described above, according to this embodiment, the reliability map generation unit 126 calculates a correction value of the reliability for each pixel according to the read areas L20a, L20b / (read areas L20a, L20b+unread areas L22a, L22b) (FIG. 21). Then, the score correction unit 127 corrects the reliability based on the correction value, so that it is possible to calculate a reliability with higher accuracy. As a result, even when the read areas L20a, L20b and the unread areas L22a, L22b occur due to sensor control, the corrected reliability values ​​can be processed uniformly, so that the recognition accuracy of the recognition process can be further improved.

[0167] (Modification 1 of the first embodiment) The information processing system 1 according to the first modification of the first embodiment differs from the information processing system 1 according to the first embodiment in that the range for calculating the correction value of the reliability can be calculated based on the receptive field of the feature. The differences from the information processing system 1 according to the first embodiment will be described below.

[0168] FIG. 30 is a schematic diagram showing the relationship between features and receptive fields. A receptive field refers to the range of an input image referenced when calculating a feature; in other words, the range of an input image viewed by a feature. Receptive field R30 in image A312 corresponds to feature region AF30 in recognition region A30 within image A312, and receptive field R32 in image A312 corresponds to feature region AF32 within recognition region A32. As shown in FIG. 31, the feature in feature region AF30 is used as the feature corresponding to recognition region A30. In this embodiment, the range in image A312 used to calculate the feature corresponding to recognition region A30 is referred to as receptive field R30. Similarly, the range in image A312 used to calculate the feature corresponding to recognition region A32 corresponds to receptive field R32.

[0169] 31 is a diagram schematically illustrating the recognition regions A30 and A32 and the receptive fields R30 and R32 in the reliability map. The score correction unit 127 according to this first modification differs from the score correction unit 127 according to the first embodiment in that it is also possible to calculate a representative value of the correction value using information on the receptive fields R30 and R32. For example, the receptive field R30 and the recognition region A30 differ in position and size within the image 312, and therefore may have different average readout areas. To more accurately reflect the influence of the readout region, it is desirable to use the range of the receptive field R30 used to calculate the feature amount.

[0170] Therefore, the score correction unit 127 corrects the detection score of the recognition region A30, for example, using a representative value of the correction values ​​within the receptive field R30. The score correction unit 127 can use a statistical value, such as the mode of the correction values ​​within the receptive field R30, as the representative value. Then, the score correction unit 127 updates the detection score of the recognition region A30 by, for example, multiplying the representative value within the receptive field R30 by the detection score. The updated detection score is used as the final reliability. Similarly, the score correction unit 127 can use a statistical value, such as the average, median, or mode of the correction values ​​within the receptive field R32, as the representative value. Then, the score correction unit 127 updates the detection score of the recognition region A30 by, for example, multiplying the representative value within the receptive field R32 by the detection score of the recognition region A30.

[0171] 31, when the detection scores are updated using the recognition areas A30 and A32, the reliability of the recognition area A30 is updated to be higher than the reliability of the recognition area A32. On the other hand, when the detection scores are updated using the receptive fields R30 and R32, for example, if the representative value is set to the mode of the receptive fields R30 and R32, the ratio between the reliability of the updated recognition area A30 and the reliability of the updated recognition area A32 will be equal. In this way, by taking into account the range of the receptive fields R30 and R3, the reliability may be updated with higher accuracy.

[0172] Figure 32 is a diagram showing the contribution of features to recognition area A30. The shading in the receptive field R30 on the right indicates a weighting value that reflects the contribution of features to recognition processing in recognition area A30 (see Figure 31). The darker the shading, the higher the contribution.

[0173] The score correction unit 127 may use such weighting values ​​to accumulate the correction values ​​in the receptive field R30 and use the result as a representative value. Since the contribution to the feature amount is reflected, the accuracy of the reliability of the updated recognition area A30 is further improved.

[0174] (Modification 2 of the first embodiment) The information processing system 1 according to the second modification of the first embodiment performs semantic segmentation as a recognition task. Semantic segmentation is a recognition technique that associates (assigns, sets, and classifies) a label or category with each pixel in an image based on the characteristics of that pixel and its surrounding pixels. It is performed, for example, by deep learning using a neural network. Semantic segmentation can recognize groups of pixels that share the same label or category based on the labels or categories associated with each pixel, and can divide an image into multiple regions at the pixel level, allowing irregularly shaped objects to be detected and clearly distinguished from surrounding objects. For example, when a semantic segmentation task is performed on a typical roadway scene, vehicles, pedestrians, signs, roadways, sidewalks, traffic lights, sky, roadside trees, guardrails, and other objects can be classified and recognized within the image by their respective categories. The types and numbers of labels and categories used for this classification can be varied depending on the dataset used for learning and individual settings. For example, it may be performed using only two labels or categories, people and background, or may use multiple, detailed labels and categories as described above, depending on the purpose and device performance. Below, differences from the information processing system 1 according to the first embodiment will be described.

[0175] FIG. 33 is a schematic diagram showing a recognition process performed on an image using general semantic segmentation. In this process, semantic segmentation is performed on the entire image, whereby a corresponding label or category is assigned to each pixel, and the image is divided into multiple regions at the pixel level by groups of pixels that form the same label or category. In general, semantic segmentation outputs the reliability of the assigned label or category for each pixel. Alternatively, for groups of pixels that form the same label or category, the average reliability of each group of pixels may be calculated, and this may be used as the reliability of the group of pixels, thereby calculating a single reliability for each group of pixels. Alternatively, a median or other value may be used instead of the average.

[0176] In the second modification of the first embodiment, the score correction unit 127 corrects the reliability calculated by a general semantic segmentation process. That is, correction is performed based on the readout area (screen average) occupied in the image, correction based on a representative value of the correction value of the recognition area, correction based on the reliability map (map integration unit 126j, readout area map generation unit 126e, readout frequency map generation unit 126f, multiple exposure map generation unit 126g, and dynamic range map generation unit 126h), and correction using the receptive field. In this way, in the second modification of the first embodiment, by applying the present invention to the recognition process by semantic segmentation, the corrected reliability is calculated, thereby enabling the reliability to be calculated with higher accuracy.

[0177] (Second embodiment) The information processing system 1 according to the second embodiment differs from the information processing system 1 according to the first embodiment in that the reliability correction value can be calculated based on the pixel readout frequency. The following describes the differences from the information processing system 1 according to the first embodiment.

[0178] Fig. 34 is a block diagram of the reliability map generating unit 126 according to the second embodiment. As shown in Fig. 34, the reliability map generating unit 126 further includes a read frequency map generating unit 126f.

[0179] 35 is a diagram showing a schematic diagram of the relationship between the recognition area A36 and the line data L36a. The upper diagram shows the line data L36a and the non-read area L36b, and the lower diagram shows a reliability map. Here, it is a read frequency map. (a) shows the line data L36a being read once, (b) shows the line data L36a being read twice, (c) shows the line data L36a being read three times, and (d) shows the line data L36a being read four times.

[0180] The read frequency map generating unit 126f performs a smoothing calculation process on the appearance frequency of pixels in the entire region of the image. For example, this smoothing calculation process is a filtering process that reduces high frequency components.

[0181] As shown in FIG. 35, in this embodiment, for example, smoothing calculation processing is performed on the entire image, for example, for each rectangular area centered on a pixel. For example, the rectangular area is a 5x5 pixel area. As a result of this processing, in FIG. 35(a), although there is variation depending on the pixel position, the correction value of each pixel is approximately halved. Meanwhile, in FIG. 35(b), the area from which line data L36a has been read shows 1 time, in FIG. 35(c), the area from which line data L36a has been read shows 3 / 2 times, and in FIG. 35(d), the area from which line data L36a has been read shows 2 times. Furthermore, in areas from which no data has been read, the read frequency is 0.

[0182] For the recognition area A36, the score correction unit 127 corrects the reliability corresponding to the recognition area A36 based on a representative value of the correction values ​​in the recognition area A36. For example, the representative value can be a statistical value such as the average value, median value, or mode of the correction values ​​in the recognition area A36.

[0183] As described above, according to this embodiment, the reliability map generation unit 126 performs a smoothing calculation process on the occurrence frequencies of pixels within a predetermined range centered on a pixel for the entire image region, and calculates a corrected reliability value for each pixel in the entire image region.The score correction unit 127 then corrects the reliability based on the correction value, making it possible to calculate a more accurate reliability that reflects the pixel readout frequency.As a result, even if there is a difference in the pixel readout frequency, the corrected reliability value can be processed uniformly, thereby further improving the recognition accuracy of the recognition process. (Third embodiment) The information processing system 1 according to the third embodiment differs from the information processing system 1 according to the first embodiment in that it is capable of calculating a reliability correction value based on the number of exposures of a pixel. The following describes the differences from the information processing system 1 according to the first embodiment.

[0184] Fig. 36 is a block diagram of the reliability map generation unit 126 according to the third embodiment 3. As shown in Fig. 36, the reliability map generation unit 126 further includes a multiple exposure map generation unit 126g.

[0185] 37 is a diagram showing a schematic diagram of the relationship between the line data L36a and the exposure frequency. The upper diagram shows the line data L36a and the non-read area L36b, and the lower diagram shows a reliability map. In this example, it is a multiple exposure map. (a) shows the line data L36a exposed twice, (b) shows the line data L36a exposed four times, and (c) shows the line data L36a exposed six times.

[0186] The read frequency map generator 126f performs a smoothing calculation process on the exposure counts of pixels within a predetermined range centered on the pixel for the entire image area, and calculates a correction value for the reliability of each pixel in the entire image area. For example, this smoothing calculation process is a filtering process that reduces high-frequency components.

[0187] As shown in FIG. 37, in this embodiment, for example, the predetermined range for which smoothing calculation processing is performed is a rectangular range corresponding to a 5×5 pixel range. Through this processing, although there is some variation depending on the pixel position, the correction value of each pixel is approximately halved in FIG. 37(a). Meanwhile, in FIG. 37(b), the area from which line data L36a is read shows a count of 1 exposure, in FIG. 37(c), the area from which line data L36a is read shows a count of 3 / 2 exposures, and in FIG. 37(d), the area from which line data L36a is read shows a count of 2 exposures. Furthermore, in areas from which no data is read, the read frequency is 0.

[0188] For the recognition area A36, the score correction unit 127 corrects the reliability corresponding to the recognition area A36 based on a representative value of the correction values ​​in the recognition area A36. For example, the representative value can be a statistical value such as the average value, median value, or mode of the correction values ​​in the recognition area A36.

[0189] As described above, according to this embodiment, the reliability map generation unit 126 performs a smoothing calculation process on the exposure counts of pixels within a predetermined range centered on the pixel for the entire image area, and calculates a corrected reliability value for each pixel in the entire image area.The score correction unit 127 then corrects the reliability based on the correction value, making it possible to calculate a more accurate reliability that reflects the exposure count of the pixel.As a result, even if there is a difference in the exposure count of a pixel, the corrected reliability value can be processed uniformly, thereby further improving the recognition accuracy of the recognition process.

[0190] (Fourth embodiment) The information processing system 1 according to the fourth embodiment differs from the information processing system 1 according to the first embodiment in that the information processing system 1 according to the fourth embodiment can calculate a correction value for reliability based on the dynamic range of pixels. The following describes the differences from the information processing system 1 according to the first embodiment.

[0191] Fig. 38 is a block diagram of the reliability map generation unit 126 according to the fourth embodiment. As shown in Fig. 38, the reliability map generation unit 126 further includes a dynamic range map generation unit 126h.

[0192] 39 is a diagram showing a schematic diagram of the relationship between the line data L36a and the dynamic range. The upper diagram shows the line data L36a and the non-read area L36b, and the lower diagram shows the reliability map. Here, it is a dynamic range map. In (a), the dynamic range of the line data L36a is 40 db, in (b), the dynamic range is 80 db, and in (c), the dynamic range is 120 db.

[0193] The dynamic range map generator 126h performs smoothing calculations on the dynamic range of pixels within a predetermined range centered on the pixel for the entire image area, and calculates a correction value for the reliability of each pixel in the entire image area. For example, this smoothing calculation is a filtering process that reduces high-frequency components.

[0194] As shown in FIG. 39, in this embodiment, for example, the predetermined range for which the smoothing calculation process is performed is a rectangular range corresponding to a 5×5 pixel range. Through this process, in FIG. 35(a), although there is some variation depending on the pixel position, the correction value for each pixel is approximately 20. Meanwhile, in FIG. 35(b), the area from which the line data L36a is read out shows a number of exposures of 40, and in FIG. 35(c), the area from which the line data L36a is read out shows a number of exposures of 80. Furthermore, in areas from which no data is read out, the read frequency is 0. The dynamic range map generator 126h normalizes the correction value to, for example, a range from 0.0 to 1.0.

[0195] For the recognition area A36, the score correction unit 127 corrects the reliability corresponding to the recognition area A36 based on a representative value of the correction values ​​in the recognition area A36. For example, the representative value can be a statistical value such as the average value, median value, or mode of the correction values ​​in the recognition area A36.

[0196] As described above, according to this embodiment, the reliability map generation unit 126 performs smoothing calculation processing on the dynamic range of pixels within a predetermined range centered on a pixel for the entire image region, and calculates a corrected reliability value for each pixel in the entire image region.The score correction unit 127 then corrects the reliability based on the correction value, making it possible to calculate a more accurate reliability that reflects the dynamic range of the pixels.As a result, even when differences occur in the dynamic range of pixels, the corrected reliability values ​​can be processed uniformly, thereby further improving the recognition accuracy of the recognition process.

[0197] (Fifth embodiment) The information processing system 1 according to the fifth embodiment differs from the information processing system 1 according to the first embodiment in that it includes a map integration unit that integrates correction values ​​of various reliabilities. The following describes the differences from the information processing system 1 according to the first embodiment.

[0198] Fig. 40 is a block diagram of the reliability map generating unit 126 according to the fifth embodiment. As shown in Fig. 40, the reliability map generating unit 126 further includes a map integrating unit 126j. The map integration unit 126j can integrate the output values ​​of the readout area map generation unit 126e, the readout frequency map generation unit 126f, the multiple exposure map generation unit 126g, and the dynamic range map generation unit 126h.

[0199] The map integration unit 126j multiplies each correction value for each pixel to integrate the correction values ​​as shown in equation (1).

number

[0200] The map integration unit 126j performs weighted addition of each correction value for each pixel, and integrates the correction values ​​as shown in equation (2).

number

[0201] As described above, according to this embodiment, the map integration unit 126j integrates the output values ​​of the readout area map generation unit 126e, the readout frequency map generation unit 126f, the multiple exposure map generation unit 126g, and the dynamic range map generation unit 126h. This makes it possible to generate correction values ​​that take into account the values ​​of each correction value, and to process the reliability values ​​after correction in a unified manner, thereby further improving the recognition accuracy of the recognition process.

[0202] (Sixth embodiment)

[0203] (6-1. Application Examples of the Technology of the Present Disclosure) Next, as a sixth embodiment, an application example of the information processing device 2 according to the first to fifth embodiments of the present disclosure will be described. Fig. 41 is a diagram showing a usage example of the information processing device 2 according to the first to fifth embodiments. Note that, in the following, when there is no particular need to distinguish between them, the information processing device 2 will be used as a representative in the description.

[0204] The information processing device 2 described above can be used in various cases, for example, as follows, for sensing light such as visible light, infrared light, ultraviolet light, and X-rays, and performing recognition processing based on the sensing results.

[0205] ·Devices that take images for viewing purposes, such as digital cameras and mobile devices with camera functions. - Devices used for traffic purposes, such as in-vehicle sensors that take pictures of the front, rear, surroundings, and interior of a vehicle for safe driving such as automatic stopping, and for recognizing the driver's condition, surveillance cameras that monitor moving vehicles and roads, and distance measuring sensors that measure distances between vehicles. A device used in home appliances such as TVs, refrigerators, and air conditioners to capture user gestures and operate the appliances in accordance with those gestures. -Devices used for medical and healthcare purposes, such as endoscopes and devices that take blood vessel images by receiving infrared light. -Devices used for security purposes, such as surveillance cameras for crime prevention and cameras for person authentication. - Cosmetic devices such as skin measuring devices that take pictures of the skin and microscopes that take pictures of the scalp. - Devices used for sports, such as action cameras and wearable cameras for sports. Agricultural equipment, such as cameras for monitoring the condition of fields and crops.

[0206] (6-2. Application examples to moving objects) The technology according to the present disclosure (the present technology) can be applied to various products. For example, the technology according to the present disclosure may be realized as a device mounted on any type of moving body, such as an automobile, an electric vehicle, a hybrid electric vehicle, a motorcycle, a bicycle, personal mobility, an airplane, a drone, a ship, or a robot.

[0207] FIG. 42 is a block diagram showing a schematic configuration example of a vehicle control system, which is an example of a mobile object control system to which the technology according to the present disclosure can be applied.

[0208] The vehicle control system 12000 includes a plurality of electronic control units connected via a communication network 12001. In the example shown in Fig. 42, the vehicle control system 12000 includes a drive system control unit 12010, a body system control unit 12020, an outside-vehicle information detection unit 12030, an inside-vehicle information detection unit 12040, and an integrated control unit 12050. Also shown as functional components of the integrated control unit 12050 are a microcomputer 12051, an audio / video output unit 12052, and an in-vehicle network I / F (interface) 12053.

[0209] The drivetrain control unit 12010 controls the operation of devices related to the drivetrain of the vehicle in accordance with various programs. For example, the drivetrain control unit 12010 functions as a control device for a drive force generating device for generating a drive force of the vehicle, such as an internal combustion engine or a drive motor, a drive force transmission mechanism for transmitting the drive force to the wheels, a steering mechanism for adjusting the steering angle of the vehicle, a braking device for generating a braking force of the vehicle, etc.

[0210] The body system control unit 12020 controls the operation of various devices equipped in the vehicle body according to various programs. For example, the body system control unit 12020 functions as a control device for a keyless entry system, a smart key system, a power window device, or various lamps such as headlamps, backup lamps, brake lamps, turn signals, and fog lamps. In this case, radio waves transmitted from a portable device that serves as a key or signals from various switches may be input to the body system control unit 12020. The body system control unit 12020 receives these radio waves or signals and controls the vehicle's door lock device, power window device, lamps, etc.

[0211] The outside-vehicle information detection unit 12030 detects information outside the vehicle equipped with the vehicle control system 12000. For example, an imaging unit 12031 is connected to the outside-vehicle information detection unit 12030. The outside-vehicle information detection unit 12030 causes the imaging unit 12031 to capture images outside the vehicle and receives the captured images. The outside-vehicle information detection unit 12030 may perform object detection processing or distance detection processing for people, cars, obstacles, signs, characters on the road surface, etc., based on the received images.

[0212] The imaging unit 12031 is an optical sensor that receives light and outputs an electrical signal according to the amount of light received. The imaging unit 12031 can output the electrical signal as an image, or as distance measurement information. The light received by the imaging unit 12031 may be visible light or invisible light such as infrared light.

[0213] The in-vehicle information detection unit 12040 detects information inside the vehicle. For example, a driver state detection unit 12041 that detects the state of the driver is connected to the in-vehicle information detection unit 12040. The driver state detection unit 12041 includes, for example, a camera that captures an image of the driver, and the in-vehicle information detection unit 12040 may calculate the degree of fatigue or concentration of the driver based on the detection information input from the driver state detection unit 12041, or may determine whether the driver is dozing off.

[0214] The microcomputer 12051 can calculate control target values ​​for the driving force generating device, steering mechanism, or braking device based on the information inside and outside the vehicle acquired by the outside-vehicle information detection unit 12030 or the inside-vehicle information detection unit 12040, and output control commands to the drivetrain control unit 12010. For example, the microcomputer 12051 can perform cooperative control aimed at realizing the functions of an ADAS (Advanced Driver Assistance System), including avoiding or mitigating collisions between vehicles, following based on the distance between vehicles, maintaining vehicle speed, warning of vehicle collisions, or warning of vehicle lane departure.

[0215] In addition, the microcomputer 12051 can perform cooperative control for the purpose of autonomous driving, which allows the vehicle to travel autonomously without relying on driver operation, by controlling the driving force generating device, steering mechanism, braking device, etc. based on information about the surroundings of the vehicle obtained by the outside vehicle information detection unit 12030 or the inside vehicle information detection unit 12040.

[0216] Furthermore, the microcomputer 12051 can output a control command to the body system control unit 12020 based on the information about the outside of the vehicle acquired by the outside information detection unit 12030. For example, the microcomputer 12051 can control the headlamps according to the position of a preceding vehicle or an oncoming vehicle detected by the outside information detection unit 12030, and perform cooperative control for the purpose of preventing glare, such as switching from high beams to low beams.

[0217] The audio / video output unit 12052 transmits at least one of audio and video output signals to an output device capable of visually or audibly notifying information to vehicle occupants or the outside of the vehicle. In the example of Fig. 36, an audio speaker 12061, a display unit 12062, and an instrument panel 12063 are exemplified as output devices. The display unit 12062 may include, for example, at least one of an on-board display and a head-up display.

[0218] FIG. 43 is a diagram showing an example of the installation position of the imaging unit 12031.

[0219] In FIG. 43, a vehicle 12100 has imaging units 12101, 12102, 12103, 12104, and 12105 as the imaging unit 12031.

[0220] The imaging units 12101, 12102, 12103, 12104, and 12105 are provided, for example, at positions such as the front nose, side mirrors, rear bumper, back door, and the top of the windshield inside the vehicle cabin of the vehicle 12100. The imaging unit 12101 provided at the front nose and the imaging unit 12105 provided at the top of the windshield inside the vehicle cabin mainly acquire images of the front of the vehicle 12100. The imaging units 12102 and 12103 provided at the side mirrors mainly acquire images of the sides of the vehicle 12100. The imaging unit 12104 provided at the rear bumper or back door mainly acquires images of the rear of the vehicle 12100. The forward images acquired by the imaging units 12101 and 12105 are mainly used to detect preceding vehicles, pedestrians, obstacles, traffic lights, traffic signs, lanes, etc.

[0221] 43 shows an example of the imaging ranges of the imaging units 12101 to 12104. Imaging range 12111 indicates the imaging range of the imaging unit 12101 provided on the front nose, imaging ranges 12112 and 12113 indicate the imaging ranges of the imaging units 12102 and 12103 provided on the side mirrors, respectively, and imaging range 12114 indicates the imaging range of the imaging unit 12104 provided on the rear bumper or back door. For example, by overlaying the image data captured by the imaging units 12101 to 12104, a bird's-eye view image of the vehicle 12100 viewed from above can be obtained.

[0222] At least one of the imaging units 12101 to 12104 may have a function of acquiring distance information. For example, at least one of the imaging units 12101 to 12104 may be a stereo camera made up of multiple imaging elements, or may be an imaging element having pixels for phase difference detection.

[0223] For example, the microcomputer 12051 can calculate the distance to each three-dimensional object within the imaging ranges 12111 to 12114 and the change in this distance over time (relative speed with respect to the vehicle 12100) based on the distance information obtained from the imaging units 12101 to 12104, thereby extracting as a preceding vehicle, in particular, the three-dimensional object that is the closest three-dimensional object on the path of the vehicle 12100 and traveling in approximately the same direction as the vehicle 12100 at a predetermined speed (for example, 0 km / h or higher). Furthermore, the microcomputer 12051 can set a vehicle-to-vehicle distance to be maintained in advance in front of the preceding vehicle, and perform automatic braking control (including follow-up stop control), automatic acceleration control (including follow-up start control), etc. In this way, cooperative control can be performed for the purpose of automatic driving, which runs autonomously without relying on driver operation.

[0224] For example, the microcomputer 12051 classifies and extracts three-dimensional object data regarding three-dimensional objects into two-wheeled vehicles, ordinary vehicles, large vehicles, pedestrians, utility poles, and other three-dimensional objects based on distance information obtained from the imaging units 12101 to 12104, and can use the data for automatic obstacle avoidance. For example, the microcomputer 12051 distinguishes obstacles around the vehicle 12100 into obstacles that are visible to the driver of the vehicle 12100 and obstacles that are difficult to see. The microcomputer 12051 then determines the collision risk, which indicates the degree of risk of collision with each obstacle, and when the collision risk is equal to or greater than a set value and a collision is possible, the microcomputer 12051 can provide driving assistance for collision avoidance by outputting an alarm to the driver via the audio speaker 12061 or the display unit 12062, or by performing forced deceleration or avoidance steering via the drivetrain control unit 12010.

[0225] At least one of the image capturing units 12101 to 12104 may be an infrared camera that detects infrared rays. For example, the microcomputer 12051 can recognize a pedestrian by determining whether or not a pedestrian is present in the images captured by the image capturing units 12101 to 12104. The pedestrian recognition is performed, for example, by extracting feature points from the images captured by the image capturing units 12101 to 12104, which are infrared cameras, and then performing pattern matching on a series of feature points that indicate the outline of an object to determine whether or not the object is a pedestrian. When the microcomputer 12051 determines that a pedestrian is present in the images captured by the image capturing units 12101 to 12104 and recognizes the pedestrian, the audio / image output unit 12052 controls the display unit 12062 to superimpose a rectangular outline on the recognized pedestrian for emphasis. The audio / image output unit 12052 may also control the display unit 12062 to display an icon or the like indicating the pedestrian at a desired position.

[0226] An example of a vehicle control system to which the technology according to the present disclosure can be applied has been described above. The technology according to the present disclosure can be applied to the imaging unit 12031 and the vehicle exterior information detection unit 12030 among the configurations described above. Specifically, for example, the sensor unit 10 of the information processing device 1 is applied to the imaging unit 12031, and the recognition processing unit 12 is applied to the vehicle exterior information detection unit 12030. The recognition result output from the recognition processing unit 12 is passed to the integrated control unit 12050 via, for example, the communication network 12001.

[0227] In this way, by applying the technology of the present disclosure to the imaging unit 12031 and the outside vehicle information detection unit 12030, it is possible to recognize both close-range objects and long-range objects, and it is also possible to recognize close-range objects with high simultaneity, thereby enabling more reliable driving assistance.

[0228] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0229] The present technology can be configured as follows:

[0230] (1) a readout unit that sets a readout unit as a part of a pixel area in which a plurality of pixels are arranged in a two-dimensional array and controls the readout of pixel signals from pixels included in the pixel area; a reliability calculation unit that calculates the reliability of a predetermined region within the pixel region based on at least one of an area, a readout count, a dynamic range, and exposure information of the region of the captured image that is set as the readout unit and read out; An information processing device comprising:

[0231] (2) The reliability calculation unit calculates a correction value of the reliability for each of the plurality of pixels based on at least one of an area of ​​a region of the captured image, a number of times read out, a dynamic range, and exposure information, and generates a reliability map in which the correction values ​​are arranged in a two-dimensional array. The information processing device according to (1) further comprises:

[0232] (3) The reliability calculation unit includes a correction unit that corrects the reliability based on a correction value of the reliability, The information processing device according to (1) or (2), further comprising:

[0233] (4) The information processing device according to (3), wherein the correction unit corrects the reliability in accordance with a representative value of the correction value based on the predetermined region.

[0234] (5) The electronic device according to (1), wherein the reading unit reads out the pixels included in the pixel area as line-shaped image data.

[0235] (6) The information processing device according to (1), wherein the reading unit reads out the pixels included in the pixel area as sampled image data in a grid or checkerboard pattern.

[0236] (7) a recognition processing execution unit that recognizes an object within the predetermined area, The information processing device according to (1) further comprises:

[0237] (8) The information processing device according to (4), wherein the correction unit calculates a representative value of the correction value based on a receptive field in which the feature amount within the specified region is calculated.

[0238] (9) The reliability map generating unit generates at least two types of reliability maps based on at least two pieces of information selected from the area, the number of readouts, the dynamic range, and the exposure information; a synthesis unit that synthesizes the at least two types of reliability maps, The information processing device according to (2) further comprises:

[0239] (10) The information processing device according to (1), wherein the predetermined area within the pixel area is an area corresponding to at least one of a label and a category associated with each pixel by semantic segmentation.

[0240] (11) a sensor unit in which a plurality of pixels are arranged in a two-dimensional array; An information processing system comprising: The recognition processing unit a readout unit that sets a readout pixel as a part of a pixel area of ​​the sensor unit and controls reading of pixel signals from pixels included in the pixel area; a recognition processing unit having a reliability calculation unit that calculates the reliability of a predetermined region within the pixel region based on at least one of an area, a readout count, a dynamic range, and exposure information of the region of the captured image that is set as the readout unit and read out; An information processing system having:

[0241] (12) A readout step of setting a readout unit as a part of a pixel area in which a plurality of pixels are arranged in a two-dimensional array and controlling the readout of pixel signals from pixels included in the pixel area; and a reliability calculation step of calculating the reliability of a predetermined area within the pixel area based on at least one of the area, the number of times readout has been performed, the dynamic range, and exposure information of the area of ​​the captured image set as the readout unit and read out. An information processing method comprising:

[0242] (13) The recognition processing unit executes a readout step of setting a readout unit as a part of a pixel area in which a plurality of pixels are arranged in a two-dimensional array and controlling the readout of pixel signals from pixels included in the pixel area; a reliability calculation step of calculating the reliability of a predetermined region within the pixel region based on at least one of an area, a readout count, a dynamic range, and exposure information of the region of the captured image set as the readout unit and read out; A program that causes a computer to execute the following. [Explanation of symbols]

[0243] 1: information processing system, 2: information processing device, 10: sensor unit, 12: recognition processing unit, 110: readout unit, 124: recognition processing execution unit, 125: reliability calculation unit, 126: reliability map generation unit, 127: score correction unit.

Claims

1. a readout unit that sets a readout unit as a part of a pixel area in which a plurality of pixels are arranged in a two-dimensional array and controls readout of pixel signals from pixels included in the pixel area; a recognition processing execution unit that performs recognition processing using DNN on image data read from the pixel area of ​​the read unit; a reliability calculation unit that calculates an evaluation value based on the recognition result by the DNN as a reliability of the recognition result of the recognition processing; Equipped with the reliability calculation unit calculates a correction value of the reliability for each of the plurality of pixels based on at least one of an area of ​​the image data region, the number of times it has been read out, a dynamic range, and exposure information which is the number of exposures in the case of multiple exposure, and generates a reliability map in which the correction values ​​are arranged in a two-dimensional array.

2. a readout unit that sets a readout unit as a part of a pixel area in which a plurality of pixels are arranged in a two-dimensional array and controls readout of pixel signals from pixels included in the pixel area; a recognition processing execution unit that performs recognition processing using DNN on image data read from the pixel area of ​​the read unit; a reliability calculation unit that calculates an evaluation value based on the recognition result by the DNN as a reliability of the recognition result of the recognition processing; Equipped with The information processing apparatus, wherein the reliability calculation unit further includes a correction unit that corrects the reliability based on a correction value of the reliability.

3. The information processing device according to claim 2 , wherein the correction unit corrects the reliability in accordance with a representative value of the correction values ​​based on a region of the image data.

4. The information processing device according to claim 1 , wherein the reading unit reads out the pixels included in the pixel region as line-shaped image data.

5. The information processing device according to claim 1 , wherein the reading unit reads out the pixels included in the pixel area as sampled image data in a grid or checkerboard pattern.

6. The information processing apparatus according to claim 1 , wherein the recognition processing is processing for recognizing an object in the image data.

7. The information processing device according to claim 3 , wherein the correction unit calculates a representative value of the correction value based on a receptive field for which the feature amount in the image data has been calculated.

8. the reliability map generation unit generates at least two types of reliability maps based on at least two pieces of information among an area of ​​the image data region, a number of times read out, a dynamic range, and exposure information which is a number of exposures in multiple exposures; a synthesis unit that synthesizes the at least two types of reliability maps, The information processing device according to claim 1 , further comprising:

9. The information processing device according to claim 1 , wherein the recognition processing is processing for associating at least one of a label and a category with each pixel by semantic segmentation.

10. a sensor unit in which a plurality of pixels are arranged in a two-dimensional array; An information processing system comprising: The recognition processing unit a readout unit that sets a readout unit as a part of a pixel area of ​​the sensor unit and controls reading of pixel signals from pixels included in the readout unit; a recognition processing execution unit that performs recognition processing using DNN on image data read from the pixel area of ​​the read unit; a reliability calculation unit that calculates an evaluation value based on the recognition result by the DNN as a reliability of the recognition result of the recognition processing; and the reliability calculation unit calculates a correction value of the reliability for each of the plurality of pixels based on at least one of an area of ​​the image data region, the number of times it has been read out, a dynamic range, and exposure information which is the number of exposures in the case of multiple exposure, and generates a reliability map in which the correction values ​​are arranged in a two-dimensional array.

11. a sensor unit in which a plurality of pixels are arranged in a two-dimensional array; An information processing system comprising: The recognition processing unit a readout unit that sets a readout unit as a part of a pixel area of ​​the sensor unit and controls reading of pixel signals from pixels included in the readout unit; a recognition processing execution unit that performs recognition processing using DNN on image data read from the pixel area of ​​the read unit; a reliability calculation unit that calculates an evaluation value based on the recognition result by the DNN as a reliability of the recognition result of the recognition processing; and The information processing system, wherein the reliability calculation unit further includes a correction unit that corrects the reliability based on a correction value of the reliability.

12. a readout step of setting a readout unit as a part of a pixel area in which a plurality of pixels are arranged in a two-dimensional array and controlling the readout of pixel signals from pixels included in the pixel area; a recognition step of performing recognition processing using a DNN on the image data read from the pixel area of ​​the read unit; a reliability calculation step of calculating an evaluation value based on the recognition result by the DNN as a reliability of the recognition result of the recognition processing; Equipped with the reliability calculation step further includes a reliability map generation step of calculating a correction value of the reliability for each of the plurality of pixels based on at least one of exposure information, which is an area of ​​the region of the image data, the number of times it has been read out, a dynamic range, and the number of exposures in the case of multiple exposure, and generating a reliability map in which the correction values ​​are arranged in a two-dimensional array.

13. a readout step of setting a readout unit as a part of a pixel area in which a plurality of pixels are arranged in a two-dimensional array and controlling the readout of pixel signals from pixels included in the pixel area; a recognition step of performing recognition processing using a DNN on the image data read from the pixel area of ​​the read unit; a reliability calculation step of calculating an evaluation value based on the recognition result by the DNN as a reliability of the recognition result of the recognition processing; Equipped with The information processing method, wherein the reliability calculation step further includes a correction step of correcting the reliability based on a correction value of the reliability.

14. The recognition processing unit executes a readout step of setting a readout unit as a part of a pixel area in which a plurality of pixels are arranged in a two-dimensional array and controlling the readout of pixel signals from pixels included in the pixel area; a recognition step of performing recognition processing using a DNN on the image data read from the pixel area of ​​the read unit; a reliability calculation step of calculating an evaluation value based on the recognition result by the DNN as a reliability of the recognition result of the recognition processing; the reliability calculation step further includes a reliability map generation step of calculating a correction value of the reliability for each of the plurality of pixels based on at least one of the area of ​​the region of the image data, the number of times read out, the dynamic range, and exposure information which is the number of exposures in multiple exposure, and generating a reliability map in which the correction values ​​are arranged in a two-dimensional array; A program that causes a computer to execute the following.

15. The recognition processing unit executes a readout step of setting a readout unit as a part of a pixel area in which a plurality of pixels are arranged in a two-dimensional array and controlling the readout of pixel signals from pixels included in the pixel area; a recognition step of performing recognition processing using a DNN on the image data read from the pixel area of ​​the read unit; a reliability calculation step of calculating an evaluation value based on the recognition result by the DNN as a reliability of the recognition result of the recognition processing; The reliability calculation step further includes a correction step of correcting the reliability based on a correction value of the reliability; A program that causes a computer to execute the following.

Citation Information

Patent Citations

  • Method and device for autofocusing

    JP2000155257A

  • Image process device, image process method and image process program

    JP2013235304A

  • Imaging apparatus and method

    JP2017112409A

  • Image recognition device, learning device, image recognition method, learning method and program

    JP2019012426A

  • Solid-state imaging apparatus and imaging apparatus

    JP2020068483A