Imaging apparatus and control method of the same

JP2024012828A5Active Publication Date: 2025-07-24CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022114572
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-07-24
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

Existing imaging devices struggle with uniform brightness adjustment across the entire image, failing to ensure appropriate brightness for both machine recognition areas and human-visual recognition areas, and lack region-specific AE transition settings.

Method used

An imaging device with pixel-level control over exposure conditions, including independent setting of exposure time and analog gain for each pixel or pixel block, and a system to determine and apply different AE transition settings based on the content and type of image recognition, prioritizing either analog gain or exposure time depending on the recognition process.

Benefits of technology

Enables accurate and high-visibility image recognition by optimizing brightness and reducing noise or motion blur in specific image regions, improving detection accuracy for both machine and human-visual recognition tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an imaging apparatus which can perform proper image recognition.SOLUTION: An imaging apparatus comprises: imaging means which has an exposure region in which an exposure condition can be set for a single or each of a plurality of pixels; acquisition means which acquires an image recognition region in which image recognition by image processing is performed; decision means which decides transition setting of the exposure condition applied to the image recognition region on the basis of the content of the image recognition; and imaging control means which performs imaging by the imaging means by using the transition setting of the exposure condition decided by the decision means.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an imaging apparatus and a control method thereof. [Background technology]

[0002] Conventionally, cameras for surveillance use are required to have AE (Auto Exposure) transition settings for surveillance that are highly visible to users. The AE transition settings for surveillance are settings that balance the exposure time that does not cause noticeable motion blur of the subject and the analog gain that does not cause little noise, and transitions to an appropriate exposure brightness. In contrast, there is a setting that transitions to an exposure that is easy for the camera to process for image recognition, and cameras with such settings are called cameras for machine recognition use. In cameras for machine recognition use, the AE transition settings may be set to an extreme balance between exposure time and analog gain to one side in order to improve the accuracy of machine recognition (image recognition). For example, when detecting edges from an image, the exposure time is set to be short in order to eliminate motion blur as much as possible, and the analog gain is set to be high in order to compensate for the exposure. Also, when comparing images using background subtraction, the analog gain is set to be low in order to suppress the amount of noise, and the exposure time is set to be long in order to compensate for the exposure. In this way, the AE transition settings for surveillance and the AE transition settings for machine recognition are different.

[0003] Furthermore, the machine recognition may be performed on a part of the captured image rather than the entire image. Patent Document 1 discloses a technique for determining exposure based on the result of detection of a subject in machine recognition. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2020-72469 A Summary of the Invention [Problem to be solved by the invention]

[0005] However, in Patent Document 1, the brightness of the entire image is changed uniformly, so that although the brightness is appropriate for the area where machine recognition is performed, the brightness is not necessarily appropriate for other areas (for example, the area that the user wants to visually recognize). In addition, in Patent Document 1, the brightness is changed uniformly (for the entire image), so there is no mention of AE transition settings for each area. In view of the above-mentioned problems, the present invention provides an imaging device capable of performing appropriate image recognition. [Means for solving the problem]

[0006] In order to solve the above problem, an imaging device according to one aspect of the present invention comprises an imaging means having an exposure area in which exposure conditions can be set for each single pixel or for each multiple pixels, an acquisition means for acquiring an image recognition area in which image recognition is performed by image processing, a determination means for determining a transition setting of the exposure conditions to be applied to the image recognition area based on the content of the image recognition, and an imaging control means for performing imaging by the imaging means using the transition setting of the exposure conditions determined by the determination means. Effect of the Invention

[0007] According to the present invention, appropriate image recognition can be performed. [Brief description of the drawings]

[0008] [Figure 1A] 1 is a diagram showing an example of the functional configuration of an imaging apparatus according to a first embodiment of the present invention. [Figure 1B] FIG. 2 is a diagram showing an example of a hardware configuration of the imaging apparatus in FIG. 1A. [Diagram 2] FIG. 1 is a diagram showing an example of the configuration of an imaging system according to a first embodiment. [Diagram 3] 5 is a flowchart showing an example of the operation of the imaging apparatus according to the first embodiment. [Figure 4] FIG. 2 is a view showing an example of a user's operation screen according to the first embodiment. [Diagram 5]FIG. 13 is a diagram showing an example of AE transition settings applied to a non-machine recognized region. [Figure 6] FIG. 13 is a diagram showing an example of AE transition settings applied to a machine recognition region for performing temporal recognition processing. [Figure 7] FIG. 13 is a diagram showing an example of AE transition settings to be applied to a machine recognition area where spatial recognition processing is performed. [Figure 8] 10 is a flowchart showing an example of the operation of an imaging system according to a second embodiment. [Figure 9A] FIG. 13 is a diagram showing an example of the arrangement of a client device according to the third embodiment. [Figure 9B] FIG. 9B is a diagram showing an example of the hardware configuration of the client device in FIG. 9A. [Figure 10] 13 is a flowchart showing an example of the operation of the imaging system according to the third embodiment. [Figure 11] 10 is a flowchart showing another example of the operation of the imaging system according to the third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Hereinafter, the embodiment for carrying out the present invention will be described in detail with reference to the attached drawings. The embodiment described below is an example of a means for realizing the present invention, and should be appropriately modified or changed depending on the configuration of the device to which the present invention is applied and various conditions, and the present invention is not limited to the following embodiment. In addition, a configuration may be made by appropriately combining parts of each embodiment described later.

[0010] <First embodiment> FIG. 1A is a block diagram illustrating a functional configuration of an imaging device 100 according to a first embodiment of the present invention. The imaging device 100 has an imaging optical system 101, an imaging unit 102, a system control unit 103, a machine recognition area acquisition unit 104, a priority determination unit 105, an AE control unit 106, an encoder unit 107, a network I / F 108, and a memory 109. AE stands for Auto Exposure. The imaging unit 102 has an imaging element 102a, an amplifier unit 102b, and an image processing unit 102c. The memory 109 includes a volatile memory such as an SRAM or DRAM, and a non-volatile memory such as a flash memory. SRAM stands for Static Random Access Memory. DRAM stands for Dynamic Random Access Memory. The imaging device 100 is, for example, a surveillance camera.

[0011] The imaging optical system 101 collects light from a subject onto a light receiving surface of an image sensor 102a. The imaging optical system 101 includes one or more lenses. For example, the imaging optical system 101 includes a zoom lens, a focus lens, and a blur correction lens.

[0012] The imaging unit 102 captures an image of a subject and generates an image. The imaging element 102a converts light from a subject collected on an imaging surface (light receiving surface) by the imaging optical system 101 into an electrical signal for each pixel and outputs the signal. The imaging element 102a has an exposure area in which the exposure time and analog gain can be set (changed) independently for each pixel or pixel block on the imaging surface. The pixel block here is a set of pixels consisting of one or more pixels, and each pixel block can have a different exposure time or analog gain. The pixel block may be composed of a single pixel or a plurality of pixels. The shape of the pixel block does not need to be rectangular, and may be any shape. In this embodiment, the pixel block is described as being rectangular (block-shaped). The imaging element 102a is an IC chip in which pixels consisting of photoelectric conversion elements are arranged in a matrix. For example, the imaging element 102a is a CCD (Charge Coupled Device) sensor or a CMOS (Complementary Metal Oxide Semiconductor) sensor. The image sensor 102a has high sensitivity mainly to visible light, and each pixel has high sensitivity to either red (R), green (G), or blue (B), but also has a certain degree of sensitivity to infrared light, so that it can clearly capture images of subjects illuminated by sunlight and subjects in places illuminated by infrared lighting.

[0013] The amplifier 102b amplifies and outputs the electrical signal output from the image sensor 102a. The signal amplification factor (analog gain) of the amplifier 102b can be set and changed for each pixel or pixel block of the image sensor 102a.

[0014] The image processing unit 102c performs A / D conversion of the electrical signal, which is an analog signal output from the amplifier unit 102b, into a digital signal. A / D is an abbreviation of Analog / Digital. The image processing unit 102c performs signal processing including demosaicing processing, white balance processing, gamma processing, etc. on the digital signal obtained by the A / D conversion to generate a digital image. The image processing unit 102c also performs brightness correction by amplifying or attenuating the digital value of the image signal corresponding to each pixel or pixel block based on the analog gain of each pixel or pixel block. The generated digital image is temporarily stored in the memory 109. At this time, the image processing unit 102c outputs an image file or a video file in a predetermined format such as JPEG, H.264, H.265, etc. to the memory 109. The memory 109 also stores teacher data used in the machine recognition process. Furthermore, the image processing unit 102c performs machine recognition processing of the image (image recognition processing by a camera). There are roughly two types of machine recognition processing of the image. One is processing that performs machine recognition of the image based on temporal changes in luminance (referred to as "temporal recognition processing"). The other is processing that performs machine recognition of the image based on spatial changes in luminance (referred to as "spatial recognition processing"). The two types of machine recognition processing will be described in detail later.

[0015] In FIG. 1A, the amplifier 102b and the image processor 102c are shown as separate from the image sensor 102a, but the image sensor 102a may include some or all of the amplifier 102b and the image processor 102c. 1A, the imaging device 100 has the imaging optical system 101, but the imaging optical system 101 may be detachably provided in the imaging device 100 like an interchangeable lens. In other words, the imaging device 100 does not always need to have the imaging optical system 101 as a component thereof.

[0016] The system control unit 103 includes a CPU (see FIG. 1B), and performs overall control of each component of the imaging device 100 and sets various parameters. The system control unit 103 performs overall control so that each component of the imaging device 100 operates in cooperation with each other. For example, the system control unit 103 controls the imaging optical system 101. More specifically, the system control unit 103 performs optical control of the optical zoom magnification, F-number, focal length, etc. The system control unit 103 performs imaging using the imaging element 102a, and may therefore be referred to as an imaging control unit.

[0017] The machine recognition area acquisition unit 104 determines pixels or pixel blocks for machine recognition and outputs them as a machine recognition area (image recognition area). The user can manually specify the machine recognition area via the network I / F 108. The imaging device 100 may automatically specify (automatically set) the machine recognition area based on the results of machine recognition. For example, the machine recognition area acquisition unit 104 of the imaging device 100 performs moving object detection in the entire image or a partial area of ​​the image, and automatically sets the area where the moving object is detected as the machine recognition area. Furthermore, the machine recognition area acquisition unit 104 performs human body identification to identify whether the moving object is a human body in the machine recognition area.

[0018] The priority determination unit 105 determines the priority of the exposure time and analog gain of the imaging unit 102. The longer the exposure time, the higher the exposure amount, but the greater the motion blur of the subject or imaging device 100. The higher the analog gain, the higher the exposure amount, but the greater the noise. The priority determination unit 105 sets different priorities for the exposure time and analog gain for non-machine recognized areas (areas recognized by humans) and machine recognized areas.

[0019] Next, the priority will be described. In the non-machine recognition area, since no mechanical recognition processing is performed, the exposure time and analog gain are each gradually transitioned in AE (automatic exposure) so that the user feels less discomfort when viewing with the human eye. In other words, in the non-machine recognition area, the AE transition is set with an emphasis on human visibility. Therefore, the exposure time and analog gain are set with the same or similar priorities. In this embodiment, AE is the amount of exposure determined by the exposure time and analog gain. In the machine recognition area, the exposure time and analog gain are prioritized according to the type (content) of machine recognition, and the AE is transitioned so that the exposure amount is appropriate. In the temporal recognition process, the AE transition is performed with priority given to the analog gain. In other words, in the temporal recognition process, after the analog gain is determined and fixed, the AE transition (adjustment of the exposure amount) is performed according to the exposure time. In spatial recognition processing, the exposure time is given priority when transitioning the AE. In other words, in spatial recognition processing, after the exposure time is determined and fixed, the AE transition (adjustment of the exposure amount) is performed by the analog gain.

[0020] The AE control unit 106 determines and controls the exposure conditions based on the luminance of each pixel or the luminance of each pixel block and the priority of the priority determination unit 105. For example, the exposure amount (exposure time + analog gain) is determined so that the average value of the luminance of each pixel or the luminance of each pixel block becomes the median value of the data gradation that can be output. Furthermore, after determining one of the exposure time and analog gain based on the priority of the priority determination unit 105, the exposure amount is adjusted to the amount determined by the other. The exposure of the image sensor 102a is controlled by the determined exposure time and analog gain. That is, the image sensor 102a captures an image using the transition setting of the determined exposure condition.

[0021] The encoder unit 107 encodes the image data processed by the image processing unit 102c into a predetermined file format such as Motion JPEG, H264, or H265 (performs encoding processing). The network I / F 108 is an interface used for communication with an external information processing device (e.g., the client device 210 in FIG. 2) or a storage device (e.g., the server 220 in FIG. 2) via the network 110. Image data encoded by the encoder unit 107 can be transmitted to the information processing device (client device 210) or the storage device (server 220) using the network I / F 108. The network I / F 108 receives a designation of an area for machine recognition and a designation of a type of machine recognition from the information processing device (client device 210). The network I / F 108 also receives control commands for panning, tilting, and zooming the imaging device 100 from the information processing device (client device 210).

[0022] The network 110 is, for example, a local area network (LAN) and is composed of a router conforming to a communication standard such as Ethernet (registered trademark). The imaging device 100, the information processing device (client device 210), and the storage device (server 220) are connected by a LAN cable or the like. The network 110 may include a wireless network such as the Internet. In this case, the imaging device 110, the information processing device (client device 210), and the storage device (server 220) may be connected by a wireless network.

[0023] Next, we will explain two types of machine recognition processing and AE transition settings. The first machine recognition process calculates the change in brightness over time. A series of multiple images are captured with the same angle of view, and the amount of change in brightness is calculated for each image or region within the image. If the amount of change in brightness over time exceeds a predetermined value (i.e., if the change in brightness in the region exceeds a predetermined amount), it can be assumed that a moving object is present. Therefore, accuracy is improved by suppressing fluctuations due to noise. Specific functions of the imaging device (camera) 100 included in the first machine recognition process include a moving object detection function that determines whether a moving object is present in an image, a detection function that detects when an object is brought into a specific region, and a detection function that detects when an object is taken away.

[0024] The second machine recognition process calculates spatial changes in luminance. The amount of spatial change in luminance in the horizontal or vertical direction is calculated for a captured image. When the amount of spatial change in luminance exceeds a predetermined value, the boundary can be detected as an edge (contour) that is a discontinuous point. It is possible to calculate feature points based on the shape of this edge. Therefore, accuracy is improved by suppressing fluctuations due to the movement of the subject. Specific functions of the imaging device (camera) 100 include human body identification, which identifies whether a person is a human body from the shape of the shoulders and head, and face recognition, which compares the positional relationship or shape of parts of the face such as the eyes, nose, and mouth with training data to evaluate the degree of match, which are included in the second machine recognition process. Training data is data on feature points that have been saved in advance. It is possible to improve accuracy by increasing the training data (number of samples) through machine learning.

[0025] The first machine recognition process (temporal recognition process) and the second machine recognition process (spatial recognition process) have different prioritized imaging conditions. Specifically, in temporal recognition process, analog gain is prioritized over exposure time, while in spatial recognition process, exposure time is prioritized over analog gain.

[0026] In temporal recognition processing, accuracy is improved by capturing images with low analog gain (low noise) and suppressing unintended brightness changes (noise) between images. In order to set the analog gain low, it is necessary to set the exposure time long, but in temporal recognition processing, spatial brightness changes, i.e., motion blur, do not have a significant effect on accuracy, so it is possible to prioritize analog gain.

[0027] In spatial recognition processing, accuracy is improved by capturing images with a short exposure time (small motion blur) and suppressing unintended brightness changes in space (motion blur). This is because motion blur is difficult to suppress using digital processing. In order to set a short exposure time, it is necessary to set the analog gain high, but it is possible to suppress the effects of noise caused by analog gain by using digital processing such as a differential filter.

[0028] Each functional block (reference numerals 102b, 102c, 103 to 108) shown in FIG. 1A is realized by software. More specifically, a program for providing the function of each functional block of FIG. 1A is stored in a memory such as a ROM (Read Only Memory) 124 shown in FIG. 1B. Then, the program is read into a RAM (Random Access Memory) 125 shown in FIG. 1B and executed by a CPU (Central Processing Unit) 123, thereby realizing the function. Note that some or all of the functional blocks may be realized by hardware. In this case, for example, a specific compiler may be used to automatically generate a dedicated circuit on an FPGA from a program for realizing the function of each functional block. FPGA is an abbreviation for Field Programmable Gate Array. Also, a gate array circuit may be formed in the same manner as an FPGA, and the function may be realized as hardware. Also, the function may be realized by an ASIC (Application Specific Integrated Circuit). Note that the functional block configuration shown in FIG. 1A is just one example, and multiple functional blocks may be configured to form one functional block, or any functional block may be divided into blocks performing multiple functions.

[0029] FIG. 1B shows an example of a hardware configuration of the imaging device 100. The imaging device 100 has an imaging optical system 121, an imaging element 122, a CPU 123, a ROM 124, a RAM 125, an imaging system control unit 126, a communication control unit 127, an A / D conversion unit 128, an image processing unit 129, an encoder unit 130, and a network I / F 131. The various units (123 to 131) of the imaging device 100 are interconnected by a system bus 132. The ROM 124 includes a flash memory. The RAM 125 includes an SRAM or DRAM. The ROM 124 and the RAM 125 correspond to the memory 109 in FIG. 1A.

[0030] The imaging optical system 121 is a group of optical members that includes a zoom lens, a focus lens, a blur correction lens, an aperture, a shutter, etc., and collects optical information of a subject. The imaging optical system 121 is connected to an image sensor 122. The imaging optical system 121 corresponds to the imaging optical system 101 in FIG. 1A. The imaging element 122 is a charge storage type solid-state imaging element such as a CMOS or a CCD that converts the light beam collected by the imaging optical system 121 into a current value (signal value). The imaging element 122 corresponds to the imaging element 102a in FIG. 1A. The imaging element 122 is connected to an A / D conversion unit 128.

[0031] The CPU 123 is a control unit that performs overall control of the operation of the imaging device 100. The CPU 123 reads instructions stored in the ROM 124 or the RAM 125, and executes processing according to the results. The CPU 123 corresponds to the system control unit 103 in FIG. 1A. The imaging system control unit 126 controls each unit of the imaging device 100 based on instructions from the CPU 123. For example, the imaging system control unit 126 controls the imaging optical system 121 in terms of focus control, shutter control, aperture adjustment, and the like. The imaging system control unit 126 corresponds to the system control unit 103 in FIG. 1A. The communication control unit 127 performs control for transmitting control commands (control signals) to each unit of the imaging device 100 from the client device 210 to the CPU 123 through communication with the client device 210. The communication control unit 127 corresponds to the system control unit 103 in FIG. 1A.

[0032] The A / D conversion unit 128 converts the current value received from the imaging element 122 into a digital signal (image data). The A / D conversion unit 128 transmits the digital signal to an image processing unit 129. The A / D conversion unit 128 corresponds to the image processing unit 102c in FIG. 1A. The image processing unit 129 performs image processing on the image data of the digital signal received from the A / D conversion unit 128. The image processing unit 129 is connected to the encoder unit 130. The image processing unit 129 corresponds to the image processing unit 102c in FIG. 1A. The encoder unit 130 converts the image data processed by the image processing unit 129 into a file format such as Motion Jpeg, H264, or H265. The encoder unit 130 is connected to a network I / F 131. The encoder unit 130 corresponds to the encoder unit 107 in FIG. 1A. The network I / F 131 is an interface used for communication with an external device such as the client device 210 via the network 110, and is controlled by the communication control unit 127. The network I / F 131 corresponds to the network I / F 108 in FIG. 1A.

[0033] The configuration of an imaging system 200 including the imaging device 100 will be described with reference to Fig. 2. Fig. 2 shows an example of the configuration of the imaging system 200. In the imaging system 200, the imaging device 100 is connected to a client device 210 and a server 220 via a network 110. The client device 210 is an information processing device such as a personal computer. The client device 210 is connected to a display device 201 and an input device 202 by wire or wirelessly. The display device 201 has a display and is a device that displays images and a user operation screen (GUI). The display is, for example, a liquid crystal display. GUI is an abbreviation for Graphic User Interface. The input device 202 includes a mouse and a keyboard, and the user can operate the input device 202 while looking at the screen of the display device 201. The GUI may be considered to be a part of the input device 202. In FIG. 2, the display device 201 and the input device 202 are provided externally to the client device 210, but the client device 210 may have at least one of the display device 201 and the input device 202 built-in.

[0034] A process executed by the imaging device 100 will be described with reference to Fig. 3. Fig. 3 shows an example of the operation of the imaging device 100. Fig. 3 describes a case where a user specifies a machine recognition area. In S300, the user sets (specifies) a machine recognition area. For example, in the case of detecting a human body, the user specifies a passageway through which a person is expected to pass as the machine recognition area. The user instructs the image capture device 100 from the client device 210 via the network 110 about the machine recognition area (sends area specification information). The client device 210 is connected to the display device 201, and is capable of displaying an image on the display device 201. The display device 201 (display 400 in FIG. 4) has a user interface function that allows the user to operate the image captured by the image capture device 100 and the buttons and slides that are superimposed and displayed, by touching or dragging. The user sets the machine recognition area for the displayed image. The user interface for setting the machine recognition area will be described later with reference to FIG. 4.

[0035] In S301, the imaging device 100 acquires a machine recognition area (image recognition area). Based on an instruction for the machine recognition area obtained via the network 110, the imaging device 100 causes the machine recognition area acquisition unit 104 to acquire the machine recognition area. In S302, the imaging device 100 controls and determines the exposure conditions (analog gain and exposure time). In this embodiment, the CPU 123 of the imaging device 100 determines transition settings of the exposure conditions to be applied to the machine recognition area based on the details of the machine recognition. In this embodiment, unless otherwise specified, the exposure conditions refer to analog gain and exposure time. The average brightness value for each region is calculated to determine the exposure conditions. The difference between this average brightness value and the target brightness is calculated, and if there is a difference, the exposure time or analog gain is transitioned. In this embodiment, the priority of the exposure time and analog gain to be transitioned is different between the non-machine recognition region and the machine recognition region. The exposure conditions are determined (set) independently for each pixel block (single or multiple pixels). Furthermore, even if the brightness is the same, the ratio of analog gain to exposure time is different between the non-machine recognition region and the machine recognition region. Details of the AE transition setting will be described with reference to the exposure condition transition diagrams shown in Figures 5, 6, and 7.

[0036] In S303, imaging is performed under the determined exposure conditions. More specifically, the system control unit 103 of the imaging device 100 performs imaging using the imaging element 102a. In S304, image processing is performed on the video signal (video data) obtained by imaging to obtain luminance information. The luminance information is the luminance value for each pixel or pixel block, the amount of temporal luminance change, and the amount of spatial luminance change. When calculating the amount of temporal luminance change, the luminance of the video signal captured one or more frames ago is stored in memory and compared with the luminance of the current frame.

[0037] In S305, the image processing unit 102c performs machine recognition processing based on the luminance information obtained in S304. In S306, the image processing unit 102c performs development processing, whereby the video data is compressed into an image such as JPEG. In S307, the compressed image is delivered (transmitted) to the client device 210 by the network I / F . The client device 210 receives the compressed image and displays the received image on the display device 201.

[0038] A method for a user to set a machine recognition area will be described with reference to Fig. 4. Fig. 4 shows an example of a user operation screen. An example of a delivery image 401 and a function selection section 404 are shown on a display 400 that serves as the operation screen. The grid shown in the delivery image 401 is the boundary of the exposure blocks of the image sensor 102a. The exposure block 402 in the upper left of the delivery image 401 is shown by a diagonal line. The user can also select an area for machine recognition. The machine recognition area specified by the user is a user-specified area 403, which is shown by a dotted area in Fig. 4. The user-specified area 403 can be specified by dragging or clicking an area of ​​the image with a mouse.

[0039] The function selection section 404 is a user interface for selecting a machine recognition function. As an example, Fig. 4 shows buttons for removal detection 405, abandonment detection 406, moving body detection 407, human body detection 408, and face recognition 409. The function selection section 404 can be used to determine the details of machine recognition (image recognition).

[0040] Removal detection is a function that determines whether a stationary object in a specified area has moved, and performs the detection based on the change in brightness or edge over time. Abandonment detection is a function that determines whether a stationary object in a specified area has not moved, and performs the detection based on the change in brightness or edge over time. Removal detection and abandonment detection are determined based on edge or brightness information for stationary objects, so motion blur does not affect the detection accuracy. Therefore, the exposure time can be set to a long second, and low-noise images can be obtained by setting the analog gain low accordingly. This improves the detection accuracy. Another example of moving object detection is a method that calculates from the change in brightness over time (detects moving objects). In this case, the criterion for judgment is whether there is a change in brightness over time (movement), so motion blur does not affect the detection accuracy. Therefore, the exposure time can be set to a long second, and low-noise images can be obtained by setting the analog gain low accordingly. This improves the detection accuracy.

[0041] Human body detection detects edges characteristic of the human body (for example, head and shoulders) from an image, and compares them with training data to calculate an evaluation value based on their similarity. If they are similar, the evaluation value is high. If they are not similar, the evaluation value is low. Since the edges of a single image are detected, it is important that there is no motion blur. Therefore, shortening the exposure time improves the accuracy of edge detection.

[0042] Like human body detection, face recognition detects edges that are characteristic of the human body (such as eyes and noses) from an image and compares them with training data to calculate an evaluation value based on similarity. If there is similarity, the evaluation value is high. If there is dissimilarity, the evaluation value is low. Since the edges of a single image are detected, it is important that there is no motion blur. Therefore, shortening the exposure time improves the accuracy of edge detection.

[0043] The transition setting of AE according to the setting of the machine recognition area will be described with reference to FIG. 5, FIG. 6, and FIG. 7. In this embodiment, the transition diagram applied to the non-machine recognition area is called the balanced AE transition diagram (FIG. 5), the transition diagram applied to the temporal machine recognition area is called the analog gain priority AE transition diagram (FIG. 6), and the transition diagram applied to the spatial machine recognition area is called the exposure time priority AE transition diagram (FIG. 7). In FIG. 5 to FIG. 7, the horizontal axis indicates the analog gain, and the vertical axis indicates the exposure time. The numbers written on the horizontal axis and the vertical axis indicate the number of steps, and the amount of exposure doubles when the number increases by one. The sum of the number of steps of the analog gain and the number of steps of the exposure time is expressed as the number of steps of exposure=EV (Exposure Value). The relationship between exposure (EV), analog gain, and exposure time is shown in Equation (1). Moreover, the value determined by the analog gain and exposure time is shown in Equation (2) as the exposure amount X. EV=0, which is the standard for EV, is the brightness at which an object captured under conditions of ISO sensitivity (analog gain) of 100, exposure time of 1 second, and aperture of F1 will be properly exposed. In this embodiment, EV and X are treated as relative values ​​that indicate the change in exposure when the analog gain or exposure time is changed, with the aperture omitted. Exposure EV (steps) = - (analog gain (steps) + exposure time (steps)) ··· formula (1) Exposure amount X (steps) = -EV (steps) = Analog gain (steps) + Exposure time (steps) ...Equation (2)

[0044] That is, the exposure condition at the bottom left of the transition diagram (the intersection of the vertical and horizontal axes) is the sum of one step of analog gain and one step of exposure time, so X=2. The exposure condition at the top right of the transition diagram is, for example, the sum of nine steps of analog gain and nine steps of exposure time, so X=18. If the value of X is the same, the amount of exposure is the same regardless of the difference between the exposure condition and the analog gain. The arrows in the diagram indicate the transition direction of the exposure condition. To change X, the analog gain or exposure time is transitioned in the direction of the arrow. In this embodiment, the transition is made one step at a time to simplify the explanation, but the value that can be transitioned at one time is not limited to one step.

[0045] In the balanced AE transition diagram in Figure 5, the exposure time and analog gain are balanced to create an image with high visibility for the user. For example, to transition from X=2 to X=6, the analog gain and exposure time are alternately transitioned one step at a time, resulting in a total of four transitions (two analog gain steps, two exposure time steps). This makes it possible to obtain an image with a good balance between noise and motion blur.

[0046] Compared to the balanced AE transition diagram in Figure 5, the analog gain priority AE transition diagram (Figure 6) applied to the temporal machine recognition area and the exposure time priority AE transition diagram (Figure 7) applied to the spatial machine recognition area have different priorities between analog gain and exposure time. We will explain each of the analog gain priority AE transition diagram and the exposure time priority AE transition diagram.

[0047] In the analog gain priority AE transition diagram applied to the temporal machine recognition region shown in Figure 6, when X is low (X = 10 steps or less in the diagram), the analog gain is prioritized and fixed at a low gain (minimum 1 step in the diagram), and the exposure is adjusted according to the exposure time. This makes it possible to obtain an image with less noise. Therefore, the amount of change in brightness over time is not buried in noise, improving detection accuracy. However, in areas where X is large (X = 11 or more in the diagram), the value of X cannot be increased any further by the exposure time alone, so the exposure is adjusted according to the analog gain.

[0048] In Fig. 6, the analog gain is changed after the exposure time reaches its maximum, but the analog gain may be changed before the exposure time reaches its maximum value. For example, the analog gain may be fixed at 1 up to X=6, and fixed at 2 up to X=8, so that stepwise control (change) may be performed. However, the priority of the analog gain is set high with respect to the balanced AE transition diagram in Fig. 5.

[0049] In the analog gain priority AE transition diagram (exposure time priority AE transition diagram) applied to the spatial machine recognition area shown in Figure 7, when X is low (X = 10 steps or less in the diagram), the exposure time is prioritized and fixed at a short second (minimum 1 step in the diagram), and the exposure is adjusted by the analog gain. This makes it possible to obtain an image with less motion blur. As a result, the amount of spatial brightness change is not buried in the motion blur, and detection accuracy is improved. However, in areas where X is large (X = 11 or more in the diagram), the value of X cannot be increased by analog gain alone, so the exposure is adjusted by the exposure time.

[0050] In Fig. 7, the exposure time is changed after the analog gain reaches its maximum, but the exposure time may be changed before the analog gain reaches its maximum value. For example, stepwise control may be performed, such as fixing the exposure time at 1 up to X=6, and fixing the exposure time at 2 up to X=8. However, the priority of the exposure time is set high for the balanced AE transition diagram (Fig. 5) applied to the non-machine recognition area.

[0051] In this embodiment, based on the type (content) of machine recognition, it is determined whether the transition setting of the exposure condition to be applied to the machine recognition area is the transition setting in Fig. 6 or the transition setting in Fig. 7. The transition setting to be applied to the non-machine recognition area (area other than the image recognition area) is the transition setting in Fig. 5, which is an exposure condition with high visibility for the user (human). In this embodiment, by setting AE transition settings (the transition settings in FIG. 6 or the transition settings in FIG. 7) with different priorities for the areas set in the machine recognition area, it is possible to improve the detection accuracy of machine recognition while outputting an image with high visibility in the non-machine recognition area. Furthermore, by selecting and determining whether to give priority to analog gain or exposure time depending on the recognition processing method of machine recognition (the contents of machine recognition), a setting suitable for the type of machine recognition can be obtained, thereby improving the detection accuracy.

[0052] 6 and 7, the transition settings of the AE for the machine recognition area have been described. By modifying (changing, adjusting) the transition settings according to the accuracy (required accuracy) required for machine recognition, it is possible to improve visibility while maintaining the accuracy required for machine recognition. For example, the accuracy required for "1. When determining whether it is a human face" and "2. When determining the age of the human face" using face recognition (spatial recognition processing) is different. In this case, "2. When determining the age of the human face" requires more detailed data, so a higher detection accuracy is required. Therefore, in "2. When determining the age of the human face", as shown in FIG. 7, the exposure time setting is fixed at 1 for X=2 to 10, and imaging is performed to suppress motion blur of the subject. In contrast, "1. When determining whether it is a human face" does not require a higher accuracy (higher accuracy) than "2. When determining the age of the human face". Therefore, in the brightness where the detection accuracy is sufficient, the AE transition is performed with emphasis on visibility, and in the brightness where the detection accuracy is not sufficient, the AE transition is performed with emphasis on machine recognition. In such a case, X=1 to 6 is specifically described as the brightness where the detection accuracy is sufficient.

[0053] For brightness levels X=1 to 6 where detection accuracy is sufficient, the transition is performed in the same manner as the balanced AE transition in Figure 5 to improve visibility. For brightness levels X=6 to 12 where detection accuracy is insufficient, the exposure time is fixed at 3 to improve detection accuracy. Although the number of steps for fixing the exposure time is different from that in Figure 7, the exposure time is prioritized (fixed) as in Figure 7. For X=12 and above, since further adjustment by analog gain is not possible as in Figure 7, exposure is adjusted by the exposure time. In this way, by changing the priority depending on the required accuracy, it is possible to improve visibility while maintaining the required detection accuracy. Also, it is possible to modify (change, adjust) the AE transition settings in the same manner depending on the type of machine recognition, regardless of the accuracy required for machine recognition.

[0054] There are also cases where a new machine recognition area is set for a non-machine recognition area. For example, when a moving object is detected (recognized) in a non-machine recognition area, a new machine recognition area is set. In this case, it is desirable to transition the exposure time and analog gain so that the AE setting in the AE transition diagram after switching (machine recognition area) is achieved regardless of the value of X. In this case, the transition may be made all at once in one frame, or may be made gradually (in stages) by several steps. When transitioning all at once, the change is made with reference to the AE setting of the target X in the AE transition diagram. When transitioning gradually by several steps, it is desirable to transition the item with priority, either the analog gain or the exposure time, first. This makes it possible to improve the detection accuracy in a short time. In this way, when the type of area (non-machine recognition area, temporal machine recognition area, spatial machine recognition area) is changed, the exposure time and analog gain are transitioned with reference to the AE transition diagram after switching.

[0055] In this embodiment, the value of the exposure amount X (=-EV) is shown as the sum of the exposure time or analog gain. However, the amount of light incident on the image sensor 102 may change due to the aperture or ND (Neutral Density) filter in the imaging optical system 101. Therefore, it is desirable to correct the value of X (=-EV) based on the luminance information and the optical information of the imaging optical system 101 and transition the exposure time or analog gain.

[0056] In the above embodiment, an example was shown in which the analog gain and the exposure time change (transition) in one step in the AE transition diagram, but it is not necessary to change in one step. For example, the transition may be made in 1 / 3 steps or in two steps.

[0057] The presence or absence of a machine recognition area for changing the transition setting of the AE can be changed by the designer or user depending on the type of machine recognition, the position of the area, the imaging situation, etc. For example, when it is desired to improve the accuracy of only face recognition among a plurality of types and contents of machine recognition (for example, the five image recognitions 405 to 409 shown in the function selection unit 404 of FIG. 4), the transition setting of the AE of the machine recognition area may be changed only in the area where face recognition is performed. In addition, the change of the transition setting of the AE does not need to be applied to all areas where machine recognition is performed. For example, even if both moving object detection and human body identification are performed simultaneously in different areas, the transition setting of the AE may be changed only for one of them. In this case, it is desirable for the designer or user to select (decide) whether to set the machine recognition area depending on the type (content) of the machine recognition. This makes it possible to improve the accuracy of the machine recognition that the designer or user considers important.

[0058] In addition, when performing different machine recognition for the same region, such as performing person identification (human body detection) after moving object detection, the AE transition setting may be changed for each machine recognition. How to change the AE transition setting when performing human body detection (person detection) after moving object detection will be specifically described below.

[0059] In motion detection (temporal recognition processing), it is desirable to suppress noise to prevent false detection (Fig. 6). In contrast, in human body detection (spatial recognition processing) that is performed after motion detection, it is desirable to shorten the exposure time to accurately calculate feature points such as human body detection and face detection (Fig. 7). For this reason, the AE transition diagram with analog gain priority in Fig. 6 is applied during the motion detection period, and the AE transition diagram with exposure time priority in Fig. 7 is applied during the person identification period. This makes it possible to reflect multiple AE transition settings in the same area, although it is limited to different timing. In other words, the accuracy of machine recognition can be improved even when different types of area settings are performed in the same area.

[0060] <Second embodiment> In the first embodiment, a case where a user specifies a machine recognition area (S300 in FIG. 3) is described. In the second embodiment, a case where the imaging device 100 automatically specifies (sets) a machine recognition area is described. In the second embodiment, face recognition is given as a specific example. To perform face recognition, moving object detection is first performed. When a moving object is detected, a machine recognition area is set for the moving object (area) and face recognition is performed. In the following description, the same configurations and processes as those in the first embodiment are given the same reference symbols and detailed description is omitted.

[0061] A second embodiment of the present invention will be described with reference to Fig. 8. Fig. 8 is a flow chart showing the procedure for carrying out the processing of this embodiment. In S800, the imaging device 100 performs a preliminary image capture. The imaging luminance of the imaging device 100 is obtained by the preliminary image capture. Thereafter, in S801, it is determined whether or not the luminance change over time is equal to or greater than a predetermined amount based on the luminance at the time of the preliminary capture. More specifically, it is determined whether or not a moving object (machine recognition target) has been detected based on the luminance change over time. In order to detect a moving object, in S801, a temporal recognition process is used, and if the luminance change over time is equal to or greater than a predetermined amount, it is determined that the area contains a moving object. At this time, the detection of the moving object is performed over the entire image. Therefore, it is desirable to set the AE transition setting during the preliminary capture period to the AE transition diagram for temporal recognition process (Fig. 6). However, if the user is to visually confirm the image, it may be set to the balanced AE transition diagram (Fig. 5).

[0062] In S801, if a moving object is not detected, the preliminary imaging (S800) is repeated, whereas if a moving object is detected, the process proceeds to S802. In S802, the area where the moving object was detected is set as the machine recognition area for performing face authentication (spatial recognition processing). At this time, the setting is performed taking into consideration detection errors and the movement of the moving object. Specifically, it is desirable to set a large area including the periphery of the actually detected area as the machine recognition area for performing spatial recognition processing. After S802, the process proceeds to S302. 3, the process proceeds to S803.

[0063] In S803, it is determined whether face authentication (machine recognition) is completed. For example, if the evaluation value of face authentication is equal to or greater than a predetermined value (if the evaluation value is equal to or greater than 80 / 100 points), it is determined that face authentication is completed. If face authentication is completed, the process proceeds to S804, where the setting of the machine recognition area is cancelled, and the process of FIG. 8 is terminated. In S803, if the evaluation value is less than 80 / 100 points, the process proceeds to S805.

[0064] In S805, an error judgment (determination of whether or not there is an error) of face authentication is performed. For example, the number of images with an evaluation value of less than 80 / 100 points is counted, and it is determined whether the count reaches a predetermined number (for example, 10). Until the number of images with an evaluation value of less than 80 / 100 points reaches 10, the judgment result of S805 is No, and the process proceeds to S302 and the processes of S302 to S803 are repeated. If the evaluation value does not reach 80 / 100 points or more even when the number of images with an evaluation value of less than 80 / 100 points reaches 10, an error is judged (S805: Yes). After it is judged as an error, the process proceeds to S804 to cancel the setting of the machine recognition area, reset the count of the number of images, and end the process of FIG. 8.

[0065] Although not shown, after S804, the AE transition setting is returned to that at the time of preliminary imaging. The above describes the method of AE transition setting in the case where the machine recognition area is automatically set in this embodiment. In this manner, even in this embodiment, by setting the AE transition setting with a different priority for the area set in the machine recognition area, it is possible to improve the detection accuracy of machine recognition while outputting an image with high visibility in the non-machine recognition area. Furthermore, by determining which of the analog gain and the exposure time is to be prioritized according to the recognition processing method of machine recognition, a setting suitable for the type of machine recognition can be obtained, thereby improving the detection accuracy. The setting of the machine recognition area may be changed based on the recognition (detection) result by machine recognition.

[0066] <Third embodiment> In the first and second embodiments, the case where machine recognition is performed by the imaging device 100 has been described. In the third embodiment, the case where machine recognition is performed by the client device 210 will be described. When performing face recognition or the like, a plurality of feature points are calculated and compared with a large amount of training data, so high computing power is required. If the client device 210 has a higher computing power than the imaging device 100, it may be desirable to perform machine recognition by the client device 210. In the following description, the same configurations and processes as those in the first and second embodiments are denoted by the same reference symbols, and detailed description thereof will be omitted.

[0067] A third embodiment of the present invention will be described with reference to Fig. 9A to Fig. 11. Fig. 9A shows the functional configuration of a client device 210. The client device 210 includes a network I / F 901, a system control unit 902, an output I / F 903, an input I / F 904, an image processing unit 905, and a memory 906. The network I / F 901 is an interface that connects the network 110 and the client device 210, and performs input and output of data. The system control unit 902 controls each module. The output I / F 903 is an interface with the display device 201. The input I / F 904 is an interface with the input device 202. The memory 906 stores images and brightness information received from the imaging device 100. The memory 906 also stores teacher data used in face recognition, programs used by the system control unit 902, and the like.

[0068] The image processing unit 905 performs machine recognition based on the image or luminance information output from the imaging device 100. When machine recognition is performed based on luminance information before compression, the same accuracy as when machine recognition is performed by the imaging device 100 is obtained. In contrast, when machine recognition is performed based on a compressed image, data is compressed and the resolution is reduced, so the detection accuracy (recognition accuracy) is lower than when machine recognition is performed by the imaging device 100. Furthermore, since the parameters of the compression process may change for each frame, even if there is no change in the subject, a temporal change may occur in the compressed image. Therefore, when performing temporal recognition processing, it is desirable to set the parameters of the machine recognition process for the change in image brightness while taking into account the change due to compression. For example, in the case of detecting a moving object, when the change in image brightness over time is judged, it is judged whether it is larger than a predetermined reference value, but when the compression rate changes or is high, it is desirable to relax the predetermined value to prevent erroneous detection.

[0069] FIG. 9B is a block diagram showing an example of the hardware configuration of the client device 210. As shown in FIG. The client device 210 includes a client CPU 911, a main memory device 912, an auxiliary memory device 913, an input I / F 914, an output I / F 915, and a network I / F 916. The elements of the client device 210 are connected to each other via a system bus 917 so as to be able to communicate with each other.

[0070] The client CPU 911 is a central processing unit that performs overall control of the operation of the client device 210. Note that the client CPU 911 may perform overall control of the imaging device 100 via the network 110. The client CPU 911 corresponds to the system control unit 902 and the image processing unit 905 in FIG. 9A. The primary storage device 912 is a storage device such as a RAM that functions as a temporary storage location for data of the client CPU 911. For example, the primary storage device 912 stores in advance patterns for pattern matching (patterns corresponding to facial features and human body features) that are used when the client device 210 performs face detection or human body detection. The primary storage device 912 corresponds to the memory 906 in FIG. 9A. The auxiliary storage device 913 is a storage device such as an HDD, ROM, or SSD that stores various programs, various setting data, and the like. The auxiliary storage device 913 may also store a database (face recognition database) that associates pre-registered face images with personal information. HDD is an abbreviation for Hard Disk Drive. SSD is an abbreviation for Solid State Drive. The auxiliary storage device 913 also corresponds to the memory 906 in FIG. 9A.

[0071] The input I / F 914 is an interface used when the client device 210 receives an input (signal) from the input device 202 etc. The input I / F 914 corresponds to the input I / F 904 in FIG. The output I / F 915 is an interface that is used when outputting information (signals) from the client device 210 to the display device 201, etc. The output I / F 915 corresponds to the output I / F 903 in FIG. The network I / F 916 is an interface used for communication with external devices such as the image capture device 100 via the network 110. The network I / F 916 corresponds to the network I / F 901 in FIG. The client CPU 911 executes processing based on a program stored in the auxiliary storage device 913, thereby implementing the processing of the client device 210 (S1000 in FIG. 10 and S1100 in FIG. 11).

[0072] Next, the process executed by the client device 210 will be described with reference to Fig. 10 and Fig. 11. Fig. 10 shows an example of an operation procedure in which the user specifies the machine recognition area (Fig. 3) as in the first embodiment, but the machine recognition process is executed by the client device 210. Fig. 11 shows an example of an operation procedure in which the machine recognition area is automatically specified (Fig. 8) as in the second embodiment, but the machine recognition process is executed by the client device 210.

[0073] First, Fig. 10 will be described. S300 to S304 in Fig. 10 are the same as in Fig. 3. S305 does not exist in Fig. 10. S306 and S307 in Fig. 10 are also the same as in Fig. 3. In S307, the imaging device 100 delivers an image to the client device 210. The image is delivered (transmitted) to the image processing unit 905 via the network I / F 901 of the client device 210. In Fig. 10, S1000 is carried out after S307. In S1000, the image processing unit 905 of the client device 210 performs machine recognition based on the delivered (transmitted) image. At this time, when performing spatial recognition processing, arithmetic and processing for machine recognition is performed on the delivered image. When performing temporal recognition processing, images are stored in the memory 906 for each frame, and arithmetic and processing for machine recognition is performed. For example, when performing face recognition, the delivered image is compared with the teacher data stored in the memory 906 to calculate an evaluation value. In this way, the machine recognition processing (imaging device) in S305 of FIG. 3 can be performed by the client device as shown in S1000.

[0074] Next, Fig. 11 will be described. S800 to S802, S302 to S304, S306, S307, and S803 to S805 in Fig. 11 are the same as those in Fig. 8. S305 does not exist in Fig. 11 either. In S307, the imaging device 100 delivers an image to the client device 210. The image is delivered (transmitted) to the image processing unit 905 via the network I / F 901 of the client device 210. In Fig. 11, the processing of S1100 is performed between S307 and S803. In S1100, the image processing unit 905 of the client device 210 performs machine recognition based on the delivered (transmitted) image. This machine recognition process is similar to S1000 (FIG. 10), so a description thereof will be omitted. In this way, the machine recognition process (imaging device) in S305 of FIG. 8 can be performed by the client device 210 as shown in S1100.

[0075] In this embodiment, a method of setting AE transitions when machine recognition processing is performed by the client device 210 has been described. In this embodiment as well, by setting AE transition settings with different priorities for areas set as machine recognition areas, it is possible to improve the detection accuracy of machine recognition while outputting an image with high visibility in non-machine recognition areas. Furthermore, by determining which of analog gain and exposure time should be prioritized according to the recognition processing method for machine recognition, a setting suitable for the type of machine recognition can be achieved, thereby improving the detection accuracy.

[0076] The types of machine recognition described above include temporal machine recognition (processing) and spatial machine recognition (processing). In addition, the machine recognition process may be performed by either the imaging device 100 or the client device 210, or by both.

[0077] <Other embodiments> The present invention can also be realized by supplying a program that realizes one or more functions of the above-mentioned embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. Therefore, the program code itself installed in the computer to realize the functional processing of the present invention by the computer also realizes the present invention. In other words, the present invention also includes the computer program itself for realizing the functional processing of the present invention. In that case, as long as it has the function of the program, it may be in the form of object code, a program executed by an interpreter, script data supplied to an OS, etc. In addition, the present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. OS is an abbreviation for Operating System.

[0078] The disclosure of the present embodiment includes the following configurations, systems, methods, and programs. Configuration 1 an imaging means having an exposure area capable of setting exposure conditions for each single pixel or for each plurality of pixels; An acquisition means for acquiring an image recognition area for performing image recognition by image processing; a determination means for determining a transition setting of the exposure condition to be applied to the image recognition area based on the content of the image recognition; an imaging control means for controlling imaging by the imaging means using the transition setting of the exposure condition determined by the determination means; An imaging device comprising: Configuration 2 the exposure conditions include an analog gain and an exposure time; 2. The imaging device according to configuration 1, wherein the determining means determines priorities of the analog gain and the exposure time in the transition setting in accordance with the content of the image recognition. Configuration 3 3. The imaging device according to configuration 2, wherein the transition setting of the exposure condition applied to the image recognition area is different from the transition setting of the exposure condition applied to an area other than the image recognition area. Configuration 4 4. The imaging device according to configuration 3, wherein when the image recognition is performed based on a change in luminance over time, the determining means sets a transition setting that prioritizes the analog gain over the exposure time. Configuration 5 4. The imaging device according to configuration 3, wherein when the image recognition is performed based on a spatial change in luminance, the determining means sets a transition setting that prioritizes the exposure time over the analog gain. Configuration 6 6. The imaging device according to any one of configurations 1 to 5, wherein the determination unit determines a transition setting of the exposure condition depending on a type of the image recognition or a precision required for the image recognition. Configuration 7 7. The imaging device according to any one of configurations 1 to 6, wherein the acquisition means acquires the image recognition area based on area designation information from an outside source. Configuration 8 a designation means for designating the image recognition area based on a result of the image recognition, 7. The imaging device according to any one of configurations 1 to 6, wherein the acquisition means acquires the image recognition area based on a designation by the designation means. Configuration 9 The imaging device according to any one of configurations 1 to 8, further comprising a change unit that changes a setting of the image recognition area based on a detection result by the image recognition. Configuration 10 The imaging device further includes an imaging optical system. The imaging device according to any one of configurations 1 to 9, characterized in that the determination means determines a transition setting of the exposure condition to be applied to the image recognition area by using the luminance of the pixel and optical information of the imaging optical system in addition to the content of the image recognition. Configuration 11 The imaging device according to any one of configurations 1 to 10, wherein the contents of the image recognition include at least one of a type of the image recognition, an object to be image-recognized, and a part of the object. Configuration 12 The imaging device according to any one of configurations 1 to 11, further comprising a processing unit that performs image processing for the image recognition. Configuration 13 The imaging device according to any one of configurations 1 to 11, wherein the imaging device transmits an image captured by the imaging means to an external information processing device, and the information processing device performs the image recognition. System 1 An imaging device according to configuration 13, an information processing device that performs image processing for the image recognition on the image transmitted from the imaging device; A system comprising: Method 1 A method for controlling an image pickup apparatus having an image pickup element having an exposure area in which exposure conditions can be set for each single pixel or for each plurality of pixels, comprising the steps of: An acquisition step of acquiring an image recognition area for performing image recognition by image processing; a determination step of determining a transition setting of the exposure condition to be applied to the image recognition area based on the content of the image recognition; an imaging step of imaging an image by the image sensor using the transition setting of the exposure condition determined by the determination step; A control method comprising the steps of: Program 1 A computer of an image pickup device having an image pickup element having an exposure area in which exposure conditions can be set for each single pixel or for each plurality of pixels, An acquisition step of acquiring an image recognition area for performing image recognition by image processing; a determination step of determining a transition setting of the exposure condition to be applied to the image recognition area based on the content of the image recognition; an imaging step of imaging an image by the image sensor using the transition setting of the exposure condition determined by the determination step; A program for executing. [Explanation of symbols]

[0079] 100 imaging device, 101 imaging optical system, 102 imaging unit, 103 system control unit, 104 machine recognition area acquisition unit, 105 priority determination unit, 106 AE control unit, 123 CPU, 126 imaging control unit, 129 image processing unit, 210 client device

Claims

1. Imaging means having an exposure area where exposure conditions can be set for each single or multiple pixels, Acquisition means for acquiring an image recognition area for performing image recognition by image processing, Determination means for determining a transition setting of the exposure conditions to be applied to the image recognition area based on the content of the image recognition, Imaging control means for performing imaging by the imaging means using the transition setting of the exposure conditions determined by the determination means, and having The exposure conditions include gain and exposure time, The determination means determines the priorities of the gain and the exposure time in the transition setting according to the content of the image recognition, and an imaging device characterized by this.

2. The transition setting of the exposure conditions to be applied to the image recognition area is different from the transition setting of the exposure conditions to be applied to areas other than the image recognition area, and the imaging device according to Claim 1 is characterized by this.

3. When performing the image recognition based on a temporal luminance change, the determination means makes a transition setting that prioritizes the gain over the exposure time, and the imaging device according to Claim 2 is characterized by this.

4. When performing the image recognition based on a spatial luminance change, the determination means makes a transition setting that prioritizes the exposure time over the gain, and the imaging device according to Claim 2 is characterized by this.

5. The determination means determines the transition setting of the exposure conditions according to the type of the image recognition or the accuracy required for the image recognition, and the imaging device according to any one of Claims 1 to 4 is characterized by this.

6. The acquisition means acquires the image recognition area based on area designation information from the outside, and the imaging device according to Claim 5 is characterized by this.

7. Further comprising designation means for designating the image recognition area based on the result of the image recognition, The acquisition means acquires the image recognition area based on the designation by the designation means, and the imaging device according to Claim 5 is characterized by this.

8. Further comprising change means for changing the setting of the image recognition area based on the detection result by the image recognition, and the imaging device according to any one of Claims 1 to 4 is characterized by this.

9. The imaging device further comprises an imaging optical system, The determination means determines the transition setting of the exposure conditions to be applied to the image recognition region by using, in addition to the content of the image recognition, the luminance of the pixels and the optical information of the imaging optical system. The imaging device according to any one of claims 1 to 4.

10. The content of the image recognition includes at least any one of the type of the image recognition, the object to be image-recognized, and the part of the object. The imaging device according to any one of claims 1 to 4.

11. The imaging device according to any one of claims 1 to 4 further includes processing means for performing image processing for the image recognition.

12. The imaging device according to any one of claims 1 to 4 is characterized in that an image captured by the imaging means is transmitted to an external information processing device, and the information processing device performs the image recognition.

13. The imaging device according to claim 12, An information processing device that performs image processing for the image recognition on the image transmitted from the imaging device, A system characterized by comprising.

14. A control method for an imaging device including an imaging element having an exposure region in which exposure conditions can be set for each single or a plurality of pixels, comprising: An acquisition step of acquiring an image recognition region for performing image recognition by image processing; A determination step of determining a transition setting of the exposure conditions to be applied to the image recognition region based on the content of the image recognition; An imaging step of performing imaging by the imaging element using the transition setting of the exposure conditions determined in the determination step, and having The exposure conditions include a gain and an exposure time, In the determination step, the priorities of the gain and the exposure time in the transition setting are determined according to the content of the image recognition. A control method characterized by this.

15. A program for causing a computer to execute the control method according to claim 14.