Information processing device, and information processing method
The information processing device enhances recognition accuracy by switching between AI models and using multiple images with varying conditions to improve the reliability of recognition results.
Patent Information
- Application Number
- JP2024011794
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2025-08-12
AI Technical Summary
Existing imaging devices struggle with low recognition accuracy due to multiple elements in the image and low resolution, as they perform only a single recognition process on captured images.
Implement an information processing device with a model switching unit that switches between a first AI model for detecting the area of a recognition target and a second AI model for recognizing the target in the detected area, using multiple captured images under different conditions to determine the final recognition result.
Improves recognition accuracy by leveraging multiple AI models and imaging conditions to enhance the reliability of the recognition outcome.
Smart Images

Figure 2025117107000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device and an information processing method. [Background technology]
[0002] In recent years, a stacked imaging device has become known in which a substrate having an imaging section in which a plurality of pixels are arranged two-dimensionally is bonded to a substrate having a signal processing section that performs signal processing on an image captured by the imaging section.
[0003] On the other hand, as use cases of imaging devices, not only are images captured by the imaging device used, but also information obtained from the captured images (so-called metadata) is increasingly used.
[0004] For example, Patent Document 1 listed below discloses an imaging device in which a first substrate having an imaging unit that captures an image and a second substrate having a recognition processing unit that performs recognition processing on the captured image output by the imaging unit are joined together. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent No. 6633216 Summary of the Invention [Problem to be solved by the invention]
[0006] However, in the imaging device disclosed in Patent Document 1, only a single recognition process is performed on the captured image in the recognition processing unit, so depending on the type of metadata, it may be difficult to obtain metadata with high recognition accuracy. For example, if the image contains many elements other than the recognition target, it may be difficult to recognize the recognition target itself due to the influence of the elements other than the recognition target. Furthermore, if the resolution of the image of the recognition target is low, the recognition accuracy of the recognition target is likely to decrease.
[0007] In light of the above circumstances, it is desirable to obtain recognition results with higher recognition accuracy by performing multiple AI processes on captured images. [Means for solving the problem]
[0008] According to the present disclosure, there is provided an information processing device comprising: a model switching unit that instructs switching of a model of AI processing performed on a captured image from a first AI model that detects an area of a recognition target included in the captured image to a second AI model that recognizes the recognition target in the area detected from the captured image; and a result determination unit that determines a recognition result to be output based on multiple recognition results of the recognition target recognized by the second AI model from multiple captured images captured under different imaging conditions.
[0009] In addition, according to the present disclosure, there is provided an information processing method by a computer, including obtaining a result of detecting an area of a recognition target included in a captured image using a first AI model, instructing to switch the model of AI processing performed on the captured image from the first AI model to a second AI model that recognizes the recognition target in the area detected from the captured image, obtaining multiple recognition results of the recognition target recognized by the second AI model from multiple captured images captured under different imaging conditions, and determining a recognition result to be output based on the multiple recognition results. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is an explanatory diagram illustrating a digital camera in which the present technology is implemented. [Figure 2] 1 is a block diagram showing an example of the configuration of a digital camera to which the present technology is applied. [Figure 3] 1 is a block diagram illustrating an example configuration of an imaging device according to an embodiment of the present disclosure. [Figure 4] FIG. 1 is a perspective view showing an example of the external configuration of an imaging device. [Figure 5] FIG. 1 is a block diagram illustrating an example of the configuration of an information processing device. [Figure 6] FIG. 10 is a flowchart showing the flow of processing in the imaging device and the information processing device. [Figure 7] FIG. 10 is an explanatory diagram showing an example of a method for determining a recognition result to be output from a plurality of recognition results. [Figure 8] FIG. 10 is an explanatory diagram showing another example of a method for determining a recognition result to be output from a plurality of recognition results. [Figure 9] FIG. 10 is a sequence diagram showing data transfer between an imaging block, a DSP, and an information processing device. [Figure 10] FIG. 10 is a flowchart showing a processing flow according to a modified example of the digital camera. [Figure 11] FIG. 10 is a flowchart showing the flow of processing in an application example using a digital camera. DETAILED DESCRIPTION OF THE INVENTION
[0011] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0012] The explanation will be given in the following order. 1. Overview 2. Configuration 2.1.Digital Camera Configuration 2.2.Configuration of the imaging device 2.3.Configuration of information processing device 3. Processing flow 4. Variations 5. Application Examples 6. Supplementary Notes
[0013] <1. Overview> First, an overview of the technology according to the present disclosure will be described with reference to Fig. 1. Fig. 1 is an explanatory diagram illustrating a digital camera 10 in which the present technology is implemented.
[0014] For example, as shown in FIG. 1, the present technology is implemented in a digital camera 10 that monitors a parking lot S where a vehicle C is parked.
[0015] The digital camera 10 captures an image of a vehicle C parked in a parking lot S, and recognizes the contents of a number plate N affixed to the vehicle C from the captured image of the vehicle C using an AI (Artificial Intelligence) model. By recognizing the registration number of the vehicle C written on the number plate N, the digital camera 10 can identify the vehicle C parked in the parking lot S.
[0016] Specifically, digital camera 10 first inputs a captured image of vehicle C into a first AI model, thereby detecting the area of license plate N affixed to vehicle C. Next, digital camera 10 inputs the area of license plate N cut out from the captured image into a second AI model, thereby being able to recognize the registration number of vehicle C written on license plate N. This allows digital camera 10 to compare the registration number of vehicle C recognized from license plate N in the captured image with the registration number of vehicle C contracted by the contractor of parking lot S, thereby making it possible to, for example, quickly detect suspicious vehicles that have entered parking lot S.
[0017] A digital camera 10 incorporating this technology can detect the area of a number sign N in a captured image using a first AI model, and then recognize the content of the number sign N in the detected area using a second AI model. This allows the digital camera 10 to perform recognition processing using an image centered on the number sign N, which is the recognition target, thereby further improving the accuracy of the recognition result of the number sign N.
[0018] Furthermore, in the digital camera 10 in which the present technology is implemented, in recognizing the number sign N using the second AI model, the number sign N is recognized in a plurality of captured images taken under different imaging conditions, and a final recognition result is determined from the respective recognition results. This allows the digital camera 10 to prevent a decrease in the recognition accuracy of the recognition target due to the imaging conditions of the captured images.
[0019] The present technology is implemented, for example, as an imaging device included in a digital camera 10, or as an information processing device that controls an imaging device. That is, the present technology may be implemented as an information processing device having only a control function for controlling an imaging device, or may be implemented as an imaging device having an imaging function, a recognition function, and a control function. Below, the configuration will be described in more detail by exemplifying a case in which the present technology is implemented as an information processing device having only a control function for controlling an imaging device.
[0020] <2. Configuration> (2.1. Digital Camera Configuration) Next, the configuration of a digital camera 10 to which the present technology is applied will be described with reference to Fig. 2. Fig. 2 is a block diagram showing an example configuration of a digital camera 10 to which the present technology is applied.
[0021] 2, the digital camera 10 includes an optical system 1, an imaging device 2, a memory 3, a signal processing unit 4, an output unit 5, a control unit 6, and an information processing device 7. The digital camera 10 is an electronic device capable of capturing both still images and moving images.
[0022] The optical system 1 includes a zoom lens, a focus lens, a diaphragm, and the like, and guides incident light from the outside to the imaging device 2.
[0023] The imaging device 2 generates image data corresponding to the incident light by photoelectrically converting the incident light guided by the optical system 1. The imaging device 2 may be, for example, a CMOS image sensor. The imaging device 2 may also detect a recognition target from the image data using a first AI model and recognize the detected recognition target using a second AI model.
[0024] The memory 3 temporarily stores the image data or the recognition result output by the imaging device 2.
[0025] The signal processing unit 4 performs signal processing such as noise removal and white balance adjustment as necessary on the image data stored in the memory 3. The image data that has been signal processed by the signal processing unit 4 is output to the output unit 5.
[0026] The output unit 5 outputs to the outside the image data that has been signal processed by the signal processing unit 4 or the recognition results stored in the memory 3. The output unit 5 may be, for example, a display device that displays an image corresponding to the image data as a through image, a writing device (drive) that stores the image data or the recognition results in a storage medium, or an output interface that transmits the image data or the recognition results to an external device.
[0027] The control unit 6 is a control device that controls each part of the digital camera 10 in accordance with user operations and the like.
[0028] The information processing device 7 instructs the imaging device 2 to switch between a first AI model and a second AI model used in the imaging device 2. Specifically, the information processing device 7 may instruct the imaging device 2 to switch between a first AI model that detects a recognition target from image data and a second AI model that recognizes the recognition target.
[0029] Furthermore, the information processing device 7 determines the recognition result to be output to the memory 3 based on multiple recognition results of the recognition target recognized by the second AI model. Specifically, the information processing device 7 may determine the recognition result to be output to the memory 3 by comparing and judging multiple recognition results of the recognition target recognized from multiple image data captured under different imaging conditions.
[0030] (2.2. Configuration of imaging device) Next, the configuration of the imaging device 2 according to this embodiment will be described with reference to FIGS.
[0031] Fig. 3 is a block diagram showing an example of the configuration of the imaging device 2 according to this embodiment. As shown in Fig. 3, the imaging device 2 includes an imaging block 20 and a signal processing block 30. The imaging block 20 and the signal processing block 30 are connected by connection lines CL1, CL2, and CL3, which are internal buses, so that data can be exchanged between them.
[0032] The imaging block 20 is a functional block that executes an imaging function, and includes an imaging unit 21 , an imaging processing unit 22 , an output control unit 23 , an output I / F 24 , and an imaging control unit 25 .
[0033] The imaging unit 21 is configured by a two-dimensional array of multiple pixels. The imaging unit 21 is driven by the imaging processing unit 22 to capture an image. Specifically, the imaging unit 21 receives light incident via the optical system 1 at each pixel and performs photoelectric conversion on the received light to generate an analog image signal corresponding to the light incident on the imaging device 2. The type of image generated by the imaging unit 21 may be, for example, an RGB (Red, Green, Blue) color image or a monochrome image containing only brightness information.
[0034] The imaging processing unit 22 performs processing related to image capture by the imaging unit 21 under the control of the imaging control unit 25. For example, the imaging processing unit 22 may drive the imaging unit 21, perform AD (Analog to Digital) conversion of analog image signals generated by the imaging unit 21, and perform imaging signal processing. Note that imaging signal processing includes, for example, processing to calculate the brightness of each predetermined small region by calculating the average value of pixel values for each small region, processing to convert the image generated by the imaging unit 21 into an HDR (High Dynamic Range) image, defect correction processing, development processing, etc.
[0035] The imaging processing unit 22 generates a digital image signal by performing AD conversion or the like on the analog image signal generated by the imaging unit 21. The generated digital image signal is output to the output control unit 23 and also to the image compression unit 35 of the signal processing block 30 via the connection line CL2.
[0036] The output control unit 23 controls selective output to the outside (for example, the memory 3 in FIG. 2) from the output I / F 24. Specifically, the output control unit 23 may control the output to the outside of the image signal output from the imaging processing unit 22, or may control the output to the outside of the recognition result of the captured image executed by the signal processing block 30 described later.
[0037] The output I / F 24 is an interface that outputs various types of information to the outside in accordance with the control of the output control unit 23. For example, the output I / F 24 may output to the outside the image signal supplied from the imaging processing unit 22 and the recognition result of the captured image executed by the signal processing block 30, which will be described later.
[0038] For example, when only output of the recognition result of the captured image (so-called metadata) is required from the imaging device 2, the output control unit 23 may control the output I / F 24 to output only the recognition result of the captured image to the outside. In such a case, the output control unit 23 can reduce the amount of data output from the output I / F 24 to the outside, thereby enabling faster data output to the outside.
[0039] The imaging control unit 25 controls the imaging processing unit 22 in accordance with the imaging information stored in the register group 27, thereby controlling the imaging of the image by the imaging unit .
[0040] The register group 27 includes a plurality of registers that store setting information related to imaging by the imaging unit 21. For example, the register group 27 may store setting information related to imaging received from the outside via the communication I / F 26, or the results of imaging signal processing by the imaging processing unit 22. Examples of setting information related to imaging include analog gain, exposure, shutter speed, frame rate, focus, white balance, and crop range.
[0041] The register group 27 may further store setting information related to output control in the output control unit 23. The output control unit 23 can control selective output to the outside in accordance with the setting information related to output control.
[0042] The communication I / F 26 is an interface for transmitting and receiving data to and from the outside (for example, the control unit 6 in FIG. 2, etc.). The communication I / F 26 may transmit and receive various setting information stored in the register group 27 to and from the outside, for example.
[0043] The signal processing block 30 is a functional block that executes a recognition function for recognizing a captured image, and includes a CPU (Central Processing Unit) 31, a DSP (Digital Signal Processor) 32, a memory 33, a communication I / F 34, an image compression unit 35, and an input I / F 36. The CPU 31, the DSP 32, the memory 33, the communication I / F 34, the image compression unit 35, and the input I / F 36 are connected via a bus, so that they can input and output information to and from each other.
[0044] The CPU 31 executes a program stored in the memory 33 to control the signal processing block 30 and input and output information to and from the imaging control unit 25 via the connection line CL1.
[0045] For example, the CPU 31 may instruct the imaging control unit 25 to capture an image via the connection line CL1. Alternatively, the CPU 31 may write setting information related to imaging to the register group 27 of the imaging control unit 25, thereby instructing the imaging block 20 to capture an image based on the written setting information. This allows the CPU 31 to instruct the imaging block 20 to capture an image while changing the imaging conditions (for example, bracket imaging with changed exposure).
[0046] The DSP 32 executes a program stored in the memory 33 to perform AI processing on the captured image supplied from the image capture processing unit 22 via the connection line CL2.
[0047] The content of the AI processing executed by the DSP 32 is determined by the content of the AI model held in the DSP 32. For example, the DSP 32 may perform AI processing to detect the position or area of a recognition target from a captured image using a first AI model. The DSP 32 may also perform AI processing to recognize the content of a recognition target detected from a captured image using a second AI model. The DSP 32 can switch the content of the AI processing to be executed by switching the content of the AI model held in accordance with an instruction from the information processing device 7.
[0048] The first AI model and the second AI model may be generated by a machine learning algorithm such as a deep neural network (DNN) or a transformer, etc. Alternatively, the first AI model and the second AI model may be generated by a rule-based algorithm.
[0049] The memory 33 stores data and the like used in the processing of the signal processing block 30. The memory 33 may be configured with a static random access memory (SRAM), a dynamic random access memory (DRAM), or the like. For example, the memory 33 may store a program received from the outside via the communication I / F 34, a captured image output from the imaging block 20, a result of AI processing executed by the DSP 32, setting information received from the input I / F 36, and the like.
[0050] The communication I / F 34 is an interface for transmitting and receiving data to and from the outside (for example, the memory 3 or the control unit 6 in FIG. 2 ). The communication I / F 34 may transmit and receive information such as a program to be executed by the CPU 31 or the DSP 32 to and from the outside. For example, the communication I / F 34 may store a program downloaded from the outside in the memory 33 to be executed by the CPU 31 or the DSP 32.
[0051] The image compression unit 35 performs compression processing on the captured image passed via the connection line CL2 from the image capture processing unit 22. This allows the image compression unit 35 to generate a compressed image with a smaller amount of data than the captured image before compression.
[0052] The AI processing in the DSP 32 may be performed on either a captured image captured by the imaging block 20 or a compressed image generated by the image compression unit 35. Because a compressed image has a smaller amount of data than a captured image, it is possible to reduce the load of the AI processing in the DSP 32 and save the storage capacity of the memory 33 that stores the compressed image. For example, a compressed image may be input to a first AI model that detects a recognition target from a captured image to reduce the load. On the other hand, a captured image may be input to a second AI model that recognizes the detected recognition target to improve recognition accuracy.
[0053] The compression process in the image compression unit 35 may be, for example, a compression process of scaling down a captured image of 12M (3968×2976) pixels to an image of VGA (640×480) size.
[0054] The input I / F 36 is an interface that receives various types of information from the outside (for example, an external sensor). For example, the input I / F 36 may receive sensing results acquired by an external sensor. The received sensing results may be subjected to AI processing by the DSP 32, for example. Examples of external sensors include a distance measurement sensor that acquires depth information, an infrared sensor that acquires infrared images, or an image sensor different from the imaging device 2.
[0055] The imaging device 2 having the above configuration can perform AI processing using a first AI model that detects a recognition target from the captured image, and AI processing using a second AI model that recognizes the content of the detected recognition target, on the captured image captured by the imaging unit 21. This allows the imaging device 2 to output the recognition result of the recognition target from the output I / F 24.
[0056] Fig. 4 is a perspective view showing an example of the external configuration of the imaging device 2. As shown in Fig. 4, the imaging device 2 is configured as a semiconductor device having a stacked structure in which a first substrate 51 and a second substrate 52 are stacked.
[0057] The first board 51 is provided with, for example, an imaging unit 21. On the other hand, the second board 52 is provided with, for example, an imaging processing unit 22, an output control unit 23, an output I / F 24, an imaging control unit 25, a CPU 31, a DSP 32, a memory 33, a communication I / F 34, an image compression unit 35, and an input I / F 36.
[0058] As one example, the first substrate 51 and the second substrate 52 may be electrically connected by a through via that penetrates the first substrate 51 and reaches the second substrate 52. Alternatively, as another example, the first substrate 51 and the second substrate 52 may be electrically connected by Cu-Cu bonding that directly bonds a Cu wiring exposed on the lower surface side of the first substrate 51 to a Cu wiring exposed on the upper surface side of the second substrate.
[0059] If it is acceptable for the area of the imaging device 2 to be large, the imaging device 2 may be configured on a single substrate.
[0060] The imaging device 2 may also be configured by stacking three or more boards. In such a case, the imaging device 2 may be configured by stacking a first board on which the imaging unit 21 is provided, a second board on which the imaging processing unit 22, the output control unit 23, the output I / F 24, the imaging control unit 25, the CPU 31, the DSP 32, the communication I / F 34, the image compression unit 35, and the input I / F 36 are provided, and a third board on which the memory 33 is provided.
[0061] (2.3. Configuration of Information Processing Device) The configuration of the information processing device 7 according to this embodiment will be further described with reference to Fig. 5. Fig. 5 is a block diagram showing an example of the configuration of the information processing device 7 according to this embodiment.
[0062] As shown in FIG. 5, the information processing device 7 includes an acquisition instruction unit 110, a stillness determination unit 120, a model switching unit 130, a crop instruction unit 140, and a result determination unit 150.
[0063] The acquisition instruction unit 110 instructs the imaging block 20 to acquire a captured image for AI processing by the DSP 32. As a result, the imaging control unit 25, imaging processing unit 22, and imaging unit 21 of the imaging block 20 operate, thereby acquiring the captured image.
[0064] Furthermore, the captured image acquired based on an instruction from the acquisition instruction unit 110 is compressed or a partial region is cut out to fit the input of the AI model. For example, in an AI process for detecting the position or region of a recognition target from a captured image, a compressed image obtained by compressing the captured image is input to a first AI model to reduce the processing load. Furthermore, in an AI process for recognizing a recognition target detected in a captured image, an image obtained by cutting out the region of the recognition target from the captured image is input to a second AI model to improve recognition accuracy and reduce the processing load.
[0065] The model switching unit 130 switches the AI model used for AI processing by the DSP32. By switching the AI model by the model switching unit 130, the DSP32 can perform various AI processing. As an example, the model switching unit 130 may switch the AI model used by the DSP32 by specifying an AI model from among multiple AI models stored in the memory 33 or the like. As another example, the model switching unit 130 may switch the AI model used by the DSP32 by uploading the AI model to the DSP32.
[0066] For example, the model switching unit 130 may switch between a first AI model that detects the area of the recognition target from the captured image and a second AI model that recognizes the recognition target detected in the captured image as the AI model used for the AI processing of the DSP32.
[0067] The stillness determination unit 120 determines whether the recognition target detected in the captured image is still. Specifically, the stillness determination unit 120 may determine that the recognition target is still when the amount of change in the area or position of the recognition target detected in the captured image over time satisfies a predetermined condition. The predetermined condition is, for example, when the amount of change in the area or position of the recognition target is equal to or less than a threshold value over a predetermined number of frames.
[0068] If the stillness determination unit 120 determines that the recognition target is still, the model switching unit 130 switches the AI model of the DSP 32 from the first AI model to the second AI model. As a result, the recognition target in the area detected by the first AI model is recognized by the second AI model.
[0069] The crop instruction unit 140 instructs an area to be cut out (cropped) from the captured image to input to the second AI model of the DSP 32. Specifically, the crop instruction unit 140 writes an instruction (crop instruction) to cut out the area of the recognition target in the captured image detected in the first processing by the first AI model to the register group 27 of the imaging block 20. This results in an image with a smaller amount of data obtained by cutting out the area of the recognition target in the captured image to be input to the second AI model of the DSP 32. Therefore, the crop instruction unit 140 can improve the recognition accuracy by the second AI model of the DSP 32 and reduce the processing load.
[0070] The crop instruction unit 140 may instruct the image capture unit 21 to cut out the area to be recognized from the captured image output by the image capture unit 21, or may instruct the image capture unit 21 to cut out and read out only the image signal of the area to be recognized.
[0071] When the model switching unit 130 switches the AI model of the DSP 32 and the crop instruction unit 140 instructs the imaging block 20 to capture images under a plurality of imaging conditions, the acquisition instruction unit 110 instructs the imaging block 20 to capture images under a plurality of imaging conditions. The imaging device 2 according to this embodiment recognizes the recognition target for each captured image captured under a plurality of imaging conditions, and the result determination unit 150 compares and determines the recognition result to be output from the output I / F 24.
[0072] The plurality of imaging conditions are, for example, conditions obtained by varying one type of imaging parameter, such as exposure, gain, contrast, color temperature, or white balance.
[0073] The result determination unit 150 determines the recognition result to be output to the memory 3 based on the recognition results of the recognition target recognized from each of a plurality of captured images captured under different imaging conditions. This allows the result determination unit 150 to determine the recognition result to be output to the memory 3 by referring to the recognition result using the captured image captured under the imaging condition that provides the highest recognition accuracy among the plurality of imaging conditions. Therefore, the information processing device 7 can improve the recognition accuracy of the recognition result to be output to the memory 3.
[0074] For example, the result determination unit 150 may determine the most reliable recognition result of the recognition target recognized from a plurality of captured images captured under different imaging conditions as the recognition result to be output to the memory 3. The reliability is an index indicating the degree of certainty of the recognition result, and is output together with the recognition result from the second AI model that recognizes the recognition target. The larger the reliability value, the higher the degree of certainty, and is expressed in a range of, for example, 0 to 1 or 0 to 100.
[0075] Alternatively, the result determination unit 150 may determine the recognition result to be output to the memory 3 by breaking down the recognition results of the recognition target recognized from each of a plurality of captured images captured under different imaging conditions into elements and taking a majority vote for each element.
[0076] <3. Processing flow> Next, the flow of processing by the imaging device 2 and information processing device 7 according to this embodiment will be described with reference to Fig. 6 to Fig. 8. Fig. 6 is a flowchart showing the flow of processing in the imaging device 2 and information processing device 7. Fig. 7 is an explanatory diagram showing an example of a method for determining a recognition result to be output from a plurality of recognition results. Fig. 8 is an explanatory diagram showing another example of a method for determining a recognition result to be output from a plurality of recognition results.
[0077] 6, first, the information processing device 7 instructs the imaging block 20 to acquire a frame of a captured image via the acquisition instructing unit 110 (S10). As a result, the imaging block 20 acquires a captured image P1 including a license plate N (recognition target) attached to the vehicle C.
[0078] The captured image P1 is compressed by a compression process that reduces the resolution, and then input to the object detection AI model M1 (first AI model) of the DSP 32. As a result, the DSP 32 detects the number sign N, which is the recognition target, using the object detection AI model M1, and outputs the coordinates of the area (bounding box) of the detected number sign N (S11).
[0079] Next, the information processing device 7 determines whether the number marker N to be recognized is stationary or not using the stillness determination unit 120 (S12). Specifically, the stillness determination unit 120 may determine that the number marker N is stationary when the amount of change in the center coordinates of the area of the number marker N to be recognized is equal to or less than a threshold value over a predetermined number of frames.
[0080] If the number plate N is not stationary (S12 / NO), the information processing device 7 returns to the operation of step S10 and instructs the imaging block 20 to acquire a frame of a captured image again.
[0081] On the other hand, if the number plate N is stationary (S12 / YES), the information processing device 7 transmits the coordinates of the area of the number plate N that is the recognition target to the outside (S13), and then instructs the model switching unit 130 to switch the AI model to the DSP 32 (S14). Specifically, the model switching unit 130 instructs the DSP 32 to switch from a first AI model that detects the area of the recognition target from the captured image to a second AI model that recognizes the content of the recognition target detected in the captured image. As a result, the AI model of the DSP 32 is switched from the first AI model to the second AI model (S15).
[0082] Next, the information processing device 7 instructs the imaging block 20 via the crop instruction unit 140 as to an area to be cut out (cropped) from the captured image (S16). Specifically, the crop instruction unit 140 instructs the imaging block 20 to cut out an area (bounding box) of the number marker N that is the recognition target. The instruction from the crop instruction unit 140 is stored in the register group 27 of the imaging block 20, whereby crop setting is performed (S17).
[0083] Next, the information processing device 7 instructs the imaging block 20, via the acquisition instructing unit 110, to perform bracket imaging, which acquires frames of captured images under different imaging conditions (S18). For example, the acquisition instructing unit 110 may instruct the imaging block 20 to perform bracket imaging, which acquires frames of captured images of the vehicle C under different exposure conditions. In other words, the acquisition instructing unit 110 may instruct the imaging block 20 to perform bracket imaging with changed exposure.
[0084] In the imaging block 20, a crop setting is performed by the process of step S17 to cut out the area of the number plate N to be recognized. Therefore, the imaging block 20 acquires a captured image P2 in which only the image signal of the area of the number plate N to be recognized is read out.
[0085] Thereafter, each of the acquired captured images P2 is input to the object recognition AI model M2 (second AI model) of the DSP 32. This enables the DSP 32 to recognize the registration number of the vehicle C marked on the license plate N, which is the recognition target, using the object recognition AI model M2. The registration number of the vehicle C recognized in each of the captured images P2 (recognition result) is output to the information processing device 7.
[0086] Next, the information processing device 7 determines, in the result determination unit 150, the registration number (recognition result) of vehicle C to be finally output to the outside based on the registration number (recognition result) of vehicle C recognized from each captured image P2 captured under different imaging conditions (S19).
[0087] 7, the object recognition AI model M2 of the DSP 32 outputs the confidence of the recognition result along with the recognition result. The confidence is an index indicating that the higher the value, the higher the certainty of the recognition result, and the range is, for example, 0 to 100. For example, assume that the confidence of the first recognition result R1 is output as "70," the confidence of the second recognition result R2 is output as "60," and the confidence of the third recognition result R3 is output as "90." In such a case, the result determination unit 150 can determine the third recognition result R3, which has the highest confidence, as the recognition result RF to be finally output from the image capture device 2.
[0088] 8, the result determination unit 150 can determine the recognition result RF that is finally output from the imaging device 2 by majority voting of the first recognition result R1, the second recognition result R2, and the third recognition result R3 captured under different imaging conditions. Specifically, the result determination unit 150 first decomposes the first recognition result R1 into the elements "TKA," "500," "NA," and "1234." Similarly, the result determination unit 150 decomposes the second recognition result R2 into the elements "TKA," "500," "HA," and "1234," and decomposes the third recognition result R3 into the elements "TKA," "503," "NA," and "1234." The result determination unit 150 can determine the recognition result RF as the recognition result including the majority elements "TKA", "500", "NA", and "1234" by taking a majority vote between the corresponding elements of the first recognition result R1, the second recognition result R2, and the third recognition result R3.
[0089] Subsequently, the information processing device 7 transmits the recognition result determined by the process of step S19 to the outside (S20).
[0090] After transmitting the recognition result to the outside, the information processing device 7 instructs the DSP 32 to switch the AI model via the model switching unit 130 (S21). Specifically, the model switching unit 130 instructs the DSP 32 to switch from the second AI model that recognizes the content of the recognition target detected in the captured image to the first AI model that detects the area of the recognition target from the captured image. As a result, the AI model of the DSP 32 is switched from the second AI model to the first AI model (S22).
[0091] By repeating the above process as one cycle, the imaging device 2 and the information processing device 7 can recognize, for example, the registration number of a vehicle C parked in a parking lot S. Therefore, the digital camera 10 can monitor the situation of the parking lot S where the vehicle C or the like is parked.
[0092] Next, the exchange of data processed by the imaging device 2 will be described with reference to Fig. 9. Fig. 9 is a sequence diagram showing the exchange of data between the imaging block 20, the DSP 32, and the information processing device 7.
[0093] 9, for example, a start instruction is input to the information processing device 7 from the console 70 operated by a user who manages the digital camera 10 (S101). As a result, an instruction to acquire a frame of a captured image is output from the information processing device 7 to the imaging block 20 (S103).
[0094] In the imaging block 20, a frame of a captured image is acquired (S105), and then image processing such as compression processing is performed on the acquired frame of the captured image (S107). The captured image that has undergone image processing is input to an object detection AI model (first AI model) of the DSP 32 (S109). As a result, the object detection AI model detects a recognition object from the captured image (S111), and coordinate data of the area (bounding box) of the recognition object is output.
[0095] The coordinate data of the detected area of the recognition target is output to the information processing device 7, where post-processing is performed (S113). The coordinate data of the area of the recognition target that has undergone post-processing is output to the console 70 (S115).
[0096] Thereafter, the information processing device 7 outputs an instruction to switch the AI model to the DSP 32 (S117). As a result, the DSP 32 switches the AI model from an object detection AI model (first AI model) that detects the area of the recognition target from the captured image to an object recognition AI model (second AI model) that recognizes the content of the recognition target detected in the captured image (S119). Furthermore, the information processing device 7 outputs a crop instruction to the imaging block 20 to crop the area of the recognition target from the captured image (S121).
[0097] Next, an instruction to acquire frames of captured images is output from the information processing device 7 to the imaging block 20 (S123). Specifically, the information processing device 7 instructs the imaging block 20 to perform bracket imaging to acquire frames of captured images under different imaging conditions (for example, exposure conditions, etc.).
[0098] In the imaging block 20, imaging conditions are changed and multiple frames of captured images are acquired (S125). The range of the frames of captured images acquired at this time may be, for example, only the area of the detected recognition target. After image processing is performed on the multiple frames of captured images acquired (S127), the captured images are input to an object recognition AI model (second AI model) of the DSP 32 (S129). As a result, the object recognition AI model recognizes the recognition target (S131), and recognition results indicating the content of the recognition target are output.
[0099] The output recognition result is output to the information processing device 7, where post-processing is performed (S133). For example, the recognition result to be finally output to the console 70 is determined based on the recognition results recognized from each of the frames of a plurality of captured images. Thereafter, the recognition result determined based on the plurality of recognition results is output to the console 70 (S135).
[0100] By transferring the data as described above, the digital camera 10 can, for example, acquire a captured image based on an instruction from the console 70 and output to the console 70 the recognition result of the recognition target included in the captured image.
[0101] <4. Modifications> Next, a modified example of the digital camera 10 according to the present embodiment will be described with reference to Fig. 10. Fig. 10 is a flowchart showing the flow of processing according to the modified example of the digital camera 10. In the modified example of the digital camera 10, between the crop instruction (S16) and frame acquisition (S18) shown in Fig. 6, AI processing for recognizing a recognition target from a captured image and outputting the reliability of the recognition result, and processing for determining whether or not to perform bracket imaging are performed are carried out.
[0102] 10, after a crop instruction instructing an area to be cut out (cropped) from the captured image is output to the imaging block 20, the information processing device 7 instructs the imaging block 20 via the acquisition instruction unit 110 to acquire a frame of the captured image. As a result, the imaging block 20 acquires a captured image in which only the image signal of the area of the number marker N to be recognized is read out (S25). The imaging conditions for the captured image acquired at this time may be the same as the imaging conditions for acquiring the frame in step S10, for example.
[0103] Next, the acquired captured image is input to the object recognition AI model M2 (second AI model) of the DSP 32. As a result, the DSP 32 recognizes the registration number of the vehicle C written on the license plate N, which is the recognition target, using the object recognition AI model M2 (S26). As a result, the DSP 32 outputs the recognized registration number of the vehicle C (recognition result) and the confidence of the recognition result to the information processing device 7. As described above, the confidence is an index indicating that the higher the value, the higher the certainty of the recognition result.
[0104] Next, the information processing device 7 determines whether the reliability of the output recognition result is equal to or greater than a threshold value (S27).
[0105] If the confidence is equal to or greater than the threshold (S27 / YES), the information processing device 7 determines that a recognition result with sufficient confidence has been obtained, and determines the recognition result obtained in step S26 as the recognition result to be finally output to the outside. Thereafter, the information processing device 7 transmits the determined recognition result to the outside of the digital camera 10 (S20).
[0106] On the other hand, if the confidence is less than the threshold (S27 / NO), the information processing device 7 determines that a recognition result with sufficient confidence has not been obtained, and instructs the imaging block 20 to perform bracket imaging to acquire frames of captured images under different imaging conditions (S18). As a result, the imaging block 20, the DSP 32, and the information processing device 7 perform the processes of steps S18 to S19 to determine the recognition result to be finally output to the outside.
[0107] In the modification of the digital camera 10 according to the present embodiment described above, the information processing device 7 causes the imaging device 2 to perform bracket imaging that varies the imaging conditions only when the reliability of the recognition result of the object recognition AI model M2 falls below a threshold. In this way, the information processing device 7 can reduce the number of times bracket imaging is performed when recognizing an object, thereby reducing the power consumption of the digital camera 10.
[0108] <5. Application Examples> In the above, the digital camera 10 detects the license plate N attached to the vehicle C as the recognition target and recognizes the registration number of the vehicle C written on the license plate N, but the present technology is not limited to such an example. The digital camera 10 according to this embodiment is capable of detecting various objects in a captured image as recognition targets and recognizing the contents of the detected objects.
[0109] As one example, digital camera 10 can detect a person's face from a captured image and perform facial expression recognition, personal recognition, gaze estimation, etc. from the detected face. As another example, digital camera 10 can detect a pattern code such as a QR code (registered trademark) or barcode attached to an item from a captured image and recognize the content of the detected pattern code.
[0110] An application example in which the digital camera 10 detects a person's face as a recognition target from a captured image and recognizes the expression or individual of the detected face will be described below with reference to Fig. 11. Fig. 11 is a flowchart showing the flow of processing in this application example by the digital camera 10.
[0111] 11, first, the information processing device 7 instructs the imaging block 20 to acquire a frame of a captured image via the acquisition instructing unit 110 (S30). As a result, the imaging block 20 acquires a captured image P3 including person H.
[0112] The captured image P3 is compressed by a compression process that reduces the resolution, and then input to the face detection AI model M3 (first AI model) of the DSP 32. As a result, the DSP 32 detects the face F of the person H using the face detection AI model M3, and outputs the coordinates of the area (bounding box) of the detected face F to the information processing device 7 (S31).
[0113] Next, the information processing device 7 determines whether the face F of the person H is still or not using the stillness determination unit 120 (S32). If the face F of the person H is not still (S32 / NO), the information processing device 7 returns to the operation of step S30 and again instructs the imaging block 20 to acquire a frame of the captured image.
[0114] On the other hand, if the face F of person H is stationary (S32 / YES), the information processing device 7 transmits the coordinates of the area of person H's face F to the outside (S33), and then instructs the model switching unit 130 to switch the AI model to the DSP32 (S34). Specifically, the model switching unit 130 instructs the DSP32 to switch from a face detection AI model M3 (first AI model) that detects person H's face F from the captured image to a face recognition AI model M4 (second AI model) that recognizes the facial expression or individual of person H's face F. This causes the AI model of the DSP32 to switch from the first AI model to the second AI model (S35).
[0115] Next, the information processing device 7 instructs the imaging block 20 via the crop instruction unit 140 as to the area to be cut out from the captured image (S36). Specifically, the crop instruction unit 140 instructs the imaging block 20 to cut out the area of the face F of person H. The instruction from the crop instruction unit 140 is stored in the register group 27 of the imaging block 20, whereby crop setting is performed (S37).
[0116] Next, the information processing device 7 instructs the imaging block 20, via the acquisition instructing unit 110, to perform bracket imaging, which acquires frames of captured images under different imaging conditions (S38). For example, the acquisition instructing unit 110 may instruct the imaging block 20 to perform bracket imaging, which acquires frames of captured images under different exposure conditions. As a result, the imaging block 20 acquires a captured image P4 by reading out only the image signal of the area of the face F of the person H from the captured images captured under different imaging conditions.
[0117] Thereafter, each of the acquired captured images P4 is input to a face recognition AI model M4 (second AI model) of the DSP 32. This enables the DSP 32 to recognize the face F of the person H using the face recognition AI model M4. The recognition result of the face F of the person H recognized from each of the captured images P4 is output to the information processing device 7.
[0118] Next, the information processing device 7 causes the result determination unit 150 to determine the recognition result of the face F of the person H to be finally output from the imaging device 2 based on the recognition results of the face F of the person H recognized from each of the captured images P4 captured under different imaging conditions (S39). Subsequently, the information processing device 7 transmits the recognition result determined by the processing of step S39 to the outside (S40).
[0119] After transmitting the recognition result to the outside, the information processing device 7 instructs the DSP 32 to switch the AI model via the model switching unit 130 (S41). Specifically, the model switching unit 130 instructs the DSP 32 to switch from a face recognition AI model M4 (second AI model) that recognizes the facial expression or individual of the face F of person H to a face detection AI model M3 (first AI model) that detects the face F of person H from the captured image. This causes the AI model of the DSP 32 to switch from the second AI model to the first AI model (S42).
[0120] By repeating the above process as one cycle, the digital camera 10 can recognize, for example, the facial expression or individual of the face F of the person H within the imaging range.
[0121] <6. Notes> Although the preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical idea described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.
[0122] For example, in the above embodiment, the imaging device 2 is provided with an imaging function and an AI processing function, but the present technology is not limited to this example.
[0123] As an example, the information processing device 7 may have an AI processing function.
[0124] As another example, when a function other than that of the imaging unit 21 of the imaging device 2 or a function of the signal processing block 30 of the imaging device 2 is executed by an image signal processor (ISP) separate from the imaging device 2, the image signal processor may have an AI processing function. In such a case, the image signal processor may further have the function of the information processing device 7.
[0125] In addition, in the above embodiment, an example of performing AI processing on a captured image is shown, but the present technology can also perform AI processing on sensing images (e.g., depth images or infrared images) acquired by an external sensor.
[0126] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.
[0127] The following configurations also fall within the technical scope of the present disclosure. (1) a model switching unit that instructs switching of a model of AI processing performed on a captured image from a first AI model that detects an area of a recognition target included in the captured image to a second AI model that recognizes the recognition target in the area detected from the captured image; a result determination unit that determines a recognition result to be output based on a plurality of recognition results of the recognition target that are respectively recognized by the second AI model from a plurality of captured images captured under different imaging conditions; An information processing device comprising: (2) The information processing device according to (1), wherein an image with reduced resolution from the captured image is input to the first AI model. (3) The information processing device according to (1) or (2), wherein the second AI model receives input of images cut out from the plurality of captured images of the areas detected by the first AI model. (4) a stillness determination unit that determines whether the recognition target is still based on a time-series detection result of the area by the first AI model; The information processing device according to any one of (1) to (3), wherein, when it is determined that the recognition target is stationary, recognition of the recognition target is performed by the second AI model. (5) The information processing device according to any one of (1) to (4), wherein the result determination unit determines the recognition result to be output based on a majority vote result for each element included in the plurality of recognition results. (6) the second AI model outputs a recognition result of the recognition target and a reliability of the recognition result; The information processing device according to any one of (1) to (4), wherein the result determination unit determines the recognition result to be output based on the reliability of each of the plurality of recognition results. (7) The information processing device described in (6), wherein the result determination unit determines the recognition result as the recognition result to be output when the reliability of the recognition result recognized by the second AI model from one captured image is equal to or greater than a threshold. (8) The information processing device described in (7), wherein, when the reliability of the recognition result is less than a threshold, the result determination unit determines the recognition result to be output based on multiple recognition results of the recognition target recognized by the second AI model from multiple captured images captured under different imaging conditions. (9) The information processing device according to any one of (1) to (8), wherein the imaging conditions are exposure conditions at the time of imaging. (10) The information processing device according to any one of (1) to (9), wherein the first AI model and the second AI model are each generated by machine learning. (11) further comprising an imaging unit that captures the captured image, The information processing device according to any one of (1) to (10), wherein the imaging unit, the model switching unit, and the result determination unit are provided on a single chip. (12) The information processing device described in (11) above, wherein the chip is constructed by joining together a first substrate on which the imaging unit is provided and a second substrate on which the model switching unit and the result determination unit are provided. (13) The information processing device according to (11) or (12), wherein the AI processing is performed within the chip provided with an imaging unit that captures the captured image. (14) The information processing device according to any one of (1) to (13), wherein the recognition target is a person's face, a pattern code attached to an object, or a license plate of a vehicle. (15) Obtaining a result of detecting a recognition target area included in the captured image using a first AI model; instructing the system to switch a model of AI processing performed on the captured image from the first AI model to a second AI model that recognizes the recognition target in the area detected from the captured image; Obtaining a plurality of recognition results of the recognition target, each of which is recognized by the second AI model from a plurality of captured images captured under different imaging conditions; determining an output recognition result based on the plurality of recognition results; A method for processing information by a computer, including: [Explanation of symbols]
[0128] 2. Imaging device 7. Information processing equipment 10 Digital Camera 20 Imaging Block 21 Imaging unit 22 Imaging processing unit 23 Output control section 25 Imaging control unit 27 Registers 30 Signal Processing Blocks 31 CPU 32 DSP 33 Memory 35 Image Compression Unit 110 Acquisition instruction section 120 Stationary judgment section 130 Model switching unit 140 Crop indicator 150 Results Determination Department
Claims
1. a model switching unit that instructs switching of a model of AI processing performed on a captured image from a first AI model that detects an area of a recognition target included in the captured image to a second AI model that recognizes the recognition target in the area detected from the captured image; a result determination unit that determines an output recognition result based on a plurality of recognition results of the recognition target that are respectively recognized by the second AI model from a plurality of captured images captured under different imaging conditions; An information processing device comprising:
2. The information processing device according to claim 1 , wherein an image with a reduced resolution is input to the first AI model from the captured image.
3. The information processing device according to claim 1 , wherein the second AI model receives as input images cut out from the plurality of captured images of the areas detected by the first AI model.
4. Further, a stillness determination unit is provided that determines whether the recognition target is still or not based on a time series detection result of the area by the first AI model, The information processing device according to claim 1 , wherein, when it is determined that the recognition target is stationary, the recognition target is recognized by the second AI model.
5. The information processing apparatus according to claim 1 , wherein the result determination unit determines the recognition result to be output based on a majority decision result for each element included in the plurality of recognition results.
6. The second AI model outputs a recognition result of the recognition target and a reliability of the recognition result; The information processing apparatus according to claim 1 , wherein the result determination unit determines the recognition result to be output based on the reliability of each of the plurality of recognition results.
7. 7. The information processing device according to claim 6, wherein the result determination unit determines the recognition result recognized by the second AI model from one captured image as the recognition result to be output when the reliability of the recognition result is equal to or greater than a threshold.
8. 8. The information processing device according to claim 7, wherein, when the reliability of the recognition result is less than a threshold, the result determination unit determines the recognition result to be output based on a plurality of recognition results of the recognition target recognized by the second AI model from a plurality of captured images captured under different imaging conditions.
9. The information processing apparatus according to claim 1 , wherein the image capturing conditions are exposure conditions at the time of capturing an image.
10. The information processing device according to claim 1 , wherein the first AI model and the second AI model are each generated by machine learning.
11. further comprising an imaging unit that captures the captured image, The information processing device according to claim 1 , wherein the imaging unit, the model switching unit, and the result determination unit are provided on a single chip.
12. The information processing device according to claim 11 , wherein the chip is configured by joining together a first substrate on which the imaging unit is provided and a second substrate on which the model switching unit and the result determination unit are provided.
13. The information processing device according to claim 11 , wherein the AI processing is performed within the chip provided with an imaging unit that captures the captured image.
14. The information processing apparatus according to claim 1 , wherein the recognition target is a person's face, a pattern code attached to an object, or a license plate of a vehicle.
15. Obtaining a result of detecting a recognition target area included in a captured image using a first AI model; instructing the system to switch a model of AI processing performed on the captured image from the first AI model to a second AI model that recognizes the recognition target in the area detected from the captured image; Obtaining a plurality of recognition results of the recognition target, each of which is recognized by the second AI model, from a plurality of captured images captured under different imaging conditions; determining an output recognition result based on the plurality of recognition results; A method for processing information by a computer, including:
Citation Information
Patent Citations
Imaging device and electronic device
JP6633216B2