Information processing device and information processing method
By employing multiple AI models to process images under different conditions, the imaging device enhances recognition accuracy for targets by addressing issues of low resolution and environmental interference.
Patent Information
- Application Number
- PCT/JP2025/002672
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-30
- Filing Date
- 2025-01-29
- Publication Date
- 2025-08-07
AI Technical Summary
Existing imaging devices face challenges in achieving high recognition accuracy for recognition targets due to the influence of elements other than the target and low image resolution, particularly when using single recognition processes.
Implementing multiple AI models to identify and determine features of a recognition target object across multiple images captured under different conditions, switching between a first AI model for target detection and a second AI model for target recognition, and aggregating results to enhance accuracy.
Improves recognition accuracy by leveraging multiple AI models to handle varying imaging conditions, reducing obscuration effects, and ensuring reliable recognition results.
Smart Images

Figure JP2025002672_07082025_PF_FP_ABST
Abstract
Description
INFORMATION PROCESSING DEVICE AND INFORMATION PROCESSING METHODCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of Japanese Priority Patent Application JP 2024-011794 filed January 30, 2024, the entire contents of which are incorporated herein by reference.
[0002] The present disclosure relates to an information processing device and an information processing method.
[0003] In recent years, there is known a stacked imaging device made by bonding together a substrate having an imaging unit in which multiple pixels are arranged two-dimensionally and a substrate having a signal processing unit that performs signal processes on an image captured by the imaging unit.
[0004] On the other hand, as use cases for imaging devices, cases of using not only images captured by imaging devices but also information obtained from captured images (what is generally called metadata) are increasing.
[0005] For example, PTL 1 listed below discloses an imaging device made by joining together a first substrate having an imaging unit that captures an image and a second substrate having a recognition processing unit that performs recognition processing on the captured image output by the imaging unit.
[0006] Japanese Patent No. 6633216Summary
[0007] However, as an example of just one technical problem addressed by the present disclosure, in the imaging device disclosed in PTL 1, since only a single recognition process is performed on the captured image by the recognition processing unit, it may be difficult to obtain metadata with high recognition accuracy, depending on the type of metadata. For example, in a case where an image contains many elements other than the recognition target, it may be difficult to recognize the recognition target, per se, due to the influence of the elements other than the recognition target. Further, in a case where the resolution of the image of the recognition target is low, the recognition accuracy for the recognition target is likely to decrease.
[0008] In view of the above circumstances, it is desirable to obtain recognition results having higher recognition accuracy, by performing multiple AI (Artificial Intelligence) processes on a captured image.Solutions to Problems
[0009] According to an embodiment of the present disclosure, and as an example of a solution that addresses one of the problems presented herein, there is provided an information processing device circuitry configured to identify from image data based on a first captured image a subportion area that contains a recognition target object, the recognition target object containing features, and determine the features of the recognition target object based on image data of a plurality of captured images that respectively include the recognition target object, and each captured under a different condition, the plurality of captured images at least include images other than the first captured image. The term AI model is implemented programmed circuitry and is sometimes referred to as an AI engine.
[0010] Further, according to the embodiment of the present disclosure, there is provided an information processing method performed by a computer, and the method includes obtaining a result of detecting an area of a recognition target included in a captured image by using a first AI model, giving an instruction to switch a model of AI processing that is performed on the captured image from the first AI model to a second AI model that recognizes the recognition target in the area detected from the captured image, obtaining multiple recognition results of the recognition target recognized by the second AI model by using multiple images captured under different imaging conditions, respectively, and determining a recognition result to be output, on the basis of the multiple recognition results.
[0011] FIG. 1 is an explanatory diagram illustrating a digital camera in which the present technology is implemented.FIG. 2 is a block diagram illustrating a configuration example of the digital camera to which the present technology is applied.FIG. 3 is a block diagram illustrating a configuration example of an imaging device according to an embodiment of the present disclosure.FIG. 4 is a perspective view illustrating an example of the external configuration of the imaging device.FIG. 5 is a block diagram illustrating a configuration example of an information processing device.FIG. 6 is a flow chart illustrating the flow of processing in the imaging device and the information processing device.FIG. 7 is an explanatory diagram illustrating an example of a method for determining a recognition result to be output from multiple recognition results.FIG. 8 is an explanatory diagram illustrating another example of the method for determining the recognition result to be output from multiple recognition results.FIG. 9 is a sequence diagram illustrating the exchange of data between an imaging block, a DSP (Digital Signal Processor), and the information processing device.FIG. 10 is a flow chart illustrating the processing flow related to a modification example of the digital camera.FIG. 11 is a flow chart illustrating the processing flow for an application example using the digital camera.Description of Embodiment
[0012] An embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings. Note that, in the present specification and the drawings, components having substantially the same functional configurations are denoted by the same reference signs and redundant description will be omitted.
[0013] Incidentally, the description will be given in the following order. 1. Overview 2. Configuration 2.1. Configuration of digital camera 2.2. Configuration of imaging device 2.3. Configuration of information processing device 3. Processing flow 4. Modification example 5. Application example 6. Supplementary note
[0014] <1. Overview> First, an overview of the technology according to the present disclosure will be described with reference to FIG. 1. FIG. 1 is an explanatory diagram illustrating a digital camera 10 in which the technology is implemented.
[0015] The present technology is implemented in the digital camera 10 that monitors a parking lot S in which a vehicle C is parked and optionally other vehicles C are parked, for example, as illustrated in FIG. 1.
[0016] The digital camera 10 captures an image of the vehicle C parked in the parking lot S, and recognizes the contents of a number sign (what is generally called a number plate in Japan, and a license plate in the U.S., for example) N attached to the vehicle C, from the captured image of the vehicle C by using an AI model. The digital camera 10 can identify the vehicle C parked in the parking lot S, by recognizing the registration number of the vehicle C written on the number sign N.
[0017] To be specific, the digital camera 10 first detects the area (subportion) of the number sign N attached to the vehicle C, by inputting a captured image of the vehicle C into a first AI model. Next, the digital camera 10 can recognize the registration number (written as alphanumeric characters) of the vehicle C written on the number sign N, by inputting the area of the number sign N cut out from (or extracted from) the captured image, into a second AI model. Thus, application of the first AI model serves as an image area detector, which detects an area of recognition target included in the captured image. According to this, the digital camera 10 can compare the registration number of the vehicle C recognized from the number sign N in the captured image with a registration number of a vehicle for which a contract is made by a contractor of the parking lot S, thereby making it possible to detect, at an early stage, a suspicious, or unauthorized, vehicle or the like that has entered the parking lot S, for example.
[0018] The digital camera 10 in which the present technology is implemented can detect the area of the number sign N in a captured image by using the first AI model and can then recognize the contents of the number sign N in the detected area by using the second AI model. In this way, the application of the second AI model serves as an image content recognition detector. As a result, the digital camera 10 can perform recognition processing by using an image centered on the number sign N to be recognized, thereby further improving the accuracy of recognition results for the contents of the number sign N.
[0019] In addition, the digital camera 10 in which the present technology is implemented recognizes each of the number signs N in multiple captured images taken under different imaging conditions (e.g. different exposure levels, angles of image capture, focal lengths, optical or Infrared wavelengths, in recognizing the contents of the number sign N by using the second AI model, and then determines a final recognition result from a compilation of each recognition result of the multiple captured images. This allows the digital camera 10 to prevent a decrease in the recognition accuracy of the recognition target due to the imaging conditions of the captured image, including whether the number sign N is covered by an filter or a coating that obscures the alphanumeric characters when viewed at an angle, but does not obscure the characters when viewed orthogonally (head-on, or 90 degrees). Such filters and coating may be popular by vehicle owners to obscure the characters in their attempt to evade toll collections, speed detectors, and / or red-light cameras.
[0020] The present technology is implemented, for example, as an imaging device included in the digital camera 10, or an information processing device that controls the imaging device. That is, the present technology may be implemented as an information processing device having only a control function for controlling an imaging device, or may be implemented as an imaging device having an imaging function, a recognition function, and a control function. Hereinafter, a configuration will be described in more detail by taking as an example a case in which the present technology is implemented as an information processing device having only a control function for controlling an imaging device.
[0021] <2. Configuration> (2.1. Configuration of digital camera) Next, the configuration of the digital camera 10 to which the present technology is applied will be described with reference to FIG. 2. FIG. 2 is a block diagram illustrating a configuration example of the digital camera 10 to which the present technology is applied.
[0022] As illustrated in FIG. 2, the digital camera 10 includes a housing that houses an optical system 1, an imaging device 2, a memory 3, a signal processing unit 4, an output unit 5, a control unit 6, and an information processing device 7. The digital camera 10 is an electronic device capable of capturing both still and moving images.
[0023] The optical system 1 includes a zoom lens, a focus lens, an aperture, and the like, and guides incident light from the outside to the imaging device 2.
[0024] The imaging device 2 performs photoelectric conversion on the incident light guided by the optical system 1, thereby generating image data corresponding to the incident light. The imaging device 2 may be a CMOS (Complementary Metal Oxide Semiconductor) image sensor, for example. Further, the imaging device 2 may also detect a recognition target from image data by using the first AI model and recognize the detected recognition target by using the second AI model. While the imaging device 2 may be referred to as a “device” it may also be referred to as circuity, and / or a “system” when describing the sensor itself, or in combination with other devices such as local or remote processing devices (e.g., microprocessors, GPUs, cloud computing resources).
[0025] The memory 3 temporarily stores the image data or the recognition result output by the imaging device 2. While the term “the memory” is used, it should be understood that it is not intended to exclusively describe a single memory device, but may be one or more memory devices.
[0026] The signal processing unit 4 performs as needed, based on the application, signal processing such as noise removal and white balance adjustment on the image data stored in the memory 3. The image data that has been subjected to signal processing by the signal processing unit 4 is output to the output unit 5. The signal processing unit 4 is implemented in circuitry. The circuitry is programmable (one or more programmable CPUs, GPUs, cloud resources or the like that are configured to perform operations by execution of computer readable code), or hardwired (e.g., application specific integrated circuitry, ASIC, or programmable array logic, PAL), or a combination of one or more programmable and / or hardwired circuitry.
[0027] The output unit 5 outputs, to the outside, the image data that has been subjected to signal processing by the signal processing unit 4 or the recognition result stored in the memory 3. The output unit 5 may be a display device that displays an image corresponding to the image data as a through image, may be a writing device (drive) that causes a storage medium to store the image data or the recognition results, or may be an output interface that transmits the image data or the recognition results to an external device, for example.
[0028] The control unit 6 is a control device that controls each part of the digital camera 10 according to user operations and the like. The control unit 6 is also implemented in circuitry (programmable and / or hardwired, as discussed above).
[0029] The information processing device 7 (which also is implemented in circuitry, as discussed above) gives an instruction to switch between the first AI model and the second AI model used in the imaging device 2. Specifically, the information processing device 7 may instruct the imaging device 2 to switch between the first AI model that detects a recognition target from image data and the second AI model that recognizes the recognition target (e.g., the alphanumeric content of a license plate).
[0030] In addition, the information processing device 7 determines the recognition result to be output to the memory 3, on the basis of multiple recognition results of the recognition target recognized by the second AI model. To be specific, the information processing device 7 may determine the recognition result to be output to the memory 3, by comparing and making a determination regarding the multiple recognition results of the recognition target recognized by using multiple pieces of image data captured under different imaging conditions, respectively. The imaging conditions include different exposure levels, different angles with respect to the main face of the license plate with respect to a boresight of the image capture device. At 90 degrees, the image capture device directly pointed at the main face of the license plate. At something different than 90 degrees, the image capture device captures the image at a non-orthogonal angle, which may give rise to some obscuration of the contents of the lettering on the license plate, especially if a filter and / or coating is applied over the license plate. Aggregating multiple images and applying them to the second AI model improves the likelihood of accurately recognizing the characters on the license plate.
[0031] (2.2. Configuration of imaging device) Next, the configuration of the imaging device 2 according to the present embodiment will be described with reference to FIGS. 3 and 4.
[0032] FIG. 3 is a block diagram illustrating a configuration example of the imaging device 2 according to the present embodiment. As illustrated in FIG. 3, the imaging device 2 includes an imaging block 20 and a signal processing block 30. The imaging block 20 and the signal processing block 30 are connected to each other via connection lines CL1, CL2, and CL3, which are internal buses, in such a manner as to be able to exchange data therebetween.
[0033] The imaging block 20 is a functional block that executes an imaging function, and includes an imaging unit 21, an imaging processing unit 22, an output control unit 23, an output I / F (Interface) 24, and an imaging control unit 25.
[0034] The imaging unit 21 includes a two-dimensional array of multiple pixels. The imaging unit 21 captures an image by being driven by the imaging processing unit 22. Specifically, the imaging unit 21 receives, at each pixel, light incident via the optical system 1 and performs photoelectric conversion on the received light to generate an analog image signal corresponding to the light incident on the imaging device 2. The type of image generated by the imaging unit 21 may be an RGB (Red, Green, Blue) color image, or a monochrome image containing only luminance information, for example.
[0035] The imaging processing unit 22 performs processing related to image capturing executed by the imaging unit 21, under the control of the imaging control unit 25. For example, the imaging processing unit 22 may drive the imaging unit 21, perform AD (Analog to Digital) conversion of an analog image signal generated by the imaging unit 21, perform the imaging signal processing, and perform other processing. Note that the imaging signal processing includes, for example, a process of calculating the brightness of each predetermined small region by calculation of the average pixel value for each predetermined small region, a process of converting the image generated by the imaging unit 21 into an HDR (High Dynamic Range) image, a defect correction process, a development process, etc.
[0036] The imaging processing unit 22 performs AD conversion or the like on the analog image signal generated by the imaging unit 21, to generate a digital image signal. The generated digital image signal is output to the output control unit 23 and also to an image compression unit 35 of the signal processing block 30 via the connection line CL2.
[0037] The output control unit 23 controls selective output from the output I / F 24 to the outside (e.g., the memory 3 in FIG. 2). To be specific, the output control unit 23 may control the external output of the image signal output from the imaging processing unit 22, and may also control the external output of the results of the captured image recognition performed by the signal processing block 30 to be described later.
[0038] The output I / F 24 is an interface that outputs various kinds of information to the outside according to the control of the output control unit 23. For example, the output I / F 24 may output, to the outside, an image signal supplied from the imaging processing unit 22 and a result of recognition of the captured image executed by the signal processing block 30 to be described later.
[0039] For example, in a case where the imaging device 2 is only required to output the recognition results of the captured image (what is generally called metadata), the output control unit 23 may control the output I / F 24 to output only the recognition results of the captured image to the outside. In such a case, the output control unit 23 can reduce the amount of data output from the output I / F 24 to the outside, thereby making it possible to output data to the outside faster.
[0040] The imaging control unit 25 controls the imaging processing unit 22 according to the imaging information stored in a register group 27, to control image capturing executed by the imaging unit 21.
[0041] The register group 27 includes multiple registers that store setting information related to image capturing executed by the imaging unit 21. For example, the register group 27 may store setting information related to image capturing received from the outside via a communication I / F 26, or the results of imaging signal processing by the imaging processing unit 22. Examples of setting information related to imaging include information for setting an analog gain, exposure, a shutter speed, a frame rate, focus, white balance, or a cropping range, for example.
[0042] Note that the register group 27 may further store setting information related to output control in the output control unit 23. The output control unit 23 can control selective output to the outside according to the setting information related to output control.
[0043] The communication I / F 26 is an interface for transmitting and receiving data to and from the outside (e.g., the control unit 6 in FIG. 2). The communication I / F 26 may transmit and receive various kinds of setting information stored in the register group 27 to and from the outside, for example.
[0044] The signal processing block 30 is a functional block that executes a recognition function for recognizing a captured image, and includes a CPU (Central Processing Unit) 31, a DSP 32, a memory 33, a communication I / F 34, the image compression unit 35, and an input I / F 36. The CPU 31, the DSP 32, the memory 33, the communication I / F 34, the image compression unit 35, and the input I / F 36 are connected via a bus, so that they can input and output information to and from each other.
[0045] The CPU 31 executes a program stored in the memory 33, to control the signal processing block 30 and to input and output information to and from the imaging control unit 25 via the connection line CL1.
[0046] For example, the CPU 31 may instruct the imaging control unit 25 to capture an image, via the connection line CL1. Further, the CPU 31 may write setting information related to imaging to the register group 27 of the imaging control unit 25 to give an instruction to perform imaging based on the written setting information. This allows the CPU 31 to instruct the imaging block 20 to perform imaging while changing the imaging conditions (for example, bracket imaging with changed exposure).
[0047] The DSP 32 executes a program stored in the memory 33, to perform AI processing on the captured image supplied from the imaging processing unit 22 via the connection line CL2.
[0048] The contents of the AI processing executed by the DSP 32 are determined according to the contents of the AI model held in the DSP 32. For example, the DSP 32 may perform AI processing to detect the position or the area of a recognition target from a captured image by using the first AI model. In addition, the DSP 32 may perform AI processing to recognize the contents of a recognition target detected from a captured image, by using the second AI model. The DSP 32 is capable of switching the contents of the AI processing to be executed (e.g., switch from executing the first AI model to executing the second AI model), by switching the contents of the stored AI model on the basis of instructions from the information processing device 7. Moreover, the switching may involve loading executable code of the second AI model so the DSP 32 may execute it after the DSP 32 has finished executing the executable code that configures the DSP 32 to implement the first AI model.
[0049] The first AI model and the second AI model may be generated by a machine learning algorithm such as a DNN (Deep Neural Network) or a Transformer. In addition, the first AI model and the second AI model may be generated by a rule-based algorithm. The first AI model may initially be trained on labeled training images that identify the recognition target object in a variety of scenarios, and then subsequently enhanced via repeated leaning on newly captured images during execution. The second AI model may be trained using images that have the alphanumeric characters obscured in various ways, and from different camera to license plate perspectives that are 90 degrees, or different angles from less than 90 degrees or greater than 90 degrees. Likewise, the images may be taken at different exposure levels, distances, color license plates, alphanumeric fonts, focal lengths, etc.
[0050] The memory 33 stores data and the like to be used in the processing of the signal processing block 30. The memory 33 may include an SRAM (Static Random Access Memory), a DRAM (Dynamic Random Access Memory), or the like. For example, the memory 33 may store a program received from the outside via the communication I / F 34, a captured image output from the imaging block 20, the results of AI processing performed by the DSP 32, setting information received from the input I / F 36, and the like.
[0051] The communication I / F 34 is an interface for transmitting and receiving data (wirelessly or via physical conductors) to and from the outside (e.g., the memory 3 or the control unit 6 in FIG. 2). The communication I / F 34 may transmit and receive such information as a program executed by the CPU 31 or the DSP 32 to and from the outside. For example, the communication I / F 34 may cause the memory 33 to store a program downloaded from the outside to be executed by the CPU 31 or the DSP 32.
[0052] The image compression unit 35 performs compression processing on the captured image passed from the imaging processing unit 22 via the connection line CL2. This enables the image compression unit 35 to generate a compressed image having a smaller amount of data than that of the captured image that has not yet been subjected to compression.
[0053] The AI processing in the DSP 32 may be performed on either the image captured by the imaging block 20 or the compressed image generated by the image compression unit 35. Since the compressed image has a smaller amount of data than that of the captured image, it is possible to reduce the load of AI processing in the DSP 32 and to save the storage capacity of the memory 33 that stores the compressed image. For example, a compressed image may be input to the first AI model that detects a recognition target from a captured image, to reduce the load. On the other hand, a captured image may be input to the second AI model that recognizes the detected recognition target, to improve recognition accuracy.
[0054] Examples of the compression process in the image compression unit 35 include a compression process for scaling down a captured image of 12M (3968 × 2976) pixels to an image of VGA (640 × 480) size, for example.
[0055] The input I / F 36 is an interface that receives various kinds of information from the outside (e.g., an external sensor). For example, the input I / F 36 may receive sensing results obtained by an external sensor. The received sensing results can be subjected to AI processing in the DSP 32, for example. Examples of external sensors include a distance measurement sensor that acquires depth information, an infrared sensor that acquires infrared images, or an image sensor different from the imaging device 2, for example.
[0056] The imaging device 2 having the above configuration can perform, on the image captured by the imaging unit 21, AI processing using the first AI model that detects a recognition target from the captured image and AI processing using the second AI model that recognizes the contents of the detected recognition target. This allows the imaging device 2 to output the recognition result of the recognition target from the output I / F 24.
[0057] FIG. 4 is a perspective view illustrating an example of the external configuration of the imaging device 2. As illustrated in FIG. 4, the imaging device 2 is configured as a semiconductor device having a laminated structure in which a first substrate 51 and a second substrate 52 are stacked.
[0058] The first substrate 51 is provided with the imaging unit 21, for example. On the other hand, the second substrate 52 is provided with, for example, the imaging processing unit 22, the output control unit 23, the output I / F 24, the imaging control unit 25, the CPU 31, the DSP 32, the memory 33, the communication I / F 34, the image compression unit 35, and the input I / F 36.
[0059] As an example, the first substrate 51 and the second substrate 52 may be electrically connected to each other by a through via that penetrates the first substrate 51 and reaches the second substrate 52. Alternatively, as another example, the first substrate 51 and the second substrate 52 may be electrically connected to each other by Cu-Cu bonding that directly bonds the Cu wiring exposed on the lower surface of the first substrate 51 to the Cu wiring exposed on the upper surface of the second substrate.
[0060] It is to be noted that, if it is acceptable for the area of the imaging device 2 to be larger, the imaging device 2 may be formed as a single substrate.
[0061] Further, the imaging device 2 may be configured by stacking three or more substrates. In such a case, the imaging device 2 may be constructed by stacking a first substrate on which the imaging unit 21 is provided, a second substrate on which the imaging processing unit 22, the output control unit 23, the output I / F 24, the imaging control unit 25, the CPU 31, the DSP 32, the communication I / F 34, the image compression unit 35, and the input I / F 36 are provided, and a third substrate on which the memory 33 is provided.
[0062] (2.3. Configuration of information processing device) Further, the configuration of the information processing device 7 according to the present embodiment will be described with reference to FIG. 5. FIG. 5 is a block diagram illustrating a configuration example of the information processing device 7 according to the present embodiment.
[0063] As illustrated in FIG. 5, the information processing device 7 includes an acquisition instructing section 110, a stillness determining section 120, a model switching section 130, a cropping instructing section 140, and a result determining section 150. Note that these respective “sections” are implemented in processing circuitry (e.g., information processing device 7), where there circuitry can be a single circuit (e.g., CPU or GPU-based computer, or multiprocessor-based circuitry with a local processor and one or more external processing circuitry devices).
[0064] The acquisition instructing section 110 instructs the imaging block 20 to acquire a captured image for AI processing performed by the DSP 32. As a result, the imaging control unit 25, the imaging processing unit 22, and the imaging unit 21 of the imaging block 20 operate to acquire a captured image.
[0065] In addition, the captured image acquired based on instructions from the acquisition instructing section 110 is compressed or a partial area of the captured image is cut out (e.g., cropped) to fit the input of the AI model. For example, in AI processing for detecting the position or the area of a recognition target from a captured image, a compressed image obtained by compressing the captured image is input to the first AI model in order to reduce the processing load. In addition, in AI processing that recognizes a recognition target detected in a captured image, an image of the recognition target area cut out from the captured image is input to the second AI model in order to improve recognition accuracy and reduce processing load.
[0066] The model switching section 130 switches the AI model used for the AI processing of the DSP 32. By switching the AI model by the model switching section 130, the DSP 32 can perform various types of AI processing. As one example, the model switching section 130 may switch the AI model to be used by the DSP 32, by designating an AI model from among multiple AI models stored in the memory 33 or the like. As another example, the model switching section 130 may switch the AI model used by the DSP 32, by uploading the AI model to the DSP 32.
[0067] For example, the model switching section 130 may switch, as the AI model used for the AI processing of the DSP 32, between the first AI model that detects the area of the recognition target from the captured image and the second AI model that recognizes the recognition target detected in the captured image.
[0068] The stillness determining section 120 determines whether the recognition target detected in the captured image is stationary or not. To be specific, the stillness determining section 120 may determine that the recognition target is stationary, in a case where the amount of time-series variation in the area or the position of the recognition target detected in the captured image satisfies a predetermined condition. The predetermined condition is satisfied, for example, in a case where the amount of change in the area or the position of the recognition target is equal to or less than a threshold value over a predetermined number of frames.
[0069] In a case where the stillness determining section 120 determines that the recognition target is stationary, the model switching section 130 switches the AI model of the DSP 32 from the first AI model to the second AI model. This allows the second AI model to carry out recognition of a recognition target, for the recognition target in the area detected by the first AI model.
[0070] The cropping instructing section 140 indicates an area to be cut out (cropped) from the captured image to input the area to the second AI model of the DSP 32. To be specific, the cropping instructing section 140 writes, in the register group 27 of the imaging block 20, an instruction (cropping instruction) to cut out the area to be recognized in the captured image detected in the first processing by the first AI model. Owing to this, an image with a smaller amount of data obtained by cutting out the area to be recognized in the captured image is input to the second AI model of the DSP 32. Therefore, the cropping instructing section 140 can reduce the processing load while improving the accuracy of recognition performed by the second AI model of the DSP 32.
[0071] Note that the cropping instructing section 140 may give an instruction to cut out the area to be recognized, from the captured image output by the imaging unit 21, or may give an instruction to cut out and read out only the image signal of the area to be recognized, from the imaging unit 21.
[0072] In a case where the model switching section 130 switches the AI model of the DSP 32 and the cropping instructing section 140 gives a cropping instruction, the acquisition instructing section 110 instructs the imaging block 20 to capture images under multiple imaging conditions. The imaging device 2 according to the present embodiment recognizes the recognition target for each of the images captured under multiple imaging conditions, and compares and makes a determination regarding recognition results in the result determining section 150, to determine the recognition result to be output from the output I / F 24.
[0073] The multiple imaging conditions are conditions in which one type of imaging parameter is varied, for example. Examples of the imaging parameter to be varied include exposure, a gain, contrast, a color temperature, or white balance, for example.
[0074] The result determining section 150 determines the recognition result to be output to the memory 3, on the basis of the recognition results of the recognition target recognized from multiple images captured under different imaging conditions, respectively. This allows the result determining section 150 to determine the recognition result to be output to the memory 3, by referring to the recognition result obtained with use of the image captured under the imaging condition that provides higher recognition accuracy among the multiple imaging conditions. Therefore, the information processing device 7 can improve the recognition accuracy of the recognition result to be output to the memory 3.
[0075] For example, the result determining section 150 may determine, as the recognition result to be output to the memory 3, the most reliable recognition result among the recognition results of the recognition target respectively recognized from multiple images captured under different imaging conditions. The degree of reliability is an index indicating the certainty factor of the recognition result, and is output from the second AI model that recognizes the recognition target, together with the recognition result. As the degree of reliability is higher, the certainty factor becomes higher, and the degree of reliability is expressed in such a range of 0 to 1 or 0 to 100, for example.
[0076] Alternatively, the result determining section 150 may decompose the recognition result of the recognition target recognized from each of multiple images captured under different imaging conditions into individual elements, and determine the recognition result to be output to the memory 3, by using a majority rule for each element (e.g., each character).
[0077] <3. Processing flow> Next, the processing flow of the imaging device 2 and the information processing device 7 according to the present embodiment will be described with reference to FIGS. 6 to 8. FIG. 6 is a flowchart illustrating the flow of processing in the imaging device 2 and the information processing device 7. FIG. 7 is an explanatory diagram illustrating an example of a method for determining a recognition result to be output from multiple recognition results. FIG. 8 is an explanatory diagram illustrating another example of the method for determining the recognition result to be output from multiple recognition results.
[0078] As illustrated in FIG. 6, first, the information processing device 7 instructs the imaging block 20 to acquire a frame of a captured image, by using the acquisition instructing section 110 (S10). As a result, the imaging block 20 acquires a captured image P1 including the number sign N (recognition target) attached to the vehicle C.
[0079] The captured image P1 is compressed by a compression process that reduces the resolution, and then input to a target detection AI model M1 (first AI model) of the DSP 32. As a result, the DSP 32 detects the number sign N, which is the recognition target, by using the target detection AI model M1, and outputs the coordinates of the area (bounding box) of the detected number sign N (S11).
[0080] Next, the information processing device 7 determines whether the number sign N to be recognized is stationary or not, by the stillness determining section 120 (S12). To be specific, the stillness determining section 120 may determine that the number sign N is stationary, in a case where the amount of change in the center coordinates of the area of the number sign N to be recognized is equal to or less than a threshold value over a predetermined number of frames.
[0081] In a case where the number sign N is not stationary (S12 / NO), the information processing device 7 returns to the operation of step S10 and instructs again the imaging block 20 to acquire a frame of a captured image.
[0082] Meanwhile, in a case where the number sign N is stationary (S12 / YES), the information processing device 7 transmits the coordinates of the area of the number sign N to be recognized to the outside (S13), and then instructs the DSP 32 to switch the AI model via the model switching section 130 (S14). To be specific, the model switching section 130 instructs the DSP 32 to switch the model from the first AI model that detects the area of the recognition target from the captured image to the second AI model that recognizes the contents of the recognition target detected in the captured image. This causes the AI model of the DSP 32 to be switched from the first AI model to the second AI model (S15).
[0083] Next, the information processing device 7 indicates, to the imaging block 20, an area to be cut out (cropped) from the captured image, by the cropping instructing section 140 (S16). Specifically, the cropping instructing section 140 instructs the imaging block 20 to cut out the area (bounding box) of the number sign N to be recognized. The instruction from the cropping instructing section 140 is stored in the register group 27 of the imaging block 20, and the cropping setting is thereby performed (S17).
[0084] Next, the information processing device 7 instructs the imaging block 20 via the acquisition instructing section 110 to perform bracket imaging for acquiring frames of captured images under different imaging conditions (S18). For example, the acquisition instructing section 110 may instruct the imaging block 20 to perform bracket imaging in which frames of captured images of the vehicle C are acquired under different exposure conditions. In other words, the acquisition instructing section 110 may instruct the imaging block 20 to perform bracket imaging whose exposure (or angle at which the image is captured) has been changed.
[0085] In the imaging block 20, a cropping setting is performed in processing in step S17 to cut out the area of the number sign N to be recognized. Therefore, the imaging block 20 acquires captured images P2 obtained by reading out only the image signal of the area of the number sign N to be recognized.
[0086] Thereafter, each of the acquired captured images P2 is input to a target recognition AI model M2 (second AI model) of the DSP 32. This enables the DSP 32 to recognize the registration number of the vehicle C marked on the number sign N, which is the recognition target, by using the target recognition AI model M2. The registration number of the vehicle C recognized in each of the captured images P2 (recognition result) is output to the information processing device 7.
[0087] Next, the information processing device 7 determines, in the result determining section 150, the registration number (recognition result) of the vehicle C to be finally output to the outside, on the basis of the registration number (recognition result) of the vehicle C recognized from each of the captured images P2 captured under different imaging conditions (S19). Moreover the alphanumeric characters detected as being present on the license plate are identified as the recognition result.
[0088] For example, as illustrated in FIG. 7, in the target recognition AI model M2 of the DSP 32, the degree of reliability (confidence) of the recognition result is assumed to be output together with the recognition result. The degree of reliability (confidence) is an index indicating that, as the value is higher, the certainty factor of the recognition result is higher, and the range is 0 to 100, for example. For example, assume that a degree of reliability (confidence) of "70" is output for a first recognition result R1, a degree of reliability (confidence) of "60" is output for a second recognition result R2, and a degree of reliability (confidence) of "90" is output for a third recognition result R3. In such a case, the result determining section 150 can determine the third recognition result R3 having the highest degree of reliability (confidence), as a recognition result RF to be finally output from the imaging device 2.
[0089] Also, as illustrated in FIG. 8, the result determining section 150 can determine the recognition result RF that is finally output from the imaging device 2, by using a majority rule for the first recognition result R1, the second recognition result R2, and the third recognition result R3 captured under different imaging conditions. Specifically, the result determining section 150 first breaks down the first recognition result R1 into the elements "TKA," "500," "NA," and "1234." Similarly, the result determining section 150 breaks down the second recognition result R2 into the elements "TKA," "500," "HA," and "1234," and breaks down the third recognition result R3 into the elements "TKA," "503," "NA," and "1234." The result determining section 150 can determine the recognition result including the elements "TKA," "500," "NA," and "1234" which are majority elements as the recognition result RF, by using a majority rule for the corresponding elements of the first recognition result R1, the second recognition result R2, and the third recognition result R3.
[0090] Next, the information processing device 7 transmits, to the outside, the recognition result determined by the process of step S19 (S20).
[0091] After transmitting the recognition result to the outside, the information processing device 7 instructs the DSP 32 to switch the AI model via the model switching section 130 (S21). To be specific, the model switching section 130 instructs the DSP 32 to switch from the second AI model that recognizes the contents of the recognition target detected in the captured image to the first AI model that detects the area of the recognition target from the captured image. This causes the AI model of the DSP 32 to be switched from the second AI model to the first AI model (S22).
[0092] By repeating the above process as one cycle, the imaging device 2 and the information processing device 7 can recognize the registration number of the vehicle C parked in the parking lot S, for example. Therefore, the digital camera 10 is capable of monitoring the situation of the parking lot S in which the vehicle C and the like are parked. Likewise, the digital camera is able to detect the license plate contents when moving on a road (e.g., speed or detector), or through an intersection (red light detector).
[0093] Next, the transfer of data processed by the imaging device 2 will be described with reference to FIG. 9. FIG. 9 is a sequence diagram illustrating the transfer of data between the imaging block 20, the DSP 32, and the information processing device 7.
[0094] As illustrated in FIG. 9, a start instruction is input to the information processing device 7 from a console 70 operated by a user who manages the digital camera 10, for example (S101). As a result, an instruction to acquire a frame of a captured image is output from the information processing device 7 to the imaging block 20 (S103).
[0095] In the imaging block 20, after the frame of the captured image is acquired (S105), image processing such as compression processing is performed on the acquired frame of the captured image (S107). The captured image that has been subjected to image processing is input to a target detection AI model (first AI model) of the DSP 32 (S109). As a result, the target detection AI model detects the target to be recognized, from the captured image (S111), and coordinate data regarding the area (bounding box) of the recognition target is output.
[0096] The coordinate data regarding the detected recognition target area is output to the information processing device 7, and post-processing is performed on the data in the information processing device 7 (S113). The coordinate data regarding the recognition target area that has been subjected to post-processing is output to the console 70 (S115).
[0097] Thereafter, an instruction to switch the AI model is output from the information processing device 7 to the DSP 32 (S117). As a result, in the DSP 32, the AI model is switched from a target detection AI model (first AI model) that detects the area of the recognition target from the captured image to a target recognition AI model (second AI model) that recognizes the contents of the recognition target detected in the captured image (S119). Further, the information processing device 7 outputs a cropping instruction to the imaging block 20 to cut out an area to be recognized, from the captured image (S121).
[0098] Next, the information processing device 7 outputs an instruction to the imaging block 20 to acquire a frame of a captured image (S123). To be specific, the information processing device 7 instructs the imaging block 20 to perform bracket imaging for acquiring frames of images captured under different imaging conditions (e.g., exposure conditions).
[0099] In the imaging block 20, imaging conditions are changed, and the multiple frames of captured images are acquired (S125). The range of the captured image frames acquired at this time may be only the area of the detected recognition target, for example. After image processing is performed on the multiple acquired frames of the captured images (S127), the captured images are input to the target recognition AI model (second AI model) of the DSP 32 (S129). As a result, the target recognition AI model recognizes the target to be recognized (S131), and recognition results indicating the contents of the recognition target are output, respectively.
[0100] The output recognition result is output to the information processing device 7, and post-processing is performed on the result in the information processing device 7 (S133). For example, the recognition result to be finally output to the console 70 is determined based on the recognition result resulting from recognition of each of the multiple captured image frames. Thereafter, a recognition result determined based on the multiple recognition results is output to the console 70 (S135).
[0101] By transmitting and receiving the above data, the digital camera 10 can obtain a captured image on the basis of an instruction from the console 70 and output, to the console 70, the recognition result of a recognition target included in the captured image, for example.
[0102] <4. Modification example> Next, a modification example of the digital camera 10 according to the present embodiment will be described with reference to FIG. 10. FIG. 10 is a flow chart illustrating the flow of processing according to a modification example of the digital camera 10. In a modification example of the digital camera 10, between the cropping instruction (S16) and frame acquisition (S18) illustrated in FIG. 6, AI processing to recognize the recognition target from the captured image and output the degree of reliability of the recognition result and processing to determine whether or not to carry out bracket imaging are performed.
[0103] To be specific, as illustrated in FIG. 10, after a cropping instruction indicating the area to be cut out (cropped) from the captured image is output to the imaging block 20, the information processing device 7 instructs the imaging block 20 to acquire a frame of the captured image, via the acquisition instructing section 110. As a result, the imaging block 20 acquires a captured image obtained by reading out only the image signal of the area of the number sign N to be recognized (S25). The imaging conditions of the captured image acquired at this time may be the same as the imaging condition applied when the frame is acquired in step S10, for example.
[0104] Next, the acquired captured image is input to the target recognition AI model M2 (second AI model) of the DSP 32. As a result, the DSP 32 recognizes the registration number of the vehicle C written on the number sign N, which is the recognition target, by using the target recognition AI model M2 (S26). Accordingly, the DSP 32 outputs the recognized registration number of the vehicle C (recognition result) and the degree of reliability (confidence) of the recognition result to the information processing device 7. As described above, the degree of reliability (confidence) is an index that indicates that, as the value is higher, the certainty factor of the recognition result is higher.
[0105] Next, the information processing device 7 determines whether or not the degree of reliability (confidence) of the recognition result having been output is equal to or greater than a threshold value (S27).
[0106] In a case where the degree of reliability (confidence) is equal to or greater than the threshold (S27 / YES), the information processing device 7 determines that a recognition result with a sufficient degree of reliability has been obtained, and determines the recognition result obtained in step S26, as the recognition result to be finally output to the outside. Then, the information processing device 7 transmits the determined recognition result to the outside of the digital camera 10 (S20).
[0107] Meanwhile, in a case where the degree of reliability (confidence) is less than the threshold value (S27 / NO), the information processing device 7 determines that a recognition result with a sufficient degree of reliability has not been obtained, and instructs the imaging block 20 to perform bracket imaging for acquiring frames of the captured images under different imaging conditions (S18). As a result, the imaging block 20, the DSP 32, and the information processing device 7 perform the processes of steps S18 and S19 to determine the recognition result to be finally output to the outside.
[0108] In the modification example of the digital camera 10 according to the present embodiment described above, the information processing device 7 causes the imaging device 2 to perform bracket imaging, which varies the imaging conditions, only in a case where the degree of reliability of the recognition result of the target recognition AI model M2 falls below a threshold value. This allows the information processing device 7 to reduce the number of bracket imaging operations when recognizing a recognition target, thereby reducing the power consumption of the digital camera 10.
[0109] <5. Application Example> In the above, the digital camera 10 detects, as the recognition target, the number sign N attached to the vehicle C, and recognizes the registration number of the vehicle C written on the number sign N, but the present technology is not limited to such examples. The digital camera 10 according to the present embodiment is capable of detecting various objects in a captured image as recognition targets and recognizing the contents of the detected objects.
[0110] As one example, the digital camera 10 can detect a person's face from a captured image and perform facial expression recognition, personal recognition, sight-line estimation, or the like from the detected face. As another example, the digital camera 10 can detect a pattern code such as a QR (Quick Response) code (registered trademark) or a barcode attached to an item from a captured image and recognize the contents of the detected pattern code.
[0111] Hereinbelow, with reference to FIG. 11, an application example in which the digital camera 10 detects a person's face as a recognition target from within a captured image and recognizes the expression or the individual of the detected face will be described. FIG. 11 is a flowchart illustrating the flow of processing of an application example by the digital camera 10.
[0112] As illustrated in FIG. 11, first, the information processing device 7 instructs the imaging block 20 to acquire a frame of a captured image via the acquisition instructing section 110 (S30). As a result, the imaging block 20 acquires a captured image P3 including a person H.
[0113] The captured image P3 is compressed by a compression process that reduces the resolution, and then input to a face detection AI model M3 (first AI model) of the DSP 32. As a result, the DSP 32 detects a face F of the person H by using the face detection AI model M3, and outputs the coordinates of the area (bounding box) of the detected face F to the information processing device 7 (S31).
[0114] Next, the information processing device 7 determines whether or not the face F of the person H is stationary, by using the stillness determining section 120 (S32). In a case where the face F of the person H is not stationary (S32 / NO), the information processing device 7 returns to the operation of step S30 and instructs the imaging block 20 to acquire a frame of a captured image again.
[0115] Meanwhile, in a case where the face F of the person H is stationary (S32 / YES), the information processing device 7 transmits the coordinates of the area of the face F of the person H to the outside (S33), and then instructs the DSP 32 to switch the AI model via the model switching section 130 (S34). To be specific, the model switching section 130 instructs the DSP 32 to switch from the face detection AI model M3 (first AI model) that detects the face F of the person H from the captured image to a face recognition AI model M4 (second AI model) that recognizes the facial expression or the individual of the face F of the person H. This causes the AI model of the DSP 32 to be switched from the first AI model to the second AI model (S35).
[0116] Next, the information processing device 7 gives an instruction to the imaging block 20 via the cropping instructing section 140 as to the area to be cut out from the captured image (S36). Specifically, the cropping instructing section 140 instructs the imaging block 20 to cut out the area of the face F of the person H. The instruction from the cropping instructing section 140 is stored in the register group 27 of the imaging block 20, and the cropping setting is thereby performed (S37).
[0117] Next, the information processing device 7 instructs the imaging block 20 to perform bracket imaging for acquiring frames of captured images under different imaging conditions, via the acquisition instructing section 110 (S38). For example, the acquisition instructing section 110 may instruct the imaging block 20 to perform bracket imaging in which frames of captured images are acquired under different exposure conditions. As a result, the imaging block 20 acquires captured images P4 obtained by reading out only the image signal of the area of the face F of the person H from the images captured under different imaging conditions.
[0118] Thereafter, each of the captured images P4 that have been acquired is input to the face recognition AI model M4 (second AI model) of the DSP 32. This enables the DSP 32 to recognize the face F of the person H by using the face recognition AI model M4. The recognition result of the face F of the person H recognized from each of the captured images P4 is output to the information processing device 7.
[0119] Next, the information processing device 7 determines, in the result determining section 150, the recognition result of the face F of the person H to be finally output from the imaging device 2, on the basis of the recognition results of the face F of the person H recognized from each of the captured images P4 captured under different imaging conditions (S39). Next, the information processing device 7 transmits the recognition result determined by the process of step S39 to the outside (S40).
[0120] After transmitting the recognition result to the outside, the information processing device 7 instructs the DSP 32 to switch the AI model via the model switching section 130 (S41). To be specific, the model switching section 130 instructs the DSP 32 to switch the AI model from the face recognition AI model M4 (second AI model) that recognizes the facial expression or the individual of the face F of the person H to the face detection AI model M3 (first AI model) that detects the face F of the person H from a captured image. This causes the AI model of the DSP 32 to be switched from the second AI model to the first AI model (S42).
[0121] By repeating the above process as one cycle, the digital camera 10 can recognize the expression or the individual of the face F of the person H within the imaging range, for example.
[0122] <6. Supplementary note> The preferred embodiment of the present disclosure has been described above in detail with reference to the attached drawings, but the technical scope of the present disclosure is not limited to these examples. It is clear that a person with ordinary knowledge in the technical field of the present disclosure can think of various alteration or modification examples within the scope of the technical ideas described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.
[0123] For example, in the above embodiment, the imaging device 2 is provided with an imaging function and an AI processing function, but the present technology is not limited to this example.
[0124] As an example, the information processing device 7 may be equipped with an AI processing function.
[0125] As another example, in a case where functions other than the imaging unit 21 of the imaging device 2, or functions of the signal processing block 30 of the imaging device 2 are performed by an ISP (Image Signal Processor) separate from the imaging device 2, the ISP may have an AI processing function. In such a case, the ISP may further include the functions of the information processing device 7.
[0126] In addition, in the above embodiment, an example of performing AI processing on a captured image is illustrated, but the present technology is also capable of performing AI processing on a sensing image (e.g., a depth image or an infrared image) acquired by an external sensor.
[0127] Further, the effects described in the present specification are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may provide other effects that are apparent to a person skilled in the art from the description in the present specification, in addition to or in place of the above effects.
[0128] Note that the following configurations are also belong to the technical scope of the present disclosure. (1) An information processing device comprising: circuitry configured to identify from image data based on a first captured image a subportion area that contains a recognition target object, the recognition target object containing features, and determine the features of the recognition target object based on image data of a plurality of captured images that respectively include the recognition target object, and each captured under a different condition, the plurality of captured images at least include images other than the first captured image. (2) The information processing device of (1), wherein the recognition target object is a license plate mounted on a vehicle and the features are alphanumeric characters on the license plate. (3) The information processing device of (2), wherein the circuitry is configured to implement a first AI engine that is trained to detect the subportion area based on image data of the first captured image. (4) The information processing device of (3), wherein the first AI engine is configured to detect the subportion area from a compressed image that has a smaller amount of image data than the first captured image. (5) The information processing device of (3), wherein the circuitry implements a second AI engine that is trained to detect alphanumeric characters on the license plate from the plurality of captured images. (6) The information processing device of (5), wherein the circuitry is configured to switch from execution of the first AI engine to execution of the second AI engine after the first AI engine has detected the subportion area. (7) The information processing device of (6), wherein the circuitry is configured to switch from the first AI engine to the second AI engine by selection of a particular AI model as the second AI engine from a plurality of AI models. (8) The information processing device of (6), wherein the circuitry is configured to switch from the first AI engine to the second AI engine via computer code uploaded from memory. (9) The information processing device of (1), wherein the different condition includes at least one of different image exposure condition, different gain, different contrast, different color temperature, or white balance condition. (10) The information processing device of (1), wherein the different condition being a different angle of image capture. (11) The information processing device of (1), further comprising: a first substrate having an image sensor that is configured to capture the first captured image and the plurality of captured images; and a second substrate comprising the circuitry, the first substrate being stacked over the second substrate. (12) The information processing device of (6), wherein the circuitry is further configured to determine whether the recognition target is stationary, and select the second AI engine based on a stationary determination result. (13) The information processing device of (5), wherein the circuitry is further configured to detect the alphanumeric content of the recognition target object from the plurality of captured images after identifying at least one degree of reliability of at least one recognition result based on the plurality of captured images taken under different imaging conditions. (14) The information processing device of (5), wherein the circuitry is further configured to detect the alphanumeric content of the recognition target object based on majority rule for each alphanumeric character on the license plate. (15) The information processing system of (5), wherein the circuitry is further configured to acquire frames of captured images according to bracket imaging under different imaging conditions. (16) The information processing device of (15), wherein the circuitry is configured to determine a degree of reliability of a recognition result provided by the second AI engine, and under a condition the recognition result is above a threshold, the circuitry is configured to transmit the recognition result without backet imaging being performed, and under another condition, the degree of reliability of the recognition result is not above the threshold, the circuitry is configured to acquire another frame of a captured image for further processing using backet imaging. (17) The information processing device of (12), wherein, under a condition the circuitry determines the recognition target is not stationary, the circuitry is configured to cause multiple cycles of respective captured images to be applied to the first AI model and respective results from the first AI model to be applied to the second AI model to enhance a likelihood of accurately detecting the alphanumeric content of the recognition target object while the recognition target object is moving. (18) The information processing device of (1), wherein the recognition target object is a face of a human and the features are features of the face. (19) An information processing system comprising: a digital camera housing; a first substrate having an image sensor that is configured to capture the first captured image and the plurality of captured images; a second substrate comprising circuitry, the first substrate being stacked over the second substrate, the circuitry configured to identify from image data based on a first captured image a subportion area that contains a recognition target object, the recognition target object containing features, and determine the features of the recognition target object based on image data of a plurality of captured images that include the recognition target object, and each captured under a different condition, the plurality of captured images at least include images other than the first captured image; an optical system that guides light from outside toward the image sensor; and a memory sized to hold image data of at least the first captured image, wherein the optical system, the memory, the first substrate and the second substrate and contained within the digital camera housing. (20) A non-transitory computer readable medium having computer readable code stored therein that upon execution by one or more processors causes the one or more processors to be configured to implement a method the method comprising: identify from image data based on a first captured image a subportion area that contains a recognition target object, the recognition target object containing features, and determine the features of the recognition target object based on image data of a plurality of captured images that respectively include the recognition target object, and each captured under a different condition, the plurality of captured images at least include images other than the first captured image.
[0129] 2: Imaging device 7: Information processing device 10: Digital camera 20: Imaging block 21: Imaging unit 22: Imaging processing unit 23: Output control unit 25: Imaging control unit 27: Register group 30: Signal processing block 31: CPU 32: DSP 33: Memory 35: Image compression unit 110: Acquisition instructing section 120: Stillness determining section 130: Model switching section 140: Cropping instructing section 150: Result determining section
Claims
1. An information processing device comprising: circuitry configured to identify from image data based on a first captured image a subportion area that contains a recognition target object, the recognition target object containing features, and determine the features of the recognition target object based on image data of a plurality of captured images that respectively include the recognition target object, and each captured under a different condition, the plurality of captured images at least include images other than the first captured image.
2. The information processing device of claim 1, wherein the recognition target object is a license plate mounted on a vehicle and the features are alphanumeric characters on the license plate.
3. The information processing device of claim 2, wherein the circuitry is configured to implement a first AI engine that is trained to detect the subportion area based on image data of the first captured image.
4. The information processing device of claim 3, wherein the first AI engine is configured to detect the subportion area from a compressed image that has a smaller amount of image data than the first captured image.
5. The information processing device of claim 3, wherein the circuitry implements a second AI engine that is trained to detect alphanumeric characters on the license plate from the plurality of captured images.
6. The information processing device of claim 5, wherein the circuitry is configured to switch from execution of the first AI engine to execution of the second AI engine after the first AI engine has detected the subportion area.
7. The information processing device of claim 6, wherein the circuitry is configured to switch from the first AI engine to the second AI engine by selection of a particular AI model as the second AI engine from a plurality of AI models.
8. The information processing device of claim 6, wherein the circuitry is configured to switch from the first AI engine to the second AI engine via computer code uploaded from memory.
9. The information processing device of claim 1, wherein the different condition includes at least one of different image exposure condition, different gain, different contrast, different color temperature, or white balance condition.
10. The information processing device of claim 1, wherein the different condition being a different angle of image capture.
11. The information processing device of claim 1, further comprising: a first substrate having an image sensor that is configured to capture the first captured image and the plurality of captured images; and a second substrate comprising the circuitry, the first substrate being stacked over the second substrate.
12. The information processing device of claim 6, wherein the circuitry is further configured to determine whether the recognition target is stationary, and select the second AI engine based on a stationary determination result.
13. The information processing device of claim 5, wherein the circuitry is further configured to detect the alphanumeric content of the recognition target object from the plurality of captured images after identifying at least one degree of reliability of at least one recognition result based on the plurality of captured images taken under different imaging conditions.
14. The information processing device of claim 5, wherein the circuitry is further configured to detect the alphanumeric content of the recognition target object based on majority rule for each alphanumeric character on the license plate.
15. The information processing system of claim 5, wherein the circuitry is further configured to acquire frames of captured images according to bracket imaging under different imaging conditions.
16. The information processing device of claim 15, wherein the circuitry is configured to determine a degree of reliability of a recognition result provided by the second AI engine, and under a condition the recognition result is above a threshold, the circuitry is configured to transmit the recognition result without backet imaging being performed, and under another condition, the degree of reliability of the recognition result is not above the threshold, the circuitry is configured to acquire another frame of a captured image for further processing using backet imaging.
17. The information processing device of claim 12, wherein, under a condition the circuitry determines the recognition target is not stationary, the circuitry is configured to cause multiple cycles of respective captured images to be applied to the first AI model and respective results from the first AI model to be applied to the second AI model to enhance a likelihood of accurately detecting the alphanumeric content of the recognition target object while the recognition target object is moving.
18. The information processing device of claim 1, wherein the recognition target object is a face of a human and the features are features of the face.
19. An information processing system comprising: a digital camera housing; a first substrate having an image sensor that is configured to capture the first captured image and the plurality of captured images; a second substrate comprising circuitry, the first substrate being stacked over the second substrate, the circuitry configured to identify from image data based on a first captured image a subportion area that contains a recognition target object, the recognition target object containing features, and determine the features of the recognition target object based on image data of a plurality of captured images that include the recognition target object, and each captured under a different condition, the plurality of captured images at least include images other than the first captured image; an optical system that guides light from outside toward the image sensor; and a memory sized to hold image data of at least the first captured image, wherein the optical system, the memory, the first substrate and the second substrate and contained within the digital camera housing.
20. A non-transitory computer readable medium having computer readable code stored therein that upon execution by one or more processors causes the one or more processors to be configured to implement a method the method comprising: identify from image data based on a first captured image a subportion area that contains a recognition target object, the recognition target object containing features, and determine the features of the recognition target object based on image data of a plurality of captured images that respectively include the recognition target object, and each captured under a different condition, the plurality of captured images at least include images other than the first captured image.
Citation Information
Patent Citations
Artificial intelligence container identification system based on channel coding
CN113610092A
Image processing apparatus, image processing method, and program
JP2009253747A
Parking lot management system
JP2015187897A
Imaging device and imaging method
JP2022132752A
Information processing system and program
JP2022181678A