Image-capturing device, control method, and program

By storing AI models and processing parameters outside the sensor unit, the device simplifies AI model switching, reducing operational complexity and maintaining performance in image-capturing devices.

WO2025182869A1PCT designated stage Publication Date: 2025-09-04SONY SEMICON SOLUTIONS CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/006257
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-26
Filing Date
2025-02-25
Publication Date
2025-09-04

Smart Images

  • Figure JP2025006257_04092025_PF_FP_ABST
    Figure JP2025006257_04092025_PF_FP_ABST
Patent Text Reader

Abstract

An image capturing device includes a sensor that captures and analyzes an image according to an artificial intelligence (AI) model stored therein. Control circuitry stores a plurality of AI models including at least a first AI model to perform object detection and a second AI model to perform object detection. The first AI model is different from the second AI model. The control circuitry switches the AI model stored in the sensor between at least the first AI model and the second AI model.
Need to check novelty before this filing date? Find Prior Art

Description

IMAGE-CAPTURING DEVICE, CONTROL METHOD, AND PROGRAM

[0001] The present technology relates to an image-capturing device configured to be capable of performing AI processing by using an AI model, with an image-captured image being an object; a control method thereof; and a program, and in particular relates to technology relating to switching an AI model being used.

[0002] There is technology for performing AI (Artificial Intelligence) processing such as object detection processing, image recognition processing, and so forth, by using an AI model subjected to machine learning, with an image-captured image being an object. For example, PTL 1 below discloses a sensor-integrated type AI processing device, in which a signal processing unit that performs AI processing is implemented in an image sensor in which a pixel array portion is formed.

[0003] WO 2023 / 090119Summary

[0004] Now, it is conceivable to switch AI models that the AI processing unit uses in order to switch inferential tasks, such as switching tasks from object detection processing to image recognition processing or the like, for example.

[0005] It is an object of the present technology to improve convenience of uses relating to switching, in cases of switching the AI models being used, as described above.

[0006] An image-capturing device according to the present technology includes: a sensor unit having an image-capturing unit that has a pixel array unit, in which a plurality of pixels having photoreceptors are arrayed two-dimensionally, and configured to perform photoelectric conversion of light from a subject and obtain an image-captured image, and having an AI processing unit configured to perform AI processing, which is processing using an AI model, with the image-captured image being an object; a memory unit that is provided outside of the sensor unit and configured to store a plurality of types of AI models; and a control unit that is provided outside of the sensor unit and configured to perform control such that the AI model used in the AI processing unit is switched among the plurality of types of AI models stored in the memory unit. According to the image-capturing device described above, a configuration is made in which the plurality of types of AI models serving as switching candidates are stored in the memory unit outside of the sensor unit, and the control unit outside of the sensor unit performs control of switching the AI models among these candidates, and accordingly when replacing the AI models that are candidates is desired, ease of work thereof can be improved.

[0007] Fig. 1 is a block diagram illustrating a schematic configuration example of an image-capturing device according to a first embodiment.Fig. 2 is an explanatory diagram of an AI processing technique according to the first embodiment.Fig. 3 is an explanatory diagram according to a modification of the first embodiment.Fig. 4 is a block diagram illustrating a schematic configuration example of an image-capturing device according to a second embodiment.Fig. 5 is an explanatory diagram regarding an example of an image-capturing scene assumed in the second embodiment.Fig. 6 is an explanatory diagram regarding picture size conversion processing to an InputTensor size.Fig. 7 is an explanatory diagram of an AI processing technique according to the second embodiment.Fig. 8 is a flowchart showing a specific processing procedure example for realizing the AI processing technique according to the second embodiment.Fig. 9 is an explanatory diagram regarding a separate example of first type objects and second type objects.Fig. 10 is a flowchart showing a specific processing procedure example for realizing an AI processing technique according to a modification of the second embodiment.

[0008] Embodiments according to the present technology will be described below in the following order, with reference to the attached drawings. <1. First Embodiment> (1-1. Configuration Example of Image-capturing Device) (1-2. AI Processing Technique According to First Embodiment) (1-3. Modification of First Embodiment) <2. Second Embodiment> (2-1. Configuration Example of Image-capturing Device) (2-2. AI Processing Technique According to Second Embodiment) (2-3. Processing Procedures) (2-4. Modification of Second Embodiment) <3. Modifications> <4. Program> <5. Summarization of Embodiments> <6. Present Technology>

[0009] <1. First Embodiment> (1-1. Configuration Example of Image-capturing Device) Fig. 1 is a block diagram illustrating a schematic configuration example of an image-capturing device 1 according to a first embodiment of the present technology. The image-capturing device 1 includes an image-capturing optical system 2, a sensor unit 3, a camera control unit 4, a communication unit 5, and an optical system driving unit 6, as illustrated.

[0010] The image-capturing optical system 2 incudes lenses such as a cover lens, a focusing lens, and so forth, and a diaphragm (iris) mechanism. Light (incident light) from a subject is guided by this image-capturing optical system 2, and is collected on a light receiving surface of the sensor unit 3.

[0011] The sensor unit 3 receives light from the subject and obtains an image-captured image. Now, in the present specification, “image-capturing” means to obtain image data by sensing of a subject. The term “image data” here is a collective expression indicating data that is made up of a plurality of pieces of pixel data. Pixel data is a concept that broadly includes data or the like indicating, for example, distance to the subject, polarization information, and information of temperature, rather than just data indicating quantity of light received from the subject in gradient values of a predetermined number. That is to say, “image data” includes gradient image data indicating gradient values of quantity of light received for each pixel, distance image data indicating information of distance to the subject for each pixel, or polarization image data indicating polarization information for each pixel, thermal image data indicating information of temperature for each pixel, or the like. Also, “image data” can include data of event images obtained by a so-called EVS (Event-based Vision Sensor). The term event as used here means change in quantity of light received that is no smaller than a predetermined quantity, and “event image” is an image indicating information of whether or not an event has occurred for each pixel.

[0012] In the following, the sensor unit 3 will be understood to be configured as a gradient image sensor that obtains gradient images, in the same way as a general image sensor, as an example. In this case, the sensor unit 3 is configured as, for example, a CCD (Charge Coupled Device) type image sensor, a CMOS (Complementary Metal Oxide Semiconductor) type image sensor, or the like.

[0013] The camera control unit 4 is includes circuitry such as a microcomputer that has a CPU (Central Processing Unit) 4a, and a memory unit 4b such as ROM (Read Only Memory), RAM (Random Access Memory), flash memory (non-volatile memory), or the like. The CPU 4a executes various types of processing in accordance with programs stored in the memory unit 4b (e.g., programs stored in the above ROM, programs loaded to the RAM), thereby performing overall control of the image-capturing device 1. The camera control unit 4 is connected to the sensor unit 3 via a bus 7, and is capable of receiving various types of data from the sensor unit 3, and transmitting various types of data to the sensor unit 3.

[0014] Also, the camera control unit 4 is connected to the communication unit 5 via the bus 7. The communication unit 5 iscircuitry that isconfigured to be capable of data communication, wired or wirelessly, with external devices (devices outside of the image-capturing device 1). The communication unit 5 can be configured to have network communication functions, and in this case, the camera control unit 4 is capable of exchanging data with a predetermined device (e.g., a server device) on a predetermined network, such as the Internet or the like, for example, via the communication unit 5.

[0015] Also, the camera control unit 4 performs driving instructions to the optical system driving unit 6, regarding a zoom lens, the focusing lens, the diaphragm mechanism, and so forth. The optical system driving unit 6 executes movement of the focusing lens and the zoom lens, opening / closing of diaphragm blades of the diaphragm mechanism, and so forth, in the image-capturing optical system 2, in accordance with these driving instructions.

[0016] Now, the camera control unit 4 also performs AF (Auto Focus) and AE (Auto Exposure) control as well. The camera control unit 4 realizes AF control by calculating a focus target value for focusing on a subject that is an object, on the basis of an image-captured image obtained by the sensor unit 3, and performing driving control of the focusing lens on the basis of the focus target value. Also, AE control is realized by the camera control unit 4 calculating target exposure values on the basis of an image-captured image obtained by the sensor unit 3, and performing opening / closing control of the diaphragm blades on the basis of the exposure value, and setting specifications regarding shutter speed and ISO sensitivity with respect to the sensor unit 3. Now, with respect to AE control, employing an exposure mode by multi-pattern metering to decide optimal exposure for the entire image, an exposure mode by central metering to decide optimal exposure for the middle of the image, an exposure mode by spot metering to decide optimal exposure for a portion of the image that is in focus, and so forth, as the exposure mode, is conceivable.

[0017] The sensor unit 3 has an image-capturing unit 31, and circuitry such as an image signal processing unit 32, an AI (Artificial Intelligence) processing unit 33, an in-sensor control unit 34, and a communication I / F (interface) 35, as illustrated.

[0018] The image-capturing unit 31 has a pixel array unit in which a plurality of pixels having photoreceptors (photoelectric converting elements) such as photodiodes or the like, for example, are arrayed two-dimensionally in a horizontal direction and a vertical direction, and obtains image-captured images by performing photoelectric conversion of light from the subject. Although omitted from illustration, the image-capturing unit 31 also includes configurations for obtaining image data as digital data, such as read circuits for reading out values (light-reception values) of the pixels, and ADC (Analog to Digital Converter) for performing digital sampling of pixel values that are analog signals, and so forth.

[0019] The image signal processing unit 32 inputs image-capturing data obtained by the image-capturing unit 31, specifically gradient image data in the present example, and subjects this to various types of image signal processing. For image signal processing here, subjecting image-capturing data input from the image-capturing unit 31 in RAW format to demosaicing processing for obtaining R (red) images, G (green) images, and B (blue) images, gain adjustment processing such as AWB (Auto White Balance) processing or the like, various types of correction processing such as shading correction processing or gamma correction processing, lens distortion correction processing, and so forth, various types of filtering processing and so forth, such as denoising processing, edge enhancement processing, and so forth, is conceivable. Also, in the case of the present example in particular, the image signal processing unit 32 is enabled to perform trimming processing on the input image, and image signal processing that is picture size conversion processing into an InputTensor size for the AI models that the AI processing unit 33 uses.

[0020] The AI processing unit 33 performs AI processing that is processing using AI models on image-captured images, input from the image signal processing unit 32, as an object. In the present example, the AI processing unit 33 is made up of a DSP (Digital Signal Processor), for example, and can switch between AI models being used. Specifically, switching of AI models is realized by switching processing parameters of the DSP.

[0021] In the first embodiment, the AI processing that the AI processing unit 33 performs broadly includes various types of inference processing that is performed on images as an object, such as, for example, object detection processing, image recognition processing, and so forth, and is not limited to particular inference processing. AI models used in the AI processing conceivably are models having a neural network such as a CNN (Convolutional Neural Network) or the like, for example.

[0022] The in-sensor control unit 34 is configured including a microcomputer that is configured having, for example, a CPU, ROM, RAM, and so forth, and centrally controls operations of the sensor unit 3. For example, the in-sensor control unit 34 performs operation control of the image-capturing unit 31. Specifically, control to start / stop image-capturing operations, and so forth, is performed. Also, the in-sensor control unit 34 performs operation control of the image signal processing unit 32 and the AI processing unit 33 as well. As for operation control of the image signal processing unit 32, the in-sensor control unit 34 is enabled to give execution instructions regarding various types of signal processing such as gain adjustment and so forth, described above, perform control processing of processing parameters, and so forth. Also, also for operation control of the AI processing unit 33, the in-sensor control unit 34 gives execution instructions for AI processing and performs control relating to switching of AI models. Now, a switching technique for AI models according to the embodiment will be described again later.

[0023] The communication I / F 35 is a communication interface for enabling data output from inside of the sensor unit 3 to outside of the sensor unit 3, and data input from outside of the sensor unit 3 to inside of the sensor unit 3, and performs data communication with the camera control unit 4 via the bus 7, following a predetermined communication data format.

[0024] The in-sensor control unit 34 can output image-captured images by the image-capturing unit 31 (including image-captured images following being subjected to at least part of processing of the image signal processing unit 32), and information of AI processing results (inference results) by the AI processing unit 33, to the camera control unit 4 via the communication I / F 35. Also, the in-sensor control unit 34 is enabled to receive various types of data transmitted from the camera control unit 4 via the communication I / F 35.

[0025] (1-2. AI Processing Technique According to First Embodiment) Now, according to the configuration of the image-capturing device 1 described above, AI models to be used by the AI processing unit 33 can be deployed from a server device that is outside of the image-capturing device 1 by communication via the communication unit 5. Due to such deploying being enabled, it is conceivable that in the image-capturing device 1, the above server device performs switching of AI models that the AI processing unit 33 uses, as a controlling entity. Specifically, in accordance with switching becoming necessary regarding an inference task to be executed by the AI processing unit 33, the server device realizes switching of the AI modules by deploying a separate AI model from the AI model that was currently in use to the image-capturing device 1.

[0026] However, in a case of employing such a technique in which the server device deploys a separate AI model, communication with the server device occurs each time the AI model is switched, leading to problems such as burdening the user in terms of communication fees, requiring time for switching, and so forth.

[0027] Conversely, it is conceivable to employ a technique for switching the AI models that the AI processing unit 33 uses by storing a plurality of types of AI models that are candidates in memory within the sensor unit 3, and switching the AI models to use among these. According to this technique, the need for communication with an external device when switching AI models can be done away with, and accordingly the above problem can be resolved. However, when replacement of at least one of the AI models that are candidates for switching becomes necessary, this replacement may become difficult. Specifically, in a case in which the image-capturing device 1 is operated completely offline (i.e., operations in which communication with the external device is not performed), realizing replacement of AI models such as described above necessitates the sensor unit 3 itself to be replaced with a separate individual unit, i.e., to exchange with a separate sensor unit 3 with a different combination of AI models stored therein. However, such exchanging of the sensor unit 3 actually involves difficult work, such as having to be performed within changing the positional relation with the image-capturing optical system 2, for example.

[0028] The present embodiment has been made in light of the above problem, and it is an object thereof to improve user convenience for the switching in a case in which switching of AI models being used is to be performed. Accordingly, in the image-capturing device 1 according to the first embodiment, the following technique is employed with regard to switching the AI model.

[0029] Fig. 2 is an explanatory diagram of the AI processing technique according to the first embodiment. First, in the image-capturing device 1 according to the first embodiment, a plurality of types of AI models that are candidates for switching are stored in the memory unit 4b in the camera control unit 4 that is provided outside of the sensor unit 3. In the illustration, an example is illustrated in which a first AI model MD1 and a second AI model MD2 are stored as the plurality of types of AI models. Note that the number of AI models stored in the memory unit 4b is not limited to two, and it is sufficient for a plurality of AI models serving as candidates to be stored.

[0030] These first AI model MD1 and second AI model MD2 serving as candidates are trained AI models that have performed machine learning such that each is capable of realizing predetermined different inference tasks, respectively.

[0031] In the image-capturing device 1 according to the first embodiment, the CPU 4a in the camera control unit 4 then acts as a main entity to perform switching control of AI models that the AI processing unit 33 uses. Specifically, the CPU 4a performs control such that the AI model being used at the AI processing unit 33 is switched between the plurality of types of AI models (first AI model MD1 and second AI model MD2 in the present example) stored in the memory unit 4b.

[0032] At this time, the CPU 4a transmits the AI model to serve as the switching destination to the in-sensor control unit 34, by communication via the communication I / F 35 and also gives instructions to set the AI model that is transmitted, to the AI processing unit 33. Thus, switching of the AI model that the AI processing unit 33 uses is realized.

[0033] As described above, in the present embodiment, a configuration is made in which a plurality of types of AI models serving as candidates are stored in the memory unit 4b outside of the sensor unit 3, and the CPU 4a (control unit) outside of the sensor unit 3 performs switching control of the AI models. Accordingly, replacement of candidate AI models in a case of operating the image-capturing device 1 offline can be performed by exchanging electronic components including at least the memory unit 4b, and specifically the electronic component serving as the camera control unit 4 (microcomputer) in the present embodiment. This can do away with the need to perform positioning with the image-capturing optical system, as in a case of exchanging the sensor unit 3. Thus, ease of work in a case of desiring to replace AI models that are candidates can be improved, and convenience of the user relating to switching of AI modules can be improved.

[0034] (1-3. Modification of First Embodiment) Now, due to characteristics of images input at the time of inference being within the range of characteristics assumed when learning, the AI models can exhibit sufficient inference performance, but when characteristics of images that are input at the time of inference happen to deviate from the range of characteristics assumed when learning, this may lead to deterioration of inference performance. Due to this point, when switching AI modules, a situation will arise in which image signal processing that the image signal processing unit 32 performs regarding input images to the AI processing unit 33 was appropriate for the AI model before switching, but is no longer appropriate for the AI model following switching.

[0035] Giving consideration to this point, in the modification of the first embodiment, along with control for switching AI models, control is performed for switching image processing parameters of the image signal processing unit 32.

[0036] Fig. 3 is an explanatory diagram according to the modification of the first embodiment. In this case, the memory unit 4b of the camera control unit 4 stores, along with the plurality of types of AI models serving as candidates (the two of first AI model MD1 and second AI model MD2 are exemplified here as well), image processing instruction information for each of these AI models. In the illustration, an example is illustrated in which, in accordance with the AI models, which are to be stored as candidates, being the two of the first AI model MD1 and the second AI model, the two of first image processing instruction information Ip1 that is image processing instruction information regarding the first AI model MD1, and second image processing instruction information Ip2 that is image processing instruction information regarding the second AI model MD2, are stored in the memory unit 4b as image processing instruction information.

[0037] Such image processing instruction information is created as information instructing image processing parameters of the image signal processing unit 32, so as to realize appropriate input image characteristics for the relevant AI model. Image processing parameters that are instructed by the image processing instruction information do not need to be parameters regarding all image processing items that the image signal processing unit 32 is capable of executed, and it is sufficient to be parameters that are at least a part of the items.

[0038] In Fig. 3, in accordance with a switching condition of the AI model that the AI processing unit 33 uses being satisfied, the CPU 4a of the camera control unit 4 switches the AI model that the AI processing unit 33 uses, to an AI model that is the object (AI model of switching destination), and also performs control to switch image processing parameters that the image signal processing unit 32 uses to image processing parameters indicated by the image processing instruction information correlated with the above AI model that is the object.

[0039] Accordingly, AI processing using the AI model after switching is performed with the image-captured image, subjected to image processing by parameters corresponding to this AI model, as an object. Thus, when switching AI models, the AI processing can be performed with the image-captured image, subjected to image processing in an appropriate form corresponding to the AI model following switching, as an object, and precision of AI processing can be improved.

[0040] Now, in the modification of the first embodiment, in a case in which the image-capturing device 1 is used online and not offline, it is conceivable to use information transmitted by an external device such as a server device or the like as each piece of the image processing instruction information. For example, a form can be conceived in which an external device such as a server device or the like transmits a plurality of types of AI models serving as candidates to the image-capturing device 1, and also transmits, along with these AI models, image processing instruction information corresponding thereto, and stores the image processing instruction information that is transmitted in the memory unit 4b.

[0041] <2. Second Embodiment> (2-1. Configuration Example of Image-capturing Device) Next, a second embodiment will be described. In the second embodiment, the AI model used by the AI processing unit 33 is switched between a first AI model that performs object detection processing with a first type object as a target, and a second AI model that performs object detection processing with a second type object that is present in a partial region of the first type object as a target.

[0042] Fig. 4 is a block diagram illustrating a schematic configuration example of an image-capturing device 1A according to the second embodiment. Note that in the following description, portions that are the same as the portions already described will be denoted by the same signs, and description will be omitted.

[0043] The image-capturing device 1A differs from the image-capturing device 1 illustrated in Fig. 1 with respect to the point that a camera control unit 4A is provided instead of the camera control unit 4. The camera control unit 4A differs from the camera control unit 4 with respect to the point of having a CPU 4aA instead of the CPU 4a. The CPU 4aA is the same as the CPU 4a with respect to the point that control is performed in which the AI model that the AI processing unit 33 uses is switched among AI models stored in the memory unit 4b, but differs with respect to the point that control and so forth relating to trimming of image-captured images such as described below is performed.

[0044] Now, the first AI model MD1 and the second AI model MD2 serving as switching candidate AI models are stored in the memory unit 4b in the image-capturing device 1A according to the second embodiment as well, but in the second embodiment, the first AI model MD1 is an AI model that performs object detection processing with a first type object as a target, and the second AI model MD2 is an AI model that performs object detection processing with a second type object that is present in a partial region of the first type object as a target.

[0045] Now, a usage of the image-capturing device 1A is assumed to be a monitoring usage of products stocked in a warehouse or the like. For example, as exemplified in Fig. 5, monitoring usage is assumed for managing, with respect to products stacked on shelves in a warehouse, types, quantities in stock, and so forth, of these products. As illustrated, product identification codes, such as one-dimensional barcodes and two-dimensional barcodes or the like, for example, are attached to predetermined positions of the individual cartons. These product identification codes enable identification of the stocked products.

[0046] In the case of the present example, the first type object = the cartons of the products, and the second type object = the product identification codes attached to the cartons. That is to say, performing object detection processing of which the cartons are the target using the first AI model MD1, and subsequently performing object detection processing of which the identification codes are the target using the second AI model MD2 with the image regions of the cartons that are detected as an object, is a premise.

[0047] In this case, reading processing of identification codes is performed as post-processing of the AI processing for detecting the regions of the identification codes such as described above. That is to say, this is processing for detecting information content (e.g., number information or the like) indicated by the identification code, from the pattern of the identification code in the image.

[0048] As an example, in the present example, such reading processing as post-processing will be understood to be performed by an external device (e.g., server device) of the image-capturing device 1A.

[0049] Fig. 6 is an explanatory diagram regarding picture size conversion processing to an InputTensor size for the AI processing unit 33. As described above, in the present example, the image signal processing unit 32 has a function of picture size conversion processing to the InputTensor size for the AI processing unit 33. Here, the image-captured image with a picture size prior to conversion to the InputTensor size will be written as “original image”. The picture size of the original image is, for example, horizontal pixel count × vertical pixel count = 4000 × 3000, as illustrated. In the present example, the picture size of the original image is the picture size of the RAW image output from the image-capturing unit 31. Also, the InputTensor size is horizontal pixel count × vertical pixel count = 640 × 480, as illustrated, for example.

[0050] Hereinafter, picture size conversion from the picture size of the original image to the InputTensor size may be referred to as “downsampling”.

[0051] (2-2. AI Processing Technique According to Second Embodiment) In light of the above premise, an example of an AI processing technique according to a second embodiment will be described with reference to Fig. 7. A first point is, in a case of performing object detection processing of identification codes using the second AI model following object detection processing of cartons using the first AI model MD1, the point that trimmed images from the original image are used as trimmed images for cartons, rather than trimmed images from images following downsampling (see Fig. 7A, Fig. 7B). Also, a second point is the point that for trimmed images of identification codes to be used for post-processing, trimmed images from the original image are used, rather than trimmed images from images following downsampling (see Fig. 7C).

[0052] As in the first point described above, using images obtained by a first type object detection region BB being trimmed from the original image enables high resolution to be realized for input images to object detection processing for identification codes (second type objects), and precision of the object detection processing for identification codes can be improved.

[0053] Also, according to the second point described above, in a case in which post-processing of which images of the second type object are the object, specifically, post-processing as processing for reading identification codes in the present example, is performed, resolution of images input to the post-processing can be raised. Accordingly, precision of the post-processing performed with images of the second type object as the object can be improved.

[0054] Fig. 7A exemplifies a case in which a plurality of cartons have been detected by object detection processing of cartons, using the first AI model MD1 in accordance with a case in which the plurality of cartons are in the image-captured image. As illustrated, a sign “BB” is given to detection regions of cartons, i.e., regarding first type object detection regions, and numerals are appended to the end of the sign for identification of the first type object detection regions BB. In the illustration, a case in which six first type object detection regions BB, from BB1 to BB6 are present is exemplified, in accordance with a case in which the count of the cartons in the image-captured image is six.

[0055] The object detection processing for identification codes using the second AI model MD2 is performed for each of these first type object detection regions BB. In Fig. 7B, the way in which object detection processing for identification codes is performed regarding the first type object detection region BB1 is exemplified. As illustrated, the first type object detection region BB1 is trimmed out from the original image, downsampling processing to the InputTensor size is performed at the image signal processing unit 32 regarding this trimmed image, and thereafter object detection processing of identification codes is performed using the second AI model MD2. Hereinafter, a detection region of a second type object obtained in object detection processing using the second AI model MD2 in this way will be written as second type object detection region BC.

[0056] In accordance with object detection processing of identification codes being performed as described above, images obtained by trimming second type object detection regions BC out from the original image are obtained, and these trimmed images are used for reading processing that is post-processing, as illustrated in Fig. 7C. That is to say, in the present example, the trimmed images are transmitted to an external device where reading processing is performed.

[0057] (2-3. Processing Procedures) Fig. 8 is a flowchart showing a specific processing procedure example for realizing the AI processing technique according to the second embodiment described above. The processing illustrated in Fig. 8 is executed by the CPU 4aA on the basis of a program stored in predetermined memory, such as ROM or the like of the camera control unit 4A, for example.

[0058] In Fig. 8, in step S101 the CPU 4aA performs setting instruction for the first AI model MD1. That is to say, the AI processing unit 33 is instructed to set the first AI model MD1 as the AI model to be used.

[0059] In step S102 following step S101, the CPU 4aA performs execution instruction of first AI processing. That is to say, this is to cause execution of object detection processing of the first type object (carton) using the first AI model MD1.

[0060] In step S103 following step S102, the CPU 4aA performs processing for acquiring AI processing results from the first AI processing. At this time, if the count of first type objects in the image-captured image is a plurality, information of a plurality of first type object detection regions BB is acquired.

[0061] In step S104 following step S103, the CPU 4aA sets a detected count = N of first type objects, and further in the subsequent step S105 sets a region identifier to n = 1. The region identifier n is an identifier for identifying the first type object detection region BB that is the object of processing.

[0062] In step S106 following step S105, the CPU 4aA performs setting instruction for the second AI model MD2, i.e. the AI processing unit 33 is instructed to set the second AI model MD2 as the AI model to be used.

[0063] In step S107 following step S106, the CPU 4aA instructs trimming from the original image for the n’th first type object. That is to say, the image signal processing unit 32 is caused to execute trimming from the original image, instructing the first type object detection region BB for the n’th first type object to be the trimming region. At this time, the trimming may be performed with the RAW image input from the image-capturing unit 31 as the object, or may be performed with an RGB image following demosaicing as the object. This point is the same regarding trimming of the second type object detection region BC that will be described later.

[0064] In step S108 following step S107, the CPU 4aA performs execution instruction of second AI processing regarding the trimmed image of the first type object. That is to say, this is to cause the AI processing unit 33 to execute object detection processing of second type objects, with the post-trimming (and also post-downsampling) image obtained by the image signal processing unit 32 as the object.

[0065] In step S109 following step S108, the CPU 4aA acquires AI processing results from the second AI processing. That is to say, information of the second type object detection region BC is acquired.

[0066] In step S110 following step S109, the CPU 4aA instructs trimming of the second type object from the original image. That is to say, the image signal processing unit 32 is caused to execute trimming from the original image, instructing the second type object detection region BC acquired in step S109 instructed to be the trimming region.

[0067] In step S111 following step S110, the CPU 4aA performs output processing of the trimmed image of the second type object. That is to say, processing is performed for acquiring the trimmed image of the second type object obtained by the processing of step S110 from the image signal processing unit 32, and transmitting the trimmed image that is acquired to a predetermined external device via the communication unit 5. As part of the output processing of step S111, or prior to the output processing of step S111, the CPU 4aA performs aspect ratio preservation of the second type object to maintain the aspect ratio of the object the same as that in the originally captured image. Such aspect ratio preservation may involve adjustment of the aspect ratio of the trimmed image of the second type object since the aspect ratio of the second type object may have become distorted during the processing of steps S101 to S110. In this case, the CPU 4aA may determine whether the second type object is distorted due to a change in aspect ratio, and adjust the aspect ratio of the trimmed image of the second type object accordingly. Such aspect ratio preservation is necessary to eliminate, or at least reduce, distortion of the second type object so that the second type object can be correctly read if, for example, the second type object is a bar code or a QR code.

[0068] In step S112 following step S111, the CPU 4aA determines whether or not n >= N holds for the region identifier. This is equivalent to determining whether or not the processing of steps S107 to S111 has been completed for all first type object detection regions BB.

[0069] In a case of determining that n >= N does not hold for the region identifier in step S112, the CPU 4aA increments the value of the region identifier by 1 (n ← n + 1) in step S113, and returns to step S107. Accordingly, the processing from step S107 to S111 is executed with the next first type object detection region BB as the object.

[0070] Conversely, in a case of determining that n >= N holds for the region identifier in step S112, the CPU 4aA ends the series of processing illustrated in Fig. 8.

[0071] Note that while an example of the post-processing with the trimmed image of the second type object detection region BC as the object is performed at an external device from the image-capturing device 1A has been described above, the post-processing conceivably may be performed inside the image-capturing device 1A, such as in the camera control unit 4A, or the like, for example. In this case, the CPU 4aA does not perform the output processing of step S111, and can perform processing for storing the trimmed image in predetermined memory as data to be used in the post-processing, or the like, for example. Either way, as for the processing in this case, it is sufficient for processing to be performed that controls such that the trimmed image is used in the post-processing.

[0072] Also, in the second embodiment, the object detection processing using the first AI model MD1 and the object detection processing using the second AI model MD2 may be performed with the same frame image as the object, and performing with a different frame image as the object (e.g., the later performs object detection processing with, as the object, the subsequent frame image from the frame image that the former object detection processing had as the object, for example) is conceivable. In a case of taking the same frame image as the object, using a frame buffer for temporarily holding the original image is sufficient.

[0073] Also, an example has been described above in which the first type object = product carton, and the second type object = identification code, but the first type object and the second type object are not limited to these, and examples such as illustrated in Fig. 9, for example, are also conceivable. Fig. 9A illustrates an example in which the first type object = full body of a person, and the second type object = face of a person, and Fig. 9B illustrates an example in which the first type object = automobile, and the second type object = license plate. In the example in Fig. 9A, an AI model that performs object detection processing with a full body of a person as a target is used as the first AI model MD1, and an image region of a full body of a person is detected as the first type object detection region BB, and also an AI model that performs object detection processing with a face of a person as a target is used as the second AI model MD2, and an image region of a face of a person is detected as the second type object detection region BC. Also, in the example in Fig. 9B, an AI model that performs object detection processing with an automobile as a target is used as the first AI model MD1, and an image region of an automobile is detected as the first type object detection region BB, and also an AI model that performs object detection processing with a license plate as a target is used as the second AI model MD2, and an image region of a license plate is detected as the second type object detection region BC.

[0074] In the example in Fig. 9A, performing identification processing of the person that is the object of the trimmed image of the face as post-processing is conceivable. Also, in the example in Fig. 9B, performing reading processing and so forth of license plate information, with respect to the trimmed image of the license plate, as post-processing, is conceivable.

[0075] In this way, examples of the first type object and the second type object can be thought of variously.

[0076] (2-4. Modification of Second Embodiment) Now, in the second embodiment, switching image processing parameters at the image signal processing unit 32 in accordance with switching of the AI models is conceivable, in the same way as the case of the first embodiment. In the second embodiment, image regions that are input to the AI processing unit 33 differ between the object detection processing using the first AI model MD1 and the object detection processing using the second AI model MD2, and accordingly switching the image processing parameters conceivably takes into consideration such difference in input image regions.

[0077] For example, portions with different image quality within an image can occur in image-captured images from the image-capturing unit 31. Specifically, due to AE control being performed, portions where brightness is appropriate, and portions where not appropriate, can occur. Also, portions with great noise and portions with little noise may occur. Alternatively, although this depends on the strength of distortion correction processing, portions with great distortion and portions with little distortion can occur. At this time, if there are portions in the first type object in which image quality is not suitable, such as portions where brightness is not appropriate, or the like, object detection processing of the second type object that is performed using the trimmed image of the first type object may not be performed appropriately.

[0078] Accordingly, a technique is proposed herein, in which, on the basis of information of a first type object detection region BB obtained by object detection processing of the first type object using the first AI model MD1, image quality of the first type object detection region BB is evaluated, and control is performed such that image signal processing by the image signal processing unit 32, performed when object detection processing of the second type object using the second AI model MD2, is performed using image processing parameters in accordance with evaluation results of the image quality.

[0079] A specific example of processing procedures will be described with reference to the flowchart in Fig. 10. Note that processing that is the same as that described with reference to Fig. 8 is denoted by the same step numbers, and detailed description will be omitted.

[0080] In this case as well, the point of performing the processing of steps S101 to S106 is the same as in Fig. 8. The CPU 4aA in the present example performs image quality evaluation processing in step S201, of which a detection region of an n’th first type object is the object, in accordance with the setting specification of step S106 having been performed. A case of making brightness to be appropriate for the first type object detection region BB will be exemplified here. Specifically, the brightness is made to be appropriate, in accordance with a case in which the brightness of the first type object detection region BB is not appropriate in conjunction with AE control being performed in which the brightness of the entire image is used as a reference, such as in the above-described multi-pattern metering.

[0081] In this case, the CPU 4aA performs evaluation relating to brightness as image quality evaluation processing of step S201. Specifically, evaluation is performed relating to difference between the brightness of the entire original image and the brightness of the first type object detection region BB. For example, a difference value between a representative value of luminance values of the first type object detection region BB (e.g., average value or median value) and a representative value of luminance values of the entire original image (e.g., average value or median value) is calculated as an image quality evaluation value. Here, an example is given in which a value obtained by subtracting the representative value of the latter from the representative value of the former is calculated as the image quality evaluation value.

[0082] In step S202 following step S201, the CPU 4aA performs setting instruction of image processing parameters in accordance with image evaluation results. Specifically, in the present example, processing is performed in which the image signal processing unit 32 is instructed regarding image processing parameters in accordance with the image quality value that is the difference value calculated in step S201. At this time, parameters for adjustment processing for brightness performed at the image signal processing unit 32 are instructed as the image processing parameters. Specifically, with regard to the parameter in this case, in a case in which the above-described difference value is a negative value, an instruction is given to make the brightness even brighter the greater the absolute value of difference value is, and in a case in which the above-described difference value is a positive value, an instruction is given to make the brightness even darker the greater the absolute value of difference value is.

[0083] In step S203 following step S202, the CPU 4aA performs execution instruction of image processing including trimming regarding the n’th first type object to the image signal processing unit 32. Accordingly, the object detection processing of the second type object using the second AI model MD2 can be made to be performed with the trimmed image of the first type object detection region BB adjusted to an appropriate image quality as the object.

[0084] The CPU 4aA in this case advances the processing to step S108, in accordance with performing the instruction of step S203. Note that processing of step S108 and thereafter, up to step S113, is the same as the case of Fig. 8, and accordingly replicative description will be omitted.

[0085] The CPU 4aA in this case advances the processing to step S201 in accordance with the value of the region identifier n being incremented in step S113. Accordingly, image quality adjustment in accordance with image quality evaluation results for the next first type object detection region BB, object detection processing of the second type object of which the trimmed image of the image-quality-adjusted first type object detection region BB is the object, and output processing of the trimmed image of the second type object detection region BC, are performed.

[0086] Now, in a case of taking noise into consideration, evaluation relating to noise is performed regarding the first type object detection region BB in the image quality evaluation processing of step S201. For example, calculating a noise amount of the first type object detection region BB as the image quality evaluation value is conceivable. In this case, as for the parameter setting instruction in step S202, an amount of the denoising of denoising processing performed at the image signal processing unit 32 adjusted in accordance with the greatness / smallness of the noise amount.

[0087] Also, in a case of taking image distortion into consideration, an amount of the image distortion of the first type object detection region BB is calculated as the image quality evaluation value in the image quality evaluation processing in step S201, and as for the parameter setting instruction in step S202, an amount of correction of distortion correction processing performed at the image signal processing unit 32 can be adjusted in accordance with the greatness / smallness of the distortion amount.

[0088] In the modification of the second embodiment, image quality adjustment of the trimmed image of the first type object detection region BB can conceivably be performed regarding a plurality of image quality items, such as performing regarding all of brightness, noise, distortion, and so forth, for example.

[0089] Note that in the modification of the second embodiment, the trimmed image of the second type object detection region BC used in the post-processing can conceivably be the image subjected to image signal processing by the image processing parameters based on the image quality results in step S201.

[0090] Now, while an example has been described above regarding the second embodiment in which one processing unit serving as the image signal processing unit 32 performs image processing that is trimming, image processing for image adjustment such as brightness or the like, and image processing that is picture size conversion processing to the InputTensor size (downsizing processing), a configuration may be made in which at least part of such image processing is performed by a separate processing unit. For example, employing a configuration in which the camera control unit 4A (CPU 4aA) performs trimming processing, and the image signal processing unit 32 performs remaining image quality processing and picture size conversion processing to the InputTensor size, is conceivable. Alternatively, employing a configuration in which the camera control unit 4A performs trimming processing, and the image signal processing unit 32 performs image processing for image quality adjustment, and a processing unit provided separately from the image signal processing unit 32 performs picture size conversion processing to the InputTensor size, is conceivable.

[0091] <3. Modifications> Note that embodiments are not limited to the specific example described above, and configurations can be employed as various modifications. For example, while an example has been described above in which the memory unit 4b that stores a plurality of types of AI models as candidates, is installed in the same microcomputer as the CPU 4a or the CPU 4aA that performs switching control of the AI models, it is sufficient for the memory unit 4b to be provided outside of the sensor unit 3, and can be provided outside of the microcomputer in which the CPU 4a or the CPU 4aA is installed.

[0092] <4. Program> Now, an embodiment can be thought of as a program to cause a computer device to realize the functions of the CPU 4a and the CPU 4aA described earlier with reference to Fig. 8, Fig. 10, and so forth. That is to say, the program according to the embodiment is a program that is readable by a computer device serving as a control unit in an image-capturing device that includes a sensor unit having an image-capturing unit that has a pixel array unit in which a plurality of pixels having photoreceptors are arrayed two-dimensionally and that is configured to perform photoelectric conversion of light from a subject and obtain an image-captured image, and an AI processing unit configured to perform AI processing that is processing using an AI model with the image-captured image as an object, a memory unit that is provided outside of the sensor unit, and that is configured to store a plurality of types of AI models, and the control unit that is provided outside of the sensor unit, and that is configured to control the sensor unit, the program causing the computer device to realize processing of performing control such that the AI model used in the AI processing unit is switched among the plurality of types of AI models stored in the memory unit. According to such a program, the functions of the CPU 4a and the CPU 4aA described earlier can be realized by the computer device.

[0093] Such a program can be recorded in advance in an HDD (Hard Disc Drive) serving as a recording medium built into equipment such as a computer device or the like, or in ROM or the like in a microcomputer that has a CPU. Alternatively, storage (recording) can be temporarily or permanently performed in removable recording media such as a flexible disk, a CD-ROM (Compact Disc Read Only Memory), an MO (Magneto Optical) disc, a DVD (Digital Versatile Disc), a Blu-ray Disc (registered trademark), a magnetic disk, semiconductor memory, a memory card, or the like. Such removable recording media can be provided as so-called packaged software. Also, such a program can, besides being installed in a personal computer or the like from a removable recording medium, be downloaded from a download site via a network such as a LAN (Local Area Network), the Internet, or the like.

[0094] Also, such a program is suitable for broadly providing the AI processing technique according to an embodiment. Specifically, installing the present program in an image-capturing device equipped with a computer device such as a CPU or the like can cause this image-capturing device to function as an image-capturing device that realizes the AI processing technique of the present disclosure.

[0095] <5. Summarization of Embodiments> As described above, an image-capturing device (1, 1A of same name) according to an embodiment includes: a sensor unit (3 of same name) having an image-capturing unit (31 of same name) that has a pixel array unit, in which a plurality of pixels having photoreceptors are arrayed two-dimensionally, and configured to perform photoelectric conversion of light from a subject and obtain an image-captured image, and having an AI processing unit (33 of same name) configured to perform AI processing, which is processing using an AI model, with the image-captured image being an object; a memory unit (4b of same name) that is provided outside of the sensor unit and configured to store a plurality of types of AI models; and a control unit (CPU 4a, 4aA) that is provided outside of the sensor unit and configured to perform control such that the AI model used in the AI processing unit is switched among the plurality of types of AI models stored in the memory unit. According to the image-capturing device described above, a configuration is made in which a plurality of types of AI models serving as switching candidates are stored in the memory unit outside of the sensor unit, and the control unit outside of the sensor unit performs control of switching AI models among these candidates, and accordingly when replacing AI models that are candidates is desired, ease of work thereof can be improved. Accordingly, user convenience relating to switching of AI models can be improved.

[0096] Also, in the image-capturing device (1 of same name) according to an embodiment, the sensor unit has an image processing unit (image signal processing unit 32) configured to perform input of the image-captured image and image processing, the AI processing unit performs AI processing, with the image-captured image that has been subjected to image processing by the image processing unit being an object, the memory unit stores, in correlation with each of the AI models, instruction information of an image processing parameter for the image processing unit, and the control unit (CPU 4a) performs control to switch the AI model that the AI processing unit uses to an object AI model, and also performs control to switch the image processing parameter that the image processing unit uses to an image processing parameter indicated by the instruction information correlated with the object AI model. Accordingly, AI processing using the AI model after switching is performed with the image-captured image, subjected to image processing by parameters corresponding to this AI model, as an object. Thus, when switching AI models, the AI processing can be performed with the image-captured image, subjected to image processing in an appropriate form corresponding to the AI model following switching, as an object, and precision of AI processing can be improved.

[0097] Further, in the image-capturing device (1A of same name) according to an embodiment, the memory unit stores a first AI model (MD1 of same name) that performs object detection processing, with a first type object being a target, and a second AI model (MD2 of same name) that performs object detection processing, with a second type object that is present in a partial region of the first type object being a target, and the control unit (CPU 4aA) performs control such that the AI model that the AI processing unit uses is switched between the first AI model and the second AI model. Thus, object detection processing of a second type object using the second AI model, with a region of a first type object detected by the first AI model as the object, can be performed. This enables object detection processing of the second type object to be performed with a region to serve as an object narrowed down, and accordingly precision of the object detection processing of the second type object can be improved.

[0098] Furthermore, in the image-capturing device according to an embodiment, when the image-captured image of a picture size, prior to conversion to an InputTensor size for the AI processing unit, is an original image, the control unit performs control such that object detection processing of the second type object by using the second AI model, is performed, with an image that is obtained by trimming a detection region of the first type object from the original image being an object. As described above, using images obtained by the detection region of the first type object being trimmed from the original image enables high resolution to be realized for input images to object detection processing for the second type objects, and precision of the object detection processing for the second type objects can be improved.

[0099] Also, in the image-capturing device according to an embodiment, when the image-captured image of a picture size, prior to conversion to an InputTensor size for the AI processing unit, is an original image, the control unit performs control such that, on the basis of detection region information of the second type object obtained in object detection processing by the second AI model, an image of which a detection region of the second type object is trimmed from the original image, is used in post-processing. Thus, in a case in which, for example, post-processing of which images of the second type object are the object, such as a case in which the second type objects are identification codes of products, and processing for reading identification codes is performed in post-processing, or the like, higher resolution can be realized for input images to the post processing. Accordingly, precision of the post-processing performed with images of the second type object as the object can be improved.

[0100] Further, in the image-capturing device (1A of same name) according to an embodiment, the sensor unit has an image processing unit (image signal processing unit 32) configured to perform input of the image-captured image and image processing, the AI processing unit performs AI processing, with the image-captured image that has been subjected to image processing by the image processing unit being an object, the control unit (CPU 4aA) evaluates image quality of the detection region of the first type object, on the basis of the detection region information of the first type object obtained in the object detection processing of the first type object by using the first AI model, and performs control such that image processing of the image processing unit, when performing object detection processing of the second type object by using the second AI model, is performed using an image processing parameter that is in accordance with evaluation results of the image quality (see Fig. 10). Accordingly, object detection processing of the second type object using the second AI model can be performed with an image, subjected to image processing with parameters that are appropriate in accordance with the image quality evaluation results of the detection region of the first type object, as the object. Thus, precision of object detection processing of the second type object can be improved.

[0101] Furthermore, in the image-capturing device according to an embodiment, the control unit performs evaluation relating to at least brightness, as evaluation of the image quality. Accordingly, image processing that makes brightness appropriate can be performed with regard to the detection region image of the first type object that is used for object detection processing of the second type object, and precision of object detection processing of the second type object can be improved.

[0102] Also, in the image-capturing device according to an embodiment, the control unit performs evaluation relating to a difference between brightness of an entirety of the original image, and brightness of the detection region of the first type object, as evaluation of the brightness. Accordingly, even in a case in which the brightness of the detection region image of the first type object is not appropriate due to AE control being performed with the brightness of the entire image-captured image as a reference, a detection region image of the first type object that has been adjusted to an appropriate brightness can be input for object detection processing of the second type object. Thus, precision of the object detection processing for the second type object can be improved.

[0103] Further, in the image-capturing device according to an embodiment, the control unit performs evaluation relating to at least noise, as evaluation of the image quality. This enables image processing that is denoising to be performed in a case in which noise is great with regard to the detection region image of the first type object used in the object detection processing for the second type object, and precision of the object detection processing for the second type object can be improved.

[0104] A control method according to an embodiment is a control method for an image-capturing device that includes a sensor unit having an image-capturing unit that has a pixel array unit, in which a plurality of pixels having photoreceptors are arrayed two-dimensionally, and configured to perform photoelectric conversion of light from a subject and obtain an image-captured image, and having an AI processing unit configured to perform AI processing, which is processing using an AI model, with the image-captured image being object, a memory unit that is provided outside of the sensor unit and configured to store a plurality of types of AI models, and a control unit that is provided outside of the sensor unit and configured to control the sensor unit, the control unit performing control such that the AI model used in the AI processing unit is switched among the plurality of types of AI models stored in the memory unit. Operations and effects that are the same as those of the image-capturing device according to an embodiment described above can be obtained by such a control method as well.

[0105] Note that the effects described in the present specification are only exemplary and not restrictive, and there may be other effects as well.

[0106] <6. Present Technology> The present technology can also assume configurations such as described below. (1) An image capturing device, comprising: a sensor configured to capture an image and to analyze the image according to an artificial intelligence (AI) model stored therein; and control circuitry configured to store a plurality of AI models including at least a first AI model to perform object detection and a second AI model to perform object detection, the first AI model being different from the second AI model, and switch the AI model stored in the sensor between at least the first AI model and the second AI model. (2) The image capturing device according to (1), wherein the first AI model identifies a target object, and the second AI model identifies a feature of the target object. (3) The image capturing device according to (1) or (2), wherein the control circuitry is further configured to provide, to the sensor, respective image processing parameters corresponding to a respective AI model when providing the respective AI model to the sensor. (4) The image capturing device according to claim 3, wherein the control circuitry is further configured to switch image processing parameters of the sensor to the respective image processing parameters when the control circuitry switches the AI model stored in the sensor to the respective AI model. (5) The image capturing device according to any one of (2) to (4), wherein prior to detecting the target object using the first AI model, the sensor downsamples the image to generate a downsampled image prior to performing object detection using the first AI model, and performs object detection on the downsampled image using the first AI model. (6) The image capturing device according to (5), wherein the sensor is further configured to crop the image to generate a cropped image prior to performing object detection with the second AI model, and performs object detection using the second AI model on the cropped image. (7) The image capturing device according to (6), wherein the cropped image is generated from the image prior to downsampling. (8) The image capturing device according to (6) or (7), wherein the cropped image includes the target object. (9) The image capturing device according to (8), wherein the sensor is further configured to determine whether an aspect ratio of the target object is different in the cropped image than an aspect ratio of the target object in the image, and to correct the aspect ratio of the target object in the cropped image to match the aspect ratio of the target object in the image. (10) The image capturing device according to any one of (6) to (9), wherein the sensor is configured to further crop the cropped image to generate a second cropped image after detection of the feature of the target object using the second AI model. (11) The image capturing device according to (10), wherein the second cropped image includes the feature of the target object. (12) The image capturing device according to (10) or (11), wherein the sensor is further configured to determine whether an aspect ratio of the feature of the target object is different in the second cropped image than an aspect ratio of the feature of the target object in the image, and to correct the aspect ratio of the feature of the target object in the second cropped image to match the aspect ratio of the feature of the target object in the image. (13) The image capturing device according to any one of (2) to (12), wherein the target object is a product in a warehouse, and the feature of the target product is a product identification code. (14) The image capturing device according to any one of (2) to (13), wherein the target object is a person and the feature of the target object is a face. (15) The image capturing device according to any one of (2) to (14), wherein the target object is a car and the feature of the target object is a license plate. (16) The image capturing device according to any one of (1) to (16), wherein the control circuitry is further configured to evaluate a quality of the image based on object detection results by the sensor using the first AI model, and update image processing parameters of the sensor based on an evaluation result of the quality of the image. (17) The image capturing device according to (16), wherein the control circuitry evaluates the quality of the image with respect to brightness. (18) The image capturing device according to (16), wherein the control circuitry evaluates the quality of the image based on an amount of noise present in the image. (19) An image capture system, comprising: a sensor configured to capture an image and to analyze the image according to an artificial intelligence (AI) model stored therein; control circuitry configured to store a plurality of AI models including at least a first AI model to perform object detection and a second AI model to perform object detection, the first AI model being different from the second AI model, and switch the AI model stored in the sensor between at least the first AI model and the second AI model; and a server to provide AI models to the control circuitry. (20) An image capturing method, comprising: storing, by control circuitry, a plurality of artificial intelligence models including at least a first AI model and a second AI model, the first AI model being different from the second AI model; switching, by the control circuitry, an AI model stored in a sensor to at least one of the first AI model and the second AI model; and capturing and analyzing, by the sensor, an image according to the AI model stored therein.

[0107] 1, 1A Image-capturing device 2 Image-capturing optical system 3 Sensor unit 4, 4A Camera control unit 4a, 4aA CPU 4b Memory unit 5 Communication unit 6 Optical system driving unit 7 Bus 31 Image-capturing unit 32 Image signal processing unit 33 AI processing unit 34 In-sensor control unit 35 Communication I / F (interface) MD1 First AI model MD2 Second AI model Ip1 First image processing instruction information Ip2 Second image processing instruction information BB1 to BB6 First type object detection region Bc Second type object detection region

Claims

1. An image capturing device, comprising:    a sensor configured to capture an image and to analyze the image according to an artificial intelligence (AI) model stored therein; and    control circuitry configured to    store a plurality of AI models including at least a first AI model to perform object detection and a second AI model to perform object detection, the first AI model being different from the second AI model, and    switch the AI model stored in the sensor between at least the first AI model and the second AI model.

2. The image capturing device according to claim 1, wherein the first AI model identifies a target object, and the second AI model identifies a feature of the target object.

3. The image capturing device according to claim 1, wherein the control circuitry is further configured to provide, to the sensor, respective image processing parameters corresponding to a respective AI model when providing the respective AI model to the sensor.

4. The image capturing device according to claim 3, wherein the control circuitry is further configured to switch image processing parameters of the sensor to the respective image processing parameters when the control circuitry switches the AI model stored in the sensor to the respective AI model.

5. The image capturing device according to claim 2, wherein prior to detecting the target object using the first AI model, the sensor downsamples the image to generate a downsampled image prior to performing object detection using the first AI model, and performs object detection on the downsampled image using the first AI model.

6. The image capturing device according to claim 5, wherein the sensor is further configured to crop the image to generate a cropped image prior to performing object detection with the second AI model, and performs object detection using the second AI model on the cropped image.

7. The image capturing device according to claim 6, wherein the cropped image is generated from the image prior to downsampling.

8. The image capturing device according to claim 6, wherein the cropped image includes the target object.

9. The image capturing device according to claim 8, wherein the sensor is further configured to determine whether an aspect ratio of the target object is different in the cropped image than an aspect ratio of the target object in the image, and to correct the aspect ratio of the target object in the cropped image to match the aspect ratio of the target object in the image.

10. The image capturing device according to claim 6, wherein the sensor is configured to further crop the cropped image to generate a second cropped image after detection of the feature of the target object using the second AI model.

11. The image capturing device according to claim 10, wherein the second cropped image includes the feature of the target object.

12. The image capturing device according to claim 11, wherein the sensor is further configured to determine whether an aspect ratio of the feature of the target object is different in the second cropped image than an aspect ratio of the feature of the target object in the image, and to correct the aspect ratio of the feature of the target object in the second cropped image to match the aspect ratio of the feature of the target object in the image.

13. The image capturing device according to claim 2, wherein the target object is a product in a warehouse, and the feature of the target product is a product identification code.

14. The image capturing device according to claim 2, wherein the target object is a person and the feature of the target object is a face.

15. The image capturing device according to claim 2, wherein the target object is a car and the feature of the target object is a license plate.

16. The image capturing device according to claim 1, wherein the control circuitry is further configured to evaluate a quality of the image based on object detection results by the sensor using the first AI model, and update image processing parameters of the sensor based on an evaluation result of the quality of the image.

17. The image capturing device according to claim 16, wherein the control circuitry evaluates the quality of the image with respect to brightness.

18. The image capturing device according to claim 16, wherein the control circuitry evaluates the quality of the image based on an amount of noise present in the image.

19. An image capture system, comprising:    a sensor configured to capture an image and to analyze the image according to an artificial intelligence (AI) model stored therein;    control circuitry configured to store a plurality of AI models including at least a first AI model to perform object detection and a second AI model to perform object detection, the first AI model being different from the second AI model, and switch the AI model stored in the sensor between at least the first AI model and the second AI model; and    a server to provide AI models to the control circuitry.

20. An image capturing method, comprising:    storing, by control circuitry, a plurality of artificial intelligence models including at least a first AI model and a second AI model, the first AI model being different from the second AI model;    switching, by the control circuitry, an AI model stored in a sensor to at least one of the first AI model and the second AI model; and    capturing and analyzing, by the sensor, an image according to the AI model stored therein.

Citation Information

Patent Citations

  • Image sensor, information processing method, and program

    WO2023218935A1