Information processing device, image capturing system, method, and program

The information processing device addresses the challenge of selecting objects or specific parts in digital cameras by using multiple trained models with varying detection granularities, enabling precise user selection and improving operability.

JP2025182055APending Publication Date: 2025-12-11CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025167589
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-03
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing object detection systems in digital cameras display vast amounts of information, making it difficult for users to visually identify and select intended objects or specific parts, leading to deteriorated user operability.

Method used

An information processing device that displays detection results from multiple trained models with varying detection granularities, allowing users to select objects or specific parts based on user operations, and determines the intended subject through a determination mechanism.

Benefits of technology

Enables users to accurately select objects or specific parts as intended, improving user operability by reducing visual clutter and enhancing selection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025182055000001_ABST
    Figure 2025182055000001_ABST
Patent Text Reader

Abstract

To provide a technique for allowing a user to select an intended object or specific part of the object in an image when selecting the object or specific part of the object.SOLUTION: An information processing device disclosed herein is configured to display on a screen a detection result of a trained model of interest among a plurality of trained models that differ in detection granularity of an object detected from an image, and determine an object to be subjected to predetermined processing on the basis of a user operation on the detection result displayed on the screen.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an imaging system, a method, and a program. [Background technology]

[0002] In recent years, advances in deep learning have led to significant improvements in the accuracy of object detection from images. Conventionally, object detection from images has been achieved by training neural networks (NNs) to recognize objects belonging to specific categories, such as faces and human bodies. Deep learning allows NNs to learn more abstract concepts than conventional methods. Deep learning enables NNs to learn object-likelihood using information about objects belonging to various categories, thereby enabling multi-object detection, which simultaneously detects objects from various categories.

[0003] Non-Patent Documents 1 to 3 describe methods for performing multi-object detection from images using deep learning. Furthermore, when a user photographs a subject, there is a need to be able to arbitrarily select the subject to be subjected to tracking processing and autofocus processing (hereinafter referred to as AF processing) from the screen of a digital camera, and the function of selecting a subject from the screen is widely implemented in existing products.

[0004] Patent Document 1 describes that a subject to be subjected to AF processing is designated according to a touch position on a touch panel, and that optimal AF processing is switched in conjunction with the designated subject. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Publication No. 2018-207309 [Non-patent literature]

[0006] [Non-Patent Document 1] Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation., Ross Girshick et al., 2014 [Non-patent document 2] SSD: Single Shot MultiBox Detector, Wei Liu et al., 2015 [Non-patent document 3] You Only Look Once: Unified, Real-Time Object Detection, Joseph Redmon et al., 2015 Summary of the Invention [Problem to be solved by the invention]

[0007] When multi-object detection that does not rely on a specific category is possible, the detection targets are objects such as people and automobiles, and specific parts that make up the objects, such as parts of people (faces and hands) and parts of automobiles (lights and tires). If all information about the detection targets is displayed on the screen of a digital camera or the like using a detection frame, the vast amount of information displayed on the screen can make it difficult for the user to visually identify the objects and specific parts. For example, the objects or specific parts that are the targets of AF processing vary depending on the user's photographic intent and preferences, making it difficult to define the objects or specific parts to be detected from the screen and automatically select and discard the vast amount of information.

[0008] On the other hand, when an object detection frame and a specific part detection frame are simultaneously displayed on the screen, the specific part of the object may be selected even if the user intends to select the object. When a huge amount of information is displayed on the screen in this way, the user is unable to select the intended object or specific part from the screen, which deteriorates the user's operability. Patent Document 1 describes an example in which AF processing is performed on an object selected by the user on a touch panel.

[0009] However, there is a problem in that it is difficult to identify whether the user has selected an object or a specific part that exists at the position touched on the touch panel.

[0010] The present invention provides a technique that enables a user to select an object or a specific portion of an object in an image as intended when selecting the object or the specific portion of the object. [Means for solving the problem]

[0011] In order to achieve the object of the present invention, an information processing device according to one embodiment of the present invention has the following configuration: That is, the information processing device is characterized by comprising: a display means for displaying on a screen a detection result by a trained model of interest among a plurality of trained models that have different detection granularity for objects detected from an image; and a determination means for determining an object for which a predetermined process is to be performed, based on a user operation on the detection result displayed on the screen. [Effects of the Invention]

[0012] According to the present invention, when a user selects an object or a specific portion of an object in an image, the user can select the object or the specific portion of the object as intended. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a diagram showing an example of a hardware configuration of an information processing apparatus. [Figure 2] FIG. 1 is a diagram showing an example of the functional configuration of an information processing apparatus according to a first embodiment. [Figure 3] 6 is a flowchart of a target subject determination process according to the first embodiment. [Figure 4] FIG. 10 is a diagram showing an example of the functional configuration of an information processing apparatus according to a second embodiment. [Figure 5] 10 is a flowchart of a target subject determination process according to the second embodiment. [Figure 6]FIG. 10 is a diagram showing an example of integrating detection frames of a plurality of specific parts. DETAILED DESCRIPTION OF THE INVENTION

[0014] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the claimed invention. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0015] (First embodiment) In the first embodiment, the detection results of an attention-trained model among multiple trained models with different detection granularity for objects detected from an image are displayed on a screen. In the first embodiment, an object to be subjected to a predetermined process is determined based on a user operation on the detection results displayed by switching between each attention-trained model. Here, the entirety of unspecified and diverse objects such as people, animals, and vehicles captured by an imaging device (e.g., a digital camera) is referred to as an "object." Meanwhile, parts of an object, such as parts of a person (hands and feet) and parts of a motorcycle (front light and tire), are referred to as "specific parts." In the first embodiment, a detection frame of the object or specific part is displayed on the screen and viewfinder of the imaging device (e.g., a digital camera), and the user selects the object or specific part from the screen.

[0016] In the first embodiment, the imaging device performs predetermined processing, such as tracking processing, AF processing, and counting processing, on an object or specific portion selected by a user on a screen. The first embodiment provides a UI (user interface) that enables a user to select an object or specific portion as intended. In the first embodiment, two trained models, one for object detection and one for specific portion detection, are maintained; however, three or more trained models with gradually changing detection granularity for objects may also be maintained. The detection granularity is defined as the size of the region of interest for the object. Furthermore, the present invention is not limited to performing tracking processing and AF processing on an object or specific portion selected by a user; a counting process may also be performed to count the number of objects or specific portions.

[0017] 1 is a diagram showing an example of the hardware configuration of an information processing device. The information processing device 100 includes a CPU 101, a memory 102, an input unit 103, a storage unit 104, a display unit 105, and a communication unit 106. The information processing device 100 is a general-purpose device capable of image processing, and includes, for example, a camera, a smartphone, a tablet, a PC, etc. The information processing device 100 may be used in combination with an imaging device (not shown) that captures an image of an object, and an imaging system (not shown) includes the imaging device and the information processing device 100.

[0018] The CPU 101 is a device that controls each part of the information processing device 100, and performs various processes by executing programs and data stored in the memory 102.

[0019] The memory 102 is a storage device that stores various data, boot programs, etc., and includes, for example, a ROM. The memory 102 provides a work area used when the CPU 101 executes various processes, and includes, for example, a RAM.

[0020] The input unit 103 is a device that accepts input of various instructions from the user, and includes, for example, a mouse, a keyboard, a joystick, and various operation buttons.

[0021] The storage unit 104 is a storage medium that stores various data and learning data for the NN, and includes, for example, an HDD, an SSD, a flash memory, an optical medium, and the like.

[0022] The display unit 105 is a device that displays various information processed by the CPU 101, and includes UIs (user interfaces) such as an LCD screen, an organic EL screen, a contact or non-contact touch panel, and an air-operated display. The display unit 105 displays on the screen images captured by an imaging device (not shown) and data received from a server (not shown). If the display unit 105 is a touch panel, the user inputs various instructions to the CPU 101 by touching the touch panel.

[0023] The communication unit 106 is a device for exchanging data between the various units in the information processing device 100, and includes, for example, a cable, a bus, a wired LAN, a wireless LAN, and the like.

[0024] 2 is a diagram showing an example of the functional configuration of the information processing device according to the first embodiment. The information processing device 100 includes a model holding unit 201, a detection unit 202, a subject determination unit 203, a display unit 204, and an input unit 205.

[0025] The model storage unit 201 stores trained models for at least two or more machine learning models. The model storage unit 201 stores, for example, two machine learning models, each with a different size of the attention area referenced when detecting an object or a part of an object (each with a different object detection granularity). Here, the machine learning model refers to a model learned using a machine learning algorithm such as deep learning (DL). Also, the trained model refers to a model that has been trained or learned in advance using appropriate teacher data for a machine learning model based on an arbitrary machine learning algorithm. However, this does not mean that the trained model does not undergo further learning beyond what has already been learned, and additional learning can also be performed.

[0026] Training data is training data used to train a machine learning model. Training data consists of pairs of input data (e.g., images) showing objects or specific parts belonging to various categories, and ground truth data in which the areas of the objects or specific parts in the images are displayed with frames. The input data are images captured in advance by an imaging device. Ground truth (GT) is ground truth data in which correct answer information is assigned in advance to the objects or specific parts in the images. The various categories are classifications of living things such as people, insects, and animals, and man-made objects such as cars and motorcycles, and include all objects to be detected.

[0027] The two trained models are realized by, for example, training a machine learning model using multiple pieces of training data with different sizes of focus areas when detecting objects, and by adjusting various hyperparameters during training. The model storage unit 201 prepares GT data A and GT data B for one piece of input data (image) as an example of multiple pieces of training data with different sizes of focus areas when detecting objects. GT data A is a GT in which a frame is added to the area of ​​each object (e.g., person and car) in the input data (image), and is used for training a model with a wide focus area for the object. GT data B is a GT in which a frame is added to the area of ​​a specific part of each object (e.g., person's face and car's tires) in the input data (image), and is used for training a model with a narrow focus area for the object.

[0028] When machine learning models are trained using input data (images) and GT data A or GT data B, trained model A trained with GT data A will detect objects, and trained model B trained with GT data B will detect specific parts. In this way, by preparing multiple pieces of training data with different sizes of focus areas when detecting objects and training the machine learning model with these training data, a trained model that detects objects or specific parts can be obtained.

[0029] The detection unit 202 detects objects or specific parts from an image using a known pattern recognition technique or a recognition technique using machine learning, and obtains the respective detection results. Here, detecting an object or a specific part means identifying the positions of objects or specific parts belonging to various categories from an image using two trained models held by the model holding unit 201.

[0030] The detection result of the object or specific part is represented by coordinate information on the image and a likelihood indicating the probability of the object or specific part being present. The coordinate information on the image is represented by the center position and size of a rectangular area on the image. Note that the coordinate information on the image may also include information regarding the rotation angle of the object or specific part.

[0031] The subject determination unit 203 determines an object or specific portion designated by the user on the screen using a detection frame of the object or specific portion detected by the trained model of the detection unit 202 and coordinate information received from the input unit 205, which will be described later. The detection frame of the object or specific portion is represented as an arbitrary shape on the image, such as a rectangle or ellipse. The display unit 204 superimposes the detection frame of the object or specific portion on the image and displays it on the screen of the display unit 105. The subject determination unit 203 saves the coordinate information of the object or specific portion selected by the user on the screen in the storage unit 104. The subject determination unit 203 also controls at least one of tracking processing, AF processing, and counting processing of the determined object or specific portion by instructing an imaging device (not shown) to perform these processes.

[0032] The display unit 204 simultaneously displays, on the screen of the display unit 105, the detection frame of the object or specific portion detected by the detection unit 202 and the target object or specific portion of interest determined by the subject determination unit 203. Here, the display unit 204 changes the thickness and color of the detection frame of the object or specific portion and the frame of the target object or specific portion of interest, and displays them on the screen in a distinguishable format.

[0033] The input unit 205 detects the position on the touch panel of the display unit 105 that is touched by the user's finger, and outputs coordinate information corresponding to this position to the subject determination unit 203 .

[0034] FIG. 3 is a flowchart of the target subject determination process according to the first embodiment.

[0035] In S301, the detection unit 202 acquires from the storage unit 104 an image in which an object appears.

[0036] In S302, the detection unit 202 selects a trained model to be used for the detection process of the target subject from the trained models related to the two machine learning models stored in the model storage unit 201. When performing the detection process of the target subject for the first time, the detection unit 202 selects the trained model with the widest region of interest for the object (the coarsest granularity of the detected object).

[0037] In addition, when the detection unit 202 judges No in the processing of S310 and performs the detection processing of the target subject for the second or subsequent time, it selects a trained model with a narrower area of ​​interest for the object (finer granularity of the detected object) than the trained model selected last time.

[0038] In S303, the detection unit 202 detects objects or specific parts belonging to various categories as objects from the image using the trained model selected in S302. The detection results of the objects or specific parts are represented by coordinate information on the image and likelihood.

[0039] In S304, the display unit 204 determines whether or not the detection process for the object on the image is the first time. If the display unit 204 determines that the detection process for the object on the image is the first time (Yes in S304), the process proceeds to S305. If the display unit 204 determines that the detection process for the object on the image is not the first time (No in S304), the process proceeds to S312.

[0040] In S305, the display unit 204 superimposes detection frames of objects or specific parts belonging to various categories detected in S303 on the image and displays them on the screen of the display unit 105. Here, the display unit 204 may display only detection frames of objects or specific parts whose likelihood exceeds a predetermined threshold, rather than superimposing detection frames of all objects or specific parts on the image and displaying them on the screen. If the display unit 204 determines that there is a lot of noise due to the detection frames of objects or specific parts, it can reduce the noise due to the detection frames of objects or specific parts by limiting the detection frames of objects or specific parts to be displayed on the screen. Note that, in the initial detection process for an object, the display unit 204 uses a trained model with the widest region of interest for the object, and therefore displays detection frames of objects belonging to various categories superimposed on the image on the screen.

[0041] In S312, the display unit 204 displays a detection frame in a state in which the peripheral area of ​​the detected object is enlarged on the screen of the display unit 105 in a superimposed manner.

[0042] In S306, the input unit 205 accepts input information from the user via the screen of the display unit 105. The user selects a detection frame corresponding to an object or a specific part on which at least one of tracking processing, AF processing, and counting processing is to be performed from among the detection frames on the image displayed by the display unit 105. The input unit 205 converts position information of the position on the touch panel where the user's finger touches into coordinate information on the image.

[0043] In S307, the detection unit 202 determines a subject of interest (object of interest or specific part of interest) using the coordinate information on the image acquired in S306 and the detection frame of the object or specific part detected in S303. The subject of interest is determined, for example, based on the detection frame of the object or specific part that has the closest Euclidean distance between the coordinate information on the image and the center coordinates of the detection frame of the object or specific part. Alternatively, the subject of interest may be determined by the user selecting one intended subject from a tree view, symbol, or the like displayed as a substitute for the detection frame of the object or specific part.

[0044] In S308, the detection unit 202 determines whether the selected trained model determined in S302 is the trained model with the narrowest region of interest for the object among the trained models in the model holding unit 201. If the detection unit 202 determines that the selected trained model is the trained model with the narrowest region of interest for the object (Yes in S308), the process proceeds to S311. If the detection unit 202 determines that the selected trained model is not the trained model with the narrowest region of interest for the object (No in S308), the process proceeds to S309.

[0045] In S309, the subject determination unit 203 determines whether or not to set the subject of interest determined in S307 as the final subject of interest. Here, the subject determination unit 203 accepts an input operation from the user regarding whether or not to end the subject of interest determination process.

[0046] In S310, the subject determination unit 203 determines whether to end the target subject determination process based on a first determination condition and a second determination condition. The first determination condition is that "the user selected to end the target subject determination process in S309." The second determination condition is that "the size of the target subject selected in S307 is smaller than the predetermined size of the target subject set in advance." If the subject determination unit 203 determines that either the first determination condition or the second determination condition is met (Yes in S310), the process proceeds to S311. If the subject determination unit 203 determines that neither the first determination condition nor the second determination condition is met (No in S310), the process returns to S302, and the target subject determination process continues.

[0047] In S311, the subject determination unit 203 determines the subject of interest found in S307 as the final subject of interest, stores the coordinate information of the subject of interest in the storage unit 104, and ends the subject of interest determination process. Thereafter, the display unit 204 superimposes a detection frame of the subject of interest on the image and displays it on the screen of the display unit 105. The subject determination unit 203 controls at least one of a tracking process, an AF process, and a counting process for the subject of interest by instructing the imaging device (not shown) to perform these processes.

[0048] (Modification 1 of the first embodiment) In S304, the display unit 204 does not need to determine whether or not the detection process for the object in the image is being performed for the first time. That is, the input unit 205 performs the process of S306 immediately after the process of S303. As a result, the display unit 204 displays the detection frame of the detected object superimposed on the original image on the screen, rather than superimposing and displaying on the screen of the display unit 105 a detection frame in a state in which the peripheral area of ​​the detected object is enlarged.

[0049] (Modification 2 of the first embodiment) In S307, the detection unit 202 may calculate an object detection frame using the detection frames of the multiple specific parts, rather than determining one target subject from the detection frames of the multiple specific parts detected in S303. The object detection frame is calculated, for example, as a larger detection frame (integrated detection result) by integrating the detection frames of the multiple specific parts. FIG. 6 is a diagram showing an example of integrating the detection frames of multiple specific parts. FIG. 6(a) shows a composite image in which the detection frames of multiple specific parts detected using the trained model B are superimposed on an image, with the multiple detection frames indicated by dashed lines representing the detection frames of the specific parts. FIG. 6(b) shows a composite image in which the object detection frame calculated by integrating the detection frames of the multiple specific parts is superimposed on an image, with the detection frame indicated by solid lines corresponding to the object detection frame. FIG. 6(c) shows a composite image in which all detection frames, including the solid-line detection frame and the dashed-line detection frame, are superimposed on an image. The detection frame indicated by solid lines in FIG. 6(c) is calculated to include all the dashed-line detection frames and the object (e.g., a car) and to be the smallest possible size.

[0050] Although an example has been described in which a detection frame for an object is calculated by integrating detection frames for multiple specific parts, a larger detection frame (integrated detection result) may be calculated by integrating multiple object detection results using a method similar to that described above. Then, the display unit 204 displays the calculated object detection frame (shown in FIG. 6(b)) or specific part detection frame (shown in FIG. 6(a)) on the screen of the display unit 105, superimposed on the image. The input unit 205 accepts input information from the user via the screen of the display unit 105. The user selects a detection frame corresponding to an object or specific part for which at least one of tracking processing, AF processing, and counting processing is to be performed from the object detection frames or specific part detection frames on the image displayed by the display unit 105.

[0051] (Modification 3 of the first embodiment) In the target subject determination process, even if the same trained model is used to detect specific parts from an image, the detection frame of the specific part changes depending on whether or not additional processing (e.g., setting a likelihood threshold when displaying the detection frame of the specific part) is performed. As described in FIG. 6, the detection unit 202 calculates the object detection frame shown in FIG. 6(b) based on the detection frame of the specific part shown in FIG. 6(a) detected using one trained model. In other words, when a new object detection frame is calculated based on multiple detection frames of specific parts that have changed due to additional processing, the size of the calculated object detection frame changes even if the trained model for detecting specific parts is the same. Therefore, the model holding unit 201 may hold only one trained model, rather than holding multiple trained models with different object detection granularities.

[0052] (Fourth modification of the first embodiment) The input unit 205 may acquire position information on an image using non-contact technology such as user gaze information and gestures, rather than acquiring coordinates from the position where the user's finger touches the touch panel. User gaze information refers to the coordinates of at least one point acquired by detecting the user's gaze directed at the display unit 105 using an imaging device or the like. Non-contact technology refers to technology that allows a user to perform input operations without touching the screen or buttons. Non-contact technology is realized by using sensors that utilize changes in infrared rays and capacitance, sensing technology such as image recognition by an imaging device or voice recognition, and wireless control technology using a mobile terminal (e.g., a smartphone or tablet). Screens used in non-contact technology further include, for example, non-contact touch panels and air-operated displays.

[0053] (Fifth Modification of the First Embodiment) When the user performs an operation to change the display on the screen of the display unit 105, the display unit 204 may change the selected trained model in accordance with the user's input and display a detection frame for an object or a specific part on the screen of the display unit 105. For example, when the user performs an input to enlarge the display of an object on the screen of the display unit 105, the detection unit 202 changes the selected trained model A for object detection to trained model B for specific part detection. The display unit 204 then displays a detection frame for the specific part detected using trained model B superimposed on the image. On the other hand, when the user performs an input to reduce the display of the specific part on the screen of the display unit 105, the detection unit 202 changes the selected trained model B for specific part detection to trained model A for object detection. The display unit 204 then displays a detection frame for the object detected using trained model A superimposed on the image.

[0054] As described above, according to the first embodiment, rather than simultaneously displaying the detection frames of the object and the specific part on the screen, the detection frames of the object or the specific part corresponding to the trained model of interest among the multiple trained models are displayed in stages. This makes it easier for the user to visually recognize the object or the specific part on the screen, and allows the user to easily select the object or the specific part. Furthermore, it is easy to determine whether the user intentionally selected the object or the specific part. According to the first embodiment, it is possible to accurately detect the object or the specific part selected by the user on the screen.

[0055] (Second embodiment) In the second embodiment, objects and specific portions are detected from an image using multiple trained models in advance, and one of the multiple trained models is set as the selected trained model. In the second embodiment, a detection frame for the object or specific portion corresponding to the selected trained model is displayed on the screen. In the second embodiment, the selected trained model is switched to another trained model by user input via a button or the like for switching the trained model. Therefore, in the second embodiment, it is possible for the user to select a specific portion by specifying coordinates once, without having to specify coordinates multiple times on the screen as in the first embodiment. Below, the differences between the second embodiment and the first embodiment will be described.

[0056] The hardware configuration of the information processing device 100 is the same as that of the first embodiment, and therefore description thereof will be omitted. Fig. 4 is a diagram showing an example of the functional configuration of the information processing device according to the second embodiment.

[0057] The information processing device 100 includes a model holding unit 401 , a detection unit 402 , a subject determination unit 403 , a display unit 404 , an input unit 405 , and a model selection unit 406 .

[0058] The model storage unit 401 has the same function as the model storage unit 201, and the input unit 405 has the same function as the input unit 205, so a description thereof will be omitted.

[0059] Similar to the detection unit 202, the detection unit 402 detects an object or a specific portion from an image, thereby acquiring a detection result of the object or the specific portion. The detection unit 402 differs from the detection unit 202 in that it uses a large number of trained models in one detection process. That is, the detection unit 402 detects an object and a specific portion from an image using all trained models held by the model holding unit 401 in one detection process. The model holding unit 401 holds the detection results of the object and the specific portion from the image by the detection unit 402. On the other hand, the detection unit 202 uses only one trained model selected in S302 of FIG. 3 as the trained model to be used in one detection process.

[0060] When the detection unit 402 receives from the model selection unit 406 a designation of a trained model to be selected from the multiple trained models held in the model holding unit 401, the detection unit 402 changes the currently selected trained model to the designated trained model. The detection unit 402 transmits the detection result of the object or specific part detected using the newly selected trained model to the subject determination unit 403.

[0061] The subject determination unit 403 determines a detection frame for an object of interest or a specific part of interest specified by the user on the image, using the detection frame for the object or specific part detected by the trained model of the detection unit 402 and the coordinate information received from the input unit 405. The detection frame for the object of interest or specific part of interest is represented as an arbitrary figure on the image, such as a rectangle or ellipse, and the display unit 204 displays the detection frame for the object or specific part on the screen of the display unit 105 by superimposing it on the image.

[0062] The display unit 404 displays on the screen of the display unit 105 the detection frame of the object or specific portion detected by the detection unit 402 and the detection frame of the target object or specific portion of interest determined by the subject determination unit 403 .

[0063] The model selection unit 406 accepts a user operation input to the information processing device 100 and outputs the accepted input to the detection unit 402. The user operation input is a selection of whether the next trained model to be selected is a trained model with a wider region of interest for an object than the currently selected trained model, or a trained model with a narrower region of interest. Upon receiving the user operation input, the model selection unit 406 transmits the user operation input to the detection unit 402. Then, the detection unit 402 changes the currently selected trained model to a new trained model in accordance with the received user operation input.

[0064] FIG. 5 is a flowchart of the target subject determination process according to the second embodiment.

[0065] In S501, the detection unit 402 acquires from the storage unit 104 an image in which an object appears.

[0066] In S502, the detection unit 402 detects objects and specific parts from the image using all trained models held in the model holding unit 401.

[0067] In S503, the display unit 404 displays on the screen of the display unit 105 a detection frame of an object or specific part detected by one trained model from among the detection results of the object and specific part detected by the detection unit 402. When performing the first detection process for an object, the model selection unit 406 selects the trained model with the widest area of ​​interest for the object. Alternatively, the size of the area of ​​interest displayed during the first object detection process may be a size set in advance by the user. Furthermore, when performing the second or subsequent object detection process after the process of S506, the model selection unit 406 selects the trained model selected in S506.

[0068] In S504, the input unit 405 or the model selection unit 406 receives input information from the user.

[0069] In S505, the detection unit 402 determines whether the input information is from the input unit 405 or the model selection unit 406. If the detection unit 402 determines that the input information is obtained from the model selection unit 406 (selection information of a trained model), the process proceeds to S506. On the other hand, if the detection unit 402 determines that the input information obtained from the input unit 405 is coordinate information on an image, the process proceeds to S507.

[0070] In S506, the model selection unit 406 uses the model selection information acquired in S504 to change the currently selected trained model to another trained model held by the model holding unit 401. The display unit 404 changes the detection frame of the object or specific part displayed on the screen of the display unit 105 according to the selected trained model, and the process returns to S503. The processes of S503 to S505 are the same as those described above, and therefore description thereof will be omitted.

[0071] In S507, the detection unit 402 detects a subject of interest using the coordinate information on the image acquired in S504 and the detection frame of the object or specific part according to the selected trained model. As in the first embodiment, the subject of interest is determined based on the detection frame of the object or specific part that has the shortest Euclidean distance between the coordinate information on the image and the center coordinates of the detection frame of the object or specific part. The subject determination unit 403 determines the subject of interest detected by the detection unit 402 as the final subject of interest, stores the coordinate information of the final subject of interest in the storage unit 104, and ends the subject of interest determination process. Thereafter, the display unit 204 superimposes the detection frame of the subject of interest on the image and displays it on the screen of the display unit 105. The subject determination unit 203 controls at least one of a tracking process, an AF process, and a counting process for the subject of interest by instructing an imaging device (not shown) to perform these processes.

[0072] (Modification 1 of the second embodiment) In S503, display unit 404 may switch the display of the detection frame of an object or a specific part displayed on the screen of display unit 105 after a predetermined time has elapsed, without receiving user input information via model selection unit 406. For example, display unit 404 displays the detection frame of an object on the screen of display unit 105, and after a predetermined time has elapsed since the display, displays the detection frames of specific parts for all objects on the screen. This allows the user to select a detection frame corresponding to an object or a specific part displayed on the screen, without switching between trained models that have different detection granularity for objects.

[0073] As described above, according to the second embodiment, by switching the selected trained model in response to a user operation, detection frames based on trained models with different object detection granularity can be displayed on the screen. This makes it possible to eliminate information other than that requested by the user from the screen, and to provide only the necessary information to the user.

[0074] (Other Examples) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0075] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0076] 10: Information processing device, 101: CPU, 102: ROM, 103: RAM, 104: Storage unit, 105: Input unit, 106: Display unit, 107: Communication unit

Claims

1. a display means for displaying on a screen a detection result of a trained model of interest among a plurality of trained models each having a different detection granularity of an object to be detected from an image; a determination means for determining an object for which a predetermined process is to be performed based on a user operation on the detection result displayed on the screen, 1. An information processing device comprising:

2. The display means displays on the screen a detection result by the attention-trained model selected by a user operation.

2. The information processing apparatus according to claim 1, wherein:

3. The display means, after a predetermined time has elapsed since the detection result by the attention-trained model was superimposed on the image and displayed, switches the attention-trained model to another attention-trained model among the plurality of trained models and displays the detection result by the other attention-trained model on the screen.

3. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.

4. the display means displays the object determined by the determination means and the detection result on the screen in a format that allows them to be distinguished from each other.

4. The information processing device according to claim 1, wherein the information processing device is a computer.

5. When the display means receives a user operation to change the display of the detection result on the screen, the display means switches the attention trained model to another attention trained model among the plurality of trained models in accordance with the change in display, and displays the detection result by the other attention trained model on the screen.

5. The information processing device according to claim 1, wherein the information processing device is a computer.

6. the detection result includes coordinate information and likelihood of the object on the image; the display means displays the detection result on the screen when the likelihood exceeds a threshold.

6. The information processing device according to claim 1, wherein the information processing device is a computer.

7. When it is determined that the detection process for the object on the image is not the first time, the display means displays on the screen an enlarged peripheral area of ​​the detection result selected by the user operation, superimposed on the image.

7. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.

8. the determination means determines the selected detection result as the object of interest when it has received a user operation to end the process of determining the object, or when it has determined that the detection result selected by the user operation is smaller than a predetermined object size.

8. The information processing device according to claim 1, wherein the information processing device is a computer.

9. The plurality of trained models include a first trained model having a coarse detection granularity for the object and a second trained model having a fine detection granularity for the object.

9. The information processing device according to claim 1, wherein the information processing device is a computer.

10. The information processing device according to claim 9, characterized in that the display means switches the display of the screen from displaying the detection results corresponding to the first trained model to displaying the detection results corresponding to the second trained model.

11. the predetermined processing includes at least one of a tracking processing, an AF processing, and a counting processing for the object determined by the determination means, a control unit that controls the imaging device to execute the predetermined processing; 11. The information processing device according to claim 1,

12. the detection result includes a detection result of at least one of the entire object and a specific part of the object.

12. The information processing device according to claim 1, wherein the information processing device is a computer.

13. the user operation includes an operation based on at least one of position information of a finger of the user touching the screen, line-of-sight information of the user, and a gesture of the user; 13. The information processing device according to claim 1, wherein the information processing device is a computer.

14. The screen includes at least one of a touch panel, a non-contact touch panel, and an air operation display; 14. The information processing device according to claim 1,

15. an imaging device that captures an image of the object; An information processing device according to any one of claims 1 to 14; An imaging system comprising:

16. a display step of displaying on a screen a detection result by a trained model of interest among a plurality of trained models each having a different detection granularity of an object to be detected from an image; a determination step of determining an object for which a predetermined process is to be performed based on a user operation on the detection result displayed on the screen. A method characterized by:

17. A program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Imaging apparatus, imaging method and program

    JP2018207309A