Information processing apparatus, method for controlling information processing, and storage medium

The information processing device addresses the challenge of controlling photography with multiple models by dynamically selecting and combining models based on shooting environment information, enhancing photography ease and efficiency.

JP2026019604APending Publication Date: 2026-02-05CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121293
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing technologies, such as those described in Patent Document 1, do not effectively facilitate the control of photography using appropriate combinations of multiple machine learning models, making it difficult for users to easily photograph intended subjects.

Method used

An information processing device that acquires shooting environment information, extracts a second group of models from a first group trained by machine learning, performs recognition processing, and executes imaging processing based on the recognition results, allowing for flexible and efficient shooting control using a combination of multiple models.

Benefits of technology

Enables users to easily photograph intended subjects by dynamically selecting and combining machine learning models suitable for the current shooting environment, improving photography efficiency and effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019604000001_ABST
    Figure 2026019604000001_ABST
Patent Text Reader

Abstract

To provide a technique for easily photographing a subject intended by a user.SOLUTION: An information processing apparatus includes an acquisition unit configured to acquire shooting environment information, an extraction unit configured to extract, based on the shooting environment information, a second model group from a first model group in which each model has been learned by machine learning, a recognition processing unit configured to execute recognition processing using the second model group, and a shooting unit configured to execute shooting processing based on a recognition result of the recognition processing unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, a control method for an information processing device, and a program. [Background technology]

[0002] In recent years, the use of machine learning technology has been progressing in various fields. The field of photography, such as cameras, is no exception, and there are many cases where machine learning models are used to control photography. Furthermore, as machine learning technology becomes more widespread, there is an increasing number of systems that allow many general users to create and use machine learning models, rather than just a limited number of companies and organizations.

[0003] The more machine learning models available, the more they can meet a variety of specific needs. However, model users must select the most appropriate model from the many options available, which takes time and effort.

[0004] In response to this, Patent Document 1 proposes a method for proposing a model suited to a purpose. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Patent No. 7068745 Summary of the Invention [Problem to be solved by the invention]

[0006] However, the technology disclosed in Patent Document 1 does not take into consideration the control of photography using an appropriate combination of multiple models, which poses a problem that it may be difficult for a user to easily photograph a subject that they intend.

[0007] The present invention has been made in view of the above-mentioned problems, and has an object to provide a technique for enabling a user to easily photograph an intended subject. [Means for solving the problem]

[0008] To achieve the above object, an information processing device according to the present invention comprises: An acquisition means for acquiring shooting environment information; an extraction means for extracting a second group of models from a first group of models, each of which has been trained by machine learning, based on the shooting environment information; a recognition processing means for executing a recognition process using the second model group; an imaging means for executing imaging processing based on the recognition result of the recognition processing means; The present invention is characterized by comprising: [Effects of the Invention]

[0009] According to the present invention, it becomes possible for a user to easily photograph an intended subject. [Brief explanation of the drawings]

[0010] [Figure 1] 4 is a flowchart showing the flow of processing performed by the information processing device according to the first embodiment. [Figure 2] FIG. 2 is a block diagram illustrating the functional configuration of the information processing device according to the first embodiment. [Figure 3] FIG. 1 is a diagram showing the hardware configuration of an information processing apparatus according to a first embodiment. [Figure 4] 4 is a flowchart showing the flow of processing performed by the information processing device according to the first embodiment. [Figure 5] FIG. 2 is a block diagram illustrating an example of the functional arrangement of an information processing device according to the first embodiment. [Figure 6] FIG. 3 is a diagram showing an example of attributes of each model in the first embodiment. [Figure 7] FIG. 3 is a diagram showing an example of a scene-model correspondence table according to the first embodiment. [Figure 8] FIG. 2 is a diagram showing an example of a screen display in the first embodiment. [Figure 9]10 is a flowchart showing the flow of processing performed by an information processing device according to a first modification of the first embodiment. [Figure 10] 10 is a flowchart showing the flow of processing performed by an information processing device according to a second modification of the first embodiment. [Figure 11] FIG. 10 is a block diagram illustrating the functional configuration of an information processing device according to a second modification of the first embodiment. [Figure 12] 10 shows an example of a history of scene recognition results in Modification 2 of the first embodiment. [Figure 13] 10 is a flowchart showing the flow of processing performed by an information processing device according to a third modification of the first embodiment. [Figure 14] FIG. 10 is a block diagram illustrating the functional configuration of an information processing device according to a third modification of the first embodiment. [Figure 15] 10 is a flowchart showing the flow of processing performed by an information processing device according to a second embodiment. [Figure 16] FIG. 10 is a block diagram illustrating the functional configuration of an information processing device according to a second embodiment. [Figure 17] FIG. 10 is a diagram showing an example of attributes of each model in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the claimed invention. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0012] (First embodiment) In the first embodiment, an example will be described in which a second model group is extracted from a first model group acquired in advance based on the scene recognition results obtained using a scene recognition model, and the second model group is used to control the shooting process.

[0013] <Hardware configuration of information processing device> FIG. 1 is a hardware configuration diagram of an information processing device 100 according to this embodiment. As shown in FIG. 1, the information processing device 100 includes a CPU 101, a ROM 102, a RAM 103, an HDD 104, a display unit 105, an operation unit 106, a communication unit 107, and an image capture unit 108. The CPU 101 is a central processing unit (CPU) that performs calculations and logical decisions for various processes and controls each component connected to a system bus 109. The ROM (Read-Only Memory) 102 is a program memory that stores programs for control by the CPU 101, including various processing procedures described below. The RAM (Random Access Memory) 103 is used as a temporary storage area such as the main memory and work area of ​​the CPU 101. Note that the program memory may be realized by loading a program into the RAM 103 from an external storage device connected to the information processing device 100.

[0014] The HDD 104 is a hard disk for storing electronic data and programs according to this embodiment. An external storage device may also be used to perform a similar function. Here, the external storage device can be realized, for example, by media (recording media) and an external storage drive for realizing access to the media. Known examples of such media include flexible disks (FDs), CD-ROMs, DVDs, USB memories, MOs, and flash memories. The external storage device may also be a server device connected via a network.

[0015] The display unit 105 is, for example, a CRT display, a liquid crystal display, or the like, and is a device that outputs images to a display screen. The display unit 105 may be an external device connected to the information processing device 100 by wire or wirelessly. The operation unit 106 may include a keyboard and a mouse, and accepts various operations by the user. The operation unit 106 may be a touch panel that allows touch input. The communication unit 107 performs two-way communication by wire or wirelessly with other information processing devices, communication devices, external storage devices, etc.

[0016] The image capturing unit 108 captures still images and videos. The captured still images and videos are recorded on the HDD 104. Alternatively, they may be transmitted to an external device via the communication unit 107. When capturing images, the images may be displayed as live view images on the display unit 105 to show the user the angle of view available for capturing images.

[0017] <Functional configuration of information processing device> 2 is a functional configuration diagram of the information processing device 100 according to this embodiment. The information processing device 100 is, for example, a camera (photographing device), and includes a photographing environment information acquisition unit 201, a model group extraction unit 202, a model group setting unit 203, a recognition processing unit 204, and a photographing unit 205.

[0018] The photographing environment information acquisition unit 201 acquires information about the photographing scene as photographing environment information. The current photographing scene is determined from a live view image acquired via the photographing unit 205. For example, if a swimming pool is shown in the live view image, the photographing scene can be determined as a swimming pool. Also, if a lion is shown in the live view image, the photographing scene can be determined as a lion-photographing scene.

[0019] The model group extraction unit 202 extracts a second model group from a first model group that has been acquired in advance, based on the imaging environment information acquired by the imaging environment information acquisition unit 201. The model group setting unit 203 selects a model from the second model group extracted by the model group extraction unit 202, taking into consideration the priority order. The recognition processing unit 204 executes subject recognition processing using the model set by the model group setting unit 203. The photographing unit 205 executes photographing processing based on the recognition result of the recognition processing unit 204.

[0020] <Shooting process> 3 is a flowchart showing the flow of processing performed by the information processing device according to this embodiment. The functions of the processing units shown in FIG. 2 are realized by the CPU 101 executing the processing in accordance with the flowchart in FIG.

[0021] 3 starts when the user starts capturing images. Here, the start of capturing images corresponds to, for example, turning on the power of the information processing device (camera) 100 to set it to a capture mode, or pressing a shutter button or the like to put the information processing device (camera) 100 in a sleep state into a state where it can capture images. Furthermore, if the information processing device (camera) 100 is a smartphone, the start of capturing images by the user may correspond to starting a camera app.

[0022] In S301, the image capturing environment information acquisition unit 201 acquires information about the image capturing scene as image capturing environment information.

[0023] In S302, the model group extraction unit 202 extracts a second model group from a first model group that has been acquired in advance, based on the image capture environment information acquired by the image capture environment information acquisition unit 201. Here, the model refers to a model that has been trained using a known machine learning technique, and in this embodiment, it is a model that has been created for the purpose of controlling image capture and supporting the user's image capture activity. When the processing of FIG. 3 is started, the first model group has been acquired in advance. Details of the process of acquiring this first model group will be described later. The second model group is a set of models that are suitable for supporting the user's image capture in the current image capture environment, and is a subset of the first model group.

[0024] The second model group can be expanded in RAM 103, which can be read at high speed, and used for shooting control. The reason for expanding the second model group in RAM 103 is that if the second model group were expanded in a storage device that takes a long time to read, shooting control would take too long for the timing at which the user wants to shoot, interfering with the user's shooting activity. Furthermore, one model in the second model group is insufficient; in fact, it is better to have more models. This is because it is difficult to control shooting with only one model in a variety of situations, and using more models allows for more flexible shooting control.

[0025] For example, consider a case where shooting control is performed using a model that detects subjects. A user may want to shoot more than one type of subject; multiple types of subjects may appear at the same time, or the type of subject may change in a short period of time. For example, in places like zoos and aquariums, there are many opportunities to shoot a variety of subjects.

[0026] In such a case, if only one model is immediately available for use, the user would have to switch the model they use every time the subject changed, which would be inconvenient for the user and could result in them missing a good opportunity to take a photo.

[0027] However, information processing devices such as cameras and smartphones have limited resources, including an upper limit on RAM capacity. It is practically difficult to keep all models stored in RAM at all times. The time required for inference processing using machine learning models must also be taken into consideration. If too many models are processed simultaneously, the processing time may become too long, potentially interfering with the user's photography.

[0028] Therefore, in this embodiment, a realistic number of models suitable for supporting the user in taking pictures in the current shooting environment are extracted, and these are used for shooting support as a second model group. Furthermore, since the shooting environment is constantly changing and the models suitable for shooting control are also constantly changing, in this embodiment, a second model group is extracted each time according to the shooting environment at the time of shooting. In this embodiment, the maximum number of the second model group will be described as three, but the present invention is not limited to this example.

[0029] In S303, the model group setting unit 203 selects a model from the second model group extracted by the model group extraction unit 202, taking into consideration the priority order, and sets the selected model in the recognition processing unit 204. For example, if the shooting scene is an aquarium pool, a model for a sea lion, a model for a dolphin, and a model for a penguin can be selected.

[0030] In S304, the recognition processing unit 204 uses the model set by the model group setting unit 203 to perform subject recognition processing.

[0031] In S305, the photographing unit 205 executes photographing processing based on the recognition result obtained by the recognition processing unit 204. For example, if a dolphin is included in the live view image, the photographing unit 205 may recognize the dolphin and automatically focus on the recognized dolphin. Then, the photographing processing of the dolphin may be executed in response to the user pressing the photographing button.

[0032] In S306, the photographing unit 205 determines whether photographing is continuing. If this step is Yes, the process returns to S301. On the other hand, if this step is No, the process ends. The processes from S301 to S305 are repeatedly executed while the user repeats photographing, and the process ends when the user finishes photographing. Here, the end of photographing corresponds to, for example, turning off the power of the information processing device (camera) 100 or switching the mode of the information processing device (camera) 100 from photographing mode to photo viewing mode. Furthermore, the end of photographing corresponds to the information processing device (camera) 100 entering a sleep state due to no operation for a set time or more, or, if the information processing device (camera) 100 is a smartphone, closing the camera app and performing another operation.

[0033] <Functional configuration of an information processing device that performs overall processing including pre-photography processing> Next, Fig. 4 is an example of a block diagram showing the functional configuration of information processing device 100 that performs the processing of Fig. 5 described below. The function of each processing unit shown in Fig. 4 is realized by CPU 101 loading a program stored in ROM 102 into RAM 103 and executing processing according to the flowchart of Fig. 5. The execution results of each processing are then stored in RAM 103. Furthermore, for example, when configuring hardware as an alternative to software processing using CPU 101, it is sufficient to configure an arithmetic unit or circuit corresponding to the function of each processing unit described here.

[0034] The same reference numerals are assigned to the same processing units as those shown in Fig. 2. In addition to the components shown in Fig. 2, the information processing device 100 further includes a model group acquisition unit 401, a model group storage unit 402, and a scene recognition model acquisition unit 403. The information processing device 100 is also connected to an input unit 404 and a display unit 405. However, the information processing device 100 is not limited to this example, and may also include the input unit 404 and the display unit 405.

[0035] The model group acquisition unit 401 acquires a first model group, which is a collection of machine learning models that individually recognize each subject in various facilities. The model group storage unit 402 stores the first model group acquired by the model group acquisition unit 401. The scene recognition model acquisition unit 403 acquires a scene recognition model, which is a machine learning model used to recognize a photographed scene. The input unit 404 accepts various input operations from the user. The display unit 405 displays various information to the user.

[0036] <Overall processing including pre-shooting processing> 5 is a flowchart showing the overall process flow including pre-shooting processes performed by the information processing device 100 in this embodiment. In addition to the processes outlined with reference to FIG. 3, processes performed by a user before shooting are also included. As a specific example, consider a situation in which a user takes pictures of animals in a facility such as a zoo or aquarium.

[0037] In S501, the model group acquisition unit 401 acquires a first model group. The first model group is specifically a collection of machine learning models created by facilities such as zoos and aquariums to individually recognize each subject within the facility. For example, there is a model that recognizes lions, a model that recognizes dolphins, etc. Each model also holds, as an attribute, information on the optimal shooting scene for photographing the recognition target within the facility. As shown in the table in FIG. 6, one model holds, as an attribute, information on one or more optimal shooting scenes. FIG. 6 shows four models with model IDs 601 of model001 to model004. Each model has three attribute values: a model name 602, an optimal shooting scene (1) 603, and an optimal shooting scene (2) 604.

[0038] 6, the model whose model ID 601 is model001 has a model name 602 of "lion recognition model," an optimal shooting scene (1) 603 of "scene recognition result is a lion image," and an optimal shooting scene (2) 604 of "none." The model whose model ID 601 is model002 has a model name 602 of "giraffe recognition model," an optimal shooting scene (1) 603 of "scene recognition result is a giraffe image," and an optimal shooting scene (2) 604 of "none."

[0039] For the model whose model ID 601 is model003, the model name 602 is a "dolphin recognition model," the optimal shooting scene (1) 603 is a "scene recognition result of a dolphin image," and the optimal shooting scene (2) 604 is a "scene recognition result of a swimming pool image." For the model whose model ID 601 is model004, the model name 602 is a "sea lion recognition model," the optimal shooting scene (1) 603 is a "scene recognition result of a sea lion image," and the optimal shooting scene (2) 604 is a "scene recognition result of a swimming pool image."

[0040] In response to an instruction from a user, the model group acquisition unit 501 downloads the first model group, for example, from the Internet. The URL from which the download is made may be communicated to the user by the facility when the user acquires an admission ticket to the facility, or may be posted within the facility using a two-dimensional code or the like, allowing the user to download the first model group at any time. Alternatively, a flash memory card on which the first model group is recorded may be distributed to the user, and the user may connect the flash memory to the information processing device 100 as an additional external storage device.

[0041] As a modified example of the process of S501, the model group acquisition unit 501 may automatically acquire the first model group without following an instruction from the user. For example, the first model group may be automatically downloaded when it is determined that the user has entered the premises of a facility based on current location information of the information processing device 100.

[0042] In S502, the model group storage unit 502 stores the first model group acquired by the model group acquisition unit 501 in the HDD 104. As described above, the HDD 104 may be cloud storage connected via a network.

[0043] In S503, the scene recognition model acquisition unit 503 acquires a scene recognition model. The scene recognition model is a machine learning model used to recognize a user's captured scene. Specifically, it is an image classification model that uses known technology to identify the type of main subject of an image and classify the image into categories. The scene recognition model classifies, for example, an image of a lion as a lion image, an image of a dolphin as a dolphin image, an image of a pool as a pool image, and an image of a sunset as a sunset image. The classification result is output as a numerical value representing the captured scene. For example, a lion image is output as scene number 1, a giraffe image as scene number 2, and so on. The scene recognition model is created by facilities such as zoos and aquariums together with the first model group, and is created so that the scene recognition result and the model can be associated, as described below.

[0044] Furthermore, in S503, the model group storage unit 502 creates a scene-model correspondence table in the HDD 104 that associates the object types that the scene recognition model can classify with each model in the first model group. An example of the scene-model correspondence table is shown in FIG. 7. In FIG. 7, the photographic scene ID 701 is the numerical value of the scene recognition result output by the scene recognition model. The optimal model ID (1) 702 and the optimal model ID (2) 703 are the IDs of the models associated with the photographic scene ID 701. The photographic scene description 704 is not essential for processing, but is a comment that describes the photographic scene in each row to make it easier for humans to understand. As shown in FIG. 7, one photographic scene may not necessarily be associated with one optimal model.

[0045] 7, for a photographing scene with a photographing scene ID 701 of "001", an optimal model ID (1) 702 is "model001", an optimal model ID (2) 703 is "none", and a photographing scene description 704 is "The scene recognition result is a lion image". For a photographing scene with a photographing scene ID 701 of "002", an optimal model ID (1) 702 is "model002", an optimal model ID (2) 703 is "none", and a photographing scene description 704 is "The scene recognition result is a giraffe image".

[0046] For a photographing scene with a photographing scene ID 701 of "003", the optimal model ID (1) 702 is "model003", the optimal model ID (2) 703 is "none", and the photographing scene description 704 is "the scene recognition result is a dolphin image". For a photographing scene with a photographing scene ID 701 of "004", the optimal model ID (1) 702 is "model004", the optimal model ID (2) 703 is "none", and the photographing scene description 704 is "the scene recognition result is a sea lion image". For a photographing scene with a photographing scene ID 701 of "005", the optimal model ID (1) 702 is "model003", the optimal model ID (2) 703 is "model004", and the photographing scene description 704 is "the scene recognition result is an image of a swimming pool".

[0047] After the processing of S503, the process proceeds to S301. In S301, the shooting environment information acquisition unit 201 performs scene recognition of the live view image using the scene recognition model acquired in S503. A live view image is a captureable image captured by the sensor of the shooting unit 108 of the information processing device 100. Specifically, for example, if the user points the shooting unit 108 at a lion and the lion is captured within a captureable angle of view, the scene recognition model classifies the live view image as a lion image, that is, recognizes it as a lion shooting scene. Similarly, if a pool is captured in the live view image, it recognizes it as a pool shooting scene. Then, the shooting environment information acquisition unit 201 acquires the recognized shooting scene (scene recognition result) as shooting environment information.

[0048] In S302, the model group extraction unit 202 extracts a second model group optimal for the current shooting scene based on the scene recognition result, which is shooting environment information. The model group extraction unit 202 compares the scene recognition result acquired as shooting environment information by the shooting environment information acquisition unit 201 with the scene-model correspondence table stored in the model group storage unit 402, and selects multiple models optimal for the shooting scene in which the user is currently located. In this embodiment, as described above, the maximum number of second model groups is three. If only two optimal models can be extracted from the scene recognition result, the remaining one may be selected from the previous extraction result, for example, the most frequently used model may be retained. Alternatively, the third model may be left as not applicable, and an arbitrary model may be added to the second model group based on user input, as described below. The model group extraction unit 202 loads the extracted second model group in RAM 103, enabling high-speed model reading during the recognition process, as described below.

[0049] Here, the display unit 405 and the input unit 404 will be described. The display unit 405 can display various information to the user in parallel with the processing flow shown in Fig. 5. The screen displayed to the user will be described with reference to Fig. 8.

[0050] Screen 800 is a screen that displays live view video, and shows the video captured by the imaging unit 108 to the user. Model column 801 displays some of the first model group in a horizontal row. Since the first model group is usually a collection of many models, only some of them are displayed on the screen, rather than all of them. In this embodiment, for example, models that have been extracted many times as part of the second model group are displayed preferentially. Models that are hidden can also be displayed sequentially in response to a user's scrolling instruction.

[0051] Frame 802 is a frame indicating the second model group. The second model group extracted from the first model group is contained inside frame 802. Figure 8 shows that three models, a sea lion model, a dolphin model, and a penguin model, have been extracted as the second model group. Frame 803 indicates the model that is currently applied to the shooting control. Frame 804 is a frame that shows the user the state of shooting control. In this embodiment, a dolphin model is used to display the state of autofocus control. Frame 804 surrounds a dolphin 805 shown in the live view video, indicating that the dolphin 805 is being automatically focused on.

[0052] The input unit 404 accepts a user instruction to change the second model group. In this embodiment, the user can replace the model by touching and dragging the rectangle representing the model with a finger via a touch panel mounted on the screen 800, thereby moving the model into or out of the frame 802.

[0053] For example, if the user drags the sea lion model 806 in Fig. 8 to the left and moves it to the left of the walrus model 807, the walrus model 807 will move rightward into the frame 802, and the sea lion model 806 and the walrus model 807 will be swapped. This operation removes the sea lion model from the second model group, and adds the walrus model to the second model group as a new model. In this way, the user can change the contents of the second model group at any time to suit the subject they want to photograph.

[0054] In S303, the model group setting unit 203 selects one or more models from the second model group extracted by the model group extraction unit 202, taking priority into consideration, and sets the selected models in the recognition processing unit 204. In this embodiment, the model extracted based on the immediately preceding scene recognition result is set as the model with higher priority. Note that a user input may be configured to accept an instruction for prioritizing the models.

[0055] In S304, the recognition processing unit 204 performs recognition processing using the model set in S303. The object is detected from the image captured by the image capture unit 108 using the set model. For example, if a dolphin model is set as shown in FIG. 8, a dolphin is detected in the image. If a dolphin cannot be detected in the image, a model with the second highest priority, such as a sea lion model, is used to attempt to detect a sea lion.

[0056] Note that if the recognition processing unit 204 is capable of performing recognition processing in parallel using multiple models, it may perform parallel processing. Furthermore, the recognition processing unit 204 may detect one subject by sequentially processing models that fulfill multiple roles. The model used by the recognition processing unit 204 may be not only a model for detecting a subject, but also a noise reduction model. In other words, a noise reduction model that is optimal for a photographed scene may be included in the second model group, and noise reduction processing of the image may be performed.

[0057] In S305, the photographing unit 205 performs photographing processing. The photographing unit 205 controls the focus based on the recognition result in S304. The photographing unit 205 also photographs and records an image according to the user's input (photographing instruction). In this embodiment, the focus is continuously adjusted to the detected subject. Note that the photographing unit 205 may control not only the focus but also the white balance. Furthermore, although an example of subject detection using an object detection model has been described so far, the subject may also be identified using a region recognition model.

[0058] In S306, the photographing unit 205 determines whether photographing is continuing. If this step is Yes, the process returns to S301. On the other hand, if this step is No, the process ends. This completes the process in FIG. 5.

[0059] As described above, according to the first embodiment, it is possible to perform photography using a combination of multiple machine learning models that are suitable for the photography environment.

[0060] [Modification 1 of the First Embodiment] As a first variant of the first embodiment, an example will be described in which the extraction results of the second model group can be adjusted to match the user's preferences by accepting correction instructions from the user regarding the model attributes of the first model group.

[0061] <Processing> FIG. 9 is a flowchart showing the flow of the model attribute correction process in this first modified example. The block diagram showing the configuration of each processing unit that performs the process in FIG. 9 is the same as that in FIG. 4, so its description will be omitted. The functions of each of these processing units are realized by CPU 101 loading a program stored in ROM 102 into RAM 103 and executing the process in accordance with the flowchart in FIG. 9. The execution results of each process are then stored in RAM 103. Furthermore, for example, when configuring hardware as an alternative to software processing using CPU 101, it is sufficient to configure an arithmetic unit or circuit that corresponds to the process of each functional unit described here.

[0062] In S901, the input unit 404 receives an instruction from the user to select a model to be edited. In this embodiment, the user can select a model by long pressing one of the models in the model column 801 on the screen 800 shown in Fig. 8. As a result of the selection, a table listing the models and their attributes, as shown in Fig. 6, is displayed on the screen.

[0063] In S902, the input unit 404 receives a model attribute selection instruction from the user. For example, the user selects one cell in the displayed table of Fig. 6 to specify the attribute to be edited. For example, the user selects the cell for "optimal shooting scene (2)" of "model004".

[0064] In S903, the input unit 404 accepts input of a new attribute value by the user. In this embodiment, a choice of values ​​that can be set by the user is displayed, and the user selects a desired attribute value from the choices, thereby inputting a new attribute value. For example, in the case of the table of FIG. 6, an edit is accepted in which the user deletes the value of "best shooting scene (2)" for "model004." This makes it possible to prevent the sea lion recognition model from being extracted as part of the second model group when the scene recognition result is a pool image.

[0065] In S904, based on the input attribute value, the model group storage unit 402 rewrites the attribute value of the first model group recorded in the HDD 104. At this time, the attribute value before the change is stored separately with a different name so that the user can later return the attribute change to the original state.

[0066] In S905, the input unit 404 determines whether or not an instruction to complete the model attribute correction work has been input by the user. If the correction is complete, the process ends. On the other hand, if the correction is to continue, the process returns to S901. This completes the process in FIG. 9.

[0067] As described above, according to the first modification of the first embodiment, the user can customize the extraction of the second model group to suit the user's preferences by editing the attribute values ​​of the first model group.

[0068] [Modification 2 of the First Embodiment] As a second variant of the first embodiment, an example will be described in which a second model group is extracted based on the history of scene recognition results rather than a single scene recognition result, thereby realizing extraction of a second model group that better matches the user's intentions for shooting, focusing on the differences from the first embodiment.

[0069] <Functional configuration of information processing device> Fig. 10 is a diagram showing an example of the functional configuration of an information processing device 100 according to Modification 2. The functions of each processing unit in Fig. 10 are realized by a CPU 101 having the hardware configuration described above executing processing in accordance with the flowchart in Fig. 11, which will be described later. In Fig. 10, an environment information storage unit 1001 is added to the components shown in Fig. 4. The image capture environment information storage unit 1101 records the scene recognition result acquired as image capture environment information together with the current time.

[0070] <Processing> Fig. 11 is a flowchart showing the main processing flow of Modified Example 2. The processes of S501 to S503 and S301 in Fig. 11 are basically the same as the processes of S501 to S503 and S301 in Fig. 5. Furthermore, the processes of S303 to S306 in Fig. 11 are basically the same as the processes of S303 to S306 in Fig. 5. Therefore, a description thereof will be omitted.

[0071] In S1101, the shooting environment information storage unit 1101 records the scene recognition result, which is the shooting environment information acquired in S301, together with the current time. Here, FIG. 12 is an example of a table recording the history of scene recognition results. In FIG. 12, the scene recognition results for the past five times are recorded in rows No. 1 to 5, and the scene recognition result acquired in the immediately preceding S301 is recorded in the bottom row No. 6. In No. 1 to 5, the scene recognition results are all "lion images," and in No. 6, the scene recognition result is "giraffe image."

[0072] In S1102, the model group extraction unit 202 extracts a second model group based on the history of the shooting environment information (scene recognition results) recorded in the shooting environment information storage unit 1001. This will be described in detail with reference to the example of FIG. 12. Since the most recent scene recognition result is a giraffe image No. 6, the process is the same as S302 in FIG. 5, up to which a giraffe model optimal for a giraffe image is extracted as the first candidate. In this modified example, since the scene recognition results for Nos. 1 to 5 are consecutive lion images, a lion model optimal for a lion image is also extracted as part of the second model group with a higher priority than other models. This is because, if a user has photographed lions with a certain frequency in the recent past, it can be determined that the user is still likely to photograph lions. Furthermore, since it is conceivable that a giraffe may accidentally appear in a shot while the user was intending to photograph a lion, a lion model optimal for a lion image is also extracted as part of the second model group with a higher priority than other models.

[0073] The degree of importance to be attached to the history of past scene recognition results is determined by preset thresholds such as frequency within a specified time range and the degree of time passage. For example, recognition models corresponding to scene recognition results recorded at a predetermined frequency or more within a time range from 30 minutes ago to the present can be preferentially extracted as the second model group. Furthermore, if there are multiple scene recognition results recorded at a predetermined frequency or more within a time range from 30 minutes ago to the present, a recognition model corresponding to the scene recognition result whose latest scene recognition result is closest to the current time can be preferentially extracted as the second model group.

[0074] As a further modification, in S301, the shooting environment information acquisition unit 201 may perform scene recognition not only on live view images but also on images already captured by the user and acquire the scene recognition results as shooting environment information. By treating the scene recognition results on captured images as part of the scene recognition result history, it becomes possible to extract a second model group that also takes into account the tendencies of images actually captured by the user.

[0075] As described above, according to the second modification of the first embodiment, the second model group can be extracted based on not only the most recent scene recognition result but also past scene recognition results, thereby enabling shooting that is closer to the user's intention.

[0076] [Modification 3 of the First Embodiment] As a third variant of the first embodiment, an example will be described in which the weight parameters of the scene recognition model are adjusted to extract a second model group that is more in line with the user's shooting intentions, focusing on the differences from the second variant of the first embodiment.

[0077] <Functional configuration of information processing device> Fig. 13 is a diagram showing an example of the functional configuration of an information processing device 100 according to Modification 3. The functions of each processing unit in Fig. 13 are realized by a CPU 101 having the hardware configuration described above executing processing in accordance with the flowchart in Fig. 14, which will be described later. In Fig. 13, a scene recognition model adjustment unit 1301 is added to the components shown in Fig. 10. The scene recognition model adjustment unit 1301 adjusts weight parameters of a scene recognition model.

[0078] <Processing> 14 is a flowchart showing the main processing flow of Modification Example 3. In S1401, the scene recognition model adjustment unit 1401 adjusts the weight parameters of the scene recognition model. Specifically, based on the history of scene recognition results up to that point, the recognition tendency of the scene recognition model is adjusted so as to make it easier to recognize photographic scenes that are more frequently photographed or captured as live view images by users.

[0079] As described above, according to the third modification of the first embodiment, it is possible to extract a second model group that is more in line with the user's intentions for taking a photograph and then take a photograph.

[0080] (Second embodiment) In the first embodiment, an example in which a scene recognition result using a scene recognition model is acquired as the shooting environment information will be described. In the second embodiment, an example in which the current location and time at the time of shooting are acquired as the shooting environment information will be described. By using the current location information and time instead of the scene recognition model, the cost of creating the scene recognition model can be reduced. The second embodiment will be described below, focusing on the differences from the first embodiment.

[0081] <Functional configuration of information processing device> Fig. 15 is a diagram showing an example of the functional configuration of an information processing device 100 according to the second embodiment. The functions of each processing unit in Fig. 15 are realized by a CPU 101 having the hardware configuration described above executing processing in accordance with the flowchart in Fig. 16, which will be described later. In Fig. 15, the information processing device 100 includes a model group acquisition unit 1501, a model group storage unit 402, a shooting environment information acquisition unit 1502, a model group extraction unit 1503, a model group setting unit 203, a recognition processing unit 204, and a shooting unit 205. Furthermore, the information processing device 100 is connected to an input unit 404 and a display unit 405, but is not limited to this example, and at least one of these may be included in the information processing device 100.

[0082] The model group acquisition unit 1501 acquires a first model group in which each model has information on the optimal shooting position and optimal shooting time as attribute values. The shooting environment information acquisition unit 1502 acquires information on the current location and / or information on the current time as shooting environment information. The model group extraction unit 1503 extracts a second model group that is optimal for the current shooting scene based on the information on the current location and / or the information on the current time acquired by the shooting environment information acquisition unit 1502. The other components have the same functions as the previously mentioned components with the same reference numerals, and therefore their description will be omitted.

[0083] <Processing> 16 is a flowchart showing the main processing flow of this embodiment. The processing other than S1601, S1602, and S1603 is basically the same as the steps with the same numbers shown in FIG. 4, so a description thereof will be omitted.

[0084] In S1601, the model group acquisition unit 1501 acquires a first model group. Each model in the first model group has information on the optimum shooting position and optimum shooting time as attribute values. Here, FIG. 17 shows an example of each model in the first model group and the attributes that each has. FIG. 17 shows a state in which four models with model IDs 1701 of model001 to model004 each have three attribute values: model name 1702, optimum shooting position 1703, and optimum shooting time 1704.

[0085] 17, for a model with model ID 1701 "model001", model name 1702 is "lion recognition model", optimal shooting position 1703 is "(x1, y1)", and optimal shooting time 1704 is "any". For a model with model ID 1701 "model002", model name 1702 is "giraffe recognition model", optimal shooting position 1703 is "(x2, y2)", and optimal shooting time 1704 is "any". For a model with model ID 1701 "model003", model name 1702 is "dolphin recognition model", optimal shooting position 1703 is "(x3, y3)", and optimal shooting time 1704 is "15:00 to 16:00". The model with model ID 1701 "model004" has model name 1702 "sea lion recognition model", optimal shooting position 1703 "(x3, y3)", and optimal shooting time 1704 "15:30-16:30". Here, the optimal shooting position 1703 for both the dolphin recognition model and the sea lion recognition model is "(x3, y3)", but this position coordinate may be the center of the spectator seats in the pool.

[0086] The optimal shooting position 1703 is information about the optimal position for photographing a target subject. For example, in the case of a lion recognition model used in a zoo, it may be the coordinates of the center of the area in front of the lion's cage where the lion can be photographed. Around that point, the user is more likely to photograph a lion than other animals, so it can be determined that the lion recognition model should be included in the second model group. Note that, in consideration of cases where the photographable area has a complex shape, it is possible to provide information about multiple optimal shooting positions, or to define the photographable area as a two-dimensional figure using an array of coordinates.

[0087] The optimal shooting time 1704 is information about the optimal time to photograph the target subject. For example, in the case of a dolphin recognition model that can be used in an aquarium, this is between the start time and the end time of a show in which dolphins appear. Since it is unlikely that a user will be attempting to photograph a dolphin outside of this time period, it can be determined that the dolphin recognition model should not be included in the second model group. Similarly, in the case of a sea lion recognition model that can be used in an aquarium, this is between the start time and the end time of a show in which sea lions appear. In the example of FIG. 17, the optimal shooting time for lions and giraffes, which have no particular time period restrictions, is set to "any," meaning that any time period is acceptable. Note that, as with the optimal shooting position, information about multiple optimal shooting times may be stored.

[0088] In S1602, the imaging environment information acquisition unit 1502 acquires information about the current location and time of the information processing device 100. Known technologies can be used to acquire the location and time. Examples of methods for acquiring location information include using GPS and so-called beacon technology, which acquires the current location by obtaining wireless information from a device installed in a facility.

[0089] In S1603, the model group extraction unit 1503 extracts a second model group that is optimal for the current shooting scene based on the current location and time information acquired by the shooting environment information acquisition unit 1502. For example, models whose distance from the optimal shooting location information of the first model group to the current location is within a threshold are extracted. Similarly, models whose difference from the optimal shooting time information of the first model group to the current time is within a threshold are extracted. Note that the second model group may be extracted using only either the location information or the time information.

[0090] As a modification of this embodiment, the second model group may be extracted using information on the shooting direction in addition to the location and time. By considering the shooting direction in addition to the shooting location, it is possible to extract a model that is more suitable for the subject that the user intends to photograph. For example, even if there are multiple subjects that can be photographed from the same location in different directions, by considering the direction in which the user intends to photograph, it is possible to preferentially extract models of subjects that the user is likely to photograph.

[0091] Furthermore, a second group of models may be extracted from the first group of models based on the history of current location information, the history of shooting direction information, and the location information associated with each model in the first group of models. This allows for preferential extraction of a model of a certain subject when it is determined that the user has reached a position where the user can shoot a certain subject and is about to shoot that subject. This allows for support in shooting the subject the user is about to shoot.

[0092] As another modification, as in modification 2 of the first embodiment, the second model group may be extracted using the history of the current location as the history of the shooting environment information. By taking into account changes in the shooting location over time, a model more suitable for the subject that the user is about to shoot can be extracted. Specifically, for example, at a zoo or aquarium, models of subjects that are likely to be in a location where the user is heading may be extracted preferentially over subjects that are in a location to which the user has already traveled.

[0093] Also, based on the user's movement history, control may be performed to not extract a model suitable for a subject in a location to which the user has already moved.Furthermore, the user's movement route may be predicted based on the user's movement history, and a model suitable for a subject in a location along the predicted route may be preferentially extracted.

[0094] Furthermore, a second group of models may be extracted from the first group of models based on a history of information on the current location and a history of information on the current time, and the location information and time information associated with each model in the first group of models. This makes it possible to predict the user's movement route from a detailed history of the user's movement over time at a zoo or aquarium, for example, and extract models of subjects that will be present at the user's destination and that are suitable for the time when a show is to start.

[0095] Furthermore, the user's operation of the information processing device 100 may also be recorded as the imaging environment information, and the second model group may be extracted based on the history of the operation. Specifically, for example, when the user turns off the power of the information processing device 100, only the history from when the power is turned on to the present may be used, without considering the history from that point on. Alternatively, when the imaging direction is suddenly changed using information from an acceleration sensor (not shown) mounted on the information processing device 100, it may be assumed that an attempt is made to capture a new subject, and the usage ratio of the imaging environment information history may be reduced.

[0096] Alternatively, of the current location and the current time, only the current location may be considered. That is, the filming environment information acquisition unit 1502 may acquire information on the current location of the information processing device 100 as the filming environment information. In this case, the model group extraction unit 1503 may extract an appropriate second model group from the first model group based on the location information associated with each model in the first model group and the information on the current location acquired by the filming environment information acquisition unit 1502.

[0097] Alternatively, of the current location and the current time, only the current time may be considered. That is, the filming environment information acquisition unit 1502 may acquire information on the current time as filming environment information. In this case, the model group extraction unit 1503 may extract an appropriate second model group from the first model group based on time information associated with each model in the first model group and information on the current time acquired by the filming environment information acquisition unit 1502. For example, if the current time falls within the time period during which a sea lion show is being held, a sea lion recognition model may be extracted.

[0098] As described above, in the second embodiment, the current location and time at the time of shooting are acquired as shooting environment information instead of the scene recognition result by the scene recognition model. This makes it possible to extract a combination of multiple machine learning models suitable for the shooting environment and perform shooting without creating a scene recognition model.

[0099] The disclosure of this specification includes the following information processing device, information processing control method, and program. (Item 1) An information processing device, An acquisition means for acquiring shooting environment information; an extraction means for extracting a second group of models from a first group of models, each of which has been trained by machine learning, based on the shooting environment information; a recognition processing means for executing a recognition process using the second model group; an imaging means for executing imaging processing based on the recognition result of the recognition processing means; An information processing device comprising: (Item 2) further comprising a storage means for storing a history of the photographing environment information; 2. The information processing device according to item 1, wherein the extraction means extracts the second model group from the first model group based on a history of the shooting environment information. (Item 3) The information processing device described in item 1 or 2, characterized in that the acquisition means recognizes the shooting scene using a scene recognition model that is a trained machine learning model, and acquires the scene recognition result as the shooting environment information. (Item 4) further comprising a storage means for storing a history of the scene recognition results; 4. The information processing device according to item 3, wherein the extraction means extracts the second model group from the first model group based on a history of the scene recognition results. (Item 5) 5. The information processing device according to item 3 or 4, wherein the acquisition means recognizes a photographed scene from a live view image using the scene recognition model. (Item 6) 4. The information processing device according to item 3, wherein the acquisition means uses the scene recognition model to recognize a photographed scene from an image that has been photographed in the past. (Item 7) further comprising an adjustment unit for adjusting weight parameters of the scene recognition model; 5. The information processing device according to item 4, wherein the adjustment means adjusts weight parameters of the scene recognition model based on a history of the scene recognition results. (Item 8) 2. The information processing device according to item 1, wherein the acquisition means acquires information about a current location of the information processing device as the imaging environment information. (Item 9) Item 9. The information processing device described in item 8, wherein the extraction means extracts the second model group from the first model group based on location information associated with each model in the first model group and information on the current location. (Item 10) Further comprising a storage means for storing a history of the current location information, 10. The information processing device according to item 8 or 9, wherein the extraction means extracts the second model group from the first model group based on a history of the current location information and location information associated with each model in the first model group. (Item 11) the acquiring means further acquires information about a current time as the photographing environment information; 10. The information processing device described in item 8 or 9, characterized in that the extraction means extracts the second model group from the first model group based on information on the current location and the current time, and position information and time information associated with each model in the first model group. (Item 12) further comprising a storage means for storing a history of the information on the current location and a history of the information on the current time; The information processing device described in item 11, characterized in that the extraction means extracts the second model group from the first model group based on the history of the current location information and the history of the current time information, and the location information and time information associated with each model in the first model group. (Item 13) the acquisition means further acquires information about a shooting direction as the shooting environment information, Item 10. The information processing device described in item 8 or 9, characterized in that the extraction means extracts the second model group from the first model group based on the information on the current location, the information on the shooting direction, and position information associated with each model in the first model group. (Item 14) further comprising a storage means for storing a history of the information on the current position and a history of the information on the photographing direction; The information processing device described in item 13, characterized in that the extraction means extracts the second model group from the first model group based on the history of information on the current location and the history of information on the shooting direction, and the position information associated with each model in the first model group. (Item 15) the acquiring means further acquires, as the photographing environment information, a history of operations of the information processing device by a user; 15. The information processing device according to any one of items 1 to 14, wherein the extraction means extracts the second model group from the first model group further based on the operation history. (Item 16) further comprising an input means for accepting edits by a user on the attributes of the first model group; 16. The information processing device according to any one of items 1 to 15, wherein the extraction means extracts the second model group from the first model group edited by the input means. (Item 17) A control method for an information processing device, comprising: an acquisition step of acquiring shooting environment information; an extraction step of extracting a second group of models from a first group of models, each of which has been trained by machine learning, based on the shooting environment information; a recognition processing step of performing recognition processing using the second model group; an imaging process for performing imaging processing based on the recognition result in the recognition processing process; 1. A method for controlling an information processing device, comprising: (Item 18) A program for causing a computer to function as the information processing device according to any one of items 1 to 16.

[0100] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0101] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0102] 201: Shooting environment information acquisition unit, 202: Model group extraction unit, 203: Model group setting unit, 204: Recognition processing unit, 205: Shooting unit

Claims

1. An information processing device, An acquisition means for acquiring shooting environment information; an extraction means for extracting a second group of models from a first group of models, each of which has been trained by machine learning, based on the shooting environment information; a recognition processing means for executing a recognition process using the second model group; an imaging means for executing imaging processing based on the recognition result of the recognition processing means; An information processing device comprising:

2. further comprising a storage means for storing a history of the photographing environment information; 2. The information processing apparatus according to claim 1, wherein the extracting means extracts the second model group from the first model group based on a history of the image-capturing environment information.

3. The information processing apparatus according to claim 1 , wherein the acquisition means recognizes the photographed scene using a scene recognition model that is a trained machine learning model, and acquires a scene recognition result as the photographing environment information.

4. further comprising a storage means for storing a history of the scene recognition results; 4. The information processing apparatus according to claim 3, wherein the extracting means extracts the second model group from the first model group based on a history of the scene recognition results.

5. 4. The information processing apparatus according to claim 3, wherein the acquisition means recognizes the photographed scene from the live view image using the scene recognition model.

6. 4. The information processing apparatus according to claim 3, wherein the acquisition means recognizes a photographed scene from a previously photographed image by using the scene recognition model.

7. further comprising an adjustment unit for adjusting weight parameters of the scene recognition model; 5. The information processing apparatus according to claim 4, wherein the adjustment means adjusts weight parameters of the scene recognition model based on a history of the scene recognition results.

8. 2. The information processing apparatus according to claim 1, wherein the acquisition means acquires information on a current position of the information processing apparatus as the image-taking environment information.

9. 9. The information processing apparatus according to claim 8, wherein the extraction means extracts the second model group from the first model group based on location information associated with each model in the first model group and information about the current location.

10. Further comprising a storage means for storing a history of the current location information, 9. The information processing apparatus according to claim 8, wherein the extraction means extracts the second model group from the first model group based on a history of information on the current location and location information associated with each model in the first model group.

11. the acquiring means further acquires information about a current time as the photographing environment information; 9. The information processing apparatus according to claim 8, wherein the extraction means extracts the second model group from the first model group based on information on the current location and the current time, and position information and time information associated with each model in the first model group.

12. further comprising a storage means for storing a history of the information on the current location and a history of the information on the current time; 12. The information processing device according to claim 11, wherein the extraction means extracts the second model group from the first model group based on a history of information on the current location and a history of information on the current time, and position information and time information associated with each model in the first model group.

13. the acquisition means further acquires information about a shooting direction as the shooting environment information, 9. The information processing device according to claim 8, wherein the extraction means extracts the second model group from the first model group based on the information on the current position, the information on the shooting direction, and position information associated with each model in the first model group.

14. further comprising a storage means for storing a history of the information on the current position and a history of the information on the photographing direction; 14. The information processing device according to claim 13, wherein the extraction means extracts the second model group from the first model group based on a history of information on the current position and a history of information on the shooting direction, and position information associated with each model in the first model group.

15. the acquisition means further acquires, as the photographing environment information, a history of operations of the information processing device by a user; 2. The information processing apparatus according to claim 1, wherein the extracting means extracts the second model group from the first model group further based on the operation history.

16. further comprising an input means for accepting edits by a user on the attributes of the first model group; 2. The information processing apparatus according to claim 1, wherein said extracting means extracts said second model group from said first model group edited by said inputting means.

17. A control method for an information processing device, comprising: an acquisition step of acquiring shooting environment information; an extraction step of extracting a second group of models from a first group of models, each of which has been trained by machine learning, based on the shooting environment information; a recognition processing step of performing recognition processing using the second model group; an imaging process for performing imaging processing based on the recognition result in the recognition processing process; 1. A method for controlling an information processing device, comprising:

18. A program for causing a computer to function as the information processing device according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Trained model proposal system, trained model proposal method, and program

    JP7068745B2