Imaging apparatus and control method for the same
The imaging device addresses the lack of automatic scene description in conventional devices by employing a multimodal AI learning model to generate and record text prompts, thereby enhancing scene description capabilities and user convenience.
Patent Information
- Application Number
- JP2023192351
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2025-05-22
AI Technical Summary
Conventional imaging devices lack the capability to automatically record text information describing the shooting scene, limiting their functionality in providing detailed scene descriptions.
An imaging device equipped with an imaging means, a generation means using a multimodal AI learning model to convert image data into text prompts, and a recording means to store these prompts, allowing for automatic generation and recording of text information describing the captured scene.
Enables the imaging device to automatically record text information that effectively describes the photographed scene, enhancing user convenience and providing a means to easily share or review image capture details.
Smart Images

Figure 2025079583000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an imaging apparatus and a control method thereof. [Background technology]
[0002] Conventionally, imaging devices have provided a function for recording a captured scene as a still image or a video. When recording image data obtained by shooting in accordance with the data format described in Non-Patent Document 1, conventional imaging devices can record information about the state of the imaging device at the time of shooting, such as the shooting position and shooting parameters, in association with the image data. However, no imaging device has provided a function for automatically recording text information describing the captured scene. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] "CIPA DC-008-2023 Exif 3.0 Image File Format Standard for Digital Still Cameras", [online], Established in May 2023, Camera & Imaging Products Association, [Retrieved October 27, 2023], Internet<URL:https: / / www.cipa.jp / std / documents / download_j.html?DC-008-2023-J> Summary of the Invention [Problem to be solved by the invention]
[0004] In one embodiment, the present invention provides an imaging device capable of automatically recording text information describing a captured scene. [Means for solving the problem]
[0005] In one aspect, the present invention provides an imaging device comprising an imaging means, a generation means for generating a prompt, which is text information describing the shooting scene of the image represented by the image data, from image data acquired by the imaging means, and a recording means for recording the prompt. [Effects of the Invention]
[0006] According to the present invention, it is possible to provide an imaging device that can automatically record text information that describes a photographed scene. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 shows a configuration of a digital camera 100 according to an embodiment. [Figure 2] FIG. 1 shows the configuration of a prompt generation unit 102 according to an embodiment. [Figure 3] Flowchart showing the photographing process of the first embodiment [Figure 4] 1 is a flowchart illustrating a prompt generation process according to a first embodiment. [Figure 5] FIG. 10 is a diagram for explaining the photographing process of the first embodiment. [Figure 6] Flowchart showing the photographing process of the second embodiment [Figure 7] 10 is a flowchart illustrating a prompt generation process according to a second embodiment. [Figure 8] FIG. 10 is a diagram illustrating the photographing process of the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] The present invention will be described in detail below based on exemplary embodiments with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the claimed invention. Furthermore, although multiple features are described in the embodiments, not all of them are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0009] In the following embodiments, the present invention will be described with reference to a digital camera. However, the present invention can also be implemented in any electronic device with an imaging function. Such electronic devices include video cameras, computer devices (personal computers, tablet computers, media players, PDAs, etc.), mobile phones, smartphones, game consoles, robots, drones, and drive recorders. These are merely examples, and the present invention can also be implemented in other electronic devices.
[0010] <First embodiment> <Example of digital camera functional configuration> FIG. 1A is a block diagram showing an example of the functional configuration of a digital camera 100 as an imaging device according to a first embodiment. Throughout the drawings, functional blocks, except for those parts that can clearly only be implemented by hardware (e.g., a lens, an image sensor, a recording medium, etc.), can be implemented by software or a combination of software and hardware. For example, a functional block may be implemented by dedicated hardware such as an ASIC. A functional block may also be implemented by one or more processors capable of executing programs, such as a CPU, executing a program stored in memory. Multiple functional blocks may also be implemented by a common configuration (e.g., a single ASIC). Hardware that implements part of the functions of one functional block may also be included in hardware that implements another functional block.
[0011] 1(A) shows only an imaging unit 101, a prompt generation unit 102, and a recording unit 103 as functional blocks of the digital camera 100 for ease of explanation and understanding of the embodiment. However, the functions of the digital camera 100 are not limited to those realized by the functional blocks shown in FIG. 1(A).
[0012] The imaging unit 101 acquires RAW image data corresponding to an optical image of a subject using a lens, an imaging element, etc. The imaging unit 101 also applies predetermined image processing to the RAW image data to generate image data according to the intended use. Here, the intended use may be, for example, recording, display, or prompt generation. Note that the image data for prompt generation may be generated by reusing image data for recording or display, or based on image data for recording or display.
[0013] The imaging unit 101 may also supply part of the auxiliary information to be recorded in a data file that stores image data for recording (for example, one or more tags related to the shooting conditions and shooting situations in the auxiliary information shown in Tables 8 and 9 of the non-patent document) to the prompt generation unit 102. The imaging unit 101 can also supply any information that can be acquired by the imaging unit 101 to the prompt generation unit 102, such as information related to the characteristics of the imaging element and evaluation values used for exposure control.
[0014] The prompt generating unit 102 generates text information (prompt) that explains the photographed scene represented by the image data from the image data supplied from the imaging unit 101 and various information.
[0015] 2 is a block diagram showing an example of the functional configuration of the prompt generation unit 102. The prompt generation unit 102 has at least an image / prompt conversion unit 201 and a prompt editing unit 202.
[0016] The image / prompt conversion unit 201 uses a multimodal AI learning model to convert image data into text information describing the captured scene of the image represented by the image data. The multimodal AI learning model may be stored in advance in the digital camera 100, or may reside in an external device with which the digital camera 100 can communicate. The multimodal AI learning model in this embodiment is a neural network trained using image data and text data such as captions and tags about the captured scene associated with the image data.
[0017] A multimodal AI learning model outputs text data for input image data. The text data is text data that describes the scene in which the image represented by the image data was captured. Such a multimodal AI learning model can be realized using publicly known techniques, such as those described in the following literature: Lili Yu and 25 others, "Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning," July 14, 2023, Meta Research website, Internet <URL:https: / / ai.meta.com / research / publications / scaling-autoregressive-multi-modal-models-pretraining-and-instruction-tuning>
[0018] Note that image / prompt conversion unit 201 may obtain the prompt from an external device with which digital camera 100 can communicate. In this case, image / prompt conversion unit 201 transmits image data for generating the prompt (and other information as necessary) to the external device. Image / prompt conversion unit 201 then receives the prompt generated by the external device for the image data.
[0019] The image / prompt conversion unit 201 acquires a prompt for each of a plurality of frames of image data and stores the prompt in the storage unit 203 .
[0020] The prompt editing unit 202 generates a final prompt (text data) based on the multiple prompts stored in the storage unit 203 and outputs the generated prompt to the recording unit 103. The prompt editing unit 202 generates one prompt from the multiple prompts by appropriately using the attached information.
[0021] The recording unit 103 records the prompt output by the prompt generation unit 102 in a recording destination. The recording destination may be a recording medium or storage device included in the digital camera 100, or may be an external device with which the digital camera 100 can communicate.
[0022] <Example of hardware configuration for a digital camera> 1B is a block diagram showing an example of the hardware configuration of the digital camera 100. Each functional block is connected to each other via a system bus 111 so that they can communicate with each other.
[0023] The CPU (Central Processing Unit) 112 and the GPU (Graphics Processing Unit) 116 are each one or more processors capable of executing programs. The GPU 116 is configured to be able to execute specific operations faster than the CPU 112, and in recent years has been often used to execute inference processing using neural networks at high speed. Instead of the GPU 116, an NPU (Neural Processing Unit) that is more specialized for executing learning and inference processing using neural networks may be used. Note that instead of using separate CPU 112 and GPU 116, an SoC (System on Chip) that integrates a CPU and a GPU (and optionally an NPU) may be used.
[0024] 1 and 2 by loading a program stored in, for example, ROM 113 into RAM 114 and executing it. CPU 112 realizes high-speed processing by using GPU 116 for processing using a neural network. Note that GPU 116 may have a dedicated RAM separate from RAM 114.
[0025] The ROM 113 is, for example, an electrically rewritable nonvolatile memory, and stores programs executable by the CPU 112, setting values, GUI data, parameters for realizing a trained neural network (training model), and the like.
[0026] The RAM 114 is used to load programs executed by the CPU 112 and to store values required during program execution.
[0027] The recording medium 115 is, for example, a semiconductor memory card or a solid-state drive (SSD), and is used as a recording destination for image data obtained by imaging. The recording medium 115 is also used as a recording destination for prompts generated by the prompt generation unit 102. When both a prompt and the image data on which the prompt is based are recorded on the recording medium 115, the two may be associated with each other. For example, a prompt generated using image data may be recorded as metadata recorded in a data file that stores the image data. When a prompt is generated based on multiple pieces of image data, information on all of the image data used to generate the prompt (for example, data file names) may be recorded together with the prompt.
[0028] The input device 117 includes a plurality of operation members such as buttons, dials, switches, and touch panels that accept operation inputs to the digital camera 100. The input device 117 may also include one or more devices (such as sensors) for acquiring additional information at the time of shooting.
[0029] for example, a GPS receiver for acquiring location information of the digital camera 100; - Clock to get the shooting date and time, A thermometer to measure the temperature of the shooting environment. A sensor (such as a gyro sensor or an acceleration sensor) that detects the magnitude and direction of movement of the digital camera 100; A microphone to capture the sound of the shooting environment, The input devices 117 may include, but are not limited to, the following.
[0030] The imaging device 118 includes, for example, an optical system unit such as a lens, an aperture, and a shutter, and an imaging element. The optical system unit may have a compound lens or a multi-lens system. The optical system unit may also be capable of changing optical characteristics such as zoom and aperture (for example, depending on the image content to be acquired). The imaging element may be, for example, a CMOS color image sensor with a primary color Bayer array color filter.
[0031] Display device 119 is, for example, a liquid crystal display provided on the surface of the housing of digital camera 100. Display device 119 may also be a touch display. The display device displays live view images or images read from recording medium 115, menu screens, and information about digital camera 100 (for example, settings, remaining battery power, remaining number of shots that can be taken, etc.).
[0032] The communication interface 120 is a circuit for communicating with an external device in accordance with one or more communication standards. It includes a connector for wired communication, an antenna for wireless communication, a transmitting / receiving circuit, and the like. The digital camera 100 can transmit image data to an external device and receive data from an external device via the communication interface 120. Typical communication standards that the communication interface 120 conforms to include, but are not limited to, HDMI (registered trademark), USB, Bluetooth (registered gazette), and wireless LAN (Wi-Fi).
[0033] The functional blocks shown in Figures 1(A) and 2 are realized by one or more pieces of hardware shown in Figure 1(B). For example, the imaging unit 101 is mainly realized by the CPU 112 and the imaging device 118. The prompt generation unit 102 is mainly realized by the CPU 112 and the GPU 116. The RAM 114 is used as a temporary storage location for data to be processed, data being processed, data resulting from processing, etc., and the ROM 113 is used for various processes as a reference destination for pre-stored setting values, programs, etc.
[0034] <Operation during shooting processing> The operation of digital camera 100 for capturing a still image will be described with reference to Figures 1 to 5. The operation of each step in the flowcharts of Figures 3 to 5 is realized by CPU 112 or GPU 116 executing a program stored in ROM 113 and controlling other hardware as necessary.
[0035] Furthermore, still image capture by digital camera 100 may be performed in response to an instruction via input device 117, or may be performed according to a predetermined condition other than an instruction. For example, still image capture may be performed continuously at regular time intervals, or still image capture may be performed when information obtained from a video being captured for live view display satisfies a predetermined condition.
[0036] The exposure conditions for still image capture can be determined based on brightness information obtained by the CPU 112 from a moving image being captured for live view display, for example. Similarly, the focal length of the lens for still image capture can be controlled based on contrast information obtained by the CPU 112 from a moving image being captured for live view display, for example.
[0037] In S301, the imaging unit 101 captures a still image and acquires image data. The imaging unit 101 may also capture a moving image and use frame images of the moving image as still image data. The imaging unit 101 outputs the acquired image data and the above-mentioned attached information to the prompt generation unit 102. The imaging unit 101 outputs image data after so-called development processing. In the image data after development processing, pixel data constituting the image data has three components (RGB or YCbCr). The imaging unit 101 may process the image data so that it is suitable for use by the prompt generation unit 102, for example, before outputting it.
[0038] When recording image data in addition to the prompt, the imaging unit 101 also outputs the image data and the associated information to the recording unit 103. The imaging unit 101 may output to the recording unit 103, for example, image data that has been subjected to processing (for example, encoding processing) according to the data format to be recorded in the recording unit 103.
[0039] In S302, the prompt generating unit 102 generates, from the image data output by the imaging unit 101, text information (prompt) that explains the captured scene of the image represented by one or more frames of image data.
[0040] The operation of prompt generator 102 will now be further described with reference to FIG. 2 and the flowchart shown in FIG. In S401, the image / prompt conversion unit 201 generates text information (prompt) that explains the captured scene of the image represented by the input image data from the input image data. The image / prompt conversion unit 201 stores the generated prompt in the storage unit 203 (RAM 114).
[0041] As described above, image / prompt conversion unit 201 can obtain a prompt by inputting image data into a learning model stored in ROM 113. Alternatively, image / prompt conversion unit 201 may transmit image data to an external device via communication interface 120 and receive a prompt from the external device.
[0042] The level of detail of the prompt generated by the image / prompt conversion unit 201 can be changed depending on the settings. For example, when the level of detail is set low, a prompt that describes only the gender of the human subject can be generated, and when the level of detail is set high, a prompt that describes the gender, age, hair color, length, etc. can be generated.
[0043] Furthermore, image / prompt conversion unit 201 may generate a prompt that includes not only a description about an element included in the image (positive prompt), but also a description about an element not included in the image (negative prompt).
[0044] 5(A) and 5(B) show examples of images represented by image data and examples of prompts generated from the image data. The image / prompt conversion unit 201 stores the prompts generated for each image data in the storage unit 203 (RAM 114).
[0045] The image / prompt conversion unit 201 generates a prompt for each frame of image data supplied from the image capture unit 101 as a prompt generation target, and stores the generated prompt in the storage unit 203 .
[0046] In S402 , the prompt editing unit 202 applies a predetermined editing process to the prompt stored in the RAM 114 to generate a final prompt, and outputs the generated prompt to the recording unit 103 .
[0047] The editing process is the operation of generating the final prompt from additional information at the time of shooting of the image data that served as the source of the prompt, prompts generated for other highly related image data, and additional information at the time of shooting.
[0048] Examples of editing processes include: (1) For one prompt, add additional information or information based on additional information at the time of capturing the image data that is the source of the prompt. (2) Generate a single prompt from the prompts generated for multiple frames of image data that are highly related to each other, taking into account additional information at the time of capture as necessary. These include, but are not limited to:
[0049] For example, the user may be able to set whether or not to perform editing processing and what type of editing processing to perform. For example, editing with the setting "every shot" corresponds to (1), and editing with the setting "every event" corresponds to (2). Other settings, such as "every 10 minutes," may also be possible. If editing processing is not performed, prompt editing unit 202 outputs the prompt generated by image / prompt conversion unit 201 to recording unit 103 as is. Note that even if editing processing is performed, an unedited prompt may be output to recording unit 103.
[0050] Highly relevant image data includes, for example, (a) image data in which the difference in at least one of the photographing date and time and the photographing location is equal to or less than a predetermined threshold; (b) Image data that shows the same or similar subjects; (c) image data in which the correlation between images is equal to or greater than a threshold; (d) image data containing the same keywords in the generated prompt; However, the image data satisfying a combination of two or more of the conditions (a) to (d) may be regarded as highly relevant image data.
[0051] For example, when editing is set to "by event," the prompt editing unit 202 can determine that image data that satisfies (a), both (a) and (b), or (a), (b), and (d) is highly relevant image data.
[0052] Here, it is assumed that editing is set to be performed for each event or at regular intervals, and that the image data of two frames shown in Figures 5(A) and 5(C), which were taken close in date and time and contain the same subject 501, are considered to be highly related image data.
[0053] In this case, the prompt editing unit 202 applies editing processing to prompt X (Figure 5(C)) and prompt Y (Figure 5(D)) generated for each image data, and generates one prompt (Figure 5(E)).
[0054] In the example shown in Figure 5, the prompt editing unit 202 generates a final prompt by applying an editing process that combines the contents of prompts X and Y taking into account the context of the shooting dates and times (shooting order) and an editing process that adds additional information related to the shooting dates and times. When synthesizing prompts, the prompt editing unit 202 can add particles, conjunctions, etc. as needed, for example, by using existing composition generation AI technology, to generate prompts with natural sentence expressions. Furthermore, when synthesizing multiple prompts, the expression of the synthesized prompt may be changed taking into account the frequency of word appearance, such as by emphasizing frequently used words in the synthesized prompt.
[0055] 3, in S303, the recording unit 103 records the prompt created by the prompt generation unit 102 in S302 on the recording medium 115. The recording unit 103 may record only the prompt, or may record the prompt in association with the image data used to generate the prompt. When the prompt is generated based on prompts and additional information for multiple frames of image data, the same prompt may be associated with each frame of image data, or the prompt may be associated only with representative image data.
[0056] The method of associating the prompt and the image data may be any known method, for example, the prompt and the image data may be included in the same file container, or the prompt and the image data may be recorded as separate files with a common file name.
[0057] The recording unit 103 may perform digital authentication processing on the prompt or image data used to create the prompt before recording. The digital authentication processing may be any processing intended to guarantee the content of the prompt or image data. For example, it may be a processing to grant an NFT (non-fungible token).
[0058] As described above, according to this embodiment, an imaging device that conventionally only provides an image capturing function can be provided with a function for generating text data (prompts) that describe a captured scene. This makes it easy to provide information about a captured scene to a third party in text form, for example. Furthermore, the generated prompts can be input to an image generation AI system and used to generate images of similar scenes. Furthermore, when reviewing images later, the prompts make it easier to recall the circumstances at the time of capture, improving convenience.
[0059] <Second embodiment> Next, a second embodiment of the present invention will be described. This embodiment can be implemented by the digital camera 100 described in the first embodiment, so a description of the contents described in the first embodiment will be omitted.
[0060] The second embodiment relates to the operation of the digital camera 100 when the image data obtained by the imaging unit 101 capturing an image under the currently set exposure conditions is not suitable for generating a prompt in the image / prompt conversion unit 201.
[0061] An example will be described using the photographic scene shown in FIG. 8(A). FIG. 8(A) shows an example of a scene in which a dog jumps to catch a ball thrown by a boy. The photographer is attempting to take a panning shot in which the dog is positioned at the center of the image and captures the moment the dog catches the ball in mid-air. The photographer has set the lens focal length and shutter speed as photographic conditions suitable for composition and panning, as shown in FIG. 8(C).
[0062] It is also assumed that the image data shown in Fig. 8(B) is obtained by photographing, and that the prompt shown in Fig. 8(D) is obtained by image / prompt conversion unit 201 from the image data shown in Fig. 8(B).
[0063] However, the prompt only describes "dog," "jump and catch," and "object," and does not include information about the "boy" or the "ball." This is because the "boy" does not appear in the image at the angle of view intended by the photographer, and panning produces blurred images of subjects other than the dog, making it impossible to identify the "ball" from the image data.
[0064] In this way, The photographer closes up part of the scene (the angle of view is narrower than the threshold), The blur range becomes larger (when using panning mode or an aperture value close to the maximum aperture (smaller than the threshold)), Camera shake is likely to occur (a shutter speed below the threshold is set), Noise is likely to increase (the shooting sensitivity is set above the threshold), - Underexposed or overexposed by more than the threshold for exposure conditions that result in proper exposure In these shooting conditions, the prompt generated from the captured image may contain less information, and the usefulness of the prompt may be reduced. Note that these are merely examples.
[0065] In this embodiment, if it is determined that the image data obtained under the current shooting conditions is not suitable for generating a prompt with sufficient information, the shooting conditions are changed to ones that are more likely to produce an image that is more suitable for generating the prompt.
[0066] The still image shooting process of digital camera 100 of this embodiment will be described below with reference to Figures 6 to 8. In Figures 6 and 7, steps that perform the same operations as in the first embodiment are given the same reference numerals as in Figures 3 and 4, and descriptions thereof will be omitted.
[0067] In S601, the image capturing unit 101 acquires the current shooting conditions. The shooting conditions acquired here are not limited to exposure conditions (particularly aperture value and shutter speed), but may also include the shooting mode and lens angle of view. In the example shown in Fig. 8, the lens focal length (angle of view) and shutter speed are used as the shooting conditions.
[0068] In S602, the image capturing unit 101 (CPU 112) determines whether image data obtained by capturing an image under the current capturing conditions is likely to be unsuitable for generating a prompt. This determination may be made by comparing a threshold value determined in advance for each capturing condition item with the current setting value. Note that the threshold value may change dynamically. For example, the shutter speed threshold value may have a value corresponding to the current focal length of the lens.
[0069] If the current shooting conditions include even one item that is not suitable for generating a prompt, the CPU 112 determines that the image data obtained by shooting under the current shooting conditions is likely to be unsuitable for generating a prompt. If the CPU 112 determines that the image data obtained by shooting under the current shooting conditions is likely to be unsuitable for generating a prompt, it executes S603; otherwise, it executes S607.
[0070] In S603, the imaging unit 101 performs imaging using imaging conditions (FIG. 8(E)) that can obtain image data suitable for generating a prompt, instead of the current imaging conditions (FIG. 8(C)). The imaging conditions that can obtain image data suitable for generating a prompt can be stored in advance in the ROM 113. The imaging conditions that can obtain image data suitable for generating a prompt are imaging conditions that can obtain a wide-angle, deep-focus image. Here, it is assumed that the imaging conditions shown in FIG. 8(E) were used to obtain the image data Z shown in FIG. 8(F).
[0071] In S604, the image capturing unit 101 stores in the RAM 114 the image data Z (FIG. 8F) acquired using the changed photographing conditions and the original photographing conditions before the change (FIG. 8C) in association with each other.
[0072] In S605, the prompt generation unit 102 generates a prompt to be recorded using the image data Z stored in RAM 114 in S604, the original shooting conditions, and the additional information. The prompt generation unit 102 outputs the generated prompt to the recording unit 103.
[0073] The operation of the prompt generating unit 102 in S605 of FIG. 6 will be described with reference to the flowchart shown in FIG. S401 is the same as that described in the first embodiment, and therefore will not be described again. As a result of the processing, it is assumed that the prompt Z shown in Fig. 8(G) is obtained. The prompt Z, the original shooting conditions, and the additional information are stored in the storage unit 203 (RAM 114).
[0074] In S701, the prompt editing unit 202 applies editing processing using the original shooting conditions (FIG. 8C) to the prompt Z to generate a prompt to be recorded. The prompt editing unit 202 outputs the generated prompt to the recording unit 103.
[0075] In the editing process, the prompt editing unit 202 applies the editing process described in S402 to the prompt Z (FIG. 8(G)) generated from the image data. However, as shown in FIG. 8(H), the shooting conditions to be added to the prompt are not the shooting conditions actually used but the shooting conditions that were originally set.
[0076] S301 and S303 are the same as those described in the first embodiment, and therefore a description thereof will be omitted.
[0077] In this manner, in this embodiment, if it is determined that the image data obtained under the current shooting conditions is not suitable for generating a prompt, shooting is performed under shooting conditions that will obtain image data suitable for generating a prompt. On the other hand, when adding shooting conditions to the prompt, the shooting conditions that were originally set before the change are added, so that the shooting conditions intended by the photographer are reflected in the prompt.
[0078] Therefore, even if the shooting conditions are set such that there is a high possibility that image data that is not suitable for generating a prompt will be obtained, it is possible to generate a prompt that includes appropriate types of information and reflects the shooting conditions that the photographer intended.
[0079] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0080] The disclosure of this embodiment includes the following imaging device, a control method thereof, and a program. (Item 1) An imaging means; a generation means for generating a prompt, which is text information explaining a shooting scene of an image represented by the image data, from the image data acquired by the imaging means; a recording means for recording said prompt; An imaging device comprising: (Item 2) 2. The imaging device according to item 1, wherein the generating means includes adding information about when the image data was captured to the prompt. (Item 3) 3. The imaging device according to item 2, wherein the information includes one or more of the exposure conditions at the time of shooting, location information, temperature, and magnitude and direction of movement of the imaging device. (Item 4) The imaging device according to any one of items 1 to 3, characterized in that the generation means generates one prompt from the prompts generated for each of multiple frames of image data acquired by the imaging means. (Item 5) 5. The imaging device according to item 4, wherein the generating means generates the single prompt by combining the prompts generated for each of the plurality of frames of image data, taking into account the order in which the image data are captured. (Item 6) The plurality of frames of image data are (a) image data in which the difference in at least one of the photographing date and time and the photographing location is equal to or less than a predetermined threshold; (b) Image data that shows the same or similar subjects; (c) image data in which the correlation between images is equal to or greater than a threshold; (d) image data containing the same keywords in the generated prompt; 6. The imaging device according to item 4 or 5, characterized in that: (Item 7) When it is determined that image data acquired according to current photographing conditions is not suitable for generating the prompt by the generating means, the imaging means acquires the image data according to predetermined photographing conditions; 7. The imaging device according to any one of items 1 to 6, wherein the generating means, when including shooting conditions in the prompt, includes the current shooting conditions. (Item 8) 8. The imaging device according to item 7, wherein the imaging means makes the determination based on any one of the angle of view, the imaging mode, the aperture value, the shutter speed, and the imaging sensitivity among the current imaging conditions. (Item 9) 9. The imaging device according to any one of items 1 to 8, wherein the recording means records the prompt in association with the image data. (Item 10) 10. The imaging device according to any one of items 1 to 9, wherein the generation means generates the prompt using a multimodal AI learning model. (Item 11) 10. The imaging device according to any one of items 1 to 9, wherein the generating means transmits the image data to an external device and acquires a prompt for the image data from the external device. (Item 12) A control method executed by an imaging device, comprising: generating a prompt, which is text information explaining a captured scene of an image represented by the image data, from the image data acquired by the imaging means; recording the prompt; 10. A method for controlling an imaging device, comprising: (Item 13) 12. A program for causing a computer to function as the generating means and the recording means of the imaging device according to any one of items 1 to 11.
[0081] The present invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Therefore, the following claims are appended to clarify the scope of the invention. [Explanation of symbols]
[0082] 100... digital camera 100, 101... imaging unit, 102... prompt generation unit, 103... recording unit
Claims
1. An imaging means; a generation means for generating a prompt, which is text information explaining a shooting scene of an image represented by the image data, from the image data acquired by the imaging means; a recording means for recording said prompt; An imaging device comprising:
2. 2. The imaging device according to claim 1, wherein the generating means adds information about when the image data was captured to the prompt.
3. 3. The imaging device according to claim 2, wherein the information includes one or more of exposure conditions at the time of shooting, position information, temperature, and magnitude and direction of movement of the imaging device.
4. 2. The imaging device according to claim 1, wherein the generating means generates one prompt from the prompts generated for each of a plurality of frames of image data acquired by the imaging means.
5. 5. The imaging device according to claim 4, wherein the generating means generates the single prompt by combining the prompts generated for each of the plurality of frames of image data, taking into account the order in which the image data are captured.
6. The plurality of frames of image data are (a) image data in which a difference in at least one of the photographing date and time and the photographing position is equal to or less than a predetermined threshold; (b) Image data showing the same or similar subjects; (c) image data in which the correlation between images is equal to or greater than a threshold; (d) image data containing the same keyword in the generated prompt; 5. The imaging device according to claim 4, wherein the imaging device is any one of the above.
7. When it is determined that image data acquired according to current photographing conditions is not suitable for generating the prompt by the generating means, the imaging means acquires the image data according to predetermined photographing conditions; The imaging device according to claim 1 , wherein the generating means, when including shooting conditions in the prompt, includes the current shooting conditions in the prompt.
8. 8. The image pickup apparatus according to claim 7, wherein the image pickup means makes the determination based on any one of an angle of view, an image pickup mode, an aperture value, a shutter speed, and an image pickup sensitivity, among the current image pickup conditions.
9. 2. The imaging device according to claim 1, wherein the recording means records the prompt in association with the image data.
10. The imaging device according to claim 1 , wherein the generating means generates the prompt using a multimodal AI learning model.
11. The imaging device according to claim 1 , wherein the generating means transmits the image data to an external device and obtains a prompt for the image data from the external device.
12. A control method executed by an imaging device, comprising: generating a prompt, which is text information explaining a shooting scene of an image represented by the image data acquired by an imaging means, from the image data acquired by the imaging means; recording said prompt; 13. A method for controlling an imaging apparatus comprising:
13. A program for causing a computer to function as the generating means and the recording means of the imaging device according to claim 1 .
Citation Information
Cited By
Imaging apparatus and control method for same
EP4808108A1
Imaging apparatus and control method for same
WO2025100249A1