Imaging apparatus and control method thereof

By introducing a multimodal AI learning model into the camera device to generate and record text information of the shooting scene, the problem that existing camera devices cannot automatically record text of the shooting scene is solved, and convenient information recording and management are realized.

CN122122889APending Publication Date: 2026-05-29CANON KK

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CANON KK
Filing Date
2024-10-24
Publication Date
2026-05-29

Smart Images

  • Figure CN122122889A_ABST
    Figure CN122122889A_ABST
Patent Text Reader

Abstract

Disclosed is an image pickup apparatus capable of automatically recording text information for describing a shooting scene. The image pickup apparatus generates a cue from image data obtained by an image pickup section, the cue being text information for describing a shooting scene of an image represented by the image data. The image pickup apparatus can record the cue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to camera equipment and its control method. Background Technology

[0002] Typically, camera devices provide the function of recording a scene as still images and moving images. When recording image data obtained through photography according to the data format described in Non-Patent Document 1, conventional camera devices can associate information related to the state of the camera device at the time of photography, such as the shooting position and shooting parameters, with the image data. However, there is currently no camera device that provides the function of automatically recording text information describing the shooting scene. Existing technical documents Non-patent literature

[0003] Non-Patent Document 1: "CIPA DC-008-2023 Digital Stereo Camera Image Photo Frame Specification Exif 3.0", [Online], formulated in May 2023, Camera & Imaging Products Association, [Searched on October 27, 2023], Internet<URL: https: / / www.cipa.jp / std / documents / download_j.html?DC-008-2023-J> Summary of the Invention The problem the invention aims to solve

[0004] The present invention provides, in its embodiments, a camera device capable of automatically recording text information used to describe the shooting scene. Solution for solving the problem

[0005] The present invention provides a camera device, characterized in that it comprises: a camera component; a generating component for generating a prompt based on image data obtained by the camera component, the prompt being text information describing the shooting scene of the image represented by the image data; and a recording component for recording the prompt. Advantages of the invention

[0006] According to the present invention, a camera device capable of automatically recording text information describing the shooting scene can be provided.

[0007] Other features and advantages of the invention will become apparent from the following description taken in conjunction with the accompanying drawings. Note that throughout the drawings, the same reference numerals denote the same or similar components. Attached Figure Description

[0008] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the specification, serve to explain the principles of the invention.

[0009] Figure 1A This is a diagram illustrating an exemplary configuration of a digital camera 100 according to an embodiment. Figure 1B This is a diagram illustrating an exemplary configuration of a digital camera 100 according to an embodiment. Figure 2 This is a diagram illustrating an exemplary configuration of the prompt generation unit 102 of an embodiment. Figure 3 This is a flowchart illustrating the shooting process of the first embodiment. Figure 4 This is a flowchart illustrating the prompt generation process of the first embodiment. Figure 5 This is a diagram used to describe the image processing of the first embodiment. Figure 6 This is a flowchart illustrating the shooting process of the second embodiment. Figure 7 This is a flowchart illustrating the prompt generation process of the second embodiment. Figure 8 This is a diagram used to describe the shooting process of the second embodiment. Detailed Implementation

[0010] In the following, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments are not intended to limit the scope of the claimed invention. Multiple features are described in the embodiments, but this does not limit the invention to requiring all such features, and multiple such features can be appropriately combined. Furthermore, in the drawings, the same reference numerals are given the same or similar configuration, and redundant descriptions are omitted.

[0011] Note that the following embodiments describe the invention in the context of a digital camera. However, the invention can also be embodied in any electronic device with camera functionality. Such electronic devices include video cameras, computer devices (personal computers, tablet computers, media players, or PDAs, etc.), mobile phone devices, smartphones, gaming devices, robots, drones, and dashcams. These are examples, and the invention can also be embodied in other electronic devices.

[0012] ·<First Embodiment> <Exemplary Functional Configuration of a Digital Camera> Figure 1AThis is a block diagram illustrating an exemplary functional configuration of a digital camera 100 as an imaging device according to a first embodiment. Throughout the figures, functional blocks can be embodied by software, or a combination of software and hardware (except for portions that can obviously only be implemented by hardware, such as lenses, image sensors, and recording media). For example, functional blocks can be implemented by dedicated hardware such as an ASIC. Furthermore, functional blocks can be implemented by one or more processors, such as a CPU, capable of executing programs stored in memory. Note that multiple functional blocks can be implemented using the same configuration (e.g., an ASIC). Additionally, hardware implementing a portion of the functionality of a certain functional block can be included in the hardware implementing other functional blocks.

[0013] For ease of description and understanding of the embodiments, Figure 1A Only the imaging unit 101, the prompt generation unit 102, and the recording unit 103, which are functional blocks of the digital camera 100, are shown. However, the functions of the digital camera 100 are not limited to those shown. Figure 1A The function block shown in the figure implements the function.

[0014] The camera unit 101 uses, for example, a lens and an image sensor to acquire RAW image data corresponding to the optical image of the subject. Furthermore, the camera unit 101 generates image data suitable for its intended purpose by applying pre-set image processing to the RAW image data. Here, the intended purpose may be, for example, recording, display, or prompting. Note that image data used for recording or display can be used as image data for prompting, or image data for prompting can be generated based on image data used for recording or display.

[0015] Furthermore, the imaging unit 101 may supply the prompt generation unit 102 with a portion of the attribute information recorded in a data file storing image data for recording (e.g., one or more labels related to the photographing conditions and shooting situation from the attribute information indicated by Tables 8 and 9 of non-patent documents). Additionally, any information obtainable by the imaging unit 101, such as information related to the characteristics of the image sensor and evaluation values ​​used in exposure control, may be supplied to the prompt generation unit 102.

[0016] The prompt generation unit 102 generates text information (prompts) describing the shooting scene indicated by the image data based on the image data supplied from the camera unit 101 and various types of information.

[0017] Figure 2 This is a block diagram illustrating an exemplary functional configuration of the prompt generation unit 102. The prompt generation unit 102 includes at least an image / prompt conversion unit 201 and a prompt editing unit 202.

[0018] The image / hint conversion unit 201 uses a multimodal AI learning model to convert image data into text information describing the shooting scene of the image represented by the image data. The multimodal AI learning model can be pre-stored, for example, in the digital camera 100, or it can exist in an external device that can communicate with the digital camera 100. According to this embodiment, the multimodal AI learning model is a neural network that has been trained using image data and text data such as captions and labels related to the shooting scene associated with the image data.

[0019] The multimodal AI learning model outputs text data corresponding to the input image data. The text data describes the scene in which the image data represents being captured. This multimodal AI learning model can be implemented using known techniques, such as those described in the following document: Lili Yu et al., “Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning,” July 14, 2023, MetaResearch website, Internet. <URL: https: / / ai.meta.com / research / publications / scaling-autoregressive-multi-modal-models-pretraining-and-instruction-tuning> .

[0020] Note that the image / cue conversion unit 201 can receive cues from an external device that can communicate with the digital camera 100. In this case, the image / cue conversion unit 201 sends the image data (and other information as needed) used to generate the cues to the external device. Then, the image / cue conversion unit 201 receives the cues that the external device has already generated for the image data.

[0021] The image / cue conversion unit 201 obtains cues for multiple image data corresponding to multiple images (frames) respectively, and stores the cues in the storage unit 203.

[0022] The prompt editing unit 202 generates a final prompt (text data) based on multiple prompts stored in the storage unit 203, and outputs the final prompt to the recording unit 103. The prompt editing unit 202 appropriately utilizes attribute information and generates a prompt based on multiple prompts.

[0023] The recording unit 103 records the prompts output from the prompt generation unit 102 to the recording destination. The recording destination may be a recording medium or storage device included in the digital camera 100, or it may be an external device that can communicate with the digital camera 100.

[0024] <Example Hardware Configuration of Digital Cameras> Figure 1B This is a block diagram illustrating an exemplary hardware configuration of a digital camera 100. The functional blocks are connected in a manner that allows them to communicate with each other via system bus 111.

[0025] Each of the CPU (Central Processing Unit) 112 and the GPU (Graphics Processing Unit) 116 is one or more processors capable of executing programs. The GPU 116 is configured to perform specific computations at a higher speed than the CPU 112 and has been commonly used in recent years, particularly for high-speed execution of inference processing using neural networks. Instead of the GPU 116, a more specialized NPU (Neural Processing Unit) can be used for the execution of learning and inference processing using neural networks. Note that instead of using the CPU 112 and GPU 116, which are distinct from each other, a SoC (System-on-a-Chip) can be used, which integrates the CPU and GPU (and in some cases further, the NPU).

[0026] The CPU 112 implements the functions shown in Figure 1 by, for example, reading a program stored in ROM 113 into RAM 114 and executing the program. Figure 2 The various functional blocks described herein. The CPU 112 achieves high-speed processing by utilizing the GPU 116 in processing that uses neural networks. Note that the GPU 116 may include dedicated RAM, unlike the RAM 114.

[0027] ROM 113 is, for example, an electrically rewritable non-volatile memory. ROM 113 stores programs that can be executed by CPU 112, settings, GUI data, parameters used to implement a trained neural network (learning model), etc.

[0028] RAM 114 is used to read the program to be executed by CPU 112 and store the values ​​required during program execution.

[0029] Recording medium 115 is, for example, a semiconductor memory card or an SSD (solid-state drive), and serves as the recording destination for image data acquired through imaging. Furthermore, recording medium 115 also serves as the recording destination for prompts generated by prompt generation unit 102. When both the prompt and the image data serving as the source of the prompt are recorded in recording medium 115, the prompt and the image data can be associated with each other. For example, the prompt generated using image data can be recorded as metadata in a data file storing the image data. When a prompt is generated based on multiple pieces of image data, information about all the multiple pieces of image data used to generate the prompt (e.g., data filenames) can be recorded along with the prompt.

[0030] Input device 117 includes multiple operating components such as buttons, dials, switches, and touch panels that accept operational input to the digital camera 100. Furthermore, input device 117 may include one or more devices (e.g., sensors) for acquiring additional information during image capture.

[0031] For example, input device 117 may include, but is not limited to: • A GPS receiver used to obtain the location information of the digital camera 100 • A clock used to obtain the shooting date and time. • A thermometer used to measure the air temperature of the shooting environment. • Sensors (gyroscope or accelerometer, etc.) used to detect the magnitude and direction of motion of the digital camera 100, and • A microphone used to capture sound from the shooting environment.

[0032] The imaging device 118 includes an optical system unit such as a lens, aperture, and shutter, as well as an image sensor. The optical system unit may include a stereo lens and a multi-view lens. Furthermore, the optical system unit may be able to change optical characteristics such as zoom and aperture (depending on, for example, the content of the image to be acquired). The image sensor may be, for example, a CMOS color image sensor including a primary color Bayer arrangement of color filters.

[0033] The display device 119 is, for example, a liquid crystal display disposed on the surface of the housing of the digital camera 100. The display device 119 may be a touch display. The display device displays live view images or images already read from the recording medium 115, menu screens, and information about the digital camera 100 (e.g., settings, remaining battery power, number of remaining images that can be taken, etc.).

[0034] Communication interface 120 is a circuit used for communicating with external devices according to one or more communication standards. Communication interface 120 includes a connector for wired communication, an antenna for wireless communication, and transmitting and receiving circuitry. Digital camera 100 can send image data to and receive data from external devices via communication interface 120. Representative communication standards compliant with communication interface 120 include, but are not limited to, HDMI. ® USB, Bluetooth ® And wireless LAN (Wi-Fi), etc.

[0035] Figure 1A and Figure 2 The functional blocks shown are composed of Figure 1B One or more hardware items are shown. For example, the camera unit 101 is mainly implemented by the CPU 112 and the camera device 118. Furthermore, the prompt generation unit 102 is mainly implemented by the CPU 112 and the GPU 116. Note that RAM 114 and ROM 113 are used in various types of processing as temporary storage for data to be processed, currently processed data, and processing result data, respectively, and as reference destinations for pre-stored settings and programs.

[0036] <Shooting and Processing Procedures> Refer to Figure 1 to Figure 5 Describe the still image capturing operation of digital camera 100. Figures 3 to 5 The steps in the flowchart are executed by CPU 112 or GPU 116, which executes the program stored in ROM 113 and controls other hardware items as needed.

[0037] Furthermore, still image capture by the digital camera 100 is performed in response to instructions received via the input device 117, and can also be performed based on conditions different from the instructions and pre-set. For example, video recording can be performed continuously at constant time intervals, or still image capture can be performed when the information obtained from moving images captured for live view display has met predetermined conditions.

[0038] Note that the CPU 112 can determine the exposure parameters for still image capture based on, for example, brightness information obtained from moving images captured for a live view display. Similarly, the CPU 112 can also control the lens focusing distance for still image capture based on, for example, contrast information obtained from moving images captured for a live view display.

[0039] In step S301, the camera unit 101 performs still image capture and obtains image data. Note that the camera unit 101 can also perform moving image capture and use the frame image of the moving image as still image data. The camera unit 101 outputs the obtained image data and the aforementioned attribute information to the prompt generation unit 102. Note that the camera unit 101 outputs so-called image data after development processing. In the image data after development processing, the pixel data constituting the image data includes three components (RGB or YCbCr). Note that the camera unit 101 may, for example, process the image data to make it suitable for use by the prompt generation unit 102 before outputting the image data.

[0040] In cases where image data is recorded in addition to prompts, the camera unit 101 also outputs image data and attribute information to the recording unit 103. For example, the camera unit 101 may output image data to the recording unit 103 that has undergone processing (e.g., encoding processing) corresponding to the data format used for recording in the recording unit 103.

[0041] In step S302, the prompt generation unit 102 generates text information (prompt) based on the image data output from the camera unit 101 to describe the shooting scene of the image indicated by one or more image data corresponding to one or more frames.

[0042] Now, will use Figure 2 and Figure 4 The flowchart shown further describes the operation of the prompt generation unit 102. In step S401, the image / hint conversion unit 201 generates text information (hints) describing the shooting scene of the image represented by the input image data. The image / hint conversion unit 201 stores the generated hints in the storage unit 203 (RAM 114).

[0043] As described above, the image / cue conversion unit 201 can obtain cues by inputting image data into a learning model stored in ROM 113. Alternatively, the image / cue conversion unit 201 can send image data to an external device via communication interface 120 and receive cues from the external device.

[0044] The level of detail in the prompts generated by the image / prompt conversion unit 201 can be changed through settings. For example, when the level of detail is set to low, prompts that only indicate gender can be generated for human subjects, while when the level of detail is set to high, prompts that indicate gender, age, hair color, and length can be generated.

[0045] Furthermore, the image / hint conversion unit 201 can generate hints that include not only descriptions (positive hints) related to elements included in the image, but also descriptions (negative hints) related to elements not included in the image.

[0046] Figure 5 Figures 5a and 5b show examples of images represented by image data and examples of prompts generated based on image data, respectively. The image / prompt conversion unit 201 stores the prompts generated for each piece of image data in the storage unit 203 (RAM 114).

[0047] The image / cue conversion unit 201 generates cues for each image data corresponding to a frame supplied from the camera unit 101 as the cue generation object, and stores the generated cues in the storage unit 203.

[0048] In step S402, the prompt editing unit 202 generates a final prompt by applying a predetermined editing process to the prompt stored in RAM 114, and outputs the final prompt to the recording unit 103.

[0049] The editing process involves generating a final prompt based on additional information captured from the image data that serves as the source of the prompt, as well as prompts generated from other highly relevant image data and their capture information.

[0050] Specific examples of editing processes include, but are not limited to: (1) Adding additional information about the image data used as the source of the prompt to a prompt, or information based on the additional information, during the recording of the image. (2) Based on the prompts that have been generated for multiple image data that are highly relevant to other multiple frames, generate a prompt, taking into account additional information at the time of shooting as needed.

[0051] For example, the user can set whether to perform editing processing and what kind of editing processing to perform. For example, editing the setting for "each shot" corresponds to (1), and editing the setting for "each event" corresponds to (2). Other settings such as "every 10 minutes" can be set. When no editing processing is performed, the prompt editing unit 202 outputs the prompt generated by the image / prompt conversion unit 201 to the recording unit 103 as is. Note that even when editing processing is performed, prompts that have not yet been edited can be output to the recording unit 103.

[0052] Examples of highly correlated multiple image data include, but are not limited to: (a) Multiple image data points differ by less than a pre-set threshold in at least one of the following aspects: shooting date and time, and shooting location. (b) Used to show multiple image data of the same or similar subjects. (c) Multiple image data sets that hold inter-image correlation equal to or higher than a threshold, and (d) Multiple image data entries containing the same keyword were generated. Note that multiple image data sets that satisfy two or more combinations of conditions (a) to (d) can be considered as highly correlated multiple image data sets.

[0053] For example, if editing for "each event" has already been set, the prompt editing unit 202 can regard multiple image data that satisfy both (a), (b), and (d), or (a), (b), and (d) as highly related multiple image data.

[0054] This assumes that settings have already been configured for editing to be performed on a per-event or per-constant-time-period basis, and by... Figure 5 Image data from two frames, indicated by 5a and 5c, which are close to each other in terms of shooting date and time and show the same subject 501, have been considered highly correlated image data.

[0055] In this case, the prompt editing unit 202 applies the editing process to the prompt X (which has already been generated regarding each piece of image data) Figure 5 5c) and prompt Y ( Figure 5 (5d) to generate a prompt ( Figure 5 (5e in the middle).

[0056] exist Figure 5 In the example shown, the prompt editing unit 202 generates the final prompt by applying editing processes to merge the contents of prompts X and Y according to the order of shooting dates and times (i.e., according to the order in which the images were shot), and editing processes to add additional information related to the shooting dates and times. When merging prompts, the prompt editing unit 202 can use, for example, existing synthetic AI techniques to add particles and conjunctions as needed, thereby generating prompts with natural written expression. Furthermore, when merging multiple prompts, the expression of the merged prompt can be changed taking into account the frequency of word occurrence; for example, words with high frequency of occurrence can be emphasized in the merged prompt.

[0057] return Figure 3In step S303, the recording unit 103 records the prompt generated by the prompt generation unit 102 in step S302 into the recording medium 115. Note that the recording unit 103 may record only the prompt, or it may record the prompt and the image data used to generate the prompt in association with each other. In cases where a prompt has already been generated based on prompts and additional information related to image data from multiple frames, the same prompt may be associated with each piece of image data, or the prompt may be associated only with representative image data.

[0058] The method for associating hints and image data with each other can be any known method. For example, hints and image data can be included in the same file container, or hints and image data can be recorded as different files with common filenames.

[0059] Note that the recording unit 103 can apply digital proof processing to the prompt and the image data used to generate the prompt, and then record the prompt and the image data. The digital proof processing is sufficient to guarantee the content of the prompt and the image data. For example, the digital proof processing could be used to grant NFTs (non-fungible tokens).

[0060] As described above, this embodiment can provide the function of generating text data (tips) to describe the shooting scene to camera devices that are typically dedicated to video recording. This makes it easy, for example, to provide information related to the shooting scene to third parties in text form. Furthermore, the generated tips can be input into image-generating AI and used to generate images of similar scenes. Moreover, when viewing images later, referring to the tips makes it easier to recall the circumstances of the shooting, thus improving convenience.

[0061] ·<Second Embodiment> Next, a second embodiment of the present invention will be described. Since this embodiment can be embodied in the digital camera 100 described in the first embodiment, the description of the content described in the first embodiment will be omitted.

[0062] The second embodiment relates to the operation of a digital camera 100 when the image data obtained by the camera unit 101 using the currently set exposure parameters is not suitable for generating a prompt in the image / prompt conversion unit 201.

[0063] Will use by Figure 8 The shooting scene indicated in 8a is provided as an example for description. Figure 8 Example 8a in the text refers to a scene where "a dog jumps and catches a ball thrown by a boy." Furthermore, the photographer attempted to capture the moment in a tracking shot when the dog is centered in the image and catches the ball in mid-air. Additionally, it is assumed that the photographer had already... Figure 8The 8c indicates that the lens focal length and shutter speed are set to shooting parameters suitable for composition and tracking shots.

[0064] It is also assumed that the information has already been obtained through video recording. Figure 8 The image data indicated by 8b in the diagram. Further assuming that the image / cue conversion unit 201 has already... Figure 8 The data of the image indicated by 8b in the image was obtained Figure 8 The prompt indicated by 8d in the text.

[0065] However, the descriptions provided are "dog," "jump and catch," and "object," and do not include information about "boy" and "ball." This is because "boy" is not shown in the image from the photographer's intended perspective, and because the subject other than the one exhibiting the same movement as the "dog" is blurred in the image obtained through tracking, it is impossible to identify the "ball" from the image data.

[0066] In this way, under the following shooting parameters, the amount of information included in the cues generated from the images obtained by video recording may be reduced, and the usefulness of the cues may decrease. • A close-up shot of a portion of the scene (with a field of view narrower than the threshold). • The blur range increases (when tracking shooting mode is set or the f-number is close to the maximum aperture (less than the threshold)). • Camera shake is likely to occur (when a slow shutter speed equal to or below the threshold is set). • Noise is easily increased (when the shooting ISO is set to be equal to or higher than the threshold). • Underexposure or overexposure above the threshold compared to the exposure parameters for achieving correct exposure. Note that these are just examples.

[0067] In this embodiment, if it is determined that the image data obtained using the current shooting parameters is not suitable for generating a prompt containing sufficient information, the current shooting parameters are changed to shooting parameters that have a high probability of generating an image that is more suitable for generating a prompt.

[0068] The following uses Figures 6 to 8 This embodiment describes the still image capture process of the digital camera 100. Figure 6 and Figure 7 In this process, steps that perform operations similar to those in the first embodiment are assigned to... Figure 3 or Figure 4 The same reference numerals are used in the accompanying drawings, and their descriptions are omitted.

[0069] In step S601, the camera unit 101 obtains the current shooting parameters. These shooting parameters are not limited to exposure parameters (especially f-stop and shutter speed), and may include the shooting mode and lens angle of view. Figure 8 In the example shown, the lens's focal length (angle of view) and shutter speed are used as shooting parameters.

[0070] In step S602, the camera unit 101 (CPU 112) determines whether there is a high probability that the image data obtained by capturing images using the current shooting parameters is unsuitable for generating a prompt. This determination can be made by comparing a pre-determined threshold for each item of the shooting parameters with the currently set value. Note that the threshold can be changed dynamically. For example, the threshold for shutter speed can have a value corresponding to the current focal length of the lens.

[0071] If the current shooting parameters include at least one item that is unsuitable for generating a prompt, the CPU 112 determines that there is a high probability that the image data obtained by shooting using the current shooting parameters is unsuitable for generating a prompt. If it has been determined that there is a high probability that the image data obtained by shooting using the current shooting parameters is unsuitable for generating a prompt, the CPU 112 executes step S603; and if this is not determined, the CPU 112 executes step S607.

[0072] In step S603, the camera unit 101 utilizes shooting parameters suitable for generating image data for prompting ( Figure 8 (8e) replaces the current shooting parameters ( Figure 8 The camera is captured using 8c). Shooting parameters suitable for generating prompts can be pre-stored in ROM 113. Assume that the shooting parameters suitable for generating prompts are those for capturing wide-angle images in deep focus mode. Here, it is assumed that parameters already used by... Figure 8 The shooting parameters indicated by 8e were obtained by Figure 8 The data Z in the image is indicated by 8f.

[0073] In step S604, the camera unit 101 uses the image data Z( obtained by using the changed shooting parameters) Figure 8 (8f) and the original shooting parameters before the change ( Figure 8 8c) are stored in RAM 114 in a related manner.

[0074] In step S605, the prompt generation unit 102 uses the image data Z and original shooting parameters, as well as additional information, already stored in RAM 114 in step S604 to generate a prompt to be recorded. The prompt generation unit 102 then outputs the generated prompt to the recording unit 103.

[0075] Will use Figure 7 The flowchart shown describes Figure 6 The operation of the prompt generation unit 102 in step S605. Step S401 is as described in the first embodiment, therefore its description is omitted. It is assumed that the result of the processing has been obtained from... Figure 8 The prompt Z indicated by 8g is stored in the storage unit 203 (RAM114), along with the original shooting parameters and additional information.

[0076] In step S701, the editing unit 202 is prompted to use the... Figure 8 The editing process of the original shooting parameters indicated by 8c is applied to prompt Z, thereby generating a prompt to be recorded. The prompt editing unit 202 outputs the generated prompt to the recording unit 103.

[0077] During the editing process, the editing unit 202 is prompted to apply the editing process described in step S402 to the image generated from the image data. Figure 8 The 8g indicator in the text indicates the Z signal. Note that, as... Figure 8 The 8h indicated in the prompt refers to the original shooting parameters, not the actual shooting conditions used.

[0078] Since steps S301 and S303 are as described in the first embodiment, their description is omitted.

[0079] As described above, according to this embodiment, when it is determined that the image data obtained using the current shooting parameters is not suitable for generating a prompt, shooting parameters suitable for generating image data for a prompt are used for recording. On the other hand, when shooting parameters need to be added to the prompt, the original shooting parameters before the change are added; in this way, the shooting parameters intended by the photographer are reflected in the prompt.

[0080] Therefore, even when shooting parameters are set that have a high probability of producing image data that is not suitable for generating prompts, it is possible to achieve the beneficial effect of generating prompts that include appropriate types of information and reflect the photographer's intentions.

[0081] (Other embodiments) This invention can be implemented by supplying a program for implementing one or more functions of the above embodiments to a system or device via a network or storage medium, and causing one or more processors in the computer of the system or device to read and execute the program. This invention can also be implemented by circuitry (e.g., an ASIC) for implementing one or more functions.

[0082] This invention is not limited to the embodiments described above, and various changes and modifications can be made without departing from the spirit and scope of the invention. Therefore, the appended claims are made to inform the public of the scope of the invention.

[0083] This application claims the benefit of priority to Japanese Patent Application 2023-192351, filed on November 10, 2023, the entire contents of which are incorporated herein by reference.

Claims

1. A camera device, characterized in that... include: Camera components; A generation component is used to generate a prompt based on image data obtained by the camera component, wherein the prompt is text information describing the shooting scene of the image represented by the image data; and A recording component is used to record the aforementioned prompt.

2. The camera device according to claim 1, characterized in that, The generating component adds information from when the image data was captured to the prompt.

3. The camera device according to claim 2, characterized in that, The information includes one or more of the following during the recording: exposure parameters, location information, air temperature, and the magnitude and direction of the movement of the camera device.

4. The camera device according to any one of claims 1 to 3, characterized in that, The generating component generates a prompt based on prompts generated for each of the multiple image data corresponding to the multiple frames obtained by the camera component.

5. The camera device according to claim 4, characterized in that, The generation component generates the single prompt by merging the prompts generated separately for the multiple image data corresponding to the multiple frames, taking into account the shooting order of the multiple image data.

6. The camera device according to claim 4 or 5, characterized in that, The multiple image data corresponding to the multiple frames are one of the following: (a) Multiple image data that differ from a preset threshold in at least one of the following aspects: shooting date and time, and shooting location; (b) Used to show multiple image data of the same or similar subjects; (c) Multiple image data sets with inter-image correlations higher than or equal to a threshold; and (d) Generate multiple image data entries with hints containing the same keywords.

7. The camera device according to any one of claims 1 to 6, characterized in that, If the image data obtained based on the current shooting parameters is determined to be unsuitable for generating the prompt by the generating component, the camera component obtains the image data based on pre-set shooting parameters, and If the prompt is to include shooting parameters, the generating component will include the current shooting parameters in the prompt.

8. The camera device according to claim 7, characterized in that, The camera component makes the determination based on one of the following current shooting parameters: angle of view, shooting mode, f-number, shutter speed, and shooting ISO.

9. The camera device according to any one of claims 1 to 8, characterized in that, The recording component records the prompt in association with the image data.

10. The camera device according to any one of claims 1 to 9, characterized in that, The generation component utilizes a multimodal AI learning model to generate the prompt.

11. The camera device according to any one of claims 1 to 9, characterized in that, The generating component sends the image data to an external device and receives a prompt corresponding to the image data from the external device.

12. A control method executed by a camera device, characterized in that... include: A prompt is generated based on image data obtained from the camera component. The prompt is text information describing the shooting scene of the image represented by the image data. Record the prompt.

13. A program for enabling a computer to function as the generating component and the recording component included in the camera device according to any one of claims 1 to 11.